Engineering antigen-binding proteins

A deep learning method using encoder-decoder architectures predicts functional antigen-binding protein pairs, addressing limitations in characterizing BCR and TCR repertoires by learning antibody features, enhancing throughput and reducing costs.

JP2025536256APending Publication Date: 2025-11-05アルケマブ セラピューティクス リミテッド
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025520820
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-10-14
Filing Date
2023-10-12
Publication Date
2025-11-05

AI Technical Summary

Technical Problem

Current methods for identifying functional BCR heavy chain-light chain pairs and TCR αβ chain pairs are limited by the lack of pairing information in bulk sequencing data, leading to challenges in characterizing the diverse BCR and TCR repertoires, especially in terms of throughput, cost, and sample requirements.

Method used

A deep learning approach using encoder-decoder architectures, such as Transformers and BERT, to predict functional antigen-binding protein pairs by learning antibody features through masked language modeling, enabling the identification of likely functional pairings based on heavy and light chain sequences.

Benefits of technology

The method achieves high recall and precision in predicting functional antigen-binding protein pairs, overcoming limitations of existing computational methods and facilitating the characterization of diverse BCR and TCR repertoires with improved throughput and reduced costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025536256000001_ABST
    Figure 2025536256000001_ABST
Patent Text Reader

Abstract

A method is described for determining whether a protein chain pair comprising a first chain and a second chain is likely to form a functional antigen-binding protein. The method includes providing a query sequence pair comprising a first protein chain sequence and a second protein chain sequence as input to a deep learning model configured to receive the protein chain sequence pair as input and generate a score as output indicative of the likelihood that the protein chain pair will form a functional antigen-binding protein, the deep learning model including an encoder module and a classifier module, and the deep learning module has been trained using paired training sequences from known antigen-binding proteins. The method finds related use in identifying antibodies from amplicon or single chain sequences. Related methods and products are described.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] FIELD OF THE INVENTION The present invention relates to methods for engineering antigen binding proteins such as B cell receptors, antibodies, T cell receptors, etc. by determining whether candidate variable chain pairings are likely to be functional, for example by determining whether a candidate heavy chain-light chain pair or a candidate α-β or γ-δ chain pair is likely to be functional, or by identifying candidate variable chains (e.g., heavy / light, α / β, γ / δ chains) that are likely to form functional pairings with input chains (e.g., light / heavy, β / α, δ / γ chains). The present invention also relates to methods for providing antigen binding proteins, e.g., therapeutic antibodies, derived from input variable chains, e.g., B cell receptor / antibody heavy or light chains. [Background technology]

[0002] Background of the Invention Effective humoral immunity requires a diverse population of B cells capable of binding to a variety of antigens via the B cell receptor (BCR). The theoretical total size of the BCR repertoire in humans is approximately 10 15 It is estimated that the variants are 9are circulating in a single individual at any given time [Rees, 2020]. The BCR is composed of two pairs of protein chains: two heavy chains and two light chains. Each B cell expresses a (likely unique) pair of heavy and light chains to form its BCR, which is expressed on its surface or secreted as an antibody. The Observed Antibody Space currently catalogs over 600 million different human heavy chain sequences and approximately 70 million light chain sequences [Kovaltsuk et al., 2018]. Characterization of an individual's BCR ensemble (also known as the individual's BCR repertoire) has proven to be a valuable tool for understanding the biology of various diseases [Vander Heiden et al., 2017, Bashford-Rogers et al., 2019, Nielsen et al., 2020, Simonich et al., 2019] and for discovering novel therapeutic antibody drugs [Krawczyk et al., 2019, Galson et al., 2020].

[0003] There are two major approaches to characterizing an individual's BCR repertoire: single B cell sequencing and bulk B cell population sequencing. Single cell sequencing is more commonly used in antibody discovery applications because it preserves the pairing information between heavy and light chains. However, single cell sequencing has limited throughput, and different platforms and protocols result in variable coverage of the BCR repertoire present within a single sample. Even the most advanced microfluidic systems typically capture approximately 10 B cells per sample. 4 Only B cell sequences can be recovered [King et al., 2021, Eccles et al., 2020, Setliff et al., 2019]. Humans typically produce approximately 10 B cells per milliliter of blood. 6Because of the limited availability of B cells [Mora and Walczak, 2019], single-cell approaches lack the ability to characterize the entire B-cell diversity even with small samples. Additionally, single-cell sequencing has very specific sample requirements (e.g., cells must typically maintain viability until processing, necessitating fresh sample processing on the day of collection or frozen processing according to specific protocols), a very high cost per sample compared to bulk sequencing (single-cell sequencing is at least an order of magnitude more expensive than bulk sequencing), and requires dedicated laboratory facilities.

[0004] Bulk B cell population sequencing is performed at approximately 10 per sample. 7 B cell sequences are more easily recoverable [Briney et al., 2019], which significantly more closely resembles the expected diversity in an individual. However, because B cells are lysed during library preparation, heavy chain-light chain pairing information is not preserved. Typically, these bulk BCR sequencing approaches focus only on heavy chains, as they play a dominant role in antigen binding and are significantly more diverse than the light chain repertoire [Kovaltsuk et al., 2018]. However, antibody discovery requires having both the heavy and light chains of an antibody to enable synthesis and functional characterization. The gap in light chain pairing information has prompted the development of computational pairing methods [Reddy et al., 2010, Zhu et al., 2013, Raybould et al., 2021, Rakocevic et al., 2021]. However, this is limited to specific datasets and a small number of specific sequences within those datasets.

[0005] Similarly, cell-mediated immunity requires a diverse population of T cells capable of binding to a variety of antigens via their T cell receptors (TCRs). The total size of the TCR repertoire in humans is approximately 10 15It is predicted to contain a unique αβ T cell receptor (TCR) pair [Carter et al., 2019]. While experimental approaches for paired αβ TCR sequencing have been developed (including single-cell approaches [Zheng et al., 2017] and multi-cell deconvolution-based approaches [Howie et al., 2015]), they remain throughput-specific and limited. Therefore, the majority of available TCR repertoire knowledge is based on bulk sequencing of single-chain repertoires, mostly β-chain repertoires. This is inherently limited, especially since both α and β TCR chains have been shown to be involved in alloreactivity and antigen specificity [Carter et al., 2019].

[0006] Thus, there remains a need for improved methods for identifying chain pairs, such as BCR heavy chain-light chain pairs and TCR αβ chain pairs, from data that do not contain this pairing information. Summary of the Invention [Means for solving the problem]

[0007] Summary of the Invention The problem of identifying BCR heavy chain-light chain pairs is by no means trivial. In fact, the diversity of the BCR repertoire creates a large search space. Furthermore, although several heavy chain-light chain combinations can generate stable BCRs (an observation that has led some to speculate that pairing may be random [Glanville et al., 2009; Jayaram et al., 2012; DeKosky et al., 2016]), only a limited number of pairings generate functional BCRs capable of binding their target antigen [Teplyakov et al., 2016; Ling et al., 2018]. This suggests that even if functional pairing is nonrandom, the determinants of functional pairing are obscured by the number of pairings that can be nonfunctional yet stable. In practice, this means that finding the right light chain for a particular heavy chain is challenging, even if stable pairings can be predicted. This is because it is likely to generate a significant number of solutions that require experimental validation and that may not be well validated if selected primarily on the basis of stability.

[0008] Several different computational approaches have been proposed, each with significant drawbacks. The first approach was based on matching the relative frequencies of BCR heavy and light chains when sequenced independently [Reddy et al., 2010]. In this study, mice were first immunized to generate a strong immune response, and then the top 4–5 most frequently occurring heavy and light chains were selected for pairing. Pairing based on relative frequency was not possible outside of these top 4–5 sequences. More recently, Rakocevic et al.

[2021] showed that this approach only worked for samples dominated by a small number of high-frequency B cells. Zhu et al.

[2013] proposed a method called phylogenetic pairing, which involves comparing the architecture of phylogenetic trees generated from heavy and light chain sequence data. This method, in this case, is limited to examining specific clonal expansions of known antiviral antibody lineages rather than the entire BCR repertoire. Raybould et al.

[2021] proposed an approach based on in silico heavy and light chain pairing structural models. The approach is inherently limited by the limited and highly distorted availability of high-quality structural templates, and at best, can only identify features associated with stability, not necessarily functionality. Furthermore, the approach was only able to pair similar sequence families, rather than specific sequences (thus potentially limiting its practical applicability—it has not been experimentally validated). Therefore, the inventors have confirmed that current computational heavy chain-light chain pairing methods are limited in that they can only be applied to specific datasets and sequences within those datasets. In fact, all existing validated approaches are only applicable to datasets where both heavy and light chain sequences are available from a sample and the data are dominated by large clonal expansions, and they only facilitate pairing of a limited number of sequences within such datasets.

[0009] The inventors further confirmed that for generalized application to antibody discovery, it would be desirable to be able to generate a practical light chain for any given heavy chain. Because BCR repertoire bulk sequencing efforts often focus on the limited resources of heavy chain sequencing, which is thought to play a more important functional role than light chains, it would be even more desirable to be able to generate this using only heavy chain information. To address this problem, the inventors hypothesized that it would be possible to use deep learning methods inspired by recent advances in natural language processing (NLP). Specifically, we hypothesized that deep learning models, including encoder-decoder architectures (e.g., Transformers [Vaswani et al., 2017]) and derived architectures such as BERT [Devlin et al., 2018] and RoBERTa [Liu et al., 2019] (encoder-only architectures) or GPT [Brown et al., 2020] and Falcon [Penedo et al., 2023], should be able to learn antibody features using masked language modeling, similar to what has been used to train such models in natural language processing. Furthermore, we hypothesized that the resulting learned representations would carry information usable by a classifier model to predict whether candidate pairs are likely to form functional pairs. Transformers have demonstrated state-of-the-art results in a wide range of NLP tasks [Vaswani et al., 2017, Devlin et al., 2018, Liu et al., 2019, Rothe et al., 2020]. To this end, we devised a method for using a pre-training encoder (termed "AntiBERTa") or decoder (termed "FAbCon") to generate fine-tuned learning representations of heavy and light chains as part of a classifier for pairing tasks. Multiple blind tests on single-cell datasets with known pairings showed that this approach predicts true pairs with high recall and precision. This approach offers a novel solution to light chain pairing and a path to filling the gap in bulk heavy chain sequencing.The inventors further confirmed that the same approach can be used to solve the problem of TCR chain pairing.

[0010] Thus, according to a first aspect, there is provided a method for determining whether a protein chain pair comprising a first chain and a second chain is likely to form a functional antigen-binding protein, the method comprising providing as input a query sequence pair comprising a first protein chain sequence and a second protein chain sequence to a deep learning model configured to receive as input the protein chain sequence pair and to generate as output a score indicative of the likelihood that the protein chain pair will form a functional antigen-binding protein, the deep learning model comprising an encoder module and a classifier module, and the deep learning module has been trained using paired training sequences from known antigen-binding proteins.

[0011] Also described according to this aspect is a method for identifying an antigen binding protein comprising a first protein chain and a second protein chain, the method comprising providing a query first protein chain, providing one or more candidate second protein chain sequences, and determining whether each protein chain pair comprising the query first protein chain and the candidate second protein chain is likely to form a functional antigen binding protein using the methods described above. Thus, also provided according to this aspect is a method for identifying antigen-binding proteins comprising a chain pair, the method comprising: providing a query first protein chain; and providing one or more candidate second protein chain sequences; and identifying second protein chains by determining whether the one or more candidate second chain sequences are likely to form a functional antigen-binding protein with the query first protein chain by providing each of the one or more query sequence pairs comprising (i) a sequence of the first protein chain and (ii) the candidate second protein chain sequences as input to a deep learning model, wherein the deep learning model is configured to receive as input a protein chain sequence pair or multiple protein sequence pairs and generate as output a score or derived information indicative of the likelihood that each protein chain pair will form a functional antigen-binding protein, wherein the deep learning model comprises an encoder module and a classifier module, and the deep learning module has been trained using paired training sequences from known antigen-binding proteins.

[0012] The method according to this embodiment may have one or more of the following features.

[0013] The score indicative of the likelihood that the protein chain pair will form a functional antigen-binding protein can be the probability that the protein chain pair will form a functional antigen-binding protein.

[0014] Identifying the second protein chain can include providing a plurality of candidate second chain sequences and obtaining a score for each protein pair comprising the first protein chain and each candidate second chain. Information derived from the scores can include a ranking of the protein chain sequence pairs, where protein sequence pairs that are more likely to form a functional antigen-binding protein are ranked higher than protein sequence pairs that are less likely to form a functional antigen-binding protein. The ranking can be based on the score, for example, by ranking by decreasing or increasing the score. Thus, also provided according to this aspect is a method for identifying an antigen-binding protein comprising a chain pair, the method comprising: identifying a second protein chain by providing a query first protein chain; and providing a plurality of candidate second protein chain sequences; and determining whether the candidate second chain sequences are likely to form a functional antigen-binding protein with the query first protein chain by providing the plurality of query sequence pairs as inputs, each query sequence pair comprising (i) a sequence of the first protein chain and (ii) a candidate second protein chain sequence, to a deep learning model, wherein the deep learning model is configured to receive the plurality of protein chain sequence pairs as inputs and generate as output respective scores or information derived therefrom (e.g., rankings of the plurality of pairs) indicative of the likelihood that each protein chain pair will form a functional antigen-binding protein, wherein the deep learning model comprises an encoder module and a classifier module, and the deep learning module has been trained using paired training sequences from known antigen-binding proteins.Also provided in accordance with this embodiment is a method for identifying functional antigen-binding proteins comprising a chain pair, the method comprising: providing a plurality of query antigen-binding proteins comprising a first protein chain and a second protein chain; and determining whether the query first and second chain sequences are likely to form a functional antigen-binding protein by providing the plurality of query sequence pairs as input to a deep learning model, the deep learning model being configured to receive the plurality of protein chain sequence pairs as input and to generate as output respective scores indicative of the likelihood of each protein chain pair to form a functional antigen-binding protein or information derived therefrom (e.g., a ranking of the plurality of pairs), the deep learning model comprising an encoder module and a classifier module, and the deep learning module being trained using paired training sequences from known antigen-binding proteins. This method may be used to prioritize query chain pairs / antigen-binding proteins based on the likelihood of forming a functional antigen-binding protein predicted using the deep learning model.

[0015] The chain pair and / or each protein chain may be referred to as a "variable chain." The phrase "known chain pair" or "known antigen-binding protein" refers to an antigen-binding protein / antigen-binding protein-derived variable chain sequence pair known to be present in an antigen-binding protein that exhibits a desired antigen-binding function or that forms part of at least one subject's B-cell or T-cell repertoire. The latter may also be referred to as a "native" chain pair. Thus, a "known protein chain / antigen-binding protein" may be a protein / chain pair that has already been identified (e.g., in a sample containing the native chain pair / protein, in an individual, etc.) and / or a protein / chain pair that has a desired function (e.g., that has been verified or can be verified by in vitro or in vivo testing, e.g., for binding affinity to a target, expression, stability, etc.). All sequences may be amino acid sequences. The first (query) chain sequence may be a heavy chain sequence and the second sequence may be a light chain sequence. The antigen-binding protein may be a B-cell receptor or an antibody, or a protein derived therefrom. Thus, the antigen-binding protein may comprise a heavy-light chain pair. The query sequence may comprise a heavy chain sequence or a light chain sequence. The match chain sequence may be a light chain sequence or a heavy chain sequence.

[0016] The antigen-binding protein may be a T cell receptor or a protein derived therefrom. The antigen-binding protein may comprise an αβ chain pair, where the first chain sequence is a β chain sequence or an α chain sequence, and the corresponding chain sequence is an α chain sequence or a β chain sequence. The first chain sequence may be a β chain sequence, and the corresponding sequence may be an α chain sequence. The antigen-binding protein may comprise a γδ chain pair, where the first chain sequence is a δ chain sequence or a γ chain sequence, and the corresponding chain sequence is a γ chain sequence or a δ chain sequence. The first chain sequence may be a δ chain sequence, and the corresponding sequence may be a γ chain sequence. The antigen-binding protein may be a T cell receptor or a protein derived therefrom. Thus, the antigen-binding protein may comprise an αβ chain pair or a γδ chain pair. Thus, the query sequence may comprise a β or δ chain sequence, or an α or γ chain sequence. The corresponding chain sequence may be an α or γ chain sequence, or a β or δ chain sequence.

[0017] The encoder module may include multiple encoders (also referred to as "encoder models"). Each encoder may receive a protein chain sequence as input. The encoder module may include one or two transformer-based encoder models. The encoder module may include one or two encoders pre-trained with training sequences of ampli?ed protein chains from known antigen-binding proteins. The encoder module may include one or two encoders of a sequence-to-sequence model. The sequence-to-sequence model may be a recurrent neural network or a transformer. The sequence-to-sequence model may be a sequence-to-sequence transformer-based model. The recurrent neural network may be a gated recurrent unit (GRU)-based model or a long short-term memory (LSTM) model. For example, the GRU-based model may include a GRU-based encoder and a GRU-based decoder. The encoder may be, for example, a four-layer bidirectional GRU with a hidden dimension of 1024. The decoder may be, for example, a 4-layer forward-only GRU with 1024 hidden dimensions. A Transformer is a deep learning model that uses an attention mechanism. A Transformer-based model may be a Transformer model with an architecture using self-attention and point-wise fully connected layers in both the encoder and decoder. The encoder and / or decoder may be composed of stacks of identical layers, for example, 6, 12, 24, or 30 layers. Each layer of the encoder may have two sublayers: a multi-head self-attention layer and a position-wise fully connected feedforward network layer. Each layer of the decoder may have three sublayers: a self-attention sublayer, a layer that performs multi-head attention on the output of the encoder stack, and a feedforward network layer.The encoder module may include a decoder model pre-trained using training sequences including paired and unpaired protein chains from known antigen-binding proteins. The decoder may be a decoder for a decoder-only Transformer-based model. The encoder module may include an embedding layer and a decoder layer (collectively sometimes referred to as the "decoder model") for a decoder-only autoregressive Transformer model. The decoder model may use flash attention, multi-query attention, and / or positional encoding. The decoder may be a 24-layer decoder with 12 attention heads, 768 embedding dimensions, and 3072 feedforward dimensions in each layer. The decoder may be a 28-layer decoder with 16 attention heads, 1024 embedding dimensions, and 4096 feedforward dimensions in each layer. The decoder may be a 56-layer decoder with 32 attention heads, 2048 embedding dimensions, and 8192 feedforward dimensions in each layer. Each such model could be used, but it has been found that the smallest of these models already has very good performance.

[0018] The encoder module may include one or two encoders that use position encoding, such as absolute position encoding or relative position encoding. In an embodiment, the encoder module includes one or two encoders that use absolute position encoding. In an embodiment, the encoder module includes one or two encoders that use relative position encoding. Relative position encoding (also referred to as relative position representation) may be implemented as described in Shaw et al. (2018). An encoder that uses relative position encoding may be an encoder that uses rotary position encoding. Rotary position encoding may be implemented as described in Su et al. (2022). Relative position encoding may improve the model's ability to capture relationships between positions in a chain. Transformer-like models with relative position coding (e.g., including encoder-only models or Transformer models) may use relative positional information as an additional component to the key and value used in the Transformer-like model's self-attention mechanism. The encoder module may include a decoder that uses positional encoding, such as ALiBi positional encoding (described in Press et al., 2021). The encoder module may include a decoder that uses rotary position embedding, as described in Su et al. (2022). For example, the Falcon (falconllm.tii.ae / ) and LlaMa2 (Touvron et al., 2023) models are decoder-only models that use rotary position embedding.

[0019] The encoder module may include one or two encoder models of a sequence-to-sequence model trained using masked language modeling. Alternatively, the encoder may be trained using span-based masked language modeling (see, e.g., Joshi et al. 2020). Training the encoder model may include training the model to randomly replace masked positions in the training amino acid sequence. The model may be trained using masked language modeling in which 15% of positions are masked during training and / or masked positions are replaced with masked tokens, random amino acids, or original amino acids at those positions. The encoder module may include one or two encoders pre-trained using training sequences comprising first and second protein chains from known antigen-binding proteins. The training sequences used for pre-training may include at least 1 million, at least 2 million, at least 5 million, at least 10 million, at least 20 million, or at least 50 million distinct sequences. The encoder module can include two copies of an encoder pre-trained with training sequences comprising ampered first and second protein chains from a known antigen-binding protein. The encoder module can include an encoder pre-trained with training sequences comprising a concatenation of a sequence from a first chain and a sequence from a second protein chain from a known antigen-binding protein, where the concatenated sequence comprises a sequence from an ampered protein chain.

[0020] The decoder may be a decoder model of a generative language model, where the decoder model has been trained using causal language modeling (i.e., a next-token prediction task) or span modeling (also referred to as "span-based masked language modeling"). The encoder module may include a decoder pre-trained with training sequences including a pair of first and second protein chains from a known antigen-binding protein and a pair of first and second protein chains from a known antigen-binding protein. The training sequences may include at least 500, 600, or 700 million individual sequences and / or at least 1, 1.5, or 2 million paired sequences. The encoder module may include a decoder pre-trained with training sequences each including a single chain from a known antigen-binding protein or a concatenation of a sequence from a first chain and a sequence from a second protein chain.

[0021] The deep learning model may further include a cross-attention module that receives the output of the encoder module as input and generates an output used by the classifier module to provide a score indicative of the likelihood that the protein chain pair will form a functional antigen-binding protein. The cross-attention module may include one or more cross-attention blocks, where each cross-attention block includes a self-attention layer and a cross-attention layer. The cross-attention module may include multiple cross-attention blocks. The cross-attention module may include one or more cross-attention blocks, including a cross-attention layer and a self-attention layer, where each cross-attention layer and self-attention layer may include multiple attention heads. Each self-attention layer may include one or more attention heads that respond to the output of the first encoder that receives the first chain sequence as input. Each cross-attention block may further include a residual connection between the output of the first encoder and the output of the self-attention layer. Each cross-attention layer may include one or more attention heads that address (i) the output of the self-attention layer or the output of a first encoder that receives the first strand sequence as input, and (ii) the output of a second encoder that receives the second strand sequence as input. Each cross-attention block may further include a residual connection between the output of the first encoder and the output of the cross-attention layer. Each self-attention layer may include 2, 3, 4, 5, 6, 12, or more attention heads. Each cross-attention layer may include 2, 3, 4, 5, 6, 12, or more attention heads. Each cross-attention block may include a self-attention layer and a cross-attention layer.Each cross-attention block may include a residual connection between the output of the self-attention layer and / or the cross-attention layer and the output of the first encoder that receives the first chain sequence as input. The output of the self-attention layer and / or the cross-attention layer may be normalized before being provided as input to a subsequent layer or module. The classification module may include a softmax layer that generates a score between 0 and 1 that can be interpreted as the probability that the input protein chain pair forms a functional antigen-binding protein. The classification module may include one or more of a dimensionality reduction layer, a regularization mechanism, and a layer having an activation function. The dimensionality reduction layer may include an attention pooling layer or an average pooling layer. The regularization layer may include a dropout mechanism. The activation function may be selected from Tanh, Leaky ReLU, GeLU, SmeLU, Swish, and ReLU. The activation function for each layer may be independently selected from Swish and ReLU. The deep learning model may receive as input a plurality of query chain sequence pairs and generate as output a respective score indicative of the likelihood that each query protein chain pair will form a functional antigen-binding protein, and / or a ranking of the plurality of query chain sequence pairs, such that query protein sequence pairs with a greater likelihood of forming a functional antigen-binding protein are ranked higher than query protein sequence pairs with a lesser likelihood of forming a functional antigen-binding protein. For example, the deep learning model may receive as input 8, 16, 32, or 64 query sequence pairs, e.g., 32 pairs. Such deep learning models may advantageously provide predictions for many candidate chain pairings with high computational efficiency.

[0022] Paired training sequences from known antigen-binding proteins may include paired training heavy and light chain sequences from single B cell sequencing data. The training data may include one or more data sets previously obtained by single B cell sequencing of samples obtained from subjects or by sequencing libraries derived therefrom. The training data may further include paired training heavy and light chain sequences from known antibodies / B cell receptors. For example, the training data may include paired training heavy and light chain sequences from one or more antibody / BCR databases, from one or more known therapeutic antibodies / BCRs, and / or from one or more antibodies / BCRs known to have desired binding functions. The training data may include paired training heavy and light chain sequences from a naive B cell receptor library. The training data may include paired training heavy and light chain sequences from an antigen-experienced B cell receptor library. Thus, the training data may include paired training heavy and light chain sequences obtained from subjects exposed to one or more specific antigens. The training first and second chain sequences from known chain pairs can include paired training α and β chain sequences from single T cell sequencing data. The training data can include one or more data sets previously obtained by single T cell sequencing of samples obtained from subjects or by sequencing libraries derived therefrom. The training data can further include paired training first and corresponding chain sequences from known T cell receptors. For example, the training data can include paired training α and β chain sequences from one or more T cell receptor databases, from one or more known therapeutic TCRs, and / or from one or more TCRs known to have desired binding functions. The training data can include paired training α and β (or δ and γ) chain sequences from a naive T cell receptor library. The training data can include paired training α and β (or δ and γ) chain sequences from an antigen-experienced T cell receptor library.Thus, the training data can include paired training α and β (or δ and γ) chain sequences obtained from subjects exposed to one or more specific antigens. The training first and second chain sequences from known chain pairs can include paired training chain sequences, each pair including chain sequences that include or consist of a V gene sequence or identifier, a J gene sequence or identifier, and a joining sequence, and optionally a D gene sequence or identifier. The training first and matched chain sequences from known chain pairs can include paired training chain sequences, each pair including chain sequences that include or consist of a V gene sequence or identifier, a J gene sequence or identifier, and a joining sequence. References to V genes or J genes can refer to the amino acid sequences corresponding to the respective genes. The training data can include at least 80,000, at least 100,000, at least 120,000, at least 150,000, at least 500,000, or at least 1,500,000 training sequence pairs, e.g., training heavy and light chain sequence pairs. Advantageously, the training data may include at least 1,500,000 training heavy and light chain sequence pairs. The training data may include chain sequence pairs from a mammal, such as a human. The training data may include mammalian heavy and / or light chain sequences. The training data may include human heavy and / or light chain sequences. The training data may include training sequence pairs from the same species as the query sequence. The training data may include at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or at least 95% sequences from the same species as the query sequence. The query sequence may be a sequence not present in the training data. The query sequence may be a sequence obtained from a sample from a subject with a desired characteristic, e.g., a desired phenotype. For example, the subject may have a particular clinical characteristic.

[0023] The training data may include simulated training sequences. The training data may include simulated paired sequences, paired simulated sequences, or simulated unpaired training sequences. The simulated data may include sequences obtained using methods for simulating antigen-binding protein sequences, such as immuneSIM (Weber et al., 2020) and / or protein sequence simulation methods, such as ProGen2 (Nijkamp et al., 2022). Known protein pairs may be referred to as a "positive set" of pairs. The training data may include a negative set of chain sequence pairs that are not predicted to form functional antibody-binding proteins. The negative set may include randomly paired first and second protein chain sequences. Alternatively or additionally, the negative set may include simulated pairs of sequences or pairs of simulated sequences. Alternatively or additionally, the negative set may include sequence pairs from antigen-binding proteins previously determined to be non-functional. Pairs may be determined to be non-functional according to one or more predetermined functionality criteria. For example, pairs that have not been experimentally expressed or that fail to bind to a target may be considered non-functional. Thus, the training data may further include a negative set comprising randomly paired first and second protein chain sequences. The randomly paired first and second protein chain sequences may have been obtained or may be obtained as part of the method by re-pairing paired training sequences from known antigen-binding proteins and / or by random pairing of unpaired training sequences. The negative set may comprise paired first and second protein chain sequences from antigen-binding proteins that are predetermined to be non-functional. The known paired training sequences (positive set) may be associated with a first label, and the randomly paired / negative set training sequences may be associated with a second label. The training data may include pairs associated with non-binary labels.The training data may include a positive set of pairs associated with a score that can take multiple values ​​(up to and including a continuum, where the continuum may be bounded, e.g., between 0 and 1). For example, pairs in the positive set may be associated with a score indicative of "nativeness" or "functionality." Such a score may reflect, for example, functional information associated with the known pair, e.g., binding affinity, or any other metric associated with binding strength. The training data may include a negative set of pairs associated with a single value (e.g., 0) or multiple values, e.g., values ​​indicative of the reliability of the pair as non-functional.

[0024] The training data may further include amp training first and / or second sequences, which may be used to pretrain the encoder or decoder of the encoder module. The amp training first and / or second strand sequences may have any of the sequence characteristics described in connection with the paired sequences. Specifically, the amp strand sequences may be the same type of sequence as the paired sequences (e.g., in which case the paired training sequences are heavy and light chain pairs and the amp training first / second strand sequences may include the amp heavy and / or light chain), may include sequences from the same organism (e.g., may include mammalian and / or human sequences, may include sequences from one or more organisms, may include sequences from a naive library and / or an exposed library, etc.), and may include the same information (e.g., gene segment identifiers, sequences, combinations thereof, etc.). The amp training sequences may include some or all of the first and / or second sequences present in the paired training sequences. Advantageously, the paired training sequences may contain more first strand sequences and / or more second strand sequences than the paired training strand sequences. The first (e.g., query) strand sequence may comprise or consist of a V gene sequence or identifier, a J gene sequence or identifier, and a junction sequence, and optionally a D gene sequence or identifier. The second (e.g., matched) strand sequence may comprise or consist of a V gene sequence or identifier, a J gene sequence or identifier, and a junction sequence. The format of the first and second strand sequences is related to the format of the training strand sequences. Thus, a deep learning model trained with training strand sequences comprising or consisting of a V gene sequence or identifier, a J gene sequence or identifier, and a junction sequence, and optionally a D gene sequence or identifier, may accept as input a strand sequence comprising or consisting of these components. Similarly, a deep learning model trained with training strand sequences comprising or consisting of a V gene sequence or identifier, a J gene sequence or identifier, and a junction sequence may accept as input a strand sequence comprising or consisting of these components. The query sequence may comprise or consist of one or more first strand CDR sequences.The second sequence may comprise or consist of one or more corresponding chain CDR sequences. The first and / or second sequence may comprise or consist of a CDR3 sequence. The first and second protein chains may be chains with different length ranges and / or different domain architectures.

[0025] Providing protein chain sequence pairs as input to a deep learning model (whether for training or prediction) may include encoding each of the protein chain sequences using a predetermined encoding scheme. According to the encoding scheme, each amino acid may be encoded individually. In embodiments, the sequences are encoded using tokens that each correspond to an individual k-mer. Providing the sequence pairs to the deep learning model may include encoding the sequences using an encoding scheme in which each amino acid is encoded individually. For example, a different token may be provided for each possible amino acid. Providing the sequence pairs to the deep learning model may include encoding the sequences using an encoding scheme in which each gene sequence identifier corresponds to an individual token. Providing the sequence pairs to the deep learning model may include encoding a query sequence using an encoding scheme in which each amino acid corresponds to an individual token. Providing the sequence pairs to the deep learning model may include encoding the query sequence using an encoding scheme in which the chain sequence is preceded or followed by a special token that indicates whether the protein chain sequence is a first or second protein chain (e.g., a heavy-H chain or a light-L chain). Providing the sequence pairs to the deep learning model may include encoding each sequence using an encoding scheme in which the sequence (i.e., the sequence available as a whole sequence rather than a gene identifier) ​​is encoded using tokens each corresponding to an individual k-mer (e.g., by using byte pair coding). Each sequence may be coded using overlapping k-mers. The k-mers may have any length that is shorter than the expected chain length. For example, the k-mers may be of length 1-100, 2-100, 2-50, 2-20, or 2-10. The k-mers may be of length 1-5. The k-mers may be of a fixed length. For example, a fixed k-mer length of 1, 2, 3, 4, or 5 may be used. A k-mer of length 1 is equivalent to encoding each character (e.g., each amino acid) individually.k-mers of length k>2 (e.g., 3) can be used as part of an encoding scheme using overlapping or non-overlapping k-mers. Overlapping k-mers can overlap by different amounts. For example, k-mers of length 3 can overlap by 1 or 2 characters. In a scheme using k=3, each token corresponds to a unique set of 3 characters (e.g., a 3-amino acid motif). Ampere training data can be filtered to exclude any sequences containing specific regions with lengths outside their respective predetermined length ranges. For example, any sequences with fewer than 20 amino acids before the CDR1 region, fewer than 10 amino acids after the junction region, CDR1 regions with lengths outside the 5-12 amino acid range, CDR2 regions with lengths outside the 1-10 amino acid range, and / or CDR3 regions with lengths outside the 5-38 amino acid range can be excluded from the amp training data. Paired training data can be filtered to exclude any pairs containing junction sequences (in the first and / or corresponding strands) outside the predetermined length ranges. In other words, the training data may not include any pairs that include a first (e.g., heavy) chain junction outside a predetermined length range and / or a second (e.g., light) chain junction outside a predetermined length range. For example, pairs that include heavy chain junction sequences of less than a predetermined length, e.g., 3, 4, 5, 6, 7, 8, 9, or 10 amino acids, may be excluded. As another example, pairs that include heavy chain junction sequences of more than a predetermined length, e.g., 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, or 35 amino acids, may be excluded. As another example, pairs that include light chain junction sequences of less than a predetermined length, e.g., 3, 4, 5, 6, 7, 8, 9, or 10 amino acids, may be excluded. As another example, pairs containing light chain joining sequences greater than a predetermined length, such as 15, 16, 17, 18, 19, 20, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 amino acids, may be excluded.The predetermined length may be the same or different between the junction sequences in the corresponding (e.g., light) and first (e.g., heavy) chains of the pair. In particular examples, it may exclude pairs containing heavy chain junction sequences of less than 7 amino acids and / or pairs containing heavy chain junction sequences of more than 30 amino acids. Alternatively or additionally, it may exclude pairs containing light chain junction sequences of less than 7 amino acids and / or pairs containing light chain junction sequences of more than 20 amino acids. The query sequence or sequence pair may include one or more gene sequence identifiers, and the method may further include replacing the one or more gene sequence identifiers with corresponding germline sequences. The deep learning model may be a Transformer-based model including an encoder pre-trained with the first and / or corresponding chain sequences for amp training and a decoder or bidirectional encoder pre-trained with the corresponding and / or first chain sequences for amp training. The encoder module may include a BERT model or a variant thereof, such as BERT, RoBERTa, DistilBERT, or RoFormer (Su et al., 2022). The encoder module may include an encoder trained using first and second strand sequences for amp training. Alternatively, the encoder module may include a model trained using a first (e.g., heavy or light) strand sequence for training and a model trained using a second (e.g., light or heavy) strand sequence. When the encoder module includes two encoders trained using first and second strand sequences for amp training, the two encoders may be the same pre-training model. Thus, the two encoders may be initialized using a pre-training model with the same parameters and architecture. The amp training strand sequence may include the full-length sequence in the variable region of the second strand. The amp training strand sequence may include the full-length sequence in the variable region of the first strand. Alternatively, the encoder module may include a decoder pre-trained using the amp training sequence and the paired training sequence.For example, the decoder may receive as input a string containing an encoding for a first strand or a second strand, preceded by a token indicating the strand type and optionally including one or more padding tokens. The decoder may also receive as input a string containing an encoding for a first strand and an encoding for a corresponding second strand (i.e., paired strand), each preceded by a token indicating the strand type and optionally including one or more padding tokens. The deep learning model may be trained using paired first and second (e.g., heavy and light) strand sequences from a known strand pair, where the sequences do not include full-length sequences in the variable regions of the second and / or first strands. In such embodiments, the deep learning model may be trained by imputing deleted sequence information to obtain paired training sequences that include full-length sequences in the variable regions of the corresponding strand and / or first strand. Imputing deleted sequence information may include replacing genetic identifiers with corresponding germline sequences. Imputing the missing sequence information can include predicting full-length sequences for each of the paired training first (e.g., heavy) and / or second (e.g., light) chain sequences from the partial sequences using a pre-training encoder. Alternatively, the amp training second (e.g., light) chain sequences and / or the amp training first (e.g., heavy) chain sequences can be converted into a format matching the format of the respective paired training sequences before pre-training the encoder.

[0026] Providing a query first protein chain or a query first protein chain of a query pair may include obtaining the sequence of the query first protein chain from a user via a user interface, from a computing device, from a sequence acquisition means or a computing device associated with the sequence acquisition means, or from a database or other computer-readable medium. Providing a query first protein chain or a query first protein chain of a query pair may include sequencing a sample containing genetic material encoding an antigen-binding molecule comprising the query sequence. Obtaining the query sequence may include performing B-cell bulk sequencing of a sample containing B cells, T-cell bulk sequencing of a sample containing T cells, or bulk sequencing of a sample containing any other cells expressing an antigen-binding molecule comprising the query sequence, or bulk sequencing of genetic material derived therefrom, for example, bulk sequencing of a B-cell receptor library or a T-cell receptor library. Providing a query first protein chain or a query first protein chain of a query pair may include obtaining a sample containing B cells, T cells, or other cells expressing an antigen-binding molecule containing the query sequence, or genetic material derived therefrom, such as a B cell receptor library or a T cell receptor library. Providing one or more candidate second protein chain sequences may include obtaining the candidate second protein chain sequences from a user via a user interface, from a computing device, from a sequence acquisition means or a computing device associated with the sequence acquisition means, or from a database or other computer-readable medium. The one or more candidate second protein chain sequences may be known second chain protein sequences or simulated second chain protein sequences. The known second chain protein sequences may be second chain sequences previously observed (e.g., in a sample, an individual, etc.) or second chain sequences known to have a predetermined function (e.g., an antigen-binding protein previously shown to bind to a particular target, expressed in a sample, etc.).Providing a query sequence (or sequence pair) may include obtaining the query sequence (or sequence pair) from a user via a user interface, from a computing device, from a sequence acquisition means or a computing device associated with the sequence acquisition means, or from a database or other computer-readable medium. Providing a query sequence or sequence pair may include sequencing a sample containing genetic material encoding an antigen-binding molecule comprising the query sequence. Providing a query sequence may include obtaining a sample containing B cells, T cells, or other cells expressing an antigen-binding molecule comprising the query sequence, or genetic material derived therefrom, e.g., a B cell receptor library or a T cell receptor library. Providing a query sequence may include sequencing a sample containing genetic material encoding an antigen-binding molecule comprising the query sequence, for example, by performing B cell bulk sequencing of a sample containing B cells (or any other cells expressing an antigen-binding molecule comprising the query sequence, or genetic material derived therefrom, e.g., a B cell receptor library). Providing a query sequence may include obtaining B cells or other cells that express an antigen-binding molecule comprising the query sequence, or a sample containing genetic material derived therefrom, for example, a B cell receptor library.

[0027] The method may include determining likelihoods for a plurality of pairs and ranking the plurality of pairs using the determined likelihoods. The method may further include providing a user via a user interface one or more identified second protein chain / pairs, portions thereof, or information derived therefrom, and / or one or more likelihoods that one or more query pairs will form a functional antigen-binding protein or information derived therefrom (e.g., ranking the query pairs by their score / likelihood to form a functional antigen-binding pair). The method may further include predicting a score indicative of the likelihood that each of the one or more query pairs comprising a respective candidate second protein chain sequence will form a functional antigen-binding protein, and identifying the candidate second protein chain sequences by applying one or more criteria to the score / likelihood. The one or more criteria may be individually selected from the score / likelihood exceeding a predetermined cutoff, the score / likelihood being the highest predicted score / likelihood of the set of candidate second protein chain sequences, and the score / likelihood being in a predetermined top percentile of predicted likelihoods of the set of candidate second protein chain sequences.

[0028] According to a second aspect, there is provided a method for providing antigen-binding protein chain pairing for a plurality of query sequences comprising a first chain sequence, the method comprising performing the method of any embodiment of the first aspect for each of the query sequences. The plurality of query sequences may be heavy chain or light chain sequences obtained by bulk B cell repertoire sequencing. The plurality of query sequences may comprise at least 10, at least 100, at least 1000, at least 10,000, or at least 100,000 sequences. The plurality of query sequences may be obtained by bulk B cell sequencing of a heavy chain or light chain repertoire in a sample, such as a sample from a subject. The plurality of sequences may be a subset of the set of sequences obtained by bulk B cell sequencing of a heavy chain or light chain repertoire in a sample. The method according to this aspect may have any of the features described in relation to the first aspect.

[0029] According to a third aspect, there is provided a method of providing an antigen binding protein with a desired property, the method comprising providing one or more query sequences comprising first chain sequences, at least one of the one or more query sequences being likely to have the desired property, and identifying a corresponding chain sequence for each of the one or more query sequences using the method of any embodiment of the first aspect. The method may have any one or more of the following features.

[0030] The method may further include obtaining one or more candidate antigen-binding proteins and one or more identified second sequences, each comprising one of the query sequences. The method may further include testing the one or more candidate antigen-binding proteins for a desired property. The method of this embodiment may have any of the features described in connection with the first or second embodiment. The one or more candidate antigen-binding proteins may be antibodies or fragments thereof. The sequences derived from the identified chain pairings may include sequences comprising the same CDRs but differing in framework regions, sequences containing one or more mutations compared to the identified chain pairing, and sequences containing one or more fragments of the identified chain pairing. Obtaining the candidate antigen-binding proteins may include identifying a coding sequence for the candidate antigen-binding protein and expressing the sequence in a suitable expression system (e.g., in a suitable host cell). The desired property may be a desired binding property (e.g., the ability to bind to one or more targets, the ability to bind to one or more targets with an affinity above one or more respective thresholds, etc.), a desired expression property (e.g., an increased expression level in one or more expression systems compared to a standard, an expression level above a predetermined level in one or more expression systems, a yield above a predetermined level in one or more expression systems, etc.), a desired stability (e.g., stability above a particular threshold under one or more conditions, etc.), or a combination thereof. The desired property may include the ability to bind to a predetermined target. Testing one or more candidate antigen-binding proteins for a desired property may include, for example, identifying one or more antigens to which the one or more candidate antigen-binding proteins bind by testing for binding to the one or more candidate antigens. Testing one or more candidate antigen-binding proteins for a desired property may include, for example, identifying one or more antigens to which the one or more candidate antigen-binding proteins are likely to bind by comparison with one or more antibodies with known targets. The antigen-binding proteins may be therapeutic antibodies, and the desired property may include binding of a therapeutic target. Antigen binding proteins are sometimes referred to herein as "immunity proteins."

[0031] Testing one or more candidate antigen-binding proteins for a desired property may include identifying the presence or absence of a desired phenotype in an organism (such as, for example, an animal model) or cell expressing the one or more candidate antigen-binding proteins. Identifying the presence of a desired phenotype may include expressing the one or more candidate antigen-binding proteins in one or more model cells (e.g., one or more cell lines) or organisms (such as, for example, one or more animal models). The method may further include optimizing the sequence of at least one of the one or more candidate antigen-binding proteins. Optimizing the sequence of the candidate antigen-binding proteins may be performed, for example, using any antibody optimization technique known in the art. Optimizing the sequence of the candidate antigen-binding proteins may be performed using information from sequence data in which strand pairings were identified, for example, by analyzing sequences similar to the input sequence in which strand pairings were identified. Methods for optimizing antigen-binding proteins are known in the art and include, inter alia, those described in Mason et al.

[2021] , Seeliger et al.,

[2015] , Warszawski et al.

[2019] , Hsiao et al.

[2019] , and Richardson et al.

[2021] . Any of these methods may be used within the context of the present invention.

[0032] The query sequence may include a heavy chain sequence (or a portion of a heavy chain sequence) of a known antibody. Thus, the first chain may be a heavy chain sequence or a portion of a heavy chain sequence of a known antibody. The query sequence may be obtained by bulk BCR sequencing of the heavy chain repertoire of one or more samples. The method may include obtaining the query sequence by bulk BCR sequencing of the heavy chain repertoire of one or more samples. The one or more samples may be from one or more subjects. The one or more subjects may be identified as having a desired characteristic, for example, a particular clinical phenotype or clinically relevant characteristic, such as a biomarker profile. For example, the one or more subjects may be resilient to a particular disease or condition. The disease or condition may be selected from cancer (e.g., breast cancer), neurodegenerative diseases (e.g., amyotrophic lateral sclerosis), and infectious diseases (e.g., COVID-19). The method may include identifying chain pairings (e.g., heavy-light pairings) of a plurality of query chain sequences (e.g., heavy chain sequences) selected from first (e.g., heavy) chain sequences identified in one or more samples, thereby obtaining a chain pairing set (e.g., heavy chain-light chain pairings). The method may further include identifying one or more targets by screening antibodies from the same source as the one or more samples against a plurality of candidate peptides. The plurality of candidate peptides may be selected based on the species from which the one or more samples originate. For example, the source of the one or more samples may be one or more human subjects, and an antibody repertoire from the same source as the one or more samples may be screened against a set of candidate peptides representative of the human peptidome to select a plurality of candidate peptides. Identifying antigens bound by one or more candidate antigen-binding proteins may include using one or more targets identified by screening antibodies from the same source as the one or more samples against a plurality of candidate peptides. The method may further include filtering the identified chain pairing sets based on one or more criteria. One or more criteria can be applied to the identity of the antigen or set of antigens to which a candidate antigen binding protein binds or is predicted to bind.Providing one or more query sequences may include providing a first query (e.g., heavy) chain sequence and a second query (e.g., heavy) chain sequence, and identifying a second (e.g., light) chain sequence for each of the one or more query sequences may include identifying one or more second (e.g., light) chain sequences for the first query sequence and one or more second (e.g., light) chain sequences for the second query sequence. The method may further include comparing the former second chain sequences with the latter corresponding chain sequences to identify one or more light chains that may be suitable for use as a common second (e.g., light) chain in a bispecific antibody comprising both first (e.g., heavy) chains. For example, one or more candidate second chain sequences may be identical or at least partially overlapping for the first and second queries, and one or more candidate second chain sequences that satisfy one or more criteria applied to the predictability of the candidate to form a functional pair with the first and second queries may be identified as suitable for use as a common second chain in a bispecific antibody. According to a fourth aspect, there is provided a method of providing a tool for predicting the likelihood of forming a functional antigen-binding protein or for identifying antigen-binding proteins comprising a chain pair, the method comprising: providing training data comprising training first and second / corresponding protein chain sequences from known antigen-binding proteins; and using the training data to train a deep learning model to receive as input one or more protein chain sequence pairs and generate a score (or information derived therefrom) indicative of the likelihood that the / each protein chain pair is part of a functional antigen-binding protein. The deep learning model comprises an encoder module and a classifier module. The method of this aspect may have any of the features described in relation to the first aspect.

[0033] The method may have any one or more of the following features: Providing training data may include providing first and second strand sequences for amp training. The first and second strand sequences for amp training may be referred to as pre-training data. The encoder module may include one or more encoders or decoders. The method may further include training a sequence-to-sequence model using the first and / or second (e.g., heavy and / or light) strand sequences for amp training and initializing an encoder of the encoder module using an encoder of the pre-training model. The encoder module may include one or two encoders, each of which may be an encoder-BERT model or a variant thereof, e.g., BERT, RoBERTa, RoFormer, SpanBERT (Joshi et al. 2020), DistilBERT, etc., or one or two decoders, each of which may be a decoder-only language model, e.g., Falcon, LlaMa, GPT3, and variants thereof. The first and second Transformer-based models may each include a RoBERTa model, a BERT model, or a RoFormer model. The encoder module may include a decoder-only model, such as a Falcon model (e.g., an embedding layer and a decoder layer of the Falcon model). The method may further include providing the training deep learning model to a user. The methods described herein are computer-implemented, unless otherwise indicated by context, for example, when obtaining, processing, or analyzing samples, or when producing, testing, or using molecules or compositions for any other purpose.

[0034] According to a fifth aspect, there is provided a system comprising a processor and a computer readable medium comprising instructions that, when executed by the processor, cause the processor to perform the method steps of any embodiment of any preceding aspect. The instructions may cause the processor to perform the method steps of any embodiment of the first to fourth aspects.

[0035] According to a sixth aspect, there is provided one or more computer-readable media comprising instructions that, when executed by one or more processors, cause the one or more processors to perform the method steps of any embodiment of any of the preceding method aspects. The instructions may cause the processor to perform the method steps of any embodiment of the first to fourth aspects.

[0036] According to a seventh aspect, there is provided a computer program product comprising instructions which, when executed by one or more processors, cause the one or more processors to perform the method steps of any embodiment of any of the preceding method aspects. The instructions may cause the processors to perform the method steps of any embodiment of the first to fourth aspects. [Brief explanation of the drawings]

[0037] BRIEF DESCRIPTION OF THE DRAWINGS [Figure 1] 1 is a flow chart that schematically illustrates a method for identifying strand pairs according to the present disclosure. [Figure 2] 1 illustrates one embodiment of a system for identifying strand pairs according to the present disclosure. [Figure 3] The training procedure for the exemplary deep learning model described herein is shown below. The method includes obtaining paired antibody sequences, pre-training an encoder model using this data and a masked language modeling (MLM) task, combining the pre-training encoder (AntiBERTa) with a cross-attention block, and fine-tuning the model including the pre-training encoder and the cross-attention block using paired antibody sequences (true and random pairs) to predict the probability of belonging to the true pair or random pair category. [Figure 4]The training procedure for the exemplary deep learning model described in Figure 3 is shown in more detail. A. Creation of training, validation, and test sets for the masked language model task. B. Setting up the pre-training procedure and feeding a "warm-up" model into subsequent steps. C. Overview of how the warm-up model is used as part of a classifier to enable prediction of the likelihood of strand pairs forming functional strand pairings. [Figure 5A] Schematic illustrating the task of masked language modeling (A) for pre-training an encoder model to learn antibody chain sequence features. [Figure 5B] Schematic illustration of the task of true-pairs versus random-pairs classification (B) used to fine-tune a classifier model including a pre-training encoder. [Figure 6A] 1 illustrates a schematic diagram of the architecture (A) of the deep learning model described herein. [Figure 6B] The architecture of the three cross-attention blocks used in such a model (B) is illustrated schematically. [Figure 7A] Figure 7 shows the results of an evaluation of the classification performance of the deep learning model described herein on an independent test dataset containing true pairs from single-cell sequencing data and decoy random pairs generated from the same dataset. A. Receiver operating characteristic (ROC) curve. [Figure 7B] B. Precision-recall (PR) curve. [Figure 8A]Figure 8 illustrates the architecture of heavy and light chains. A. Architecture of an Ig heavy chain. The approximate boundaries of the V, D, and J genes are marked, along with the junction boundaries. The segment of the V gene near the N-terminus is shown with a dotted boundary because the read lengths from many NGS methods are too short to cover this region and / or because some primers used for NGS insert only slightly within the V region. However, it is still possible to infer the V gene using the sequence within the solid boundary. [Figure 8B] Same as BA, but for the light chain. [Figure 9A] 9 illustrates a schematic diagram of the training of an exemplary deep learning model described herein. A. Pre-training of a generative transformer model (FAbCon) with a next token prediction task. [Figure 9B] B. FabCon small model architecture. [Figure 9C] C. Fine-tuning of the FabCon model to identify strand pairs. [Figure 9D] D. Architecture of the deep learning model described herein for identifying strand pairs. DETAILED DESCRIPTION OF THE INVENTION

[0038] Detailed Description of the Invention In describing the present invention, the following terminology will be utilized and is intended to be defined as indicated below.

[0039] B cell receptors are transmembrane proteins expressed on the surface of B cells. They contain a membrane-bound immunoglobulin molecule (also called an antibody) that recognizes the cognate antigen and a binding portion (also called an "antigen-binding subunit" or "membrane immunoglobulin" or "mIg") that contains a signal transduction moiety. The membrane-bound immunoglobulin molecule contains two immunoglobulin light chains and two immunoglobulin heavy chains and is identical to the corresponding secreted antibody, except for the integral membrane domain. The signal transduction portion is a heterodimer called Ig-α / Ig-β (CD79) linked to the immunoglobulin by disulfide bridges. Antibodies (Abs) or immunoglobulins (Igs) are immune proteins that belong to one of a defined set of isotypes (IgA, IgD, IgE, IgG, or IgM) and contain an antigen-binding site and a constant region that mediates interactions with other components of the immune system. In humans and most mammals, antibodies comprise four polypeptide chains connected by disulfide bonds: two identical heavy chains and two identical light chains. The light chains typically contain one variable domain, V L and one constant domain C L whereas the heavy chain typically consists of one variable domain V H and 3-4 constant domains C H 1. C H2, .... The variable domains form the antigen-binding region and can also be referred to as FV regions. Each variable domain contains three hypervariable regions called complementarity-determining regions (CDRs), which together form the antigen-binding site. The variable region of each immunoglobulin heavy or light chain is encoded by several pieces known as gene segments (subgenes): Ig heavy chains contain variable (V), diversity (D), and joining (J) segments, and Ig light chains contain V and J segments. Multiple copies of V, D, and J gene segments exist in the genome, and developing B cells assemble Ig variable regions by (nearly) randomly selecting and combining one V, one D, and one J gene segment (or one V and one J segment for light chains) in a process called V(D)J recombination. This process involves the formation of a double-stranded break between the required segments, which forms a hairpin loop that then joins together. The joining process is imprecise, resulting in the addition or deletion of nucleotides between the V and J (light chain) or V, DJ, and D, J (heavy chain) segments, generating a great deal of sequence diversity at the junctions between the segments (referred to as "junction sequences"). The V(D)J recombination process generates novel amino acid sequences in the antigen-binding region of Ig, thereby generating a vast diversity of antigen-recognition capabilities. As a result of this process, each Ig heavy chain variable region contains a V segment, a D segment, and a J segment, along with junction sequences spanning the junctions between these segments (as illustrated in Figure 8A). Similarly, each Ig light chain hypervariable region contains a V segment and a J segment, along with junction sequences spanning the junctions between these segments (as illustrated in Figure 8B). Within the variable region, CDR1 and CDR2 are found in the V segment, and CDR3 comprises part of the V, all of the D (in the case of the heavy chain), and part of the J segment.

[0040] T cell receptors are membrane-anchored proteins expressed on the surface of T cells. They contain a pair of protein chains that together form a binding moiety that recognizes cognate antigens. In mammals, they are expressed in a complex with the constant T cell coreceptor chain CD3, which includes the CD3γ chain, the CD3δ chain, and two CD3ε chains. The constant chains associate with the T cell receptor and the constant ζ chain, which together form a TCR complex that can generate a signal when antigen binds to the T cell receptor. TCRs are heterodimeric proteins containing two highly variable chains, α and β chains (in the majority of T cells) or alternative γ and δ chains (in a minority of T cells). Each chain contains two extracellular domains: a variable region (or variable domain) and a constant region (or constant domain, proximal to the cell membrane), a transmembrane region, and a short intracytoplasmic tail. The variable regions together bind peptides (antigens) within the context of MHC (major histocompatibility complex) molecules, in the case of αβ TCRs. Each variable domain contains three hypervariable regions called complementarity-determining regions (CDRs, designated CDR1, CDR2, and CDR3 in each chain), which together form the antigen-binding site. TCRs are members of the immunoglobulin superfamily, which also includes BCRs and antibodies. In a process similar to that described above, the variable region of each TCR chain is encoded by several pieces known as gene segments (subgenes): β and δ chains contain variable (V), diversity (D), and joining (J) segments, and α and γ chains contain V and J segments. Multiple copies of V, D, and J gene segments exist in the genome, and developing T cells assemble TCR chain variable regions by (nearly) randomly selecting and combining one V, one D, and one J gene segment (or one V and one J segment for α / γ) in a process called V(D)J recombination. This process involves the formation of a double-stranded break between the required segments, which forms a hairpin loop and is then ligated together.The joining process is imprecise, resulting in the addition or deletion of nucleotides between the V and J (α / γ chain) or V, DJ, and D and J (β / δ chain) segments, resulting in a great deal of sequence diversity at the junctions between the segments (referred to as "junction sequences"). The V(D)J recombination process generates novel amino acid sequences in the antigen-binding region of the TCR, thereby generating a vast diversity of antigen recognition capabilities. As a result of this process, each β / δ chain variable region contains a V segment, a D segment, and a J segment, along with junction sequences spanning the junctions between these segments. Similarly, each α / γ chain hypervariable region contains a V segment and a J segment, along with junction sequences spanning the junctions between these segments. Within the variable region, CDR1 and CDR2 are found in the V segment, and CDR3 comprises part of the V, all of the D (in the case of heavy chains), and part of the J segment.

[0041] As used herein, a "variable chain" of an antigen-binding protein (also referred to herein simply as a "chain") refers to the antigen-binding protein chain involved in antigen recognition, or a portion thereof containing at least a portion of the variable region of that chain. The variable chain comprises the variable region responsible for the diverse antigen-recognition repertoire within the antigen-binding protein. The variable chain can be a BCR heavy or light chain, an antibody heavy or light chain, a TCR α or β chain, a TCR γ or δ chain, or a portion of any such chain containing at least a portion of one or more variable regions within these chains.

[0042] The B cell receptor repertoire (or corresponding antibody repertoire) present in a sample can be studied using sequencing approaches. As explained above, two main sequencing approaches are used: single B cell sequencing and bulk B cell population sequencing. Because the BCR signaling portion and the antigen-binding portion transmembrane domain are not variable, these techniques focus on the common portions between the B cell repertoire and the corresponding antibody repertoire. Therefore, in the context of this disclosure, references to BCR sequences, BCR repertoires, BCR heavy chain sequences, BCR light chain sequences, and portions of any thereof are used interchangeably with corresponding antibody sequences, antibody repertoires, antibody heavy chain sequences, antibody light chain sequences, and corresponding portions. For example, reference to sequencing the BCR heavy chain variable region is equally equivalent to sequencing the corresponding antibody heavy chain variable region, and the two terms may be used interchangeably. The term "antigen-binding protein" is used herein to mean a BCR protein, a TCR protein, an antigen-binding portion of a BCR protein, an antibody, or any portion thereof that maintains the antigen-binding properties of the original BCR protein, TCR protein, or antibody. It should be noted that the repertoire of antibodies circulating in an individual's blood may not match the B cell receptor repertoire present in a sample at the same time. This is because antibodies produced by B cells that are no longer present in the individual (e.g., because they have died) may be present in the sample. Thus, the term "corresponding antibody repertoire" refers to the repertoire of antibodies that can be expressed by B cells present in the sample, rather than the repertoire of antibodies (proteins) actually present in the sample.

[0043] Single B-cell sequencing can maintain correspondence between heavy and light chain sequences. Two major approaches can be used to achieve this. The first approach is physical linking of VH and VL [DeKosky et al., 2016]. The second approach is cell barcoding (e.g., as provided by 10x Genomics) [King et al., 2021]. The physical linking approach has higher throughput than the cell barcoding approach, but it is more difficult to recover the entire sequence. In contrast, cell barcoding has lower throughput but allows for easier recovery of the entire sequence. Regardless of the approach, single B-cell sequencing is limited (to various degrees) in terms of throughput, as explained above. Some single B-cell sequencing technologies are further limited in terms of the length of the sequences recovered. As a result, BCR / antibody sequences identified using some single B-cell sequencing methods may be limited to studying a single CDR region, e.g., CDR3, in both the heavy and light chains (in other words, flanking V and J segments may be identified, but they may not be sequenced sufficiently to yield the sequences of CDR1 and CDR2 in the V segment). In other words, datasets from single B-cell sequencing methods may vary in the extent to which heavy and light chains are identified. Within the sequenced region, it may be impractical to sequence (or record) every single base of the V(D)J segment; thus, sequencing efforts may be focused on obtaining sufficient information to identify the junction sequences and the V, D, and J genes. As a result, such methods can provide information including the identities of the V, D, and J segments for the heavy chain (e.g., in the form of a V / D / J gene segment identifier), the sequence of the joining segment in the heavy chain, the identities of the V and J segments for the light chain (e.g., in the form of a V / J gene identifier), and the sequence of the joining segment in the light chain. The identities of each segment can be used to retrieve the corresponding germline sequence from a database.However, if the data only contains the identity of each segment, it may not capture any mutations (e.g., somatic mutations) that may exist in a particular chain compared to a reference germline sequence. In contrast, bulk B cell population sequencing does not maintain the pairing between heavy chain sequences and light chain sequences, but is less limited in terms of sequencing capacity within the heavy chain and light chain repertoires, respectively (particularly, the sequencing depth of the BCR repertoire). Bulk B cell population sequencing may include sequencing the heavy chain repertoire, the light chain repertoire, or both of the B cell population. However, as mentioned above, even when both the light chain and heavy chain repertoires are sequenced, due to the bulk nature of the process, it is not possible to maintain pairing information during the sequencing process. Such sequencing may generate information as sparse as that obtained with single-cell B sequencing, or more detailed information, e.g., the entire CDR sequence, the sequence of multiple CDRs, the entire variable region sequence, or the entire variable region sequence and enough constant region sequence to determine the isotype of the sequence. Similar considerations apply to sequencing T cell repertoires. In particular, many of the processes and limitations described above in the context of testing B cell receptors and antibodies (and in particular in the context of sequencing these repertoires) also apply to T cell repertoires.

[0044] As used herein, the term "variable chain sequence" encompasses the terms "heavy chain sequence," "light chain sequence," "α chain sequence," "β chain sequence," "γ chain sequence," and "δ chain sequence," and refers to any information obtainable from B-cell or T-cell sequencing technology within the combination of one or more gene segment identifiers and / or junction sequences at one end to the entire chain sequence at the other end. Specifically, the terms "heavy chain sequence" and "light chain sequence" refer to any information obtainable from B-cell sequencing technology within the combination of one or more gene segment identifiers and / or junction sequences at one end to the entire chain sequence at the other end. Furthermore, the terms "variable chain sequence," "heavy chain sequence," "light chain sequence," "α chain sequence," "β chain sequence," "γ chain sequence," and "δ chain sequence" interchangeably refer to amino acid sequences or corresponding nucleic acid coding sequences. Similarly, a variable chain pairing or pair (e.g., heavy chain-light chain pairing or pairing) refers to a combination of heavy and light chain sequences, α and β chain sequences, or γ and δ chain sequences as defined herein, each within the combination of one or more gene segment identifiers and / or junction sequences at one end to the entire chain sequence at the other end.

[0045] Within the context of providing a desired antibody or antigen-binding protein, such as a therapeutic antibody, the term "antibody" (Ab) includes monoclonal antibodies, polyclonal antibodies, multispecific antibodies (e.g., bispecific antibodies), and antibody fragments (e.g., scFvs) that exhibit a desired biological activity and that comprise a heavy chain-light chain pairing identified as described herein or a heavy chain-light chain pairing derived from a heavy chain-light chain pairing identified as described herein (e.g., by further optimization, affinity maturation, etc.).

[0046] As used herein, a "sample" can be a cell or tissue sample, biological fluid, or extract (e.g., a DNA or RNA extract obtained from a subject) from which B cell genomic material (e.g., RNA or DNA) can be obtained for genomic analysis, for example, by sequencing (e.g., whole genome sequencing, whole exome sequencing, targeted / capture sequencing, RNA-seq, etc.). A sample can be a cell, tissue, or biological fluid sample (e.g., a biopsy) obtained from a subject. Such a sample may be referred to as a "subject sample." Specifically, a sample can be a blood sample, lymph node sample, spleen sample, or tumor sample, or a sample derived therefrom (e.g., by B cell purification, T cell purification, RNA extraction, etc.). As used herein, the terms "genomic material," "genomic sequencing," etc., encompass both material / sequences present in the genome and transcriptome of a sample, unless the context dictates otherwise. The sample may be freshly obtained from a subject, or may have been processed and / or stored (e.g., frozen, fixed, or subjected to one or more purification, enrichment, or extraction steps) prior to genomic analysis. The sample may be a cell or tissue culture sample. Thus, a sample as used herein may refer to any type of sample that includes B cells or genomic material derived therefrom, whether from a biological sample obtained from a subject or from a sample obtained from a cell line, etc. The sample is preferably from a mammal (e.g., a mammalian cell sample, a sample from a mammalian subject such as a cat, dog, horse, donkey, sheep, pig, goat, cow, mouse, rat, rabbit, guinea pig, etc.), preferably from a human (e.g., a human cell sample, a sample from a human subject, etc.).Furthermore, samples may be transported and / or stored, and collection may occur at a location remote from the location of sequence data acquisition (e.g., sequencing), and / or any computer-implemented method steps described herein may occur at a location remote from the location of sample collection and / or remote from the location of genomic data acquisition (e.g., sequencing) (e.g., computer-implemented method steps may be performed using a network-connected computer, e.g., using a "cloud" provider).

[0047] The term "sequence data" refers to information indicative of the presence of genomic (DNA or RNA) or proteomic material in a sample having a particular sequence. Thus, sequence data can include one or more nucleotide sequences and / or one or more amino acid sequences. Such information can be obtained using sequencing technologies, such as next-generation sequencing (NGS), such as whole exome sequencing (WES), whole genome sequencing (WGS), whole transcriptome sequencing (RNA-seq), or capture genomic locus sequencing (targeted or panel sequencing). When using NGS technologies, sequence data can include a count of the number of sequencing reads having a particular sequence. Sequence data can be mapped to a reference sequence, e.g., a reference genome, using methods known in the art, such as Bowtie (Langmead et al., 2009). Thus, the count of a sequencing read or equivalent non-digital signal can be associated with a particular location or locus (where "location" refers to the location in the reference genome or transcriptome to which the sequence data is mapped). Furthermore, a position may contain mutations, and in that case, the count of sequencing reads or equivalent non-digital signals may be associated with each possible variant (also referred to as an "allele") at a particular position. The process of identifying the presence of a mutation at a particular position in a sample is called "variant calling" and can be performed using methods known in the art (e.g., general-purpose NGS variant callers, such as GATK HaplotypeCaller, gatk.broadinstitute.org / hc / en-us / articles / 360037225632-HaplotypeCaller, tools specifically designed for immune sequences, such as IgBLAST, www.ncbi.nlm.nih.gov / igblast / , [Ye et al., 2013], etc.). Genomic sequence data can be converted to amino acid sequences by translating the coding regions in silico (either directly from the mRNA sequence or from identified coding regions in the genomic sequence), as is known in the art.

[0048] As used herein, "treatment" means reducing, alleviating, or eliminating one or more symptoms of the disease being treated compared to the symptoms before treatment. "Prevention" (or prophylaxis) means delaying or preventing the onset of disease symptoms. Prevention may be absolute (so that the disease does not occur at all) or may only be effective in some individuals or for a limited amount of time. The compositions described herein may be pharmaceutical compositions that additionally comprise a pharmaceutically acceptable carrier, diluent, or excipient. The pharmaceutical composition may optionally comprise one or more additional pharmaceutically active polypeptides and / or compounds. Such formulations may, for example, be in a form suitable for intravenous infusion.

[0049] As used herein, the term "computer system" includes hardware, software, and data storage devices for implementing the systems or performing the methods according to the above-described embodiments. For example, a computer system may include a central processing unit (CPU), a graphical processing unit (GPU), input means, output means, and data storage, which may be embodied as one or more connected computing devices. Preferably, a computer system includes a computing device with a display or a display for providing a visual output display (e.g., in business process design). Data storage may include RAM, a disk drive, or other computer-readable media. A computer system may include multiple computing devices connected by a network and capable of communicating with each other over the network. It is expressly contemplated that a computer system may consist of or include a cloud computer. The term "processor" encompasses any processing unit or combination of processing units, particularly including a CPU and a GPU. As used herein, the term "computer-readable medium" includes, but is not limited to, any one or more non-transitory media directly readable and accessible by a computer or computer system. Media may include, but are not limited to, magnetic storage media such as floppy disks, hard disk storage media, magnetic tape, optical storage media such as optical disks and CD-ROMs, electrical storage media such as memory, including RAM, ROM, and flash memory, and hybrids and combinations of the above, for example, magnetic / optical storage media.

[0050] Identification of variable chain pairs The present disclosure provides methods for identifying functional variable chain pairs and / or predicting whether a variable chain pair is likely to function. An exemplary method is described with reference to FIG. 1. FIG. 1 illustrates an embodiment in which a heavy or light chain sequence of a B cell receptor or antibody is used to identify a heavy chain-light chain pair. In other words, FIG. 1 illustrates an embodiment in which the variable chain sequences are the heavy and light chains of an antibody from a BCR / BCR. However, the method described with reference to FIG. 1 is applicable to embodiments in which a TCR α, β, γ, or δ chain sequence is used to identify an αβ (when the query chain or chain pair is an α chain or β chain or an αβ chain pair) or γδ (when the query chain or chain pair is a γ chain or δ chain or a γδ chain pair) chain pair. In optional step 10, a sample containing B cell genomic material (typically in the form of RNA, in which case RNA encoding BCRs expressed by the cells from which the B cell genomic material is derived can be extracted and sequenced) can be obtained from a subject. Similarly, a sample containing T cell genomic material may be used in embodiments in which TCR chain pairs are identified. In optional step 12, the BCR repertoire in the sample may be sequenced using bulk BCR sequencing. This may include sequencing the heavy chain BCR repertoire in the sample and / or the light chain BCR repertoire in the sample. Similarly, the TCR repertoire in the sample may be sequenced using bulk TCR sequencing. This may include sequencing the β chain repertoire and / or the α chain repertoire in the sample. In step 14, a query chain sequence or sequence pair is provided. In exemplary embodiments, the query sequence is a heavy chain sequence. In other embodiments, the query chain sequence may be a light chain sequence. In other embodiments, the query may be a chain sequence pair. Providing a query sequence may include selecting a query sequence in step 14A, such as one of the heavy chain sequences sequenced in step 12. Providing a query sequence pair may include selecting in step 14A a query sequence, such as one of the heavy chain sequences sequenced in step 12, and a query sequence, such as one of the light chain sequences sequenced in step 12 or obtained from a database or other source.Providing a query sequence or sequence pair may include providing a sequence (or a pair of sequences each comprising) comprising a V gene sequence or identifier, a J gene sequence or identifier, and a junction sequence in step 14B. For example, step 14B may include extracting the V gene sequence or identifier, the J gene sequence or identifier, and the junction sequence from a bulk BCR sequencing dataset to serve as the selection sequence (or each of the sequence pairs). In the context of TCR pairing, similar steps may be performed, for example, using a query β chain sequence. When providing a single query sequence, step 14 may include step 14C of selecting one or more candidate sequences for pairing with the query sequence. These may be obtained from a database, a computing device (including, for example, by simulation), or a user interface. In optional step 16, a deep learning model is provided, where the deep learning model is configured to receive as input a query variable chain sequence pair (which may include a query sequence and candidate sequences for pairing, or a query sequence pair) and generate as output a score indicative of the likelihood that the sequence pair will form a functional pair. In exemplary embodiments, the query sequence pair is a heavy chain-light chain sequence pair, and thus the deep learning model is configured to receive as input a query heavy chain sequence and a query light chain sequence and generate as output a score indicative of the likelihood that the query heavy chain sequence and the query light chain sequence form a functional pair. In other embodiments, the query chain sequence may be a heavy chain sequence, and thus the deep learning model may be configured to receive as input a query heavy chain sequence and a plurality of candidate light chain sequences and generate as output a respective score indicative of the likelihood that the query heavy chain sequence and each of the candidate light chain sequences form a functional pair, or information derived therefrom, such as a ranking of a plurality of pairs or a plurality of candidate light chain sequences based on their respective scores.In yet other embodiments, the query chain sequence may be a β chain sequence (or an α, δ, or γ chain sequence), and thus the deep learning model may be configured to receive as input the query β chain sequence (or an α, δ, or γ chain sequence) and one or more candidate α chain sequences (or β, γ, or δ chain sequences), and generate as output a score or information derived therefrom (e.g., a ranking of multiple candidate pairs) that indicates the likelihood that the query β chain sequence and each of the candidate α chain sequences (or each of the corresponding candidate β, γ, or δ chain sequences) will form a functional pair. The deep learning model may have been pre-trained using training variable chain sequences from known variable chain pairs, e.g., in exemplary embodiments, training heavy chain and light chain sequences from known heavy chain-light chain pairs. Thus, providing the deep learning model may simply involve accepting or otherwise retrieving a training deep learning model from a computer-readable medium, such as a memory associated with a processor executing the method. Training a deep learning model is described in more detail below. Alternatively, the deep learning model may have been partially trained as part of the method using training variable chain sequences from known heavy chain-light chain pairs, e.g., in an exemplary embodiment, using training heavy and light chain sequences from known heavy chain-light chain pairs, and optionally using training variable chain sequences from known amp heavy and light chains.

[0051] In step 18, sequence pairs including a query strand sequence and candidate matched strand sequences or query strand sequence pairs are provided to a deep learning model. Step 18 may include optional step 18A of encoding the sequences in the sequence pairs using a predetermined encoding scheme. The encoding scheme used may be predefined based on the content of the training variable strand sequences (e.g., heavy and light) used to train the deep learning model. Step 18 may include optional step 18B of selecting a sequence pair associated with the highest likelihood (e.g., highest score / highest probability or highest ranking) of forming a functional pair among multiple sequence pairs. In optional step 20, results of any of the above steps (particularly step 18) may be provided to a user, e.g., via a user interface. These results may be used, for example, to provide therapeutic antibodies, as described further below. The method may be repeated for multiple query sequences, which may include repeating steps 14-18.

[0052] The deep learning model includes an encoder module. The encoder module refers to a machine learning model trained to receive sequence data (protein chain sequences) as input and generate an embedded representation of the sequence data as output, or to generate a decoded or generated sequence from the learned embedded representation of the sequence data as output. Therefore, the sequence encoder can be an encoder model for a bidirectional encoder-only model, an encoder model for an encoder-decoder model, or a decoder model for a decoder-only model, such as a generative model. For example, generative models with architectures such as LlaMa (Touvron et al., 2023), Falcon LLM (falconllm.tii.ae / ), and GPT (e.g., GPT-3, Brown et al., 2020) can be used. Alternatively, a transformer-encoder-based architecture such as AntiBERTa (see Examples and Leem et al., 2022) can be used. When a decoder model is used, the embedded representation can be obtained as an embedding from the final decoder layer of the decoder. The encoder module may be referred to herein as an "encoder," but its architecture is not limited to that of an encoder. The embedding representation is typically a representation in a latent space configured to enable the original sequence data to be reconstructed from the embedding representation. The encoder module is a deep learning model. The encoder module may be trained as part of a sequence-to-sequence model. The encoder module may be a transformer-based encoder or decoder. A deep learning model including a sequence encoding module may be trained using a masked language modeling (MLM) task or a causal language modeling (CLM, next token prediction) task. The former may be particularly used for encoder-only models and encoder-decoder models, and the latter may be particularly used for generative models.

[0053] Training of the deep learning model is now described with reference to optional steps 10'-16'. In step 10', training data is provided that includes at least training variable chain sequences from known variable chain pairs. In exemplary embodiments, the training data includes heavy and light chain sequences from known heavy and light chain pairs. The training data may include at least 20,000 training chain pairs, at least 30,000, at least 40,000, at least 50,000, at least 60,000, at least 70,000, at least 80,000, at least 90,000, at least 100,000, at least 120,000, or at least 150,000 training chain pairs. In embodiments related to B cell receptors / antibodies, the training data may include at least 80,000, at least 100,000, at least 120,000, or at least 150,000 training heavy and light chain sequence pairs. In embodiments, the training data includes at least 1 million, 1.5 million, or 2 million paired sequences. The training data may further include Ampere training sequences, which in exemplary embodiments are heavy and light chain sequences (although they may be any first and second chains described herein). The Ampere training chain sequences may be referred to as "pre-training data." In such cases, the paired training data may be referred to as "fine-tuning data." In embodiments, the paired training data is also used for pre-training. Thus, the training data may include fine-tuning training data (including paired chain sequences, specifically paired heavy and light chain sequences in exemplary embodiments) and pre-training data (including Ampere chain sequences, specifically Ampere heavy and light chain sequences in exemplary embodiments, and optionally paired training sequences).The pre-training data may include at least 100,000, at least 200,000, at least 300,000, at least 400,000, at least 500,000, at least 600,000, at least 700,000, at least 800,000, at least 900,000, at least 1 million (or at least 500, 1000, 1500, 2000, 2500, 3000, 3500, or 40 million) individual training strand sequences of the first type and / or the corresponding type. In embodiments, the training data includes at least 500, 600, or 700 million individual sequences. The training data may include at least 1 million (or at least 500, 1000, 1500, 2000, 2500, 3000, 3500, or 40 million) amp-training heavy chain sequences and at least 1 million (or at least 500, 1000, or 15 million) amp-training light chain sequences. As one of skill in the art will appreciate, the amount of training and / or pre-training data may be limited by the amount of suitable data available and may change as more data becomes available. The pre-training and / or training data may include simulated data. The simulated data may include paired or amp-trained protein sequences obtained using methods for simulating antigen-binding protein sequences, such as immuneSIM (Weber et al., 2020) and / or ProGen2 (Nijkamp et al., 2022). Pre-training data may be collected and / or used for pre-training as described in Leem et al., 2022. For example, Leem et al., 2022 used Ampere chain sequences containing approximately 42 million heavy chains and 15 million light chains. More data can be advantageously used if available. Furthermore, the amount of data available can depend on the particular use case, e.g., the identity of the first and corresponding chain sequences (e.g., more data may be available for the αβ TCR than the rarer γδ TCR), the criteria used when filtering the data (see step 12'), etc.The provided numbers may be applied to the data before and / or after any filtering is applied. In step 12', the training data is filtered. For example, the training data may be filtered to eliminate any chains that contain several amino acids within any predetermined chain region outside a predetermined length range. For example, the pre-training data may be filtered to eliminate chain sequences that do not contain at least 20 amino acids before the CDR1 region, sequences that do not contain at least 10 amino acids after the junction sequence, sequences that do not contain 5-12 residues in the CDR1 region, sequences that do not contain 1-10 residues in the CDR2 region, and / or sequences that do not contain 5-38 residues in the CDR3 region. As another example, the training data may be filtered based on any characteristic of the data, including, for example, the cell type or organism from which the data originated, whether the data is from a naive library, or whether the data is from a subject immunized with a specific antigen, etc. In other words, the training data may be filtered to ensure that the training data contains only data with one or more characteristics of interest (selection filter) or does not contain any (exclusion filter). In step 14', negative training data including non-native strand pairs, also referred to as "negative training data," may be provided. Non-native strand pairs may be obtained by randomly reassigning strands within pairs in a paired dataset. Alternatively, non-native strand pairs may be obtained by randomly pairing one or more strands from an Ampere dataset that includes both strand types (e.g., Ampere heavy chains and Ampere light chains). For example, a heavy chain from bulk heavy chain sequencing data may be randomly paired with a light chain from bulk light chain sequencing data. Instead of or in addition to using randomly generated negative pairs, negative pairs may be selected as pairs known to be non-functional, such as from prior experiments. Pairs known to be non-functional may include one or more of: pairs that failed to express in one or more prior experiments; pairs that did not have one or more desired functional properties in one or more prior experiments (e.g., lack of binding to a target);

[0054] In step 16', the training data is used to train a deep learning model that receives as input (in an exemplary embodiment) a query pair including a heavy chain sequence and a light chain sequence, and generates as output a score indicative of the likelihood that the pair is functional (e.g., in an exemplary embodiment, a probability that the pair is functional) or information derived therefrom (e.g., a ranking of multiple input pairs based on the score indicative of the likelihood that each pair is functional). Training the deep learning model may include first using the training strand sequences to train a model including an encoder module, such as an encoder-decoder-based model or a bidirectional encoder model (also referred to as a "sequence-to-sequence" model), or a decoder-based model (e.g., a generative language model), to obtain a pre-training encoder model, and using the pre-training model to initialize the encoder module of a deep learning model that includes an encoder module and a classification module trained for paired sequence prediction (fine-tuning). Alternatively, training the deep learning models may include training first and second sequence-to-sequence models using a first type and a second type of amp training strand sequences, respectively (in an exemplary embodiment, the first type is a heavy chain and the second type is a light chain), and initializing each encoder module of the deep learning models with a pre-training encoder module of each of these models.Alternatively, training the deep learning model may include training a sequence-to-sequence or generative language model using concatenated training strand sequences of a first type and a second type (in exemplary embodiments, the first type is a heavy chain and the second type is a light chain), and initializing an encoder module of the deep learning model that receives the concatenated sequence pair as input using a pre-training encoder module of the model (which may be an encoder or a decoder depending on the model architecture). Training the deep learning model may include, when strand sequences from the known strand pair do not include the full-length sequences of the variable regions, imputing missing sequence information to obtain training sequences that include the full-length sequences of the variable regions of the second type strand and / or the first type strand (e.g., in exemplary embodiments, the light chain and / or the heavy chain). Training the sequence-to-sequence model may include training the model using MLM. Training the generative language model may include training the model using CLM. Step 16' may include defining one or more encoding schemes for the training data by deriving a vocabulary for encoding the training chain sequences, in an exemplary embodiment, specifically a vocabulary for encoding the training heavy chain sequences and a vocabulary for encoding the training light chain sequences. The encoding scheme may encode every single amino acid as a separate token, or may encode a subset of the sequences as individual tokens (e.g., using complete region, k-mers within a region, or byte pair encoding). Defining the encoding scheme may include excluding from the vocabulary constructed based on the content of the training chain sequences any tokens that are used less than a predetermined threshold (e.g., 2) in the training data.The inventors have found that, at least in the case of predicting antibody / BCR light-heavy chain pairing, an encoding scheme that represents each amino acid as a token is practical when training data is available at sufficient sequence resolution and in a sufficiently large amount to constrain models using such granular representations.

[0055] system FIG. 2 illustrates one embodiment of a system for identifying a variable chain pair for an input variable chain or predicting whether an input variable chain pair is likely to function according to the present disclosure. The system includes a computing device 1, which includes a processor 101 and a computer-readable memory 102. In the illustrated embodiment, the computing device 1 also includes a user interface 103, illustrated as a screen, but may include any other means of communicating information to a user, for example, via an audible or visual signal. The computing device 1 is communicatively connected, for example, via a network 6, to a sequence data acquisition means 3, such as a sequencing machine, and / or to one or more databases 2 that store sequence data. The one or more databases 2 may additionally store other types of information, such as reference sequences, parameters, etc., that can be used by the computing device 1. The computing device may be a smartphone, server, tablet, personal computer, or other computing device. The computing device is configured to implement a method for identifying a variable chain pair or predicting whether a variable chain pair is likely to function (preferably a heavy chain-light chain pair or an αβ chain pair, advantageously a heavy chain-light chain pair), as described herein. In an alternative embodiment, computing device 1 is configured to communicate with a remote computing device (not shown) and is itself configured to implement the method of identifying variable chain pairs or predicting whether a variable chain pair is likely to be functional, as described herein. In such a case, the remote computing device may also be configured to send results of the method to the computing device. Communication between computing device 1 and the remote computing device may be via a wired or wireless connection and may take place over a local or public network, such as over the public Internet or WiFi. Sequence data acquisition means 3 may be wired to computing device 1 or may communicate via a wireless connection, for example over network 6 as illustrated.The connection between the computing device 1 and the sequence data acquisition means 3 can be direct or indirect (e.g., via a remote computer). The sequence data acquisition means 3 is configured to acquire sequence data from a nucleic acid sample, for example, a genomic DNA sample or RNA sample extracted from B cells or T cells purified from a fluid and / or tissue sample (e.g., peripheral blood, spleen, lymph node, tumor tissue, or any other type of sample containing B cells or T cells). In some embodiments, the sample may have been subjected to one or more pre-processing steps, such as DNA / RNA purification, fragmentation, library preparation, targeted sequence capture (e.g., exon capture and / or panel sequence capture). Any sample preparation process suitable for use in determining B cell receptor sequences or repertoires may be used within the context of the present invention. The sequence data acquisition means 3 is preferably a next-generation sequencer. The sequence data acquisition means 3 may be directly or indirectly connected to one or more databases 2 in which sequence data (raw or partially processed) can be stored.

[0056] Applicable The above methods apply to any situation in which it is desirable to identify antibodies or BCRs that are likely to bind to their target from limited information on heavy chains, light chains, amplicon heavy and light chains, or portions thereof (e.g., V genes, J genes, junction sequences, etc.). This is often the case in the context of the antibody therapeutic discovery process. Antibody therapeutics have been shown to be a successful approach for a wide range of diseases, from neurodegenerative diseases to cancer. As such, the approaches described herein find use in the context of providing therapeutics in each of these clinical settings. Furthermore, the methods described herein can be used to identify potentially functional antibodies or BCRs from any input heavy / light chains or portions thereof, whether the input information is generated de novo for a specific purpose (e.g., from a patient or sample identified as having a desired phenotype) or from existing / legacy datasets (e.g., to mine or remine an existing dataset to discover new therapies or identify immune proteins that may explain why a particular clinical phenotype persists).

[0057] Thus, the present invention also provides methods for providing an antibody therapeutic, the method comprising identifying a heavy chain-light chain pairing using any of the methods described herein, or the therapeutic is derived (e.g., by further optimization, mutation, etc.) from a heavy chain-light chain pairing identified using any of the methods described herein. The heavy chain-light chain pairing can be obtained for input heavy chain sequences obtained by bulk BCR sequencing of the heavy chain repertoire in one or more samples. The heavy chain-light chain pairing can be obtained for input heavy chain sequences obtained by bulk BCR sequencing of the heavy chain repertoire in one or more samples and one or more candidate light chain sequences obtained by bulk BCR sequencing of the light chain repertoire in one or more samples (in which case the one or more samples can be the same or different samples as those used to analyze the heavy chain BCR repertoire). The one or more samples can be from one or more subjects. The one or more subjects can be identified as having a desired characteristic, such as a particular clinical phenotype or clinically relevant characteristic, such as a biomarker profile. For example, one or more subjects may be resilient to a particular disease or condition. The disease or condition may be selected from cancer (e.g., breast cancer), a neurodegenerative disease (e.g., amyotrophic lateral sclerosis), or an infectious disease (e.g., COVID-19). The method may include obtaining a heavy chain-light chain pairing set by identifying heavy chain-light chain pairings for a plurality of input heavy chain sequences selected from heavy chain sequences identified in one or more samples. The method may further include identifying a target (or putative target or target set) of each heavy chain-light chain pairing in the heavy chain-light chain pairing or heavy chain-light chain pairing set. The method may further include identifying one or more targets by screening antibodies from the same source as the one or more samples against a plurality of candidate peptides. The plurality of candidate peptides may be selected based on the species of origin of the one or more samples.For example, the source of one or more samples can be one or more human subjects, and an antibody repertoire from the same source as the one or more samples can be screened against a set of candidate peptides representative of the human peptidome. Identifying a target (or putative target or set of targets) for each heavy-chain-light chain pairing in a heavy-chain-light chain pairing or set of heavy-chain-light chain pairings can include using one or more targets identified by screening antibodies from the same source as the one or more samples against a plurality of candidate peptides. The method can further include filtering the set of heavy-chain-light chain pairings based on one or more criteria. The one or more criteria can apply to the identity of the putative target or set of targets for which the heavy-chain-light chain pairings were identified. The method can further include obtaining an antibody or fragment thereof comprising the identified heavy-chain-light chain pairing or a heavy-chain-light chain pairing derived from the identified heavy-chain-light chain pairing. Obtaining the antibody or fragment thereof can include identifying a coding sequence for the antibody or fragment thereof and expressing the sequence in a suitable expression system (e.g., in a suitable host cell). The method may further include identifying one or more antigens to which the antibody or fragment thereof binds by testing binding to one or more candidate antigens. The method may further include optimizing the sequence of the antibody or fragment thereof. Optimizing the sequence of the antibody or fragment thereof may be performed using any antibody optimization technique known in the art. Optimizing the sequence of the antibody or fragment thereof may be performed using information from sequence data in which heavy chain-light chain pairings were identified, for example, by analyzing sequences similar to the input sequence in which heavy chain-light chain pairings were identified.

[0058] The present invention also provides a method of providing an immunotherapeutic composition, the method comprising identifying a heavy chain-light chain pairing as described herein and generating an immunotherapeutic composition comprising an antibody comprising the heavy chain-light chain pairing or an antibody derived from the heavy chain-light chain pairing (e.g., by further optimization, mutation, etc.).

[0059] The methods described herein can also be used in the context of providing bispecific antibodies. For example, the methods described herein can be used to identify light chains that may be suitable for pairing with two different heavy chains of interest. Thus, the present invention also provides methods for providing bispecific antibodies, wherein the method comprises identifying a consensus light chain pairing for each of two heavy chains using any of the methods described herein, or wherein the bispecific antibody is a combination of two heavy chains and a consensus light chain derived (e.g., by further optimization, mutation, etc.) from a heavy chain-light chain pairing identified using any of the methods described herein. In such embodiments, it can be advantageous to use a deep learning model to predict scores indicative of the likelihood of functional pairing for multiple candidate light chain or heavy chain sequences for each query sequence. For example, a deep learning model can be used to predict the score / likelihood of a first query heavy chain forming a functional pair with each of multiple first candidate light chain sequences, and to predict the score / likelihood of a second query heavy chain sequence forming a functional pair with each of multiple second candidate light chain sequences. The plurality of first and second candidate light chain sequences are preferably identical or at least partially overlapping. The predicted scores / likelihoods for the plurality of first and second light chain sequences can then be compared to identify one or more light chains that may be suitable for use as the common light chain in a bispecific antibody comprising both heavy chains. For example, a candidate light chain that is relatively likely to form a functional pair with both heavy chains can be used. For example, the candidate light chains can be ranked based on the sum of their scores / likelihoods (or any other combination metric that combines both likelihoods) of forming a functional pair with the first and second heavy chains.

[0060] The methods described herein can also be used in the context of antibody optimization. For example, the methods described herein can be used to identify heavy chain-light chain pairings that have one or more advantageous properties (e.g., improved functionality, developability, etc.) compared to the original pairing, e.g., for the heavy chains of the pairing. For example, a paired heavy chain can be paired with multiple candidate light chain pairs, and a score indicative of the likelihood of the candidate to form a functional pair with the heavy chain can be determined. The candidate pairs can then be ranked or otherwise prioritized based on these scores / likelihoods and, optionally, other criteria applied to functionality or developability. Thus, the present invention also provides methods for providing improved antibodies, wherein the method comprises identifying a heavy chain-light chain pairing from an input heavy chain (or light chain) of the original antibody using any of the methods described herein, or wherein the improved antibody is a heavy chain-light chain pairing derived (e.g., by further optimization, mutation, etc.) from a heavy chain-light chain pairing identified using any of the methods described herein. The methods described herein also apply to any situation in which it is desirable to identify a TCR likely to bind to its target from information limited to the β chain, α chain (or less commonly the γ or δ chain), or portions thereof (e.g., V genes, J genes, joining sequences, etc.). This is often the case in the context of the discovery process of cellular therapeutics, such as engineered T cells. Thus, the present invention also provides methods for providing a TCR-based therapeutic, such as an engineered T cell, expressing a particular TCR, wherein the method comprises identifying an αβ or γδ chain pairing using any of the methods described herein, or wherein the therapeutic is derived (e.g., by further optimization, mutation, etc.) from an αβ or γδ chain pairing identified using any of the methods described herein. Thus, the methods described herein may also be used in the context of T cell receptor optimization, in a manner similar to that described above for antibodies.

[0061] The following are presented as examples and should not be construed as limitations on the scope of the claims. [Example]

[0062] Example These examples describe how to identify heavy chain-light chain pairings and / or predict the likelihood that a chain pair is functional according to the invention, and how to validate it using single-cell data sets containing known pairings.

[0063] Example 1 - AntiBERTa for pairing prediction method Dataset The model was trained in two steps (illustrated in Figures 3 and 4): a pre-training step (Step 1 in Figure 3, Steps A-B in Figure 4) in which a bidirectional Transformer Encoder model based on the RoBERTa architecture (Liu et al., 2019) was trained using Ampere heavy and light chain antibody sequences; and a fine-tuning step (Steps 2-3 in Figure 3, Step C in Figure 4) in which a model including the Transformer pre-training encoder and classifier modules was trained using known pairs (positive examples) and random pairs (negative examples). The resulting model was then tested on independent testing data. The following dataset was used: The pre-training step is described in Leem et al., 2022. The pre-trained model is referred to as "AntiBERTa." In Leem et al., 2022, the pre-training model was fine-tuned for paratope prediction using human antibody sequences from the Structural Antibody Database (SAbDab, Dunbar et al., 2014). In this example, we directly used the pre-training model for fine-tuning the heavy-light chain pairings as described below. However, it is also possible to use the fine-tuning model described in Leem et al. 2022 (i.e., fine-tuned for paratope prediction) as a starting point for fine-tuning the pairings described herein.

[0064] Pre-training: Human antibody sequences spanning 61 studies were downloaded from the OAS database (Kovaltsuk et al., 2018). As instructed by OAS, antibody sequences were first filtered to remove any sequencing errors. Sequences were also required to have at least 20 residues before CDR1 and 10 residues after CDR3. Finally, sequences were filtered to have 5–12 residues in CDR1, 1–10 residues in CDR2, and 5–38 residues in CDR3. This resulted in a maximum sequence length of 148 residues. The entire collection of 71.98M unique sequences (52.89M heavy chains and 19.09M light chains) was then disjointly split into a training set, validation set, and test set using an 80:10:10 ratio. In total, the MLM training set contained 42.3M heavy chain and 15.3M light chain sequences, while the MLM validation and test sets consisted of 5.3M heavy chain and 1.9M light chain sequences, respectively. AntiBERTa is a single model trained on both heavy and light chains.

[0065] Fine-tuning: Fine-tuning was performed using an in-house dataset of single-cell B-cell receptor sequencing data [referred to as the "Alchemab paired BCR dataset"]. These represent positive examples (known pairs). The classifier model requires the heavy and light chains separately as input. For training, users must provide a value of 1 or 0, with 1 indicating that the heavy and light chains are a true, authentic pair, and 0 indicating that the heavy and light chains are not a true pair. Since a "negative" dataset of VH-VL pairs is not available, this was approximated by randomly shuffling the Alchemab paired BCR dataset, assuming that randomly generated pairs cannot be paired. Thus, strictly speaking, the classifier predicts whether a pair is likely to be a native pair or a random pair (the former is assumed to be functional, and the latter is assumed to be non-functional or at least unable to be paired). A total of 1,121,212 paired sequences were used as the "positive" set, and an additional 1,121,212 were randomly paired to generate the "negative" set. Together, there are a total of 2242424 sequence pairs. Furthermore, a quality control step was applied to the 2242424 sequence set (including both native / positive and random / negative pairs) to specifically identify (i) erroneous sequences that were too long (a cutoff deemed reasonable based on expert knowledge was used in this example, set at 250 residues, although other values ​​may be equally suitable based on the specific data used), or (ii) sequences with stop codons in the reads (in the amino acid sequences provided as output by the single-cell sequencing analysis software, the cutoff was set at 250 residues). 「*」 (iii) sequences with heavy and light chain V gene pairs occurring too rarely (fewer than 10 pairs in total) were removed. These criteria were applied to the combined set of native and random pairs.

[0066] After quality control, a total of 2,239,117 pairs remained (i.e., approximately 0.15% of pairs were filtered out), of which 1,791,293 were used for training, 223,912 for validation, and 223,912 for testing. Pairs were randomly assigned to the training, validation, and test datasets. Other approaches are possible, such as sampling from negative and positive pairs separately to ensure a balanced representation of the two categories in each data subset. Validation helps select the neural network state that should have optimal performance, and the test set is used to test the network's generalization performance.

[0067] Testing: A final independent evaluation of model performance was performed using publicly available datasets of single-cell B-cell receptor sequencing data (10x single-cell data) from King et al. 2021 (7 donors, 30,222 unique heavy-light chain pairs), Eccles et al. 2020 (1 donor, 741 unique heavy-light chain pairs), and Setliff et al. 2019 (2 donors, 4,944 unique heavy-light chain pairs). These were used exclusively in the final testing and therefore represent an unbiased assessment of model performance.

[0068] Array Tokenization The pre-training dataset contained full-length heavy and light chain sequences, which were tokenized as single amino acids. The fine-tuning and testing data also contained full-length heavy and light chain sequences, which were also tokenized as single amino acids.

[0069] If any of the data used did not contain full-length heavy and light chain sequences, for example, because the sequencing method used to generate this data did not cover the entire amino acid sequence of the light / heavy chains, it may still be possible to use the full-length sequences tokenized as single amino acids. For example, some datasets contain (i) for each heavy chain: a V gene identifier, a junction sequence, a J gene identifier, and a D gene identifier, and (ii) for each light chain: a V gene identifier, a junction sequence, and a J gene identifier. In such cases, each entry may include a combination of gene identifiers and sequences, such as IGHV3-23 / CAR...DYW / IGHJ6-IGKV3-20 / CQQ... / IGKJ2. Such sequences can be converted to full amino acid sequences by imputing germline sequences for the corresponding portions (e.g., germline sequences of the corresponding V and J gene identifiers).

[0070] As mentioned above, in this example, all sequences were tokenized as single amino acids. The vocabulary used consisted of 25 tokens: 20 standard amino acids and 5 special tokens ( <s>、< / s> , <pad> 、 <unk>, and <mask>), where each amino acid acts as a token and no byte pair encoding was used. Each chain begins with a start ( <s> ) token and end (< / s> ) token, <pad>The tokens are used to pad out tensors to fit the maximum array length of the mini-batch. <unk>Tokens are used for ambiguous amino acids such as X. A maximum length of 150 is allowed because, together with the start and end tokens, this covers the maximum sequence length of the pre-training dataset (148). Briefly, the advantage of this is that the training set for the OAS is covered while ensuring that the amount of unnecessary padding is minimized.

[0071] Other tokenization schemes are possible and are explicitly contemplated. For example, a model can be trained using tokenized sequences corresponding to V gene identifiers, junction sequences, and J gene identifiers. This can be particularly useful when only partial sequence information is available, as explained above. To accommodate this data format, a custom encoding method can be used, in which each V gene constitutes a single token, each J gene constitutes a single token, and the junction amino acid sequence is tokenized as a single amino acid as above. Junction sequences are the most diverse regions of the sequence and are likely to mediate most of the binding functionality, resulting in increased granularity in the tokenization of this sequence. Such a scheme can use tokens with the smallest number of occurrences, e.g., two in the training set. Other possible schemes for tokenizing junction amino acid sequences (or entire sequences) include, for example, byte pair encoding or overlapping k-mer (e.g., 3-mer) tokens. Schemes such as byte pair encoding can be particularly useful when sequence data for more entire sequences is available, e.g., when single-cell data with hundreds of thousands or even millions of sequences is used to train the model. When using byte pair encoding, several non-overlapping amino acids are encoded as tokens using a dictionary that is automatically defined in a data-driven manner. Byte pair encoding is a subword tokenization scheme that replaces common pairs of consecutive bytes (in this case, consecutive amino acids) with bytes that do not appear in the data. Any pairs that occur more than once in the data are replaced by the corresponding token.

[0072] In the context of this example, a "sentence" is a tokenized representation of a heavy or light chain sequence. Each sentence is represented by a special token <s> , followed by a token for each amino acid (or a token representing a heavy or light chain V gene, an overlapping 3-mer token, a J gene token, or a token for each consecutive amino acid pair), then a special token< / s> Any sentence with fewer tokens than the maximum length (150 in the single amino acid encoding case, 34 in the custom encoding case described above) will have the special <pad>Padded the array with tokens.

[0073] Model We pretrained a bidirectional transformer model based on the RoBERTa architecture (Liu et al., 2019) as described in Leem et al., 2022. As described above, we pretrained the model using a dataset containing both heavy and light chain sequences. Note that separate RoBERTa models could be pretrained for heavy and light chain sequences. This was found to be unnecessary, as the model can distinguish between the two types of chains and learn the features of light and heavy chain sequences separately. This pretrained RoBERTa model trained on ampere antibody sequences is referred to herein as "AntiBERTa," and the training procedure is described in Figure 4 (Steps A and B).

[0074] Two copies of the AntiBERTa model (encoder-only) were then used to process the heavy and light chains, respectively, and these outputs were provided to the cross-attention and classification modules, as further described below. The resulting model was fine-tuned using the dataset described above (2,242,424 pair sequences, including 1,121,212 random pairs that formed the negative set, assigned the label "0," and 1,121,212 real pairs that formed the positive set, assigned the label "1" (Step C in Figure 4). In the fine-tuning step, parameters were shared between the two AntiBERTa models (i.e., the two models were "Siamese" models). The use of pre-trained so-called "checkpoints" has been shown to be a powerful strategy in the context of NLP [Rothe et al., 2020]. Specifically, the BERT-to-BERT architecture was shown to perform well for NMT in Rothe et al.

[2020] .

[0075] If the paired training data used to fine-tune the full model (including two copies of the pre-training AntiBERTa model) did not include full-length BCR heavy and light chain sequences, one of two alternative approaches can be used to generate an equivalent paired dataset containing full-length sequences. In the first approach, full-length sequences can be obtained by replacing V and J gene identifiers with their corresponding germline sequences. In the second approach, the pre-training AntiBERTa model (or any other such "checkpoint" model, e.g., the GPT-2 model) can be used to predict full-length sequences for the training set independently for the heavy and light chains (using each model) based on known portions of the chains. Predictions from the "checkpoint" model can be obtained using some or all of the known portions of the chains (e.g., gene segment identifiers, partial sequences, etc.), optionally in combination with some information obtained from the germline sequence of any segment for which the full sequence is not available (e.g., the identities of some of the amino acids of the segment, e.g., the first k amino acids of the segment (where k can be, for example, 1, 2, 3, 5, 10, etc.)). In other words, an AntiBERTa model trained on the Amp full-length heavy and light chain data can be used to predict the full-length sequences of heavy chains in the training data from the V gene identifiers, J gene identifiers, and junction sequences provided in the data. Similarly, an AntiBERTa model trained on the Amp full-length heavy and light chain data can be used to predict the full-length sequences of light chains in the training data from the V gene identifiers, J gene identifiers, and junction sequences provided in the data. Using these same approaches, any limited paired training data can be mapped to a more expanded format of data that may have been available for training the "checkpoint" model. Alternatively, the data used to train the "checkpoint" model can be converted to a limited format that matches the format of the paired training data.This could still benefit from the potential additional information gleaned by pre-training models from the vast number of available Ampere sequences, but would not fully take advantage of the range of information available in such Ampere sequence data. In this example, full-length sequences were available for both pre-training and fine-tuning and therefore were not required.

[0076] Model Architecture - Pre-training A 12-layer Bidirectional Transformer model based on the RoBERTa architecture (Liu et al., 2019) was pretrained as described in Leem et al., 2022. This model has 768 embedding dimensions, 3072 feedforward dimensions, 12 encoder layers, and 12 self-attention heads, which are predefined and accepted standards. In total, the model has 86 million learnable parameters. The model has a maximum sequence length of 150. It is first trained on a masked language modeling (MLM) task, similar to "fill-in-the-blank." The MLM process used is illustrated schematically in Figure 5A. During MLM pretraining, 15% of the residues in the input sequence are perturbed, where the amino acid is replaced by a mask (gray square in Figure 5A) in 80% of cases, the original amino acid in 10% of cases, and a random amino acid in the remaining 10% of cases (as described in Liu et al., 2019). The model:

number

[0077] Model Architecture - Fine Tuning The pre-training model is then used to construct a neural network with the architecture illustrated in Figure 6A (where the AntiBERTa encoder is the encoder of the pre-training model described above) and trained for the classification task illustrated in Figure 5B.

[0078] In the encoder module, the model processes the input heavy and light chains separately using each copy of the pre-training encoder. This generates an L × 768 tensor for the heavy chain sequences (where L is the length of the heavy chain sequence) and an L' × 768 tensor for the light chain sequences (where L' is the length of the light chain sequence). The 768 corresponds to the 768 numbers generated by the pre-training encoder (the embedding dimension). These can be understood as a set of 768 numbers that describe some contextuality of amino acids in the heavy / light chain sequences. In practice, we process multiple heavy chains and multiple light chains at a time to improve processing efficiency (reducing computational time per processed chain pair). This generates a B × max(L) × 768 tensor (where B is the number of heavy chain sequences and max(L) is the maximum length of a heavy chain sequence among all heavy chain sequences in the batch). Similarly, we generate a B' x max(L') x 768 tensor for the B' number of light chains. B and B' are often identical, and max(L) and max(L') are not necessarily exactly the same. These are the outputs of the encoder module.

[0079] The output of the encoder module is used as input to the cross-attention module. Specifically, the two tensors (B × max(L) × 768 and B' × max(L') × 768) are then processed through six cross-attention blocks. The first three cross-attention blocks are visualized in more detail in Figure 6B. The subsequent blocks follow the same architecture. Each cross-attention block contains a self-attention layer and a cross-attention layer. The self-attention layer and cross-attention layer each contain 12 attention heads. In the self-attention layer, only the heavy chain sequence embeddings are scored using the multi-head self-attention scoring mechanism described in Vaswani et al. (2017):

number

[0080] The concept behind the cross-attention block is that the self-attention layer learns pairwise importance between positions within the heavy chain, while the cross-attention layer learns pairwise importance between each heavy chain position relative to the light chain position. The output of the cross-attention module is used as input to the classifier module. Specifically, after six cross-attention blocks, the output is processed by an attention pooling mechanism. Attention pooling is described in Safari et al. (2020). Essentially, the concept is to compress a variable-length tensor B × max(L) × 768 into a fixed form B × 768. Attention pooling can be conceptualized as a weighted average. Other pooling mechanisms can be used. For example, pooling can be performed by embedding the start token. However, this is considered less advantageous as it results in a large information loss. Alternatively, an unweighted average can be used. This retains more information than using only the start token embedding, but does not have the same ability to place different emphasis on different positions as using a weighted average. More complex pooling mechanisms may be used, but are more difficult to train. Thus, the selection of a suitable pooling mechanism may be made to strike a good balance between preserving information and the need for additional parameters to be trained, depending in part on the amount of training data available and the computing power available for training.The attention pool tensor is then processed by a three-layer neural network that includes (1) a sigmoid linear unit (also known as a Swish unit, described in Ramachandran et al., 2017) that reduces the input dimensionality from B × 768 to B × 256, followed by a dropout rate of 0.25 (note that other dropout rates are possible and are typically empirically evaluated, with a dropout rate of 0.1–0.3 currently found to be suitable), (2) a rectified linear unit (ReLU) that further reduces the dimensionality of the Swish activation input from B × 256 to B × 64, followed by a dropout rate of 0.25, and (3) the ReLU processed output is then passed through a sigmoid layer that computes a score between 0 and 1. It should be noted that other selections of activation functions can be used instead of or in addition to the combination of Swish and ReLU, such as using only Swish layers, using only ReLU layers, or using TanH layers instead of or in combination with Swish and / or ReLU layers. These activation functions advantageously introduce nonlinearity via trainable parameters, providing flexibility to best fit the data. For this reason, the selection of the activation function can be guided, at least to some extent, by the expected behavior of the data, and different possible selections are typically empirically evaluated to identify the selection that best fits the data. For practical purposes, the score can be considered as the likelihood of heavy chain pairing with light chains. In particular, given the nature of generating negative datasets, numbers closer to 0 suggest pairs more reminiscent of pairs found in randomly shuffled datasets, while numbers closer to 1 suggest pairs more reminiscent of pairs found in true pair datasets, such as from single-cell sequencing. As illustrated in Figure 5B, the training data includes true native pairs (training labels of "1") and random pairs (training labels of "0"), and the model is trained to predict scores as close as possible to 1 for the latter and as close as possible to 0 for the former.The training data may also include pairs associated with non-binary labels, e.g., continuous scores indicative of “nativeness” or “functionality.” Such scores may, for example, reflect functional information associated with the known pairs, e.g., binding affinity, or any other metric associated with binding strength.

[0081] The full classifier model was trained using the following regime. Training was conducted for a total of 10 epochs using a weight decay of 0.1 and a cosine learning rate schedule, resulting in a peak learning rate of 3e-5. For the first five epochs, the parameters of the encoder module, which generates the embeddings, were "frozen." This means that only the cross-attention block, attention pooling, and final three-layer neural network were optimized during the first five epochs, while the encoder block remained unchanged. This advantageously allows for better generalization. For the remainder of training, the parameters of the encoder module were varied (although the variations were applied simultaneously to both copies of the pre-training encoder; i.e., they remained exact copies of each other). This approach allowed training to focus on training the classification portion of the model initially, while only fine-tuning the encoding portion of the model was performed in later stages, when the learning rate was slower. For training on the classification task, binary cross-entropy was used as the loss function.

[0082] Model Architecture - Alternatives We designed an alternative approach that does not use the cross-attention module described above. In this approach, the pre-training encoder module has a longer maximum length than the one described above and receives a concatenation of heavy and light chain sequences as input. Such a model jointly embeds both the heavy and light chains, as opposed to generating separate embeddings. This is then used as input to the classifier module described above, which includes an attention pooling layer, Swish, ReLU, and softmax layers. Pre-training the encoder module can use masked language modeling in a manner similar to that described above. Note that the encoder already contains an attention mechanism that can attend to the heavy and light chains, eliminating the need for an additional cross-attention module. Fine-tuning of the full model (including the pre-training encoder and classifier modules) can be performed as described above, i.e., using a first period in which the encoder module parameters are frozen and training focuses on the classifier module, followed by a second period in which variations are also made to the encoder module parameters.

[0083] result Training a classifier model to predict whether two chains are likely to pair imposes some limitations on the amount of training data that can be obtained, since paired heavy-light chain data is required. This type of data is available in relatively limited amounts and is further limited by providing V and J gene identifiers instead of full sequences (thereby effectively limiting predictions to germline sequences in these sections). To circumvent these limitations, we constructed a model that includes an encoder trained on a fairly large dataset of Ampere heavy and light chain sequences (42.3 million and 15.3 million sequences, respectively).

[0084] To evaluate the model, we used data sets from King et al., Setliff et al., and Eccles et al. Although an in-house test set of paired data (made up of 223,912 pairs) could be used for this purpose, we chose these three sets to evaluate because they are outside the scope of such samples and provide a more rigorous reflection of model generalization. We expect the results on the in-house test set to be at least as good as those shown below for independent data.

[0085] To approximate the intended use case of pairing bulk heavy chain sequencing datasets with bulk light chain sequencing datasets, we generated more "negative" data in the same way as for training the model, and repeated the randomization process multiple times for each test set. For this purpose, random paired sequences were mixed with true paired sequences. Performance was measured using the area under the receiver operating curve (ROC) score. An ROC score of 0.5 suggests that the model generates random predictions, while an ROC score of 1.0 suggests that the model generates perfect predictions. Using the same randomization protocol described above, ROC scores for the different sets are provided in Table 1. The ROC curve for the largest of these datasets is shown in Figure 7A. Figure 7B shows the corresponding precision-recall (PR) curve. The PR curve illustrates the tradeoff between precision (positive predictive value, the fraction of true pairs compared to predicted pairs) and recall (sensitivity, the fraction of true pairs predicted as pairs) as the classification threshold is varied from 0 to 1. The area under the PR curve (AUPR) provides a useful indicator of the model's performance in finding true pairs in unbalanced cases (i.e., when more unpaired candidates are expected than truly paired candidates). AUPR should be compared to the 50% positive fraction in the illustrative example. Thus, the data show that the model performs very well in distinguishing true pairs from random pairs (significantly better than random). Additionally, the ROC curve can be used to identify a probability threshold that provides a desired tradeoff between false positives and true positives. The advantageous tradeoff may depend on circumstances, such as the importance of not missing true positive pairings versus the ability of any subsequent testing steps (whether in vitro or in silico) to accommodate a larger number of candidate pairs. For example, a probability threshold of approximately 0.7, 0.75, 0.8, 0.85, 0.9, or 0.95 may be advantageous. In an embodiment, a threshold of approximately 0.8 is used.

[0086] [Table 1]

[0087] The inventors performed experiments with a larger proportion of negative pairs, and the results are shown in Table 2.

[0088] [Table 2]

[0089] The data show that predictive performance is stable even when attempting to identify true pairs among a larger number of random pairs. This suggests that the model will likely perform well in real-world situations where there are expected to be more random pairs than true pairs among the set of candidate pairs. The data further show stable performance across three test datasets, suggesting that the model exhibits good generalizability, at least within the same species. Thus, the true vs. random pair probabilities provided by the model can be used to compare candidate pairs, for example, by ranking or otherwise prioritizing the candidate pairs based on this probability.

[0090] Example 2 - FAbCon for pairing prediction In this example, we used an alternative architecture that uses a generative language model: a decoder-only model pre-trained for the next token prediction task (rather than an encoder-only model pre-trained for the masked language modeling task as in Example 1).

[0091] FabCon is an antibody-specific large-scale language model (LLM) based on the Falcon LLM from natural language processing (Penedo, G. et al. 2023). FAbCon is pre-trained using causal language modeling (CLM) on 779.4 million unpaired antibody sequences. It is then fine-tuned for the pairing task.

[0092] method Dataset. To pretrain the FAbCon model using the next-token prediction task, we collected a dataset of 823.7 million sequences (821.2 million amperes, 2.5 million pairs). This was split into 95% (777 million amperes and 2.4 million pair sequences) for pretraining the FAbCon model, and the remaining 5% (44.3 million amperes and 0.1 million pair sequences) was used to evaluate the pretraining progress.

[0093] More specifically, we downloaded the Observed Antibody Space (OAS) database on February 23, 2023. We prefiltered samples using a procedure similar to that previously discussed (Bachas et al. 2022). For example, we did not use B cell receptor sequences derived from pre-B cell samples. However, we retained all sequences, regardless of count. In total, we used 1.47 billion ampere sequences (i.e., heavy or light chain only). Additionally, we supplemented this corpus with a set of proprietary B cell receptor sequences from 376 libraries, bringing the total number of sequences to 1.54 billion. We then clustered the dataset at 90% sequence identity across the VH or VL domains using Linclust (Steinegger, M. & Soding, J). After clustering, we retained 777.8 million sequences from the OAS and 43.4 million sequences from our proprietary data. In addition to the Ampere data, we also used heavy-light chain paired sequences for pretraining. We first combined 1.5 million publicly available paired sequences (Jaffe et al. 2022) with an in-house dataset of 1.4 million paired sequences. Due to the scarcity of paired data, we applied a 99% redundancy criterion to the paired data. This resulted in a total of 2.5 million new sequences, of which 1.4 million were from the Jaffe et al. dataset and 1.1 million were from our in-house data. The final dataset, containing 823.7 billion sequences (821.2 million Ampere sequences, 2.5 million pairs), was split into 95% (777 million Ampere sequences and 2.4 million paired sequences) and 5% (443 million Ampere sequences and 100,000 paired sequences) for training and evaluating the CLM, respectively.

[0094] 1,791,293 sequences were used for fine-tuning: this included 895,765 sequences from single B-cell sequencing, which we consider "true" pairs. Random matching of heavy and light chains in the single B-cell sequencing data generated an additional 895,528 sequences—we consider these "false" pairs. Following training, the fine-tuning model was evaluated against three separate external single B-cell sequencing datasets.

[0095] Tokenization. The sequences are tokenized at the amino acid level, where each amino acid is a token. The antibody heavy chain is

number

number

number

number

[0096] Pre-training. FAbCon is a generative Transformer model based on the Falcon architecture (Penedo et al. 2023), available at falconllm.tii.ae / falcon-models.html and https: / / huggingface.co / docs / transformers / main / model_doc / falcon. It is a decoder-only autoregressive Transformer model with an architecture based on GPT-3 (Brown et al. 2020), ALiBi positional encoding (Press et al. 2021), and FlashAttention (Dao et al. 2022). The Falcon model uses multi-query attention, which shares key and value embeddings across attention heads, reducing the memory cost of inference and pre-training. We pretrained three FAbCon variants (Figure 9A) for the next-token prediction task: a 144 million parameter variant, a 300 million parameter variant, and a 2.4 billion parameter variant. All FAbCon variants were pretrained using 48 NVIDIA A100s (80GB) from the NVIDIA DGX SuperCloud (Cambridge-1 environment). The configuration of each variant is listed in Table 3. For next-token prediction, the model is tasked with predicting the amino acid sequence from left to right. Only the N-terminal leading residue is used to inform the prediction of the subsequent amino acid. Each FAbCon model has a maximum context window of 256 (i.e., a maximum sequence length of 256). During pretraining, both unpaired and paired antibody sequences are allowed to form one mini-batch.

[0097] FAbCon variants are pre-trained using the CLM objective, where the model autoregressively predicts each residue by attending only to its previous residue in the sequence. More formally, the model predicts the probability of a position, γ, t : P(γ t |γ1,γ2,...,γ t-1) The task was to predict the probability of amino acid residues in the sequence. During pre-training, the model parameters were

number

number

[0098] As mentioned above, all three FAbCon variants feature flash attention and multi-query attention to increase pre-training speed and reduce memory footprint. In each variant, we employed a "deep narrow" architecture, where we increased the number of layers and maintained a relatively lower embedding dimension. Optimization used the Fused AdamW optimizer with a peak learning rate of 2e-5, a cosine learning rate schedule, 200k steps, a weight decay of 0.01, and gradient norm clipping of 1.0.

[0099] [Table 3]

[0100] Fine-tuning. Figures 9B and 9D show different architectures of FAbCon for pre-training (Figure 9B) and for the classifier model (Figure 9D) that identifies whether an input sequence is a true or false pairing. The difference between FAbCon-Small (a generative language model) and FAbCon-Small Pairing (a pairing classifier model) is the final "head" layer at the end. FAbCon-Small Pairing uses the representation from the decoder layer to output the probability that the input sequence forms a pair. This architecture is equally applicable to FAbCon-Medium and FAbCon-Large; for practical reasons, we only show the fine-tuning results for the FAbCon-Small model.

[0101] FAbCon Small was fine-tuned over 10 epochs with a peak learning rate of 5e-5 after a 5% warm-up. Dropout was set to 0.1 and weight decay to 0.01. We used the model checkpoint with the lowest validation set loss, corresponding to the 5th epoch. Model freezing was not implemented here. Figure 9C provides an overview of how the model is pre-trained and then fine-tuned for pairing purposes.

[0102] result The results of the process described above are shown below in Tables 4 and 5. These showed very good prediction accuracy, although slightly lower than Example 1 (although it is expected that similar performance can be achieved with further optimization).

[0103] [Table 4]

[0104] [Table 5]

[0105] The same model was also fine-tuned for antibody-antigen binding prediction (classification of sequence pairs as binders or non-binders for a specific antigen; data not shown) and performed exceptionally well (average typical accuracy was 0.851 across three different test antigens, increasing to 0.851 and 0.883 for the medium and large models), indicating that the model discriminates binders from non-binders well and further demonstrating that the encoding module of the model described herein can learn information to distinguish functional from non-functional pairs.

[0106] Example 3 - Discussion This example describes a machine learning, NLP-inspired approach to the BCR heavy chain-light chain pairing problem. The deep learning-based approach described herein offers benefits based on covering the BCR repertoire as thoroughly as possible while eliminating the need for bulk light chain sequencing or single-cell sequencing. Furthermore, this approach has the ability to learn general characteristics of pairings from a training dataset and use this learning to predict previously unseen chain pairings. While this will likely be advantageous in many cases given the extreme diversity of the BCR repertoire, it may be particularly advantageous in the context of applications such as identifying specific antibodies that may underlie a desired phenotype, such as individual or other rare antibodies.

[0107] Through the combined use of a larger training dataset of ampli?ed full-length sequences for pre-training the language model and a smaller paired dataset with potentially less extensive sequence coverage, the present approach ensures that most of the available training data is subjected to ?ne-tuning of the ?nal model, Which includes an encoding module (this is the part of the language model that provides the latent representation of the input sequences, e.g., an encoder or decoder, depending on the model architecture) and a classi?er module.

[0108] Additionally, the present approach increases the variety of functional pairs identified by providing a method for evaluating candidate strands as "authentic" strands. For example, the present approach would be ideally suited to situations in which repertoires (or at least subsets of repertoires) of both pairs have been identified, but pairing information between the two repertoires is unavailable. Indeed, in such cases, the present approach can take advantage of knowledge that the native partner of the candidate strand is likely to be present in the candidate set. In contrast, approaches that generate strands for pairing "de novo" may generate predictions that will not fold properly for this type of strand or that will not form a functional pair with any strand. Note that approaches that generate de novo predictions can be used in combination with the present approach, e.g., by using the former to provide candidates for evaluation by the latter.

[0109] The application of bulk heavy chain repertoire analysis to antibody discovery remains related to the challenge of light chain pairing. The deep learning-based approach described herein allows for the identification of potential light chains that pair with any heavy chain of interest by evaluating known or otherwise emerging candidate light chains for pairing with a high in silico accuracy. This approach therefore has the potential to fill the gap in light chain pairing information, thereby enabling therapeutic antibody discovery and a better understanding of the immune system.

[0110] Finally, although this approach has been described in the context of identifying light chain pairings for heavy chain queries, it can also be applied to the inverse problem of identifying heavy chain pairings for light chain queries, as well as to pairings of other immune molecule dimers. This is a less frequent problem because antibodies are widespread therapeutic, diagnostic, and research tools, and particularly in the context of antibodies, heavy chain sequencing is more common and the heavy chain is thought to play a more important role in determining specificity and affinity.

[0111] References Vander Heiden et al., 2017. Dysregulation of B Cell Repertoire Formation in Myasthenia Gravis Patients Revealed through Deep Sequencing. J Immunol. 2017 Feb 15; 198(4):1460-1473. Bashford-Rogers et al, 2019. Analysis of the B cell receptor repertoire in six immune-mediated diseases. Nature volume 574, pages122-126(2019). Nielsen et al., 2020. Human B Cell Clonal Expansion and Convergent Antibody Responses to SARS-CoV-2. bioRxiv. Preprint. 2020 Jul 9. doi: 10.1101 / 2020.07.08.194456. Simonich et al., 2019. Kappa chain maturation helps drive rapid development of an infant HIV-1 broadly neutralizing antibody lineage. Nature Communications volume 10, Article number: 2190 (2019). Krawczyk et al., 2019. Looking for therapeutic antibodies in next-generation sequencing repositories. mAbs. Volume 11, 2019 - Issue 7, Pages 1197-1205. Galson et al., 2020. Deep Sequencing of B Cell Receptor Repertoires From COVID-19 Patients Reveals Strong Convergent Immune Signatures. Front. Immunol., 15 December 2020. doi.org / 10.3389 / fimmu.2020.605170. Mora and Walczak, 2019. How many different clonotypes do immune repertoires contain? Current Opinion in Systems Biology. Volume 18, December 2019, Pages 104-110 Kovaltsuk et al., 2018. Observed Antibody Space: A Resource for Data Mining Next-Generation Sequencing of Antibody Repertoires. J Immunol October 15, 2018, 201 (8) 2502-2509. Teplyakov et al., 2016. Structural diversity in a human antibody germline library. MAbs. Aug-Sep 2016; 8(6):1045-63. Glanville et al., 2009. Precise determination of the diversity of a combinatorial antibody library gives insight into the human immunoglobulin repertoire. PNAS December 1, 2009 106 (48) 20216-20221. Jayaram et al., 2012.Germline VH / VL pairing in antibodies. Protein Engineering, Design and Selection, Volume 25, Issue 10, October 2012, Pages 523-530. Ling et al., 2018. Effect of VH-VL Families in Pertuzumab and Trastuzumab Recombinant Production, Her2 and FcγIIA Binding. Front. Immunol., 12 March 2018. doi.org / 10.3389 / fimmu.2018.00469 DeKosky et al., 2016. Large-scale sequence and structural comparisons of human naive and antigen-experienced antibody repertoires. PNAS May 10, 2016 113 (19) E2636-E2645. King et al., 2021. Single-cell analysis of human B cell maturation predicts how antibody class switching shapes selection dynamics. Science Immunology 12 Feb 2021. Vol. 6, Issue 56, eabe6291 Eccles et al., 2020. T-bet+ Memory B Cells Link to Local Cross-Reactive IgG upon Human Rhinovirus Infection. Cell Reports Volume 30, Issue 2, 14 January 2020, Pages 351-366.e7 Setliff et al., 2019.High-Throughput Mapping of B Cell Receptor Sequences to Antigen Specificity. Cell Volume 179, Issue 7, 12 December 2019, Pages 1636-1646.e15 Reddy et al., 2010. Monoclonal antibodies isolated without screening by analyzing the variable-gene repertoire of plasma cells. Nature Biotechnology volume 28, pages965-969(2010). Zhu et al., 2013. Mining the antibodyome for HIV-1-neutralizing antibodies with next-generation sequencing and phylogenetic pairing of heavy / light chains. PNAS. 2013 Apr 16; 110(16):6470-5. Raybould et al., 2021 Public Baseline and shared response structures support the theory of antibody repertoire functional commonality. PLoS Comput Biol 17(3): e1008781. Rakocevic et al., 2021.The landscape of high-affinity human antibodies against intratumoral antigens. bioRxiv. 8 Feb 2021. doi.org / 10.1101 / 2021.02.06.430058 Vaswani et al., 2017. Attention Is All You Need. arXiv: 1706.03762 Devlin, Jacob, et al. "Bert: Pre-training of deep bidirectional transformers for language understanding.” arXiv preprint arXiv: 1810.04805 (2018). Radford et al., 2019. Language Models are Unsupervised Multitask Learners. https: / / openai.com / blog / better-language-models / Liu et al., 2019. RoBERTa: A Robustly Optimized BERT Pretraining Approach. arXiv: 1907.11692 Rothe et al., 2020.Leveraging Pre-trained Checkpoints for Sequence Generation Tasks. arXiv: 1907.12461 Dunbar and Deane, 2016. ANARCI: antigen receptor numbering and receptor classification. Bioinformatics. 2016 Jan 15; 32(2):298-300. Rees, 2020. Understanding the human antibody repertoire. MAbs. Jan-Dec 2020; 12(1):1729683. Ye et al., 2013. IgBLAST: an immunoglobulin variable domain sequence analysis tool. Nucleic Acids Res. 2013 Jul; 41(Web Server issue):W34-40. Carter JA et al. Single T Cell Sequencing Demonstrates the Functional Role of αβ TCR Pairing in Cell Lineage and Antigen Specificity. Frontiers in Immunology. Vol. 10. 2019, p. 1516. Zheng GXY, Terry JM, Belgrader P, Ryvkin P, Bent ZW, Wilson R, et al. Massively parallel digital transcriptional profiling of single cells. Nat Commun. (2017) 8:14049. Howie B, Sherwood AM, Berkebile AD, Berka J, Emerson RO, Williamson DW, et al. High-throughput pairing of T cell receptor α and β sequences. Sci Transl Med. (2015) 7:301ra131. Eve Richardson, Jacob D. Galson, Paul Kellam, Dominic F. Kelly, Sarah E. Smith, Anne Palser, Simon Watson & Charlotte M. Deane (2021) A computational method for immune repertoire mining that identifies novel binders from different clonotypes, demonstrated by identifying anti-pertussis toxoid antibodies, mAbs, 13:1. Yi-Chun Hsiao, Yonglei Shang, Danielle M. DiCara, Angie Yee, Joyce Lai, Si Hyun Kim, Diego Ellerman, Racquel Corpuz, Yongmei Chen, Sharmila Rajan, Hao Cai, Yan Wu, Dhaya Seshasayee & Isidro Hotzel (2019) Immune repertoire mining for rapid affinity optimization of mouse monoclonal antibodies, mAbs, 11:4, 735-746. Warszawski S, Borenstein Katz A, Lipsh R, Khmelnitsky L, Ben Nissan G, Javitt G, et al. (2019) Optimizing antibody affinity and stability by the automated design of the variable light-heavy chain interfaces. PLoS Comput Biol 15(8): e1007207. Seeliger D, Schulz P, Litzenburger T, Spitz J, Hoerer S, Blech M, Enenkel B, Studts JM, Garidel P, Karow AR. Boosting antibody developability through rational sequence optimization. MAbs. 2015;7(3):505-15. doi: 10.1080 / 19420862.2015.1017695. Mason, D.M., Friedensohn, S., Weber, C.R. et al. Optimization of therapeutic antibodies by predicting antigen specificity from antibody sequence via deep learning. Nat Biomed Eng (2021). Leem, J., Mitchell, L.S., Farmery, James H.R., Barton, J., Galson, J.D. Deciphering the language of antibodies using self-supervised learning. Patterns 3, 100513. July 8, 2022. Safari, Pooyan, Miquel India, and Javier Hernando. "Self-attention encoding and pooling for speaker recognition.” arXiv preprint arXiv: 2008.01077 (2020). Ramachandran, Prajit, Barret Zoph, and Quoc V. Le. "Searching for activation functions.” arXiv preprint arXiv: 1710.05941 (2017). Shaw, Peter, Jakob Uszkoreit, and Ashish Vaswani. "Self-attention with relative position representations.” arXiv preprint arXiv: 1803.02155 (2018). Su, Jianlin, et al. "Roformer: Enhanced transformer with rotary position embedding.”arXiv preprint arXiv: 2104.09864 (2021). Nijkamp, Erik, et al. "ProGen2: exploring the boundaries of protein language models.” arXiv preprint arXiv: 2206.13517 (2022). Cedric R Weber, Rahmad Akbar, Alexander Yermanos, Milena Pavlovi-, Igor Snapkov, Geir K Sandve, Sai T Reddy, Victor Greiff, immuneSIM: tunable multi-feature simulation of B- and T-cell receptor repertoires for immunoinformatics benchmarking, Bioinformatics, Volume 36, Issue 11, June 2020, Pages 3594-3596. Penedo, G. et al. The RefinedWeb Dataset for Falcon LLM: Outperforming Curated Corpora with Web Data, and Web Data Only. arXiv (2023) doi: 10.48550 / arxiv.2306.01116. Steinegger, M. & Soding, J. Clustering huge protein sequence sets in linear time. Nat. Commun. 9, 2542 (2018). Jaffe, D. B. et al. Functional antibodies exhibit light chain coherence. Nature 611, 352-357 (2022). Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al. Language models are few-shot learners. Advances in neural information processing systems, 33:1877-1901, 2020. Dao, T., Fu, D. Y., Ermon, S., Rudra, A., and Re, C. Flashattention: Fast and memory-efficient exact attention with io-awareness. In Advances in Neural Information Processing Systems, 2022. Press, O., Smith, N., and Lewis, M. Train short, test long: Attention with linear biases enables input length extrapolation. In International Conference on Learning Representations, 2021. Touvron H. et al. 2023. Llama 2: Open Foundation and Fine-Tuned Chat Models arXiv: 2307.09288v2. 19 July 2023 Mandar Joshi, Danqi Chen, Yinhan Liu, Daniel S. Weld, Luke Zettlemoyer, Omer Levy. SpanBERT: Improving Pre-training by Representing and Predicting Spans. arXiv: 1907.10529v3. 18 Jan 2020

[0112] All references cited herein are incorporated by reference in their entirety for all purposes to the same extent as if each individual publication or patent or patent application was specifically and individually indicated to be incorporated by reference in its entirety.

[0113] The specific embodiments described herein are offered by way of example and not by way of limitation. Various modifications and variations of the compositions, methods, and uses of the present technology will be apparent to those skilled in the art without departing from the scope and spirit of the described technology. Any subheadings herein are included for convenience only and should not be construed as limiting the disclosure in any way. Other aspects and embodiments of the present invention provide those aspects and embodiments described above, where the term "comprising" is replaced with the term "consisting of" or "consisting essentially of," unless the context dictates otherwise.

[0114] The methods of any of the embodiments described herein may be provided as a computer program or as a computer program product or a computer readable medium carrying a computer program arranged to perform the method described above when run on a computer.

[0115] Unless the context dictates otherwise, the feature descriptions and definitions set forth above are not limited to any particular aspect or embodiment of the invention, but apply equally to all aspects and embodiments described.

[0116] Throughout this specification and claims, the following terms shall have the meanings expressly associated therewith unless the context clearly dictates otherwise. As used herein, the phrase "in one embodiment" does not necessarily refer to the same embodiment, although it may be the same. Furthermore, as used herein, the phrase "in another embodiment" does not necessarily refer to different embodiments, although it may be different. Thus, as described below, various embodiments of the invention may be readily combined without departing from the scope or spirit of the invention.

[0117] It should be noted that, as used in this specification and the appended claims, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. Ranges may be expressed herein as from one particular value preceded by "about" and / or to another particular value preceded by "about." When such a range is expressed, another embodiment includes from the one particular value and / or to the other particular value. Similarly, when values ​​are expressed as approximations, by use of the preceding "about," it will be understood that the particular value forms another embodiment. The term "about" in connection with numerical values ​​is arbitrary and means, for example, ±10%.

[0118] Throughout this specification (including the claims that follow), unless the context otherwise requires, the words "comprise" and "include," as well as variations such as "comprises," "comprising," and "including," shall be understood to imply the inclusion of a specified integer or step or group of integers or steps, but not the exclusion of any other integer or step or group of integers or steps. As used herein, "and / or" shall be considered a specific disclosure of each of the two specifically stated features or components, with or without the other. For example, "A and / or B" shall be considered a specific disclosure of each of (i) A, (ii) B, and (iii) A and B, as if each were individually specified.< / pad> < / unk> < / pad> < / mask> < / unk> < / pad>

Claims

1. 1. A computer-implemented method for determining whether a protein chain pair comprising a first chain and a second chain is likely to form a functional antigen binding protein, the method comprising: providing a query sequence pair comprising the sequence of the first protein chain and the sequence of the second protein chain as input to a deep learning model configured to receive the protein chain sequence pair as input and generate as output a score indicative of the likelihood that the protein chain pair will form a functional antigen binding protein; Including, The method, wherein the deep learning model includes an encoder module and a classifier module, and the deep learning module is trained using paired training sequences from known antigen-binding proteins.

2. 1. A computer-implemented method for identifying an antigen binding protein comprising a first protein chain and a second protein chain, the method comprising: providing a query first protein chain; providing one or more candidate second protein chain sequences; and providing one or more query sequence pairs, each comprising (i) a sequence of the first protein chain and (ii) a candidate second protein chain sequence, as input to a deep learning model to determine whether the one or more candidate second chain sequences are likely to form a functional antigen-binding protein with the query first protein chain; identifying the second protein chain by the deep learning model is configured to receive as input one or more protein chain sequence pairs and generate as output a score or information derived therefrom indicative of the likelihood of each protein chain pair to form a functional antigen-binding protein, the deep learning model comprising an encoder module and a classifier module, and the deep learning module has been trained using paired training sequences from known antigen-binding proteins, optionally identifying the second protein chain comprises providing a plurality of candidate second chain sequences, and the information derived from the score comprises a ranking of the protein chain sequence pairs, wherein protein sequence pairs that are more likely to form a functional antigen-binding protein are ranked higher than protein sequence pairs that are less likely to form a functional antigen-binding protein.

3. the antigen-binding protein is (i) a heavy chain-light chain pair, wherein the first chain is a heavy chain or a light chain and the second chain is a light chain or a heavy chain, and optionally, the first chain is a heavy chain and the second chain is a light chain; or (ii) an αβ chain pair, wherein the first chain is a β chain or an α chain, and the second chain is an α chain or a β chain, and optionally, the first chain is a β chain and the second chain is an α chain; or (ii) a γδ chain pair, wherein the first chain is a δ chain or a γ chain, and the second chain is a γ chain or a δ chain, and optionally, the first chain is a δ chain and the second chain is a γ chain; 3. The method of claim 1 or claim 2, comprising:

4. 4. The method of any one of claims 1 to 3, wherein the encoder module comprises one or two encoders pre-trained with training sequences from amplicon protein chains from known antigen binding proteins, or a decoder pre-trained with training sequences comprising amplicon and amplicon protein chains from known antigen binding proteins.

5. the encoder module: one or two encoders of a sequence-to-sequence model or a decoder of a generative language model, optionally said sequence-to-sequence model or generative language model being a Transformer-based model; or 5. The method of any one of claims 1 to 4, comprising one or two transformer-based encoder or decoder models.

6. 6. The method of claim 1, wherein the encoder module comprises one or two encoder models of a sequence-to-sequence model, the encoder models being trained using masked language modeling, and optionally, training the encoder models comprises replacing random mask positions of the training amino acid sequences, and / or 15% of the positions are masked during training, and / or the masked positions are replaced with mask tokens, random amino acids, or the original amino acid at that position; or wherein the encoder module comprises a decoder model of a generative language model, the decoder model being trained using causal language modeling.

7. 7. The method of any one of claims 1 to 6, wherein the encoder module comprises one or two encoders pre-trained with training sequences comprising ampered first and second protein chains from known antigen binding proteins, and optionally the training sequences comprise at least 1 million, at least 2 million, at least 5 million, at least 10 million, at least 20 million, or at least 50 million distinct sequences; or wherein the encoder module comprises a decoder pre-trained with training sequences comprising ampered first and second protein chains from known antigen binding proteins and paired first and second protein chains from known antigen binding proteins, and optionally the training sequences comprise at least 500, 600, or 700 million distinct sequences, and / or at least 1, 1.5, or 2 million paired sequences.

8. 8. The method of any one of claims 1 to 7, wherein the encoder module comprises two copies of an encoder pre-trained with training sequences comprising first and second protein chains from known antigen binding proteins.

9. 8. The method of any one of claims 1 to 7, wherein the encoder module comprises an encoder pre-trained with training sequences comprising a concatenation of a sequence from a first chain and a sequence from a second protein chain from a known antigen binding protein, wherein the concatenated sequence comprises a sequence from an ampere protein chain, or wherein the encoder module comprises a decoder pre-trained with training sequences comprising a concatenation of a sequence from a single chain or a first chain and a sequence from a second protein chain, respectively, from a known antigen binding protein.

10. 9. The method of any one of claims 1 to 8, wherein the deep learning model further comprises a cross-attention module that receives as input the output of the encoder module and generates an output that is used by the classifier module to provide a score indicative of the likelihood that the protein chain pair will form a functional antigen binding protein.

11. The method of claim 10 , wherein the cross-attention module includes one or more cross-attention blocks, each cross-attention block including a self-attention layer and a cross-attention layer, and / or the cross-attention module includes multiple cross-attention blocks.

12. The method of claim 11 , wherein the cross-attention module includes one or more cross-attention blocks including a cross-attention layer and a self-attention layer, and each cross-attention layer and self-attention layer includes multiple attention heads.

13. The method of any one of claims 10 to 12, wherein each self-attention layer comprises one or more attention heads attending to the output of a first encoder that receives the first strand sequence as input, and / or each cross-attention block further comprises a residual connection between the output of the first encoder and the output of the self-attention layer.

14. 14. The method of claim 10, wherein each cross-attention layer comprises one or more attention heads that attend to (i) the output of the self-attention layer or the output of a first encoder that receives the first strand sequence as input, and (ii) the output of a second encoder that receives the second strand sequence as input, and / or each cross-attention block further comprises a residual connection between the output of a first encoder and the output of the cross-attention layer.

15. 15. The method of any one of claims 1 to 14, wherein the classification module comprises a softmax layer that generates a score between 0 and 1 that is interpretable as a probability that an input protein chain pair forms a functional antigen-binding protein, and / or wherein the classification module comprises one or more of a dimensionality reduction layer, a regularization mechanism, and a layer with an activation function, optionally wherein the dimensionality reduction layer comprises an attention pooling layer, and / or the regularization layer comprises a dropout mechanism, and / or the activation function is independently selected from Tanh, Leaky ReLU, SmeLU, GeLU, Swish, and ReLU.

16. 16. The method of any one of claims 1 to 15, wherein the deep learning model receives as input a plurality of query chain sequence pairs and produces as output a respective score indicative of the likelihood of each query protein chain pair to form a functional antigen-binding protein, and / or a ranking of the plurality of query chain sequence pairs, such that query protein sequence pairs that are more likely to form a functional antigen-binding protein are ranked higher than query protein sequence pairs that are less likely to form a functional antigen-binding protein.

17. 17. The method of any one of claims 1 to 16, wherein the paired training sequences from known antigen-binding proteins comprise paired training heavy and light chain sequences from single B-cell sequencing data, and / or the training data further comprises a negative set comprising randomly paired first and second protein chain sequences, optionally obtained by re-pairing paired training sequences from known antigen-binding proteins and / or by random pairing of unpaired training sequences, and / or the negative set comprises paired first and second protein chain sequences from antigen-binding proteins that are predetermined to be non-functional, and / or the paired training sequences are associated with a first label and the random paired training sequences are associated with a second label.

18. 18. The method of any one of claims 1 to 17, wherein providing protein chain sequence pairs as inputs to the deep learning model comprises encoding each of the protein chain sequences using a predetermined encoding scheme, optionally wherein each amino acid is encoded individually, or wherein sequences are encoded using tokens each corresponding to an individual k-mer, and / or wherein each protein chain sequence is preceded or followed by a special token indicating whether the protein chain sequence is a first or second protein chain.

19. providing a query first protein chain or a query first protein chain of a query pair, obtaining the sequence of said query first protein chain from a user via a user interface, from a computing device, from a sequence obtaining means or a computing device associated with a sequence obtaining means, from a database or other computer readable medium; and / or Sequencing a sample containing genetic material encoding an antigen-binding molecule comprising the query sequence, optionally obtaining the query sequence comprises performing B cell bulk sequencing of a sample containing B cells, T cell bulk sequencing of a sample containing T cells, or bulk sequencing of a sample containing any other cells that express an antigen-binding molecule comprising the query sequence, or bulk sequencing of genetic material derived therefrom, for example, bulk sequencing of a B cell receptor library or a T cell receptor library; and / or Obtaining a sample containing B cells, T cells, or other cells that express an antigen-binding molecule containing the query sequence, or genetic material derived therefrom, such as a B cell receptor library or a T cell receptor library; and / or Including, Providing one or more candidate second protein chain sequences comprises obtaining said candidate second protein chain sequences from a user via a user interface, from a computing device, from a sequence obtaining means or a computing device associated with a sequence obtaining means, from a database or other computer readable medium; and / or The method of any one of claims 1 to 18, wherein the one or more candidate second chain protein sequences are known second chain protein sequences or simulated second chain protein sequences.

20. 20. The method of any one of claims 1 to 19, further comprising providing to a user via a user interface one or more identified second protein chains, portions thereof, or information derived therefrom, and / or one or more scores indicative of the likelihood that one or more query pairs will form a functional antigen-binding protein, or information derived therefrom, and / or further comprising predicting a score indicative of the likelihood that each of one or more query pairs comprising a respective candidate second protein chain sequence will form a functional antigen-binding protein, and identifying candidate second protein chain sequences by applying one or more criteria to the score, optionally wherein the one or more criteria are selected from: the score exceeding a predetermined cutoff; the score being the highest predicted score of a set of candidate second protein chain sequences; and the score being in a predetermined top percentile of predictability of a set of candidate second protein chain sequences; and / or the score is a probability that the protein chain pair will form a functional antigen-binding protein.

21. 21. A method of providing antigen-binding protein chain pairings for a plurality of query sequences comprising a first chain sequence, the method comprising performing the method of any one of claims 2 to 20 for each of the query sequences, optionally wherein the plurality of query sequences are heavy chain or light chain sequences obtained by bulk B-cell repertoire sequencing.

22. 1. A method for providing an antigen binding protein having desired properties, said method comprising: providing one or more query sequences comprising first strand sequences, at least one of the one or more query sequences likely to have the desired property; Identifying a second strand sequence for each of the one or more query sequences using the method of any one of claims 2 to 20; A method comprising:

23. 1. A method for providing a tool for predicting the likelihood of forming a functional antigen binding protein, said method comprising: providing training data comprising training first and second protein chain sequences from known antigen binding proteins; training a deep learning model to receive as input a protein chain sequence pair and generate as output a score indicative of the likelihood that the protein chain pair will form a functional antigen binding protein, the deep learning model comprising an encoder module and a classifier module; A method comprising:

24. a processor; a computer readable medium comprising instructions which, when executed by said processor, cause said processor to perform the steps of the method of any one of claims 1 to 23; A system including:

25. One or more computer readable media comprising instructions that, when executed by one or more processors, cause said one or more processors to perform the steps of the method of any one of claims 1 to 23.