Engineering of antigen binding proteins

Deep learning model predicts whether protein chain pairs may form functional antigen-binding proteins, which solves the difficulty of identifying BCR and TCR chain pairs in large-scale sequencing data, and achieves efficient and accurate functional pairing prediction.

CN120239885APending Publication Date: 2025-07-01ALCHEMAB THERAPEUTICS LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380080611.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-10-14
Filing Date
2023-10-12
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

Existing methods are difficult to effectively identify B-cell receptor (BCR) heavy-chain pairs or T-cell receptor (TCR) αβ strand pairs from data that do not contain chain pair information, especially in large-scale sequencing data, resulting in difficulty in determining functional pairing.

Method used

Using deep learning models, especially encoder-decoder structures such as AntiBERTa or FAbCon, predict whether protein chain pairs may form functional antigen-binding proteins through masking language modeling and cross-attention mechanisms, model training and fine-tuning is used to use training sequences of known antigen-binding proteins.

Benefits of technology

High recall and accurate prediction of functional BCR or TCR chain pairs is achieved, filling the gap in the lack of light chain pairing information in large-scale heavy chain sequencing, and improving the efficiency and accuracy of antibody and TCR identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120239885A_ABST
    Figure CN120239885A_ABST
Patent Text Reader

Abstract

Methods for determining whether a pair of protein chains comprising a first chain and a second chain is likely to form a functional antigen binding protein are described. The method comprises: providing a query sequence pair comprising a sequence of a first protein chain and a sequence of a second protein chain to a deep learning model as input, the deep learning model is configured to take a pair of protein chain sequences as an input and to generate as an output a score indicative of a probability that the pair of protein chains forms a functional antigen binding protein, where the deep learning model comprises an encoder module and a classifier module, and wherein the deep learning module has been trained using pairs of training sequences from known antigen binding proteins. The methods can be used to identify antibodies from non-paired or single chain sequences. Related methods and products are described.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to methods for engineering antigen-binding proteins such as B cell receptors, antibodies, and T cell receptors, which are achieved by determining whether a candidate variable chain pairing is likely to be functional, such as determining whether a candidate heavy chain-light chain pair or a candidate α-β or γ-δ chain pair is likely to be functional, or identifying candidate variable chains (such as heavy / light chains, α / β chains, γ / δ chains) that are likely to form a functional pairing with an input chain (such as a light / heavy chain, β / α chain, δ / γ chain). The present invention also relates to methods for providing antigen-binding proteins (such as therapeutic antibodies) derived from input variable chains (such as B cell receptor / antibody heavy or light chains). Background Art

[0002] Effective humoral immunity requires a diverse repertoire of B cells capable of binding different antigens through their B cell receptors (BCRs). It is estimated that the theoretical total size of the BCR repertoire in humans is as high as approximately 10 15 variants, of which approximately 10 9 variants circulate in a single individual at any given time [Rees, 2020]. The BCR consists of two protein chains in two pairs: two heavy chains and two light chains. Each B cell expresses a heavy chain and a light chain pair (which may be unique) to form its BCR, which is expressed on its surface or secreted as an antibody. More than 600 million different human heavy chain sequences and approximately 70 million light chain sequences are currently cataloged in the Observed Antibody Space [Kovaltsuk et al., 2018]. Characterizing an individual's BCR repertoire (also known as an individual's BCR library) has proven to be a valuable tool for understanding the biology of various diseases [Vander Heiden et al., 2017; Bashford-Rogers et al, 2019; Nielsen et al., 2020; Simonich et al., 2019] and discovering new therapeutic antibody drugs [Krawczyk et al., 2019; Galson et al., 2020].

[0003] There are two main methods for characterizing an individual's BCR repertoire: sequencing of single B cells and sequencing of bulk B cell populations. Single cell sequencing is more commonly used in antibody discovery applications because it preserves the pairing information between the heavy and light chains. However, single cell sequencing has limited throughput, and different platforms and protocols vary in their coverage of the BCR repertoire present in a single sample. Even the most advanced microfluidic systems typically can only recover approximately 10 4Sequences of B cells / sample [King et al., 2021; Eccles et al., 2020; Setliff et al., 2019]. Humans typically have approximately 10 6 B cells / ml of blood [Mora and Walczak, 2019], which means that single-cell methods cannot characterize the full B cell diversity even in small samples. Additionally, compared to bulk sequencing, single-cell sequencing has very specific sample requirements (e.g., cells generally must remain viable until processed, so fresh samples that need to be processed on the day of collection, or fresh samples frozen according to a specific protocol), very high cost / sample (single-cell sequencing is at least an order of magnitude more expensive than bulk sequencing), and requires dedicated laboratory equipment.

[0004] Sequencing large populations of B cells can more easily recover approximately 10 7 B cell sequences / sample [Briney et al., 2019], which is significantly closer to the expected diversity in an individual. However, since B cells are lysed during library preparation, the heavy-chain-light-chain pairing information is not retained. Generally, these bulk BCR sequencing methods focus only on the heavy chain because it plays a dominant role in antigen binding and is more diverse than the light-chain repertoire [Kovaltsuk et al., 2018]. However, for antibody discovery, it is necessary to have both the heavy and light chains of the antibody so that the antibody can be synthesized and functionally characterized. The gap in light-chain pairing information has facilitated the development of computational pairing methods [Reddy et al., 2010, Zhu et al., 2013, Raybould et al., 2021, Rakocevic et al., 2021]. However, these are limited to specific datasets and some specific sequences within those datasets.

[0005] Similarly, cellular immunity requires a diverse repertoire of T cells that can bind different antigens through their T cell receptors (TCRs). It is estimated that the total size of the TCR repertoire in humans contains up to approximately 10 15A unique αβ T cell receptor (TCR) pair [Carter et al., 2019]. Although experimental methods for paired αβ TCR sequencing have been developed (including single-cell methods [Zheng et al., 2017] and methods based on multicellular deconvolution [Howie et al., 2015]), these methods remain specialized and limited in throughput. Thus, most of the available knowledge of TCR repertoires is based on bulk sequencing of single-chain repertoires (primarily β-chain repertoires). This is inherently limited, especially since it has been shown that both α TCR chains and β TCR chains are involved in alloreactivity and antigen specificity [Carter et al., 2019].

[0006] Accordingly, there remains a need for improved methods for identifying chain pairs from data that do not contain such pairing information, such as BCR heavy-chain-light-chain pairs or TCR αβ chain pairs. SUMMARY OF THE INVENTION

[0007] The problem of identifying BCR heavy-chain-light-chain pairs is far from trivial. In fact, the diversity of the BCR repertoire creates a vast search space. Additionally, while several heavy-chain-light-chain combinations can give rise to stable BCRs (this observation has led some to speculate that pairing may be random [Glanville et al., 2009; Jayaram et al., 2012; DeKosky et al., 2016]), only a limited number of pairings result in functional BCRs that can bind their target antigens [Teplyakov et al., 2016; Ling et al., 2018]. This suggests that functional pairing is non-random, but the determinants of functional pairing are obscured by the number of pairings that may be stable but non-functional. In practice, this means that even if stable pairs can be predicted, finding the correct light chain for a particular heavy chain is challenging because it will generate a large number of scenarios that need to be experimentally verified, and if selection is based primarily on stability, it is expected to be difficult to validate.

[0008] A variety of different computational methods have been proposed, each with several significant drawbacks. The first method is based on matching the relative frequencies of BCR heavy and light chains upon independent sequencing [Reddy et al., 2010]. In this study, mice were first immunized to generate a strong immune response, and then the top 4 to 5 most common heavy and light chains were selected for pairing. Beyond these top 4 to 5 sequences, pairing based on relative frequencies was not possible. Recently, Rakocevic et al.

[2021] showed that this method is only effective when the sample is dominated by a small number of high-frequency B cells. Zhu et al.

[2013] proposed a method called phylogenetic pairing, which involves comparing the structures of phylogenetic trees generated from heavy and light chain sequence data. This method is limited to examining specific clonal expansions; in this case, for known antiviral antibody lineages, rather than the entire BCR repertoire. Raybould et al.

[2021] proposed a method based on a computational structural model of heavy and light chain pairing. This method is inherently limited by the limited and severely skewed availability of high-quality structural templates and can at most identify features related to stability that do not necessarily translate into function. Additionally, this method can only pair families of similar sequences and not specific sequences (which would limit its practical applicability - this has not been experimentally verified). Thus, the present inventors have determined that current methods for computational heavy-light chain pairing are limited because they are only applicable to specific datasets and the sequences within those datasets. In fact, any of the existing validated methods are only applicable to datasets in which both heavy and light chain sequences can be obtained from the sample, where the data is dominated by large clonal expansions, and only facilitate the pairing of a limited number of sequences within those datasets.

[0009] The present inventors have also identified a general application for antibody discovery, with the expectation of being able to generate a viable light chain for any given heavy chain. It is also desirable to be able to achieve this result using only heavy chain information, as BCR library bulk sequencing efforts typically focus limited resources on sequencing heavy chains, which are thought to play a more important functional role than light chains. To address these issues, the present inventors hypothesized that it might be possible to use deep learning methods inspired by recent advances in natural language processing (NLP). Specifically, they hypothesized that deep learning models incorporating an encoder-decoder architecture (such as the transformer [Vaswani et al., 2017]) and derivative architectures such as BERT [Devlin et al., 2018] and RoBERTa [Liu et al., 2019] (encoder-only architectures) or GPT [Brown et al., 2020] and Falcon [Penedo et al., 2023] should be able to learn the features of antibodies in a manner similar to that used to train such models for natural language processing, using masked language modelling. They also hypothesized that the resulting learned representation would carry information that could be used by a classifier model to predict whether a candidate pair is likely to form a functional pair. Transformers have shown state-of-the-art results in a wide variety of NLP tasks [Vaswani et al., 2017; Devlin et al., 2018; Liu et al., 2019; Rothe et al., 2020]. Therefore, the present inventors designed methods for generating learned representations of heavy and light chains using a pre-trained encoder (referred to as 'AntiBERTa') or decoder (referred to as 'FAbCon'), which are fine-tuned as part of a classifier for the pairing task. After multiple blind tests on single-cell datasets with known pairings, they showed that the method predicts true pairs with high recall and precision. The method provides a new solution for light chain pairing, as well as a way to fill in the gaps in large-scale heavy chain sequencing. The present inventors have also determined that the same method can be used to solve the TCR chain pairing problem.

[0010] Accordingly, in a first aspect, there is provided a method for determining whether a protein chain pair comprising a first chain and a second chain is likely to form a functional antigen-binding protein, the method comprising: providing a query sequence pair comprising the sequence of the first protein chain and the sequence of the second protein chain as input to a deep learning model, the deep learning model being configured to take a protein chain sequence pair as input and produce as output a score indicative of the probability that the protein chain pair forms a functional antigen-binding protein, wherein the deep learning model comprises an encoder module and a classifier module, and wherein the deep learning module has been trained using paired training sequences from known antigen-binding proteins.

[0011] In accordance with this aspect, a method for identifying an antigen-binding protein comprising a first protein chain and a second protein chain is also described, the method comprising: providing a query first protein chain, providing one or more candidate second protein chain sequences; and using the method as described above to determine whether each protein chain pair comprising the query first protein chain and a candidate second protein chain is likely to form a functional antigen-binding protein. Accordingly, in accordance with this aspect, there is also provided a method for identifying an antigen-binding protein comprising a chain pair, the method comprising: providing a query first protein chain, and identifying a second protein chain by: providing one or more candidate second protein chain sequences; and determining whether one or more candidate second chain sequences are likely to form a functional antigen-binding protein with the query first protein chain by: providing each query sequence pair among one or more query sequence pairs as input to a deep learning model, each query sequence pair comprising (i) the sequence of the first protein chain and (ii) a candidate second protein chain sequence, wherein the deep learning model is configured to take a protein chain sequence pair or multiple protein sequence pairs as input and produce as output a score indicative of the probability that each protein chain pair forms a functional antigen-binding protein or information derived therefrom, wherein the deep learning model comprises an encoder module and a classifier module, and wherein the deep learning module has been trained using paired training sequences from known antigen-binding proteins.

[0012] The method according to this aspect may have one or more of the following features.

[0013] The score indicative of the probability that the protein chain pair forms a functional antigen-binding protein may be the probability that the protein chain pair forms a functional antigen-binding protein.

[0014] Identifying a second protein chain may include providing a plurality of candidate second chain sequences, and obtaining a score for each protein pair comprising the first protein chain and a corresponding candidate second chain. Information derived from the scores may include an ordering of the protein chain sequence pairs, wherein protein sequence pairs that are more likely to form a functional antigen-binding protein are ordered higher than protein sequence pairs that are less likely to form a functional antigen-binding protein. The ordering may be based on the scores, for example, ordered by decreasing or increasing score. Accordingly, in accordance with this aspect, there is also provided a method of identifying an antigen-binding protein comprising a chain pair, the method comprising: providing a query first protein chain, and identifying a second protein chain by: providing a plurality of candidate second protein chain sequences; and determining whether a candidate second chain sequence is likely to form a functional antigen-binding protein with the query first protein chain by: providing a plurality of query sequence pairs as input to a deep learning model, each query sequence pair comprising (i) a sequence of the first protein chain and (ii) a candidate second protein chain sequence, wherein the deep learning model is configured to take as input a plurality of protein chain sequence pairs and produce as output a corresponding score indicative of the probability that each protein chain pair forms a functional antigen-binding protein or information derived therefrom (e.g., an ordering of the plurality of pairs), wherein the deep learning model comprises an encoder module and a classifier module, and wherein the deep learning module has been trained using paired training sequences from known antigen-binding proteins. Similarly, in accordance with this aspect, there is also provided a method of identifying a functional antigen-binding protein comprising a chain pair, the method comprising: providing a plurality of query antigen-binding proteins comprising a first protein chain and a second protein chain; and determining whether the query first chain sequence and the second chain sequence are likely to form a functional antigen-binding protein by: providing a plurality of query sequence pairs as input to a deep learning model, wherein the deep learning model is configured to take as input a plurality of protein chain sequence pairs and produce as output a corresponding score indicative of the probability that each protein chain pair forms a functional antigen-binding protein or information derived therefrom (e.g., an ordering of the plurality of pairs), wherein the deep learning model comprises an encoder module and a classifier module, and wherein the deep learning module has been trained using paired training sequences from known antigen-binding proteins. The method can be used to prioritize query chain pairs / antigen-binding proteins based on the likelihood of the query chain pair / antigen-binding protein forming a functional antigen-binding protein as predicted using a deep learning model.

[0015] A chain pair and / or each protein chain may be referred to as a "variable chain". The terms "known chain pair" or "known antigen-binding protein" refer to an antigen-binding protein / variable chain sequence pair from an antigen-binding protein that is known to be present in an antigen-binding protein that exhibits a desired antigen-binding function or in an antigen-binding protein that forms part of the B-cell or T-cell repertoire of at least one subject. The latter may also be referred to as a "natural" chain pair. Thus, a "known protein chain / antigen-binding protein" can be a protein / chain pair that has been previously identified (e.g., in a sample, subject, etc. containing a natural chain pair / protein) and / or a protein / chain pair that has a desired function (e.g., verified or verifiable by in vitro or in vivo testing, such as binding affinity to a target, expression, stability, etc.). All sequences can be amino acid sequences. The first (query) chain sequence can be a heavy chain sequence, and the second sequence can be a light chain sequence. The antigen-binding protein can be a B-cell receptor or an antibody, or a protein derived therefrom. Thus, the antigen-binding protein can comprise a heavy chain-light chain pair. The query sequence can comprise a heavy chain sequence or a light chain sequence. The corresponding chain sequence can be a light chain sequence or a heavy chain sequence.

[0016] The antigen-binding protein can be a T-cell receptor or a protein derived therefrom. The antigen-binding protein can comprise an αβ chain pair, wherein the first chain sequence is a β chain sequence or an α chain sequence, and the corresponding chain sequence is an α chain sequence or a β chain sequence. The first chain sequence can be a β chain sequence, and the corresponding sequence can be an α chain sequence. The antigen-binding protein can comprise a γδ chain pair, wherein the first chain sequence is a δ chain sequence or a γ chain sequence, and the corresponding chain sequence is a γ chain sequence or a δ chain sequence. The first chain sequence can be a δ chain sequence, and the corresponding sequence can be a γ chain sequence. The antigen-binding protein can be a T-cell receptor or a protein derived therefrom. Thus, the antigen-binding protein can comprise an αβ chain pair or a γδ chain pair. Thus, the query sequence can comprise a β or δ chain sequence, or an α or γ chain sequence. The corresponding chain sequence can be an α or γ chain sequence, or a β or δ chain sequence.

[0017] The encoder module may include multiple encoders (also referred to as "encoder models"). Each encoder may take a protein chain sequence as input. The encoder module may include one or two Transformer-based encoder models. The encoder module may include one or two encoders that have been pre-trained using training sequences from unpaired protein chains from known antigen-binding proteins. The encoder module may include one or two encoders of a sequence-to-sequence model. The sequence-to-sequence model may be a recurrent neural network or a Transformer. The sequence-to-sequence model may be a sequence-to-sequence Transformer-based model. The recurrent neural network may be a gated recurrent unit (GRU)-based model or a long short-term memory (LSTM) model. For example, a GRU-based model may include a GRU-based encoder and a GRU-based decoder. The encoder may be a 4-layer bidirectional GRU, for example with a hidden dimension of 1024. The decoder may be a 4-layer forward-only GRU, for example with a hidden dimension of 1024. A Transformer is a deep learning model that uses an attention mechanism. A Transformer-based model may be a Transformer model with a structure for both the encoder and decoder that uses self-attention and point-wise fully connected layers. The encoder and / or decoder may be composed of a stack of the same number of layers (e.g., 6, 12, 24, or 30 layers). Each layer of the encoder may have two sub-layers: a multi-head self-attention layer and a position-wise fully connected feed-forward network layer. Each layer of the decoder may have three sub-layers: a self-attention sub-layer, a layer that performs multi-head attention on the output of the encoder stack, and a feed-forward network layer. The encoder module may include such a decoder model that has been pre-trained using training sequences that include paired and unpaired protein chains from known antigen-binding proteins. The decoder may be a decoder of a decoder-only Transformer-based model. The encoder module may include an embedding layer and a decoder layer of a decoder-only autoregressive Transformer model (which may be collectively referred to as the "decoder model"). The decoder model may use flash attention, multi-query attention, and / or positional encoding.The decoder can be a 24-layer decoder with 12 attention heads per layer, an embedding dimension of 768, and a feed-forward dimension of 3072. The decoder can be a 28-layer decoder with 16 attention heads per layer, an embedding dimension of 1024, and a feed-forward dimension of 4096. The decoder can be a 56-layer decoder with 32 attention heads per layer, an embedding dimension of 2048, and a feed-forward dimension of 8192. Each such model is available, but the smallest of these models has been found to have very good performance.

[0018] The encoder module can include one or two encoders using positional encoding (such as absolute positional encoding or relative positional encoding). In some embodiments, the encoder module includes one or two encoders using absolute positional encoding. In some embodiments, the encoder module includes one or two encoders using relative positional encoding. Relative positional encoding (also known as relative position representation) can be implemented as described in Shaw et al. (2018). The encoder using relative positional encoding can be an encoder using rotary positional encoding. Rotary positional encoding can be implemented as described in Su et al. (2022). Relative positional encoding can improve the model's ability to capture the relationships between positions in the chain. Transformer-like models with relative positional encoding (including, for example, encoder-only models or Transformer models) can use relative position information as an additional component of the keys and values used in the self-attention mechanism of the Transformer-like model. The encoder module can include a decoder using positional encoding (such as ALiBi positional encoding as described in Press et al., 2021). The encoder module can include a decoder using rotary position embeddings as described in Su et al. (2022). For example, the Falcon (falconllm.tii.ae / ) and Llama2 (Touvron et al., 2023) models are decoder-only models using rotary position embeddings.

[0019] The encoder module may include one or two encoder models of a sequence-to-sequence model trained using masked language modeling. Alternatively, the encoder may have been trained using span-based masked language modeling (see, e.g., Joshi et al. 2020). Training the encoder model may include training the model to replace randomly masked positions in the training amino acid sequence. The model may have been trained using masked language modeling, where 15% of the positions are masked during training and / or where the masked positions are replaced with a mask token, a random amino acid, or the original amino acid at that position. The encoder module may include one or two such encoders that have been pre-trained using training sequences that include unpaired first and second protein chains from known antigen-binding proteins. The training sequences for pre-training may include at least 1 million, at least 2 million, at least 5 million, at least 10 million, at least 20 million, or at least 50 million independent sequences. The encoder module may include two copies of such an encoder that have been pre-trained using training sequences that include unpaired first and second protein chains from known antigen-binding proteins. The encoder module may include an encoder that has been pre-trained using training sequences that include a concatenation of a sequence from the first chain from a known antigen-binding protein and a sequence from the second protein chain from a known antigen-binding protein, where the concatenated sequences include sequences from unpaired protein chains.

[0020] The decoder may be a decoder model of a generative language model, where the decoder model has been trained using causal language modeling (i.e., the next token prediction task) or span modeling (also known as "span-based masked language modeling"). The encoder module may include a decoder that has been pre-trained using training sequences that include unpaired first and second protein chains from known antigen-binding proteins and paired first and second protein chains from known antigen-binding proteins. The training sequences may include at least 500 million, 600 million, or 700 million independent sequences and / or at least 1 million, 1.5 million, or 2 million paired sequences. The encoder module may include a decoder that has been pre-trained using training sequences that each include a single chain from a known antigen-binding protein or a concatenation of a sequence from the first chain from a known antigen-binding protein and a sequence from the second protein chain from a known antigen-binding protein.

[0021] The deep learning model may also include a cross-attention module that takes the output of the encoder module as input and produces an output that is used by the classifier module to provide a score indicating the probability that a protein chain pair forms a functional antigen-binding protein. The cross-attention module may include one or more cross-attention blocks, where each cross-attention block includes a self-attention layer and a cross-attention layer. The cross-attention module may include multiple cross-attention blocks. The cross-attention module may include one or more cross-attention blocks that include a cross-attention layer and a self-attention layer, and each cross-attention layer and self-attention layer may include multiple attention heads. Each self-attention layer may include one or more attention heads that attend to the output of the first encoder with the first chain sequence as input. Each cross-attention block may also include a residual connection between the output of the first encoder and the output of the self-attention layer. Each cross-attention layer may include one or more attention heads that attend to: (i) the output of the self-attention layer or the output of the first encoder with the first chain sequence as input; and (ii) the output of the second encoder with the second chain sequence as input. Each cross-attention block may also include a residual connection between the output of the first encoder and the output of the cross-attention layer. Each self-attention layer may include 2, 3, 4, 5, 6, 12 or more attention heads. Each cross-attention layer may include 2, 3, 4, 5, 6, 12 or more attention heads. Each cross-attention block may include a self-attention layer and a cross-attention layer. Each cross-attention block may include a residual connection between the output of the self-attention layer and / or the cross-attention layer and the output of the first encoder with the first chain sequence as input. The output of the self-attention layer and / or the cross-attention layer may be layer-normalised before being provided as input to a subsequent layer or module. The classification module may include a softmax layer that produces a score between 0 and 1, which may be interpreted as the probability that the input protein chain pair forms a functional antigen-binding protein. The classification module may include one or more of the following: a dimensionality reduction layer, a regularisation mechanism, and a layer with an activation function. The dimensionality reduction layer may include an attention pooling layer or an average pooling layer.The regularization layer may include a dropout mechanism. The activation function may be selected from Tanh, Leaky ReLU, GeLU, SmeLU, Swish, and ReLU. The activation function in each layer may be independently selected from Swish and ReLU. The deep learning model may take as input multiple query chain sequence pairs and produce as output a corresponding score indicating the probability that each query protein chain pair forms a functional antigen-binding protein and / or a ranking of the multiple query chain sequence pairs such that the query protein sequence pairs that are more likely to form a functional antigen-binding protein are ranked higher than the query protein sequence pairs that are less likely to form a functional antigen-binding protein. For example, the deep learning model may take as input 8, 16, 32, or 64 query sequence pairs (such as 32 pairs). Such a deep learning model may advantageously be able to provide predictions for many candidate chain pairings with high computational efficiency.

[0022] Paired training sequences from known antigen-binding proteins can include paired training heavy and light chain sequences from single B cell sequencing data. The training data can include one or more data sets, each previously obtained by single B cell sequencing of a sample obtained from a subject or by sequencing of a library derived therefrom. The training data can also include paired training heavy and light chain sequences from known antibodies / B cell receptors. For example, the training data can include paired training heavy and light chain sequences from one or more antibody / BCR databases, from one or more known therapeutic antibodies / BCRs, and / or from one or more known antibodies / BCRs having a desired binding function. The training data can include paired training heavy and light chain sequences from an initial B cell receptor library. The training data can include paired training heavy and light chain sequences from a B cell receptor library that has experienced antigen. Thus, the training data can include paired training heavy and light chain sequences obtained from a subject that has been exposed to one or more specific antigens. The training first chain sequence and second chain sequence from a known chain pair can include paired training alpha and beta chain sequences from single T cell sequencing data. The training data can include one or more data sets, each of which was previously obtained by single T cell sequencing of a sample obtained from a subject or by sequencing of a library derived therefrom. The training data can also include paired training first chain sequences and corresponding chain sequences from known T cell receptors. For example, the training data can include paired training alpha and beta chain sequences from one or more T cell receptor databases, from one or more known therapeutic TCRs, and / or from one or more known TCRs having a desired binding function. The training data can include paired training alpha and beta (or delta and gamma) chain sequences from an initial T cell receptor library. The training data can include paired training alpha and beta (or delta and gamma) chain sequences from a T cell receptor library that has experienced antigen. Thus, the training data can include paired training alpha and beta (or delta and gamma) chain sequences obtained from a subject that has been exposed to one or more specific antigens. The training first chain sequence and second chain sequence from a known chain pair can include paired training chain sequences, where each pair includes a chain sequence that contains or consists of: a V gene sequence or identifier, a J gene sequence or identifier, and a joining sequence, and optionally a D gene sequence or identifier. The training first chain sequence and corresponding chain sequence from a known chain pair can include paired training chain sequences, where each pair includes a chain sequence that contains or consists of: a chain sequence that contains or consists of: a V gene sequence or identifier, a J gene sequence or identifier, and a joining sequence. A reference to a V gene or J gene can refer to the amino acid sequence corresponding to the respective gene. The training data can include at least 80,000, at least 100,000, at least 120,000, at least 150,000, at least 500,000, or at least 1,500,000 pairs of training sequences, such as training heavy and light chain sequences.Advantageously, the training data can include at least 1,500,000 pairs of training heavy and light chain sequences. The training data can include mammalian, e.g., human, chain sequence pairs. The training data can include mammalian heavy and / or light chain sequences. The training data can include human heavy and / or light chain sequences. The training data can include training pairs of sequences from the same species as the query sequence. The training data can include at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or at least 95% of sequences from the same species as the query sequence. The query sequence can be a sequence that is not present in the training data. The query sequence can be a sequence obtained from a sample from an object having a desired characteristic (e.g., a desired phenotype). For example, the object can have a specific clinical characteristic.

[0023] Training data may include a simulated training sequence. The training data may include a simulated paired sequence, a paired simulated sequence, or a simulated unpaired training sequence. The simulated data may include sequences obtained using methods for simulating antigen-binding protein sequences (such as immuneSIM (Weber et al., 2020)) and / or methods for simulating protein sequences (such as ProGen2 (Nijkamp et al., 2022)). Known protein pairs may be referred to as the "positive set" of the pairs. The training data may include a negative set of chain sequence pairs that are not expected to form a functional antibody-binding protein. The negative set may include randomly paired first and second protein chain sequences. Alternatively or in addition, the negative set may include simulated sequence pairs or pairs of simulated sequences. Alternatively or in addition, the negative set may include sequence pairs from antigen-binding proteins previously determined to be non-functional. Pairs may be determined to be non-functional according to one or more predetermined functional criteria. For example, pairs that cannot be experimentally expressed or fail to bind a target may be considered non-functional. Thus, the training data may also include a negative set containing randomly paired first and second protein chain sequences. The randomly paired first and second protein chain sequences may have been obtained or may be obtained as part of a method by re-pairing paired training sequences from known antigen-binding proteins and / or by randomly pairing unpaired training sequences. The negative set may include paired first and second protein chain sequences from antigen-binding proteins previously determined to be non-functional. Known paired training sequences (positive set) may be associated with a first label, while randomly paired / negative set training sequences may be associated with a second label. The training data may include pairs associated with non-binary labels. The training data may include a positive set of pairs associated with a score that can take multiple values (up to and including a continuous score, where the continuous score can be, for example, between 0 and 1). For example, pairs in the positive set may be associated with a score indicating "naturalness" or "functionality". Such a score may, for example, reflect functional information related to known pairs, such as binding affinity, or any other metric related to binding strength. The training data may include a negative set of pairs associated with a single value (such as 0) or multiple values (such as values indicating the confidence that the pair is non-functional).

[0024] The training data may also include unpaired training first sequences and / or second sequences. These sequences can be used to pre-train the encoder or decoder of the encoder module. The unpaired training first chain sequences and / or second chain sequences can have any of the characteristics of the sequences described for the paired sequences. In particular, the unpaired chain sequences can be sequences of the same type as the paired sequences (e.g., when the paired training sequences are heavy and light chain pairs, the unpaired training first chain sequence / second chain sequence can include unpaired heavy chains and / or light chains), can include sequences from the same organism (e.g., can include mammalian and / or human sequences, can include sequences from one or more organisms, can include sequences from an initial library and / or an antigen-exposed library, etc.), can include the same information (e.g., such as gene segment identifiers, sequences, and combinations thereof). The unpaired training sequences can include some or all of the first sequences and / or second sequences present in the paired training sequences. Advantageously, the unpaired training sequences can include more first chain sequences and / or more second chain sequences than the paired training chain sequences. The first (e.g., query) chain sequence can include or consist of: a V gene sequence or identifier, a J gene sequence or identifier, and a joining sequence, and optionally a D gene sequence or identifier. The second chain (e.g., corresponding) sequence can include or consist of: a V gene sequence or identifier, a J gene sequence or identifier, and a joining sequence. The forms of the first chain sequence and the second chain sequence are related to the form of the training chain sequence. Thus, a deep learning model trained using a training chain sequence that includes or consists of a V gene sequence or identifier, a J gene sequence or identifier, and a joining sequence, and optionally a D gene sequence or identifier can accept as input a chain sequence that includes or consists of these components. Similarly, a deep learning model trained using a training chain sequence that includes or consists of a V gene sequence or identifier, a J gene sequence or identifier, and a joining sequence can accept as input a chain sequence that includes or consists of these components. The query sequence can include or consist of one or more first chain CDR sequences. The second sequence can include or consist of one or more corresponding chain CDR sequences. The first sequence and / or the second sequence can include or consist of a CDR3 sequence. The first protein chain and the second protein chain can be chains with different length ranges and / or different domain structures.

[0025] Providing protein chain sequence pairs as input to a deep learning model (whether for training or prediction) can include encoding each protein chain sequence using a predetermined encoding scheme. According to the encoding scheme, each amino acid can be encoded individually. In some embodiments, the sequences are encoded using tokens that each correspond to an individual k-mer. Providing the sequence pairs to the deep learning model can include encoding the sequences using an encoding scheme in which each amino acid is encoded individually. For example, different tokens can be provided for each possible amino acid. Providing the sequence pairs to the deep learning model can include encoding the sequences using an encoding scheme in which each gene sequence identifier corresponds to an individual token. Providing the sequence pairs to the deep learning model can include encoding the query sequence using an encoding scheme in which each amino acid corresponds to an individual token. Providing the sequence pairs to the deep learning model can include encoding the query sequence using an encoding scheme in which there is a special token (e.g., heavy chain - H or light chain - L) before or after the chain sequence indicating whether the protein chain sequence is the first or second protein chain. Providing the sequence pairs to the deep learning model can include encoding each sequence using an encoding scheme in which the sequences are encoded using tokens that each correspond to an individual k-mer (e.g., using byte - pair encoding), i.e., sequences that can be obtained as the full sequence rather than gene identifiers. Each sequence can be encoded using overlapping k-mers. The k-mer can have any length shorter than the expected chain length. For example, the k-mer can have a length of 1 to 100, 2 to 100, 2 to 50, 2 to 20, or 2 to 10. The k-mer can have a length of 1 to 5. The k-mer can have a fixed length. For example, fixed k-mer lengths of 1, 2, 3, 4, or 5 can be used. A k-mer of length 1 is equivalent to encoding each character (e.g., each amino acid) individually. A k-mer with length k>2 (e.g., 3) can be used as part of an encoding scheme using overlapping or non - overlapping k-mers. The overlapping k-mers can overlap to different degrees. For example, a k-mer of length 3 can overlap by 1 or 2 characters. In a scheme with k = 3, each token corresponds to a unique group of 3 characters (e.g., a motif of 3 amino acids). Unpaired training data may have been filtered to exclude any sequences that contain a particular region with a length outside the corresponding predetermined length range. For example, any sequence that has fewer than 20 amino acids before the CDR1 region, fewer than 10 amino acids after the joining region, a CDR1 region length outside the range of 5 to 12 amino acids, a CDR2 region length outside the range of 1 to 10 amino acids, and / or a CDR3 region length outside the range of 5 to 38 amino acids can be excluded from the unpaired training data. Paired training data may have been filtered to exclude any pairs that contain a joining sequence (in the first chain and / or the corresponding chain) outside the predetermined length range.In other words, the training data may not include any pairs containing a first (e.g., heavy) chain linkage outside a predetermined length range and / or a second (e.g., light) chain linkage outside a predetermined length range. For example, pairs containing a heavy chain linkage sequence shorter than a predetermined length (e.g., 3, 4, 5, 6, 7, 8, 9, or 10 amino acids) may already have been excluded. As another example, pairs containing a heavy chain linkage sequence longer than a predetermined length, e.g., 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, or 35 amino acids, may already have been excluded. As another example, pairs containing a light chain linkage sequence shorter than a predetermined length, e.g., 3, 4, 5, 6, 7, 8, 9, or 10 amino acids, may already have been excluded. As another example, pairs containing a light chain linkage sequence longer than a predetermined length, e.g., 15, 16, 17, 18, 19, 20, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 amino acids, may already have been excluded. For the linkage sequences in the corresponding (e.g., light) chain and the first (e.g., heavy) chain of a pair, the predetermined lengths may be the same or different. In a specific example, pairs containing a heavy chain linkage sequence of fewer than 7 amino acids may already have been excluded, and / or pairs containing a heavy chain linkage sequence of more than 30 amino acids may already have been excluded. Alternatively or in addition, pairs containing a light chain linkage sequence of fewer than 7 amino acids may already have been excluded, and / or pairs containing a light chain linkage sequence of more than 20 amino acids may already have been excluded. The query sequence or sequence pair may include one or more gene sequence identifiers, and the method may further include replacing one or more gene sequence identifiers with corresponding germline sequences. The deep learning model may be a transformer-based model that includes an encoder pre-trained using unpaired training first chain sequences and / or corresponding chain sequences, and a decoder or bidirectional encoder pre-trained using unpaired training corresponding chain sequences and / or first chain sequences. The encoder module may include a BERT model or a variant thereof, e.g., BERT, RoBERTa, DistilBERT, or RoFormer (Su et al., 2022). The encoder module may include an encoder trained using unpaired training first chain sequences and second chain sequences. Alternatively, the encoder module may include a model trained using training first (e.g., heavy or light) chain sequences and a model trained using second (e.g., light or heavy) chain sequences. When the encoder module includes two encoders trained using unpaired training first chain sequences and second chain sequences, the two encoders may be the same pre-trained model. Thus, the two encoders may be initialized using a pre-trained model with the same structure and the same parameters. The unpaired training chain sequences may include the full-length sequence of the variable region of the second chain. The unpaired training chain sequences may include the full sequence of the variable region of the first chain. Alternatively, the encoder module may include a decoder pre-trained using unpaired training sequences and paired training sequences.For example, a decoder may take as input a string that contains an encoding of a first or second chain, preceded by a token indicating the chain type, and optionally contains one or more padding tokens. The decoder may also take as input a string that contains an encoding of a first chain and an encoding of a corresponding second chain (i.e., a pair of chains, each preceded by a token indicating the chain type, and optionally containing one or more padding tokens). The deep learning model may have been trained using paired first and second (e.g., heavy and light) chain sequences from known chain pairs, where the sequences do not contain the full-length sequences of the variable regions of the second and / or first chains. In such an embodiment, the deep learning model may have been trained by obtaining paired training sequences that contain the full-length sequences of the variable regions of the corresponding chain and / or first chain via inputting missing sequence information. The inputting missing sequence information may include replacing gene identifiers with corresponding germline sequences. The inputting missing sequence information may include predicting the full-length sequences of each paired training first (e.g., heavy) and / or second (e.g., light) chain sequence from partial sequences using a pre-trained encoder. Alternatively, prior to the pre-trained encoder, unpaired training second (e.g., light) chain sequences and / or unpaired training first (e.g., heavy) chain sequences may have been converted into a format that matches the format of the respective paired training sequences.

[0026] Providing the query first protein chain or the query first protein chain in a query pair may include obtaining the sequence of the query first protein chain as follows: obtaining from a user via a user interface, obtaining from a computing device, obtaining from a sequence acquisition device or a computing device associated with a sequence acquisition device, obtaining from a database or other computer-readable medium. Providing the query first protein chain or the query first protein chain in a query pair may include sequencing a sample containing genetic material encoding an antigen-binding molecule containing the query sequence. Obtaining the query sequence may include bulk sequencing of B cells in a sample containing B cells, bulk sequencing of T cells in a sample containing T cells, or bulk sequencing of a sample containing any other cells expressing an antigen-binding molecule containing the query sequence or genetic material derived therefrom (e.g., a B cell receptor library or a T cell receptor library). Providing the query first protein chain or the query first protein chain in a query pair may include obtaining a sample containing B cells, T cells, or other cells expressing an antigen-binding molecule containing the query sequence or genetic material derived therefrom (e.g., a B cell receptor library or a T cell receptor library). Providing one or more candidate second protein chain sequences may include obtaining the sequence of the candidate second protein chain as follows: obtaining from a user via a user interface, obtaining from a computing device, obtaining from a sequence acquisition device or a computing device associated with a sequence acquisition device, obtaining from a database or other computer-readable medium. The one or more candidate second protein chain sequences may be known second chain protein sequences or simulated second chain protein sequences. The known second chain protein sequences may be sequences of second chains that have been previously observed (e.g., in a sample, an individual, etc.) or are known to have a predetermined function (e.g., an antigen-binding protein that has previously been shown to bind to a specific target, will be expressed in a sample, etc.). Providing the query sequence (or sequence pair) may include obtaining the query sequence (or sequence pair) as follows: obtaining from a user via a user interface, obtaining from a computing device, obtaining from a sequence acquisition device or a computing device associated with a sequence acquisition device, obtaining from a database or other computer-readable medium. Providing the query sequence or sequence pair may include sequencing a sample containing genetic material encoding an antigen-binding molecule containing the query sequence. Providing the query sequence may include obtaining a sample containing B cells, T cells, or other cells expressing an antigen-binding molecule containing the query sequence or genetic material derived therefrom. Providing the query sequence may include sequencing a sample containing genetic material encoding an antigen-binding molecule containing the query sequence, e.g., by bulk sequencing of B cells in a sample containing B cells (or any other cells expressing an antigen-binding molecule containing the query sequence, or genetic material derived therefrom, e.g., a B cell receptor library). Providing the query sequence may include obtaining a sample containing B cells, or other cells expressing an antigen-binding molecule containing the query sequence, or genetic material derived therefrom such as a B cell receptor library.

[0027] The method may include determining probabilities for a plurality of pairs and ranking the plurality of pairs using the determined probabilities. The method may further include providing to a user, via a user interface, one or more identified second protein chains / pairs, a portion thereof, or information derived therefrom, and / or one or more probabilities that one or more query pairs form a functional antigen-binding protein or information derived therefrom (e.g., ranking of query pairs according to the probability / score that a query pair forms a functional antigen-binding pair). The method may further include predicting a score indicative of the probability that each one or more query pairs comprising a respective candidate second protein chain sequence forms a functional antigen-binding protein, and identifying candidate second protein chain sequences by applying one or more criteria to the scores / probabilities. The one or more criteria may be independently selected from: the score / probability being higher than a predetermined cut-off value; the score / probability being the highest predicted score / probability of a group of candidate second protein chain sequences; and the score / probability being in a predetermined top percentile of the predicted probabilities of a group of candidate second protein chain sequences.

[0028] According to a second aspect, there is provided a method of providing antigen-binding protein chain pairing for a plurality of query sequences comprising a first chain sequence, the method comprising: performing the method of any embodiment of the first aspect for each query sequence. The plurality of query sequences may be heavy or light chain sequences obtained by high-throughput B cell repertoire sequencing. The plurality of query sequences may comprise at least 10, at least 100, at least 1000, at least 10,000 or at least 100,000 sequences. The plurality of query sequences may have been obtained by high-throughput B cell sequencing of a heavy or light chain repertoire in a sample (e.g., a sample from a subject). The plurality of sequences may be a subset of a group of sequences obtained by high-throughput B cell sequencing of a heavy or light chain repertoire in a sample. The method according to the present aspect may have any of the features described in relation to the first aspect.

[0029] According to a third aspect, there is provided a method of providing an antigen-binding protein having desired properties, the method comprising: providing one or more query sequences comprising a first chain sequence, wherein at least one of the one or more query sequences may have a desired property, and identifying a corresponding chain sequence for each of the one or more query sequences using the method of any embodiment of the first aspect. The method may have any one or more of the following features.

[0030] The method may further comprise obtaining one or more candidate antigen-binding proteins, each candidate antigen-binding protein comprising one of the query sequences and one or more identified second sequences. The method may further comprise testing one or more candidate antigen-binding proteins for desired properties. The methods of this aspect may have any of the features described with respect to the first or second aspect. The one or more candidate antigen-binding proteins may be antibodies or fragments thereof. Sequences derived from the identified chain pairings may include: sequences having the same CDRs but different framework regions, sequences having one or more mutations compared to the identified chain pairing, and sequences having one or more fragments of the identified chain pairing. Obtaining the candidate antigen-binding proteins may include identifying the coding sequences of the candidate antigen-binding proteins and expressing the sequences in a suitable expression system (such as in a suitable host cell). The desired properties may be desired binding properties (such as the ability to bind to one or more targets, the ability to bind to one or more targets with an affinity higher than one or more corresponding thresholds, etc.), desired expression properties (such as an increased expression level compared to a standard in one or more expression systems, an expression level higher than a predetermined level in one or more expression systems, a yield higher than a predetermined level in one or more expression systems, etc.), desired stability properties (such as stability higher than a specific threshold under one or more conditions), or combinations thereof. The desired properties may include the ability to bind to a predetermined target. Testing one or more candidate antigen-binding proteins for desired properties may include identifying one or more antigens that bind to the one or more candidate antigen-binding proteins, for example by testing binding to one or more candidate antigens. Testing one or more candidate antigen-binding proteins for desired properties may include identifying one or more antigens that may bind to the one or more candidate antigen-binding proteins, for example by comparison with one or more antibodies having known targets. The antigen-binding protein may be a therapeutic antibody, and the desired properties may include binding to a therapeutic target. The antigen-binding protein may also be referred to herein as an "immunoprotein".

[0031] Testing for the presence or absence of a desired phenotype in an organism (such as an animal model) or cell expressing one or more candidate antigen-binding proteins may be included in determining a desired property of one or more candidate antigen-binding proteins. Identifying the presence of a desired phenotype may include expressing one or more candidate antigen-binding proteins in one or more model cells (such as one or more cell lines) or organisms (such as one or more animal models). The method may also include optimizing the sequence of at least one of the one or more candidate antigen-binding proteins. Optimizing the sequence of a candidate antigen-binding protein may be carried out, for example, using any antibody optimization technique known in the art. Optimizing the sequence of a candidate antigen-binding protein may be carried out using information from sequence data in which chain pairing is identified (such as by analyzing sequences similar to the input sequences in which chain pairing is identified). Methods for optimizing antigen-binding proteins are known in the art and include those described in Mason et al.

[2021] , Seeliger et al.,

[2015] , Warszawski et al.

[2019] , Hsiao et al.

[2019] and Richardson et al.

[2021] among others. Any of these methods may be used within the scope of the present invention.

[0032] The query sequence may comprise the heavy chain sequence (or a portion of the heavy chain sequence) of a known antibody. Thus, the first chain may be the heavy chain sequence of a known antibody or a portion of the heavy chain sequence of a known antibody. The query sequence may have been obtained by massive BCR sequencing of the heavy chain repertoire in one or more samples. The method may comprise the step of obtaining the query sequence by massive BCR sequencing of the heavy chain repertoire in one or more samples. The one or more samples may be from one or more subjects. The one or more subjects may have been identified as having a desired characteristic, such as a particular clinical phenotype or clinically relevant characteristic, such as a biomarker profile. For example, the one or more subjects may be resistant to a particular disease or disorder. The disease or disorder may be selected from cancer (such as breast cancer), neurodegenerative diseases (such as amyotrophic lateral sclerosis), and infectious diseases (such as COVID-19). The method may comprise identifying chain pairings (such as heavy-light pairings) for a plurality of query chain sequences (such as heavy chain sequences) that are the first (such as, heavy) chain sequences identified in the one or more samples, thereby obtaining a set of chain pairings (such as heavy chain-light chain pairings). The method may further comprise identifying one or more targets by screening antibodies against a plurality of candidate peptides from the same source as the one or more samples. The plurality of candidate peptides may be selected based on the species from which the one or more samples are derived. For example, the source of the one or more samples may be one or more human subjects, and the antibody repertoire from the same source as the one or more samples may be screened against a set of candidate peptides representing the human peptidome to select the plurality of candidate peptides. Identifying an antigen that binds to one or more candidate antigen-binding proteins may comprise using one or more targets that are identified by screening antibodies against a plurality of candidate peptides from the same source as the one or more samples. The method may further comprise filtering the identified set of chain pairings based on one or more criteria. The one or more criteria may be applied to the identity of the antigen or group of antigens to which the candidate antigen-binding protein binds or is predicted to bind. Providing one or more query sequences may comprise providing a first query (such as, heavy) chain sequence and a second query (such as, heavy) chain sequence, and identifying the second (such as, light) chain sequence for each of the one or more query sequences may comprise identifying one or more second (such as, light) chain sequences for the first query sequence and one or more second (such as, light) chain sequences for the second query sequence. The method may further comprise comparing a previous second chain sequence and a subsequent corresponding chain sequence to identify one or more light chains that may be suitable as a common second (such as, light) chain for a bispecific antibody comprising the two first (such as, heavy) chains.For example, one or more candidate second chain sequences can be the same or at least partially overlapping for a first query and a second query, and one or more candidate second chain sequences that meet one or more criteria applied to the predicted probability of the candidate sequence forming a functional pair with the first query and the second query can be identified as suitable for use as a common second chain of a bispecific antibody. According to a fourth aspect, there is provided a method of providing a tool for predicting whether a pair of protein chains is likely to form a functional antigen-binding protein or for identifying an antigen-binding protein comprising the pair of chains, the method comprising: providing training data comprising training first and second / corresponding protein chain sequences from known antigen-binding proteins, and training a deep learning model to take as input one or more pairs of protein chain sequences and generate a score (or information derived therefrom) indicative of the probability that the / each pair of protein chain sequences is part of a functional antigen-binding protein using the training data. The deep learning model comprises an encoder module and a classifier module. The method of this aspect can have any of the features described with respect to the first aspect.

[0033] The method can have any one or more of the following features. Providing the training data can include providing unpaired training first chain sequences and second chain sequences. The unpaired training first chain sequences and second chain sequences can be referred to as pre-training data. The encoder module can comprise one or more encoders or decoders. The method can also include training a sequence-to-sequence model using the unpaired training first and / or second (e.g., heavy and / or light) chain sequences, and initializing the encoder of the encoder module using the encoder of the pre-trained model. The encoder module can comprise one or two such encoders, which can each be an encoder of: a BERT model or a variant thereof, such as BERT, RoBERTa, RoFormer, SpanBERT (Joshi et al. 2020) or DistilBERT; or one or two decoders, which can each be a decoder of a decoder only language model (e.g., Falcon, Llama, GPT3 and their variants). The first and second transformer-based models can each comprise a RoBERTa model, a BERT model or a RoFormer model. The encoder module can comprise a decoder only model, such as a Falcon model (e.g., the embedding layer and decoder layer of the Falcon model). The method can also include providing the trained deep learning model to a user. Unless the context otherwise indicates, the methods described herein are computer-implemented, such as in the case of obtaining, processing, analyzing a sample, or producing, testing a molecule or composition, or using a molecule or composition for any other purpose.

[0034] According to a fifth aspect, there is provided a system comprising: a processor; and a computer-readable medium containing instructions which, when executed by the processor, cause the processor to perform the steps of the method of any embodiment of any of the foregoing aspects. The instructions may cause the processor to perform the steps of the method of any embodiment of the first to fourth aspects.

[0035] According to a sixth aspect, there is provided one or more computer-readable media containing instructions which, when executed by one or more processors, cause the one or more processors to perform the steps of the method of any embodiment of any of the foregoing method aspects. The instructions may cause the processor to perform the steps of the method of any embodiment of the first to fourth aspects.

[0036] According to a seventh aspect, there is provided a computer program product containing instructions which, when executed by one or more processors, cause the one or more processors to perform the steps of the method of any embodiment of any of the foregoing method aspects. The instructions may cause the processor to perform the steps of the method of any embodiment of the first to fourth aspects. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 is a flowchart schematically showing a method of an identification chain pair according to the present disclosure.

[0038] Figure 2 shows an embodiment of a system for an identification chain pair according to the present disclosure.

[0039] Figure 3 shows a training procedure for an exemplary deep learning model described herein. The method includes: obtaining unpaired antibody sequences, pre-training an encoder model using this data and a masked language modeling (MLM) task, combining the pre-trained encoder (AntiBERTa) with a cross-attention block, and fine-tuning a model containing the pre-trained encoder and the cross-attention block using paired antibody sequences (true pairs and random pairs) to predict the probability of belonging to a true pair or a random pair class.

[0040] Figure 4 shows in more detail Figure 3 the training procedure for the exemplary deep learning model described in. A. Creation of training sets, validation sets, and test sets for the masked language model task. B. Establishment of the pre-training procedure and how the "warmed up" model proceeds to subsequent steps. C. An overview of how the warmed-up model can be used as part of a classifier to predict the likelihood of a chain pair forming a functional chain pairing.

[0041] FIG. 5 schematically shows a masked language modeling task (A) for pre-training an encoder model to learn antibody chain sequence features, and a true pair versus random pair classification task (B) for fine-tuning a classifier model comprising the pre-trained encoder.

[0042] FIG. 6 schematically shows the structure of the deep learning model described herein (A) and the structure of three cross-attention blocks used in such a model (B).

[0043] Figure 7 Shows the classification performance evaluation results of the deep learning model described herein on an independent test dataset that includes true pairs from single-cell sequencing data and decoy random pairs generated from the same dataset. A. Receiver operator characteristic (ROC) curve. B. Precision-Recall (PR) curve.

[0044] Figure 8 Schematically shows the structures of the heavy and light chains. A. Structure of an Ig heavy chain. The approximate boundaries of the V, D, and J genes are marked along with the boundaries of the junction. The segment of the V gene towards the N-terminus is within the dashed boundary due to the short read lengths from many NGS methods not covering this region and / or some primers used for NGS slightly inserting into the V region. However, it is still possible to infer the V gene using the sequence within the solid boundary. B. The same as A., but for the light chain.

[0045] FIG. 9 schematically shows the training of the exemplary deep learning model described herein. A. Pre-training a generative transformer model (FAbCon) using a next token prediction task. B. Structure of the FabCon-small model. C. Fine-tuning the FabCon model to identify chain pairs. D. Structure of the deep learning model described herein for identifying chain pairs. DETAILED DESCRIPTION

[0046] In describing the present invention, the following terms will be employed and are intended to be defined as follows.

[0047] The B cell receptor is a transmembrane protein expressed on the surface of B cells. The B cell receptor comprises a binding portion (also referred to as the "antigen-binding subunit" or "membrane immunoglobulin", "mIg") and a signal transduction portion, the binding portion comprising a membrane-bound immunoglobulin molecule (also referred to as an antibody) that recognizes a cognate antigen. The membrane-bound immunoglobulin molecule comprises two immunoglobulin light chains and two immunoglobulin heavy chains and is identical to the corresponding secreted antibody except for the intact membrane domain. The signal transduction portion is a heterodimer called Ig-α / Ig-β (CD79), which is bound together by a disulfide bridge and binds to the immunoglobulin. An antibody (Ab) or immunoglobulin (Ig) is an immunoprotein that comprises an antigen-binding site and a constant region belonging to one of a limited set of isotypes (IgA, IgD, IgE, IgG, or IgM) and mediates interactions with other components of the immune system. In humans and most mammals, an antibody comprises four polypeptide chains: two identical heavy chains and two identical light chains that are linked by disulfide bonds. The light chain generally consists of a variable domain V L and a constant domain C L whereas the heavy chain generally comprises a variable domain V H and three to four constant domains C H 1, C H 2…. The variable domains form the antigen-binding region and may also be referred to as the F V region. Each variable domain contains three hypervariable regions called complementarity-determining regions (CDRs) that together form the antigen-binding site. The variable region of each immunoglobulin heavy or light chain is encoded by several segments (pieces) - called gene segments (subgenes): the Ig heavy chain contains variable (V), diversity (D), and joining (J) segments, while the Ig light chain contains V and J segments. Multiple copies of the V, D, and J gene segments are present in the genome, and developing B cells assemble the Ig variable region by (almost) randomly selecting and combining one V, one D, and one J gene segment (or one V and one J segment in the light chain) in a process called V(D)J recombination. This process involves the formation of double-strand breaks between the required segments, which form hairpin loops that are subsequently joined together. The joining process is imprecise, resulting in variable addition or deletion of nucleotides between the V and J (light chain) or V and DJ and D and J (heavy chain) segments, generating a large diversity in the sequence at the junctions between the segments (called "junctional sequences"). The V(D)J recombination process generates new amino acid sequences in the antigen-binding region of the Ig, resulting in a large number of different antigen recognition capabilities. As a result of this process, each Ig heavy chain variable region contains: a V segment, a D segment, and a J segment, and a junctional sequence spanning the junctions between these segments (as Figure 8as shown in A). Similarly, each Ig light chain hypervariable region contains: a V segment and a J segment, and a joining sequence spanning the join between these segments (as Figure 8 shown in B). Within the variable region, CDR1 and CDR2 are present within the V segment, and CDR3 contains some of V, all of D (in the heavy chain), and some of the J segment.

[0048] T cell receptors are membrane-anchored proteins expressed on the surface of T cells. T cell receptors comprise pairs of protein chains that together form the binding portion that recognizes cognate antigens. In mammals, these are expressed as a complex with the invariant T cell co-receptor chain CD3, which comprises the CD3γ chain, the CD3δ chain, and two CD3ε chains. The invariant chains associate with the T cell receptor and the invariant ζ chain to form the TCR complex, which together are able to generate a signal when an antigen binds to the T cell receptor. The TCR is a heterodimeric protein that comprises two highly variable chains, the α and β chains (in most T cells), or alternatively the γ and δ chains (in a minority of T cells). Each chain comprises two extracellular domains: a variable region (or variable domain) and a constant region (or constant domain, proximal to the cell membrane), a transmembrane region, and a short cytoplasmic tail. In the case of the αβ TCR, within the context of MHC (major histocompatibility complex) molecules, the variable regions together bind to a peptide (antigen). Each variable domain contains three hypervariable regions called complementarity-determining regions (CDRs, called CDR1, CDR2, and CDR3 respectively on each chain), which together form the antigen-binding site. The TCR is a member of the immunoglobulin superfamily, which includes the BCR and antibodies. In a process similar to that explained above, the variable region of each TCR chain is encoded by several segments - called gene segments (subgenes): the β and δ chains contain variable (V), diversity (D), and joining (J) segments, while the α and γ chains contain V and J segments. Multiple copies of the V, D, and J gene segments are present in the genome, and developing T cells assemble the variable region of the TCR chain by (almost) randomly selecting and combining one V, one D, and one J gene segment (or one V and one J segment in the α / γ chains) in a process called V(D)J recombination. This process involves the formation of double-strand breaks between the required segments, which form hairpin loops that are subsequently joined together. The joining process is imprecise, resulting in variable addition or deletion of nucleotides between the V and J (α / γ chains) or V and DJ and D and J (β / δ chains) segments, generating great diversity in the sequence at the junction between the segments (called the "junction sequence"). The V(D)J recombination process generates new amino acid sequences in the antigen-binding region of the TCR, resulting in a large number of different antigen recognition capabilities. As a result of this process, each variable region of the β / δ chain contains: a V segment, a D segment, and a J segment, as well as the junction sequence spanning the junction between these segments. Similarly, each hypervariable region of the α / γ chain contains: a V segment and a J segment, as well as the junction sequence spanning the junction between these segments. Within the variable region, CDR1 and CDR2 are present within the V segment, and CDR3 contains some of the V, all of the D (in the heavy chain), and some of the J segment.

[0049] The "variable chain" of an antigen-binding protein (also simply referred to herein as "chain") used in this text refers to the chain of the antigen-binding protein involved in antigen recognition, or a part thereof, i.e., at least a part of the variable region containing that chain. The variable chain contains the variable regions responsible for the diverse repertoire of antigen recognition properties within the antigen-binding protein. The variable chain can be a BCR heavy or light chain, an antibody heavy or light chain, a TCR α or β chain, a TCR γ or δ chain, or any part of such a chain containing at least a part of one or more variable regions of these chains.

[0050] Sequencing methods can be used to study the B cell receptor repertoire (or the corresponding antibody repertoire) present in a sample. As described above, two main sequencing methods are used: single B cell sequencing and sequencing of large populations of B cells. Since the BCR signaling portion and the transmembrane domain of the antigen-binding portion are invariant, these techniques focus on the part common between the B cell repertoire and the corresponding antibody repertoire. Thus, in the context of the present disclosure, references to BCR sequences, BCR repertoires, BCR heavy chain sequences, BCR light chain sequences, and any part thereof can be used interchangeably with the corresponding antibody sequences, antibody repertoires, antibody heavy chain sequences, antibody light chain sequences, and their corresponding parts. For example, a reference to equally sequencing the variable region of the BCR heavy chain is equivalent to sequencing the corresponding variable region of the antibody heavy chain, and these two terms can be used interchangeably. The term "antigen-binding protein" as used herein refers to a BCR protein, a TCR protein, the antigen-binding portion of a BCR protein, an antibody, or any part thereof that retains the antigen-binding properties of the original BCR protein. A TCR protein or an antibody. Note that the antibody repertoire circulating in an individual's blood may not match the B cell receptor repertoire present in a sample at the same time point. This is because antibodies produced by B cells that are no longer present in the individual (e.g., because they have died) may be present in the sample. Thus, the term "corresponding antibody repertoire" refers to the antibody repertoire expressed by the B cells present in the sample, rather than the library of antibodies (proteins) actually present in the sample.

[0051] Single B cell sequencing can maintain the correspondence between heavy and light chain sequences. Two main methods can be used to do this. The first method is the physical ligation of VH and VL [DeKosky et al., 2016]. The second method is cell barcoding (e.g., as provided by 10×Genomics) [King et al., 2021]. The physical ligation method has higher throughput compared to the cell barcoding method, but it is more difficult to recover full sequences. In contrast, cell barcoding has lower throughput but allows for easier recovery of full sequences. Regardless of the method, single B cell sequencing is (to varying degrees) limited in terms of throughput, as described above. Some single B cell sequencing techniques are additionally limited in terms of the length of the sequences recovered. Thus, BCR / antibody sequences identified using some single B cell sequencing methods may be limited to studying a single CDR region, e.g., CDR3 in both the heavy and light chains (in other words, although the flanking V and J segments may be identified, they may not be fully sequenced to obtain the sequences of CDR1 and CDR2 in the V segment). In other words, datasets from single B cell sequencing methods may vary in the extent to which heavy and light chain sequences are identified. Within the sequenced regions, it may also be impractical to sequence (or record) each single base of the V(D)J segments, and thus sequencing efforts may focus on obtaining junction sequences and sufficient information to identify the V, D, and J genes. Thus, the information that such methods can provide includes: the identity of the V, D, and J segments of the heavy chain (e.g., in the form of V- / D- / J gene segment identifiers), the sequence of the junction segment in the heavy chain, the identity of the V and J segments of the light chain (e.g., in the form of V- / J gene identifiers), and the sequence of the junction segment in the light chain. The identity of the corresponding segments can be used to retrieve the corresponding germline sequences from a database. However, in cases where the data only contains the identity of the corresponding segments, any mutations (e.g., somatic mutations) that may be present in a particular chain may not be captured compared to the reference germline sequence. In contrast, sequencing of large B cell populations cannot maintain the pairing between heavy and light chain sequences, but is less limited in terms of the sequencing capabilities (especially the sequencing depth of the BCR repertoire) within the heavy and light chain repertoires separately. Sequencing of large B cell populations can include sequencing of the heavy chain repertoire, the light chain repertoire, or both of the B cell population. However, as described above, due to the large nature of the process, even if both the light and heavy chain repertoires are sequenced, it is not possible to maintain pairing information during the sequencing process. Such sequencing can produce information that is as sparse as that obtained from single cell B sequencing, or can produce more detailed information (including, for example, full CDR sequences, multiple CDR sequences, full variable region sequences, or full variable region sequences and sufficient constant regions) to determine the isotype of the sequence. Similar considerations apply to sequencing of the T cell repertoire. In particular, many of the processes and limitations described above related to B cell receptor and antibody studies (and especially related to the sequencing of these repertoires) apply to the T cell repertoire.

[0052] As used herein, the term "variable chain sequence" encompasses the terms "heavy chain sequence", "light chain sequence", "alpha chain sequence", "beta chain sequence", "gamma chain sequence", and "delta chain sequence" and refers to any information obtainable from B cell sequencing or T cell sequencing techniques, ranging from a combination of one or more gene segment identifiers and / or joining sequences at one end to the full chain sequence at the other end. In particular, the terms "heavy chain sequence" and "light chain sequence" refer to any information obtainable from B cell sequencing techniques, ranging from a combination of one or more gene segment identifiers and / or joining sequences at one end to the full chain sequence at the other end. In addition, the terms "variable chain sequence", "heavy chain sequence", "light chain sequence", "alpha chain sequence", "beta chain sequence", "gamma chain sequence", and "delta chain sequence" may interchangeably refer to an amino acid sequence or the corresponding nucleic acid coding sequence. Similarly, variable chain pairing or variable chain pairs (e.g., heavy chain-light chain pairing or heavy chain-light chain pairs) refer to a combination of a heavy chain sequence and a light chain sequence, an alpha chain sequence and a beta chain sequence, or a gamma chain sequence and a delta chain sequence as defined herein, each ranging from a combination of one or more gene segment identifiers and / or joining sequences at one end to the full chain sequence at the other end.

[0053] In the context of providing a desired antibody or antigen-binding protein (e.g., a therapeutic antibody), the term "antibody" (Ab) includes monoclonal antibodies, polyclonal antibodies, multispecific antibodies (e.g., bispecific antibodies), and antibody fragments (e.g., scFv) that exhibit the desired biological activity and contain a heavy chain-light chain pairing as identified herein or a heavy chain-light chain pairing derived from a heavy chain-light chain pairing as identified herein (e.g., by further optimization, affinity maturation, etc.).

[0054] As used herein, a "sample" can be a cell or tissue sample, a biological fluid, an extract (e.g., a DNA or RNA extract obtained from a subject) from which B cell genomic material (e.g., RNA or DNA) can be obtained for genomic analysis, e.g., by sequencing (e.g., whole genome sequencing, whole exome sequencing, targeted / capture sequencing, RNA-seq, etc.). The sample can be a cell, tissue, or biological fluid sample obtained from a subject (e.g., a biopsy). Such a sample can be referred to as an "object sample". In particular, the sample can be a blood sample, a lymph node sample, a spleen sample, or a tumor sample, or a sample derived therefrom (e.g., by B cell purification, T cell purification, RNA extraction, etc.). Unless the context otherwise indicates, terms such as "genomic material", "genomic sequencing", etc. as used herein encompass the material / sequences present in the genome and transcriptome of the sample. The sample can be a sample freshly obtained from a subject, or can be a sample that has been processed and / or stored prior to genomic analysis (e.g., frozen, fixed, or subjected to one or more purification, enrichment, or extraction steps). The sample can be a cell or tissue culture sample. Thus, the sample described herein can refer to any type of sample containing B cells or genomic material derived therefrom, whether a biological sample obtained from a subject or a sample obtained from, e.g., a cell line. The sample is preferably from a mammal (e.g., a mammalian cell sample or a sample from a mammalian subject (e.g., a cat, dog, horse, donkey, sheep, pig, goat, cow, mouse, rat, rabbit, or guinea pig)), more preferably from a human (e.g., a human cell sample or a sample from a human subject). In addition, the sample can be transported and / or stored, and collection can occur at a location remote from the location of sequence data collection (e.g., sequencing), and / or any of the computer-implemented method steps described herein can occur at a location remote from the sample collection location and / or from the genomic data collection (e.g., sequencing) location (e.g., the computer-implemented method steps can be performed by a networked computer, e.g., via a "cloud" provider).

[0055] The term "sequence data" refers to information indicating the presence of genomic material (DNA or RNA) or proteomic material in a sample having a specific sequence. Thus, sequence data can comprise one or more nucleotide sequences and / or one or more amino acid sequences. Such information can be obtained using the following sequencing techniques: for example, next generation sequencing (NGS), such as whole exome sequencing (WES), whole genome sequencing (WGS), whole transcriptome sequencing (RNAseq), or sequencing of captured genomic loci (targeted or panel sequencing). When using NGS techniques, sequence data can comprise a count of the number of sequencing reads having a specific sequence. Sequence data can be mapped to a reference sequence, such as a reference genome, using methods known in the art (e.g., such as Bowtie (Langmead et al., 2009)). Thus, a count of sequencing reads or an equivalent non-numeric signal can be associated with a specific position or locus (where "position" refers to the position to which the sequence data is mapped in the reference genome or transcriptome). Additionally, the position can contain a mutation, in which case the count of sequencing reads or an equivalent non-numeric signal can be associated with each possible variant (also referred to as an "allele") of the specific position. The process of identifying the presence of a mutation at a specific position in a sample is referred to as "variant calling" and can be performed using methods known in the art (e.g., such as general purpose NGS variant callers, such as GATK HaplotypeCaller, gatk.broadinstitute.org / hc / en-us / articles / 360037225632-HaplotypeCaller or tools designed specifically for immune sequences, such as IgBLAST, www.ncbi.nlm.nih.gov / igblast / , [Ye et al., 2013]). As known in the art, genomic sequence data can be converted into an amino acid sequence by translating the coding region (in silico) (either directly from the mRNA sequence or from the coding region identified in the genomic sequence).

[0056] As used herein, "treatment" refers to a reduction, alleviation, or elimination of one or more symptoms of a disease being treated, relative to the symptoms prior to treatment. "Prevention" (or prophylaxis) refers to delaying or preventing the onset of disease symptoms. Prevention can be absolute (such that no disease occurs), or can be effective only in some individuals or for a limited amount of time. The compositions described herein can be pharmaceutical compositions that additionally comprise a pharmaceutically acceptable carrier, diluent, or excipient. The pharmaceutical compositions can optionally comprise one or more additional pharmaceutically active polypeptides and / or compounds. Such formulations can be, for example, in a form suitable for intravenous infusion.

[0057] As used herein, the term "computer system" includes hardware, software, and data storage devices for embodying a system according to the above-described embodiments or for performing a method according to the above-described embodiments. For example, a computer system can comprise a central processing unit (CPU), a graphical processing unit (GPU), an input device, an output device, and a data memory, which can be embodied as one or more connected computing devices. Preferably, the computer system has a display or comprises a computing device having a display to provide a visual output display (e.g., in the design of a business process). The data memory can comprise RAM, a disk drive, or other computer-readable media. The computer system can comprise a plurality of computing devices connected via a network and capable of communicating with each other via the network. It is expressly contemplated that the computer system can consist of or comprise cloud computers. The term "processor" encompasses any processing unit or combination of processing units, particularly including the CPU and GPU. As used herein, the term "computer-readable media" includes, but is not limited to, any non-transitory media or media directly readable and accessible by a computer or computer system. The media can include, but is not limited to, magnetic storage media, such as floppy disks, hard disk storage media, and magnetic tape; optical storage media, such as optical discs or CD-ROMs; electrical storage media, such as memories, including RAM, ROM, and flash memory; and hybrids and combinations thereof, such as magnetic / optical storage media. Identifying variable chain pairs

[0058] The present disclosure provides methods for identifying functional variable chain pairs and / or predicting whether a variable chain pair is likely to be functional. Illustrative methods will be described with reference to Figure 1 to illustrate the methods. Figure 1 Embodiments are shown in which the heavy or light chain sequences of a B cell receptor or antibody are used to identify heavy chain-light chain pairs. In other words, Figure 1 Embodiments are shown in which the variable chain sequences are BCR / heavy and light chains of antibodies from the BCR. However, with reference to Figure 1The described method is applicable to embodiments in which TCRα, β, γ, or δ chain sequences are used to identify αβ (if the query chain or chain pair is an α or β chain or an αβ chain pair) or γδ (if the query chain or chain pair is a γ or δ chain or a γδ chain pair) chain pairs. In optional step 10, a sample containing B cell genomic material can be obtained from an object (usually in the form of RNA, where the RNA encoding the BCR expressed by the cells from which the B cell genomic material is derived can be extracted and sequenced). Similarly, a sample containing T cell genomic material can be used in embodiments in which TCR chain pairs are identified. In optional step 12, bulk BCR sequencing can be used to sequence the BCR repertoire in the sample. This can include sequencing the heavy chain BCR repertoire in the sample and / or the light chain BCR repertoire in the sample. Similarly, bulk TCR sequencing can be used to sequence the TCR repertoire in the sample. This can include sequencing the β chain repertoire and / or the α chain repertoire in the sample. In step 14, a query chain sequence or sequence pair is provided. In the illustrated embodiment, the query sequence is a heavy chain sequence. In other embodiments, the query chain sequence can be a light chain sequence. In other embodiments, the query can be a chain sequence pair. Providing the query sequence can include, in step 14A, selecting the query sequence as one of the heavy chain sequences sequenced in step 12. Providing the query sequence pair can include, in step 14A, selecting the query sequence as one of the heavy chain sequences sequenced in step 12 and selecting the query sequence as one of the light chain sequences sequenced in step 12 or a light chain sequence obtained from a database or other source. Providing the query sequence or sequence pair can include, in step 14B, providing a sequence (or a sequence pair each containing a V gene sequence or identifier, a J gene sequence or identifier, and a joining sequence) containing a V gene sequence or identifier, a J gene sequence or identifier, and a joining sequence. For example, step 14B can include extracting, from a bulk BCR sequencing dataset, the V gene sequence or identifier, the J gene sequence or identifier, and the joining sequence for a selected sequence (each in the sequence pair). Similar steps can be performed in the case of TCR pairing, for example, using a query β chain sequence. When a single query sequence is provided, step 14 can include step 14C, which is to select one or more candidate sequences for pairing with the query sequence. These can be obtained from a database, a computing device (including, for example, by simulation), or a user interface. In optional step 16, a deep learning model is provided, where the deep learning model is configured to take as input a query variable chain sequence pair (which can include the query sequence and a candidate sequence for pairing, or a query sequence pair) and produce as output a score indicating the probability that the sequence pair forms a functional pair. In the illustrated embodiment, the query sequence pair is a heavy chain-light chain sequence pair, and thus the deep learning model is configured to take as input the query heavy chain sequence and the query light chain sequence and produce as output a score indicating the probability that the query heavy chain sequence and the query light chain sequence form a functional pair.In some other embodiments, the query chain sequence can be a heavy chain sequence, and thus the deep learning model can be configured to take as input a query heavy chain sequence and a plurality of candidate light chain sequences, and produce as output a corresponding score indicating the probability that the query heavy chain sequence and each candidate light chain sequence form a functional pair, or information derived therefrom (e.g., a ranking of multiple pairs or multiple candidate light chain sequences based on their respective scores). In some other embodiments, the query chain sequence can be a β chain sequence (or an α, δ, or γ chain sequence), and thus the deep learning model can be configured to take as input a query β chain sequence (or an α, δ, or γ chain sequence) and one or more candidate α chain sequences (or β, γ, or δ chain sequences), and produce as output a score indicating the probability that the query β chain sequence and each candidate α chain sequence (or each corresponding candidate β, γ, or δ chain sequence) form a functional pair, or information derived therefrom (e.g., a ranking of multiple candidate pairs). The deep learning model can be pre-trained using training variable chain sequences from known variable chain pairs, such as the training heavy and light chain sequences from known heavy chain-light chain pairs in the illustrated embodiments. Thus, providing the deep learning model can simply include retrieving the pre-trained deep learning model from a computer-readable medium (e.g., a memory associated with the processor executing the method), or otherwise receiving the pre-trained deep learning model. The training of the deep learning model is explained in more detail below. Alternatively, the deep learning model can be trained at least in part using training variable chain sequences from known heavy chain-light chain pairs (such as the training heavy and light chain sequences from known heavy chain-light chain pairs in the illustrated embodiments) and optionally using training variable chain sequences from known unpaired heavy and light chains, as part of the method of the present invention.

[0059] In step 18, a sequence pair comprising a query chain sequence and a candidate corresponding chain sequence, or a query chain sequence pair, is provided to the deep learning model. Step 18 can include optional step 18A of encoding the sequences in the sequence pair using a predetermined encoding scheme. The encoding scheme used can have been predefined based on the content of the training variable chain sequences (e.g., heavy and light chain sequences) used to train the deep learning model. Step 18 can include optional step 18B of selecting the sequence pair among the plurality of sequence pairs associated with the highest probability of forming a functional pair (e.g., the highest score / highest probability or highest ranking). In optional step 20, the results of any previous step (and particularly step 18) can be provided to the user, for example via a user interface. These results can be used, for example, to provide a therapeutic antibody, as will be further described below. The method can be repeated for a plurality of query sequences. This can include repeating steps 14 to 18.

[0060] The deep learning model includes an encoder module. The encoder module refers to a machine learning model that has been trained to take sequence data (protein chain sequence) as input and produce an embedded representation of the sequence data as output, or produce a decoded or generated sequence as output from the learned embedded representation of the sequence data. Thus, the sequence encoder can be the encoder model of a bidirectional encoder-only model, the encoder model of an encoder-decoder model, or the decoder model of a decoder-only model (e.g., a generative model). For example, generative models with architectures such as those in LlaMa (Touvron et al., 2023), Falcon LLM (falconllm.tii.ae / ), and GPT (e.g., GPT-3, Brown et al. 2020) can be used. Alternatively, a structure based on a Transformer encoder, such as AntiBERTa (see the examples and Leem et al. 2022), can be used. When using a decoder model, the embedded representation can be obtained as an embedding from the final decoder layer of the decoder. The encoder module may also be referred to as an "encoder" in this document, but its structure is not limited to that of an encoder. The embedded representation is a representation in the latent space, and is typically configured such that the original sequence data can be reconstructed from the embedded representation. The encoder module is a deep learning model. The encoder module can be trained as part of a sequence-to-sequence model. The encoder module can be a Transformer-based encoder or decoder. The deep learning model containing the sequence encoding module may have been trained using a masked language modeling (MLM) task or a causal language modeling (CLM, next token prediction) task. The former can be particularly used for encoder-only models and encoder-decoder models, while the latter can be particularly used for generative models.

[0061] The training of the deep learning model will now be explained with reference to optional steps 10'-16'. In step 10', training data is provided that includes at least training variable chain sequences from known variable chain pairs. In the illustrated embodiment, the training data includes heavy and light chain sequences from known heavy-light chain pairs. The training data can include at least 20,000 training chain pairs, at least 30,000, at least 40,000, at least 50,000, at least 60,000, at least 70,000, at least 80,000, at least 90,000, at least 100,000, at least 120,000, or at least 150,000 training chain pairs. In embodiments related to B cell receptors / antibodies, the training data can include at least 80,000 pairs, at least 100,000 pairs, at least 120,000 pairs, or at least 150,000 pairs of training heavy and light chain sequences. In some embodiments, the training data includes at least 1 million, 1.5 million, or 2 million paired sequences. The training data can also include unpaired training sequences, which in the illustrated embodiment are heavy and light chain sequences (but can be any first and second chains described herein). The unpaired training chain sequences can be referred to as "pre-training data". In such cases, the paired training data can be referred to as "fine-tuning data". In some embodiments, the paired training data is also used for pre-training. Thus, the training data can include training data for fine-tuning (including paired chain sequences, particularly paired heavy and light chain sequences in the illustrated embodiment), and pre-training data (including unpaired chain sequences, particularly unpaired heavy and light chain sequences in the illustrated embodiment, and optionally also including paired training sequences). The pre-training data can include at least 100,000, at least 200,000, at least 300,000, at least 400,000, at least 500,000, at least 600,000, at least 700,000, at least 800,000, at least 900,000, at least 1 million (or at least 5 million, 10 million, 15 million, 20 million, 25 million, 30 million, 35 million, or 40 million) unpaired training chain sequences of the first type and / or corresponding type. In some embodiments, the training data includes at least 500 million, 600 million, or 700 million independent sequences. The training data can include at least 1 million (or at least 5 million, 10 million, 15 million, or 20 million) unpaired training heavy chain sequences and at least 1 million (or at least 5 million, 10 million, or 15 million) unpaired training light chain sequences. As will be understood by those skilled in the art, the amount of training and / or pre-training data can be limited by the amount of suitable data available and can change as more data becomes available. The pre-training and / or training data can include simulated data.Simulated data may include paired or unpaired protein sequences obtained using methods with simulated antigen-binding protein sequences, such as immuneSIM (Weber et al., 2020) and / or ProGen2 (Nijkamp et al., 2022). Pretraining data may be collected and / or used for pretraining as described in Leem et al., 2022. For example, in Leem et al., 2022, unpaired chain sequences containing approximately 42 million heavy chains and 15 million light chains were used. If available, more data can be advantageously used. Additionally, the amount of available data may depend on the specific use case, such as the identity of the first chain sequence and the corresponding chain sequence (e.g., more data is available for αβ TCRs compared to the rarer γδ TCRs), the criteria used when filtering the data (see step 12’), etc. The numbers provided may apply to the data before and / or after any filtering is applied. In step 12’, the training data is filtered. For example, the training data may be filtered to exclude any chains that contain a number of amino acids outside a predetermined length range within any predetermined region of the chain. For example, the pretraining data may be filtered to exclude chain sequences that do not contain at least 20 amino acids before the CDR1 region, sequences that do not contain at least 10 amino acids after the joining sequence, sequences that do not contain 5 to 12 residues in the CDR1 region, sequences that do not contain 1 to 10 residues in the CDR2 region, and / or sequences that do not contain 5 to 38 residues in the CDR3 region. As another example, the training data may be filtered based on any feature of the data, including, for example, the cell type, organism from which the data is derived, whether the data is from an initial library, whether the data is from an object immunized with a specific antigen, etc. In other words, the training data may be filtered to ensure that the training data contains only (inclusion filter) or does not contain (exclusion filter) data with one or more desired features. In step 14’, negative training data, also known as “negative training data,” containing non-native chain pairs may be provided. Non-native chain pairs may be obtained by randomly reassigning the chains in the pairs in a paired dataset. Alternatively, non-native chain pairs may be obtained by randomly pairing chains from one or more unpaired datasets containing two types of chains (e.g., unpaired heavy chains and unpaired light chains). For example, heavy chains from a large amount of heavy chain sequencing data may be randomly paired with light chains from a large amount of light chain sequencing data. As an alternative or supplement to using randomly generated negative pairs, negative pairs may be selected as pairs known to be non-functional, such as based on previous experiments. Pairs known to be non-functional may include one or more of the following: pairs that could not be expressed in one or more previous experiments, pairs that did not possess one or more desired functional features (e.g., lack of binding to a target) in one or more previous experiments.

[0062] In step 16', the deep learning model is trained using training data to take as input a query pair (in the illustrated embodiment) comprising a heavy chain sequence and a light chain sequence, and to produce as output a score indicating the probability that the pair is functional (e.g., in the illustrated embodiment, the probability that the pair is functional) or information derived therefrom (e.g., the ranking of multiple input pairs based on scores indicating the probability that each pair is functional). Training of the deep learning model may include first training a model comprising an encoder module, such as an encoder-decoder based model or a bidirectional encoder model (also referred to as a "sequence-to-sequence" model) or a decoder-based model (e.g., a generative language model), using unpaired training chain sequences to obtain a pre-trained encoder model, and initializing (fine-tuning) the encoder module of a deep learning model comprising an encoder module and a classification module that is trained for paired sequence prediction using the pre-trained model. Alternatively, training of the deep learning model may include training first and second sequence-to-sequence models using, respectively, first type and second type unpaired training chain sequences (in the illustrated embodiment, the first type is the heavy chain and the second type is the light chain), and initializing the corresponding encoder modules of the deep learning model using the pre-trained encoder modules of these respective models. Alternatively, training of the deep learning model may include training a sequence-to-sequence model or a generative language model using first type training chain sequences and second type training chain sequences in series (in the illustrated embodiment, the first type is the heavy chain and the second type is the light chain), and initializing the encoder module of a deep learning model that takes as input a sequence pair in series using the pre-trained encoder module of such a model (which may be an encoder or a decoder, depending on the model structure). Training of the deep learning model may include obtaining training sequences by inputting missing sequence information, the training sequences comprising the full-length sequences of the variable regions of the second type chain and / or the first type chain (e.g., the light chain and / or the heavy chain in the illustrated embodiment), which is done in cases where the chain sequences from known chain pairs do not contain the full-length sequences of the variable regions. Training of the sequence-to-sequence model may include training the model using MLM. Training of the generative language model may include training the model using CLM. Step 16' may include defining one or more encoding schemes for the training data by obtaining a vocabulary for encoding the training chain sequences (in particular, in the illustrated embodiment, a vocabulary for encoding the training heavy chain sequences and a vocabulary for encoding the training light chain sequences). The encoding scheme may encode each individual amino acid as a separate token, or may encode subsets of sequences as separate tokens (e.g., full regions, k-mers within a region, or using byte pair encoding). Defining the encoding scheme may include excluding any tokens that occur fewer than a predetermined threshold number of times (e.g., 2) in the training data from the vocabulary constructed based on the content of the training chain sequences.The present inventors have found that a coding scheme representing each amino acid as a token is practical at least in the case of predicting antibody / BCR light chain - heavy chain pairing, where the training data has sufficient sequence resolution and is of a large enough quantity to constrain a model using such a granular representation. System

[0063] Figure 2An embodiment of a system for identifying a variable chain pair of an input variable chain or for predicting whether an input variable chain pair is likely to be functional according to the disclosure of the present invention is shown. The system includes a computing device 1, and the computing device 1 includes a processor 101 and a computer-readable memory 102. In the illustrated embodiment, the computing device 1 further includes a user interface 103, which is shown as a screen, but may include any other device for transmitting information to the user, such as by auditory or visual signals. The computing device 1 is communicatively connected to a sequence data acquisition device 3 (such as a sequencer) and / or one or more databases 2 storing sequence data, for example, via a network 6. One or more databases may additionally store other types of information that can be used by the computing device 1, such as reference sequences, parameters, and the like. The computing device can be a smartphone, a server, a tablet, a personal computer, or other computing devices. The computing device is configured to implement the method described herein for identifying a variable chain pair or predicting whether a variable chain pair is likely to be functional (suitably, a heavy-light chain pair or an αβ chain pair, advantageously, a heavy-light chain pair). In some alternative embodiments, the computing device 1 is configured to communicate with a remote computing device (not shown), which itself is configured to implement the method described herein for identifying a variable chain pair or predicting whether a variable chain pair is likely to be functional. In such a case, the remote computing device may also be configured to send the results of the method to the computing device. The communication between the computing device 1 and the remote computing device can be carried out through a wired or wireless connection and can be via a local or public network, such as via the public Internet or WiFi. The sequence data acquisition device 3 can be wired to the computing device 1 or may be capable of communicating via a wireless connection, such as via the network 6 shown. The connection between the computing device 1 and the sequence data acquisition device 3 can be direct or indirect (such as through a remote computer). The sequence data acquisition device 3 is configured to obtain sequence data from a nucleic acid sample, such as a genomic DNA sample or an RNA sample extracted from B cells or T cells purified from a fluid and / or tissue sample (such as peripheral blood, spleen, lymph node, tumor tissue, or any other type of sample containing B cells or T cells). In some embodiments, the sample may have undergone one or more pretreatment steps, such as DNA / RNA purification, fragmentation, library preparation, target sequence capture (such as exon capture and / or panel sequence capture). Any sample preparation method suitable for determining the B cell receptor sequence or repertoire can be used within the context of the present invention. The sequence data acquisition device is preferably a next-generation sequencer. The sequence data acquisition device 3 can be directly or indirectly connected to one or more databases 2, on which sequence data (raw or partially processed) can be stored. Application

[0064] The above methods can be applied to any situation where it is desired to identify antibodies or BCRs that may bind to their targets from limited information on heavy chains, light chains, unpaired heavy and light chains, or portions thereof (such as V genes, J genes, and joining sequences). This situation often occurs in the context of the antibody therapeutic discovery process. Antibody therapeutics have been shown to be a successful approach for a wide variety of diseases ranging from neurodegenerative diseases to cancer. Thus, the methods described herein can be used to provide therapeutics in each of these clinical settings. In addition, the methods described herein can be used to identify potential functional antibodies or BCRs from any input heavy / light chain or portion thereof, whether the input information is newly generated for a specific purpose (e.g., from a patient or sample identified as having a desired phenotype) or from an existing / historical data set (e.g., to mine or re-mine an existing data set to discover new therapeutics or to identify immunoproteins that may explain why certain clinical phenotypes persist).

[0065] Accordingly, the present invention also provides a method of providing an antibody therapeutic agent, the method comprising using any of the methods described herein to identify a heavy chain-light chain pairing, or a heavy chain-light chain pairing derived from a heavy chain-light chain pairing identified using any of the methods described herein (e.g., by further optimization, mutation, etc.). A heavy chain-light chain pairing is obtainable for an input heavy chain sequence that has been obtained by massive BCR sequencing of a heavy chain repertoire in one or more samples. A heavy chain-light chain pairing is obtainable for an input heavy chain sequence and one or more candidate light chain sequences, the input heavy chain sequence having been obtained by massive BCR sequencing of a heavy chain repertoire in one or more samples, the one or more candidate light chain sequences having been obtained by massive BCR sequencing of a light chain repertoire in one or more samples (wherein the one or more samples may be the same samples used for analyzing the heavy chain BCR repertoire, or different samples). The one or more samples may be from one or more subjects. The one or more subjects may have been identified as having a desired characteristic, such as a particular clinical phenotype or a clinically relevant characteristic, such as a biomarker profile. For example, the one or more subjects may be resistant to a particular disease or disorder. The disease or disorder may be selected from cancer (e.g., breast cancer), neurodegenerative diseases (e.g., amyotrophic lateral sclerosis), or infectious diseases (e.g., COVID-19). The method may include identifying heavy chain-light chain pairings for a plurality of input heavy chain sequences selected from the heavy chain sequences identified in the one or more samples, thereby obtaining a set of heavy chain-light chain pairings. The method may also include identifying the target (or putative target or group of targets) for each heavy chain-light chain pairing in the heavy chain-light chain pairing or set of heavy chain-light chain pairings. The method may also include identifying one or more targets by screening antibodies against a plurality of candidate peptides from the same source as the one or more samples. The plurality of candidate peptides may be selected based on the species from which the one or more samples are derived. For example, the source of the one or more samples may be one or more human subjects, and an antibody library from the same source as the one or more samples may be screened against a set of candidate peptides representing the human peptidome. Identifying the target (or putative target or group of targets) for each heavy chain-light chain pairing in the heavy chain-light chain pairing or set of heavy chain-light chain pairings may include using one or more targets that have been identified by screening antibodies against a plurality of candidate peptides from the same source as the one or more samples. The method may also include filtering the set of heavy chain-light chain pairings based on one or more criteria. The one or more criteria may be applied to the identity of the putative target or the group of targets identified for the heavy chain-light chain pairing. The method may also include obtaining an antibody or a fragment thereof that comprises the identified heavy chain-light chain pairing or a heavy chain-light chain pairing derived from the identified heavy chain-light chain pairing. Obtaining the antibody or a fragment thereof may include identifying the coding sequence of the antibody or a fragment thereof, and expressing the sequence in a suitable expression system (e.g., in a suitable host cell).The method may further comprise identifying one or more antigens that bind to the antibody or fragment thereof, e.g., by testing binding to one or more candidate antigens. The method may further comprise optimizing the sequence of the antibody or fragment thereof. Optimizing the sequence of the antibody or fragment thereof can be carried out using any antibody optimization technique known in the art. Optimizing the sequence of the antibody or fragment thereof can be carried out using information from sequence data from which heavy-chain-light-chain pairings are identified, e.g., by analyzing sequences similar to the input sequences from which heavy-chain-light-chain pairings are identified.

[0066] The present invention also provides a method for providing an immunotherapeutic composition, the method comprising identifying a heavy-chain-light-chain pairing as described herein, and generating an immunotherapeutic composition comprising an antibody containing the heavy-chain-light-chain pairing or an antibody derived from the heavy-chain-light-chain pairing (e.g., by further optimization, mutation, etc.).

[0067] The methods described herein can also be used in the context of providing bispecific antibodies. For example, the methods described herein can be used to identify a light chain that will be suitable for pairing with two different heavy chains of interest. Accordingly, the present invention also provides a method for providing a bispecific antibody, the method comprising identifying a common light-chain pairing for each of two heavy chains using any of the methods described herein, or identifying a combination of a common light chain and two heavy chains, the combination being derived from a heavy-chain-light-chain pairing identified using any of the methods described herein (e.g., by further optimization, mutation, etc.). In such an embodiment, it can be advantageous to use a deep learning model to predict scores indicative of the functional pairing probabilities of multiple candidate light-chain or heavy-chain sequences for each query sequence. For example, a deep learning model can be used to predict the score / probability that a first query heavy chain forms a functional pair with each of a first plurality of candidate light-chain sequences, and to predict the score / probability that a second query heavy-chain sequence forms a functional pair with each of a second plurality of candidate light-chain sequences. The first and second plurality of candidate light-chain sequences can advantageously be the same or can at least partially overlap. Then, the scores / probabilities predicted for the first and second plurality of light-chain sequences can be compared to identify one or more light chains that are likely to be suitable for use as a common light chain for a bispecific antibody comprising the two heavy chains. For example, candidate light chains with a relatively high probability of forming a functional pair with the two heavy chains can be used. For example, candidate light chains can be ranked based on the sum of the scores / probabilities of forming a functional pair with the first and second heavy chains (or any other combined metric that combines the two probabilities).

[0068] The methods described herein can also be used in the context of antibody optimization. For example, the methods described herein can be used to identify heavy-chain-light-chain pairings that have one or more advantageous properties (such as improved functional or developability properties) compared to an original pairing (such as the heavy chain of the pairing). For example, the heavy chain in a pair can be paired with multiple candidate light chains, and a score indicating the probability that the candidate forms a functional pair with the heavy chain can be determined. Then, based on these scores / probabilities and optionally other criteria applicable to functional or developability properties, the candidate pairs can be ranked or otherwise prioritized. Accordingly, the present invention also provides a method for providing an improved antibody, the method comprising identifying a heavy-chain-light-chain pairing from an input heavy chain (or light chain) of an original antibody using any of the methods described herein, or a heavy-chain-light-chain pairing derived from: a heavy-chain-light-chain pairing identified using any of the methods described herein (such as by further optimization, mutagenesis, etc.). The methods described herein can also be applied to any context in which it is desirable to identify a TCR that may bind its target from information limited to the β-chain, α-chain (or, less commonly, the γ or δ chain) or portions thereof (such as the V gene, J gene, and joining sequences). This often occurs in the context of the discovery process of cell therapeutics such as engineered T cells. Accordingly, the present invention also provides a method for providing a TCR-based therapeutic agent such as an engineered T cell expressing a specific TCR, the method comprising using any of the methods described herein to identify an αβ or γδ chain pairing, or an αβ or γδ chain pairing derived from: an αβ or γδ chain pairing identified using any of the methods described herein (such as by further optimization, mutagenesis, etc.). Accordingly, the methods described herein can also be used in the context of T cell receptor optimization, in a manner similar to that described above for antibodies.

[0069] The following is presented by way of example and is not to be construed as limiting the scope of the claims. Examples

[0070] These examples describe methods for identifying heavy-chain-light-chain pairings and / or predicting the probability that a chain pair is functional according to the present invention, and validate them using a single-cell dataset with known pairings. Example 1 - AntiBERTa for pairing prediction Method Dataset

[0071] The model was trained in two steps (as shown in Figure 3 and Figure 4 ): a pre-training step, in which a bidirectional transformer encoder model based on the RoBERTa architecture (Liu et al., 2019) was trained using unpaired heavy-chain and light-chain antibody sequences (step 1 in Figure 3 , Figure 4step A-B) therein; and a fine-tuning step, in which a model comprising a pre-trained Transformer encoder and a classifier module is trained using known pairs (positive examples) and random pairs (negative examples). Figure 3 step 2-3 therein, Figure 4 step C) therein. The resulting model is then tested using independent test data. The following datasets are used. The pre-training step is described in Leem et al., 2022. The pre-trained model is called "AntiBERTa". In Leem et al., 2022, the pre-trained model was fine-tuned using human antibody sequences from the Structural Antibody Database (SAbDab, Dunbar et al., 2014) for paratope prediction. In this example, the pre-trained model is directly used for fine-tuning for heavy-chain-light-chain pairing, as described below. However, the fine-tuning model described in Leem et al. 2022 (i.e., fine-tuned for paratope prediction) can also be used as a starting point for fine-tuning for pairing as described herein.

[0072] Pre-training: Human antibody sequences across 61 studies were downloaded from the OAS database (Kovaltsuk et al., 2018). First, any sequencing errors in the antibody sequences were filtered out as indicated by OAS. In addition, the sequences were required to have at least 20 residues before CDR1 and at least 10 residues after CDR3. Finally, filtering was performed to obtain sequences with 5 to 12 residues in CDR1, 1 to 10 residues in CDR2, and 5 to 38 residues in CDR3. This resulted in a maximum sequence length of 148 residues. Thereafter, the entire set of 71.98M unique sequences (52.89M unpaired heavy chains and 19.09M unpaired light chains) was split into disjoint training, validation, and test sets using a ratio of 80:10:10. Overall, the MLM training set contains 42.3M heavy-chain sequences and 15.3M light-chain sequences, while the MLM validation set and the MLM test set each consist of 5.3M heavy chains and 1.9M light chains. AntiBERTa is a single model trained on both heavy and light chains.

[0073] Fine-tuning: Fine-tuning was performed using an internal dataset of single-cell B cell receptor sequencing data, called the "Alchemab paired BCR dataset". These represent positive instances (known pairs). The classifier model requires the heavy chain and the light chain as inputs respectively. For training, the user must provide a number of 1 or 0, where 1 indicates that the heavy chain and the light chain are a true, genuine pair, and 0 indicates that the heavy chain and the light chain are not a true pair. Since there is no "negative" dataset of VH-VL pairs, it was approximated by randomly shuffling the Alchemab paired BCR dataset, and it was assumed that the randomly generated pairs are not pairable. Thus, strictly speaking, the classifier predicts whether a pair is likely a natural pair or a random pair (the former is assumed to be functional, while the latter is assumed to be non-functional or at least not pairable). A total of 1,121,212 paired sequences were used as the "positive" set, and another 1,121,212 pairs were randomly paired to generate the "negative" set. In total, there were 2,242,424 pairs of sequences. A further quality control step was performed on this set of 2,242,424 pairs of sequences (containing both natural / positive pairs and random / negative pairs), specifically removing sequences that (i) were incorrectly too long (using a cut-off value determined to be reasonable based on expertise, set to 250 residues in this example, but other values may be equally suitable based on the specific data used), or (ii) had a stop codon in the read (identified as "*" in the amino acid sequence provided as output by the single-cell sequencing analysis software - but information from the original RNA sequence reads could be used instead), or (iii) the frequency of occurrence of the heavy chain V gene and the light chain V gene pair was too low (less than 10 pairs in total). These criteria apply to the combined set of natural and random pairs.

[0074] After quality control, a total of 2,239,117 pairs remained (i.e., approximately 0.15% of the pairs were filtered out), of which 1,791,293 pairs were used for training, 223,912 pairs were used for validation, and 223,912 pairs were used for testing. These pairs were randomly assigned to the training, validation, and test datasets. Other methods are also feasible, such as sampling from the negative and positive pairs respectively to ensure an equal representation of the two classes in each data subset. Validation helps to select the state of the neural network that should have the best performance, and the test set is used to test the generalization performance of the network.

[0075] Testing: The final independent evaluation of model performance was carried out using publicly available datasets of single-cell B cell receptor sequencing data (10× single-cell data) from King et al. 2021 (7 donors, 30,222 unique heavy-light chain pairs), Eccles et al. 2020 (1 donor, 741 unique heavy-light chain pairs), and Setliff et al. 2019 (2 donors, 4,944 unique heavy-light chain pairs). These were only used for the final test and thus represent an unbiased evaluation of model performance. Sequence tokenisation

[0076] The pre-training dataset contains full-length heavy and light chain sequences, and these are tokenised at the single amino acid level. The fine-tuning and test data also contain full-length heavy and light chain sequences, which are tokenised at the single amino acid level.

[0077] If any data to be used (whether for training, testing, or model use) does not contain full-length heavy and light chain sequences, for example, because the sequencing method used to generate this data cannot recover the complete amino acid sequence of the light / heavy chain, a model trained with full-length sequences tokenised at the single amino acid level can still be used. For example, some datasets contain: (i) for each heavy chain: V gene identifier, joining sequence, J gene identifier, and D gene identifier; and (ii) for each light chain: V gene identifier, joining sequence, and J gene identifier. In such cases, each entry may contain a combination of gene identifiers and sequences, such as IGHV3-23 / CAR...DYW / IGHJ6-IGKV3-20 / CQQ... / IGKJ2. Such sequences can be converted to full amino acid sequences by inputting the germline sequences of the corresponding parts (e.g., the germline sequences of the corresponding V and J gene identifiers).

[0078] As described above, in this example, all sequences are tokenised at the single amino acid level. The vocabulary used consists of 25 tokens: the standard 20 amino acids and five special tokens ( <s>、< / s> , <pad> 、 <unk>And <mask>)。Each amino acid acts as a token, and byte pair encoding is not used. Each chain is encoded with a start token ( <s>) and end marker (< / s> ).; <pad>The token is used to pad the maximum sequence length in a minibatch for tensors. <unk>Tokens are used for ambiguous amino acids, such as X. The maximum allowed length is 150, as this covers the maximum sequence length in the pre-training dataset (148), as well as the start and end tokens. Briefly, the advantage of this is that it covers the training set in OAS while ensuring that the amount of non-essential padding is minimized.

[0079] Other tokenization schemes can also be used and are explicitly contemplated. For example, a tokenized sequence corresponding to the V gene identifier, the joining sequence, and the J gene identifier can be used to train the model. As described above, this is particularly useful when only partial sequence information is available. To process this data format, a custom encoding method can be used, where each V gene constitutes a single token, each J gene constitutes a single token, and the joining amino acid sequence is tokenized as single amino acids as described above. The joining sequence is the most diverse region of the sequence and is thought to mediate most of the binding function, thus increasing the granularity of the tokenization of this sequence. In such a scheme, tokens can be used if they occur at least, for example, 2 times in the training set. Other viable schemes for tokenization of the joining amino acid sequence (or the full sequence) include, for example, byte pair encoding, or tokenization for overlapping k-mers (such as 3-mers). If more full sequence data is available, for example, if single-cell data using approximately hundreds of thousands or even millions of sequences is used to train the model, a scheme such as byte pair encoding can be particularly useful. Using byte pair encoding, several non-overlapping amino acids are encoded as tokens using a dictionary automatically defined in a data-driven manner. Byte pair encoding is a subword tokenization scheme that replaces common consecutive byte pairs (in this case, consecutive amino acids) with bytes not present in the data. All pairs that occur more than once in the data will be replaced by the corresponding token.

[0080] In the context of this embodiment, a "sentence" is the tokenized representation of a heavy or light chain sequence. Each sentence begins with a special token <s>Start, followed by markers for each amino acid (or markers representing the V genes of the heavy or light chains, overlapping 3-mer markers, J gene markers, or markers for each consecutive amino acid pair), and then a special marker< / s> . For any sentence whose tokens are less than the maximum length (150 in the case of single amino acid encoding, 34 in the case of the custom encoding described above), the sequence uses a special <pad>The label is filled in. model

[0081] As described in Leem et al., 2022, a bidirectional transformer model based on the RoBERTa architecture (Liu et al., 2019) is pre-trained. As described above, the model is pre-trained using a dataset containing both heavy-chain sequences and light-chain sequences. It should be noted that a separate RoBERTa model could have been pre-trained on heavy-chain and light-chain sequences separately. It was found that this was not necessary because the model was able to distinguish between the two types of chains and learn the features of light-chain and heavy-chain sequences separately. This pre-trained RoBERTa model trained on unpaired antibody sequences is referred to as "AntiBERTa" in this article, and the training procedure is as Figure 4 (Steps A and B) described.

[0082] Then, two copies of the AntiBERTa model (only the encoder) are used to process the heavy chain and the light chain separately, and the outputs of these are provided to the cross-attention module and the classification module, as further described below. In Figure 4 Step C, the generated model is fine-tuned using the dataset described above (2,242,424 pairs of paired sequences, including 1,121,212 random pairs (which form a negative set assigned the label "0") and 1,121,212 true pairs (which form a positive set assigned the label "1")). In the fine-tuning step, the parameters are shared between the two AntiBERTa models (i.e., the two models are "Siamese" models). It has been shown that using pre-trained so-called "checkpoints" is a powerful strategy in the NLP environment [Rothe et al., 2020]. In particular, Rothe et al.

[2020] showed that the BERT-to-BERT architecture works well for NMT.

[0083] If the paired training data used for fine-tuning the full model (including two copies of the pre-trained AntiBERTa model) does not contain full-length BCR heavy and light chain sequences, one of two alternative methods can be used to generate an equivalent paired dataset containing full-length sequences. In the first method, full-length sequences can be obtained by replacing the V and J gene identifiers with their corresponding germline sequences. In the second method, the pre-trained AntiBERTa model (or any other such "checkpoint" model, such as the GPT-2 model) can be used to predict the full-length sequences of the training set, independent of the heavy and light chains based on the known parts of the chains (using the respective models). The predictions from the "checkpoint" model can be obtained using some or all of the known parts of the chains (such as gene segment identifiers, partial sequences, etc.), optionally combined with some information obtained from the germline sequences of any segments that are not available for the full sequence (such as, for example, the identity of some amino acids of the segment, such as the first k amino acids of the segment, where k can be, for example, 1, 2, 3, 5, 10, etc.). In other words, the AntiBERTa model trained on unpaired full-length heavy and light chain data can be used to predict the full-length sequence of the heavy chain in the training data from the V gene identifier, J gene identifier, and joining sequence provided in the data. Similarly, the AntiBERTa model trained on unpaired full-length heavy and light chain data can be used to predict the full-length sequence of the light chain in the training data from the V gene identifier, J gene identifier, and joining sequence provided in the data. The same two methods can be used to map any limited paired training data into a more extended format of data that can be used to train the "checkpoint" model. Alternatively, the data used to train the "checkpoint" model can be converted into a limited format that matches the format of the paired training data. This may still benefit from the potential additional information collected by the pre-trained model from the large amount of available unpaired sequences. However, it may not fully utilize the range of information available in such unpaired sequence data. In this embodiment, full-length sequences can be used for both pre-training and fine-tuning, and thus this is not necessary. Model Structure - Pre-training

[0084] The 12-layer bidirectional transformer model based on the RoBERTa structure (Liu et al., 2019) was pre-trained as described in Leem et al., 2022. The model has an embedding dimension of 768, a feed-forward dimension of 3072, and 12 encoder layers and 12 self-attention heads. These are all pre-defined accepted standards. The model has a total of 86 million learnable parameters. The maximum sequence length of the model is 150. It was first trained with the masked language modeling (MLM) task, which is similar to a "fill-in-the-blank". Figure 5A The MLM process used is schematically shown. During MLM pre-training, 15% of the residues in the input sequence are perturbed, where the amino acid is replaced by a mask ( Figure 5A the gray square in in 80% of the cases, replaced by a random amino acid in the remaining 10% of the cases (as described in Liu et al., 2019). The model is trained to predict the correct amino acid at the perturbed position using a loss. For a sequence S = {s1, s2, … sl} in batch B with a perturbed position M, the loss is provided by the following: Model Structure - Fine Tuning

[0085] Then, the pre-trained model is used to construct a neural network with the Figure 6A structure shown (where the AntiBERTa encoder is the encoder of the above pre-trained model) and trained for the Figure 5B classification task shown.

[0086] In the encoder module, the model processes the input heavy chain and input light chain separately using corresponding copies of a pre-trained encoder. This generates an L×768 tensor for the heavy chain sequence (where L is the length of the heavy chain sequence) and an L’×768 tensor for the light chain sequence (where L’ is the length of the light chain sequence). 768 corresponds to the 768 numbers (embedding dimension) generated by the pre-trained encoder. These can be understood as a set of 768 numbers that describe some contextual properties of the amino acids in the heavy / light chain sequence. In practice, multiple heavy chains and multiple light chains are processed at once to improve processing efficiency (reduce the computational time for processing each pair of chains). This generates a B×max(L)×768 tensor, where B is the number of heavy chain sequences and max(L) is the maximum length of the heavy chain sequences among all heavy chain sequences in the batch. Similarly, a B’×max(L’)×768 tensor is generated for B’ number of light chains. B and B’ are usually the same, and max(L) and max(L’) are not necessarily the same. These are the outputs of the encoder module.

[0087] The outputs of the encoder module are used as inputs to the cross-attention module. Specifically, the two tensors (B×max(L)×768 and B’×max(L’)×768) are then processed through six cross-attention blocks. The first three cross-attention blocks are visualized in more detail in Figure 6B The subsequent blocks follow the same structure. Each cross-attention block contains a self-attention layer and a cross-attention layer. Both the self-attention layer and the cross-attention layer contain 12 attention heads. In the self-attention layer, only the heavy chain sequence embeddings are scored using the multi-head self-attention scoring mechanism described in Vaswani et al. (2017): where the query Q, key K (dimension d k ), and value V (dimension d v ) are heavy chain tensors. This is because the model is mainly designed to utilize heavy chain features to identify light chains, as heavy chain sequences are usually more readily available and are considered more important in determining function. However, the roles of the light chain and heavy chain can be interchanged. The heavy chain sequence embeddings processed by self-attention are then layer-normalized and a residual connection is formed with the non-attention-processed heavy chain input embeddings (i.e., the heavy chain input embeddings before applying attention and normalization). This is thought to compensate for the signal reduction that may occur through layer normalization. Layer normalization and residual connections are described in Vaswani et al., 2017 for the encoder stack. The residual connection is a simple sum, in Figure 6B In the figure, it is represented by the "+" sign. The cross-attention layer performs the same calculations as the self-attention, except that the Q tensor comes from the heavy-chain embedding processed by the self-attention, while the K and V come from the light chain. This cross-attention layer itself is similar to the attention mechanism in the decoder of the Vaswani et al. 2017 Transformer method, but there are also some key differences. In particular, the cross-attention layer does not use any masking and there is no layer for generating tokens. In fact, contrary to the model in Vaswani et al., this model is not used for generating the next position prediction. In the case of using a decoder in this method, the masking (designed to ensure that the model does not "look ahead" to positions beyond the current position for the next token prediction) and the layer for generating tokens are useless in the current situation. After the cross-attention, the output goes through layer normalization and forms a residual connection with the non-attention-processed heavy-chain input again. This residual, normalized, cross-attention-processed output is used as the "heavy-chain" input for the next cross-attention block, and the process is repeated. The output is a B×max(L)×768 tensor.

[0088] The idea behind the cross-attention block is that the self-attention layer learns the pairwise importance between positions in the heavy chain, while the cross-attention layer learns the pairwise importance between each heavy-chain position relative to the light-chain positions.

[0089] The output of the cross-attention module is used as the input to the classifier module. In particular, after six cross-attention blocks, the output is processed by an attention pooling mechanism. Attention pooling is described in Safari et al. (2020). In essence, the idea is to compress a tensor of variable length B×max(L)×768 into a fixed shape B×768. Attention pooling can be conceptualized as a weighted average. Other pooling mechanisms can be used. For example, pooling can be done by taking the embedding of the start token. However, this is considered less favorable as it results in a large loss of information. Alternatively, an unweighted average can be used. This retains more information than just using the embedding of the start token, but does not allow for different weights to be applied to different positions as well as using a weighted average. More complex pooling mechanisms can also be used, although they are more computationally intensive to train. Therefore, choosing the appropriate pooling mechanism requires a balance between retaining information and the need for additional training parameters, which depends in part on the amount of available training data and the computational power available for training. The attention pooling tensor is then processed by a three-layer neural network that consists of: (1) a SiLU unit (also known as a Swish unit, described in Ramachandran et al., 2017) that reduces the dimensionality of the input from B×768 to B×256 and then drops out at a rate of 0.25 (note that other dropout rates can be used and are typically evaluated empirically; in this case, dropout rates between 0.1 and 0.3 were found to be appropriate); (2) a rectified linear unit (ReLU) that further reduces the dimensionality of the input activated by Swish from B×256 to B×64, followed by a dropout rate of 0.25; (3) the output processed by ReLU is then passed to a sigmoid layer that computes a score between 0 and 1. Note that instead of or in addition to the combination of Swish and ReLU, other choices of activation functions can be used, such as using only a Swish layer, only a ReLU layer, using a TanH layer instead of Swish and / or ReLU layers or in combination with Swish and / or ReLU layers, etc. These activation functions advantageously introduce non-linearity through trainable parameters, thus providing flexibility to best fit the data. Therefore, the choice of activation function can be guided at least to some extent by the expected behavior of the data and different possible choices are typically evaluated empirically to determine the choice that best fits the data. In fact, the score can be regarded as the probability of the heavy chain and light chain pairing. Specifically, given the nature of the negative dataset generated, the closer the value is to 0, the more the pair resembles a pair found in a randomly scrambled dataset; while the closer the value is to 1, the more the pair resembles a pair found in a true paired dataset from, for example, single-cell sequencing. As Figure 5B As shown, the training data includes true native pairs (training label '1') and random pairs (training label '0'), and the model is trained to predict the latter with a score as close to 1 as possible and the former with a score as close to 0 as possible. The training data can also include pairs associated with non-binary labels (such as continuous scores indicating "nativeness" or "functionality"). Such scores can, for example, reflect functional information related to known pairs, such as binding affinity, or any other metric related to binding strength.

[0090] The complete classifier model is trained using the following scheme. The training is carried out for a total of 10 epochs, with an allowed peak learning rate of 3e-5, a weight decay of 0.1, and a cosine learning rate schedule. For the first 5 epochs, the parameters of the encoder module that generates the embeddings are "frozen". This means that in the first five epochs, only the cross-attention block, attention pooling, and the final three-layer neural network can be optimized, while the encoder block remains unchanged. This is beneficial for achieving better generalization. For the rest of the training, the parameters of the encoder module are allowed to vary (although the two copies of the pre-trained encoder vary simultaneously, i.e., they remain exact copies of each other). This approach enables the training to initially focus on training the classification part of the model and only allows adjustments to the encoding part of the model at a lower learning rate later. The training of the classification task uses binary cross-entropy as the loss function. Model Structure - Alternative

[0091] An alternative method is designed that does not use the cross-attention module described above. In this method, the maximum length of the pre-trained encoder module is longer than that in the above, and the concatenation of the heavy-chain sequence and the light-chain sequence is used as the input. Such a model jointly embeds both the heavy chain and the light chain instead of generating separate embeddings. Then, it is used as the input to the classifier module described above, which includes an attention pooling layer, Swish, ReLU, and Softmax layers. The pre-training of the encoder module can use masked language modeling in a similar manner to the above. It should be noted that the encoder already includes an attention mechanism that can focus on the heavy chain and the light chain, so an additional cross-attention module can be dispensed with. The fine-tuning of the complete model (including the pre-trained encoder and the classifier module) can be carried out in a similar manner to the above, i.e., using a first stage where the parameters of the encoder module are frozen and the training focuses on the classifier module, followed by a second stage where the parameters of the encoder module are also allowed to vary. Results

[0092] Training the classifier model to predict whether two chains are likely to pair is limited by the amount of training data available because paired heavy-light chain data is required. The amount of such data available is relatively limited, and its content is further restricted by providing V and J gene identifiers instead of the full sequences (thus effectively restricting the prediction to the germline sequences of these segments). To overcome these limitations, a model containing an encoder was constructed that was trained on a larger dataset of unpaired heavy chain sequences and light chain sequences (42.3 million and 15.3 million sequences, respectively).

[0093] To evaluate the model, the inventors used datasets from King et al., Setliff et al., and Eccles et al. While the internal paired data test set (consisting of 223,912 pairs) could also be used for this purpose, the inventors chose to evaluate on the above three datasets because these datasets were outside the scope of their samples and provided a more stringent reflection of model generalization. The results of the internal test set were expected to be at least as good as the results of the independent data shown below.

[0094] To approximate the envisioned use case of pairing a large heavy chain sequencing dataset with a large light chain sequencing dataset, more "negative" data was generated in the same manner as the model was trained, and the randomization process was repeated multiple times for each test set. Thus, randomly paired sequences were mixed with truly paired sequences. Performance was measured using the area under the receiver operating characteristic curve (ROC) score. An ROC score of 0.5 indicates that the model generates random predictions, while an ROC score of 1.0 indicates that the model generates perfect predictions. Table 1 provides the ROC scores for different sets using the same randomization scheme as above. Figure 7 The ROC curve for the largest of these datasets (King) is shown in A. Figure 7 B shows the corresponding precision-recall (PR) curve. The PR curve shows the trade-off between precision (positive predictive value, the fraction of true pairs in the predicted pairs) and recall (sensitivity, the fraction of true pairs predicted as pairs) when the classification threshold varies between 0 and 1. The area under the PR curve (AUPR) provides an available indication of the performance of the model in finding true pairs in an imbalanced situation (i.e., where the expected number of unpaired candidates is much larger than the true paired candidates). The AUPR will be compared with the positive class score (50% in the illustrated embodiment). Thus, the data indicates that the model performs very well (significantly better than random) in distinguishing true pairs from random pairs. Additionally, the ROC curve can be used to determine the probability threshold that results in the desired trade-off between false positives and true positives. A favorable trade-off may depend on circumstances such as the importance of not losing true positive pairs vs. the ability to accommodate a larger number of candidate pairs in any subsequent test step (either on a computer or in vitro). For example, probability thresholds of about 0.7, 0.75, 0.8, 0.85, 0.9, or 0.95 may be favorable. In some embodiments, a threshold of about 0.8 is used. Dataset Dataset size AUROC Eccles 741 positive, 741 negative 0.802 King 30332 positive, 30332 negative 0.814 Setliff 4944 positive, 4944 negative 0.835 Table 1. ROC scores for the test dataset (50% random pairs).

[0095] The inventors also experimented with a larger proportion of negative pairs, and the results are shown in Table 2. Dataset Dataset size AUROC Eccles 741 positive, 1770 negative 0.805 King 30332 positive, 73759 negative 0.813 Setliff 4944 positive, 24720 negative 0.831 Table 2. ROC scores for the test dataset (70% random pairs).

[0096] The data indicates that even when attempting to identify true pairs among a larger number of random pairs, the prediction performance remains stable. This indicates that in real-life situations, when the number of random pairs in a set of candidate pairs is expected to be larger than the true pairs, the model is likely to perform well. The data further shows stable performance across 3 test datasets, indicating that the model exhibits good generalization ability at least within the same species. Thus, the probability of true pairs vs. random pairs provided by the model can be used to compare candidate pairs, e.g., by ranking the candidate pairs based on this probability or otherwise prioritizing them. Example 2 - FAbCon for pairing prediction

[0097] In this embodiment, an alternative architecture is used that employs a generative language model. This is a decoder-only model pre-trained for the next token prediction task (instead of the encoder-only model pre-trained for the masked language modeling task in Example 1).

[0098] FabCon is an antibody-specific large language model (LLM) based on the Falcon LLM from natural language processing (Penedo, G. et al. 2023). FAbCon is pre-trained using causal language modeling (CLM) on 779.4 million unpaired and paired antibody sequences. It is then fine-tuned for paired tasks. Method

[0099] Dataset. For pre-training the FAbCon model using the next token prediction task, a dataset of 823.7 million sequences (821.2 million unpaired, 2.5 million paired) was collected. This was split into 95% (777 million unpaired sequences and 2.4 million paired sequences) for pre-training the FAbCon model and the remaining 5% (44.3 million unpaired sequences and 0.1 million paired sequences) for evaluating the progress of pre-training.

[0100] More specifically, the Observed Antibody Space (OAS) database was downloaded on February 23, 2023. Samples were pre-filtered following a similar procedure as discussed previously (Bachas et al. 2022). For example, B cell receptors from pre-B cell samples were not used. However, all sequences were retained regardless of the count. A total of 1.47 billion unpaired sequences (i.e., only heavy or light chains) were used. Additionally, a set of proprietary B cell receptor sequences from 376 libraries was used to supplement the corpus, bringing the total number of sequences to 1.54 billion. Then, Linclust (Steinegger, M. & J) was used to cluster the dataset with 90% sequence identity in the VH or VL domain. After clustering, 777.8 million sequences were retained from the OAS and 43.4 million sequences were retained from the proprietary data. In addition to the unpaired data, paired heavy-chain and light-chain sequences were also used for pre-training. First, 1.5 million publicly available paired sequences (Jaffe et al. 2022) and an internal dataset of 1.4 million paired sequences were merged. Due to the scarcity of paired data, a 99% redundancy criterion was applied to the paired data. This resulted in a new total of 2.5 million sequences, of which 1.4 million were from the dataset of Jaffe et al. and 1.1 million were from the internal data. The final dataset contains 823.7 million sequences (821.2 million unpaired sequences, 2.5 million paired sequences), which was split into 95% (777 million unpaired sequences and 2.4 million paired sequences) and 5% (44.3 million unpaired sequences and 0.1 million paired sequences) for training and evaluating the CLM, respectively.

[0101] 1,791,293 sequences were used for fine-tuning: This included 895,765 sequences from single B-cell sequencing, which were considered 'true' pairs. Another 895,528 sequences were generated by random matching of heavy and light chains in single B-cell sequencing data and were considered 'false' pairs. After training, the fine-tuned model was evaluated on three independent external single B-cell sequencing datasets.

[0102] Tokenization. Sequences were tokenized at the amino acid level, where each amino acid is a token. The antibody heavy chain starts with token, while the light chain is prefixed with token. FAbCon can accept both unpaired and paired chains as input. In the case of paired chains, the input is actually a longer single amino acid sequence containing both token and token (e.g., ). FAbCon accepts 26 tokens: 20 amino acids, <|endoftext|>, <|unk|>, <|pad|>, and <|mask|>.

[0103] Pretraining. FAbCon is a generative transformer model based on the Falcon architecture (Penedo et al., 2023), which is available at falconllm.tii.ae / falcon - models.html and https: / / huggingface.co / docs / transformers / main / mode_doc / falcon. This is a decoder - only autoregressive transformer model with a GPT - 3 - based architecture (Brown et al., 2020), ALiBi positional encoding (Press et al., 2021), and FlashAttention (Dao et al., 2022). The Falcon model uses multi - query attention, which shares key and value embeddings among attention heads, thus reducing the memory cost of both pre - training and inference. Three FAbCon variants have been pre - trained for the next token prediction task (Figure 9A): a variant with 144 million parameters, a variant with 300 million parameters, and a variant with 2.4 billion parameters. All FAbCon variants were pre - trained using 48 NVIDIA A100 (80GB) from NVIDIA's DGX SuperCloud (Cambridge - 1 environment). The configuration of each variant is listed in Table 3. For next token prediction, the task of the model is to predict the amino acid sequence in a left - to - right manner. Only the previous residues at the N - terminus are used to provide information for the prediction of subsequent amino acids. The maximum context window of each FAbCon model is 256 (i.e., the maximum sequence length is 256). During pre - training, both unpaired antibody sequences and paired antibody sequences can form a mini - batch.

[0104] The FAbCon variants were pre - trained using the CLM objective, where the task of the model is to autoregressively predict each residue by only looking at the previous residues in the sequence. More formally, the task of the model is to predict the probability r of the amino acid residue at position t t : P(r t |r1, r2, …, r t-1 ).

[0105] During pre - training, the model parameters are updated via gradient - based optimization to minimize the cross - entropy loss given as follows: where N is the batch size, T is the sequence length, V is the vocabulary size, and L t,i is the label of the residue at position t, and the i-th word in the vocabulary (here including 20 amino acids, <|endoftext|>, <|unk|>, <|pad|>, <|mask|>, and (26 words for the prefix).

[0106] As described above, all three FAbCon variants are characterized by flash attention and multi-query attention, thus improving the pre-training speed and reducing memory footprint. For each variant, a "deep and narrow" structure was adopted, where the number of layers was increased and the embedding dimension was kept relatively low. The optimization used the Fused AdamW optimizer with a peak learning rate of 2e-5, a cosine learning rate schedule, 200k steps, a weight decay of 0.01, and a gradient norm clipping of 1.0. Table 3. Configuration of the FAbCon model

[0107] Fine-tuning. Figures 9B and 9D show the different structures of the FAbCon for pre-training (Figure 9B) and the classifier model for distinguishing whether the input sequence is a true pair or a false pair (Figure 9D). The difference between FAbCon-small (generative language model) and FAbCon-small-pairing (pairing classifier model) lies in the final "head" layer. FAbCon-small-pairing uses the representation from the decoder layer to output the probability that the input sequence forms a pair. This structure can be applied in a similar way to FAbCon-medium and FAbCon-large; for practical reasons, only the results of fine-tuning the FAbCon-small model are shown.

[0108] The FAbCon-small model was fine-tuned for 10 epochs with a peak learning rate of 5e-5 after 5% warm-up. The dropout rate was set to 0.1, and the weight decay was set to 0.01. The checkpoint of the model with the lowest validation set loss, corresponding to the 5th epoch, was used. Model freezing was not implemented here. Figure 9C provides an overall view of how the model was pre-trained and then fine-tuned for pairing purposes. Results

[0109] The results of the above process are shown in Tables 4 and 5 below. These results show very good prediction accuracy, although slightly lower than that of Example 1 (but similar performance is expected to be achieved through further optimization). Dataset Dataset size AUROC Eccles 741 positive, 741 negative 0.790 King 30332 positive, 30332 negative 0.803 Setliff 4944 positive, 4944 negative 0.828 Table 4. ROC scores for the test dataset (50% random pairs). Dataset Dataset size AUROC Eccles 741 positive, 1770 negative 0.797 King 30332 positive, 73759 negative 0.802 Setliff 4944 positive, 24720 negative 0.825 Table 4. ROC scores for the test dataset (70% random pairs).

[0110] The same model was also fine-tuned for antibody-antigen binding prediction (classifying sequence pairs as binders or non-binders to a specific antigen; data not shown), and the results showed extremely strong performance (average precision between 0.851 for the three different antigens tested, and improved to 0.851 and 0.883 for the medium and large models). This indicates that the model can well distinguish binders from non-binders, further demonstrating that the encoding module of the model described herein can learn information to distinguish functional pairs from non-functional pairs. Example 3 - Discussion

[0111] This example describes a machine learning NLP-inspired method to solve the BCR heavy-light chain pairing problem. The deep learning-based method described herein provides the benefit of covering the BCR repertoire as deeply as possible, while eliminating the need for extensive light chain sequencing or single cell sequencing. In addition, the method is able to learn general features of pairing from the training dataset and use this learning to predict pairings of previously unseen chains. This can be advantageous in many situations considering the extreme diversity of the BCR repertoire, but especially so in application scenarios such as identifying specific antibodies or other rare antibodies that may underlie a desired phenotype in an individual.

[0112] The method fully utilizes the available training data through the following combination: pre-training the language model using a larger unpaired full-length sequence training dataset, and fine-tuning the final model using a smaller paired dataset that may have a smaller sequence coverage. The final model includes an encoding module (this is the part of the language model that provides a latent representation of the input sequence, such as an encoder or decoder, depending on the model structure) and a classifier module.

[0113] In addition, the method also provides a way to evaluate candidate chains as "true" chains, thus increasing the variation of functional pairs being identified. For example, the method is ideally suited for situations where the libraries of two pairs (or at least a subset of the library) have been identified, but the pairing information between the two libraries is not available. In fact, in such a situation, the method can utilize the knowledge that the natural partner of a candidate chain is likely to be present in the group of candidate chains. In contrast, "de novo" methods for generating chains for pairing may generate predictions that cannot properly fold such chains or cannot form functional pairs with any chain. It should be noted that the method for generating de novo predictions can be used in combination with the method of the present invention, for example, using the former to provide candidates for evaluation by the latter.

[0114] For applying a large number of heavy chain repertoire analyses to antibody discovery, the light chain pairing problem remains relevant. The deep learning-based method described herein enables the identification of light chains that are likely to pair with any given heavy chain with high computational accuracy by evaluating the pairing of known or otherwise generated candidate light chains. Thus, the method has the potential to fill the gap in light chain pairing information, enabling the discovery of therapeutic antibodies and a better understanding of the immune system.

[0115] Finally, while the method is described in the context of identifying light chain pairings for heavy chain queries, it is also applicable to the reverse problem of identifying heavy chain pairings for light chain queries, as well as the pairing of other immune molecule dimers. This is a less common problem as antibodies are widely used therapeutic, diagnostic, and research tools, and in the specific case of antibodies, heavy chain sequencing is more common and the heavy chain is considered to play a more important role in determining specificity and affinity. References Vander Heiden et al.,2017.Dysregulation of B Cell RepertoireFormation in Myasthenia Gravis Patients Revealed through Deep Sequencing.JImmunol.2017 Feb 15;198(4):1460-1473. Bashford-Rogers et al,2019.Analysis of the B cell receptor repertoirein six immune-mediated diseases.Nature volume 574,pages122-126(2019). Nielsen et al.,2020.Human B Cell Clonal Expansion and ConvergentAntibody Responses to SARS-CoV-2.bioRxiv.Preprint.2020 Jul 9.doi:10.1101 / 2020.07.08.194456. Simonich et al., 2019. Kappa chain maturation helps drive rapid development of an infant HIV-1 broadly neutralizing antibody lineage. Nature Communications volume 10, Article number: 2190 (2019). Krawczyk et al., 2019. Looking for therapeutic antibodies in next-generation sequencing repositories, mAbs. Volume 11, 2019-Issue 7, Pages 1197-1205. Galson et al., 2020. Deep Sequencing of B Cell Receptor Repertoires From COVID-19 Patients Reveals Strong Convergent Immune Signatures. Front. Immunol., 15 December 2020. doi.org / 10.3389 / fimmu.2020.605170. Mora and Walczak, 2019. How many different clonotypes do immune repertoires contain?Current Opinion in Systems Biology. Volume 18, December 2019, Pages 104-110 Kovaltsuk et al., 2018. Observed Antibody Space: A Resource for Data Mining Next-Generation Sequencing of Antibody Repertoires. J Immunol October 15, 2018, 201(8)2502-2509. Teplyakov et al.,2016.Structural diversity in a human antibodygermline library.MAbs.Aug-Sep 2016;8(6):1045-63. Glanville et al.,2009.Precise determination of the diversity of acombinatorial antibody library gives insight into the human immunoglobulinrepertoire.PNAS December 1,2009 106(48)20216-20221. Jayaram et al.,2012.Germline VHNL pairing in antibodies.ProteinEngineering,Design and Selection,Volume 25,Issue 10,October 2012,Pages 523-530. Ling et al.,2018.Effect of VH-VL Families in Pertuzumab andTrastuzumab Recombinant Production,Her2 and FcγllA Binding.Front.Immunol.,12March 2018.doi.org / 10.3389 / fimmu.2018.00469 DeKosky et al.,2016.Large-scale sequence and structural comparisonsof human naive and antigen-experienced antibody repertoires.PNAS May 10,2016113(19)E2636-E2645. King et al., 2021. Single-cell analysis of human B cell maturation predicts how antibody class switching shapes selection dynamics. Science Immunology 12 Feb 2021. Vol. 6, Issue 56, eabe6291 Eccles et al., 2020. T-bet+ Memory B Cells Link to Local Cross-Reactive IgG upon Human Rhinovirus Infection. Cell Reports Volume 30, Issue 2, 14 January 2020, Pages 351 - 366.e7 Setliff et al., 2019. High-Throughput Mapping of B Cell Receptor Sequences to Antigen Specificity. Cell Volume 179, Issue 7, 12 December 2019, Pages 1636 - 1646.e15 Reddy et al., 2010. Monoclonal antibodies isolated without screening by analyzing the variable-gene repertoire of plasma cells. Nature Biotechnology volume 28, pages 965 - 969(2010). Zhu et al., 2013. Mining the antibodyome for HIV-1-neutralizing antibodies with next-generation sequencing and phylogenetic pairing of heavy / light chains. PNAS. 2013 Apr 16;110(16):6470 - 5. Raybould et al.,2021.Public Baseline and shared response structuressupport the theory of antibody repertoire functional commonality.PLoS ComputBiol 17(3):e1008781. Rakocevic et al.,2021.The landscape of high-affinity human antibodiesagainst intratumoral antigens,bioRxiv.8 Feb 2021.doi.org / 10.1101 / 2021.02.06.430058 Vaswani et al.,2017.Attention Is All You Need.arXiv:1706.03762 Devlin,Jacob,et al.″Bert:Pre-training of deep bidirectionaltransformers for language understanding."arXiv preprint arXiv:1810.04805(2018). Radford et al.,2019.Language Models are Unsupervised MultitaskLearners.https: / / openai.com / blog / better-language-models / Liu et al.,2019.RoBERTa:A Robustly Optimized BERT PretrainingApproach.arXiv:1907.11692 Rothe et al.,2020.Leveraging Pre-trained Checkpoints for SequenceGeneration Tasks.arXiv:1907.12461 Dunbar and Deane, 2016. ANARCI: antigen receptor numbering and receptor classification. Bioinformatics. 2016 Jan 15;32(2):298 - 300. Rees, 2020. Understanding the human antibody repertoire. MAbs. Jan - Dec 2020;12(1):1729683. Ye et al., 2013. IgBLAST: an immunoglobulin variable domain sequence analysis tool. Nucleic Acids Res. 2013 Jul;41(Web Server issue):W34 - 40. Carter JA et al. Single T Cell Sequencing Demonstrates the Functional Role of αβ TCR Pairing in Cell Lineage and Antigen Specificity. Frontiers in Immunology. Vol. 10. 2019, p. 1516. Zheng GXY, Terry JM, Belgrader P, Ryvkin P, Bent ZW, Wilson R, et al. Massively parallel digital transcriptional profiling of single cells. Nat Commun. (2017) 8:14049. Howie B, Sherwood AM, Berkebile AD, Berka J, Emerson RO, Williamson DW, et al. High - throughput pairing of T cell receptor α and β sequences. Sci Transl Med. (2015) 7:301ra131. Eve Richardson,Jacob D.Galson,Paul Kellam,Dominic F.Kelly,SarahE.Smith,Anne Palser,Simon Watson&Charlotte M.Deane(2021)A computationalmethod for immune repertoire mining that identifies novel binders fromdifferent clonotypes,demonstrated by identifying anti-pertussis toxoidantibodies,mAbs,13:1. Yi-Chun Hsiao,Yonglei Shang,Danielle M.DiCara,Angie Yee,Joyce Lai,SiHyun Kim,Diego Ellerman,Racquel Corpuz,Yongmei Chen,Sharmila Rajan,Hao Cai,Yan Wu,Dhaya Seshasayee&Isidro (2019)Immune repertoire mining for rapidaffinity optimization of mouse monoclonal antibodies,mAbs,11:4,735-746. Warszawski S,Borenstein Katz A,Lipsh R,Khmelnitsky L,Ben Nissan G,Javitt G,et al.(2019)Optimizing antibody affinity and stability by theautomated design of the variable light-heavy chain interfaces.PLoS ComputBiol 15(8):e1007207. Seeliger D, Schulz P, Litzenburger T, Spitz J, Hoerer S, Blech M, Enenkel B, Studts JM, Garidel P, Karow AR. Boosting antibody developability through rational sequence optimization. MAbs. 2015;7(3):505-15. doi:10.1080 / 19420862.2015.1017695. Mason, D.M., Friedensohn, S., Weber, C.R. et al. Optimization of therapeutic antibodies by predicting antigen specificity from antibody sequence via deep learning. Nat Biomed Eng (2021). Leem, J., Mitchell, L.S., Farmery, James H.R., Barton, J., Galson, J.D. Deciphering the language of antibodies using self-supervised learning. Patterns 3, 100513. July 8, 2022. Safari, Pooyan, Miquel India, and Javier Hernando. "Self-attention encoding and pooling for speaker recognition." arXiv preprint arXiv:2008.01077 (2020). Ramachandran, Prajit, Barret Zoph, and Quoc V. Le. ″Searching for activation functions." arXiv preprint arXiv:1710.05941 (2017). Shaw, Peter, Jakob Uszkoreit, and Ashish Vaswani. “Self-attention with relative position representations.” arXiv preprint arXiv:1S03.02155 (2018). Su, Jianlin, et al. “Roformer: Enhanced transformer with rotary position embedding.” arXiv preprint arXiv:2104.09864 (2021). Nijkamp, Erik, et al. “ProGen2: exploring the boundaries of protein language models.” arXiv preprint arXiv:2206.13517 (2022). Cédric R Weber, Rahmad Akbar, Alexander Yermanos, Milena Igor Snapkov, Geir K Sandve, Sai T Reddy, Victor Greiff, immuneSIM: tunable multi-feature simulation of B-and T-cell receptor repertoires for immunoinformatics benchmarking, Bioinformatics, Volume 36, Issue 11, June 2020, Pages 3594 - 3596. Penedo, G. et al. The RefinedWeb Dataset for Falcon LLM: Outperforming Curated Corpora with Web Data, and Web Data Only. arXiv (2023) doi:10.48550 / arxiv.2306.01116. Steinegger, M. & J. Clustering huge protein sequence sets in linear time. Nat. Commun. 9, 2542 (2018). Jaffe, D.B. et al. Functional antibodies exhibit light chain coherence. Nature 611, 352 - 357 (2022). Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J.D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al. Language models are few-shot learners. Advances in neural information processing systems, 33: 1877 - 1901, 2020. Dao, T., Fu, D.Y., Ermon, S., Rudra, A., and Re, C. Flash attention: Fast and memory-efficient exact attention with io-awareness. In Advances in Neural Information Processing Systems, 2022. Press, O., Smith, N., and Lewis, M. Train short, test long: Attention with linear biases enables input length extrapolation. In International Conference on Learning Representations, 2021. Touvron H. et al. 2023. Llama 2:Open Foundation and Fine-Tuned Chat Models arXiv:2307.09288v2. 19 July 2023 Mandar Joshi, Danqi Chen, Yinhan Liu, Daniel S. Weld, Luke Zettlemoyer, Omer Levy. SpanBERT: Improving Pre-training by Representing and Predicting Spans. arXiv:1907.10529v3. 18 Jan 2020

[0116] All references cited herein are incorporated herein by reference in their entirety for all purposes to the same extent as if each individual publication or patent or patent application was specifically and individually indicated to be incorporated herein by reference in its entirety.

[0117] The specific embodiments described herein are provided by way of example and not limitation. Various modifications and variations of the uses, compositions, and methods of the described technology will be apparent to those skilled in the art without departing from the scope and spirit of the technology. Any subheadings included herein are for convenience only and should not be construed as limiting the disclosure in any way. Unless the context otherwise indicates, other aspects and embodiments of the invention provide the above aspects and embodiments with the term "comprising" replaced by the terms "consisting of" or "consisting essentially of".

[0118] The methods of any of the embodiments described herein can be provided as a computer program or as a computer program product or a computer-readable medium carrying a computer program, the computer program being arranged to perform the above methods when run on a computer.

[0119] Unless the context otherwise indicates, the descriptions and limitations of the above features are not limited to any particular aspect or embodiment of the invention and equally apply to all aspects and embodiments described.

[0120] Throughout the specification and claims, unless the context otherwise clearly indicates, the following terms have the meanings specifically associated herein. As used herein, the phrase "in one embodiment" does not necessarily refer to the same embodiment, but it may. Additionally, as used herein, the phrase "in another embodiment" does not necessarily refer to a different embodiment, but it may refer to a different embodiment. Thus, as described below, various embodiments of the invention can be readily combined without departing from the scope or spirit of the invention.

[0121] It must be noted that, as used in the specification and the appended claims, a noun without a quantifier means one or more, unless the context clearly indicates otherwise. Ranges can be expressed herein as from "about" a particular value and / or to "about" another particular value. When expressing such a range, another embodiment includes from a particular value and / or to another particular value. Similarly, when a value is expressed as an approximation by use of the foregoing "about", it will be understood that the particular value forms another embodiment. The term "about" associated with a numerical value is optional and means, for example, + / - 10%.

[0122] Throughout this specification (including the appended claims), unless the context requires otherwise, the words "comprise" and "include" and variations thereof will be understood to imply the inclusion of the stated integer or step or group of integers or steps but not the exclusion of any other integer or step or group of integers or steps. As used herein, "and / or" is considered to specifically disclose each of the two indicated features or components with or without the other. For example, "A and / or B" shall be considered to specifically disclose (i) A, (ii) B, and (iii) each of A and B, as if each were individually recited herein.< / pad> < / unk> < / pad> < / mask> < / unk> < / pad>

Claims

1. A computer-implemented method for determining whether a pair of protein chains comprising a first chain and a second chain is likely to form a functional antigen-binding protein, the method comprising: Providing a query sequence pair comprising the sequence of the first protein chain and the sequence of the second protein chain as input to a deep learning model, the deep learning model being configured to take a pair of protein chain sequences as input and produce a score indicative of the probability that the pair of protein chains forms a functional antigen-binding protein as output, wherein the deep learning model comprises an encoder module and a classifier module, and wherein the deep learning module has been trained using paired training sequences from known antigen-binding proteins.

2. A computer-implemented method for identifying an antigen-binding protein comprising a first protein chain and a second protein chain, the method comprising: Providing a query first protein chain, and Identifying the second protein chain by: Providing one or more candidate second protein chain sequences; and And Determining whether the one or more candidate second chain sequences are likely to form a functional antigen-binding protein with the query first protein chain by: providing one or more query sequence pairs as input to a deep learning model, each query sequence pair comprising (i) the sequence of the first protein chain and (ii) a candidate second protein chain sequence, wherein the deep learning model is configured to take one or more pairs of protein chain sequences as input and produce a score indicative of the probability that each pair of protein chains forms a functional antigen-binding protein or information derived therefrom as output, wherein the deep learning model comprises an encoder module and a classifier module, and wherein the deep learning module has been trained using paired training sequences from known antigen-binding proteins, optionally wherein identifying the second protein chain comprises providing a plurality of candidate second chain sequences and the information derived from the score comprises a ranking of the pairs of protein chain sequences, wherein pairs of protein sequences that are more likely to form a functional antigen-binding protein are ranked higher than pairs of protein sequences that are less likely to form a functional antigen-binding protein.

3. The method according to claim 1 or claim 2, wherein the antigen-binding protein comprises: (i) a heavy chain-light chain pair, wherein the first chain is a heavy chain or a light chain, and the second chain is a light chain or a heavy chain, optionally wherein the first chain is a heavy chain and the second chain is a light chain; or (ii) an αβ chain pair, wherein the first chain is a β chain or an α chain, and the second chain is an α chain or a β chain, optionally wherein the first chain is a β chain and the second chain is an α chain; or (ii) a γδ chain pair, wherein the first chain is a δ chain or a γ chain, and the second chain is a γ chain or a δ chain, optionally wherein the first chain is a δ chain and the second chain is a γ chain.

4. The method according to any of the preceding claims, wherein the encoder module comprises one or two such encoders, or a decoder, the encoder having been pre-trained using training sequences from unpaired protein chains from known antigen-binding proteins, the decoder having been pre-trained using training sequences comprising paired and unpaired protein chains from known antigen-binding proteins.

5. The method according to any one of the preceding claims, wherein the encoder module comprises: one or two encoders of a sequence-to-sequence model, or a decoder of a generative language model, optionally wherein the sequence-to-sequence model or the generative language model is a Transformer-based model; or one or two Transformer-based encoder or decoder models.

6. The method according to any one of the preceding claims, wherein the encoder module comprises one or two encoder models of a sequence-to-sequence model, wherein the encoder model has been trained using masked language modeling, optionally wherein training the encoder model comprises training the model to replace randomly masked positions in a training amino acid sequence and / or wherein 15% of the positions are masked during training and / or wherein the masked positions are replaced with a mask token, a random amino acid, or the original amino acid at that position, or wherein the encoder module comprises a decoder model of a generative language model, wherein the decoder model has been trained using causal language modeling.

7. The method according to any one of the preceding claims, wherein the encoder module comprises one or two such encoders that have been pre-trained using a training sequence comprising unpaired first and second protein chains from a known antigen-binding protein, optionally wherein the training sequence comprises at least 1 million, at least 2 million, at least 5 million, at least 10 million, at least 20 million, or at least 50 million independent sequences; or wherein the encoder module comprises such a decoder that has been pre-trained using a training sequence comprising unpaired first and second protein chains from a known antigen-binding protein and paired first and second protein chains from a known antigen-binding protein, optionally wherein the training sequence comprises at least 500 million, 600 million, or 700 million independent sequences and / or at least 1 million, 1.5 million, or 2 million paired sequences.

8. The method according to any one of the preceding claims, wherein the encoder module comprises two copies of such an encoder that have been pre-trained using a training sequence comprising unpaired first and second protein chains from a known antigen-binding protein.

9. The method according to any one of claims 1 to 7, wherein the encoder module comprises such an encoder that has been pre-trained using a training sequence comprising a concatenation of a sequence from a first chain from a known antigen-binding protein and a sequence from a second protein chain from a known antigen-binding protein, wherein the concatenated sequences comprise sequences from unpaired protein chains; or wherein the encoder module comprises such a decoder that has been pre-trained using a training sequence comprising a single chain from a known antigen-binding protein or a concatenation of a sequence from a first chain from a known antigen-binding protein and a sequence from a second protein chain from a known antigen-binding protein, respectively.

10. The method according to any one of claims 1 to 8, wherein the deep learning model further comprises a cross-attention module, the cross-attention module taking the output of the encoder module as input and producing an output which is used by the classifier module to provide a core indicating the probability that the protein chain pair forms a functional antigen-binding protein.

11. The method according to claim 10, wherein the cross-attention module comprises one or more cross-attention blocks, each cross-attention block comprising a self-attention layer and a cross-attention layer, and / or wherein the cross-attention module comprises a plurality of cross-attention blocks.

12. The method according to claim 11, wherein the cross-attention module comprises one or more cross-attention blocks, the cross-attention blocks comprising a cross-attention layer and a self-attention layer, and each cross-attention layer and self-attention layer comprising a plurality of attention heads.

13. The method according to any one of claims 10 to 12, wherein each self-attention layer comprises one or more attention heads, the attention heads attending to the output of the first encoder taking the first chain sequence as input, and / or wherein each cross-attention block further comprises a residual connection between the output of the first encoder and the output of the self-attention layer.

14. The method according to any one of claims 10 to 13, wherein each cross-attention layer comprises one or more attention heads, the attention heads attending to: (i) the output of the self-attention layer or the output of the first encoder taking the first chain sequence as input, and (ii) the output of the second encoder taking the second chain sequence as input, and / or wherein each cross-attention block further comprises a residual connection between the output of the first encoder and the output of the cross-attention layer.

15. The method according to any one of the preceding claims, wherein the classification module comprises a softmax layer, the softmax layer producing a score from 0 to 1, the score being interpretable as the probability that the input protein chain pair forms a functional antigen-binding protein; and / or wherein the classification module comprises one or more of the following: a dimensionality reduction layer, a regularization mechanism, and a layer with an activation function, optionally wherein the dimensionality reduction layer comprises an attention pooling layer and / or wherein the regularization layer comprises a dropout mechanism, and / or wherein the activation function is independently selected from Tanh, Leaky ReLU, SmeLU, GeLU, Swish, and ReLU.

16. The method according to any one of the preceding claims, wherein the deep learning model takes a plurality of query chain sequence pairs as input and produces corresponding scores indicating the probability that each query protein chain pair forms a functional antigen-binding protein and / or a ranking of the plurality of query chain sequence pairs as output, the ranking being such that query protein sequence pairs more likely to form a functional antigen-binding protein are ranked higher than query protein sequence pairs less likely to form a functional antigen-binding protein.

17. The method according to any one of the preceding claims, wherein the paired training sequences from a known antigen-binding protein comprise paired training heavy-chain sequences and light-chain sequences from single B cell sequencing data, and / or wherein the training data further comprises a negative set containing randomly paired first protein chain sequences and second protein chain sequences, optionally wherein the randomly paired first protein chain sequences and second protein chain sequences are obtained by re-pairing the paired training sequences from a known antigen-binding protein and / or by randomly pairing unpaired training sequences, and / or wherein the negative set comprises paired first protein chain sequences and second protein chain sequences from antigen-binding proteins previously determined to be non-functional, and / or wherein the paired training sequences are associated with a first label and the randomly paired training sequences are associated with a second label.

18. The method according to any one of the preceding claims, wherein providing a protein chain sequence pair as input to the deep learning model comprises: Encoding each of the protein chain sequences using a predetermined encoding scheme, optionally wherein each amino acid is encoded individually, or wherein the sequences are encoded using markers each corresponding to an individual k-mer, and / or wherein there is a special marker indicating whether the protein chain sequence is a first protein chain or a second protein chain before or after each protein chain sequence.

19. The method according to any one of the preceding claims, wherein providing a query first protein chain or the query first protein chain in a query pair comprises: obtaining the sequence of the query first protein chain as follows: obtaining from a user via a user interface, obtaining from a computing device, obtaining from a sequence acquisition device or a computing device associated with a sequence acquisition device, obtaining from a database or other computer-readable medium; and / or sequencing a sample containing genetic material encoding an antigen-binding molecule containing the query sequence, optionally wherein obtaining the query sequence comprises bulk sequencing of B cells in a sample containing B cells, bulk sequencing of T cells in a sample containing T cells, or bulk sequencing of a sample containing any other cell expressing an antigen-binding molecule containing the query sequence or genetic material derived therefrom, such as a B cell receptor library or a T cell receptor library; and / or obtaining a sample containing B cells, T cells or other cells expressing an antigen-binding molecule containing the query sequence, or genetic material derived therefrom, such as a B cell receptor library or a T cell receptor library; and / or wherein providing one or more candidate second protein chain sequences comprises obtaining the sequence of the candidate second protein chain as follows: obtaining from a user via a user interface, obtaining from a computing device, obtaining from a sequence acquisition device or a computing device associated with a sequence acquisition device, obtaining from a database or other computer-readable medium; and / or wherein the one or more candidate second protein chain sequences are known second chain protein sequences or simulated second chain protein sequences.

20. The method according to any one of the preceding claims, the method further comprising providing to a user, via a user interface, information about one or more identified second protein chains, a portion thereof, or derived therefrom, and / or one or more scores indicating the probability of one or more query pairs forming a functional antigen-binding protein, or information derived therefrom; and / or the method further comprising predicting a score indicating the probability of each one or more query pairs comprising a corresponding candidate second protein chain sequence forming a functional antigen-binding protein, and identifying candidate second protein chain sequences by applying one or more criteria to the scores, optionally wherein the one or more criteria are selected from: the score is higher than a predetermined cut-off value, the score is the highest predicted score of a group of candidate second protein chain sequences, and the score is in a predetermined top percentile of the predicted probabilities of a group of candidate second protein chain sequences; and / or wherein the core is the probability of the protein chain pair forming a functional antigen-binding protein.

21. A method for providing antigen-binding protein chain pairing for a plurality of query sequences comprising a first chain sequence, the method comprising: Performing the method according to any one of claims 2 to 20 on each of the query sequences, optionally wherein the plurality of query sequences are heavy or light chain sequences obtained by bulk B cell repertoire sequencing.

22. A method of providing an antigen-binding protein having desired properties, the method comprising: providing one or more query sequences comprising a first chain sequence, wherein at least one of the one or more query sequences may have the desired properties, and using the method according to any one of claims 2 to 20 to identify a second chain sequence for each of the one or more query sequences.

23. A method of providing a tool for predicting whether a protein chain pair is likely to form a functional antigen-binding protein, the method comprising: providing training data comprising a training first protein chain sequence and a training second protein chain sequence from a known antigen-binding protein; and training a deep learning model to take a protein chain sequence pair as input and produce as output a score indicating the probability of the protein chain pair forming a functional antigen-binding protein, wherein the deep learning model comprises an encoder module and a classifier module.

24. A system, the system comprising: a processor; and a computer-readable medium comprising instructions that, when executed by the processor, cause the processor to perform the steps of the method according to any one of claims 1 to 23. One or more computer-readable media comprising instructions that, when executed by one or more processors, cause the one or more processors to perform the steps of the method according to any one of claims 1 to 23.