Methods for identifying cognate immunoglobulin pairs from individual b cells
A high-throughput method using barcoded primers and multiwell substrates addresses the challenge of identifying VH/VL pairs in B cells, enabling rapid and cost-effective antibody development.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-09-04
- Publication Date
- 2026-03-12
AI Technical Summary
Identifying cognate pairs of variable heavy (VH) and variable light (VL) domains from B cells in the immune repertoire is time-consuming and expensive due to the vast number of unique antibodies produced by B cells, necessitating methods like single-cell sequencing.
A cost-effective, high-throughput method involving depositing B cells into multiwell substrates, using barcoded primers to amplify and sequence polynucleotides, calculating pairing scores based on intra-well and inter-well abundance, and identifying cognate pairs of VH and VL domains.
Enables rapid and low-cost identification of novel antibodies by efficiently pairing VH and VL domains, facilitating their development as potential therapeutics.
Smart Images

Figure US2025044873_12032026_PF_FP_ABST
Abstract
Description
METHODS FOR IDENTIFYING COGNATE IMMUNOGLOBULIN PAIRS FROM INDIVIDUAL B CELLS RELATED APPLICATIONS
[0001] This application claims the benefit under 35 U.S.C. § 119(e) of U.S. Provisional Application No.63 / 690,413, filed on September 4, 2024, the entire contents of which is hereby incorporated by reference in their entirety. REFERENCE TO AN ELECTRONIC SEQUENCE LISTING
[0002] The contents of the electronic sequence listing (A136170025WO00-SEQ-EAS.xml; Size: 170,859 bytes; and Date of Creation: September 3, 2025) are herein incorporated by reference in its entirety. BACKGROUND
[0003] High-throughput sequencing of an immune repertoire has emerged as a critical step in understanding adaptive responses following infection or vaccination or in autoimmunity. Yet, identifying cognate pairs of immune receptors or native antibody variable heavy chains and variable light chains within this immune repertoire remains a major challenge, in part because methods for studying the million or so immune cells that represent the diversity of the immune repertoire are limited. SUMMARY
[0004] Antibody (also called immunoglobulin) production by B cells is a critical part of the humoral immune response. It is estimated that adult humans have between 109and 1015distinct B cell clones, each clone producing antibodies with a unique variable heavy (VH) chain and variable light (VL) chain cognate pair (also referred to as VH / VL cognate pairs). This number fluctuates upon antigen recognition by somatic hypermutation, a process by which B cells randomly induce mutations in immunoglobulin coding regions to ultimately enhance the affinity of antibodies they produce for their cognate antigen. Immunization with an antigen of interest is an efficient way to generate a diversity of antibodies with therapeutic potential. Due to the vast number of unique antibodies produced by B cells, identifying individual VH / VL1#14358394v1cognate pairs requires time-consuming and expensive techniques, such as single-cell sequencing, to ensure that the identified VH and VL sequences are derived from the same B cell clone. Provided herein are cost-effective, high-throughput methods for identifying VH / VL cognate pairs from individual B cell clones in a heterogeneous B cell population. Such methods allow for rapid, low-cost identification of novel antibodies for further development as potential therapeutics.
[0005] Accordingly, in some aspects, the disclosure provides a method of identifying cognate pairs of VH domains and VL domains from individual B cells comprising (a) depositing B cells into wells of a multiwell substrate at a concentration of about 2-100 B cells per well, wherein each of the B cells comprises (i) a polynucleotide encoding a polypeptide comprising a VH domain and (ii) a polynucleotide encoding a polypeptide comprising a VL domain; (b) amplifying the polynucleotide of (i) and the polynucleotide of (ii) using barcoded primers, each primer coded to a respective well of the multiwell substrate, to produce subsets of amplified barcoded polynucleotides, each subset coded to a respective well of the multiwell substrate; (c) sequencing the amplified barcoded polynucleotides to produce barcoded records; (d) calculating a pairing score for respective barcoded records, wherein calculating the pairing score comprises calculating an intra-well abundance score and an inter-well abundance score; and (e) identifying cognate pairs of VH domains and VL domains based on the pairing score.
[0006] In some aspects, the disclosure provides a method of identifying cognate pairs of heavy chain variable (VH) domains and light chain variable (VL) domains from individual B cells comprising (a) sequencing amplified barcoded polynucleotides to produce barcoded records, wherein the amplified barcoded polynucleotides are produced by depositing B cells into wells of a multiwell substrate at a concentration of about 2-100 cells per well, wherein each of the B cells comprises (i) a polynucleotide encoding a polypeptide comprising a VH domain and (ii) a polynucleotide encoding a polypeptide comprising a VL domain, and amplifying the polynucleotide of (i) and the polynucleotide of (ii) using barcoded primers, each primer coded to a respective well of the multiwell substrate, to produce subsets of amplified barcoded polynucleotides, each subset coded to a respective well of the multiwell substrate; (b) calculating a pairing score for respective barcoded records, wherein calculating the pairing score comprises calculating an intra-well abundance score and an inter-well abundance score; and (c) identifying cognate pairs of VH domains and VL domains based on the pairing score.
[0007] In some aspects, the disclosure provides a method of identifying cognate pairs of heavy chain variable (VH) domains and light chain variable (VL) domains from individual B2#14358394v1cells comprising (a) calculating a pairing score for respective barcoded records, wherein calculating the pairing score comprises calculating an intra-well abundance score and an inter- well abundance score, and wherein the barcoded records are produced by depositing B cells into wells of a multiwell substrate at a concentration of about 2-100 cells per well, wherein each of the B cells comprises (i) a polynucleotide encoding a polypeptide comprising a VH domain and (ii) a polynucleotide encoding a polypeptide comprising a VL domain, amplifying the polynucleotide of (i) and the polynucleotide of (ii) using barcoded primers, each coded to a respective well of the multiwell substrate to subsets of produce amplified barcoded polynucleotides, each subset coded to a respective well of the multiwell substrate, and sequencing amplified barcoded polynucleotides to produce the barcoded records; and (b) identifying cognate pairs of VH domains and VL domains based on the pairing score.
[0008] In some aspects, the disclosure provides a method of identifying cognate pairs of heavy chain variable (VH) domains and light chain variable (VL) domains from individual B cells comprising (a) sequencing a set of amplified barcoded polynucleotides to produce barcoded records, wherein the set of amplified barcoded polynucleotides comprises (i) a first subset of 2-100 polynucleotides encoding respective polypeptides, each comprising a respective VH domain and 2-100 polynucleotides encoding respective polypeptides, each comprising a respective VL domain, wherein the polynucleotides of the first subset are coded to a respective well of a multiwell plate from which the polynucleotides of the first subset are obtained; (ii) one or more additional subsets, each subset comprising 2-100 polynucleotides encoding respective polypeptides, each comprising a respective VH domain and 2-100 polynucleotides encoding respective polypeptides, each comprising a respective VL domain, wherein the polynucleotides of each of the one or more additional subsets are coded to a respective well of a multiwell plate from which the polynucleotides of each of the respective one or more additional subsets are obtained; (b) calculating a pairing score for respective barcoded records, wherein calculating the pairing score comprises calculating an intra-well abundance score and an inter-well abundance score; and (c) identifying cognate pairs of VH domains and VL domains based on the pairing score.
[0009] In some embodiments of methods provided herein, calculating the intra-well abundance score comprises, for each of the subsets of amplified barcoded polynucleotides, (i) identifying a copy number for the amplified barcoded polynucleotides that encode the VH domain and (ii) identifying a copy number for the amplified barcoded polynucleotides that encode the VL domain. In some embodiments, calculating the inter-well abundance score3#14358394v1comprises (i) for each VH domain, identifying the number of subsets comprising an amplified barcoded polynucleotide encoding the VH domain and (ii) for each VL domain, identifying the number of subsets comprising an amplified barcoded polynucleotide encoding the VL domain. In some embodiments, the inter-well abundance score is a positive inter-well abundance score when the number of subsets for the VH domain is at least 2 and the number of subsets for the VL domain is at least 2. In some embodiments, the intra-well abundance score is a positive intra-well abundance score when the copy number for the amplified barcoded polynucleotides that encode the VH domain and the copy number for the amplified barcoded polynucleotides that encode the VL domain is at a ratio of about 1:20 to 1:100. In some embodiments, the ratio is about 1:20 to 1:30. In some embodiments, the ratio is about 1:24.
[0010] In some embodiments, calculating the pairing score comprises identifying barcoded records having (i) a positive inter-well abundance score and (ii) a positive intra-well abundance score.
[0011] In some embodiments, a method provided herein further comprises calculating a mutation score for respective barcoded records. In some embodiments, calculating the mutation score comprises (i) aligning the polynucleotide encoding a polypeptide comprising a VH domain with a first germline polynucleotide, and (ii) identifying the number of mutations in the polynucleotide encoding a polypeptide comprising a VH domain compared to the first germline polynucleotide. In some embodiments, the first germline polynucleotide encodes a heavy chain polypeptide or a portion thereof. In some embodiments, calculating the mutation score comprises (i) aligning the polynucleotide encoding a polypeptide comprising a VL domain with a second germline polynucleotide, and (ii) identifying the number of mutations in the polynucleotide encoding a polypeptide comprising a VL domain compared to the second germline polynucleotide. In some embodiments, the second germline polynucleotide encodes a light chain polypeptide or a portion thereof. In some embodiments, the mutation score is a positive mutation score when the ratio of the number of mutations in the polynucleotide encoding a polypeptide comprising a VH domain to the number of mutations in the polynucleotide encoding a polypeptide comprising a VH domain is 1.5 to 5.
[0012] In some embodiments, a method provided herein further comprises calculating a developability score for respective barcoded records. In some embodiments, calculating the developability score comprises identifying one or more liability factors. In some embodiments, the liability factors are selected from N-linked glycosylation sites, single cysteine residues,4#14358394v1VH2 gene association, oxidation potential, deamidation potential, isomerization potential, and proteolytic cleavage potential.
[0013] In some embodiments, sequencing comprises next generation sequencing. In some embodiments, next generation sequencing comprises nanopore sequencing.
[0014] In some embodiments, the concentration is about 20-30 B cells per well. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] FIG.1 is an exemplary schematic of VH and VL sequences present in different wells of a multiwell substrate.
[0016] FIGs.2A-2B are exemplary results of copy numbers of VL sequences (FIG.2A), VH sequences (FIG.2B), and combined sequences (FIG.2C) associated with different barcoded primers and thus particular wells of a multiwell substrate.
[0017] FIG.3 is a graph depicting target binding activity of exemplary antibody heavy and light chain pairs relative to a reference antibody. B cells from mice immunized with a target protein were subjected to barcoded sequencing using the methods described herein. Antibody heavy and light chain pairs found to co-occur in 2 or more wells were produced recombinantly in CHO cells, purified by protein A, and tested for binding by ELISA compared to a reference antibody known to bind that same target.
[0018] FIGs. 4A-4C depict on and off rates of three exemplary antibodies produced using the methods described herein. B cells from mice immunized with three different target proteins were subjected to barcoded sequencing (Protein Target 1 (FIG. 4A), Protein Target 2 (FIG. 4B), and Protein Target 3 (FIG.4C). Antibody heavy and light chain pairs found to co-occur in 2 or more wells were produced recombinantly in CHO cells, purified by protein A, and tested for binding by Carterra LSA (SPRi) to determine on and off rates to their designated targets. Each spot represents the on and off rates of a single purified monoclonal antibody. DETAILED DESCRIPTION
[0019] Immunization of a host with an antigen of interest is an efficient way to generate a diversity of antibodies that can be further developed for therapeutic purposes. However, it can be costly and time-consuming to identify the unique variable heavy (VH) and variable light (VL) chain cognate pairs that make up the antibodies produced by different B cell clones,5#14358394v1usually requiring the use of single-cell B cell receptor (BCR) sequencing or single-cell next- generation sequencing (NGS). Single-cell sequencing is typically required to accurately identify VH / VL cognate pairs, to ensure the sequenced VH and VL chains originate from the same cell and thus belong to the same antibody. Due to the vast number of unique B cell clones and antibodies, single-cell sequencing substantially limits the degree of antibody identification that can be achieved. To overcome such limits, provided herein, in some embodiments, are high-throughput methods to identify cognate VH / VL pairs from individual B cells in a population of B cells. B Cells and Antibodies
[0020] Naturally occurring antibodies, also called immunoglobulins, are proteins typically made up of two identical heavy chains and two identical light chains. The specificity of an antibody for a given antigen is determined by the variable region of the heavy and light chains (variable heavy, VH, and variable light, VL, domains). The pairing of the VH and VL domains confers affinity and avidity of an antibody for its cognate antigen. Thus, identifying novel antibodies (e.g., antibodies produced upon immunization with an antigen) requires identifying both the VH and VL domains.
[0021] Antibodies are produced by B cells, a critical effector of the humoral response. During B cell development, immature B cells undergo V(D)J recombination to rearrange the variable (V), diversity (D), and joining (J) germline gene segments encoding the heavy chain of the B cell receptor (BCR), a portion of which corresponds to the antibodies produced by that B cell. Germline genes include genes that are inherited. Immature B cells additionally undergo VJ recombination to rearrange the V and J germline gene segments encoding the light chain of the BCR. V(D)J and VJ recombination of germline gene segments results in a diversity of unique BCRs expressed by naïve B cells that, upon encounter of antigen, become activated and undergo somatic hypermutation to improve the affinity of their BCR for its cognate antigen. Activated B cells ultimately become long-lived plasma cells, which secrete antibodies to confer long-term immunological protection.
[0022] B cells include, for example, plasma B cells, memory B cells, B1 cells, B2 cells, marginal-zone B cells, and follicular B cells. B-cells express immunoglobulins (antibodies, B cell receptor). In some aspects, a method of identifying cognate pairs of VH domains and VL domains (e.g., from individual B cells) comprises depositing B cells into wells of a multiwell substrate at a concentration of about 2-100 B cells per well (e.g., about 2, about 3, about 4,6#14358394v1about 5, about 6, about 7, about 8, about 9, about 10, about 15, about 20, about 25, about 30, about 35, about 40, about 45, about 50, about 55, about 60, about 65, about 70, about 75, about 80, about 85, about 90, about 95, or about 100 B cells per well), wherein each of the B cells comprises (i) a polynucleotide encoding a polypeptide comprising a VH domain and (ii) a polynucleotide encoding a polypeptide comprising a VL domain. “Polynucleotide” is used interchangeably with “nucleic acid” herein and refers to a polymer of (two or more) nucleotides (i.e., ribonucleotides and deoxyribonucleotides).
[0023] In some embodiments, B cells are obtained from a sample of B cells derived from a subject (e.g., isolated from a subject). In some embodiments, a sample of B cells includes at least 1,000 B cells, at least 10,000 B cells, at least 100,000 B cells, or more. In some embodiments, a sample includes a number of B cells in the range of 1,000 to 1,000,000 B cells. In some embodiments, B cells are enriched in a sample.
[0024] In some embodiments, at least 2 B cells are deposited per well. In some embodiments, at least 3, 4, 5, 6, 7, 8, 9, or 10 B cells are deposited per well. In some embodiments, at least 15 B cells are deposited per well. In some embodiments, at least 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, or at least 100 B cells are deposited per well. In some embodiments, no more than 100 B cells are deposited per well. In some embodiments, no more than 95, 90, 85, 80, 75, 70, 65, 60, 55, 50, 45, 40, 35, or no more than 30 B cells are deposited per well. In some embodiments, at least 3 B and no more than 100 B cells are deposited per well. In some embodiments, about 15 to about 35 B cells are deposited per well. In some embodiments, at least 15 B cells but no more than 35 B cells are deposited per well. In some embodiments, at least 20 B cells but no more than 30 B cells are deposited per well. In some embodiments, about 20–30 B cells are deposited per well. In some embodiments, about 20–25 B cells are deposited per well. In some embodiments, about 25–30 B cells are deposited per well. In some embodiments, 20–30 B cells are deposited per well. In some embodiments, 20– 25 B cells are deposited per well. In some embodiments, 25–30 B cells are deposited per well.
[0025] In some embodiments, the exact number of B cells deposited per well may be 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100. While the disclosure refers to wells of a multiwell plate, other substrates with multiple depressions, for example, arranged in a grid pattern, could be used. Thus, the term “well” includes any depression in a7#14358394v1substrate that can hold a cell and preferably cell media or other liquid containing the cell. It should be appreciated that a multiwell plate can include any number of wells.
[0026] The number of wells analyzed to identify cognate pairs of immunoglobulin heavy chain variable regions and light chain variable regions can be selected based on the number of cells per well, for example. In some embodiments, a multiwell substrate comprises one or more multi well plates. In some embodiments, a multiwell substrate comprises at least one multiwell plate comprising 96 wells. In some embodiments, a multiwell substrate comprises at least one multiwell plate comprising 384 wells. In some embodiments, a multiwell substrate comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more multiwell plates.
[0027] Since the genetic information of cognate pairing is present in the DNA of each subject’s adaptive immunity cells as well as their associated RNA transcripts, either RNA or DNA can be sequenced in the methods of the provided disclosure. In some embodiments, a recombined sequence from a T-cell or B-cell encoding a T cell receptor or immunoglobulin molecule, or a portion thereof, is referred to as a clonotype. The DNA or RNA corresponds to sequences from immunoglobulin (Ig) genes that encode antibodies. In some embodiments, the DNA and RNA analyzed in the methods of the disclosure correspond to sequences encoding VH domains and VL domains.
[0028] In some aspects, the disclosure provides a method of identifying cognate pairs of VH domains and VL domains (e.g., from individual B cells) comprising: (a) depositing B cells into wells of a multiwell substrate at a concentration of about 2-100 (e.g., 20-30) B cells per well, wherein each of the B cells comprises (i) a polynucleotide encoding a polypeptide comprising a VH domain and (ii) a polynucleotide encoding a polypeptide comprising a VL domain; (b) amplifying the polynucleotide of (i) and the polynucleotide of (ii) using barcoded primers, each primer coded to respective well of the multiwell substrate, to produce subsets of amplified barcoded polynucleotides, each subset coded to a respective well of the multiwell substrate; (c) sequencing the amplified barcoded polynucleotides to produce barcoded records; (d) calculating a pairing score for respective barcoded records, wherein calculating the pairing score comprises calculating an intra-well abundance score and an inter-well abundance score; and (e) identifying cognate pairs of VH domains and VL domains based on the pairing score.
[0029] In some aspects, the disclosure provides a method of identifying cognate pairs of VH domains and VL domains (e.g., from individual B cells) comprising: (a) sequencing amplified barcoded polynucleotides to produce barcoded records, wherein the amplified barcoded polynucleotides are produced by depositing B cells into wells of a multiwell substrate at a8#14358394v1concentration of about 2-100 (e.g., 20-30)cells per well, wherein each of the B cells comprises (i) a polynucleotide encoding a polypeptide comprising a VH domain and (ii) a polynucleotide encoding a polypeptide comprising a VL domain, and amplifying the polynucleotide of (i) and the polynucleotide of (ii) using barcoded primers, each primer coded to a respective well of the multiwell substrate, to produce subsets of amplified barcoded polynucleotides, each subset coded to a respective well of the multiwell substrate; (b) calculating a pairing score for respective barcoded records, wherein calculating the pairing score comprises calculating an intra-well abundance score and an inter-well abundance score; and (c) identifying cognate pairs of VH domains and VL domains based on the pairing score.
[0030] In some aspects, the disclosure provides a method of identifying cognate pairs of heavy chain variable (VH) domains and light chain variable (VL) domains (e.g., from individual B cells) comprising: (a) calculating a pairing score for respective barcoded records, wherein calculating the pairing score comprises calculating an intra-well abundance score and an inter-well abundance score, and wherein the barcoded records are produced by depositing B cells into wells of a multiwell substrate at a concentration of about 2-100 (e.g., 20-30)cells per well, wherein each of the B cells comprises (i) a polynucleotide encoding a polypeptide comprising a VH domain and (ii) a polynucleotide encoding a polypeptide comprising a VL domain, amplifying the polynucleotide of (i) and the polynucleotide of (ii) using barcoded primers, each coded to a respective well of the multiwell substrate to subsets of produce amplified barcoded polynucleotides, each subset coded to a respective well of the multiwell substrate, and sequencing amplified barcoded polynucleotides to produce the barcoded records; and (b) identifying cognate pairs of VH domains and VL domains based on the pairing score.
[0031] In some aspects, the disclosure provides a method of identifying cognate pairs of heavy chain variable (VH) domains and light chain variable (VL) domains (e.g., from individual B cells) comprising: (a) sequencing a set of amplified barcoded polynucleotides to produce barcoded records, wherein the set of amplified barcoded polynucleotides comprises (i) a first subset of 2-100 (e.g., 20-30)polynucleotides encoding respective polypeptides, each comprising a respective VH domain and 2-100 (e.g., 20-30) polynucleotides encoding respective polypeptides, each comprising a respective VL domain, wherein the polynucleotides of the first subset are coded to a respective well of a multiwell plate from which the polynucleotides of the first subset are obtained; (ii) one or more additional subsets, each subset comprising 2-100 (e.g., 20-30) polynucleotides encoding respective polypeptides, each comprising a respective VH domain and 2-100 (e.g., 20-30) polynucleotides encoding9#14358394v1respective polypeptides, each comprising a respective VL domain, wherein the polynucleotides of each of the one or more additional subsets are coded to a respective well of a multiwell plate from which the polynucleotides of each of the respective one or more additional subsets are obtained; (b) calculating a pairing score for respective barcoded records, wherein calculating the pairing score comprises calculating an intra-well abundance score and an inter-well abundance score; and (c) identifying cognate pairs of VH domains and VL domains based on the pairing score. Immunization
[0032] In some embodiments, the methods of the disclosure involve immunizing a subject with an antigen. In some embodiments, a subject is a mouse, rat, rabbit, goat, sheep, donkey, guinea pig, cow, pig, dog, horse, or any other mammal that responds to the antigen by developing an immune response. In some embodiments, a subject is a mouse. In some embodiments, a mouse is an engineered mouse. In some embodiments, an engineered mouse produces a human antibody. Non-limiting examples of mouse strains include ATX-GK, ATX- GL, CD1, BALB / c, and / or C57 / Bl6. In some embodiments, an antigen is a protein, peptide, and / or peptide fragment. In some embodiments, an antigen is a non-amino acid-based molecule. In some embodiments, antigens are naturally occurring, genetically engineered variants of the protein, and / or codon optimized for expression in a subject. Further, antigens of the present disclosure, can include modifications, such as deletions, additions and substitutions, generally conservative in nature, to the naturally occurring sequence, so long as the protein maintains its ability to elicit an immune response. In some embodiments, an antigen is human epidermal growth factor receptor 2 (HER2). In some embodiments, an antigen is a soluble protein, a synthetic peptide, and DNA molecule, and RNA molecules, a cell, a virus-like particle, or a fraction of a cell or cell lysate. In some embodiments, an antigen is an antigen associated with a cancer. In some embodiments, an antigen is an antigen associated with infectious disease. In some embodiments, an antigen is an antigen associated with an autoimmune disease. In some embodiments, an antigen is an antigen associated with an inflammatory disease. It should be appreciated that the methods described herein are not limited by a specific antigen.
[0033] In some embodiments, a sample is obtained from an immunized subject. In some embodiments, a sample is a biological sample. In some embodiments, samples used in the methods of the disclosure are from one or more tissues of the subject. For example, a sample10#14358394v1is from tumor tissue, blood and blood plasma, bone marrow, lymph fluid, cerebrospinal fluid surrounding the brain and the spinal cord, synovial fluid surrounding bone joints, and the like. In some embodiments, a sample is a blood sample. In some embodiments, a sample is a tumor biopsy. Non-limiting examples of tumor biopsy include biopsies from tumor of the brain, liver, lung, heart, colon, kidney, or bone marrow. In some embodiments, a sample is a biopsy, such as a biopsy of skin, brain, liver, lung, heart, colon, kidney, or bone marrow. Any biopsy technique used by those skilled in the art is used for isolating a sample from a subject. In some embodiments, a sample is obtained from bodily material of a subject and includes discarded material such as waste materials, shed skin cells, blood, teeth or hair. Barcoding
[0034] In some aspects, methods of identifying cognate VH and VL domains provided herein utilize barcoded primers to amplify polynucleotides encoding polypeptides comprising VH domains or VL domains, thereby producing amplified barcoded polynucleotides, which are then sequenced to produce barcoded records. In some embodiments, barcoded records are ultimately used to identify cognate VH and VL domains (e.g., cognate VH / VL pairs), as described elsewhere herein. Barcoded primers
[0035] In some aspects, methods provided herein comprise amplifying a polynucleotide encoding a polypeptide comprising a VH domain and a polynucleotide encoding a polypeptide comprising a VL domain using barcoded primers, wherein the barcoded primers are coded to respective wells of a multiwell substrate, each well additionally comprising 20-30 B cells of interest (e.g., B cells obtained from an immunized subject). Barcoded primers include primers (e.g. primers for polymerase chain reaction, PCR) comprising a unique nucleotide sequence that is coded to a specific well, which facilitates the identification of polynucleotides obtained from the same well. In some embodiments, barcoded primers target VH domain gene segments or portions of VH domain gene segments. In some embodiments, barcoded primers target at least one VH domain gene, or a portion thereof, selected from genes of any one of variable heavy family 1 genes (VH1), VH2, VH3, VH4, VH5, VH6, VH7, diversity heavy family 1 genes (VD1), VD2, VD3, VD4, VD5, VD6, VD7, joining heavy family 1 genes (VJ1), VJ2, VJ3, VJ4, VJ5, or VJ6. In some embodiments, barcoded primers target VL domain gene segments or portions of VL domain gene segments. In some embodiments, barcoded primers target at least one VL11#14358394v1domain gene, or a portion thereof, selected from genes of any one of kappa variable family 1 genes (Vκ1), Vκ2, Vκ3, Vκ4, Vκ5, Vκ6, Vκ7, kappa joining family 1 genes (Jκ1), Jκ2, Jκ3, Jκ4, Jκ5, lambda variable family 1 genes (Vλ1), Vλ2, Vλ3, Vλ4, Vλ5, Vλ6, Vλ7, Vλ8, Vλ9, Vλ10, Vλ11, lambda joining family 1 genes (Jλ1), Jλ2, Jλ3, Jλ4, Jλ5, Jλ6, and Jλ7. Exemplary primer sequences are shown in Tables 1 and 2. Any suitable barcodes or method of barcoding primers can be used and are contemplated herein. Table 1. Exemplary forward primer sequences.12#14358394v113#14358394v1Table 2. Exemplary reverse primer sequences.14#14358394v115#14358394v1Barcoded polynucleotides
[0036] In some aspects, methods provided herein comprise amplifying a polynucleotide encoding a polypeptide comprising a VH domain and a polynucleotide encoding a polypeptide comprising a VL domain using barcoded primers to produce subsets of amplified barcoded polynucleotides. As used herein, amplified barcoded polynucleotides include polynucleotides derived from B cells that have been amplified with barcoded primers and thus correspond to individual wells. In some embodiments, subsets of amplified barcoded polynucleotides comprise at least 2, at least 3, at least 10, at least 30, at least 100, at least 300, at least 1000, at16#14358394v1least 3000, at least 10,000, at least 30,000, at least 100,000, at least 300,000, at least 1,000,000, at least 3,000,000, at least 10,000,000, at least 30,000,000, or more amplified barcoded polynucleotides. In some embodiments, subsets of amplified barcoded polynucleotides encode cognate VH and VL pairs derived from at least 1, at least 2, at least 3, at least 10, at least 30, at least 100, at least 300, at least 1000, at least 10,000, at least 100,000, at least 1,000,000, at least 10,000,000, at least 1,000,000,000 or more B cell clones (e.g., individual B cells). Barcoded records
[0037] In some aspects, a method of identifying cognate pairs of VH domains and VL domains (e.g., from individual B cells) comprises sequencing amplified barcoded polynucleotides to produce barcoded records (i.e., the recorded sequences of amplified barcoded polynucleotides). As used herein, a barcoded record refers to a sequence (e.g., written record) of a polynucleotide comprising a polynucleotide derived from a barcoded primer (i.e. a unique barcode sequence) and a sequence encoding a polypeptide comprising a VH domain or a VL domain. Barcoded records are used to calculate a pairing score, described elsewhere herein, to identify cognate VH and VL pairs. Sequencing
[0038] Methods provided herein comprise sequencing amplified barcoded polynucleotides to produce barcoded records. In some embodiments, amplified barcoded polynucleotides from multiple wells of a multiwell substrate are pooled prior to sequencing. Any suitable technique for sequencing nucleic acids can be used. In some embodiments, sequencing reads are first demultiplexed according to well-specific barcodes incorporated during amplification.
[0039] DNA sequencing techniques include classic dideoxy sequencing reactions (Sanger method) using labeled terminators or primers and gel separation in slab or capillary electrophoresis. In some embodiments, next generation sequencing (NGS) platforms are used in the present disclosure. NGS sequencing refers to any post-classic Sanger type sequencing method which is capable of high throughput, multiplex sequencing of large numbers of samples simultaneously. Current NGS sequencing platforms, generate reads from multiple distinct nucleic acids in the same sequencing run. Throughput is varied, with 100 million bases to 600 giga bases per run, and throughput is rapidly increasing due to improvements in technology. The principle of operation of different NextGen sequencing platforms is also varied and includes: sequencing by synthesis using reversibly terminated labeled nucleotides, pyrosequencing, 454 sequencing, allele specific hybridization to a library of labeled17#14358394v1oligonucleotide probes, sequencing by synthesis using allele specific hybridization to a library of labeled clones that is followed by ligation, real time monitoring of the incorporation of labeled nucleotides during a polymerization step, colony sequencing, single molecule real time sequencing, and SOLiD sequencing. In some embodiments, sequencing comprises nanopore sequencing. In some embodiments, sequencing comprises NGS. In some embodiments, sequencing comprises sequencing by synthesis (SBS), semiconductor sequencing, single molecule real-time sequencing (SMRT), LoopSeq, DNA nanoball sequencing, stLFR, and other related approaches. Pairing Score
[0040] In some aspects, methods provided herein comprise calculating a pairing score for barcoded records, wherein calculating the pairing score comprises calculating an intra-well abundance score and an inter-well abundance score. Each barcoded record can be traced back to a source well based on the unique barcode sequence (derived from the barcoded primers) it comprises, because each barcode sequence is associated with a single “source” well from which the polynucleotide amplified by the barcoded primers was obtained. Sequences of polynucleotides encoding VH domains and sequences of polynucleotides encoding VL domains that comprise the same barcode sequence, but different VH domain-encoding or VL domain-encoding polynucleotides, indicate that multiple unique B cell clones were deposited into the well corresponding to the barcode sequence (see, e.g., exemplary wells A and B in FIG. 1, in which multiple VH and VL sequences are found). Conversely, sequences of polynucleotides encoding VH domains and sequences of polynucleotides encoding VL domains that comprise different barcode sequences (i.e., corresponding to multiple wells) but the same VH domain-encoding or VL domain-encoding polynucleotides indicate that the same B cell clones were deposited into multiple wells (see, e.g., exemplary wells B and C in FIG.1, which both contain VH4 and VL4 sequences). Thus, where two or more polynucleotides encoding the same VH domain comprise two or more barcodes, there must also be polynucleotides encoding the cognate VL domains comprising the same barcodes. Consequently, the frequency with which a polynucleotide encoding a VL domain is identified with the same barcode sequences (corresponding to a single well and / or multiple wells) as a polynucleotide encoding a VH domain, the higher the likelihood that said VL domain and said VH domain constitute a cognate pair. In some embodiments, calculating the pairing score comprises identifying barcoded records having (i) a positive inter-well abundance score and18#14358394v1(ii) a positive intra-well abundance score. Inter-well abundance score and intra-well abundance score are defined elsewhere herein.
[0041] In some embodiments, a pairing score is (P) is calculated as a weighted sum using the following non-limiting exemplary equation: ^= ^ ∙ ^ + ^ ∙ ^ + ^ ∙ ^
[0042] where: • ^ = intra-well abundance score • ^ = inter-well abundance score • ^ = abundance metric •^, ^, ^ = weighting factors chosen based on empirical performance (e.g., α = 0.5, β =0.4, γ = 0.1).
[0043] Other equations known in the art for calculating a pairing score are also contemplated for use with the methods described herein.
[0044] In some embodiments, bioinformatics methods are used to identify groups of sequences forming clonal families and subfamilies, and thereby immunoglobulin sequences of interest. Such bioinformatics methods involve measurements of sequence similarity.
[0045] In some embodiments, the heavy chains and / or light chain from B cells derived from a common progenitor are clonally related. Therefore, a heavy chain clonal family is associated with a light chain clonal family by observing the correlation across wells. Once an association is established between the heavy chains of a clonal family and the light chains of a clonal family, pairs are assigned in each well by selecting the heavy chain that is a member of the heavy chain clonal family and a light chain that is a member of the light chain clonal family.
[0046] In some embodiments, related immunoglobulin heavy and / or light chain sequences are identified through computational phylogenetic analysis of the homology between the immunoglobulin heavy chain and / or light chain V(D)J sequences. In some aspects, standard classification methods (i.e., clustering) of the sequences representing the individual or combinations of the V, D, and / or J gene segments and / or other sequences derived from the immunoglobulin heavy chain and / or light chain are used to identify clonal families or subfamilies (for example, by using ClustalX).
[0047] In some embodiments, the performance of a pairing score may be assessed using different thresholds. In some embodiments, as a non-limiting example, a precision-recall19#14358394v1analysis is performed to determine the proportion of cognate VH / VL pairs that are correct using the following exemplary equation: Precision =True Positives True Positives + False Positives
[0048] In some embodiments, a pairing score threshold (τ) is too low and admits many false- positive assignments. In some embodiments, a pairing score threshold (τ) is too high and reduces recall by excluding true cognates with borderline pairing scores. In some embodiments, a pairing score threshold (τ) is between 0.4 and 0.80. In some embodiments, a pairing score threshold (τ) is between 0.55 and 0.75. In some embodiments, a pairing score threshold is (τ) is about 0.65. It should be appreciated that a pairing score threshold may be adjusted based on the desired use case. For example, in some embodiments, a lower pairing score threshold may be used if high recall and low precision is desired. For example, in some embodiments, a higher pairing score threshold may be used if stringent selection is desired.
[0049] In some embodiments, reproducibility of replicate results of pairing scores may be assessed to determine overlap in top-ranked pairs. In some embodiments, reproducibility is calculated using the Jaccard index (J):
[0050] where ^ and ^ are the sets of cognate VH / VL pairs identified in each replicate.
[0051] Other equations known in the art for calculating reproducibility are also contemplated for use with the methods described herein.
[0052] In some embodiments, the accuracy of a pairing score is determined using spike-in validation. For example, in some embodiments, control wells are seeded with known monoclonal antibody-producing cells. Abundance Score Intra-well abundance score
[0053] In some embodiments, methods provided herein comprise calculating an intra-well abundance score as part of calculating a pairing score. An intra-well abundance score is20#14358394v1calculated by identifying the number of times within a given well that a polynucleotide encoding a particular VH domain and / or a polynucleotide encoding a particular VL domain is observed. For example, in the exemplary schematic shown in FIG.1, multiple copies of VH4 and VL4are found in well C, corresponding to a specific barcode. Primers containing the specific barcode corresponding to each well – e.g., the well C primers in the case of well C – are deposited in the wells. Upon amplification of the VH and VL sequences with the barcoded primers, the multiple copies of the VH4and VL4sequences originating from well C will be identified as being derived from the same well based on the specific barcode in the well C primers (the same is true of, e.g., the VH4 and VL4 sequences from well B that are amplified with the well B primers). That is, an intra-well abundance score is calculated by identifying the copy number of polynucleotides encoding a particular VH domain and having the same barcode sequence (i.e., coming from the same well), as well as by identifying the copy of number of polynucleotides encoding a particular VL domain and having the same barcode sequence (i.e., coming from the same well). Using the methods provided herein, it has been observed that, on average, there are 24 copies of polynucleotides encoding a given VL domain for every one copy of a polynucleotide encoding a cognate VH domain. In some embodiments, calculating the intra-well abundance score comprises, for each of the subsets of amplified barcoded polynucleotides, (i) identifying a copy number for the amplified barcoded polynucleotides that encode the VH domain and (ii) identifying a copy number of the amplified barcoded polynucleotides that encode the VL domain. In some embodiments, an intra-well abundance score is a positive intra-well abundance score when the copy number for the amplified barcoded polynucleotides that encode the VH domain and the copy number for the amplified barcoded polynucleotides that encode the VL domain is at a ratio of about 1:2 to 1:100. In some embodiments, a ratio is 1:2, 1:3, 1:4, 1:5, 1:6, 1:7, 1:8, 1:9, or 1:10. In some embodiments, a ratio is about 1:15, about 1:20, about 1:25, about 1:30, about 1:35, about 1:40, about 1:45, about 1:50, about 1:55, about 1:60, about 1:65, about 1:70, about 1:75, about 1:80, about 1:85, about 1:90, about 1:95, or about 1:100. In some embodiments, a ratio is 1:20 to 1:30. In some embodiments, a ratio 1:20 to 1:30. In some embodiments, a ratio is about 1:20 to 1:25. In some embodiments, a ratio is about 1:25 to 1:30. In some embodiments, a ratio is about 1:23 to 1:25. In some embodiments, a ratio is 1:24. Cognate VH and VL pairs assigned a positive intra-well abundance score are prioritized for further analysis and development.
[0054] In some embodiments, an intra-well abundance score (I) is calculated using the following exemplary equation:21#14358394v1%&
[0055] For each well w, the ratio of VH to VL sequences was calculated as # ='$In some embodiments, a positive intra-well abundance score is assigned when #$falls within the range of approximately 1:20–1:30, centered at 1:24. In some embodiments, the intra-well abundance score (I) for a candidate pair is defined as the number (or fraction) of wells in which the pair exhibited a positive ratio. Inter-well abundance score
[0056] In some embodiments, methods provided herein comprise calculating an inter-well abundance score as part of calculating a pairing score. An inter-well abundance score is calculated based on the number of unique barcodes (i.e., corresponding to different wells) found in polynucleotides encoding the same VH domain and / or polynucleotides encoding the same VL domain. For example, in the exemplary schematic shown in FIG. 1, VH1 and VL1 are found in multiple different wells – wells A, B, and D – each corresponding to different barcode. Amplification of the exemplary sequence with barcoded primers – i.e., with well A primers deposited in well A; well B primers deposited in well B; and well C primers deposited in well C – indicates that the copies of the VH1and VL1sequences are derived from different wells, which informs the VH1and VL1inter-well abundance score. The identification of polynucleotides encoding a given VH domain and / or a given VL domain comprising multiple unique barcodes indicates that the B cell clone from which said domains are derived has expanded. Expansion of a given B cell clone is a hallmark of antigen-specificity and B cell activation. The numeric value of the inter-well abundance score corresponds to the number of unique barcodes identified in polynucleotides encoding a given VH domain and / or a given VL domain. In some embodiments, calculating the inter-well abundance score comprises (i) for each VH domain, identifying the number of subsets comprising an amplified barcoded polynucleotide encoding the VH domain and (ii) for each VL domain, identifying the number of subsets comprising an amplified barcoded polynucleotide encoding the VL domain. In some embodiments, an inter-well abundance score is a positive inter-well abundance score when the number of subsets for the VH domain is at least 2 and the number of subsets for the VL domain is at least 2. Cognate VH and VL pairs assigned a positive inter-well abundance score are prioritized for further analysis and development.22#14358394v1
[0057] In some embodiments, an inter-well abundance score (R) is defined as the number of distinct wells in which the same VH and VL sequences co-occur. In some embodiments, a positive inter-well abundance score (R) is required at least two distinct wells.
[0058] In some embodiments, a positive inter-well abundance score is defined as recurrence (e.g., inter-well recurrence) in at least two distinct wells. Mutation Score
[0059] In some embodiments, a method provided herein further comprises calculating a mutation score for respective barcoded records. A mutation score can be calculated by assigning a positive score to cognate VH and VL pairs that have undergone somatic hypermutation (e.g. comprise mutations in the VH and / or VL domains), resulting in underlying DNA changes that encode for amino acid substitutions that can increase affinity of the antibody binding to its target. In particular, it is expected that mutations (e.g., substitutions, insertions, and / or deletions) will affect the affinity of an antibody binding its target if said mutations occur in the complementary determining regions (CDRs) of the VH and / or VL domains. CDRs are regions of hypervariability within VH and VL domains that confer antigen specificity to antibodies. Mutations in the regions between the CDRs, called the framework regions, can also contribute to the functionality of an antibody (e.g., binding affinity), as such mutations may affect the overall antibody structure. In some embodiments, a barcoded record receives a positive mutation score when at least one mutation is identified in a CDR of a given VH domain or a given VL domain. In some embodiments, a barcoded record receives a positive mutation score when mutations are present in both a CDR and a framework of a given VH domain or a given VL domain, and the ratio of CDR mutations to framework mutations is at least 1:1.
[0060] In some embodiments, calculating the mutation score comprises (i) aligning the polynucleotide encoding a polypeptide comprising a VH domain with a first germline polynucleotide, and (ii) identifying the number of mutations in the polynucleotide encoding a polypeptide comprising a VH domain compared to the first germline polynucleotide. In some embodiments, the mutations are mutations that are predicted to change the amino acid sequence of the VH domain (e.g. non-conservative mutations). A first germline polynucleotide includes a first inherited gene segment, e.g., an inherited gene segment that encodes a VH domain or a portion thereof. In some embodiments, a first germline polynucleotide encodes a heavy chain polypeptide or a portion thereof.23#14358394v1
[0061] In some embodiments, calculating the mutation score comprises (i) aligning the polynucleotide encoding a polypeptide comprising a VL domain with a second germline polynucleotide, and (ii) identifying the number of mutations in the polynucleotide encoding a polypeptide comprising a VL domain compared to the second germline polynucleotide. In some embodiments, the mutations are mutations that are predicted to change the amino acid sequence of the VL domain (e.g. non-conservative mutations). A second germline polynucleotide includes a second inherited gene segment, e.g., an inherited gene segment that encodes a VL domain or a portion thereof. In some embodiments, a second germline polynucleotide encodes a light chain polypeptide or a portion thereof.
[0062] In some embodiments, a mutation score is a positive mutation score when the ratio of the number of mutations in the polynucleotide encoding a polypeptide comprising a VH domain to the number of mutations in the polynucleotide encoding a polypeptide comprising a VL domain is 1.5 to 5. In some embodiments, a ratio is 1.5 to 4. In some embodiments, a ratio is 1.5 to 3. In some embodiments, a ratio is 1.5 to 2. In some embodiments, a ratio is 2. In some embodiments, a ratio is 1.5. Cognate VH and VL pairs receiving a positive mutation score are prioritized for further analysis and development.
[0063] In some embodiments, a mutation score (M) is calculated for each VH / VL pair as: )=*+,-. where *+,-is the total number of observed mutations across the VH and VL sequences, and . is the total number of nucleotides or amino acids compared. In some embodiments, VH / VL pairs with higher mutation scores are more likely to represent affinity-matured antibodies, consistent with antigen-driven somatic hypermutation. It should be appreciated that other equations known in the art for calculating mutation scores of VH / VL pairs are contemplated for use with the methods described herein Developability Score
[0064] In some embodiments, methods provided herein further comprise calculating a developability score for respective barcoded records. A developability score can be calculating by penalizing VH and VL domains with liability factors. Liability factors are defined as characteristics that may comprise the structure, functionality, or further use (e.g., conjugation24#14358394v1to an additional antibody or drug) of the cognate VH and VL pair. Cognate VH and VL pairs with mutations that change germline cysteine residues in either the VH or VL domain receive a negative developability score. Cognate VH and VL pairs with mutations that create a single additional cysteine residue anywhere in the VH or VL domain, but not a new cysteine pair, receive a negative score. Cognate VH and VL pairs with mutations that create an N-linked glycosylation site (e.g. NXS or NXT, where X is any amino acid except proline, P) receive a negative score. Cognate VH and VL pairs derived from gene segments encoding VH and VL domains known to be difficult to manufacture (e.g., VH2 genes) receive a negative score. Cognate VH and VL pairs that contain consensus sequences known to have potential for oxidation, deamidation, isomerization, and proteolytic cleavage receive a negative score.
[0065] In some embodiments, calculating the developability score comprises identifying one or more liability factors. In some embodiments, the liability factors are selected from N-linked glycosylation sites, single cysteine residues, VH2 gene association, oxidation potential, deamidation potential, isomerization potential, and proteolytic cleavage potential. Liability factors can include any characteristics known to increase the difficulty of producing, testing, improving, or modifying an antibody. Liability factors have been characterized and are known in the art (see, e.g., Jain et al., Identifying developability risks for clinical progression of antibodies using high-throughput in vitro and in silico approaches. Mabs, 2023; 15(1): 2200540; see also, e.g., Jain et al., Biophysical properties of the clinical-stage antibody landscape. App Biol Sci, 2017; 114(5): 944-949.)
[0066] In some embodiments, developability score (D) is assigned according to the following relationship:where *2345323-367is the number of identified liability features, and *864-,967is the total number of features assessed. In some embodiments, a developability score is within 0 (least developable, liabilities in all assessed features) to 1 (most developable, no liabilities). Examples of liabilities include but are not limited to unpaired cysteines, oxidation, deamidation, and glycosylation sites. It should be appreciated that other equations known in the art for calculating developability scores of VH / VL pairs are contemplated for use with the methods described herein.25#14358394v1Integrated Score
[0067] In some embodiments, a method provided herein further comprises calculating an integrated score for respective barcoded records. To prioritize cognate VH / VL pairs with favorable properties, an integrated score (S) may be calculated using the following exemplary formula: := ; ∙ ) + (1 − ;) ∙ / where: • ) = mutation score (fractional divergence from germline), • / = developability score (fraction of features free of liabilities), • ; = weighting factor (0 ≤ λ ≤ 1) selected according to desired balance between affinity maturation and developability.
[0068] In some embodiments, λ is set to 0.5, yielding equal weight to mutation and developability. In some embodiments, λ is set to a value that may be biased toward mutation (e.g., λ = 0.7) during early discovery to favor affinity. In some embodiments, λ is set to a value that may be biased toward developability (e.g., λ = 0.3) during lead selection. Thus, in some embodiments, a weighting factor ; is varied depending on the desired balance between affinity maturation and developability for a particular use case.
[0069] It should be appreciated that other equations known in the art for calculating developability scores of VH / VL pairs are contemplated for use with the methods described herein. Development of Identified Antibodies
[0070] Once cognate pairs of immunoglobulin light chain and heavy chain regions are determined by the methods described herein, they are synthetically generated by DNA synthesis for cloning into expression vectors or incorporation into transcriptionally active PCR products. In some embodiments, the sequence used for the synthesis are derived directly from the high-throughput NGS sequences. In some embodiments, variable regions of Ig genes are cloned by DNA synthesis and incorporated into a vector containing the appropriate constant region using restriction enzymes and standard molecular biology. In some embodiments, the exact nucleotide sequence of the immunoglobulin variable regions is codon optimized and / or further mutated to increase protein expression levels. Restriction sites are also added for the26#14358394v1purpose of cloning. In some embodiments, the amplified V(D)J regions are inserted into vectors that already contain either the kappa, lambda, gamma or other heavy chain isotype constant regions.
[0071] In some embodiments, antibody or antigen binding molecules are produced with the identified candidate cognate pairs of immunoglobulin heavy chain and light chain variable regions. Libraries of candidate cognate pairs are generated and screened according to methods known in the art. Methods for screening antibodies of the disclosure include using display strategies such as phage display strategies, yeast display strategies, and / or ribosome display strategies. In some embodiments, affinity of the antibodies against the antigen utilized in the methods described herein is measured. In these screening methods, the antigen is fixed and panning is used to remove weakly bound clones, leaving those clones bound intact to the antigen. This method is characterized by repeated selection, proliferation, and enrichment of positive clones to enable the processing of large libraries.
[0072] In some embodiments, screening methods also include multiple rounds of selection to enrich for one or more antibodies with improved antigen binding. In some embodiments, at each round of selection, further amino acid mutations are introduced into the antibodies using art recognized methods. In some embodiments, at each round of selection the stringency of binding to the antigen is increased to select for antibodies with increased affinity for a desired target.
[0073] Screening methods for increasing the stringency of the binding to the antigen are also performed. Any art recognized methods of increasing the stringency of an antibody-antigen interaction assay are useful herein. In some embodiments, one or more of the assay conditions are varied (for example, the salt concentration of the assay buffer) to reduce the affinity of the antibody molecules for the desired antigen. In some embodiments, the length of time permitted for the antibodies to bind to the desired target is reduced. In some embodiments, a competitive binding step is added to the antibody-antigen interaction assay.
[0074] Once candidates with desired binding and / or functional properties are identified, they are advanced to relevant assays and other assessments based on the desired, downstream, product profile. For therapeutic antibodies intended for use in passive immunization, candidates are advanced to assays and preclinical testing models to determine the best candidates for clinical testing in humans or for use in animal health, including, but not limited, assessments of properties such as stability and aggregation, formulation and dosing ease, protein expression and manufacturing, species selectivity, pharmacology, pharmacokinetics,27#14358394v1safety and toxicology, absorption, metabolism and target-antibody turnover, as well as immunogenicity. For diagnostics, antibodies prepared by identification of cognate pairs of immunoglobulin chains are used as diagnostic probes for biomarkers or provide immune system information about the disease state of, or effect of treatment on, a human or animal. EXAMPLES
[0075] Example 1: High-throughput identification of cognate VH / VL pairs.
[0076] To identify cognate variable heavy domain / variable light domain (VH / VL) pairs of novel antibodies for potential therapeutic applications, an antigen is suspended in a suitable carrier and injected into a mouse. After at least one day, cells are harvested from the blood, bone marrow, secondary lymphoid organs (i.e., spleen and lymph nodes) or other organs of the mouse and B cells are isolated from the harvested tissue. B cells may be enriched and then are deposited into a multiwell plate at a concentration of approximately 2-100 B cells per well (e.g., approximately 20 B cells per well). B cells are lysed and reverse transcription is performed. Following reverse transcription, the resulting cDNA is amplified using barcoded primers designed against VH and VL genes, where each set of barcoded primers contains a unique barcode sequence corresponding to a specific well. Amplified sequences are subjected to next-generation sequencing. Pairing scores are calculated for barcoded, sequenced VH and VL polynucleotides based on the inter-well and intra-well abundance scores, which are determined from the number of copies of each sequences and which wells they were derived from. Positive intra-well abundance scores are calculated when unique VH and VL polynucleotides are found within the same wells at a copy number ratio of approximately 1:100 (e.g., approximately 1:24) of VH polynucleotide copies to VL polynucleotide copies. Positive inter-well abundance scores are calculated when the same VH and VL polynucleotides are identified as being derived from multiple different wells.
[0077] Upon identification of VH / VL cognate pairs using the pairing scores, VH / VL cognate pairs are further subjected to analysis to determine mutation scores and developability scores. The mutation score is calculated by comparing the germline VH gene sequences and VL gene sequences with the sequenced VH and VL polynucleotides, respectively, and identifying the number of mutations (insertions, deletions, or substitutions) in the VH and VL polynucleotides that are not present in the germline sequences. VH / VL cognate pairs with more mutations receive a positive mutation score, as they are predicted to have a greater affinity for antigen A due to having arisen from B cells in the mouse undergoing somatic hypermutation after28#14358394v1encountering antigen A. The developability score is calculated by analyzing the sequenced VH / VL cognate pair polynucleotides and identifying regions of the polynucleotides that include N-linked glycosylation sites, single cysteine residues, are associated with VH2 genes, and / or affect oxidation potential, deamidation potential, isomerization potential, and proteolytic cleavage potential. The presence of regions known to increase the difficulty of developing, testing, modifying, or improving antibodies decreases the developability score.
[0078] Taken together, the methods disclosed herein allow for high-throughput identification of cognate VH / VL pairs belonging to novel antibodies of high affinity for antigens of interest that can be developed for therapeutic applications. Example 2. Overview of Examples 3-17
[0079] The following Examples collectively illustrate an exemplary end-to-end antibody discovery workflow, beginning with immunization and B-cell harvest, proceeding through reverse transcription, barcoding, sequencing, and computational pairing, and extending to post- pairing analyses and functional validation. Using this workflow, thousands of candidate VH / VL pairs have been identified, scored, and prioritized. Examples 3–15 describe exemplary embodiments of individual components of the process, and Example 16 provides an exemplary integrated summary and workflow schematic highlighting key performance metrics across all stages. Data provided in Example 16 corresponds to an exemplary discovery program wherein hundreds of candidate VH / VL pairs were identified, scored and prioritized to express multiple antibodies, for which 29 antibodies demonstrated measurable, dose-dependent binding to Antigen A with nanomolar affinities (FIG.3) Example 17 provides additional exemplary data (FIGs. 4A-4C) demonstrating that the methods described herein can be generalized across antigens. Example 3. Immunization and B Cell Isolation
[0080] The following provides an exemplary protocol for immunization and B-cell isolation that may be employed in certain embodiments to carry out the methods of the present disclosure. Protocols substantially identical or functionally similar to the one described herein have been performed successfully by the Applicant across numerous antibody discovery campaigns. A person of ordinary skill in the art, upon reading the present disclosure, will understand how to implement immunization and B-cell isolation using the representative29#14358394v1protocol described, with routine modifications as appropriate for a given antigen or host or system.
[0081] In some embodiments, antibody discovery is performed using the ATX-Gx™ humanized transgenic mice antibody generation platform available from ALLOY THERAPEUTICS, though comparable platforms may be used. In certain embodiments, a cohort of mice (e.g., with 5 female animals per group) are immunized with antigen formulated in adjuvant according to a multi-dose protocol. In one exemplary implementation, immunization follows a standard RIMMS protocol with dosing of 10 µg antigen emulsified in complete Freund’s adjuvant on Day 0, followed by 10 µg antigen emulsified in incomplete Freund’s adjuvant on Days 7, 14, and 21. On Day 24, spleen, lymph node, bone marrow, and peripheral blood are collected. In several embodiments, approximately 5 × 10⁷ total B cells are recovered across the cohort. B cells may be diluted to 2 × 10⁵ cells / mL and plated at ~20 cells per well into 96-well or 384-well PCR plates. An exemplary cohort and plating summary is shown in Table 3.
[0082] Table 3. Exemplary Cohort and Plating SummaryExample 4. Primer Architecture and Barcoding
[0083] The following provides an exemplary protocol for reverse transcription, amplification, and sequencing that may be employed in certain embodiments to carry out the methods of the present disclosure. Protocols substantially identical or functionally similar to the one described herein have been performed successfully by the Applicant across numerous antibody discovery campaigns. A person of ordinary skill in the art, upon reading the present disclosure, will understand how to implement reverse transcription, amplification, and sequencing using the representative protocol described, with routine modifications as appropriate for a given sample preparation or sequencing platform.
[0084] In some embodiments, B cells plated as described in Example 3 are lysed in-well and cDNA synthesized. Ig variable regions are amplified using multiplex primer pools. In certain embodiments, forward primers anneal to framework region 1 (FR1) of VH, VK, and VL gene families, and reverse primers annealed to constant regions (CH1 or CL). Each primer may include a 5′ extension comprising: (i) a well-specific barcode; (ii) a molecular identifier30#14358394v1(UMI) for deduplication; and (iii) a universal sequencing adapter tail. In one exemplary implementation, barcodes corresponded to SEQ ID NOs: 1–192disclosed herein, although other barcodes may be employed.
[0085] In certain embodiments, reverse transcription and amplification are performed in 20 μL reactions using a one-step RT-PCR kit (e.g., Superscript III, Invitrogen 12574026). In some embodiments, up to 20 ng of total RNA template is added per reaction. In some embodiments, the primer cocktail contains forward primers covering FR1 regions of human VH, VK, and VL gene families and reverse primers in CH1 and CL regions. Exemplary cycling conditions may include: • 50 °C, 30 minutes (reverse transcription) • 95 °C, 15 minutes (enzyme activation) • 25 cycles of: 94 °C, 30 seconds; 62 °C, 1 minute; 72 °C, 1 minute • 72 °C, 5 minutes; 4 °C hold
[0086] In some embodiments, following amplification, amplified products are evaluated by agarose gel electrophoresis (2% gel). Wells with visible bands may be advanced directly to library preparation. In some embodiments, wells lacking sufficient product are subjected to an additional “rescue” amplification (Round 1.5) using ExTaq DNA Polymerase (e.g., Takara RR001A) and the same primer cocktail. In these cases, 1 μL of the initial RT-PCR reaction is used as template in three 25 μL reactions, which are pooled after cycling. Exemplary rescue PCR conditions include: • 95 °C, 3 minutes • 12 cycles of: 94 °C, 30 seconds; 62 °C, 1 minute; 72 °C, 1 minute • 72 °C, 5 minutes; 4 °C hold
[0087] In several embodiments, PCR products are pooled and quantified. Concentration was measured by spectrophotometric (Nanodrop A260) or fluorometric (Ribogreen) assay.
[0088] Sequencing libraries may be prepared using a ligation-based protocol (e.g., Oxford Nanopore SQK-LSK114) including a DNA repair step (e.g., NEBNext Companion Module, NEB E1780S). Final libraries may be eluted in 25 μL of elution buffer, and 12 μL loaded onto a MinION flow cell. In exemplary embodiments, each run generates at least 8 million total reads.31#14358394v1Example 5. Sequencing and Demultiplexing
[0089] The following provides an exemplary protocol for sequencing and demultiplexing that may be employed in certain embodiments to carry out the methods of the present disclosure. Protocols substantially identical or functionally similar to the one described herein have been performed successfully by the Applicant across numerous antibody discovery campaigns. A person of ordinary skill in the art, upon reading the present disclosure, will understand how to implement sequencing and demultiplexing using the representative protocol described, with routine modifications as appropriate for a given antigen or host or system.
[0090] Sequencing reads are first demultiplexed according to the well-specific barcodes incorporated during amplification. Consensus sequences are generated for each unique molecule.
[0091] In some embodiments, the resulting high-quality sequences may be aligned to VH and VL germline reference repertoires to determine the closest germline gene family assignments. Exemplary sequencing quality metrics are shown in Table 4. Table 4. Exemplary Sequencing Quality MetricsQ30 = a sequencing quality parameter indicating 99.9% of base calling is accurate Example 6. Pairing Score Calculation
[0092] The following provides an exemplary protocol for calculating a pairing score that may be employed in certain embodiments to carry out the methods of the present disclosure. Protocols substantially identical or functionally similar to the one described herein have been performed successfully by the Applicant across numerous antibody discovery campaigns. A person of ordinary skill in the art, upon reading the present disclosure, will understand how to implement a pairing score calculation using the representative protocol described, with routine modifications as appropriate for a given antigen or host or system.
[0093] In some embodiments, cognate VH / VL pairs are identified using an intra-well abundance score (I), an inter-well abundance score (R), and an abundance metric (A).
[0094] Intra-well abundance score (I): In some embodiments, for each well w, the ratio of %& VH to VL sequences are calculated as #'$= %('. A positive intra-well abundance score is32#14358394v1assigned when #$falls within the range of approximately 1:2 to 1:100. In some embodiments, a positive intra-well abundance score is assigned when #$falls within the range of approximately 1:20–1:30, for example, e.g., centered at 1:24. In some embodiments, the intra- well abundance score (I) for a candidate pair is defined as the number (or fraction) of wells in which the pair exhibits a positive ratio.
[0095] Inter-well abundance score (R): In some embodiments, is defined as the number of distinct wells in which the same VH and VL sequences co-occur. In some embodiments, a positive inter-well abundance score (R) requires at least two distinct wells.
[0096] Abundance metric (A): In some embodiments, is defined as the normalized total read depth across all wells for a candidate pair, scaled between 0 and 1.
[0097] In some embodiments, the overall pairing score (P) is calculated as a weighted sum: ^= ^ ∙ ^ + ^ ∙ ^ + ^ ∙ ^where: • ^ = intra-well abundance score; • ^ = inter-well abundance score; • ^ = abundance metric; •^, ^, ^ = weighting factors chosen based on empirical performance (e.g., α = 0.5, β =0.4, γ = 0.1).
[0098] In some embodiments, pairs are designated as cognate when ^ ≥ = (e.g., = = 0.65)and separated from the next highest-scoring alternative by a margin Δ (e.g., ≥0.05). Exemplary pairing scores are shown in Table 5. Table 5. Pairing Scores
[0099] It should be appreciated that other mathematical formulations known in the art may be employed to combine intra-well and inter-well abundance scores to determine pairing scores. Thus, the above equation is provided as one non-limiting example for determining pairing scores.33#14358394v1Example 7. Spike-in Validation
[0100] The following provides an exemplary protocol for spike-in validation that may be employed in certain embodiments to carry out the methods of the present disclosure. Protocols substantially identical or functionally similar to the one described herein have been performed successfully by the Applicant across numerous antibody discovery campaigns. A person of ordinary skill in the art, upon reading the present disclosure, will understand how to implement spike-in validation using the representative protocol described, with routine modifications as appropriate for a given antigen or host or system.
[0101] In some embodiments, to validate the accuracy of the pairing score framework, control wells are seeded with known monoclonal antibody-producing cells (e.g., spike-in validation). In some embodiments, these wells contain B cells expressing VH / VL pairs of defined identity. In some embodiments, the standard workflow (lysis, reverse transcription, amplification with barcoded primers, sequencing, and calculation of intra-well and inter-well abundance scores) is applied identically to both spike-in control wells and experimental wells.
[0102] In some embodiments, pairing scores are then calculated as described in Example 6. In some embodiments, predicted VH / VL assignments are compared to the known true pairings. In an exemplary embodiment, 18 of 20 spiked-in wells yielded the correct VH / VL assignment, corresponding to an accuracy of 90%. In some embodiments, errors consist of either (i) insufficient read depth leading to no call, or (ii) misassignment where the true pair scored just below threshold.
[0103] This exemplary embodiment demonstrates that the pairing score, derived from intra- well abundance scores (VH:VL ratio) and inter-well abundance scores (recurrence across multiple wells), reliably identifies true cognate VH / VL pairs under the same ~20-cells-per-well occupancy conditions used for discovery. Exemplary spike-in validation is shown below in Table 6. Table 6. Exemplary Spike-in Validation34#14358394v1Example 8. Intra-Well Ratio Distribution
[0104] The following provides an exemplary protocol for intra-well ratio distribution that may be employed in certain embodiments to carry out the methods of the present disclosure. Protocols substantially identical or functionally similar to the one described herein have been performed successfully by the Applicant across numerous antibody discovery campaigns. A person of ordinary skill in the art, upon reading the present disclosure, will understand how to implement and / or analyze intra-well ratio distribution using the representative protocol described, with routine modifications as appropriate for a given antigen or host or system.
[0105] In some embodiments, the distribution of VH:VL copy ratios across wells is analyzed. In an exemplary experiment, a positive intra-well abundance score was defined as a ratio of approximately 1:20–1:30 (e.g., centered around 1:24). Exemlary analysis of >100 wells showed that 85% of candidate cognate pairs fell within this interval, with a median VH:VL ratio of 1:24. In contrast, exemplary non-cognate combinations were broadly distributed and did not cluster near the expected ratio. Exemplary intra-well ratio values are shown in Table 7. Table 7. Exemplary Intra-well Ratio SummaryExample 9. Inter-Well Recurrence
[0106] The following provides an exemplary protocol for analyzing inter-well recurrence that may be employed in certain embodiments to carry out the methods of the present disclosure. Protocols substantially identical or functionally similar to the one described herein have been performed successfully by the Applicant across numerous antibody discovery campaigns. A person of ordinary skill in the art, upon reading the present disclosure, will understand how to implement and / or analyze inter-well recurrence using the representative protocol described, with routine modifications as appropriate for a given antigen or host or system.
[0107] In some embodiments, the frequency with which candidate VH / VL pairs are observed across multiple wells is analyzed. In some embodiments, a positive inter-well abundance score is defined as recurrence (e.g., inter-well recurrence) in at least two distinct wells.
[0108] In an exemplary embodiment, across 384 wells plated at ~20 cells per well, true cognate VH / VL pairs identified by pairing score thresholds (e.g., Example 7) were found to35#14358394v1recur in multiple wells with high frequency. In contrast, exemplary non-cognate combinations (generated by random reassignment of VH and VL sequences) rarely recurred more than once.
[0109] This exemplary result demonstrates that recurrence across wells provides a reliable signature of clonal expansion and supports the designation of a positive inter-well abundance score when the same VH and VL sequences are observed in two or more distinct wells. Exemplary inter-well recurrence frequences are shown in Table 8. Table 8. Exemplary Inter-well Recurrence Frequencies (Scaffold)Example 10. Threshold Optimization
[0110] The following provides an exemplary protocol for threshold optimization that may be employed in certain embodiments to carry out the methods of the present disclosure. Protocols substantially identical or functionally similar to the one described herein have been performed successfully by the Applicant across numerous antibody discovery campaigns. A person of ordinary skill in the art, upon reading the present disclosure, will understand how to implement threshold optimization using the representative protocol described, with routine modifications as appropriate for a given antigen or host or system.
[0111] In some embodiments, to evaluate the performance of a pairing score (exemplified in Example 6) across different thresholds, a precision–recall analysis is performed. In some embodiments, candidate VH / VL pairs are scored using intra-well abundance scores, inter-well abundance scores, and abundance metrics as described herein. Spike-in controls of known VH / VL pairs (exemplified in Example 7) are used as ground truth for true positives, while shuffled VH / VL combinations serve as negative controls.
[0112] In some embodiments, Precision is defined as the proportion of called cognate pairs that were truly correct: Precision =True Positives True Positives + False Positives36#14358394v1
[0113] In some embodiments, pairing score thresholds (τ) ranging from 0.40 to 0.80 are tested. At each threshold, precision, recall, and F1 score are calculated.
[0114] In an exemplary embodiment, results demonstrated that thresholds below 0.55 admitted too many false-positive assignments, while thresholds above 0.75 reduced recall by excluding true cognates with borderline scores. In an exemplary embodiment, a threshold of τ = 0.65 achieved the best balance, yielding 91% precision and 87% recall in the spike-in validation set, with an F1 score of 0.89. Exemplary threshold optimization values are shown in Table 9. Table 9. Exemplary Threshold OptimizationExample 11. Replicate Reproducibility
[0115] The following provides an exemplary protocol for analyzing replicate reproducibility that may be employed in certain embodiments to carry out the methods of the present disclosure. Protocols substantially identical or functionally similar to the one described herein have been performed successfully by the Applicant across numerous antibody discovery campaigns. A person of ordinary skill in the art, upon reading the present disclosure, will understand how to implement and / or analyze replicate reproducbility using the representative protocol described, with routine modifications as appropriate for a given antigen or host or system.
[0116] In some embodiments, to evaluate the reproducibility of the method, replicate plates containing approximately 20 B cells per well are processed independently, from cell lysis through sequencing and analysis. In some embodiments, candidate VH / VL pairs are identified in each replicate using the exemplary pairing score framework described in Example 6.
[0117] In some embodiments, pairing results are compared between replicates by calculating the overlap in top-ranked pairs. In an exemplary experiment, the top 30 scoring pairs in each replicate were analyzed. In some embodiments, reproducibility is measured using the Jaccard index (J), defined as:37#14358394v1where ^ and ^ are the sets of cognate VH / VL pairs identified in each replicate.
[0118] In an exemplary embodiment, the Jaccard index across replicate plates was 0.72, corresponding to 72% overlap among the top 30 identified pairs. In an exemplary embodiment, broader analysis across all pairs above the threshold (τ = 0.65) yielded consistent overlap values (0.70–0.75). These exemplary results demonstrate that the method produces reproducible identification of cognate VH / VL pairs across independent replicates. Exemplary replicate reproducibility is shown in Table 10. Table 10. Exemplary Replicate ReproducibilityExample 12. Mutation Score Analysis of Cognate VH / VL Pairs
[0119] The following provides an exemplary protocol for mutation score analysis of cognate VH / VL pairs that may be employed in certain embodiments to carry out the methods of the present disclosure. Protocols substantially identical or functionally similar to the one described herein have been performed successfully by the Applicant across numerous antibody discovery campaigns. A person of ordinary skill in the art, upon reading the present disclosure, will understand how to implement mutation score analysis of cognate VH / VL pairs using the representative protocol described, with routine modifications as appropriate for a given antigen or host or system.
[0120] In some embodiments, cognate VH / VL pairs identified using the pairing score framework (as exemplified in Example 6) are further analyzed for mutation burden relative to their respective germline sequences. In some embodiments, germline references are obtained from the IMGT repertoire. For each VH and VL sequence, the number of nucleotide substitutions, insertions, or deletions relative to the assigned germline gene is recorded.
[0121] In some embodiments, a mutation score (M) is calculated for each pair as:38#14358394v1where *+,-is the total number of observed mutations across the VH and VL sequences, and . is the total number of nucleotides or amino acids compared.
[0122] In some embodiments, pairs with higher mutation scores are considered more likely to represent affinity-matured antibodies, consistent with antigen-driven somatic hypermutation.
[0123] In an exemplary embodiment, 29 antibodies exhibiting measurable binding to Antigen A (FIG.3) were analyzed. The median mutation score among binders was 0.08 (8% divergence from germline), compared to 0.02 (2% divergence) among non-binding candidate pairs. Exemplary mutation score analysis is shown in Table 11. Table 11. Exemplary Mutation Score AnalysisExample 13. Developability Analysis
[0124] The following provides an exemplary protocol for developability analysis that may be employed in certain embodiments to carry out the methods of the present disclosure. Protocols substantially identical or functionally similar to the one described herein have been performed successfully by the Applicant across numerous antibody discovery campaigns. A person of ordinary skill in the art, upon reading the present disclosure, will understand how to implement developability analysis using the representative protocol described, with routine modifications as appropriate for a given antigen or host or system.
[0125] In some embodiments, cognate VH / VL pairs identified using the pairing score framework (exemplified in Example 6) are analyzed for sequence features predictive of antibody developability. In some embodiments, for each sequence, regions known to present liabilities are identified, including: • N-linked glycosylation motifs (N-X-S / T),39#14358394v1• Unpaired cysteine residues, • VH2 family usage, and • Residues associated with chemical instability (oxidation-prone methionine, deamidation-prone asparagine, isomerization-prone aspartate, proteolysis-prone motifs).
[0126] In some embodiments, a developability score (D) is assigned according to the following relationship:where *2345323-367is the number of identified liability features, and *864-,967is the total number of features assessed.
[0127] In some embodiments, scores range from 0 (least developable, liabilities in all assessed features) to 1 (most developable, no liabilities).
[0128] In an exemplary embodiment, sequences exhibiting strong binding to Antigen A but containing multiple liabilities (e.g., unpaired cysteines and glycosylation sites) received reduced developability scores, while sequences lacking such liabilities scored higher. Exemplary developability score analysis is shown in Table 12. Table 12. Exemplary Developability Score AnalysisExample 14. Integrated Scoring of Cognate VH / VL Pairs
[0129] The following provides an exemplary protocol for integrated scoring of cognate VH / VL pairs that may be employed in certain embodiments to carry out the methods of the40#14358394v1present disclosure. Protocols substantially identical or functionally similar to the one described herein have been performed successfully by the Applicant across numerous antibody discovery campaigns. A person of ordinary skill in the art, upon reading the present disclosure, will understand how to implement integrated scoring of cognate VH / VL pairs using the representative protocol described, with routine modifications as appropriate for a given antigen or host or system.
[0130] In some embodiments, cognate VH / VL pairs identified by the pairing score (exemplified in Example 6) are further evaluated using both the mutation score (M) (exemplified in Example 12) and the developability score (D) (exemplified in Example 13). In some embodiments, to prioritize antibodies with favorable properties, an integrated score (S) is calculated as: := ; ∙ ) + (1 − ;) ∙ / where: • ) = mutation score (fractional divergence from germline), • / = developability score (fraction of features free of liabilities), • ; = weighting factor (0 ≤ λ ≤ 1) selected according to desired balance between affinity maturation and developability.
[0131] In some embodiments, λ is set to 0.5, yielding equal weight to mutation and developability. In other embodiments, λ may be biased toward mutation (e.g., λ = 0.7) during early discovery to favor affinity, or toward developability (e.g., λ = 0.3) during lead selection. It should be appreciated that different weighting factors may be selected based on the desired balance between affinity maturation and developability for a particular use case.
[0132] An exemplary application of this integrated framework to 29 antigen-binding antibodies revealed that top-ranked candidates consistently combined moderate-to-high mutation scores (≥0.05) with high developability scores (≥0.80). Antibodies with high mutation scores but multiple liabilities ranked lower, demonstrating that the method disclosed herein distinguishes between affinity maturation and developability risks. Exemplary integrated scoring of antigen-binding antibodies is shown in Table 13. Table 13. Exemplary Integrated Scoring of Antigen-Binding Antibodies (scaffold)41#14358394v1Example 15. Functional Confirmation of Selected Antibodies
[0133] The following provides an exemplary protocol for functional confirmation of selected antibodies that may be employed in certain embodiments to carry out the methods of the present disclosure. Protocols substantially identical or functionally similar to the one described herein have been performed successfully by the Applicant across numerous antibody discovery campaigns. A person of ordinary skill in the art, upon reading the present disclosure, will understand how to implement functional confirmation of selected antibodies using the representative protocol described, with routine modifications as appropriate for a given antigen or host or system.
[0134] In some embodiments, to validate that cognate VH / VL pairs identified by a method described herein can be expressed as functional antibodies, a panel of exemplary antibodies is expressed and tested for binding to Antigen A. In some embodiments, antibodies are chosen to represent a broad range of mutation scores (as exemplified in Example 12), developability scores (as exemplified in Example 13), and integrated scores (as exemplified in Example 14).
[0135] In some embodiments, VH and VL sequences are cloned into human IgG1 / kappa expression vectors and transiently expressed in Expi293F cells (Thermo Fisher). Supernatants are harvested (e.g., after 5 days), and antibodies are purified (e.g., by Protein A chromatography).
[0136] In some embodiments, binding activity is assessed (e.g., by ELISA). Plates are coated with recombinant Antigen A (1 µg / mL), blocked with 1% BSA, and incubated with serial dilutions (e.g., 10-point serial dilutions) of purified antibodies. Bound antibodies are detected using anti-human IgG-HRP secondary antibody and developed with TMB substrate. Absorbance is measured at 450 nm, and binding affinities are estimated as half-maximal effective concentration (EC50).
[0137] From an exemplary embodiment tested, 29 antibodies demonstrated measurable, dose-dependent binding to Antigen A, with EC50 values ranging from ~0.5–20 nM. Antibodies with higher integrated scores (≥0.75) generally exhibited stronger binding (lower EC50) and higher expression yields, whereas antibodies with low integrated scores (<0.50) were less likely to bind or were poorly expressed. Exemplary functional binding results are shown in Table 14.42#14358394v1Table 14. Functional Binding ResultsExample 16. Exemplary End-to-End Workflow on Antibody Panel
[0138] The following provides an exemplary protocol for end-to-end workflow on an antibody panel that may be employed in certain embodiments to carry out the methods of the present disclosure. Protocols substantially identical or functionally similar to the one described herein have been performed successfully by the Applicant across numerous antibody discovery campaigns. A person of ordinary skill in the art, upon reading the present disclosure, will understand how to implement the end-to-end workflow on an antibody panel using the representative protocol described, with routine modifications as appropriate for a given antigen or host or system.
[0139] The complete antibody discovery workflow described in Examples 3–15 was applied to an exemplary panel of antibodies against Antigen A. This integrated example illustrates the sequence of operations from B-cell isolation through functional validation of expressed antibodies. 1. Immunization and B-cell harvest (Example 3): Mice were immunized with Antigen A, and ~20 B cells were deposited per well into 96-well plates. 2. Reverse transcription and amplification with barcoded primers (Example 4): Ig variable regions were reverse-transcribed and amplified using FR1 / CH1 and FR1 / CL primers incorporating well-specific barcodes. 3. Sequencing and demultiplexing (Example 5): Products were sequenced, demultiplexed by barcode, and collapsed by consensus to yield high-quality VH and VL repertoires. 4. Pairing score calculation (Example 6): Intra- and inter-well abundance scores were used to assign cognate VH / VL pairs.43#14358394v15. Validation and optimization (Examples 7-11): Pairing accuracy was validated with spike-in controls, intra- and inter-well metrics were quantified, thresholds optimized, reproducibility measured, and barcode collision ruled out. 6. Post-pairing sequence analyses (Examples 12-14): Cognate pairs were evaluated for mutation burden, developability liabilities, and integrated scoring. 7. Functional confirmation (Example 15): 29 antibodies were confirmed to bind Antigen A with nanomolar affinities (e.g., a positive binding call).
[0140] Together, this exemplary end-to-end workflow demonstrates a fully integrated discovery pipeline: starting from immunized animals, through sequencing and computational pairing, to validated antigen-specific antibodies with supporting developability metrics. Exemplary end-to-end antibody discovery results are shown in Table 15. Table 15. End-to-End Antibody Discovery ResultsExample 17. Generalization Across Antigens
[0141] The following provides an exemplary embodiment that the methods of the present disclosure may be generalized across antigens. Protocols substantially identical or functionally similar to the one described herein have been performed successfully by the Applicant across numerous (e.g., hundreds) of antibody discovery campaigns. A person of ordinary skill in the art, upon reading the present disclosure, will understand how the representative protocol44#14358394v1described herein may be used to generalize across antigens, with routine modifications as appropriate for a given antigen or host or system.
[0142] To demonstrate the general applicability of the disclosed workflow, the same end- to-end process (Examples 3–17) was applied to immunization campaigns against multiple antigens, including exemplary Antigen A (Protein Target 1), exemplary Antigen B (Protein Target 2), and exemplary Antigen C (Protein Target 3) shown herein. In each case, B cells were isolated from immunized animals, processed at ~20 cells per well, and analyzed using the barcoded primer, sequencing, and pairing workflow described above.
[0143] Pairing and sequencing performance were comparable across all antigens (Protein Targets 1, 2, and 3), with ~300–500 cognate VH / VL pairs identified per campaign. Inter-well recurrence rates, intra-well ratio distributions, and threshold optimization results were consistent with those observed for exemplary Antigen A, indicating that the pairing score framework described in Example 6 is antigen-independent. Thus, it should be appreciated that the pairing score framework described herein could be applied to produce antibodies to any antigen.
[0144] Functional confirmation was obtained by expressing and testing a panel of antibodies from each campaign. In the case of exemplary Antigen B, 22 antibodies demonstrated measurable binding (EC50 0.5–15 nM), and for exemplary Antigen C, 31 antibodies showed measurable binding (EC500.7–25 nM). In each instance, antibodies with high integrated scores (≥0.75) were enriched among confirmed binders, validating the predictive value of the scoring framework.
[0145] On and off rates were determined for purified monoclonal antibodies to Protein Target 1 (FIG.4A), Protein Target 2 (FIG.4B), and Protein Target 3 (FIG.4C). Exemplary functional results are shown in Table 16. Table 16. Exemplary Functional Results Across Antigens
[0146] Taken together, the foregoing Examples demonstrate that the disclosed workflow enables reproducible identification of cognate VH / VL pairs from small populations of B cells,45#14358394v1accurate assignment of cognate partners using intra-well and inter-well abundance scores, and robust prioritization of candidates through mutation and developability scoring. In an exemplary embodiment, the integrated pipeline was validated functionally by the expression and characterization of more than 300 antibodies against exemplary Antigen A, of which twenty-nine exhibited measurable binding in the nanomolar range, and further generalized across additional antigens (e.g., exemplary Antigen B and exemplary Antigen C), confirming the broad applicability of the method. These results establish that the methods described herein are not limited to a single antigen or application but instead provide a general framework for high-throughput antibody discovery suitable for diverse therapeutic targets. EQUIVALENTS AND SCOPE
[0147] While several inventive embodiments have been described and illustrated herein, those of ordinary skill in the art will readily envision a variety of other means and / or structures for performing the function and / or obtaining the results and / or one or more of the advantages described herein, and each of such variations and / or modifications is deemed to be within the scope of the inventive embodiments described herein. More generally, those skilled in the art will readily appreciate that all parameters, dimensions, materials, and configurations described herein are meant to be exemplary and that the actual parameters, dimensions, materials, and / or configurations will depend upon the specific application or applications for which the inventive teachings is / are used. Those skilled in the art will recognize or be able to ascertain using no more than routine experimentation, many equivalents to the specific inventive embodiments described herein. It is, therefore, to be understood that the foregoing embodiments are presented by way of example only and that, within the scope of the appended claims and equivalents thereto, inventive embodiments may be practiced otherwise than as specifically described and claimed. Inventive embodiments of the present disclosure are directed to each individual feature, system, article, material, kit, and / or method described herein. In addition, any combination of two or more such features, systems, articles, materials, kits, and / or methods, if such features, systems, articles, materials, kits, and / or methods are not mutually inconsistent, is included within the inventive scope of the present disclosure.
[0148] The indefinite articles “a” and “an,” as used herein in the specification and in the claims, unless clearly indicated to the contrary, should be understood to mean “at least one.”46#14358394v1
[0149] It should also be understood that, unless clearly indicated to the contrary, in any methods claimed herein that include more than one step or act, the order of the steps or acts of the method is not necessarily limited to the order in which the steps or acts of the method are recited.
[0150] In the claims, as well as in the specification above, all transitional phrases such as “comprising,” “including,” “carrying,” “having,” “containing,” “involving,” “holding,” “composed of,” and the like are to be understood to be open-ended, i.e., to mean including but not limited to. Only the transitional phrases “consisting of” and “consisting essentially of” shall be closed or semi-closed transitional phrases, respectively, as set forth in the United States Patent Office Manual of Patent Examining Procedures, Section 2111.03.
[0151] The terms “about” and “substantially” preceding a numerical value mean ±10% of the recited numerical value.
[0152] Where a range of values is provided, each value between and including the upper and lower ends of the range are specifically contemplated and described herein.47#14358394v1
Claims
What is claimed is: CLAIMS 1. A method of identifying cognate pairs of heavy chain variable (VH) domains and light chain variable (VL) domains from individual B cells comprising: (a) depositing B cells into wells of a multiwell substrate at a concentration of about 2- 100 B cells per well, wherein each of the B cells comprises (i) a polynucleotide encoding a polypeptide comprising a VH domain and (ii) a polynucleotide encoding a polypeptide comprising a VL domain; (b) amplifying the polynucleotide of (i) and the polynucleotide of (ii) using barcoded primers, each primer coded to a respective well of the multiwell substrate, to produce subsets of amplified barcoded polynucleotides, each subset coded to a respective well of the multiwell substrate; (c) sequencing the amplified barcoded polynucleotides to produce barcoded records; (d) calculating a pairing score for respective barcoded records, wherein calculating the pairing score comprises calculating an intra-well abundance score and an inter-well abundance score; and (e) identifying cognate pairs of VH domains and VL domains based on the pairing score.
2. A method of identifying cognate pairs of heavy chain variable (VH) domains and light chain variable (VL) domains from individual B cells comprising: (a) sequencing amplified barcoded polynucleotides to produce barcoded records, wherein the amplified barcoded polynucleotides are produced by depositing B cells into wells of a multiwell substrate at a concentration of about 2-100 cells per well, wherein each of the B cells comprises (i) a polynucleotide encoding a polypeptide comprising a VH domain and (ii) a polynucleotide encoding a polypeptide comprising a VL domain, and amplifying the polynucleotide of (i) and the polynucleotide of (ii) using barcoded primers, each primer coded to a respective well of the multiwell substrate, to produce subsets of amplified barcoded polynucleotides, each subset coded to a respective well of the multiwell substrate;48#14358394v1(b) calculating a pairing score for respective barcoded records, wherein calculating the pairing score comprises calculating an intra-well abundance score and an inter-well abundance score; and (c) identifying cognate pairs of VH domains and VL domains based on the pairing score.
3. A method of identifying cognate pairs of heavy chain variable (VH) domains and light chain variable (VL) domains from individual B cells comprising: (a) calculating a pairing score for respective barcoded records, wherein calculating the pairing score comprises calculating an intra-well abundance score and an inter-well abundance score, and wherein the barcoded records are produced by depositing B cells into wells of a multiwell substrate at a concentration of about 2-100 cells per well, wherein each of the B cells comprises (i) a polynucleotide encoding a polypeptide comprising a VH domain and (ii) a polynucleotide encoding a polypeptide comprising a VL domain, amplifying the polynucleotide of (i) and the polynucleotide of (ii) using barcoded primers, each coded to a respective well of the multiwell substrate to subsets of produce amplified barcoded polynucleotides, each subset coded to a respective well of the multiwell substrate, and sequencing amplified barcoded polynucleotides to produce the barcoded records; and (b) identifying cognate pairs of VH domains and VL domains based on the pairing score.
4. A method of identifying cognate pairs of heavy chain variable (VH) domains and light chain variable (VL) domains from individual B cells comprising: (a) sequencing a set of amplified barcoded polynucleotides to produce barcoded records, wherein the set of amplified barcoded polynucleotides comprises (i) a first subset of 2-100 polynucleotides encoding respective polypeptides, each comprising a respective VH domain and 2-100 polynucleotides encoding respective polypeptides, each comprising a respective VL domain,49#14358394v1wherein the polynucleotides of the first subset are coded to a respective well of a multiwell plate from which the polynucleotides of the first subset are obtained; (ii) one or more additional subsets, each subset comprising 2-100polynucleotides encoding respective polypeptides, each comprising a respective VH domain and 2-100 polynucleotides encoding respective polypeptides, each comprising a respective VL domain, wherein the polynucleotides of each of the one or more additional subsets are coded to a respective well of a multiwell plate from which the polynucleotides of each of the respective one or more additional subsets are obtained; (b) calculating a pairing score for respective barcoded records, wherein calculating the pairing score comprises calculating an intra-well abundance score and an inter-well abundance score; and (c) identifying cognate pairs of VH domains and VL domains based on the pairing score.
5. The method of any one of the preceding claims, wherein calculating the intra-well abundance score comprises, for each of the subsets of amplified barcoded polynucleotides, (i) identifying a copy number for the amplified barcoded polynucleotides that encode the VH domain and (ii) identifying a copy number for the amplified barcoded polynucleotides that encode the VL domain.
6. The method of any one of the preceding claims, wherein calculating the inter-well abundance score comprises (i) for each VH domain, identifying the number of subsets comprising an amplified barcoded polynucleotide encoding the VH domain and (ii) for each VL domain, identifying the number of subsets comprising an amplified barcoded polynucleotide encoding the VL domain.
7. The method of any one of the preceding claims, wherein the inter-well abundance score is a positive inter-well abundance score when the number of subsets for the VH domain is at least 2 and the number of subsets for the VL domain is at least 2.50#14358394v18. The method of any one of the preceding claims, wherein the intra-well abundance score is a positive intra-well abundance score when the copy number for the amplified barcoded polynucleotides that encode the VH domain and the copy number for the amplified barcoded polynucleotides that encode the VL domain is at a ratio of about 1:2 to 1:100, optionally about 1:20 to 1:
30.
9. The method of claim 8, wherein the ratio is about 1:
24.
10. The method of any one of the preceding claims, wherein calculating the pairing score comprises identifying barcoded records having (i) a positive inter-well abundance score and (ii) a positive intra-well abundance score.
11. The method of any one of the preceding claims, wherein the method further comprises calculating a mutation score for respective barcoded records.
12. The method of claim 11, wherein calculating the mutation score comprises (i) aligning the polynucleotide encoding a polypeptide comprising a VH domain with a first germline polynucleotide, and (ii) identifying the number of mutations in the polynucleotide encoding a polypeptide comprising a VH domain compared to the first germline polynucleotide.
13. The method of claim 12, wherein the first germline polynucleotide encodes a heavy chain polypeptide or a portion thereof.
14. The method of any one of claims 11-13, wherein calculating the mutation score comprises (i) aligning the polynucleotide encoding a polypeptide comprising a VL domain with a second germline polynucleotide, and (ii) identifying the number of mutations in the polynucleotide encoding a polypeptide comprising a VL domain compared to the second germline polynucleotide.
15. The method of claim 14, wherein the second germline polynucleotide encodes a light chain polypeptide or a portion thereof.51#14358394v116. The method of any one of claims 11-15, wherein the mutation score is a positive mutation score when the ratio of the number of mutations in the polynucleotide encoding a polypeptide comprising a VH domain to the number of mutations in the polynucleotide encoding a polypeptide comprising a VH domain is 1.5 to 5.
17. The method of any one of the preceding claims, wherein the method further comprises calculating a developability score for respective barcoded records.
18. The method of claim 17, wherein calculating the developability score comprises identifying one or more liability factors.
19. The method of claim 18, wherein the liability factors are selected from N-linked glycosylation sites, single cysteine residues, VH2 gene association, oxidation potential, deamidation potential, isomerization potential, and proteolytic cleavage potential.
20. The method of any one of claims 1-19, wherein sequencing comprises next generation sequencing.
21. The method of claim 20, wherein next generation sequencing comprises nanopore sequencing.
22. The method of any one of claims 1-20, wherein the concentration is about 20-30 B cells per well.52#14358394v1
Citation Information
Patent Citations
Identification of antigen-specific B cell receptors
US10428325B1
Identification of polynucleotides associated with a sample
US11098302B2
Uniquely tagged rearranged adaptive immune receptor genes in a complex gene set
US20160024493A1