Sequence-based high throughput method generating camelid antibodies to cover broad epitopes with high-resolution

The high-throughput method for generating camelid antibodies by enriching antigen-specific B cells and creating an NGS library addresses the challenges of insufficient epitope coverage in current antibody technologies, achieving comprehensive and effective epitope targeting.

JP2025084951APending Publication Date: 2025-06-03ZHEJIANG NANOMAB TECH CENT CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025033577
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2019-02-27
Filing Date
2025-03-04
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

Current antibody technologies face challenges in targeting functional epitopes of disease targets due to insufficient epitope coverage, redundant selection, and low efficacy, largely attributed to random and sporadic binder generation.

Method used

A high-throughput method is developed to generate camelid antibodies by enriching and proliferating antigen-specific B cells from immunized camelids, creating an antibody next-generation sequencing (NGS) library containing VHH, VH, and VL chain sequences, and selecting representative sequences based on strain priority factors to ensure comprehensive epitope coverage.

Benefits of technology

This method enables the systematic and rational targeting of functional epitopes, achieving broad and high-resolution epitope coverage, thereby improving the efficacy of antibody-based diagnostic and therapeutic applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025084951000017
    Figure 2025084951000017
  • Figure 2025084951000018
    Figure 2025084951000018
  • Figure 2025084951000019
    Figure 2025084951000019
Patent Text Reader

Abstract

To provide a method for generating a camelid antibody specific for an antigen.SOLUTION: The present invention provides a method comprising: a) enriching and proliferating B-cells from immunized camelids specific to an antigen; b) generating antibody NGS libraries comprising VHH2, VHH3, and VH1 chain sequences from the antigen-specific B-cells; c) grouping antibody sequences of VHH2, VHH3, and VH1 chain in the NGS libraries by lineages; d) ranking the lineages from step c) by one or more lineage priority factors; e) selecting a representative sequence from a lineage comprising a VHH2 or VHH3 with a top ranking from step d) in the NGS library; and f) testing an antibody comprising the selected sequence from step e) to determine if the antibody binds to the antigen or a portion thereof.SELECTED DRAWING: None
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Each target has hundreds or thousands of epitopes, and the number of epitopes involved in biological functions is very limited. Therefore, with current antibody technology, targeting the functional epitopes of disease targets is a major challenge. However, since current technology generates binders randomly and sporadically, insufficient coverage of epitopes, redundant selection, and low efficacy are bottlenecks.

[0002] Dromedary camels and Bactrian camels belong to the family Camelidae of the Old World, and alpacas belong to the family Camelidae of the New World. Only the family Camelidae (commonly known as camelids) has a dichotomous adaptive humoral immune system that includes both conventional antibodies and heavy-chain antibodies (HcAbs). Furthermore, HcAbs have evolved a comprehensive paratope architecture as one of the driving factors for recognizing a very broad range of epitopes of antigens, and IgG1 antibodies complement the HcAb binding architecture for more diverse recognition.

[0003] The camelidae family has a unique humoral immune system consisting of two types of HcAbs, IgG2 and IgG3, which have short and long hinge regions. Phylogenetic analysis has confirmed that HcAbs branched off from the conventional antibody IgG1 as a result of recent adaptive changes. IgG1 and IgG3 neutralize West Nile virus, while IgG2 has been reported to be less effective in infected or vaccinated animals (Daley LP, Clin. Vaccine Immunol. 17:239-46, 2010). Furthermore, the epitope ranges sampled by HcAbs and IgG1 can overlap, but HcAbs can also reach sites inaccessible to IgG1. The understanding of the exact roles and functions of the various camelidae IgG isotypes is still in its infancy. However, the diverse paratope architectures of HcAbs (IgG2 and IgG3), such as oblong, convex, concave, protruding, and flat surfaces, provide the greatest opportunity to develop antibodies against the target in question, especially for diagnostic and therapeutic applications. The simplicity of HcAbs without light chain pairing also makes gene cloning and antibody manipulation much easier. Additionally, conventional IgG1, which contributes 25-50% of the total IgG in camelidae animals, shows a different recognition pattern from the HcAb repertoire of immunized human dromedaries or llamas (McCoy LE, J. Exp. Med. 2012), and specific unique epitopes or drugable target hotspots are accessible to IgG1 due to high affinity and desired functionality (Cristina Basilico, The Journal of Clinical Investigation, Volume 124 Number 7 July, 2014; Bas van der Woninga, MABS, VOL. 8, NO. 6, 1126-1135, 2016), thus playing an important role in expanding the antigen-binding repertoire.Camelidae animals have two types of light chains (Vκ or Vλ, which pair with VH1 to form conventional IgG1, and whose germline organization has recently been revealed (Laura M. Griffin, Journal of Immunological Methods Volume 405, Pages 35 - 46, March 2014; Alex Klarenbeek, mAbs 7:4, 693 - 706; 2015).

[0004] Extensive somatic hypermutation and gene conversion are significantly higher between VHHs than between VHs within the primary VHH B cell repertoire (30% vs. 1.5%), further supporting the further diversification of the HcAb repertoire to complement the lack of light chains. Equally importantly, the VHH domain of HcAb expands the overall antigen-binding repertoire, for example, by creating an oblong (rugby ball-shaped) structure with a convex paratope surface, which makes it very suitable for insertion into cavities or grooves (such as active sites and allosteric sites) on the surface of the antigen. In contrast, the VH-VL domains of conventional IgG contain a flatter or concave paratope surface. The following mechanisms of B cell repertoire diversification contribute significantly to the unique binding properties of VHH: (i) Most VHHs contain a framework region 2 (FR2) with hydrophilic amino acid substitutions compared to conventional FR2 (Val37→Phe / Tyr, Gly44→Glu, Leu45→Arg, and Trp47→Gly), which are involved in light chain binding. (ii) The expanded CDR1 region is associated with extensive somatic hypermutation in immune B cells within residues 27-30 according to Kabat numbering. (iii) Additional disulfide bonds between CDR1-CDR3 (camel) or FR2-CDR3 (llama and alpaca) in most VHHs. (iv) Additional disulfide bonds within CDR1 and CDR3 in specific parts of VHH. (v) Longer CDR3 loops are also identified, perhaps by additional non-template nucleotide insertions in some VHHs (Adhdi Arbabi-Ghahroudi, Frontiers in Immunology, Vol 8, 2017; Viet Khong Nguyen, The EMBO Journal Vol.19 No.5 2000; Mehdi Arbabi-Ghahroudi et al, Front. Immunol., November 20, 2017; Nguyen VK, Immunogenetics 54:39-47, 2002; Conrath KE, Dev Comp Immunol 27:87-103, 2003).There is a set of non-classical VHHs (without FR2 hydrophilic amino acids) derived from the same IGHV3 or IGHV4, D, and J gene loci as conventional IgG1. It has also been found that these heavy-chain antibodies can recognize the same or similar epitopes as IgG1 because both categories of antibodies share the same or similar CDR3s involved in epitope recognition (Conrath KE Dev Comp Immunol 27:87-103, 2003; Nick Deschacht, The Journal of Immunology. 184(10)5696-5704, 2010). The composition of the HcAb germline and the VHH structure are shown in FIGS. 1A and B.

[0005] Functional and physicochemical advantages such as high affinity, specificity, simple gene cloning, high expression yield, ease of purification, and a highly soluble and stable single-domain fold form the basis of HcAb technology. Furthermore, the antigen-binding repertoire expanded by conventional IgG1 enables even broader epitope coverage. Additionally, the close homology with the human counterparts of VHH, VH, Vκ, and Vλ brings great advantages to humanization and the development of therapeutic methods. By utilizing the unique antibody tissues of camels and NGS technology to capture the entire B-cell antibody repertoire, a novel method has been developed here to generate hundreds or thousands of diverse antibodies that comprehensively cover a wide range of target epitopes at high resolution, thereby enabling the targeting of these important and functional epitopes in a systematic and rational manner.

Summary of the Invention

Problems to be Solved by the Invention

[0006] The present invention discloses a high-throughput method for generating camelid antibodies against an antigen, the method comprising: a) enriching and proliferating antigen-specific B cells derived from an immunized camelid; and b) generating an antibody next-generation sequencing (NGS) library containing VHH 2 , VHH 3 and VH1-chain sequences; and c) VHHs within the NGS library2 , VHH 3 , and VH 1 grouping the sequences of VHH, VH, and VHH by strain, and d) ranking the strains containing VHH heavy chains (VHH 2 , VHH 3 ) by one or more strain priority factors, and e) selecting representative sequences from the strains of the top-ranked VHH heavy chains (VHH 2 and VHH 3 ) in the NGS database library, and f) testing the antibodies containing the selected VHH heavy chain sequences to determine whether the antibodies bind to the antigen or a part thereof. In one embodiment, the antigen comprises a plurality of epitopes.

[0007] In one embodiment, the minimum CDR3 distance of a particular CDR3 is 1 or less between groups of CDR3s from the strains, and the minimum CDR3 distance of a particular CDR3 is the shortest Hamming distance of this CDR3 compared to all other CDR3s of the same length.

[0008] In some embodiments, the strain priority factors are selected from the group consisting of strains from high sequence abundance to low sequence abundance, strains from high amplification factor to low amplification factor after enrichment and proliferation of B cells in vitro, changes in the abundance of strain sequences during the immune process, changes in the abundance of strain sequences before and after depletion of specific unwanted B cells, strains sharing the same naive B cell origin between VHH and VH, avoidance of developability liability sequences, and combinations thereof.

[0009] In some embodiments, the VHH 2 and / or VHH 3 strains are selected from the top 100 strains in e).

[0010] In some embodiments, the method further comprises repeating e) - f) to generate camelid antibodies, and representative sequences are selected from the top 101 - 200, 201 - 300, 301 - 400, 401 - 500, 501 - 600, 601 - 700, 701 - 800, 801 - 900, 901 - 1000, 1001 - 1100, 1101 - 1200, 1201 - 1300, 1301 - 1400, 1401 - 1500, 1501 - 1600, 1601 - 1700, 1701 - 1800, 1801 - 1900, or 1901 - 2000 ranked lineages. In some embodiments, the method further comprises repeating e) - f) to generate camelid antibodies, and representative sequences are selected from the top 2000 - 10,000 ranked lineages.

[0011] In some embodiments, the test antibody is expressed by prokaryotic or eukaryotic cells.

[0012] In preferred embodiments, the method further comprises monitoring the immune responses of IgG2, 3 (HcAb) and IgG1 (conventional IgG).

[0013] In some embodiments, sequences within the same lineage group of antibodies of only the selected IgG2 or IgG3 heavy chain can be selected for antibody optimization by repeating e) - f) of the method.

[0014] In some embodiments, the antigen or immunogen can be a cell, tissue, or biological fluid.

[0015] In some embodiments, the antigen can be a complex immunogen, and the method further comprises identifying individual antigens contained in the complex immunogen by protein array, cell / tissue antigen cDNA library, or mass spectrometry-based immunoprecipitation using the antibody determined to bind to the complex immunogen in step (f).

[0016] In one aspect of the present invention, the method includes: i) an FR2 hydrophilic region; ii) an extended CDR1; iii) an additional disulfide bond between CDR1-CDR3 or FR2-CDR3; iv) an additional disulfide bond within CDR3; v) a long CDR3 (≧15aa); vi) an additional disulfide bond within CDR1; vii) a non-classical VHH having the same V and J germline as conventional IgG1; viii) a non-classical VHH having a specific predetermined sequence signature; ix) a specific predetermined canonical binding loop structure; x) a convergent motif or sequence signature between individuals of the same immune group; xi) the length of CDR2; xii) the length of CDR3; xiii) the length and identity of CDR3; xiv) the presence of three or more positive charges within the CDR3 region; xv) the number of cysteines in the amino acid sequence; and xvi) further includes lineage subgrouping having specific VHH characteristics selected from the group consisting of 2-4 amino acid motifs found in the CDR region. The motif is identified from the three-dimensional structure of the ligand / receptor complex.

[0017] In another aspect of the present invention, a method for generating camelid antibodies against an antigen in high throughput is provided. The method includes: a) concentrating and proliferating antigen-specific B cells from immunized camelids; b) generating an antibody NGS library containing VHH 2 , VHH 3 , VH 1 and VL1 chain sequences; c) grouping the VHH 2 , VHH 3 , VH 1 , and VL 1 NGS sequences by lineage; d) pairing VH 1 / VL 1 lineages according to the anchor binders generated by single B cell sorting and the hetero-hybridoma approach; e) ranking the lineages and lineage pairs of steps c) and d) by lineage preference factors; f) the top-ranked VHH 2 , VHH 3 lineages and VH 1 / VL 1selecting a representative sequence or sequence pair from a set of lineages; and g) testing an antibody comprising the selected heavy chain / light chain sequence pair or heavy chain VHH 2 , VHH 3 sequence to determine whether the antibody binds to the antigen or a portion thereof. In one embodiment, the antigen comprises a plurality of epitopes.

[0018] In one embodiment, the minimum CDR3 distance of a particular CDR3 is 1 or less between groups of CDR3s from lineages, and the minimum CDR3 distance of a particular CDR3 is the shortest Hamming distance of this CDR3 compared to all other CDR3s of the same length.

[0019] In one embodiment, the ranking of the lineage pairs in step e) is based on the lineage priority factor of the VH1 lineage of the lineage pair.

[0020] In some embodiments, the lineage priority factor is selected from the group consisting of lineages from high sequence abundance to low sequence abundance, enrichment and proliferation of B cells in vitro, lineages from high amplification factor to low amplification factor later, changes in the abundance of lineage sequences during the immune process, changes in the abundance of lineage sequences before and after depletion of specific unwanted B cells, lineages sharing the same naive B cell origin between VHH and VH, avoidance of developability liability sequences, and combinations thereof.

[0021] In some embodiments, the anchor of the IgG1 repertoire is generated by single B cell sorting and the heterohybridoma approach.

[0022] In some embodiments, the test antibody is expressed by prokaryotic or eukaryotic cells.

[0023] In some embodiments, one representative sequence or one representative pair VH of VHHs from each of the top 100 lineages or lineage pairs 1 / VL 1is selected. In some embodiments, 100 lineages / lineage pairs are 95, 90, 85, 80, 75, 70, 65, 60, 55, 50, 45, 40, 35, 30, 25, 20, 15, or 10 VHH lineages and 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, or 95 VH 1 / VL 1 each containing a lineage pair, and VL 1 includes Vκ and Vλ.

[0024] In some embodiments, the method further comprises repeating f) - g) to generate camelid antibodies, and representative sequences are selected from the top 101 - 200, 201 - 300, 301 - 400, 401 - 500, 501 - 600, 601 - 700, 701 - 800, 801 - 900, 901 - 1000, 1001 - 1100, 1101 - 1200, 1201 - 1300, 1301 - 1400, 1401 - 1500, 1501 - 1600, 1601 - 1700, 1701 - 1800, 1801 - 1900, or 1901 - 2000 ranked lineages. In some embodiments, the method further comprises repeating f) - g) to generate camelid antibodies, and representative sequences are selected from the top 2000 - 10,000 ranked lineages.

[0025] In some embodiments, the criteria for ranking / selecting lineages are lineages from high sequence abundance to low sequence abundance, lineages from high amplification factor to low amplification factor, changes in lineage sequence abundance during the immune process, changes in lineage sequence abundance before and after depletion of specific unwanted B cells, lineages sharing the same naive B cell origin between VHH and VH 1 and those sharing the same naive B cell origin between VHH and VH, avoidance of liability sequences for development, and combinations thereof.

[0026] In some embodiments, the antigen or immunogen can be a cell or tissue, and the generated VHH is used to identify individual corresponding antigens by protein array or cell / tissue antigen cDNA library or mass spectrometry based on immunoprecipitation.

[0027] In another aspect of the invention, a method for generating camelid antibodies against multiple epitopes of a specific antigen in high throughput comprises: i) FR2 hydrophilic region; ii) extended CDR1; iii) additional disulfide bonds within CDR1 and / or CDR3; iv) long CDR3 (≧15aa); v) VHH 2 VHH 3 and VH 1 sequences sharing the same naive B cell origin; vi) sequence-based prediction of antigen-binding loop structures; x) convergent motifs or sequence signatures between individuals of the same immunological group; xi) length of CDR2; xii) length of CDR3; xiii) length and identity of CDR3; xiv) presence of three or more positive charges within the CDR3 region; xv) number of cysteines in the amino acid sequence; and xvi) further comprising subgrouping VHH lineages having characteristics selected from the group consisting of 2-4 amino acid motifs found in the CDR regions. The motif is identified from the three-dimensional structure of the ligand / receptor complex.

[0028] In some embodiments, sequences within the same lineage group of antibodies tested in the first round can be selected for antibody optimization by repeating f)-g) in the second round.

[0029] In some embodiments, a method for generating camelid antibodies against an antigen in high throughput further comprises applying the selected VHH sequence to guide the selection of VH1-VL1 pairs of these clones sharing the same naive B cell origin, where the selection criteria are: 1) the difference between CDR1 and CDR2; 2) the differences in FR1, 2, 3, 4.

[0030] In another aspect of the invention, a method for generating humanized VHH antibodies comprises: a) enriching and proliferating antigen-specific B cells from immunized camelids; b) VHH from antigen-specific B cells 2 VHH 3 VH 1 and VL 1Generating an antibody NGS library comprising a lock array; and c) VHH 2 , VHH 3 , VH 1 Grouping the NGS sequences by lineage; d) parental VHH 2 antibody, VHH 3 antibody or VH sharing the same naïve B cell origin 1 Identifying the replaceable positions within by comparing the amino acid sequences of a plurality of related antibodies, each of which binds to the same epitope as the parental antibody within the same lineage, with the amino acid sequence thereof; and e) parental VHH 2 antibody or VHH 3 Replacing the amino acid at one or more replaceable positions of the antibody with the amino acid at the corresponding position in a human antibody; f) Testing an antibody comprising the substituted residue within the selected sequence to determine whether the antibody binds to an antigen or a portion thereof.

[0031] In one embodiment, the replacement positions of the humanized VHH antibody are within the FR region. In one embodiment, the replacement positions of the humanized VHH antibody are within the CDR region.

[0032] In one embodiment, the parental antibody is a camelid antibody. In one embodiment, the parental antibody is a humanized camelid antibody.

[0033] In another aspect of the invention, an isolated camelid antibody or antigen-binding portion comprising an antibody sequence generated by the invention.

[0034] In another aspect of the invention, a pharmaceutical composition comprising the camelid antibody of the invention and a pharmaceutically acceptable carrier. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Exemplary embodiments are shown in the referenced figures. It is intended that the embodiments and figures disclosed herein be considered illustrative rather than limiting.

Figure 1A

Figure 1B

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19A

Figure 19B

Figure 20A

Figure 20B

Figure 21A

Figure 21B

Figure 22

Figure 23

Figure 24

Figure 25

Figure 26

Figure 27

Figure 28

Figure 29

Figure 30

Figure 31

Figure 32

Figure 33

Figure 34

Figure 35

Figure 36

Figure 37

Figure 38

Figure 39

Figure 40

Figure 41

Figure 42

[0036] The following embodiments and their aspects are not limiting, but are intended to be illustrative and diagrammatic, and are described and illustrated in conjunction with systems, compositions, and methods. Definitions

[0037] As used herein, the terms "comprising" or "comprise" are used with respect to compositions, methods, and respective components thereof that, while useful in one embodiment, can include unspecified elements whether useful or not. In general, it is understood by those skilled in the art that the terms used herein are generally intended as "open" terms (e.g., the term "including" should be construed as "including but not limited to"; the term "having" should be construed as "having at least"; the term "include" should be construed as "including but not limited to").

[0038] Unless otherwise specified, the terms "a", "an", "the", and similar references used in the context of describing particular embodiments of this application (especially in the claims) can be construed to cover both singular and plural forms. The recitation of a range of values herein is intended merely as a convenient method of referring individually to each separate value within the range. Unless otherwise indicated herein, each separate value is incorporated herein as if it were individually recited herein. All methods described herein can be performed in any suitable order unless otherwise indicated herein or otherwise clearly contradicted by context. The use of any and all examples or exemplary language (e.g., "such as") provided with respect to particular embodiments herein is intended merely to better clarify the use and is not intended to limit the scope of the claimed use in any other way. The abbreviation "e.g." is derived from the Latin "exempli gratia" and is used herein to indicate non-limiting examples. Thus, the abbreviation "e.g." is synonymous with the term "for example". No language in this application should be construed as indicating any non-claimed element as essential to the practice of this application.

[0039] The term "plurality" refers to one or more, for example, two or more, about five or more, about ten or more, about twenty or more, about fifty or more, about one hundred or more, about two hundred or more, about five hundred or more, about one thousand or more, about two thousand or more, about five thousand or more, about ten thousand or more, about twenty thousand or more, and usually about two hundred thousand or less. A "population" includes a plurality of items.

[0040] As used herein, the term "about" refers to measurable values such as amounts, durations, etc., and includes variations of ±20%, ±10%, ±5%, ±1%, ±0.5% or ±0.1% from the specified value.

[0041] As used herein, the term "epitope" can include any protein determinant that can specifically bind to an immunoglobulin or T cell receptor. Epitope determinants are usually composed of surface groups of chemically active molecules such as amino acids or sugar side chains, and usually have specific three-dimensional structural characteristics and specific charge characteristics. An antibody is said to specifically bind to an antigen when the equilibrium dissociation constant is ≦1 μM, preferably ≦100 nM, and most preferably ≦10 nM.

[0042] The term "K D " can refer to the equilibrium dissociation constant of a particular antibody-antigen interaction.

[0043] As used herein, the term "immune response" can refer to the action of, for example, lymphocytes, antigen-presenting cells, phagocytes, granulocytes, and soluble macromolecules (such as antibodies, cytokines, and complement) produced by the above cells or the liver, resulting in selective damage, destruction, or elimination of invading pathogens, pathogen-infected cells or tissues, cancerous cells, or in the case of autoimmunity or pathological inflammation, normal biological cells or tissues from the organism.

[0044] As used herein, "antigen-specific T cell response" can refer to a response by T cells resulting from the stimulation of T cells by an antigen to which the T cells are specific. Non-limiting examples of responses by T cells upon antigen-specific stimulation include proliferation and cytokine production (e.g., IL-2 production).

[0045] As used herein, the term "antibody" refers to an intact immunoglobulin, or a monoclonal or polyclonal antigen-binding fragment having an Fc (crystallizable fragment) region, or an FcRn-binding fragment of the Fc region referred to herein as the "Fc fragment" or "Fc region". Antigen-binding fragments can be produced by recombinant DNA techniques or by enzymatic or chemical cleavage of intact antibodies. Antigen-binding fragments include, among others, Fab, Fab’, F(ab’)2, Fv, dAb, and complementarity-determining region (CDR) fragments, single-chain antibodies (scFv), single-domain antibodies, chimeric antibodies, diabodies, and polypeptides comprising at least a portion of an immunoglobulin sufficient to confer specific antigen binding to the polypeptide. The Fc region includes a portion of two heavy chains that contribute to two or three classes of antibodies. The Fc region can be produced by recombinant DNA techniques or by enzymatic (e.g., papain cleavage) or through chemical cleavage of intact antibodies.

[0046] As used herein, the term "antibody fragment" refers to a protein fragment that includes only a portion of an intact antibody, generally including the antigen-binding site of the intact antibody, and thus retains the ability to bind to an antigen. Examples of antibody fragments included in this definition are: (i) Fab fragments having VL, CL, VH, and CH1 regions; (ii) Fab' fragments that are Fab fragments having one or more cysteine residues at the C-terminus of the CH1 region; (iii) Fd fragments having VH and CH1 regions; (iv) Fd' fragments having VH and CH1 regions and one or more cysteine residues at the C-terminus of the CHI region; (v) Fv fragments having VL and VH regions of one arm of an antibody; (vi) dAb fragments consisting of VH regions (Ward et al., Nature 341, 544-546 (1989)); (vii) isolated CDR regions; (viii) F(ab')2 fragments, bivalent fragments comprising two Fab' fragments linked by disulfide bridges in the hinge region; (ix) single-chain antibody molecules (e.g., single-chain Fv; scFv) (Bird et al., Science 242:423-426 (1988); and Huston et al., PNAS (USA) 85:5879-5883 (1988)); (x) "diabodies" having two antigen-binding sites, comprising a heavy-chain variable region (VH) connected to a light-chain variable region (VL) within the same polypeptide chain (see, e.g., European Patent No. 404,097; International Publication No. 93 / 11161; and Hollinger et al., Proc. Natl. Acad. Sci. USA, 90:6444-6448 (1993)); (xi) "linear antibodies" comprising a pair of tandem Fd segments (VH-CH1-VH-CH1) that together with complementary light-chain polypeptides form a pair of antigen-binding regions (Zapata et al., Protein Eng. 8(10):1057-1062 (1995); and U.S. Patent No. 5,641,870).

[0047] As used herein, the term "single-chain variable fragment", "single-chain antibody variable fragment" or "scFv" antibody refers to an antibody form that contains only the variable regions of the heavy chain (VH) and the light chain (VL) and is connected by a linker peptide. The scFv can be expressed as a single-chain polypeptide. The scFv retains the specificity of the intact antibody from which it is derived. The light and heavy chains may be in any order, for example, VH-linker-VL or VL-linker-VH, as long as the specificity of the scFv for the target antigen is retained.

[0048] As used herein, the term "isolated antibody" can refer to an antibody that is substantially free of other antibodies having different antigen specificities (e.g., an isolated antibody that specifically binds to the TRAIL protein can be substantially free of antibodies that specifically bind to antigens other than the TRAIL protein). However, an isolated antibody that specifically binds to the human TRAIL protein may be cross-reactive with other antigens such as TRAIL proteins from other species. Furthermore, an isolated antibody may not be substantially free of other cellular materials and / or chemical substances.

[0049] As used herein, the term "monoclonal antibody" or "monoclonal antibody composition" can refer to a preparation of antibody molecules of a single molecular composition. A monoclonal antibody composition exhibits a single binding specificity and affinity for a particular epitope.

[0050] As used herein, the term "recombinant human antibody" refers to all human antibodies prepared, expressed, made, or isolated by recombinant means, e.g., antibodies isolated from transgenic or transchromosomal animals (e.g., mice) for human immunoglobulin genes or hybridomas prepared therefrom (described below), (b) antibodies isolated from host cells transformed to express human antibodies (e.g., from transfectomas), (c) antibodies isolated from recombinant combinatorial human antibody libraries, and (d) antibodies prepared, expressed, made, or isolated by any other means including splicing of human immunoglobulin gene sequences to other DNA sequences. Such recombinant human antibodies have variable regions where the framework and CDR regions are derived from human germline immunoglobulin sequences. However, in certain embodiments, such recombinant human antibodies can be subjected to in vitro mutagenesis (or, if transgenic animals for human Ig sequences are used, in vivo somatic mutagenesis), and thus, the amino acid sequences of the VH and VL regions of the recombinant antibody are related to, but may not naturally exist within the human germline repertoire in vivo, sequences derived from human germline VH and VL sequences.

[0051] The term "isotype" can refer to the antibody class (e.g., IgM or IgG1) encoded by the heavy chain constant region gene. Antibodies can be or be derived from immunoglobulin G (IgG), IgM, IgE, IgA, or IgD molecules.

[0052] "VHH 2 ", "VHH 3 " and "VH 1 " refer to the heavy chains of the three camelid IgG isotypes IgG2, IgG3, and IgG1, respectively. VL 1 represents the light chain of camelid IgG1. Camelid VL 1 includes, but is not limited to, Vκ and Vλ.

[0053] As used interchangeably herein, the terms "corresponding amino acids" and "corresponding amino acid" refer to amino acid residues that are in the same position (i.e., opposite one another) when two or more amino acid sequences are aligned. Methods for aligning and numbering antibody sequences are described in great detail in Chothia, Kabat, supra, and others. As is known in the art (see, e.g., Kabat 1991 Sequences of Proteins of Immunological Interest, DHHS, Washington, DC), in order to achieve an alignment, one or both of the antibody amino acids may sometimes be made to have one, two, or three gaps and / or insertions of up to one, two, three, or four residues, or up to about 15 residues (especially in light and heavy chain CDR3).

[0054] The term "native" antibody refers to antibodies in which the heavy and light chains are made by the immune system of a multicellular organism and pair to form the antibody. The spleen, lymph nodes, bone marrow, blood, and other lymphoid tissues are examples of tissues that contain cells that produce native antibodies. For example, an antibody produced by B cells isolated from a first animal immunized with an antigen is a native antibody. Native antibodies contain heavy and light chains that pair naturally.

[0055] The term "naturally paired" refers to heavy and light chain sequences that are paired by the immune system of a multicellular organism.

[0056] As used herein, the term "mixture" refers to a combination of elements, e.g., cells, that are interspersed and not in any particular order. A mixture is homogeneous and not spatially separated into distinct components. Examples of mixtures of elements include multiple different cells that are spatially undressed and present in the same aqueous solution.

[0057] The term "evaluating" includes all forms of measurement and also includes determining whether an element is present or not. The terms "determining", "measuring", "evaluating", "assessing" and "assaying" are used synonymously and can include quantitative and / or qualitative determinations. Evaluating can be relative or absolute. "Evaluating for presence" includes determining the amount of what is present and / or determining whether it is present or absent.

[0058] The term "enriched" is intended to refer to a component of a composition (e.g., a particular type of cell) that is more concentrated (e.g., at least 2-fold, at least 5-fold, at least 10-fold, at least 50-fold, at least 100-fold, at least 500-fold, at least 1,000-fold) compared to other components (e.g., other cells) in the sample than before enrichment. In some cases, the enriched material can represent a substantial proportion (e.g., greater than 2%, greater than 5%, greater than 10%, greater than 20%, greater than 50%, or more, usually about 90% - 100% or less) of the sample in which it is present.

[0059] The term "enriching" is intended to mean any method by which antigen-specific cells can be obtained from a larger B cell population. As described in more detail below, enriching can be accomplished, for example, by panning using beads or cell sorting.

[0060] The term "obtaining" in the context of obtaining an element (e.g., a cell or a sequence) is intended to include receiving the element as well as physically producing the element.

[0061] The term "peripheral blood mononuclear cell" or "PBMC" refers to blood cells that have a single, mostly circular nucleus (as opposed to a segmented nucleus) and include lymphocytes (T cells, B cells, and NK cells), monocytes, and macrophages. PBMCs can be enriched from whole blood using a Ficoll gradient.

[0062] The term "antigen-specific B cell" refers to memory B cells that have antibodies that specifically bind to an antigen on their surface, and their progenitor cells.

[0063] When a cell or its progeny is obtained from a host, the cell "derives" from the host. The progeny of a progenitor cell are derived from the progenitor cell.

[0064] The term "support bearing an antigen" includes any type of support (e.g., solid or semi-solid supports such as plates and beads) that contains an antigen or a portion thereof immobilized thereon. The antigen can be immobilized on the support directly or indirectly, e.g., via a linker, via a biotin-streptavidin interaction, or e.g., via a cell. Such supports are utilized in methods for enriching antigen-specific B cells by panning or using beads.

[0065] The term "panning" is used to refer to a method in which B cells are applied to a container (e.g., a plate) having one or more surfaces coated with an antigen or a portion thereof. Unbound cells can be removed by washing the surface after the cells are applied.

[0066] The term "bead-based enrichment" is used to refer to a method in which B cells are mixed with beads conjugated to an antigen or a portion thereof, e.g., magnetic beads.

[0067] The term "cell sorting" is used to refer to a method in which B cells are mixed with an antigen detectable in solution (e.g., an antigen detectable fluorescently). In cell sorting methods, cells bound to the antigen are sorted from unbound cells. Fluorescence-activated cell sorting (FACS) is an example of a cell sorting method.

[0068] The term "complex immunogen" is intended to refer to an immunogen containing multiple antigens. A complex immunogen can be composed of multiple different antigens made separately and then mixed together, or they can be naturally complexed during immunization (e.g., as when using whole cells and tissues or a fraction thereof).

[0069] The term "activate" refers to the stimulation of B cells to a) proliferate, b) differentiate into blasts and / or plasma cells, and c) secrete antibodies. Activation of B cells can be done by contacting the B cells with an antigen, T cells expressing CD40L, and cytokines, although other methods are known (see, for example, Wykes, Imm. Cell. Biol. 2003 81:328-331).

[0070] The term "activated B cell" refers to a cell population including the progeny of activated B cells. As noted above, activation causes B cells to proliferate, and the progeny of such cells are referred to herein as activated B cells.

[0071] The term "collect" refers to the act of separating cells in a culture medium from a substrate. Collection can be done, for example, by pipetting or decantation.

[0072] The term "immunized with an antigen" and its grammatical equivalents (e.g., "immunized animal") are intended to refer to any animal (human, rabbit, mouse, rat, sheep, cow, chicken, camel) in which an immune response to the antigen has been initiated. The animal can be exposed to a foreign antigen, for example, via exposure to an infectious agent, vaccination, or administration of an antigen and an adjuvant (e.g., by injection). The term "immunized with an antigen" is also intended to include animals in which an immune response to a "self" antigen has been initiated, i.e., animals having an autoimmune disease.

[0073] The terms "ranking" and "ranked abundance order" refer to the order of arrangement when they are enumerated by abundance. That is, the most abundant sequence first, the second most abundant sequence second, the third most abundant sequence third, and so on. In some cases, the sequences can be ranked by creating a frequency distribution and then sorting the sequences by their frequencies.

[0074] The term "corresponding rank" or "correspondingly ranked" refers to two sequences having the same position in two ranks. For example, the first, second, and third positions in the first rank correspond to the first, second, and third positions in the second rank, respectively.

[0075] The term "lineage rank" refers to the order of lineages enumerated by a priority factor. The priority factors include, but are not limited to, the abundance of the lineage sequence, the amplification factor, the dynamic change of the lineage sequence before and after depletion of specific unwanted B cells, the dynamic change of the lineage sequence abundance during the immune process, lineages sharing the same naive B cell origin between VHH and VH, avoidance of developability liability sequences, and combinations thereof.

[0076] The term "Hamming distance" refers to the number of positions at which corresponding symbols differ between two sequences of the same length.

[0077] As used herein, the terms "lineage-grouped antibodies", "lineage-related antibodies", and "lineage-associated antibodies", and their grammatical equivalent variants, are antibodies produced by cells sharing a common B cell ancestor. Lineage-associated antibodies bind to the same epitope of an antigen and typically have very similar sequences, particularly in the CDR3 of the light and heavy chains. The CDR3s of both the heavy and light chains of lineage-related antibodies can have the same length and nearly the same sequence (i.e., differ by at most 5, i.e., 0, 1, 2, 3, 4, or 5 residues). Among the group of CDR3s from a lineage, the minimum CDR3 distance of a particular CDR3 is the shortest Hamming distance of this CDR3 compared to all other CDR3s of the same length. In some embodiments, the minimum CDR3 distance is 1 or less. In some cases, the B cell ancestor contains a genome with a rearranged light chain VIC region and a rearranged heavy chain VDJ region and produces antibodies that have not undergone affinity maturation. "Naive" or "virgin" B cells present within spleen tissue are exemplary B cell common ancestors.

[0078] Related antibodies are related through antibodies produced within a common antibody ancestor, e.g., the ancestor of naive B cells. The term "lineage-related antibodies" is intended to describe a group of antibodies produced by cells arising from the same ancestral B cell. A "lineage group" includes a group of antibodies that are related to one another by lineage.

[0079] As used herein, the term "at least CDR3" or "at least CDR3 sequence" refers to the CDR3 sequence alone, the CDR3 sequence in combination with the CDR1 and / or CDR2 sequences, or the sequence of at least 50 contiguous amino acids (up to the full length of the variable domain) of the variable domain. Here, the sequence includes the CDR3 sequence.

[0080] As used herein, the term "phylogenetic tree" refers to a diagram obtained from a phylogenetic analysis depicting a virtual branching array of lineages leading to the individual species of interest. The branching points within a phylogenetic tree are called nodes.

[0081] As used herein, the term "constructing a phylogenetic tree" refers to the computational act of creating a phylogenetic tree from an array.

[0082] As used herein, the term "lineage" refers to a theoretical descent line. "Lineage" is used synonymously with "group", and a group of antibodies related by lineage may be referred to as a "lineage group". The terms "group" or "lineage" are exclusive in that an array can belong to only one group or lineage.

[0083] As used herein, the term "subgrouping" refers to the further grouping of arrays within a lineage based on unique features or signatures. A "subgroup" is not exclusive, and an array can be included in different subgroups. For example, an array can have two, three, four, five, or six unique features simultaneously. "Subgrouping" is only for VHHs. Applying the VHH array signature can assist in selecting / narrowing down lineages (representative arrays) for testing in a better way. This can result in better biological function / bioactivity results.

[0084] As used herein, the term "lineage analysis" refers to analyzing the theoretical descent line of an antibody, which is typically done by analyzing a phylogenetic tree.

[0085] As used herein, the term "sequence read" refers to the sequence of nucleotides determined by a sequencer, and the determination is made, for example, by base calling software related to the technology.

[0086] As used herein, the term "obtaining an amino acid sequence" refers to obtaining a file containing the amino acid sequence. As is well known, a nucleic acid sequence can be translated into an amino acid sequence in silico.

[0087] As used herein, the term "most abundantly expressed" with respect to a protein sequence means being the most abundant in a sample. The abundance of a protein can be determined, for example, by counting the sequence reads encoding that protein. The protein encoded by the most sequence reads is the most abundant protein.

[0088] The terms "anchor" and "anchor binder" as used interchangeably herein refer to conventional antibodies produced by single B cell sorting or heterohybridomas having natural H and L pairing, such that after encountering an epitope of an antigen, the heavy and light chain lineages consisting of groups of sequences derived from clonal expansion of naive B cell H and L sequences can be positioned / paired. The lineages can be "anchored" considering the amino acid sequences of the heavy and light chains known to pair with each other. In these embodiments, the branches rotate around their nodes until a minimum number of crossovers (e.g., no crossovers) exist between the anchored sequences. After "aligning" the tree by tanglegram analysis, the leaves known to pair can be connected by edges. When the leaves known to pair are connected by edges, in theory, intervening leaves can pair with each other as long as they do not create edge or crossover events with each other.

[0089] The phrases "monoclonal antibody that recognizes an epitope on an antigen", "antibody that recognizes an antigen", and "antibody specific for an antigen" are used interchangeably herein with the term "antibody that specifically binds to an antigen".

[0090] The term "specific binding" refers to the ability of an antibody to preferentially bind to a particular antigen present in a homogeneous mixture of different molecules. In certain embodiments, the specific binding interaction discriminates between the desired and undesired molecules in a sample, and in some embodiments, exceeds about 10 to 100-fold, or for example, exceeds about 1000-fold or 10,000-fold.

[0091] As used herein, the term "not substantially bind to" a protein or cell means that it cannot bind to or does not bind to the protein or cell with high affinity. That is, KD is 2x10 -6 M or more, more preferably 1x10 -5 M or more, more preferably 1x10 -4 M or more, more preferably 1x10 -3 M or more, even more preferably 1x10 -2 M or more and binds to a protein or cell having the same.

[0092] The term "high affinity" of an IgG antibody refers to an antibody having a KD of 1x10 -6 M or less, preferably 1x10 -7 M or less, more preferably 1x10 -8 M or less, even more preferably 1x10 -9 M or less, or even more preferably 1x10 -10 M or less, for the target antigen. However, "high affinity" binding may vary for other antibody isotypes.

[0093] The term "pharmaceutical preparation" refers to a preparation in a form such that the biological activity of the active ingredient contained therein is effective and does not contain any additional ingredients that are toxic and unacceptable to the subject to which the preparation is administered.

[0094] An "effective amount" of an agent, such as a pharmaceutical preparation or cell, refers to an amount effective at the dosage and for the period required to achieve the desired therapeutic result, such as the treatment of a disease, condition or disorder, and / or the pharmacokinetic or pharmacodynamic effect of the treatment. The effective amount may vary depending on factors such as the condition of the subject, age, gender, and weight, and the population of cells administered. In some embodiments, the provided methods include administering cells and / or compositions in an effective amount, such as a therapeutically effective amount.

[0095] A "CDR-grafted antibody" is an antibody that includes one or more CDRs derived from an antibody of a particular species or isotype and a framework of another antibody of the same or a different species or isotype.

[0096] A "humanized antibody" has a sequence that is different from the sequence of an antibody derived from a non-human species by one or more amino acid substitutions, deletions, and / or additions such that, when administered to a human subject, the humanized antibody is less likely to induce an immune response and / or induces an immune response of lower severity as compared to the non-human species antibody. In one embodiment, specific amino acids within the framework and constant regions of the heavy and / or light chains of the non-human species antibody are mutated to produce the humanized antibody. In another embodiment, the constant region(s) from a human antibody is fused to the variable region(s) of the non-human species. In another embodiment, the humanized antibody is a CDR-grafted antibody that includes one or more CDRs derived from an antibody of a particular species or isotype and a framework of a human antibody. In another embodiment, one or more amino acid residues in one or more CDR sequences of the non-human antibody are altered to reduce the potential immunogenicity of the non-human antibody when administered to a human subject. Here, the altered amino acid residues are not critical for the immunospecific binding of the antibody to its antigen or the amino acid sequence changes made are conservative changes such that the binding of the humanized antibody to the antigen is not significantly worse than the binding of the non-human antibody to the antigen. Examples of methods for making humanized antibodies can be found in U.S. Patent Nos. 6,054,297, 5,886,152, and 5,877,293.

[0097] The term "chimeric antibody" refers to an antibody that includes one or more regions from one antibody and one or more regions from one or more other antibodies. In one embodiment, one or more of the CDRs are derived from a human antibody. In another embodiment, all of the CDRs are derived from a human antibody. In another embodiment, CDRs from two or more human antibodies are mixed and coincide in the chimeric antibody. For example, a chimeric antibody can include CDR1 from the light chain of a first human antibody, CDR2 and CDR3 from the light chain of a second human antibody, and CDRs from the heavy chain of a third antibody. Other combinations are possible.

[0098] The term "biparatopic antibody" refers to an antibody that binds to two non-overlapping epitopes of an antigen. In some embodiments, the biparatopic antibody includes only the VHH of the heavy chain without the light chain. In some embodiments, the biparatopic antibody includes only the VHH of the heavy chain and a conventional VH 1 / VL 1 pair of both. In some embodiments, the biparatopic antibody includes two conventional VH 1 / VL 1 pairs. In some embodiments, the biparatopic antibody has a first heavy chain and a first light chain from a monoclonal antibody that targets one epitope, and additional antibody heavy and light chains that target another epitope. In some embodiments, the additional light or heavy chain can be different from the first light or heavy chain.

[0099] The binding of the disclosed invention's antibodies to an antigen can be evaluated using one or more techniques well established in the art. For example, in a preferred embodiment, the antibody can be tested by an ELISA assay using, for example, a recombinant antigen protein. Further suitable binding assays include, but are not limited to, flow cytometry assays in which the antibody reacts with a cell line expressing a human antigen such as HEK293 cells. Additionally or alternatively, the binding of the antibody, such as the binding kinetics (e.g., KD value), can be tested with a BIAcore binding assay, Octet Red96 (Pall), etc.

[0100] The term "single B cell sorting" refers to sorting isolated and separated single B cells based on antigen specificity. Techniques for single cell separation, isolation, and sorting include, but are not limited to, FACS (e.g., fluorescence-activated cell sorting that isolates cells binding to an antigen using a fluorescently tagged antigen), ISAAC (immunospot array assay on a chip), LCM (laser capture microdissection), microengraving, and droplet microfluidics.

[0101] A method for generating only the heavy chain antibody of a camelid or its binding portion for recognizing an antigen, particularly for therapeutic use, the method comprising: a) concentrating and amplifying antigen-specific B cells from an immunized camelid; b) generating an antibody NGS library containing VHH 2 VHH 3 VH 1 and VL 1 chain sequences from the antigen-specific B cells; c) grouping VHH 2 VHH 3 and VH 1 sequence data as phylogenetic clonotypes based on CDR3 - for example, grouping CDR3 amino acid sequences that differ by 0 or 1 amino acid and have the same length; d) ranking the lineages containing the VHH heavy chain (VHH 2 VHH 3 ) by a lineage preference factor; e) selecting representative sequences from the lineages of VHH 2 VHH 3 and ranking them at the top in the antibody sequence library according to the lineage preference factor; f) testing the antibody containing the selected representative sequence to determine whether the antibody binds to the antigen or a part thereof. In one embodiment, the antigen comprises a plurality of epitopes.

[0102] In one embodiment, it can be tested by selecting sequences from multiple lineages and repeating steps f) and g). In one embodiment, f) may include: 1) synthesizing the DNA of the selected representative sequence; 2) constructing a vector containing the DNA sequence; 3) expressing the vector in cells; and 4) performing affinity and bioactivity tests against specific antigens. In one embodiment, b) may include: 1) preparing cDNA from a concentrated population of antigen-specific B cells; 2) sequencing the cDNA to obtain multiple VHH 2 、VHH 3 、VH 1 heavy chain sequences, and multiple VL 1 (Vκ and Vλ) light chain sequences, and generating camelid IgG2 (HcAb), IgG3 (HcAb), and IgG1 (conventional Ab) libraries. In one embodiment, the generated camelid antibodies include IgG2. In one embodiment, the generated camelid antibodies include IgG3.

[0103] Another aspect of the present invention is a method for generating a camelid antibody or a binding portion thereof for recognizing an antigen, particularly for therapeutic use. The method includes: a) concentrating and amplifying antigen-specific B cells of an immunized camelid; b) generating an antibody NGS library containing VHH 2 、VHH 3 、VH 1 and VL 1 chain sequences from the antigen-specific B cells; c) grouping the VHH 2 、VHH 3 、VH 1 、and VL 1 sequence data as phylogenetic chronotypes based on CDR3; d) pairing the IgG1 VH 1 / VL 1 lineages according to the anchor binders generated by single B cell sorting and the hetero-hybridoma approach; e) ranking the lineages and lineage pairs of steps c) and d) by lineage preference factors; f) the top-ranked VHH 2 、VHH 3lineage and VH 1 / VL 1 selecting a representative sequence from a pair of lineages of, and g) the selected VH 1 / VL 1 pair sequence or VHH 2 , VHH 3 testing an antibody comprising the sequence of, to determine whether the antibody binds to the antigen or a part thereof. In one embodiment, the VL1 chain comprises Vκ and Vλ. In one embodiment, the antigen comprises a plurality of epitopes.

[0104] In one embodiment, it can be tested by selecting sequences from multiple lineages and repeating steps f) and g). In one embodiment, f) may include: 1) synthesizing DNA encoding the selected representative amino acid sequence; 2) constructing a vector containing the DNA sequence; 3) expressing the vector in cells; 4) performing affinity and bioactivity tests against a specific antigen. In one embodiment, b) may include: 1) preparing cDNA from an enriched population of antigen-specific B cells; 2) sequencing the cDNA to obtain multiple VHH 2 , VHH 3 , VH 1 heavy chain sequences, and multiple VL 1 (Vκ and Vλ) light chain sequences to generate camelid IgG2 (HcAb), IgG3 (HcAb), and IgG1 (conventional Ab) libraries. In one embodiment, the generated camelid antibodies include IgG2. In one embodiment, the generated camelid antibodies include IgG3. In one embodiment, the generated camelid antibodies include conventional IgG1. In one embodiment, the ranking of the lineage pairs in step e) is based on the lineage preference factor of the VH 1 lineage of the lineage pair. In one embodiment, the ranking of the lineage pairs in step e) is based on the lineage preference factor of the VL 1 lineage of the lineage pair. This method may further include a test to determine whether the antibody that binds to the antigen inhibits the binding of the antigen to another protein, for example, whether the antibody binding inhibits the specific binding of a ligand to its cognate receptor.

[0105] A method for generating a humanized camelid antibody targeting an antigen comprises: a) enriching and proliferating antigen-specific B cells from an immunized camelid; and b) generating an antibody NGS library containing VHH 2 , VHH 3 , and VH 1 chain sequences; c) grouping the sequences of VHH 2 , VHH 3 , and VH 1 by lineage; d) identifying replaceable positions within a parental VHH antibody or VH 1 sharing the same naive B cell origin by comparing the amino acid sequences of a plurality of related antibodies, each of which binds to the same epitope as the parental antibody within the same lineage; e) substituting the amino acids at one or more replaceable positions of the parental VHH antibody or VH1 antibody with the amino acids at the corresponding positions in a human antibody; and f) testing the antibody containing the substituted residues within the selected sequences to determine whether the antibody binds to the antigen or a portion thereof. In one embodiment, the antigen comprises a plurality of epitopes. In one embodiment, the replaceable positions are within the CDR regions. In one embodiment, the replaceable positions are within the FR regions.

[0106] The conventional IgG1 antibodies of camelids generated in the present invention can be humanized by substituting the amino acids at one or more replaceable positions of the parental IgG1 antibody with the amino acids at the corresponding positions in a human antibody.

[0107] Embodiments of a method for the immunization of Camelidae animals and the isolation / expansion of antigen-specific B cells are schematically shown in FIG. 2. Such embodiments include 1) immunizing Camelidae with DNA, or immunizing small molecules with carrier proteins, or immunizing peptides with antigen conjugates such as carrier proteins or proteins or cells or tissues, 2) separately monitoring the immune responses of IgG2, IgG3, and IgG1, 3) obtaining a sample of cells containing B cells from PBMC, spleen, lymph nodes, lymphoid tissue, where B cells refer to memory cells, plasmablasts, and different stages of B cells having cell membrane 1gG2 (HcAb), IG3 (HcAb), and IgG1 (conventional IgG), 4) enriching antigen-specific B cells with cell surface antibodies via either physical surface antigen panning, or magnetic bead isolation or flow sorting, and 5) activating the enriched B cells for cell expansion in the presence of antigen, Camelidae CD40-L expressing cells, and growth factors. The activation step selectively stimulates memory B cells to differentiate into plasma cells. These plasma cells divide rapidly and express large amounts of antibody. In some embodiments, the immune responses of IgG2, IgG3, and IgG1 in the antiserum are monitored by a) purifying IgG2, IgG3, and IgG1 using Protein A and Protein G columns under different pH elution conditions, b) analyzing the immune response titers of IgG2, 3, and IgG1, or c) testing the biological activities of IgG2, 3, and IgG1 in a desired immunoassay.

[0108] The activation step of this method of growing only B cells having surface-bound antibodies actually bound to an antigen has three effects: (1) By the activation step, only B cells specifically binding to the antigen are grown, thereby increasing the relative concentration of these cells compared to cells non-specifically bound to the support. (2) By the activation step of this method, expression of HcAb or conventional IgG heavy and light chain mRNAs is induced only within B cells specifically binding to the antigen. (3) It is recognized that these “rare” antigen-specific B cells, which express antibodies having high affinity or specificity, or “rare” epitopes are amplified and the signal-to-noise ratio is significantly improved.

[0109] In some embodiments, antigens used for enrichment include, but are not limited to, the following: a) An immunogen; b) A desired domain / epitope of the immunogen; c) A complex immunogen: An animal can be immunized with multiple antigens or cells or tissues, or biological fluids, and antigen-specific B cells for each of the multiple antigens can be enriched separately or as a whole. Next, the antigen-specific B cells can be activated and collected separately or as a whole. The simplicity of VHH provides advantages for high-throughput VHH cloning and expression, and it is relatively easy to identify each corresponding antigen of VHH by deconvoluting the complex immunogen. The complex immunogen can be deconvoluted by methods such as, but not limited to, protein arrays or mass spectrometry based on immunoprecipitation or cell, tissue antigen-cDNA library screening methods.

[0110] Sometimes, it may be desirable to deplete unwanted B cells by panning prior to antigen-specific panning to improve the purity of B cells.

[0111] VHH 2 、VHH 3 、VH 1、Vκ and Vλ NGS libraries can be prepared from B cells expressing IgG2 (HcAb), IgG3 (HcAb), and IgG1 (conventional Ab), respectively.

[0112] IgG1 (conventional Ab), IgG2 (HcAb), and IG3 (HcAb) of camelids each have a unique gene constitution, allowing the design of specific primer sets to individually amplify cDNA (e.g., Figure 3). cDNA amplification can be performed in the following steps: a) Total RNA extraction (TRIOL) and purification (RNeasy kit); b) Quantification of RNA and optional storage at -20 °C. c) mRNA capture / RT-PCR using isotype-specific primer sets (VHH 2 、VHH 3 、VH 1 、Vκ and Vλ).

[0113] VHH 2 、VHH 3 、VH 1 、Vκ and Vλ NGS libraries can be prepared, for example, by Nextera Library via PCR with the addition of NGS adapters and library indexing.

[0114] The cDNA of the NGS library can be sequenced by high-throughput sequencing of the library using, for example, an Illumina MiSeq300x2 instrument.

[0115] The sequences can be structured by bioinformatics processes: quality assessment using NGSQCTookit, assembly of R1 / R2 reads, translation, and identification of CDR1, 2, and 3.

[0116] VHH 2 、VHH 3 、VH 1, Vκ and Vλ antibodies can be grouped, for example, using NGS data of CDR3 amino acid sequences, and phylogenetic lineages can be constructed.

[0117] A lineage is defined by a group of sequences derived from the same naive B cell (same V and J arrangements), and a lineage can be defined as a group having an amino acid sequence difference of 1 amino acid or less within the CDR3 region (Hamming distance of 1 or less, or the same CDR3 sequence if the total number of amino acids is within 5aa). The amount of lineages is presumed to reflect the amount of naive B cells in the library and the number of epitopes recognized by these antibodies (Figure 4).

[0118] Generally, lineage size correlates with antibody maturation and clonal expansion. Bioinformatics techniques enable the structuring and visualization of data for a rational approach to candidate antibody selection. For each NGS library, up to 10,000 lineages can be identified by structured sequences through a bioinformatics process including QC using NGSQCTookit, assembly of R1 / R2 reads, translation, identification of CDR1, 2, 3, and lineage grouping based on CDR3 similarity.

[0119] VHH (VHH 2 , VHH 3 ) sequences can be further grouped (subgrouped) by unique sequence signatures.

[0120] Camelids have evolved multiple mechanisms to further diversify their VHH B cell repertoires and expand antigen-binding capabilities. The sequence "signatures" resulting from these mechanisms allow for further grouping of lineages. Such additional criteria for further grouping of lineages reflect the subtly different recognition of antibodies and assist in identifying epitopes with unique VHH recognition patterns (Figure 5). Examples of VHH signatures include, but are not limited to: i) FR2 hydrophilic region: In most VHH antibodies, FR2 has amino acid substitutions characteristic of conventional IgG: 37Phe / Tyr, 44Glu, 45Arg, and 47Gly. ii) Extended CDR1: Many VHHs have an additional hypervariable region (residues 27 - 30 according to Kabat numbering) adjacent to CDR1. VHHs use this region together with the long CDR3 to increase the surface area interacting with the antigen; iii) Additional disulfide bonds between CDR1 - CDR3 or FR2 - CDR3: In camels and dromedaries, 82% of VHHs have a disulfide bond between CDR1 - CDR3, and in llamas and alpacas, 74% of VHHs have a disulfide bond between FR2 - CDR3. iv) Long CDR3: In the case of these VHHs with a long CDR3 (≥15 aa), additional disulfide bonds are often seen. v) Additional disulfide bonds within CDR3: Approximately 5 - 10% of VHHs have an additional disulfide bond within CDR3. This may exhibit more conformational recognition patterns. vi) Additional disulfide bonds within CDR1. vii) Non - classical VHHs having the same V and J germline segments as conventional IgG: A VHH lineage group that shares the same naive B - cell origin (same V and J configuration) as conventional IgG1. This shows the same or similar epitope recognition by both VHH and IgG1. viii) Non - classical VHHs with unique sequence signatures: For example, conserved Trp118 replaced by Arg118 and / or a lower hydrophobicity profile within FR3. ix) Novel canonical binding loop structures: Hypermutation hotspots present at important sites for determining the canonical loop structure create interesting possibilities for diversifying the VHH structural repertoire. Crystallographic studies often emphasize that the CDR1 and CDR2 loops of camel VHHs deviate from the known canonical structures of conventional VH1. Sequence - based prediction of novel Ag - binding loop conformations also supports further grouping of lineages. x) Length of CDR2 xi) Length of CDR3 xii) Presence of three or more positive charges in the CDR3 region xiii) Number of cysteines in the amino acid sequence. xiv) 2 - 4 amino acid motifs found in the CDR regions. The motifs are identified from the three - dimensional structure of the ligand / receptor complex. xv) Length and identity of CDR3 and xvi) Convergent motifs or sequence signatures among camels of the same immunological group.

[0121] VHH (VHH 2 、VHH 3 )、VH 1 -Vκ, and VH 1 -Vλ humanization can be guided by phylogenetic analysis.

[0122] The present invention provides a method for identifying positions in an antibody that can be modified without significantly reducing the binding activity of the antibody. In some embodiments, the method includes identifying replaceable positions in a parental antibody by comparing its amino acid sequence with the amino acid sequences of a plurality of related antibodies that bind to the same antigen and epitope as the parental antibody within the same lineage.

[0123] In some embodiments, the amino acids at the replaceable positions can be replaced with different amino acids without significantly affecting the activity of the antibody. Using the subject method, the amino acid sequence of the CDR can be altered without significantly reducing the affinity of the antibody.

[0124] In humanization methods, or other antibody engineering methods, the present invention finds use in various therapeutic and diagnostic applications.

[0125] Bispecific / bioparatopic antibodies or antigen-binding fragments can be produced by various methods such as fusion of hybridomas or ligation of Fab’ fragments. See, for example, Songsivilai & Lachmarm, Clin. Exp. Immunol. 79:315-321 (1990), Kostelny et al., J. Immunol. 148:1547-1553 (1992). Further, bispecific antibodies can be formed as “diabodies” or “Janusins”. Multiple VHH variable domains can also be connected by linkers to form divalent and multivalent antibodies.

[0126] VH 1 and VL 1 (Vκ or Vλ) pairing lineages were found to pair with each other within antibodies secreted by heterohybridomas and / or flow-sorted single B cells VH 1 -VL 1 can be identified by considering the “anchor” amino acid sequences (shown in Figure 6).

[0127] The challenge of conventional IgG1 development in camels using NGS technology is the method of identifying the original native H and L pairs. Typically, two approaches are used to establish the anchor H and L (κ and λ) lineages: (1) heterohybridomas and (2) single B cell sorting. After the lineage pairs are further grouped by these anchors, representative sequence H / L pairs from each lineage pair are selected for DNA synthesis, binding screening, and bioactivity testing as VHH antibodies.

[0128] I. Heterohybridoma method Isolate lymphocytes from PBMC or isolate spleen or lymphocytes from immunized camels; fuse the lymphocytes with a mouse myeloma fusion partner cell line such as SP / 20 to generate heterohybridomas; screen the supernatants of the heterohybridomas using ELISA and bioactivity assays; VH from the selected heterohybridomas 1 and VL1 perform array determination; these pairs of VH and VL are used as anchors for pairing VH 1 lineages and VL 1 lineages from the IgG1 NGS library.

[0129] II. Single B cell NGS method a) Recover antigen-specific B cells after panning and proliferation; b) Perform single B cell sorting; c) Amplify amplicons by VH 1 -VL 1 linkage PCR; d) Use NGS to identify VH 1 -VL 1 pairs (VH 1 -Vκ or VH 1 -Vλ) as anchors through bioinformatics.

[0130] This method can capture the entire antigen-specific B cell repertoire from immunity, such as conventional IgG1 using HcAb and NGS, and utilizes the simplicity of HcAb: a) IgG2 / HcAb, IgG3 / HcAb, IgG1 / κ and IgG1 / λ b) Lineages with different CDR3_different epitopes c) VHH lineages with sequence signatures_different epitopes d) Pairing of VH-Vκ or VH-Vλ lineages with VH-VL anchors generated from hetero-hybridomas

[0131] Lineages of antibodies that recognize epitopes in a broad spectrum of antigens can be selected (shown in Figure 7).

[0132] Each lineage or lineage pair is estimated to recognize one unique epitope, i.e., one representative sequence (VHH) or one representative pair (VH 1 -VL 1 ) from the top 100 lineages / lineage pairs (for example, 70 sequences in the case of VHH, VH 1 -Vκ and VH 1For -Vλ, it is selected for 30 arrays), gene synthesis, binder screening, and bioactivity testing. The strain selection criteria (priority factors) include, but are not limited to, the following: 1) Strains from high-abundance sequences to low-abundance sequences: The total number of unique cDNAs (sequence abundance) from each strain ranges from 2 to 50,000, and the strain with the most abundant sequence may show the broadest clone expansion after antigen stimulation. 2) Strains from high amplification factors to low amplification factors, dynamic changes in sequence abundance before and after enrichment / proliferation of B cells (the amplification factor can range from 5 to 1,000). 3) Changes in the abundance of strain sequences during the immune process. This indicates the enrichment of antigen-specific sequences and antibody affinity maturation (the change in sequence abundance / number of unique cDNAs can range from 2 to 1,000). 4) When applicable, changes in strain sequence abundance before and after depleting specific unwanted B cells (the change in sequence abundance / number of unique cDNAs can range from 2 to 1,000). 5) VHH and VH 1 Strains that share the same naive B cell origin (same V and J arrangements) between and. 6) Avoidance of sequences responsible for development potential: There are sequences consisting of several amino acids that can cause development potential problems, such as thermal stability (hydrophobic core, charge cluster residues, etc.), chemical stability (deamidation and isomerization), solubility (surface hydrophobicity, etc.), and heterogeneity (glycosylation) (Tomoyuki Igawa et al., mAbs, 2011). It is necessary to avoid selecting these strains or strain pairs.

[0133] The selection criteria may be a combination of the above priority factors.

[0134] The selected strain sequences or sequence pairs are used for DNA synthesis and constructed into expression vectors such as VHH, scFv, Fab, HcAb, camelid IgG1, and human Fc chimeras.

[0135] In some cases, the selected VHH strain and the selected VH 1 / VL 1 Both pairs of the system may share the same naive B cell origin.

[0136] In another aspect, more pairs (e.g., 70 sequences for VHH, 30 sequence pairs for VH 1 -Vκ or VH 1 -Vλ) within the same top-ranked pair of lineages in the first selection round are selected for gene synthesis, binder screening, and bioactivity testing. This is because representative sequence pairs of VH 1 -Vκ or VH 1 -Vλ require more combinatorial tests before identifying the optimal pair.

[0137] If the desired results are not obtained with the first 100 antibodies, more sequences and sequence pairs (e.g., 70 sequences for VHH, 30 sequence pairs for VH 1 -Vκ or VH 1 -Vλ) of the next top-ranked 100 lineages are selected for gene synthesis, binder screening, and bioactivity testing.

[0138] Due to the significance of this method, representative sequences can be systematically and relevantly selected from each lineage for testing to comprehensively cover a wide range of epitopes at high resolution. This improves antibody discovery in the following situations. a) Discovery of therapeutic antibodies using a large pool of candidates for the best affinity, specificity, and developability; b) Discovery of companion diagnostic antibodies for parallel and other applications; c) Construction and development of bivalent and multivalent antibodies; d) Discovery of pairs of antibody heavy and light chains; e) Antibodies that bind to the same epitope can be identified by lineage-related sequences (shown in Figure 8).

[0139] While CDR1 and CDR2 are somewhat involved in determining other binding characteristics, it is well recognized that the CDR3 sequence is the main determinant of binding to the epitope. After screening leads (e.g., marked with @ and + in Figure 8) from each lineage, more candidates with different characteristics such as affinity, specificity, functionality, productivity, and developability can be identified, and the most desirable antibodies can be tested and selected. This is because antibodies from the same lineage are supposed to recognize the same or similar epitopes. This step also helps to build a large pool of candidates for further antibody drug development.

[0140] About 10 - 20% of VHH and VH that share the origin of the same naive B cells 1 may exist. The sequences of VHH selected from the first selection round can identify VH that has the same V(D)J arrangement as VHH 1 and can help further test VH-VL pairs for further selection. The selection criteria include 1) Difference in CDR1 and / or CDR2 > 2 aa (amino acids); 2) Difference in FR1 and / or 2 and / or 3 and / or 4 > 2 aa; 3) Sequences that share the same naive B cell origin between VHH and conventional VH; 4) VH sequences that can pair with both Vλ and Vκ, and the like.

[0141] The selected sequences or sequence pairs are used for DNA synthesis and constructed into expression vectors such as VHH, scFv, Fab, HcAb, camelid IgG1, and human Fc chimeras.

[0142] More sequences and sequence pairs can be selected until the best antibodies are identified, and the remaining clones are retained as a pool of further candidates.

[0143] Phylogenetic analysis of VHH (VHH 2 and VHH 3 ), VH 1 -Vκ and VH 1Humanization of -Vλ.

[0144] It is a non-classical VHH gene that shares the same naive cells as the conventional VH, which is useful for (1) subgrouping the VHH lineage, (2) selecting HcAbs and conventional IgGs that recognize the same or similar epitopes, and (3) facilitating the humanization of both HcAbs and conventional IgGs.

[0145] VHH domains typically show high sequence identity with human type 3 VH domains (VH3), which explains their low immunogenicity (Cortez-Retamozo V, Int J Cancer. 98(3):456-62, 2002). Furthermore, the camelid VH 1 , Vλ, and Vκ domains of conventional antibodies also show significant homology with their human counterparts in both sequence and structure (Alex Klarenbeek et al., mAbs 7:4, 693--706; 2015). Figure 9. As is well known, sequences within the same lineage group share the same or similar CDR3 sequences and recognize the same epitope. By functional screening, these amino acids within variable regions such as CDRs and FRs can be identified, and since these amino acids constitute the same biological function even if they differ within the lineage, they are substitutable. Therefore, these resistant positions can be substituted with human germline antibody amino acids, and the substitutable amino acids can even be within the CDR region for better humanization.

[0146] Furthermore, as mentioned above, these non-classical VHHs (except for FR2 hydrophilic amino acids) are derived from the same IGHV3 or IGHV4, D, and J gene loci as the conventional VH 1 and the phylogenetic structure between the VHHs and VH 1 sequences within these groups is similar, and their humanization designs can be further supported through phylogenetic analysis. Pharmaceutical formulation

[0147] In another aspect, the invention provides a composition, e.g., a pharmaceutical composition, comprising one or a combination of the monoclonal antibodies of the invention, or an antigen-binding portion(s) thereof, formulated together with a pharmaceutically acceptable carrier. Such compositions can include one or a combination of antibodies (e.g., two or more different), or an immunoconjugate or bispecific molecule of the invention. For example, the pharmaceutical compositions of the invention can include a combination of antibodies (or immunoconjugates or bispecific antibodies) that bind to different epitopes on a target antigen or have complementary activities. Examples

[0148] Example 1: Identification of a group of VHH antibodies specifically binding to an antigen using B cell isolation and amplification (BIA) / NGS sequence analysis and single B cell method.

[0149] 1A. BIA / NGS Materials and methods BIA

[0150] A1. Construction of CD40L-expressing EL4.IL-2-C The cell line TIB-181 (EL4.IL-2) was obtained from the American Type Culture Collection and stably transfected with a pCMV-6-based vector containing cDNA encoding human CD40L, which resulted in the expression of human CD40L. Stable cell lines were selected and treated with mitomycin as feeder cells.

[0151] A2. Alpaca conditioning medium Non-immunized alpacas were sacrificed and splenocytes were isolated for the preparation of alpaca conditioning medium. Activation medium containing 10% FBS, phytohemagglutinin (PHA), and phorbol myristate acetate (PMA) was prepared. 4x10 8 Splenocytes were suspended in a T175 flask in activation medium and incubated at 37 °C, 5% CO 2 for 48 hours. After incubation, the supernatant was collected and filtered as alpaca conditioning medium.

[0152] A3. Immunity of Animals The antigen in 0.5 mL PBS (for example, 400 μg keyhole limpet cyanin (“KLH”)) is emulsified with 0.5 mL of complete Freund's adjuvant. The emulsified antigen is subcutaneously injected along the neck and back of the alpaca. Approximately 200 μL (or less) of 5 injections are performed. Immunization is carried out 3 times at 14-day intervals.

[0153] A4. Isolation of Lymphocytes from Different Organs To isolate PBMCs from blood, an EDTA-containing blood sample from an immunized alpaca is diluted 2-fold with 1xDPBS containing 2% FBS. Next, for density centrifugation, the diluted blood is slowly placed on Ficoll-Paque PLUS density gradient medium. The upper layer is removed, and the lymphocyte layer is transferred to a clean centrifuge tube. Next, the PBMCs are washed twice with 1xDPBS.

[0154] To isolate lymphocytes from the spleen, the spleen from an immunized alpaca is washed with 1xDPBS and placed in a clean dish. The spleen is inflated with a balloon by injecting medium until most of the lymphocytes are released. Next, the bottom of a 20 cc syringe is used to finely crush the spleen. Using too much force during crushing helps to obtain the highest possible yield of lymphocytes. All of the released cells are collected by gentle centrifugation (for example, 1400 rpm). After aspirating the supernatant, 5 times the volume of the pellet of erythrocyte lysis buffer is added, and the sample is left standing for at least 4 minutes. Next, RPMI1640 medium is added to terminate the lysis. Next, the lymphocytes are washed twice with 1xDPBS.

[0155] To isolate lymphocytes from lymph nodes, collect mesenteric lymph nodes and inguinal lymph nodes from immunized alpacas. Lymphocytes are released by grinding the lymph nodes in RPMI1640 medium. The cells are passed through a cell mesh and collected by centrifugation. Add five times the volume of red blood cell lysis buffer to the pellet, let the sample stand for at least 4 minutes to remove RBCs. Next, add RPMI1640 to stop the lysis of red blood cells. Then wash the lymphocytes twice with 1xDPBS.

[0156] To isolate lymphocytes from bone marrow, open the tibia and radius of immunized alpacas at both ends of the bone and collect the bone marrow. Cells are released by grinding the bone marrow in RPMI1640 medium. The cells are passed through a cell mesh and collected by centrifugation. Add five times the volume of red blood cell lysis buffer to the pellet, let the sample stand for at least 4 minutes to remove RBCs. Next, add RPMI1640 medium to stop the lysis and wash the lymphocytes twice with 1xDPBS.

[0157] A5. Depletion of non-specific cells Resuspend the collected cells in RPMI1640 medium containing 10% FBS, 1% penicillin-streptomycin, and 0.05 mM 2-mercaptoethanol to obtain a cell suspension of 1M / mL. Incubate the cells in a 6-well culture plate at 37 °C for 1 hour to non-specifically adhere macrophages and monocytes to the plate surface. After pre-incubation, collect and count the non-bound cells.

[0158] A6. B cell panning Resuspend the collected cells in 10% FBS and 1% penicillin-streptomycin to obtain a cell suspension of 1 million / mL. Incubate the cells at 37 °C for 1.5 hours at 5 million per 10 cm dish pre-coated with an antigen for specific B cell panning. After incubation, wash the seeded cells 2 - 10 times with RPMI1640 medium until only a few floating cells are found. Collect and count all non-bound cells and designate the count as Al.

[0159] A7. In vitro B cell culture Add 20 mL of B cell medium to a dish containing B cells retained by panning. The B cell medium should contain 10% FBS, 1% penicillin-streptomycin, 10% alpaca-conditioned medium, various growth factors (e.g., one or more interleukins) (at a concentration of 1 - 50 ng / ml), and 2.5 μM MMC-pretreated feeder cells (EL4.IL-2-C3, expressing alpaca CD40L). As a control in the quality control test, also culture dishes with B cells only and feeder cells only. The cells are cultured in 5% CO 2 for 8 - 10 days.

[0160] A8. Cell collection and QC On day 10, perform a test for antibody secretion using 50 μL of the supernatant of the B cell culture in ELISA. Collect and count all the cells in the co-culture dish. This number is regarded as A2. Collect and count the cells in the dish with feeder cells only. This number is regarded as A3. The B cell amplification factor (BCAF) is calculated by the following formula: BCAF = (A2 - A3) / (5M - A1)

[0161] A9. Construction of NGS library Use B cells from each panning to prepare template RNA for constructing an NGS cDNA library. A library using B cells before panning was also constructed to calculate sequence enrichment. i. RNA isolation

[0162] Extract RNA from co-cultured B cells using the TRIzol method.

[0163] The cultured B cells are lysed with TRIzol reagent by repeated pipetting or passing through a syringe and needle. 0.5 - 1x10 6Use 1 ml of reagent per cell. Add 20% chloroform and stir the tube for about 15 seconds. Carefully remove the aqueous phase using a pipette. Add an equal volume of isopropanol to the aqueous phase and mix gently. Centrifuge the sample at maximum speed (12,000 rpm) for 10 minutes. Remove the isopropanol, wash the pellet with 1 ml of 75% ethanol / DEPC-treated water, and mix gently. Centrifuge the pellet again at 7,000 rpm for 1 minute and recover the RNA in approximately 70 μl of RNase-free water.

[0164] ii. Reverse transcription For 0.2 mL of PCR mix, use different primers (random 6mer, oligo dT, and AI.CH2+AI.CH2.2) to amplify 8 reactions respectively. Heat the RNA- primer mix at 65 °C for 5 minutes and then incubate on ice for at least 1 minute. Prepare the RT reaction mix and perform reverse transcription at 30 °C for 10 minutes, 42 °C for 50 minutes, and 75 °C for 15 minutes.

[0165] iii. cDNA amplification PCR Approximately 2 μl of cDNA from reverse transcription is used as a template for a 50 μl PCR reaction containing 25 μl of 2x Primer star mix, 18 μl of RNase-free dH20, 1 μl of each primer NGS-leader1_L (GCAGTGGCTGCAGGTGTCCACTCG - SEQ ID NO: 63), NGS-leader2_L (GCAGGTCCCCAAGGTGTCCTGTCC - SEQ ID NO: 64), NGS-leader3_L (GGTGGTCCTGGCTGCTCT - SEQ ID NO: 65), NGS-hinge l_L (TTGTGGTTTTGGTGTCTTGGG - SEQ ID NO: 66), and NGS-hinge2_L (GGGGTCTTCGCTGTGGTGCGC - SEQ ID NO: 67), and the following cycles are performed: 98 °C for 3 minutes, 20 cycles (98 °C for 15 seconds, 58 °C for 30 seconds, 72 °C for 30 seconds), 72 °C for 3 minutes. The resulting amplicons are purified using the TIANGEN® PCR purification kit according to the manufacturer's instructions with a size cut-off of 300 bp or more.

[0166] iv. Index PCR Each amplicon sample is individually barcoded in a second "tagging" 50 μl PCR reaction that includes 25 μl of 2x Primer star mix, 18 μl of RNase Free dH20, 1 μl of each primer pair (such as P5-seqF and P7-indexl-seqR), and 100 ng of template as the first round of PCR. Next, the following cycles were performed: 30 seconds at 98 °C, 12 cycles (10 seconds at 98 °C, 30 seconds at 65 °C, 30 seconds at 72 °C); 5 minutes at 72 °C. The last three amplicons (derived from three different reverse primers) are pooled and purified from a 1.5% (w / v) agarose gel using the TIANGEN® PCR purification kit.

[0167] A10. NGS Data Analysis The amplified cDNA is sequenced using the MiSeq Sequencing System (Illumina, Miseq, 300x2). From each sample, 1 to 3 million reads are generated. The data is quality-checked using the NGSQC toolkit and assembled using FLASH. The assembled sequences are converted to protein sequences, and the CDR1, CDR2, and CDR3 regions (based on IMGT numbering) are computationally identified. The sequences are clustered into lineages / groups based on the same CDR3 length, a CDR3 Hamming distance of 1 or less, and the same mapped V / J germline segments as shown in Figure 4. The sequences are further sub-grouped and clustered based on the same CDR3 length and 80% or more CDR3 identity. For example, a lineage can be defined by clones that have a CDR3 length of at least 12 amino acids, a CDR3 Hamming distance of 0 or 1 (compared to a reference sequence), and the same V / J region amino acid sequence. The enrichment score for a sequence or sequence group is calculated using the ratio of frequencies between libraries constructed from B cells before and after panning.

[0168] Typical clones (20 or more) are selected from different lineages based on lineage preference factors as shown in Figure 7, and the antibodies they produce are tested in various binding assays such as ELISA and FACS. Clones selected from different lineages are typically found to bind to different epitopes of the target protein. To optimize existing clones, additional clones can be selected from the same lineage as shown in Figure 8. Clones selected within a lineage are typically shown to bind to different parts of the same epitope.

[0169] Option A: Select the clone with the highest number within the cluster. As shown in Table 1 below, Cluster C328 has a 357-fold increase in frequency after enrichment. NBL505-A1L1-P3R3_355, which is the top clone within the cluster, is a good option for selection for further testing.

Table 1

[0170] Option B: Select the clone with the highest enrichment score. By comparing the sequences within the library before and after enrichment (e.g., by panning or fluorescence-activated cell sorting), the enrichment score can be calculated based on the sequence frequencies before and after enrichment. To increase the chance of selecting clones that secrete functional VHHs, the enrichment score can be used to prioritize clone selection. In Table 2 below, clone NBL505-A1L2-P3R2_5559 is selected based on the enrichment rate.

Table 2

[0171] 1B. Single B cell B1. Sorting of Ag-specific single B cells from immunized alpaca PBMC

[0172] Peripheral blood mononuclear cells (PBMCs) were obtained from immunized alpacas by Ficoll density gradient centrifugation (GE) and divided into tubes containing 200x10 6 cells for immunostaining. The cells were incubated with 200 μL of KLH-biotin diluted to 5 μg / mL in MACS buffer (PBS + 2% FBS + 2 mM EDTA) at 4 °C for 30 minutes and then washed twice with 5 mL of ice-cold MACS buffer. Next, the cells were stained with rabbit anti-lama IgG (FI&L), APC-streptavidin, and a live / dead dye. Next, the stained samples were collected on a Moflo Cell Sorter Cytometer (Beckman), and single IgG+, KLFI+, live+ cells were collected into individual PCR tubes containing 8 μL of lysis buffer (Tiandz), 1 mM dNTP (Takara), 10 μL of buffer containing 3.75 μM random 6mer (Takara), and 1.25 μM oligo dT primer (Takara) per well.

[0173] B2. Single-cell RT-PCR and B-cell cloning The collected antigen-specific alpaca B cells were lysed in a collection tube and heated at 65 °C for 5 minutes. After cooling to 4 °C, total RNA from the lysed single cells was reverse transcribed at 42 °C for 50 minutes after an initial step of 30 °C for 10 minutes in 4 μL of 5xPrimeScript II buffer (Takara), 20 U of RNase inhibitor (Takara), 200 U of PrimeScript II RTase (Takara), and 4.5 μL of RNase-free water (Takara) in a final volume of 20 μL to allow random hexamer hybridization. The reaction was stopped by incubating at 72 °C for 15 minutes.

[0174] Next, the variable regions of the rearranged heavy chain (HC) locus, lambda (LCλ) or kappa (LCκ) light chain locus are individually amplified from each single cell cDNA by 2 rounds of nested PCR. For the variable segments, for 3 μL of cDNA, the first round of PCR at 98°C for 5 min, 98°C for 15 s, 55°C for 1 min (62°C for LCκ, 58°C for LCλ), and 72°C for 1 min is performed for 40 cycles, and then the final extension step is performed at 72°C for 7 min in 2X PrimeStar MAX buffer (Takara) and 100 nM primer / 40 μL reaction volume. 4 μL of the first amplification product was further amplified by the second PCR round. The second round of the PCR protocol consists of a denaturation step at 98°C for 5 min, 40 amplification cycles (98°C for 30 s, 58°C for 30 s for HC, 62°C for LCκ, 58°C for LCλ, and 72°C for 1 min), and a final step at 72°C for 7 min, and is performed using 2X PrimeStar MAX buffer (Takara) and 100 nM primer / 50 μL reaction volume. The PCR products from each single cell are detected on a 1.5% agarose GelRed gel. The PCR products from each well were filtered and purified using a commercially available purification kit (Tiangen).

[0175] Ligation is performed in a total volume of 20 μL containing 10 μL of Genbuilder and cloning-ligase (Genscript), 100 ng of digested and purified PCR product, and 100 ng of linearized vector. Electrocompetent E. coli TOP10 bacteria are transformed with 20 μl of the ligation product. Colonies are screened by PCR using PET-SEQ-F (TGCTGGTCTGCTGCTCCTCGC - SEQ ID NO: 68) as the forward primer and PET-SEQ-R (ACCGTCTATCAGGGCGATGG - SEQ ID NO: 69) as the reverse primer. The expected insert band is approximately 700 bp in length. Plasmid DNA from 10 colonies is isolated and sequenced for each plate to ensure that the consensus variable gene sequence is reliably identified.

[0176] Using the above single B cell sorting method and the purified epidermal growth factor receptor as an antigen, a group of 11 clones expressing VHH antibodies that specifically bind to EGFR with high affinity were identified. See Figure 10.

[0177] Using the above-described panning B cell enrichment method and mesothelin (MSLN) as an antigen, a set of clones was selected from NGS data including lineage and CDR3 length using an identity grouping method. Of the 12 most abundant clones in the library, 7 clones were shown to be strong MSLN binders (see Table 3 below). Importantly, MSLN is composed of three domains, and the antibodies of 5 MSLN-binding clones specifically bind to the epitope of domain 1 of MSLN. Furthermore, one antibody of the selected clones specifically binds to domain 2, and one antibody of the clones specifically binds only to domain 3. Thus, in one operation of lineage grouping and selection, clones that recognize a broad spectrum of epitopes of the full-length antigen can be identified. This will provide more opportunities to select clones for the treatment of different conditions of antigen binding, or more options for bispecific combinations. Furthermore, this method can also identify clones that have been shown to behave as blockers or non-blockers of antigen-ligand complexes. For example, among the 7 MSLN-binding clones, antibodies from 2 clones that bind to either domain 2 or 3 of MSLN but do not inhibit the binding of CA125 to MSLN are identified. In contrast, the 5 domain 1 epitope binders have been shown to prevent CA125 from binding to MSLN. The method of the present disclosure can efficiently select antibodies systematically and relevantly for testing in order to cover a wide range of epitopes with high resolution.

Table 3

[0178] In a further example of the identification of clones that produce antigen-specific antibodies, clones that produce anti-KLH antibodies were discovered by the following procedure.

[0179] For the isolation of antigen-specific B cells and further experiments, one alpaca (name #009) was immunized once with 200 μg KLH (keyhole limpet hemocyanin) in complete Freund's adjuvant, and then immunized every two weeks with 100 μg KLH (sigma) in incomplete Freund's adjuvant (Sigma) according to a standard immunization regimen. After evaluation of the antiserum by ELISA of serial dilution samples, 100 - 200 mL of blood was collected to isolate peripheral blood mononuclear cells (PBMC) using the Ficoll-Paque density gradient method (GE) according to the manufacturer's instructions. The isolated PBMC (viability > 95%) were resuspended in complete RPMI1640 medium to obtain a cell suspension of 10 6 / mL. 4 mL of the PBMC suspension was added to each well of a 6-well plate and incubated for 1 hour to capture non-specific binding cells. Next, the non-bound cells were collected and resuspended in complete RPMI1640 medium to 10 6 / mL. 5 mL of the cell suspension was added to a 10 cm high-binding Petri dish (protein binding capacity > 500 ng / cm 2 ) pre-coated with antigen with gentle shaking at 50 rpm and 37 °C for 1.5 hours. After incubation, the non-bound cells were washed away with 1xDPBS 2 - 10 times to remove non-specific binding cells. Mitomycin-treated EL4.IL-2-C3 feeder cells (stably expressing alpaca CD40L) resuspended in B cell medium were added to the dish at 0.5x10 6 cells / mL for B cell in vitro co-culture. The final volume was 20 mL per 10 cm dish. The B cell co-culture medium contains 10% FBS (fetal bovine serum), 1% penicillin-streptomycin, 10% alpaca-conditioned medium from the culture of alpaca blank PBMC, and various growth factors, such as one or more interleukins at a concentration of 1 - 50 ng / ml.

[0180] After 10 days of co-culture, 50 μL of the supernatant is used for ELISA to test for antibody secretion and the specificity of binding to the KLH immunogen. During that time, the co-cultured cells are collected and the total cell count is determined. The amplification of B cells after isolation and amplification of B cells is calculated by subtracting the cell count of the feeder cell control from the total cell count. To compare this concentrated cell count to the amount of cells initially added, the amplification of antigen-specific B cells can be calculated. As shown in Figure 11, B cells from five different co-culture dishes showed consistent antigen-specific VHH antibody secretion. Antigen-specific B cells in PBMC were amplified 4 - 6-fold after isolation and amplification of B cells (BIA).

[0181] After BIA, three NGS libraries are constructed using the mRNA of the co-cultured cells. For reverse transcription, oligo dT, random hexamers, and CH2-specific primers are used respectively. (Maass DR, Sepulveda J, Pernthaner A, Shoemaker CB. Alpaca (Lama pacos) as a convenient source of recombinant camelid heavy chain antibodies (VHHs). J Immunol Methods. 2007;324(1-2):13-25.). Next, these cDNAs are subjected to two rounds of PCR reactions. All three of these libraries are sent to Genscript (Nanjing, China) and sequenced on a MiSeq Sequencing System (Illumina, Miseq, 300x2) with 30% PhiX genomic DNA spike-in at 4 to 5 million reads per sample. The data is quality-checked using the NGSQC Toolkit and assembled with FLASH. Sequences are clustered into lineages / groups based on the same CDR3 length, a CDR3 Hamming distance of 1 or less, and the same mapped V / J germline sequences as shown in Figure 4. Sequences are further sub-grouped and clustered based on the same CDR3 length and 80% or more CDR3 identity. For each library, more than 800 groups are generated. Here, several lineage preference factors: sequence abundance, classical VHH vs. non-classical VHH, and CDR3 length are applied to select clones from these groups. Twenty clones are selected from 20 different groups and their synthesis, expression, and purification are performed. The bioinformatics data related to the selected clones is shown in Tables 4 to 6 below. The complete amino acid sequence of the antibody expressed by each of the selected clones is obtained as the contiguous sequence of each domain (FR1, CDR1, etc.).

Table 4

Table 5

Table 6

[0182] Verification of the antigen-specific binding of the antibody is performed by ELISA. The library is investigated based on subgrouping signatures (abundance of sequences, classical VHH vs. non-classical VHH, and length of CDR3). In this example, only one clone is selected from each cluster. In this example, most of the selected clones generate classical VHH rather than non-classical VHH. In this example, amino acid sequences with both long and short hinge sequences are included in the selected clones, as well as one clone with an additional pair of disulfide bonds within CDR3 compared to the normal VHH sequence. Such a clone selection strategy achieved a 100% success rate in the selection of clones with binding activity. As shown in the ELISA assay, all 20 selected clones showed specific binding activity against the KLH antigen KLH (shown in Figure 13). Eighteen of them had an average binding EC of 0.465 nm (excluding clone numbers 6 and 9 as outliers). Three clones (numbers 1, 3, and 13) were shown to be potent KLH binders with EC50s of 67, 77, and 72 pmol / L, respectively. Most of the leads were shown to have sub-nanomolar potencies. Therefore, cloned B cells secreting potent KLH-binding antibodies can be identified by BIA and NGS. To confirm the correlation results between the length of CDR3 and binding activity, clones with different CDR3 lengths are included in the list. The shortest CDR3 length seen is 9 amino acids, while the longest CDR3 length is 21. The average CDR3 length of the 20 VHH antibodies is 16 amino acids. As shown in Figure 42, a positive correlation is seen between the length of CDR3 and the clone ELISA activity.

Table 7

[0183] Example 2: Discovery of antibodies that block PD-1:PD-L1 complex formation by epitope prediction Overview of the experimental approach A workflow for discovering antibodies that block ligand:receptor binding interactions (complex formation) is shown in FIG. 14. This workflow starts with the identification of the three-dimensional structures of the ligand and receptor, and / or the structure of their complex, by computational or crystallographic methods. Both the receptor and ligand portions that form the binding interface can be determined by examining the interface of the "docked" proteins. The amino acids that interact to establish or stabilize the bound complex can be determined by examining the structure. Data from experiments that change amino acids at the interface and examine the effect of such changes on the binding strength of the complex can contribute to the determination of the amino acids that interact to form and stabilize the complex.

[0184] A short linear amino acid sequence portion of 2 to 4 amino acids of the ligand or receptor at the interface of the complex is selected, and then the NGS amino acid sequence library is searched to identify antibody cDNA clones that encode the short amino acids selected within the CDR portion of the sequence, preferably the CDR3 portion of the antibody sequence. A screening peptide sequence 2 to 4 amino acids in length is set as a keyword for searching the CDR3 sequence from the NGS database to obtain satisfactory VHH sequences and their abundances.

[0185] The sequences identified as present in the NGS library above a selected abundance threshold are selected for gene synthesis and expressed in HEK293 cells fused with a human IgG4-Fc tag. The expressed proteins are purified and subjected to functional tests, such as functional tests for antigen binding and inhibition of PD-1:PD-L1 complex formation.

[0186] PD-1 / PD-L1 structure analysis The structure of the PD-1 / PD-L1 complex was downloaded from the PDB database (PDB ID: 4ZQK). PYMOL software was used for structural analysis. The structure showed that PD-L1 covers two peptides of PD-1 (Figure 15). The loop sequence close to PD-L1 is SFVLNWYRMSPSNQTDKLAA (SEQ ID NO: 138), and the loop sequence close to PD-1 is YLCGAISLAPKAQIKESLR (SEQ ID NO: 139). The polar contact residue regions observed in the PD-1 / PD-L1 crystal structure were selected. The polar contacts between PD-1 and PD-L1 are displayed in PYMOL using "actions-find-polar contacts-to others excluding solvent".

[0187] Figure 15 shows the interface of the PD-1:PD-L1 complex (the complex data is publicly available - Protein Data Bank ID number: 4ZQK) and two adjacent peptides of PD-1, which are "blocking" peptides that may inhibit complex formation. Figure 16 shows the results of the interaction analysis between the PD-1 protein and the two adjacent peptides. The amino acids that interact between the SFV....SLR peptide of PD-1 and the AFT...RIT peptide of PD-L1 at the interface of those complexes are identified by connecting lines.

[0188] Based on the above analysis, the polar contact residues of PD-L1 were determined, and they are F-D-Q-ADYKR (SEQ ID NO: 144). Since there is a very long amino acid interval (>2 amino acids) in F-D-Q-A, only the peptide ADYKR (SEQ ID NO: 143) meets the selection strategy. For the selection of VHH, peptides of 2 - 4 amino acids are selected. Next, the peptides ADYK (SEQ ID NO: 68), DYKR (SEQ ID NO: 142), ADY, DYK, YKR, AD, DY, YK, and KR are set as screening criteria from the NGS database (Table 4).

[0189] Additional structures of PD-1 complexed with various antibodies are available, downloaded from the PDB database and analyzed in the same manner as above. Table 8 shows the complexes of PD-1 with various anti-PD-1 antibodies and the short peptides identified as potential blocking peptides by the interaction analysis of the above complexes.

Table 8

[0190] CDR3 within VHH is the main binding region to the antigen. In VHH antibodies, the CDR3 region is longer than the CDR3 region in conventional (VHVL) antibodies. The longer the CDR3, the larger the binding region provided by the VHH antibody. Therefore, attention has been focused on the CDR3 sequence. CDR3 sequences containing screening peptides were extracted from the NGS database, and their abundances were counted to eliminate repetitive sequences. It has been found that 2 - 4 amino acids are appropriate for searching the NGS database. Two amino acids may be too short (resulting in an unmanageable number of hits), and four amino acids may overly limit the number of options. Therefore, a three - amino - acid string was selected to search the NGS database.

[0191] Previous publications have described protocols for antibody discovery by panning phage surface VHH display libraries. However, with this method, only VHH clones with high abundances in the library tend to be obtained, and as a result, VHHs with low abundances are lost. Based on keyword searches using various screening peptide strings, a series of VHH sequences with various abundances (the proportion of all clones holding the DNA encoding the selected amino acid sequence) were selected from the NGS database. Table 9 shows the CDR3 regions of the amino acid sequences encoded within the selected clones with different abundances.

Table 9

[0192] Selection of CDR3 sequences from the NGS database A series of screening peptides with amino acid lengths of 2 - 4 were determined. The CDR3 was considered to contain the screening peptide and be able to bind to PD-1 and block PD-L1. Since the CDR3 in VHH is the main binding region to the antigen, it is longer than that of conventional antibodies, and the longer the CDR3, the larger the binding region provided to the VHH. The CDR3 sequences containing the screening peptide were extracted from the NGS database and their abundances were counted to eliminate repetitive sequences. Due to different screening peptides, it was found that 2 - 4 amino acids were appropriate. It should be noted that 2AA might be too short and the number of selections could be limited with 4AA, so the screening peptide with a length of 3 amino acids was the best choice.

[0193] Expression and purification of the antibody The nucleic acid encoding the selected VHH antibody sequence was synthesized and inserted into the pCDNA3.4 vector fused with an IgG4-FC tag containing a (G4S) 3 linker. The recombinant plasmid was confirmed by sequencing. Next, the VHH-FC construct was transformed into HEK293 cells. The transformed cells were cultured for 5 - 7 days to obtain the recombinant protein. Then, the recombinant VHH antibody was purified from the filtered culture supernatant. The protein concentration of the obtained antibody was measured by UV absorbance at 280 nm. The purity of the purified recombinant antibody was evaluated by Coomassie staining of sodium dodecyl sulfate-polyacrylamide gel (SDS-PAGE) and high performance liquid chromatography (HPLC).

[0194] ELISA binding assay For the ELISA binding assay, 96-well plates are coated overnight with 2 μg / ml of PD-1 protein. Recombinant VHH-Fc protein is added to the wells and left for a certain period of time. Then, 1 μg / ml of HRP-conjugated anti-Fc antibody is added as the detection antibody, and the assay reagent is added. The absorbance is read at 450 nm. The binding of nivolumab functions as a positive control, and the plates coated with BSA are used as the negative control group.

[0195] ELISA blocking assay Each well of a 96-well microtiter plate is coated with 2 μg / ml of PD-1 protein / PBS and blocked with 1% BSA. Purified VHH-Fc protein at a concentration of 5 μg / ml is added to each well.

[0196] Next, 2 μg / ml of biotin-labeled PD-L1 protein is added to each well. The PD-L1 protein is detected using HRP-conjugated anti-human IgG and TMB as the substrate. The intensity of the developed color is measured at 450 nm. The wells without added PD-1 and without added VHH-Fc are the negative control and positive control, respectively.

[0197] Thirty clones are selected, and VHH antibodies are purified from the selected clones. As described above, thirty clones are selected for expression and purification. Nineteen clones showed good expression of the recombinant VHH antibody.

[0198] The VHH antibodies purified from the nineteen clones are evaluated by ELISA for binding to PD-L1 and inhibition of PD-1:PD-L1 complex formation as described above. The results are shown in Table 10 and Table 11, respectively.

Table 10

Table 11

[0199] 1194-z0-IgG4 is a positive control antibody known to have activity blocking PD-1:PD-L1 complex formation. NBL507-BMK2-H4-IgG4 is an irrelevant antibody used as a negative control.

[0200] Seven of the selected clones produced VHH antibodies with substantial activity binding to PD-1, and five with the strongest binding were selected to test for activity blocking PD-1:PD-L1 complex formation. Clone SS5 has been found to exhibit stronger inhibition of complex formation than the positive control 1194-z0-IgG4.

[0201] From the results obtained, a collection of clones expressing antibodies with CDR3 sequences represented among all clones selected within at least 10 copies is likely to contain at least one clone expressing a VHH antibody that will specifically bind to the antigen and block antigen complex formation with a specific protein binding partner. For example, in the experiments disclosed herein, 100% of the clones with CDR3 sequences exceeding 10 copies among the selected clones expressed VHH antibodies binding to PD-1, and 33% of them expressed VHH antibodies that would block the binding of PD-1 to PD-L1. In contrast, only 25% of the clones with CDR3 sequences present in less than 10 copies among the selected clones expressed VHH antibodies binding to PD-1, and none of the selected clones expressed VHH antibodies that would block the binding of PD-1 to PD-L1.

[0202] In additional experiments conducted as described above, 30 clones were selected as described above for expression and purification. 19 clones were shown to have well-expressed recombinant VHH antibodies. Seven of these showed strong binding to human PD-1. The overall success rate of NGS-guided clone selection to identify clones secreting antigen-binding VHH antibodies is 37%.

[0203] Example 3: Development of Array Signatures and Clone Selection Rules Convergent (overlapping) sequences among animals immunized with BCMA. In this experiment, two alpacas named 507-A1 (Al) and 507-A2 (A2) are immunized with human recombinant BCMA protein. Animals A1 and A2 receive the same amount of recombinant human BCMA using the same immunization regimen. 93 and 40 VHH sequences with unique CDR3 amino acid sequences are discovered as promising leads from animals A1 and A2, respectively. When clones that are double-blinded with respect to immunization and clone identification are selected among the animals, there are 20 sequences that have CDR3 sequences shared by both animals. Table 12 shows the 20 unique CDR3 sequences shared by both animals. [Table 12]

[0204] The relationship between the CDR3 sequences of animals A1 and A2 is shown as a Venn diagram in Figure 18.

[0205] Antibody sequences shared between VHH2 (long hinge) sequences and VHH3 (short hinge) sequences. In alpaca A1, six VHH sequences with either a long region or a short hinge region can be found. As shown in Figure 19A, an original pool of 249 sequences is seen, of which 135 are found to have a hinge sequence. Among these, 26 non-redundant sequences with a long hinge sequence and 89 short hinge sequences are seen. Of these, 19 unique CDR3 sequences have a long hinge sequence and 53 unique CDR3 sequences have a short hinge sequence. It can be seen that 6 CDR3 sequences are shared among these two groups.

[0206] In a similar consideration of the amino acid sequences of clones from animal A2, two sequences shared by animal A2 in both the long hinge pool and the short hinge pool are identified as shown in Figures 20A and 20B.

[0207] Overall, the eight sequences of A1 and A2 animals are found to be within both the long hinge sequence pool and the short hinge sequence pool.

[0208] Convergent sequences between conventional antibodies of alpaca (VH) and VHH2 / VHH3. 19 single-chain antibodies have been identified from two animals showing good binding to human BCMA. Surprisingly, two VHH sequences shared by the long hinge pool and the short hinge pool named 1A1 and 1D2 are also shared by leads from the traditional VH repertoire. The number of sequences shared by different classes of antibodies is shown in Figure 21A. Eight VHH CDR3 sequences shared by either the long hinge or the short hinge are shown in Figure 21B, and the highlighted sequences are shared by all VH / VHH2 and VHH3.

[0209] The shared sequences have been shown to be strong BCMA binders. To test whether the overlapping sequences are a preferred signature, antibodies expressed by eight convergent VHH clones shared by VHH2 and VHH3 are purified. By either ELISA or flow cytometry, all were found to be strong BCMA binders. The sequences highlighted in Figure 21B named 1A1 and 1D2 (SEQ ID NOs: 117 and 118) are shared by both VH antibodies and VHH antibodies. The eight convergent VHH antibodies bind to cells of the tumor cell line RPMI8226 that overexpresses BCMA, but do not bind to 293T cells with negative BCMA expression. In ELISA, similar results were shown using either His or Fc conjugate-bound human BCMA as the coating antigen (Table 13).

Table 13

[0210] Repeated array signature. Finally, the two arrays shared by VH / VHH2 / VHH3 are also seen in both animals A1 and A2 (see Fig. 22). Therefore, selecting these arrays found in multiple animals / VH / VHH germlines is thought to be a useful signature for antibodies that exhibit strong and specific binding to their antigens.

[0211] Data supporting additional signatures. Examine the affinity of classical and non-classical VHHs for binding to human BCMA. Classical VHHs have a higher affinity than non-classical VHHs. The results are shown in Fig. 23; P < 0.05.

[0212] Additional signature statistics FR2 hydrophilic region: In most VHH antibodies, FR2 has unique amino acid substitutions compared to conventional IgG: 37Phe / Tyr, 44Glu, 45Arg, and 47Gly. Fig. 24 shows the percentage of clones with these FR2-unique amino acids of VHH antibodies from three libraries (each against a different antigen (NBL501 (anti-MSLN), NBL504 (anti-PD1), and NBL602 (anti-KLH))). Surprisingly, this specific FR2 substitution pattern was found to be included in up to 8% of the entire VHH repertoire of clones from these three libraries. The frequency of individual substitutions at positions 37, 44, 45, and 47 (Kabat numbering) of the FR2-unique sequence was also investigated, and these data are shown in Fig. 25.

[0213] Interestingly, the native sequence replacements are highly diversified. At positions 44 and 45, Glu and Arg are the dominant amino acids, while they are variable at positions 37 and 47. In the NBL501 anti-MSLN project, Tyr 70% and Phe 10% occupied position 37. In contrast, in the anti-PD1 (NBL504) and anti-KLH projects (NBL602), it was found that 37Phe reached 70% and 37Tyr was about 25%. A consistently high proportion of Leu was observed instead of Gly at position 47, and clones with Phe and Trp at this position exist at a significant proportion. The signature amino acid Gly occupies only less than 10% of the frequency at this position. Therefore, the signature of alpaca VHHFR2 is proposed as 37Phe / Tyr, 44Glu, 45Arg, and 47Gly / Leu / Phe. Extended CDR1 and CDR2:VHH have an additional hypervariable region (residues 27 - 30 according to Kabat numbering) adjacent to CDR1. In VHH antibodies with this region and a long CDR3 region, the surface area interacting with the antigen increases. However, antibodies with this signature are often not identified.

[0214] The CDR2 domain usually has a length of 5 - 9 amino acids. However, many clones contain a "long CDR2" with a length of 14 - 17 amino acids. Importantly, it has been found that VHHs containing a long CDR2 have a higher binding affinity for their antigens than VHHs containing a short CDR2. ELISA data from antibodies from three different libraries are shown in Figure 27.

[0215] Additional disulfide bonds within CDR3: Approximately 5 - 10% of VHH antibodies have additional disulfide bonds within the CDR3 domain. This may indicate that the epitope to which the antibody binds is a more "conformational" recognition site formed from the three-dimensional structure of the antigen rather than a short linear amino acid sequence. Approximately 2 - 19% of CDR3 contains intra-disulfide bonds. Figure 28 shows the ratios of three libraries of antibodies having additional disulfide bonds within the CDR3 domain (additional Cys residues, i.e., identified as two cysteine amino acids within CDR3). The NBL504 anti-PD1 library contains a high proportion of antibodies with long CDR3 domains and contains significantly more intra-CDR3 disulfide bonds than the other two libraries, which have approximately 2% of such clones.

[0216] Additional disulfide bonds between CDR1 and CDR2: Although at low frequency, antibodies having additional disulfide bonds between CDR1 and CDR2 can be found as additional Cys amino acids in CDR1 or CDR2. See Figures 29 and 30.

[0217] Additional disulfide bonds between CDR1-CDR3 or FR2-CDR3: In camels and dromedaries, 82% of the VHH antibodies have a disulfide bond between CDR1 and CDR3, while in llamas and alpacas, 74% of the VHH antibodies have a disulfide bond between FR2 and CDR3. 70-80% of the VHH antibody sequences do not contain additional disulfide bonds. 10-25% of the VHH sequences may contain additional disulfide bonds (4 cysteines within the sequence). Interestingly, 5-10% of the VHH antibody sequences contain unpaired cysteines. It is unclear whether these VHH antibodies can pair with another VHH antibody or not by forming disulfide bonds between the unpaired cysteines. Within a single VHH sequence, a total of 7 or fewer cysteines have been observed. These data suggest that up to 3 intradomain disulfide bonds can be formed with the remaining unpaired cysteines. Figure 31 shows the analysis of the number of cysteine residues within the amino acid sequences of VHH antibodies selected from 3 different libraries.

[0218] Additional disulfide bonds can be seen between CDR1 and CDR3, but this is rare. See Figure 32.

[0219] Most of the additional disulfide bonds in VHH antibodies are between FR2 / CDR2-CDR3. Consistent with the calculation of the total number of cysteines, most of the additional disulfide bonds are between FR2 / CDR2 (by either IMGT or Kabat nomenclature) and CDR3. The additional paired cysteines are 2-9% of all VHH sequences in the NGS database. See Figure 33.

[0220] A significant negative correlation is observed between the OD value of the supernatant by clone and the number of cysteines in its amino acid sequence (see Figure 34). Therefore, when selecting a clone, it is usually recommended to avoid an odd number of cysteines in the VHH amino acid sequence. Furthermore, it is preferable to avoid selecting a VHH sequence containing more than one pair of cysteine sequences. Additional disulfide bonds can affect the expression yield of VHH or have an adverse effect on the binding affinity. Therefore, additional disulfide bonds can have an adverse effect on the development potential of downstream production of nanobodies.

[0221] In the long CDR3:PD1 NGS library-(NBL504), up to 86% of the VHH clones have a "long CDR3" domain with more than 15 amino acids. See Figure 35.

[0222] Figures 36, 37, and 38 show a positive correlation between the length of the CDR3 domain and the antibody affinity assayed by FACS for VHH anti-BCMA antibodies binding to the surface of cells from two different cell lines expressing BCMA and by ELISA using the supernatant. The p-value for all correlations is 0.001 or less. It has been consistently shown that the length of CDR3 has a significant correlation with the binding affinity of the antibody to its antigen. Therefore, when there are a large number of clones that may be selected, it is preferable to select an antibody with a CDR3 domain length longer than 14 amino acids.

[0223] CDR3 variants with similar lengths have been observed to be likely to bind to similar epitopes of a single antigen. For a population of VHH antibodies binding to almost the same epitope or the same epitope of a specific antigen, a narrow range of CDR3 lengths was observed. See Figure 39.

[0224] Furthermore, antibodies that bind to the same epitope can be identified by related sequences. For example, clones with a CDR3 length that differs by 0 or 1 amino acid. Clones 1182, 1202, and 1734 were selected from anti-PD1 antibodies of the same lineage, all of which compete with Keytruda for PD1 binding (see Figure 40). It is confirmed that clones selected from related lineages with the same length of CDR3 are in the same bin. In the experiments disclosed herein, it is suggested that related clones with the same length of CDR3 domain have a high chance of binding to exactly the same epitope. When selecting clones based on NGS data, it is preferable to reduce the number of such redundant clones in the selected pool.

[0225] Non-classical VHHs having the same V and J germline segments as conventional IgG: A group of VHH lineages sharing the same naive B cell origin (same V and J configuration) as conventional IgG1 is shown to recognize the same or similar epitopes, either by VHH antibodies or IgG1 antibodies. Among the 507 libraries, two out of 81 sequences shared by VHH2 / VHH3 / VH are present, which is approximately 2.5% of the clones.

[0226] Non-classical VHHs with a unique sequence signature - conserved Trp118 is replaced by Arg118 or has a low hydrophobic profile in FR3: Trp118 can be seen at a maximum rate of 3% within each of the NBL501, NBL504, and NBL602 libraries. See Figure 41.

[0227] Overall, the above data suggest the following clone selection rules. 1. When immunizing one animal, select the VHH sequences shared by VHH2 and VHH3. When investigating VHVL, select those shared between VH and VHH. 2. Convergent motif or sequence signature: Different animals within the same experimental group can convergently generate the same motif or sequence signature via the same VDJ arrangement. The paratopes encoded by these motifs or sequence signatures can target functional epitopes. When multiple animals are immunized, select convergent sequences shared among the animals. 3. FR2 hydrophilic region: In most VHH antibodies, FR2 has amino acid substitutions specific to conventional IgG: 37Phe / Tyr, 44Glu, 45Arg, and 47Gly / Leu / Phe. 4. Non-classical VHH with a unique sequence signature: Conserved Trp118 substituted with Arg118 and / or a lower hydrophobic profile in FR4. 5. Classical VHH has a higher affinity than non-classical. When selecting clones, preferably select classical VHH. Non-classical VHH does not contain the FR2 signature. 6. Avoid selecting clones with an odd number of cysteines in the sequence: One clone was selected for synthesis using three cysteines. However, it did not express. Some screened clones with three cysteines have below-average expression. 7. In conventional antibodies, it has been found that folding problems occur due to the continued presence of positive or negative charges within CDR3. For VHH, one clone with three Rs in CDR3 was selected, but the clone did not express. Some screened clones with three consecutive positive charges (K or R or a mixture thereof) have below-average expression. Avoid selecting clones with more than three consecutive positive charges within CDR3. 8. To avoid clones with a positive charge at the N-terminus, two such clones were selected, but neither expressed. There are few screened clones with a positive charge at the N-terminus. 9. Some alpacas have VHH with a long CDR2 (17aa instead of 8 / 9), and clones with such a long CDR2 are thought to have high binding affinity. 10. In some projects, the length of CDR3 has a positive correlation with binding affinity, and longer CDR3 clusters are preferentially selected. 11. In the case of related clusters with similar CDR3 lengths, since it has been found that homologous CDRs of the same length lead to binding to similar epitopes, it is avoided to select many redundant candidates. 12. Convergent motifs or sequence signatures: Different animals within the same experimental group can convergently generate the same motif or sequence signature via the same VDJ arrangement. The paratopes encoded by these motifs or sequence signatures can target functional epitopes. 13. Novel canonical binding loop structures: Hypermutation hotspots present at important sites for determining canonical loop structures create interesting possibilities for diversifying the VHH structure repertoire. Crystallographic studies often emphasize that the CDR1 and CDR2 loops of camelid VHHs deviate from the known canonical structures of conventional VHs. Sequence-based prediction of novel Ag-binding loop conformations is supposed to support further grouping of the lineages (Laura S. Mitchell, Lucy J. Colwell, Comparative analysis of nanobody sequence and structure data, Proteins. 2018;86:697 - 706).

[0228] References Daley LP, Kutzler MA, et al. Effector functions of camelid heavy-chain antibodies in immunity to West Nile virus. Clin. Vaccine Immunol. 17:239 - 46, 2010. McCoy LE, et al. Potent and broad neutralization of HIV-1 by a llama antibody elicited by immunization. J. Exp. Med.2012. Cristina Basilico, et al. Four individually druggable MET hotspots mediate HGF-driven tumor progression, The Journal of Clinical Investigation, Volume 124 Number 7 July, 2014. Bas van der Woninga, et al. DNA immunization combined with scFv phage display identifies antagonistic GCGR specific antibodies and reveals new epitopes on the small extracellular loops, MABS, VOL. 8,NO. 6, 1126-1135, 2016. Laura M. Griffin et al. Analysis of heavy and light chain sequences of conventional camelid antibodies from Camelus dromedarius and Camelus bactrianus species, Journal of Immunological Methods Volume 405, Pages 35-46, March 2014. Adhdi Arbabi-Ghahroudi, et al. camelid Single-Domain Antibodies: Historical Perspective and Future Outlook, Frontiers in Immunology, Vol 8, 2017. Viet Khong Nguyen, et al. Camel heavy-chain antibodies: diverse germline VHH and specific mechanism enlarge the antigen-binding repertoire The EMBO Journal Vol. 19 No.5 2000 Mehdi Arbabi-Ghahroudi. Camelid Single-Domain Antibodies: Historical Perspective and Future Outlook. Front. Immunol., 20 November 2017. Nguyen VK, et al. Heavy-chain antibodies in Camelidae; a case of evolutionary innovation. Immunogenetics 54:39-47, 2002. Conrath KE, et al. Emergence and evolution of functional heavy-chain antibodies in Camelidae. Dev Comp Immunol 27:87-103, 2003. Nick Deschacht, et al. A Novel Promiscuous Class of Camelid Single-Domain Antibody Contributes to the Antigen-Binding Repertoire, The Journal of Immunology. 184 (10) 5696-5704, May 2010. Cortez-Retamozo V, et al. Efficient tumor targeting by single-domain antibody fragments of camels. Int J Cancer. 98(3):456-62, 2002. Alex Klarenbeek, et al. Camelid Ig V genes reveal significant human homology not seen in therapeutic target genes, providing for a powerful therapeutic antibody platform, mAbs 7:4, 693-706; 2015. Tomoyuki Igawa, et al. Engineering the variable region of therapeutic IgG antibodies. mAbs 3:3, 243-252; 2011. Laura S. Mitchell, Lucy J. Colwell, Comparative analysis of nanobody sequence and structure data, Proteins. 2018; 86:697-706. Maass DR, Sepulveda J, Pernthaner A, Shoemaker CB. Alpaca (Lama pacos) as a convenient source of recombinant camelid heavy chain antibodies (VHHs). J Immunol Methods. 2007;324(l-2):13-25.

Claims

1. 1. A method for producing camelid antibodies specific to an antigen, comprising the steps of: a) enriching and expanding B cells from an immunized camelid specific for said antigen; b) isolating VHHs from the antigen-specific B cells 2 , V.H.H. 3 , and VH 1 generating an antibody NGS library comprising the chain sequences; c) VHHs in the NGS library 2 , V.H.H. 3 , and VH 1 Grouping the antibody sequences of the chains by phylogenetic order; d) ranking the lineages of step c) by one or more lineage preference factors; e) the top-ranked VHHs of step d) in said NGS library 2 or VHH 3 Selecting representative sequences from the lineage comprising f) testing the antibody comprising the selected sequence from step e) to determine whether the antibody binds to the antigen or portion thereof; A method comprising:

2. 2. A method for generating a camelid antibody as described in claim 1, wherein the minimum CDR3 distance of a particular CDR3 is equal to or less than 1 among said group of CDR3s from a lineage, and the minimum CDR3 distance of a particular CDR3 is the shortest Hamming distance of such CDR3 compared to all other CDR3s of the same length.

3. A method for producing a camelid antibody according to claim 1 or 2, comprising the steps of: i) FR2 hydrophilic region, ii) extended CDR1; iii) an additional disulfide bond between CDR1-CDR3 or FR2-CDR3; iv) an additional disulfide bond within CDR3; v) a long CDR3 (≧15 aa); vi) an additional disulfide bond within CDR1; vii) non-classical VHHs with the same V and J germline as conventional IgG1; viii) non-classical VHHs with defined sequence signatures; ix) novel canonical binding loop structures, and x) Convergent motifs or sequence signatures The method further comprises subgrouping the lineages with VHH-specific characteristics selected from the group consisting of:

4. The lineage preference factors include lineages from high sequence abundance to low sequence abundance, lineages from high amplification factor to low amplification factor, changes in lineage sequence abundance during the immunization process, changes in lineage sequence abundance before and after depletion of specific unwanted B cells, VHH and VH 1 4. A method for producing a camelid antibody according to any one of claims 1 to 3, wherein the antibody is selected from the group consisting of lineages sharing the same naive B cell origin between the camelid and the camelid, avoidance of sequences responsible for developmental potential, and combinations thereof.

5. VHH 2 and V.H.H. 3 A method for producing camelid antibodies according to any one of claims 1 to 4, wherein the first top 100 lineages of are selected in e).

6. A method for producing a camelid antibody according to any one of claims 1 to 5, wherein said antibody in step f) is expressed by a prokaryotic or eukaryotic cell.

7. 7. A method for producing camelid antibodies according to any one of claims 1 to 6, further comprising monitoring the immune response of camelid antibodies IgG2 (HcAb), IgG3 (HcAb) and IgG1 (conventional IgG).

8. Repeat steps e) to f) to obtain the selected VHH 2 or VHH 3 A method for producing a camelid antibody according to any one of claims 1 to 7, further comprising optimising sequences within said same phylogenetic group of heavy chain only antibodies.

9. 1. A method for producing camelid antibodies specific to an antigen, comprising the steps of: a) enriching and expanding B cells from an immunized camelid specific for said antigen; b) isolating VHHs from the antigen-specific B cells 2 , V.H.H. 3 , V.H. 1 and VL 1 generating an antibody NGS library comprising the chain sequences; c) VHHs in the NGS library 2 , V.H.H. 3 , V.H. 1 , and VL 1 Grouping the sequences by lineage and d) VH according to anchor binders generated by single B cell sorting and heterohybridoma approaches 1 / VL 1 Pairing lineages; e) ranking the lineages and lineage pairs of steps c) and d) by one or more lineage preference factors; f) selecting, within said NGS library, the top-ranked VHHs in step e) 2 or VHH 3 The lineage and VH 1 / VL 1 selecting a representative sequence or sequence pair from each of the phylogenetic pairs; g) testing an antibody comprising the selected sequence pair or sequence from step f) to determine whether the antibody binds to the antigen or portion thereof; A method comprising:

10. 10. A method for producing camelid antibodies as claimed in claim 9, wherein the anchor of the IgG1 repertoire is produced by single B cell sorting and heterohybridoma approaches.

11. The ranking of phylogenetic pairs in step e) is based on the VH 1 11. A method for producing a camelid antibody according to claim 9 or claim 10 based on a lineage preference factor of the lineage.

12. One representative sequence or one representative pair of VHHs from the top 100 lineages or lineage pair groups 1 / VL 1 A method for producing a camelid antibody as claimed in claim 10, wherein: is selected as the anchor.

13. The top 100 lineages or lineage pair groups are 70 VHH lineage groups and 30 VHH 1 / Vκ or VH 1 13. A method for producing a camelid antibody as claimed in claim 12, comprising a Vλ / Vλ phylogenetic group pair.

14. The lineage preference factors include lineages from high sequence abundance to low sequence abundance, lineages from high amplification factor to low amplification factor, changes in lineage sequence abundance during the immunization process, changes in lineage sequence abundance before and after depletion of specific unwanted B cells, VHH and VH 1 14. A method for producing a camelid antibody according to any one of claims 10 to 13, wherein the antibody is selected from the group consisting of lineages sharing the same naive B cell origin between the camelid and the camelid, avoidance of sequences responsible for developmental potential, and combinations thereof.

15. A method for producing a camelid antibody according to any one of claims 9 to 14, further comprising repeating steps f) to g), wherein the VHH clone comprises: i) FR2 hydrophilic region, ii) extended CDR1; iii) an additional disulfide bond between CDR1-CDR3 or FR2-CDR3; iv) an additional disulfide bond within CDR3; v) a long CDR3 (≧15 aa); vi) an additional disulfide bond within CDR1; vii) non-classical VHHs with the same V and J germline as conventional IgG; viii) non-classical VHHs with defined sequence signatures; ix) novel canonical binding loop structures, and x) Convergent motifs or sequence signatures A method having a feature selected from the group consisting of:

16. Antigen-specific humanized VHH 2 or VHH 3 1. A method for producing an antibody, comprising: a) enriching and expanding B cells from an immunized camelid specific for said antigen; b) VHH from antigen-specific B cells 2 , V.H.H. 3 , V.H. 1 generating an antibody NGS library comprising the chain sequences; c) VHHs in the NGS library 2 , V.H.H. 3 , V.H. 1 Grouping the sequences by lineage and d) By comparing its amino acid sequence with that of multiple related antibodies within the same lineage, each of which binds to the same epitope as the parent antibody, the parent VHH that share the same naive B cell origin 2 , V.H.H. 3 Antibody or VH 1 identifying a substitutable position in e) said parent VHH 2 , V.H.H. 3 replacing an amino acid at one or more of said substitutable positions of the antibody with an amino acid at a corresponding position in a human antibody; f) testing antibodies containing the substituted residues in the selected sequences to determine whether the antibodies bind to the antigen or portion thereof; A method comprising:

17. A method for generating a humanized VHH antibody according to claim 16, wherein the substitutable positions are within the CDR regions.

18. A method for generating a humanized VHH antibody according to claim 16, wherein the substitutable positions are within the FR region.

19. the antigen is a composite immunogen; 19. A method for producing a camelid antibody according to any one of claims 1 to 18, further comprising using antibodies determined to bind to the composite immunogen in step f) to identify individual antigens contained in the composite immunogen by protein array, cell / tissue antigen cDNA library or mass spectrometry based immunoprecipitation.

20. The selected VHH sequences are applied to VHH sequences of these clones that share the same naive B cell origin. 1 -VL 1 The method further includes guiding the selection of pairs, the selection criteria being 1) VHH and VH 1 2) the same CDR3 sequence between CDR1 and CDR2; 3) differences in FR1, FR2, FR3 and FR4.

21. A method for producing a camelid antibody according to any one of claims 1 to 8, further comprising repeating steps e) to f) to produce the antibody.

22. A method for producing a camelid antibody according to any one of claims 9 to 15, further comprising repeating f) to g) to produce the camelid antibody.

23. A method for producing a camelid antibody according to any one of claims 1 to 28, wherein the expressing cells of the tested antibodies comprise eukaryotic cells.

24. 30. An isolated camelid antibody or antigen-binding portion produced by the method of any one of claims 1 to 29.