Method for discovery of alternative antigen specific antibody variants
By integrating B-cell cloning with NGS, the method identifies antibody variants with improved developability by avoiding post-translational modification hotspots, addressing the limitations of current screening methods and enhancing antibody characterization.
Patent Information
- Application Number
- US19/029982
- Authority / Receiving Office
- US · United States
- Patent Type
- Applications(United States)
- Current Assignee / Owner
- Priority Date
- 2017-03-07
- Filing Date
- 2025-01-17
- Publication Date
- 2025-08-21
AI Technical Summary
Current methods for generating antibodies are limited by the need for extensive screening and fail to characterize the full diversity of B-cell responses, leading to a loss of characterized immune response diversity and the identification of antibodies with developability issues such as unpaired Cys-residues, unusual glycosylation sites, and degradation hotspots.
A method combining B-cell cloning with next-generation sequencing (NGS) to identify variant binders by sequencing antibody variable domains, aligning them with reference sequences, and selecting variants that improve developability by avoiding post-translational modification hotspots.
Enables the identification of antibody variants with improved developability and reduced post-translational modification risks, providing a comprehensive profile of the antibody response and enabling the production of functional antibodies without extensive screening.
Smart Images

Figure US20250263778A1-D00001 
Figure US20250263778A1-D00002 
Figure US20250263778A1-D00003
Abstract
Description
CROSS REFERENCE TO RELATED APPLICATIONS
[0001] This application is a continuation of U.S. application Ser. No. 16 / 562,412, filed Sep. 5, 2019, which is a continuation of International Application No. PCT / EP2018 / 055278, filed Mar. 5, 2018, which claims priority to European Patent Application No. 17159617.4, filed Mar. 7, 2017, each of which are incorporated herein by reference in its entirety.SEQUENCE LISTING
[0002] This application contains a Sequence Listing which has been submitted electronically in XML format and is hereby incorporated by reference in its entirety. Said XML copy, created on Jan. 17, 2025, is named P34157-US-1_Seq_List.xml and is 118 kilobytes in size.FIELD OF THE INVENTION
[0003] The current invention is in the field of antibody technology. More precisely herein is reported a method combining the versatility of B-cell cloning (BCC) with the power of next generation sequencing (NGS) to identify variant binders of a reference binder present within the B-cell population obtained from one or more immunized animals.BACKGROUND OF THE INVENTION
[0004] Today antibodies are generally generated either by phage display or by immunizing laboratory animals and isolating the antibody producing B-cells therefrom. In the latter case the number of B-cells to be processed is reduced based on the properties of the B-cell or the respective secreted antibody. Thereafter the sequence information is obtained. Thus, candidate selection is done mostly based on the binding and functional properties of the antibodies but “blinded” with respect to the amino acid sequence, and thereby with respect e.g. to developability aspects, such as unpaired Cys-residues, unusual glycosylation sites, degradation hotspots (Asp, Asn, Met etc.).
[0005] In WO 2015 / 070191 a systems and methods for detection of genomic variants are reported. WO 2015 / 155035 reports methods for identifying and mapping the epitopes targeted by an antibody response. In WO 2015 / 164757 methods of viral neutralizing antibody epitope mapping are reported. WO 2016 / 023962 reports consensus-based allele detection. In WO 2016 / 118883 the detection of rare sequence variants, methods and compositions therefore are reported.
[0006] The huge number of B-cell obtained during current immunization processes prevents the characterization of all isolated antibodies produced thereby in detail. A selection and reduction of clone numbers has to be made, reducing the characterized immune response diversity.
[0007] Wu et al. disclosed the focused evolution of HIV-1 neutralizing antibodies revealed by structures and deep sequencing (Science 333 (2011) 1593-1602).
[0008] Zhu et al. disclosed the de novo identification of VRCO1 class HIV-1-neutralizing antibodies by next-generation sequencing of B-cell transcripts (Proc. Natl. Acad. Sci. USA 110 (2013) E4088-E4097).
[0009] Wu et al. and Zhu et al. each disclose a method for the identification of “new” cognate antibody variants for a given reference binder sequences from the same or different species as a reference binder.
[0010] Fridy et al. disclosed a robust pipeline for rapid production of versatile nanobody repertoires (Nat. Meth. 11 (2014) 1253-1260+SI). Like Wu et al. and Zhu et al. do the methods as done by Fridy et al. not disclose any result or method or approach using sequence repertoire for searching antibody variants.
[0011] Glanville et al. disclosed what insight deep sequencing in library selection projects brings (Curr. Opin. Struct. Biol. 33 (2015) 146-160).
[0012] Thus, there is a need to provide methods for identifying based on sequence information contained in the immune response diversity additional or variant antibodies binding to the same antigen but having different properties.SUMMARY OF THE INVENTION
[0013] However, the present inventors have found that the challenge presented by the aforementioned loss of characterized immune response diversity can be overcome by exploiting the quantitative nature of next-generation sequencing.
[0014] It has been found by the current inventors that by combining the versatility of B-cell cloning (BCC) with the power of next generation sequencing (NGS) it is possible to identify variant binders of a reference binder present within the B-cell population obtained from one or more immunized animals with respect to the same antigen.
[0015] The method as reported herein merges the efficiency of B-cell cloning with the power of next generation sequencing, providing a highly streamlined approach for the identification of antibody variants of a reference antibody without the need to do an extensive (immuno- or cellular-) assay based screening.
[0016] It has been found that with the methods as reported herein it is possible to identify for antibody variable domains and also complete VH / VL pairs that have at least one developability hot-spot, i.e. that have at least one amino acid residue prone to post-translational modification, a variant that does not comprise said developability hot-spot, i.e. that has said at least one amino acid residue prone to post-translational modification changed to a different amino acid residue not prone to the same post-translational modification.
[0017] One other result of the herein reported method is the ability to provide a unique profile of the antibody response of one individual animal or of a group of animals immunized with the same antigen. This profile can be a source of valuable information.
[0018] One aspect as reported herein is a method for selecting a variant of a reference antibody variable domain encoding nucleic acid, wherein either the variant antibody variable domain amino acid sequence encoded by said variant of the reference antibody variable domain encoding nucleic acid, has improved developability compared to the reference antibody variable domain amino acid sequence encoded by said reference encoding nucleic acid, or wherein in the variant antibody variable domain amino acid sequence encoded by said variant of the reference antibody variable domain encoding nucleic acid, at least one amino acid residue that is post-translationally modified has been changed compared to the reference antibody variable domain amino acid sequence encoded by said reference encoding nucleic acid, wherein the variant and the reference antibody variable domain when paired with the respective other domain form an antibody binding site that bind to the same antigen,
[0019] the method comprising the following steps:
[0020] (i) providing a multitude of DNA-containing samples each including one or more antibody variable domain encoding nucleic acids;
[0021] (ii) performing PCR amplification of the antibody variable domain encoding nucleic acids of the multitude of (i) using consensus sequence-specific primers to obtain amplification products;
[0022] (iii) sequencing a plurality of the amplification products obtained in step (ii) in order to determine the relative proportion of each nucleotide at each position (in a sequencing read);
[0023] (iv) performing a sequence alignment between the sequencing (read) results of (iii) and the reference antibody variable domain encoding nucleic acid;
[0024] (v) performing a sequence-identity or homology-based ranking of the antibody variable domain encoding nucleic acids in said sequence alignment of (iv) with the reference antibody variable domain encoding nucleic acid being the template sequence; and
[0025] (vi) selecting the variant antibody variable domain encoding nucleic acid from one of the top 10 sequences of the sequence ranking of step (v);
[0026] whereby the variant selected in step (vi) is selected so that the developability has improved and / or at least one amino acid residue that is post-translationally modified has been changed.
[0027] One aspect as reported herein is a method for selecting a variant of a reference antibody variable domain, wherein the reference antibody variable domain has at least one amino acid residue that is post-translationally modified, the method comprising the following steps:
[0028] receiving sequencing data produced by sequencing nucleic acids from a multitude of B-cell clones each producing an antibody specifically binding to the same target as the reference antibody;
[0029] aligning the sequencing data with the sequence of the reference antibody being the template sequence;
[0030] selecting a sequence that has the highest structural / functional identity / similarity to the reference sequence but not having the at least one amino acid residue that is post-translationally modified of the reference antibody variable domain.
[0031] One aspect as reported herein is a method for identifying a variant antibody of a reference antibody specifically binding to the same target / antigen comprising the following steps:
[0032] i) providing
[0033] α) the amino acid sequence or nucleic acid sequence of a reference antibody, whereby the reference antibody specifically binds to a target / antigen,
[0034] β) at least one biological property of said reference antibody and optionally one or more assays to determine the at least one biological property,
[0035] γ) a multitude of amino acid sequences or nucleic acid sequences of variant antibodies specifically binding to the same target as the reference antibody, which have been determined by next generation sequencing,
[0036] ii) generating one or more sets of (related) VH or / and VL sequences, wherein the sequences are aligned based on
[0037] α) identical length of VH or / and VL, or / and
[0038] β) identical length of all β-sheet framework regions, or / and
[0039] γ) identical length of all HVRs / CDRs, or / and
[0040] δ) identical HVR3 / CDR3 sequence, mutations in frameworks and / or other HVRs / CDRs with up to 3 amino acid exchanges allowed, or / and
[0041] ε) homologous HVR / CDR3 sequence with up to 2 amino acid exchanges allowed and identical HVR / CDR1 and 2, or / and
[0042] ζ) mutations in frameworks and HVR / CDRs are allowed;
[0043] iii) ranking the aligned sequences in the one or more sets of ii) by
[0044] α) the number of, and / or
[0045] β) the position(s) of, and / or
[0046] γ) the change of the physico-chemical properties resulting from the, and / or
[0047] δ) the difference of VH / VL orientation resulting from the amino acid difference(s) to the reference antibody sequence;
[0048] iv) identifying one of the best 10 aligned and ranked antibodies as a variant antibody of a reference antibody, whereby the variant antibody does have at least one amino acid residue less that is post-translationally modified as the reference antibody.
[0049] In one embodiment of this aspect step ii) further comprises annotating the sequence with the same numbering scheme, which is the Wolfguy numbering scheme.
[0050] In one embodiment of this aspect differences in the sequence are annotated in the form: reference antibody amino acid residue-position-variant antibody amino acid residue and are grouped into a mutation tuple.
[0051] In one embodiment of this aspect the change of the physico-chemical properties is determined by the change in charge, hydrophobicity and / or size.
[0052] In one embodiment of this aspect the change of the physico-chemical properties is determined using a mutation risk score.
[0053] In one embodiment of this aspect the mutation risk score is determined based on the following Table, wherein residues that are not explicitly given in this Table are weighted with the value one:WolfguyWolfguyIndexWeightIndexWeight1010.215121021.11522.610301531.21040.51542.31050.21551.91060.81563.71070.515741080.815841090.419341100.21944111019541120.41963.311301973.91140.11982.61150.61993.41160.12513.81170.12521.91180.225341190.22543.81201.22553.612102564122428741230288412422893.91250.52903.420142913.72022.62922.120322933.62041.72942.320512952.32060296120702971.22080.72982.42091.62990.5210235142111.33523.52123.13533.52130.935432143.135533010.135633021.235733031.735833040.335933051.83603306036133072.43623308036333091.23643334136533351366333613673337138233100.838333110.938433120.838533132.538633140.138733151.53883316038933171.539033180.839133190.139233201.439333210.239433220.33953.53230.239633241.83973.53250.63981.53262.2399333313270.632833292.533043312.93322.84013.540224030.540414050.34060.140714080.140914100.14110.1with positions 101 to 125 corresponding to heavy chain variable domain framework 1, positions 151 to 199 corresponding to CDR-H1, positions 201 to 214 corresponding to heavy chain variable domain framework 2, positions 251 to 299 corresponding to CDR-H2, positions 301 to 332 corresponding to heavy chain variable domain framework 3, positions 351 to 399 corresponding to CDR-H3, positions 401 to 411 corresponding to heavy chain variable domain framework 4.
[0054] In one embodiment of this aspect the multitude of amino acid sequences or nucleic acid sequences of variant antibodies specifically binding to the same target as the reference antibody are obtained from B-cells from the same immunization campaign as the reference antibody, wherein the B-cells have been enriched for antigen-specific antibody expressing B-cells.
[0055] In one embodiment of all aspects one or more of the following is removed in the variant: i) unpaired Cys-residues in the variable domain or the HVR, ii) glycosylation sites, and iii) degradation hot-spots (Asp, Asn or Met).
[0056] One aspect as reported herein is a method for selecting a variant of a reference antibody variable domain, wherein the reference antibody variable domain has at least one amino acid residue that is post-translational modified, the method comprising the steps of the method according to any one of the preceding aspects.
[0057] One aspect as reported herein is a method for producing an antibody comprising the following steps:
[0058] cultivating a cell comprising the nucleic acid obtained with a method according to any one of the previous aspects and all other nucleic acids required for the expression of a functional antibody,
[0059] recovering the antibody from the cell or the cultivation medium.
[0060] One aspect as reported herein is a cell comprising the nucleic acid obtained with the method according to any one of the previous aspects.
[0061] One aspect as reported herein is a method for identifying a variant antibody of a reference antibody (that has comparable biological properties as the reference antibody) specifically binding to the same target / antigen comprising the following steps:
[0062] i) providing
[0063] α) the amino acid sequence or nucleic acid sequence of a reference antibody, whereby the reference antibody specifically binds to a target / antigen,
[0064] β) at least one biological property of said reference antibody and optionally one or more assays to determine the at least one biological property,
[0065] γ) a multitude (at least 10, at least 100, at least 1,000, at least 10,000) of amino acid sequences or nucleic acid sequences of variant antibodies specifically binding to the same target as the reference antibody (optionally obtained in / from the same immunization campaign as the reference antibody), which have been determined by next generation sequencing;
[0066] ii) generating one or more sets of related VH or / and VL sequences, wherein the sequences are aligned based on
[0067] α) identical length of VH or / and VL, or / and
[0068] β) identical length of all β-sheet framework regions, or / and
[0069] γ) identical length of all HVRs / CDRs, or / and
[0070] δ) identical HVR3 / CDR3 sequence, 1 or 2 (1 to 3 amino acid exchanges) mutations in frameworks and / or HVRs / CDRs allowed, or / and
[0071] ε) highly homologous HVR / CDR3 sequence (1 or 2 amino acid exchanges allowed) and identical HVR / CDR1 and 2 (with mutations in frameworks allowed), or / and
[0072] ζ) mutations in frameworks and HVR / CDRs are allowed;
[0073] iii) ranking the aligned sequences in the one or more sets of ii) by
[0074] α) the number of, and / or
[0075] β) the position(s) of, and / or
[0076] γ) the change of the physico-chemical properties resulting from the, and / or
[0077] δ) the difference of VH / VL orientation (angle) resulting from the amino acid difference(s) to the reference antibody sequence;
[0078] iv) identifying one of the best 10 (or best 5, or best 3, or the best) aligned and ranked antibodies as a variant antibody of a reference antibody.
[0079] In one embodiment step ii) further comprises annotating the sequence with the same numbering scheme and annotating differences. In one embodiment the numbering scheme is the Kabat EU numbering or Wolfguy numbering scheme. In one preferred embodiment the numbering scheme is the Wolfguy numbering scheme. In one embodiment the differences are annotated in the form: reference antibody amino acid residue-position-variant antibody amino acid residue. In one embodiment the mutations of a variant antibody with respect to the reference antibody are grouped in a mutation tuple.
[0080] In one embodiment step ii) further comprises removing sequences that are identical in sequence to the reference antibody and / or one of the variant antibodies, or that are identical or similar with regard to the number, location and / or type of amino acid difference with respect to the reference antibody.
[0081] In one embodiment the change of the physico-chemical properties is determined by the change in (overall) charge, hydrophobicity and / or size. In one embodiment the change of the physico-chemical properties is determined using a mutation risk score. In one embodiment the risk score is determined based on the following Table, wherein residues that are not explicitly given in this Table are weighted with the value one:WolfguyWolfguyIndexWeightIndexWeightFramework 11010.2CDR-H115121021.11522.610301531.21040.51542.31050.21551.91060.81563.71070.515741080.815841090.419341100.21944111019541120.41963.311301973.91140.11982.61150.61993.41160.1CDR-H22513.81170.12521.91180.225341190.22543.81201.22553.612102564122428741230288412422893.91250.52903.4Framework 220142913.72022.62922.120322933.62041.72942.320512952.32060296120702971.22080.72982.42091.62990.52102CDR-H335142111.33523.52123.135342130.93543.52143.13553.5Framework 33010.135633021.235733031.735833040.335933051.83603306036133072.43623308036333091.23643334136533351366333613673337138233100.838333110.938433120.838533132.538633140.138733151.53883316038933171.539033180.839133190.139233201.439333210.239433220.33953.53230.239633241.83973.53250.63981.53262.2399333313270.632833292.533043312.93322.8Framework 44013.540224030.540414050.34060.140714080.140914100.14110.1with WolfGuy Index positions 101 to 125 corresponding to heavy chain variable domain framework 1, positions 151 to 199 corresponding to CDR-H1, positions 201 to 214 corresponding to heavy chain variable domain framework 2, positions 251 to 299 corresponding to CDR-H2, positions 301 to 332 corresponding to heavy chain variable domain framework 3, positions 351 to 399 corresponding to CDR-H3, positions 401 to 411 corresponding to heavy chain variable domain framework 4.
[0082] In one embodiment the mutation risk score takes negative values, whereby larger negative values indicate a larger risk of loss of function.
[0083] In one embodiment the multitude (at least 10, at least 100, at least 1,000, at least 10,000) of amino acid sequences or nucleic acid sequences of variant antibodies specifically binding to the same target as the reference antibody are obtained from B-cells from the same immunization campaign as the reference antibody, wherein the B-cells have been enriched for antigen-specific antibody expressing B-cells. In one embodiment the enrichment is by antigen-specific sorting or / and cell panning or / and non-antigen specific antibody producing B-cells depletion.
[0084] One aspect as reported herein is a method for identifying a variant antibody of a reference antibody (that has comparable biological properties as the reference antibody) specifically binding to the same target / antigen comprising the following steps:
[0085] i) determining / generating / measuring
[0086] α) the amino acid sequence or nucleic acid sequence of a reference antibody, whereby the reference antibody specifically binds to a target / antigen,
[0087] β) at least one biological property of the reference antibody,
[0088] γ) by next generation sequencing for a multitude (at least 10, at least 100, at least 1,000, at least 10,000) of variant antibodies the amino acid sequence or nucleic acid sequence, whereby the variant antibodies specifically bind to the same target / antigen as the reference antibody and optionally are obtained in / from the same immunization campaign as the reference antibody;
[0089] ii) aligning the VH or / and VL sequences based on
[0090] α) identical length of VH or / and VL, or / and
[0091] β) identical length of all 3-sheet framework regions, or / and
[0092] γ) identical length of all HVRs / CDRs, or / and
[0093] δ) identical HVR3 / CDR3 sequence, mutations in frameworks and / or HVRs / CDRs 1 or 2 (1 to 3 amino acid exchanges) allowed, or / and
[0094] ε) highly homologous HVR3 / CDR3 sequence (1 or 2 amino acid exchanges allowed) and identical HVR / CDR 1 and 2 (with mutations in frameworks allowed), or / and
[0095] ζ) mutations in frameworks and HVRs / CDRs are allowed to generated one or more sets of related VH / VL sequences;
[0096] iii) ranking the aligned sequences in the one or more sets of ii) by
[0097] α) the number of, and / or
[0098] β) the position(s) of, and / or
[0099] γ) the change of the physico-chemical properties resulting from the, and / or
[0100] δ) the difference of VH / VL orientation (angle) resulting from the amino acid difference(s) to the reference antibody sequence;
[0101] iv) identifying one of the 10 (or 5 or 3 or the) best aligned and ranked antibodies as a variant antibody of a reference antibody.
[0102] In one embodiment step ii) further comprises annotating the sequence with the same numbering scheme and annotate differences. In one embodiment the numbering scheme is the Kabat EU numbering or the Wolfguy numbering scheme. In one preferred embodiment the numbering scheme is the Wolfguy numbering scheme. In one embodiment the differences are annotated in the form: reference antibody amino acid residue-position-variant antibody amino acid residue. In one embodiment the mutations of a variant antibody with respect to the reference antibody are grouped in a mutation tuple.
[0103] In one embodiment step ii) further comprises removing sequences that are identical in sequence to the reference antibody and / or one of the variant antibodies, or that are identical or similar with regard to the number, location and type of amino acid difference with respect to the reference antibody.
[0104] In one embodiment the change of the physico-chemical properties is determined by the change in charge, hydrophobicity and / or size. In one embodiment the change of the physico-chemical properties is determined using a mutation risk score. In one embodiment the risk score is determined based on the following Table, wherein residues that are not explicitly given in this Table are weighted with the value one:WolfguyWolfguyIndexWeightIndexWeightFramework 11010.2CDR-H115121021.11522.610301531.21040.51542.31050.21551.91060.81563.71070.515741080.815841090.419341100.21944111019541120.41963.311301973.91140.11982.61150.61993.41160.1CDR-H22513.81170.12521.91180.225341190.22543.81201.22553.612102564122428741230288412422893.91250.52903.4Framework 220142913.72022.62922.120322933.62041.72942.320512952.32060296120702971.22080.72982.42091.62990.52102CDR 35142111.33523.52123.13533.52130.935432143.13553Framework 33010.135633021.235733031.735833040.335933051.83603306036133072.43623308036333091.23643334136533351366333613673337138233100.838333110.938433120.838533132.538633140.138733151.53883316038933171.539033180.839133190.139233201.439333210.239433220.33953.53230.239633241.83973.53250.63981.53262.2399333313270.632833292.533043312.93322.8Framework 44013.540224030.540414050.34060.140714080.140914100.14110.1 indicates data missing or illegible when filedwith Wolf-Guy Index positions 101 to 125 corresponding to heavy chain variable domain framework 1, positions 151 to 199 corresponding to CDR-H1, positions 201 to 214 corresponding to heavy chain variable domain framework 2, positions 251 to 299 corresponding to CDR-H2, positions 301 to 332 corresponding to heavy chain variable domain framework 3, positions 351 to 399 corresponding to CDR-H3, positions 401 to 411 corresponding to heavy chain variable domain framework 4 (positions 501 to 523 corresponding to light chain variable domain framework 1, positions 551 to 599 corresponding to CDR-L1, positions 601 to 615 corresponding to light chain variable domain framework 2, positions 651 to 699 corresponding to CDR-L2, positions 701 to 734 corresponding to light chain variable domain framework 3, positions 751 to 799 corresponding to CDR-L3, positions 801 to 810 corresponding to light chain variable domain framework 4).
[0105] In one embodiment the mutation risk score takes negative values, whereby larger negative values indicate a larger risk of loss of function.
[0106] In one embodiment the multitude (at least 10, at least 100, at least 1,000, at least 10,000) of amino acid sequences or nucleic acid sequences of variant antibodies specifically binding to the same target as the reference antibody are obtained from B-cells from the same immunization campaign as the reference antibody, wherein the B-cells have been enriched for antigen-specific antibody expressing B-cells. In one embodiment the enrichment is by antigen-specific sorting or / and cell panning or / and non-antigen specific antibody producing B-cells depletion.
[0107] One aspect as reported herein is a method for identifying a variant antibody of a reference antibody (that has comparable biological properties as the reference antibody) specifically binding to the same target / antigen comprising the following steps:
[0108] a) providing a multitude of B-cells producing / expressing (a multitude of) different antibodies binding to the same target / antigen as the reference antibody (obtained from the same animal species as the reference antibody);
[0109] b) isolating (and amplifying) from the multitude of B-cells the antibody encoding nucleic acids;
[0110] c) sequencing the antibody encoding nucleic acids by means of a next generation sequencing method;
[0111] d) generating a sequence alignment of the antibody encoding nucleic acid sequences by
[0112] i) aligning the VH or / and VL sequences based on
[0113] α) identical length of VH or / and VL, or / and
[0114] β) identical length of all 3-sheet framework regions, or / and
[0115] γ) identical length of all HVRs / CDRs, or / and
[0116] δ) identical HVR3 / CDR3 sequence, mutations in frameworks and / or HVRs / CDRs 1 or 2 (1 to 3 amino acid exchanges) allowed, or / and
[0117] ε) highly homologous HVR3 / CDR3 sequence (1 or 2 amino acid exchanges allowed) and identical HVR / CDR 1 and 2 (with mutations in frameworks allowed), or / and
[0118] ζ) mutations in frameworks and HVRs / CDRs are allowed to generated one or more sets of related VH / VL sequences;
[0119] ii) ranking the aligned sequences in the one or more sets of ii) by
[0120] α) the number of, and / or
[0121] β) the position(s) of, and / or
[0122] γ) the change of the physico-chemical properties resulting from the, and / or
[0123] δ) the difference of VH / VL orientation (angle) resulting from the amino acid difference(s) to the reference antibody sequence;
[0124] e) identifying one or more variant antibodies of the reference antibody;
[0125] f) determining one or more biological properties of the variant antibodies; and
[0126] g) selecting a variant antibody with comparable or improved one or more biological properties and thereby identifying a variant antibody of a reference antibody.
[0127] One aspect as reported herein is a method for identifying a variant antibody of a reference antibody (that has comparable biological properties as the reference antibody) specifically binding to the same target / antigen comprising the following steps:
[0128] (i) providing one or more DNA-containing samples that comprise a multitude of antibody encoding nucleic acids;
[0129] (ii) performing PCR amplification of regions of the antibody encoding nucleic acids in each of the samples of (i) using consensus sequence-specific primers, wherein the consensus sequence-specific primers bind to consensus sequences that are common to a plurality of genes within the multitude of antibody encoding nucleic acids, thereby generating a pool of amplification products;
[0130] (iii) sequencing said amplification products in order to determine the relative proportion of each nucleotide at each position (in a sequencing read);
[0131] (iv) performing a sequence alignment (based on sequence identity) between the sequencing (read) results of (iii) and at least one reference sequence, which reference sequence corresponds to an antibody having desirable properties; and
[0132] (v) identifying one or more variant antibodies of the reference antibody.
[0133] One aspect as reported herein is a method for identifying a variant antibody of a reference antibody (that has comparable biological properties as the reference antibody) specifically binding to the same target / antigen comprising the following steps:
[0134] a) amplifying one or more regions of interest from a biological sample comprising nucleic acids encoding antibody variable domains, wherein a plurality of amplicons for each region of interest are generated;
[0135] b) attaching an adapter and a random component to each amplicon generated in (a) and amplifying each of said extended amplicon;
[0136] c) sequencing the amplicons comprising the random component generated in (b), wherein redundant reads are generated and wherein the redundant reads are grouped by the random component, and identifying a consensus sequence;
[0137] d) comparing the consensus sequence to a reference sequence, wherein a consensus sequence that differs from the reference sequence comprises a mutation / variation; and
[0138] e) identifying one or more variant antibodies based on the results of step d).
[0139] In one embodiment steps a) to c) are
[0140] a) hybridizing a primer pool comprising one or more primer pairs specific to one or more regions of interest from a biological sample comprising nucleic acids encoding antibody variable domains, extending from an upstream primer of the primer pair to a downstream primer of the primer pair, and ligating the extension product to the downstream primer of the primer pair, wherein products comprising the regions of interest flanked by sequences required for amplification are generated;
[0141] b) attaching an adapter comprising a random component and attaching an adapter comprising an index sequence to the products from (a) and amplifying each of said extended sequences;
[0142] c) sequencing the products comprising the random component generated in (b), wherein redundant reads are generated and wherein the redundant reads are grouped by the random component and identifying a consensus sequence.
[0143] One aspect as reported herein is a method for selecting a variant of a reference antibody variable domain encoding nucleic acid, wherein the reference antibody variable domain amino acid sequence encoded by said encoding nucleic acid has at least one developability hot-spot (i.e. comprises at least one amino acid residue that is post-translationally modified resulting in a change of the biological properties (reduction of the binding affinity to its target / antigen) of the reference antibody comprising said reference antibody variable domain in one of its binding sites), the method comprising the following steps:
[0144] (i) providing a multitude of DNA-containing samples (genomic material of antibody secreting B-cell) each including one or more antibody variable domain encoding nucleic acids;
[0145] (ii) performing PCR amplification of the antibody variable domain encoding nucleic acids of (i) using consensus sequence-specific primers to obtain amplification products (wherein said consensus sequence-specific primers bind to consensus sequences that are common to a plurality of genes within the nucleic acids, thereby generating a pool of amplification products);
[0146] (iii) sequencing a plurality of the amplification products obtained in step (ii) in order to determine the relative proportion of each nucleotide at each position (in a sequencing read);
[0147] (iv) performing a sequence alignment (based on sequence identity) between the sequencing (read) results of (iii) and the reference antibody variable domain encoding nucleic acid;
[0148] (v) performing a sequence-identity / homology-based ranking of the antibody variable domain encoding nucleic acids in the sequence alignment with the reference antibody variable domain encoding nucleic acid being the perfect / template / reference sequence; and
[0149] (vi) selecting the variant antibody variable domain encoding nucleic acid based on the sequence ranking of step (v),
[0150] whereby the variant antibody variable domain selected in step (vi) is selected so that the developability hot-spot is removed (i.e. so that it comprises at least one amino acid residue less that is post-translationally modified compared to the reference antibody resulting in a reduced change of the biological properties (reduced reduction of binding affinity to its target / antigen) of the reference antibody comprising said reference antibody variable domain in one of its binding sites).
[0151] In one embodiment of all aspects the developability hot-spot (the amino acid reside that is post-translationally modified) is selected from the group consisting of one or more unpaired Cys-residues in the variable domain or in one of the HVRs, one or more glycosylation sites, an amino acid residue in an N- or O-glycosylation site, or / and one or more degradation hot-spots (Asp, Asn or Met).
[0152] One aspect as reported herein is a method for selecting a variant of a reference antibody variable domain, wherein the reference antibody variable domain has at least one developability hot-spot, the method comprising the following steps:
[0153] receiving sequencing data produced by sequencing nucleic acids from a multitude of B-cell clones each producing an antibody specifically binding to the same target as the reference antibody;
[0154] aligning the sequencing data with the sequence of the reference antibody variable domain being the perfect / template / reference sequence;
[0155] selecting a sequence that has the highest structural / functional identity / similarity to the reference antibody variable domain sequence but not having the developability hot-spot.
[0156] In one embodiment the developability hot-spot (the amino acid reside that is post-translationally modified) is selected from the group consisting of one or more unpaired Cys-residues in the variable domain or in one of the HVRs, one or more glycosylation sites, an amino acid residue in an N- or O-glycosylation site, or / and one or more degradation hot-spots (Asp, Asn or Met).
[0157] One aspect as reported herein is a method for selecting a variant of a reference antibody variable domain, wherein the reference antibody variable domain has at least one developability hot-spot, the method comprising the steps of one of the method as reported herein.
[0158] One aspect as reported herein is a method for producing an antibody comprising the following steps:
[0159] cultivating a cell comprising the nucleic acid encoding a variant antibody of a reference antibody, wherein the nucleic acid has been obtained with one of the methods as reported herein and all other nucleic acids required for the expression of a functional antibody,
[0160] recovering the antibody from the cell or the cultivation medium.
[0161] One aspect as reported herein is a cell comprising the nucleic acid obtained with one of the methods as reported herein.
[0162] It is expressly stated that each aspect may also be an embodiment of a different aspect.
[0163] The methods as reported herein can be used for the identification of alternative antigen specific antibody variants of a reference antibody without developability issues.
[0164] The method as reported herein can even be worked with polyclonal sera as by the use of NGS these are “monoclonalized”.DETAILED DESCRIPTION OF THE INVENTION
[0165] It will be readily understood that the embodiments, as generally described herein, are exemplary. The following more detailed description of various embodiments is not intended to limit the scope of the present disclosure, but is merely representative of various embodiments. Moreover, the order of steps or actions of the methods disclosed herein may be changed by those skilled in the art without departing from the scope of the present disclosure. In other words, unless a specific order of steps or actions is required for proper operation of the embodiment, the order or use of specific steps or actions may be modified.Definitions
[0166] General information regarding the nucleotide sequences of human immunoglobulins light and heavy chains is given in: Kabat, E. A., et al., Sequences of Proteins of Immunological Interest, 5th ed., Public Health Service, National Institutes of Health, Bethesda, MD (1991).
[0167] As used herein, the amino acid positions of all constant regions and domains of the heavy and light chain are numbered according to the Kabat numbering system described in Kabat, et al., Sequences of Proteins of Immunological Interest, 5th ed., Public Health Service, National Institutes of Health, Bethesda, MD (1991) and is referred to as “numbering according to Kabat” herein. Specifically, the Kabat numbering system (see pages 647-660) of Kabat, et al., Sequences of Proteins of Immunological Interest, 5th ed., Public Health Service, National Institutes of Health, Bethesda, MD (1991) is used for the light chain constant domain CL of kappa and lambda isotype, and the Kabat EU index numbering system (see pages 661-723) is used for the constant heavy chain domains (CH1, Hinge, CH2 and CH3, which is herein further clarified by referring to “numbering according to Kabat EU index” in this case).
[0168] It must be noted that as used herein and in the appended claims, the singular forms “a”, “an”, and “the” include plural reference unless the context clearly dictates otherwise. Thus, for example, reference to “a cell” includes a plurality of such cells and equivalents thereof known to those skilled in the art, and so forth. As well, the terms “a” (or “an”), “one or more” and “at least one” can be used interchangeably herein. It is also to be noted that the terms “comprising”, “including”, and “having” can be used interchangeably.
[0169] To a person skilled in the art procedures and methods are well known to convert an amino acid sequence, e.g. of a polypeptide, into a corresponding nucleic acid sequence encoding this amino acid sequence. Therefore, a nucleic acid is characterized by its nucleic acid sequence consisting of individual nucleotides and likewise by the amino acid sequence of a polypeptide encoded thereby.
[0170] The use of recombinant DNA technology enables the generation derivatives of a nucleic acid. Such derivatives can, for example, be modified in individual or several nucleotide positions by substitution, alteration, exchange, deletion or insertion. The modification or derivatization can, for example, be carried out by means of site directed mutagenesis. Such modifications can easily be carried out by a person skilled in the art (see e.g. Sambrook, J., et al., Molecular Cloning: A laboratory manual (1999) Cold Spring Harbor Laboratory Press, New York, USA; Hames, B. D., and Higgins, S. G., Nucleic acid hybridization—a practical approach (1985) IRL Press, Oxford, England).
[0171] Useful methods and techniques for carrying out the current invention are described in e.g. Ausubel, F. M. (ed.), Current Protocols in Molecular Biology, Volumes I to III (1997); Glover, N. D., and Hames, B. D., ed., DNA Cloning: A Practical Approach, Volumes I and 11 (1985), Oxford University Press; Freshney, R. I. (ed.), Animal Cell Culture—a practical approach, IRL Press Limited (1986); Watson, J. D., et al., Recombinant DNA, Second Edition, CHSL Press (1992); Winnacker, E. L., From Genes to Clones; N.Y., VCH Publishers (1987); Celis, J., ed., Cell Biology, Second Edition, Academic Press (1998); Freshney, R. I., Culture of Animal Cells: A Manual of Basic Technique, second edition, Alan R. Liss, Inc., N.Y. (1987).
[0172] The term “about” denotes a range of + / −20% of the thereafter following numerical value. In one embodiment the term about denotes a range of + / −10% of the thereafter following numerical value. In one embodiment the term about denotes a range of + / −5% of the thereafter following numerical value.
[0173] The term “glycan” denotes a polysaccharide, or oligosaccharide. Glycan is also used herein to refer to the carbohydrate portion of a glycoconjugate, such as a glycoprotein, glycolipid, glycopeptide, glycoproteome, peptidoglycan, lipopolysaccharide or a proteoglycan. Glycans usually consist solely of β-glycosidic linkages between monosaccharides. Glycans can be homo- or heteropolymers of monosaccharide residues, and can be linear or branched.
[0174] The term “antibody” herein is used in the broadest sense and encompasses various antibody structures, including but not limited to monoclonal antibodies, polyclonal antibodies, and multispecific antibodies (e.g., bispecific antibodies) so long as they exhibit the desired antigen-binding activity.
[0175] The term “next generation sequencing” as used herein denotes a method comprising massive parallel sequencing providing a sequence output much higher than that of traditional sequencing (e.g. traditional Sanger sequencing). This term is also often defined as “deep sequencing” or “second-generation sequencing” (Metzker, M. L., Nat. Rev. Genet. 11 (2010) 31-46; Mardis, E. R., Annu. Rev. Genom. Hum. Genet. 9 (2008) 387-402). The use of a next generation sequencing method allows obtaining of sequencing data in a very short time. Different technologies are commercially available for performing next generation sequencing, such as, for example, pyrosequencing (454 Life Sciences, Roche Diagnostics Corp., Basel, Switzerland), sequencing by synthesis (HiSeq™ and MiSeq™ Illumina, Inc., San Diego, CA), sequencing by ligation (SOLiD™, Life Technologies Corp. Logan, UT), Polonator sequencing, ion semiconductor sequencing (Ion PGM™, and Jon Proton™, Life Technologies Corp., Logan, UT), Ion Torrent sequencing, nanopore sequencing, single-molecule real-time sequencing (SMRT™, Pacific Biosciences, Menlo Park, CA), HeliScope Single Molecule sequencing, tunneling currents sequencing, sequencing by hybridization, mass spectrometry sequencing, microfluidic Sanger sequencing, RNA polymerase (RNAP) sequencing and others.
[0176] The term “sequencing platform” denotes a system for sequencing nucleic acids, including genomic DNA (gDNA), complementary DNA (cDNA) and RNA. The system may include one or more machines or apparatuses (e.g., amplification machines, sequencing machines, detection devices, etc.), data storage and analytical devices (e.g., hard drives, remote storage systems, processors, etc.), reagents (e.g., primers, probes, linkers, tags, NTPs, etc.) and particular methods for their use. For example, sequencing by synthesis and pyrosequencing and different platforms.
[0177] The term “read” denotes a single instance of determining the identity of a nucleotide at a particular position or the sequence of nucleotides in a particular polynucleotide. If a nucleotide or polynucleotide sequence is determined X times in a sequencing assay, there are “X reads” or a “read depth of X” or “read coverage of X” for that nucleotide or polynucleotide.
[0178] The term “hypervariable region” or “HVR”, as used herein, refers to each of the regions of an antibody variable domain which are hypervariable in sequence (“complementarity determining regions” or “CDRs”) and / or form structurally defined loops (“hypervariable loops”), and / or contain the antigen-contacting residues (“antigen contacts”). Generally, antibodies comprise six HVRs; three in the VH (H1, H2, H3), and three in the VL (L1, L2, L3).
[0179] HVRs herein include
[0180] (a) hypervariable loops occurring at amino acid residues 26-32 (L1), 50-52 (L2), 91-96 (L3), 26-32 (H1), 53-55 (H2), and 96-101 (H3) (Chothia, C. and Lesk, A. M., J. Mol. Biol. 196 (1987) 901-917);
[0181] (b) CDRs occurring at amino acid residues 24-34 (L1), 50-56 (L2), 89-97 (L3), 31-35b (H1), 50-65 (H2), and 95-102 (H3) (Kabat, E. A. et al., Sequences of Proteins of Immunological Interest, 5th ed. Public Health Service, National Institutes of Health, Bethesda, MD (1991), NIH Publication 91-3242.);
[0182] (c) antigen contacts occurring at amino acid residues 27c-36 (L1), 46-55 (L2), 89-96 (L3), 30-35b (H1), 47-58 (H2), and 93-101 (H3) (MacCallum et al. J. Mol. Biol. 262: 732-745 (1996)); and
[0183] (d) combinations of (a), (b), and / or (c), including HVR amino acid residues 46-56 (L2), 47-56 (L2), 48-56 (L2), 49-56 (L2), 26-35 (H1), 26-35b (H1), 49-65 (H2), 93-102 (H3), and 94-102 (H3).
[0184] Unless otherwise indicated, HVR residues and other residues in the variable domain (e.g., FR residues) are numbered herein according to Kabat et al., supra.
[0185] An “isolated” antibody is one, which has been separated from a component of its natural environment. In some embodiments, an antibody is purified to greater than 95% or 99% purity as determined by, for example, electrophoretic (e.g., SDS-PAGE, isoelectric focusing (IEF), capillary electrophoresis) or chromatographic (e.g., ion exchange or reverse phase HPLC). For review of methods for assessment of antibody purity, see, e.g., Flatman, S. et al., J. Chromatogr. B 848 (2007) 79-87.
[0186] An “isolated” nucleic acid refers to a nucleic acid molecule that has been separated from a component of its natural environment. An isolated nucleic acid includes a nucleic acid molecule contained in cells that ordinarily contain the nucleic acid molecule, but the nucleic acid molecule is present extrachromosomally or at a chromosomal location that is different from its natural chromosomal location.
[0187] The term “monoclonal antibody” as used herein refers to an antibody obtained from a population of substantially homogeneous antibodies, i.e., the individual antibodies comprising the population are identical and / or bind the same epitope, except for possible variant antibodies, e.g., containing naturally occurring mutations or arising during production of a monoclonal antibody preparation, such variants generally being present in minor amounts. In contrast to polyclonal antibody preparations, which typically include different antibodies directed against different determinants (epitopes), each monoclonal antibody of a monoclonal antibody preparation is directed against a single determinant on an antigen. Thus, the modifier “monoclonal” indicates the character of the antibody as being obtained from a substantially homogeneous population of antibodies, and is not to be construed as requiring production of the antibody by any particular method. For example, the monoclonal antibodies to be used in accordance with the present invention may be made by a variety of techniques, including but not limited to the hybridoma method, recombinant DNA methods, phage-display methods, and methods utilizing transgenic animals containing all or part of the human immunoglobulin loci.
[0188] The “class” of an antibody refers to the type of constant domain or constant region possessed by its heavy chain. There are five major classes of antibodies: IgA, IgD, IgE, IgG, and IgM, and several of these may be further divided into subclasses (isotypes), e.g., IgG1, IgG2, IgG3, IgG4, IgA1, and IgA2. The heavy chain constant domains that correspond to the different classes of immunoglobulins are called α, δ, ε, γ, and μ, respectively.
[0189] The term “N-linked oligosaccharide” denotes oligosaccharides that are linked to the peptide backbone at an asparagine amino acid residue, by way of an asparagine-N-acetyl glucosamine linkage. N-linked oligosaccharides are also called “N-glycans.” All N-linked oligo saccharides have a common pentasaccharide core of Man3GlcNAc2. They differ in the presence of, and in the number of branches (also called antennae) of peripheral sugars such as N-acetyl glucosamine, galactose, N-acetyl galactosamine, fucose and sialic acid. Optionally, this structure may also contain a core fucose molecule and / or a xylose molecule. N-linked oligosaccharides are attached to a nitrogen of asparagine or arginine side-chains. N-glycosylation motifs, i.e. N-glycosylation sites, comprise an Asn-X-Ser / Thr consensus sequence, where X is any amino acid except proline. Thus, an amino acid residue in an N-glycosylation site can be any amino acid residue in the Asn-X-Ser / Thr consensus sequence, where X is any amino acid except proline. In one embodiment is the amino acid residue in an N-glycosylation site Asn, Ser or Thr.
[0190] The term “O-linked oligosaccharide” denotes oligosaccharides that are linked to the peptide backbone at a threonine or serine amino acid residue. In one embodiment is the amino acid residue in an O-glycosylation site Ser or Thr.
[0191] The term “glycosylation state” denotes a specific or desired glycosylation pattern of an antibody. A “glycoform” is an antibody comprising a particular glycosylation state. Such glycosylation patterns include, for example, attaching one or more sugars at position N-297 of the Fc-region of an antibody (numbering according to Kabat), wherein said sugars are produced naturally, recombinantly, synthetically, or semi-synthetically. The glycosylation pattern can be determined by many methods known in the art. For example, methods of analyzing carbohydrates on proteins have been reported in US 2006 / 0057638 and US 2006 / 0127950 (the disclosures of which are hereby incorporated by reference in their entirety).
[0192] The term “variable region” or “variable domain” refers to the domain of an antibody heavy or light chain that is involved in binding the antibody to antigen. The variable domains of the heavy chain and light chain (VH and VL, respectively) of a native antibody generally have similar structures, with each domain comprising four conserved framework regions (FRs) and three hypervariable regions (HVRs). (See, e.g., Kindt, T. J. et al. Kuby Immunology, 6th ed., W.H. Freeman and Co., N.Y. (2007), page 91) A single VH or VL domain may be sufficient to confer antigen-binding specificity. Furthermore, antibodies that bind a particular antigen may be isolated using a VH or VL domain from an antibody that binds the antigen to screen a library of complementary VL or VH domains, respectively. See, e.g., Portolano, S. et al., J. Immunol. 150 (1993) 880-887; Clackson, T. et al., Nature 352 (1991) 624-628).
[0193] Alignment of a variant amino acid sequence or variant nucleic acid sequence with respect to a reference amino acid sequence or reference nucleic acid sequence can be done based on the “percent (%) sequence identity”. The “percent (%) sequence identity” is defined as the percentage of residues in a variant sequence that are identical with the residues in the reference sequence, after aligning the sequences and introducing gaps, if necessary, to achieve the maximum percent sequence identity. Alignment can be achieved in various ways that are within the skill in the art, for instance, using publicly available computer software such as BLAST, BLAST-2, ALIGN or Megalign (DNASTAR) software. Those skilled in the art can determine appropriate parameters for aligning sequences, including any algorithms needed to achieve maximal alignment over the full length of the sequences being compared. For purposes herein, however, % sequence identity values are generated using the sequence comparison computer program ALIGN-2. The ALIGN-2 sequence comparison computer program was authored by Genentech, Inc., and the source code has been filed with user documentation in the U.S. Copyright Office, Washington D.C., 20559, where it is registered under U.S. Copyright Registration No. TXU510087. The ALIGN-2 program is publicly available from Genentech, Inc., South San Francisco, California, or may be compiled from the source code. The ALIGN-2 program should be compiled for use on a UNIX operating system, including digital UNIX V4.0D. All sequence comparison parameters are set by the ALIGN-2 program and do not vary.
[0194] In situations where ALIGN-2 is employed for sequence comparisons, the % sequence identity of a given (variant) sequence A to, with, or against a given (reference) sequence B (which can alternatively be phrased as a given (variant) sequence A that has or comprises a certain % sequence identity to, with, or against a given (reference) sequence B) is calculated as follows:100 times the fraction X / Y where X is the number of residues scored as identical matches by the sequence alignment program ALIGN-2 in that program's alignment of A and B, and where Y is the total number of residues in B. It will be appreciated that where the length of sequence A is not equal to the length of sequence B, the % sequence identity of A to B will not equal the % sequence identity of B to A.The term “glycostructure” as used within this application denotes a single, defined N- or O-linked oligosaccharide at a specified amino acid residue. Thus, the term “antibody with a G1 glycostructure” denotes an antibody comprising at the asparagine amino acid residue at about amino acid position 297 according to the Kabat numbering scheme or in the FAB region a biantennary oligosaccharide comprising only one terminal galactose residue at the non-reducing ends of the oligosaccharide. The term “oligosaccharide” as used within this application denotes a polymeric saccharide comprising two or more covalently linked monosaccharide units.
[0196] The term “developability hot-spot” denotes an amino acid residue within the amino acid sequence of a polypeptide, such as e.g. an antibody variable domain, that is post-translationally modified, i.e. that is prone to post-translational modification. This post-translational modification results in a change, e.g. a reduction or loss, of at least one biological property, such as e.g. antigen binding in case of an antibody variable domain.
[0197] The term “post-translational modification” denotes a covalent modification of amino acid residues within a polypeptide following biosynthesis. Post-translational modifications can occur on the amino acid side chains by modifying an existing functional group or introducing a new one. Common post-translational modifications for Ala is N-acetylation; for Arg is deimination, methylation; for Asn is deamidation, N-linked glycosylation; for Asp is isomerization; for Cys is disulfide-bond formation, oxidation, palmitoylation, N-acetylation, S-nitrosylation; for Gln is cyclization; for Glu is cyclization, gamma-carboxylation; for Gly is N-myristoylation, N-acetylation; for His is phosphorylation; for Lys is acetylation, ubiquitination, SUMOylation, methylation, hydroxylation; for Met is N-acetylation, oxidation; for Pro is hydroxylation; for Ser is phosphorylation, O-linked glycosylation, N-acetylation; for Thr is phosphorylation, O-linked glycosylation, N-acetylation; for Trp is oxidation, formation of Kynurenine; for Tyr is sulfation, phosphorylation; for Val is N-acetylation.
[0198] The term “binding site of an antibody” denotes the pair of a light chain variable domain and a heavy chain variable domain.Antibody Glycosylation
[0199] Human antibodies are mainly glycosylated at the asparagine residue at about position 297 (Asn297) of the heavy chain CH2 domain or in the FAB region with a more or less fucosylated biantennary complex oligosaccharide (antibody amino acid residue numbering according to Kabat, supra). The biantennary glycostructure can be terminated by up to two consecutive galactose (Gal) residues in each arm. The arms are denoted (1,6) and (1,3) according to the glycoside bond to the central mannose residue. The glycostructure denoted as G0 comprises no galactose residue. The glycostructure denoted as G1 contains one or more galactose residues in one arm. The glycostructure denoted as G2 contains one or more galactose residues in each arm (Raju, T. S., Bioprocess Int. 1 (2003) 44-53). Human constant heavy chain regions are reported in detail by Kabat, supra, and by Brueggemann, M., et al., J. Exp. Med. 166 (1987) 1351-1361; Love, T. W., et al., Methods Enzymol. 178 (1989) 515-527. CHO type glycosylation of antibody Fc-regions is e.g. described by Routier, F. H., Glycoconjugate J. 14 (1997) 201-207.
[0200] An antibody in general comprises two so called full length light chain polypeptides (light chain) and two so called full length heavy chain polypeptides (heavy chain). Each of the full length heavy and light chain polypeptides contains a variable domain (variable region) (generally the amino terminal portion of the full length polypeptide chain) comprising binding regions, which interact with an antigen. Each of the full length heavy and light chain polypeptides comprises a constant region (generally the carboxyl terminal portion). The constant region of the full length heavy chain mediates the binding of the antibody i) to cells bearing a Fc gamma receptor (FcTR), such as phagocytic cells, or ii) to cells bearing the neonatal Fc receptor (FcRn) also known as Brambell receptor. It also mediates the binding to some factors including factors of the classical complement system such as component (Clq). The variable domain of a full length antibody's light or heavy chain in turn comprises different segments, i.e. four framework regions (FR) and three hypervariable regions (CDR). A “full length antibody heavy chain” is a polypeptide consisting in N-terminal to C-terminal direction of an antibody heavy chain variable domain (VH), an antibody constant domain 1 (CH1), an antibody hinge region, an antibody constant domain 2 (CH2), an antibody constant domain 3 (CH3), and optionally an antibody constant domain 4 (CH4) in case of an antibody of the subclass IgE. A “full length antibody light chain” is a polypeptide consisting in N-terminal to C-terminal direction of an antibody light chain variable domain (VL), and an antibody light chain constant domain (CL). The full length antibody chains a linked together via inter-polypeptide disulfide bonds between the CL-domain and the CH1 domain and between the hinge regions of the full length antibody heavy chains.
[0201] Is has been reported in recent years that the glycosylation pattern of antibodies, i.e. the saccharide composition and multitude of attached glycostructures, has a strong influence on the biological properties (see e.g. Jefferis, R., Biotechnol. Prog. 21 (2005) 11-16). Antibodies produced by mammalian cells contain 2-3% by mass oligosaccharides (Taniguchi, T., et al., Biochem. 24 (1985) 5551-5557). This is equivalent e.g. in an antibody of class G (IgG) to 2.3 oligosaccharide residues in an IgG of mouse origin (Mizuochi, T., et al., Arch. Biochem. Biophys. 257 (1987) 387-394) and to 2.8 oligosaccharide residues in an IgG of human origin (Parekh, R. B., et al., Nature 316 (1985) 452-457), whereof generally two are located in the Fc-region at Asn297 and the remaining in the variable region (Saba, J. A., et al., Anal. Biochem. 305 (2002) 16-31).
[0202] For the notation of the different N- or O-linked oligosaccharides the individual sugar residues are listed from the non-reducing end to the reducing end of the oligosaccharide molecule. The longest sugar chain is chosen as basic chain for the notation. The reducing end of an N- or O-linked oligosaccharide is the monosaccharide residue, which is directly bound to the amino acid of the amino acid backbone of the antibody, whereas the end of an N- or O-linked oligosaccharide, which is located at the opposite terminus as the reducing end of the basic chain, is termed non-reducing end.
[0203] All oligosaccharides are described with the name or abbreviation for the non-reducing saccharide (i.e., Gal), followed by the configuration of the glycosidic bond (a or p), the ring bond (1 or 2), the ring position of the reducing saccharide involved in the bond (2, 3, 4, 6 or 8), and then the name or abbreviation of the reducing saccharide (i.e., GlcNAc). Each saccharide is preferably a pyranose. For a review of standard glycobiology nomenclatures see, Essentials of Glycobiology Varki et al. eds., 1999, CSHL Press.
[0204] The term “defined glycostructure” denotes within this application a glycostructure in which the monosaccharide residue at the non-reducing ends of the glycostructure is of a specific kind. The term “defined glycostructure” denotes within this application a glycostructure in which the monosaccharide residue at the non-reducing end of glycostructures are defined and of a specific kind.Post-Translational Modification Prone Amino Acid Residues
[0205] Asn and Asp residues share a common degradation pathway that precedes via the formation of a cyclic succinimide intermediate. Succinimide formation results from an intramolecular rearrangement after deamidation of Asn or dehydration of Asp by nucleophilic attack of the backbone nitrogen of the succeeding amino acid on the Asn / Asp side chain γ-carbonyl group. The metastable cyclic imide can hydrolyze at either one of its two carbonyl groups to form aspartyl or iso-aspartyl linkages in different ratios, depending on hydrolysis conditions and conformational restraints. In addition, alternative degradation mechanisms were proposed such as nucleophilic attack by the backbone carbonyl oxygen to form a cyclic isoimide or direct water-assisted hydrolysis of Asn to Asp. Several analytical methods, mostly charge-sensitive methods such as ion exchange chromatography or isoelectric focusing, are known to a person skilled in the art to detect either of the degradation products, i.e. succinimide, Asp or isoAsp. Most suitable for the quantification and the localization of degradation sites in proteins is the analysis via liquid chromatography tandem mass spectrometry (LC-MS / MS).Wolfguy Numbering Scheme
[0206] The Wolfguy numbering defines CDR regions as the set union of the Kabat and Chothia definition. Furthermore, the numbering scheme annotates CDR loop tips based on CDR length (and partly based on sequence) so that the index of a CDR position indicates if a CDR residue is part of the ascending or the descending loop. A comparison with established numbering schemes is shown in the following Table.TABLENumbering of CDR-L3 and CDR-H3 using Chothia / Kabat (Ch-Kb), Honegger and Wolfguy numbering schemes. The latterhas increasing numbers from the N-terminal basis to theCDR peak and decreasing ones starting from the C-terminalCDR end. Kabat schemes fix the two last CDR residues andintroduce letters to accommodate for the CDR length. Incontrast to Kabat nomenclature, the Honegger numberingdoes not use letters and is common for VH and VL.326881028473032789103857313289010486732329911058773333092C887343319310789751332941089075235195109917533529611092754353971119375535498112947563559911395757356100 114 95a758357100a115 95b759358100b116 95c760359100c117 95d761360100d118 95e762361100e119 95f763362100f 120764363100g121765364100h122766384100i 123784385100j 124785386100k125786387100l 126787388127788389128789390129790391130791392131792393132793394133794395134795396135796397136797398101 13796798399102 13897799401103 F W98801402104 14099802403105 141100 803404106 142101 804Wolfguy VHCh-KbHoneggerCh-KbWolfguy VL
[0207] Wolfguy is designed such that structurally equivalent residues (i.e. residues that are very similar in terms of conserved spatial localization in the Fv structure) are numbered with equivalent indices as far as possible. This is illustrated in FIG. 1.
[0208] An example for a Wolfguy-numbered full-length VH and VL sequence can be found in the following Table.TABLEVH (left) and VL (right) sequence of the crystal structurewith PDB ID 3PP4 (21), numbered with Wolfguy, Kabatand Chothia. In Wolfguy, CDR-H1-H3, CDR-L2 and CDR-L3are numbered depending only on length, while CDR-L1is numbered depending on loop length and canonicalcluster membership. The latter is determined by calculatingsequence similarities to different consensus sequences.Here, we only give a single example of CDR-L1 numbering,as it is of no importance for generating our VH-VLorientation sequence fingerprint.WolfguyKabatChothiaPDB ID 3PP4 VHFramework 1101Q 1Q 1Q102V 2V 2V103Q 3Q 3Q104L 4L 4L105V 5V 5V106Q 6Q 6Q107S 7S 7S108G 8G 8G109A 9A 9A110E10E10E111V11V11V112K12K12K113K13K13K114P14P14P115G15G15G116S16S16S117S17S17S118V18V18V119K19K19K120V20V20V121S21S21S122C22C22C123K23K23K124A24A24A125S25S25SCDR-H1151G26G26G152Y27Y27Y153A28A28A154F29F29F155S30S30S156Y31Y31Y157.32S 31a.158.33W 31b.193.34I 31c.194.35N 31d.195. 35a. 31e.196S 35b.32S197W 35c.33W198I 35d.34I199N 35e.35NFramework 2201W36W36W202V37V37V203R38R38R204Q39Q39Q205A40A40A206P41P41P207G42G42G208Q43Q43Q209G44G44G210L45L45L211E46E46E212W47W47W213M48M48M214G49G49GCDR-H2251R50R50R252I51I51I253F52F52F254P 52aP 52aP255G 52b. 52b.256. 52c. 52c.286. 52d. 52d.287.53G53G288D54D54D289G55G55G290D56D56D291T57T57T292D58D58D293Y59Y59Y294N60N60N295G61G61G296K62K62K297F63F63F298K64K64K299G65G65GFramework 3301R66R66R302V67V67V303T68T68T304I69I69I305T70T70T306A71A71A307D72D72D308K73K73K309S74S74S310T75T75T311S76S76S312T77T77T313A78A78A314Y79Y79Y315M80M80M316E81E81E317L82L82L318S 82aS 82aS319S 82bS 82bS320L 82cL 82cL321R83R83R322S84S84S323E85E85E324D86D86D325T87T87T326A88A88A327V89V89V328Y90Y90Y329Y91Y91Y330C92C92C331A93A93A332R94R94RCDR-H3351N95N95N352V96V96V353F97F97F354D98D98D355G99G99G356.100 Y100 Y357.100aW100aW358.100bL100bL359.100c.100c.360.100d.100d.361.100e.100e.362.100f.100f.363.100g.100g.364.100h.100h.365.100i .100i .385.100j .*.386.100k.*.387.100l .*.388. 100m.*.389.100n.*.390.100o.*.391.100p.*.392.100q.*.393.100r.*.394.100s.*.395Y100t .*.396W100u.*.397L100v.*.398V101V101V399Y102Y102YFramework 4401W103W103W402G104G104G403Q105Q105Q404G106G106G405T107T107T406L108L108L407V109V109V408T110T110T409V111V111V410S112S112S411S113S113SPDB ID 3PP4 VLFramework 1501D 1D 1D502I 2I 2I503V 3V 3V504M 4M 4M505T 5T 5T506Q 6Q 6Q507T 7T 7T508P 8P 8P509L 9L 9L510S10S10S511L11L11L512P12P12P513V13V13V514T14T14T515P15P15P516G16G16G517E17E17E518P18P18P519A19A19A520S20S20S521I21I21I522S22S22S523C23C23CCDR-L1551R24R24R552S25S25S553S26S26S556K27K27K561S 27aS28S562L 27bL29L563L 27cL30L581H 27dH 30aH582S 27eS 30bS583N28N 30cN594G29G 30dG595I30I 30eI596T31T31T597Y32Y32Y598L33L33L599Y34Y34YFramework 2601W35W35W602Y36Y36Y603L37L37L604Q38Q38Q605K39K39K606P40P40P607G41G41G608Q42Q42Q609S43S43S610P44P44P611Q45Q45Q612L46L46L613L47L47L614I48I48I615Y49Y49YCDR-L2651Q50Q50Q652.*.*.653.*.*.692.*.*.693.*.*.694M51M51M695S52S52S696N53N53N697L54L54L698V55V55V699S56S56SFramework 3701G57G57G702V58V58V703P59P59P704D60D60D705R61R61R706F62F62F707S63S63S708G64G64G709S65S65S710G66G66G711S67S67S712G68G68G713.*.*.714.*.*.715T69T69T716D70D70D717F71F71F718T72T72T719L73L73L720K74K74K721I75I75I722S76S76S723R77R77R724V78V78V725E79E79E726A80A80A727E81E81E728D82D82D729V83V83V730G84G84G731V85V85V732Y86Y86Y733Y87Y87Y734C88C88CCDR-L3751A89A89A752Q90Q90Q753N91N91N754L92L92L755E93E93E756.94L94L757.95P95P758. 95a. 95a.793. 95b. 95b.794. 95c. 95c.795. 95d. 95d.796L 95e. 95e.797P 95f. 95f.798Y96Y96Y799T97T97TFramework 4801F98F98F802G99G99G803G100 G100 G804G101 G101 G805T102 T102 T806K103 K103 K807V104 V104 V808E105 E105 E809I106 I106 I810K107 / 106K107 KNext Generation Sequencing (NGS)(a) Sample Preparation
[0209] In certain aspects, the method as reported herein comprises, in part, amplifying one or more regions of interest from a biological sample comprising nucleic acid. The amplification generates a plurality of amplicons for each region of interest.
[0210] Amplification takes place in the presence of one or more primer pairs. A first primer of the primer pair comprises a sequence complementary to an upstream portion of the region of interest and a second primer of the primer pair comprises a sequence complementary to a downstream portion of the region of interest. The primer pairs are designed to anneal to complementary strands of nucleic acid (i.e. one primer of the primer pair anneals to the sense strand and one primer of the primer pair anneals to the antisense strand). The complementary sequence may be altered based on the region of interest to be amplified. The complementary sequences of the primer pair may comprise about 10 to about 100 nucleotides complementary to the region of interest.
[0211] One or more primer pairs are contacted with a sample comprising nucleic acid. Nucleic acid may be, for example, RNA or DNA. Modified forms of RNA or DNA may be used. In one exemplary embodiment, the nucleic acid is cDNA.
[0212] In general, amplification of the region of interest is carried out using polymerase chain reaction (PCR). A PCR reaction may comprise sample comprising nucleic acid, one or more primer pairs, polymerase, water, buffer, and deoxynucleotide triphosphates (dNTPs) in a single reaction vial. PCR may be performed according to standard methods in the art. By way of non-limiting example, the PCR reaction may comprise denaturation, followed by about 15 to about 30 cycles of denaturation, annealing and extension, followed by a final extension.(b) Sequencing Library Preparation
[0213] In certain aspects, the method as reported herein comprises, in part, attaching an adapter, and / or index sequence to each amplicon or product generated in Section (a).
[0214] In one embodiment, the nucleotide sequence comprising an adapter, a random component and / or an index sequence is attached to an amplicon or product via PCR. For example, the amplicon or product may be contacted with a nucleotide sequence comprising an adaptor and an index sequence and a PCR reaction is conducted. Then, this product is contacted with a nucleotide sequence comprising an adaptor and a PCR reaction is conducted. The resulting product is a nucleotide sequence comprising an adaptor, a region of interest, an index sequence and a downstream adaptor. Alternatively, the amplicon or product may be contacted with a nucleotide sequence comprising an adaptor and a random component and a PCR reaction is conducted. Then, this product is contacted with a nucleotide sequence comprising an adaptor and an index sequence and a PCR reaction is conducted. The resulting product is a nucleotide sequence comprising an adaptor, an index sequence, a region of interest, a random component and a downstream adaptor.
[0215] The products or amplicons comprising an adapter, a random component and / or an index sequence are then subjected to exponential PCR. In an embodiment, an exponential PCR reaction may comprise the products or amplicons comprising an adapter, a random component and / or an index sequence, primers, polymerase, water, buffer, and deoxynucleotide triphosphates (dNTPs) in a single reaction vial. Exponential PCR may be performed according to standard methods in the art. By way of non-limiting example, the exponential PCR reaction may comprise denaturation, followed by about 15-30 cycles of denaturation, annealing and extension, followed by a final extension.
[0216] Upon performing exponential PCR, the products or amplicons comprising an adapter, and / or an index sequence are amplified. The exponential PCR products comprise: an adapter, a region of interest, a downstream adapter and an index sequence.(c) Sequencing
[0217] In certain aspects, the method as reported herein comprises, in part, sequencing the exponential PCR product. The sequencing of the exponential PCR product generates redundant reads. The redundant reads are grouped by random component and a consensus sequence is identified such that the redundant reads mitigate sequence errors.
[0218] Sequencing may be performed according to standard methods in the art. Sequencing is preferably performed on a massively parallel sequencing platform, many of which are commercially available including, but not limited to Illumina, Roche / 454, Ion Torrent, and PacBIO. In an exemplary embodiment, Illumina sequencing is used.
[0219] Reads may be separated by the index sequence and trimmed to remove primer sequences. Reads may be grouped by the random component. In certain embodiments, groups of reads with less than three, less than four, or less than five reads may be removed. To eliminate ambiguous sequences, the random components may be sorted by abundance and clustered at an identity of about 85%. Alternatively, the random components may be sorted by abundance and clustered at an identity of about 65% to about 95%. The random components may be clustered from most abundant to least abundant. Given that most sequencing errors are random and that the correct sequence should occur more often than a variant with sequencing errors, the abundance-weighted clustering provides a means to eliminate spurious random components that are most likely due to sequencing errors while retaining the more abundant (and most likely true positive) random components.
[0220] This redundant sequencing of each amplicon or product allows the error-correction of each amplicon or product. For example, a consensus sequence is generated for each random component group by scoring and weighing the nucleotide at each base position. Sequences with a consensus sequence that is identical to the most abundant sequence associated with the same random component are kept; this process is called quality filtering. Specifically, at every position, the nucleotides called by each sequence read are compared and a consensus nucleotide is called if there is at least about 90% agreement between the reads. If there is less than about 90% agreement, an “N” is called in the consensus sequence at that position.(d) Comparison to Reference Sequence
[0221] After an error-corrected consensus sequence (ECCS) has been identified, the ECCS may be compared to a reference sequence to determine the presence of one or more differences. A reference sequence may be a sequence of an antibody for which variants with the same biological properties but without negative features, such as developability issues (hot-spots), are searched for.The Method as Reported Herein
[0222] In the art a back-up / replacement candidate is generally identified either by introducing mutations in the amino acid sequence to address, e.g., developability liabilities of the original candidate or by performing a de-novo screening of de-selected antibodies. In the first case it cannot be excluded that by removing the developability liability also the binding and / or therapeutic properties of the antibody can be affected. In the latter case it is questionable if a true replacement candidate can be found.
[0223] By using next generation sequencing (NGS) the number of B-cell clones that can be sequenced is dramatically increased compared to previous methods. But only clones showing promising properties in the early screening will be further pursued. For the other clones the sequence data will simply be stored. That is, not all of the sequenced clones will be expressed and fully characterized.
[0224] It has now been found that this unused potential of the NGS data obtained in a project can be used to identify one or more replacement candidates for a lead candidate, wherein the replacement candidates do not have e.g. the developability liability / liabilities of the original lead antibody, such as e.g. amino acid residues prone to post-translational modification or isolated cysteine-residues. This identification and selection is achieved by using the NGS sequence information to identify an antibody closely related to the selected candidate based on the amino acid / nucleotide sequence and at the same time maintaining at least one of the biological / binding properties of the lead candidate, i.e. the reference antibody comprising the developability hot-spot.
[0225] In one embodiment of all aspects, a high-diversity library is included within the original antigen-specific library, in order to increase library diversity and allow a better sequencing. For example, the genomic PhiX library, sold by Illumina Inc., may be used.
[0226] In one embodiment of all aspects, both the amplification and the sequencing steps are performed using the Illumina technology. In this case, during the amplification step specific subsequences, called adaptors and required for the subsequent step of sequencing, are added to the nucleotide fragments according to Illumina's instructions.
[0227] In one preferred embodiment of all aspects two PCRs are performed before sequencing.
[0228] Sequencing is in one embodiment performed using the commercially available MiSeq kit (Illumina Inc.).
[0229] In one embodiment of all aspects sequencing is performed using a paired-end sequencing method, in which both ends of a nucleotide fragment are simultaneously sequenced. This allows for a more rapid sequencing and is particularly useful in the case of long fragments.
[0230] In one embodiment of all aspects a multiplexing technique is used so that different samples can be sequenced during the same run, thus allowing a rapid and simultaneous analysis of many combinations of sample / library. To perform multiplexing, typically, specific sequences (barcodes or indexes) are added during an amplification step using specific PCR primers.
[0231] For example, the commercially available MiSeq kit provides for both paired-end sequencing and multiplexing.
[0232] Analysis of the sequencing data can be performed by a software able to process the data generated by the sequencing. Suitable software tools are commercially available. An example of a suitable software is BWA (Burrows-Wheeler Aligner), see the work of Li H. and Durbin R. (2009) (Li H. and Durbin R. (2009) Fast and accurate short read alignment with Burrows-Wheeler Transform. Bioinformatics, 25:1754-60). Another suitable software is FastQC (see website: www.bioinformatics.babraham.ac.uk / projects / fastqc / ), which is a tool for analyzing and controlling high throughput sequence data.
[0233] In other cases, in accordance with the methods as reported herein, one or more of the consensus sequence-specific primers may further comprise a tag and / or adaptor.
[0234] The current method comprises in one specific embodiment the following general steps:
[0235] assembly of paired (Illumina MiSeq) reads, e.g. by a computer program (optionally including corrections: Molecular Identifier Group-based Error Correction (MiGec) software pipeline that allows UMI-barcode extraction and error correction based on consensus building of the complete VH sequence);
[0236] extraction of antibody variable domains, optionally including correction: sequence replicas correction, considering only n≥3 CDR3 clusters, signal peptide detection and quality assessment;
[0237] structure-guided clustering and ranking of antibody variants; and
[0238] ranking mutations versus a reference sequence for their ability to retain VH / VL pairing, binding and function, optionally the ranking is based on matrices and depends also on the conservation of the mutation(s).Structure-Guided Clustering and Ranking of Antibody Variants (SCaRAb)Antibody Variable Region Variant Characterization
[0239] The method as reported herein is based on a given reference antibody for which the heavy chain variable region (VH) sequence and the light chain variable region (VL) sequence are known (subsequently denoted as the “reference” or “reference antibody”).
[0240] In addition, there exists a set of known VH and / or VL sequences that can be considered to be related to the reference antibody's VH and / or VL sequences (subsequently denoted as the “variants” or “variant antibodies”), e.g. derived from B-cells obtained in the same immunization campaign as the reference antibody. The relation between the reference and its variants is / can be based, for example, on i) identical lengths of VH and VL sequence, ii) identical lengths of all 3-sheet framework regions, or / and iii) identical lengths of all complementarity-determining regions (CDRs).
[0241] The criteria for the alignment / selection of the respective sequences depends on the available data pool. If a limited number of variants is available the criteria should be less stringent, whereas when a high number of variants is available the criteria can be / should be more stringent (in order to identify the best possible variant). For example, the number of mutations in the entire VH and / or VL could be used as criterion with less than 10, less than 9, less than 8, less than 7, less than 6, less than 5, less than 4 allowed. But it has to be noted that with the number of allowed mutations the likelihood that the variant has not the same / comparable binding properties as the reference antibody is increasing. Alternatively, it is possible to use criteria such as less than “x” mutations in HVR3 / CDR3, less than “y” mutations in all HVRs / CDRs, less than “z” mutations in all FRs, less than “n” mutations in total, or combinations thereof on amino acid or / and nucleic acid level.
[0242] If the criterion for the selection is the removal of developability hot-spots the alignment / selection criteria could be:
[0243] without (free) cysteine in the HVRs / CDRs
[0244] without degradation prone Asp, Asn, Met residues / motifs in the variable domains.
[0245] Both the reference sequences as well as their variants are then annotated with an established annotation scheme, such as e.g. the Wolfguy antibody numbering scheme (Bujotzek, A., et al., Prot. Struct. Funct. Bioinform., 83 (2015) 681-695). This is used throughout the method, i.e. it is used for all and any annotation of positions in the method as reported herein. An example for a Wolfguy-annotated VH reference sequence and its variants can be found in FIG. 1.
[0246] Once the annotated VH and VL sequences have been aligned based on the residue indices, the variants can be described in terms of a compact “mutation tuple”, which describes only the sequence differences with regard to the reference (i.e., type and location of substituted amino acids). This principle is illustrated in the following Table using VH variants of antibody 763.TABLEMutation tuple notation for six VH variants of reference antibody 763(see also FIG. 1). The mutation tuple is a concatenated series of strings,each specifying i) the amino acid type found in the reference sequence(1-letter code), ii) the Wolfguy position at which the amino acid islocated, and iii) the amino acid type of the sequence variant (1-lettercode), if it not matches the one of the reference sequence.Number ofexchangedVH variant nameamino acidsMutation tuple1101:11801:208941T322P1101:1836:149713Y293D, S295N, T322P1114:12508:88295Q102P, S103P, S125P, Y196S,R306K1114:28111:186805S103L, N197K, T322P, F329V,V331G2103:7194:163682E211A, R306K2106:2910:183421R306KClustering
[0247] Because a given set of sequence variants can be large and redundant, it is meaningful to perform a clustering that identifies variants that are identical or similar with regard to number, location and type of the respective amino acid substitutions. For this purpose, e.g., the well-established k-medoids clustering algorithm can be used.
[0248] For the current case (antibody variable domains), a suitable lower boundary for k is the number of different mutation tuple lengths in the dataset. For example, if the dataset to cluster consists of the variants (T322P), (R306K), (E211A,R306K), and (Q102P,S103P,S125P,Y196S,R306K), the minimum value for k should be set to 3, with allows for separate clusters for tuple length one, two, and five. If desired, the value of k can be increased further to realize a finer clustering based on where the amino acid substitutions are located.
[0249] To perform a k-medoids-based clustering of sequence variants described by mutation tuples, a novel distance metric that incorporates both the antibody-specific location of the amino acid substitution, as well as the approximate physico-chemical properties of the exchange, can be devised. The latter is quantified using the BLOSUM62 amino acid substitution matrix (Henikoff, S. and Henikoff, J. G., Proc. Nat. Acad. Sci., 89 (1992) 10915-10919).
[0250] This is done as outlined in the following.
[0251] Given the two exemplary mutation tuples (E211A,R306K) and (Y293D,S295N,T322P):
[0252] 1. Calculate the location-based distance matrix between the two tuples (distance in residue numbers: 293−211=82 etc.):E211AR306KY293D8213S295N8411T322P111162. Pick the mutation pair with the minimum distance: S295N and R306K (residue number distance of 11).
[0254] The distance contribution for this pair is (11+abs(BLOSUM62(S,N)−BLOSUM62(R,K)))2.
[0255] Update the distance matrix so that the taken pair is removed:E211AR306KY293D 82infS295NinfinfT322P111infRepeat procedure (step 2) until all mutations have been paired.3. If the number of mutations per tuple does not match, unpaired mutations will remain (in this example, T322P). To account for the unpaired mutation T322P, add to the distance(310+abs(BLOSUM62(T,T)−BLOSUM62(T,P)))2 The value 310 is the theoretical maximum distance in terms of VH Wolfguy indices (411-101). Repeat procedure (3) until all unpaired mutations are accounted for.In this example, the distance between the two mutation tuples is:mutation_tuple_dist[(E211A,R306K),(Y293D,S295N,T322P)]=sqrt((11+abs (BLOSUM62(S,N)-BLOSUM62(R,K)))2+(82+abs (BLOSUM62(Y,D)-BLOSUM62(E,A)))2+(310+abs (BLOSUM62(T,T)-BLOSUM62(T,P)))2)=sqrt(122+842+3162)≈327.194Mutation Risk ScoreThe mutation risk score is a means to quantify the risk that an antibody variant, specified by a mutation tuple as defined above, will lose the function that is displayed by the reference antibody (where the function is typically binding affinity towards a given target / antigen). The mutation risk score always takes negative values, and larger negative values indicate a larger risk of loss of function.The mutation risk score incorporates the BLOSUM62 matrix to rate the severity of the amino acid substitution in terms of the resulting changes in biochemical properties such as charge, hydrophobicity and size. Furthermore, the mutation risk score contains an antibody variable region position-specific weighting factor to account for the location where the amino acid substitution occurs. For example, residues belonging to the CDRs typically involved in antigen binding are weighted higher than peripheral residues that are unlikely to be involved in antigen binding.
[0262] For a given mutation tuple consisting of n mutation strings of the form XiposiYi, the mutation risk score is calculated as follows:Mutant Risk Score= ∑ni=1-1BLOSUM62(Xi,Yi)*weight(posi)}BLOSUM62(Xi,Yi)>0-2*weight (posi)}BLOSUM62(Xi,Yi)=0-2+BLOSUM62(Xi,Yi)*weight (posi)}BLOSUM62(Xi,Yi)<0
[0263] The following Table specifies the antibody-specific weights for the conserved positions of the VH domain. Residues that are not explicitly given in this Table are weighted with the value one. Residues are numbered with the Wolfguy numbering scheme.TABLEAntibody-specific weights for the conserved positions of theVH domain for the framework (left) and CDR regions (right).WolfguyWolfguyIndexWeightIndexWeightFramework 11010.2CDR-H115121021.11522.610301531.21040.51542.31050.21551.91060.81563.71070.515741080.815841090.419341100.21944111019541120.41963.311301973.91140.11982.61150.61993.41160.1CDR-H22513.81170.12521.91180.225341190.22513.81201.22553.612102564122428741230288412422893.91250.52903.4Framework 220142913.72022.62922.120322933.62041.72942.320512952.32060296120702971.22080.72982.42091.62990.52102CDR-H335142111.33523.52123.135342130.93543.52143.13553.5Framework 3010.135633021.235733031.735833040.335933051.83603306036133072.43623308036333091.23643334136533351366333613673337138233100.838333110.938433120.838533132.538633140.138733151.53883316038933171.539033180.839133190.139233201.439333210.239433220.33953.53230.239633241.83973.53250.63981.53262.2399333313270.632833292.533043312.93322.8Framework 44013.540224030.540414050.34060.140714080.140914100.14110.1 indicates data missing or illegible when filedABangle Distance
[0264] In addition to quantifying the mutation risk for a given variant, it is meaningful to assess if the variant is likely to preserve the VH-VL orientation of the reference antibody. The aim is to obtain antibody variants that retain the same antigen-binding properties and stability as the reference antibody. For this purpose, the VH-VL orientation has been characterized using the six ABangle orientation measures HL, HC1, LC1, HC2, LC2, and dc (five angular and one linear distance measure) defined by Dunbar et al. (Prot. Eng. Des. Sel. 26 (2013) 611-620). The individual ABangle values for the reference antibody and its variants are predicted using the machine learning-based approach described by Bujotzek et al. (Bujotzek, A., et al., Prot. Struct. Funct. Bioinform., 83 (2015) 681-695). In this approach, the parameters of VH-VL orientation are predicted from a sequence fingerprint of influential residues at the domain interface between VH and VL domain.The ABangle Concept
[0265] When making a comparison between any two amino acid based structures, generally distance-based metrics such as the root-mean-square deviation (RMSD) of equivalent atoms are used.
[0266] To characterize the orientation between any two three-dimensional objects, it is necessary to define:
[0267] a frame of reference on each object.
[0268] axes to measure orientation parameters about.
[0269] terminology to describe and quantify these parameters.
[0270] The ABangle concept is a method which fully characterizes VH-VL orientation in a consistent and absolute sense using five angles (HL, HC1, LC1, HC2 and LC2) and a distance (dc). The pair of variable domains of an antibody, VH and VL, is denoted collectively as an antibody Fv fragment.
[0271] In a first step antibody structures are extracted from a data bank (e.g. the protein data bank, PDB). Chothia antibody numbering (Chothia and Lesk, 1987) is applied to each of the antibody chains. Chains that are successfully numbered are paired to form Fv regions. This is done by applying the constraint that the H37 position Cα coordinate of the heavy chain (alpha carbon atom of the amino acid residue at heavy chain variable domain position 37) must be within 20 Å of the L87 position Cα coordinate of the light chain. A non-redundant set of antibodies is created using CDHIT (Li, W. and Godzik, A. Bioinformatics, 22 (2006) 1658-1659), applying a sequence identity cut-off over the framework of the Fv region of 99%.
[0272] The most structurally conserved residue positions in the heavy and light domains are used to define domain location. These positions are denoted as the VH and VL coresets. These positions are predominantly located on the β-strands of the framework and form the core of each domain. The coreset positions are given in the following Table:light chainlight chainheavy chainheavy chainL44L35H35H17L19L37H12H72L69L74H38H92L14L88H36H84L75L38H83H91L82L18H19H90L15L87H94H20L21L17H37H21L47L86H11H85L20L85H47H25L48L46H39H24L49L70H93H86L22L45H46H89L81L16H45H88L79L71H68H87L80L72H69H22L23L73H71H23L36H70
[0273] The coreset positions are used to register frames of reference onto the antibody Fv region domains.
[0274] The VH domains in the non-redundant dataset are clustered using e.g. CDHIT, applying a sequence identity cut-off of 80% over framework positions in the domain. One structure is randomly chosen from each of the 30 largest clusters. This set of domains is aligned over the VH coreset positions e.g. using Mammoth-mult (Lupyan, D., et al., Bioinf 21 (2005) 3255-3263). From this alignment the Cα coordinates corresponding to the eight structurally conserved positions H36, H37, H38, H39, H89, H90, H91 and H92 in the 0-sheet interface are extracted. Through the resulting 240 coordinates a plane is fitted. For the VL domain positions L35, L36, L37, L38, L85, L86, L87 and L88 are used to fit the plane.
[0275] The procedure described above allows mapping the two reference frame planes onto any Fv structure. Therefore, the measuring of the VH-VL orientation can be made equivalent to measuring the orientation between the two planes. To do this fully and in an absolute sense requires at least six parameters: a distance, a torsion angle and four bend angles. These parameters must be measured about a consistently defined vector that connects the planes. This vector is denoted C in the following. To identify C, the reference frame planes are registered onto each of the structures in the non-redundant set as described above and a mesh placed on each plane. Each structure therefore has equivalent mesh points and, thus, equivalent VH-VL mesh point pairs. The Euclidean distance is measured for each pair of mesh points in each structure. The pair of points with the minimum variance in their separation distance is identified. The vector which joins these points is defined as C.
[0276] The coordinate system is fully defined using vectors, which lie in each plane and are centered on the points corresponding to C. H1 is the vector running parallel to the first principal component of the VH plane, while H2 runs parallel to the second principal component. L1 and L2 are similarly defined on the VL domain. The HL angle is a torsion angle between the two domains. The HC1 and LC1 bend angles are equivalent to tilting-like variations of one domain with respect to the other. The HC2 and LC2 bend angles describe twisting-like variations of one domain to the other.
[0277] To describe the VH-VL orientation six measures are used, a distance and five angles. These are defined in the coordinate system as follows:
[0278] the length of C, dc,
[0279] the torsion angle, HL, from H1 to L1 measured about C,
[0280] the bend angle, HC1, between H1 and C,
[0281] the bend angle, HC2, between H2 and C,
[0282] the bend angle, LC1 between L1 and C, and
[0283] the bend angle, LC2, between L2 and C.
[0284] The term “VH-VL orientation” is used in accordance with its common meaning in the art as it would be understood by a person skilled in the art (see, e.g., Dunbar et al., Prot. Eng. Des. Sel. 26 (2013) 611-620; and Bujotzek, A., et al., Proteins, Struct. Funct. Bioinf, 83 (2015) 681-695). It denotes how the VH and VL domains orientate with respect to one another.
[0285] Thus the VH-VL orientation is defined by
[0286] the length of C, dc,
[0287] the torsion angle, HL, from H1 to L1 measured about C,
[0288] the bend angle, HC1, between H1 and C,
[0289] the bend angle, HC2, between H2 and C,
[0290] the bend angle, LC1 between L1 and C, and
[0291] the bend angle, LC2, between L2 and C,wherein reference frame planes are registered by i) aligning the Cα coordinates corresponding to the eight positions H36, H37, H38, H39, H89, H90, H91 and H92 of VH and fitting a plane through them and ii) aligning the Cα coordinates corresponding to the eight positions L35, L36, L37, L38, L85, L86, L87 and L88 of VL and fitting a plane through them, iii) placing a placed on each plane, whereby each structure has equivalent mesh points and equivalent VH-VL mesh point pairs, and iv) measuring the Euclidean distance for each pair of mesh points in each structure, whereby the vector C joins the pair of points with the minimum variance in their separation distance,wherein H1 is the vector running parallel to the first principal component of the VH plane, H2 is the vector running parallel to the second principal component of the VH plane, L1 is the vector running parallel to the first principal component of the VL plane, L2 is the vector running parallel to the second principal component of the VL plane, the HL angle is the torsion angle between the two domains, the HC1 and LC1 are the bend angles equivalent to tilting-like variations of one domain with respect to the other, and the HC2 and LC2 bend angles are equivalent to the twisting-like variations of one domain to the other.
[0292] The positions are determined according to the Chothia index.
[0293] The vector C was chosen to have the most conserved length over the non-redundant set of structures. The distance, dc, is this length. It has a mean value of 16.2 Å and a standard deviation of only 0.3 Å.
[0294] The following Table lists the top 10 positions and residues identified by the random forest algorithm as being important in determining each of the angular measures of VH-VL orientation.TABLEX represents the variableL36V / L38E / L42H / L43L / L44F / L45T / L46G / L49G / L95HAngletop 10 important input variablesHLL87F L42G / L43T L44V H61D L89L H43Q H43N H44KH62K / H89V L55H L53RHC1X L56P L41D L89A L97V L94N L34H L34N L96W L100AHC2H62S H62K / H89V H43K H50W H46K / H62D H35S H61QH43Q H33W H58TLC1L91W L89A X L97V L94N L50G H43Q L56P H62Sb L55ALC2L50Y L42G / L43T L44V L42Q L55H H99Y L93T L94L L53R L85T(for more detailed information see Dunbar, J., et al., Protein Eng. Des. Sel., 26 (2013) 611-620 and Bujotzek, A., et al., Prot. Struct. Funct. Bioinf. 83 (2015) 681-695, which are incorporated by reference in their entirety herewith).
[0295] Thereby a fast sequence-based predictor that predicts VH-VL-interdomain orientation is provided. The VH-VL-orientation is described in terms of the six absolute ABangle parameters to precisely separate the different degrees of freedom of VH-VL-orientation. The deviation between two sequences / structures is shown by the average root-mean-square deviation (RMSD) of the carbonyl atoms of the amino acid backbone.
[0296] In one embodiment of all aspects as reported herein the VH / VL orientation is determined as follows:
[0297] generating from the multitude of variant antibody sequences and for the reference antibody Fv fragments,
[0298] determining the VH-VL-orientation for the reference Fv fragment and for each of the variant antibody Fv fragments of the multitude of variant antibody Fv fragments based on a sequence fingerprint of the antibody Fv fragment,
[0299] identifying / selecting / obtaining / ranking those variant antibody Fv fragments that have the smallest difference in the VH-VL-orientation compared to the reference antibody's VH-VL-orientation.
[0300] In one embodiment the method comprising the following step:
[0301] identifying / selecting / obtaining / ranking those variant antibody Fv fragments that have the highest (structural) similarity in the VH-VL-interdomain angle compared to the reference antibody's VH-VL-interdomain angle.
[0302] In one embodiment a VH-VL-interface residue is an amino acid residue whose side chain atoms have neighboring atoms of the opposite chain with a distance of less than or equal to 4 Å (in at least 90% of all superimposed Fv structures).
[0303] In one embodiment the set of VH-VL-interface residues comprises residues 210, 296, 610, 612, 733 (numbering according to Wolfguy index).
[0304] In one embodiment the set of VH-VL-interface residues comprises residues 199, 202, 204, 210, 212, 251, 292, 294, 295, 329, 351, 352, 354, 395, 396, 397, 398, 399, 401, 403, 597, 599, 602, 604, 609, 610, 612, 615, 651, 698, 733, 751, 753, 796, 797, 798 (numbering according to Wolfguy index).
[0305] In one embodiment the set of VH-VL-interface residues comprises residues 197, 199, 208, 209, 211, 251, 289, 290, 292, 295, 296, 327, 355, 599, 602, 604, 607, 608, 609, 610, 611, 612, 615, 651, 696, 698, 699, 731, 733, 751, 753, 755, 796, 797, 798, 799, 803 (numbering according to Wolfguy index).
[0306] In one embodiment the set of VH-VL-interface residues comprises residues 197, 199, 202, 204, 208, 209, 210, 211, 212, 251, 292, 294, 295, 296, 327, 329, 351, 352, 354, 355, 395, 396, 397, 398, 399, 401, 403, 597, 599, 602, 604, 607, 608, 609, 610, 611, 612, 615, 651, 696, 698, 699, 731, 733, 751, 753, 755, 796, 796, 797, 798, 799, 801, 803 (numbering according to Wolfguy index).
[0307] In one embodiment the set of VH-VL-interface residues comprises residues 199, 202, 204, 210, 212, 251, 292, 294, 295, 329, 351, 352, 354, 395, 396, 397, 398, 399, 401, 403, 597, 599, 602, 604, 609, 610, 612, 615, 651, 698, 733, 751, 753, 796, 797, 798, 801 (numbering according to Wolfguy index).
[0308] In one embodiment the set of VH-VL-interface residues comprises residues 197, 199, 202, 204, 208, 209, 210, 211, 212, 251, 292, 294, 295, 296, 327, 329, 351, 352, 354, 355, 395, 396, 397, 398, 399, 401, 403, 597, 599, 602, 604, 607, 608, 609, 610, 611, 612, 615, 651, 696, 698, 699, 731, 733, 751, 753, 755, 796, 797, 798, 799, 801, 803 (numbering according to Wolfguy index).
[0309] In one embodiment the identifying / selecting / obtaining / ranking is based on the top 80% variant antibody Fv fragments regarding VH-VL-orientation.
[0310] In one embodiment the identifying / selecting / obtaining / ranking is of the top 20% variant antibody Fv fragments regarding VH-VL-orientation.
[0311] In one embodiment the VH-VL-orientation is determined by calculating the six ABangle VH-VL-orientation parameters.
[0312] In one embodiment the VH-VL-orientation is determined by calculating the ABangle VH-VL-orientation parameters using a random forest method.
[0313] In one embodiment the VH-VL-orientation is determined by calculating the ABangle VH-VL-orientation parameters using one random forest method for each ABangle.
[0314] In one embodiment the VH-VL-orientation is determined by calculating the habitual torsion angle, the four bend angles (two per variable domain), and the length of the pivot axis of VH and VL (HL, HC1, LC1, HC2, LC2, dc) using a random forest model.
[0315] In one embodiment the random forest model is trained only with complex antibody structure data.
[0316] In one embodiment the highest structural similarity is the lowest average root-mean-square deviation (RMSD). In one embodiment the RMSD is the RMSD determined for all Calpha atoms (or carbonyl atoms) of the amino acid residues of the non-human or parent antibody to the corresponding Calpha atoms of the variant antibody.
[0317] In one embodiment a model assembled from template structures aligned on either consensus VH or VL framework, followed by VH-VL reorientation on a consensus Fv framework is used for determining the VH-VL-orientation.
[0318] In one embodiment a model aligned on the β-sheet core of the complete Fv (VH and VL simultaneously) is used for determining the VH-VL-orientation.
[0319] In one embodiment a model in which the antibody Fv fragment is reoriented on a consensus Fv framework is used for determining the VH-VL-orientation.
[0320] In one embodiment a model using template structures aligned onto a common consensus Fv framework and VH-VL orientation not being adjusted in any form is used for determining the VH-VL-orientation.
[0321] In one embodiment a model assembled from template structures aligned on either consensus VH or VL framework, followed by VH-VL reorientation on a VH-VL orientation template structure chosen based on similarity is used to determine the VH-VL-orientation.
[0322] Once the ABangle values for the reference antibody and its variants have been determined, one can rank the variants according to their similarity with regard to the VH-VL orientation of the reference antibody. In order to compare similarity in ABangle space, we define a set of ABangle parameters as the tuple θ:=(HL, HC1, LC1, HC2, LC2, dc):=(ϑ1, ϑ2, ϑ3, ϑ4, ϑ5, ϑ6). The Euclidean distance between two sets of ABangle parameters is thendistABangle(θa,θb)=∑ i=16(ϑia-ϑib)2.
[0323] As distABangle mingles angular (HL, HC1, LC1, HC2, LC2) with linear (dc) distance measures, they cannot be interpreted as factual distance in angular space, but serve only as an abstract distance measure.Filtering
[0324] Depending on the application of SCaRAb, the mutation tuples can be screened for certain sequence features (e.g., removal of a potential glycosylation site that is present in the reference antibody) or liabilities (e.g., introduction of a new free cysteine residue that is not present in the reference antibody) and filters can be applied accordingly.Example AIsolation and Properties of B-Cell Cloning Binders (BCC Binders, Binding ELISA)
[0325] At day 6 after the 3rd immunization 15 ml blood were harvested from the immunized rabbit R176 and 108.6×10E6 PBMC were isolated. For the NGS of VHs 4.2×10E6 PBMC were resuspended in RLT Buffer, whereas the remaining PBMCs were further processed (macrophage depletion, enrichment on antigen) for isolation of antigen specific B-cell clones by B-cell cloning process (see e.g. Seeber, S., et al, PLoS One, 9 (2014) e86184).
[0326] In total, during this B-cell cloning process 504 single B-cells of bleed 1 of animal R176 were deposited and cultivated after macrophage / KLH-binder depletion and 840 single B-cells were deposited and cultivated after enrichment on the antigen (LRP8, Low-density lipoprotein receptor-related protein 8, UniProtKB—Q14114).
[0327] After macrophage depletion, the primary screening identified 279 IgG-secreting B-cell clones. Of these clones 4 supernatants bound to the antigen. After antigen enrichment 341 IgG-positive supernatants could be identified. Among them 55 B-cell supernatants bound to the antigen (see following Table).TABLEIsolation and properties of binders identified by B-cell cloningrbIgGantigenantigenwellsIgG+[% totalantigen[% total[% IgG+Cell Treatmenttotalwellswells][n]wells]wells]Macrophage5042795540.81.4depletion(SA_KLHbiot. neg.)antigen-84034141556.516.1specificenrichmentSA_KLH biot.neg.,SA_antigenpos.Identification in NGS Data Pool Variants of Reference Antibody (BCC Binder Variant)
[0328] Totally 4 binders identified by B-cell cloning were chosen for identification in NGS repertoires of VHs variants; all clones were isolated by B-cell cloning after antigen specific enrichment, all exhibited specificity for the antigen (EC50 below 20 ng / ml) and two of them revealed cross-reactivity to the murine antigen (EC50 below 20 ng / ml) (see following Table).TABLEProperties of B-cell clones selected for NGS variants analysisbindingEC50EC50specificityhu-antigenmu-antigenclone(ELISA)[ng / ml][ng / ml]BCC.755hu<20>2000BCC.763hu; mu<20<20BCC.770hu<20>2000BCC.776hu; mu<20<20
[0329] NGS repertoire from PBMCs and from antigen-enriched B-cells was analyzed for identification of VHs variants with ≤6 amino acid replacements in the entire VH (FR1 to FR4, mutations in frameworks and HVR / CDRs are allowed) compared to the VH of the reference B-cell binders. Totally 441 diverse VH variants could be identified, with different distribution for the 4 binders and for the NGS sample delivering the variants, as shown in the following Table.TABLERelated VHs, identified in NGS samplesnumber of related VHs,identified in NGSsamples, with ≤6 AAreplacements in VHclone(FR1 to FR4)fromBCC.755322antigen enriched B-cellsBCC.7637PBMCBCC.77075antigen enriched B-cellsBCC.77637PBMC and antigen enrichedB-cellsStructure-Guided Clustering and Ranking of Antibody Variants (SCaRAb)
[0330] As outlined above a clustering of the identified variants had been performed. The results are presented in the following tables. Note that variants representing cluster medoids, i.e., the variant in the cluster whose average dist_ABangle to all other variants in the cluster is minimal, have been highlighted with a grey background. The generally known k-medoids-based-clustering has been extended in the method as reported herein by using a distance-related function. These variants can be interpreted as the representative or exemplar of a given cluster of variants.VH variantnumber ofriskABanglefreeClusternamemutationsmutation tuplescoredistancecysteineindexVH variants of antibody 755 (40 clusters)2107:3548:3S307T, F329C, A394G−18.400.19Y016900 1:N:0:32114:20056:3S307T, C330V, A394G−20.400.00Y025187 1:N:0:32116:13408:3F302C, S307T, A394G−13.200.00Y01588 1:N:0:31104:8159:3S307T, C330R, A394G−28.400.00Y014989 1:N:0:41107:18832:3S307T, C330W, A394G−24.400.00Y06130 1:N:0:41110:14978:3S307T, C330Y, A394G−24.400.00Y09932 1:N:0:41111:19785:3S307T, S319R, A394C−8.700.00Y01275 1:N:0:42112:23323:3Y293C, S307T, A394G−22.800.00Y01570 1:N:0:42113:10625:3W296C, S307T, A394G−12.400.28Y013290 1:N:0:42105:4626:3S307T, S319R, A394G−8.700.00N05248 1:N:0:12111:25558:3D314N, S319T, A394G−6.200.00N019928 1:N:0:11113:18172:3S307T, A331S, A394G−11.300.00N023110 1:N:0:22103:22100:3S307T, A318V, A394G−10.000.00N05261 1:N:0:21101:9654:3S307T, S308T, A394G−8.400.00N014883 1:N:0:31103:11925:3S307T, F329V, A394G−15.900.19N07788 1:N:0:31103:18092:3S307T, T311N, A394G−10.200.00N023861 1:N:0:31104:18405:3F302L, S307T, A394G−10.800.00N016109 1:N:0:31106:26206:3I304N, S307T, A394G−9.900.00N07674 1:N:0:31106:22579:3S307T, A318D, A394G−11.600.00N020023 1:N:0:31107:23439:3S307T, L315Q, A394G−14.400.00N05644 1:N:0:31107:18066:3S307T, E323A, A394G−9.000.00N09061 1:N:0:31109:10987:3F302V, S307T, A394G−12.000.00N015579 1:N:0:31111:4272:3S307T, T325P, A394G−10.200.00N05585 1:N:0:31111:22960:3S307T, D324A, A394G−15.600.00N014878 1:N:0:31115:9261:3S307T, M317L, A394G−9.150.00N011684 1:N:0:31117:8419:3S307T, E323D, A394G−8.500.00N021526 1:N:0:32101:21589:3F302S, S307T, A394G−13.200.00N09999 1:N:0:32104:28355:3S307T, T327A, A394G−9.600.36N08492 1:N:0:32104:10919:3S307T, L320I, A394G−9.100.00N017808 1:N:0:32105:5190:3S307T, T321A, A394G−8.800.00N016629 1:N:0:32106:3273:3S307T, L320P, A394G−15.400.00N011170 1:N:0:32106:2261:3S307T, S319G, A394G−8.600.00N014587 1:N:0:32108:8221:3S307T, R332G, A394G−19.600.00N022107 1:N:0:32110:7760:3S307T, T327S, A394G−9.000.36N022092 1:N:0:32111:22747:3S307T, A326D, A394G−17.200.00N014477 1:N:0:32117:13720:3S307T, A326S, A394G−10.600.00N012053 1:N:0:31102:19406:3S307T, D324E, A394G−9.300.00N03778 1:N:0:41102:25932:3S305T, S307T, A394G−10.200.00N014332 1:N:0:41105:10046:3S307T, T325A, A394G−9.600.00N021225 1:N:0:41108:3516:3R301L, S307T, A394G−8.800.00N011106 1:N:0:41110:29049:3S307T, D324N, A394G−10.200.00N09990 1:N:0:41111:15125:3S307T, S309T, A394G−9.600.00N08276 1:N:0:41112:15902:3S307T, D314E, A394G−8.450.00N011470 1:N:0:42101:14383:3S307T, T325K, A394G−10.200.00N09526 1:N:0:42101:22864:3S307T, S319T, A394G−8.500.00N022719 1:N:0:42102:7971:3S307T, V313L, A394G−10.900.00N07987 1:N:0:42102:11457:3S307T, D314V, A394G−8.900.00N011982 1:N:0:42103:27670:3I304T, S307T, A394G−9.300.00N06132 1:N:0:42103:5454:3S307T, L315R, A394G−14.400.00N017222 1:N:0:42105:17113:3S307T, Y328H, A394G−9.900.00N05256 1:N:0:42108:19861:3S307T, D314G, A394G−8.700.00N03706 1:N:0:42108:27997:3F302I, S307T, A394G−10.800.00N09341 1:N:0:42110:8959:3R301G, S307T, A394G−8.800.00N012221 1:N:0:42110:8552:3T303S, S307T, A394G−10.100.00N014278 1:N:0:42113:10355:3S307T, D314K, A394G−8.700.00N019267 1:N:0:42116:3784:3S307T, R332S, A394G−16.800.00N09465 1:N:0:42105:19779:3T291K, S307T, A394G−19.500.00N01879 1:N:0:21101:6229:3K298T, S307T, A394G−15.600.00N08183 1:N:0:31104:13381:3W296G, S307T, A394G−12.400.28N024822 1:N:0:31105:13132:3T290N, S307T, A394G−15.200.15N024432 1:N:0:31106:8967:3Y293D, S307T, A394G−26.400.00N021515 1:N:0:31107:21991:3N295T, S307T, A394G−13.000.13N04193 1:N:0:31110:20533:3A297S, S307T, A394G−9.600.00N02171 1:N:0:31111:6033:3Y293F, S307T, A394G−9.600.00N018835 1:N:0:31112:4770:3K298I, S307T, A394G−20.400.00N017947 1:N:0:31118:25241:3T291N, S307T, A394G−15.800.00N010586 1:N:0:31118:15224:3K298N, S307T, A394G−13.200.00N020798 1:N:0:32102:20616:3T291P, S307T, A394G−19.500.00N018534 1:N:0:32105:24474:3A297D, S307T, A394G−13.200.00N08007 1:N:0:32109:12894:3T291S, S307T, A394G−12.100.00N05370 1:N:0:32113:16261:3A297E, S307T, A394G−12.000.00N020702 1:N:0:32117:19754:3Y292S, S307T, A394G−16.800.21N04077 1:N:0:31101:21506:3S288R, S307T, A394G−20.400.00N014114 1:N:0:41105:16900:3G299S, S307T, A394G−9.400.00N024181 1:N:0:41107:16371:3A297G, S307T, A394G−10.800.00N023843 1:N:0:41109:15849:3N295Y, S307T, A394G−17.600.00N020837 1:N:0:41112:15276:3W296R, S307T, A394G−13.400.30N022867 1:N:0:41117:27611:3N295K, S307T, A394G−13.000.00N05863 1:N:0:42105:22497:3A294V, S307T, A394G−13.000.34N025166 1:N:0:42108:13105:3W296L, S307T, A394G−12.400.28N04949 1:N:0:42110:13214:3K298Q, S307T, A394G−10.800.00N04092 1:N:0:42112:3904:3Y292T, S307T, A394G−16.801.13N012465 1:N:0:42114:8546:3K298R, S307T, A394G−9.600.00N013514 1:N:0:42105:20717:5V104L, N155S, D196Y, −45.200.07N117101 1:N:0:4R197V, G199S2117:26776:5I152F, D153S, N155S, −37.700.17N118522 1:N:0:3D196Y, R197A1107:22265:1C122R−20.000.00Y217831 1:N:0:42109:15764:1Q102P−3.300.00N210046 1:N:0:32109:11666:1V104L−0.500.00N222695 1:N:0:32115:18433:1T123I0.000.00N220332 1:N:0:31110:4244:1V112D−2.000.00N213646 1:N:0:42114:5742:1L118M−0.100.00N218170 1:N:0:41103:11280:5S307T, S352C, A394G, −34.400.38Y323651 1:N:0:4F397C, N398D1117:15314:4V104L, T113M, S307T, −8.900.00N517723 1:N:0:4A394G2103:5849:4V104L, P117T, S307T, −9.200.00N510044 1:N:0:4A394G2105:28283:4R110G, N155K, S307T, −13.000.00N512169 1:N:0:4A394G2113:9804:4V104L, N155S, S307T, −10.800.00N521625 1:N:0:4A394G1101:25753:4S307T, C330W, A331S, −27.300.00Y76416 1:N:0:2A394G1101:15616:4S307T, L320M, C330W−25.100.00Y711852 1:N:0:4A394G2108:8173:4S307T, A326D, C330R, −37.200.00Y723018 1:N:0:4A394G2103:21863:4S307T, C330G, Y351D, −48.400.27Y77444 1:N:0:4A394G2117:7027:4S307T, C330G, N354Y, −42.400.08Y712367 1:N:0:4A394G1109:12708:4S305T, S307T, T321P, −10.800.00N724556 1:N:0:2A394G2105:10840:4S305A, S307T, A318D, −13.400.00N718512 1:N:0:3A394G2119:17571:4S307T, T322N, F329L, −14.000.19N78893 1:N:0:3A394G1101:27737:4S307T, A318G, A326D, −18.800.00N714174 1:N:0:4A394G1103:15990:4S307T, D314V, R332G, −20.100.00N716033 1:N:0:4A394G1103:15010:4S305T, S307T, S319R, −10.500.00N716228 1:N:0:4A394G1109:28899:4S305P, S307T, S319R, −14.100.00N717376 1:N:0:4A394G1114:14635:4S307T, S319G, T325P, −10.400.00N723670 1:N:0:4A394G2108:14119:4S307T, D314E, T321A, −8.850.00N74451 1:N:0:4A394G1110:18648:4Y293S, F302V, S307T, −26.400.00N72202 1:N:0:2A394G2105:16971:4G299S, S307T, L315P, −16.900.00N75020 1:N:0:2A394G2114:12873:4N295K, S307T, D314Y, −13.500.00N715745 1:N:0:2A394G1102:21561:4T291H, S307T, D314G, −23.500.00N72675 1:N:0:3A394G1103:15418:4A294G, F302Y, S307T, −13.400.36N710210 1:N:0:3A394G1116:13513:4Y293D, F302Y, S307T, −26.800.00N717404 1:N:0:3A394G1117:14122:4S288R, S305P, S307T, −25.800.00N724446 1:N:0:3A394G1118:12162:4A297E, F302I, S307T, −14.400.00N722228 1:N:0:3A394G2101:28389:4K298N, F302L, S307T, −15.600.00N716276 1:N:0:3A394G2106:26242:4K298N, R301L, S307T, −13.600.00N721852 1:N:0:3A394G2113:6396:4N295T, F302I, S307T, −15.400.13N76995 1:N:0:3A394G1113:7787:4K298E, F302I, S307T, −13.200.00N714034 1:N:0:4A394G2108:16555:4S288G, S307T, D314E, −16.450.00N720565 1:N:0:4A394G2116:26684:4K298N, F302I, S307T, −15.600.00N76924 1:N:0:4A394G2114:14700:4W296G, F302T, S307T, −17.200.28N723845 1:N:0:5A394G1106:4127:4N295H, K298R, S307I, −19.100.00N715400 1:N:0:3A394G1106:16406:4N295K, A297S, S307T, −14.200.00N715444 1:N:0:3A394G1113:12273:4S288R, T290N, S307T, −27.200.15N721196 1:N:0:3A394G1106:9726:4A294S, G299A, S307T, −11.700.35N722782 1:N:0:4A394G2107:10746:4T291P, N295T, S307T, −24.100.13N722307 1:N:0:4A394G2101:25197:4S307T, S319R, N354K, −15.700.03N73708 1:N:0:3A394G2105:4942:4S307T, E323A, Y351L, −21.000.17N77894 1:N:0:3A394G2112:4654:4S307T, E323G, N354I, −26.700.08N76406 1:N:0:3A394G1101:27729:4S307T, T322P, T353P, −21.300.00N714154 1:N:0:4A394G1109:26370:4S307T, L320R, S355R, −24.500.22N74240 1:N:0:5A394V2117:9077:4S307T, T325L, N354D, −13.700.17N716439 1:N:0:5A394G1107:23167:5P117L, Y292S, S307T, −20.000.21N83797 1:N:0:3T311P, A394G1113:10964:1P403Q−1.500.70N914412 1:N:0:31118:15879:1A394T−6.000.00N920254 1:N:0:32115:26912:5T116G, P117T, T123I, −9.100.00N1016126 1:N:0:4S307T, A394G1119:1879:4E105V, S307T, C330Y, −25.200.00Y1113689 1:N:0:4A394G1104:18713:4G108R, S307T, T325A, −12.800.00N112981 1:N:0:3A394G1114:19267:4S103L, F302Y, S307T, −8.800.00N1119071 1:N:0:3A394G2110:27515:4G108R, S307T, Y328D, −26.600.00N1112485 1:N:0:4A394G2115:9097:4V104L, S307T, T321K, −9.500.00N114698 1:N:0:4A394G1107:23186:4P117L, K298Q, S307T, −11.300.00N113788 1:N:0:3A394G1114:11441:4S103L, T291A, S307T, −15.800.00N1120252 1:N:0:3A394G2115:11654:4T123I, K298N, S307T, −13.200.00N1110133 1:N:0:3A394G1114:25898:4S103L, S288R, S307T, −20.400.00N1120164 1:N:0:4A394G2104:14172:5V104L, N155S, S305A, −12.600.00N1222246 1:N:0:3S307T, A394G1105:10148:2S307T, A394C−8.400.00Y1419074 1:N:0:42116:23124:2A331S, V409G−7.900.00N144613 1:N:0:32118:4192:2A326V, T405S−4.700.00N1417694 1:N:0:41101:2409:2S319R, N354I−17.800.08N1416339 1:N:0:31108:2650:2S307T, A394V−8.400.00N1410812 1:N:0:31113:10966:2Y351D, P403Q−21.500.78N1414429 1:N:0:31117:18531:2S3071, A394G−15.600.00N1419527 1:N:0:41111:4052:2P117H, E208G−3.200.09N1519173 1:N:0:41107:7761:2V124A, I152F−9.200.00N157540 1:N:0:32112:22110:2Q102K, N155H−3.000.00N1523423 1:N:0:41101:23578:3V124A, S307T, A394G−12.400.00N163670 1:N:0:31102:17479:3V124I, S307T, A394G−9.070.00N164575 1:N:0:31102:17819:3Q102P, S307T, A394G−11.700.00N167207 1:N:0:31103:14565:3T123I, S307T, A394G−8.400.00N169892 1:N:0:31110:5480:3V104L, S307T, A394G−8.900.00N169843 1:N:0:31112:19316:3G108W, S307T, A394G−11.600.00N1619410 1:N:0:31114:2899:3S103L, S307T, A394G−8.400.00N168564 1:N:0:31115:19772:3L111P, S307T, A394G−8.400.00N167379 1:N:0:31116:6924:3V112A, S307T, A394G−9.200.00N166457 1:N:0:32102:22485:3E105A, S307T, A394G−9.000.00N168517 1:N:0:32105:16948:3E106K, S307T, A394G−9.200.00N1618514 1:N:0:32106:20355:3V104M, S307T, A394G−8.900.00N1622838 1:N:0:31102:17194:3G115R, S307T, A394G−10.800.00N169315 1:N:0:41106:14178:3L118P, S307T, A394G−9.400.00N1611856 1:N:0:41108:15665:3P114L, S307T, A394G−8.900.00N162204 1:N:0:42108:4840:3G108R, S307T, A394G−11.600.00N1619320 1:N:0:41111:28917:3D153N, S307T, A394G−9.600.00N1612309 1:N:0:12114:3507:3N155T, S307T, A394G−12.200.00N1615988 1:N:0:31108:15399:3N155K, S307T, A394G−12.200.00N1617590 1:N:0:42108:28274:3N155S, S307T, A394G−10.300.00N1615542 1:N:0:42104:17031:3V251G, V253F, T303P−36.100.26N198630 1:N:0:31114:13316:5E208G, S307T, S355G, −32.200.30Y205473 1:N:0:4A394G, F397C1112:23132:5S307T, D324E, C33OR, −37.700.00Y2124993 1:N:0:3R332S, A394G2112:6426:5A294V, K298N, R301C,−18.300.34Y216616 1:N:0:4S307T, A394G1102:14155:5S307T, S309P, A318D, −15.900.00N2117596 1:N:0:3L320M, A394G2111:15176:5S307T, K316I, T327A, −21.600.36N217641 1:N:0:3Y328S, A394G1108:5048:5S307T, E323A, A331S, −20.300.00N2119993 1:N:0:4R332S, A394G2109:22371:5S307T, T325K, F329Y, −25.030.22N214136 1:N:0:4R332I, A394G2117:25398:5N295K, F302I, S307T, −15.450.00N216020 1:N:0:3D314E, A394G1110:14855:5N295K, W296R, F302S, −22.800.30N212972 1:N:0:4S307T, A394G2114:9301:5A297P, K298N, S305T, −18.600.00N215365 1:N:0:4S307T, A394G2101:12534:5S305P, S307T, E323G, −25.100.09N2115438 1:N:0:3S352R, A394G2119:15108:5S307T, L320R, S352I, −36.000.08N212242 1:N:0:3T353A, A394G2105:23971:4G209V, G214V, S307T, −31.900.66N227940 1:N:0:3A394G2104:13560:4D153E, R197G, S307T, −24.600.05N2223700 1:N:0:3A394G2105:19440:4N155I, V251A, S307T, −25.500.23N2221172 1:N:0:3A394G1111:10546:3S307T, A394G, W401C−22.400.00Y2314497 1:N:0:41107:23593:3S307T, A394G, T405A−9.000.00N2321806 1:N:0:21113:2847:3S307T, A394G, T408S−8.500.00N2316642 1:N:0:32104:4801:3S307T, A394G, L406M−8.450.00N2314088 1:N:0:32114:4974:3S307T, A394G, W401R−25.900.00N2320311 1:N:0:31115:6244:3S307T, S355R, A394G−18.900.22N2310522 1:N:0:31118:5143:3S307T, S352I, A394G−22.400.08N235064 1:N:0:32105:22801:3S307T, A394G, Y395F−9.570.11N2310153 1:N:0:31116:17720:3S307T, N354T, A394G−15.400.05N239959 1:N:0:42102:17926:3S307T, S355I, A394G−22.400.08N2322781 1:N:0:42103:22390:3S307T, A394G, F397L−15.400.32N2316654 1:N:0:42104:8414:3S307T, A394G, F397V−18.900.28N236123 1:N:0:42109:21276:3S307T, A394G, A396D−20.400.09N2316256 1:N:0:42116:11277:3S307T, S355G, A394G−15.400.08N239861 1:N:0:42119:29036:3S307T, N354M, A394G−22.400.03N239615 1:N:0:42106:27137:3S307T, A394G, N398D−9.900.13N239548 1:N:0:51116:17833:3S352N, A394G, A396S−12.500.14N2325032 1:N:0:41117:14435:1C330W−16.000.00Y2423923 1:N:0:32112:17648:1F329C−10.000.19Y2413944 1:N:0:32111:13503:1Y351C−16.000.19Y2416690 1:N:0:31102:7658:1R332S−8.400.00N244808 1:N:0:31103:2392:1F329L−5.000.19N2412352 1:N:0:31107:28592:1A326S−2.200.00N2414825 1:N:0:32104:18926:1F302V−3.600.00N2418007 1:N:0:32105:25911:1E323D−0.100.00N2414756 1:N:0:32113:14965:1R301Q−0.100.00N247732 1:N:0:31113:6574:1F302L−2.400.00N2413048 1:N:0:42107:16580:1Y328H−1.500.00N245388 1:N:0:42115:20613:1S319R−0.300.00N245973 1:N:0:41116:5071:1Y293D−18.000.00N2415508 1:N:0:31108:9327:1K298E−2.400.00N249648 1:N:0:42101:26499:1K298N−4.800.00N249008 1:N:0:42111:18866:1A294P−6.900.66N249328 1:N:0:42114:12755:1A297S−1.200.00N2423738 1:N:0:41118:21354:1N354T−7.000.05N2417316 1:N:0:31112:20714:5S255I, W296L, G299C, −29.300.28Y2524469 1:N:0:4S307T, A394G1101:28349:5E208G, K306I, S307T, −11.500.09N2516016 1:N:0:3S319R, A394G2105:9076:5I213N, A294D, S307T, −22.100.34N2519470 1:N:0:3K316Q, A394G2102:26984:5T254H, N295K, S307T, −28.500.00N255065 1:N:0:4S319R, A394G1111:26510:5V253A, T291P, K298N, −32.300.00N257666 1:N:0:3S307T, A394G1106:10308:5T254S, S307T, A318T, −12.800.16N252576 1:N:0:1N354D, N398S1106:18041:5T254S, S307T, A318P, −13.600.16N2523428 1:N:0:4N354D, N398S1101:18416:2D314V, C330G−20.500.00Y263880 1:N:0:31110:23623:2Y293C, E323D−14.500.00Y263014 1:N:0:42105:11893:2D314E, E323A−0.650.00N267688 1:N:0:41102:9925:2G299R, F302I−4.400.00N266735 1:N:0:41114:21228:5S103L, V251A, A294S, −18.300.38N276267 1:N:0:3S307T, A394G1102:22954:4V251G, F302C, S307T, −32.200.26Y281504 1:N:0:4A394G2109:26695:4D196E, S307T, C330G, −30.050.00Y284723 1:N:0:4A394G1106:26861:4G209V, S307T, L320R, −22.000.66N2811322 1:N:0:3A394G2112:7769:4T254A, S307T, T327P, −17.800.36N2824277 1:N:0:3A394G1106:14651:4V253F, S307T, D314E, −20.450.00N289283 1:N:0:4A394G1106:25426:4V251G, I304S, S307T, −28.600.26N2812341 1:N:0:4A394G2102:13541:4E208G, S288G, S307T, −19.200.09N2810881 1:N:0:4A394G2106:17372:4V253G, S307T, L320R, −34.000.00N2824874 1:N:0:4A394G2109:7533:4E208A, T291K, S307T, −21.600.06N2820614 1:N:0:4A394G2113:15085:4G199S, K298N, S307T, −20.000.14N285712 1:N:0:4A394G1103:22282:4T322A, E323A, C330S, −16.100.00Y2919571 1:N:0:3A331S2117:21062:4S288R, T291A, T312K, −36.800.00N2911524 1:N:0:4Y328D2101:3566:3I213N, S307T, A394G−12.900.00N3014679 1:N:0:32101:10234:3E208K, S307T, A394G−9.100.55N3016082 1:N:0:32106:29060:3Q204L, S307T, A394G−15.200.23N3012886 1:N:0:32106:25797:3V202G, S307T, A394G−21.400.22N3016513 1:N:0:32112:16946:3E211V, S307T, A394G−13.600.17N309038 1:N:0:32117:7810:3E208A, S307P, A394G−15.300.06N304136 1:N:0:31102:3683:3I213S, S307T, A394G−12.000.00N3011930 1:N:0:41104:29232:3I213T, S307T, A394G−11.100.00N3010626 1:N:0:41106:14147:3W201S, S307T, A394G−28.400.00N3011888 1:N:0:41109:16238:3R203H, S307T, A394G−12.400.00N3015426 1:N:0:41102:19092:3V251G, S307T, A394G−27.400.26N3021785 1:N:0:31104:15958:3G199V, S307T, A394G−25.400.07N306342 1:N:0:31109:13663:3T254S, S307T, A394G−12.200.00N3015593 1:N:0:32101:22839:3D196E, S307T, A394G−10.050.00N304783 1:N:0:32102:27837:3I252L, S307T, A394G−9.350.00N307005 1:N:0:32107:23079:3R197G, S307T, A394G−24.000.05N3021136 1:N:0:31114:14583:3D196N, S307T, A394G−11.700.00N306584 1:N:0:42107:10102:3M198V, S307T, A394G−11.000.00N3020753 1:N:0:42118:21108:3V253G, S307T, A394G−28.400.00N3017365 1:N:0:42118:6440:4S307T, S308A, A394G, −8.500.00N3118144 1:N:0:4S411A1107:24338:4V251G, S307T, A394G, −27.800.26N317886 1:N:0:4S410F1108:5019:4S307T, A394G, L399F, −18.400.02N3218802 1:N:0:3G404R1110:24949:4S307T, Y328V, A394G, −29.400.09N325805 1:N:0:4A396D2103:24484:4S307T, S355R, A394G, −32.900.22N3222076 1:N:0:4W401G18045 1:N:0:3A394G, V407A2112:27082:3T290S, L320P, F329C−20.400.22Y397147 1:N:0:32105:19134:3Y293D, K298N, K306N−22.800.00N3923323 1:N:0:4VH variants of antibody 763 (4 clusters)1114:28111:5S103L, N197K, T322P, −30.700.07N018680 1:N:0:1F329V, V331G2106:2910:1R306K0.000.00N118342 1:N:0:1VH variants of antibody 770 (10 clusters)1119:25404:3T116G, P117T, A124V−4.700.00N018606 1:N:0:41106:8006:3T113M, A124V, S156T−7.700.00N017093 1:N:0:32118:26397:3F152S, E211D, Y292S−19.450.26N07206 1:N:0:32104:2150:1L406M−0.050.00N110503 1:N:0:32116:18898:1V313A−5.000.00N116116 1:N:0:31106:12880:1S319G−0.200.00N125089 1:N:0:42103:27776:1I213L−0.450.00N111682 1:N:0:42118:19944:1T405N−0.600.00N12771 1:N:0:41110:19477:1S295N−2.300.17N124960 1:N:0:31118:19369:1Y292S−8.400.19N16823 1:N:0:32118:25093:1W296G−4.000.26N113289 1:N:0:31110:22596:1Y293D−18.000.00N120720 1:N:0:41115:13276:1T291K−11.100.00N112164 1:N:0:41115:26242:1T291A−7.400.00N117606 1:N:0:42107:14816:1D352V−17.500.10N15418 1:N:0:31102:27934:5V251F, S254R, D289A, −58.300.34N26512 1:N:0:3T291P, A326D1104:21790:5V104L, V252I, R301P, −21.930.00N29830 1:N:0:4R332S, A353K1118:12558:5V104L, V252I, R301P, −18.630.00N22046 1:N:0:4T303P, A353K2104:22350:5V104L, V252I, R301P, −14.130.00N222387 1:N:0:4T321K, A353E2105:12294:5V104L, V252I, R301P, −15.330.00N218712 1:N:0:4T325M, A353K2110:9944:5V104L, V252I, R301P, −16.430.00N215110 1:N:0:4A331S, A353K2118:6852:5V104L, V252I, R301P, −14.430.00N218638 1:N:0:4D324E, A353K1112:26205:5V104L, V252I, D289G, −25.230.14N214133 1:N:0:3R301P, A353K2103:3741:5V104L, V252I, G299D, −15.030.00N216519 1:N:0:4R301P, A353K2118:8662:5V104L, V252I, A294E, −20.430.17N222920 1:N:0:4R301P, A353K2112:9951:5V2521, A294E, S295T, −22.230.31N211009 1:N:0:3R301P, A353K1101:3865:5V104L, V252I, R301P, −33.530.21N219178 1:N:0:4Y351D, A353K2114:9668:5V104L, V252I, R301P, −16.530.11N214956 1:N:0:4A353K, A396S2119:20444:5C101Q, Q102E, S103Q, −2.850.00Y37869 1:N:0:3V104L, E105V2112:7171:5V104M, A124V, G197V, −33.400.15N35737 1:N:0:4M198V, G199S2105:3795:3C330G, R332S, V407A−30.400.00Y49269 1:N:0:42111:23175:3V252I, R301P, A353K−13.030.00N414951 1:N:0:31113:15979:4A124V, S156R, G197A, −43.300.17N520411 1:N:0:3G199I1112:18176:4A124V, Y196F, G197A, −23.100.16N524346 1:N:0:4G199D2107:13654:4A124V, V252I, R301P, −17.030.00N518311 1:N:0:3A353K2116:13810:4V104M, V252I, R301P, −13.530.00N57307 1:N:0:4A353K1111:21746:2V104L, T116P−0.800.00N65128 1:N:0:31113:9389:2I322N, R332G−12.700.00N622434 1:N:0:32103:16223:2K208Q, L315P−8.200.86N623064 1:N:0:32105:3619:2V2021, I213V−1.170.29N615924 1:N:0:32106:2580:2S308A, S410A−0.100.00N614766 1:N:0:41106:23729:2I213S, W296V−8.600.26N62056 1:N:0:42105:4288:2T291N, I322S−8.600.00N618963 1:N:0:42104:6077:2V251G, Y292F−19.700.51N614499 1:N:0:41119:11057:2T321K, Y351A−16.600.20N610711 1:N:0:32110:10010:2R332S, Y395S−22.400.17N62769 1:N:0:42111:27401:2G355R, D394V−29.000.24N69934 1:N:0:41103:19179:1G207V0.000.00N713949 1:N:0:31106:18574:1S103P0.000.00N716375 1:N:0:31114:23251:1S103L0.000.00N75565 1:N:0:31118:25289:1A124T−4.000.00N712604 1:N:0:31118:20338:1V104M−0.500.00N715736 1:N:0:32104:11879:1V104L−0.500.00N79882 1:N:0:32110:10996:1P206L0.000.00N72922 1:N:0:32102:5275:1Q204H−3.400.15N722352 1:N:0:41115:28168:4G209A, I322S, Y328S, −30.400.37N818911 1:N:0:4D352A1119:16081:4T290N, T291M, G355M,−42.400.40N822926 1:N:0:3F397L1115:27661:5A124V, S156N, G197A, −32.300.24Y98185 1:N:0:4G199S, F329C1105:27684:5V104L, V112F, V252I, −14.730.00N915476 1:N:0:4R301P, A353K1114:18185:5S103L, V104L, V252I, −13.530.00N911772 1:N:0:4R301P, A353K2103:14314:5V104L, Q204R, V252I, −15.230.15N917267 1:N:0:4R301P, A353K2106:8058:5V104L, R110G, V252I, −14.330.00N99673 1:N:0:4R301P, A353K2114:26957:5P117A, A124V, V252I, −17.330.00N99620 1:N:0:4R301P, A353K2115:2997:5V104L, T123I, V252I, −13.530.00N99890 1:N:0:4R301P, A353KVH variants of antibody 776 (7 clusters)2109:5738:3Q102R, T322I, C330W−18.000.00Y021603 1:N:0:31111:2345:3Q102R, T322I, D324N−3.800.00N017106 1:N:0:21115:10228:3Q102R, W201S, T322I−22.000.00N022509 1:N:0:21117:25724:3L104V, T322I, S410P−1.700.00N06418 1:N:0:22101:26149:3Q102R, I304N, T322I−3.500.00N012052 1:N:0:22115:14482:3L104V, T123I, T322I−1.400.00N021163 1:N:0:21113:27497:3Q102R, V313E, T322I−12.000.00N017512 1:N:0:32104:4023:3Q102R, V202D, T322I−15.000.25N013062 1:N:0:32110:8209:3Q102R, T119P, T322I−2.600.00N074511:N:0:31115:9216:3Q102R, T3221, Y328F−3.000.00N078421:N:0:41112:5455:3Q102R, T322I, T325P−3.800.00N08335 1:N:0:52104:16530:3Q102R, I321K, T322I−3.000.00N03365 1:N:0:52104:13191:3Q102R, T322I, W401R−19.500.00N017970 1:N:0:52115:14608:3Q102R, T123I, T322I−2.000.00N023050 1:N:0:52109:26527:3Q102R, A297T, T322I−4.400.00N012633 1:N:0:11109:17671:3Q102R, S156I, T322I−16.800.00N016866 1:N:0:21116:15403:3Q102R, N199T, T322I−8.800.17N09100 1:N:0:21111:24306:3Q102R, I251S, T322I−17.200.37N015673 1:N:0:41109:15266:3Q102R, W296R, T322I−7.000.31N010649 1:N:0:51113:4091:3Q102R, R253S, T322I−14.000.00N016678 1:N:0:51117:17821:3Q102R, K298E, T322I−4.400.00N020921 1:N:0:51119:22336:3Q102R, T322I, P399L−17.000.07N024364 1:N:0:41107:18024:5T113M, T116G, P117S, −5.400.00N113209 1:N:0:3T155A, T322I1106:12016:1Y292N−8.400.27N211024 1:N:0:32102:18888:1F397V−10.500.11N22945 1:N:0:31114:22845:4Q102R, I251L, T322I, −13.900.29Y36504 1:N:0:3F329C2117:27848:4Q102R, R203C, N290K, −18.800.06Y310451 1:N:0:5T322I1112:9308:4Q102R, G108R, T155M,−10.900.00N36605 1:N:0:2T322I2112:8356:4Q102R, T322I, T325K, −8.800.11N623198 1:N:0:3F329L
[0331] For the final selection of NGS variants for DNA synthesis and HEK transient transfection with parental VK, all “medoids” VHs were selected that had more than 1 amino acid replacement, which resulted in AB angle Distance≤0.5 and had no free cysteine. Totally 31 VH have been selected for clone 755, 4 for clone 763, 10 for clone 770 and 9 for clone 776 (see Table below). As shown in the following Table selected variants were delivered either by the pool of PBMC or by the pool of antigen-enriched PBMC.TABLEFinal selection of NGS VH variants for gene synthesis and transient transfection with parental VK.VH variantABanglenumber ofriskfreeclusterclone variantfromnameDistancemutationsmutationsscorecysteineindexBCC.755-1AP1101:18779:191640.005Q102P, V104L,−13.10N121:N:0:4S307T, D324E,A394GBCC.755-2AP1101:19661:14240.002L120F, L154P−13.90N151:N:0:4BCC.755-3AP1101:21776:240810.002G299A, S305P−6.40N261:N:0:4BCC.755-4AP1102:23311:83750.093E208G, S307T,−11.20N301:N:0:4A394GBCC.755-5AP1104:29248:106130.004I213T, S307T,−11.70N281:N:0:4E323A, A394GBCC.755-6AP1104:5907:199270.133S288G, N295T,−13.20N391:N:0:3E323ABCC.755-7BP1110:21663:115080.055M198L, S255K, S307T,−19.90N351:N:0:1A394G, N398TBCC.755-8BP1111:20770:145820.002S307T, A394G−8.40N141:N:0:1BCC.755-9AP1112:17135:112760.004E105A, T123P,−9.00N51:N:0:3S307T, A394GBCC.755-10AP1113:7949:22670.005S307T, S319I, T322P,−12.60N211:N:0:4A331S, A394GBCC.755-11AP1115:13289:249280.295E208G, N295V, W296G,−28.00N341:N:0:4K298T, G299VBCC.755-12AP1115:4917:177810.165S305T, S307T,−29.20N361:N:0:3Y351N, A394G, L399VBCC.755-13AP1116:8054:134800.005V104L, T116G,−9.60N101:N:0:4P117S, S307T,A394GBCC.755-14AP1117:21523:92440.003T116G, P117T,−5.90N41:N:0:4I152FBCC.755-15AP1117:26048:174400.003S305P, S307T,−13.80N01:N:0:4A394GBCC.755-16AP1117:9244:192950.365Q102R, S307T,−10.40N81:N:0:4S319R, T327S,A394GBCC.755-17AP1118:21273:12320.004S307T, T312P,−12.40N71:N:0:4A318X, A394GBCC.755-18AP2101:5252:214330.003T113M, S307T,−8.40N161:N:0:3A394GBCC.755-19AP2101:5682:222040.002I213L, T291P−11.55N61:N:0:4BCC.755-20AP2102:17483:25320.163S307T, A394G,−14.40N231:N:0:3A396GBCC.755-21AP2102:28708:121700.001E323A−0.60N241:N:0:3BCC.755-22AP2103:7871:54140.005I152F, D153S,−17.90N381:N:0:3N155S, S307T,A394GBCC.755-23AP2110:7379:50320.004R110H, S307T,−12.00N111:N:0:4A318D, A394GBCC.755-24AP2111:21535:59800.001P117L−0.50N21:N:0:3BCC.755-25AP2112:28984:96870.021L399F−6.00N91:N:0:4BCC.755-26AP2113:22155:94320.255N155S, S156V,−62.40N11:N:0:4D196Y, R197G, G199TBCC.755-27AP2114:20534:116040.005S255G, S307T, A318D,−23.20N251:N:0:4A326G, A394GBCC.755-28AP2115:16505:50700.004S305P, K306T,−11.00N291:N:0:4K316T, L320RBCC.755-29AP2115:3097:180450.005T123I, R301P,−10.80N331:N:0:3S307T, A394G,V407ABCC.755-30AP2116:5079:188870.065T123P, E208A,−15.30N271:N:0:3S307T, S309L,A394GBCC.755-31AP2119:9570:25970.264D153A, V251G, S307T,−32.20N221:N:0:4A394GBCC.763-1BP1101:11801:208940.001T322P−0.90N11:N:0:1BCC.763-2BP2103:7194:163680.152E211A, R306K−3.90N31:N:0:1BCC.763-3BP1114:12508:88290.005Q102P, S103P,−18.00N01:N:0:1S125P, Y196S,R306KBCC.763-4BP1101:1836:149710.183Y293D, S295N, T322P−21.20N21:N:0:1BCC.770-1AP1101:11665:42350.004V104L, V252I,−13.53N51:N:0:3R301P, A353KBCC.770-10AP2117:26896:175980.122Y292H, R301P−1.45N61:N:0:4BCC.770-2AP1108:27622:111160.005V104L, L118Q,−14.33N91:N:0:4V252I, R301P,A353KBCC.770-3AP1108:9880:91080.001K298N−4.80N11:N:0:3BCC.770-4AP1109:17331:32060.213T321K, R332G, Y351D−31.80N41:N:0:3BCC.770-5AP1115:10270:147560.175T123A, A124V, S156T,−22.30N31:N:0:3G197A, G199SBCC.770-6AP2101:15918:205910.154M317I, I322T,−32.40N81:N:0:4R332V, Y351NBCC.770-7AP2101:7723:109030.005V104L, V252I,−15.33N21:N:0:4R301P, T325P,A353KBCC.770-8AP2111:12893:243510.003A124V, F152I,−11.60N01:N:0:4S153DBCC.770-9AP2115:12441:160840.001T123I0.00N71:N:0:3BCC.776-1AP1106:12016:110240.271Y292N−8.40N21:N:0:3BCC.776-2BP1109:2895:140570.001T322I−0.90N21:N:0:2BCC.776-3AP1111:24306:156730.373Q102R, I251S,−17.20N01:N:0:4T322IBCC.776-4BP1112:3580:161330.002Q102R, T322I−2.00N41:N:0:1BCC.776-5BP1117:25709:64290.003L104V, I304S,−2.60N01:N:0:2T322IBCC.776-6AP2107:19369:120390.001L104V−0.50N51:N:0:3BCC.776-7BP2111:7168:172450.005R110D, T113K, T116A,−2.20N11:N:0:5P117S, T322IBCC.776-8AP2115:18009:168580.004Q102R, T123I,−2.05N31:N:0:4D314E, T322IBCC.776-9BP2116:22283:221160.004Q102R, R301Q, T322I,−2.15N61:N:0:2L406MAP: After Panning cell pool;BP: PBMC cell pool before panningDose Response Curve Based Binding Analysis of NGS-Variants
[0332] Variant VH and parental VL plasmids were transiently co-transfected into HEK293 cells. Additionally, the reference B-cell clones were also transfected and used as reference. After purification of supernatant all clones were analyzed for binding to human and murine antigen as described in the experimental section below. The EC50-based analysis was carried out in replicates at different occasions to warrant statistical accuracy.
[0333] FIG. 3 shows a correlation plot of human and murine antigen binding displaying the reference antibody (binder identified by screening after panning) and the respective clones selected by NGS. All biochemical and mutation score data of NGS variants and BCC references are consolidated in the following Table below. The plot data show good correlations regarding binding behavior and relative ranking (see Table data). Most of the variants (60-80%) display binding EC50 values to the human antigen comparable to the relative control, and some even a slightly improved EC50 value. The same was observed for EC50 values for binding to murine antigen for those clones displaying cross reactivity to murine antigen. Variants for BCC 763 were identified in PBMC NGS repertoire (before enrichment) and 80% of the variants displayed binding properties comparable to the reference; for BCC 755 and BCC 776 most of the variants with comparable behavior to the reference were recovered from the enriched pool of B-cells, yet still a couple of good binders could be obtained from total PBMC NGS repertoire. Variants for BCC 770 were recovered from NGS repertoire of enriched pool. It can be seen that for each reference antibody a variant antibody with lower EC50 value could be identified with the method s reported herein.TABLEConsolidation of biochemical and mutation score data of NGS variants and BCC references,sequences by EC50 value with human antigen (lowest first, highest last).EC50EC50humanmurineantigenantigenABanglenumber ofallriskclone variant[ng / ml][ng / ml]fromDistancemutationsmutationsscoreVariants BCC 755BCC.755-14<20>2000AP0.003T116G, P117T,−5.90I152FBCC.755-7<20>2000BP0.055M198L, S255K,−19.90S307T, A394G,N398TBCC.755-24<20>2000AP0.001P117L−0.50BCC.755-6<20>2000AP0.133S288G, N295T,−13.20E323ABCC.755<20>2000ref 1BCC.755-22<20>2000AP0.005I152F, D153S,−17.90N155S, S307T,A394GBCC.755-2<20>2000AP0.002L120F, L154P−13.90BCC.755-4<20>2000AP0.093E208G, S307T,−11.20A394GBCC.755-13<20>2000AP0.005V104L, T116G,−9.60P117S, S307T,A394GBCC.755-21<20>2000AP0.001E323A−0.60BCC.755-27<20>2000AP0.005S255G, S307T, A318D,−23.20A326G, A394GBCC.755-11<20>2000AP0.295E208G, N295V, W296G,−28.00K298T, G299VBCC.755-3<20>2000AP0.002G299A, S305P−6.40BCC.755<20>2000ref 2BCC.755-17<50>2000AP0.004S307T, T312P,−12.40A318X, A394GBCC.755-8<50>2000BP0.002S307T, A394G−8.40BCC.755-25<50>2000AP0.021L399F−6.00BCC.755-16<50>2000AP0.365Q102R, S307T,−10.40S319R, T327S,A394GBCC.755-9<50>2000AP0.004E105A, T123P,−9.00S307T, A394GBCC.755-19<50>2000AP0.002I213L, T291P−11.55BCC.755-10<50>2000AP0.005S307T, S319I, T322P,−12.60A331S, A394GBCC.755-5<50>2000AP0.004I213T, S307T, E323A,−11.70A394GBCC.755-18<50>2000AP0.003T113M, S307T,−8.40A394GBCC.755-15<50>2000AP0.003S305P, S307T,−13.80A394GBCC.755-30<50>2000AP0.065T123P, E208A,−15.30S307T, S309L,A394GBCC.755-29<50>2000AP0.005T123I, R301P, S307T,−10.80A394G, V407ABCC.755-2350-100>2000AP0.004R110H, S307T,−12.00A318D, A394GBCC.755-1>100>2000AP0.005Q102P, V104L,−13.10S307T, D324E,A394GBCC.755-20>100>2000AP0.163S307T, A394G,−14.40A396GBCC.755-12>2000>2000AP0.165S305T, S307T,−29.20Y351N, A394G, L399VBCC.755-26>2000>2000AP0.255N155S, S156V,−62.40D196Y, R197G, G199TBCC.755-28>2000>2000AP0.004S305P, K306T,−11.00K316T, L320RBCC.755-31>2000>2000AP0.264D153A, V251G, S307T,−32.20A394GVariants for BCC 763BCC.763-1<20<20BP0.001T322P−0.90BCC.763<20<20ref. 1BCC.763<20<20ref. 2BCC.763-2<20<20BP0.152E211A, R306K−3.90BCC.763-4<50<20BP0.183Y293D, S295N,−21.2T322PBCC.763-3>100>50BP0.005Q102P, S103P,−18.00S125P, Y196S,R306KVariants for BCC 770BCC.770-1<20>2000AP0.004V104L, V252I,−13.53R301P, A353KBCC.770<20>2000ref. 1BCC.770-9<20>2000AP0.001T123I0.00BCC.770-2<20>2000AP0.005V104L, L118Q,−14.33V252I, R301P,A353KBCC.770<20>2000ref. 2BCC.770-8<20>2000AP0.003A124V, F152I,−11.60S153DBCC.770-3<20>2000AP0.001K298N−4.80BCC.770-7>100>2000AP0.005V104L, V252I,−15.33R301P, T325P,A353KBCC.770-10>100>2000AP0.122Y292H, R301P−1.45BCC.770-4>2000>2000AP0.213T321K, R332G,−31.80Y351DBCC.770-5>2000>2000AP0.175T123A, A124V,−22.30S156T, G197A,G199SBCC.770-6>2000>2000AP0.154M317I, I322T,−32.40R332V, Y351NVariants for BCC 776BCC.776-8<20<20AP0.004Q102R, T123I,−2.05D314E, T322IBCC.776-1<20<20AP0.271Y292N−8.40BCC.776-9<20<20BP0.004Q102R, R301Q,−2.15T322I, L406MBCC.776<20<20ref. 1BCC.776-2<20<20BP01T322I−0.9BCC.776<20<20ref. 2BCC.776-6<20<20AP0.001L104V−0.50BCC.776-3<20<20AP0.373Q102R, I251S,−17.20T322IBCC.776-7<20<20BP0.005R110D, T113K,−2.20T116A, P117S,T322IBCC.776-4<20<20BP0.002Q102R, T322I−2.00BCC.776-5<20<20BP0.003L104V, I304S,−2.60T322IAP: After Panning cell pool;BP: PBMC cell pool before panning
[0334] Thus, with this procedure antigen specific binders could be identified with binding properties comparable (or even improved) to the antigen specific B-cell clones isolated as described. This has been demonstrated by DNA synthesis, recombinant expression and biochemical analysis of sequence-based identified variants. This opens up the way to a completely new application of NGS data.Example BNGS Variants for B-Cell Cloning (BCC) Binders Bearing Developability Hot-Spots
[0335] Totally for seven B-cell cloning binders clonally related binders were identified in the NGS repertoires with the method as reported herein; all clones were isolated by B-cell cloning after hu-CDCP1 specific enrichment (enrichment type is indicated in the Table below) all exhibited specificity for hu-CDCP1 (EC50 / IC50 abs (ng / ml) range below 200 ng / ml) and all VHs bore cysteine or N-glycosylation site spots in HCDRs.TABLEproperties of B cell clones selected for NGS variants analysisEC / IC50absList Dev.clone[ng / ml]SpotsCDR2 PEPCDR3 PEPCDCP1_1059.87CysCIYAGSGRIKYASWAKGHCDR2(SEQ ID NO: 106)CDCP1_22341.23CysCIYAGSGGATYYASWAKGHCDR2(SEQ ID NO: 107)CDCP1_23631.59N{P}[ST]IINTSGNTYYANWAKGHCDR2(SEQ ID NO: 108)CDCP1_28412.69CysFIGSSGTTYCATWAKGHCDR2(SEQ ID NO: 109)CDCP1_212197.93CysGGYACDLHCDR3(SEQ ID NO: 110)CDCP1_08846.76N{P}[ST]IINTSGNTYYANWAKGHCDR2(SEQ ID NO: 111)CDCP1_23427.57N{P}[ST]IFYVATNITWYASWAKGHCDR2(SEQ ID NO: 112)cloneenrichment TypeCDCP1_088on plates coated CDCP1 proteinCDCP1_105on plates coated CDCP1 proteinCDCP1_234antigen specific sort using biotinylated CDCP1 with the followingsortgates: rbIgM⊖ / rbIgG⊕ / CDCP1⊕CDCP1_212antigen specific sort using biotinylated CDCP1 with the followingsortgates: rbIgM⊖ / rbIgG⊕ / CDCP1⊕CDCP1_223antigen specific sort using biotinylated CDCP1 with the followingsortgates: rbIgM⊖ / rbIgG⊕ / CDCP1⊕CDCP1_236antigen specific sort using biotinylated CDCP1 with the followingsortgates: rbIgM⊖ / rbIgG⊕ / CDCP1⊕CDCP1_284antigen specific sort using biotinylated CDCP1 with the followingsortgates: rbIgM⊖ / rbIgG⊕ / CDCP1⊕
[0336] NGS repertoire from PBMCs and from antigen enriched B-cells was analyzed for identification of VHs variants with ≤11 amino acid replacements in the entire VH (FR1 to FR4) compared to the VH of the reference B-cell binders. Those VH cognate variants with improved developability properties (no Cys and / or N-glycosylation site spots any longer in HCDRs) were gene synthesized, co-transfected with the parental light chain in HEK cells and the expressed antibodies were evaluated for binding properties in comparison with BCC references.Dose Response Curve Based Binding Analysis of NGS-Variants
[0337] Variant VHs and parental VL plasmids were transiently co-transfected into HEK293 cells. Additionally, the seven parental antibodies were transfected and used as reference. After purification of supernatant all clones were analyzed for binding to human CDCP1 antigen as described in the Examples section.
[0338] FIG. 4 shows binding of NGS variants to human CDCP1 in comparison to the respective parental clones. For each antibody, a different number of NGS sequence variants was tested. At least one variant for each clone could be identified that shows EC50 values in the same range as the reference antibody (marked by *).
[0339] In the following Table for each BCC clone the VH variants identified, the B-cell source of NGS sample providing the variants, the total number of amino acid replacements in entire VH and the absolute EC50 / IC50 values are shown.TABLEVH variants identified with improved in silico developability properties: the B cell source of NGS sample providing the variants, the total number of amino acid replacementsin entire VH and the absolute EC50 / IC50 values.Total AAMutationsSpot originalEC / IC50 abscompared toSample Nameclone(ng / ml)Sourcereference VHCDCP1 105Cys CDR29.87CDCP1 105-126.88Enriched B cells3CDCP1 105-2binding lostEnriched B cells10CDCP1 105-4binding lostEnriched B cells6CDCP1 105-5binding lostEnriched B cells8CDCP1 105-6binding lostEnriched B cells11CDCP1 105-7binding lostEnriched B cells6CDCP1 105-8233.92Enriched B cells3CDCP1 105-9binding lostEnriched B cells9CDCP1 105-1023.56Enriched B cells2CDCP1 105-1117.66Enriched B cells2CDCP1 105-12binding lostEnriched B cells3CDCP1 105-13binding lostEnriched B cells2CDCP1 105-1442.93Enriched B cells2CDCP1 105-15binding lostEnriched B cells6CDCP1 105-1647.46Enriched B cells5CDCP1 105-17133.83Enriched B cells3CDCP1 105-18binding lostEnriched B cells6CDCP1 105-19binding lostEnriched B cells4CDCP1 105-20binding lostEnriched B cells7CDCP1 223Cys CDR246.59CDCP1 223-1binding lostPBMC5CDCP1 223-246.11PBMC4CDCP1 223-3991.24PBMVC6CDCP1 236N{P}[ST] CDR231.59CDCP1 236-539.48PBMC10CDCP1 284Cys CDR212.69CDCP1 284-114.23Enriched B cells11CDCP1 284-212.11Enriched B cells11CDCP1 284-310.25Enriched B cells11CDCP1 284-410.56Enriched B cells11CDCP1 284-514.34Enriched B cells11CDCP1 284-616.35Enriched B cells11CDCP1 284-714.67Enriched B cells11CDCP1 284-813.19Enriched B cells11CDCP1 284-98.69Enriched B cells11CDCP1 284-1014.02Enriched B cells6CDCP1 284-118.29Enriched B cells6CDCP1 284-1217.05Enriched B cells5CDCP1 284-1315.32Enriched B cells7CDCP1 284-1416.21Enriched B cells6CDCP1 284-1511.01Enriched B cells6CDCP1 212Cys CDR3197.93CDCP1 212-1448.61Enriched B cells3CDCP1 212-2630.55Enriched B cells2CDCP1 212-3binding lostEnriched B cells5CDCP1 212-4binding lostEnriched B cells3CDCP1 212-5189.14Enriched B cells7CDCP1 088N{P}[ST] CDR246.76CDCP1 088-1binding lostPBMC3CDCP1 088-2binding lostPBMC5CDCP1 088-354.38PBMC10CDCP1 234N{P}[ST] CDR227.57CDCP1 234-143.11Enriched B cells9CDCP1 234-227.16Enriched B cells4CDCP1 234-317.57Enriched B cells3CDCP1 234-422.37Enriched B cells3CDCP1 234-522.04Enriched B cells3CDCP1 234-629.79Enriched B cells3CDCP1 234-825.45Enriched B cells2CDCP1 234-923.09Enriched B cells3CDCP1 234-1022.24Enriched B cells3CDCP1 234-1119.85Enriched B cells3CDCP1 234-1214.61Enriched B cells3CDCP1 234-1321.18Enriched B cells3CDCP1 234-1418.72Enriched B cells3CDCP1 234-1529.18Enriched B cells3 indicates data missing or illegible when filed
[0340] The CDRH3 sequences from confirmed binders found in the NGS pool showed quite distinct features, providing a basis to classify sequences. Some of the ‘best’ binders showed strong homology in the CDRH3 region. It has been found that the NGS repertoire pool can be used to find such variants of good binders.
[0341] Analyses with sufficient sequencing depth and optimized normalization conditions are the basis for such a process. Most importantly the complex sequence diversity has to be analyzed on the DNA level to group together sequences with comparable properties, which are likely of the same phylogenetic origin. Especially with rabbits, that not only use somatic hypermutation but also gene conversion during clonal expansion it is very difficult to make this grouping to allow the identification of clonally related VHs.
[0342] The NGS sequence pools were screened with the method as reported herein for VH variants that possess identical CDRH3 sequences as the above indicated four BCC binders. The analysis of the complete VH region on a DNA level showed that, in addition to exactly identical VHs, variants can be identified varying from the reference BCCs by numerous mutations on the V region outside CDRH3 (see Table below). Alignments of the sequences indicated that the mutations occurred at various different positions throughout the sequence and were not concentrated to a certain region (data not shown).TABLEDetailed analysis of selected CDRH3 sequences and variants withintheir VH.Number ofMutations in totalNumber of mutationsVH (compared toin total VH of all NGSclosest Germline)Samples (compared toCDRH3 Sequencein BCCclosest Germline)DHDTGSHPYNYENMDV65, 6, 7, 8, 10(SEQ ID NO: 01)DHDTGSHPYSYENMDV66, 7(SEQ ID NO: 02)DHDTGSSPYNYDNMDV77, 8, 9, 11, 12, 18(SEQ ID NO: 03)DSLSYGYAYATNYFNI84, 5, 6, 8, 9, 10, 11, 12, 14
[0343] Beside sequences with identical CDRH3 (and mutations in the VH), the NGS pool can also be screened to identify variants with CDRH3 regions highly homologous to those of reference BCC binders. This can be done by aligning the CDRH3 of the binders with the total NGS repertoire and selecting sequences with high homology (e.g. max. one mutation on the peptide level). The following Table shows an extract of the alignment done with the CDRH3 sequences of the ‘best’ binders 9-11. The CDRH3 sequences S1-S25 found in the NGS pools are closely related to these ‘best’ binders.TABLEPeptide CDR3 sequences of identified binders9-11 aligned with an extract of CDR3s of the same length and a high degree of homology (S1-S25).9S12DHDT--GSHP-YN-YENMDVDHDT--GSHP-YS-YENIDV(SEQ ID NO: 01)(SEQ ID NO: 15)10S13DHDT--GSHP-YS-YENMDVDNDT--GSHP-YS--YENMDV(SEQ ID NO: 02)(SEQ ID NO: 16)11S14DHDT--GSSP-YN-YDNMDVDHDN--GSHP-YS--YENMDV(SEQ ID NO: 03)(SEQ ID NO: 17)S1S15DHDT--GSSQ-YN-YDNMDVDHDT--GSHP-YS--HENMDV(SEQ ID NO: 04)(SEQ ID NO: 18)S2S16DHDT--GSNP-YN--YDNMDVDHDT--GCHP-YS--YENMDV(SEQ ID NO: 05)(SEQ ID NO: 19)S3S17DHDT--GSSP-YN--YDNMDVDHDT--GNHP-YS--YENMDV(SEQ ID NO: 06)(SEQ ID NO: 20)S4S18DHDT--GSHP-YK--YANMDVDHDT--GSHP-YS--YENMDV(SEQ ID NO: 07)(SEQ ID NO: 21)S5S19EHDT--GSHP-YC--YENMDVDYDT--GSHP-YN--YENMDV(SEQ ID NO: 08)(SEQ ID NO: 22)S6S20DHDT--GNHP-YN--YENMDVDHDT--GSHP-YN--YENLDV(SEQ ID NO: 09)(SEQ ID NO: 23)S7S21DHDT--GSHP-YS--YENMYVDHDT--GSHP-YN--YENRDV(SEQ ID NO: 10)(SEQ ID NO: 24)S8S22DHDT--GSHP-YS--YENMDFDHDT--GSHP-YN--NENMDV(SEQ ID NO: 11)(SEQ ID NO: 25)S9S23DHET--GSHP-YS--YENMDVDHDT--GSHP-YN--YENMDV(SEQ ID NO: 12)(SEQ ID NO: 26)S10S24EHDT--GSHP-YS--YENMDVDHDT--GSSP-DN--YDNMEV(SEQ ID NO: 13)(SEQ ID NO: 27)S11S25DHDT--GSHP-YS-YDNMDVDHDA--GSSP-YN-YDNMDV(SEQ ID NO: 14)(SEQ ID NO: 28)
[0344] As mentioned above, four BCC CDRH3s have been found in more than one animal in the NGS samples. It has been shown in previous studies that the more shared sequences CDR3 are found in animals after immunization suggesting these sequences to be antigen specific (Galson, J. D., et al., Crit. Rev. Immunol. 35 (2015) 463-478).
[0345] As can be seen from the above the combination of data generated by NGS in combination with sequence data from BCC can be used to identify alternative antigen specific variant antibodies to a reference antibody.
[0346] NGS data can be screened for sequences with similar or identical CDRH3 but higher mutation rate within other regions of the VH compared to i.e. BCC binders to find antibodies with improved properties.VL
[0347] If a transgenic animal expressing a common light chain is used for immunization, solely analysis of the VH-repertoire is sufficient. In other transgenic models or wild-type animals VH / VL pairs need to be identified as belonging together and contributing both in equal manner to antigen specificity. For this, a wide range of methods can be used, ranging from single-cell sorting combined with single-cell cloning to special RNA capturing. For example, fusion PCR is suited quite well for combination with the UMI error correction method. The paired-chain information is retained through physically attachment of the alpha and beta (or heavy and light) transcripts from each single cell. The fusion of transcripts can be accomplished by (multiplexed) overlap-extension PCR.
[0348] The following examples and figures are provided to aid the understanding of the present invention, the true scope of which is set forth in the appended claims. It is understood that modifications can be made in the procedures set forth without departing from the spirit of the invention.DESCRIPTION OF THE FIGURES
[0349] FIGS. 1A&FIG. 1B VH (top) and VL (bottom) sequences of the reference antibody 763 and six variants of its VH domain. Framework and CDR classification follows Wolfguy nomenclature. In the VH variants, amino acid substitutions with regard to the reference are shown in light grey. CDRs are shaded grey in the numbering.
[0350] FIGS. 2A & 2B VH (top) and VL (bottom) sequences of the reference antibody 763 and six variants of its VH domain. Residues forming part of the VH-VL orientation fingerprint have been marked with a gray background. Amino acid substitutions in this region are likely to induce a change VH-VL orientation as compared to the reference antibody. CDRs are shaded grey in the numbering.
[0351] FIGS. 3A, 3B, 3C & 3D EC50 (M) correlation plot; human versus murine LRP8 binding of all the NGS variants. The controls, the clones identified by screening and panning, (red and blue), and clones, identified by NGS (grey), were plotted; A: Variants for BCC 763, B: Variants for BCC 776, C: Variants for BCC 770, D: Variants for BCC 755.
[0352] FIGS. 4A, 4B, 4C, 4D, 4E, 4F & 4G Absolute EC50 [nM] values show binding of all NGS variants (white bars) to human CDCP1. Parental controls are shown in black. Successful NGS variants are marked with (*). na=not available, meaning that no EC50 value could be derived from the binding curve due to weak binding of the variant. A: Variants for CDCP1_105; B: Variants for CDCP1_223; C: Variants for CDCP1_236; D: Variants for CDCP1_284; E: Variants for CDCP1_212; F: Variants for CDCP1_088; G: Variants for CDCP1_234.
[0353] FIG. 5 Scheme of the BCC and NGS workflow.EXAMPLESExample 1Immunization of Rabbits
[0354] A KLH conjugate of a human LRP8 was used for the immunization of the New Zealand White rabbit. Each rabbit was immunized with 500 μg of the immunogen, emulsified with complete Freund's adjuvant, at day 0 by intradermal application and 500 μg each at days 7, 14, 28, 42 by alternating intramuscular and subcutaneous applications. Thereafter, rabbits received monthly subcutaneous immunizations of 500 μg, and small samples of blood were taken 7 days after immunization for the determination of serum titers. A larger blood sample (10% of estimated total blood volume) was taken during the third, fourth and fifth month of immunization (at 5-7 days after immunization), and peripheral mononuclear cells were isolated, which were used as a source of antigen-specific B-cells in the B-cell cloning process.Example 2Determination of Serum Titers (ELISA)
[0355] Biotinylated human LRP8 was immobilized on a 96-well streptavidin-coated plate at 0.5 μg / ml, 100 μl / well, in PBS, followed by blocking of the plate with 2% CroteinC in PBS, 200 μl / well. Thereafter 100 μl / well serial dilutions of antisera, in duplicates, in 0.5% CroteinC in PBS were applied. The detection was done with HRP-conjugated donkey anti-rabbit IgG antibody (Jackson Immunoresearch / Dianova, Cat. No. 711-036-152; 1 / 16 000), each diluted in 0.5% CroteinC in PBS, 100 μl / well. For all steps, plates were incubated for 1 h at 37° C. Between all steps plates were washed 3 times with 0.05% Tween 20 in PBS. Signal was developed by addition of BM Blue POD (peroxidase)-substrate soluble (Roche Diagnostics GmbH, Mannheim, Germany), 100 μl / well; and stopped by addition of 1 M HCl, 100 μl / well. Absorbance was read out at 450 nm, against 690 nm as reference. Titer was defined as dilution of antisera resulting in half-maximal signal.Example 3Isolation of Rabbit Peripheral Blood Mononuclear Cells (PBMCs)
[0356] Blood samples were taken of immunized wild-type rabbits (NZW). EDTA containing whole blood was diluted twofold with 1×PBS (PAA, Pasching, Austria) before density centrifugation using lympholyte mammal (Cedarlane Laboratories, Burlington, Ontario, Canada) according to the specifications of the manufacturer. The PBMCs were washed twice with 1×PBS.Example 4Depletion of Macrophages / Monocytes
[0357] The PBMCs were seeded on sterile KLH-coated SA-6-well-plates to deplete macrophages and monocytes through unspecific adhesion and to remove cell binding to KLH. Each well was filled at maximum with 4 ml medium and up to 6×10E6 PBMCs from the immunized rabbit and were allowed to bind for 1 h at 37° C. and 5% CO2. The cells in the supernatant (peripheral blood lymphocytes (PBLs)) were used for the antigen panning step.Example 5Enrichment of B-Cells
[0358] Sterile streptavidin coated 6-well plates (Microcoat, Bernried, Germany) were coated either with 2 μg / ml of the biotinylated KLH protein or the biotinylated LRP8 / CDCP1 protein in PBS for 3 h at room temperature. Prior to the panning step these 6-well plates were washed three times with sterile PBS. Coated plates were seeded with up to 6×10E6 PBLs per 4 ml medium and allowed to bind for 1 h at 37° C. and 5% CO2. Non-adherent cells were removed by carefully washing the wells 1-2 times with 1×PBS. The remaining sticky cells were detached by trypsin for 10 min. at 37° C. and 5% CO2. Trypsination was stopped with EL-4 B5 medium. The cells were kept on ice until the immune fluorescence staining.EL-4 B5 Medium
[0359] RPMI 1640 (Pan Biotech, Aidenbach, Germany) supplemented with 10% FCS (Hyclone, Logan, UT, USA), 2 mM Glutamine, 1% penicillin / streptomycin solution (PAA, Pasching, Austria), 2 mM sodium pyruvate, 10 mM iEPES (PAN Biotech, Aidenbach, Germany) and 0.05 mM beta-mercaptoethanol (Gibco, Paisley, Scotland).Example 6Immune Fluorescence Staining and Flow Cytometry
[0360] An anti-IgG antibody FITC conjugate (AbD Serotec, Düsseldorf, Germany) was used for single cell sorting. For surface staining, B-cells pre-treated with a depletion step and an enrichment step (Example 4 and 5) were incubated with the anti-IgG antibody FITC conjugate in PBS (phosphate buffered saline solution) and incubated for 45 min. in the dark at 4° C. After staining the cells were washed two times with ice cold PBS. Finally, the labelled B-cells were resuspended in ice cold PBS and immediately subjected to the FACS analyses. Propidium iodide in a concentration of 5 μg / ml (BD Pharmingen, San Diego, CA, USA) was added prior to the FACS analyses to discriminate between dead and live cells.
[0361] A Becton Dickinson FACSAria equipped with a computer and the FACSDiva software (BD Biosciences, USA) were used for single cell sort.Example 7B-Cell Cultivation
[0362] The cultivation of the single sorted B-cells was done according to a method described by Seeber et al. (Seeber, S., et al., PLoS One, 9 (2014) e86184.). Briefly, single sorted rabbit B-cells were incubated in 96-well plates with 200 l / well EL-4 B5 medium containing Pansorbin Cells (1:100,000) (Calbiochem (Merck), Darmstadt, Germany), 5% rabbit thymocyte supernatant (MicroCoat, Bernried, Germany) and gamma-irradiated murine EL-4 B5 thymoma cells (5×10E5 cells / well) for 7 days at 37° C. in the incubator. The supernatants of the B-cell cultivation were removed for screening and the remaining cells were harvested immediately and frozen at −80° C. in 100 μl RLT buffer (Qiagen, Hilden, Germany).Example 8Enzyme-Linked Immunosorbent Assay (ELISA)Human Antigen:
[0363] The antigen, biotinylated human LRP8, was incubated with 5 μL sample containing the anti-LRP8 antibody at a concentration of 250 ng / mL in a total volume of 25 μL in PBS, 0.5% BSA and 0.05% Tween in a 384 w microtiterplate (Maxisorb (with Streptavidin, Nunc). After 1.5 hrs. incubation at 25° C. unbound antibody was removed by washing 6 times with 90 μL PBS (dispense and aspiration). The antigen-antibody complex was detected by an anti-rabbit antibody conjugated to POD (ECL anti-rabbit IgG-POD, Cat. No. NA9340V; POD=peroxidase). 20-30 min after adding 35 μL POD-substrate 3,3′,5,5′-tetramethyl benzidine (TMB; Piercenet, Cat. No. 34021) the optical density was determined at 370 nm. The EC50 value was calculated with a four parameter logistic model using GraphPad Prism 6.0 software.Murine Antigen:
[0364] The antigen, biotinylated murine LRP8, was incubated with 5 μL sample containing the anti-LRP8 antibody at a concentration of 250 ng / mL in a total volume of 25 μL in PBS, 0.5% BSA and 0.05% Tween in a 384 w microtiterplate (Maxisorb (with Streptavidin, Nunc). After 1.5 hrs. incubation at 25° C. unbound antibody was removed by washing 6 times with 90 μL PBS (dispense and aspiration). The antigen-antibody complex was detected by an anti-rabbit antibody conjugated to POD (ECL anti-rabbit IgG-POD, Cat. No. NA9340V). 20-30 min after adding 35 μL POD-substrate 3,3′,5,5′-tetramethyl benzidine (TMB, Piercenet, Cat. No. 34021) the optical density was determined at 370 nm. The EC50 value was calculated with a four parameter logistic model using GraphPad Prism 6.0 software.Example 9Pcr Amplification of V-Domains for SLIC Cloning
[0365] Total RNA was prepared from B-cell lysates (resuspended in RLT buffer (Qiagen, Cat. No. 79216) using the NucleoSpin 8 / 96 RNA kit (Macherey & Nagel; Cat. No. 740709.4, 740698) according to manufacturer's protocol. RNA was eluted with 60 μl RNase free water. 6 μl of RNA was used to generate cDNA by reverse transcriptase reaction using the Superscript III First-Strand Synthesis SuperMix (Invitrogen, Cat. No. 18080-400) and an oligo dT-primer according to the manufacturer's instructions. All steps were performed on a Hamilton ML Star System. 4 μl of cDNA were used to amplify the immunoglobulin heavy and light chain variable regions (VH and VL) with the AccuPrime SuperMix (Invitrogen, Cat. No. 12344-040) in a final volume of 50 μl using the primers rbHC.up and rbHC.do for the heavy chain and rbLC.up and rbLC.do for the light chain:rbHC.upAAGCTTGCCACCATGGAGACTGGGCTGCGCTGGCTTC(SEQ ID NO: 30)rbHC.doCCATTGGTGAGGGTGCCCGAG(SEQ ID NO: 31)rbLC.upAAGCTTGCCACCATGGACAYGAGGGCCCCCACTC(SEQ ID NO: 32)rbLC.doCAGAGTRCTGCTGAGGTTGTAGGTAC(SEQ ID NO: 33)
[0366] All forward primers were specific for the signal peptide (of respectively VH and VL) whereas the reverse primers were specific for the constant regions (of respectively CH1 and CL). The PCR conditions for the RbVH+RbVL were as follows: hot start at 94° C. for 5 min.; 35 cycles: 20 sec. at 94° C.; 20 sec. at 70° C.; 45 sec. at 68° C.; final extension at 68° C. for 7 min.
[0367] 8 μl of the 50 μl PCR solution were loaded on a 48 E-Gel 2% (Invitrogen, Cat. No. G8008-02). Positive PCR reactions were purified using the NucleoSpin Extract II kit (Macherey & Nagel; Cat. No. 740609250) according to manufacturer's protocol and eluted in 50 μl elution buffer. All purification steps were performed on a Hamilton ML Starlet System. 5 μl of purified VH and VL PCR solutions were used for DNA-sequencing.Example 10Ngs VH-PCR from PBMCs and Antigen-Enriched B-Cells
[0368] 4.2×10E6 PBMCs and 1.2×10E6 antigen-enriched B-cells were resuspended in 300 μl RLT Buffer (Qiagen; Cat. No. 79216). Total RNA was prepared from B-cell lysates using RNeasy Mini or Micro kit (Qiagen; Cat. No. 74134) according to manufacturer's protocol. RNA was eluted in 50 μl and 30 μl RNase free water, respectively, for PBMCs and antigen-enriched B-cells. 6 μl of RNA was used to generate cDNA by reverse transcriptase reaction using the Superscript III First-Strand Synthesis SuperMix (Invitrogen, Cat. No. 18080-400) and an oligo dT-primer according to the manufacturer's instructions.
[0369] 50-80 ng of cDNA were used to amplify the immunoglobulin heavy chain variable regions (VH) with the AccuPrime SuperMix (Invitrogen, Cat. No. 12344-040) in a final volume of 50 μl using the primers rbHCfinal_FS.up and rbHC_shortCH1_fs2.do:rbHCfinal_FS.upATGGAGACTGGGCTGCGCTGGCTTC(SEQ ID NO: 34)rbHC_shortCH1_fs2.doGGGAAGACTGATGGAGC(SEQ ID NO: 35)
[0370] The forward primer is specific for the signal peptide VH whereas the reverse primers specific for the constant regions is. The PCR conditions were as follows: Hot start at 94° C. for 3 min.; 22 cycles: 20 sec. at 94° C.; 20 sec. at 68° C.; 40 sec. at 68° C.; final extension at 68° C. for 5 min. Totally 6 PCR reactions were performed on each cell pool sample. 8 μl of one PCR reaction were loaded on a 12 E-Gel 2% (Invitrogen, Cat. No. G521802).
[0371] All PCR reactions respectively for the 2 B-cell libraries (PBMC-library; antigen-enriched B-cell-library) were purified with one column using the NucleoSpin Extract II kit (Macherey & Nagel; Cat. No. 740609) according to manufacturer's protocol and eluted in 50 μl elution buffer. 5 μl of cleaned VH PCR solutions were used for DNA-MiSeq sequencing.Example 11Template Preparation for NGS Sequencing
[0372] Paired-Ends Run 2×300 Base: Minimal DNA amount for Samples: 100 ng, good 500 ng
[0373] The NGS sequencing was run on MiSeq from Illumina. After purification on AMPure XP beads PCR templates were assessed on a DNA1000 Agilent BioAnalyzer Chip. The library preparation was performed using the TruSeq Nano DNA Sample Preparation Kit according to manufacturer's protocol.
[0374] The final libraries were quantified using qPCR technology. qPCR reactions were prepared according to the KAPA SYBR FAST qPCR protocol and run using the Roche Light Cycler 480. The samples were pooled and contrasted with PhiX.
[0375] In more detail, the libraries were analyzed by a paired-end Illumina MiSeq sequencing run with Illumina sequencing primers.
[0376] All reagents were thawed at RT just before starting experiment. The reagent cartridge was thawed in a water bath. The cartridge was inverted several times to ensure mixing of reagents and all air bubbles were removed by hitting the cartridge on the bench. 1 mL of 0.2 M NaOH was prepared by adding 200 μl 1 M NaOH to 800 μL laboratory-graded water. The prepared solution was vortexed, spun down and stored on ice. Flow cell was brought to RT, and carefully washed with laboratory-graded water, dried using kimtech precision wipes and inserted into the sequencer following the instructions. 5 μl of 4 nM DNA library pool was mixed with freshly diluted 0.2 M NaOH, vortexed briefly and spun down on a table top centrifuge. The solution was incubated for 5 min. at Room temperature and 990 μL pre-chilled HT1 was added and mixed by briefly vortexing. The resulting 20 pM denatured library in 1 mM NaOH was stored on Ice until further use. To obtain 600 μl of a 12 pM library, 360 μL of the 20 pM denatured library was diluted with 240 μl pre-chilled HT1, inverted several times to mix and then pulse centrifuged. The resulting 12 pM library was stored on ice until further use. To prepare 4 nM PhiX library, 2 μL of the 10 nM PhiX library control was added to 3 μl of 10 mM Tris-HCl, pH 8.5 with 0.1% Tween 20. The dilution was briefly vortexed and pulse centrifuged. To denature the PhiX Control 5 μl freshly diluted 0.2 M NaOH was added to the 5 μL of the prepared 4 nM PhiX library, vortexed briefly and spun down on a table top centrifuge. The solution was incubated for 5 min. at Room temperature and 990 μL pre-chilled HT1 was added and mixed by briefly vortexing. To obtain a 12.5 pM PhiX library, 375 μL of the 20 pM denatured PhiX solution was mixed with 225 μL Pre-chilled HT1, briefly vortexed and pulse centrifuged. 520 μL 12 pM Sample library and 80 μL 12.5 pM PhiX were combined to create a library with 15% PhiX control spike-in. The combined sample library and PhiX control were stored on ice until loaded onto the MiSeq reagent cartridge.Example 12Bioinformatics Analysis of NGS Sequences for Identification of Clonally Related VH Variants
[0377] Data from Illumina MiSeq consist of two paired and usually overlapping reads per sequence. All data have been analyzed using the following workflow:
[0378] Assembly of paired reads by FLASH (any other software tool should work as well)
[0379] FLASH available from http: / / ccb.jhu.edu / software / FLASH /
[0380] Using Flash v1.2.10 with DEFAULT PARAMETERS (no outies, min overlap 10 bp, max overlap 65 bp)
[0381] Result: Overlapped sequences without Illumina adaptors
[0382] Extraction of antibody variable domains:
[0383] Translating DNA to all 6 ORFs
[0384] For each ORF:
[0385] Searching peptide sequence for FW1 by comparing to a consensus FW1 sequence and counting the difference. Continuing if that value is above a defined threshold.
[0386] Alike searching for FW2, trying to connect to FW1 (area in between is CDR1)
[0387] Alike searching for FW3, trying to connect to FW2 (area in between is CDR2)
[0388] Alike searching for FW4, trying to connect to FW3 (area in between is CDR3)
[0389] Usually, in just 1 of the 6 ORFs a variable domain with the above described procedure can be identified. If multiples are found, a score that described the distance to the consensus is calculated and best ORF is selected.
[0390] For variable domains found, several values are. Most importantly, the closest germlines were detected by simply aligning the variable domain sequence to the available germline repertoire provided by IMGT and report the best hit. By this it also reports per sequence the V / D / J Germlines that are most likely to be the origin of those sequences.
[0391] Result: Table with one row per sequence containing all information about the contained variable domain.
[0392] Additional Analysis performed: Calculated #Mutations for each CDR / FR on DNA / PEP level compared to the reference sequences.Example 13Transient Transfection of NGS VH Variants with Parental VL
[0393] For recombinant expression of NGS variants, PCR-products coding for parental VL of B-cell clones were cloned as cDNA into expression vectors by the overhang cloning method (Haun, R. S., et al., BioTechniques 13 (1992) 515-518; Li, M. Z., et al., Nature Methods 4 (2007) 251-256) in an expression cassette containing the rabbit constant region to accept the VL region. The expression vectors contained an expression cassette consisting of a 5′ CMV promoter including intron A, and a 3′ BGH poly adenylation sequence. In addition to the expression cassette, the plasmids contained a pUC18-derived origin of replication and a beta-lactamase gene conferring ampicillin resistance for plasmid amplification in E. coli. Furthermore, the expression vector contained the rabbit kappa LC constant region to accept the VL regions.
[0394] Linearized expression plasmids coding for the kappa constant region and VL inserts were amplified by PCR using overlapping primers. Purified PCR products were incubated with T4 DNA-polymerase which generated single-strand overhangs. The reaction was stopped by dCTP addition. In the next step, plasmid and insert were combined and incubated with recA which induced site specific recombination. The recombined plasmids were transformed into E. coli. The next day the grown colonies were picked and tested for correct recombined plasmid by plasmid preparation, restriction analysis and DNA-sequencing.
[0395] Selected NGS VH variants were synthesized (Gene Art, Regensburg, Germany) and cloned as cDNA into expression vectors. The expression vectors contained an expression cassette consisting of a 5′ CMV promoter including intron A, and a 3′ BGH poly adenylation sequence. In addition to the expression cassette, the plasmids contained a pUC18-derived origin of replication and a beta-lactamase gene conferring ampicillin resistance for plasmid amplification in E. coli. Furthermore, the expression vector contained the rabbit IgG constant region designed to accept the VH regions.
[0396] For antibody expression, 500 ng of the isolated HC and LC plasmids were transiently co-transfected into 2 ml (96-well plate) of FreeStyle HEK293-F cells (Invitrogen, Cat. No. R790-07) by using 239-Free Transfection Reagent (Novagen) following procedure suggested by Reagent supplier. After 1-week cultivation the HEK supernatants were harvested, filtered (1.2 μm Supor-PALL) and purified with MabSelectSuRe (50 μl, GE Healthcare). Columns were equilibrated with 1×PBS. Samples were eluted with 2.5 mM HCl, pH 2.6 and neutralized with 10×PBS.Example 14Staining Procedure for Antigen (CDCP1) Specific Sort
[0397] The cells from the macrophage depletion step were used to perform the antigen specific sort. In a first round the cells were incubated with the biotinylated CDCP1 antigen at a concentration of 5 μg / ml on ice. After two washing steps the biotinylated and cell-bound CDCP1 was detected with an A647-streptavidin conjugate (Invitrogen). In parallel, the anti-IgG FITC (AbD Serotec, Dusseldorf, Germany) and the anti-IgM PE (BD Pharmingen) antibodies were added. The stained cells were washed two times. Finally, the PBMCs were resuspended in ice cold PBS and immediately subjected to the FACS analyses. The cell gate used for the single cell sorting was: rbIgM⊖ / rbIgG ⊕ / CDCP1⊕.Example 15Screening Hu CDCP1 Binders
[0398] Nunc Maxisorb streptavidin coated plates (MicroCoat, Cat. No. #11974998001) were coated with 25 μl / well biotinylated human CDCP1-AviHis fusion protein at a concentration of 200 ng / ml and incubated at 4° C. over night. After washing (2×90 μl / well with PBST-buffer (Phosphate Buffered Saline Tween-20)) 25 μl anti-CDCP1 antibody samples were added and incubated for one hour at RT. After washing (3×90 μl / well with PBST-buffer) 25 μl / well goat-anti-human IgG-HRP conjugate (Millipore, Cat. No. AP502P) was added in 1:1,000 dilution and incubated at RT for one hour on a shaker. After washing (4×90 μl / well with PBST-buffer) 25 μl / well TMB substrate (Roche Diagnostics GmbH, Cat. No. 11835033001) was added and incubated until OD reached 1.5-2.5. The reaction was stopped by the addition of 25 μl / well 1 N HCl. Measurement took place at 370 / 492 nm.Example 16Ngs VH-PCR from PBMCs and Antigen-Enriched (Panning Sample) B-Cells
[0399] 4*10E6 PBMCs and 352 antigen-enriched B-cells were resuspended in 350 μl RLT Buffer (Qiagen, Cat. No. 79216). Total RNA was prepared from B-cells lysate using RNeasy Mini or Micro kit (Qiagen, Cat. No. 74134) according to manufacturer's protocol. RNA was eluted in 50 μl and 30 μl RNase free water, respectively, for PBMCs and antigen-enriched B-cells. 6 μl of RNA was used to generate cDNA by reverse transcriptase reaction using the Superscript III First-Strand Synthesis SuperMix (Invitrogen, Cat. No. 18080-400) and an oligo dT-primer according to the manufacturer's instructions.
[0400] 50-80 ng of cDNA were used to amplify the immunoglobulin heavy chain variable regions (VH) with the AccuPrime SuperMix (Invitrogen, Cat. No. 12344-040) in a final volume of 50 μl using the primers rbHCfinal_FS.up and rbHC_shortCH1_fs2.do:rbHCfinal_FS.upATGGAGACTGGGCTGCGCTGGCTTC(SEQ ID NO: 34)rbHC_shortCH1_fs2.doGGGAAGACTGATGGAGC(SEQ ID NO: 35)
[0401] The forward primer is specific for the signal peptide VH whereas the reverse primers specific for the constant regions is. The PCR conditions were as follows: hot start at 94° C. for 3 min; 20 and 29 cycles (respectively for PBMC and antigen-enriched samples) of 20 sec. at 94° C.; 20 sec. at 68° C.; 40 sec. at 68° C.; final extension at 68° C. for 5 min. Totally 6-8 PCR reactions were performed each cell pool sample.CDCP1: 3 wt-rabbits (5571, 5565, 5566)Sample IDDescriptionAnimalPCRsCyclesG1 (5571)PBMC55716 × 50 μl PCR 20(Lympholite)a 1 μl cDNAG2 (5565)PBMC55656 × 50 μl PCR 20(Lympholite)a 1 μl cDNAG3 (5566)PBMC55666 × 50 μl PCR 20(Lympholite)a 1 μl cDNAG4 (5565)MΦ dep. + AG Sort55654 × 50 μl PCR 29a 4 μl cDNAG5 (5566)MΦ dep. + AG Sort55664 × 50 μl PCR 29a 4 μl cDNAG6 (5571)MΦ dep. + AG Sort55714 × 50 μl PCR 29a 4 μl cDNASequencing results: number of rabbit VH sequences and clustersH3totalnon-bad VHH3ClustersnamesequencespairablesequenceOKClustersn > 3G11,021,345149,11044,692827,54328,4878,506G21,040,062146,08047,609846,37321,8424,544G31,406,058245,44575,1061,085,50730,9476,876G41,019,453147,60141,484830,3684,7741,349G51,166,355171,69860,861933,7964,8471,055G61,166,807173,88855,718937,2013,753806Example 17Hek Transient Transfection of NGS VH Variants with Parental VLFor recombinant expression of NGS variants, PCR-products coding for parental VL of B-cell clones were cloned as cDNA into expression vectors by the overhang cloning method (Haun, R. S., et al., BioTechniques 13 (1992) 515-518; Li, M. Z., et al., Nature Methods 4 (2007) 251-256). The expression vectors contained an expression cassette consisting of a 5′ CMV promoter including intron A, and a 3′ BGH poly adenylation sequence. In addition to the expression cassette, the plasmids contained a pUC18-derived origin of replication and a beta-lactamase gene conferring ampicillin resistance for plasmid amplification in E. coli. Furthermore, the expression vector contained the rabbit kappa LC constant region to accept the VL regions.Linearized expression plasmids coding for the kappa constant region and VL inserts were amplified by PCR using overlapping primers. Purified PCR products were incubated with T4 DNA-polymerase which generated single-strand overhangs. The reaction was stopped by dCTP addition. In the next step, plasmid and insert were combined and incubated with recA which induced site specific recombination. The recombined plasmids were transformed into E. coli. The next day the grown colonies were picked and tested for correct recombined plasmid by plasmid preparation, restriction analysis and DNA-sequencing.
[0405] Selected NGS VH variants were synthesized (Gene Art, Regensburg, Germany) and cloned as cDNA into expression vectors. The expression vectors contained an expression cassette consisting of a 5′ CMV promoter including intron A, and a 3′ BGH poly adenylation sequence. In addition to the expression cassette, the plasmids contained a pUC18-derived origin of replication and a beta-lactamase gene conferring ampicillin resistance for plasmid amplification in E. coli. Furthermore, the expression vector contained the rabbit IgG constant region designed to accept the VH regions.
[0406] For antibody expression, 500 ng of the isolated HC and LC plasmids were transiently co-transfected into 2 ml (96-well plate) of HEK293-F cells (Invitrogen, Cat. No. R790-07) by using 239-Free Transfection Reagent (Novagen) following procedure suggested by Reagent supplier. After 1-week cultivation the HEK supernatants were harvested, filtered (1.2 μm Supor-PALL) and purified with MabSelectSuRe (50 μl, GE Healthcare). Columns were equilibrated with 1×PBS. Samples were eluted with 2.5 mM HCl, pH 2.6 and neutralized with 10×PBS.Example 18Human CDCP1 Binding ELISA
[0407] Nunc Maxisorb streptavidin coated plates (MicroCoat, Cat. No. #11974998001) were coated with 25 μl / well biotinylated human CDCP1-AviHis fusion protein at a concentration of 200 ng / ml and incubated at 4° C. over night. After washing (2×90 μl / well with PBST-buffer) 25 μl anti-CDCP1 antibody samples were added in a 1:2 dilution series starting at 5 μg / ml. Plates were incubated one hour at RT. After washing (3×90 μl / well with PBST-buffer) 25 μl / well of a mix of goat-anti-human IgG-HRP conjugate (Jackson, Cat. No. 109-036-098) and donkey-anti-rabbit IgG (GE Healthcare, Cat. No. NA9340V, Lot #389592,) was added in 1:9,000 dilution and incubated at RT for one hour on a shaker. After washing (4×90 μl / well with PBST-buffer) 25 μl / well TMB substrate (Roche, Cat. No. 11835033001) was added and incubated until OD reached 1.5-2.5. The reaction was stopped by addition of 25 μl / well 1 N HCl. Measurement took place at 370 / 492 nm.SEQ ID NOSEQUENCE36CQSVEESGGRLVTPGTPLTLTCTASGFSLSSYNMNWVRQAPGKGLEWIGYINKGGSAYYASWAKGRFTISRTSTTVDLKMTSPTTEDTATYFCVRSGGGGNLNLWGQGTLVTVSS37CQSVEESGGRLVTPGTPLTLTCTASGFSLSSYNMNWVRQAPGKGLEWIGYINKGGSAYYASWAKGRFTISRTSTTVDLKMTSPTPEDTATYFCVRSGGGGNLNLWGQGTLVTVSS38CQSVEESGGRLVTPGTPLTLTCTASGFSLSSYNMNWVRQAPGKGLAWIGYINKGGSAYYASWAKGRFTISKTSTTVDLKMTSPTTEDTATYFCVRSGGGGNLNLWGQGTLVTVSS39CQSVEESGGRLVTPGTPLTLTCTASGFSLSSYNMNWVRQAPGKGLAWIGYINKGGSAYYASWAKGRFTISKTSTTVDLKMTSPTTEDTATYFCVRSGGGGNLNLWGQGTLVTVSS40CQSVEESGGRLVTPGTPLTLTCTASGFSLSSYNMNWVRQAPGKGLAWIGYINKGGSAYYASWAKGRFTISKTSTTVDLKMTSPTTEDTATYFCVRSGGGGNLNLWGQGTLVTVSS41CQSVEESGGRLVTPGTPLTLTCTASGFSLSSYNMNWVRQAPGKGLAWIGYINKGGSAYYASWAKGRFTISKTSTTVDLKMTSPTTEDTATYFCVRSGGGGNLNLWGQGTLVTVSS42CQLVEESGGRLVTPGTPLTLTCTASGFSLSSYKMNWVRQAPGKGLEWIGYINKGGSAYYASWAKGRFTISRTSTTVDLKMTSPTPEDTATYVCGRSGGGGNLNLWGQGTLVTVSS43CQSVEESGGRLVTPGTPLTLTCTASGFSLSSYNMNWVRQAPGKGLEWIGYINKGGSAYYASWAKGRFTISKTSTTVDLKMTSPTTEDTATYFCVRSGGGGNLNLWGQGTLVTVSS44CQSVEESGGRLVTPGTPLTLTCTVSGIDLSRSAVGWFRQAPGKGLEYIGFIGSSGTTYCATWAKGRFTISKASTTVALKITSPTTEDTATYFCASRNYDDYTFDPWGPGTLVTVSS45CQSVEESGGRLVTPGTPLTLTCTVSGIDLSSYAVGWFRQAPGKGLEYIGFIGSSGTTYYATWAKGRFTISKASTTVSLKMTSPTTEDTATYFCASRNYDDYTFDPWGPGTLVTVSS46CQSVEESGGRLVTPGTPLTLTCTVSGIDLSRFAVGWFRQAPGKGLEYIGFIGSSGSTYYASWAKGRFTISKSSTTVDLKIPGPTTEDTATYFCASRNYDDYSFDSWGPGTLVTVAS47CQSVEESGGRLVTPGTPLTLTCTVSGIDLSRFAVGWFRQAPGKGLEYIGFIGSSGSTYYASWAKGRFTISKSSTTVDLKMPGPTTEDTATYFCASRNYDDYSFDSWGPGTLVTVSS48CQSVEESGGRLVTPGTPLTLTCTVSGIDLSRFAVGWFRQAPGKGLEYIGFIGSSGSTYYASWAKGRFTISKASTTVDLKMPGPTTEDTATYFCASRNYDDYSFDSWGPGTLVTVAS49CQSVEESGGRLVTPGTPLTLTCTVSGIDLSRFAVGWFRQAPGKGLEYIGFIGSSGSTYYASWAKGRFTISKSSTTVDLKMPGPTTEDTATYFCASRNYDDYSFDPWGPGTLVTVAS50CQSVEESGGRLVTPGTPLTLTCTVSGIDLSRFAVGWFRQAPGKGLEYIGFIGSSGSTYYASWAKGRFTISKSSTTVDLKMPSPTTEDTATYFCASRNYDDYSFDSWGPGTLVTVAS51CQSVEESGGRLVTPGTPLTLTCTVSGIDLSRFAVGWFRQAPGKGLEYIGFIGSSGSTYYASWAKGRFTISKSSTTVDLKMTGPTTEDTATYFCASRNYDDYSFDSWGPGTLVTVAS52CQSVEESGGRLVTPGTPLTLTCTVSGIDLSSYAVGWLRQAPGKGLEYIGFIGSSGTTYYATWAKGRFTISKASTTVSLKMTSPTTEDTATYFCASRNYDDYTFDPWGPGTLVTVSS53CQSVEESGGRLVTPGTPLTLTCTVSGIDLSSYAVGWFRQAPGKGLEYIGFFGSSGTTYYATWAKGRFTISKASTTVSLKMTSPTTEDTATYFCASRNYDDYTFDPWGPGTLVTVSS54CQSVEESGGRLVTPGTPLTLTCTVSGIDLSRSAVGWFRQAPGKGLEYIGFIGSSGSTYYASWAKGRFTISKSSTTVDLKMPGPTTEDTATYFCASRNYDDYSFDSWGPGTLVTVAS55CQSVEESGGRLVTPGTPLTLTCTVSGIDLSSYAVGWFRQAPGKGLEYIGFIGSSGTTYYATWGKGRFTISNASTTVSLKMTSPTTEDTATYFCASRNYDDYTFDPWGPGTLVTVSS56CQSVEESGGRLVTPGTPLTLTCTVSGIDLSRFAVGWFRQAPGKGLEYIGFIGSSGSTYYASWAKGRFTISKSSTTVDLKMTSLTTEDTATYFCASRNYDDYSFDSWGPGTLVTVAS57CQSVEESGGRLVTPGTPLTLTCTVSGIDLSSYAVGWFRQAPGKGLEYIGFIGTSGTTYYATWAKGRFTISKASTTVSLKMTSPTTEDTATYFCASRNYDDYTFDPWGPGTLVTVSS58CQSVEESGGRLVTPGTPLTLTCTVSGIDLSSYAVGWFRQAPGKGLEYIGFIGSSGTTYYANWAKGRFTISKASTTVSLKMTSPTTEDTATYFCASRNYDDYTFDPWGPGTLVTVSS59CQSVEESGGRLVTPGTPLTLTCTVSGIDLSRFAVGWFRQAPGKGLEYIGFIGSSGTTYYASWAKGRFTISKSSTTVDLKMPGPTTEDTATYFCASRNYDDYSFDSWGPGTLVTVAS60CQSVEESGGRLVTPGTPLTLTCTVSGFSLSAYVVSWVRQVPGEGLEWIGSLIFDSNRYYASWAKGRFTISKTSTTVDLTITSPTTEDTATYFCARGGYACDLWGQGTLVTVSS61CQSVEESGGRLVAPGTPLTLTCTVSGFSLSAYVVSWVRQVPGEGLEWIGSLIFDSNRYYASWAKGRFTISKTSTTVDLTITSPTIEDTATYFCARGGYASDLWGQGTLVTVSS62CQSVEESGGRLVTPGTPLTLTCTVSGFSLSAYVVSWVRQVPGEGLEWIGSLIFDSNRYYASWAKGRFTISKTSTTVDLTITSPTIEDTATYFCARGGYASDLWGQGTLVTVSS63CQSVEESGGRLVTPGTPLTLTCTVSGFSLSAYVVSWVRQVPGEGLEWIGSLIFDSNRYYASWAKGRFTISKTSTTVDLTITSPTIEDTATYFCARGWTYLDLWGQGTLVTVSS64CQSVEESGGRLVTPGTPLTLTCTVSGFSLSAYVVSWVRQVPGEGLEWIGSLIFDSNRYYASWAKGRFTISKTSTTVDPTITSPTIEDTATYFCARGGYASDLWGQGTLVTVSS65CQSVEESGGRLVTPGTPLTLTCTVSGFSLSAYVVSWVRQVPGEGLEWIGSLVFDTNTFYASWAKGRFTISKTSPTVDLTITSPTTEDTATYFCTRGGYASDLWGQGTLVTVSS67CQSVEESGGRLVTPGTPLTLTCTVSGIDLSRSAVGWFRQAPGKGLEYIGFIGSSGTTYCATWAKGRFTISKASTTVALKITSPTTEDTATYFCASRNYDDYTFDPWGPGTLVTVSS68CQSLEESGGRLVTPGTPLTLTCTVSGIDLNNDYMTWVRQAPGKGLEWIGIFYVATEITWYASWAKGRFTISKTSTTVDLKITSPTTEDTATYFCGRDGGYTGDGYAFELWGQGTLVTVSS69CQSLEESGGRLVTPGTPLTLTCTVSGIDLNNDYMTWVRQAPGKGLEWIGIFYVATEITWYASWAKGRFTISKTSTTVDLKITSPTTEDTATYFCGRDGGYTGDGYAFELWGQGTPVTVSS70SQSLEESGGRLVTPGTPLTLTCTVSGIDLNNDYMTWVRQAPGKGLEWIGIFYVATEITWYASWAKGRFTISKTSTTVDLKITSPTTEDTATYFCGRDGGYTGDGYAFELWGQGTLVTVSS71CQSLEESGGRLVTPGASLTLTCTVSGIDLNNDYMTWVRQAPGKGLEWIGIFYVATEITWYASWAKGRFTISKTSTTVDLKITSPTTEDTATYFCGRDGGYTGDGYAFELWGQGTLVTVSS72CQSLEESGGRLVTPGTPLTLTCTVSGIDLNNDYMTWVRQAPGKGLEWIGIFYVATEITWYASWAKGRFTISKTSTTVDLKITSTTTEDTATYFCGRDGGYTGDGYAFELWGQGTLVTVSS73CQSLEESGGDLVKPGASLTLTCTASGIDLNNDYMTWVRQAPGKGLEWIGIFYVATEITWYASWAKGRFTISKTSTTVDLKITRPTTEDTATYICGRDGGYTGDGYAFELWGQGTLVTVSS74CQSLEESGGRLVTPGTPLTLTCTVSGIDLNNDYMTWVRQAPGKGLEWIGIFYVATEITWYASWAKGRFTIAKTSTTVDLKITSPTTEDTATYFCGRDGGYTGDGYAFELWGQGTLVTVSS75CQSLEESGGRLVTPGTPLTLTCTVSGIDLTNDYMTWVRQAPGKGLEWIGIFYVATEITWYASWAKGRFTISKTSTTVDLKITSPTTEDTATYFCGRDGGYTGDGYAFELWGQGTLVTVSS76CQSLEESGGRLVTPGTPLTLTCAVSGIDLNNDYMTWVRQAPGKGLEWIGIFYVATEITWYASWAKGRFTISKTSTTVDLKITSPTTEDTATYFCGRDGGYTGDGYAFELWGQGTLVTVSS77CQSLEESGGRLVTPGTPLTLTCTVSGIDLNNDYMTWVRQAPGKGLEWIGIFYVETEITWYASWAKGRFTISKTSTTVDLKITSPTTEDTATYFCGRDGGYTGDGYAFELWGQGTLVTVSS78CQSLEESGGRLVTPGTPLTLTCTVSGIDLNNDYMTWVRQAPGKGLEWIGIFYVATEITWYASWTKGRFTISKTSTTVDLKITSPTTEDTATYFCGRDGGYTGDGYAFELWGQGTLVTVSS79CQSLEESGGRLVTPGTPLTLTCTVSGIDLNNDYMTWVRQAPGKGLEWIGIFYVATEITWYASWAKGRFTISKTSTAVDLKITSPTTEDTATYFCGRDGGYTGDGYAFELWGQGTLVTVSS80CQSLEESGGRLVTPGTPLTLTCTVSGIDLNDDYMTWVRQAPGKGLEWIGIFYVATEITWYASWAKGRFTISKTSTTVDLKITSPTTEDTATYFCGRDGGYTGDGYAFELWGQGTLVTVSSSequences taken from the drawingsSEQ ID NOSEQUENCE 81CQSVEESGGRLVTPGTPLTLTCTASGFSLSSYNMNWVRQAPGKGLEWIGYINKGGSAYYASWAKG 82CQSVEESGGRLVTPGTPLTLTCTASGFSLSSYNMNWVRQAPGKGLEWIGYINKGGSAYDANWAKG 83CPPVEESGGRLVTPGTPLTLTCTAPGFSLSSSNMNWVRQAPGKGLEWIGYINKGGSAYYASWAKG 84CQLVEESGGRLVTPGTPLTLTCTASGFSLSSYKMNWVRQAPGKGLEWIGYINKGGSAYYASWAKG 85CQSVEESGGRLVTPGTPLTLTCTASGFSLSSYNMNWVRQAPGKGLAWIGYINKGGSAYYASWAKG 86AAVLTQTPSPVSAAVGGTVTISCQSSPNILGNYLSWFQQKPGQPPKLLIYYTSTLASGVPSRFKG 87RFTISRTSTTVDLKMTSPTTEDTATYFCVRSGGGGNLNLWGQGTLVTVSS 88RFTISRTSTTVDLKMTSPTPEDTATYFCVRSGGGGNLNLWGQGTLVTVSS 89RFTISKTSTTVDLKMTSPTTEDTATYFCVRSGGGGNLNLWGQGTLVTVSS 90RFTISRTSTTVDLKMTSPTPEDTATYVCGRSGGGGNLNLWGQGTLVTVSS 91SGSGTQFTLTISDVQCDDAATYYCLGVYRSDSDNVFGGGTEVVVK 92CQSVEESGGRLVTPGTPLTLTCTASGFSLSSYNMNWVRQAPGKGLEWIGYINKGGSAYYASWAKG 93CQSVEESGGRLVTPGTPLTLTCTASGFSLSSYNMNWVRQAPGKGLEWIGYINKGGSAYDANWAKG 94CPPVEESGGRLVTPGTPLTLTCTAPGFSLSSSNMNWVRQAPGKGLEWIGYINKGGSAYYASWAKG 95CQLVEESGGRLVTPGTPLTLTCTASGFSLSSYKMNWVRQAPGKGLEWIGYINKGGSAYYASWAKG 96CQSVEESGGRLVTPGTPLTLTCTASGFSLSSYNMNWVRQAPGKGLAWIGYINKGGSAYYASWAKG 97AAVLTQTPSPVSAAVGGTVTISCQSSPNILGNYLSWFQQKPGQPPKLLIYYTSTLASGVPSRFKG 98RFTISRTSTTVDLKMTSPTTEDTATYFCVRSGGGXGNLNLWGQGTLVTVSS 99RFTISRTSTTVDLKMTSPTPEDTATYFCVRSGGGXGNLNLWGQGTLVTVSS100RFTISKTSTTVDLKMTSPTTEDTATYFCVRSGGGXGNLNLWGQGTLVTVSS101RFTISRTSTTVDLKMTSPTPEDTATYVCGRSGGGXGNLNLWGQGTLVTVSS102SGSGTQFTLTISDVQCDDAATYYCLGVYRSDSDNVFGGGTEVVVK103CQSVEESGGRLVTPGTPLTLTCTVSGFSLSAYVVSWVRQVPGEGLEWIGSLIFDSNRYYASWAKGRFTISKTSTTVDLTITSPITEDTATYFCARGGYACDLWGQGTLVTVSS104CQSVEESGGRLVTPGTPLTLTCTVSGFSLSAYVVSWVRQVPGEGLEWIGSLIFDSNRYYASWAKGRFTISKTSTTVDPTITSPTIEDTATYFCARGWTYLDLWGQGTLVTVSS105CQSLEESGGRLVTPGTPLTLTCTVSGIDLNNDYMTWVRQAPGKGLEWIGIFYVATNITWYASWAKGRFTISKSSTTVDLKITSPTTEDTATYFCGRDGGYTGDGYAFELWGQGTLVTVSS
Claims
1. -13. (canceled)14. A method for producing an antibody comprising an improved variant of a reference antibody variable domain, wherein the reference antibody variable domain comprises one or more of the following developability liabilities: (i) unpaired Cys-residues in the variable domain or the HVR, (ii) glycosylation sites, and (iii) degradation hot spots (Asp, Asn, Met) in the variable domain, and wherein the improved variant and the reference antibody variable domain when paired with the respective other domain bind to a target antigen, the method comprising:a) immunizing one or more animals with the target antigen;b) obtaining B cells from the one or more animals immunized with the target antigen;c) performing an assay to determine binding of antibodies produced by the obtained B cells to the target antigen;d) from the determining of c), selecting a reference antibody comprising a reference antibody variable domain that binds to the target antigen, wherein the reference antibody variable domain comprises one or more of the following developability liabilities: (i) unpaired Cys-residues in the variable domain or the HVR, (ii) glycosylation sites, and (iii) degradation hot spots (Asp, Asn, Met);e) performing PCR amplification of antibody variable domain encoding nucleic acids of the obtained B cells using consensus-specific primers to obtain amplification products;f) sequencing the amplification products;g) performing a sequence-identity / homology-based ranking of the antibody variable domain encoding nucleic acids based on a sequence alignment of the sequencing results of (f) to a nucleic acid sequence encoding the reference antibody variable domain;(h) identifying an improved variant encoded by one of the top 10 sequences of the sequence ranking of step (g) that does not comprise one or more of the following i) unpaired Cys-residues in the variable domain or the HVR, ii) glycosylation sites, and iii) degradation hot-spots (Asp, Asn or Met); andi) producing an antibody comprising the improved variant of the reference antibody variable domain.
15. The method of claim 14, wherein the reference antibody variable domain and variable domains of the antibodies produced by the obtained B cells are annotated according to the Wolfguy numbering scheme.
16. The method according to claim 14, wherein differences in the sequence are annotated in the form: reference antibody amino acid residue-position-variant antibody amino acid residue, and wherein the sequences are grouped into a mutation tuple.
17. The method according to claim 15, wherein differences in the sequence are annotated in the form: reference antibody amino acid residue-position-variant antibody amino acid residue, and wherein the sequences are grouped into a mutation tuple18. The method according to claim 14, wherein the homology-based ranking comprises ranking aligned sequences based on change of the physico-chemical properties resulting from amino acid differences of variable domain of antibodies produced by the obtained B cells to the reference antibody variable domain, wherein the change of the physico-chemical properties is determined using a mutation risk score, wherein the mutation risk score is determined based on the following Table, wherein residues that are not explicitly given in this Table are weighted with the value one:Wolfguy Wolfguy IndexWeightIndexWeight1010.215121021.11522.610301531.21040.51542.31050.21551.91060.81563.71070.515741080.815841090.419341100.21944111019541120.41963.311301973.91140.11982.61150.61993.41160.12513.81170.12521.91180.225341190.22543.81201.22553.612102564122428741230288412422893.91250.52903.420142913.72022.62922.120322933.62041.72942.320512952.32060296120702971.22080.72982.42091.62990.5210235142111.33523.52123.135342130.93543.52143.13553.53010.135633021.235733031.735833040.335933051.83603306036133072.43623308036333091.23643334136533351366333613673337138233100.838333110.938433120.838533132.538633140.138733151.53883316038933171.539033180.839133190.139233201.439333210.239433220.33953.53230.239633241.83973.53250.63981.53262.2399333313270.632833292.533043312.93322.84013.540224030.540414050.34060.140714080.140914100.14110.1with positions 101 to 125 corresponding to heavy chain variable domain framework 1, positions 151 to 199 corresponding to CDR-H1, positions 201 to 214 corresponding to heavy chain variable domain framework 2, positions 251 to 299 corresponding to CDR-H2, positions 301 to 332 corresponding to heavy chain variable domain framework 3, positions 351 to 399 corresponding to CDR-H3, positions 401 to 411 corresponding to heavy chain variable domain framework 4.
19. The method of claim 14, wherein the obtaining B cells of step (b) comprises enriching for antigen-specific antibody expressing B-cells.
20. A method for producing an improved variant of a reference antibody variable domain, wherein the reference antibody variable domain comprises one or more of the following developability liabilities: (i) unpaired Cys-residues in the variable domain or the HVR, (ii) glycosylation sites, and (iii) degradation hot spots (Asp, Asn, Met) in the variable domain, and wherein the improved variant and the reference antibody variable domain when paired with the respective other domain bind to a target antigen, the method comprising:a) immunizing one or more animals with the target antigen;b) obtaining B cells from the one or more animals immunized with the target antigen;c) performing an assay to determine binding of antibodies produced by the obtained B cells to the target antigen;d) from the determining of c), selecting a reference antibody comprising a reference antibody variable domain that binds to the target antigen, wherein the reference antibody variable domain comprises one or more of the following developability liabilities: (i) unpaired Cys-residues in the variable domain or the HVR, (ii) glycosylation sites, and (iii) degradation hot spots (Asp, Asn, Met);e) performing PCR amplification of antibody variable domain encoding nucleic acids of the obtained B cells using consensus-specific primers to obtain amplification products;f) sequencing the amplification products;g) producing an antibody comprising an improved variant of a reference antibody variable domain,wherein the improved variant is encoded by one of top ten sequences identified by a sequence-identity / homology-based ranking of the antibody variable domain encoding nucleic acids based on a sequence alignment of the sequencing results of (f) to the reference antibody variable domain encoding nucleic acid sequence,and wherein the improved variant does not comprise one or more of the following i) unpaired Cys-residues in the variable domain or the HVR, ii) glycosylation sites, and iii) degradation hot-spots (Asp, Asn or Met).
21. The method of claim 20, wherein the reference antibody variable domain and variable domains of the antibodies produced by the obtained B cells are annotated according to the Wolfguy numbering scheme.
22. The method according to claim 20, wherein differences in the sequence are annotated in the form: reference antibody amino acid residue-position-variant antibody amino acid residue, and wherein the sequences are grouped into a mutation tuple.
23. The method according to claim 21, wherein differences in the sequence are annotated in the form: reference antibody amino acid residue-position-variant antibody amino acid residue, and wherein the sequences are grouped into a mutation tuple24. The method according to claim 20, wherein the homology-based ranking comprises ranking aligned sequences based on change of the physico-chemical properties resulting from amino acid differences of variable domain of antibodies produced by the obtained B cells to the reference antibody variable domain, wherein the change of the physico-chemical properties is determined using a mutation risk score, wherein the mutation risk score is determined based on the following Table, wherein residues that are not explicitly given in this Table are weighted with the value one:Wolfguy Wolfguy IndexWeightIndexWeight1010.215121021.11522.610301531.21040.51542.31050.21551.91060.81563.71070.515741080.815841090.419341100.21944111019541120.41963.311301973.91140.11982.61150.61993.41160.12513.81170.12521.91180.225341190.22543.81201.22553.612102564122428741230288412422893.91250.52903.420142913.72022.62922.120322933.62041.72942.320512952.32060296120702971.22080.72982.42091.62990.5210235142111.33523.52123.135342130.93543.52143.13553.53010.135633021.235733031.735833040.335933051.83603306036133072.43623308036333091.23643334136533351366333613673337138233100.838333110.938433120.838533132.538633140.138733151.53883316038933171.539033180.839133190.139233201.439333210.239433220.33953.53230.239633241.83973.53250.63981.53262.2399333313270.632833292.533043312.93322.84013.540224030.540414050.34060.140714080.140914100.14110.1with positions 101 to 125 corresponding to heavy chain variable domain framework 1, positions 151 to 199 corresponding to CDR-H1, positions 201 to 214 corresponding to heavy chain variable domain framework 2, positions 251 to 299 corresponding to CDR-H2, positions 301 to 332 corresponding to heavy chain variable domain framework 3, positions 351 to 399 corresponding to CDR-H3, positions 401 to 411 corresponding to heavy chain variable domain framework 4.
25. The method of claim 20, wherein the obtaining B cells of step (b) comprises enriching for antigen-specific antibody expressing B-cells.
26. A method for producing an improved variant of a reference antibody variable domain, wherein the reference antibody variable domain comprises one or more of the following developability liabilities: (i) unpaired Cys-residues in the variable domain or the HVR, (ii) glycosylation sites, and (iii) degradation hot spots (Asp, Asn, Met) in the variable domain, and wherein the improved variant and the reference antibody variable domain when paired with the respective other domain bind to a target antigen, the method comprising:a) performing PCR amplification of antibody variable domain encoding nucleic acids of B cells using consensus-specific primers to obtain amplification products, the B cells having been obtained from one or more animals immunized with the target antigen;b) sequencing the amplification products;c) producing an antibody comprising an improved variant of a reference antibody variable domain, wherein the improved variant is encoded by one of top ten sequences identified by a sequence-identity / homology-based ranking of the antibody variable domain encoding nucleic acids based on a sequence alignment of the sequencing results of (b) to the reference antibody variable domain encoding nucleic acid sequence, and wherein the improved variant does not comprise one or more of the following i) unpaired Cys-residues in the variable domain or the HVR, ii) glycosylation sites, and iii) degradation hot-spots (Asp, Asn or Met).
27. The method of claim 26, wherein the reference antibody variable domain and variable domains of the antibodies produced by the obtained B cells are annotated according to the Wolfguy numbering scheme.
28. The method according to claim 26, wherein differences in the sequence are annotated in the form: reference antibody amino acid residue-position-variant antibody amino acid residue, and wherein the sequences are grouped into a mutation tuple.
29. The method according to claim 27, wherein differences in the sequence are annotated in the form: reference antibody amino acid residue-position-variant antibody amino acid residue, and wherein the sequences are grouped into a mutation tuple30. The method according to claim 26, wherein the homology-based ranking comprises ranking aligned sequences based on change of the physico-chemical properties resulting from amino acid differences of variable domain of antibodies produced by the obtained B cells to the reference antibody variable domain, wherein the change of the physico-chemical properties is determined using a mutation risk score, wherein the mutation risk score is determined based on the following Table, wherein residues that are not explicitly given in this Table are weighted with the value one:Wolfguy Wolfguy IndexWeightIndexWeight1010.215121021.11522.610301531.21040.51542.31050.21551.91060.81563.71070.515741080.815841090.419341100.21944111019541120.41963.311301973.91140.11982.61150.61993.41160.12513.81170.12521.91180.225341190.22543.81201.22553.612102564122428741230288412422893.91250.52903.420142913.72022.62922.120322933.62041.72942.320512952.32060296120702971.22080.72982.42091.62990.5210235142111.33523.52123.135342130.93543.52143.13553.53010.135633021.235733031.735833040.335933051.83603306036133072.43623308036333091.23643334136533351366333613673337138233100.838333110.938433120.838533132.538633140.138733151.53883316038933171.539033180.839133190.139233201.439333210.239433220.33953.53230.239633241.83973.53250.63981.53262.2399333313270.632833292.533043312.93322.84013.540224030.540414050.34060.140714080.140914100.14110.1with positions 101 to 125 corresponding to heavy chain variable domain framework 1, positions 151 to 199 corresponding to CDR-H1, positions 201 to 214 corresponding to heavy chain variable domain framework 2, positions 251 to 299 corresponding to CDR-H2, positions 301 to 332 corresponding to heavy chain variable domain framework 3, positions 351 to 399 corresponding to CDR-H3, positions 401 to 411 corresponding to heavy chain variable domain framework 4.
31. The method of claim 26, wherein the obtaining B cells of step (b) comprises enriching for antigen-specific antibody expressing B-cells.
32. The method of claim 14, wherein the B-cells are obtained from the same immunization campaign as a B cell expressing the reference antibody.
33. The method of claim 20, wherein the B-cells are obtained from the same immunization campaign as a B cell expressing the reference antibody.
34. The method of claim 26, wherein the B-cells are obtained from the same immunization campaign as a B cell expressing the reference antibody.
35. A cell comprising a nucleic acid encoding the improved variant of the reference antibody variable domain of claim 14.
36. A cell comprising a nucleic acid encoding the improved variant of the reference antibody variable domain of claim 20.
37. A cell comprising a nucleic acid encoding the improved variant of the reference antibody variable domain of claim 26.