Natural and artificial intelligent methods for discovering humanlike nanobodies with low immunogenicity
By integrating in silico and experimental techniques with generative ML models, the method effectively addresses the challenge of discovering nanobodies with low immunogenicity, enhancing the efficiency and effectiveness of nanobody development.
Patent Information
- Application Number
- PCT/CN2024/130797
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-08
- Filing Date
- 2024-11-08
- Publication Date
- 2025-05-15
AI Technical Summary
Current methods for discovering nanobodies with low immunogenicity are limited by the need for extensive humanization processes, which can introduce affinity and solubility issues, and rely on transgenic animals with limited repertoire diversity.
The combination of in silico and experimental methods to discover low immunogenicity nanobodies from natural immune repertoires of camelids, synthetic libraries generated by generative ML models, and de novo design of VHHs for specific targets.
This approach enables the efficient generation of nanobodies with reduced immunogenicity, overcoming the limitations of traditional methods by leveraging natural and artificial intelligence to identify low-risk antigen-binding sites.
Smart Images

Figure PCTCN2024130797-FTAPPB-I100001 
Figure PCTCN2024130797-FTAPPB-I100002 
Figure PCTCN2024130797-FTAPPB-I100003
Abstract
Description
NATURAL AND ARTIFICIAL INTELLIGENT METHODS FOR DISCOVERING HUMANLIKE NANOBODIES WITH LOW IMMUNOGENICITY
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims the benefit of U.S. Provisional Patent Application No. 63 / 597, 124, filed November 8, 2023, which is incorporated by reference in its entirety for all purposes.BACKGROUND OF THE INVENTION
[0003] The present application generally relates to nanobody discovery. Specifically, the application describes methods of discovering humanlike camelids nanobodies (VHH) with low immunogenicity for a specific target from natural immune repertoires of camelids, synthetic libraries constructed by generative ML models and de novo design.
[0004] Monoclonal antibodies have been extensively used in the areas of therapeutics (Heo, 2022) . Since first drug approval in 1986, more than 100 antibody-based drugs have been approved by the FDA. However, many failed to reach approval status. Recent review study (Sun et al., 2022) listed thirteen reasons for the failure, one of which is immunogenicity. In 2016, Pfizer terminated a Phase III trial of bococizumab for treating patients with high LDL because of immunogenicity issue.
[0005] Heavy-chain-only antibodies (HCAbs) exist naturally in the immune repertoire of camelids and cartilaginous fish. Camelids HCAbs with homodimer form consist of one variable region (VHH) and two constant domains. VHH, sufficient for antigen binding, has dimensions in the nanometer range with about 12-15kD molecular weight. It is also known as a nanobody because of its nanometer size or single-domain antibody. In 2019, the FDA approved the first nanobody based medicine caplacizumab for the treatment of acquired thrombotic thrombocytopenic purpura in adults. More recently, envafolimab, a single domain anti-PD-L1 antibody, was approved in China for the treatment of MSI-H or dMMR advanced solid tumors.
[0006] It was found that there are sets of non-classical VHH (without FR2 hydrophilic amino acids) which are derived from the same gene locus, IGHV1, IGHV3 or IGHV4, D and J as conventional IgG1 do. These VH-like single domains with an IGHV1, IGHV3 or IGHV4 imprints contain a conventional FR2 GLEW motif and account for approximately 10%and more of total HCAb. These non-classic VHH nanobodies offer a great advantage over classic VHH nanobodies as therapeutic leads because the lack of hydrophilic amino acids (F / Y42, E / K49, R50, and G / L52) featured in VHHs may greatly reduce immunogenicity risk.
[0007] Some unique features of nanobodies (VHH) suggest that VHH may have low immunogenicity in human (Rossotti et al., 2022) : small size which decreases the number of potentially immunogenic epitopes; stable and good solubility which reduces the chance of immunogenic aggregate formation, rapid blood clearance in non-half-life extended cases and high sequence homology to human IGHV3 family germlines. However, some other unique features may suggest that VHH are inherently immunogenic: exposed preexisting VH / VL interface region of VHH due to lack of VL, long CDR3 sequences and possible novel CDR3 structure fold in nanobodies.
[0008] To reduce potential immunogenicity of VHH, several methods and / or strategies have been developed. Humanization, the process of replacing the xenogeneic sequences with human sequences in nanobody framework region, is the most common method to reduce the immunogenicity of nanobodies. Despite the expected low immunogenicity of nanobodies, humanization is performed routinely as part of the development process. However, such engineering process may introduce additional issues to the nanobody, such as affinity reduction, solubility and other developability issues. The more changes introduced, the more likely that engineered nanobodies will have those issues. Synthetic humanized nanobody libraries like NaLi-H1 (Moutel et al., 2016) have been developed to generate humanized nanobodies by-passing humanization step. However, nanobodies from such libraries may have non-favorable biophysical properties and further engineering may be needed because of lack of B-cells positive and negative selections during B-cells development in germinal center in vivo. To overcome such issues in in vitro systems, several transgenic animals (Clarke et al., 2019, Teng et al., 2020) have been developed to produce fully human nanobodies using complicated genetic engineering processes. However, repertoire diversity of these animals may be limited compared to natural repertoires in camelids.
[0009] Recent advances in machine learning (ML) methods, especially in the area of generative ML methods and large protein language models, make it possible to generate or design protein, antibody sequences in silico (Ferruz et al., 2022, Tileli et al., 2020) . More recently, a generative model of protein backbones (RFdiffusion) has been developed to design protein binders (Joseph et al, 2022) . In combination with proteinMPNN (Dauparas et al., 2022) , a deep learning-based tool to design sequences from structures, de novo protein design becomes a reality.SUMMARY OF THE INVENTION
[0010] The present invention combines in silico and experimental methods to discover low immunogenicity nanobodies (VHH) from natural immune repertoires of camelids, synthetic libraries generated by generative ML models and de novo design of low immunogenicity VHHs for a specific target.
[0011] Thus, in one aspect, the present application discloses a method of generating a low immunogenicity nanobody (VHH) specific to an antigen from natural immune repertoires of camelids. The method comprises the steps of:
[0012] a) selecting camelid germlines having predicted low immunogenicity;
[0013] b) constructing libraries from camelid repertoires;
[0014] c) selecting VHH sequences from the libraries in (b) that are derived from the camelid germlines in (a) ; and
[0015] d) generating a nanobody specific to the antigen from the VHH sequences of (c) .
[0016] In another aspect, the present application discloses a method of generating a low immunogenicity nanobody (VHH) specific to an antigen using a generative machine learning model. The method comprises the steps of:
[0017] a) selecting VHH sequences having predicted low immunogenicity;
[0018] b) training the generative machine learning model with the VHH sequences of (a) ;
[0019] c) using the trained generative machine learning model of (b) to generate synthetic VHH sequences;
[0020] d) constructing libraries from the synthetic VHH sequences generated in (c) ; and
[0021] e) generating a nanobody specific to the antigen from the libraries in (d) .
[0022] In yet another aspect, the present application discloses a method of generating a low immunogenicity nanobody (VHH) specific to an antigen using de novo design. The method comprises the steps of:
[0023] a) selecting VHH templates having predicted low immunogenicity;
[0024] b) using a first machine learning model to generate protein backbone structures that specifically bind to the antigen based on the structures of the VHH templates in (a) and the structure of the antigen;
[0025] c) using a second machine learning model to generate synthetic VHH sequences based on the protein backbone structures in (b) ; and
[0026] d) synthesize a nanobody specific to the antigen having the synthetic VHH sequence in (c) .BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Fig. 1 shows the evolution of in silico immunogenicity analysis methods.
[0028] Fig. 2 is a graph that shows the correlation of homology score vs ADA incidence rate.
[0029] Fig. 3 is a graph that shows the correlation of rarity score vs ADA incidence rate.
[0030] Fig. 4 is a graph that shows the correlation of 9-mer score vs ADA incidence rate.
[0031] Figs. 5A-D are graphs that show the immunogenicity differences between classical (A, C) and non-classical (B, D) VHHs as analyzed by rarity score (A, B) and 9-mer score (C, D) . These VHHs are all from Alpaca.
[0032] Figs. 6A and 6B are graphs that show the immunogenicity analysis of Alpaca germlines using rarity score (A) and 9-mer score (B) .
[0033] Figs. 7A and 7B are graphs that show the immunogenicity analysis of Llama germlines using rarity score (A) and 9-mer score (B) .
[0034] Figs. 8A and 8B are graphs that show the immunogenicity analysis of Bactrian germlines using rarity score (A) and 9-mer score (B) .
[0035] Figs. 9A and 9B are graphs that show the immunogenicity analysis of Dromedary germlines using rarity score (A) and 9-mer score (B) .
[0036] Figs. 10A and 10B are graphs that show the immunogenicity analysis of Mouse germlines using rarity score (A) and 9-mer score (B) .
[0037] Figs. 11A-D are graphs that show the rarity score (A, B) and 9-mer score for Alpaca VHHs from low immunogenicity germlines (A, C) comparing to VHHs from high immunogenicity germlines (B, D) .
[0038] Figs. 12A-C are graphs that show the correlations of 9-mer scores with mismatch values for VHHs from Alpaca (A) , Llama (B) and Bactrian (C) .
[0039] Figs. 13A and 13B are graphs that show the 9-mer score distribution for VHHs from low immunogenicity germlines and with mismatch score smaller than 11 (A) and 13 (B) .
[0040] Fig. 14 is a diagram that shows the experimental methods to build libraries with low immunogenicity VHHs.
[0041] Fig. 15 is a graph that shows the multiple sequence alignments of the first 180 bps of top 15 VH3 genes as listed in table 4. Boxed area is the region used to design capture oligo.
[0042] Figs. 16A-D are graphs that show in silico immunogenicity scores distribution differences between VHHs having a tryptophan (W) at IMGT position 118 (A, C) and VHHs having an arginine (R) at IMGT position 118 (B, D) . VHHs are analyzed by rarity score (A, B) and 9-mer score (C, D) . These VHHs are all from Alpaca.
[0043] Fig. 17 is a diagram that shows the experimental methods to build libraries enriching VHHs with R at IMGT position 118.
[0044] Fig. 18 is a graph that shows the 9-mer score differences between libraries built using primers in Table 6 vs primers amplifying all VHHs.
[0045] Fig. 19 is a diagram that shows the experimental methods to build libraries with low immunogenicity VHHs combining two methods together.
[0046] Figs. 20A and 20B are diagrams that show the generative ML methods for building libraries with low immunogenicity VHHs.
[0047] Fig. 21 is a diagram that shows a reinforced learning agent utilizing a generated ML that has been pre-trained with low immunogenicity VHH sequences.
[0048] Fig. 22 is a diagram that shows the use of RFdiffusion and ProteinMPNN to VHHs with low immunogenicity for a specific target.
[0049] Figs. 23A-D are graphs that show the immunogenicity scores for synthetic VHH sequences generated by machine learning methods based on template VHHs classical-3Q-OKT3 and non-classical- 2Q-OKT3. Figs. 23 A and C show the rarity scores and 9-mer scores of classical-3Q-OKT3-based synthetic VHH sequences, and Figs. 23 B and D show the rarity scores and 9-mer scores of non-classical-2Q-OKT3-based synthetic VHH sequences.
[0050] Fig. 24 is a graph that shows the binding activities of synthetic VHH sequences generated by machine learning methods based on template VHHs classical-3Q-OKT3 and non-classical-2Q-OKT3 in Jurkat cell binding experiment.
[0051] Fig. 25 is a graph that shows the immunogenicity scores for VHH sequences generated from generative AI models trained using the total VHH dataset vs VHH dataset with predicted low immunogenicity.DETAILED DESCRIPTION OF THE INVENTION
[0052] Definitions
[0053] The term “homology score” refers to a measure of similarity between VHH sequences and human germline sequences. In some embodiments, the value is the blastp bit score of top matched alignment between VHH sequences and human germline sequences.
[0054] The term “rarity score” refers to a measure of similarity between VHH sequences and human germline sequences. In some embodiments, the value is calculated based on framework regions of germlines. Before calculation, a profile of each residue usage percentage in each length of 4 framework regions is determined based on all human IGHV germlines. For each VHH sequence, the residue in each position of 4 framework regions is compared to the profile of that framework region of same length and the rarity score for each position of framework region is calculated based on usage percentage of that residue divided by the top usage percentage of same position. The rarity score for a sequence is an average of rarity scores of all framework region residues.
[0055] The term “9-mer score” refers to a measure of similarity between VHH sequences and human antibody sequences found in human immune repertoires and with considerations of natural secreted proteins in human body and Tregitopes. In some embodiments, the value is calculated by traversing the VHH sequences in sliding windows of 9-mer length and computes the frequency of occurrence of each 9-mer peptide in the human population. The peptide score then be adjusted, through comparison with human secreted protein dataset and Tregitope dataset. Subsequently, the scores from individual 9-mers are then averaged to generate an overall score, which provides a quantitative measure of predicted VHH immunogenicity.
[0056] The term “enriched” is intended to refer to component of a composition (e.g., a particular type of cells) that is more concentrated (e.g., at least 2x, at least 5x, at least 10x, at least 50x, at least 100x, at least 500x, at least 1,000x) , relative to other components in the sample (e.g., other cells) than prior to enrichment. In some cases, something that is enriched may represent a significant percent (e.g., greater than 2%, greater than 5%, greater than 10%, greater than 20%, greater than 50%, or more, usually up to about 90%-100%) of the sample in which it resides.
[0057] The term VHH sequence or camelid germline having a “predicted low immunogenicity VHH” refers to a VHH sequence that does not or is less likely to elicit a host immune response against it in humans, or produces or is likely to produce a low risk host immune response against it in humans, and a camelid germline that is likely to produce such VHH sequences. Immunogenicity or potential immunogenicity can be measured by any method known in the art for assessing the immunogenicity of antibodies in humans, such as episax analysis, Epibase Assay or anti-drug-antibody (ADA) assay. In the present application, a camelid germline may be considered having a predicted low immunogenicity if the sequence has a rarity score no lower than 88%, preferably no lower than 89%, and more preferably no lower than 90%and / or a 9-mer score no higher than 7, preferably no higher than 6, and more preferably no higher than 5. In a preferred embodiment, a camelid germline having a predicted low immunogenicity if the germline sequence has a rarity score no lower than 88%and a 9-mer score no higher than 7. In the present application, a VHH sequence may be considered having a predicted low immunogenicity if the sequence has a rarity score higher than 85%, preferably higher than 88%, and more preferably higher than 90%and / or a 9-mer score lower than 40, preferably lower than 35, and more preferably lower than 30. In a preferred embodiment, a VHH sequence having a predicted low immunogenicity if the sequence has a rarity score higher than 85%and a 9-mer score lower than 40. Potential immunogenicity of germlines and VHH sequences can also be predicted by other methods known in the art, for example, homology to human germlines.
[0058] The term “mismatch score” refers to the measure for determining SHM rate. It is calculated as an average number of mismatches in 100 bp alignment with the best matched germline gene.
[0059] The term “protein backbone” refers to a continuous chain of atoms that runs throughout the length of a protein, consisting of a repeated sequence of three atoms (nitrogen, alpha-carbon, carbon) .
[0060] Provided in the present application are methods of generating a low immunogenicity nanobody (VHH) specific to an antigen from natural immune repertoires of camelids, a generative machine learning model, and / or de novo design.
[0061] In some embodiments, the method is directed to generating a low immunogenicity nanobody (VHH) specific to an antigen, comprising:
[0062] a) selecting camelid germlines having predicted low immunogenicity;
[0063] b) constructing libraries from camelid repertoires;
[0064] c) selecting VHH sequences from the libraries in (b) that are derived from the camelid germlines in (a) ; and
[0065] d) generating a nanobody specific to the antigen from the VHH sequences of (c) .
[0066] In some embodiments, the libraries are constructed from repertoires produced by camelids having high percentage of VHHs from the camelid germlines that have predicted low immunogenicity or comprising the nanobodies from such camelid germlines.
[0067] To assess immunogenicity of antibody sequences, several in silico analysis methods have been developed as summarized in Fig. 1. These methods include simple homology of variable region sequences to human germlines, profile-based rarity score analysis and more recently human repertoire-based metrics (Prihoda et al., 2022) . We developed a 9-mer score method by combining human repertoire-based metrics, Tregitope and self-epitope information. All three methods showed significant correlation with the clinical reported anti-drug antibody (ADA) incidence rate (Fig. 2, 3, 4) with 9-mer score showing the highest correlation. When applying rarity score and 9-mer score analysis methods to Alpaca VHHs, we noticed that non-classical VHHs tend to have higher rarity score and lower 9-mer score as comparing to classical VHHs (Fig. 5) , indicating that non-classical VHHs will have lower immunogenicity.
[0068] To analyze germlines used in camelids species, we collected germlines for Alpaca, Dromedary and Bactrian from public databases (https: / / www. imgt. org / , https: / / www. ncbi. nlm. nih. gov / ) and determined the germlines for Llama by sequencing of PCR amplified fragments of variable region genes from 7 animals. The result is summarized in Tables 1 and 2. Based on the existence of hallmark residues in the FR2 region, these germlines can be classified into two categories: VH germlines, mainly for conventional antibodies and VHH germlines, mainly used for nanobodies.
[0069] Table 1. Camelids heavy chain variable germline genes and alleles collected from literature, extracted from genomes and sequenced from our own efforts
[0070] Table 2. Camelids heavy chain constant genes and corresponding hinge sequences collected from literature, extracted from genomes and sequenced from our own efforts
[0071] Non-classical VHHs used VH germlines, the ones without hallmark residues in FR2 region. To understand why non-classical VHHs have lower predicted immunogenicity than classical VHHs, we applied in silico immunogenicity analysis to camelids germlines. As results showed in Fig. 6, 7, 8, 9, all 4 species of camelids include a group of VH germlines with low predicted immunogenicity. For comparison, we also analyzed the immunogenicity of Mouse germlines (Fig. 10) . Overall Mouse germlines displayed higher predicted immunogenicity than those in camelids. Based on the distributions of Alpaca germline rarity scores and 9-mer scores (Fig. 6) , we grouped germlines into three immunogenicity categories as summarized in Table 3: low (rarity score >= 88%and 9-mer score <= 7) , medium (rarity score >= 88%or 9-mer score <= 7 and not in low category) and high (rarity score < 88%and 9-mer score > 7) .
[0072] Table 3. Alpaca germlines are classified into three categories: low (rarity score >= 88%and 9-mer score <= 7) , medium (rarity score >= 88%or 9-mer score <= 7 and not in low category) and high (rarity score ≤ 88%and 9-mer score ≥ 7)
[0073] In another embodiment, the following low immunogenicity (rarity score >= 88%and 9-mer score <= 7) Llama germlines are identified: IGHV3_1_llama (SEQ ID NO: 94) , IGHV3_105_llama (SEQ ID NO: 95) , IGHV3_106_llama (SEQ ID NO: 96) , IGHV3_107_llama (SEQ ID NO: 97) , IGHV3_108_llama (SEQ ID NO: 98) , IGHV3_11_llama (SEQ ID NO: 99) , IGHV3_114_llama (SEQ ID NO: 100) , IGHV3_116_llama (SEQ ID NO: 101) , IGHV3_12_llama (SEQ ID NO: 102) , IGHV3_120_llama (SEQ ID NO: 103) , IGHV3_121_llama (SEQ ID NO: 104) , IGHV3_122_llama (SEQ ID NO: 105) , IGHV3_123_llama (SEQ ID NO: 106) , IGHV3_124_llama (SEQ ID NO: 107) , IGHV3_126_llama (SEQ ID NO: 108) , IGHV3_129_llama (SEQ ID NO: 109) , IGHV3_134_llama (SEQ ID NO: 110) , IGHV3_137_llama (SEQ ID NO: 111) , IGHV3_139_llama (SEQ ID NO: 112) , IGHV3_14_llama (SEQ ID NO: 113) , IGHV3_140_llama (SEQ ID NO: 114) , IGHV3_141_llama (SEQ ID NO: 115) , IGHV3_147_llama (SEQ ID NO: 116) , IGHV3_15_llama (SEQ ID NO: 117) , IGHV3_17_llama (SEQ ID NO: 118) , IGHV3_22_llama (SEQ ID NO: 119) , IGHV3_25_llama (SEQ ID NO: 120) , IGHV3_27_llama (SEQ ID NO: 121) , IGHV3_29_llama (SEQ ID NO: 122) , IGHV3_3_llama (SEQ ID NO: 123) , IGHV3_31_llama (SEQ ID NO: 124) , IGHV3_32_llama (SEQ ID NO: 125) , IGHV3_33_llama (SEQ ID NO: 126) , IGHV3_34_llama (SEQ ID NO: 127) , IGHV3_35_llama (SEQ ID NO: 128) , IGHV3_36_llama (SEQ ID NO: 129) , IGHV3_37_llama (SEQ ID NO: 130) , IGHV3_38_llama (SEQ ID NO: 131) , IGHV3_39_llama (SEQ ID NO: 132) , IGHV3_40_llama (SEQ ID NO: 133) , IGHV3_43_llama (SEQ ID NO: 134) , IGHV3_44_llama (SEQ ID NO: 135) , IGHV3_48_llama (SEQ ID NO: 136) , IGHV3_5_llama (SEQ ID NO: 137) , IGHV3_50_llama (SEQ ID NO: 138) , IGHV3_52_llama (SEQ ID NO: 139) , IGHV3_55_llama (SEQ ID NO: 140) , IGHV3_59_llama (SEQ ID NO: 141) , IGHV3_6_llama (SEQ ID NO: 142) , IGHV3_60_llama (SEQ ID NO: 143) , IGHV3_61_llama (SEQ ID NO: 144) , IGHV3_62_llama (SEQ ID NO: 145) , IGHV3_64_llama (SEQ ID NO: 146) , IGHV3_66_llama (SEQ ID NO: 147) , IGHV3_68_llama (SEQ ID NO: 148) , IGHV3_7_llama (SEQ ID NO: 149) , IGHV3_70_llama (SEQ ID NO: 150) , IGHV3_71_llama (SEQ ID NO: 151) , IGHV3_73_llama (SEQ ID NO: 152) , IGHV3_74_llama (SEQ ID NO: 153) , IGHV3_76_llama (SEQ ID NO: 154) , IGHV3_77_llama (SEQ ID NO: 155) , IGHV3_80_llama (SEQ ID NO: 156) , IGHV3_81_llama (SEQ ID NO: 157) , IGHV3_82_llama (SEQ ID NO: 158) , IGHV3_83_llama (SEQ ID NO: 159) , IGHV3_86_llama (SEQ ID NO: 160) , IGHV3_89_llama (SEQ ID NO: 161) , IGHV3_90_llama (SEQ ID NO: 162) , IGHV3_91_llama (SEQ ID NO: 163) , IGHV3_93_llama (SEQ ID NO: 164) , IGHV3_94_llama (SEQ ID NO: 165) , IGHV3_95_llama (SEQ ID NO: 166) , IGHV3_96_llama (SEQ ID NO: 167) , IGHV3_99_llama (SEQ ID NO: 168) , IGHV3-1*01_lama (SEQ ID NO: 169) and IGHV3S6_lama (SEQ ID NO: 170) .
[0074] In another embodiment, the following low immunogenicity (rarity score >= 88%and 9-mer score <= 7) Dromedary germlines are identified: IGHV1S1*02_Dromedary (SEQ ID NO: 171) , IGHV1S18*01_Dromedary (SEQ ID NO: 172) , IGHV1S22*01_Dromedary (SEQ ID NO: 173) , IGHV1S25*01_Dromedary (SEQ ID NO: 174) , IGHV3S1*01_Dromedary (SEQ ID NO: 175) , IGHV3S1*02_Dromedary (SEQ ID NO: 176) , IGHV3S1*03_Dromedary (SEQ ID NO: 177) , IGHV3S1*04_Dromedary (SEQ ID NO: 178) , IGHV3S11*01_Dromedary (SEQ ID NO: 179) , IGHV3S12*01_Dromedary (SEQ ID NO: 180) , IGHV3S12*02_Dromedary (SEQ ID NO: 181) , IGHV3S18*01_Dromedary (SEQ ID NO: 182) , IGHV3S19*01_Dromedary (SEQ ID NO: 183) , IGHV3S2*01_Dromedary (SEQ ID NO: 184) , IGHV3S20*01_Dromedary (SEQ ID NO: 185) , IGHV3S22*01_Dromedary (SEQ ID NO: 186) , IGHV3S23*01_Dromedary (SEQ ID NO: 187) , IGHV3S24*01_Dromedary (SEQ ID NO: 188) , IGHV3S25*01_Dromedary (SEQ ID NO: 189) , IGHV3S26*01_Dromedary (SEQ ID NO: 190) , IGHV3S27*01_Dromedary (SEQ ID NO: 191) , IGHV3S28*01_Dromedary (SEQ ID NO: 192) , IGHV3S29*01_Dromedary (SEQ ID NO: 193) , IGHV3S3*01_Dromedary (SEQ ID NO: 194) , IGHV3S30*01_Dromedary (SEQ ID NO: 195) , IGHV3S31*01_Dromedary (SEQ ID NO: 196) , IGHV3S32*01_Dromedary (SEQ ID NO: 197) , IGHV3S33*01_Dromedary (SEQ ID NO: 198) , IGHV3S35*01_Dromedary (SEQ ID NO: 199) , IGHV3S36*01_Dromedary (SEQ ID NO: 200) , IGHV3S38*01_Dromedary (SEQ ID NO: 201) , IGHV3S4*01_Dromedary (SEQ ID NO: 202) , IGHV3S5*01_Dromedary (SEQ ID NO: 203) , IGHV3S6*01_Dromedary (SEQ ID NO: 204) , IGHV3S7*01_Dromedary (SEQ ID NO: 205) and IGHV3S9*01_Dromedary (SEQ ID NO: 206) .
[0075] In yet another embodiment, the following low immunogenicity (rarity score >= 88%and 9-mer score <= 7) Bactrian camel germlines are identified: IGHV3|cb_vh_63 (SEQ ID NO: 207) , IGHV3S40_Bactrian (SEQ ID NO: 208) , IGHV3S51_Bactrian (SEQ ID NO: 209) , IGHV3|cb_vh_79 (SEQ ID NO: 210) , IGHV3|cb_vh_45 (SEQ ID NO: 211) , IGHV3|cb_vh_29 (SEQ ID NO: 212) , IGHV3|cb_vh_34 (SEQ ID NO: 213) , IGHV3|cb_vh_52 (SEQ ID NO: 214) , IGHV3|cb_vh_100 (SEQ ID NO: 215) , IGHV3|cb_vh_57 (SEQ ID NO: 216) , IGHV3|cb_vh_97 (SEQ ID NO: 217) , IGHV3|cb_vh_12 (SEQ ID NO: 218) , IGHV3|cb_vh_37 (SEQ ID NO: 219) , IGHV3|cb_vh_38 (SEQ ID NO: 220) , IGHV3|cb_vh_27 (SEQ ID NO: 221) , IGHV3|cb_vh_39 (SEQ ID NO: 222) , IGHV3|cb_vh_19 (SEQ ID NO: 223) , IGHV3|cb_vh_109 (SEQ ID NO: 224) , IGHV3|cb_vh_6 (SEQ ID NO: 225) , IGHV3|cb_vh_10 (SEQ ID NO: 226) , IGHV3|cb_vh_112 (SEQ ID NO: 227) , IGHV3S29_Bactrian (SEQ ID NO: 228) , IGHV3|cb_vh_25 (SEQ ID NO: 229) , IGHV3S25_Bactrian (SEQ ID NO: 230) , IGHV3|cb_vh_14 (SEQ ID NO: 231) , IGHV3-2_Bactrian (SEQ ID NO: 232) , IGHV3|cb_vh_72 (SEQ ID NO: 233) , IGHV3|cb_vh_32 (SEQ ID NO: 234) , IGHV3|cb_vh_49 (SEQ ID NO: 235) , IGHV3|cb_vh_23 (SEQ ID NO: 236) , IGHV3|cb_vh_70 (SEQ ID NO: 237) , IGHV3|cb_vh_82 (SEQ ID NO: 238) , IGHV3|cb_vh_78 (SEQ ID NO: 239) , IGHV3|cb_vh_40 (SEQ ID NO: 240) , IGHV3|cb_vh_15 (SEQ ID NO: 241) , IGHV3|cb_vh_36 (SEQ ID NO: 242) , IGHV3S40*1_Bactrian (SEQ ID NO: 243) , IGHV3-1_Bactrian (SEQ ID NO: 244) , IGHV3S42_Bactrian (SEQ ID NO: 245) , IGHV3|cb_vh_20 (SEQ ID NO: 246) , IGHV3|cb_vh_30 (SEQ ID NO: 247) , IGHV3|cb_vh_89 (SEQ ID NO: 248) , IGHV3|cb_vh_81 (SEQ ID NO: 249) , IGHV3|cb_vh_48 (SEQ ID NO: 250) , IGHV3|cb_vh_103 (SEQ ID NO: 251) , IGHV3|cb_vh_56 (SEQ ID NO: 252) , IGHV3|cb_vh_35 (SEQ ID NO: 253) , IGHV3S6_Bactrian (SEQ ID NO: 254) , IGHV3|cb_vh_85 (SEQ ID NO: 255) , IGHV3|cb_vh_59 (SEQ ID NO: 256) , IGHV3|cb_vh_66 (SEQ ID NO: 257) , IGHV3|cb_vh_110 (SEQ ID NO: 258) , IGHV3|cb_vh_77 (SEQ ID NO: 259) , IGHV3|cb_vh_68 (SEQ ID NO: 260) , IGHV3|cb_vh_33 (SEQ ID NO: 261) , IGHV3|cb_vh_24 (SEQ ID NO: 262) , IGHV3|cb_vh_92 (SEQ ID NO: 263) , IGHV3|cb_vh_83 (SEQ ID NO: 264) , IGHV3|cb_vh_108 (SEQ ID NO: 265) , IGHV3|cb_vh_46 (SEQ ID NO: 266) , IGHV3|cb_vh_4 (SEQ ID NO: 267) , IGHV3|cb_vh_80 (SEQ ID NO: 268) , IGHV3|cb_vh_75 (SEQ ID NO: 269) , IGHV3|cb_vh_28 (SEQ ID NO: 270) , IGHV3|cb_vh_91 (SEQ ID NO: 271) , IGHV3|cb_vh_96 (SEQ ID NO: 272) , IGHV3|cb_vh_101 (SEQ ID NO: 273) , IGHV3|cb_vh_18 (SEQ ID NO: 274) , IGHV3|cb_vh_22 (SEQ ID NO: 275) , IGHV3|cb_vh_51 (SEQ ID NO: 276) , IGHV3|cb_vh_67 (SEQ ID NO: 277) , IGHV3|cb_vh_111 (SEQ ID NO: 278) , IGHV3|cb_vh_88 (SEQ ID NO: 279) , IGHV3|cb_vh_98 (SEQ ID NO: 280) , IGHV3|cb_vh_90 (SEQ ID NO: 281) , IGHV3|cb_vh_104 (SEQ ID NO: 282) , IGHV3|cb_vh_105 (SEQ ID NO: 283) , IGHV3|cb_vh_115 (SEQ ID NO: 284) , IGHV3|cb_vh_8 (SEQ ID NO: 285) , IGHV3|cb_vh_3 (SEQ ID NO: 286) , IGHV3|cb_vh_50 (SEQ ID NO: 287) , IGHV3|cb_vh_61 (SEQ ID NO: 288) , IGHV3|cb_vh_69 (SEQ ID NO: 289) , IGHV3|cb_vh_86 (SEQ ID NO: 290) and IGHV3|cb_vh_11 (SEQ ID NO: 291) .
[0076] The sequences of the remaining germline for Llama, Dromedary and Bactrian camel that are not deemed to have low immunogenicity are provided as SEQ ID Nos: 352 through 597 in the Sequence Listing entitled “24A927 SEQ ST26 EN. xml” , which is concurrently filed with and incorporated into this application.
[0077] With such criteria, no mouse germlines will be in low immunogenicity category. IGHV1, IGHV4 and many VHH germlines are in the high immunogenicity category. Alpaca VHHs mapped to low immunogenicity germlines and to high immunogenicity germlines showed clear difference in predicted immunogenicity (Fig. 11) . Interestingly, the 9-mer score distribution of VHHs from low immunogenicity germlines showed two peaks (Fig. 11C) . To further understand this phenomenon, we analyzed the correlation between VHH 9-mer score and its mismatches. Based on NGS sequences of VHHs from Alpaca, Llama and Bactrian, we saw significant correlation between 9-mer scores and mismatches (Fig. 12) . Such results suggest that VHHs from low immunogenicity germlines with high mismatches have higher predicted immunogenicity possibility. Indeed, using cutoff of mismatches 11 (Fig. 13A) and mismatches 13 (Fig. 13B) , we see single peak distribution of 9-mer scores for VHH from low immunogenicity germlines, although the distribution of 9-mer scores for mismatch cutoff of 13 is wider than the one with mismatch cutoff of 11.
[0078] Thus, in one embodiment, the VHH sequences generated in (c) are excluded from the libraries if their mismatch score is higher than 12.
[0079] In one embodiment, the camelid repertoires for constructing the libraries comprise camelid germlines regardless of the level of their predicted immunogenicity. In another embodiment, the repertoires are enriched for camelid germlines having predicted low immunogenicity.
[0080] In one embodiment, an experimental method was developed to build libraries with low immunogenicity VHHs enriched. As Table 3 showed that all IGHV1 and IGHV4 germlines and many VHH germlines have high predicted immunogenicity, a method employing IGHV3 specific primer for only amplifying VHHs from IGHV3 germlines and VH capture oligos for capturing VHHs from low immunogenicity VH germlines were developed to enrich low immunogenicity VHHs during library construction. Fig. 14 showed a representative diagram of such method. Germline multiple sequence alignment was used for the design of IGHV3 FR1 specific primers (Table 5) . For designing capture oligos, we performed multiple sequence alignment of commonly used IGHV3 germlines (Table 4) .
[0081] Table 4. Top 15 VH3 germlines used in Alpaca VHHs and their predicted immunogenicity classifications
[0082] Multiple alignment results (Fig. 15) showed that the FR2 region (boxed area) is the best region to design such capture oligos. Table 5 listed examples of primer and oligo sequences from the four species for such method.
[0083] Table 5. Sequences for capture oligos and IGHV3 specific primers
[0084] Based on the Alpaca VHH sequences, we noticed that VHHs with the first residue of FR4 (IMGT position 118) as arginine (R) is more likely to have low predicted immunogenicity than those with tryptophan (W) at that position (21.5%VHHs having 9-mer score < 40 vs 9.5%, Fig. 16) . In one embodiment, a second experimental method was developed to build libraries with low immunogenicity VHHs enriched by using primers specific for IGHV3 and FR4 with an arginine (R) at IMGT position 118. A diagram of such method is shown in Fig. 17 and examples of the sequences from the four species for such method are listed in Table 6. Indeed, libraries built with such method (R libraries) showed statistically significantly lower average 9-mer score than those including all VHHs (Fig. 18) .
[0085] Table 6. Sequences for FR4 and IGHV3 specific primers used in Fig. 15
[0086] In one embodiment, a third experiment method was developed to build libraries with low immunogenicity VHHs enriched by combining the above two experimental methods. A diagram of such method is shown in Fig. 19.
[0087] In one embodiment, the sequences of the nanobodies are aligned with germlines and best aligned germlines are determined.
[0088] In one embodiment, the sequences of the nanobodies are aligned with germlines and mismatch scores are determined based on best aligned germlines.
[0089] In one embodiment, the sequences of the nanobodies are aligned with VHH including gene-conversion like event with humanlike pseudogenes usage on best aligned germlines.
[0090] In some embodiments, the method is directed to generating a low immunogenicity nanobody (VHH) specific to an antigen using a generative machine learning model, comprising:
[0091] a) selecting VHH sequences having predicted low immunogenicity;
[0092] b) training the generative machine learning model with the VHH sequences of (a) ;
[0093] c) using the trained generative machine learning model of (b) to generate synthetic VHH sequences;
[0094] d) constructing libraries from the synthetic VHH sequences generated in (c) ; and
[0095] e) generating a nanobody specific to the antigen from the libraries in (d) .
[0096] In one embodiment, a set of low immunogenicity VHH sequences is identified from VHH NGS sequences using rarity scores and / or 9-mer scores. A generative machine learning model such as GPT as shown in Fig. 20B can be built and trained using the identified low immunogenicity VHH set. The trained generative model can then be used to generate new VHH sequences with predicted low immunogenicity (Fig. 20A) . Generated VHH sequences can be used to build VHH libraries like phage libraries by synthesizing individual VHH sequences. Libraries can be screened to identify VHHs against specific targets.
[0097] In another embodiment, all VHH sequences are used to train a generative model such as GPT. The trained model is then used as pre-training for reinforced learning agent (RL agent) . The neural network architecture of the pre-training model and the RL agent model are the same. In each step, the RL agent model is used to generate VHH sequences. The generated sequences are scored using rarity score and / or 9-mer score to generate an immunogenicity score. This score is combined with sequence generating likelihood score. The combined score is used as feedback to the agent. With multiple rounds of such training, RL agent model can generate low immunogenicity sequences as shown in Fig. 21. Such sequences can be used to build libraries which can be screened to identify VHHs against specific targets.
[0098] In some embodiments, the method is directed to generating a low immunogenicity nanobody (VHH) specific to an antigen using de novo design of VHH sequences, comprising:
[0099] a) selecting VHH templates having predicted low immunogenicity;
[0100] b) using a first machine learning model to generate protein backbone structures that specifically bind to the antigen based on the structures of the VHH templates in (a) and the structure of the antigen;
[0101] c) using a second machine learning model to generate synthetic VHH sequences based on the protein backbone structures in (b) ; and
[0102] d) synthesize a nanobody specific to the antigen having the synthetic VHH sequence in (c) .
[0103] In one embodiment, using VHH template sequences predicted to be low immunogenicity and target structure, backbone structures can be generated using tools like RFdiffusion (Joseph et al, 2022) . From generated backbone structures, sequences can be generated using tools such as proteinMPNN (Dauparas et al., 2022) . Generated sequences can then be further filtered using various criteria, such as PI, immunogenicity score, naturalness score, solubility score, etc. Filtered VHH sequences can be modeled and docking between VHH models and target structure can be performed to further validate designed sequences. A diagram of such method is shown in Fig. 22.
[0104] Examples of VHH template sequences predicted to be low immunogenicity include:
[0105] a) FR1: EVQLVESGGGLVQPGGSLRLSCAAS (SEQ ID NO: 304)
[0106] FR2: MSWFRQAPGKEREGVSA (SEQ ID NO: 305)
[0107] FR3: YYADSVKGRFTISRDNSKNTLYLQMNSLRAEDTAVYYC (SEQ ID NO: 306)
[0108] FR4: WGQGTLVTVSS (SEQ ID NO: 307) ;
[0109] b) FR1: EVQLVESGGGLVQPGGSLRLSCAAS (SEQ ID NO: 308)
[0110] FR2: MSWYRQAPGKEREGVSA (SEQ ID NO: 309)
[0111] FR3: YYADSVKGRFTISRDNSKNTLYLQMNSLRAEDTAVYYC (SEQ ID NO: 310)
[0112] FR4: WGQGTLVTVSS (SEQ ID NO: 311) ;
[0113] c) FR1: EVQLVESGGGLVQPGGSLRLSCAAS (SEQ ID NO: 312)
[0114] FR2: MSWVRQAPGKGLEWVSA (SEQ ID NO: 313)
[0115] FR3: YYADSVKGRFTISRDNSKNTLYLQMNSLRAEDTAVYYC (SEQ ID NO: 314)
[0116] FR4: RGQGTQVTVSS (SEQ ID NO: 315) ;
[0117] d) FR1: QVQLVESGGGLVQPGGSLRLSCAAS (SEQ ID NO: 316)
[0118] FR2: MSWFRQAPGKEREWVS (SEQ ID NO: 317)
[0119] FR3: YADSVKGRFTISRDNSKNTLYLQMNSLRAEDTAVYYC (SEQ ID NO: 318)
[0120] FR4: WGQGTLVTVSS (SEQ ID NO: 319) ;
[0121] e) FR1: QVQLVESGGGLVQPGGSLRLSCAAS (SEQ ID NO: 320)
[0122] FR2: MSWYRQAPGKEREWVS (SEQ ID NO: 321)
[0123] FR3: YADSVKGRFTISRDNSKNTLYLQMNSLRAEDTAVYYC (SEQ ID NO: 322)
[0124] FR4: WGQGTLVTVSS (SEQ ID NO: 323) ;
[0125] f) FR1: QVQLVESGGGLVKPGGSLRLSCAAS (SEQ ID NO: 324)
[0126] FR2: MSWVRQAPGKGLEWVS (SEQ ID NO: 325)
[0127] FR3: YADSVKGRFTISRDNSKNTLYLQMNSLRAEDTAVYYC (SEQ ID NO: 326)
[0128] FR4: RGQGTLVTVSS (SEQ ID NO: 327) ;
[0129] g) FR1: QVQLVESGGGLVQPGGSLRLSCAAS (SEQ ID NO: 328)
[0130] FR2: MSWFRQAPGKEREWVS (SEQ ID NO: 329)
[0131] FR3: YADSVKGRFTISRDNSKNTLYLQMNSLRAEDTAVYYC (SEQ ID NO: 330)
[0132] FR4: WGQGTQVTVSS (SEQ ID NO: 331) ;
[0133] h) FR1: QVQLVESGGGLVQPGGSLRLSCAAS (SEQ ID NO: 332)
[0134] FR2: MSWYRQAPGKEREWVS (SEQ ID NO: 333)
[0135] FR3: YADSVKGRFTISRDNSKNTLYLQMNSLRAEDTAVYYC (SEQ ID NO: 334)
[0136] FR4: WGQGTQVTVSS (SEQ ID NO: 335) ; and
[0137] i) FR1: QVQLVESGGGLVKPGGSLRLSCAAS (SEQ ID NO: 336)
[0138] FR2: MSWVRQAPGKGLEWVS (SEQ ID NO: 337)
[0139] FR3: YADSVKGRFTISRDNSKNTLYLQMNSLRAEDTAVYYC (SEQ ID NO: 338)
[0140] FR4: RGQGTQVTVSS (SEQ ID NO: 339) .
[0141] Preferred embodiments are described in the following examples. Other embodiments within the scope of the claims herein will be apparent to one skilled in the art from consideration of the specification or practice of the invention as disclosed herein. It is intended that the specification, together with the examples, be considered exemplary only, with the scope and spirit of the invention being indicated by the claims, which follow the examples.
[0142] Example 1. Design CD3 VHHs with predicted low immunogenicity
[0143] To test the method of using AI tools to design VHHs with predicted low immunogenicity, we chose to design VHHs binding to CD3, which is a target for T cell engager immunotherapy. Several T cell engagers have been approved for several cancers. The design is based on OKT3 / CD3 complex (PDB ID: 1SY6) . PISA (Krissinel et al., 2007) was used to identify contact residues of OKT3 heavy chain on CD3ε: G32, S33, E34, L36, N43, G45, G46, D47, E48, S55, E57, R79, G80, S81, K82, P83, D85. These residues are considered as the epitope of OKT3 and used for designing structures by RFdiffusion. Two VHH templates were generated using one OKT3 heavy CDR sequence and were used as shown in Table 7.
[0144] Table 7. Template sequences used by RFdiffusion for structure design
[0145] Residues that are marked bold are from the CDRs of the OKT3 heavy chain. The remaining residues are from template sequences.
[0146] 1000 protein backbone structures were generated using RFdiffusion for each VHH template above and 200 synthetic VHH sequences for each protein backbone structure were generated using proteinMPNN to convert structures to sequences. We used AbNativ score (Ramon et al., 2024) with cutoff of 0.8 to select sequences for further analysis. After this filtering step, 3405 and 1507 sequences were obtained from non-classical-2Q-OKT3 and classical-3Q-OKT3 VHH template, respectively. Fig. 23 showed histograms of two immunogenicity scores for these synthetic VHH sequences. Sequences from the classical-3Q-OKT3 VHH template have an average of 93%for rarity score and an average of 27 for 9-mer score, while sequences from the non-classical-2Q-OKT3 VHH template sequences have an average of 95%for rarity score and an average of 18 for 9-mer score. Such results suggest that these synthetic VHH sequences have very low predicted immunogenicity. To select VHH sequences for testing, these sequences were further screened using one or more of the following criteria:
[0147] 1) Sequences with sequence liabilities (for example, N-linked glycosylation and free cysteine) were excluded;
[0148] 2) Sequences with PI < 7.0 or PI > 8.7 were excluded;
[0149] 3) Sequences that bind to at least four residues of the target epitope based on HADDOCK (Ambrosetti et al., 2020) docking results using Alphafold2 (Jumper et al., 2021) predicted structures were included.
[0150] 10 sequences as shown in Table 8 were selected for DNA synthesis, protein expression and functional testing.
[0151] Table 8.10 Synthetic VHH sequences selected for further testing
[0152] These synthetic VHHs were expressed in 1 ml CHO cells. 4 VHHs showed high expression (Table 9) .
[0153] Table 9. Expression of 2 VHH template sequences and 10 synthetic VHH sequences
[0154] Jurkat cell binding experiment using flow cytometry showed that several VHHs have weak binding activities (Fig. 24) .
[0155] Example 2. Generate VHHs with predicted low immunogenicity using generative AI method
[0156] A generative model was built using GPT2 architecture (Radford et al., 2019) using the following configuration parameters:
[0157] {
[0158] ″activation_function″ : ″gelu_new″ ,
[0159] ″architectures″ : [
[0160] ″GPT2LMHeadModel″
[0161] ] ,
[0162] ″attn_pdrop″ : 0.1,
[0163] ″bos_token_id″ : 20,
[0164] ″embd_pdrop″ : 0.1,
[0165] ″eos_token_id″ : 21,
[0166] ″initializer_range″ : 0.02,
[0167] ″layer_norm_epsilon″ : 1e-05,
[0168] ″model_type″ : ″gpt2″ ,
[0169] ″n_embd″ : 768,
[0170] ″n_head″ : 8,
[0171] ″n_inner″ : null,
[0172] ″n_layer″ : 8,
[0173] ″n_positions″ : 146,
[0174] ″output_hidden_states″ : true,
[0175] ″reorder_and_upcast_attn″ : false,
[0176] ″resid_pdrop″ : 0.1,
[0177] ″scale_attn_by_inverse_layer_idx″ : false,
[0178] ″scale_attn_weights″ : true,
[0179] ″summary_activation″ : null,
[0180] ″summary_first_dropout″ : 0.1,
[0181] ″summary_proj_to_labels″ : true,
[0182] ″summary_type″ : ″cls_index″ ,
[0183] ″summary_use_proj″ : true,
[0184] ″torch_dtype″ : ″float32″ ,
[0185] ″transformers_version″ : ″4.36.2″ ,
[0186] ″use_cache″ : true,
[0187] ″vocab_size″ : 23
[0188] }
[0189] The models were trained using either whole VHH dataset that comprises 27 million sequences (the “27M model” ) or the VHH dataset with predicted low immunogenicity that comprises 1.8 million sequences with a rarity score of no less than 86% (the “lowImmuno model” ) . The models were trained up to 30 epochs and the model checkpoints with lowest validation loss were used to generate new VHH sequences. 10,000 VHH sequences were generated from each model. Sequences generated from the lowImmuno model have significantly lower predicted immunogenicity than sequences generated from the 27M model (P < 0.0001 for both rarity score and 9-mer score; see Fig. 25) .
[0190] Example 3. VHH from high immunogenicity germline has a higher ADA rate than those from low immunogenicity germlines
[0191] As part of the IIT (investigator initialized trial) for CAR-T therapy development, several VHHs in the format of CAR-T have been tested in terminal cancer patients and their anti-drug antibody (ADA) rates were measured using ELISA and equivalent methods in those patients.
[0192] Clone 1444, a VHH binding to mesothelin, is mapped to IGHV3S57 which is classified as having high immunogenicity (see Table 3) . The CAR-T using this VHH has a >37%ADA rate (54 patients) . Clone 555, a VHH binding to MUC1 and clone 024, a VHH binding to mesothelin, are mapped to IGHV3S6 and IGHV3S25 respectively, both of which are classified as having low immunogenicity category (see Table 3) . The CAR-T using these two VHHs (bispecific design) has a 20%ADA rate (5 patients) .
[0193] EMBODIMENTS:
[0194] Embodiment 1. A method of generating a low immunogenicity nanobody (VHH) specific to an antigen, comprising:
[0195] a) selecting camelid germlines having predicted low immunogenicity;
[0196] b) constructing libraries from camelid repertoires;
[0197] c) selecting VHH sequences from the libraries in (b) that are derived from the camelid germlines in (a) ; and
[0198] d) generating a nanobody specific to the antigen from the VHH sequences of (c) .
[0199] Embodiment 2. The method of Embodiment 1, wherein the camelid germlines having predicted low immunogenicity have a rarity score no lower than 88%and / or a 9-mer score no higher than 7.
[0200] Embodiment 3. The method of Embodiment 1 or 2, wherein the camelid is alpaca and wherein the germlines having predicted low immunogenicity comprise IGHV3S1*01, IGHV3S35*01, IGHV3S40*01, IGHV3-1*01, IGHV3S2*01, IGHV3S10*01, IGHV3S42*01, IGHV3S41*01, IGHV3S39*01, IGHV3S6*01, IGHV3S44*01, IGHV3S7*01, IGHV3S25*01, IGHV3S15*01, IGHV3S16*01, IGHV3S37*01, IGHV3S36*01, IGHV3S11*01, IGHV3S31*01, IGHV3S12*01, IGHV3S30*01, IGHV3-2*01, IGHV3S8*01, IGHV3S26*01, IGHV3S33*01, IGHV3S32*01, IGHV3S5*01, IGHV3S38*01, IGHV3S9*01 or IGHV3S67*01.
[0201] Embodiment 4. The method of Embodiment 1 or 2, wherein the camelid is llama and wherein the germlines having predicted low immunogenicity comprise IGHV3_1_llama, IGHV3_105_llama, IGHV3_106_llama, IGHV3_107_llama, IGHV3_108_llama, IGHV3_11_llama, IGHV3_114_llama, IGHV3_116_llama, IGHV3_12_llama, IGHV3_120_llama, IGHV3_121_llama, IGHV3_122_llama, IGHV3_123_llama, IGHV3_124_llama, IGHV3_126_llama, IGHV3_129_llama, IGHV3_134_llama, IGHV3_137_llama, IGHV3_139_llama, IGHV3_14_llama, IGHV3_140_llama, IGHV3_141_llama, IGHV3_147_llama, IGHV3_15_llama, IGHV3_17_llama, IGHV3_22_llama, IGHV3_25_llama, IGHV3_27_llama, IGHV3_29_llama, IGHV3_3_llama, IGHV3_31_llama, IGHV3_32_llama, IGHV3_33_llama, IGHV3_34_llama, IGHV3_35_llama, IGHV3_36_llama, IGHV3_37_llama, IGHV3_38_llama, IGHV3_39_llama, IGHV3_40_llama, IGHV3_43_llama, IGHV3_44_llama, IGHV3_48_llama, IGHV3_5_llama, IGHV3v50_llama, IGHV3_52_llama, IGHV3_55_llama, IGHV3_59_llama, IGHV3_6_llama, IGHV3_60_llama, IGHV3_61_llama, IGHV3_62_llama, IGHV3_64_llama, IGHV3_66_llama, IGHV3_68_llama, IGHV3_7_llama, IGHV3_70_llama, IGHV3_71_llama, IGHV3_73_llama, IGHV3_74_llama, IGHV3_76_llama, IGHV3_77_llama, IGHV3_80_llama, IGHV3_81_llama, IGHV3_82_llama, IGHV3_83_llama, IGHV3_86_llama, IGHV3_89_llama, IGHV3_90_llama, IGHV3_91_llama, IGHV3_93_llama, IGHV3_94_llama, IGHV3_95_llama, IGHV3_96_llama, IGHV3_99_llama, IGHV3-1*01_lama or IGHV3S6_lama.
[0202] Embodiment 5. The method of Embodiment 1 or 2, wherein the camelid is dromedary and wherein the germlines having predicted low immunogenicity comprise IGHV1S1*02_Dromedary, IGHV1S18*01_Dromedary, IGHV1S22*01_Dromedary, IGHV1S25*01_Dromedary, IGHV3S1*01_Dromedary, IGHV3S1*02_Dromedary, IGHV3S1*03_Dromedary, IGHV3S1*04_Dromedary, IGHV3S11*01_Dromedary, IGHV3S12*01_Dromedary, IGHV3S12*02_Dromedary, IGHV3S18*01_Dromedary, IGHV3S19*01_Dromedary, IGHV3S2*01_Dromedary, IGHV3S20*01_Dromedary, IGHV3S22*01_Dromedary, IGHV3S23*01_Dromedary, IGHV3S24*01_Dromedary, IGHV3S25*01_Dromedary, IGHV3S26*01_Dromedary, IGHV3S27*01_Dromedary, IGHV3S28*01_Dromedary, IGHV3S29*01_Dromedary, IGHV3S3*01_Dromedary, IGHV3S30*01_Dromedary, IGHV3S31*01_Dromedary, IGHV3S32*01_Dromedary, IGHV3S33*01_Dromedary, IGHV3S35*01_Dromedary, IGHV3S36*01_Dromedary, IGHV3S38*01_Dromedary, IGHV3S4*01_Dromedary, IGHV3S5*01_Dromedary, IGHV3S6*01_Dromedary, IGHV3S7*01_Dromedary or IGHV3S9*01_Dromedary.
[0203] Embodiment 6. The method of Embodiment 1 or 2, wherein the camelid is Bactrian camel and wherein the germlines having predicted low immunogenicity comprise IGHV3|cb_vh_63, IGHV3S40_Bactrian, IGHV3S51_Bactrian, IGHV3|cb_vh_79, IGHV3|cb_vh_45, IGHV3|cb_vh_29, IGHV3|cb_vh_34, IGHV3|cb_vh_52, IGHV3|cb_vh_100, IGHV3|cb_vh_57, IGHV3|cb_vh_97, IGHV3|cb_vh_12, IGHV3|cb_vh_37, IGHV3|cb_vh_38, IGHV3|cb_vh_27, IGHV3|cb_vh_39, IGHV3|cb_vh_19, IGHV3|cb_vh_109, IGHV3|cb_vh_6, IGHV3|cb_vh_10, IGHV3|cb_vh_112, IGHV3S29_Bactrian, IGHV3|cb_vh_25, IGHV3S25_Bactrian, IGHV3|cb_vh_14, IGHV3-2_Bactrian, IGHV3|cb_vh_72, IGHV3|cb_vh_32, IGHV3|cb_vh_49, IGHV3|cb_vh_23, IGHV3|cb_vh_70, IGHV3|cb_vh_82, IGHV3|cb_vh_78, IGHV3|cb_vh_40, IGHV3|cb_vh_15, IGHV3|cb_vh_36, IGHV3S40*1_Bactrian, IGHV3-1_Bactrian, IGHV3S42_Bactrian, IGHV3|cb_vh_20, IGHV3|cb_vh_30, IGHV3|cb_vh_89, IGHV3|cb_vh_81, IGHV3|cb_vh_48, IGHV3|cb_vh_103, IGHV3|cb_vh_56, IGHV3|cb_vh_35, IGHV3S6_Bactrian, IGHV3|cb_vh_85, IGHV3|cb_vh_59, IGHV3|cb_vh_66, IGHV3|cb_vh_110, IGHV3|cb_vh_77, IGHV3|cb_vh_68, IGHV3|cb_vh_33, IGHV3|cb_vh_24, IGHV3|cb_vh_92, IGHV3|cb_vh_83, IGHV3|cb_vh_108, IGHV3|cb_vh_46, IGHV3|cb_vh_4, IGHV3|cb_vh_80, IGHV3|cb_vh_75, IGHV3|cb_vh_28, IGHV3|cb_vh_91, IGHV3|cb_vh_96, IGHV3|cb_vh_101, IGHV3|cb_vh_18, IGHV3|cb_vh_22, IGHV3|cb_vh_51, IGHV3|cb_vh_67, IGHV3|cb_vh_111, IGHV3|cb_vh_88, IGHV3|cb_vh_98, IGHV3|cb_vh_90, IGHV3|cb_vh_104, IGHV3|cb_vh_105, IGHV3|cb_vh_115, IGHV3|cb_vh_8, IGHV3|cb_vh_3, IGHV3|cb_vh_50, IGHV3|cb_vh_61, IGHV3|cb_vh_69, IGHV3|cb_vh_86 or IGHV3|cb_vh_11.
[0204] Embodiment 7. The method of any one of Embodiments 1 to 6, wherein the camelid repertoires in (b) are enriched for camelid germlines having predicted low immunogenicity prior to constructing the libraries.
[0205] Embodiment 8. The method of any one of Embodiments 1 to 7, wherein the libraries are constructed using primers specific for camelid germlines having predicted low immunogenicity.
[0206] Embodiment 9. The method of Embodiment 8, wherein the primers comprise the sequences of
[0207] SAGKTGCAGCTSGTGGAGTCT (Alpaca) , SAGGTGCAGCTSGTGGAGTCT (Llama) ,
[0208] SAGGTGCAGCTGGTGGAGTCT (Dromedary) , or SAGGTGCAGCTGGTGGAGTCT (Bactrian) .
[0209] Embodiment 10. The method of any one of Embodiments 1 to 8, wherein the libraries are constructed using primers specific for nanobodies having an arginine at the first residue of FR4 (IMGT position 118) .
[0210] Embodiment 11. The method of Embodiment 10, wherein primers comprise the sequences of
[0211] TGAGGAGACGGTGACCTGGGTCCCCTGGCCCCK (Alpaca) ,
[0212] TGAGGAGACGGTGACCTGGGTCCCCTGGCCCCK (Llama) ,
[0213] TGAGGAGACGGTGACCTGGGTCCCCTGGCCCCK (Dromedary) , or
[0214] TGAGGAGACGGTGACCTGGGTCCCCTGGCCCCK (Bactrian) .
[0215] Embodiment 12. The method of any one of Embodiments 1 to 8, wherein libraries are constructed using capture oligos targeting the FR2 region of camelid germlines having predicted low immunogenicity.
[0216] Embodiment 13. The method of Embodiment 12, wherein the capture oligos comprise the sequences of
[0217] AGCTGGGTCCGCCAGGCTCCAGGAAAGGGGCTC (Alpaca) ,
[0218] AGCTGGGTCCGCCAGGCTCCAGGAAAGGGGCTC (Llama) ,
[0219] CTGGGTCCGCCAGGCTCCAGGGAAGGGGCT (Dromedary) , or
[0220] AGCTGGGTCCGCCAGGCTCCAGGGAAGGGGCTC (Bactrian) .
[0221] Embodiment 14. The method of any one of Embodiments 1 to 8, wherein the libraries are constructed using primers specific for nanobodies having an arginine at the first residue of FR4 (IMGT position 118) and capture oligos targeting the FR2 region of camelid germlines having predicted low immunogenicity.
[0222] Embodiment 15. The method of any one of Embodiments 1 to 14, wherein the VHH sequences are generated from the libraries using next generation technology (NGS) .
[0223] Embodiment 16. The method of any one of Embodiments 1 to 15, further comprising determining the mismatch score of the VHH sequences in (c) , wherein the VHH sequence having a mismatch score greater than 12 is excluded.
[0224] Embodiment 17. The method of Embodiment 1, wherein the repertoires are generated from camelids immunized with a peptide, a protein, an mRNA, a DNA and / or a cell.
[0225] Embodiment 18. The method of any one of Embodiments 1 to 17, wherein the nanobody specific to the antigen is identified by phage panning or screening, B cell panning or screening, and / or NGS methods.
[0226] Embodiment 19. A method of generating a low immunogenicity nanobody (VHH) specific to an antigen using a generative machine learning model, comprising:
[0227] a) selecting VHH sequences having predicted low immunogenicity;
[0228] b) training the generative machine learning model with the VHH sequences of (a) ;
[0229] c) using the trained generative machine learning model of (b) to generate synthetic VHH sequences;
[0230] d) constructed libraries from the synthetic VHH sequences generated in (c) ; and
[0231] e) generating a nanobody specific to the antigen from the libraries in (d) .
[0232] Embodiment 20. The method of Embodiment 19, wherein the VHH sequences having a predicted low immunogenicity have a rarity score higher than 85%and / or a 9-mer score lower than 40.
[0233] Embodiment 21. The method of Embodiment 19 or 20, where generative machine learning model comprises VAE, GAN or GPT.
[0234] Embodiment 22. The method of any one of Embodiments 19 to 21, where generative machine learning model is pretrained with protein or antibody sequences.
[0235] Embodiment 23. The method of any one of Embodiments 19 to 22, where the libraries are constructed using DNA synthesized based on the synthetic VHH sequences in (c) .
[0236] Embodiment 24. The method of any one of Embodiments 19 to 23, wherein the nanobody specific to the antigen is identified by phage panning, and / or in silico methods.
[0237] Embodiment 25. The method of any one of Embodiments 19 to 24, wherein the VHH sequences having a predicted low immunogenicity have a rarity score higher than 85%and / or a 9-mer score lower than 40.
[0238] Embodiment 26. A method of generating a low immunogenicity nanobody (VHH) specific to an antigen using one or more machine learning models, comprising:
[0239] a) selecting VHH templates having predicted low immunogenicity;
[0240] b) using a first machine learning model to generate protein backbone structures that specifically bind to the antigen based on the structures of the VHH templates in (a) and the structure of the antigen;
[0241] c) using a second machine learning model to generate synthetic VHH sequences based on the protein backbone structures in (b) ; and
[0242] d) synthesize a nanobody specific to the antigen having the synthetic VHH sequence in (c) .
[0243] Embodiment 27. The method of Embodiment 26, wherein the VHH templates having a predicted low immunogenicity have a rarity score higher than 85%and / or a 9-mer score lower than 40.
[0244] Embodiment 28. The method of Embodiment 26 or 27, where the structure of the antigen is a 3D structure determined experimentally or a modelled structure determined by protein modeling tools.
[0245] Embodiment 29. The method of any one of Embodiments 26 to 28, where the VHH templates comprises the sequences of:
[0246] a) FR1: EVQLVESGGGLVQPGGSLRLSCAAS
[0247] FR2: MSWFRQAPGKEREGVSA
[0248] FR3: YYADSVKGRFTISRDNSKNTLYLQMNSLRAEDTAVYYC
[0249] FR4: WGQGTLVTVSS;
[0250] b) FR1: EVQLVESGGGLVQPGGSLRLSCAAS
[0251] FR2: MSWYRQAPGKEREGVSA
[0252] FR3: YYADSVKGRFTISRDNSKNTLYLQMNSLRAEDTAVYYC
[0253] FR4: WGQGTLVTVSS;
[0254] c) FR1: EVQLVESGGGLVQPGGSLRLSCAAS
[0255] FR2: MSWVRQAPGKGLEWVSA
[0256] FR3: YYADSVKGRFTISRDNSKNTLYLQMNSLRAEDTAVYYC
[0257] FR4: RGQGTQVTVSS;
[0258] d) FR1: QVQLVESGGGLVQPGGSLRLSCAAS
[0259] FR2: MSWFRQAPGKEREWVS
[0260] FR3: YADSVKGRFTISRDNSKNTLYLQMNSLRAEDTAVYYC
[0261] FR4: WGQGTLVTVSS;
[0262] e) FR1: QVQLVESGGGLVQPGGSLRLSCAAS
[0263] FR2: MSWYRQAPGKEREWVS
[0264] FR3: YADSVKGRFTISRDNSKNTLYLQMNSLRAEDTAVYYC
[0265] FR4: WGQGTLVTVSS;
[0266] f) FR1: QVQLVESGGGLVKPGGSLRLSCAAS
[0267] FR2: MSWVRQAPGKGLEWVS
[0268] FR3: YADSVKGRFTISRDNSKNTLYLQMNSLRAEDTAVYYC
[0269] FR4: RGQGTLVTVSS;
[0270] g) FR1: QVQLVESGGGLVQPGGSLRLSCAAS
[0271] FR2: MSWFRQAPGKEREWVS
[0272] FR3: YADSVKGRFTISRDNSKNTLYLQMNSLRAEDTAVYYC
[0273] FR4: WGQGTQVTVSS;
[0274] h) FR1: QVQLVESGGGLVQPGGSLRLSCAAS
[0275] FR2: MSWYRQAPGKEREWVS
[0276] FR3: YADSVKGRFTISRDNSKNTLYLQMNSLRAEDTAVYYC
[0277] FR4: WGQGTQVTVSS; or
[0278] i) FR1: QVQLVESGGGLVKPGGSLRLSCAAS
[0279] FR2: MSWVRQAPGKGLEWVS
[0280] FR3: YADSVKGRFTISRDNSKNTLYLQMNSLRAEDTAVYYC
[0281] FR4: RGQGTQVTVSS.
[0282] Embodiment 30. The method of any one of Embodiments 26 to 29, wherein the first machine learning model is RFdiffusion.
[0283] Embodiment 31. The method of any one of Embodiments 26 to 30, wherein the second machine learning model is proteinMPNN.
[0284] Embodiment 32. The method of any one of Embodiments 26 to 31, wherein the synthetic VHH sequences in (c) are further screened based on their physical and / or biochemical properties.
[0285] Embodiment 33. The method of any one of Embodiments 1-32, wherein the nanobodies are expressed by prokaryotic or eukaryotic cells.
[0286] In view of the above, it will be seen that several objectives of the invention are achieved and other advantages attained.
[0287] As various changes could be made in the above methods and compositions without departing from the scope of the invention, it is intended that all matter contained in the above description and shown in the accompanying drawings shall be interpreted as illustrative and not in a limiting sense.
[0288] All references cited in this specification, including but not limited to patent publications and non-patent literature, and references cited therein, are hereby incorporated by reference. The discussion of the references herein is intended merely to summarize the assertions made by the authors and no admission is made that any reference constitutes prior art. Applicants reserve the right to challenge the accuracy and pertinence of the cited references.
[0289] As used herein, in particular embodiments, the terms “about” or “approximately” when preceding a numerical value indicates the value plus or minus a range of 10%. Where a range of values is provided, it is understood that each intervening value, to the tenth of the unit of the lower limit unless the context clearly dictates otherwise, between the upper and lower limit of that range and any other stated or intervening value in that stated range is encompassed within the disclosure. That the upper and lower limits of these smaller ranges can independently be included in the smaller ranges is also encompassed within the disclosure, subject to any specifically excluded limit in the stated range. Where the stated range includes one or both of the limits, ranges excluding either or both of those included limits are also included in the disclosure.
[0290] The indefinite articles “a” and “an, ” as used herein in the specification and in the embodiments, unless clearly indicated to the contrary, should be understood to mean “at least one. ”
[0291] The phrase “and / or, ” as used herein in the specification and in the embodiments, should be understood to mean “either or both” of the elements so conjoined, i.e., elements that are conjunctively present in some cases and disjunctively present in other cases. Multiple elements listed with “and / or” should be construed in the same fashion, i.e., “one or more” of the elements so conjoined. Other elements can optionally be present other than the elements specifically identified by the “and / or” clause, whether related or unrelated to those elements specifically identified. Thus, as a non-limiting example, a reference to “A and / or B” , when used in conjunction with open-ended language such as “comprising” can refer, in one embodiment, to A only (optionally including elements other than B) ; in another embodiment, to B only (optionally including elements other than A) ; in yet another embodiment, to both A and B (optionally including other elements) ; etc.
[0292] As used herein in the specification and in the embodiments, “or” should be understood to have the same meaning as “and / or” as defined above. For example, when separating items in a list, “or” or “and / or” shall be interpreted as being inclusive, i.e., the inclusion of at least one, but also including more than one, of a number or list of elements, and, optionally, additional unlisted items. Only terms clearly indicated to the contrary, such as “only one of” or “exactly one of, ” or, when used in the embodiments, “consisting of, ” will refer to the inclusion of exactly one element of a number or list of elements. In general, the term “or” as used herein shall only be interpreted as indicating exclusive alternatives (i.e., “one or the other but not both” ) when preceded by terms of exclusivity, such as “either, ” “one of, ” “only one of, ” or “exactly one of.” “Consisting essentially of, ” when used in the embodiments, shall have its ordinary meaning as used in the field of patent law.
[0293] As used herein in the specification and in the embodiments, the phrase “at least one, ” in reference to a list of one or more elements, should be understood to mean at least one element selected from any one or more of the elements in the list of elements, but not necessarily including at least one of each and every element specifically listed within the list of elements and not excluding any combinations of elements in the list of elements. This definition also allows that elements can optionally be present other than the elements specifically identified within the list of elements to which the phrase “at least one” refers, whether related or unrelated to those elements specifically identified. Thus, as a non-limiting example, “at least one of A and B” (or, equivalently, “at least one of A or B, ” or, equivalently “at least one of A and / or B” ) can refer, in one embodiment, to at least one, optionally including more than one, A, with no B present (and optionally including elements other than B) ; in another embodiment, to at least one, optionally including more than one, B, with no A present (and optionally including elements other than A) ; in yet another embodiment, to at least one, optionally including more than one, A, and at least one, optionally including more than one, B (and optionally including other elements) ; etc.
[0294] Camelid Germline Sequences
[0295] REFERENCE
[0296] 1. Ambrosetti, F., Jiménez-García, B., Roel-Touris, J., &Bonvin, A. M. J. J. (2020) . Modeling Antibody-Antigen Complexes by Information-Driven Docking. Structure, 28 (1) , 119-129. e2. https: / / doi.org / 10.1016 / j.str.2019.10.011
[0297] 2. Clarke SC, Ma B, Trinklein ND, Schellenberger U, Osborn MJ, Ouisse LH, Boudreau A, Davison LM, Harris KE, Ugamraj HS, Balasubramani A, Dang KH, Jorgensen B, Ogana HAN, Pham DT, Pratap PP, Sankaran P, Anegon I, van Schooten WC, Brüggemann M, Buelow R, Force Aldred S. Multispecific Antibody Development Platform Based on Human Heavy Chain Antibodies. Front Immunol. 2019 Jan 7; 9: 3037.
[0298] 3. Dauparas, J., Anishchenko, I., Bennett, N., Bai, H., Ragotte, R. J., Milles, L. F., Wicky, B. I. M., Courbet, A., de Haas, R. J., Bethel, N., Leung, P. J. Y., Huddy, T. F., Pellock, S., Tischer, D., Chan, F., Koepnick, B., Nguyen, H., Kang, A., Sankaran, B., …Baker, D. (2022) . Robust deep learning-based protein sequence design using ProteinMPNN. Science, 378 (6615) , 49-56.
[0299] 4. Fesseha, H. (2020) . Monoclonal Antibody and its Diagnostic Application-Review. Biomedical Journal of Scientific &Technical Research, 30 (4) .
[0300] 5. Ferruz, N., Schmidt, S., & B. (2022) . ProtGPT2 is a deep unsupervised language model for protein design. Nature Communications, 13 (1) .
[0301] 6. Joseph L. Watson, David Juergens, Nathaniel R. Bennett, Brian L. Trippe, Jason Yim, Helen E. Eisenach, Woody Ahern, Andrew J. Borst, Robert J. Ragotte, Lukas F. Milles, Basile I. M. Wicky, Nikita Hanikel, Samuel J. Pellock, Alexis Courbet, William Sheffler, Jue Wang, Preetham Venkatesh, Isaac Sappington, Susana Vázquez Torres, Anna Lauko, Valentin De Bortoli, Emile Mathieu, Regina Barzilay, Tommi S. Jaakkola, Frank DiMaio, Minkyung Baek, David Baker. Broadly applicable and accurate protein design by integrating structure prediction networks and diffusion generative models. bioRxiv 2022. 12. 09. 519842.
[0302] 7. Jumper, J., Evans, R., Pritzel, A. et al. Highly accurate protein structure prediction with AlphaFold. Nature 596, 583-589 (2021) . https: / / doi. org / 10.1038 / s41586-021-03819-2
[0303] 8. Krissinel, Evgeny, and Kim Henrick. “Inference of Macromolecular Assemblies from Crystalline State. ” Journal of Molecular Biology 372, no. 3 (September 21, 2007) : 774-97. https: / / doi.org / 10.1016 / j. jmb.2007.05.022.
[0304] 9. Moutel S, Bery N, Bernard V, Keller L, Lemesre E, de Marco A, Ligat L, Rain JC, Favre G, Olichon A, Perez F. NaLi-H1: A universal synthetic library of humanized nanobodies providing highly functional antibodies and intrabodies. Elife. 2016 Jul 19; 5: e16228.
[0305] 10. Prihoda, D., Maamary, J., Waight, A., Juan, V., Fayadat-Dilman, L., Svozil, D., &Bitton, D. A. (2022) . BioPhi: A platform for antibody design, humanization, and humanness evaluation based on natural antibody repertoires and deep learning. MAbs, 14 (1) .
[0306] 11. Radford, Alec and Wu, Jeff and Child, Rewon and Luan, David and Amodei, Dario and Sutskever. (2019) . Language Models are Unsupervised Multitask Learners. https: / / cdn. openai. com / better-language-models / language_models_are_unsupervised_multitask_learners. pdf
[0307] 12. Ramon, A., Ali, M., Atkinson, M., Saturnino, A., Didi, K., Visentin, C., Ricagno, S., Xu, X., Greenig, M., &Sormanni, P. (2024) . Assessing antibody and nanobody nativeness for hit selection and humanization with AbNatiV. Nature Machine Intelligence, 6 (1) , 74-91. https: / / doi.org / 10.1038 / s42256-023-00778-3
[0308] 13. Rossotti, M.A., Bélanger, K., Henry, K.A. and Tanha, J. (2022) , Immunogenicity and humanization of single-domain antibodies. FEBS J, 289: 4304-4327.
[0309] 14. Sun A, Benet L, Z: Late-Stage Failures of Monoclonal Antibody Drugs: A Retrospective Case Study Analysis. Pharmacology 2020; 105: 145-163.
[0310] 15. Teng Y, Young JL, Edwards B, Hayes P, Thompson L, Johnston C, Edwards C, Sanders Y, Writer M, Pinto D, Zhang Y, Roode M, Chovanec P, Matheson L, Corcoran AE, Fernandez A, Montoliu L, Rossi B, Tosato V, Gjuracic K, Nikitin D, Bruschi C, McGuinness B, Sandal T, Romanos M. Diverse human VH antibody fragments with bio-therapeutic properties from the Crescendo Mouse. N Biotechnol. 2020 Mar 25; 55: 65-76.
[0311] 16. Tileli Amimeur, Jeremy M. Shaver, Randal R. Ketchem, J. Alex Taylor, Rutilio H. Clark, Josh Smith, Danielle VanCitters, Christine Siska, Pauline Smidt, Megan Sprague, Bruce A. Kerwin, Dean Pettit. Designing Feature-Controlled Humanoid Antibody Discovery Libraries Using Generative Adversarial Networks bioRxiv 2020.04.12.024844; doi: https: / / doi.org / 10.1101 / 2020.04.12.024844.
Claims
1.A method of generating a low immunogenicity nanobody (VHH) specific to an antigen, comprising:a) selecting camelid germlines having predicted low immunogenicity;b) constructing libraries from camelid repertoires;c) selecting VHH sequences from the libraries in (b) that are derived from the camelid germlines in (a) ; andd) generating a nanobody specific to the antigen from the VHH sequences of (c) .2.The method of Claim 1, wherein the camelid germlines having predicted low immunogenicity have a rarity score no lower than 88%or a 9-mer score no higher than 7.3.The method of Claim 1 or 2, wherein the camelid is alpaca and wherein the germlines having predicted low immunogenicity comprise IGHV3S1*01, IGHV3S35*01, IGHV3S40*01, IGHV3-1*01, IGHV3S2*01, IGHV3S10*01, IGHV3S42*01, IGHV3S41*01, IGHV3S39*01, IGHV3S6*01, IGHV3S44*01, IGHV3S7*01, IGHV3S25*01, IGHV3S15*01, IGHV3S16*01, IGHV3S37*01, IGHV3S36*01, IGHV3S11*01, IGHV3S31*01, IGHV3S12*01, IGHV3S30*01, IGHV3-2*01, IGHV3S8*01, IGHV3S26*01, IGHV3S33*01, IGHV3S32*01, IGHV3S5*01, IGHV3S38*01, IGHV3S9*01 or IGHV3S67*01.4.The method of Claim 1 or 2, wherein the camelid is llama and wherein the germlines having predicted low immunogenicity comprise IGHV3_1_llama, IGHV3_105_llama, IGHV3_106_llama, IGHV3_107_llama, IGHV3_108_llama, IGHV3_11_llama, IGHV3_114_llama, IGHV3_116_llama, IGHV3_12_llama, IGHV3_120_llama, IGHV3_121_llama, IGHV3_122_llama, IGHV3_123_llama, IGHV3_124_llama, IGHV3_126_llama, IGHV3_129_llama, IGHV3_134_llama, IGHV3_137_llama, IGHV3_139_llama, IGHV3_14_llama, IGHV3_140_llama, IGHV3_141_llama, IGHV3_147_llama, IGHV3_15_llama, IGHV3_17_llama, IGHV3_22_llama, IGHV3_25_llama, IGHV3_27_llama, IGHV3_29_llama, IGHV3_3_llama, IGHV3_31_llama, IGHV3_32_llama, IGHV3_33_llama, IGHV3_34_llama, IGHV3_35_llama, IGHV3_36_llama, IGHV3_37_llama, IGHV3_38_llama, IGHV3_39_llama, IGHV3_40_llama, IGHV3_43_llama, IGHV3_44_llama, IGHV3_48_llama, IGHV3_5_llama, IGHV3_50_llama, IGHV3_52_llama, IGHV3_55_llama, IGHV3_59_llama, IGHV3_6_llama, IGHV3_60_llama, IGHV3_61_llama, IGHV3_62_llama, IGHV3_64_llama, IGHV3_66_llama, IGHV3_68_llama, IGHV3_7_llama, IGHV3_70_llama, IGHV3_71_llama, IGHV3_73_llama, IGHV3_74_llama, IGHV3_76_llama, IGHV3_77_llama, IGHV3_80_llama, IGHV3_81_llama, IGHV3_82_llama, IGHV3_83_llama, IGHV3_86_llama, IGHV3_89_llama, IGHV3_90_llama, IGHV3_91_llama, IGHV3_93_llama, IGHV3_94_llama, IGHV3_95_llama, IGHV3_96_llama, IGHV3_99_llama, IGHV3_1*01_lama or IGHV3S6_lama.5.The method of Claim 1 or 2, wherein the camelid is dromedary and wherein the germlines having predicted low immunogenicity comprise IGHV1S1*02_Dromedary, IGHV1S18*01_Dromedary, IGHV1S22*01_Dromedary, IGHV1S25*01_Dromedary, IGHV3S1*01_Dromedary, IGHV3S1*02_Dromedary, IGHV3S1*03_Dromedary, IGHV3S1*04_Dromedary, IGHV3S11*01_Dromedary, IGHV3S12*01_Dromedary, IGHV3S12*02_Dromedary, IGHV3S18*01_Dromedary, IGHV3S19*01_Dromedary, IGHV3S2*01_Dromedary, IGHV3S20*01_Dromedary, IGHV3S22*01_Dromedary, IGHV3S23*01_Dromedary, IGHV3S24*01_Dromedary, IGHV3S25*01_Dromedary, IGHV3S26*01_Dromedary, IGHV3S27*01_Dromedary, IGHV3S28*01_Dromedary, IGHV3S29*01_Dromedary, IGHV3S3*01_Dromedary, IGHV3S30*01_Dromedary, IGHV3S31*01_Dromedary, IGHV3S32*01_Dromedary, IGHV3S33*01_Dromedary, IGHV3S35*01_Dromedary, IGHV3S36*01_Dromedary, IGHV3S38*01_Dromedary, IGHV3S4*01_Dromedary, IGHV3S5*01_Dromedary, IGHV3S6*01_Dromedary, IGHV3S7*01_Dromedary or IGHV3S9*01_Dromedary.6.The method of Claim 1 or 2, wherein the camelid is Bactrian camel and wherein the germlines having predicted low immunogenicity comprise IGHV3|cb_vh_63, IGHV3S40_Bactrian, IGHV3S51_Bactrian, IGHV3|cb_vh_79, IGHV3|cb_vh_45, IGHV3|cb_vh_29, IGHV3|cb_vh_34, IGHV3|cb_vh_52, IGHV3|cb_vh_100, IGHV3|cb_vh_57, IGHV3|cb_vh_97, IGHV3|cb_vh_12, IGHV3|cb_vh_37, IGHV3|cb_vh_38, IGHV3|cb_vh_27, IGHV3|cb_vh_39, IGHV3|cb_vh_19, IGHV3|cb_vh_109, IGHV3|cb_vh_6, IGHV3|cb_vh_10, IGHV3|cb_vh_112, IGHV3S29_Bactrian, IGHV3|cb_vh_25, IGHV3S25_Bactrian, IGHV3|cb_vh_14, IGHV3-2_Bactrian, IGHV3|cb_vh_72, IGHV3|cb_vh_32, IGHV3|cb_vh_49, IGHV3|cb_vh_23, IGHV3|cb_vh_70, IGHV3|cb_vh_82, IGHV3|cb_vh_78, IGHV3|cb_vh_40, IGHV3|cb_vh_15, IGHV3|cb_vh_36, IGHV3S40*1_Bactrian, IGHV3-1_Bactrian, IGHV3S42_Bactrian, IGHV3|cb_vh_20, IGHV3|cb_vh_30, IGHV3|cb_vh_89, IGHV3|cb_vh_81, IGHV3|cb_vh_48, IGHV3|cb_vh_103, IGHV3|cb_vh_56, IGHV3|cb_vh_35, IGHV3S6_Bactrian, IGHV3|cb_vh_85, IGHV3|cb_vh_59, IGHV3|cb_vh_66, IGHV3|cb_vh_110, IGHV3|cb_vh_77, IGHV3|cb_vh_68, IGHV3|cb_vh_33, IGHV3|cb_vh_24, IGHV3|cb_vh_92, IGHV3|cb_vh_83, IGHV3|cb_vh_108, IGHV3|cb_vh_46, IGHV3|cb_vh_4, IGHV3|cb_vh_80, IGHV3|cb_vh_75, IGHV3|cb_vh_28, IGHV3|cb_vh_91, IGHV3|cb_vh_96, IGHV3|cb_vh_101, IGHV3|cb_vh_18, IGHV3|cb_vh_22, IGHV3|cb_vh_51, IGHV3|cb_vh_67, IGHV3|cb_vh_111, IGHV3|cb_vh_88, IGHV3|cb_vh_98, IGHV3|cb_vh_90, IGHV3|cb_vh_104, IGHV3|cb_vh_105, IGHV3|cb_vh_115, IGHV3|cb_vh_8, IGHV3|cb_vh_3, IGHV3|cb_vh_50, IGHV3|cb_vh_61, IGHV3|cb_vh_69, IGHV3|cb_vh_86 or IGHV3|cb_vh_11.7.The method of any one of Claims 1 to 6, wherein the camelid repertoires in (b) are enriched for camelid germlines having predicted low immunogenicity prior to constructing the libraries.8.The method of any one of Claims 1 to 7, wherein the libraries are constructed using primers specific for camelid germlines having predicted low immunogenicity.9.The method of Claim 8, wherein the primers comprise the sequences of SAGKTGCAGCTSGTGGAGTCT (Alpaca) , SAGGTGCAGCTSGTGGAGTCT (Llama) , SAGGTGCAGCTGGTGGAGTCT (Dromedary) , or SAGGTGCAGCTGGTGGAGTCT (Bactrian) .10.The method of any one of Claims 1 to 8, wherein the libraries are constructed using primers specific for nanobodies having an arginine at the first residue of FR4 (IMGT position 118) .11.The method of Claim 10, wherein primers comprise the sequences of TGAGGAGACGGTGACCTGGGTCCCCTGGCCCCK (Alpaca) , TGAGGAGACGGTGACCTGGGTCCCCTGGCCCCK (Llama) , TGAGGAGACGGTGACCTGGGTCCCCTGGCCCCK (Dromedary) , or TGAGGAGACGGTGACCTGGGTCCCCTGGCCCCK (Bactrian) .12.The method of any one of Claims 1 to 8, wherein libraries are constructed using capture oligos targeting the FR2 region of camelid germlines having predicted low immunogenicity.13.The method of Claim 12, wherein the capture oligos comprise the sequences of AGCTGGGTCCGCCAGGCTCCAGGAAAGGGGCTC (Alpaca) , AGCTGGGTCCGCCAGGCTCCAGGAAAGGGGCTC (Llama) , CTGGGTCCGCCAGGCTCCAGGGAAGGGGCT (Dromedary) , or AGCTGGGTCCGCCAGGCTCCAGGGAAGGGGCTC (Bactrian) .14.The method of any one of Claims 1 to 8, wherein the libraries are constructed using primers specific for nanobodies having an arginine at the first residue of FR4 (IMGT position 118) and capture oligos targeting the FR2 region of camelid germlines having predicted low immunogenicity.15.The method of any one of Claims 1 to 14, wherein the VHH sequences are generated from the libraries using next generation technology (NGS) .16.The method of any one of Claims 1 to 15, further comprising determining the mismatch score of the VHH sequences in (c) , wherein the VHH sequence having a mismatch score greater than 12 is excluded.17.The method of claim 1, wherein the repertoires are generated from camelids immunized with a peptide, a protein, an mRNA, a DNA or a cell.18.The method of any one of Claims 1 to 17, wherein the nanobody specific to the antigen is identified by phage panning or screening, B cell panning or screening, or NGS methods.19.A method of generating a low immunogenicity nanobody (VHH) specific to an antigen using a generative machine learning model, comprising:a) selecting VHH sequences having predicted low immunogenicity;b) training the generative machine learning model with the VHH sequences of (a) ;c) using the trained generative machine learning model of (b) to generate synthetic VHH sequences;d) constructed libraries from the synthetic VHH sequences generated in (c) ; ande) generating a nanobody specific to the antigen from the libraries in (d) .20.The method of Claim 19, wherein the VHH sequences having a predicted low immunogenicity have a rarity score higher than 85%or a 9-mer score lower than 40.21.The method of Claim 19 or 20, where generative machine learning model comprises VAE, GAN or GPT.22.The method of any one of Claims 19 to 21, where generative machine learning model is pretrained with protein or antibody sequences.23.The method of any one of Claims 19 to 22, where the libraries are constructed using DNA synthesized based on the synthetic VHH sequences in (c) .24.The method of any one of Claims 19 to 23, wherein the nanobody specific to the antigen is identified by phage panning, or in silico methods.25.The method of any one of Claims 19 to 24, wherein the VHH sequences having a predicted low immunogenicity have a rarity score higher than 85%or a 9-mer score lower than 40.26.A method of generating a low immunogenicity nanobody (VHH) specific to an antigen using one or more machine learning models, comprising:a) selecting VHH templates having predicted low immunogenicity;b) using a first machine learning model to generate protein backbone structures that specifically bind to the antigen based on the structures of the VHH templates in (a) and the structure of the antigen;c) using a second machine learning model to generate synthetic VHH sequences based on the protein backbone structures in (b) ; andd) synthesize a nanobody specific to the antigen having the synthetic VHH sequence in (c) .27.The method of Claim 26, wherein the VHH templates having a predicted low immunogenicity have a rarity score higher than 85%or a 9-mer score lower than 40.28.The method of Claim 26 or 27, where the structure of the antigen is a 3D structure determined experimentally or a modelled structure determined by protein modeling tools.29.The method of any one of Claims 26 to 28, where the VHH templates comprises the sequences of:a) FR1: EVQLVESGGGLVQPGGSLRLSCAASFR2: MSWFRQAPGKEREGVSAFR3: YYADSVKGRFTISRDNSKNTLYLQMNSLRAEDTAVYYCFR4: WGQGTLVTVSS;b) FR1: EVQLVESGGGLVQPGGSLRLSCAASFR2: MSWYRQAPGKEREGVSAFR3: YYADSVKGRFTISRDNSKNTLYLQMNSLRAEDTAVYYCFR4: WGQGTLVTVSS;c) FR1: EVQLVESGGGLVQPGGSLRLSCAASFR2: MSWVRQAPGKGLEWVSAFR3: YYADSVKGRFTISRDNSKNTLYLQMNSLRAEDTAVYYCFR4: RGQGTQVTVSS;d) FR1: QVQLVESGGGLVQPGGSLRLSCAASFR2: MSWFRQAPGKEREWVSFR3: YADSVKGRFTISRDNSKNTLYLQMNSLRAEDTAVYYCFR4: WGQGTLVTVSS;e) FR1: QVQLVESGGGLVQPGGSLRLSCAASFR2: MSWYRQAPGKEREWVSFR3: YADSVKGRFTISRDNSKNTLYLQMNSLRAEDTAVYYCFR4: WGQGTLVTVSS;f) FR1: QVQLVESGGGLVKPGGSLRLSCAASFR2: MSWVRQAPGKGLEWVSFR3: YADSVKGRFTISRDNSKNTLYLQMNSLRAEDTAVYYCFR4: RGQGTLVTVSS;g) FR1: QVQLVESGGGLVQPGGSLRLSCAASFR2: MSWFRQAPGKEREWVSFR3: YADSVKGRFTISRDNSKNTLYLQMNSLRAEDTAVYYCFR4: WGQGTQVTVSS;h) FR1: QVQLVESGGGLVQPGGSLRLSCAASFR2: MSWYRQAPGKEREWVSFR3: YADSVKGRFTISRDNSKNTLYLQMNSLRAEDTAVYYCFR4: WGQGTQVTVSS; ori) FR1: QVQLVESGGGLVKPGGSLRLSCAASFR2: MSWVRQAPGKGLEWVSFR3: YADSVKGRFTISRDNSKNTLYLQMNSLRAEDTAVYYCFR4: RGQGTQVTVSS.30.The method of any one of Claims 26 to 29, wherein the first machine learning model is RFdiffusion.31.The method of any one of Claims 26 to 30, wherein the second machine learning model is proteinMPNN.32.The method of any one of Claims 26 to 31, wherein the synthetic VHH sequences in (c) are further screened based on their physical or biochemical properties.33.The method of any one of Claims 1-32, wherein the nanobodies are expressed by prokaryotic or eukaryotic cells.
Citation Information
Patent Citations
Sequence-based high throughput methods to produce camel antibodies to cover wide range of epitopes with high resolution
CN114126646A
A method for protein design
WO2023198726A1
Selection of nanobodies using sequence features
WO2024094096A1