Compositions and methods for identifying nanobodies and nanobody affinity

CN116457368BActive Publication Date: 2026-09-04UNIV OF PITTSBURGH OF THE COMMONWEALTH SYST OF HIGHER EDUCATION
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202180045907.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-05-01
Filing Date
2021-04-29
Publication Date
2026-09-04
Estimated Expiration
2041-04-29

AI Technical Summary

Technical Problem

特异性主要由互补决定区(CDR)决定,其中CDR3环可能很长,使得难以进行确信的MS分析

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116457368B_ABST
    Figure CN116457368B_ABST
Patent Text Reader

Abstract

Provided herein are methods of identifying a set of complementarity determining region (CDR) 3, 2, and / or 1 nanobody amino acid sequences (CDR3, CDR2, and / or CDR1 sequences) in which a reduced number of said CDR3, CDR2, and / or CDR1 sequences are false positives as compared to a control; methods of determining antigen affinity of a nanobody peptide sequence; and related methods of training a deep learning model.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-references to related applications

[0002] This application claims the benefit of U.S. Provisional Application No. 63 / 018,559, filed May 1, 2020, which is expressly incorporated herein by reference in its entirety. Background Technology

[0003] Nanobodies (Nb) are V-type antibodies derived from camel-derived heavy chain antibodies (HcAb). H Natural antigen-binding fragments of the H domain. They are characterized by their small size and robust structure, excellent solubility and stability, ease of bioengineering and manufacturing, low immunogenicity in humans, and rapid tissue penetration. For these reasons, Nb has become a promising agent for cutting-edge biomedical, diagnostic, and therapeutic applications (Muyldermans, 2013; Beghein, 2017; Rasmussen, 2011; Jovcevska, I. and Muyldermans, S., 2020).

[0004] Display-based techniques have been developed for Nb discovery (Lauwereys, 1998; Pardon, 2014; McMahon, 2018; Egloff, 2019). These methods typically produce small amounts of target-synthesized Nb that bind to specific targets with moderate affinity and do not directly analyze naturally circulating antigen-specific HcAb / Nb libraries. Recently, mass spectrometry-based proteomics has emerged as a promising technique for Nb discovery (Fridy, 2014). However, large-scale, sensitive, and reliable analysis of antigen-specific Nb proteomes remains a significant challenge for at least several reasons: (a) the diversity and dynamic range of circulating antibodies are orders of magnitude greater than that of any cellular proteome; (b) Nb sequence databases obtained from immunized camelids often contain millions of unique sequences, posing a challenge to accurate database searches (Savitski, 2015); and (c) this vast database is overrepresented by conserved Nb framework sequences that offer little specificity for identification. Specificity is primarily determined by the complementarity-determining region (CDR), where the CDR3 loop can be very long, making it difficult to perform reliable MS analysis. (d) Current methods are limited by the availability of effective protocols and informatics capable of accurately quantifying and classifying large Nb libraries. Summary of the Invention

[0005] This article provides a method for identifying a set of complementarity-determining region (CDR) 3, 2, and / or 1 nanobody amino acid sequences (CDR3, CDR2, and / or CDR1 sequences), wherein a reduced number of CDR3, CDR2, and / or CDR1 sequences compared to a control is a false positive. The method comprises: (a) obtaining a blood sample from a camel immunized with an antigen; (b) obtaining a nanobody cDNA library from the blood sample; (c) identifying the sequence of each cDNA in the library; (d) isolating the nanobody from the same or a second blood sample from a camel immunized with the antigen; and (e) using trypsin. (f) Digesting the nanobody with an enzyme or chymotrypsin to produce a set of digestion products; (g) performing mass spectrometry analysis on the digestion products to obtain mass spectrometry data; (h) selecting sequences identified in step c. that are relevant to the mass spectrometry data; (h) identifying sequences from the CDR3, CDR2, and / or CDR1 regions of the sequences from step g.; and (i) selecting from the CDR3, CDR2, and / or CDR1 region sequences of step h. those sequences having a percentage of fragmentation coverage equal to or greater than the desired percentage; wherein the selected sequences in step (i) comprise a group with a reduced number of false-positive CDR3, CDR2, and / or CDR1 sequences. In some embodiments, step (d) includes obtaining plasma from a blood sample and isolating the nanobody using one or more affinity separation methods. In some aspects, the one or more affinity separation methods in step (d) include one or more of protein G agarose affinity chromatography and protein A agarose affinity chromatography. In some aspects, step (d) further includes a function selection step, which includes selecting antigen-specific nanobodies using antigen-specific affinity chromatography and eluting antigen-specific nanobodies at different stringency levels to generate different nanobodies fractions, and performing steps (e) to (i) separately for each fraction, and estimating the affinity of each different step (i) CDR3, CDR2 and / or CDR1 region sequence for the antigen based on the relative abundance of the CDR3, CDR2 and / or CDR1 region sequences in each nanobodies fraction.

[0006] In some embodiments, a set of complementarity-determining region (CDR) 3 nanobody amino acid sequences (CDR2 sequences) are used, wherein a reduced number of CDR3 sequences compared to a control are considered false positives. The method includes: (a) obtaining a blood sample from a camel immunized with an antigen; (b) obtaining a nanobody cDNA library using the blood sample; (c) identifying the sequence of each cDNA in the library; (d) isolating the nanobody from the same or a second blood sample from the camel immunized with the antigen; (e) digesting the nanobody with trypsin or chymotrypsin to produce a set of digestion products; (f) performing mass spectrometry analysis on the digestion products to obtain mass spectrometry data; (g) selecting sequences identified in step c. that are relevant to the mass spectrometry data; (h) identifying sequences from the CDR3 regions in the sequences from step g.; and (i) selecting from the CDR3 region sequences from step h. those sequences having a percentage of fragmentation coverage equal to or greater than the desired percentage; wherein the selected sequences in step (i) comprise the group with a reduced number of false positive CDR3 sequences. In some embodiments, step (d) includes obtaining plasma from a blood sample and isolating nanobodies using one or more affinity separation methods. In some aspects, the one or more affinity separation methods in step (d) include one or more of protein G agarose affinity chromatography and protein A agarose affinity chromatography. In some aspects, step (d) further includes a function selection step, which includes selecting antigen-specific nanobodies using antigen-specific affinity chromatography and eluting the antigen-specific nanobodies at different stringency levels to produce different nanobody fractions, and performing steps (e) through (i) individually for each fraction, and estimating the affinity of the CDR3 region sequence for the antigen in each different step (i) based on the relative abundance of the CDR3 region sequence in each nanobody fraction.

[0007] In some embodiments, a set of complementarity-determining region (CDR) 2 nanobody amino acid sequences (CDR2 sequences) are used, wherein a reduced number of CDR2 sequences compared to a control is considered a false positive. The method includes: (a) obtaining a blood sample from a camel immunized with an antigen; (b) obtaining a nanobody cDNA library using the blood sample; (c) identifying the sequence of each cDNA in the library; (d) isolating the nanobody from the same or a second blood sample from the camel immunized with the antigen; (e) digesting the nanobody with trypsin or chymotrypsin to produce a set of digestion products; (f) performing mass spectrometry analysis on the digestion products to obtain mass spectrometry data; (g) selecting sequences identified in step c. that are relevant to the mass spectrometry data; (h) identifying sequences from the CDR2 regions in the sequences from step g.; and (i) selecting from the CDR2 region sequences from step h. those sequences having a percentage of fragmentation coverage equal to or greater than the desired percentage; wherein the selected sequences in step (i) comprise the group with a reduced number of false positive CDR2 sequences. In some embodiments, step (d) includes obtaining plasma from a blood sample and isolating nanobodies using one or more affinity separation methods. In some aspects, the one or more affinity separation methods in step (d) include one or more of protein G agarose affinity chromatography and protein A agarose affinity chromatography. In some aspects, step (d) further includes a function selection step, which includes selecting antigen-specific nanobodies using antigen-specific affinity chromatography and eluting the antigen-specific nanobodies at different stringency levels to produce different nanobody fractions, and performing steps (e) through (i) individually for each fraction, and estimating the affinity of the CDR2 region sequence for the antigen in each different step (i) based on the relative abundance of the CDR2 region sequence in each nanobody fraction.

[0008] In some embodiments, a set of complementarity-determining region (CDR) 1 nanobody amino acid sequences (CDR1 sequences) are used, wherein a reduced number of CDR1 sequences compared to a control is considered a false positive. The method includes: (a) obtaining a blood sample from a camel immunized with an antigen; (b) obtaining a nanobody cDNA library using the blood sample; (c) identifying the sequence of each cDNA in the library; (d) isolating the nanobody from the same or a second blood sample from the camel immunized with the antigen; (e) digesting the nanobody with trypsin or chymotrypsin to produce a set of digestion products; (f) performing mass spectrometry analysis on the digestion products to obtain mass spectrometry data; (g) selecting sequences identified in step c. that are relevant to the mass spectrometry data; (h) identifying sequences from the CDR1 regions in the sequences from step g.; and (i) selecting from the CDR1 region sequences from step h. those sequences having a percentage of fragmentation coverage equal to or greater than the desired percentage; wherein the selected sequences in step (i) comprise the group with a reduced number of false positive CDR1 sequences. In some embodiments, step (d) includes obtaining plasma from a blood sample and isolating nanobodies using one or more affinity separation methods. In some aspects, the one or more affinity separation methods in step (d) include one or more of protein G agarose affinity chromatography and protein A agarose affinity chromatography. In some aspects, step (d) further includes a function selection step, which includes selecting antigen-specific nanobodies using antigen-specific affinity chromatography and eluting the antigen-specific nanobodies at different stringency levels to produce different nanobody fractions, and performing steps (e) through (i) individually for each fraction, and estimating the affinity of the CDR1 region sequence for the antigen in each different step (i) based on the relative abundance of the CDR1 region sequence in each nanobody fraction.

[0009] In some embodiments, antigen-specific affinity chromatography uses a resin that binds to the antigen. In some embodiments, antigen-specific affinity chromatography uses a resin coupled to a protein tag and an antigen. In some embodiments, antigen-specific affinity chromatography uses a resin coupled to a maltose-binding protein and an antigen.

[0010] Some aspects further include generating a CDR3, CDR2, or CDR1 peptide having the sequence identified in step (i). Some aspects further include generating a CDR3, CDR2, and / or CDR1 region having the sequence identified in step (i).

[0011] This article also includes nanobodies containing amino acid sequences selected from SEQ ID NO: 1-2536 and SEQ ID NO: 2665-2667.

[0012] This document further provides a computer-implemented method comprising: (a) receiving a nanobody peptide sequence; (b) identifying a plurality of complementarity-determining regions (CDRs) of the nanobody peptide sequence, the CDRs including CDR3, CDR2, and / or CDR1 regions; (c) applying a fragmentation filter to discard one or more false-positive CDR3, CDR2, and / or CDR1 regions of the nanobody peptide sequence; (d) quantifying the abundance of one or more undiscarded CDR3, CDR2, and / or CDR1 regions of the nanobody peptide sequence; and (e) inferring antigen affinity based on the quantitative abundance of one or more undiscarded CDR3, CDR2, and / or CDR1 regions of the nanobody peptide sequence.

[0013] In some embodiments, the computer-implemented method further includes classifying one or more undiscarded CDR3, CDR2, and / or CDR1 regions of the nanobody peptide sequence as having low antigen affinity, intermediate antigen affinity, or high antigen affinity.

[0014] In some embodiments, the computer-implemented method further includes assembling one or more undiscarded CDR3, CDR2, and / or CDR1 regions of a nanobody peptide sequence that are classified as having high antigen affinity into a nanobody protein.

[0015] In some aspects of the computer-implemented method, the fragmentation filter is configured to require a minimum computational fragmentation coverage percentage. In other or additional aspects, the minimum computational fragmentation coverage percentage is about 30%. In some aspects, the minimum computational fragmentation coverage percentage for trypsin-treated samples is about 50%, and for chymotrypsin-treated samples, the percentage is about 40%.

[0016] In some implementations, the computer-implemented method further includes receiving a plurality of nanobody peptide sequences; and comparing each of the nanobody peptide sequences with a database to separate the nanobody peptide sequences into excluded and non-excluded subgroups, wherein no excluded subgroup nanobody peptide sequence is found in the database, and wherein the CDR region is identified only in the non-excluded subgroup nanobody peptide sequences.

[0017] In some embodiments of the computer-implemented method, the abundance of one or more undiscarded CDR3, CDR2, and / or CDR1 regions of the nanobody peptide sequence is quantified based on the relative MS1 ion signal intensity. In some embodiments, antigen affinity is inferred using k-means clustering based on epitope similarity.

[0018] This paper also provides a method for training a deep learning model, the method comprising: creating a dataset using the computer-implemented method described above; and using the dataset to train a deep learning model to classify nanobody peptide sequences with low antigen affinity and nanobody peptide sequences with high antigen affinity, wherein the dataset includes multiple nanobody peptide sequences and corresponding antigen affinity markers. In some embodiments, the deep learning model is a convolutional neural network.

[0019] This paper further provides a method for determining the antigen affinity of a nanobody peptide sequence, the method comprising: receiving a nanobody peptide sequence; inputting the nanobody peptide sequence into a trained deep learning model; and using the trained deep learning model to classify the nanobody peptide sequence as having low antigen affinity or high antigen affinity. In some embodiments, the deep learning model is a convolutional neural network. In some embodiments, the trained deep learning model is trained according to the method described above for training the deep learning model. Attached Figure Description

[0020] Figure 1 (AK). Computer simulation analysis of the NGS Nb database reveals the advantages of chymotrypsin in Nb proteomics. (A) Nb crystal structure (PDB: 4QGY). CDR rings are color-coded. (B) Sequence length distribution of CDRs in the database. (C) Cumulative plot of computer simulation digestion of the Nb database by two proteases and corresponding peptide masses. (D) Length distribution of CDR3 peptides digested by trypsin and chymotrypsin. (E) Complementarity of trypsin and chymotrypsin based on simulated Nb mapping. 10,000 Nb samples with unique CDR3 sequences were randomly selected and computer simulation digestion was performed to produce CDR3 peptides. Peptides with molecular weights of 0.8–3 kDa and sufficient CDR3 coverage (≥30%) were used for Nb mapping. (FG) Identification of unique CDR3 peptides based on the percentage of CDR3 fragment ions matched in MS / MS spectra (1F: trypsin; 1G: chymotrypsin). CDR3 peptides were identified using database searches using either the “Target” database (light orange) or the “Decoy” database (gray). (HK) 3D plot of normalized CDR3 peptide identification, CDR3 fragmentation percentage, and CDR3 length from the Target database search. FDR: False Detection Rate. FDRs of CDR3 identification are colored on the 3D plot. Color bars show the proportion of FDRs. FDRs below 5% are shown in gradient red. (1H: Trypsin analysis; 1I: Chymotrypsin analysis). (JL). Representative high-quality MS / MS spectra of CDR3 peptides digested by trypsin and chymotrypsin. The sequence in Figure 1K is NTVYLEMNSLKPEDTAVYSCAAGVSDYGCYR (SEQ ID NO: 2656). The sequence in Figure 1L is YCAAAEGLASGSY (SEQ ID NO: 2657).

[0021] Figure 2 (AG). Schematic diagram of a hybrid proteomics pipeline for reliable and in-depth analysis of the antigen-binding Nb proteome. (A) Schematic diagram of the Nb proteomics pipeline. The pipeline consists of three main parts: purification of Nb for camel immunization and antigen-specific Nb, proteomics analysis of Nb (facilitated by dedicated software Augur Llama and deep learning), and high-throughput comprehensive structural analysis of antigen-Nb complexes. (B) ELISA measurement of camel immunization responses to three antigens: GST, HSA, and PDZ. (C) Identification of unique CDR combinations and unique CDR3 sequences for different antigens. (D) Trypsin and chymotrypsin in high-quality Nb... GST Comparison of CDR3 plots. (E) Nb of three different proteases (gluC, trypsin, and chymotrypsin). GSTComparison of CDR3 identification. Results are based on three independent experiments. (F) Solubility of randomly selected antigen-specific Nb. (G) Validation of antigen binding of selected Nb.

[0022] Figure 3 (AL). Classification of Nb libraries for GST, HSA, and PDZ binding. (A) Chymotrypsin against CDR3 GST Label-free MS quantification and thermographic analysis of fingerprints. (B) Chymotrypsin on label-free CDR3. GST Reproducibility and accuracy of peptide quantification. (C) Percentage of different Nb affinity clusters classified by quantitative proteomics. (D) Nb ELISA affinity (LogIC50 at OD450 nm) versus SPRK. D Linear correlation of measurements (R) 2 =0.85). (E) Box plots of ELISA affinity for different Nb clusters. p-values ​​are calculated based on the Student's t-test. * indicates p-value < 0.05, ** indicates < 0.01, *** indicates < 0.001, **** indicates < 0.0001, ns indicates not significant. (F) Summary of 25 Nb clusters. HSA (Circle) ELISA affinity graph at OD450nm. Based on the top 14 Nb K values ​​in ELISA... D Affinity was measured using the SPR (triangle) method. (G) Summary of 11 soluble Nb PDZ A graph of ELISA affinity. (H) Representative Nb from three different affinity clusters. GST SPR kinetic analysis. For G60(C1), Ka(1 / Ms)=4.9e3, Kd(1 / s)=5.9e-3, K D =1.3μM; for G95(C2), Ka(1 / Ms) = 1.4e4, Kd(1 / s) = 1.1e-3, K D =77nM; For G13(C3), Ka(1 / Ms) = 4.74e5, Kd(1 / s) = 1.7e-4, K D =360pM. (I) High affinity Nb HSA Representative SPR kinetic measurements. For H14, Ka(1 / Ms) = 2.5e5, Kd(1 / s) = 5.75e-6, K D =22.3 pM. (J)Nb PDZ SPR kinetic measurements of P10. For P10, Ka(1 / Ms) = 2.06e6, Kd(1 / s) = 9.03e-6, K D=4.4 pM. (K) Immunoprecipitation of GST (1 nM) by different Nb-conjugated dynabeads and GSH resin. (L) Schematic diagram of the PDZ domain of mammalian mitochondrial outer membrane protein 25. Nb PDZ Fluorescence microscopy analysis of P10. Nb was combined with Alexa Fluor 647 for immunostaining of native mitochondria in the COS-7 cell line. Mitotracker was used as a positive control.

[0023] Figure 4 (AK). Structural landscape of the HSA-specific Nb proteome revealed by a comprehensive structural approach. (A) Sequence changes in pI and hydrophilicity between human and camelid serum albumin (top). Heatmap of major epitopes mapped by structural docking (bottom). (B) Cartoon representation of the four major HSA epitopes. HSA is shown in gray. E1, E2, and E3 are shown in light orange, orange, and cyan, respectively. (C) Surface representation showing the colocalization of the electrostatic potential surface with the three major epitopes. (D) HSA epitopes and their fractions (%) based on the polymerization crosslinking model (E1: residues 57-62, 135-169; E2: 322-331, 335, 356-365, 395-410; E3: 29-37, 86-91, 117-123, 252-290; E4: 566-585, 595, 598-606 and E5: 188-208, 300-306, 463-468). (EG) Representative crosslinking model of the HSA-Nb complex. The best scoring model is presented. Satisfactory DSS or EDC crosslinks are shown as blue bars. (H) The putative salt bridge between glutamic acid 400 (HSA) and arginine 108 of Nb CDR3 is presented. Local sequence alignment between HSA and camel albumin is shown. (I) ELISA affinity screening of 19 different Nb species with wild-type HSA and point mutant (E400R) (heatmap). * indicates decreased affinity. (J) RMSD (root mean square deviation) plot of HSA-Nb crosslinking model. (K) Bar chart showing the percentage of all DSS and EDC crosslinks of HSA-Nb that satisfy the model.

[0024] Figure 5 (AK). Mechanism of Nb affinity maturation. (A) High affinity (dark) and low affinity (light) Nb GST and Nb HSA (A) CDR3 length distribution. (B) Comparison of pI for different Nb. (C) Comparison of pI and hydrophilicity of CDRs among different Nb. (E) CDR3 sequence diagram. Alignment was based on 1,000 randomly selected unique CDR3 sequences of 15 residues each with the same length. Schematic diagram of CDR3 architecture: The highly variable "head" is dark gray, and the semi-variable "torso" is light gray. (F) CDR3 head (NbGST and Nb HSA ) and CDR2(Nb GST A pie chart showing the amino acid composition of (G)Nb. Only the top 6 most abundant residues are shown. GST and Nb HSA The relative changes in the abundant amino acids at the CDR3 head are shown. This reveals the positively charged residues of K (lysine) / R (arginine) / H (histidine), the negatively charged residues of D (aspartic acid) / E (glutamic acid), the aromatic residues of Y (tyrosine), and the small, flexible amino acids of G (glycine) / S (serine). (H) High affinity vs. low affinity Nb HSA The relative abundances of Y, G, and S at the CDR3 head were compared. Their relative abundances were plotted as a function of the relative positions of the individual residues. A representative structure of the antigen-Nb complex (PDB: 5F1O) with two tyrosine residues at the CDR3 head was shown inserted into the deep pocket of the antigen. (I) ELISA affinity with Nb HSA A correlation diagram showing the number of specific amino acids on the CDR3 head. Pearson correlation coefficients and statistics are displayed. (J) ELISA Affinity and Nb GST Correlation plot of the number of positively charged residues on CDR2. (K) Sequence markers of two representative convolutional CDR3 filters learned by the deep learning model (filter 14 for high affinity Nb). HSA Filter 3 is used for low-affinity Nb HSA The sequence in the upper part of Figure 5K is SEQ ID NO: 2661 (YXXXXXX, residue 2 can be Y, L, D, R or I; residue 3 can be K or G; residue 4 can be R, Y, T or D; residue 5 can be P, D or R; residue 6 can be E, Y, V, P, W or D; residue 7 can be G, W, D or P). The sequence in the lower part of Figure 5K is SEQ ID NO: 2662 (YXXXLXX, residue 2 can be D, P, K or A; residue 3 can be F, P, D or A; residue 4 can be H, T or G; residue 6 can be G or N; residue 7 can be R, P, D or Y).

[0025] Figure 6 (AH): The remarkable versatility of Nb in antigen binding. (A) Electrostatic potential surface of the PDZ domain and major E2 epitopes (PDB: 2JIK; E1: 7-8, 35-36, 43, 99-100 and E2: 25-26, 45-46, 48, 78-79, 82-83, 85-86). (B) High-affinity Nb PDZDocking model of P10 with long CDR3 (deep orange). (C) Comparison of crystal structure of PDZ-peptide ligand complex (PDB: 1EB9) and docking model of PDZ-Nb complex. Conserved ligand binding sites are shown in cyan. Side chains of CDR3 and peptide ligand are shown. (D) Heatmap showing ELISA affinity of 11 different Nb with wild-type or mutant (R46E: K48D) PDZ. * indicates a 10-100,000-fold decrease in ELISA affinity. (E) Different Nb (high affinity Nb) HSA 、Nb GST 、Nb PDZ Comparison of CDR3 length (top) and pI (bottom) of Nb from sequence databases. Data were smoothed using a Gaussian function. (F) Comparison of pI and hydrophilicity among different Nb species. (G) Pie chart of the six most abundant amino acids on the Nb CDR3 head. (H) Schematic model of Nb binding antigen.

[0026] Figure 7(AF). NGS Nb database analysis and identification of representative false-positive CDR3 peptides. (A) Normalized variability of Nb sequences. Approximately 500,000 unique Nb sequences were compared based on the IMGT numbering scheme to generate a map. Amino acids were grouped based on their properties (i.e., positive, negative, polar, and nonpolar) and color-coded. (B) Quality distribution of approximately 1.5 million human protein peptides identified from PeptideAtlas. (C) Computer-simulated digestion of the Nb NGS database with different proteases (AspN, GluC, LysC, trypsin, and chymotrypsin) and peptide quality mapping. (D) Overlap between the target Nb sequence database of immunized llamas and another local llama bait database. Each database includes approximately 500,000 sequences. (E) Representative low-quality / false-positive MS / MS spectra (HCD) of trypsin CDR3 peptides. (F) Representative low-quality / false-positive MS / MS spectra (HCD) of chymotrypsin CDR3 peptides. Few high-resolution fragmented ions were matched in the spectra. The sequences in Figure 7E are NTVYLQMNSLKPE (SEQ ID NO: 2658) and DTSIYYCAATPVFQSMSTMATESVYDYWGQGTQVTVSSEPK (SEQ ID NO: 2659). The sequence in Figure 7F is CAAGSGVGLY (SEQ ID NO: 2660).

[0027] Figure 8 (AJ). Informatics pipeline of “Augur Llama” for Nb proteomics and Nb binding agent validation. (A) Schematic diagram of the informatics pipeline. Three modules are presented, including 1) peptide identification, 2) Nb peptide and protein quality control, and 3) quantification and classification. Nb proteomics data are first searched against a search engine. Initial identification through the search engine is automatically annotated and evaluated based on different quality filters at the peptide and protein levels. High-quality fingerprint peptides passing through the quality filters are quantified and clustered. (B) Illustration of Nb CDR3 spectra and covering quality filters. (C) Illustration of peptide classification methods. (D) Identified Nb PDZ Phylogenetic tree and Web logo analysis of 230 unique CDR3s. (E) PCR amplification of HcAb variable domains (V) from B lymphocytes of camels. H (H) Schematic diagram. (F) V from a cDNA library prepared from immune bone marrow / blood. H DNA gel electrophoresis of HPCR amplicon. (G) Fractional separation of Nb based on different fractionation schemes. GST SDS-PAGE analysis of (H)Nb. PDZ SDS-PAGE analysis. Maltose-binding protein (MBP) tags fused to the PDZ domain, and the fusion protein served as an affinity stalk for separation. MBP was used as a negative control for quantification. (I) Identification of unique Nb for different antigens. (J) Comparison of antigen-specific Nb identified by chymotrypsin-based or trypsin-based methods. The Y-axis represents the percentage of positive hits selected for validation by random selection.

[0028] Figure 9(AD). Nb GST Proteomics quantification, biochemical validation, and affinity measurement. (A) Nb based on different fractionation methods. GST (A) Proteomics quantification and heatmap analysis. (B) Pearson correlation of LC retention times for Nb peptide samples separated by different fractions. (C) Representative GST bead binding analysis. GST-coupled resin was used to specifically isolate recombinant Nb from E. coli lysates. Red arrows indicate enriched Nb. Inactivating resin was used as a negative control. (D) 10 representative Nb species. GST SPR dynamics measurement.

[0029] Figure 10 (AB). Characterization of high-quality HSA and PDZ Nb. (A) Representative high-affinity Nb HSA SPR kinetic measurements. (B) Selected high-quality Nb PDZ Bead binding analysis. Recombinant MBP fused with PDZ was used as an affinity stem for the isolation of Nb from E. coli lysates. MBP-coupled resin was used as a negative control. I: E. coli lysate input, B: bead control, P: PDZ affinity pull-out.

[0030] Figure 11 (AG). Mixed structure analysis of GST-Nb complexes. (A) Structural docking heatmap analysis of 64,670 GST-Nb complexes, showing three polymerization epitopes (E1: 75-88, 143-148; E2: 33-43, 107-127; E3: 158-200, 213-220). (B) Cartoon representation of the three major GST epitopes. GST dimers are shown in gray. E1, E2, and E3 are shown in pale yellow, orange, and dark blue-green, respectively. (C) Surface representation showing the colocalization of electrostatic surfaces with the three major epitopes. (D) GST epitopes and their abundance (%) based on the polymerization crosslinking model are shown in different colors.

[0031] Figure 12 (AH). CDR sequence analysis of different Nb and sequence conservation of camel and human albumin. (AB) Comparison of amino acid abundance at the CDR3 head between high-affinity and low-affinity Nb. (CF) Comparison of CDR1 and CDR2 of different Nb. (G) Comparison of the relative positions of tyrosine (Y), glycine (G), and serine (S) at the CDR3 head of GST Nb. (H) Sequence alignment of human serum albumin and llama serum albumin. Conserved amino acids are highlighted.

[0032] Figure 13(AF). Comparison between different antigenic epitopes. (A) Comparison of the geometry of the major epitopes of three different antigens (i.e., E2 of PDZ, E3 of GST dimer, and E3 of HSA). Different epitopes are color-coded on the antigenic structure. (B) Surface electrostatic potential of PDZ domain and E1 epitope. (C) Solvent-accessible area plot of different epitopes. The y-axis represents the area of ​​different epitopes in square angstroms. (D) Net formal charge of epitopes. (E) Relative abundance of different amino acids in the CDR3 head. DB: NGSNb sequence database. (F) Comparison of pI of CDR1 and CDR2 in different antigen-specific Nb.

[0033] Figure 14 Examples of computing systems that perform the methods and procedures described in some embodiments of this disclosure are depicted.

[0034] Figure 15(AB) shows the results of an amino acid sequence filter derived from a deep learning method. The sequence filter can be used to accurately separate high-affinity and low-affinity HSA Nb. The sequence in Figure 15A is SEQ ID NO: 2663 (LXYRXXX, residue 2 can be N, Y, V, or G; residue 5 can be L or W; residue 6 can be E, G, N, T, or S; residue 7 can be D or E). The sequence in Figure 15B is SEQ ID NO: 2664 (XXXXXXX, residue 1 can be C, F, Q, S, H, K, L, Y, or R; residue 2 can be G, P, A, or N; residue 3 can be E, S, G, T, P, V, Y, H, or A; residue 4 can be C, A, S, P, or D; residue 5 can be I, W, V, T, or A; residue 6 can be M, Q, or H; residue 7 can be K, Y, Q, V, or W).

[0035] Figure 16(AC) shows the results of an amino acid sequence filter derived from a deep learning method. The sequence filter can be used to accurately separate high-affinity and low-affinity HSA Nb. The sequence in Figure 16A is SEQ ID NO: 2665(TXXLXX; residue 2 can be D, P, K, or A; residue 3 can be F, P, L, D, or A; residue 4 can be H, T, or G; residue 6 can be G, E, N, or R; residue 7 can be R, P, G, D, or Y). The sequence in Figure 16B is SEQ ID NO: 2666(XXRXXXX; residue 1 can be E, G, W, D, or I; residue 2 can be N, G, or C; residue 4 can be A, H, or D; residue 5 can be E, R, Y, A, or T; residue 6 can be G, A, or P; residue 7 can be L, S, or Y). The sequence of Figure 16C is SEQ ID NO: 2667 (XXGAQXW; residue 1 can be R or A; residue 2 can be K or L; residue 6 can be L, G, Y or W). Detailed Implementation

[0036] This paper reports a comprehensive proteomics platform for the in-depth discovery, classification, and high-throughput structural characterization of antigen-binding Nb libraries. The sensitivity and robustness of these techniques were validated using antigens spanning three orders of magnitude in immune responses, including small, weakly immunogenic antigens from the mitochondrial membrane. Tens of thousands of highly diverse specific Nb families were confidently identified and quantified based on their physicochemical properties; a large fraction exhibited sub-nM affinity. Using high-throughput structural modeling, structural proteomics, and deep learning, >100,000 antigen-Nb complexes were systematically investigated, significantly advancing the understanding of immunogenicity and maturation of Nb affinity. The research reveals the remarkable efficiency, specificity, diversity, and versatility of the mammalian humoral immune system.

[0037] the term

[0038] Unless the context clearly specifies otherwise, as used in this specification and claims, the singular forms “a / an” and “the” include plural references. For example, the term “a cell” includes a plurality of cells, including mixtures thereof.

[0039] As used herein, the term “about” is intended to cover variations of ±20%, ±10%, ±5%, or ±1% relative to a measurable value, such as a quantity or percentage.

[0040] "Administration" to a subject or "administration" includes any route by which a drug is introduced or delivered to a subject. Administration may be performed via any suitable route, including oral, intravenous, intraperitoneal, intranasal, inhalation, etc. Administration includes self-administration and administration by another person.

[0041] The terms “antibody” and “antibodies” are used broadly herein and include polyclonal antibodies, monoclonal antibodies, and bispecific antibodies. In addition to complete immunoglobulin molecules, the term “antibody” also includes fragments or polymers of immunoglobulin molecules, as well as human or humanized forms of immunoglobulin molecules or fragments thereof. Antibodies are typically heterotetrameric glycoproteins of approximately 150,000 Daltons, composed of two identical light (L) chains and two identical heavy (H) chains. Each heavy chain has a variable domain (V) at one end. H Then there are multiple constant domains. Each light chain has a variable domain (V0) at one end. L ), and has a constant domain at the other end.

[0042] As used herein, the terms “antigen” and “immunogen” are used interchangeably and refer to a substance capable of inducing an immune response in a subject, typically a protein, nucleic acid, polysaccharide, toxin, or lipid. The term also refers to an immunologically active protein, i.e., a protein that, once administered to a subject (either directly or by administration to the subject of a nucleotide sequence or carrier encoding the protein), can elicit or target a humoral and / or cellular immune response.

[0043] The terms “antigenic determinant” and “epitope” are used interchangeably herein and refer to a location on an antigen or target recognized by an antigen-binding molecule (e.g., the nanobody of the present invention). An epitope can be formed from consecutive amino acids (“linear epitopes”) or from non-continuous amino acids arranged side-by-side through the tertiary folding of a protein. The latter type of epitope (an epitope generated from at least some non-continuous amino acids) is described herein as a “conformational epitope.” An epitope typically comprises at least three, and more usually at least five, or eight to ten amino acids in a unique spatial conformation. Methods for determining the spatial conformation of an epitope include, for example, X-ray crystallography and two-dimensional nuclear magnetic resonance. See, for example, Epitope Mapping Protocols in Methods in Molecular Biology, Vol. 66, ed. Glenn E. Morris (1996).

[0044] The terms “antigen binding site,” “binding site,” and “binding domain” refer to specific elements, portions, or amino acid residues of a polypeptide, such as a nanobody, that bind to an antigenic determinant or epitope.

[0045] As used herein, the term "biological sample" means a sample of biological tissue or bodily fluid. Such samples include, but are not limited to, tissues isolated from animals. Biological samples may also include sections of tissues such as biopsy and autopsy samples, frozen sections obtained for histological purposes, blood, plasma, serum, sputum, feces, tears, mucus, hair, and skin. Biological samples also include explants derived from patient tissues and primary and / or transformed cell cultures. Biological samples can be provided by taking cell samples from animals, but may also be provided by using previously isolated cells (e.g., cells isolated by another person, at another time, and / or for another purpose), or by performing the methods disclosed herein in vivo. Archived tissues, such as those with a history of treatment or outcomes, may also be used.

[0046] The term “cDNA library” in this article refers to a combination of different cDNA fragments that constitute a portion of the transcriptome of a given organism.

[0047] The terms "CDR" and "complementarity-determining region" are used interchangeably and refer to a portion of the antibody variable chain involved in binding to the antigen. Therefore, a CDR is part of an "antigen-binding site," or simply an "antigen-binding site." In some embodiments, a nanobody comprises three CDRs that together form the antigen-binding site.

[0048] As used herein, the term "comprising" and its variations are used synonymously with the term "including" and its variations, and are open-ended and non-restrictive terms. Although the terms "comprising" and "including" are used herein to describe various embodiments, the terms "consistently made up of" and "comprises" may be used in place of "comprising" and "including" to provide more specific embodiments, and are also disclosed.

[0049] "Composition" means any pharmaceutical agent that has a beneficial biological effect. Beneficial biological effects include both therapeutic effects (e.g., treatment of a disease or other adverse physiological condition) and preventive effects (e.g., prevention of a disease or other adverse physiological condition). The term also includes pharmaceutically acceptable pharmacologically active derivatives of the beneficial pharmaceutical agents specifically mentioned herein, including but not limited to bacteria, carriers, polynucleotides, cells, salts, esters, amides, prodrugs, active metabolites, isomers, fragments, analogs, etc. When the term "composition" is used, or when a particular composition is specifically identified, it should be understood that the term includes the composition itself as well as pharmaceutically acceptable pharmacologically active carriers, polynucleotides, salts, esters, amides, prodrugs, conjugates, active metabolites, isomers, fragments, analogs, etc.

[0050] A "control" is a substitute object or sample used for comparison purposes in an experiment. Controls can be "positive" or "negative".

[0051] "Effective amount" includes, but is not limited to, amounts that can improve, reverse, alleviate, prevent, or diagnose symptoms or signs of a medical condition or ailment (such as cancer). Unless otherwise expressly stated or provided in the context, "effective amount" is not limited to the minimum amount sufficient to improve the condition. The severity of the disease or ailment and the ability to treat, prevent, treat, or alleviate the disease or ailment can be measured by biomarkers or clinical parameters without implying any limitation. In some embodiments, the term "effective amount of recombinant nanobody" refers to an amount of recombinant nanobody sufficient to prevent, treat, or alleviate cancer.

[0052] A "fragment" or "functional fragment," whether or not it is linked to another sequence, can include the insertion, deletion, substitution, or other selected modifications of a specific region or specific amino acid residue, provided that the activity of the fragment is not significantly altered or impaired compared to the unmodified peptide or protein. These modifications can provide additional properties, such as the removal or addition of amino acids capable of disulfide bonding, increasing its biological lifespan, altering its secretion characteristics, etc. In any case, the functional fragment must possess biologically active properties, such as binding to HSA and / or improving cancer.

[0053] The term "fragmentation coverage percentage" refers to the percentage obtained using the following formula:

[0054] f(x, enzyme) is a function that calculates the fragmentation coverage (%) of the peptide digested by the enzyme.

[0055] x is the length of CDR3 in peptide mapping.

[0056] f(x, chymotrypsin) = 0.0023x 2 -0.0497x + 0.7723, x [5, 30]

[0057] f(x, trypsin) = 0.00006x 2 -0.00444x+0.9194,x[5,30].

[0058] In some implementations, a minimum calculated fragmentation coverage percentage is required. In other or additional aspects, the required minimum calculated fragmentation coverage percentage is about 30%. In some aspects, the required minimum calculated fragmentation coverage percentage is about 50% when trypsin is the enzyme and about 40% when chymotrypsin is the enzyme.

[0059] As used herein, the “functional selection step” is a method of dividing nanobodies into different parts or groups based on functional characteristics. In some embodiments, the functional characteristic is the affinity of the nanobodies for CD3, CD2, or CD1 region antigens. In other embodiments, the functional characteristic is the thermal stability of the nanobodies. In other embodiments, the functional characteristic is the intracellular permeability of the nanobodies. Therefore, the present invention includes a method for identifying a set of complementarity-determining region (CDR) 3, 2, or 1 region nanobodies amino acid sequences (CDR3, CDR2, or CDR1 sequences), wherein a reduced number of CDR3, CDR2, or CDR1 sequences compared to a control is a false positive, the method comprising: obtaining a blood sample from a camel immunized with an antigen; obtaining a nanobodies cDNA library using the blood sample; identifying the sequence of each cDNA in the library; isolating the nanobodies from the same or a second blood sample from the camel immunized with the antigen; and performing the function. The steps include: selecting the nanobody and digesting it with trypsin or chymotrypsin to produce a set of digestion products; performing mass spectrometry analysis on the digestion products to obtain mass spectrometry data; selecting the sequences identified in step c. that are relevant to the mass spectrometry data; identifying sequences from the CDR3, CDR2, or CDR1 regions in the sequences from step g.; and excluding sequences from the CDR3, CDR2, or CDR1 regions in step h. that have a fraction less than the calculated fragmentation coverage percentage; wherein the unexcluded sequences include groups with a reduced number of false positive CDR3, CDR2, or CDR1 sequences. It should be understood that the method steps following the function selection step can be performed separately for each different portion or group resulting from the function selection.

[0060] The "half-life" of the amino acid sequence, compound, or peptide of the present invention can generally be defined as the time required for the serum concentration of the amino acid sequence, compound, or peptide to decrease by 50% in vivo, for example, due to degradation of the sequence or compound and / or clearance or isolation of the sequence or compound by natural mechanisms. The in vivo half-life of the nanobodies, amino acid sequences, compounds, or peptides of the present invention can be determined in any known manner, for example by pharmacokinetic analysis. These are, for example, Kenneth, A. et al., Chemical Stability of Pharmaceuticals: A Handbook for Pharmacists; Peters, et al., Pharmacokinete analysis: A Practical Approach (1996); "Pharmacokinetics", M. G. Babaldi and D. Perron, published by Marcel Dekker, 2nd revised edition (1982).

[0061] The terms “identity” or “homology” should be interpreted as the percentage of nucleotide bases or amino acid residues in a candidate sequence that are identical to those in the corresponding sequence after sequence alignment. Vacancies may be introduced, if necessary, to achieve the maximum percentage of identity across the entire sequence, and any conserved substitutions are not considered as part of sequence identity. A polynucleotide or polynucleotide region (or polypeptide or polypeptide region) having a certain percentage (e.g., 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or higher) of “sequence identity” with another sequence means that the percentage of bases (or amino acids) are the same when comparing the two sequences. This alignment and homology percentage or sequence identity can be determined using software programs known in the art. Such alignment can be provided using, for example, the method of Needleman et al. (1970) J. Mol. Biol. 48: 443-453, which is conveniently implemented by a computer program such as the Align program (DNAstar, Inc.). In some embodiments, the percentage of identity is determined along the entire length of the sequences being compared.

[0062] As used herein, the terms “increased” or “increase” generally mean an increase in a static significant amount; for the avoidance of any doubt, “increase” means an increase of at least 10% compared to a reference level, such as an increase of at least about 20%, or at least about 30%, or at least about 40%, or at least about 50%, or at least about 60%, or at least about 70%, or at least about 80%, or at least about 90%, or an increase of up to and including 100%, or any increase between 10 and 100%, or an increase of at least about 2 times, or at least about 3 times, or at least about 4 times, or at least about 5 times, or at least about 10 times, or any increase between 2 times and 10 times or more compared to a reference level.

[0063] As used herein, the term “isolated” means isolated from a biological sample, i.e., blood, plasma, tissue, foreign body, or cell. As used herein, the term “isolated” when used in the context of, for example, nucleic acids, means nucleic acids of interest that are at least 60% free of, at least 75% free of, at least 90% free of, at least 95% free of, at least 98% free of, and even at least 99% free of other components associated with the nucleic acids prior to isolation.

[0064] The term "mass spectrometry" refers to the measurement of the mass-to-charge ratio (m / z) of one or more molecules present in a sample. "Mass spectrometry data" refers to the mass, charge, mass-to-charge ratio, molecular weight, and / or amino acid identity or sequence of one or more molecules present in a sample. In some embodiments, mass spectrometry data are the amino acid sequences of molecules present in the sample. Sequences "associated" with mass spectrometry data include cDNA sequences having the expected identical or highly similar amino acid sequences determined in the mass spectrometry step of the method. In some embodiments, sequences are associated with mass spectrometry data when there is approximately 80%, approximately 85%, approximately 90%, approximately 91%, approximately 92%, approximately 93%, approximately 94%, approximately 95%, approximately 96%, approximately 97%, approximately 98%, or approximately 99% similarity or identity. In some embodiments, sequences are associated with mass spectrometry data when there is approximately 90-100% similarity or identity.

[0065] As used in this article, the terms "nanobody" and "V" are used to refer to... H H”, V H "H antibody fragment" is used indiscriminately to refer to a variable domain of a single heavy chain of an antibody of the type found in camelids, which contains no light chain, such as light chains derived from camelids, as described in PCT Publication WO 94 / 04678, which is incorporated herein by reference in its entirety. As used herein, "single-domain antibody" means nanobodies and Fc domains.

[0066] As used herein, the term "nucleic acid" refers to a polymer composed of nucleotides, such as deoxyribonucleotides (DNA) or ribonucleotides (RNA). As used herein, the terms "ribonucleic acid" and "RNA" refer to polymers composed of ribonucleotides. As used herein, the terms "deoxyribonucleic acid" and "DNA" refer to polymers composed of deoxyribonucleotides.

[0067] As used herein, “operably linked” refers to the arrangement of polypeptide segments within a single polypeptide chain, wherein each polypeptide segment may be, but is not limited to, a protein, a fragment thereof, a linker peptide, and / or a signal peptide. The term “operably linked” may also refer to the direct fusion of different individual polypeptides within a single polypeptide or its fragments, wherein there are no intermediate amino acids between the different segments, and when individual polypeptides are linked to each other via a “connector” containing one or more intermediate amino acids.

[0068] As used herein, the terms “reduced,” “reduce,” “reduction,” or “decrease” generally mean a reduction in a statistically significant amount. However, for the avoidance of ambiguity, “reduction” means a reduction of at least 5% compared to a reference level, such as a reduction of at least about 10%, or at least about 20%, or at least about 30%, or at least about 40%, or at least about 50%, or at least about 60%, or at least about 70%, or at least about 80%, or at least about 90%, or at most and including a reduction of 100% (i.e., a level that is not present compared to the reference sample), or any reduction between 10% and 100%.

[0069] The terms "polynucleotide" and "oligonucleotide" are used interchangeably and refer to a polymeric form of nucleotides of any length (deoxyribonucleotides or ribonucleotides or their analogues). Polynucleotides can have any three-dimensional structure and can perform any known or unknown function. The following are non-limiting examples of polynucleotides: genes or gene fragments, exons, introns, messenger RNA (mRNA), transfer RNA, ribosomal RNA, ribozymes, cDNA, recombinant polynucleotides, branched polynucleotides, plasmids, vectors, isolated DNA of any sequence, isolated RNA of any sequence, nucleic acid probes, and primers. Polynucleotides may contain modified nucleotides, such as methylated nucleotides and nucleotide analogues. If present, modifications to the nucleotide structure may be conferred before or after polymer assembly. The nucleotide sequence may be interspersed with non-nucleotide components. Polynucleotides may be further modified after polymerization, for example, by binding to a labeled component. The term also refers to double-stranded and single-stranded molecules. Unless otherwise specified or required, any polynucleotide embodiment of the invention includes each of two complementary single-stranded forms known or intended to constitute a double-stranded form.

[0070] The term "peptide," in its broadest sense, refers to a compound consisting of two or more subunit amino acids, amino acid analogs, or peptide mimics. The subunits may be linked by peptide bonds. In another embodiment, the subunits may be linked by other bonds, such as esters, ethers, etc. As used herein, the term "amino acid" refers to natural and / or non-natural or synthetic amino acids, including glycine and its D or L optical isomers, as well as amino acid analogs and peptide mimics. Peptides with three or more amino acids are generally called oligopeptides if the peptide chain is short. If the peptide chain is long, the peptide is generally called a polypeptide or protein. The terms "peptide," "protein," and "polypeptide" are used interchangeably herein.

[0071] The term "recombinant" in this document refers to a combination of two or more polypeptides that are not naturally occurring.

[0072] The term "specificity" refers to the number of different types of antigens or antigenic determinants that a particular antigen-binding molecule (e.g., the nanobody of the present invention) can bind to. Nanobodies with low specificity bind to multiple different epitopes (or polypeptide regions) via a single antigen-binding site or binding domain, while nanobodies with high specificity bind to one or more epitopes (or polypeptide regions) via a single antigen-binding site or binding domain. In some embodiments, a few epitopes (or polypeptide regions) are similar or highly similar, such as cross-species epitopes. As used herein, the term "specific binding" with respect to nanobodies means that the nanobody preferentially binds to one epitope (or polypeptide region) compared to other epitopes (or polypeptide regions). Specific binding can depend on binding affinity and the stringency of the conditions under which binding occurs. In one example, the nanobody specifically binds to an epitope when high-affinity binding is present under stringent conditions. In some embodiments, the HSA-binding polypeptide or nanobody described herein specifically binds to human serum albumin.

[0073] It should be understood that the specificity of antigen-binding molecules (e.g., HSA-binding peptides, the nanobodies of this invention) can be determined based on affinity and / or cohesion. The equilibrium constant (Ka) for the dissociation of the antigen from the antigen-binding molecule is used. D Affinity, represented by K, is a measure of the binding strength between an antigenic determinant and an antigen-binding site on an antigen-binding molecule. D The smaller the value, the stronger the binding strength between the antigenic determinant and the antigen-binding molecule (or, affinity can also be expressed as the affinity constant (K)). A ), which is 1 / K DMethods for determining affinity are well known to those skilled in the art. Affinity is a measure of the strength of binding between an antigen-binding molecule (e.g., an HSA-binding peptide and the nanobody of the present invention) and the associated antigen. Affinity is related to the affinity between the antigenic determinant and the antigen-binding site on the antigen-binding molecule, as well as the number of associated binding sites present on the antigen-binding molecule. Typically, antigen-binding proteins (e.g., the HSA-binding peptide and nanobody of the present invention) will have an affinity of 10... -5 Up to 10 -12 mol / L or less, preferably 10 -7 Up to 10 -12 moles per liter or less, and more preferably 10 -8 Up to 10 -12 dissociation constant (K) per mole / liter D (That is, with 10) 5 Up to 10 12 liters / moles or more, preferably 10 7 Up to 10 12 liters / moles or more, more preferably 10 8 Up to 10 12 Association constant (K) per liter / molar A It binds to its antigen. In some implementations, Ka (association rate, 1 Ms) is approximately 10. 5 10 6 10 7 10 8 10 9 10 10 Or 10 11 In some implementations, Ka is approximately 10. 7 In some implementations, Kd (dissociation rate, s) is approximately 10. -5 10 -6 10 -7 10 -8 10 -9 10 -10 Or 10 -11 In some implementations, K D For about 10 -7 In some implementations, the antigen-binding protein disclosed herein is expressed at a concentration of less than about 10. -9 moles / liter of K D It binds to its antigen. Any K+ greater than 10 μM D The value is generally considered to represent nonspecific binding. The dissociation constant can be the actual or apparent dissociation constant, as will be clear to those skilled in the art.

[0074] The term "subject" is defined herein as including animals, such as mammals, including but not limited to primates (e.g., humans), cattle, sheep, goats, horses, dogs, cats, rabbits, rats, mice, etc. In some implementations, the subject is a human.

[0075] Compositions and methods

[0076] In some respects, this document discloses methods for identifying a set of complementarity-determining region (CDR) 3, 2, or 1 amino acid sequences (CDR3, CDR2, or CDR1 sequences) of nanobodies, wherein a reduction in the number of CDR3, CDR2, or CDR1 sequences compared to a control is considered a false positive. The term "false positive" as used herein refers to a result indicating the presence of something when it is not actually present. The phrase "a sequence is a false positive" as used herein refers to a CDR3, CDR2, and / or CDR1 sequence that does not specifically bind to the test antigen, or a CDR3, CDR2, and / or CDR1 sequence contained in a nanobody that cannot specifically bind to the test antigen. It should be understood that the number or amount of false-positive CDR3, CDR2, and / or CDR1 sequences can be reduced using the methods disclosed herein, wherein the fragmentation filter is set to at least about 30% (e.g., at least about 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 99%) for trypsin-treated samples and / or at least about 30% (e.g., at least about 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 99%) for chymotrypsin-treated samples. In some instances, false-positive CDR3, CDR2, and / or CDR1 sequences can be substantially removed using the methods disclosed herein, wherein the fragmentation filter is set to about 50% for trypsin-treated samples and / or about 40% for chymotrypsin-treated samples.

[0077] Therefore, compared to a control, the disclosed method for identifying CDR3, CDR2, and / or CDR1 sequences reduces the number of false positive CDR3, CDR2, and / or CDR1 sequences. The reduction can be, for example, at least about 2-fold, at least about 3-fold, at least about 4-fold, at least about 5-fold, at least about 10-fold, at least about 20-fold, at least about 50-fold, or at least about 100-fold compared to the number of false positive CDR3, CDR2, and / or CDR1 sequences identified without using the methods described herein.

[0078] In some implementations, the method includes:

[0079] a. Obtaining blood samples from camels immunized with antigens;

[0080] b. Obtain a nanobody cDNA library using the blood sample;

[0081] c. Identify the sequence of each cDNA in the cDNA library;

[0082] d. Isolate nanobodies from the same or second blood sample of the camel immunized with the antigen;

[0083] e. Digest the nanobody with trypsin or chymotrypsin to produce a set of digestion products;

[0084] f. Perform mass spectrometry analysis on the digestion products to obtain mass spectrometry data;

[0085] g. Select the sequence identified in step c. that is associated with the mass spectrometry data;

[0086] h. Identify the sequences from the CDR3, CDR2, and / or CDR1 regions in the sequence from step g; and

[0087] i. Select from the CDR3, CDR2 and / or CDR1 region sequences of step h. those sequences that have a percentage of fragmentation coverage equal to or greater than the desired percentage; wherein the selected sequences comprise a group with a reduced number of false positive CDR3, CDR2 and / or CDR1 sequences.

[0088] In some implementations, the method includes:

[0089] a. Obtaining blood samples from camels immunized with antigens;

[0090] b. Obtain a nanobody cDNA library using the blood sample;

[0091] c. Identify the sequence of each cDNA in the library;

[0092] d. Isolate nanobodies from the same or second blood sample of the camel immunized with the antigen;

[0093] e. Digest the nanobody with trypsin or chymotrypsin to produce a set of digestion products;

[0094] f. Perform mass spectrometry analysis on the digestion products to obtain mass spectrometry data;

[0095] g. Select the sequence identified in step c. that is associated with the mass spectrometry data;

[0096] h. Identify the sequences from the CDR3, CDR2, and / or CDR1 regions in the sequence from step g; and

[0097] i. Select from the CDR3, CDR2, and / or CDR1 region sequences of step h. those sequences having a fragmentation coverage percentage equal to or greater than the desired fragmentation coverage percentage; wherein when chymotrypsin is used in step e., the fragmentation coverage percentage is determined by the formula f(x, chymotrypsin) = 0.0023x² - 0.0497x + 0.7723, x [5, 30], or when trypsin is used in step e., the fragmentation coverage percentage is determined by the formula f(x, trypsin) = 0.00006x² - 0.00444x + 0.9194, x [5, 30], and where x is the length of the CDR3, CDR2, and / or CDR1 region sequences; and

[0098] j. Wherein the selected sequences in step i. comprise a group with a reduced number of false positive CDR3, CDR2 and / or CDR1 sequences.

[0099] In some aspects, the selected CDR3, CDR2, and / or CDR1 region sequences in step i. have a minimum desired fragmentation coverage percentage of about 30%. In some aspects, the selected CDR3, CDR2, and / or CDR1 region sequences in step i. have a minimum desired fragmentation coverage percentage of about 50%, and trypsin is used in step e. In some embodiments, the selected CDR3, CDR2, and / or CDR1 region sequences in step i. have a minimum desired fragmentation coverage percentage of about 40%, and chymotrypsin is used in step e.

[0100] It should be understood that the nanobody cDNA library in step b. is obtained from a biological sample (e.g., blood or bone marrow) of an immunized subject. In some embodiments, the cDNA library is obtained from B cells. A cDNA (cloned cDNA or complementary DNA) library is a combination of cDNAs generated from mRNA in a biological sample (e.g., blood or bone marrow sample) using reverse transcription technology. Methods for generating cDNA libraries are well known in the art. Therefore, in some embodiments, step b. further includes the steps of isolating mRNA from the biological sample (e.g., blood or bone marrow sample) and / or reverse transcribing the isolated mRNA into cDNA.

[0101] The generated cDNA is then sequenced as described in step c. In some embodiments, step c further includes amplifying the camel IgG heavy chain cDNA sequence from the variable domain to the CH2 domain using specific primers (e.g., SEQ ID NO: 2646 and SEQ ID NO: 2647), and separating the CH1-deficient V from conventional IgG (with the CH1 domain) using DNA gel electrophoresis. HThe steps include: amplifying the H gene from frame 1 to frame 4 using a second forward primer (e.g., SEQ ID NO: 2648) and a second reverse primer (e.g., SEQ ID NO: 2649); purifying the amplicon from this second PCR (e.g., using a PCR cleanup kit or isolation kit); and adding an adaptor for sequencing analysis (e.g., using forward primer SEQ ID NO: 2650 and reverse primer SEQ ID NO: 2651) for further sequencing analysis (e.g., MiSeq sequencing analysis). Sequencing analysis methods may include, for example, single-molecule real-time (SMRT) sequencing, nanopore DNA sequencing, massively parallel signature sequencing (MPSS), polymerase cloning sequencing (polony sequencing), 454 pyrosequencing, Illumina (Solexa) sequencing, combinatorial probe anchor synthesis (cPAS), SOLiD sequencing, or MiSeq sequencing.

[0102] Step d above may be performed simultaneously with, before, or after steps a, b, and / or c. In some instances, step d may also include obtaining plasma from a blood sample and separating the nanobodies using one or more affinity separation methods. The affinity separation method may be any affinity separation method known in the art, including, for example, protein G agarose affinity chromatography, protein A agarose affinity chromatography, hydroxyapatite chromatography, gel electrophoresis, or dialysis. Protein G agarose affinity chromatography and protein A agarose affinity chromatography are two well-known affinity chromatography methods (Grodzki AC, Berenstein E. (2010) Antibody Purification: Affimity Chromatography - Protein A and Protein G Sepharose. Oliver C., Jamur M. (eds.) Immunocytochemical Methods and Protocols. Methods in Molecular Biology (Methods and Protocols), Vol. 588. Humana Press). The method relies on the reversible interaction between the protein and a specific ligand immobilized in the chromatographic matrix. The sample is applied under conditions favorable to ligand-specific binding due to electrostatic and hydrophobic interactions, van der Waals forces, and / or hydrogen bonding. After washing away unbound material, the bound protein is recovered by changing the buffer conditions to favor desorption. Protein A agarose affinity chromatography and protein G agarose affinity chromatography are commonly used for antibody purification due to the high binding affinity and specificity of protein A or G to the Fc region of the antibody. In some embodiments, step d. includes one or more of protein G agarose affinity chromatography and protein A agarose affinity chromatography.

[0103] In some instances, step d. further includes a function selection step, which includes selecting antigen-specific nanobodies using antigen-specific affinity chromatography and eluting the antigen-specific nanobodies at different stringency levels to generate different nanobody fractions, and performing steps e. to i. individually for each fraction, and estimating the affinity of the antigen for each different step i. of the CDR3, CDR2, and / or CDR1 region sequences based on the relative abundance of the CDR3, CDR2, and / or CDR1 region sequences in each of the nanobody fractions. In some embodiments, antigen-specific affinity chromatography is a resin that binds to the antigen. In some embodiments, antigen-specific affinity chromatography is a resin coupled to maltose-binding proteins and the antigen.

[0104] It should be understood and considered herein that the term "rigor" refers to salt buffers of varying concentrations (e.g., about 0.1 M to about 20 M MgCl2 in a neutral pH buffer, preferably about 1 M to about 10 M MgCl2 in a neutral pH buffer, or preferably about 1 M to about 4.5 M MgCl2 in a neutral pH buffer), alkaline solutions with varying pH values ​​(e.g., 1-100 mM NaOH, about pH 11, 12, and 13), acidic solutions with varying pH values ​​(e.g., 0.1 M glycine, about pH 3, 2, and 1), or combinations thereof. It should also be understood that the terms "different nanobody fractions" or "different biochemical fractions" refer to different fractions of nanobodies eluted from an antigen-conjugated solid support (e.g., resin) under varying degrees of rigor. Nanobodies most resistant to high salt, high acidity, or high alkalinity conditions exhibit the highest affinity for the antigen.

[0105] In this document, for example in step e, the term "digestion product" refers to the peptide mixture following a digestion step using enzymes, including, for example, trypsin, chymotrypsin, LysC, GluC, and AspN. In some embodiments, nanobodies are digested with trypsin (e.g., Pierce). TM Trypsin, MS grade, catalog number: 90057), chymotrypsin (e.g., Pierce) TM chymotrypsin (TLCK treated), MS grade, catalog number: 90056), LysC (or Lys-C protease, such as Pierce) TM Lys-C protease, MS grade, catalog number: 90051), GluC (or Glu-C protease, such as Pierce) TM Glu-C protease, MS grade, catalog number: 90054) and / or AspN (or Asp-N protease, such as Pierce) TM Asp-N protease (MS grade, catalog number: 90053) digests the nanobody to produce the corresponding digestion products. Trypsin, chymotrypsin, LysC, GluC, and AspN are enzymes that digest proteins. The cleavage rules for these enzymes digesting nanobodies are:

[0106] Trypsin: C-terminus to K / R, without P.

[0107] Chymotrypsin: C-terminus to W / F / L / Y, without P following it.

[0108] GluC: C-terminus to D / E, without P-terminus.

[0109] AspN: N-terminus to D

[0110] LysC: C-terminus to K

[0111] The digestion step can be carried out at temperatures from about 2°C to about 60°C (e.g., at about 2°C, 4°C, 6°C, 8°C, 10°C, 12°C, 14°C, 16°C, 18°C, 20°C, 22°C, 24°C, 26°C, 28°C, 30°C, 32°C, 34°C, 36°C, 38°C, 40°C, 42°C, 44°C, 46°C, 48°C, 50°C, 52°C, 54°C, 56°C, 58°C, or 60°C) for about 5 minutes, 10 minutes, 30 minutes, 45 minutes, 1 hour, 2 hours, 4 hours, 6 hours, 8 hours, 10 hours, 12 hours, 14 hours, 16 hours, 18 hours, 20 hours, 22 hours, 24 hours, 36 hours, 48 ​​hours, or 72 hours.

[0112]

[0113] Step f includes mass spectrometry analysis of the digested products to obtain mass spectrometry data. Methods for peptide analysis using mass spectrometry are well known in the art. In some embodiments, the mass spectrometry analysis described herein is performed in combination with: gas chromatography-mass spectrometry (GC-MS), liquid chromatography-mass spectrometry (LC-MS), capillary electrophoresis (CE-MS), ion mobility spectrometry-mass spectrometry (IMS / MS or IMMS), matrix-assisted laser desorption / ionization (MALDI-TOF), surface-enhanced laser desorption / ionization (SELDI-TOF), or tandem MS (MS-MS). This step can identify the sequence of nanobodies or a fraction of nanobodies in the sample based on the quality of amino acids and a sequence homology search in a peptide database translated from the cDNA library of step b. In some instances, mass spectrometry is used to analyze and generate digested product spectra from each nanobody fraction separately. In some instances, the spectra of the digested products refer to electron ionization data presented as an intensity versus m / z (mass-to-charge ratio) plot.

[0114] It should be understood in this article that the determination of nanobody sequences is not solely based on mass spectrometry. It is determined by matching / correlating the sequences identified by mass spectrometry with the sequences identified by sequencing of a cDNA library. The matching sequences are then selected. Therefore, step g. includes selecting the sequences identified in step c. that are correlated with the mass spectrometry data, and step h includes identifying the CDR3 region sequence from the sequence in step g.

[0115] Step i. comprises selecting sequences from the CDR3, CDR2, and / or CDR1 region sequences of step h. that have a fragmentation coverage percentage equal to or greater than the desired fragmentation coverage percentage. In some embodiments, for trypsin-treated samples, the fragmentation coverage percentage is equal to or greater than about 30% (e.g., about 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 99%). In some embodiments, for chymotrypsin-treated samples, the fragmentation coverage percentage is equal to or greater than about 30% (e.g., at least about 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 99%). In some embodiments, the fragmentation coverage percentage of the trypsin-treated sample is about 50%, and the fragmentation coverage percentage of the chymotrypsin-treated sample is about 40%.

[0116] In some embodiments, the method described herein further includes generating nanobodies comprising CDR3, CDR2, and / or CDR1 regions having the sequences identified in step i. The nanobodies gene is cloned into a vector, which is then transformed into competent cells for nanobodies protein expression, extraction, and purification.

[0117] In some embodiments, the nanobody comprises at least 80% (e.g., at least about 80%, 85%, 90%, 95%, 98%, or 99%) of an amino acid sequence selected from the group consisting of SEQ ID NO: 1-157. In some embodiments, the nanobody has a sequence selected from the group consisting of SEQ ID NO: 1-157. In some embodiments, the nanobody comprises at least 80% (e.g., at least about 80%, 85%, 90%, 95%, 98%, or 99%) of an amino acid sequence selected from the group consisting of SEQ ID NO: 158-2536. In some embodiments, the nanobody has a sequence selected from the group consisting of SEQ ID NO: 158-2536. In some embodiments, the nanobody comprises at least 80% (e.g., at least about 80%, 85%, 90%, 95%, 98%, or 99%) of an amino acid sequence selected from the group consisting of SEQ ID NO: 2665-2667. In some embodiments, the nanobody has a sequence selected from the group consisting of SEQ ID NO: 2665-2667.

[0118] This paper discloses a PDZ-specific nanobody comprising an amino acid sequence selected from the group consisting of SEQ ID NO: 158-2536. This paper also discloses a PDZ-specific nanobody comprising an amino acid sequence selected from the group consisting of SEQ ID NO: 143-157. As used herein, "PDZ" refers to an 80-100 amino acid domain found in signal transduction proteins, also known as the DHR (Dlg homology region) or GLGF (glycine-leucine-glycine-phenylalanine) domain. PDZ domains bind to a short region at the C-terminus of other specific proteins. PDZ domains are generally classified into three distinct categories based on the chemical properties of their ligands. The difference between the different ligand categories lies in the penultimate binding residue found at the extreme COOH of the target protein. Type I domains recognize the sequence XS / TX-Φ* (where X = any amino acid, Φ = hydrophobic amino acid, *COOH terminus). Type II domains bind to ligands with the sequence X-Φ-X-Φ*. Type III domains interact with the sequence XXC*. Binding specificity within each domain category can be conferred by variant (X) residues and residues outside the canonical binding motif. Furthermore, some PDZ domains do not belong to any of these specific categories. Proteins containing PDZ domains include, but are not limited to, Erbin, GRIP, Htra1, Htra2, Htra3, PSD-95, SAP97, CARD10, CARD11, CARD14, PTP-BL, and SYNJ2BP. In some embodiments, the PDZ domain is derived from SYNJ2BP.

[0119] This document discloses GST-specific nanobodies containing the amino acid sequences listed in Table 4. It also discloses GST-specific nanobodies containing amino acid sequences selected from the group consisting of SEQ ID NO: 1-98. "Glutathione S-transferase" or "GST" herein refers to glutathione-S-transferase (GST), a family of phase II detoxification enzymes that catalyze the binding of glutathione (GSH) to various endogenous and exogenous electrophilic compounds. In some embodiments, the GST peptide is a peptide in the pGEX6p-1 vector.

[0120] This document discloses HSA-specific nanobodies, wherein the HSA-specific nanobodies comprise the amino acid sequences listed in Table 5. This document also discloses HSA-specific nanobodies, wherein the HSA-specific nanobodies comprise amino acid sequences selected from the group consisting of SEQ ID NO: 99-142. "Human serum albumin" or "HSA" herein refers to a polypeptide encoded by the ALB gene. In some embodiments, the HSA polypeptide is identified in one or more publicly available databases including: HGNC: 399, Entrez Gene: 213, Ensembl: ENSG00000163631, OMIM: 103600, UniProtKB: P02768. In some embodiments, the HSA polypeptide comprises the sequence of SEQ ID NO: 2668, or a polypeptide sequence having equal or greater than about 80%, about 85%, about 90%, about 95%, or about 98% homology to SEQ ID NO: 2668 or a polypeptide comprising a portion of SEQ ID NO: 2668. The HSA peptide of SEQ ID NO: 2668 may represent an immature or pre-processed form of mature HSA, and therefore, this document includes the mature or processed portion of the HSA peptide of SEQ ID NO: 2668.

[0121] Here, a robust proteomics pipeline was developed for large-scale quantitative analysis and epitope mapping of antigen-bound Nb proteomes based on high-throughput structural characterization of antigen-Nb complexes.

[0122] Example

[0123] Example 1. Advantages of chymotrypsin in large-scale Nb proteomics analysis.

[0124] HcAb(V HThe variable domains of the H / Nb cDNA library were amplified from B lymphocytes of two lama glama, and 13.6 million unique Nb sequences were recovered from the database via next-generation genome sequencing (NGS) (DeKosky, 2013). Approximately 500,000 Nb sequences were aligned to generate sequence markers (Fig. 1A, 7A). The CDR3 loop exhibited the greatest sequence diversity and length variation, providing excellent specificity for Nb identification (Fig. 1B, 1C). Computational simulation analysis of the Nb database showed that, due to the limited number of trypsin cleavage sites on Nb, trypsin primarily produces large CDR3 peptides (Fig. 1A). Therefore, most CDR3 residues (77%) are covered by large trypsin peptides greater than 2.5 kDa (Fig. 1D, 1E), which is suboptimal for proteomics analysis (Fig. 7B). In contrast, chymotrypsin, which is rarely used for proteomics cleavage of specific aromatic and hydrophobic residues, appears to be more suitable (Methods, Fig. 1A, 7B). 91% of the CDR3 sequence was covered by chymotrypsin peptides smaller than 2.5 kDa (Figures 1D and 1E). Random selection and simulation confirmed that chymotrypsin covered significantly more CDR3 sequence than trypsin (Figure 1F). Furthermore, there was a small overlap (approximately 9%) between the two enzymes, indicating good complementarity for efficient Nb analysis.

[0125] Due to the large database size and unusual Nb sequence structure, the estimated false discovery rate (FDR) for CDR3 identification may be exaggerated. To test this, antigen-specific HcAbs were proteased with trypsin or chymotrypsin, and identification was performed using two different databases with a state-of-the-art search engine: a specific “target” database from immunized llamas, and a similarly sized “decoy” database from unrelated llamas, which did not actually contain identical sequences (Fig. 7D). Therefore, any CDR3 peptide identified from the decoy database search was considered a false positive (Elias, JE, and Gygi, SP, 2007). A large number of false positive CDR3 peptides were nonspecifically identified from the decoy database search. These false peptide profile matches were found to often contain MS / MS fragmentation on poor CDR3 fingerprint sequences (Figs. 7E, 7F). The vast majority (95%) of these mismatches could be removed using a simple fragmentation filter we have implemented, which requires at least 50% (by trypsin, Fig. 1G) and 40% (by chymotrypsin, Fig. 1H) coverage of the CDR3 high-resolution diagnostic ion in the MS2 spectrum (Fig. 1K, 1L). The filter was further optimized based on the CDR3 length before being integrated into the new open-source software “Augur Llama” (Fig. 8A-8C) for reliable Nb proteomic analysis (Fig. 1I, 1J).

[0126] Example 2. Development of a comprehensive proteomics pipeline for Nb discovery and characterization.

[0127] This article demonstrates a robust platform (methods, methods, and techniques) for comprehensive quantitative Nb proteomics and high-throughput structural characterization of antigen-Nb complexes. Figure 2A Domestic camelids were immunized with the antigen of interest. Nb cDNA libraries were then prepared from the blood and / or bone marrow of the immunized camelids (Fridy, 2014). NGS was performed to create >10 7 A rich database of unique Nb protein sequences (Figures 8E and 8F). Simultaneously, antigen-specific V proteins were affinity-isolated from serum. H HcAbs were eluted using a stepwise gradient of salts or pH buffers. Fractionated HcAbs were efficiently digested with trypsin or chymotrypsin to release Nb CDR peptides for identification and quantification by nanofluid chromatography coupled with high-resolution MS. Initial candidates from database searches were annotated for CDR identification. CDR3 fingerprints were filtered to remove false positives, their abundance from different biochemical fractions was quantified to infer Nb affinity, and they were assembled into Nb proteins—all of the above steps were automated by Augur Llama. This pipeline enables the identification and characterization of an unprecedented scale of diverse, specific, and high-quality Nb. Simultaneously, to enable structural analysis of tens of thousands of antigen-Nb interactions, a robust approach has been developed to integrate high-throughput computational docking (Schneidman-Duhovny, 2005), crosslinking, and mass spectrometry (CXMS) (Chait, 2016; Rout, 2019; Yu, 2018; Leitner, 2016) and mutagenesis. Furthermore, a deep learning approach was developed to learn latent features associated with the Nb library.

[0128] Example 3. Robust, in-depth and high-quality identification of antigen-specific Nb.

[0129] To validate this pipeline, three benchmark antigens were selected: glutathione S-transferase (GST), human serum albumin (HSA)—an important drug target (Larsen, 2016), and a small PDZ domain derived from mitochondrial outer membrane protein 25. These antigens span three orders of magnitude of immune responses, with PDZ exhibiting only weak immunogenicity (Figure 2B), and are ideal for assessing the robustness of our technology.

[0130] Here, 64,670 unique Nb species were identified. GST Sequences (9,915 unique CDR combinations from 3,453 CDR3 Nb families), 34,972 unique Nb sequences HSA(7,749 unique CDRs from 2,286 unique CDR3 Nb families) and 2,379 high-quality Nb PDZ A smaller cohort of sequences (495 unique CDRs from 230 CDR3 families) (Methods, Fig. 2C, 8G). Chymotrypsin has been shown to provide the most useful fingerprint information for identifying Nb from the various proteases tested (Fig. 2D, 2E). The Nb library exhibits unusual CDR3 diversity (Fig. 8D).

[0131] One group of 146 Nb were randomly selected from three antigen-specific Nb groups and expressed in *E. coli*. One group of 130 Nb (89%) showed excellent solubility and could be easily purified in large quantities (Fig. 2F). Complementary methods, including immunoprecipitation, ELISA, and SPR, were used to assess antigen binding (Methods, Fig. 2G, 9C, 9D, 10, Tables 1-3). The quality of the Nb identified by trypsin and chymotrypsin was quite high (Fig. 8H). GST, HSA, and PDZ confirmed 86.2% (CI). 95% 6.8% and 90.5% (CI) 95% The results demonstrate the high sensitivity and specificity of this method. (11.5%) and 100% genuine Nb binder.

[0132] Example 4. Accurate large-scale quantification and clustering of the Nb proteome.

[0133] Different strategies were evaluated for accurate affinity-based classification of Nb. Briefly, antigen-specific HcAbs were affinity-separated from serum and eluted via stepwise high-salt gradients, high-pH buffers, or low-pH buffers (Methods, Fig. 8I, 8J). Different HcAb fractions were accurately quantified by label-free quantitative proteomics (Zhu, 2010; Cox, J. and Mann, M, 2008). CDR3 peptides (and corresponding Nb) were then clustered into three groups based on their relative ionic strength (Fig. 3A, 3B, 9A, and 9B). This classification was performed using a high-pH method to separate 31% of Nb. GST and 47% of Nb HSA Assigned to the C3 high-affinity group (Fig. 3C). Numerous Nb molecules with unique CDR3 sequences from each cluster were randomly expressed. GST And through ELISA and SPR(R) 2=0.85 (Figure 3D, Table 1) to evaluate different fractional separation methods. While low-pH methods did not provide sufficient resolution to separate different affinity groups, salt gradients, particularly high-pH methods, enabled significant and reproducible Nb separation based on Nb affinity (Figure 3E). Nb from high-pH clusters 1 and 2 (C1, C2) typically exhibited low and intermediate affinities ranging from μM to tens of nM, respectively, while over 50% of C3 were ultra-high affinity, sub-nM binders (Figure 3H, 9D). To further validate these results, a random set of 25 Nb groups was purified from C3. HSA (They have different CDR3 values), and their ELISA affinities were ranked (Figure 3F, Table 2). The top 14 Nb were selected. HSA SPR measurements were performed, revealing 11 Nb molecules with affinities ranging from tens to hundreds of pM and exhibiting distinct binding kinetics. The remaining three Nb molecules... HSA Showing single-digit nM K D (Figures 3I and 10A). Thirteen soluble Nb molecules were purified. PDZ Their high affinity was confirmed by ELISA and immunoprecipitation (Figures 3G, 10B, and Table 3). Representative highly soluble Nb PDZ P10 of K D It is 4.4 pM (Figure 3J).

[0134] Natural mitochondrial immunoprecipitation (Nb GST ) and fluorescence imaging (Nb PDZ The high affinity of Nb for Nb (Figures 3K and 3L) was further positively evaluated. Quantitative methods enable large-scale and accurate classification of the Nb proteome based on desired properties (such as affinity).

[0135] Example 5. Landscape of the antigen-binding Nb proteome revealed by a comprehensive structural assay method.

[0136] The identification and classification of a large library of high-quality Nb allows for the study of the overall structural landscape of humoral immune responses to antigen binding. 34,972 Nb HSAStructural docking and clustering revealed three major HSA epitopes (Fig. 4A). The presence of abundant native serum albumin (76% identical to HSA, Fig. 12H) allowed for the study of the specificity of humoral immunity in camelids. The two albumin sequences were compared, and their variations were calculated based on pI and hydrophilicity (Methods, Fig. 4A). All three epitopes were co-localized with peaks of pI and hydrophilicity corresponding to large sequence differences. This result illustrates the specificity of Nb for antigen recognition. It appears that Nb preferentially binds to stable helical secondary structures (Fig. 4B). The epitopes were found to be highly charged. E2 and E3 were predominantly negative (net formal charge of -4 and -5, respectively, Fig. 13D), while E1 was more heterogeneous, with a mixed charge of -2 net formal charge (Fig. 4C).

[0137] Nineteen HSA-Nb complexes (Shi, 2014; Kim, 2018) were crosslinked to verify the epitopes identified via docking. Overall, the model satisfied 92% crosslinking, and the median RMSD of the model was [value missing]. (Figs. 4J, 4K). Crosslinking confirmed the docking results and identified two highly aggregated epitopes (E2, E3) (65% and 20%, respectively) (Fig. 4D, Table 2). E1 was identified by low-abundance (5%) crosslinking. Crosslinking also identified two additional minor epitopes not revealed by docking (Fig. 4D). High shape complementarity was observed between HSA and Nb, including convex Nb complementary sites and concave HSA epitopes (Figs. 4E-4G). To further confirm the major E2, we introduced a single-point mutation with minimal impact on the overall structure (Pires, 2016) on HSA, E400R. The resulting mutation reversed the surface charge to mimic the positive charge of the orthologous site in E2 of camel albumin, potentially disrupting the salt bridge formed between it and arginine in NbCDR3 (Fig. 4H). Nineteen high-affinity binders were then selected, and this point mutation on the HSA-Nb interaction was evaluated by ELISA (Fig. 4I, Table 2). E400R almost completely eliminated the binding of 5 out of the 19 Nb tested (26%), indicating that E2 is the true dominant epitope.

[0138] This method was further used to map epitopes of 64,670 GST-Nb complexes. Three major epitopes on GST were accurately identified (Figs. 11A, 11B, 11F, 11G) and verified by crosslinking, with relative abundances of 18.75%, 31.25%, and 50% for E1, E2, and E3, respectively (Figs. 11D, 11E). E1 and E3 contain negatively charged surface patches. E2 overlaps with the GST dimerization cavity (Fig. 11C); in the model presented in this paper, E2Nb inserts its CDR3 into this cavity. Similar to HSA, a preference for charged surface residues and high shape complementarity of Nb were confirmed. In summary, these results indicate that Nb can bind to diverse protein surfaces and has a preference for highly charged cavities on antigens.

[0139] Example 6. Exploring the mechanism of Nb affinity maturation.

[0140] Based on the most reliable high-pH dataset, the physicochemical and structural characteristics that distinguish high-affinity (mature) and low-affinity Nb were investigated. Shorter CDR3s exhibited different distributions of high-affinity binders for HSA and GST, respectively (Fig. 5A), thereby reducing the entropy of antigen binding. A significant increase in pI was observed (Fig. 5B), from the slightly acidic nature of low-affinity Nb to the relative basicity of high-affinity Nb.

[0141] The contributions of CDR to the pI and hydrophilicity of Nb were compared, and CDR3 was determined. HSA Mainly responsible for Nb HSA The polarity shift, and CDR1 GST and CDR2 GST Mainly responsible for Nb GST The polarity transition (Fig. 5C). A slightly higher hydrophilicity was observed in the high-affinity Nb (Fig. 5D).

[0142] The structure of CDR3 can be viewed as having a "head" region with the highest sequence variability and a "body" region with lower specificity (Finn, 2016) (Fig. 5E). Certain residues are enriched in the CDR3 head, including aspartic acid and arginine (forming strong electrostatic interactions) (Tiller, 2017), small and flexible residues of glycine and serine, hydrophobic residues such as alanine and leucine, and aromatic residues of tyrosine (Figs. 5F and 12). Comparison of Nb with different affinity groups revealed three main differences. First, high-affinity Nb is rich in charged residues (Mitchell, LS and Colwell, LJ, 2018) (Methods, Fig. 5G). Second, complex differences were identified for different antigens: high-affinity Nb… HSA Electrostatic enhancement tends to be achieved by increasing positively charged residues (39%) and decreasing (46%) negatively charged residues on the CDR3 head. High affinity Nb GSTThe main change was altering their charges on other CDRs. Positively charged residues increased by 29.2% and 117.2% on CDR1 and CDR2, respectively, while negatively charged residues decreased by 44.2% and 21.5%. This charge variation may increase the physicochemical complementarity between Nb and episites. Third, for high-affinity Nb... HSA Tyrosine (51%), glycine, and serine (58%) were more enriched at the CDR3 head. For high-affinity Nb GST Tyrosine levels increased in the CDR3 head (73%), but the percentages of glycine and serine were almost unaffected.

[0143] To further explore the putative role of these residues in enhancing HSA binding affinity, their localization frequencies along the CDR3 head were calculated (Fig. 5H). Tyrosine residues are more commonly found in high-affinity Nb residues. HSA The CDR3 head center allows its large aromatic side chains to insert into specific epitope pockets (Desmyter, 1996; Li, 2016). Glycine and serine tend to be located away from the CDR3 center, providing additional flexibility and facilitating the orientation of the tyrosine side chains within the antigen pocket. These results were confirmed by correlation analysis between the number of these residue groups and the ELISA affinity of our purified Nb (Figures 5I, 5J).

[0144] A deep learning model was developed to learn the latent features (method) that enable Nb affinity classification. The most information-rich Nb for classifying high-affinity binders... HSA The CDR3 filter revealed patterns of consecutive lysine and arginine, and tyrosine and glycine (Figure 5K, Table 4). For low-affinity binders, the most informative filter preferentially selected phenylalanine, histidine, and two consecutive aspartic acid pairs. Furthermore, this analysis revealed a trend towards consecutive negative and positive charge pairs for high-affinity and low-affinity binders, respectively.

[0145] Example 7. The excellent versatility and flexibility of Nb for antigen recognition.

[0146] Hundreds of different high-affinity Nb to the weakly immunogenic PDZ domain CDR3 The identification of the family has facilitated research into the structural basis of such interactions. Two putative epitopes were identified based on docking (Figs. 6A, 13B). E2 may be the major epitope because it has a large positively charged surface area (Figs. 6A, 6B) and it is more likely to have an α-helix and two β-chains. E2 overlaps with conserved ligand binding sites common to many PDZ-interacting proteins (Sheng, 2001; Doyle, 1996) (Fig. 6C). Notably, Nb PDZAffinities >100,000 times higher than those of natural PDZ ligands (in μM affinity) have been obtained (Niethammer, 1998) (Fig. 3J). Such high affinity is likely achieved through extensive electrostatic and hydrophobic interactions formed by long CDR3 rings wrapped around small, shallow epitopes (Figs. 6C, 13A). Modeling results indicate that R46 and K48 of the second β chain in the PDZ epitope are associated with Nb. PDZ The corresponding residues in the protein form a salt bridge. A double mutant PDZ (R46E:K48D) was generated, and its effect on Nb was assessed by ELISA. PDZ Affinity. Most (8 / 11) Nb PDZ The mutants exhibited significantly reduced affinity or no affinity, confirming that E2 is indeed the major epitope (Figure 6D).

[0147] Nb PDZ Several other observations were also made. First, the distribution of CDR3 ring lengths formed a main peak with a median of approximately 20 aa, exceeding the upper limit of its natural distribution (Fig. 6E). Second, Nb PDZ It is quite acidic, with a median pI of 4.9 (Fig. 6F), which is mainly contributed by CDR3 (Fig. 6E, 13F). Again, although Nb... PDZ It is acidic, but due to the compensation of hydrophobic residues, Nb PDZ There appears to be no significant change in hydrophilicity (Figs. 6G, 13E). Finally, negatively charged aspartic acid, glycine, and serine significantly increased, accounting for half of the CDR3 head residues; and high affinity for Nb... GST and Nb HSA In contrast, the reduction in large-volume tyrosine residues was also significant, reflecting the relatively shallow binding pocket of E2 (Figs. 7C, 7E). Overall, these results demonstrate the remarkable versatility of Nb in antigen binding.

[0148] This study reports the development of a robust platform integrating proteomics, informatics, and structural modeling techniques for analyzing the Nb proteome of antigen-binding antigens. This pipeline enables sensitive and reliable identification of a large number of high-quality Nb against diverse challenging antigens. It can also accurately classify circulating Nb based on their physicochemical properties. Our technique identified thousands of ultra-high affinity Nb. Combining computational docking and structural proteomics, this study characterized, mapped, and validated the major epitopes of 102,673 antigen-Nb complexes. This "big data" analysis allows for, for the first time, a global-scale proteomics and structural dissection of humoral immune responses.

[0149] These results reveal, with unprecedented depth, the efficiency, specificity, diversity, and versatility of antigen-binding Nb, which together shape the epic landscape of antibody immunity in camelids (Fig. 6H).

[0150] Efficiency: Nb binds efficiently using shape and electrostatic complementarity. Specific residues, such as charged aspartic and arginine, aromatic tyrosine, and small, flexible glycine and serine, allow for ring flexibility that generates high-affinity Nb. Complex and fine-tuned interactions specific to different CDRs were revealed. Furthermore, the presence of multiple dominant epitopes for Nb binding was confirmed, which could serve as a general mechanism for efficient pathogen recognition (Akram, A. and Inman, RD, 2012).

[0151] Specificity and diversity: Thousands of highly different Nb were discovered, which evolved to recognize specific HSA surface bags with some of the most significant sequence variations (Figure 4A) to ensure a specific, effective and safe immune response.

[0152] Multifunctionality: For antigens that tend to evade immune responses, such as PDZ, Nb can dramatically alter the size and physicochemical properties of the complementary site to mimic the binding of natural ligands with excellent affinity and specificity. The study demonstrates the fascinating and rapid evolution of protein-protein interactions.

[0153] Nb is highly effective in neutralizing viruses and inhibiting enzyme activity (Lauwereys, 1998; Desmyter, 1996; Acharya, 2013; Arabia, 2017). These findings suggest that these highly robust and efficient camel HcAbs are evolutionarily advantageous for their survival in arid natural habitats and the challenges of invasive pathogens, and the driving forces behind this incredible selection and adaptation remain a mystery (Flajnik, 2011).

[0154] These technologies have wide applications in challenging biomedical fields such as cancer biology, brain research, and virology. These informatics tools for Nb proteomics are freely available to the research community. High-quality Nb datasets can serve as blueprints for antibody-antigen studies and facilitate computational antibody design (Sircar, 2011; Baran, 2017; Chevalier, 2017).

[0155] Example 8. Method

[0156] Animal immunization. Two llamas were immunized with an initial dose of 1 mg HSA and a combination of GST and GST fusion PDZ domain of mitochondrial outer membrane protein 25 (OMP25), followed by three booster immunizations of 0.5 mg every 3 weeks. Hemorrhage and bone marrow aspirates were collected from the animals 10 days after the last booster immunization. All the above procedures were performed by Capralogics, Inc. in accordance with IACUC protocols.

[0157] mRNA isolation and cDNA preparation. Approximately 1-3 × 10⁻⁶ mRNA and cDNA were isolated from 350 ml of immune blood using a Ficoll gradient (Sigma). 9 5-9 × 10⁶ peripheral mononuclear cells were isolated from 30 ml of bone marrow aspirate. 7 Plasma cells. mRNA was isolated from each cell using an RNase kit (NEB) and Maxima was used. TM H Minus cDNA Synthesis Master Mix (Thermo) was used to reverse transcribe mRNA into cDNA. Cameloidea IgG heavy chain cDNA sequences (Abrabi, 1997) from the variable domain to the CH2 domain were specifically amplified using primers CALL001 (GTCCTGGCTGCTCTTCTACAAGG, SEQ ID NO: 2646) and CH2FORTA4 (CGCCATCAAGGTACCAGTTGA, SEQ ID NO: 2647). V lacking the CH1 domain... H The H gene was isolated from conventional IgG and purified by DNA gel electrophoresis (Qiagen), and subsequently re-amplified from frame 1 to frame 4 using a second forward (ATCTACACTCTTTCCCTACACGACGCTCTTCCGATCTNNNNNNNNATGGCT[C / G]A[G / T]GTGCAGCTGGTGGAGTCTGG, SEQ ID NO: 2648, where N represents A, T, C, or G) and a second reverse (GTGACTGGAGTTCAGACGTGTGCTCTTCCGATCTNNNNNNNNGGAGACGGTGACCTGGGT, SEQ ID NO: 2649, where N represents A, T, C, or G). Random octamer replacement adapter sequences were added to aid in cluster identification using Illumina MiSeq. The amplicons (approximately 450–500 bp) from the second PCR were purified using the Monarch PCR Cleanup Kit (NEB). A final round of PCR was performed using primers MiSeq-F (AATGATACGGCGACCACCGAGATCTACACTCTTTCCCTA, SEQ ID NO: 2650) and MiSeq-R (CAAGCAGAAGACGGCATACGAGATTTCTGAATGTGACTGGAGTTCA, SEQ ID NO: 2651) to add indexed P5 / P7 adaptors prior to MiSeq sequencing.

[0158] Next-generation sequencing using Illumina MiSeq. Sequencing is based on the Illumina MiSeq platform with a 300bp paired-end model. Each database generates over 30 million reads. Read QC tools in FastQC v0.11.8 (www.bioinformatics.babraham.ac.uk / projects / fastqc / ) are used for quality checking and control of FASTQ data. Raw Illumina reads are processed by software tools in the BBMap project (github.com / BioInfoTools / BBMap / ). Repeated reads and DNA barcode sequences are continuously removed before converting nucleotide sequences to amino acid sequences.

[0159] V was isolated from immune serum and separated by biochemical fractionation. H H antibody. Approximately 175 ml of plasma was isolated from 350 ml of immune blood using a Ficoll gradient (Sigma). Camelidae single-chain V antibody. H H antibodies were separated from plasma supernatant using a two-step purification procedure using Protein G and Protein A agarose beads (Marvelgent), followed by acid elution, neutralization, and dilution in 1×PBS buffer to a final concentration of 0.1–0.3 mg / ml. To purify antigen-specific V... H H antibody, GST or HSA conjugated with CNBr resin and V H The H mixture was incubated together at 4°C for 1 hour and then thoroughly washed with high-salt buffer (1×PBS and 350 mM NaCl) to remove non-specific binders. Specific V was then released from the resin using one of the following elution conditions. H H antibody: eluted with alkaline (1-100 mM NaOH, pH 11, 12, and 13), acidic (0.1 M glycine, pH 3, 2, and 1) or salt (1 M-4.5 M MgCl2 in neutral pH buffer). For purification of PDZ-specific V... H H, a fusion protein of MBP-PDZ was produced (in which maltose-binding protein / MBP was fused to the N-terminus of the PDZ domain to avoid steric hindrance of the small PDZ after coupling) and used as an affinity stem. MBP coupling resin was used as a control (Figure 6J). All eluted V H H was neutralized and dialyzed into 1×DPBS.

[0160] Proteolytic analysis of antigen-specific Nb and nanofluid chromatography (nLC / MS) coupled with mass spectrometry. For GST and HSA V... H H, process each elution according to the following protocol. For PDZ-specific V HH, only the most stringent biochemical eluates (i.e., pH 13, pH 1, MgCl2 3M, and 4.5M) and the corresponding nonspecific MBP binders from different fractions (negative controls) were pooled for protein hydrolysis. For example, for PDZ-specific V eluted with pH 13 buffer... H H, Nonspecific MBP-bound Nb was pooled from pH 11, pH 12, and pH 13 fractions for use as a negative control, to improve the stringency of our downstream LC / MS quantification. V H H was reduced at 57°C in 8M urea buffer (containing 50mM ammonium bicarbonate, 5mM TCEP, and DTT) for 1 hour, and then alkylated in the dark at room temperature with 30mM iodoacetamide for 30 minutes. The alkylated sample was then aliquoted and digested in solution with trypsin or chymotrypsin. For trypsin-digested samples, 1:100 (w / w) trypsin and Lys-C were added and digested overnight at 37°C, followed by a 4-hour incubation with 1:100 trypsin at 37°C the next morning. For chymotrypsin-digested samples, 1:50 (w / w) chymotrypsin was added and digested at 37°C for 4 hours. After protein hydrolysis, the peptide mixture was desalted using a self-filled stage-tips or Sep-pak C18 column (Waters) and treated with Q Exactive. TM HF-XHybrid Quadrupole Orbitrap TM Analysis was performed using a Thermo Fisher nano-LC 1200 mass spectrometer coupled online. In short, the desalted Nb peptide was loaded onto an analytical column (C18, 1.6 μm particle size). The instrument was run on an IonOpticks HPLC system with a pore size of 75 μm × 25 cm and eluted using a 90-minute liquid chromatography gradient (5% B - 7% B, 0–10 min; 7% B - 30% B, 10–69 min; 30% B - 100% B, 69–77 min; 100% B, 77–82 min; 100% B - 5% B, 82 min–82 min 10 s; 5% B, 82 min 10 s–90 min; mobile phase A consisted of 0.1% formic acid (FA), and mobile phase B consisted of 80% acetonitrile (ACN) containing 0.1% FA). The flow rate was 300 nl / min. The QE HF-X instrument was operated in data-dependent mode, where the top 12 most abundant ions (mass range 350–2,000, charge states 2–8) were fragmented by high-energy collisional dissociation (HCD). The target resolution for MS is 120,000, and the target resolution for tandem MS (MS / MS) analysis is 7,500. The quadrupole isolation window is 1.6Th, and the maximum injection time for MS / MS is set to 80ms.

[0161] Nb DNA synthesis and cloning. The Nb gene was codon-optimized for expression in E. coli, and the nucleotides were synthesized in vitro (Synbiotech). After verification by Sanger sequencing, the Nb gene was cloned into the pET-21b(+) vector at BamHI and XhoI (for GST Nb) or EcoRI and NotI restriction sites (for HSA and PDZ Nb).

[0162] Purification of recombinant proteins. DNA constructs were transformed into BL21(DE3) competent cells according to the manufacturer's instructions and plated overnight at 37°C with 50 μg / ml ampicillin on agar. Single colonies were inoculated into LB medium containing ampicillin and cultured overnight at 37°C. Cultures were then inoculated 1:100 (v / v) into fresh LB medium and shaken at 37°C until OD600nm reached 0.4–0.6. GST, GST-PDZ, and Nb were induced with 0.5 mM IPTG, while MBP and MBP-PDZ were induced with 0.1 mM IPTG. Induction was performed overnight at 16°C. Cells were then harvested, briefly sonicated, and lysed on ice with lysis buffer (1×PBS, 150 mM NaCl, 0.2% TX-100, and protease inhibitors). After lysis, soluble protein extracts were collected at 15,000 x g for 10 min. GST and GST-PDZ were purified using GSH resin and eluted with glutathione. MBP (maltose-binding protein) and MBP-PDZ fusion proteins were purified using amylose resin and eluted with maltose according to the manufacturer's instructions. Nb was purified using His-Cobalt resin and eluted with imidazole. The eluted proteins were then dialyzed in dialysis buffer (e.g., 1×DPBS, pH 7.4) and stored at -80°C before use.

[0163] Nb immunoprecipitation analysis. Following Nb induction and cell lysis, cell lysates were run on SDS-PAGE to estimate Nb expression levels. Recombinant Nb from cell lysates was diluted in 1×DPBS (pH 7.4) to final concentrations of approximately 5 μM (for GST Nb) and approximately 50 nM (for PDZ Nb). To test for specific interactions between Nb and antigens, different antigens were conjugated to CNBr resin. Inactivated or MBP-conjugated CNBr resin was used as a control. The antigen-conjugated resin or control resin was incubated with Nb lysates at 4°C for 30 min. The resin was then washed three times with washing buffer (1×DPBS containing 150 mM NaCl and 0.05% Tween 20) to remove non-specific binding. Specific antigen-bound Nb was then eluted from the resin with hot LDS buffer containing 20 mM DTT and run on SDS-PAGE. The intensity of Nb on the gel was compared between the antigen-specific signal and the control signal to identify false positive bindings.

[0164] ELISA (Enzyme-Linked Immunosorbent Assay). Indirect ELISA was performed to assess the camel-like immune response to the antigen and quantify the relative affinity of antigen-specific Nb. The antigen was spread onto 96-well ELISA plates (R&D system) overnight at 4°C in spread buffer (15 mM sodium carbonate, 35 mM sodium bicarbonate, pH 9.6) at a rate of approximately 1–10 ng per well. The wells were then blocked for 2 hours at room temperature with blocking buffer (DPBS, 0.05% Tween 20, 5% milk). To test the immune response, immune serum was serially diluted 5-fold in blocking buffer. The diluted serum was incubated with the antigen-coated wells at room temperature for 2 hours. HRP-bound anti-lamb Fc secondary antibody (Bethyl) was diluted 1:10,000 in blocking buffer and incubated with each well at room temperature for 1 hour. For Nb affinity testing, scrambled Nb that did not bind to the antigen of interest was used as a negative control. The Nb of the two specific binding agents used for testing and scrambling negative controls was serially diluted 10-fold from 10 μM to 1 pM in blocking buffer. HRP-binding secondary antibodies against the His-tag (Genscript) or T7-tag (Thermo) were diluted 1:5,000 or 1:10,000 in blocking buffer and incubated at room temperature for 1 hour. Three washes with 1×PBST (DPBS, 0.05% Tween 20) were performed to remove nonspecific absorbance between incubations. After the final wash, the sample was further incubated for 10 minutes in the dark at room temperature with freshly prepared w3,3′,5,5′-tetramethylbenzidine (TMB) substrate to generate a signal. After STOP solution (R&D system), the plate was read at multiple wavelengths (450 nm and 550 nm) on a plate reader (Multiskan GO, Thermo Fisher). A false-positive Nb binder is defined as one that meets either of the following two criteria: i) The ELISA signal is detectable only at a concentration of 10 μM, but not at a concentration of 1 μM. ii) At a concentration of 1 μM, a significant signal reduction (more than 10-fold) is detected compared to the signal at 10 μM, while the signal is undetectable at lower concentrations. Raw data were processed using Prism 7 (GraphPad) to fit 4PL curves and calculate logIC50.

[0165] Nb affinity was measured using surface plasmon resonance (SPR). SPR, a Biacore 3000 system (GE Healthcare), was used to measure Nb affinity. The antigen protein was immobilized on an activated CM5 sensor chip using the following steps: Protein analytes were diluted to 10–30 μg / ml in 10 mM sodium acetate, pH 4.5, and injected into the SPR system at a rate of 5 μl / min for 420 s. The sensor surface was then blocked with 1 M ethanolamine-HCl (pH 8.5). For each Nb analyte, a series of dilutions (spanning three orders of magnitude) were injected at a flow rate of 20–30 μl / min into HBS-EP+ running buffer (GE Healthcare) containing 2 mM DTT for 120–180 s, followed by a dissociation time of 5–20 min based on the dissociation rate. Between each injection, the sensor chip surface was regenerated with a low-pH buffer (pH 1.5–2.5) containing 10 mM glycine-HCl or a high-pH buffer (pH 12–13) containing 20–40 mM NaOH. Regeneration was performed for 30 seconds at a flow rate of 40–50 μl / min. The measurements were repeated, and only highly reproducible data were analyzed. The binding sensing map for each Nb was processed and analyzed using BIAevaluation by fitting a 1:1 Langmuir model or a 1:1 Langmuir model with mass transfer.

[0166] Crosslinking and mass spectrometry analysis of antigen-nanobody complexes. Prior to crosslinking, different Nb groups were incubated with equimolar concentrations of the antigen of interest in amine-free buffer (e.g., 1×DPBS and 2 mM DTT) at 4 °C for 1–2 h. Amine-specific disuccinimide octanoic acid (DSS) or heterobifunctional linker 1-ethyl-3-(3-dimethylaminopropyl)carbodiimide hydrochloride (EDC) were added to the antigen-Nb complex at final concentrations of 1 mM or 2 mM, respectively. For DSS crosslinking, the reaction was carried out at 23 °C with continuous stirring for 25 min. For EDC crosslinking, the reaction was carried out at 23 °C for 60 min. The reactants were quenched with 50 mM Tris-HCl (pH 8.0) for 10 min at room temperature. After protein reduction and alkylation, the crosslinked samples were separated by 4–12% SDS-PAGE gels (NuPAGE, Thermo Fisher). As previously described (Shi, 2014; Shi, 2015), regions corresponding to cross-linked species were cleaved and digested in a gel with trypsin and Lys-C. Following proteolysis, the peptide mixture was desalted and digested with Q Exactive. TM HF-X Hybrid Quadrupole-Orbitrap TMThe analysis was performed using a Thermo Fisher nano-LC 1200 mass spectrometer coupled with a Thermo Fisher microarray. The cross-linked peptides were loaded onto a picochip column (C18, 3 μm particle size). The pore size was 50 μm × 10.5 cm; New Objective), and elution was performed using a 60-minute LC gradient: 5% B - 8% B, 0–5 min; 8% B - 32% B, 5–45 min; 32% B - 100% B, 45–49 min; 100% B, 49–54 min; 100% B - 5% B, 54 min–54 min 10 s; 5% B, 54 min 10 s–60 min 10 s; Mobile phase A consisted of 0.1% formic acid (FA), and mobile phase B consisted of 80% acetonitrile containing 0.1% FA. The QE HF-X instrument was operated in data-dependent mode, where the top 8 most abundant ions (mass range 380–2,000, charge states 3–7) were fragmented by high-energy collisional dissociation (normalized collision energy 27). The target resolution for MS was 120,000, while the target resolution for MS / MS analysis was 15,000. The quadrupole isolation window was 1.8Th, and the maximum injection time for MS / MS was set to 120 ms. Following MS analysis, data was searched using pLink2 to identify cross-linked peptides (Chen, 2019). MS and MS / MS quality precisions were specified as 10 and 20 p.pm, respectively. Other search parameters included cysteine ​​ureomethylation as a fixed modification and methionine oxidation as a variable modification. A maximum of three trypsin deletion cleavage sites were allowed. Initial search results were obtained using a default 5% false discovery rate, estimated using a target-decoy search strategy. Cross-linking spectra were then manually examined to remove false positives, essentially as described previously (Shi, 2014; Kim, 2018; Shi, 2015).

[0167] Site-directed mutagenesis. Mammalian expression plasmids for HSA were obtained from Addgene. An E400R point mutation was introduced into the HSA sequence using primers HSA-F (GGTGTTCGACCGGTTCAAGCCTCTGG, SEQ ID NO: 2652) and HSA-R (TTGGCGTAGCACTCGTGA, SEQ ID NO: 2653) using the Q5 site-directed mutagenesis kit (NEB). After sequence verification via Sanger sequencing, plasmids carrying wild-type HSA and the mutant were transfected into HeLa cells using the Lipofectamine 3000 transfection kit (Thermo) and Opti-MEM (Gibco), according to the manufacturer's protocol. Cells were cultured overnight, and then the medium was replaced with DMEM without FBS supplementation to remove BSA. After 48 hours of culture at 37°C and 5% CO2, the HSA-expressing medium was collected and stored at -20°C. Protein expression was confirmed by SDS-PAGE and Western blotting analysis of the medium.

[0168] The PDZ domain (in the pGEX6p-1 vector) was obtained from General Biosystems. A two-point mutant of PDZ (i.e., R46E:K48D) was introduced using specific primers for PDZ-F (TGATGAAAATGGCGCAGCC, SEQ ID NO: 2654) and PDZ-R (ATTTCACTCACATAGATACCACTATCATTACTAACATAC, SEQ ID NO: 2655) via the Q5 site-directed mutagenesis kit. After verification by Sanger sequencing, the mutant vector was transformed into BL21(DE3) cells for expression. The GST fusion PDZ mutant protein was purified using GSH resin as previously described.

[0169] Fluorescence microscopy. COS-7 cells were plated on glass-bottomed culture dishes at 60-70% initial confluence and cultured overnight to allow cell attachment. Cells were incubated with MitoTracker Orange CMTMRos (1:4000) at 37°C for 30 minutes, washed once with PBS, and fixed for 10 minutes with pre-chilled methanol / ethanol (1:1). After washing with PBS, cells were blocked with 5% BSA for 1 hour. Then, Alexa Fluor... TM647-bound Nb (1:100) was added to cells and incubated at room temperature for 15 minutes. Two-color wide-field fluorescence images were acquired using our custom system on an Olympus IX71 inverted microscope frame with 561 nm and 642 nm excitation lasers (MPB Communications, Pointe-Claire, Quebec, Canada) and 100X oil immersion objectives (NA=1.4, UPLSAPO100XO; Olympus).

[0170] Text-based CDR (Complementary Defined Region) annotation. The CDR annotation method is modified from (Fridy, 2014). [*] indicates any residue.

[0171] CDR1 Notes: First, search for the short sequence motif “SC”, located between residues 20 and 26 of the Nb sequence. The start of the CDR1 sequence is defined as residue 5, followed by the “SC” motif. Once the first residue is identified, we then search for another sequence motif, “W[*]R”, located between residues 32 and 40 of the Nb sequence, and define the end of the CDR1 sequence as the first residue preceding the “W[*]R” motif.

[0172] CDR2 Note: The start of the CDR2 sequence is defined as the 14th residue, followed by the "W[*]R" motif. Once the first residue is identified, the motif "RF" located between residues 63 and 72 of Nb is then identified, and the end of the CDR2 sequence is defined as the 8th residue before the "RF" motif.

[0173] CDR3 Notes: First, search for the motif "Y[*]C" or "YY[*]", located between residues 90 and 105 of Nb. The start of the CDR3 sequence is defined as the 3rd residue, followed by the "Y[*]C" or "YY[*]" motif. Once the first residue of CDR3 is identified, any of the following sequence motifs ("WG[*]G", "WGQ[*]", "W[*]Q[*]", "[*]GQG", "[*][*]GQ", and "WG[*][*]") are then used to locate the end of CDR3. These motifs are located within the last 14 residues of the C-terminal Nb sequence. CDR3 ends one residue before the sequence motif. More information can be found in the Augur Llama script.

[0174] Cleavage rules of Nb digestion by different proteases in computer simulation:

[0175] Trypsin: C-terminus to K / R, without P.

[0176] Chymotrypsin: C-terminus to W / F / L / Y, without P following it.

[0177] GluC: C-terminus to D / E, without P-terminus.

[0178] AspN: N-terminus to D

[0179] LysC: C-terminus to K

[0180] Sequence alignment for the Nb database: Software ANACI was used (Dunbar, J. and Deane, CM, 2016). Three CDRs (CDR1-CDR3) and four frame sequences (FR1-FR4) were annotated according to the IMGT numbering scheme (Lefranc, 2003). Alignments below the threshold e-value of 100 were removed, and the remaining sequences were plotted by WebLogo (Crooks, 2004).

[0181] Computer-simulated digestion of the Nb database by different proteases and Nb CDR3 mapping analysis. Based on the aforementioned cleavage rules, computer-simulated digestion of a high-quality database containing approximately 500,000 unique Nb sequences was performed using different enzymes, including trypsin, chymotrypsin, LysC, GluC, and AspN. Peptides containing CDR3 were obtained to calculate sequence coverage. The CDR3 coverage was then summed to generate Figures 1D and 7B. The CDR3 peptide length distribution (via trypsin and chymotrypsin) was plotted to generate Figure 1E.

[0182] MS mapping assisted by trypsin and chymotrypsin to simulate Nb. 10,000 Nb sequences with unique CDR3 fingerprint sequences were randomly selected from a database. The selected Nb was then subjected to computer-simulated digestion (without allowing erroneous cleavage sites) using trypsin or chymotrypsin to generate CDR3 peptides. The following criteria were applied to these peptides to better simulate Nb identification by MS: 1) Peptides of a size suitable for bottom-up proteomics (between 850 and 3,000 Da) were first selected. 2) Peptides containing the highly conserved C-terminal FR4 motif WGQGQVTS were further discarded. Based on our observations, such peptides are generally predominantly fragmented with C-terminal γ ions, while fragmentation on the CDR3 sequence is poor, which is crucial for definitive CDR3 peptide identification. 3) CDR3 peptides with limited Nb fingerprint information (containing less than 30% CDR3 sequence coverage) were removed. As a result, 2,111 unique trypsin peptides and 5,154 unique chymotrypsin peptides were obtained. These peptides were then used to map the Nb protein. After protein assembly, only Nb identification with sufficiently high CDR3 fingerprint sequence coverage (≥60%) was used to generate the Venn diagram in Figure 1F.

[0183] Phylogenetic analysis of Nb CDR3 sequences. Phylogenetic trees were generated by Clustal Omega (Sievers, 2014) using unique Nb CDR3 sequences as input and additional side-joint sequences (i.e., YYCAA to the N-terminus of the CDR3 sequence and WGQG to the C-terminus of the sequence) to assist alignment. Data were plotted by ITol (Interactive Tree of Life) (Letunic, I. and Bork, P., 2007). The isoelectric point and hydrophobicity of Nb CDR3 were calculated using the BioPython library. Sequence alignments were visualized using Jalview (Waterhouse, 2009).

[0184] The reproducibility of Nb peptide quantification was assessed. Peptide identification shared between different LC runs was used to assess the reproducibility of the label-free quantification method. For a typical 90-minute LC gradient, the peptide peak width or full width at half maximum (FWHM) is typically less than 5 seconds. Differences in peptide retention times between different LC runs were calculated to generate the nuclear density estimation plot in Figure 3B. Peptide retention times from different LC runs were used to calculate Pearson correlations and plotted in Figure 9B.

[0185] Sequence alignment and analysis of HSA and llama serum albumin. Llama (Camelus Ferus) serum albumin sequences were extracted using tblastn (NCBI) and aligned with HSA. The isoelectric point (pI) and hydrophilicity values ​​of individual amino acids were obtained online from (www.peptide2.com / N_peptide_hydrophobicity_hydrophilicity.php). These values ​​were normalized between 0 and 1.0, and the sequence variation (pairwise difference in pI and hydrophilicity) between the two albumins was calculated for each alignment position. For a specific alignment residue position, a value of 0 indicates the presence of identical residues in both sequences, while 1.0 indicates the largest sequence variation, such as a charge reversal from the negatively charged residue glutamate 400 in HSA to the positively charged residue arginine at the corresponding alignment position in the llama albumin. A value of 0.5 was assigned to positions identifying amino acid insertions or deletions. The sequence variations in pI and hydrophilicity between HSA and llama serum albumin were thus plotted. The graph was further smoothed using a Gaussian function to generate Figure 4A.

[0186] The relative abundance of amino acids on the Nb CDRs was analyzed. The amino acid frequencies of each CDR (including the CDR1, CDR2, and CDR3 heads) were calculated and normalized to generate the bar and pie charts in Figures 6, 7, 12, and 13. The CDR3 head sequence was obtained by removing the four semi-conserved C-terminal residues of CDR3. The CDR residue frequencies of both high-affinity and low-affinity Nb were normalized based on the sum of the CDR residues for each affinity group.

[0187] The amino acid positions at the CDR3 head were analyzed. The relative positions of the residues at the CDR3 head were calculated, where a value of 0 represents the N-terminus of the CDR3 head, and 1.0 represents the last residue. The CDR3 head sequence was then divided into 20 bins, each 0.05 pixels wide. Within each bin, the occurrence of specific types of amino acids (e.g., tyrosine, glycine, or serine) was counted and normalized to the sum of the residues at the CDR3 head. The distribution of different amino acids, including their relative positions and abundance, is plotted in Figures 5H and 12G.

[0188] Proteomics database search for candidate Nb peptides. Raw MS data were searched against an internally generated Nb sequence database using Sequest HT embedded in ProteomeDiscoverer 2.1 (Thermo Fisher), with FDR estimated using a standard target-decoy strategy. The quality precision for MS1 and MS2 was specified as 10 ppm and 0.02 Da, respectively. Other search parameters included cysteine ​​ureomethylation as a fixed modification and methionine oxidation as a variable modification. For trypsin and chymotrypsin-treated samples, a maximum of one or two missing cleavage sites were allowed, respectively. Initial search results were filtered based on q-values ​​(Kall, 2007) with an FDR of 0.01 (strict). After the database search, peptide matching (PSM) was exported, processed, and analyzed by Augur Llama, as follows:

[0189] a. Nanobody identification

[0190] i) Quality assessment of CDR3 fingerprints

[0191] Candidate peptides were first annotated as CDR or FR peptides. To definitively identify CDR3 fingerprint peptides, we implemented a filter / algorithm that required sufficient coverage of high-resolution CDR3 fragment ions in the PSM (see illustration in Figure 8B). The filter was evaluated using a target sequence database containing approximately 500,000 unique Nb sequences and a similarly sized non-overlapping decoy database. The target and decoy Nb sequence databases used in this paper were obtained from different llamas. Any peptide identification from the decoy database was considered a false positive. The FDR was defined as the percentage of peptide identification in the decoy database compared to that in the target database. CDR3 length was also considered capable of developing sensitive CDR3 peptide filters. CDR3 fragmentation coverage was defined as the percentage of CDR3 residues that matched the fragment ion (b or y ion) within a quality precision window. Spectra of the same peptide were merged for evaluation. Only CDR3 peptides that passed this filter (5% FDR) were selected for downstream Nb assembly.

[0192] ii) Nanobody sequence assembly

[0193] CDR peptides, including the confirmed CDR3 peptide, are used for Nb protein assembly. Two additional criteria must be met before Nb can be identified. These criteria include: 1) Both CDR1 and CDR2 peptides must be usable for Nb assembly. 2) For any Nb identification, at least 50% combined CDR coverage is required.

[0194] b. Quantification and classification of antigen-specific Nb libraries

[0195] Raw MS data were accessed using MSFileReader 3.1 SP4 (ThermoFisher) and the Python library pymsfilereader (github.com / frallain / pymsfilereader). Reliable CDR3 peptides passing through a quality filter were quantified by label-free LC / MS.

[0196] i) Quantitative analysis of CDR3 peptide

[0197] To achieve accurate label-free quantification of CDR3 peptides across different LC runs, different retention time windows were specified for peptide peak extraction. For peptides directly identifiable by search engines based on MS / MS spectroscopy, a small quantization window with a retention time (RT) offset of + / - 0.5 minutes was used for peak extraction. For peptides not directly identified from a particular LC run (due to peptide complexity and random ion sampling), their RTs were predicted based on the RTs of adjacent LC runs and adjusted using the median RT difference of peptides co-identified between the two LC runs. In this case, a relaxed RT window of + / - 2.0 minutes (for a typical 90-minute LC gradient) was applied, where approximately 95% of all identified peptides can be matched between two LC runs to facilitate peak extraction. Both m / z and z of the peptides were used for peak extraction, with a quality precision window of + / - 10 ppm. Peptide peaks were extracted and smoothed using a Gaussian function. Their AUCs (area under the curve) were calculated and averaged across repeated LC runs to infer the CDR3 peptide intensity.

[0198] ii) Classification of Nb

[0199] To achieve accurate classification, for example based on Nb affinity, the relative ionic strength (AUC) of CDR3 fingerprint peptides in three biochemically fractionated Nb samples (F1, F2, and F3) was quantified as I1, I2, and I3. Based on the quantification results, the CDR3 peptides were arbitrarily divided into three clusters (C1, C2, and C3) using the following criteria:

[0200] 1) For the C3 (high affinity) cluster: I3 > I1 + I2 (indicating that Nb is more specific to F3)

[0201] 2) For the C2 (medium affinity) cluster: I2 > I1 + I3 (indicating that Nb is more specific to F2)

[0202] 3) For the C1 (low affinity) cluster:

[0203] I1 > I2 + I3 (indicating that Nb is more specific to F1 or may be a non-specific binder), or, if I1 < I2 + I3 and I2 < I1 + I3 and I3 < I1 + I2, then these Nb identifications may be non-specific and are also classified as C1. See the explanation in Figure 8C.

[0204] The above method was used to classify HSA and GST Nb. Several modifications were made to the quantification and characterization of high-affinity PDZ Nb. Specifically, an additional control, “F_control” (ionic strength as I_control), was added for MBP-interacting Nb for quantification. A high-affinity cluster Nb (represented by its unique CDR3 peptide) was defined when the total I2 and I3 intensities of the Nb CDR3 peptide were 20 times higher than the I_control (i.e., 20 * I_control < I2 + I3). For Nb quantified using more than one unique CDR3 peptide, the classification results between different CDR3 peptides from the same Nb must be consistent; otherwise, they are removed before reporting the final results.

[0205] Heatmap analysis of the relative intensities of CDR3 peptides. The identified CDR3 peptides were quantified based on their relative MS1 ionic intensities and subsequently clustered using a script in Augur Llama. Z-scores were calculated based on the relative ionic intensities and used to generate the heatmap in Figure 3A for visualization.

[0206] Structural modeling of the antigen-Nb complex. Structural models of Nb were obtained using the MODELLER multi-template comparison modeling protocol (Webb, B. and Sali, A., 2014). Next, the CDR 3-loop was refined, and the top 5 scoring loop conformations were selected for downstream docking. Each Nb model was then docked to the corresponding antigen using the antibody-antigen docking protocol in PatchDock software, which prioritizes the CDR search (Schneidman-Duhovny, 2005). The models were then re-scored using statistical potential SOAP (Dong, 2013). Based on the SOAP score, antigen interface residues from the 10 best-scoring models were used. Epitopes were identified. Once epitopes were defined, Nb was clustered based on epitope similarity using k-means clustering. The clusters revealed the most immunogenic surface patches on the antigen. Antigen-Nb complexes with CXMS data were modeled using the distance-constrained PatchDock protocol, which optimized constraint satisfaction (Schneidman-Duhovny, 2020; Russel, 2012). If the Ca-Ca distance between the crosslinking residues of the DSS and EDC crosslinkers is respectively within... and If the crosslinking is within a certain range, the constraint is considered satisfied (Shi, 2014; Fernandez-Martinez, 2016). In the case of ambiguous constraints, such as GST dimers, one of the crosslinkings must be satisfied.

[0207] Machine learning analysis of the Nb library. A deep neural network was trained to distinguish between low-affinity and high-affinity Nb, characterized by accurate high-pH fractionation and quantitative proteomics. The model consists of a convolutional layer with batch normalization and ReLU activation, followed by a max-pooling layer ending with a fully connected layer to integrate extracted features into a logits layer that leads to classifier predictions. The convolutional layer consists of 20 1D filters, representing local receptive fields of 7 amino acids each, long enough to capture relevant CDRs and short enough to avoid overfitting the data. During forward propagation, each filter slides along the protein sequence with a fixed stride, performs element-wise multiplication with the current sequence window, and then adds them to generate the filter response. The model achieved a classification accuracy of 92%.

[0208] To understand the physicochemical characteristics learned by the network to distinguish between low-affinity and high-affinity binders, the activation path from prediction back to the activated filters is computed through the network. Similar to the backpropagation algorithm, a backward iteration is performed from the last two layers of the fully connected network, extracting the output signal for each sequence and finding the highest peak with the largest weight contributing to classification. In the same way, the contribution of each upstream filter to these peaks is computed. Additionally, filter activity in the CDR is analyzed to extract region-specific master filters. This network interpretation process yields the unique contribution of each filter for each sequence. Each filter is activated along the sequence downsampled in the max-pooling layer. For each filter, its highest peak is then selected for classification. Finally, the filter with the largest contribution in each sequence is determined, and we also obtain an interesting filter that contributes more than 30% in these regions of interest.

[0209] Computer implementation method

[0210] It should be understood that the logical operations described in this paper relative to the figures can be implemented as (1) operating on a computing device (e.g., Figure 14 The logical operations described herein are (1) a sequence of computer-implemented actions or program modules (i.e., software) on a computing device described herein, (2) interconnected machine logic circuits or circuit modules (i.e., hardware) within the computing device, and / or (3) a combination of the software and hardware of the computing device. Therefore, the logical operations discussed herein are not limited to any particular combination of hardware and software. Implementation is a matter of choice depending on the performance of the computing device and other requirements. Therefore, the logical operations described herein are referred to in various ways as operations, structural devices, actions, or modules. These operations, structural devices, actions, and modules can be implemented in software, firmware, special-purpose digital logic, and any combination thereof. It should also be understood that more or fewer operations than those shown in the figures and described herein may be performed. These operations may also be performed in a different order than those described herein.

[0211] See Figure 14An exemplary computing device 500 is shown on which the methods described herein can be implemented. It should be understood that the exemplary computing device 500 is merely one example of a suitable computing environment on which the methods described herein can be implemented. Optionally, the computing device 500 may be a well-known computing system, including but not limited to personal computers, servers, handheld or laptop devices, multiprocessor systems, microprocessor-based systems, networked personal computers (PCs), minicomputers, mainframe computers, embedded systems, and / or distributed computing environments, including any of the above systems or devices. A distributed computing environment enables remote computing devices connected to communication networks or other data transmission media to perform various tasks. In a distributed computing environment, program modules, applications, and other data may be stored on local and / or remote computer storage media.

[0212] In the most basic configuration of the computing device 500, the computing device typically includes at least one processing unit 506 and a system memory 504. Depending on the exact configuration and type of the computing device, the system memory 504 may be volatile (such as random access memory (RAM)), non-volatile (such as read-only memory (ROM), flash memory, etc.), or a combination of both. This most basic configuration in... Figure 14 The image is shown as dashed line 502. Processing unit 506 may be a standard programmable processor that performs arithmetic and logical operations necessary for operating the computing device 500. The computing device 500 may also include a bus or other communication mechanism for transmitting information between various components of the computing device 500.

[0213] The computing device 500 may have additional features / functions. For example, the computing device 500 may include additional storage devices, such as removable storage device 508 and non-removable storage device 510, including but not limited to magnetic or optical disks or tapes. The computing device 500 may also include a network connection 516 that allows the device to communicate with other devices. The computing device 500 may also have input devices 514, such as a keyboard, mouse, touch screen, etc. It may also include output devices 512, such as a display, speaker, printer, etc. Additional devices may be connected to a bus to facilitate data communication between components of the computing device 500. All of these devices are well known in the art and need not be described in detail herein.

[0214] Processing unit 506 can be configured to execute program code encoded in a tangible computer-readable medium. A tangible computer-readable medium is any medium capable of providing data that causes computing device 500 (i.e., the machine) to operate in a particular manner. Various computer-readable media can be used to provide instructions to processing unit 506 for execution. Exemplary tangible computer-readable media may include, but are not limited to, volatile media, non-volatile media, removable media, and non-removable media implemented in any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. System memory 504, removable storage device 508, and non-removable storage device 510 are all examples of tangible computer storage media. Exemplary tangible computer-readable recording media include, but are not limited to, integrated circuits (e.g., field-programmable gate arrays or application-specific integrated circuits), hard disks, optical disks, magneto-optical disks, floppy disks, magnetic tapes, holographic storage media, solid-state devices, RAM, ROM, electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROMs, digital versatile discs (DVDs) or other optical storage devices, magnetic cassettes, magnetic tapes, disk storage devices, or other magnetic storage devices.

[0215] In an example implementation, processing unit 506 may execute program code stored in system memory 504. For example, a bus may carry data to system memory 504, and processing unit 506 may receive and execute instructions from said system memory. Data received from system memory 504 may optionally be stored on removable storage device 508 or non-removable storage device 510 before or after execution by processing unit 506.

[0216] It should be understood that the various techniques described herein may be implemented in combination with hardware or software, or in combination thereof, where appropriate. The methods and apparatuses, or certain aspects or portions thereof, of the subject matter disclosed herein may take the form of program code (i.e., instructions) implemented in a tangible medium such as a floppy disk, CD-ROM, hard disk drive, or any other machine-readable storage medium, wherein when the program code is loaded into a machine (such as a computing device) and executed by said machine, said machine becomes an apparatus for practicing the subject matter disclosed herein. Where the program code executes on a programmable computer, the computing device typically includes a processor, a storage medium readable by the processor (including volatile and non-volatile memory and / or storage elements), at least one input device, and at least one output device. One or more programs may implement or utilize the processes described in connection with the subject matter disclosed herein, for example, by using an application programming interface (API), reusable controls, etc. Such programs may be implemented in high-level procedural or object-oriented programming languages ​​to communicate with a computer system. However, programs may be implemented in assembly or machine language, if desired. In any case, the language may be a compiled or interpreted language, and it may be combined with hardware implementations.

[0217] As described above, the logical operations described herein, such as those described in Embodiment 8, can be implemented in hardware, software, or, where appropriate, a combination thereof. For example, they can be implemented using, for instance... Figure 14 The logical operations are implemented using one or more computing devices of the computing device 500. The logical operations described in Example 8 include, but are not limited to, methods for determining the antigen affinity of nanobody peptide sequences, methods for training deep learning models, and deep learning-based methods for inferring the antigen affinity of nanobody peptide sequences. These operations have been described in detail above.

[0218] In some implementations, the computer-implemented method includes:

[0219] Receive nanobody peptide sequences;

[0220] Identification of multiple CDR regions of the nanobody peptide sequence, including the CDR3 region:

[0221] A fragmentation filter is applied to discard one or more false-positive CDR3 regions of the nanobody peptide sequence;

[0222] Quantitative analysis of the abundance of one or more non-discarded CDR3 regions of nanobody peptide sequences; and

[0223] Antigen affinity is inferred based on the quantitative abundance of one or more non-discarded CDR3 regions of the nanobody peptide sequence.

[0224] In some implementations, a method for training a deep learning model includes:

[0225] Create a dataset containing multiple nanobody peptide sequences and corresponding antigen affinity tags; and

[0226] The dataset was used to train a deep learning model to classify nanobody peptide sequences with low antigen affinity and nanobody peptide sequences with high antigen affinity.

[0227] In some implementations, methods for determining the antigen affinity of nanobody peptide sequences include:

[0228] Receive nanobody peptide sequences;

[0229] The nanobody peptide sequence is input into a trained deep learning model; and

[0230] The trained deep learning model is used to classify the nanobody peptide sequences as having low antigen affinity or high antigen affinity.

[0231] Table 1. Summary of GST Nb and its biophysical and physicochemical properties

[0232]

[0233]

[0234]

[0235]

[0236]

[0237]

[0238]

[0239]

[0240]

[0241]

[0242]

[0243]

[0244]

[0245]

[0246]

[0247]

[0248]

[0249]

[0250]

[0251]

[0252]

[0253]

[0254]

[0255]

[0256]

[0257]

[0258]

[0259]

[0260]

[0261]

[0262]

[0263]

[0264]

[0265]

[0266]

[0267]

[0268]

[0269] Table 2. Summary of HSA Nb and its biophysical and physicochemical properties.

[0270]

[0271]

[0272]

[0273]

[0274]

[0275]

[0276]

[0277]

[0278]

[0279]

[0280]

[0281]

[0282]

[0283]

[0284]

[0285]

[0286]

[0287]

[0288]

[0289]

[0290] Table 3. Summary of PDZ Nb and its biophysical and physicochemical properties

[0291]

[0292]

[0293]

[0294] Table 4. GST Summary: Amino Acid Sequence Filters Derived from Deep Learning Methods

[0295]

[0296] Table 5. HSA Summary: Amino Acid Sequence Filters Derived from Deep Learning Methods

[0297]

[0298] References

[0299] 1. Muyldermans, S. Nanobodies: natural single-domain antibodies. Annu Rev Biochem 82, 775-797(2013).

[0300] 2. Beghein, E.&Gettemans, J. Nanobody Technology: A Versatile Toolkit for Microscopic Imaging, Protein-Protein Interaction Analysis, and Protein Function Exploration. Front Immunol 8, 771(2017).

[0301] 3. Rasmussen, S.G. et al. Structure of a nanobody-stabilized active state of the beta(2)adrenoceptor. Nature 469, 175-180(2011).

[0302] 4. Jovcevska, I.&Muyldermans, S. The Therapeutic Potential of Nanobodies. BioDrugs 34, 11-26(2020).

[0303] 5. Lauwereys, M. et al. Potent enzyme inhibitors derived from dromedary heavy-chain antibodies. The EMBO journal 17, 3512-3520(1998).

[0304] 6. Pardon, E. et al. A general protocol for the generation of Nanobodies for structural biology. Nature protocols 9, 674-693(2014).

[0305] 7.McMahon,C.et al.Yeast surface display platform for rapid discoveryof conformationally selective nanobodies.Nature structural&molecular biology25,289-296(2018).

[0306] 8.Egloff,P.et al.Engineered peptide barcodes for in-depth analyses ofbinding protein libraries.Nature methods 16,421-428(2019).

[0307] 9.Fridy,P.C.et al.A robust pipeline for rapid production of versatilenanobody repertoires.Nature methods 11,1253-1260(2014).

[0308] 10.Savitski,M.M.,Wilhelm,M.,Hahne,H.,Kuster,B.&Bantscheff,M.AScalable Approach for Protein False Discovery Rate Estimation in LargeProteomic Data Sets.Molecular&cellular proteomics:MCP 14,2394-2404(2015).

[0309] 11.DeKosky,B.J.et al.High-throughput sequencing of the paired humanimmunoglobulin heavy and light chain repertoire.Nature biotechnology 31,166-169(2013).

[0310] 12.Elias,J.E.&Gygi,S.P.Target-decoy search strategy for increasedconfidence in large-scale protein identifications by mass spectrometry.Naturemethods 4,207-214(2007).

[0311] 13.Schneidman-Duhovny,D.,Inbar,Y.,Nussinov,R.&Wolfson,H.J.PatchDockand SymmDock:servers for rigid and symmetric docking.Nucleic acids research33,W363-W367(2005).

[0312] 14.Chait,B.T.,Cadene,M.,Olinares,P.D.,Rout,M.P.&Shi,Y.RevealingHigher Order Protein Structure Using Mass Spectrometry.Journal of theAmerican Society for Mass Spectrometry 27,952-965(2016).

[0313] 15.Rout,M.P.&Sali,A.Principles for Integrative Structural BiologyStudies.Cell 177,1384-1403(2019).

[0314] 16.Yu,C.&Huang,L.Cross-Linking Mass Spectrometry:An EmergingTechnology for Interactomics and Structural Biology.Analytical Chemistry 90,144-165(2018).

[0315] 17.Leitner,A.,Faini,M.,Stengel,F.&Aebersold,R.Crosslinking and MassSpectrometry:An Integrated Technology to Understand the Structure andFunction of Molecular Machines.Trends in biochemical sciences 41,20-32(2016).

[0316] 18.Larsen,M.T.,Kuhlmann,M.,Hvam,M.L.&Howard,K.A.Albumin-based drugdelivery:harnessing nature to cure disease.Mol Cell Ther 4,3(2016).

[0317] 19.Zhu,W.H.,Smith,J.W.&Huang,C.M.Mass Spectrometry-Based Label-FreeQuantitative Proteomics.J Biomed Biotechnol(2010).

[0318] 20.Cox,J.&Mann,M.MaxQuant enables high peptide identification rates,individualized p.p.b.-range mass accuracies and proteome-wide proteinquantification.

[0319] Nature biotechnology 26,1367-1372(2008).

[0320] 21.Shi,Y.et al.Structural characterization by cross-linking revealsthe detailed architecture of a coatomer-related heptameric module from thenuclear pore complex.Molecular&cellular proteomics:MCP 13,2927-2943(2014).

[0321] 22.Kim,S.J.et al.Integrative structure and functional anatomy of anuclear pore complex.Nature 555,475-482(2018).

[0322] 23.Pires,D.E.V.,Ascher,D.B.&Blundell,T.L.mCSM:predicting the effectsof mutations in proteins using graph-based signatures.Bioinformatics(Oxford,England)30,335-342(2014).

[0323] 24.Finn,J.A.et al.Improving Loop Modeling of the AntibodyComplementarity-Determining Region 3 Using Knowledge-Based Restraints.PloSone 11,e0154811(2016).

[0324] 25.Tiller,K.E.et al.Arginine mutations in antibody complementarity-determining regions display context-dependent affinity / specificity trade-offs.The Journal of biological chemistry 292,16638-16652(2017).

[0325] 26.Mitchell,L.S.&Colwell,L.J.Analysis of nanobody paratopes revealsgreater diversity than classical antibodies.Protein Eng Des Sel 31,267-275(2018).

[0326] 27.Desmyter,A.et al.Crystal structure of a camel single-domain VHantibody fragment in complex with lysozyme.Nat Struct Biol 3,803-811(1996).

[0327] 28.Li,T.et al.Immuno-targeting the multifunctional CD38 usingnanobody.Scientific reports 6(2016).

[0328] 29.Sheng,M.&Sala,C.PDZ domains and the organization of supramolecularcomplexes.Annu Rev Neurosci 24,1-29(2001).

[0329] 30.Doyle,D.A.et al.Crystal structures of a complexed and peptide-freemembrane protein-binding domain:Molecular basis of peptide recognition byPDZ.Cell 85,1067-1076(1996).

[0330] 31.Niethammer,M.et al.CRIPT,a novel postsynaptic protein that bindsto the third PDZ domain of PSD-95 / SAP90.Neuron 20,693-707(1998).

[0331] 32.Akram,A.&Inman,R.D.Immunodominance:A pivotal principle in hostresponse to viral infections.Clin Immunol 143,99-115(2012).

[0332] 33.Bar-On,Y.M.,Phillips,R.&Milo,R.The biomass distribution onEarth.Proceedings of the National Academy of Sciences of the United States ofAmerica 115,6506-6511(2018).

[0333] 34.Chaplin,D.D.Overview of the immune response.J Allergy Clin Immun125,S3-S23(2010).

[0334] 35.Acharya,P.et al.Heavy chain-only IgG2b llama antibody effectsnear-pan HIV-1 neutralization by recognizing a CD4-induced epitope thatincludes elements of coreceptor-and CD4-binding sites.J Virol 87,10173-10181(2013).

[0335] 36.Arabi,Y.M.et al.Middle East Respiratory Syndrome.New Engl J Med376,584-594(2017).

[0336] 37.Flajnik,M.F.,Deschacht,N.&Muyldermans,S.A Case Of Convergence:WhyDid a Simple Alternative to Canonical Antibodies Arise in Sharks and Camels?PLoS biology 9(2011).

[0337] 38.Sircar,A.,Sanni,K.A.,Shi,J.&Gray,J.J.Analysis and modeling of thevariable region of camelid single-domain antibodies.J Immunol 186,6357-6367(2011).

[0338] 39.Baran,D.et al.Principles for computational design of bindingantibodies.Proceedings of the National Academy of Sciences of the UnitedStates of America 114,10900-10905(2017).

[0339] 40.Chevalier,A.et al.Massively parallel de novo protein design fortargeted therapeutics.Nature 550,74-79(2017).

[0340] 41.Arbabi Ghahroudi,M.,Desmyter,A.,Wyns,L.,Hamers,R.&Muyldermans,S.Selection and identification of single domain antibody fragments from camelheavy-chain antibodies.FEBS letters 414,521-526(1997).

[0341] 42.Shi,Y.et al.A strategy for dissecting the architectures of nativemacromolecular assemblies.Nature methods 12,1135-1138(2015).

[0342] 43.Chen,Z.L.et al.A high-speed search engine pLink 2 with systematicevaluation for proteome-scale identification of cross-linked peptides.Naturecommunications 10,3404(2019).

[0343] 44.Dunbar,J.&Deane,C.M.ANARCI:antigen receptor numbering and receptorclassification.Bioinformatics(Oxford,England)32,298-300(2016).

[0344] 45.Lefranc,M.P.et al.IMGT unique numbering for immunoglobulin and Tcell receptor variable domains and Ig superfamily V-like domains.Dev CompImmunol 27,55-77(2003).

[0345] 46.Crooks,G.E.,Hon,G.,Chandonia,J.M.&Brenner,S.E.WebLogo:a sequencelogo generator.Genome research 14,1188-1190(2004).

[0346] 47.Sievers,F.&Higgins,D.G.Clustal Omega,accurate alignment of verylarge numbers of sequences.Methods in molecular biology 1079,105-116(2014).

[0347] 48.Letunic,I.&Bork,P.Interactive Tree Of Life(iTOL):an online toolfor phylogenetic tree display and annotation.Bioinformatics(Oxford,England)23,127-128(2007).

[0348] 49.Waterhouse,A.M.,Procter,J.B.,Martin,D.M.,Clamp,M.&Barton,G.J.Jalview Version 2--a multiple sequence alignment editor and analysisworkbench.Bioinformatics(Oxford,England)25,1189-1191(2009).

[0349] 50.Kall,L.,Canterbury,J.D.,Weston,J.,Noble,W.S.&MacCoss,M.J.Semi-supervised learning for peptide identification from shotgun proteomicsdatasets.Nature methods 4,923-925(2007).

[0350] 51.Webb,B.&Sali,A.Comparative Protein Structure Modeling UsingMODELLER.Curr Protoc Bioinformatics 47,561-32(2014).

[0351] 52.Dong,G.Q.,Fan,H.,Schneidman-Duhovny,D.,Webb,B.&Sali,A.Optimizedatomic statistical potentials:assessment of protein interfaces andloops.Bioinformatics(Oxford,England)29,3158-3166(2013).

[0352] 53.Schneidman-Duhovny,D.&Wolfson,H.J.Modeling of MultimolecularComplexes.Methods in molecular biology 2112,163-174(2020).

[0353] 54.Russel,D.et al.Putting the pieces together:integrative modelingplatform software for structure determination of macromolecularassemblies.PLoS biology 10,e1001244(2012).

[0354] 55.Fernandez-Martinez,J.et al.Structure and Function of the NuclearPore Complex Cytoplasmic mRNA Export Platform.Cell 167,1215-1228 e1225(2016).

Claims

1. A method for identifying a set of complementarity-determining regions (CDR3), CDR2, and / or CDR1) amino acid sequences of nanobody, wherein a reduction in the number of said CDR3, CDR2, and / or CDR1 sequences compared to a control is a false positive, the method comprising: a. Obtaining blood samples from camels immunized with antigens; b. Obtain a nanobody cDNA library using the blood sample; c. Identify the sequence of each cDNA in the library; d. Isolate nanobodies from the same or second blood sample of the camel immunized with the antigen; e. Digest the nanobody with trypsin or chymotrypsin to produce a set of digestion products; f. Perform mass spectrometry analysis on the digestion products to obtain mass spectrometry data; g. Select the sequence identified in step c. that is associated with the mass spectrometry data; h. Identify the sequences from the CDR3, CDR2, and / or CDR1 regions in the sequence from step g; and i. Select from the CDR3, CDR2, and / or CDR1 region sequences of step h. those sequences having a fragmentation coverage percentage equal to or greater than the desired fragmentation coverage percentage; wherein, when chymotrypsin is used in step e., the fragmentation coverage percentage is determined by the formula f(x, chymotrypsin) = 0.0023x 2 -0.0497x+0.7723,x[5,30] is determined, or when trypsin is used in step e, the fragmentation coverage percentage is determined by the formula f(x,trypsin)=0.00006x 2 - 0.00444x+0.9194, x[5,30] is determined, where x is the length of the CDR3, CDR2, or CDR1 region sequence; and j. Wherein the selected sequences in step i. comprise a group with a reduced number of false positive CDR3, CDR2 and / or CDR1 sequences.

2. The method of claim 1, wherein the desired fragmentation coverage percentage is approximately 30%.

3. The method of claim 1, wherein the desired fragmentation coverage percentage is about 50% and trypsin is used in step e.

4. The method of claim 1, wherein the desired fragmentation coverage percentage is about 40% and chymotrypsin is used in step e.

5. The method of any one of claims 1 to 4, wherein step d. comprises obtaining plasma from the blood sample and isolating nanobodies using one or more affinity separation methods.

6. The method of claim 5, wherein the one or more affinity separation methods in step d. comprise one or more of protein G agarose affinity chromatography and protein A agarose affinity chromatography.

7. The method of any one of claims 1 to 4, wherein step d. further comprises a function selection step, the step comprising selecting antigen-specific nanobodies using antigen-specific affinity chromatography and eluting the antigen-specific nanobodies at different stringency levels to generate different nanobodies fractions, and performing steps e. to i. separately for each fraction, and estimating the affinity of the antigen for each different step i. of the CDR3, CDR2, and / or CDR1 region sequences based on the relative abundance of the CDR3, CDR2, and / or CDR1 region sequences in each of the nanobodies fractions.

8. The method of claim 7, wherein the antigen-specific affinity chromatography is a resin that binds to the antigen.

9. The method of claim 7, wherein the antigen-specific affinity chromatography is a resin coupled with maltose-binding protein and the antigen.

10. The method of any one of claims 1 to 4, further comprising generating CDR3, CDR2 and / or CDR1 peptides having the sequences identified in step i.

11. The method of any one of claims 1 to 4, further comprising generating nanobodies comprising CDR3, CDR2 and / or CDR1 regions having the sequences identified in step i.

Citation Information

Patent Citations

  • Improvements in the Manufacture of Rubber Articles.

    GB103600A

  • Immunoglobulins devoid of light chains

    WO1994004678A1