Compositions and methods for identifying nanobodies and nanobody affinities
By using a method that includes digesting nanobodies with trypsin or chymotrypsin and mass spectrometry, followed by antigen-specific affinity chromatography and computational analysis, the method addresses the challenges of nanobody proteome analysis, achieving accurate and efficient nanobody identification with reduced false positives.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-10-17
- Publication Date
- 2026-03-10
AI Technical Summary
Current methods for large-scale, sensitive, and reliable analysis of nanobody (Nb) proteomes face challenges due to the high diversity and dynamic range of circulating antibodies, limited database availability, and difficulties in accurately identifying complementarity-determining regions (CDRs) like CDR3, CDR2, and CDR1, leading to false positives and inefficiencies in nanobody identification.
A method involving obtaining a blood sample from a camelid, creating a cDNA library, digesting nanobodies with trypsin or chymotrypsin, performing mass spectrometry analysis, and selecting sequences that correlate with mass spectrometry data to identify specific CDR regions, combined with antigen-specific affinity chromatography and a computer-implemented method to reduce false positives and enhance specificity.
This approach allows for accurate quantification and classification of large nanobody repertoires with reduced false positives, enabling efficient identification of nanobodies with high antigen affinity.
Smart Images

Figure 2026041706000001_ABST
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application is a continuation of U.S. Provisional Application No. 63 / 018,559, filed May 1, 2020. The benefit of this U.S. provisional application is claimed, and the entirety of this U.S. provisional application is expressly incorporated herein by reference. do. [Background technology]
[0002] Nanobodies (Nb) are the V-chain heavy chain antibodies (HcAb) of camelids. H Derived from the H domain Nb is a naturally occurring antigen-binding fragment that binds to the IgG1 antigen. Nb is characterized by its small size and exceptional structural properties. Robustness, excellent solubility and stability, ease of bioengineering and manufacturing, low immunogenicity in humans For these reasons, Nb has the properties of: They have emerged as promising agents for cutting-edge biomedical, diagnostic, and therapeutic applications (M uyldermans, 2013;Beghein, 2017;Rasmussen , 2011;Jovcevska, I. & Muyldermans, S, 2 020).
[0003] Display-based techniques have been developed for Nb discovery (Lauwereys, 1998;Pardon, 2014;McMahon, 2018;Egloff, These methods typically target a small number of molecules that bind to a specific target with moderate affinity. Generate targeted synthetic Nb and directly isolate naturally circulating antigen-specific HcAb / Nb repertoires. Recently, mass spectrometry-based proteomics has emerged as a promising approach for Nb discovery. (Fridy, 2014). However, for at least several reasons, Significant challenges remain for large-scale, sensitive, and reliable analysis of specific Nb proteomes. (a) The diversity and dynamic range of circulating antibodies is greater than that of any cellular proteome. (b) Nb sequence database obtained from immunized camelids. typically contain millions of unique sequences that pose challenges to accurate database searches. (Savitski, 2015). (c) This large database is stored The Nb framework sequences account for a large proportion of the Nb fragments, providing little specificity for identification. The specificity is mainly determined by the complementarity-determining regions (CDRs), among which CDRs The DR3 loop can be long, making reliable MS analysis difficult. (d) Current The method is an efficient protocol that allows for accurate quantification and classification of large Nb repertoires. Limited by the availability of databases and informatics. Summary of the Invention
[0004] Provided herein are nanoparticles with 3, 2, and / or 1 complementarity determining regions (CDRs). Identifying a group of CDR3, CDR2 and / or CDR1 amino acid sequences, The identified CDR3, CDR2 and / or CDR1 sequences are false positives compared to the control. 1. A method comprising: (a) obtaining a blood sample from a camelid immunized with an antigen; (b) obtaining a cDNA library of nanobodies using blood samples; (c) identifying the sequence of each cDNA in the library; and (d) immunizing the antigen. isolating nanobodies from the same or a second blood sample from the camelid; e) digesting the nanobodies with trypsin or chymotrypsin to generate digestion products; (f) performing mass spectrometry analysis of the digestion products to obtain mass spectrometry data; and (g). (h) selecting the sequences identified in step c that correlate with the mass spectrometry data; Identifying the sequences of the CDR3, CDR2 and / or CDR1 regions within the sequence of the fragment (i) from the sequences of the CDR3, CDR2 and / or CDR1 regions of step h, and selecting sequences that are at least as fragmented as the fragment coverage percentage; The selected sequences have reduced false positive CDR3, CDR2 and / or CDR1 sequences. In some embodiments, step (d) comprises: obtaining plasma from the cells and isolating the nanobodies using one or more affinity isolation methods. In some embodiments, the one or more affinity isolation methods of step (d) include Protein G Sepharose affinity chromatography and Protein A Sepharose affinity chromatography In some embodiments, step (d) comprises one or more of: Selecting antigen-specific nanobodies using antigen-specific affinity chromatography; The antigen-specific nanobodies were eluted under various degrees of stringency, thereby and producing a nanobody fraction comprising step (e) to step (f). i) is carried out separately for each fraction, and each different step (i) for the antigen The affinity of the CDR3, CDR2 and / or CDR1 region sequences of the nanobody, respectively The relative sequence of the CDR3, CDR2 and / or CDR1 regions in each fraction It further includes a functional selection step that predicts based on pair abundance.
[0005] In some embodiments, the Nanobody amino acid sequence of complementarity determining region (CDR) 3 (CDR) 4 is A method for identifying a group of CDR3 sequences (CDR2 sequences) in which the reduced CDR3 sequences are false positives compared to the control. A method comprising: (a) obtaining a blood sample from a camelid immunized with an antigen; (b) obtaining a cDNA library of nanobodies using a blood sample; (c) identifying the sequence of each cDNA in the library; and (d) identifying the antigen-immunized laminin. isolating nanobodies from the same or a second blood sample from the cucurbit; and (e ) digesting the nanobody with trypsin or chymotrypsin to generate a digestion product population; (f) performing mass spectrometry analysis of the digestion products to obtain mass spectrometry data; and (g) performing mass spectrometry analysis of the digestion products. (h) selecting the sequences identified in step c that correlate with the quantitative analysis data; (i) identifying the sequence of the CDR3 region within the sequence of step g; and (ii) determining the CDR3 region of step h. selecting sequences from the sequence having at least a required fragmentation coverage percentage; The selected sequences of step (i) include a group having reduced false positive CDR3 sequences. In some embodiments, step (d) comprises extracting plasma from the blood sample. and isolating the nanobodies using one or more affinity isolation methods. In some embodiments, the one or more affinity isolation methods of step (d) include using Protein G Se. Sepharose affinity chromatography and Protein A Sepharose affinity chromatography In some embodiments, step (d) comprises one or more of the following: Selecting antigen-specific nanobodies using affinity chromatography and The antigen-specific nanobodies are eluted under a given stringency, thereby allowing the identification of different nanobodies. and producing a fraction, wherein steps (e) through (i) are performed for each fraction. The fractions are run separately and the CDR3 regions of each different step (i) are analyzed against the antigen. The affinity of the CDR3 region sequences was calculated by the relative affinity of the CDR3 region sequences in each of the nanobody fractions. It further includes a functional selection step that predicts based on abundance.
[0006] In some embodiments, the Nanobody amino acid sequence of complementarity determining region (CDR) 2 (CDR) 3 is a method for identifying a group of CDR2 sequences in which the reduced CDR2 sequences are false positives compared to the control; A method comprising: (a) obtaining a blood sample from a camelid immunized with an antigen; (b) obtaining a cDNA library of nanobodies using a blood sample; (c) identifying the sequence of each cDNA in the library; and (d) identifying the antigen-immunized laminin. isolating nanobodies from the same or a second blood sample from the cucurbit; and (e ) digesting the nanobody with trypsin or chymotrypsin to generate a digestion product population; (f) performing mass spectrometry analysis of the digestion products to obtain mass spectrometry data; and (g) performing mass spectrometry analysis of the digestion products. (h) selecting the sequences identified in step c that correlate with the quantitative analysis data; (i) identifying the sequence of the CDR2 region within the sequence of step g; and (ii) determining the CDR2 region of step h. selecting sequences from the sequence having at least a required fragmentation coverage percentage; The selected sequences of step (i) include a group having reduced false positive CDR2 sequences. In some embodiments, step (d) comprises extracting plasma from the blood sample. and isolating the nanobodies using one or more affinity isolation methods. In some embodiments, the one or more affinity isolation methods of step (d) include using Protein G Se. Sepharose affinity chromatography and Protein A Sepharose affinity chromatography In some embodiments, step (d) comprises one or more of the following: Selecting antigen-specific nanobodies using affinity chromatography and The antigen-specific nanobodies are eluted under a given stringency, thereby allowing the identification of different nanobodies. and producing a fraction, wherein steps (e) through (i) are performed for each fraction. The fractions are run separately and the CDR2 regions of each different step (i) are analyzed against the antigen. The affinity of the CDR2 region sequences was calculated by the relative affinity of the CDR2 region sequences in each of the nanobody fractions. It further includes a functional selection step that predicts based on abundance.
[0007] In some embodiments, the Nanobody amino acid sequence of Complementarity Determining Region (CDR) 1 (CDR) 2 a method for identifying a group of CDR1 sequences in which the reduced CDR1 sequences are false positives compared to the control; A method comprising: (a) obtaining a blood sample from a camelid immunized with an antigen; (b) obtaining a cDNA library of nanobodies using a blood sample; (c) identifying the sequence of each cDNA in the library; and (d) identifying the antigen-immunized laminin. isolating nanobodies from the same or a second blood sample from the cucurbit; and (e ) digesting the nanobody with trypsin or chymotrypsin to generate a digestion product population; (f) performing mass spectrometry analysis of the digestion products to obtain mass spectrometry data; and (g) performing mass spectrometry analysis of the digestion products. (h) selecting the sequences identified in step c that correlate with the quantitative analysis data; (i) identifying the sequence of the CDR1 region within the sequence of step g; and (ii) determining the CDR1 region of step h. selecting sequences from the sequence having at least a required fragmentation coverage percentage; The selected sequences of step (i) include a group having reduced false positive CDR1 sequences. In some embodiments, step (d) comprises extracting plasma from the blood sample. and isolating the nanobodies using one or more affinity isolation methods. In some embodiments, the one or more affinity isolation methods of step (d) include using Protein G Se. Sepharose affinity chromatography and Protein A Sepharose affinity chromatography In some embodiments, step (d) comprises one or more of the following: Selecting antigen-specific nanobodies using affinity chromatography and The antigen-specific nanobodies are eluted under a given stringency, thereby allowing the identification of different nanobodies. and producing a fraction, wherein steps (e) through (i) are performed for each fraction. The fractions are run separately and the CDR1 regions of each different step (i) are analyzed against the antigen. The affinity of the CDR1 region sequences was calculated by the relative affinity of the CDR1 region sequences in each of the nanobody fractions. It further includes a functional selection step that predicts based on abundance.
[0008] In some embodiments, antigen-specific affinity chromatography is performed using a method comprising: In some embodiments, antigen-specific affinity chromatography is performed using a resin coated with the antibody. is a resin coupled to a protein tag and an antigen. In some embodiments, the antigen specific Heterologous affinity chromatography is performed using a resin coupled to maltose binding protein and antigen. is.
[0009] Some embodiments include a CDR3, CDR2, or CDR3 fragment having the sequence identified in step (i). In some embodiments, step (i) further comprises generating a CDR1 peptide. Nanobodies containing CDR3, CDR2, and / or CDR1 regions having the sequences identified in The method further includes creating a
[0010] Also included herein are SEQ ID NOs: 1-2536 and SEQ ID NO: 26 Nanobodies comprising an amino acid sequence selected from: 65-2667.
[0011] Also provided herein is a computer-implemented method, comprising: (a) a nanoscale image processing system; (b) receiving a nanobody peptide sequence; and (b) determining multiple complementarity of the nanobody peptide sequence. and identifying CDR regions, wherein the CDR regions are CDR3, CDR2 and CDR3. and (c) applying a fragmentation filter to identify fragments containing CDR1 regions. and one or more false positive CDR3, CDR2 and / or CDR3 of the nanobody peptide sequence. (d) discarding one or more of the discarded nanobody peptide sequences; (e) quantifying the abundance of CDR3, CDR2 and / or CDR1 regions that are not present in the DNA; one or more undiscarded CDR3, CDR2 and / or CDR3 of the Nanobody peptide sequence and predicting antigen affinity based on the quantified abundance of the CDR1 region. It is a computer-implemented method.
[0012] In some embodiments, the computer-implemented method comprises: The undiscarded CDR3, CDR2 and / or CDR1 regions on the The antigen affinity of the antibody can be further classified as having low, moderate, or high antigen affinity. include.
[0013] In some embodiments, the computer-implemented method comprises: one or more undiscarded CDR3, CDR2 and / or CDR3s of the selected Nanobody peptide sequence or CDR1 region into a Nanobody protein.
[0014] In some aspects of the computer-implemented method, the fragmentation filter is In other embodiments, the fragmentation coverage ratio is calculated based on the fragmentation coverage ratio. In further embodiments, the minimum calculated fragmentation coverage percentage is about 30%. In some embodiments, the minimum calculated fragmentation coverage percentage is determined based on the trypsin treatment. Approximately 50% for the sample and approximately 40% for the chymotrypsin-treated sample. be.
[0015] In some embodiments, the computer-implemented method comprises: and comparing each of the nanobody peptide sequences to a database to identify the nanobody. Separation of dipeptide sequences into excluded and non-excluded subgroups and further comprising: adding the nanobody peptide sequences of the excluded subgroup to a database. The CDR regions are not found in the nanobody peptide sequences of the non-excluded subgroup. It is identified only by the column.
[0016] In some embodiments of the computer-implemented method, one or more of the Nanobody peptide sequences The abundance of the undiscarded CDR3, CDR2 and / or CDR1 regions of the In some embodiments, antigen affinity is quantified based on ion signal intensity. , as inferred using k-means clustering based on epitope similarity.
[0017] Also provided herein is a method for training a deep learning model, comprising: creating a data set using the data implementation method; and using the data set to Nanobody peptide sequences with high antigen affinity and nanobody peptide sequences with high antigen affinity and training a deep learning model to classify columns and a training array containing multiple Nanobody peptide sequences and corresponding antigen affinity labels; In some embodiments, the deep learning model is a convolutional It is a neural network.
[0018] Further provided herein are methods for determining the antigen affinity of Nanobody peptide sequences. By receiving nanobody peptide sequences and applying them to a trained deep learning model, By inputting the nanobody peptide sequence and using a pre-trained deep learning model, Nanobody peptide sequences were classified as having low or high antigen affinity. In some embodiments, the deep learning model comprises: In some embodiments, a trained convolutional neural network The deep learning model is trained according to the method for training a deep learning model described above. will be done. [Brief explanation of the drawings]
[0019] [Figure 1A]In silico analysis of the NGS Nb database reveals the superiority of chymotrypsin for Nb proteomics. Nb crystal structure (PDB: 4QGY). CDR loops are color-coded. [Figure 1B] In silico analysis of NGS Nb database reveals the superiority of chymotrypsin over Nb proteomics. Sequence length distribution of CDRs in the database. [Figure 1C] In silico analysis of the NGS Nb database reveals the superiority of chymotrypsin for Nb proteomics. In silico digestion of the Nb database with two proteases and cumulative plot of the corresponding peptide masses. [Figure 1D] In silico analysis of the NGS Nb database reveals the superiority of chymotrypsin for Nb proteomics. Length distribution of trypsin- and chymotrypsin-digested CDR3 peptides. [Figure 1E] In silico analysis of the NGS Nb database reveals the superiority of chymotrypsin for Nb proteomics. Complementarity of trypsin and chymotrypsin for simulation-based Nb mapping. 10,000 Nbs with unique CDR3 sequences were randomly selected and digested in silico to generate CDR3 peptides. Peptides with molecular weights between 0.8 and 3 kDa and sufficient CDR3 coverage (≥30%) were used for Nb mapping. [Figure 1F] In silico analysis of NGS Nb databases reveals the superiority of chymotrypsin for Nb proteomics. Evaluation of unique CDR3 peptide identifications (1F: trypsin; 1G: chymotrypsin) based on the percentage of matched CDR3 fragment ions in MS / MS spectra. CDR3 peptides were identified by database searches using either the "target" database (salmon) or the "decoy" database (gray). [Figure 1G]In silico analysis of NGS Nb databases reveals the superiority of chymotrypsin for Nb proteomics. Evaluation of unique CDR3 peptide identifications (1F: trypsin; 1G: chymotrypsin) based on the percentage of matched CDR3 fragment ions in MS / MS spectra. CDR3 peptides were identified by database searches using either the "target" database (salmon) or the "decoy" database (gray). [Figure 1H] In silico analysis of NGS Nb databases reveals the superiority of chymotrypsin for Nb proteomics. Figure 1 shows a 3D plot of normalized CDR3 peptide identifications, CDR3 fragment fraction, and CDR3 length from targeted database searches. FDR is the false discovery rate. The FDR of CDR3 identifications is color-coded on the 3D plot. The color bar indicates the FDR scale. FDRs below 5% are displayed in a red gradient (1H: analysis with trypsin; 1I: analysis with chymotrypsin). Figures 1J–L show representative high-quality MS / MS spectra of CDR3 peptides digested with trypsin and chymotrypsin. The sequence in Figure 1K is NTVYLEMNSLKPEDTAVYSCAAGVSDYGCYR (SEQ ID NO: 2656). The sequence in Figure 1L is YCAAAEGLASGSY (SEQ ID NO: 2657). [Figure 1I]In silico analysis of NGS Nb databases reveals the superiority of chymotrypsin for Nb proteomics. Figure 1 shows a 3D plot of normalized CDR3 peptide identifications, CDR3 fragment fraction, and CDR3 length from targeted database searches. FDR is the false discovery rate. The FDR of CDR3 identifications is color-coded on the 3D plot. The color bar indicates the FDR scale. FDRs below 5% are displayed in a red gradient (1H: analysis with trypsin; 1I: analysis with chymotrypsin). Figures 1J–L show representative high-quality MS / MS spectra of CDR3 peptides digested with trypsin and chymotrypsin. The sequence in Figure 1K is NTVYLEMNSLKPEDTAVYSCAAGVSDYGCYR (SEQ ID NO: 2656). The sequence in Figure 1L is YCAAAEGLASGSY (SEQ ID NO: 2657). [Figure 1J] In silico analysis of NGS Nb databases reveals the superiority of chymotrypsin for Nb proteomics. Figure 1 shows a 3D plot of normalized CDR3 peptide identifications, CDR3 fragment fraction, and CDR3 length from targeted database searches. FDR is the false discovery rate. The FDR of CDR3 identifications is color-coded on the 3D plot. The color bar indicates the FDR scale. FDRs below 5% are displayed in a red gradient (1H: analysis with trypsin; 1I: analysis with chymotrypsin). Figures 1J–L show representative high-quality MS / MS spectra of CDR3 peptides digested with trypsin and chymotrypsin. The sequence in Figure 1K is NTVYLEMNSLKPEDTAVYSCAAGVSDYGCYR (SEQ ID NO: 2656). The sequence in Figure 1L is YCAAAEGLASGSY (SEQ ID NO: 2657). [Figure 1K]In silico analysis of NGS Nb databases reveals the superiority of chymotrypsin for Nb proteomics. Figure 1 shows a 3D plot of normalized CDR3 peptide identifications, CDR3 fragment fraction, and CDR3 length from targeted database searches. FDR is the false discovery rate. The FDR of CDR3 identifications is color-coded on the 3D plot. The color bar indicates the FDR scale. FDRs below 5% are displayed in a red gradient (1H: analysis with trypsin; 1I: analysis with chymotrypsin). Figures 1J–L show representative high-quality MS / MS spectra of CDR3 peptides digested with trypsin and chymotrypsin. The sequence in Figure 1K is NTVYLEMNSLKPEDTAVYSCAAGVSDYGCYR (SEQ ID NO: 2656). The sequence in Figure 1L is YCAAAEGLASGSY (SEQ ID NO: 2657). [Figure 2A] Figure 1. Schematic diagram of a hybrid proteomics pipeline for reliable and in-depth analysis of antigen-bound Nb proteomes. Schematic diagram of the pipeline for Nb proteomics. The pipeline consists of three main components: immunization of camelids and purification of antigen-specific Nbs, proteomic analysis of Nbs (facilitated by the dedicated software Augur Llama and deep learning), and high-throughput integrated structural analysis of antigen-Nb complexes. [Figure 2b] Figure 1. Schematic of the hybrid proteomics pipeline for reliable and in-depth analysis of antigen-binding Nb proteomes. ELISA measurements of camelid immune responses to three antigens: GST, HSA, and PDZ. [Figure 2c] Schematic of the hybrid proteomics pipeline for reliable and in-depth analysis of antigen-binding Nb proteomes: Identification of unique CDR combinations and unique CDR3 sequences for different antigens. [Figure 2d] Figure 1. Schematic of a hybrid proteomics pipeline for reliable and in-depth analysis of antigen-binding Nb proteomes. Comparison of trypsin and chymotrypsin for CDR3 mapping of high-quality NbGSTs. [Figure 2e] Schematic diagram of the hybrid proteomics pipeline for reliable and in-depth analysis of antigen-binding Nb proteomes. Comparison of Nb GST CDR3 identification by three different proteases (gluC, trypsin, and chymotrypsin). Results are based on three independent experiments. [Figure 2F] Schematic of the hybrid proteomics pipeline for reliable and in-depth analysis of antigen-binding Nb proteomes. Solubility of randomly selected antigen-specific Nbs. [Figure 2G] Figure 1. Schematic of the hybrid proteomics pipeline for reliable and in-depth analysis of antigen-binding Nb proteomes. Validation of selected Nbs for antigen binding. [Figure 3a] Classification of Nb repertoires for GST, HSA, and PDZ binding. Label-free MS quantification and heatmap analysis of chymotrypsin-based CDR3 GST fingerprints. [Figure 3b] Classification of Nb repertoires for GST, HSA, and PDZ binding. Reproducibility and accuracy of label-free CDR3 GST peptide quantification by chymotrypsin. [Figure 3c] Classification of Nb repertoires for GST, HSA, and PDZ binding. Proportions of different Nb affinity clusters classified by quantitative proteomics. [Figure 3d] Classification of Nb repertoires for GST, HSA, and PDZ binding. Linear correlation (R2 = 0.85) between Nb ELISA affinity (LogIC50 at OD450nm) and SPRKD measurements. [Figure 3e] Classification of Nb repertoires for GST, HSA, and PDZ binding. Box plots of ELISA affinities of different Nb clusters. p-values were calculated based on Student's t-test. * indicates p-value < 0.05, ** indicates p-value < 0.01, *** indicates p-value < 0.001, **** indicates p-value < 0.0001, and ns indicates not significant. [Figure 3f]Classification of Nb repertoire for GST, HSA, and PDZ binding. Plot summarizing ELISA affinity of 25 Nb HSA (circles), OD at 450 nm. KD affinity of the top 14 Nbs ranked by ELISA was measured by SPR (triangles). [Figure 3g] 1. Classification of Nb repertoire for GST, HSA, and PDZ binding. 2. Summary plot of ELISA affinities of 11 soluble NbPDZs. [Figure 3h] Classification of Nb repertoires for GST, HSA, and PDZ binding. SPR kinetic analysis of representative NbGSTs from three different affinity clusters. For G60 (C1), Ka(1 / Ms) = 4.9e3, Kd(1 / s) = 5.9e-3, KD = 1.3 μM; for G95 (C2), Ka(1 / Ms) = 1.4e4, Kd(1 / s) = 1.1e-3, KD = 77 nM; for G13 (C3), Ka(1 / Ms) = 4.74e5, Kd(1 / s) = 1.7e-4, KD = 360 pM. [Figure 3i] Classification of Nb repertoires for GST, HSA, and PDZ binding. Representative SPR kinetics measurements of the high-affinity Nb HSA. For H14, Ka(1 / Ms) = 2.5e5, Kd(1 / s) = 5.75e-6, and KD = 22.3 pM. [Figure 3j] Classification of Nb repertoire for GST, HSA, and PDZ binding. SPR kinetics measurement of NbPDZP10. For P10, Ka(1 / Ms) = 2.06e6, Kd(1 / s) = 9.03e-6, KD = 4.4 pM. [Figure 3k] Classification of Nb repertoires for GST, HSA, and PDZ binding. Immunoprecipitation of GST (1 nM) with different Nb-conjugated Dynabeads and GSH resin. [Figure 3l]Classification of Nb repertoires for GST, HSA, and PDZ binding. Schematic diagram of the PDZ domain of mammalian mitochondrial outer membrane protein 25. Fluorescence microscopy analysis of the NbP DZP10. Nbs were conjugated with Alexa Fluor 647 for native mitochondrial immunostaining in COS-7 cell lines. Mitotracker was used as a positive control. [Figure 4a] Structural landscape of the HSA-specific Nb proteome revealed by integrative structural approaches. Sequence variations in pI and hydropathy between human and camel serum albumins (top panel). Heatmap of major epitopes mapped by structural docking (bottom panel). [Figure 4b] Structural landscape of the HSA-specific Nb proteome revealed by an integrative structural approach. Ribbon representation of the four dominant HSA epitopes. HSA is shown in gray. E1, E2, and E3 are salmon, orange, and cyan, respectively. [Figure 4c] Structural landscape of the HSA-specific Nb proteome revealed by integrative structural approaches. Surface representation showing the electrostatic potential surface and colocalization of three major epitopes. [Figure 4d] Structural landscape of the HSA-specific Nb proteome revealed by an integrative structural approach. HSA epitopes and their fractions (%) based on a convergent cross-linking model (E1: residues 57–62, 135–169; E2: 322–331, 335, 356–365, 395–410; E3: 29–37, 86–91, 117–123, 252–290; E4: 566–585, 595, 598–606; and E5: 188–208, 300–306, 463–468). [Figure 4e] Structural landscape of the HSA-specific Nb proteome revealed by integrative structural approaches. Representative cross-linking models of HSA-Nb complexes. The best scoring models are presented. Satisfactory DSS or EDC cross-links are displayed as blue bars. [Figure 4f]Structural landscape of the HSA-specific Nb proteome revealed by integrative structural approaches. Representative cross-linking models of HSA-Nb complexes. The best scoring models are presented. Satisfactory DSS or EDC cross-links are displayed as blue bars. [Figure 4g] Structural landscape of the HSA-specific Nb proteome revealed by integrative structural approaches. Representative cross-linking models of HSA-Nb complexes. The best scoring models are presented. Satisfactory DSS or EDC cross-links are displayed as blue bars. [Figure 4h] Structural landscape of the HSA-specific Nb proteome revealed by integrative structural approaches. The putative salt bridge between glutamic acid 400 (HSA) and arginine 108 of Nb CDR3 is shown. Local sequence alignment between HSA and camelid albumins is shown. [Figure 4i] Structural landscape of the HSA-specific Nb proteome revealed by an integrative structural approach. ELISA affinity screening (heat map) of 19 different Nbs for binding to wild-type HSA and a point mutant (E400R). * indicates reduced affinity. [Figure 4j] Structural landscape of the HSA-specific Nb proteome revealed by integrative structural approaches. Root mean square deviation (RMSD) plot of the HSA-Nb crosslinked model. [Figure 4k] Structural landscape of the HSA-specific Nb proteome revealed by the integrative structural approach. Bar plot showing the percentage of all DSS and EDC crosslinks of HSA-Nbs that satisfy the model. [Figure 5a] Mechanism of Nb affinity maturation. CDR3 length distribution of high affinity (dark) and low affinity (light) NbGST and NbHSA. [Figure 5b] Mechanism of Nb affinity maturation. Comparison of pI of different Nbs. [Figure 5c] Mechanism of Nb affinity maturation. Comparison of CDR pI and hydropathy between different Nbs. [Figure 5d] Mechanism of Nb affinity maturation. Comparison of CDR pI and hydropathy between different Nbs. [Figure 5e] Mechanism of Nb affinity maturation. Plot of CDR3 sequences. The alignment is based on a random selection of 1,000 unique CDR3 sequences, each with the same 15-residue length. Schematic of the CDR3 architecture: the hypervariable "head" is in dark gray, the semivariable "torso" is in light gray. [Figure 5f] Mechanism of Nb affinity maturation. Pie chart of the amino acid composition of the CDR3 head (NbGST and NbHSA) and CDR2 (NbGST). Only the top six abundant residues are displayed. [Figure 5g] Mechanism of Nb affinity maturation. Relative changes in the abundance of amino acids in the CDR3 heads of both NbGST and NbHSA. Positively charged residues K (lysine) / R (arginine) / H (histidine), negatively charged residues D (aspartic acid) / E (glutamic acid), aromatic residues Y (tyrosine), and small flexible amino acids G (glycine) / S (serine) are shown. [Figure 5h] Mechanism of Nb affinity maturation. Comparison of the relative abundance of Y, G, and S on the CDR3 head between high-affinity and low-affinity NbHSAs. Their relative abundance is plotted as a function of the relative position of each residue. A representative structure (PDB:5F1O) of an antigen-Nb complex showing two tyrosines in the CDR3 head inserted into a deep pocket of the antigen. [Figure 5i] Mechanism of Nb affinity maturation. Correlation plot of ELISA affinity and the number of specific amino acids on the CDR3 head of NbHSA. Pearson correlation coefficients and statistics are displayed. [Figure 5j] 1. Mechanism of Nb affinity maturation. Correlation plot of ELISA affinity and the number of positively charged residues on CDR2 of NbGST. [Figure 5k]The mechanism of Nb affinity maturation. Sequence logos of two representative convolutional CDR3 filters (filter 14 for high-affinity NbHSA; filter 3 for low-affinity NbHSA trained by a deep learning model) are shown. The sequence in the top panel of Figure 5K is SEQ ID NO: 2661 (YXXXXXX, residue 2 can be Y, L, D, R, or I; residue 3 can be K or G; residue 4 can be R, Y, T, or D; residue 5 can be P, D, or R; residue 6 can be E, Y, V, P, W, or D; residue 7 can be G, W, D, or P). The sequence in the bottom panel of Figure 5K is SEQ ID NO: 2662 (YXXXLXX, residue 2 can be D, P, K, or A; residue 3 can be F, P, D, or A; residue 4 can be H, T, or G; residue 6 can be G, N; residue 7 can be R, P, D, or Y). [Figure 6] The excellent versatility of Nbs for antigen binding is shown in Figure 1. A shows the electrostatic potential surface of the PDZ domain and the dominant E2 epitope (PDB: 2JIK; E1: 7-8, 35-36, 43, 99-100, and E2: 25-26, 45-46, 48, 78-79, 82-83, 85-86). B shows a docking model of the high-affinity Nb PDZP10 with its long CDR3 (DeepSalmon). C shows a comparison of the crystal structure of the PDZ-peptide ligand complex (PDB: 1EB9) with the docking model of the PDZ-Nb complex. The conserved ligand-binding site is shown in cyan. The side chains of both the CDR3 and the peptide ligand are indicated. D shows a heat map showing the ELISA affinity of 11 Nbs for binding to wild-type or mutant (R46E:K48D) PDZ. * indicates a 10- to 100,000-fold decrease in ELISA affinity. E. Plot comparison of both CDR3 length (top row) and pI (bottom row) for different Nbs (high affinity NbHSA, NbGST, NbPDZ, and Nb from the sequence database). Data are smoothed with a Gaussian function. F. Comparison of pI and hydropathy between different Nbs. G. Pie chart of the top six most abundant amino acids in Nb CDR3 heads. H. Schematic model of antigen binding by Nbs. [Figure 7] Analysis of the NGS Nb database and identification of representative false-positive CDR3 peptides. A shows the normalized variability of Nb sequences. Approximately 0.5 million unique Nb sequences were aligned based on the IMGT numbering scheme, and a plot was generated. Amino acids were grouped and color-coded based on their properties (positive, negative, polar, and non-polar). B shows the mass distribution of approximately 1.5 million human proteins identified in PeptideAtlas. C shows in silico digestion of the Nb NGS database with different proteases (AspN, GluC, LysC, trypsin, and chymotrypsin) and a plot of peptide masses. D shows the overlap between the target Nb sequence database from an immunized llama and a decoy database from another native llama. Each database contained approximately 0.5 million sequences. E shows a representative low-quality / false-positive MS / MS spectrum (HCD) of a tryptic CDR3 peptide. F shows that of a chymotryptic CDR3 peptide. There were few matching high-resolution fragment ions in the spectra. The sequences in Figure 7E are NTVYLQMNSLKPE (SEQ ID NO: 2658) and DTSIYYCAATPVFQSMSTMATESVYDYWGQGTQVTVSSEPK (SEQ ID NO: 2659). The sequence in Figure 7F is CAAGSGVGLY (SEQ ID NO: 2660). [Figure 8a] Figure 1. "Augur Llama" informatics pipeline for Nb proteomics and Nb binder validation. Schematic representation of the informatics pipeline. Three modules are presented, including 1) peptide identification, 2) quality control of Nb peptides and proteins, and 3) quantification and classification. Nb proteomics data are first searched against a search engine. Initial identifications that pass the search engine can be automatically annotated and evaluated based on various quality filters at the peptide and protein level. High-quality fingerprint peptides that pass the quality filters can be quantified and clustered. [Figure 8b]Figure 1. Informatics pipeline of "Augur Llama" for Nb proteomics and Nb binder validation. Figure 2. Nb CDR3 spectra and coverage quality filters. [Figure 8c] Figure 1. Informatics pipeline of "Augur Llama" for Nb proteomics and Nb binder validation. Diagram of peptide classification method. [Figure 8d] Figure 1. Informatics pipeline of "Augur Llama" for Nb proteomics and validation of Nb binders. Phylogenetic tree and weblogo analysis of 230 unique CDR3s of identified NbPDZs. [Figure 8e] Figure 1. Informatics pipeline of "Augur Llama" for Nb proteomics and Nb binder validation. Schematic diagram of PCR amplification of HcAb variable domains (VHHs) from camelid B lymphocytes. [Figure 8f] "Augur Llama" informatics pipeline for Nb proteomics and Nb binder validation. DNA gel electrophoresis of VHH PCR amplicons from cDNA libraries prepared from immunized bone marrow / blood. [Figure 8g] Figure 1. Informatics pipeline of "Augur Llama" for Nb proteomics and validation of Nb binders. Figure 2. SDS-PAGE analysis of fractionated NbGSTs based on different fractionation protocols. [Figure 8h] Augur Llama informatics pipeline for Nb proteomics and Nb binder validation. SDS-PAGE analysis of NbPDZs. A maltose-binding protein (MBP) tag was fused to the PDZ domain, and the fusion protein was used as an affinity handle for isolation. MBP was used as a negative control for quantification. [Figure 8i] "Augur Llama" informatics pipeline for Nb proteomics and Nb binder validation. Unique Nb identification against different antigens. [Figure 8j] Figure 1. Informatics pipeline of "Augur Llama" for Nb proteomics and validation of Nb binders. Comparison of antigen-specific Nbs identified by either chymotrypsin- or trypsin-based methods. Y-axis is the percentage of positive hits randomly selected for validation. [Figure 9a] Proteome quantification, biochemical validation, and affinity measurement of NbGST. Proteome quantification and heat map analysis of NbGST based on different fractionation methods. [Figure 9b] Proteomic quantification, biochemical validation, and affinity measurements of NbGST. Pearson correlation of LC retention times of different fractionated Nb peptide samples. [Figure 9c] Proteomic quantification, biochemical validation, and affinity measurements of NbGST. Representative GST bead-binding assay. GST-conjugated resin was used to specifically isolate recombinant Nb from E. coli lysate. Red arrows indicate enriched Nb. Inactivated resin was used as a negative control. [Figure 9d-1] Proteome quantification, biochemical validation, and affinity measurements of NbGSTs. SPR kinetic measurements of 10 representative NbGSTs. [Figure 9d-2] Proteome quantification, biochemical validation, and affinity measurements of NbGSTs. SPR kinetic measurements of 10 representative NbGSTs. [Figure 10a-1] Characterization of high-quality HSA and PDZ Nb. SPR kinetics of a representative high-affinity NbHSA. [Figure 10a-2] Characterization of high-quality HSA and PDZ Nb. SPR kinetics of a representative high-affinity NbHSA. [Figure 10b] Characterization of high-quality HSA and PDZ Nbs. Bead-binding assay of selected high-quality Nb-PDZs. Recombinant MBP-fused PDZs were used as affinity handles to isolate Nbs from E. coli lysates. MBP-bound resin was used as a negative control. I: E. coli lysate input, B: bead control, P: affinity pull-out with PDZ. [Figure 11] Hybrid structural analysis of GST-Nb complexes. A. Heatmap analysis of structural docking of 64,670 GST-Nb complexes showing three convergent epitopes (E1: 75-88, 143-148; E2: 33-43, 107-127; E3: 158-200, 213-220). B. Ribbon representation of the three major GST epitopes. GST dimers are displayed in gray. E1, E2, and E3 are light yellow, orange, and dark blue-green, respectively. C. Surface representation showing electrostatic surface colocalization with the three major epitopes. D. GST epitopes and their abundance (%) based on the convergent crosslinking model are displayed in different colors. [Figure 12a] Analysis of the CDR sequences of different Nbs and the sequence conservation of camelid and human albumins. Comparison of the amount of amino acids on the CDR3 head between high-affinity and low-affinity Nbs. [Figure 12b] Analysis of the CDR sequences of different Nbs and the sequence conservation of camelid and human albumins. Comparison of the amount of amino acids on the CDR3 head between high-affinity and low-affinity Nbs. [Figure 12c] Analysis of the CDR sequences of different Nbs and sequence conservation between camelid and human albumins. Comparison of CDR1 and CDR2 of different Nbs. [Figure 12d] Analysis of the CDR sequences of different Nbs and sequence conservation between camelid and human albumins. Comparison of CDR1 and CDR2 of different Nbs. [Figure 12e] Analysis of the CDR sequences of different Nbs and sequence conservation between camelid and human albumins. Comparison of CDR1 and CDR2 of different Nbs. [Figure 12f] Analysis of the CDR sequences of different Nbs and sequence conservation between camelid and human albumins. Comparison of CDR1 and CDR2 of different Nbs. [Figure 12g]Analysis of the CDR sequences of different Nbs and the sequence conservation between camelid and human albumins. Comparison of the relative positions of tyrosine (Y), glycine (G), and serine (S) on the CDR3 heads of GST Nbs. [Figure 12h] Analysis of the CDR sequences of different Nb and sequence conservation in camelid and human albumins. Sequence alignment of human serum albumin and llama serum albumin. Conserved amino acids are highlighted. [Figure 13] Comparison between different antigen epitopes. A: Comparison of the shapes of the major epitopes of three different antigens (i.e., E2 of PDZ, E3 of GST dimer, and E3 of HSA). Different epitopes are color-coded on the antigen structures. B: Surface electrostatic potential vs. E1 epitope of PDZ domain. C: Plot of the solvent-accessible area of different epitopes. The y-axis represents the area of different epitopes in square angstroms. D: Net charge of the epitope. E: Relative abundance of various amino acids on the CDR3 head. DB: NGS Nb sequence database. F: Comparison of pI between CDR1 and CDR2 between different antigen-specific Nbs. [Figure 14] 1 illustrates an example computing system that can perform the methods and procedures described in certain embodiments of the present disclosure. [Figure 15]Figures 15A-B show the results of an amino acid sequence filter derived from a deep learning approach. The sequence filter can be used to accurately separate low-affinity binding HSA Nbs from high-affinity binding HSA Nbs. The sequence in Figure 15A is SEQ ID NO:2663 (LXYRXXX, residue 2 can be N, Y, V, or G; residue 5 can be L or W; residue 6 can be E, G, N, T, or S; residue 7 can be D or E). The sequence of Figure 15B is SEQ ID NO:2664 (XXXXXXX, residue 1 can be C, F, Q, S, H, K, L, Y, or R; residue 2 can be G, P, A, or N; residue 3 can be E, S, G, T, P, V, Y, H, or A; residue 4 can be C, A, S, P, or D; residue 5 can be I, W, V, T, or A; residue 6 can be M, Q, or H; residue 7 can be K, Y, Q, V, or W). [Figure 16] Figures 16A-C show the results of an amino acid sequence filter derived from a deep learning approach. Using the sequence filter, we can accurately separate low-affinity binding HSA Nbs from high-affinity binding HSA Nbs. The sequence in Figure 16A is SEQ ID NO: 2665 (TXXXLXX; residue 2 can be D, P, K, or A; residue 3 can be F, P, L, D, or A; residue 4 can be H, T, or G; residue 6 can be G, E, N, or R; residue 7 can be R, P, G, D, or Y). The sequence in Figure 16B is SEQ ID NO: 2666 (XXRXXXX; residue 1 can be E, G, W, D, or I; residue 2 can be N, G, or C; residue 4 can be A, H, or D; residue 5 can be E, R, Y, A, or T; residue 6 can be G, A, or P; residue 7 can be L, S, or Y). The sequence in Figure 16C is SEQ ID NO:2667 (XXGAQXW; residue 1 can be R or A; residue 2 can be K or L; residue 6 can be L, G, Y, or W). DETAILED DESCRIPTION OF THE INVENTION
[0020] Reported here is a detailed discovery, classification, and hybridization of the antigen-associated Nb repertoire. This is an integrated proteomic platform for high-throughput structural characterization. The sensitivity and robustness of the technique has been demonstrated by the use of immunoglobulins containing small, weakly immunogenic antigens derived from mitochondrial membranes. Validated using antigens spanning three orders of magnitude in the immune response. Tens of thousands of highly diverse and specific The Nb family was clearly identified and quantified according to its physicochemical properties. The reaction had sub-nM affinity. Over 100,000 antigen-Nb complexes using genome-wide sequencing and deep learning These proteins have been systematically investigated to significantly advance our understanding of immunogenicity and Nb affinity maturation. Studies of the mammalian humoral immune system have revealed its remarkable efficiency, specificity, diversity, and versatility. It became clear.
[0021] term As used in this specification and claims, the singular forms "a," "an," and "the" " includes plural referents unless the context clearly dictates otherwise. For example, the term "a "Cell" includes a plurality of cells, including mixtures thereof.
[0022] As used herein, the term "about" when referring to a measurable value, such as an amount, percentage, or the like, is used to refer to the amount of the measurement. This means that the deviations are within ±20%, ±10%, ±5%, or ±1% of the determinable value. do.
[0023] "Administration" or "administering" to a subject includes introducing or delivering an agent to a subject. Administration may be by any route, including oral, intravenous, intraperitoneal, intranasal, inhalation, etc. Administration can be by any suitable route. Administration can include self-administration and administration by another person. Examples include:
[0024] The term "antibody" is used broadly herein and includes polyclonal antibodies, monoclonal antibodies, and In addition to intact immunoglobulin molecules, "antibody" molecules include clonal antibodies and bispecific antibodies. Also included in the term "antibody" are fragments or polymers of those immunoglobulin molecules, and human or humanized forms of immunoglobulin molecules or fragments thereof. is usually composed of about 150 identical light (L) chains and two identical heavy (H) chains. It is a heterotetrameric glycoprotein of 2,000 daltons. Each heavy chain has a variable domain at one end. Main (V H ) followed by several constant domains. Each light chain has at one end variable domain (V L ) and at its other end a constant domain.
[0025] As used herein, the term "antigen" or "immunogen" refers to a substance that elicits an immune response in a subject. Substances capable of inducing ATP, typically proteins, nucleic acids, polysaccharides, toxins, or lipids The term is also used interchangeably to refer to proteins (directly or administering to a subject a nucleotide sequence or vector encoding the protein; When administered to a subject (by Refers to a protein that is immunologically active in the sense that it is capable of eliciting an immune response.
[0026] The terms "antigenic determinant" and "epitope" are also used interchangeably herein. and on the antigen recognized by the antigen-binding molecule (such as the Nanobody of the invention). An epitope can be a sequence of adjacent amino acids (a "linear epitope"), or a sequence of adjacent amino acids (a "linear epitope"), or a sequence of adjacent amino acids (a "linear epitope"), or a "single-stranded epitope" (a "single-stranded ... can be formed both from juxtaposed non-adjacent amino acids by tertiary folding of proteins. The latter epitopes are those formed by at least several non-contiguous amino acids. These are referred to herein as "conformational epitopes." Epitopes are usually at least At least three, more commonly at least five or eight to ten amino acids in a unique space. Methods for determining the spatial structure of an epitope include, for example, X-ray crystallography. and 2D nuclear magnetic resonance. For example, Epitope Mapping Pro tocols in Methods in Molecular Biology, Vol. 66, Glenn E. Morris, Ed. (1996) sea bream.
[0027] The terms "antigen-binding site," "binding site," and "binding domain" refer to an antigenic determinant or A specific element, part, or antigen of a polypeptide, such as a Nanobody, that binds to an epitope. It refers to amino acid residues.
[0028] As used herein, the term "biological sample" refers to a sample of biological tissue or biological fluid. "sample" means a sample. Such samples include tissues isolated from animals, Biological samples include, but are not limited to, biopsy and autopsy samples, histological samples, and the like. Frozen sections, blood, plasma, serum, sputum, stool, tears, mucus, hair, and other samples collected for the purpose Biological samples may also include tissue sections such as skin. Biological samples may include explants derived from patient tissue, and primary and / or transformed cell cultures. Biological samples may be derived from animals or Alternatively, a sample of cells can be obtained by removing the cells from a previously isolated (e.g. using cells (e.g., isolated by another person at a different time and / or for a different purpose) or by performing the methods disclosed herein in vivo. This can also be achieved by using archival tissue that has a history of treatment or outcome. It is also possible.
[0029] The term "cDNA library" is used herein to refer to a collection of transcripts from a given organism. It refers to the combination of different cDNA fragments that make up part of a genome.
[0030] The terms "CDR" and "complementarity determining region" are used interchangeably and refer to an antibody that is CDRs refer to the portions of the variable chains of an antibody that are involved in binding to the antigen. In some embodiments, the nanobody is part of an "antigen binding site." Each antibody contains three CDRs which collectively form the antigen-binding site.
[0031] As used herein, the term "comprising" and variations thereof ", "including" and variations thereof, and "comprising" and "including" are non-limiting terms. While the term "(ng)" is used herein to describe various embodiments, Instead of "comprising" and "including," "consisting essentially of" and " Use the term "consisting of" to specify more specific implementations. Forms may be provided and are disclosed.
[0032] "Composition" refers to any agent that has a beneficial biological effect. is a compound that provides a therapeutic effect, e.g., treatment of a disorder or other undesirable physiological condition, and a The term "prophylactic" includes both therapeutic and preventative effects, such as prevention of a disorder or other undesirable physiological condition. These terms also refer to bacteria, vectors, polynucleotides, cells, salts, esters, amides, etc. These include, but are not limited to, pro-agents, active metabolites, isomers, fragments, analogs, etc. The pharmaceutically acceptable, pharmacologically active compounds of the beneficial agents specifically mentioned herein are not When the term "composition" is used, it refers to a specific composition. Where a composition is specifically identified, the term refers to the composition itself as well as to a pharmaceutically acceptable carrier. Pharmacologically active vectors, polynucleotides, salts, esters, amides, proagents It is understood that this includes compounds, conjugates, active metabolites, isomers, fragments, analogs, etc. I want to be done that.
[0033] A "control" is another subject or sample used in an experiment for comparison purposes. It can be "positive" or "negative."
[0034] An "effective amount" includes, but is not limited to, an amount that alleviates symptoms of a medical condition or disorder (e.g., cancer) or It includes an amount that can ameliorate, cure, mitigate, prevent, or diagnose a condition or symptom. Unless otherwise indicated by, "effective amount" is limited to the smallest amount sufficient to improve the condition. The severity of the disease or disorder, and the ability to prevent, treat, or alleviate the disease or disorder The ability of a treatment to achieve this is not limited by biomarkers or clinical parameters. In some embodiments, the term "recombinant nanobody" can be used to describe "An effective amount" refers to an amount of recombinant Nanobody sufficient to prevent, treat, or alleviate cancer. Point.
[0035] A "fragment" or "functional fragment" is a fragment whose activity is comparable to that of the unmodified peptide. Other sequences, unless significantly altered or reduced compared to the peptide or unmodified protein Insertion, deletion, or deletion of specific regions or specific amino acid residues, regardless of whether they are linked to These modifications may include substitutions or other selected modifications that may alter disulfide bonds. removing or adding amino acids that can synthesize it, extending its biological life, It may provide some additional properties, such as altering secretory properties. In this case, the functional fragment may have bioactive properties such as binding to HSA and / or ameliorating cancer. It is necessary to have a certain quality.
[0036] The term "fragmentation coverage percentage" refers to the percentage obtained using the following formula: This is what is meant. f(x, enzyme) is the fragmentation coverage (%) of the peptide digested by the enzyme. is a function that calculates x is the length of the CDR3 to which the peptide is mapped. f(x, chymotrypsin) = 0.0023 × 2 -0.0497x+0.7723, x [5,30] f(x, trypsin) = 0.00006x 2 -0.00444x+0.9194, x[ 5,30] In some embodiments, a minimum calculated fragmentation coverage percentage is required. In other or further aspects, the minimum calculated fragmentation required is The coverage rate is about 30%. In some embodiments, when trypsin is the enzyme, The minimum calculated fragmentation coverage percentage required is approximately 50%, When liposin is the enzyme, it is about 40%.
[0037] As used herein, a "functional selection step" refers to the selection of nanobodies based on their functional properties. In some embodiments, the method is to separate the soluble matter into different fractions or groups. The functional properties are determined by the antigen affinity of the nanobody or the CD3, CD2, or CD1 domains. In another embodiment, the functional property is the thermal stability of the Nanobody. The functional property is the intracellular penetration of the nanobody. (CDR) 3, 2 or 1 region of the Nanobody amino acid sequence (CDR3, CDR2 or A reduced number of CDR3, CDR2 or CDR1 sequences are used to identify a group of CDR3, CDR2 or CDR1 sequences. a method for detecting a false positive result compared to a false positive result obtained by taking a blood sample from an immunized camelid animal; The blood samples were used to obtain nanobody cDNA libraries. The sequence of each cDNA in the library is identified and the antigen is immunized. Isolating nanobodies from the same or a second blood sample from the camelid and A potent selection step is performed and the nanobodies are digested with trypsin or chymotrypsin. and performing mass spectrometry of the digested products to generate a group of digested products. and selecting the sequences identified in step c that correlate with the mass spectrometry data. and identifying the sequence of the CDR3, CDR2, or CDR1 region within the sequence of step g. and the sequence of the CDR3, CDR2 or CDR1 region in step h is used to calculate the fragment length. and excluding sequences that are less than the fragmented coverage percentage, wherein the non-excluded sequences are The method includes a group having false positive CDR3, CDR2 or CDR1 sequences. The method steps following the feature selection step are: It should be understood that this may be performed separately for each section or group.
[0038] The "half-life" of an amino acid sequence, compound or polypeptide of the invention is generally determined by, e.g., Degradation of the sequence or compound, and / or clearance of the sequence or compound by natural mechanisms For detection or sequestration, serum concentrations of amino acid sequences, compounds, or polypeptides may be increased in vivo. The amino acid sequence of the nanobody of the invention can be defined as the time it takes for the amino acid sequence to decrease by 50%. The in vivo half-life of an acid sequence, compound, or polypeptide can be determined, for example, by the method of Kenneth et al., supra. , A et al., Chemical Stability of Pharma ceuticals: A Handbook for Pharmacists;Pe ters et al., Pharmacokinete analysis: A Practical Approach (1996);“Pharmacokinet ics”, M Gibaldi & D Perron, published by Marcel Dekker, 2nd Rev. edition (1982) It can be determined by any known method, such as pharmacokinetic analysis.
[0039] The terms "identity" or "homology" are used to refer to the degree of homology achieved by a sequence. After aligning the sequences and introducing gaps, if necessary, the sequence identity is calculated as and bases or residues that are identical to the bases or residues of the corresponding sequences being compared, without any consideration of conservative substitutions. shall be interpreted to mean the percentage of nucleotide bases or amino acid residues in a candidate sequence. A specific percentage (e.g., 61%, 62%, 63%, 64%, 65%) of the sequence is used for another sequence. , 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75% ,76%,77%,78%,79%,80%,81%,82%,83%,84%,85% ,86%,87%,88%,89%,90%,91%,92%,93%,94%,95% 96%, 97%, 98%, 99% or more) or polynucleotide region (or polypeptide or polypeptide region) When two sequences are compared, the percentage of bases (or amino acids) that are This alignment and percent homology or sequence identity can be determined using software programs known in the art Such alignments are described, for example, by Needleman et al. 0) Provided using the method of J. Mol. Biol. 48: 443-453 This can be done using computer programs such as the Align program (DNAstar, Inc.). In some embodiments, percent identity is determined by a ratio. It is determined along the entire length of the sequences being compared.
[0040] As used herein, the terms "increase" or "increase" generally refer to a statically significant For the avoidance of doubt, "increased" means an increase by a significant amount compared to baseline levels. at least a 10% increase, e.g., at least about 20%, or at least about 30% , or at least about 40%, or at least about 50%, or at least about 60%, or at least about a 70%, or at least about an 80%, or at least about a 90% increase or an increase of up to (and including) 100%, or an increase of 10 to 100% compared to the baseline level Any increase between 100%, or at least about 2-fold, or less than 100% compared to baseline levels at least about three times, or at least about four times, or at least about five times, or at least It means about a 10-fold increase, or any increase between 2-fold and 10-fold or more.
[0041] As used herein, the term "isolate" refers to the isolation of a biological sample, i.e., blood. As used herein, "refers to isolation from plasma, tissue, exosomes, or cells." The term "isolated," when used in the context of nucleic acids, for example, refers to the state in which the nucleic acids are bound prior to isolation. At least 60%, at least 75%, at least 90% less of the other ingredients used These terms refer to nucleic acids that are 95%, at least 98%, and even at least 99% free of the nucleic acid of interest. .
[0042] The term "mass spectrometry" refers to the measurement of the mass-to-charge ratio (m / z). "Mass spectrometry data" means the measurement of one or more molecules present in a sample. The mass, charge, mass-to-charge ratio, molecular weight, and / or amino acid identity or amino acid sequence of the molecule In some embodiments, the mass spectrometry data is a sequence of molecules present in a sample. The amino acid sequence of the clone. The sequences, including the cDNA sequences, that "correlate" with the mass spectrometry data are The predicted identical or very similar amino acid sequences determined in the mass spectrometry step of the method In some embodiments, the sequence is about 80%, about 85%, about 90%, about 91%. %, approximately 92%, approximately 93%, approximately 94%, approximately 95%, approximately 96%, approximately 97%, approximately 98%, or Correlation with mass spectrometry data occurs when there is approximately 99% similarity or identity. In terms of morphology, sequences are considered to be approximately 90-100% similar or identical to mass spectrometry data. correlates with.
[0043] As used herein, "nanobody," "V H H," "V H H antibody fragment The terms are used interchangeably and refer to PCT Publication No. WO 2004 / 023094, which is incorporated by reference in its entirety. and those derived from camelids, such as those described in 94 / 04678, which have no light chains at all. As used herein, the term "antibody" refers to the variable domain of a single heavy chain of an antibody of the type found in the family Camelidae. "Single domain antibody" as used herein refers to a nanobody and an Fc domain.
[0044] As used herein, the term "nucleic acid" refers to a nucleotide, e.g., a deoxyribonucleic acid. refers to a polymer composed of nucleotides (DNA) or ribonucleotides (RNA) As used herein, the terms "ribonucleic acid" and "RNA" refer to ribonucleic acid. As used herein, "deoxyribonucleic acid" and "deoxyribonucleic acid" refer to a polymer composed of deoxyribonucleic acids. The term "DNA" refers to a polymer composed of deoxyribonucleotides .
[0045] As used herein, "operably linked" means within a single polypeptide chain. The term "polypeptide segment" refers to an arrangement of polypeptide segments, and individual polypeptide segments may be, but are not limited to, However, proteins, fragments thereof, connecting peptides, and / or signal peptides The term operably linked means that there are amino acids between the different segments. Direct separation of different individual polypeptides within a single polypeptide or fragment thereof It refers to a fusion of individual polypeptides, and further refers to a "linker" that includes one or more intervening amino acids. It can also refer to the case where devices are connected to each other via a "network."
[0046] As used herein, "reduced," "reducing," "reduction," or "reducing" The term generally refers to a statistically significant decrease. In addition, "reduced" means a decrease of at least 5% compared to the baseline level, e.g., at least About 10%, or at least about 20%, or at least about 30%, or at least about 40%, or at least about 50%, or at least about 60%, or at least about 7 0%, or at least about 80%, or at least about 90% reduction, or up to 100% a decrease (i.e., disappearance level compared to the reference sample) in (including 100%) of This refers to any decrease between 10 and 100% compared to the baseline level.
[0047] The terms "polynucleotide" and "oligonucleotide" are used interchangeably. Deoxyribonucleotides or ribonucleotides or their analogs are used A polynucleotide refers to a polymeric form of nucleotides of any length. They can have any three-dimensional structure and can perform any function, known or unknown. The following are non-limiting examples of polynucleotides: genes or gene fragments Segment, exon, intron, messenger RNA (mRNA), transfer RNA, ribosomal RNA, ribozyme, cDNA, recombinant polynucleotide, branched polynucleotide oligonucleotides, plasmids, vectors, isolated DNA of any sequence, Isolated RNA, nucleic acid probes, and primers. Polynucleotides are methylated. Nucleotides may include modified nucleotides, such as nucleotides and nucleotide analogs. Modifications to the polymer structure, if any, can be imparted before or after assembly of the polymer. The sequence of nucleotides may be interrupted by non-nucleotide components. can be further modified after polymerization, such as by conjugation with a labeling component. Applies to both double-stranded and single-stranded molecules. Any embodiment of the present invention that is a nucleotide may comprise a double-stranded form and a double-stranded form thereof. Each of the two complementary single-stranded forms known or predicted to be Contains.
[0048] The term "polypeptide" is used in its broadest sense to mean a compound consisting of two or more subunits. The subunits are amino acids, amino acid analogs, or peptidomimetics. , peptide bonds. In another embodiment, the subunits may be linked by other bonds, For example, it may be linked by an ester, ether, etc. The term "amino acid" refers to glycine and both the D or L optical isomers, as well as amino acid analogs. Any of natural and / or unnatural or synthetic amino acids, including peptide mimetics and peptide mimetics. Peptides with three or more amino acids are generally called oligopeptides if the peptide chain is short. When the peptide chain is long, the peptide is generally called a polypeptide or protein. The terms "peptide," "protein," and "polypeptide" are used herein to refer to Used interchangeably.
[0049] "Recombinant" as used herein with respect to polypeptides refers to polypeptides that are not naturally occurring. Refers to a combination of two or more polypeptides.
[0050] The term "specificity" refers to the ability of a particular antigen-binding molecule (such as a Nanobody of the invention) to bind. Nanobodies with low specificity refer to the number of different types of antigens or antigenic determinants they can target. Multiple distinct epitopes (or polypeptides) via antigen-binding sites or binding domains While highly specific nanobodies bind to a single antigen-binding site or binding domain, It binds to one or a few epitopes (or polypeptide regions) via a In some embodiments, the small number of epitopes (or polypeptide regions) may be, for example, cross-species epitopes. Similar or very similar, such as a tope. The term "heterologously binds" when used herein in reference to Nanobodies means that the The nanobody is able to target an epitope (or polypeptide region) compared to the target epitope (or polypeptide region). Specific binding refers to binding that is preferential to a specific peptide region. This may depend on the stringency of the conditions under which it is carried out. A nanobody specifically binds to an epitope when there is high affinity binding at the In some embodiments, the HSA-binding polypeptides or Nanobodies described herein are derived from human It specifically binds to serum albumin.
[0051] The specificity of an antigen-binding molecule (e.g., HSA-binding polypeptide, Nanobody of the invention) can be determined by the parent It should be understood that affinity can be determined based on affinity and / or avidity. The equilibrium constant for dissociation of the antigen from the antigen-binding molecule (K D ) and the antigenic determinant and the antigen-binding molecule It is a measure of the binding strength between the antigen-binding site. D The smaller the value, the closer the antigenic determinant and antigen The binding strength between the binding molecules increases (or the affinity is determined by the affinity constant (K A ) This can also be expressed as 1 / K D Methods for determining affinity are well known to those skilled in the art. The binding activity of the antigen-binding molecules (HSA-binding polypeptides and nanobodies of the present invention) Avidity is a measure of the strength of binding between an antigenic determinant and an antigen-binding molecule. The affinity between the antigen-binding site on the molecule and the number of related binding sites present on the antigen-binding molecule. Typically, antigen-binding proteins (HSA-binding polypeptides, and the Nanobodies of the invention) -5 ~10 -12 moles / liter or less, preferably is 10 -7 ~10 -12 moles / liter or less, more preferably 10 -8 ~10 -12 mole / liter dissociation constant (K D ) (i.e., 10 5 ~10 12 Liters / mole or more preferred Or 10 7 ~10 12 liters / mole or more, more preferably 108 ~10 12 liter / mol binding constant (K A In some embodiments, the Ka (On rate, 1Ms) is about 10 5 , 10 6 , 10 7 , 10 8 , 10 9 , 10 10 ,Ma or 10 11 In some embodiments, the Ka is about 10 7 Some implementations In this form, the Kd (off rate, s) is approximately 10 -5 , 10 -6 , 10 -7 , 10 -8 , 1 0 -9 , 10 -10 , or 10 -11 In some embodiments, K D is about 10 -7 In some embodiments, the antigen binding proteins disclosed herein have a nucleotide sequence of about 10 -9 moles / liter less than K D binds to its antigen with a K greater than 10 μM D The value is It is generally considered to represent non-specific binding. As will be appreciated by those skilled in the art, the dissociation constant is It may be the actual dissociation constant or the apparent dissociation constant.
[0052] The term "subject" as used herein includes primates (e.g., humans), bovine, ovine, guinea pigs, and the like. Mammals, including but not limited to dogs, horses, dogs, cats, rabbits, rats, mice, etc. In some embodiments, the subject is a human.
[0053] Compositions and Methods In some embodiments, the present invention provides a method for identifying a complementarity determining region (CDR) of 3, 2, or 1. Identifying a group of nanobody amino acid sequences (CDR3, CDR2, or CDR1 sequences) in the region The reduced CDR3, CDR2 and / or CDR1 sequences are false positives compared to the control. The term "false positive" as used herein refers to a false positive that occurs when something is not present. As used herein, "a sequence is a false positive" refers to a result that indicates the presence of a gene despite the fact that the gene is not present. The phrase "is" refers to a CDR3, CDR2 and / or CDR3 fragment that does not specifically bind to the test antigen. DR1 sequence, or C contained in a nanobody that cannot specifically bind to the test antigen False positive CDR3, CDR2 and / or CDR1 sequences. The number or amount of CDR1 sequences and / or CDR2 sequences can be determined by trypsinizing the fragmentation filter. For samples, at least about 30% (e.g., at least about 30%, 35%, 40% , 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90% , 95%, or 99%, and / or at least at least about 30% (e.g., at least about 30%, 35%, 40%, 45%, 50%, 55%) %, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 99% ) and can be reduced using the methods disclosed herein. In some embodiments, the false positive CDR3, CDR2 and / or CDR1 sequences are Fragmented filters were split to approximately 50% for trypsinized samples and / or was set at approximately 40% for chymotrypsin-treated samples and the method disclosed herein can be largely eliminated using
[0054] Thus, the disclosed methods for identifying CDR3, CDR2 and / or CDR1 sequences The method determines the number of CDR3, CDR2 and / or CDR1 sequences that are false positives compared to the control. This reduction can be achieved, for example, by identifying a compound that is not identified using the methods described herein. at least About 2 times, at least about 3 times, at least about 4 times, at least about 5 times, at least about 10 times , resulting in a decrease of at least about 20-fold, at least about 50-fold, or at least about 100-fold. obtain.
[0055] In some embodiments, the method comprises: a. Obtaining a blood sample from a camelid immunized with the antigen; b. Obtaining a cDNA library of nanobodies using a blood sample; c. Identifying the sequence of each cDNA in the cDNA library; d. Nanobody from the same or a second blood sample from an antigen-immunized camelid. and isolating e. Digesting the nanobody with trypsin or chymotrypsin to generate a digestion product population. And, f. performing mass spectrometry analysis of the digestion products to obtain mass spectrometry data; g. Selecting the sequences identified in step c that correlate with the mass spectrometry data; h. Identifying the sequences of the CDR3, CDR2 and / or CDR1 regions within the sequence of step g. And, i. From the sequences of the CDR3, CDR2 and / or CDR1 regions of step h, the necessary fragments are selected. and selecting sequences that are equal to or greater than the segmentation coverage rate, and the selected sequences are a group of selected CDR3, CDR2 and / or CDR1 sequences, including a group of selected CDR3, CDR2 and / or CDR1 sequences having false positives; This includes:
[0056] In some embodiments, the method comprises: a. Obtaining a blood sample from a camelid immunized with the antigen; b. Obtaining a cDNA library of nanobodies using a blood sample; c. identifying the sequence of each cDNA in the library; d. Nanobody from the same or a second blood sample from an antigen-immunized camelid. and isolating e. Digesting the nanobody with trypsin or chymotrypsin to generate a digestion product population. And, f. performing mass spectrometry analysis of the digestion products to obtain mass spectrometry data; g. Selecting the sequences identified in step c that correlate with the mass spectrometry data; h. Identifying the sequences of the CDR3, CDR2 and / or CDR1 regions within the sequence of step g. And, i. From the sequences of the CDR3, CDR2 and / or CDR1 regions of step h, the necessary fragments are selected. selecting sequences that are greater than or equal to the fragmentation coverage rate, If chymotrypsin is used in step e, the ratio of =0.0023x2-0.0497x+0.7723, determined by x[5,30] , or if trypsin is used in step e, the formula f(x, trypsin)=0.00 006x2-0.00444x+0.9194, determined by x[5,30], x is the length of the sequence of the CDR3, CDR2 and / or CDR1 region, and , including j. The selected sequences of step i are subjected to reduced false positive CDR3, CDR2 and / or includes a group having a CDR1 sequence.
[0057] In some embodiments, the selected CDR3, CDR2 and / or CDR3 in step i The CDR1 region sequence must meet the minimum required fragmentation coverage rate of approximately 30%. In some embodiments, the selected CDR3, CDR2 and The CDR1 and / or CDR2 region sequences are required to achieve a minimum fragmentation coverage of approximately 50%. In some embodiments, trypsin is used in step e. The selected CDR3, CDR2 and / or CDR1 region sequences in the fragment are approximately 40 % of the minimum required fragmentation coverage, and is used.
[0058] The nanobody cDNA library in step b is a biological sample of the subject to be immunized. It should be understood that the sample is obtained from a blood sample or bone marrow. In this embodiment, the cDNA library is obtained from B cells. cDNA (or complementary DNA) libraries are created using reverse transcription techniques to extract DNA from biological samples. A set of cDNAs generated from mRNA in a sample (e.g., a blood or bone marrow sample). Methods for preparing cDNA libraries are well known in the art. Thus, in some embodiments, step b comprises the step of: isolating mRNA from the sample (a human or bone marrow sample) and / or the isolated mRNA The method further comprises the step of reverse transcribing the mRNA into cDNA.
[0059] The generated cDNA is then sequenced as described in step c. In one embodiment, step c comprises using a specific primer (e.g., SEQ ID NO:26 46 and SEQ ID NO: 2647) to obtain the CH2 domain from the variable domain. Steps for amplifying camelid IgG heavy chain cDNA sequences and DNA gel electrophoresis Using V lacking the CH1 domain H The H gene is a conventional IgG (having a CH1 domain) a second forward primer (e.g., SEQ ID NO: 2648) and a second reverse primer (e.g., SEQ ID NO: 2649) This step involves reamplified frameworks 1 to 4 using the 2 PCR amplicons (e.g., using a PCR clean-up kit or isolation kit) purification (e.g., using a forward primer for sequencing analysis), Reverse primer SEQ ID NO: 2650 and reverse primer SEQ ID NO: 26 51) for sequencing analysis (e.g., MiSeq sequencing analysis). This further involves another PCR step using primers that add adapters for the Methods of sequencing analysis include, for example, single molecule real-time (SMRT) sequencing. sequencing, nanopore DNA sequencing, massively parallel signature sequencing (MPS) S), Polony sequencing, 454 pyrosequencing, Illumina ( Solexa) sequencing, combinatorial probe anchor synthesis (cPAS) , SOLiD sequencing, or MiSeq sequencing.
[0060] Step d above may be performed simultaneously with steps a, b, and / or c. It can be performed before steps a, b, and / or c, or after steps a, b, and / or c. In some embodiments, step d comprises obtaining plasma from the blood sample; and isolating the Nanobodies using one or more affinity isolation methods. For example, Protein G Sepharose affinity chromatography, Protein A Sepharose affinity chromatography, Rhodium affinity chromatography, hydroxylapatite chromatography, gel electrophoresis The method can be any affinity separation method known in the art, including electrophoresis, or dialysis. Protein G Sepharose affinity chromatography and Protein A Sepharose affinity Two well-known affinity chromatography methods are odzki AC, Berenstein E. (2010) Antibod Purification: Affinity Chromatography - Protein A and Protein G Sepharose. In: Oliver C., Jamur M. (eds) Immunocytoche mical Methods and Protocols. Methods in Molecular Biology (Methods and Protocols) ), vol. 588. Humana Press.) This method It relies on reversible interactions between specific ligands immobilized on a chromatographic matrix. The sample is subjected to electrostatic and hydrophobic interactions, van der Waals forces, and / or is applied under conditions that favor specific binding to the ligand as a result of hydrogen bonding. After washing away unbound substances, the buffer conditions are changed to favor desorption. The bound protein is recovered by Protein A Sepharose affinity chromatography. Protein G Sepharose affinity chromatography and protein G Sepharose affinity chromatography are used to isolate the Fc region of an antibody. Protein A or G is commonly used for antibody purification due to its high binding affinity and specificity for the In some embodiments, the one or more affinity isolation methods of step d are , Protein G Sepharose affinity chromatography and Protein A Sepharose parenteral The methods include one or more of the following:
[0061] In some embodiments, step d also includes the step of using antigen-specific affinity chromatography. The selection of antigen-specific nanobodies was performed using the ELISA kit and the antibody was then subjected to various levels of stringency. Elution of specific nanobodies from the original sample, thereby creating different nanobody fractions. and performing steps e through i separately for each fraction; of the CDR3, CDR2 and / or CDR1 region sequences of each different step i for the antigen. The affinity was measured for the CDR3, CDR2 and CDR3 fragments of each nanobody fraction, respectively. and / or a functional selection step based on the relative abundance of CDR1 region sequences. In some embodiments, antigen-specific affinity chromatography includes In some embodiments, the antigen-specific affinity chromatography is performed using a conjugated resin. Raffy is a resin bound to maltose binding protein and antigen.
[0062] The term "degree of stringency" refers to the degree of stringency of different concentrations of salt buffer (e.g., neutral pH). About 0.1 M to about 20 M MgCl in H buffer, preferably about 1 M to about 20 M MgCl in a neutral pH buffer. 10 M MgCl, or preferably about 1 M to about 4.5 M MgCl in a neutral pH buffer 2), alkaline solutions of different pH values (e.g., 1 to 100 mM NaOH, pH approximately 11, 12 and 13), acidic solutions of different pH values (e.g., 0.1 M glycine, pH about 3, 2 and and 1), or a combination thereof, are understood to refer to and are contemplated herein. The terms "different nanobody fractions" or "different biochemical fractions" should be used to refer to different The nucleic acid that is eluted from the antigen-bound solid support (e.g., resin) under a certain degree of stringency is It should also be understood that the term "high salt, high acid, or high alkaline" refers to different fractions of a serotonin. The nanobody that is most resistant to these conditions will have the highest affinity for the antigen.
[0063] The term "digestion products" as used herein, such as in step e, refers to the digestion of Digestion step with trypsin, chymotrypsin, LysC, GluC, and AspN In some embodiments, the nanobody is purified by trypsin (Pi erce™ Trypsin Protease, MS Grade, Catalog Number: 90057 etc.), chymotrypsin (Pierce™ chymotrypsin protease (TLCK process) (Processed), MS grade, catalog number: 90056, etc.) , LysC (or Pierce™ Lys-C Protease, MS Grade, Catalytic Lys-C protease (e.g., log number 90051), GluC (or Pierce Glu-C Protease, MS Grade, Catalog Number: 90054 uC protease), and / or AspN (or Pierce™ Asp- Asp-N-Protease, MS Grade, Catalog Number: 90053 trypsin, chymotrypsin, LysC, GluC, and AspN are enzymes that digest proteins. The cleavage rules for nanobody digestion by are as follows: Trypsin: No K / R or P residues at the C-terminus Chymotrypsin: W / F / L / Y, P not continuing from C-terminus GluC: No D / E or P continues from the C-terminus AspN: N-terminus to D LysC: C-terminus to K The digestion step is carried out at temperatures between about 2°C and about 60°C (e.g., about 2°C, 4°C, 6°C, 8°C, 10°C). , 12℃, 14℃, 16℃, 18℃, 20℃, 22℃, 24℃, 26℃, 28℃, 30℃ , 32℃, 34℃, 36℃, 38℃, 40℃, 42℃, 44℃, 46℃, 48℃, 50℃ , 52°C, 54°C, 56°C, 58°C, or 60°C) for approximately 5 minutes, 10 minutes, or 30 minutes. ,45 minutes,1 hour,2 hours,hours,4 hours,6 hours,8 hours,10 hours,12 hours,1 4 hours, 16 hours, 18 hours, 20 hours, 22 hours, 24 hours, 36 hours, 48 hours, or or 72 hours. [Table 1]
[0064] Step f includes performing mass spectrometry of the digestion products to obtain mass spectrometry data. Methods for using mass spectrometry for peptide analysis are well known in the art. In embodiments, the mass spectrometry herein may be performed by gas chromatography (GC-MS), liquid chromatography (LC-MS), or the like. Chromatography (LC-MS), Capillary Electrophoresis (CE-MS), Ion Transport Analysis by ion beam spectrometry (IMS / MS or IMMS), matrix-assisted laser desorption ionization MALDI-TOF, surface-enhanced laser desorption / ionization (SELDI-TOF), This step is performed in combination with tandem MS (MS-MS). The mass of the amino acid and the data of the polypeptide translated from the cDNA library in step b Based on a sequence homology search in the database, the nanobody or nanobodies in the sample were identified. In some examples, a separate sequence from each Nanobody fraction can be identified. Mass spectrometry is used to analyze and generate spectra of the digestion products. In some examples, the spectrum of the digestion products is shown as an intensity vs. m / z (mass-to-charge ratio) plot and represents the electron ionization data present as a function of the
[0065] It is noted herein that the sequencing of nanobodies is not solely based on mass spectrometry. It should be understood that the sequence identified by mass spectrometry may be used in sequencing. Determined by matching / correlating with sequences from cDNA libraries identified by sequencing The matched sequences are then selected. Thus, step g is a step of analyzing the mass spectrometry data. and step h comprises selecting the sequence identified in step c that correlates with step The method includes identifying the sequence of the CDR3 region in the sequence from g.
[0066] Step i comprises: This involves selecting sequences that achieve a required percentage of fragmentation coverage or higher. In terms of morphology, the percentage of fragmentation coverage was approximately 30 for trypsin-treated samples. % (e.g., about 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 7 0%, 75%, 80%, 85%, 90%, 95%, or 99%) or more. In this embodiment, the fragmentation coverage percentage is at least about 30% (e.g., at least about 30%, 35%, 40%, 45%, 50%, 55% , 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 99%) In some embodiments, the fragmentation coverage percentage is determined by trypsin treatment. Approximately 50% for the digested samples and approximately 40% for the chymotrypsin-treated samples is.
[0067] In some embodiments, the methods described herein comprise: Furthermore, nanobodies can be generated that contain CDR3, CDR2 and / or CDR1 regions that The nanobody gene is cloned into a vector, which is then transfected with the nanobody gene. The cells are transformed into competent cells for protein expression, extraction, and purification.
[0068] In some embodiments, the nanobody is selected from the group consisting of SEQ ID NOs: 1-157. and at least 80% (e.g., at least about 80%, 85%, 90%) of a sequence selected from In some embodiments, the amino acid sequence is 95%, 98%, or 99% identical to the amino acid sequence of the first amino acid sequence. The nanobody has a sequence selected from the group consisting of SEQ ID NOs: 1-157. In some embodiments, the nanobody is selected from the group consisting of SEQ ID NOs: 158-2536. and at least 80% (e.g., at least about 80%, 85%) of a sequence selected from the group consisting of 90%, 95%, 98% or 99% identical amino acid sequence. In embodiments, the nanobody is selected from the group consisting of SEQ ID NOs: 158 to 2536. In some embodiments, the nanobody has a sequence selected from SEQ ID NO: 2665 to 2667 and at least 80% (e.g., at least amino acid sequences that are about 80%, 85%, 90%, 95%, 98%, or 99% identical to In some embodiments, the nanobody comprises SEQ ID NOs:2665-26 67.
[0069] The present invention relates to an amino acid sequence selected from the group consisting of SEQ ID NOs: 158 to 2536. Also disclosed herein are PDZ-specific Nanobodies comprising the amino acid sequence SEQ ID NO: 1. PDZ-specific nanoparticles comprising an amino acid sequence selected from the group consisting of NO: 143 to 157 As used herein, "PDZ" refers to a DHR (Dlg homologous region). domain) or GLGF (glycine-leucine-glycine-phenylalanine) domain It refers to an 80-100 amino acid domain found in signaling proteins called P DZ domains bind to short C-terminal regions of other specific proteins. PDZ domains , which are conventionally divided into three distinct classes classified according to the chemical nature of the ligands The different ligand classes target the penultimate COOH residues found on the terminal COOH of the target protein. They are distinguished by differences in binding residues. Type I domains have the sequence XS / TX-Φ* (herein where X = any amino acid, Φ = hydrophobic amino acid, *COOH terminus). The type III domain binds to a ligand with the sequence X-Φ-X-Φ*. The binding specificity within each domain class is determined by the variant (X) residues. It can be provided by groups and residues outside the standard binding motif. DZ domains do not fall into any of these specific classes. Proteins involved include Erbin, GRIP, Htra1, Htra2, Htra3, and PSD. -95, SAP97, CARD10, CARD11, CARD14, PTP-BL, and In some embodiments, PDZ includes, but is not limited to, SYNJ2BP. The domain is derived from SYNJ2BP.
[0070] Disclosed herein are GST-specific Nanobodies comprising the amino acid sequences of Table 4. The specification also provides an amino acid sequence selected from the group consisting of SEQ ID NOs: 1-98. GST-specific nanobodies containing "glutathione S-transferase" or or "GST" as used herein refers to glutathione-S-transferase (GST). This refers to the interaction of a wide variety of endogenous and exogenous electrophilic compounds with glutathione (GSH). In some embodiments, the GST polypeptide is a family of phase 2 detoxification enzymes that catalyze conjugation. The polypeptide is in the pGEX6p-1 vector.
[0071] Disclosed herein are HSA-specific Nanobodies comprising the amino acid sequences of Table 5. The specification also provides an amino acid sequence selected from the group consisting of SEQ ID NOs: 99-142. Disclosed are HSA-specific nanobodies comprising a sequence. " refers herein to the polypeptide encoded by the ALB gene. In this embodiment, the HSA polypeptide is identified in one or more publicly available databases. The following were identified: HGNC:399, Entrez Ge ne:213, Ensembl:ENSG00000163631, OMIM:1036 00, UniProtKB:P02768. In some embodiments, the HSA polynucleotide is The peptide may have the sequence of SEQ ID NO: 2668, or SEQ ID NO: 266 8, a polypeptide having about 80%, about 85%, about 90%, about 95%, or about 98% homology with SEQ ID NO: 2668. The HSA polypeptide of EQ ID NO:2668 is an immature form or precursor to mature HSA. The HSA portion of SEQ ID NO:2668 is used herein to represent the process form. The mature or processed portion of a polypeptide may be included.
[0072] Here, we present a large-scale quantitative analysis of the antigen-bound Nb proteome and the antigen-Nb complexes. A robust process for epitope mapping and characterization based on high-throughput structural characterization of A teomics pipeline was developed. [Example]
[0073] Example 1. Superiority of chymotrypsin in large-scale Nb proteomics analysis HcAb(V H The variable domains of the H / Nb cDNA library were cloned into two laminae The genomic DNA was amplified from B lymphocytes of glamas and subjected to next-generation genome sequencing (NGS). Kosky, 2013) recovered 13.6 million unique Nb sequences in the database. Approximately 500,000 Nb sequences were aligned to generate sequence logos (Figs. 1A and 7A). The CDR3 loop has both the greatest sequence diversity and sequence length variation, making it a valuable tool for Nb identification. It provides excellent specificity (Figures 1B and 1C). In silico analysis of the Nb database shows that Nb Due to the limited number of trypsin cleavage sites on the peptide, trypsin primarily digests the large CDR3 peptides. As a result, most of the CDR3 residues (77 % is covered by large tryptic peptides of >2.5 kDa (Fig. 1D, 1E) and therefore were not optimal for proteomic analysis (Fig. 7B). This results in the discovery of a protein that is rarely used in proteomics, which cleaves specific aromatic and hydrophobic residues. Motrypsin appears to be more suitable (Methods, Fig. 1A, 7B). 91% can be covered by chymotryptic peptides smaller than 2.5 kDa (Figure 1D , 1E). Random selection and simulations showed that chymotrypsin was more effective than trypsin. We confirmed that the two methods could cover significantly more CDR3 sequences (Fig. 1F). There is little overlap between the enzymes (approximately 9%), demonstrating excellent complementarity for efficient Nb analysis. Ta.
[0074] The estimated false discovery rate (FDR) for CDR3 identification is low due to the large database size and Nb distribution. The abnormal column structure may be causing the increase in size. Specific HcAbs were proteolyzed with trypsin or chymotrypsin and analyzed for identification. Using the Edge search engine, we searched two different databases: immunized and unimmunized llamas. The specific "target" database from which it is derived, and unrelated sequences that do not have the same sequence literally A similarly sized "decoy" database from a variety of llamas was used (Figure 7D). , all CDR3 peptides identified from the decoy database search were considered false positives (E lias, JE & Gygi, SP, 2007). Decoy database search From the search, a large number of false-positive CDR3 peptides were nonspecifically identified. Spectral matching is generally performed by using MS / MS fragments on the CDR3 fingerprint sequence. These mismatches were found to be poorly segmented (Figures 7E and 7F). The majority (95%) of the CDR3 fragments were identified by high-resolution CDR3 analysis in the MS2 spectra (Figures 1K and 1L). 50% (by trypsin, Figure 1G) and 40% (by chymotrypsin, Figure 1G) of the fragments We implemented a simple fragmentation filter that requires a minimum coverage of 1H. The filter is a new tool for reliable Nb proteomics analysis. Before integration into the open source software "Augur Llama" (Figures 8A-8C). The fragments were then further optimized based on the length of the CDR3 (Figures 1I and 1J).
[0075] Example 2. Development of an integrated proteomic pipeline for Nb discovery and characterization Comprehensive quantitative Nb proteomics and high-throughput structural characterization of antigen-Nb complexes A robust platform for the evaluation of lactic acid bacteria is presented here (Methods, Figure 2A). A camelid is immunized with an antigen of interest. The blood and / or prepared a Nb cDNA library from bone marrow (Fridy, 2014). Run S and 10 7 Creating a rich database of over 100 unique Nb protein sequences On the other hand, antigen-specific V H H is affinity isolated from serum and treated with salt or p Elution was performed using a step gradient of H buffer. Nanoflow liquid chromatography coupled with high-resolution MS For chromatographic identification and quantification, the fractionated HcAb was purified by trypsin or The Nb CDR peptides were efficiently digested with chymotrypsin to release the Nb CDR peptides. The first candidates that passed the search were annotated for CDR identification. The image prints were filtered to remove false positives and to identify these various biochemical fractions. The abundance of ATP was quantified to estimate Nb affinity and assemble into Nb proteins. The above steps were automated using Augur Llama. This pipeline allows: This will enable unprecedented levels of diverse, specific, and high-quality Nb identification and characterization. High-throughput computational methods are being used to enable the structural analysis of tens of thousands of antigen-Nb interactions. Docking (Schneidman-Duhovny, 2005), cross-linking Mass spectrometry (CXMS)(Chait, 2016;Rout, 2019;Yu, 20 18; Leitner, 2016), and developed a robust method that integrates mutagenesis. Furthermore, we use deep learning to learn latent features related to Nb repertoires. A learning approach was developed.
[0076] Example 3. Robust, detailed, and high-quality identification of antigen-specific Nbs To validate this pipeline, three benchmark antigens were selected: Glutathione S-transferase (GST), human serum albumin (HSA) (important (Larsen, 2016)) and mitochondrial outer membrane protein 25 These antigens are small PDZ domains. Although PDZ alone is weakly immunogenic, This extends to the immune response (Figure 2B), making it ideal for assessing the robustness of the technology.
[0077] Here, 64,670 unique Nb GST Sequence (3,453 CDR3 Nb families) 9,915 unique CDR combinations from Lee), 34,972 unique Nb HSA (7,749 unique CDRs from 2,286 unique CDR3 Nb families), and A smaller cohort of 2,379 high-quality Nb PDZ Sequence (230 CDR3 fragments) We identified 495 unique CDRs from the genus Mycobacterium (Methods, Figures 2C, 8G). Among various proteases, chymotrypsin provided the most useful fingerprint information for Nb identification. The Nb repertoire was found to provide exceptionally high CDR3 diversity (Fig. 2D, 2E). The results showed a variety of patterns (Figure 8D).
[0078] A random set of 146 Nbs was selected from three antigen-specific Nb groups and tested against E.coli A group of 130 Nbs (89%) showed excellent solubility and were easily and abundantly expressed. The purified antibody was then purified (Figure 2F). To assess antigen binding, immunoprecipitation and ELISA were performed. We employed complementary approaches including SA and SPR (Methods, Figs. 2G, 9C, 9D, 1 0, Tables 1-3). Nb identified by trypsin and chymotrypsin were of equal quality. (Figure 8H). 86.2% (CI 95% :6.8%), 90.5% (CI 95% : 11.5%), and 100% pure Nb binders with GST, HSA, and PDZ, respectively. These results demonstrate the high sensitivity and specificity of this approach. are.
[0079] Example 4. Accurate large-scale quantification and clustering of Nb proteomes To accurately classify Nbs based on their affinity, various strategies were evaluated. , antigen-specific HcAb is affinity isolated from serum and purified using a stepwise high salt gradient, high pH buffer, or low pH buffer (Methods, Fig. 8I, 8J). The activity was accurately quantified by label-free quantitative proteomics (Zhu, 2014). 010; Cox, J. & Mann, M, 2008). Then, the CDR3 peptide The ions (and corresponding Nb) were classified into three groups based on their relative ionic intensities. This classification was confirmed by the high pH method. G ST 31% of Nb HSA 47% of the total were assigned to the C3 high affinity group (Figure 3C). Some Nbs with unique CDR3 sequences from STAR GST is randomly expressed, and These affinities were analyzed by ELISA and SPR ( 2 = 0.85, Figure 3D, Table 1). Various fractionation methods were evaluated. The low pH method provided sufficient resolution to separate the different affinity groups. Although the salt gradient method and especially the high pH method did not provide significant and reproducible separation of Nb, , based on their affinity (Figure 3E). High pH clusters 1 and 2 (C1, Nb from C2) generally have low and mediocre affinities ranging from μM to tens of nM, respectively. However, more than 50% of C3 was an ultra-high affinity, sub-nM binder (Figs. 3H, 9D). To further verify this result, 25 Nb HSA Random set of various CDs R3) were purified from C3 and ranked for their ELISA affinities ( Fig. 3F , Table 2). The top 14 Nb HSA were selected for SPR measurements. Eleven of them showed various results. The remaining three Nb had affinities in the tens to hundreds of pM range with combined reaction rates. HSA teeth, Single-digit nMK D (Figs. 3I, 10A). 13 soluble Nb PDZ and refine it. The high affinity of these antibodies was confirmed by ELISA and immunoprecipitation (Fig. 3G, 10B, and Table 3). ) Representative highly soluble Nb PDZ P10 K D was 4.4 pM (Fig. 3J).
[0080] Immunoprecipitation and fluorescence imaging of native mitochondria (Nb PDZ ) for ultra-high affinity Nb(Nb GST ) (Figures 3K and 3L) were further evaluated. This allows for large-scale and accurate characterization of Nb proteomes based on desirable properties such as affinity. It can be classified as:
[0081] Example 5. Randomization of antigen-binding Nb proteome revealed by integrated structure determination methods Doscape Identification and classification of a large repertoire of high-quality Nb allows for antigen-mediated humoral immunity. This allows investigation into the overall structural landscape of the response. 34,972 Nb H SA Structural docking and clustering of the three major HSA epitopes The abundant natural serum albumin (76% identical to HSA, Figure 1H) ) has made it possible to investigate the specificities of humoral immunity in camelids. The albumin sequences were aligned and their variations were calculated based on pI and hydropathy. All three epitopes correspond to large sequence differences (Methods, Fig. 4A). This result indicates that the antigenicity of Nb is consistent with that of the major peaks of pI and hydropathy. Nb exhibits exceptional specificity of recognition. Nb preferentially binds to stable helical secondary structures. The epitope appears to be highly charged (Fig. 4B). E1 and E3 were predominantly negative (net formal charges of -4 and -5, respectively, Figure 13D), whereas E1 was more heterogeneous with mixed charge (net formal charge of −2) (Fig. 4C).
[0082] 19 HSA-Nb complexes (Shi, 2014; Kim, 2018) were cross-linked. The epitopes identified by docking were validated. Overall, 92% of the cross-links were model The cross-linking was satisfied by the cleavage of the hydroxyl group, with a median RMSD of 5.6 Å (Figures 4J and 4K). The docking results were confirmed by the docking method, and two closely spaced epitopes (E2 and E3) (6 E1 was identified by low abundance cross-linking. Cross-linking revealed two additional cleavage sites not revealed by docking. We also identified minor epitopes in the convex Nb paratope and the concave HSA epitope (Fig. 4D). High shape complementarity was observed between HSA and Nb, including the loop (Fig. 4E-4G). To further characterize E2, single-point immunoprecipitation was performed on HSA with minimal impact on the overall structure. The mutation E400R was introduced (Pires, 2016). The resulting mutation was , surface charge to mimic the positive charge at the orthologous position of E2 in camelized animal albumin It is possible to reverse the ATPase activity and disrupt the salt bridge formed between it and the arginine of Nb CDR3. 19 high-affinity binders were then selected and analyzed for their ability to bind HSA-Nb. The effect of this point mutation on the activity of the E400R was assessed by ELISA (Figure 4I, Table 2). almost completely abolished binding of 5 of 19 Nbs (26%) tested, and E2 This showed that it was a genuine major epitope.
[0083] This approach was applied to map epitopes of 64,670 GST-Nb complexes. Three major epitopes on GST were precisely identified (Fig. 11A, 11B, 11F, 11G), and 18.7 for E1, E2, and E3, respectively. The cross-linking was performed at relative abundances of 5%, 31.25%, and 50% (Fig. 11D, 11 E). E1 and E3 contain negatively charged surface patches. E2 is a GST dimerization cavity. In the model presented here, E2 Nb occupies this cavity with its Insert CDR3. Similar to HSA, it has a preference for charged surface residues and a high shape factor for Nb. Collectively, these results demonstrate that Nb binds to a variety of protein surfaces. , indicating a preference for highly charged cavities on the antigen.
[0084] Example 6. Investigation of the mechanism of Nb affinity maturation Based on the high pH data set, the most robust classification was high affinity (mature) and low affinity The physicochemical and structural characteristics that distinguish Nb from HSA and GST were investigated. A shorter CDR3 with a different distribution of high-affinity binders (Figure 5A) exhibits enhanced antigen binding The entropy of the low-affinity Nb ranges from slightly acidic to relatively salty for the high-affinity Nb. A significant increase in pI was observed up to baseline (Figure 5B).
[0085] The contribution of CDRs to the pI and hydropathy of Nb was compared, and CDR3 HSA is Nb HSA is the main cause of polarity shift in CDR1 GST and CDR2 GST is Nb GST We determined that the hydrophilic Nb was the main cause of the polarity shift (Fig. 5C). was observed to be slightly higher (Fig. 5D).
[0086] The structure of the CDR3 consists of a "head" region that contains the highest sequence variability and a "head" region that contains the lower specificity. The posterior cingulate cortex can be thought of as having a "torso" region (Finn, 2016) (Figure 5E). Aspartic acid and arginine (which form strong electrostatic interactions) (Tiller, 2017), small, flexible residues such as glycine and serine, and hydrophobic residues such as alanine and leucine. Certain residues are enriched in the CDR3 head, including aromatic residues, as well as tyrosine. Comparison of Nbs from different affinity groups revealed three major differences (Fig. 5F and Fig. 12). First, high-affinity Nbs were more abundant in charged residues (Mitchell et al., 2004). ll, LS & Colwell, LJ, 2018) (Methods, Figure 5G). No. Second, we identified complex differences for various antigens. HSA is the CDR3 header Increase positively charged residues on the nucleotide sequence (39%) and decrease negatively charged residues (46%) This tends to strengthen the electrostatic charge. GST is mainly the power of other CDRs. In CDR1 and CDR2, 29.2% and 30.2% of the positively charged residues were positively charged, respectively. There was a 117.2% increase in the number of negatively charged residues and a 44.2% and 21.5% decrease in the number of negatively charged residues. The charge change may increase the physicochemical complementarity between the Nb and the epitope. Third, tyrosine (51%), glycine, and serine (58%) are high-affinity Nb HSA of High affinity Nb was more enriched in the CDR3 head. GST So, CDR3 head Ciro The glycine and serine fractions were almost unaffected. There wasn't.
[0087] To further explore the putative role of these residues in enhancing HSA binding affinity Then, their positional frequencies were calculated along the CDR3 head (Figure 5H). compatibility Nb HSA It is more frequently found in the center of the CDR3 head of into the specific epitope pocket(s) (Desmyter , 1996; Li, 2016). Glycine and serine are located away from the center of CDR3. This provides additional flexibility and allows for orientation of the tyrosine side chain within the antigen pocket. These results are based on the number of residues and the ELISA parental control of the purified Nb in this study. This was confirmed by correlation analysis between the serotonin levels and the serotonin levels (Figures 5I and 5J).
[0088] Develop a deep learning model to learn latent features that enable Nb affinity classification The most informative Nbs for high-affinity binder classification were HSA CDR3 Filter revealed a pattern of consecutive lysines and arginines, tyrosines, and glycines. (Figure 5K, Table 4). For low affinity binders, the most beneficial filter was phenylalanine. The analysis also prioritizes amino acids, histidine, and two consecutive aspartic acids. , negative and positive charge pairs for high affinity and low affinity binders, respectively. This revealed a tendency for A to be continuous.
[0089] Example 7. Superior versatility and resilience of Nbs for antigen recognition Hundreds of divergent, high-affinity Nbs against weakly immunogenic PDZ domains CDR3 Fami The identification of Li prompted investigation of the structural basis of such interactions. Two putative epitopes were identified (Fig. 6A, 13B). E2 is a large, positively charged It has a smooth surface (Fig. 6A, 6B) and is more structured with an α-helix and two β-strands. E2 has a large number of PDZ-interacting proteins, making it a potential major epitope. overlapped with a conserved ligand-binding site shared by the ; Doyle, 1996) (Figure 6C). Surprisingly, Nb PDZ is a natural PDZ It has 100,000 times higher affinity (μM affinity) than the ligand (Nie Such high affinity is due to the small and shallow epitaxial structure (Fig. 3J). This is achieved by the long CDR3 loop wrapping around the tope, resulting in extensive electrostatic interactions and The modeling results indicated that P R46 and K48 of the second β-strand of the DZ epitope are PDZ The corresponding remainder The double mutant PDZ (R46E:K48D) was generated and Nb PDZIts affinity for Nb was assessed by ELISA. PDZ The majority of (8 / 11) showed significantly reduced or no affinity for the mutants, and E2 This confirmed that it was indeed the major epitope (Fig. 6D).
[0090] Nb PDZ There are several other observations about the CDR3 loop length. The distribution forms one major peak, with the median at about 20°C, pushing the upper limit of its natural distribution. aa (Fig. 6E). Second, Nb PDZ is slightly acidic with a median pI of 4.9 (Fig. 6F), to which CDR3 contributes significantly (Fig. 6E, 13F). Despite their acidic nature, Nb PDZ is hydropathic at the expense of hydrophobic residues. However, the negative effects of α-glucan on the ATP-dependent ATP synthesis did not appear to appreciably alter the ATP-dependent ATP synthesis (Figs. 6G, 13E). The positively charged aspartic acid and small glycine and serine were significantly increased, resulting in the CDR3 header. High affinity Nb GST and Nb HSA Compared to the bulky Tyrosh A decrease in ATP was also evident, reflecting a fairly shallow pocket in E2 for binding (Figure 7 C, 7E). Collectively, these results demonstrate the remarkable versatility of Nbs for antigen binding. is doing.
[0091] In this study, we used proteomics, informatics, and report the development of a robust platform that integrates analytical and structural modeling technologies. The pipeline offers a broad repertoire of high-quality Nb with high sensitivity against a variety of challenging antigens. It also allows for accurate and reliable identification of circulating Nb based on its physicochemical properties. Thousands of ultra-high affinity Nbs have been identified using this technique. In this paper, we combined computational docking and structural proteomics to identify 102,673 antibodies. The protease-Nb complex was structurally characterized and mapped, and the major epitopes were verified. "Big data" analysis provides the first global proteomic and structural analysis of the humoral immune response. Make it possible.
[0092] These results provide unprecedented depth into the vast landscape of camelid antibody immunity. The efficiency, specificity, diversity, and versatility of the antigen-binding Nb formed together were revealed (Figure 6 H).
[0093] Efficiency: Nb efficiently utilizes both shape and electrostatic complementarity for binding. aspartic acid and arginine, aromatic tyrosine, and small, flexible glycine and Specific residues such as serine allow flexibility in the loop resulting in high affinity Nbs. We have revealed a complex and finely tuned interaction specific to the CDR. Multiple preferential binding sites for Nb act as a general mechanism for efficient recognition. The presence of a specific epitope was confirmed (Akram, A. & Inman, RD, 2 012).
[0094] Specificity and diversity: To ensure a specific, effective, and safe immune response, several Thousands of proteins have evolved to recognize specific HSA surface pockets with the most prominent sequence variations We found as many highly branched Nb as possible (Fig. 4A).
[0095] Versatile: For antigens that tend to evade immune responses, such as PDZ, Nb can be used as a paratope The size and physicochemical properties of the This work highlights the fascinating and rapid evolution of protein-protein interactions. It shows.
[0096] Nb is highly potent in virus neutralization and inhibition of enzyme activity (Lauweer ys, 1998;Desmyter, 1996;Acharya, 2013;Ar abi, 2017). These findings highlight the importance of these highly robust and efficient camelid H cAb has the potential to survive both in dry natural habitats and under aggressive pathogenic challenges. It shows that there is an evolutionary advantage to such incredible selection and adaptation. The driving force(s) behind this remain a mystery (Flajnik, 2011).
[0097] These techniques address challenging biomedical applications such as cancer biology, brain research, and virology. These methods for Nb proteomics can find wide application in various applications. Informatics tools are freely available to the research community. The dataset serves as a blueprint for studying antibody-antigen interactions, allowing for computational antibody design. (Sircar, 2011; Baran, 2017; C hevalier, 2017).
[0098] Example 8. Methods animal immunization Two llamas were analyzed for HSA and mitochondrial outer membrane protein 25 (OMP25), respectively. The initial dose of 1 mg of GST and the GST-fused PDZ domain was administered, followed by Three consecutive boosts of 0.5 mg were administered every 3 weeks. Blood samples and bone marrow aspirates were collected from the final The animals were extracted 10 days after the boost. All the above procedures were carried out according to the IACUC protocol. Implemented by Capralogics, Inc. in accordance with Coll.
[0099] mRNA isolation and cDNA preparation Approximately 1~3×10 9 Peripheral mononuclear cells were isolated from 350 ml of immune blood, with 5-9 x 10 7 Plasma cells were isolated from 30 ml of bone marrow aspirate using a Ficoll gradient (Sigma). The mRNA was isolated from each cell using the RNeasy kit (NEB). and lysed it with Maxima™ H Minus cDNA Synthesis Master Mix (T The cDNA was reverse transcribed using the ELISA kit (hermo). Camelid IgG heavy chain cDNA sequences were amplified using primer CALL001 (GTCCTGGC TGCTCTTCTACAAGG, SEQ ID NO: 2646) and CH2FORT A4(CGCCATCAAGGTACCAGTTGA, SEQ ID NO:2647) The V lacking the CH1 domain was specifically amplified using the . H H The gene was isolated from conventional IgG and purified by DNA gel electrophoresis (Qiagen). , then the second forward (ATCTACACTCTTTCCCTACACGACG CTCTTCCGATCTNNNNNNNNATGGCT[C / G]A[G / T]GTG CAGCTGGTGGAGTCTGG, SEQ ID NO:2648, N is A, T, C or G) and second reverse (GTGACTGGAGTTCAGACGTGT GCTCTTCCGATCTNNNNNNNNGGAGACGGTGACCTGGGT, SEQ ID NO: 2649, where N represents A, T, C, or G) to identify the Frameworks 1 to 4 were re-amplified. Cluster identification on Illumina MiSeq To aid in this, random 8-mer replacement adapter sequences were added. The amplicons (approximately 450-500 bp) were purified using the Monarch PCR cleanup kit. The primers were purified using a kit (NEB). CGACCACCGAGATCTACACTCTTTCCTA, SEQ ID NO: 2650) and MiSeq-R (CAAGCAGAAGACGGCATACGAGATT TCTGAATGTGACTGGAGTTCA, SEQ ID NO: 2651 A final round of PCR was performed to isolate the indexed PNPs prior to MiSeq sequencing. 5 / P7 adapter added.
[0100] Next-generation sequencing with Illumina Miseq Sequencing was performed on an Illumina MiSeq platform with a 300bp paired-end model. Each database generated over 30 million leads. FastQC v0.11.8 was used for quality check and management of FASTQ data. Lead QC tool (www.bioinformatics.babraham.ac The raw Illumina reads were analyzed using BB The software tools of the Map project (github.com / BioInfoTo The nucleotide sequence was converted to an amino acid sequence using the . Previously, duplicated reads and DNA barcode sequences were sequentially removed.
[0101] V from immune serum H Isolation and biochemical fractionation of H antibodies Approximately 175 ml of plasma was extracted from 350 ml of immunized blood by Ficoll gradient (Sigma). It was isolated from the liquid of camelids. H H antibody is a serotype of Protein G and Protein A. Plasma supernatant was purified by a two-step procedure using alpharose beads (Marvelgent). After isolation and elution with acid, the product was neutralized with 1x PBS buffer and diluted to a final concentration of 0.1-0.3 m Antigen-specific V H To purify H antibodies, GST or HSA conjugates were used. The CNBr resin was then H Incubate with H mixture at 4°C for 1 hour. Wash thoroughly with high salt buffer (1x PBS and 350 mM NaCl) to remove nonspecific The specific V was then extracted using one of the following elution conditions: H H antibody from resin That is, the cellulose was released from alkaline (1-100 mM NaOH, pH 11, 12, and 13), acidic (0.1 M glycine, pH 3, 2, and 1) or salt elution (neutral pH buffer PDZ-specific V H For the purification of H, MBP- PDZ fusion protein (P to avoid steric hindrance of small PDZ after coupling) The parent DZ domain was fused to maltose-binding protein (MBP) at its N-terminus. As a control, MBP-bound resin was used (Fig. 6J). Before the chromatographic analysis, all eluted V H H were neutralized and dialyzed separately into 1x DPBS.
[0102] Nanoflow liquid chromatography coupled to proteolysis and mass spectrometry of antigen-specific Nb Graphy (nLC / MS) analysis GST and HSA V H For H, each elution was treated separately according to the following protocol: . PDZ-specific V H For H, the most stringent biochemical elution (i.e. pH 13, pH 1, MgCl2 3M and 4.5M) and from different fractions, respectively Only these non-specific MBP binders (negative controls) were pooled for proteolysis. For example, PDZ-specific V eluted by pH 13 buffer H In the case of H, nonspecific MBP Bound Nb was pooled from the pH 11, pH 12 and pH 13 fractions and subjected to downstream LC Improved stringency of V / MS quantification. H H in 8M urea buffer (50 mM bicarbonate The mixture was reduced in ammonium chloride, 5 mM TECEP, and DTT at 57°C for 1 hour in the dark. The mixture was alkylated with 30 mM iodoacetamide at room temperature for 30 minutes. The diluted sample was divided into two halves and purified in solution using trypsin or chymotrypsin. For trypsin-digested samples, trypsin and Lys- Add 1:100 trypsin and digest overnight at 37 °C. Add 1:100 trypsin the next morning and digest at 37 °C. The samples were digested in a water bath at 1:50 (w / w) for 4 hours. Trypsin was added and digested for 4 hours at 37°C. After proteolysis, the peptide mixture was Desalting was performed using a self-packed stage tip or a Sep-pak C18 column (Waters). , Q Exactive(TM) HF-X Hybrid Quadrupole Or It was coupled online to a Bitrap™ mass spectrometer (Thermo Fisher). The desalted Nb peptides were analyzed using a nano-LC 1200. Analysis column (C18, particle size 1.6 μm, pore size 100 Å, 75 μm × 25 cm, Load onto IonOpticks and run a 90-minute liquid chromatography gradient (5% B to 7%B, 0~10 minutes; 7%B~30%B, 10~69 minutes; 30%B~100%B, 69~ 77 minutes; 100%B, 77~82 minutes; 100%B~5%B, 82 minutes~82 minutes 10 seconds; 5% B, 82 min 10 s–90 min; mobile phase A consisted of 0.1% formic acid (FA), and mobile phase B Elution was performed using a solution consisting of 0.1% FA in 80% acetonitrile (ACN). The flow rate was 300 nl / min. The QE HF-X instrument was operated in data-dependent mode. The top 12 most abundant ions (mass range 350–2,000, charge states 2–8) were analyzed. The target resolution was determined by MS. The number was set to 120,000 for tandem MS (MS / MS) analysis and 7,500 for tandem MS (MS / MS) analysis. The quadrupole isolation window was 1.6 Th, and the maximum injection time for MS / MS was set to 80 ms. It was determined.
[0103] Synthesis and cloning of Nb DNA. Nb gene in Escherichia coli. The expression of the vector was codon-optimized and nucleotides were synthesized in vitro (Synbiotic). After verification by Sanger sequencing, the Nb gene was cloned into pET-21b(+) BamHI and XhoI (for GST Nb), or EcoRI and NotI restriction The nucleotide sequences were cloned into the nucleotide sequences (for HSA and PDZ Nbs).
[0104] Recombinant protein purification Transform the DNA construct into BL21(DE3) competent cells according to the manufacturer's instructions. The cells were plated on agar plates containing 50 μg / ml ampicillin overnight at 37°C. A single colony was inoculated into LB medium containing ampicillin for overnight culture at 1°C. After this, the culture was inoculated into fresh LB medium at 1:100 (v / v) until OD600nm reached 1. The mixture was shaken at 37°C until the pH reached 0.4-0.6. GST, GST-PDZ, and Nb were added at 0 Induction with 0.5 mM IPTG, and MBP and MBP-PDZ with 0.1 mM IPTG. Induction was carried out overnight at 16°C. Cells were then harvested, briefly sonicated, and stored on ice. Lysis buffer (1x PBS, 150 mM NaCl, 0.2% protease inhibitors) After lysis, the soluble protein extract was centrifuged at 15,000 × g for 10 min. The GST and GST-PDZ were purified using GSH resin and then purified with glutathione. MBP (maltose binding protein) and MBP-PDZ fusion protein were eluted by HPLC. Proteins were purified by using amylose resin and maltodextrin according to the manufacturer's instructions. Nb was purified by His-cobalt resin and eluted using imidazole. The eluted protein was then eluted using dialysis buffer (e.g., 1x DPBS, p H7.4) and stored at -80°C until use.
[0105] Nb immunoprecipitation assay After Nb induction and cell lysis, the cell lysates were subjected to SDS-PAGE to estimate the Nb expression levels. The recombinant Nb in the cell lysis solution was diluted to a final concentration of approximately 5 μM in 1×DPBS (pH 7.4). The antigens were diluted to approximately 50 nM (for GST Nbs) and 50 nM (for PDZ Nbs). Various antigens were conjugated to CNBr resin to test for specific interactions with Inactivated or MBP-conjugated CNBr resin was used as a control. The resin was then incubated with Nb lysate for 30 min at 4°C. The resin was then washed with washing buffer ( Wash three times with 1x DPBS containing 150mM NaCl and 0.05% Tween 20 Then, specific antigen-binding Nb was added to the PBS containing 20 mM DTT. The resulting product was eluted from the resin with hot LDS buffer and subjected to SDS-PAGE. The intensity of b is compared between the antigen-specific signal and the control signal to derive false positive binding. Ta.
[0106] ELISA (enzyme-linked immunosorbent assay) To assess camelid immune responses to antigens and to quantify the relative affinity of antigen-specific Nbs Indirect ELISA was performed using the antigen in a 96-well ELISA plate (R&D system). Add approximately 1-10 ng of the coating buffer (15 mM sodium carbonate) to each well. The mixture was coated overnight at 4°C in 35 mM sodium bicarbonate (35 mM thorium, pH 9.6). Next, the well surface was coated with blocking buffer (DPBS, 0.05% Tween 20, 5% The cells were blocked with PBS (milk) for 2 hours at room temperature. The supernatant was serially diluted 5-fold in blocking buffer. The diluted serum was incubated at room temperature for 2 hours on the antigen-treated plate. HRP conjugate to llama Fc (Bethyl) was incubated with the covered wells. Gated secondary antibodies were diluted 1:10,000 in blocking buffer and mixed with each well. Both were incubated at room temperature for 1 hour. In the Nb affinity test, A scrambled Nb was used as a negative control. The specific binder Nb was serially diluted 10-fold from 10 μM to 1 pM in blocking buffer. Diluted. His tag (Genscript) or T7 tag (Thermo) HRP-conjugated secondary antibodies were added at 1:5,000 or 1:10,000 in buffer. The diluted solution was incubated at room temperature for 1 hour. Nonspecific absorbance was removed between incubations. Wash three times with 1x PBST (DPBS, 0.05% Tween 20) to remove After the final wash, the sample was washed with freshly prepared 3,3',5,5'-tetramethyl Further incubation with rubenzidine (TMB) substrate for 10 minutes at room temperature in the dark After the signal was developed, the plate was read using a stop solution (R&D Systems). kan GO, Thermo Fisher) at multiple wavelengths (450 nm and 550 nm ) and the plate was read. False positive Nb was detected if either of the following two criteria was met: i) ELISA signal was only detectable at a concentration of 10 μM. ii) At a concentration of 1 μM, the signal was significantly higher than at 10 μM. A significant signal reduction (up to 10-fold) was detected at 1000 mg / mL of HCl, but no signal was detected at lower concentrations. The raw data was processed using Prism 7 (GraphPad) to generate 4PL songs. A line was fitted and the logIC50 calculated.
[0107] Nb affinity measurement by SPR Surface plasmon resonance (SPR), Biacore 3000 system, GE Health Nb affinity was measured using the following steps: The antigen protein was immobilized on the sensor chip. Dilute to 10-30 μg / ml with sodium, pH 4.5, and inject into the SPR system at 5 μl / min. The sensor surface was then infused with 1M ethanolamine-HCl (pH 8.0) for 420 seconds. For each Nb sample, the antibody was blocked with HBS-EP + Runny No. 1 containing 2mM DTT. A dilution series (spanning three orders of magnitude) was prepared in 20-30 ml of PBS containing 100% EDTA. Inject for 120-180 seconds at a flow rate of 0 μl / min, and then dissociate for 5-20 minutes based on the dissociation rate. Between each injection, a 10 mM glycine-HCl (pH 1.5-2.5) solution was added. low pH buffer containing 20-40 mM NaOH (pH 12-13) The sensor chip surface was regenerated at a flow rate of 40-50 μl / min for 30 seconds. Measurements were performed in duplicate, and only highly reproducible data were used for analysis. The rhamnosyltransferase was processed and analyzed using a 1:1 Langmuir model or BIAevaluation. was analyzed by fitting with a 1:1 Langmuir model with mass transfer.
[0108] Cross-linking and mass spectrometry of antigen-nanobody complexes Prior to cross-linking, different Nbs were incubated in an amine-free buffer (1x DTT containing 2 mM DTT). PBS) at 4°C for 1-2 hours with an equimolar concentration of the antigen of interest. Amine-specific disuccinimidyl suberate (DSS) or heterobifunctional linkers 1-ethyl-3-(3-dimethylaminopropyl)carbodiimide hydrochloride (EDC ) was added to the antigen-Nb complex at a final concentration of 1 mM or 2 mM, respectively. For cross-linking, the reaction was carried out at 23°C for 25 minutes with constant stirring. EDC cross-linking The reaction was carried out for 60 min at 23°C. For 10 min at room temperature, 50 mM Tris-HCl The reaction was quenched with HCl (pH 8.0). After reduction and alkylation of the protein, Cross-linked samples were run on a 4–12% SDS-PAGE gel (NuPAGE, Thermo Scientific) The region corresponding to the cross-linked species was excised and purified as previously described. The DNA was then in-gel digested with pusin and Lys-C (Shi, 2014; Shi, 2015). After proteolysis, the peptide mixture was desalted and purified using Q Exactive™ HF-X Hybrid Quadrupole-Orbitrap™ Mass Spectrometer (Ther A nano-LC1200 (Thermo Fisher) was coupled to a nano-LC1200 (Thermo Fisher) The cross-linked peptides were analyzed on a PicoTip column (C18, particle size 3 μm, pore size (300Å diameter, 50 μm × 10.5 cm, New Objective) and LC gradient (5% B to 8% B, 0 to 5 min; 8% B to 32% B, 5 to 45 min; 32% B to 100%B, 45~49 minutes; 100%B, 49~54 minutes; 100%B~5%B, 54 minutes~ 54 min 10 sec; 5% B, 54 min 10 sec to 60 min 10 sec; mobile phase A is 0.1% formic acid (FA) Mobile phase B consisted of 0.1% FA in 80% acetonitrile (ACN). The QE HF-X instrument was operated in data-dependent mode, with the top The eight most abundant ions (mass range 380–2,000, charge states 3–7) were analyzed with high energy -Fragmented by collision-induced dissociation (normalized collision energy 27). Target fragmentation The capacity was set to 120,000 for MS and 15,000 for MS / MS analysis. The quadrupole isolation window was 1.8 Th, and the maximum injection time for MS / MS was set to 120 ms. After MS analysis, the data were searched by pLink2 for the identification of cross-linked peptides. (Chen, 2019). Mass accuracy was 1.0 for MS and MS / MS, respectively. 0 and 20 p.m. Other search parameters included the system as a fixed modifier. Carbamidomethylation of methylamine and oxidation of methionine were included as variable modifications. The initial search results were obtained with a default false discovery rate of 5%. The crosslinked spectra were then obtained using the α- and β-paired α-paired β ... were manually checked to remove false positive identifications essentially as previously described (Shi, 20 14;Kim, 2018;Shi, 2015).
[0109] Site-directed mutagenesis The mammalian expression plasmid for HSA was obtained from Addgene. The E400R point mutation , primer HSA-F(GGTGTTCGACCGGTTCAAGCCTCTGG, S EQ ID NO:2652) and HSA-R(TTGGCGTAGCACTCGTGA , SEQ ID NO: 2653) was used to generate the Q5 site-directed mutagenesis kit (N After sequence verification by Sanger sequencing, Lipofectamine 3000 transfection was performed according to the manufacturer's protocol. Wild-type HS was isolated using a fusion kit (Thermo) and Opti-MEM (Gibco). Plasmids containing A and mutants were transfected into HeLa cells. The cells were cultured overnight. After incubation, the medium was replaced with DMEM without FBS supplements to remove BSA. After 48 hours of incubation at 5% CO2, the HSA-expressing medium was collected and stored at -20°C. The medium was analyzed by SDS-PAGE and Western blotting to confirm protein expression. I acknowledged it.
[0110] The PDZ domain (in the pGEX6p-1 vector) was obtained from General Biosystems. The double point mutant of PDZ (i.e., R46E:K48D) was obtained from PDZ. -F(TGATGAAAATGGCGCAGCCGCC, SEQ ID NO:2654 ) and PDZ-R(ATTTCACTCACATAGATACCACTATCATTAC Q was analyzed using specific primers for ATP (TAACATAC, SEQ ID NO: 2655). 5 was introduced using a site-directed mutagenesis kit. After verification, the mutant vector was transformed into BL21(DE3) cells and expressed. The DZ mutant proteins were purified by GSH resin as previously described.
[0111] Fluorescence microscope COS-7 cells were plated on glass-bottom dishes at 60–70% initial confluence. The cells were cultured overnight to allow them to adhere to the dish. with MTMRos (1:4000) for 30 min at 37°C, washed once with pre-chilled PBS The sections were fixed in methanol / ethanol (1:1) for 10 minutes. After washing with PBS, the sections were Cells were blocked with 100% BSA for 1 hour. Then, AlexaFluor™ 647 Conjugated Nb (1:100) was added to the cells and incubated for 15 min at room temperature. Two-color wide-field fluorescence images were taken using 561 nm and 642 nm excitation lasers (MPB Commu Pointe-Claire, Quebec, Canada ) and a 100X oil immersion objective lens (NA=1.4, UPLSAPO 100XO; Olympus Custom-built system on an Olympus IX71 inverted microscope frame with a 1000-megapixel (PUS) was obtained using
[0112] Text-based CDR (complementarity determining region) annotation The CDR annotation method was modified from (Fridy, 2014). [*] indicates optional means a residue of
[0113] CDR1 annotation: A short sequence motif located between residues 20 and 26 of the Nb sequence The first search was for the "SC" motif. The start of the CDR1 sequence was the fifth amino acid sequence followed by the "SC" motif. Once the first residue is identified, the next step is to identify the residues located between Nb residues 32 and 40. We searched for another sequence motif, "W[*]R," which corresponds to the CDR1 sequence, and then we found that the end of the CDR1 sequence was a "W[*]R" motif. is defined as the first residue before the
[0114] CDR2 annotation: The start of the CDR2 sequence is followed by a "W[*]R" motif. Once the first residue is identified, the next step is to determine the residue between Nb residues 63 and 72. Identify the localized motif "RF" and align the end of the CDR2 sequence to the 8th position before the "RF" motif. The residues were defined as the first residue.
[0115] CDR3 annotation: First, the Y[*]C region located between Nb residues 90 and 105 The start of the CDR3 sequence was searched for as "Y[*]" or "YY[*]". The first CDR3 residue is defined as the third residue following a "C" or "YY[*]" motif. Once the residues were identified, the following sequence motifs were then identified: "WG[*]G", "WGQ[*]", W[*]Q[*], "[*]GQG", "[*][*]GQ and "[WG[*][*] ") were used to identify the ends of the CDR3. These motifs are located at the C-terminal N b Located within the last 14 residues of the sequence. CDR3 ends one residue before the sequence motif. For more information, see the Augur Llama script.
[0116] Cleavage rules for in silico digestion of Nb with various proteases: Trypsin: No K / R or P residues at the C-terminus Chymotrypsin: W / F / L / Y, P not continuing from C-terminus GluC: No D / E or P continues from the C-terminus AspN: N-terminus to D LysC: C-terminus to K
[0117] Sequence alignment of Nb database: The sequences of Nb were aligned using the software ANARCI ( Dunbar, J. & Deane, CM, 2016) Three CDRs (CDR1 to CDR3) and four framework sequences (FR1 to FR 4) and annotated according to the IMGT numbering scheme (Lefranc, 2003). Alignments with e-values below a threshold of 100 were removed, and the remaining sequences were compared with WebLo Plotted by go (Crooks, 2004).
[0118] In silico digestion of Nb database with different proteases and Nb CDR3 mapping Analysis of the A high-quality database containing approximately 500,000 unique Nb sequences was analyzed by truncation according to the above cutting rules. Uses a variety of enzymes including trypsin, chymotrypsin, LysC, GluC, and AspN The CDR3-containing peptides were obtained and sequence coverage was calculated. The CDR3 coverage was then summed to generate Figure 1D and Figure 7B. The distribution of trypsin and chymotrypsin-induced protein lengths was plotted to produce Figure 1E.
[0119] Simulation of trypsin- and chymotrypsin-assisted MS mapping of Nb A total of 10,000 Nb sequences with unique CDR3 fingerprint sequences were collected from the database. The selected Nbs were then purified using either trypsin or chymotrypsin. In silico digestion was performed using either (no uncleaved sites allowed) and the CDR3 peptide was To better simulate Nb identification by MS, the following criteria were applied: These peptides were analyzed using the following methods: 1) Size (850) suitable for bottom-up proteomics 2) The highly conserved WGQGQVT peptide was initially selected. Peptides containing the C-terminal FR4 motif of S were further discarded. The peptides were predominantly fragmented by C-terminal y-ions, but had no clear CDR3 fragments. The fragmentation of ions on the CDR3 sequence, which is essential for peptide identification, is often insufficient. 3) CDR3 peptides with limited Nb fingerprint information (less than 30% C As a result, 2,111 unique triplicates were removed. The total number of chymotryptic peptides and 5,154 unique chymotryptic peptides were obtained. The peptides were used to map the Nb protein. Only Nb identifications with high CDR3 fingerprint sequence coverage (≥ 60%) were used. This generated the Venn diagram in Figure 1F.
[0120] Phylogenetic analysis of Nb CDR3 sequences The phylogenetic tree includes unique Nb CDR3 sequences and additional fragments to aid alignment. The CDR3 sequence was inserted with a YYCAA at the N-terminus and a WGQG at the C-terminus. The dataset was generated using Clustal Omega (Sievers, 2014) with the input data. The data is then transferred to the Interactive Tree of Life (ITol) Plotted by BioPy The thon library was used to calculate the isoelectric point and hydrophobicity of Nb CDR3. Alignment is viewed using Jalview (Waterhouse, 2009). It became awakened.
[0121] Assessment of the reproducibility of Nb peptide quantification Reproducibility of label-free quantification methods using shared peptide identifications across different LC runs A typical 90-minute LC gradient allowed the peptide peak width or full width at half maximum (F The WHM was generally less than 5 seconds. The differences in peptide retention times between different LC runs were calculated. The Kernel Density Estimation plot in Figure 3B was generated. Peptide retention from different LC runs Using time, Pearson correlations were calculated and plotted in Figure 9B.
[0122] Sequence alignment and analysis of HSA and llama serum albumin Llama (Camelus Ferus) serum albumin sequence was obtained and analyzed by tblastn(N The isoelectric points (pI) and pIs of individual amino acids were determined by CBI. Hydropathy values can be found at (www.peptide2.com / N_peptide_hy drophobicity_hydrophilicity.php) online These values were normalized between 0 and 1.0 to estimate the sequence variation ( Pairwise differences in pI and hydropathy) were calculated for each aligned position. For a particular aligned residue position, a value of 0 indicates that no identical residues are found between the two sequences. The value 1.0 indicates that the negatively charged residue glutamic acid 400 of HSA is Charge reversal to the positively charged residue arginine at the corresponding alignment position in albumin The maximum sequence variation is indicated by a value of 0.5 at the position where an amino acid insertion or deletion was confirmed. In this way, the pI and hydropathicity between HSA and llama serum albumin were determined. The plot was further smoothed by a Gaussian function to obtain the variability of both sequences. , which generated Figure 4A.
[0123] Analysis of the relative abundance of amino acids on Nb CDRs Calculate the amino acid frequency in each CDR (including CDR1, CDR2, and CDR3 heads) The results were then normalized to generate the bar graphs and pie charts shown in Figures 6, 7, 12, and 13. The head sequence was obtained by removing the semi-conserved C-terminal four residues of CDR3. The CDR residue frequencies of both high-affinity and low-affinity Nbs were calculated by summing the CDR residues in each affinity group. was normalized based on
[0124] Analysis of amino acid positions on the CDR3 head The relative positions of residues on the CDR3 head were calculated, where a value of 0 indicates exactly The CDR3 head sequence is then split into bins with a bin width of 0.0, where 1.0 indicates the N-terminus and 1.0 indicates the last residue. The sample was sliced into 20 bins of 5 mm. Within each bin, a specific type of amino acid (tyrosine, glycine) was The occurrence of nucleotides (e.g., nucleotides, nucleotides, or serines) is counted and compared to the total residues on the CDR3 head. The distribution of different amino acids, including their relative positions and abundances, is shown in Figures 5H and 12G. was plotted.
[0125] Proteomics database of Nb peptide candidates The raw MS data were analyzed using Proteome Discoverer 2.1 (Therm Sequest HT embedded in Fisher (FDR estimation) A standard target-decoy strategy was used to target an in-house generated Nb sequence database. The mass accuracy was 10 ppm and 0 ppm for MS1 and MS2, respectively. Other search parameters included carbamazepine and cysteine as fixed modifications. We included midomethylation and oxidation of methionine as variable modifications. Lipsin-treated samples were allowed to have a maximum of one or two uncleaved sites, respectively. The initial search results were filtered based on the q-value with a 0.01 (strict) FDR percolator. After the database search, the data were filtered (Kall, 2007). r Llama exports peptide spectrum matching (PSM) data in the following steps: were analyzed and processed.
[0126] a. Nanobody identification i) Quality assessment of CDR3 fingerprints Peptide candidates were first annotated as either CDR or FR peptides. To clearly identify the CDR3 fingerprint peptides, the CDR3 fingerprints in the PSM were analyzed. Filter / Alpha required sufficient coverage of high-resolution CDR3 fragment ions We implemented a filter algorithm (see illustration in Figure 8B). The filter is based on a unique Nb sequence of approximately 500,000. A target sequence database containing columns and a non-overlapping decoy database of similar size The Nb sequence database of targets and decoys used in this paper was used for evaluation. The peptides were obtained from different llamas. Peptide identifications from the decoy database were considered false positives. FDR is the ratio of peptide identifications from the decoy data compared to those from the target database. The CDR3 length was also a factor in determining the sensitivity of the peptides. This was taken into consideration to enable the development of a new CDR3 peptide filter. The mass coverage is the mass of a fragment ion (b ion or y ion) within the mass accuracy window. The same was used for evaluation. The peptide spectra were combined. The CDRs that passed this filter (5% FDR) were Only three peptides were selected for downstream Nb assembly.
[0127] ii) Nanobody sequence assembly CDR peptides containing authentic CDR3 peptides were used for Nb protein assembly To identify Nb, two additional criteria must be met. These include: 1) Both CDR1 and CDR2 peptides are used for Nb assembly. 2) For any Nb identification, a minimum of 50% composite CDR coverage must be achieved. A rate was made mandatory.
[0128] b. Quantification and classification of antigen-specific Nb repertoires MS raw data were read using MSFileReader 3.1 SP4 (ThermoFis her), and the pymsfilereader python library (github b.com / frallain / pymsfilereader) Highly reliable CDR3 peptides that have passed the quality filter are analyzed by label-free LC / MS. was quantified by.
[0129] i) Quantification of CDR3 peptides Enables accurate, label-free quantification of CDR3 peptide identification across different LC runs To achieve this, different retention time windows were specified for peptide peak extraction. For peptides that can be directly identified by the search engine based on MS spectra, peak extraction is performed. For detection, a small quantification window of + / - 0.5 min retention time (RT) shift was used. Peptides not directly identified from a particular LC run (peptide and stochastic ion samples) For the LCs (due to the complexity of the coding), their RTs were predicted based on the RTs of the adjacent LCs. , adjusted using the median RT difference of commonly identified peptides between the two LC runs In this case, approximately 100% of all identified peptides were extracted to facilitate extraction of peptide peaks. 95% agreement between two LC runs + / - 2.0 min (for a typical 90 min LC gradient) A relaxed RT window of + / - 10 ppm was applied. The peaks were extracted using both the m / z and z of the peptides. The data were extracted and smoothed using a Gaussian function. Their AUC (area under the curve) was calculated and replicated. The AUCs from the combined LC runs were averaged to infer CDR3 peptide intensities.
[0130] ii) Classification of Nb For example, to enable accurate classification based on Nb affinity, the relative ion intensities (AUC) of CDR3 fingerprint peptides between three different biochemically defined Nb samples (F1, F2, and F3) were quantified as I1, I2, and I3. Based on the quantification results, the CDR3 peptides were arbitrarily classified into three clusters (C1, C2, and C3) using the following criteria.
[0131] 1) For the C3 (high affinity) cluster: I3 > I_{1}+I_{2} (indicating that Nb is specific to F3) 2) For the C2 (medium affinity) cluster: I2 > I_{1}+I_{3} (indicating that Nb is specific to F2)
[0132] 3) For the C1 (low affinity) cluster: I1 > I_{2}+I_{3} (indicating that Nb is more specific to F1 or has a high probability of being a non-specific binder). Alternatively, if I1 < I_{2}+I_{3} and I2 < I_{1}+I_{3} and I3 < I_{1}+I_{2}, these Nb identifications are likely to be non-specifically identified and grouped into C1. See Figure 8C. Using the above method, HSA and GST Nb were classified. Some modifications were made for the quantification and characterization of high-affinity PDZ Nb. Specifically, an additional control "F_control" (ion intensity of I_control) for MBP interaction Nb was included for quantification. When the sum of the intensities of I2 and I3 of the Nb CDR3 peptide is higher than 20 times that of I_control (i.e., 20 * I_control < I2 + I3), high Defined affinity clusters Nb (represented by their unique CDR3 peptides) . For Nbs where multiple unique CDR3 peptides were used for quantification, the classification results between different CDR3 peptides from the same Nb should be consistent, and if not, they were removed before the final results were reported.
[0133] Heatmap analysis of the relative intensities of CDR3 peptides The identified CDR3 peptides were quantified based on their relative MS1 ion intensities and then clustered using the Augur Llama script. Z-scores were calculated based on the relative ion intensities to generate the heatmap in Figure 3A for visualization. This was used for
[0134] Structure modeling of the antigen-Nb complex The structural model of Nb was obtained using the multi-template comparative modeling protocol of MODELLER (Webb , B. & Sali, A, 2014). Next, the CDR3 loop was refined and the top 5 scoring loop structures were selected for downstream docking. Then, each Nb model was docked to its respective antigen using the antibody-antigen docking protocol of PatchDock software, which focused on CDR search (Schneidman-Duhovny, 2005 ). The models were then re-scored by the statistical potential SOAP (Dong, 2013). The antigen interface residues (distance < X Å from the Nb atoms) among the 10 best-scoring models by SOAP score were used to determine the epitope. After defining the epitope, k-means clustering was used to group the epitopes based on their similarity. Nbs were clustered by clustering. Clusters represent the most immunogenic surface patches on the antigen. CXMS data reveal that antigen-Nb complexes exhibit distances that optimize the achievement of binding. Modeled by the constraint-based PatchDock protocol (Schneidma n-Duhovny, 2020;Russel, 2012). C between cross-linked residues When the a-Ca distance is within 25 Å and 20 Å for DSS and EDC crosslinkers, respectively, considered that restraint had been achieved (Shi, 2014;Fernandez-Mart In the case of ambiguous constraints, such as GST dimers, one of the bridges may be successful. You need to be standing.
[0135] Machine learning analysis of Nb repertoire Deep neural networks combined with accurate high-pH fractionation and quantitative proteomics The NIH-10 ... This model has one convolutional layer with batch normalization and ReLU activation functions, and It consists of a max-pooling layer followed by a fully connected layer, which processes the extracted features in a This is integrated into a logit layer that leads to class prediction. The convolutional layer consists of 20 1D filters. long enough to capture relevant CDRs and avoid data overfitting We construct a local receptive field with a window size of 7 amino acids, which is short enough to avoid stumbling. During the forward pass, each filter slides along the protein sequence with a fixed stride. and performs an element-wise multiplication with the current array window, then sums and fills it. The classification accuracy of the model was 92%.
[0136] The network is trained to distinguish between low-affinity and high-affinity binders. To understand the learned physicochemical features, we extract activation fields from the predicted ones through the network. The activation path to the router was calculated. The process is repeated backward from the last two layers of the fully connected network, and for each sequence Extract the output signal and look for the highest peaks that give the most weight to the classification. The contribution of each filter to the upstream filter activity was calculated. The network activity was analyzed to extract domain-specific dominant filters. The interpretation process results in a unique contribution per filter per sequence. The filter is activated along the downsampled sequence in the max pooling layer. For each filter, its highest peak was selected, which led to the classification. We then determined the most contributing filters for each sequence, and found that the filters contributed more than 30% of the total number of sequences in the region of interest. An interesting filter with the above contributions was also obtained.
[0137] Computer-implemented methods The logical operations described herein with respect to the various figures are: (1) a computing device; A computer running on a computer system (e.g., the computing device described in Figure 14) A sequence of data-implemented actions or program modules (i.e., software), (2 ) Interconnected machine logic circuits or circuit modules within a computing device ( (i.e., hardware), (3) the software and hardware of a computing device It should be understood that the present invention may be implemented as a combination of hardware and software. The logical operations described are not limited to any particular combination of hardware and software. It is a matter of choice that depends on the performance and other requirements of the computing device. Accordingly, the logical operations described herein may be operations, structural devices, acts, or models. These operations, structural devices, acts, and modules are called software. Implemented in software, firmware, dedicated digital logic, or any combination thereof More or fewer operations than those shown in the figures and described herein may be performed. It should also be understood that these operations may be performed differently than those described herein. The steps may be performed in any order.
[0138] Referring to FIG. 14, an exemplary computing device capable of implementing the methods described herein is shown. An exemplary computing device 500 is shown. Please understand that this document is merely an example of a suitable computing environment in which the methods described herein may be implemented. It should be understood that, optionally, the computing device 500 is a personal computer. computers, servers, handheld or laptop devices, multiprocessor systems, Microprocessor-based systems, networked personal computers (PCs), Minicomputers, mainframe computers, embedded systems, and / or the above This includes distributed computing environments that include multiple systems or devices. The computing system may be any known computing system, including but not limited to the above. In a computing environment, remote Distributed computing allows multiple computing devices to perform different tasks. In an operating environment, program modules, applications, and other data are It may be stored on a local and / or remote computer storage medium.
[0139] In its most basic configuration, computing device 500 typically includes at least It includes one processing unit 506 and system memory 504. Depending on the exact configuration and type of memory, system memory 504 may be volatile (randomly accessed) or memory (RAM), non-volatile (read-only memory (ROM), flash memory This most basic configuration may be either a The structure is shown in FIG. 14 by dashed line 502. The processing unit 506 A standard programmer performs the arithmetic and logic operations necessary for the operation of the computing device 500. The computing device 500 may also be a A bus or other communication system for communicating information between the various components of the computing device 500. The mechanism may include:
[0140] Computing device 500 may have additional features / functionality. The computing device 500 may include, but is not limited to, a magnetic or optical disk or tape. Such storage may include, but is not limited to, removable storage 508 and non-removable storage 510. Computing device 500 may include any additional storage. It may include network connection(s) 516 that allow the device to communicate with other devices. The computing device 500 can also include a keyboard, mouse, touchscreen, and The device may have input device(s) 514 such as a display, screen, etc. It may also include output device(s) 512 such as a speaker, printer, etc. Additional devices may be used to facilitate data communication between components of the mobile device 500. The devices can be connected to a bus. All of these devices are well known in the art and are described herein. There's no need to go into detail.
[0141] The processing unit 506 may be configured to process program code encoded on a tangible computer-readable medium. The tangible computer-readable medium may be configured to execute a Any medium capable of providing data that causes the device 500 (i.e., machine) to act in a particular way various computer systems to provide instructions to the processing unit 506 for execution. Any computer-readable medium may be utilized. Examples of tangible computer-readable media include computer A computer program for storing information such as readable instructions, data structures, program modules, or other data. Volatile, non-volatile, removable, and and non-removable media, including but not limited to: system memory 504; Removable storage 508 and non-removable storage 510 are all tangible components. Examples of tangible computer-readable recording media include integrated circuits (e.g., For example, field programmable gate arrays or application specific ICs), hard disks disks, optical disks, magneto-optical disks, floppy disks, magnetic tapes, holographic Storage media, solid state devices, RAM, ROM, electrically erasable programmable read only memory Read-only memory (EEPROM), flash memory or other memory technologies, CD-R OM, Digital Versatile Disc (DVD) or other optical storage, magnetic cassette, magnetic tape, magnetic disk storage or other magnetic storage devices, Not limited to:
[0142] In an exemplary implementation, the processing unit 506 executes programs stored in the system memory 504. For example, the bus may transfer data to system memory 504. The instructions can be carried by the processing unit 506 from which they are received and executed. The data received by the library 504 may be processed by the processing unit 506 either before or after execution. Optionally stored on removable storage 508 or non-removable storage 510 It can be done.
[0143] The various techniques described herein may be implemented in connection with hardware or software, or It should be understood that the present invention may be practiced in connection with any of the above methods, or any combination thereof, where appropriate. Thus, the methods and apparatus of the presently disclosed subject matter, or particular aspects or portions thereof, may be stored on a floppy disk, CD-ROM, hard drive, or any other machine-readable storage medium. takes the form of program code (i.e., instructions) embodied in a tangible medium such as a storage medium. The program code can be loaded into a machine such as a computing device When executed, the machine becomes an apparatus for practicing the presently disclosed subject matter. When executing program code on a mobile computer, the computing device Generally, a processor, a processor-readable storage medium (volatile and non-volatile) memory and / or storage elements), at least one input device, and at least One or more programs may be, for example, applications. Through the use of programming interfaces (APIs), reusable controls, etc. The processes described in connection with the subject matter of this disclosure can be implemented or utilized as needed. Such programs use high-level procedural language to communicate with the computer system. Alternatively, it can be implemented in an object-oriented programming language. Depending on the program, the program(s) can be implemented in assembly language or machine language. However, the language may be a compiled or interpreted language, and the hardware It can be combined with the implementation.
[0144] As mentioned above, the logical operations described herein, such as those described in Example 8, are hardware-based. It can be implemented in hardware, software, or a combination of both, as appropriate. For example, the logical operation may be performed by one or more processors, such as the computing device 500 of FIG. The logic operations described in Example 8 can be performed using a computing device. The calculations include methods to determine the antigen affinity of nanobody peptide sequences, training deep learning models, and A deep learning-based method for screening and predicting antigen affinity of nanobody peptide sequences These operations include, but are not limited to, the methods described in detail above. is doing.
[0145] In some embodiments, the computer-implemented method comprises: receiving a nanobody peptide sequence; Identifying multiple CDR regions of the nanobody peptide sequence, wherein the CDR regions are C and identifying a DR3 region. Apply a fragmentation filter to identify one or more false positives in the nanobody peptide sequence. Discarding the CDR3 region; Quantify the abundance of one or more undiscarded CDR3 regions in the nanobody peptide sequence To do, Quantified presence of one or more undiscarded CDR3 regions in the nanobody peptide sequence and predicting antigen affinity based on the abundance of the antibody.
[0146] In some embodiments, a method for training a deep learning model comprises: A dataset containing multiple nanobody peptide sequences and corresponding antigen affinity labels was generated. To achieve Using the dataset, we identified nanobody peptide sequences with low antigen affinity and those with high antigen affinity. We trained a deep learning model to classify nanobody and peptide sequences with and without nucleotide sequences. This includes:
[0147] In some embodiments, a method for determining the antigen affinity of a Nanobody peptide sequence teeth, receiving a nanobody peptide sequence; inputting the nanobody peptide sequence into a trained deep learning model; Using a trained deep learning model, we identified nanobody peptide sequences with low antigen affinity. and classifying the antibody as having high affinity or high antigen affinity.
[0148] [Table 2-1] [Table 2-2] [Table 2-3] Table 2-4 Table 2-5 Table 2-6 Table 2-7 Table 2-8 Table 2-9 Table 2-10 Table 2-11 Table 2-12 Table 2-13 Table 2-14 Table 2-15 Table 2-16 Table 2-17 Table 2-18 Table 2-19 Table 2-20 Table 3-1 Table 3-2 Table 3-3 Table 3-4 Table 3-5 Table 3-6 Table 3-7 Table 3-8 Table 3-9 Table 3-10 Table 3-11 Table 3-12 Table 3-13 Table 4-1 Table 4-2 Table 4-3 Table 5 Table 6
[0149] References 1.Muyldermans,S.Nanobodies:natural singl e-domainantibodies.Annu Rev Biochem82,77 5-797 (2013). 2.Beghein,E.& Gettemans,J.NanobodyTechno logy:A Versatile Toolkit for Microscopic Imaging,Protein-Protein Interaction Analyze lysis,and Protein Function Exploration.F ront Immunol 8,771 (2017). 3.Rasmussen,S.G.et al.Structure of a nan obody-stabilized active state of the bet a(2)adrenoceptor.Nature 469,175-180(2011 ). 4.Jovcevska,I.& Muyldermans,S.The Therap eutic Potential of Nanobodies.BioDrugs34 ,11-26(2020). 5.Lauwereys,M.etal.Potent enzyme inhibit ors derived from dromedary heavy-chain a ntibodies.The EMBO journa l17,3512-3520( 1998). 6.Pardon,E.etal.Ageneral protocol for th e generation of Nanobodies for structura l biology.Nature protocols 9,674-693(201 4). 7.McMahon,C.et al.Yeast surface display platform for rapid discovery of conforma tionally selective nanobodies.Nature str uctural & molecular biology 25,289-296(2 018). 8.Egloff,P.etal.Engineered peptide barco des for in-depth analyses of binding pro tein libraries.Nature methods 16,421-428 (2019). 9.Fridy,P.C.etal.A robust pipeline for r apid production of versatile nanobody re pertoires.Nature methods 11,1253-1260(20 14). 10.Savitski,M.M.,Wilhelm,M.,Hahne,H.,Kus ter,B.& Bantscheff,M.A Scalable Approach for Protein False Discovery Rate Estima tion in Large Proteomic Data Sets.Molecu lar & cellular proteomics:MCP14,2394-240 4(2015). 11.DeKosky,B.J.et al.High-throughput seq uencing of the paired human immunoglobul in heavy and light chain repertoire.Natu re biotechnology 31,166-169(2013). 12.Elias,J.E.& Gygi,S.P.Target-decoy sea rch strategy for increased confidence in large-scale protein identifications by mass spectrometry.Nature methods 4,207-2 14(2007). 13.Schneidman-Duhovny,D.,Inbar,Y.,Nussin ov,R.& Wolfson,H.J.PatchDock and SymmDoc k:servers for rigid and symmetric dockin g.Nucleic acids research 33,W363-W367(20 05). 14.Chait,B.T.,Cadene,M.,Olinares,P.D.,Ro ut,M.P.& Shi,Y.Revealing Higher Order Pr otein Structure Using Mass Spectrometry. Journal of the American Society for Mass Spectrometry 27,952-965(2016). 15.Rout,M.P.& Sali,A.Principles for Inte grative Structural Biology Studies.Cell 177,1384-1403(2019). 16.Yu,C.& Huang,L.Cross-Linking Mass Spe ctrometry:An Emerging Technology for Int eractomics and Structural Biology.Analyt ical Chemistry 90,144-165(2018). 17.Leitner,A.,Faini,M.,Stengel,F.& Aeber sold,R.Cross linking and Mass Spectromet ry:An Integrated Technology to Understan d the Structure and Function of Molecula r Machines.Trends in biochemical science s 41,20-32(2016). 18.Larsen,M.T.,Kuhlmann,M.,Hvam,M.L.&How ard,K.A.Albumin-based drug delivery:harn essing nature to cure disease.Mol Cell T her 4,3(2016). 19.Zhu,W.H.,Smith,J.W.& Huang,C.M.Mass S pectrometry-Based Label-Free Quantitativ e Proteomics.J Biomed Biotechnol(2010). 20.Cox,J.& Mann,M.MaxQuant enables high peptide identification rates,individuali zedp.p.b.-range mass accuracies and prot eome-wide protein quantification.Nature biotechnology 26,1367-1372(2008). 21.Shi,Y.et al.Structural characterizati onby cross-linking reveals the detailed architecture of a coatomer-related hepta meric module from the nuclear pore compl ex.Molecular & cellular proteomics:MCP13 ,2927-2943(2014). 22.Kim,S.J.et al.Integrative structure a n functional anatomy of a nuclear pore c omplex.Nature555,475-482(2018). 23.Pires,D.E.V.,Ascher,D.B.& Blundell,T. L.mCSM:predicting the effects of mutatio ns in proteins using graph-based signatu res.Bio informatics(Oxford,England)30,33 5-342(2014). 24.Finn,J.A.et al.Improving Loop Modelin g of the Antibody Complementarity-Deter mining Region 3 Using Knowledge-Based Re straints.PloS one11,e0154811(2016). 25.Tiller,K.E.et al.Arginine mutations i n antibody complementarity-determining r egions display context-dependent affinit y / specificity trade-offs.The Journal of biological chemistry 292,16638-16652(201 7). 26.Mitchell,L.S.& Colwell,L.J.Analysis o f nanobody paratopes reveals greater div ersity than classical antibodies.Protein Eng Des Sel 31,267-275(2018). 27.Desmyter,A.etal.Crystal structure of a camel single-domain VH antibody fragme nt in complex with lysozyme.Nat Struct B iol3,803-811(1996). 28.Li,T.et al.Immuno-targeting the multi functional CD38 using nanobody.Scientifi c reports 6(2016). 29.Sheng,M.& Sala,C.PDZ domains and the organization of supramolecular complexes .Annu Rev Neurosci 24,1-29(2001). 30.Doyle,D.A.et al.Crystal structures of acomplexed and peptide-free membrane pr otein-binding domain:Molecular basis of peptide recognition by PDZ.Cell 85,1067- 1076(1996). 31.Niethammer,M.et al.CRIPT,a novel post synaptic protein that binds to the third PDZ domain of PSD-95 / SAP90.Neuron 20,69 3-707(1998). 32.Akram,A.& Inman,R.D.Immunodominance:A pivotal principle in host response to vi ral infections.Clin Immunol 143,99-115(2 012). 33.Bar-On,Y.M.,Phillips,R.& Milo,R.The b iomass distribution on Earth.Proceedings of the National Academy of Sciences of the United States of America 115,6506-65 11(2018). 34.Chaplin,D.D.Overview of the immune re sponse.J Allergy Clin Immun125,S3-S23(20 10). 35.Acharya,P.et al.Heavy chain-only IgG2 b llama antibody effects near-pan HIV-1 neutralization by recognizing a CD4-indu ced epitope that includes elements of co receptor-and CD4-binding sites.J Virol87 ,10173-10181(2013). 36.Arabi,Y.M.et al.Middle East Respirato rySyndrome.New Engl J Med 376,584-594(20 17). 37.Flajnik,M.F.,Deschacht,N.& Muylderman s,S.A Case Of Convergence:Why Did a Simp le Alternative to Canonical Antibodies A rise in Sharks and Camels? PLoS biology 9(2011). 38.Sircar,A.,Sanni,K.A.,Shi,J.& Gray,J.J .Analysis and modeling of the variable r egion of camelid single-domain antibodie s.J Immunol 186,6357-6367(2011). 39.Baran,D.et al.Principles for computat ional design of binding antibodies.Proce edings of the National Academy of Scienc es of the United States of America 114,1 0900-10905(2017). 40.Chevalier,A.et al.Massively parallel denovo protein design for targeted thera peutics.Nature 550,74-79(2017). 41.Arbabi Ghahroudi,M.,Desmyter,A.,Wyns, L.,Hamers,R.& Muyldermans,S.Selection an d identification of single domain antibo dy fragments from camel heavy-chain anti bodies.FEBS letters 414,521-526(1997). 42.Shi,Y.et al.A strategy for dissecting the architectures of native macromolecu larassemblies.Nature methods 12,1135-113 8(2015). 43.Chen,Z.L.etal.A high-speed search eng ine pLink 2 with systematic evaluation f or proteome-scale identification of cros s-linked peptides.Nature communications 10,3404(2019). 44.Dunbar,J.& Deane,C.M.ANARCI:antigen r eceptor numbering and receptor classific ation.Bioinformatics(Oxford,England)32,2 98-300(2016). 45.Lefranc,M.P.et al.IMGT unique numberi ngfor immunoglobulin and T cell receptor variable domains and Ig superfamily V-l ike domains.Dev Comp Immunol 27,55-77(20 03). 46.Crooks,G.E.,Hon,G.,Chandonia,J.M.& Br enner,S.E.WebLogo:a sequence logo genera tor.Genome research 14,1188-1190(2004). 47.Sievers,F.& Higgins,D.G.Clustal Omega ,accurate alignment of very large number s of sequences.Methods in molecular biol ogy 1079,105-116(2014). 48.Letunic,I.& Bork,P.Interactive Tree O f Life(iTOL):an online tool for phylogen etic tree display and annotation.Bioinfo rmatics(Oxford,England)23,127-128(2007). 49.Waterhouse,A.M.,Procter,J.B.,Martin,D .M.,Clamp,M.& Barton,G.J.Jalview Version 2--a multiple sequence alignment editor and analysis workbench.Bioinformatics(Ox ford,England)25,1189-1191(2009). 50.Kall,L.,Canterbury,J.D.,Weston,J.,Nob le,W.S.& MacCoss,M.J.Semi-supervised lea rning for peptide identification from sh otgun proteomics datasets.Nature methods 4,923-925(2007). 51.Webb,B.& Sali,A.Comparative Protein S tructure Modeling Using MODELLER.Curr Pr otoc Bioinformatics 47,561-32(2014). 52.Dong,G.Q.,Fan,H.,Schneidman-Duhovny,D .,Webb,B.& Sali,A.Optimized atomic stati stical potentials:assessment of protein interfaces and loops.Bioinformatics(Oxfo rd,England)29,3158-3166(2013). 53.Schneidman-Duhovny,D.& Wolfson,H.J.Mo deling of Multimolecular Complexes.Metho ds in molecular biology 2112,163-174(202 0). 54.Russel,D.et al.Putting the pieces tog ether:integrative modeling platform soft ware for structure determination of macr omolecular assemblies.PLoS biology10,e10 01244(2012). 55.Fernandez-Martinez,J.et al.Structure and Function of the Nuclear Pore Complex Cytoplasmic mRNA Export Platform.Cell 1 67,1215-1228 e1225(2016).
Claims
1. Nanobody amino acid sequences of complementarity determining regions (CDRs) 3, 2 and / or 1 (CDR3 , CDR1 and / or CDR2 sequences) to identify the reduced CDR3, CDR4 and / or CDR5 sequences. 2 and / or CDR1 sequences are false positives compared to a control, a. obtaining a blood sample from a camelid immunized with an antigen; b. Using the blood sample to obtain a cDNA library of nanobodies. and, c. identifying the sequence of each of the cDNAs in the library; d. from the same or a second blood sample from said camelid immunized with said antigen. Isolating the nanobody; e. Digesting the Nanobody with trypsin or chymotrypsin to generate a digestion product population. To do, f. performing mass spectrometry of the digestion products to obtain mass spectrometry data; g. Selecting the sequences identified in step c that correlate with the mass spectrometry data; h. Identifying the sequences of the CDR3, CDR2 and / or CDR1 regions within the sequence of step g. To do, i. from the sequences of the CDR3, CDR2 and / or CDR1 regions of step h, selecting sequences with a fragmentation coverage rate of at least If chymotrypsin is used in step e, the percentage of cytosine coverage is calculated using the formula f(x, chymotrypsin) Lipsin) = 0.0023x 2 -0.0497x + 0.7723, x [5, 30] or, if trypsin is used in step e, by the formula f(x, trypsin) = 0.00006x 2 Determined by -0.00444x + 0.9194, x [5, 30] and x is the length of the sequence of the CDR3, CDR2 or CDR1 region, respectively. selecting, j. The selected sequences of step i are and / or a group having a CDR1 sequence.
2. The method of claim 1 , wherein the required fragmentation coverage percentage is about 30%.
3. The required fragmentation coverage percentage is about 50%, and in step e., trypsin The method of claim 1 , wherein
4. The required fragmentation coverage percentage is about 40%, and in step e., chymotriptyline The method of claim 1 , wherein syn is used.
5. Step d comprises obtaining plasma from said blood sample and using one or more affinity isolation methods. and isolating the nanobody using the method according to any one of claims 1 to 4. method.
6. The one or more affinity isolation methods of step d may include using a Protein G Sepharose affinity chromatograph. and one or more of: chromatographic analysis and protein A sepharose affinity chromatography The method of claim 5 , comprising:
7. Step d. isolates the antigen-specific nanobody using antigen-specific affinity chromatography. and lyse said antigen-specific nanobodies under varying degrees of stringency. and extracting the nanobody fractions, thereby creating different nanobody fractions. Steps e through i are carried out separately for each fraction, and each different antibody against said antigen is obtained. The affinity of the CDR3, CDR2 and / or CDR1 region sequences of step i is determined by , the CDR3, CDR2 and / or CDR3 fragments in each of the Nanobody fractions Further comprising a functional selection step, which is based on the relative abundance of CDR1 region sequences; The method according to any one of claims 1 to 6.
8. The antigen-specific affinity chromatography is performed on a resin conjugated to the antigen. The method of claim 7 .
9. The antigen-specific affinity chromatography is performed by subjecting a maltose binding protein and the antigen to 8. The method of claim 7, wherein the resin is bonded to
10. CDR3, CDR2 and / or CDR1 peptides having the sequences identified in step i. The method of any one of claims 1 to 9, further comprising creating a
11. the CDR3, CDR2 and / or CDR1 regions having the sequences identified in step i. The method of any one of claims 1 to 9, further comprising producing a nanobody comprising 。
12. From SEQ ID NO: 1 to 2536 and SEQ ID NO: 2665 to 2667 A Nanobody comprising a selected amino acid sequence.
13. 1. A computer-implemented method comprising: receiving a nanobody peptide sequence; by identifying multiple complementarity determining region (CDR) regions of said Nanobody peptide sequence; wherein the CDR regions include CDR3, CDR2 and / or CDR1 regions. Identifying and Applying a fragmentation filter to identify one or more false positives of said Nanobody peptide sequence discarding the CDR3, CDR2 and / or CDR1 regions of interest; one or more of the non-discarded CDR3, CDR2 and / or CDR3 of said Nanobody peptide sequence or quantifying the abundance of the CDR1 region; the one or more non-discarded CDR3, CDR2 and CDR3s of the Nanobody peptide sequence and predicting antigen affinity based on said quantified abundance of CDR1 and / or CDR2 regions. and, the computer-implemented method comprising:
14. the one or more non-discarded CDR3, CDR2 and CDR3s of the Nanobody peptide sequence and / or CDR1 region, and 14. The computer-implemented method of claim 13, further comprising classifying the object as having a specific property. method.
15. said one or more of said Nanobody peptide sequences classified as having high antigen affinity The undiscarded CDR3, CDR2 and / or CDR1 regions of the nanobody protein The computer-implemented method of claim 14 further comprising assembling the data into a plurality of rows.
16. The fragmentation filter is configured to determine the minimum calculated fragmentation coverage ratio. A computer-implemented method according to any one of claims 13 to 15, configured to request method.
17. 17. The method of claim 16, wherein the minimum calculated fragmentation coverage percentage is about 30%. The computer-implemented method described herein.
18. The minimum calculated fragmentation coverage percentage was 0.01 for the trypsin-treated sample. and about 50% for chymotrypsin-treated samples and about 40% for chymotrypsin-treated samples.
18. The computer-implemented method of claim 17.
19. receiving a plurality of Nanobody peptide sequences; comparing each of the Nanobody peptide sequences to a database; Separating the peptide sequences into excluded and non-excluded subgroups and wherein the Nanobody peptide sequences of the excluded subgroup are included in the database. and the CDR regions are not found in the nanobodies of the non-excluded subgroup. said comparing said peptides identified only in the dipeptide sequence; The computer-implemented method of any one of claims 13 to 18, further comprising:
20. the one or more non-discarded CDR3, CDR2 and CDR3s of the Nanobody peptide sequence The abundance of the CDR1 and / or CDR2 regions is quantified based on relative MS1 ion signal intensity. The computer-implemented method of any one of claims 13 to 19.
21. The antigen affinity is determined using k-means clustering based on epitope similarity. The computer-implemented method of any one of claims 13 to 20, wherein
22. 1. A method of training a deep learning model, comprising: A computer-implemented method according to any one of claims 13 to 21 is used to generate a dataset. and Using this dataset, Nanobody peptide sequences with low antigen affinity and those with high antigen affinity were identified. A deep learning model was trained to classify nanobody-peptide sequences with affinity. and performing a sequence analysis of the data set, the sequence comprising a plurality of Nanobody peptide sequences and corresponding the training comprising an antigen affinity label; The method comprising:
23. 23. The method of claim 22, wherein the deep learning model is a convolutional neural network. method.
24. 1. A method for determining the antigen affinity of a Nanobody peptide sequence, comprising: receiving a nanobody peptide sequence; inputting the nanobody peptide sequence into a trained deep learning model; The trained deep learning model is used to identify the nanobody peptide sequence. classifying as having antigen affinity or high antigen affinity; The method comprising:
25. 25. The method of claim 24, wherein the deep learning model is a convolutional neural network. method.
26. The trained deep learning model is trained according to claim 22.
26. The method of claim 24 or claim 25.