Method for assembling protein sequence in complex protein mixture

By generating an antibody-derived peptide sequence library and utilizing mass spectrometry and a proteomics search engine, combined with a cross-linking agent to connect distant regions, the problem of accurate antibody chain assembly in complex antibody mixtures was solved, achieving efficient identification and assembly of antibody chains in antibody mixtures and improving antibody binding efficiency.

CN121969936APending Publication Date: 2026-05-01RAPID NOVOR INC
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
RAPID NOVOR INC
Filing Date
2024-07-25
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing technologies struggle to correctly assemble multiple complete antibody chains from complex antibody mixtures, especially polyclonal antibody mixtures, resulting in significant batch-to-batch variations and inconsistent antibody affinities.

Method used

A method was employed to generate an antibody-derived peptide sequence library by contacting antibody samples with different proteases and chemical proteolytic agents. Short antibody-derived peptides were then determined using mass spectrometry, and a candidate chain sequence library was generated using computer analysis. The antibody chain sequences were then identified using mass spectrometry analysis and a proteomics search engine. Long-distance regions were connected using cross-linking agents to improve the reliability of the assay.

Benefits of technology

This improves the accuracy and consistency of antibody chain assembly, reduces batch-to-batch variability, and ensures efficient binding of antibodies to target antigens.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121969936A_ABST
    Figure CN121969936A_ABST
Patent Text Reader

Abstract

Antibodies are major effectors of the adaptive immune system. They bind to the unique characteristics of specific molecules (referred to as antigens) provide numerous tools and strategies for diagnostic, research and clinical applications. The ability to sequence several antibodies in a polyclonal mixture, or at least several major forms thereof (or subpopulations with specific binding characteristics), may result in a faster recombinant antibody production process. So far, due to the complexity of the task, there are few attempts to carry out sequencing on the polyclonal antibody. The present application relates to methods of combinatorially assembling several intact antibody chains using methods including crosslinking, intact chain separation, from middle to bottom proteomics and / or assembling a database using combinatorial chains with a proteomic search engine (based on separation, biochemistry and bioinformatics). The methods may also include recombinantly expressing the identified candidate antibodies to test their binding to the target antigen.
Need to check novelty before this filing date? Find Prior Art

Description

Cross-references to applications related to methods for assembling protein sequences in complex protein mixtures.

[0001] This application claims the benefit of U.S. Provisional Patent Application No. 63 / 517,222, filed August 2, 2023, which is incorporated herein by reference. Sequence List

[0002] The sequence list submitted with this document is an XML file named G17294-00027_Seq listing.xml, created on July 25, 2024, and 192,940 bytes in size. The contents of the above file are incorporated herein by reference in their entirety. Technical Field

[0003] This invention relates generally to the field of antibodies, and more specifically to the identification and sequencing of antibodies in antibody mixtures. Background Technology

[0004] Antibodies are major effectors of the adaptive immune system. Their unique properties of binding to specific molecules (called antigens) provide numerous tools and strategies for diagnostics, research, and clinical applications. Antibodies are among the fastest-growing biomolecules in clinical trials, with a global market valued at approximately $130.9 billion in 2020, projected to grow to $223.7 billion by the end of 2025. 1 .

[0005] From a usage perspective, antibodies can be divided into three main categories: monoclonal antibodies (mAbs), produced by a single plasma cell, have identical heavy and light chain sequences, and bind to unique portions of an antigen (i.e., epitopes). On the other hand, polyclonal antibodies (pAbs) consist of antibody libraries derived from different B cells; these antibodies recognize a mixture of similar and different epitopes of the same antigen. Recombinant antibodies (rAbs) are monoclonal antibodies produced in vitro by cloning their genes into expression vectors. They are then introduced into an expression host to produce the associated antibody protein. Monoclonal antibodies have been developed and used for a variety of biological applications, including diagnostic and clinical applications such as the treatment of autoimmune diseases, cancer, and infectious diseases. The next logical advance in immunotherapy is the combination of multiple monoclonal antibodies targeting different epitopes. 2 Since the physiological immune response to infection is more of a polyclonal response, it would be beneficial to develop oligoclonal / polyclonal immunotherapy strategies.

[0006] The choice between producing and using mAbs or pAbs is influenced by various factors, including application type, production costs and time, and technical expertise. Both types of antibodies have their advantages and disadvantages. Polyclonal antibodies can be produced faster than monoclonal antibodies and require less technical skill. All that is needed is to inoculate one or a few animals with the target antigen and adjuvant. Furthermore, pAbs are heterogeneous, ensuring they can recognize a given epitope under different conformations or minor variations. Additionally, pAbs are more flexible than mAbs in terms of buffer usage and epitope conformational changes. However, pAbs produced by different animals can have significantly different affinities for specific antigens. Moreover, significant differences in affinity have been observed in different blood samples from the same animal, making batch-to-batch variability a significant concern for polyclonal antibodies. A common strategy is to mix large quantities of pAbs enriched from different animals to produce a product with average performance that is least prone to batch-to-batch variability. The resulting product is a mixture of high-performing and low-performing antibodies, thus providing a relatively reproducible but average product.

[0007] Unlike pAbs, mAbs are a homogeneous population of a single type with low batch-to-batch variability, but their high specificity can sometimes limit their use. For example, small differences in epitope structure or composition can significantly affect antibody-antigen affinity. mAbs are produced by the same immortalized B cells; fusion B cells and myeloma cells can maintain their production in vitro. However, hybridoma cells may suffer from gene loss, gene mutation, and cell line genetic drift. These potential problems sometimes encountered with mAbs can be overcome by using recombinant antibodies (rAbs). For recombinant antibody rAbs, the antibody sequence must first be determined to synthesize the immunoglobulin (Ig) L and H chain genes and generate expression constructs. The constructs are then transfected into high-yielding cell lines such as CHO or HEK293.

[0008] Similar strategies can be applied to pAbs. The ability to sequence several antibodies from a polyclonal mixture, or at least several of their major forms (or subsets with specific binding properties), could lead to faster production of high-quality recombinant antibodies. Furthermore, generating selected recombinant forms to produce more defined complex mixtures, such as recombinant oligoclonal mixtures, is an attractive solution that combines the advantages of mAbs and rAbs (eliminating batch-to-batch variability and avoiding hybridoma loss) with pAbs (a response more closely resembling the natural immune system).

[0009] To date, there has been very little work on sequencing pAbs, likely due to the complexity of the task.

[0010] Cheung et al. 3A sequencing approach is proposed, which first enriches antibodies from immunized animals and then analyzes a reference database created by next-generation sequencing (NGS) of B cell Ig libraries from immunized animals using a standard proteomics mass spectrometry (MS) method. (Wine et al.) 4 A relatively similar approach was used. Gilchuk et al. 5 An example of discovering antibodies from human blood by combining single-cell mRNA sequencing and proteomics was published. De novo sequencing from polyclonal protein mixtures is rare. De novo polyclonal antibody sequencing may be advantageous because obtaining B cells from the same animal used to produce pAbs is typically impossible. Furthermore, NGS applied to B cell libraries from animals other than humans and mice is more exploratory and requires some unconventional effort to develop and test suitable primers.

[0011] Assembling small amounts of pure recombinant antibodies from artificial or natural polyclonal mixtures requires overcoming the following three main challenges:

[0012] Challenge 1: Correctly assembling different complementary determination regions (CDRs)

[0013] Challenge 2: Correctly assemble the separate heavy (H) chain and light (L) chain.

[0014] Challenge 3: Correctly pair a given L chain with a given H chain.

[0015] Assembling proteins from peptide cleavages of complex protein mixtures has been a subject of much research, often referred to as the inference problem. 6 This problem of inference arises under specific experimental conditions, such as shotgun bottom-up proteomics. In this case, complex protein mixtures are initially digested by one or more proteases; then, peptides are isolated and analyzed by LC-MS. The initial proteins can be identified by matching these peptide fragments against a database of known proteins. Under standard proteomics conditions, sample mixtures are quite heterogeneous, with only a very small number of similar proteins sharing similar peptides. Despite the challenge, this problem is addressed in most standard proteomics facilities. The main challenge of de novo assembly of relatively similar single antibodies from polyclonal mixtures is a very complex one that, to our knowledge, is not routinely performed in any proteomics lab: (1) most proteins currently share very similar sequence fragments, and (2) there is an additional challenge that some sequence regions vary more and have no equivalents in genomic databases, requiring a de novo approach that does not rely on the possible prior knowledge of the components present in the mixture (i.e., mismatch with sequence databases).

[0016] Guthals et al. 7An attempt was made to sequence antibodies from a mixture derived from human serum. The authors managed to sequence several heavy and light chains, focusing their efforts on the four light chains (LC) and seven heavy chains (HC) with the highest confidence. Of the 28 antibodies expressed, only two showed affinity for the antigen. Notably, the affinity of these two antibodies was several orders of magnitude lower than that of the polyclonal antibodies found in the original serum, suggesting potential errors in their sequences. Their approach involved generating a given chain (heavy or light chain) using a method based on assembling tightly overlapping peptides and expanding these reads. Overlapping peptides were generated by placing the polyclonal mixture in different proteases and using LC-MS-based proteomics to isolate the peptides and generate de novo sequencing fragments. These de novo sequenced peptides were assembled together, and shared sequence reads were generated by combining information from these short reads into long fragments using overlapping “contigs.” Constructing long sequences using a single method is highly computationally intensive and prone to long-distance assembly errors. With this neighborhood-based assembly method, the confidence in correctly assembling two regions rapidly decreases as the distance between them increases.

[0017] More specifically, let's assume a protein contains n consecutive regions R_1, R_2, ..., R_n along its chain. In a complex mixture of proteins, the sequence of each region may have multiple choices. For correct protein assembly, de novo sequencing must ensure that all regions originate from the same protein. Based on the nearest neighbor assembly method, assume the confidence level of R_k and R_(k+1) belonging to the same protein is "p". Then, the probability that all n regions belong to the same protein decreases exponentially to p. (n-1) .

[0018] For example, if p=0.9, the probability of correctly assembling 8 regions is only 0.9. 7 The probability is approximately 0.478. In practice, the confidence level for assembling two adjacent regions may be much lower. For example, when p=0.5, the probability of correctly assembling 8 regions is approximately 0.5. 7 The value is approximately 0.008, less than 1%. Therefore, the construction of long-chain antibodies using only the proximity method is severely limited. Figure 1 graphically illustrates the limitations of assembling long chains using only the proximity method.

[0019] Therefore, new methods are needed to correctly assemble several complete antibody chains from complex mixtures.

[0020] This specification relates to several documents, the entire contents of which are incorporated herein by reference. Summary of the Invention

[0021] In various aspects and implementation schemes, this disclosure provides the following items 1 to 28:

[0022] 1. A method for determining the amino acid sequence of one or more antibodies or antibody chains present in an antibody mixture, the method comprising:

[0023] (A) Generate an antibody-derived peptide sequence library through the following steps:

[0024] (i) Contact multiple samples from the mixture with a reducing agent;

[0025] (ii) Optionally, the plurality of samples are contacted with an agent that prevents disulfide bond formation or modifies cysteine ​​residues into lysine analogs;

[0026] (iii) Contacting multiple samples with one or more proteases and / or chemical proteolytic agents to obtain antibody-derived peptides, wherein each sample is contacted with a different protease, chemical proteolytic agent or combination thereof, thereby obtaining multiple antibody-derived peptide digests;

[0027] (iv) The amino acid sequence and intensity of short antibody-derived peptides present in the antibody-derived peptide digest were determined by mass spectrometry using de novo sequencing, wherein the length of the short antibody-derived peptides was less than 50 amino acids, preferably less than 40, 30 or 20 amino acids.

[0028] (v) Assign antibody-derived peptide sequences to specific complementarity-determining regions (CDR1, CDR2, or CDR3) or framework regions (FR1, FR2, FR3, or FR4).

[0029] (vi) A library of candidate antibody chain sequences is generated on a computer by combining antibody-derived peptide sequences from different antibody regions identified in (v);

[0030] (B) Perform at least one of (a) through (e) to identify the amino acid sequence of one or more antibodies or antibody chains present in a candidate antibody chain sequence library:

[0031] (a)

[0032] (i) Optionally, the sample is contacted with an immunoglobulin (Ig) domain-isolating protease to obtain Fab, F(ab')2, and Fc fragments; optionally, the Fc fragment is removed (e.g., using protein A / G beads).

[0033] (ii) Perform a separation step on the sample in the mixture to separate the antibodies, antibody chains or antibody fragments Fab, (F(ab')2 and Fc) present in the mixture into multiple fractions;

[0034] (iii) Optionally, the sample from the antibody mixture is contacted with a reducing agent and / or a reagent that modifies cysteine ​​residues to prevent disulfide bond formation or to modify it into a lysine analog.

[0035] (iv) Contact the fraction with one or more proteases and / or chemical protein hydrolysants to obtain a digestion fraction containing antibody-derived peptides, wherein the peptides correspond to different regions of an antibody or antibody chain;

[0036] (v) Perform mass spectrometry (MS) analysis on the digested fractions to obtain MS and / or MS / MS spectra;

[0037] (vi) Using the antibody-derived peptide sequence library obtained in (A), analyze MS and / or MS / MS spectra using a proteomics search engine; and

[0038] (vii) Based on their co-elution curves in different fractions, assemble antibody-derived peptides produced from the same antibody or antibody chain;

[0039] (b)

[0040] (i) Optionally incubate samples from antibody mixtures with a reagent that modifies free cysteine ​​residues;

[0041] (ii) Incubate the sample under denaturing, non-reducing conditions;

[0042] (iii) Contacting the sample with one or more proteases and / or chemical protein hydrolysants to obtain digested antibody-derived peptides, wherein the digested antibody-derived peptides comprise dimers of peptides from distant antibody regions linked by disulfide bonds;

[0043] (iv) Perform mass spectrometry (MS) analysis on the digested fractions to obtain MS and / or MS / MS spectra;

[0044] (v) Analyze MS and / or MS / MS spectra using a proteomics search engine, the proteomics search engine being adapted to identify cross-linked peptides using the antibody-derived peptide sequences identified in (A);

[0045] (c)

[0046] (i) Incubate samples from antibody mixtures with a protein cross-linking agent;

[0047] (ii) Incubate the sample under denaturing and reducing conditions;

[0048] (iii) Optionally, the sample is contacted with a reagent that modifies cysteine ​​residues to prevent disulfide bond formation or to convert cysteine ​​into a lysine analogue;

[0049] (iv) Contact the sample with one or more proteases and / or chemical protein hydrolysants to obtain digested antibody-derived peptides, wherein the digested antibody-derived peptides comprise cross-linked dimers of peptides from distant antibody regions;

[0050] (v) Perform mass spectrometry (MS) analysis on the digested fractions to obtain MS and / or MS / MS spectra;

[0051] (vi) Analyze MS and / or MS / MS spectra using a proteomics search engine, which is adapted to identify cross-linked peptides using the antibody-derived peptide sequences identified in (A);

[0052] (d)

[0053] (i) A candidate antibody chain sequence library is generated on a computer by combining antibody-derived peptide sequences from different antibody regions identified in (A) and (v);

[0054] (ii) The amino acid sequence and intensity of the long antibody-derived peptides present in the digests of multiple antibody-derived peptides in (A) and (iii) are determined by mass spectrometry, wherein the length of the long antibody-derived peptides is greater than 20 amino acids, preferably greater than 30, 40 or 50 amino acids.

[0055] (iii) Compare long antibody-derived peptide sequences with a library of candidate antibody chain sequences to identify antibody chain sequences present in the antibody mixture;

[0056] (e)

[0057] (i) Assess the overlap between antibody-derived peptides present in the sample, where a high level of overlap indicates that the antibody-derived peptides belong to the same antibody or antibody chain.

[0058] 2. The method according to item 1, wherein the separation step includes chromatography or gel separation.

[0059] 3. The method according to paragraph 2, wherein the chromatography is hydrophobic interaction chromatography (HIC), and wherein the gel separation is natural gel, isoelectric focusing (IEF) gel, or 2D gel separation.

[0060] 4. The method according to any one of claims 1 to 3, wherein the separation step is performed under non-reducing conditions.

[0061] 5. The method according to any one of claims 1 to 4, wherein the protease is pepsin, trypsin, chymotrypsin, AspN, LysC, GluC, or any combination thereof.

[0062] 6. The method according to any one of claims 1 to 5, comprising incubating a sample from an antibody mixture with a protein cross-linking agent.

[0063] 7. The method according to claim 6, wherein the protein crosslinking agent comprises bis(sulfosuccinimide) octanoate (BS3).

[0064] 8. The method according to any one of claims 1 to 7, wherein the protein sequence database search engine is Mascot, Sequest, Novor-Cloud, or Maxquant.

[0065] 9. The method according to any one of claims 1 to 8, comprising contacting a sample with a reagent that modifies cysteine ​​residues into lysine analogs.

[0066] 10. The method according to paragraph 9, wherein the agent for modifying the cysteine ​​residue to prevent disulfide bond formation includes iodoacetamide, and / or the agent for modifying the cysteine ​​residue to a lysine analogue includes electrophilic ethylamine, such as 2-bromoethylamine hydrobromide (BEA).

[0067] 11. The method according to any one of claims 1 to 10, comprising contacting the sample with a reagent that prevents the formation of disulfide bonds from cysteine ​​residues.

[0068] 12. The method according to claim 11, wherein the agent preventing the formation of disulfide bonds from cysteine ​​residues comprises N-ethylmaleimide.

[0069] 13. The method according to any one of claims 1 to 12, wherein the Ig domain isolates the protease comprising IdeS and / or IdeZ.

[0070] 14. The method according to any one of claims 1 to 13, wherein the short antibody-derived peptide has a length of 5 to 20, 5 to 30, 5 to 40, or 5 to 50 amino acids.

[0071] 15. The method of any one of claims 1 to 14, wherein the length of the long antibody-derived peptide is 20 to 100, 30 to 100, 40 to 100, or 50 to 100 amino acids.

[0072] 16. The method according to claim 15, wherein the length of the long antibody-derived peptide is 40 to 80 amino acids.

[0073] 17. The method according to any one of claims 1 to 16, wherein the level of overlap between two antibody-derived peptides is determined by calculating an overlap score, and wherein the score of overlap between amino acids located at the amino and carboxyl termini of the antibody-derived peptide is lower than that between amino acids located at internal positions of the antibody-derived peptide.

[0074] 18. The method according to any one of claims 1 to 17, wherein the method includes performing at least (a).

[0075] 19. The method according to any one of claims 1 to 18, wherein the method includes performing at least (b).

[0076] 20. The method according to any one of claims 1 to 19, wherein the method includes performing at least (c).

[0077] 21. The method according to any one of claims 1 to 20, wherein the method includes performing at least (d).

[0078] 22. The method according to any one of claims 1 to 21, wherein the method includes performing at least (e).

[0079] 23. The method according to any one of claims 1 to 22, wherein the antibody mixture is a mixture of polyclonal antibodies or a mixture of monoclonal antibodies.

[0080] 24. The method according to any one of claims 1 to 23, wherein the MS is liquid chromatography-MS (LC-MS) and the MS / MS is tandem MS.

[0081] 25. The method according to any one of claims 1 to 24 further comprises expressing one or more recombinant antibodies or antibody fragments, said recombinant antibodies or antibody fragments corresponding to one or more antibodies identified by the method according to any one of claims 1 to 14.

[0082] 26. The method according to paragraph 25 further includes evaluating the binding of the recombinant antibody or antibody fragment to the target antigen.

[0083] 27. The method according to claim 26, wherein assessing the binding of the antibody or antibody fragment includes performing an immunoassay.

[0084] 28. The method according to item 27, wherein the immunoassay is an enzyme-linked immunosorbent assay (ELISA).

[0085] Other objects, advantages, and features of this disclosure will become more apparent after reading the following non-limiting description of specific embodiments given by way of example only with reference to the accompanying drawings. Attached Figure Description

[0086] In the attached diagram:

[0087] Figures 1A-B are graphs depicting the confidence of two paired regions as a function of the distance separating them. This graph is a schematic diagram illustrating that the confidence that two regions belong to the same chain decreases as they become more distant. Figure 1A: Assuming a 90% confidence that two adjacent specific peptides are in the same chain, the confidence of R1 and R6 drops to 59%, and the confidence of R1 and R14 drops to approximately 25%. Figure 1B: Shows the probability fractions (90%, 80%, and 50%) of two adjacent portions belonging to the same chain and the effect of distance between a given sequence and any other sequence in the chain.

[0088] Figure 2A shows an explanation of the nomenclature used to illustrate the naming of the region (CDR and frame region for light and heavy chains).

[0089] Figure 2B illustrates the assembly process of peptide regions after digestion. Using different proteases (with different amino acid specificities) on the same sample can produce short, overlapping reads (an important step in proper assembly).

[0090] Figure 2C shows examples of long-distance regions of a given antibody heavy chain in different antibody mixtures. Correctly paired CDRs can be performed in pairs (2.1, 2.2, and 2.3) and the overall association of a particular CDR1-CDR2-CDR3 can be verified (2.4).

[0091] Figure 3A depicts distant regions linked by natural covalent bonds (disulfide bonds) or artificially via cross-linking agents (e.g., lysine). From a linear perspective, CDR1 and CDR3 are “distant” regions. However, their covalent linkage via disulfide bonds, or potentially via cross-linking agents, increases the confidence in associating these two distant regions.

[0092] Figure 3B is a schematic diagram illustrating that using only the "nearest neighbor" (VA) method to construct the entire strand results in lower confidence as the strand distance increases. Identifying the connection between CDR1 and CDR3 using disulfide bonds or crosslinks can improve confidence, thus affecting the overall confidence of the entire strand sequencing. Black indicates high confidence, and white indicates low confidence.

[0093] Figure 4 depicts a specific example of a given strain of rabbit antibody (IGHV1S1, see IMGT), with cysteine ​​(C) and lysine (K) highlighted. CDR1 and CDR3 are underlined. If this antibody is digested with lysC (unreduced), two peptides linked together by disulfide bonds should be obtained, containing the specific CDR1 and CDR3. This given peptide can be analyzed “as is” using a mid-to-bottom proteomics approach, or it can be isolated in a gel, or subsequently analyzed with another protease (such as pepsin, gluC, or AspN). In this example, lysC followed by pepsin digestion is shown. The application here is to correlate a specific CDR1 with a specific CDR3 in a mixture of other antibodies and to increase the confidence in assembling distant regions within a given antibody.

[0094] Figure 5A depicts a general procedure for assembling a given antibody in the presence of other antibodies. Step 1: Digest the polyclonal mixture with different proteases. In Step 2, these peptides are converted into sequence information. Step 3: Assemble these regions using a Fasta format assembly method. These fragments can be complete antibodies or Fab regions to simplify short reads (assembled FRs and CDRs are rare). This assembly database can be used to search the generated initial dataset (Step 1). Alternatively, a subset of it can be used to search a specific experimental dataset (Step 4), using a smaller dataset than in Step 3. Steps 3 through 5 can be repeated several times under different experimental conditions. Finally, these regions are assembled into chains, which can be recombinantly expressed, and their affinity can be tested in functional assays.

[0095] Figure 5B: Generating a list of possible regions using the de novo method can be done using standard methods, such as the software Novor and Peaks (Bioinformatics Solutions). Some phylogenetic and more conserved regions can be identified using standard search engines and publicly available sequence databases (imgt.org, Uniprot, GenBank, RefSeq).

[0096] Figure 5C: These regions can be assembled using combinatorial methods and used as a database for a proteomics search engine. The main scoring strategy (i.e., ranking possible hits and removing noise) is performed using overlap scores (rather than total identified peptides). In short, the confidence score is weighted based on the quality of the overlapping peptide sequences rather than the total identified peptides, retaining only sequences showing a high level of overlap and discarding the rest (see Figure 8). This process can be performed in a modular manner.

[0097] Figure 5D: A series of experiments allows the possible candidate list to be narrowed down to a very small number and reduces the number of false positives.

[0098] Figure 6 depicts the HIC-HPLC separation and fractionation of a mixture of two mouse antibodies (designated “P13” and “P14”) from Absolute Antibody. Four rectangular areas highlight the four fractions analyzed by LC-MS, representing reduction, digestion, and fractionation.

[0099] Figure 7 illustrates the sequence naming strategy used in this work and how the relationship between the sequence name and the sequence itself exists. In this embodiment, the region data of the alpaca antibody from Example 2 are used. Table 3 shows the different possible FR1, CDR1, FR2, etc., which can be arranged in a matrix structure or a vector. According to Table 3, the vector FR1[1] is the sequence EVHLVESGGGLVQPGSLRLSCVVS (i.e., the first sequence in the table), and so on. The name “seq-1-2-3-1-2-7-1” refers to the combination FR1[1]-CDR1[2]-FR2[3]-CDR2[1]-FR3[2]-CDR3[7]-FR4[1]. The letters “a”, “b”, and “c” are used to design any possible combination (i.e., “a” for FR1 indicates that any of the five possible FR1 sequences can be used from Table 3) or a specific number for a particular sequence (i.e., seq-1-bcdefg indicates FR[1], and the remaining fab region is any possible combination). Finally, a limited range of combinations can also be used, such as seq-1-(2-3)-cdefg (i.e., FR1[1], followed by CDR1[2] or CDR1[3], etc.). This naming allows for the rapid generation of combined sequences or the acquisition of a list of possible sequences. Furthermore, the selected possible sequence groups obtained after analyzing the database search can reveal possible patterns (all possible sequences have the same FR1, CDR1, etc.).

[0100] Figure 8 shows that typical search engines like Novor Cloud (or any other possible search engine such as Mascot, Sequest, GPM, etc.) rank and score proteins based on the number of possible experimental peptides supporting a given sequence. In this study, an additional “quality score” was developed, specifically for cases where overlapping peptides are generated using multiple different proteases. Each peptide matching a given protein sequence is converted into a vector of the same amino acid length, carrying a discrete value incrementing from 0 (N and C ends) to a maximum of 3 (each step further from the N and C ends). This numerical vector is located at the same position in the sequence as the identified peptide. All peptides are then summed from the N end to the C end of the protein. In this case, although Case A shows more peptides to explain the FASTA sequence “Case A” (i.e., it is likely to rank highly by any search engine), the minimum overlap is “0”, while Case B, with significantly fewer peptides, has a minimum score of 1. Therefore, if any search engine is used to score protein rankings (based on the total number of identified peptides), Case A would be listed as the most likely case (i.e., a larger number of identified total peptides), but using the strategy developed in this study based on minimum overlap scores, Case B is the most likely case (a smaller total number of peptides than A, but with higher quality overlapping peptides).

[0101] Figure 9A depicts the overlap score plot of the four best candidates for the unique CDR3[7]. As shown in the upper part of the figure, all four sequences from FR1 to CDR3 are similar (i.e., seq-1-2-3-1-2-7), differing only in the FR4 / J region and the hinge region. It can be seen that the coverage of the entire sequence (FR1 to FR3) is similar, and it can be noted that the C-terminus of CDR3 is better covered by FR4[1] and the hinge[2] (i.e., G03).

[0102] Figure 9B shows an SDS-page gel image obtained under non-reducing conditions in an immunoprecipitation assay using rabbit IgG as the antigen and a natural alpaca polyclonal antibody. Several bands were identified, one above 150 kDa representing typical alpaca IgG1a,b, and another at 15 kDa representing a fragment of the atypical form of IgG2b,c naturally present in the sample.

[0103] Figure 10A shows a heatmap of the top 10 candidate sequences. White areas are associated with low-score overlap, while darker areas are associated with high-score overlap. CDR3

[10] showed better overlap scores. The four sequences containing CDR3

[10] are highlighted.

[0104] Figure 10B is an overlap score plot of the top 4 selected sequences containing selected CDR3

[10] .

[0105] Figure 11A depicts the structure of the recombinant VHH antibody. The VHH domain forms a dimer with human Fc IgG1 and the human hinge region.

[0106] Figure 11B depicts the ELISA plots of the two recombinants and two negative controls presented in this study. Natural alpaca anti-rabbit IgG could not be plotted in the same ELISA as the two recombinants prepared with human Fc, thus requiring different secondary antibodies, although the midpoint of the curves was approximately 0.2–0.4 nM.

[0107] Figures 12A-B show the combined heavy chain CDR3 of the rabbit antibody. H [3] (Figure 12A) and CDR3 L

[10] Minimum fractional distribution of light chains (Fig. 12B).

[0108] Figure 13 shows the overlap scores of the top 3 best heavy chain sequences identified after non-reductive digestion and combinatorial database search (rabbit polyclonal). Their main difference lies in the CDR3 and FR4 chains.

[0109] Figure 14 shows the overlap scores of the top three best light chain sequences identified after non-reductive digestion and combinatorial database search. Their main differences lie in the FR2 and CDR2 regions.

[0110] Figure 15 shows the ELISA plots of recombinant antibody PD108_R1 and total polyclonal mixture, with the midpoints of both in the low nM range.

[0111] Figure 16 shows the natural gel separation of naturally occurring polyclonal human antibody against RBD isolated from vaccinated individuals. The polyclonal antibody was digested with IdeS to produce the fab2 and Fc fragments.

[0112] Figures 17A-C are plots representing seven different contigs of the heavy chains CDR1 (Figure 17A), CDR2 (Figure 17B), and CDR3 (Figure 17C). In each plot (i.e., representing a given contig), the normalized intensity of the different peptide moieties of the given contig is plotted.

[0113] Figure 18 shows the three contigs obtained and selected for germline gene alignment. After filling the gaps with additional overlapping de novo peptides and correcting Ile and Leu, the final sequence of the complete heavy chain variable region was obtained. Detailed Implementation

[0114] Unless otherwise stated herein or clearly contradicted by the context, the use of the terms “a,” “an,” and “the,” as well as similar pronouns, in the context of describing the technology (particularly in the context of the following claims) should be interpreted as encompassing both singular and plural forms.

[0115] Unless otherwise stated, the terms “comprising,” “having,” “including,” and “containing” should be interpreted as open-ended terms (i.e., meaning “including but not limited to”).

[0116] Unless otherwise stated in this document or clearly contradicted by the context, all methods described herein may be performed in any suitable order.

[0117] The use of any and all examples or exemplary language (“e.g.”, “such as”) provided herein is intended only to better illustrate implementations of the claimed technology and, unless otherwise required, does not constitute a limitation on the scope.

[0118] Nothing in the specification should be construed as indicating that any unclaimed element is essential to the practice of implementing the claimed technology.

[0119] In this document, the term “approximately” has its general meaning. The term “approximately” is used to indicate that a value includes the inherent error variation of the equipment or method used to determine the value, or covers values ​​close to said value, such as within 10% of said value (or range of values).

[0120] Unless otherwise stated herein, references to ranges of values ​​herein are intended only as a shorthand for referring to each individual value falling within that range, and each individual value is included in the specification as if it were individually referenced herein. All subsets of values ​​within the range are also included in the specification as if they were individually referenced herein.

[0121] In the case of the description of features or aspects of this disclosure in accordance with a list of Markush groups or alternatives, those skilled in the art will recognize that this disclosure is therefore also described in accordance with any single member or subgroup of the list of Markush groups or alternatives.

[0122] Unless otherwise expressly defined, all technical and scientific terms used herein should be regarded as having the meaning commonly understood by one of ordinary skill in the art (e.g., in stem cell biology, cell culture, molecular genetics, immunology, immunohistochemistry, protein chemistry, and biochemistry).

[0123] Unless otherwise stated, the recombinant proteins, cell culture, and immunotherapy techniques used in this disclosure are standard procedures well known to those skilled in the art. These techniques are described and explained in the literature, such as J. Perbal, *A Practical Guide to Molecular Cloning*, John Wiley and Sons (1984); J. Sambrook et al., *Molecular Cloning: A Laboratory Manual*, Cold Spring Harbour Laboratory Press (1989); TA Brown (ed.), *Essential Molecular Biology: A Practical Approach*, Volumes 1 & 2, IRL Press (1991); DM Glover and BD Hames (ed.), *DNA Cloning: A Practical Approach*, Volumes 1–4, IRL Press (1995 and 1996); FMAusubel et al. (ed.), *Current Protocols in Molecular Biology*, Greene Pub. Associates and Wiley-Interscience (1988, including all updates to date); Ed Harlow and David Lane (ed.), *Antibodies: A Laboratory Manual*, Cold Spring Harbour Laboratory (1988); and JE Coligan et al. (ed.), *Current Protocols*. in Immunology, John Wiley & Sons (including all updates to date).

[0124] All antibody presentation regions in this article are based on the Chothia numbering scheme.

[0125] This paper reports research demonstrating that combining experimental evidence from several complementary chains can significantly improve the nearest neighbor-based chain assembly for correct assembly of single or multiple chains. This evidence increases confidence in the correct pairing of chains over long distances.

[0126] In a first aspect, the present invention provides a method for determining the amino acid sequence of one or more antibodies or antibody chains present in an antibody mixture, the method comprising:

[0127] (A) Generate an antibody-derived peptide sequence library through the following steps:

[0128] (i) Contact multiple samples from the mixture with a reducing agent;

[0129] (ii) Optionally, the plurality of samples are contacted with an agent that prevents disulfide bond formation or modifies cysteine ​​residues into lysine analogs;

[0130] (iii) Contacting multiple samples with one or more proteases and / or chemical proteolytic agents to obtain antibody-derived peptides, wherein each sample is contacted with a different protease, chemical proteolytic agent or combination thereof, thereby obtaining multiple antibody-derived peptide digests;

[0131] (iv) Using de novo sequencing, the amino acid sequence and intensity of short antibody-derived peptides present in the antibody-derived peptide digest were determined by mass spectrometry, wherein the length of the short antibody-derived peptides was less than 50, 40, 30 or 20 amino acids.

[0132] (v) Assign antibody-derived peptide sequences to specific complementarity-determining regions (CDR1, CDR2, or CDR3) or framework regions (FR1, FR2, FR3, or FR4).

[0133] (vi) A library of candidate antibody chain sequences is generated on a computer by combining antibody-derived peptide sequences from different antibody regions identified in (v);

[0134] (B) Perform at least one of (a) through (e) to identify the amino acid sequence of one or more antibodies or antibody chains present in a candidate antibody chain sequence library:

[0135] (a)

[0136] (i) Optionally, the sample is contacted with an immunoglobulin (Ig) domain-isolating protease to obtain Fab, F(ab')2, and Fc fragments; optionally, the Fc fragment is removed (e.g., using protein A / G beads).

[0137] (ii) Perform a separation step on the sample in the mixture to separate the antibodies, antibody chains or antibody fragments Fab, (F(ab')2 and Fc) present in the mixture into multiple fractions;

[0138] (iii) Optionally, the sample from the antibody mixture is contacted with a reducing agent and / or a reagent that modifies cysteine ​​residues to prevent disulfide bond formation or to modify it into a lysine analog.

[0139] (iv) Contact the fraction with one or more proteases and / or chemical protein hydrolysants to obtain a digestion fraction containing antibody-derived peptides, wherein the peptides correspond to different regions of an antibody or antibody chain;

[0140] (v) Perform mass spectrometry (MS) analysis on the digested fractions to obtain MS and / or MS / MS spectra;

[0141] (vi) Using the antibody-derived peptide sequence library obtained in (A), analyze MS and / or MS / MS spectra using a proteomics search engine; and

[0142] (vii) Based on their co-elution curves in different fractions, assemble antibody-derived peptides produced from the same antibody or antibody chain;

[0143] (b)

[0144] (i) Optionally incubate samples from antibody mixtures with a reagent that modifies free cysteine ​​residues;

[0145] (ii) Incubate the sample under denaturing, non-reducing conditions;

[0146] (iii) Contacting the sample with one or more proteases and / or chemical protein hydrolysants to obtain digested antibody-derived peptides, wherein the digested antibody-derived peptides comprise dimers of peptides from distant antibody regions linked by disulfide bonds;

[0147] (iv) Perform mass spectrometry (MS) analysis on the digested fractions to obtain MS and / or MS / MS spectra;

[0148] (v) Analyze MS and / or MS / MS spectra using a proteomics search engine that is suitable for identifying cross-linked peptides using the antibody-derived peptide sequences identified in (A);

[0149] (c)

[0150] (i) Incubate samples from antibody mixtures with a protein cross-linking agent;

[0151] (ii) Incubate the sample under denaturing and reducing conditions;

[0152] (iii) Optionally, the sample is contacted with a reagent that modifies cysteine ​​residues to prevent disulfide bond formation or to convert cysteine ​​into a lysine analogue;

[0153] (iv) Contact the sample with one or more proteases and / or chemical protein hydrolysants to obtain digested antibody-derived peptides, wherein the digested antibody-derived peptides comprise cross-linked dimers of peptides from distant antibody regions;

[0154] (v) Perform mass spectrometry (MS) analysis on the digested fractions to obtain MS and / or MS / MS spectra;

[0155] (vi) Analyze MS and / or MS / MS spectra using a proteomics search engine that is suitable for identifying cross-linked peptides using the antibody-derived peptide sequences identified in (A);

[0156] (d)

[0157] (i) A candidate antibody chain sequence library is generated on a computer by combining antibody-derived peptide sequences from different antibody regions identified in (A) and (v);

[0158] (ii) The amino acid sequence and intensity of the long antibody-derived peptides present in the digests of multiple antibody-derived peptides in (A) and (iii) are determined by mass spectrometry, wherein the length of the long antibody-derived peptides is greater than 20 amino acids, preferably greater than 30, 40 or 50 amino acids.

[0159] (iii) Compare long antibody-derived peptide sequences with a library of candidate antibody chain sequences to identify antibody chain sequences present in the antibody mixture;

[0160] (e)

[0161] (i) Assess the overlap between antibody-derived peptides present in the sample, where a high level of overlap indicates that the antibody-derived peptides belong to the same antibody or antibody chain.

[0162] In one implementation, at least (a) is performed. In one implementation, at least (b) is performed. In one implementation, at least (c) is performed. In one implementation, at least (d) is performed. In one implementation, at least (e) is performed.

[0163] In one implementation, at least two of (a)-(e) are performed. In one implementation, two of (a)-(e) are performed. In one implementation, (a) and (b) are performed. In another implementation, (a) and (c) are performed. In another implementation, (a) and (d) are performed. In another implementation, (a) and (e) are performed. In another implementation, (b) and (c) are performed. In another implementation, (b) and (d) are performed. In another implementation, (b) and (e) are performed. In another implementation, (c) and (d) are performed. In another implementation, (c) and (e) are performed. In another implementation, (d) and (e) are performed.

[0164] In one implementation, at least three of (a)-(e) are performed. In one implementation, three of (a)-(e) are performed. In one implementation, (a), (b), and (c) are performed. In one implementation, (a), (b), and (d) are performed. In one implementation, (a), (b), and (e) are performed. In one implementation, (a), (c), and (d) are performed. In one implementation, (a), (c), and (e) are performed. In one implementation, (a), (d), and (e) are performed. In one implementation, (b), (c), and (d) are performed. In one implementation, (b), (c), and (e) are performed. In one implementation, (c), (d), and (e) are performed.

[0165] Proximity or neighboring regions refer to at least two adjacent peptides in an antibody, but can extend to adjacent regions, such as those with FR1 and CDR1, CDR1 and FR2, FR2 and CDR2, and other conformations shown in Figure 2A. Distant regions correspond to two regions separated by at least one additional region. The following eight examples are considered distant regions:

[0166] CDR1 and CDR2

[0167] CDR2 and CDR3

[0168] CDR1 and CDR3

[0169] FR1 and FR2,

[0170] FR2 and FR3,

[0171] FR1 and FR3,

[0172] FR1 with CDR2 or CDR3,

[0173] FR2 and CDR3, etc.

[0174] This article uses the following terminology to distinguish between heavy chains and light chains: FR1 H and FR1 L These two examples correspond to frames 1 of the heavy and light chains, respectively. Assuming a mixture of "i" different antibodies, the terminology can be expanded as follows: FR1Hi and FR1Li represent the heavy and light chains of frame 1 for a given antibody "i," respectively. A specific CDR1H1 is part of the same chain as CDR2H1 and CDR3H1.

[0175] It is important to remember that, in the case of antibodies, for single-chain assembly (heavy or light chain), the function of a given antibody depends heavily on the correct assembly of three specific far-field regions: the CDRs (CDR1, CDR2, and CDR3), and the correct heavy and light chain pairings. While the frame regions (FRs) can be mutated, their loosely defined combinations can be defined depending on the lineage used. However, their exact composition is also important for antibody activity.

[0176] The challenge that this disclosure seeks to address is to correctly assemble specific CDR1, CDR2, and CDR3 within the chain by combining multiple methods (based on isolation, biochemistry, and bioinformatics), and to extend this to include FR1, FR2, FR3, and FR4, thereby assembling several complete antibody chains.

[0177] Figures 2A-C illustrate the naming and assembly of individual regions and closed regions being assembled, while Figure 2C highlights the pairing of distant regions within a given chain.

[0178] To improve the accuracy of these long sequences, in addition to Guthals et al. 7 In addition to the proposed proximity assembly method, combinations of different methods are suggested. The proposed methods are based on the unique structural characteristics inherent in each antibody or other experimental methods, primarily protein isolation methods and data analysis.

[0179] These specific methods fall into four main categories:

[0180] Category 1 Use covalent bonds other than peptide bonds:

[0181] This category includes strategies for assembling long-distance regions in a given sequence using 1) natural disulfide bonds and 2) artificial cross-linking (e.g., using protein cross-linking agents). This approach allows the detection of distant regions covalently linked due to proximity in the antibody's higher-order structure. For example, in the heavy and light chains of IgG, the first and second cysteine ​​residues at the N-terminus are typically linked, forming a loop. This loop includes CDR1, CDR2, and CDR3 (see Figure 3). As shown in Figure 3A, under non-reducing conditions, a specific peptide representing FR1-CDR1 can bind to its FR3-CDR3. Similarly, artificial cross-linking of primary amines (i.e., lysine residues) followed by digestion with a specific protease or combination of proteases can also highlight the binding of specific long-distance regions correctly. In Figure 3B, the overall confidence of assembling chains using only the proximity method (VA) is compared to a combination of methods such as proximity method + disulfide bond information. By using these additional targeting experiments, the confidence of pairing distant regions can be improved. Using this method, a unique FR1-CDR1 can be associated with a unique FR3-CDR3 in other very similar mixtures of FR1, CDR1, FR3, and CDR3.

[0182] Figure 4 illustrates the results of using a specific protease on the rabbit heavy chain from germline IGHV1S1 under non-reducing conditions. Using the protease Lys-C under non-reducing conditions should allow the separation of long peptides consisting of multiple regions, including CDR1 paired with a CDR3 linked by a native disulfide bond. This peptide can then be analyzed directly by LC-MS or by a second digestion with a different protease. The disulfide-linked peptide can then be separated, which will include the signature peptide sequences from a specific CDR1 and another specific CDR3, thus allowing the pairing of a given unique CDR1 with a given unique CDR3 in a mixture of several other CDR1, CDR2, and CDR3 regions present. Artificial cross-linking also allows for the pairing of specific, unique, distant regions within the antibody. The selected examples presented in this disclosure are performed between cross-linked lysine residues, but any suitable protein cross-linking agent can be used. The protein cross-linking agent can be a bifunctional or heterofunctional reagent. The reactive functional groups of protein crosslinking agents can be NHS ester compounds (reacting with primary amines in lysine side chains), maleimide compounds (reacting with thiol-containing molecules in cysteine ​​side chains), hydrazide compounds (reacting with aldehyde-containing molecules), or carbodiimide compounds, such as EDC (reacting with carboxylates). These two functional groups can be separated by spacers, such as alkyl or polyethylene glycol (PEG) chains. Reagents for inducing protein cross-linking are well known in the art, including, for example, glutaraldehyde, DSG, disuccinimide octanoate (DSS), disuccinimide octanoate sodium sulfonate (BS3), bis(succinimide)penta(ethylene glycol (BS(PEG)5), TSAT, DSP, DTSSP, DST, BSOCOES, EGS, sulfonated EGS, DMA, DMP, DMS, DTBP, DDFNB, SIA, SMAP, SIAB, sulfonated SIAB, AMAS, BMPS, GMBS, sulfonated GMBS, MBS, sulfonated MBS, SMCC, sulfonated SMCC, SMBP, sulfonated SMBP, SMPH, LC-SMCC, sulfonated KMUS, SPDP, LC-SPDP, sulfonated LC-SPDP, or SMPT. Analysis of cross-linked peptides can be performed using suitable tools, such as pLink. 8 ECL 9 xQuest 10,11 ProteinProspector 12,13 Kojak 14 OpenPepXL 15 Or MS-Annika 16For this method, a search engine suitable for proteomics-level identification of cross-linked peptides is used. Examples of such search engines include xQuest, StavroX, pLink / pLink2, xPhrophet, Protein Prospector, Kojak, Xi, Xilmass, MetaMorpheusXL, and Xolik. 17 In one implementation, the search engine is pLink / pLink 2.

[0183] Category 2 Complete chain separation can be achieved using any chemical / biochemical separation strategy (gel, liquid chromatography, HPLC) and information clustering based on LC-MS peak intensity.

[0184] In this category, the chains are separated intact based on their physicochemical properties. Complex intact polyclonal mixtures are then separated into multiple fractions, whether intact antibodies or reduced, using HPLC, standard liquid chromatography, gel electrophoresis (natural or 2D gel electrophoresis), or any analytical separation strategy. Antibody domains (Fab / Fab2 rather than intact antibodies) can also be well separated using specific proteases such as IdeS, IdeZ, papain, or even pepsin. These different fractions are then digested using one or more proteases and analyzed using LC-MS. Distant regions from the same sequence can then be assembled based on the intensity distribution of the different fractions. This can be achieved by correlating peak intensities, as distant regions of the same antibody exhibit similar elution profiles in different fractions.

[0185] Category 3 Middle-down proteomics: Longer fragments containing several regions within a single peptide confirm specific pairings of adjacent CDR regions (MS / MS fragments combined with specific peptide qualities). A database of different CDR and FR combinations was generated based on multiple protease digestions and neighboring assembly. Experimental middle-down datasets were matched against the generated database to identify possible combinations of CDRs and FRs.

[0186] Category 4 This approach utilizes a proteomics search engine with a combinatorial chain assembly database and scores the quality of the chain assembly. This method relies on generating different possible regions and then exhaustively combining them, or limiting possible chain assemblies using information inferred from experimental observations. The resulting database is then used with the proteomics search engine. The selected potential candidates are primarily based on the quality of the overall assembly (i.e., the quality of overlap between different peptides obtained using different proteases), and less on the total number of peptides identified for a given potential candidate sequence. The use of quality overlap instead of total peptides has not been extensively explored in the literature on proteomics analysis.

[0187] Due to the variability of antibody sequence characteristics, none of the proposed methods are universal. Combining these methods can confirm chain assembly with higher confidence.

[0188] In the study reported in the examples, two different strategies were used to validate the four methods described above.

[0189] Strategy 1: Use known standard antibody mixtures. The antibody standards have already been sequenced; therefore, the correct assembly of distant regions can be easily checked against the known standard sequences.

[0190] Strategy 2: Method validation via protein expression. Recombinant expression in the form of sequencing can validate a procedure for identifying specific sequences in a natural polyclonal pAb library. The affinity of the recombinant antibody for the natural polyclonal antibody is then tested. This procedure can be validated by obtaining recombinant antibodies with high affinity for the same target from a de novo project.

[0191] Figure 5A illustrates a general procedure for sequencing and validating polyclonal antibodies.

[0192] Step 1. Polyclonal antibodies were reduced and alkylated by digestion with several proteases in tandem or parallel. Peptide fragments were analyzed by LC-MS.

[0193] Step 2. MS and MS / MS data are used to generate different regions, namely CDR and FR regions, using de novo sequencing methods with software such as Novor or Peaks (Bioinformatics Solution). Some germline and more conserved regions can be identified using standard search engines with publicly available antibody sequences (e.g., from IMGT, Uniprot, GenBank, RefSeq).

[0194] Step 3. Use the combination method to assemble these regions into sequences, and allow the construction of a FASTA database of combined sequences.

[0195] In step 4, targeted experiments are performed to identify the correct long-range pairings, including but not limited to protease digestion under non-reducing conditions, protein isolation, cross-linking agent use, and analysis of longer peptide fragments (middle-to-bottom proteomics).

[0196] Step 5. Perform data analysis using a database search engine. Use pLink or any software capable of analyzing cross-linked or non-reducing peptides. 18For experiments not involving any non-reducing or cross-linked peptides, Novor Cloud or any database-based sequence matching search engine (i.e., Mascot, Sequest, Maxquant, and similar search engines) can be used. Heavy and light chain pairing can be performed using different experimental conditions (e.g., see Le Bihan et al.). 19 ).

[0197] In step 6, based on these experiments and data analysis, a series of recombinant proteins were prepared and their binding affinity was tested. A more detailed process is shown in Figures 5B-D. Alternative adjustments based on this procedure were used, such as manually assembling several regions / contigs, then performing complete protein separation using native gels, assembling distant regions based on elution curves, and then using germline matching for gap filling.

[0198] Figure 5B highlights two key elements important for generating a list of potential candidate sequences. These elements are: 1) the raw data (MS / MS) and 2) a de novo search of this raw data to generate a list of multiple distinct regions (frameworks and complementarity-determining regions). These short reads can be obtained using any de novo software (such as Novor Cloud, Peak, etc.) or manual de novo sequencing. Some phylogenetic and more conserved regions can be identified using standard search engines with publicly available sequences (e.g., IMGT, Uniprot, GenBank, RefSeq).

[0199] In Figure 5C, a combinatorial approach was used to assemble different regions into unique sequences. These sequences were assembled into a database (CA-FASTA DBS) and then used in conjunction with a search engine to search the raw LC-MS data. Lists of peptides and potential candidates were generated using standard proteomics search engines (Mascot, Sequest, Novor Cloud, etc.). An algorithm primarily weighted by peptide overlap quality was used to reduce the list of potential candidate sequences. This algorithm reduces the number of potential candidates based on the quality of different peptide overlaps and can be used in two different ways:

[0200] 1) High-throughput: Evaluate the minimum overlap score (MO score) for all possible sequences. Sequences with an MO score of zero will be discarded.

[0201] 2) A more detailed analysis based on the overall quality of the overlap spectrum can rank potential candidates during visual inspection.

[0202] The present invention also provides a method for assembling complete or partial antibody protein sequences from a peptide mixture and / or from a defined list of regions (e.g., multiple CDR1, 2, 3, multiple FR1, 2, and 3, and JC regions) obtained from a variety of unique antibodies (i.e., a polyclonal mixture), the method comprising any single method or combination of the following:

[0203] Evidence of long-range region pairing (two peptides linked by non-peptide bonds) for LC-MS analysis was collected using a combination of digestion or cross-linking under non-reducing conditions, and peptide matching was performed using a search engine that could process LC-MS data from cross-linking experiments (natural or chemical cross-linking via disulfide bonds).

[0204] Isolate intact or reduced antibodies that retain the whole chain. These chains can be separated into fractions by gel electrophoresis or chromatography, then the fractions are digested with one or more proteases, and the peptides are identified and quantified by LC-MS analysis. Near-field and far-field regions are then assembled based on the spectral similarity between different fractions to complete the assembly of the entire chain; and

[0205] Peptides that span long distances (e.g., peptides including CDR1, FR2, and CDR2) are generated using specific proteases or chemical cleavage methods. Sequencing of these peptides is performed either by de novo sequencing or using a proteomics engine in conjunction with the FASTA database.

[0206] The present invention also provides a method for assembling a complete antibody protein sequence from a list of multiple mixtures of peptides and / or defined regions (e.g., multiple CDR1, 2, 3, multiple FR1, 2, and 3 and JC regions) containing a variety of unique antibodies, wherein the method comprises:

[0207] (1) Using a combination of digestion or cross-linking under non-reducing conditions, evidence of long-distance pairing is collected by LC-MS and peptide matching, using a search engine that can process cross-linking data (natural or chemical cross-linking via disulfide bonds).

[0208] (2) Isolate intact or reduced antibodies that still retain intact chains or domains (Fab / Fab2), and separate these chains into fractions by gel electrophoresis or chromatography. Then, digest with one or more proteases and analyze by LC-MS to identify and quantify the peptides. Assemble near-field and far-field regions based on the spectral similarity of different fractions to complete the entire chain assembly; and

[0209] (3) Use specific proteases with chemical cleavage methods to produce peptides long enough to span long distances (e.g., peptides including CDR1, FR2, and CDR2). Sequencing of these peptides is performed either by de novo sequencing or by using a search engine in conjunction with the FASTA database.

[0210] This disclosure also provides a method for selecting potential sequence candidates based on the overlap quality of different shorter peptide sequences during merging, which yields longer sequences. These shorter peptides are the result of digestion with various proteases on a mixture of polyclonal antibodies. The method includes:

[0211] Peptides were analyzed using LC-MS and the FASTA database. The FASTA database was generated using combinatorial or more limited combinatorial methods to generate antibody sequences from multiple lists of FR1, 2, 3 and CDR1, 2, 3.

[0212] Generate peptide and protein sequence candidates with confidence values ​​from a list of identified peptides discovered using standard proteomics search engines (Novor Cloud, Sequest, Mascot, and Maxquant).

[0213] The identified short peptide sequences are vectorized by assigning an intensity factor of 0, 1, 2, or 3 to each amino acid that makes up the short peptide sequence (0 for the C and N ends, 1 for the penultimate amino acid at the C and N ends, and 3 for all other amino acids in the peptide sequence).

[0214] The overlap quality of each amino acid in a candidate sequence is calculated by summing all the values ​​obtained from the vectorization of all possible peptides that make up the candidate sequence, while discarding 3 to 5 amino acids at the N and C ends of each candidate sequence (i.e., reducing the importance of edge effects); and all candidate sequences are ranked based on the highest minimum overlap score.

[0215] If a candidate protein sequence contains a value of "0" (typically indicating a lack of overlap), it is discarded. If multiple candidate sequences are retained, an overlap score spectrum is plotted based on amino acid positions, and the overlap quality is manually evaluated to select one or more possible candidates.

[0216] This disclosure also provides a method for selecting potential sequence candidates based on the overlap quality of different peptide sequences that match a given sequence, the given sequence being generated by parallel digestion with various proteases on a polyclonal mixture, wherein the method includes:

[0217] Peptides were analyzed using LC-MS, and a FASTA database was generated using combinatorial or targeted methods to generate antibody sequences from a list of multiple FR1, 2, 3 and CDR1, 2, 3, JC regions, optionally including hinge regions.

[0218] Use standard proteomics search engines (such as Novor cloud, Sequest, Mascot, and Maxquant) to generate peptide and protein sequence candidates with confidence values ​​from a list of peptides;

[0219] The identified peptide sequence is vectorized by assigning an intensity factor of 0, 1, 2 or 3 to each amino acid that makes up the peptide sequence (0 for the C and N ends, 1 for the penultimate amino acid at the C and N ends, and 3 for all other amino acids in the peptide sequence).

[0220] The overlap mass of each amino acid in the candidate sequence is calculated by summing all the values ​​obtained by vectorizing all possible peptides that make up the candidate sequence.

[0221] The N and C ends of each sequence candidate are discarded. All sequence candidates are ranked based on the highest minimum overlap score. If a protein candidate contains a value of 0 in its sequence, it is discarded. If multiple candidates are retained, an overlap score spectrum is plotted and the overlap quality is manually evaluated to select one or more possible candidates.

[0222] Antibodies or antibody chains in a sample can be cleaved using suitable reagents to produce digestible peptides. Reagents for cleaving proteins include chemical reagents such as cyanogen bromide (CNBr) cleaved at methionine (Met) residues; skatole (BNPS) cleaved at tryptophan (Trp) residues; formic acid cleaved at the aspartic-proline (Asp-Pro) peptide bond; hydroxylamine cleaved at the asparagine-glycine (Asn-Gly) peptide bond; 2-nitro-5-thiocyanobenzoic acid (NTCB) cleaved at the cysteine ​​(Cys) residue; and enzymes such as proteases. In one embodiment, antibodies or antibody chains in a sample are cleaved using any suitable enzyme (e.g., proteases) or combinations thereof. Enzymes that can be used for proteolytic cleavage include trypsin, pepsin, thermophilic proteases, chymotrypsin, AspN, LysargiNase, LysC, LysN, GluC, ArgC, Pro / Ala protease, Sap9, KEX2, or any combination thereof. The method may also include a step of modifying cysteine ​​residues to convert them to lysine analogs, for example using electrophilic ethylamines, such as 2-bromoethylamine hydrobromide (BEA), which alkylates cysteine ​​residues (converting them to aminoethylcysteine ​​residues), thus creating cleavage sites for proteases such as Lys-C and trypsin. Digestion can also be performed using a combination of chemical reagents and proteases. Digestion can be allowed to proceed completely so that the protein is cleaved at all bonds that the digestive reagent can cleave; or digestion conditions can be adjusted so that fragmentation is not intentionally incomplete to produce larger fragments, which may be particularly helpful in determining the sequence of distant regions in an antibody chain; or digestion conditions can be modulated so that the protein is partially digested into domains. Conditions that can be modified to adjust the degree of digestion include duration, temperature, pressure, pH, presence or absence of protein denaturing agents, specific protein denaturing agents (e.g., urea, guanidine hydrochloride, detergents, acid cleavage detergents, methanol, acetonitrile, other organic solvents), concentration of denaturing agents, amount or concentration of cleaving agents or their weight ratio relative to the protein to be digested, etc.

[0223] The methods described herein can be used to identify one or more antibodies or antibody chains from an antibody mixture, such as a polyclonal antibody mixture or a monoclonal antibody mixture. This mixture may be derived from a biological sample (e.g., a blood sample) of a subject, such as a subject who has been vaccinated or infected with a microorganism (e.g., a virus). The polyclonal antibody mixture may be collected from tissue culture supernatants of animals or B cells. In various embodiments of the non-limiting methods of this disclosure, the polyclonal antibody mixture may contain, for example, at least two different immunoglobulins, or at least three, five, ten, twenty, fifty, one hundred, or five hundred different immunoglobulins.

[0224] For example, antibody mixtures collected from tissue culture supernatants of animals or B cells can be passed through a Protein A or Protein G Sepharose™ column, which separates the antibodies from other serum proteins. Furthermore, the collected polyclonal antibodies can undergo antigen affinity purification to enrich antibodies with high specific activity. Purification steps, such as antigen affinity purification, can be performed to reduce the complexity of the polyclonal mixture and ultimately reduce the potential number of false positives. The collected polyclonal antibodies can be concentrated or buffer-exchanged before or after purification, or both.

[0225] The antibodies that can be identified by the methods described herein can be conventional IgG (paired heavy and light chain sequences) or atypical antibodies (heavy chain antibodies from camelids such as alpacas, rheas, and camels), as well as IgY from birds (such as chickens).

[0226] In one embodiment, the method further includes evaluating the binding of a recombinant antibody or antibody fragment generated from the sequence information obtained by the methods described herein to its target antigen. This step can be used to confirm whether the putative antibody identified by the method forms a functional antibody or antibody fragment, or, in the case of identifying more than one putative / candidate antibody, to allow confirmation of which putative / candidate antibody is a functional antibody or antibody fragment capable of binding to the target antigen. For example, this can be achieved by recombinantly expressing an antibody or antibody fragment corresponding to the identified putative antibody and evaluating the binding of the recombinant antibody or antibody fragment to its target antigen. This can be achieved by introducing nucleic acids encoding the heavy and light chains of the putative antibody into a suitable expression system, such as CHO cells, HEK293 cells, yeast cells, or E. coli cells, culturing the cells to conditions suitable for producing antibodies or antibody fragments, and evaluating the binding of the antibody or antibody fragment to the antigen (e.g., by immunoassays such as ELISA or surface plasmon resonance (SPR)). In one embodiment, the binding of an antibody or antibody fragment to its target antigen can be evaluated by expressing the antibody or antibody fragment on the surface of a phage and evaluating the binding of the phage to the antigen (phage panning).

[0227] Example

[0228] This disclosure is further described in detail by way of the following non-limiting examples.

[0229] Example 1: Sequencing of two antibodies in an artificial mixture

[0230] Example 1 demonstrates one embodiment of the method disclosed herein, in which two commercially available antibodies having known sequences are mixed, separated under non-reducing conditions, fragmented using different proteases, sequenced using LC-MS and MS / MS, and assembled to produce an accurate and complete antibody sequence.

[0231] In this embodiment, two standard antibodies are separated by chromatography under non-reducing conditions, the fractions are then digested with proteases, and chain assembly is performed based on peptide identification, peak intensity, and intensity correlation between different peptides.

[0232] The antibody mixture was separated into fractions using hydrophobic interaction chromatography (HIC), which showed the correct assembly of the two mouse IgG2a molecules in the artificial mixture (Figure 6).

[0233] Chromatography was used to separate two standard mixed antibodies under non-reducing conditions. The fractions were then subjected to proteolytic digestion and chain assembly based on peptide identification, peak intensity, and intensity correlations between different peptides. Hydrophobic interaction chromatography (HIC) was used to mix and separate 50 μg of each intact antibody, yielding two prominent peaks. Subsequent fractions were reduced with dithiothreitol (DTT), alkylated with iodoacetamide (IAA), and digested with chymotrypsin and pepsin.

[0234] The digested fractions were then analyzed using liquid chromatography-mass spectrometry (LC-MS). Peptides in each fraction were identified using the Novor search engine, which contains an internal database of antibody sequences. Peaks associated with selected sequences representing any region of heavy complementarity determination (CDR) of the two mixed antibodies were then extracted and evaluated for each fraction. Intensity data showed that CDRs clustered together in specific fractions according to their origin (Table 1). Some peaks marked as “contamination” were caused by sample residues, accounting for approximately 3–5% of the main signal, and were therefore negligible. Paired scatter plots were constructed based on the data from the inverse hyperbolic sine function (arcsinh) transformation, and correlations between different CDRs were extracted (Table 2). Table 1 Table 2

[0235] This example demonstrates that by using a mixture of two monoclonal antibodies, the different CDR regions of each antibody can be easily and correctly assembled using a separation method (hydrophobic interaction chromatography in this example).

[0236] Example 2: Assembly of heavy chain-based antibodies from a natural mixture (alpaca) - generating a single region library.

[0237] method

[0238] This article describes a method for assembling heavy-chain-based antibodies from a natural mixture (alpaca). A commercially available polyclonal antibody mixture from alpaca was used. It contained typical IgG (IgG1a, b) and atypical IgG (IgG2b, c) containing only the heavy chain. Although there is some inconsistency in the nomenclature of alpaca antibodies, the IMGT nomenclature (https: / / www.imgt.org / ) was chosen. The antibody was labeled PD106 (internal reference) and prepared for 5-enzyme digestion. The sample was diluted with HPLC-grade water and then heated with DTT at 95°C for 15 min to a final concentration of 30 mM. After cooling to room temperature, the sample was alkylated with iodoacetamide (IAA) at a final concentration of 50 mM for 30 min in the dark at room temperature. The sample was precipitated by adding three times the sample volume of acetone and then incubated at -20°C for 1 h. The precipitate was centrifuged at 23000 xg for 10 min at 4°C; the acetone was carefully decanted to avoid destroying the precipitate, and the precipitate was then dried in a low-pressure centrifuge (SpeedVac™). The precipitate was resuspended in 10 µL of 4M urea at 37°C for 10 min to ensure complete dissolution, then diluted to 100 µL with HPLC-grade water, and 20 µL was taken into five tubes for five different digestions. Pepsin, trypsin, chymotrypsin, AspN, and LysC enzymes were used for digestion. The sample was reconstituted in 40 µL of 0.1% formic acid (FA) and loaded with 5 µg onto Evotips. The sample was analyzed on a Thermo Orbitrap™ Exploris 240 using a 44-minute method, and selected samples (AspN, LysC, and trypsin) were injected in a top-down proteomics mode. The dried peptides were resuspended in 0.1% FA at a concentration of 1 µg / µl and 5 µg was loaded into HPLC vials. Two HPLC column systems were used: a Thermo Scientific PepMap 75µm x 15cm, 3µm, 100Å column and a WATERS C18 column, nanoEASE™ M / Z peptide BEH C18, 300Å, 1.7µm, 300µm x 100mm. Mass spectrometry was performed on an Orbitrap™ Eclipse (Thermo Scientific) with a full elemental scan at 60,000 Hz resolution. Dynamic exclusion was performed for 15 seconds, and an HCD with a relative collision energy of 35 Hz was applied to the separated precursor in charge states 2–11. Product ions were detected in the Orbitrap™ at 30,000 Hz resolution with a maximum injection time of 150 ms.

[0239] sequencing

[0240] De novo sequencing using the Novor Cloud search engine and germline searches revealed several regions, as shown in Table 3. These regions were separated based on information found in IMGT. However, some experimentally discovered sequences were not present in the IMGT database. Most FR1 sequences ended with a "CAAS" at the serine "S". Meanwhile, CDR1 started with a glycine "G". The division between regions can sometimes be ambiguous and can be defined by analyzing overlapping peptides. For FR2, a dominant "QAP" of approximately 4–6 amino acids was found at the N-terminus of FR2. Frame 3 (FR3) was easily defined by the cysteine ​​residue at the third C-terminal position. The short de novo sequencing regions shown in Table 3 can be generated using any de novo sequencing software, such as the Novor search engine or Peak (BioinformaticsSolution). Table 3

[0241]

[0242] EVHLVESGGGLVQPGGSLRLSCVVS (SEQ ID NO:7); KVQLVESGGGLVQAGGSLRLSCAAS (SEQ ID NO:8); QVQLVESGGGLVQAGDSLRLSCAVS (SEQ ID NO:9);QVQLVESGGGLVQTGGSLRLSCAAS (SEQ ID NO:10); QVQLVESGGGLVQTGGSLRLSCALS (SEQ ID NO:11); GFDFDDW (SEQ ID NO:12); GFRFSFY (SEQ ID NO:13); GFTFDDF (SEQ ID NO:14); GGTFNAY (SEQ ID NO:15); GVDSISDMS (SEQ ID NO:16); DMSWFRQAPGKEREGVSCI (SEQ ID NO:17); FMAWFRQAPEKEREFVTRI (SEQ ID NO:18); QMSWVRQAPGKLEWLATI (SEQ ID NO:19); TAGWFRQAPGKAREFLASI (SEQ ID NO:20); TIGWFRQAPGKEREPVSCI (SEQ ID NO:21); NNNGDS (SEQ ID NO:22); NWSGKF (SEQ ID NO:23); SKRDGL (SEQ ID NO:24); SRHDDM (SEQ ID NO:25); SRRDGR (SEQ ID NO:26); INYADSVKGRFTIGRDTAKNTVYLQMNSLKPEDSAVYYCAA (SEQ ID NO:27); PRYPDSAEGRFTISRDNAKNTLYLQMNSLKPEDTAVYCAK (SEQ IDNO:28); TRYGDSVKGRFTISRDNAKEMAFLQMNSLKPEDTAIYYCVA (SEQ ID NO:29); TYYADSVKGRFTFSSDNAKRTVYLQMNSLKPEDTAVYYCAA (SEQ ID NO:30); TYYTDSVKDRFTISRDNANKVVFLQMNGLKPEDTAVYYCAA (SEQ ID NO:31); DEPPYRCSDYWEPWREY (SEQ ID NO:32);DEPPYRCSGGWEPWREY (SEQ ID NO:33);DEPPYRCSSSWDPWREY (SEQ ID NO:34);DEPPYRCSSSWTPWREY (SEQ ID NO:35); GSTWGVSGRRVPDYDY (SEQ ID NO:36);GTTWGVSGRRVPDYDY (SEQ ID NO:37); NO:39); RILCPMDWNSREYTEEVVGS (SEQ ID NO:40); RLQGLSAEAEEYDF (SEQ ID NO:41); WGQGAQVTVSS (SEQ ID NO:42); WGQGTQVTVSS (SEQID NO:43); AHHSEDPSSKCPKCP (SEQ ID NO:45). ;

[0243] A list of potential Fab sequences was generated using a combinatorial method. All sequences had a similar structure and were based on information extracted from Table 3. The naming strategy was as follows: Seq-abcdefg, where “a” is the vector “FR1”; therefore, if a=1, then FR1 is (from information obtained from Table 3, the sequence is the first entry of the FR1 vector “EVHLVESGGGLVQPGSLRLSCVVS” (SEQ ID NO:7)). This corresponds to the first entry in the “vector” FR1.

[0244] Therefore, the sequence Seq-1-2-3-1-2-7-1 is the first entry in vector FR1, the second entry in vector CDR1, then the third entry in FR2, the first entry in vector CDR2, the second entry in FR3, the seventh entry in vector CDR3, and finally the first entry in region FR4. Figure 7 illustrates the principles and connections between sequence naming and how they are assembled. This naming convention allows for the association of sequence names with a given sequence and allows for easy observation of trends in the sequence (based solely on the naming strategy used), thereby facilitating analysis. Although several CDR3s were identified, only 10 sequences are described in Table 3. In addition, 10 effective antibodies based on heavy chains were identified. This example demonstrates how to construct only two candidate sequences. The two complete constructs are named “R1” and “R2”. Sequence R1 will be based on CDR3[7] (PKFPRLDQWVTWDELDY, SEQ ID NO:38), and R2 will be based on CDR3

[10] (RLQGLSAEEEYDF, SEQ ID ID:41). The decision was made to construct the antibody based on its more unique aspect, namely CDR3.

[0245] First, a global FASTA database was generated from these seven regions (FR1, CDR1, FR2, CDR2, FR3, CDR3, and FR4). Hinge regions were included later to reduce the number of combinations to be explored. The overall size of the database can be difficult to manage; in this example, there are at least 5×5×5×5×5×10×2 possibilities, or 62,500 sequences. If the analysis were limited to each CDR3 region, the size of the database to be generated would be reduced by a factor of 10 (6,250 sequences). For each CDR3, a combination database containing all possible combinations was built (the use of phylogenetics was not considered in this particular embodiment, but it is one way to reduce the number of possible combinations). Therefore, for these two sequences, R1 and R2, we get: for R1: Seq-abcde-7-g and R2: Seq-abcde-10-g, where a, b, c, d, e, and g can take any value within their respective “region vectors” (i.e., “a” in FR1 will take sequence values ​​from 1 to 5, as will b, c, d, and e, with only one CDR3 and two J / FR4 regions at a time).

[0246] Each specific CDR3 generated a total of 5×5×5×50×5×1×2=6250 sequences (ranked as "7" or "10" in the vector CDR3 described in this paper). These sequences were generated and assembled using the R (i386 3.6.3) Stringr and Sequiner packages, along with internal scripts for generating FASTA files and analysis. Each FASTA dataset was then generated and used for the initial experimental dataset (a mixture of multiple LC-MS runs with trypsin, pepsin, chymotrypsin, AspN, and LysC). Data was searched using the online tool Novor Cloud (https: / / app.novor.cloud / home), although any database search engine such as Mascot, Sequest, and Maxquant could be used. A major advantage of using Novor Cloud is its ability to handle most combinations of different protease digestions, which is advantageous for the final assembly for this specific application. Novor Cloud's search criteria are as follows:

[0247] FDR was set to 1%, MS tolerance was set to 5 ppm, MSMS tolerance was set to 100 ppm, and all modifications were set to variable modifications: Pyro-Glu (E), Pyro-Glu (Q), carbamoyl methylation (C), and deamidation (NQ).

[0248] A major problem with constructing and searching these same or similar experimental datasets using the FASTA database, which is based on combinations of the most potent / abundant peptides found in the samples, is that the results favor any protein construct with the most experimental evidence, even if they are incorrect. To address this issue, a method is proposed to rank protein candidates based on the quality of peptide overlap, as shown in Figure 8, where the same protein is digested with different specific proteases; therefore, peptide overlap should be obtained.

[0249] Here, in Case A, using traditional proteomics identification software (such as Novor Cloud, Mascot, and Sequest), the hypothesized protein had more experimental evidence (more identified peptides) than in Case B. Using any of these search engines directly, Case A would rank better than Case B. This study used a different approach that prioritizes the quality of overlap between different peptides. Therefore, the identified peptide sequences were vectorized by assigning a “strength” factor of 0, 1, 2, or 3 to each amino acid that makes up the peptide sequence (0 for the C and N ends, 1 for the penultimate amino acid at the C and N ends, and 3 for all other amino acid positions in the peptide sequence). The peptide vectors were projected onto the correct positions in the sequence candidates, and then the overlap quality of each amino acid in the sequence was calculated by summing all the values ​​obtained from the vectorization of a specific amino acid position. This method allows us to see “0” values ​​in the sequence, indicating poor overlap in Case A. Protein matching using overlap calculations allows protein matches to be ranked based on the quality of peptide overlap rather than total peptide abundance. It is not based on clicks from high peptide counts. In the embodiment of Figure 8, Case B will have a higher “rank” due to the better overlap quality of the peptides, even though the protein finds fewer peptides.

[0250] Therefore, under standard proteomics conditions, "Case A" would appear as the best hit in any standard proteomics search engine; however, using the proposed method, Case B would rank higher based on overlap score. The selection of peptide vectors containing maximum values ​​of 0, 1, 2, and 3 is based on basic empirical observations; the maximum value "3" is chosen to avoid giving too much weight to longer peptides, which typically contain less sequence information than shorter peptides. Total overlap is calculated by projecting the total score for each amino acid position. This strategy has two different applications:

[0251] Method 1: Filter protein candidates by a minimum score close to zero (then eliminate proteins with gaps in sequence coverage). This method allows for rapid elimination of a large number of protein candidates in a high-throughput manner.

[0252] Method 2: Observe the quality of sequence or region coverage. Proteins can then be ranked based on overlap quality rather than total peptide matching. In this example, Case A ranks higher among total peptides identified using any conventional search engine. However, Case B has a higher overlap score than Case A based on minimum overlap. At the experimental level, terminal amino acids of the protein are discarded because their value is "0" (edge ​​effect). The assessment of minimum overlap depends on the sample, and the parameters used are described in all the different examples using this overlap quality strategy. Method 1 was used as a high-throughput primary screening method to eliminate meaningless sequences (discarding hundreds to thousands of sequences). Method 2 was used for a number of selected potentially suitable candidates. The following sections illustrate how two heavy chain antibodies from a natural mixture of alpaca antibodies were sequenced.

[0253] After generating experimental datasets of natural polyclonal mixtures using multiple proteases (chymotrypsin, pepsin, AspN, trypsin, Lys-C), the same experimental data were used to search two different FASTA files: R1: seq-abcde-7-g and R2: seq-abcde-10-g.

[0254] Preliminary search for CDR3[7], R1 and CDR3

[10] , R2

[0255] For CDR3[7] (referred to as R1), 1562 combinatorially constructed proteins were identified from a database of 6250 sequences using Novor cloud. An edge effect is encountered when the protein terminal value is “0” (since all peptide values ​​at the N-terminus and C-terminus are “0”). The analysis did not extend to the J region, which reduced the confidence of the measured FR4 overlap. Due to the edge effect, any sequence from amino acid positions 1 to 4 (N-terminus) and after position 99 (referred to as “high”) or after position 93 (referred to as “low”) is removed from the overlap calculation. These limits are artificially defined to determine the minimum within the variable portion of the protein sequence, regardless of the sequence terminals.

[0256] Of the 1562 sequences preserved in the Novor cloud, only 66 combinatorially constructed proteins passed the criterion of having no null values ​​within artificially defined limits. In terms of total peptides discovered, the top two candidates with the best quality overlap ranked 57th and 844th, respectively, in the Novor cloud. They correspond to Seq-1-2-3-1-2-7-1 and Seq-1-2-3-1-2-7-2, both with a minimum overlap (MO) score of 6. A list of the 66 combinatorially constructed proteins is shown in Table 4. Interestingly, the top two candidates share the same FR1, CDR1, FR2, CDR2, and FR3 sequences, and apparently also the same CDR3, differing only in the FR4 region. Table 4

[0257] For CDR3

[10] (referred to as R2), a total of 571 proteins were identified. Similarly, to avoid edge effects, as part of the minimum overlap calculation, amino acids at positions 1 to 4 and 107 and above were deleted. Of these 571 sequences, only 9 selected sequences did not show gaps in their sequences. All selected forms were predominantly FR4 "2". The sequence with the best MO score, ranked 361st in the Novor cloud, was Seq-3-4-4-2-3-10-2 with an MO score of 5. A list of the 9 sequences is shown in Table 5. To confirm pairing of distant regions, complementation analysis was performed. Table 5

[0258] For sample digestion and disulfide bond pairing analysis under non-reducing conditions, 20 µg of PD106 was dried using a low-pressure centrifugation (Speedvac™) in duplicate. Both samples were then reconstituted in 10 µL of 8 M urea, with one parallel sample treated with NEM to a final concentration of 2 mM, followed by shaking incubation at 37°C for 2 hours. Subsequently, both samples were diluted to 40 µL with 50 mM ammonium bicarbonate at pH 8 to achieve a final urea concentration of 2 M, allowing for efficient digestion by trypsin (Promega) at a 1:20 enzyme-to-protein ratio. The samples were then incubated overnight at 37°C. The next day, the digested samples were diluted to 80 µL with 50 mM ammonium bicarbonate at pH 8 to achieve a final urea concentration of 1 M, thus adapting to the urea tolerance range of GluC (Promega). GluC was added at a 1:20 enzyme-to-protein ratio, followed by shaking incubation for 4 hours. After drying the samples under low pressure (SpeedVac™), the samples were reconstituted in 40 µL of 0.1% FA and 5 µg was loaded onto Evotips according to the manufacturer's instructions. The samples were then run on a Thermo Orbitrap™ Exploris 240 using a 44-minute gradient method. Peptides containing disulfide bonds were identified using pLink v2. The parameters used for software identification were as follows: flow type: disulfide bond (HCD-SS); enzyme: LysC-AspN, Try, or Try-GluC; peptide mass: 300-9000; peptide length: 3-90; fixed modification: Gln->pyro-Glu; and variable modifications: “N-ethylmaleimide [C]”, “oxidation [M]”, “deamidated [N]”, “deamidated [Q]”, and “Acetyl [Protein N-term]”. The protein database contained the target sequences and corresponding reverse sequences of the target pAb or mAb mixture. Table 6 shows some identified peptides, with peptide assignments limited to specific regions (FR1, CDR1, FR2, etc.). In some cases, the information obtained from the peptides is insufficient to specifically identify a single given region, thus a range is proposed. In other cases, distant pairings of specific regions can yield precise information, such as in peptides #2 and #4, where the unique FR1 and CDR1 in Table 3 pair with the unique FR3 and CDR3. Table 6

[0259] DTAVYYCAK (SEQ ID NO:46); LSCVVSGFR (SEQ ID NO:47); DTAVYYCAKPK (SEQ ID NO:48); NTLYLQMNSLKPEDTAVYYCAK (SEQ ID NO:49); NTLYLQMNSLKPEDTAVYYCAKPK (SEQ ID NO:50); ID NO:51); LSCAVSGGTFNAYTAGWFR (SEQ ID NO:52); VVFLQMNGLKPEDTAVYYCAADEPPYR (SEQ ID NO:53); LSCAASGFDFDDWTLGWFR (SEQ ID NO:54); (SEQ ID NO:56); LLCPMDWNSR (SEQ ID NO:57);EGVSCLSR (SEQ ID NO:58); EREGVSCLSR (SEQ ID NO:59); LLCPMDWNSRR (SEQ ID NO:60); LLCPLDWSPR (SEQ ID NO:61);

[0260] The disulfide bond information in Table 6, combined with previous searches, confirmed two sequences, R1 (seq-1-2-3-1-2-7-g) and R2 (seq-3-4-4-2-3-10-2). The disulfide bond (SS) search could not completely confirm CDR3

[10] , but proposed two other possibilities, CDR3[8] and CDR3[9].

[0261] Regarding R1, different decisions had to be made regarding the appropriate J fragment for two reasons. First, neither of the two possible J peptides possesses a "protease-friendly cleavage site," meaning that the proteases used would not produce a peptide containing a unique element of one of the two possible J regions, resulting in low coverage of the J fragment. To ensure good coverage of the J region, it was crucial to extend the database from the J region to the hinge region. Second, this polyclonal sample mixture comprised a combination of IgG1a,b (standard antibodies with both heavy and light chains) and IgG2b,c (atypical, heavy-chain-only type, which is the focus of this study). The typical antibody sequence transitions from the J region to the C region, which begins with the CH1 domain; therefore, the resulting peptide is from J to CH1, with the following structure… "TVSSASTK"… The lysine (K) is the trypsin and lysC site, making this peptide extremely abundant in IgG1a,b. For the atypical antibody (the antibody targeted in this study, as the alpaca antibody lacks a light chain), the transition is from the J region to the hinge region. Two different peptides from alpacas defined the hinge (see Table 3): EPKTPKPQPQPQPQPNPTTESKCPKCP (SEQ ID NO:44) as “FR5(1)” and AHHSEDPSSKCPKCP (SEQ ID NO:45) as “FR-5(2)”. These peptides were added to a small database of possible sequences: R1: SEQ-1-2-3-1-2-7-gh, R2: Seq-3-4-2-3-(8-10)-gh, where “g” is J / FR4 (with two possible values) and “h” is “FR5” (as the hinge peptide), also with two possible values. The rest of the sequence has been confirmed. It is important to remember that the sample was not graded and contained a combination of all digested peptides from typical and atypical antibodies, which could interfere with correct allocation. Adding the hinge region as a strategy will help verify the correct J region.

[0262] R1 final sequencing: Seq-1-2-3-1-2-7-gh

[0263] In Figure 9A, overlap scores for four different sequences (seq-1-2-3-1-2-7-gh, where g and h can have two different values) were plotted for the original experimental dataset 1 of the peptide. There were no significant differences between the four sequences from FR1 to FR3. However, differences can be seen when comparing the transition from FR3 to CDR3 (highlighted by two arrows), depending on the choice of J / C and hinge region peptide. The best sequence overlap was generated with g=1 and h=2. Under these conditions, the peptide “LDQWVTWDELDYWGQGAQVTVVSSAHSE” (SEQ ID NO:62) of 1082.1584 amu was detected with an intensity of 7.3e9. This peptide covered CDR3[7], FR4[1] and hinge region FR5[2] very well. Although the peptide is a non-trypsin cleavage product, the actual native sample showed some degradation of the VHH protein, with a band of about 15 kDa (Figure 9B). This corresponds to the VHH fragment that binds to the antigen (rabbit IgG), which ends in …VSSAHHSE at the C-terminus; therefore, the R1 sequence is most likely Seq-1-2-3-1-2-7-1. The next section will examine the 15kDa small sequence found in the sample.

[0264] R1's middle-to-bottom sequence confirmation:

[0265] Using the same search criteria as described above, the LC-MS dataset generated for the study was searched in the database containing all possibilities of CDR3[7] for R1. Any long peptide covering at least 3 regions was searched to meet the criteria of pairing distant regions. The results are shown in Table 7. For R1, Table 7 reports 4 different peptides, each covering 3 to 4 regions, including FR1, CDR1, FR2, CDR2, and FR3. For example, CDR1 can be associated with CDR2 peptides using this “middle-down” approach. Table 7

[0266] EVHLVESGGGLVQPGGSLRLSC(Cam)VVSGFRFSFYQMSWVRQAPGK (SEQ ID NO:63); YQMSWVRQAPGKGLEWLATINNNGDSPRYPDSAEGRF (SEQ ID NO:64); ATINNNGDSPRYPDSAEGRFTISRDNAKNTLY (SEQ ID NO:65); ID NO: 66)

[0267] Experimental description of the binding of natural alpaca fraction to antigen (i.e., rabbit IgG).

[0268] The aim of this experiment was to attach rabbit IgG, rabbit IgG Fc, and rabbit IgG F(ab)2 to agarose beads and then use immunoprecipitation to assess how a commercial alpaca polyclonal antibody binds to its antigen (rabbit IgG). To prepare the antigen fragments, 100 μl of rabbit IgG at a concentration of 10 mg / μl (Sigma I8140-10 mg, lot number SLBK40780) was mixed with 100 μl of PBS and 20 μl of IdeS (50 u / μl, Promega) and incubated at 37°C for one hour. To generate the F(ab)2 and Fc fragments, 0.3 mL of Genscript™ Protein A Resin FF (catalog number L00464-5) was used on a Sigma (C2728) fritted column. The resin was equilibrated twice with 0.75 mL of PBS; then, the IdeS-digested sample was loaded. The sample was reloaded three times and kept flow-through (FT) (corresponding to Fab2 fraction). The column was then washed with 0.3 mL PBS, and the wash buffer was mixed with FT. This was followed by three more washes with 0.5 mL PBS and then discarded. Fc was eluted from the protein A beads using 0.1 mL of glycine at pH 2.7 in 0.124 mL of tris 1M neutralization buffer. Buffer exchange was performed using an Amicon™ 3kDa filter (Sigma). A Pierce NHS-activated Pierce tube was split into three tubes fitted with Mobicol™ filters with 10 μm pores (Bocascientific). The three forms of antigen were conjugated with NHS beads:

[0269] 1) Total rabbit IgG,

[0270] 2) Rabbit F(ab)2,

[0271] 3) Rabbit Fc.

[0272] A second Pierce NHS activated agarose spin column was used as a negative control, neutralized with only 1M tris. Four tubes were incubated overnight with their respective antigens, then washed with water and neutralized with 20 μl of 1M tris. 25 μg of alpaca anti-rabbit polyclonal antibody was added to each of the four columns, incubated for one hour, and then eluted and washed with 50 μl of water. The binding fractions were eluted using SDS-PAGE loading buffer and heated at 95°C for 5 minutes. Gel electrophoresis was performed (see Figure 9B). Each gel band was trypsinized using the standard protocol (in gel digestion) and run on a Thermo Orbitrap™ Exploris 240 at a 44-minute Evosep™ gradient. Results showed that a portion of the alpaca antibody bound to the beads (Tris E lane), and two significant bands were observed at 150 kDa (labeled 1) and 15 kDa (labeled 7) in most of the other three samples (total rabbit antigen, rabbit Fab2, and rabbit Fc). The 15 kDa band conforms to the variable domain of VHH, and the trypsin digest shows the same peptide as the aforementioned peptide LDQWVTWDELDYWGQGAQVTVSSAHHSE (SEQ ID NO:62), suggesting that this domain appears to be cleaved from the chain via an undefined mechanism. This peptide, produced by trypsin digestion, is a hemitrypsin with a mass of 1082.1598 amu and a charge state of 3+, derived from cdr3 (LDQWVTWDELDY, SEQ ID NO:67), the J region (WGQGAQVTVVSS, SEQ ID NO:42), and the IgG2c hinge (AHHSE, SEQ ID NO:68); therefore, it lacks the CH1 domain, a characteristic feature of atypical alpaca antibodies.

[0273] Final sequencing of R2: Seq-3-4-4-2-3-(8,9,10)-gh

[0274] From twelve possibilities (3×2×2), Novor cloud identified ten of the most plausible sequences, as shown in Figure 10A. The heatmap shows the overlap scores for amino acid positions, with white indicating low overlap scores and black indicating high overlap scores. The two arrows at the top of the heatmap point to the N-terminus and C-terminus of the CDR3 region, respectively, indicating that CDR3

[10] (RLQGLSAEEEYDF, SEQ ID NO:41) is the most likely of the three selected CDR3s, with four candidate sequences containing this CDR3 sequence (positions 2, 4, 6, and 10). Figure 10B shows an illustration of these four candidates with this particular CDR3

[10] , containing two different FR4s and two different hinge peptides. G02 and G04 are the most likely sequences because they both have overlap scores above zero in the CDR3 region. Evidence suggests the presence of both peptides in the sample (Seq-3-4-4-2-3-10-2-1 and Seq-3-4-4-2-3-10-2-2, which differ only in the hinge region and are therefore distinct isotypes). Thus, the most probable sequence for R2 is a combination of the two hinge possibilities: Seq-3-4-4-2-3-10-2.

[0275] For its antigen test recombinant sequences R1 and R2:

[0276] Recombinant forms of human IgG1 heavy chain Fc and alpaca variable regions PD106_R1 (Seq-1-2-3-1-2-7-1) and PD106_R2 (Seq-3-4-4-2-3-10-2) were created. The CH2 domain of the human Fc IgG1 is “DKTHTCPPAPELL…” (SEQ ID NO: 217), the complete sequence of which can be viewed on IMGT (accession number J00228). The affinity of the recombinants for rabbit IgG was assessed using ELISA with anti-human HPR species as secondary antibodies. Figure 11B shows that PD106_R1 and R2 have strong affinity for the antigen, with 50% OD 450nm less than 1 nM. Two negative controls were used: a human IgG isotype and a recombinant non-binding agent called PD106_R10.

[0277] Example 3: Sequencing of rabbit polyclonal antibodies to obtain paired heavy and light chains (PD108)

[0278] Heavy and light chains from a mixture of polyclonal antibodies from two rabbits immunized with the short peptide sequence FPPSSEEL (SEQ ID NO: 68) were sequenced. The method described in PCT / CA2022 / 051194 was used to pair the heavy and light chains. Fourteen recombinant forms were generated, only five of which showed similar affinity to the original natural polyclonal antibody (named PD108). The experiments focused on a monomeric form called PD108_r1, which combines a given heavy chain (named Heavy 1) with a given Kappa chain (named Kappa 1). The aim of this example is to demonstrate how to perform de novo sequencing of individual heavy and light chains in complex mixtures of similar antibodies. A list of different regions (FRs and CDRs shown in Table 8) was generated by fully digesting the polyclonal mixture, combined with de novo sequencing and a database search using common rabbit strains (IMGT). The correct regions were then combined to generate the functional chain; this process involved combinatorial database generation, sequence overlap scoring, and digestion of unreduced samples to reduce the number of combinations that pair distant regions. The sample-PD108 is a rabbit anti-FPPSSEEL polyclonal antibody, purchased from Vivitide. Table 8 QSVEESGGRLVTPGGSLTLTCTVS (SEQ ID NO:69); QSVEESGGRLVTPGTPLTLTCTVS(SEQ ID NO:70); GFSLNNY (SEQ ID NO:71); GFSLNTY (SEQ ID NO:72); GFSLSSY (SEQID NO:73); GFSLSTY (SEQ ID NO:74); AMGWVRQAPEKGLEYIGII (SEQ ID NO:75);AMGWVRQAPGKGLEYIGII (SEQ ID NO:76); GVSWVRQAPGKGLEFI (SEQ ID NO:77);PMAWVRQAPGKGLEWIGWI (SEQ ID NO:78); PMGWVRQAPGKGLEWIGWI (SEQ ID NO:79);SMSWVRQAPGKGLEWIGII (SEQ ID NO:80); ANSGY (SEQ ID NO:81); GFIRASGS (SEQ IDNO:82); GPISG (SEQ ID NO:83); GSSGN (SEQ ID NO:84); GTITG (SEQ ID NO:85); ARYANWAKGRFTISRTSTTVDLKMTSLTTEDTATYFCAR (SEQ ID NO:86); ARYASWAKGRFTISKTSTTVDLKMTSLTTEDTATYFCAR (SEQ ID NO:87); AWYASWVKGRFTISKTSTTVDLKITSPTTEDTATYFCTR(SEQ ID NO:88); RYYASWTKGRFTISKTSTTVDLKVTSPTTEDTATYFCAT (SEQ ID NO:89); TYYASWAKGRFTISKTSTTVDLKITSPTTEDTATYFCAR (SEQ ID NO:90); DGDSINGAVMDF (SEQ ID NO:91); DGDTTTGAVMDL (SEQ ID NO:92); DGNSAYNSGVNL (SEQ ID NO:93); NLNAVSATGHTFDP(SEQ ID NO:94); NLNVVSSTGHAFDP (SEQ ID NO:95); SLFGSGDVDNL (SEQ ID NO:96);WGPGTLVTVSSGQPK (SEQ ID NO:97);WGQGTLVTVSSGQPK (SEQ ID NO:98);WGRGTLVTVSSGQPK (SEQ ID NO:99); ADIVMTQTPASVEAAVGGTVTIKC (SEQ ID NO:100);ADVVMTQTPSPVSAAVGGTVSISC (SEQ ID NO:101); ALVMTQTPASVEAAVGGTVTINC (SEQ ID NO:102); AYDMTQTPASVEAAVGGTVTIKC (SEQ ID NO:103); DVVMTQTPSSVEAAVGGTVTIKC (SEQ ID NO:104); QAVVTQTPSSVSAAVGGTVTISC (SEQ ID NO:105); QVLTQTASPVSAAVGGTVTINC (SEQ ID NO:106); QVLTQTASPVSAAVGSTVTINC (SEQ ID NO:107); QASQSISGSYLA (SEQ ID NO:108); QASQSISNLLA (SEQ ID NO:109); QASQSISNQLS (SEQ ID NO:110);QASQSISSYLA (SEQ ID NO:111); QASQSVYNNDRLS (SEQ ID NO:112); QASQSVYNNNLA (SEQ ID NO:113); QSSKSVDKDNRLA (SEQ ID NO:114); QSSKSVYNHNRLS (SEQ ID NO:115);WFQQKPGQRPKLLIY (SEQ ID NO:116); WLQQKPGQPPKRLIY (SEQ ID NO:117);WVQQKPGQPPKRLIY (SEQ ID NO:118); WYQQKPGQPPKLLIS (SEQ ID NO:119);WYQQKPGQPPKLLIY (SEQ ID NO:120); WYQQKPGQPPKRLIY (SEQ ID NO:121);WYQQKPGQPPKVLIY (SEQ ID NO:122); EASKLAS (SEQ ID NO:123); EASKVAS (SEQ ID NO:124); KASTLAS (SEQ ID NO:125); KTSTLAS (SEQ ID NO:126); RASTLAS (SEQ ID NO:127);RTSDLAS(SEQ ID NO:128); SASTLAS (SEQ ID NO:129); YASTLAS (SEQ ID NO:130); GVPSRFKGSGFGTQFTLTISDVQCDDVATYYC (SEQ ID NO:131); GVPSRFKGSGSGTEFTLTVSDLECADAATYYC (SEQ ID NO:132); GVPSRFSGSGSGTQFTLTTISDLECDAATYYC (SEQ ID NO:133); GVSSRFKGSGSGTDFTLTIRDLECADAATYYC (SEQ ID NO:134); GVSSRFKGSGSGTQFTLTISDVQCDDAATYHC (SEQ ID NO:135); GVSSRFKGSGSGTQFTLTISDVQCDDAATYYC (SEQ ID NO:136); GVSSRFKGSGSGTQFTLTISGVECADAATYYC (SEQ ID NO:137); GVSSRFKGSGSGTQLTLTISDLECADAATYYC (SEQ ID NO:138); AGGYSGSSDLCV (SEQ ID NO:139); AGGYSSVSDTA (SEQID NO:140); AGGYSSVTDTA (SEQ ID NO:141); AGGYSSVVDTA (SEQ ID NO:142); LGGYDSSSGDRWK (SEQ ID NO:143); LGIYEAGSDTS (SEQ ID NO:144); QCTYGSATTRIYGEA(SEQ ID NO:145); QCTYVDSSYIGG (SEQ ID NO:146); QQGYTGNNIDNP (SEQ ID NO:147);QSYDYGGGSYGNS (SEQ ID NO:148); FGGGTEVAVK(SEQ ID NO:149); FGGGTEVVVK (SEQ IDNO:150)。;

[0279] PD108 was digested with several different proteases under reduction and alkylation conditions. 140 µg of PD108 was dissolved in 769 µL and first heated at 95 °C for 15 min, then reduced with DTT to a final concentration of 30 mM. The sample was then aliquoted: 30% for cysteine ​​modification with 2-bromoethylamine hydrobromide (BEA), and 70% for alkylation with iodoacetamide (IAA). IAA was added to a final concentration of 50 mM at room temperature in the dark for 30 min. The sample was precipitated by adding three times the sample volume of acetone and then incubated at -20 °C for 1 h. The sample was then centrifuged at 4 °C for 23000 x g for 10 min; acetone was carefully decanted to avoid disrupting the precipitate; the precipitate was then dried under low pressure in a SpeedVac™. The precipitate was resuspended in 10 µL of 4M urea at 37 °C for 10 min to ensure complete dissolution, then diluted to 100 µL with HPLC-grade water. 20 µL of each precipitate was then transferred to five separate tubes for digestion with LysC, trypsin, pepsin, chymotrypsin, and AspN. Pepsin digestion was performed with shaking at 37 °C for 15 min; then the enzyme was inactivated by heating at 95 °C for 3 min, and dried under low pressure using SpeedVac™. The remaining sample was digested overnight at 37 °C, and the mixture was dried under low pressure the next day using SpeedVac™. Finally, the reconstituted digest was analyzed by LC-MS / MS in 40 µL of 0.1% FA solution.

[0280] Add BEA to the fraction reserved for cysteine ​​modification, add 320 µL of BEA (dissolved in 100 mM tris pH 8), and immediately add 60 µL of 1 M tris pH, then add hourly for 3 hours to maintain the reaction near neutral pH. Incubate the reaction at 25°C for 4 hours, then add TCA to a final concentration of 20% and incubate overnight at 4°C. Precipitate the protein by centrifugation at 17,000 rpm for 30 minutes at 4°C, decant the TCA, wash the precipitate twice with 500 µL of acetone, and centrifuge at 17,000 rpm for 10 minutes; then decant the acetone from the precipitate. Dry the precipitate under low pressure in a SpeedVac™ container, then reconstitute it in 10 µL of 4 M urea and shake at 37°C for 10 minutes. Dilute the sample to 40 µL with HPLC-grade water and then aliquot into two tubes. Digestion was performed by adding 30 µL ammonium bicarbonate (pH 8) + 1 µg of enzyme (one tube with trypsin and the other with lysC) and incubating overnight at 37°C; then, the sample was dried under low pressure using SpeedVac™ the next day. After acidification, 5 µg of the digested sample was loaded onto Evotips according to the manufacturer's instructions and run in a gradient at 44 minutes using a Thermo Orbitrap™ Exploris 240. All enzymes were purchased from Promega.

[0281] Analysis of the samples using Novor de novo sequencing revealed several new regions. These regions were determined based on a combination of de novo sequencing methods and matching with a rabbit germline protease digestion database. The major heavy chain lineages were identified as IGHV1S69, IGHV1S40, IGHV1S44, and IGHV1S55, while the light chain lineages were IGKV1S50, 15, 17, 10, 32, and 1. The λ chain was not detected in the samples. Table 8 lists several major forms, some of which showed mutations not found in the IMGT database. The light chain population was found to be more diverse than the heavy chain population.

[0282] Preliminary search of CDR3 heavy and light chains in PD108_r1

[0283] Table 8 shows the different regions for light and heavy chains. This equates to 21,600 potential combinations for heavy chains and 573,440 possibilities for light chains.

[0284] To reduce the size of the database, a combined database was generated for each individual, more exclusive CDR3 region. In this embodiment, using CDR3H[3] = DGNSAYNSGVNL (SEQ ID NO:93) and CDR3L

[10] = QSYDYGGSYGNS (SEQ ID NO:148), 3600 heavy chain sequences (2×4×6×5×5×1×3) and 57344 light chain sequences (8×8×7×8×8×1×2) were obtained, respectively.

[0285] Regarding light chains, the number of combinations was manually reduced by testing two different FR4s of the target sequence QSYDYGGSYGNS (SEQ ID NO:148): Seq-abcde-10-1: QSYDYGJGSYGNS (SEQ ID NO:148) + FGGGTEVAVK (SEQ ID NO:149), and no overlapping peptides were detected in the initial digestion of PD108A. However, for the sequence Seq-abcde-10-2: QSYDYGGSYGNS (SEQ ID NO:148) + FGGGTEVVVK (SEQ ID ID:150), the peptide sequence QSYDYGGGSYGNSSFGGTEVVVK (SEQ ID NO:151) conformed to several experimental sequences found in MS analysis. A 2+ with an intensity of 1.77e10 was found at 1164.5244 amu. Therefore, for the light chain combinatorial database, we now have:

[0286] Seq-abcde-10-2

[0287] Where “a”FR1 is any one of the eight possibilities.

[0288] any of the eight possibilities for “b”CDR1

[0289] any of the seven possibilities for “c”FR2

[0290] “d” is any of the eight possibilities of CDR2.

[0291] "e", any of the eight possibilities of FR3

[0292] The only possibility that “f”CDR3

[10] is fixed as QSYDYGGSYGNS (SEQ ID NO:148)

[0293] The only possibility that “g”, FR4[2] is fixed as FGGGTEVVVK (SEQ ID NO:150)

[0294] The κ light chain has 28,672 possibilities (8×8×7×8×8×1×2), and the heavy chain has 3,600 possibilities. Then, these two FASTA databases were generated:

[0295] Light chain: Seq-abcde-10-2;

[0296] Heavy chain: Seq-abcde-3-g.

[0297] Two FASTA databases were used in conjunction with the initial experimental set (7 different digestion conditions), which was searched using Novor cloud. The search parameters included five variable modifications: Cetyl (C), Pyro-Glu (Q), carbamoyl methylation (C), deamidation (NQ), and Pyro-Glu (E), where the Cetyl modification corresponds to a 43.042199 amu ethanolamine-like modification on cysteine. The precursor quality tolerance was 5 ppm, the fragment tolerance was 100 ppm, and the FDR was 1%. After searching the databases, the protein list obtained from Novor cloud did not decrease significantly. For example, for the light chain Seq-abcde-10-2, 26,026 proteins were reported from 28,672 sequences (a database reduction of less than 10%). For the heavy chain: Seq-abcde-3-g, 1,212 sequences were reported from 3,600 sequences (a reduction of 2 / 3). These reported proteins represent approximately 33% of the entire heavy chain database and 91% of the entire light chain database. To reduce the list of potential candidates, all identified peptides were extracted from both searches, and peptide redundancy was removed while retaining any instances of similar peptides with different modifications. Subsequently, the list of candidate protein sequences was reduced by discarding sequences showing low overlap between amino acid positions 6 and 120 in the heavy chain and between positions 6 and 110 in the light chain.

[0298] Based on peptide counting, the two candidate proteins ranked 1st and 74th in Novo Cloud, and showed the highest level of "minimum overlap score," with MO scores of 14 and 6, respectively (as shown in Figure 12A). Their respective sequences are as follows:

[0299] G001, minimum overlap score is 14: seq-2-4-6-1-4-3-2

[0300] G074, minimum overlap score is 6: seq-2-4-6-1-4-3-3

[0301] The two sequences showed a high degree of similarity, with differences only in the "FR4" region.

[0302] A FASTA database was generated containing 68 heavy chain protein candidates with a minimum overlap score greater than 0. This database includes two potential candidate sequences that differ only in the FR4 or J region, and an additional 66 sequences, all with a minimum overlap score of 3. These 68 proteins were selected from 1212 sequences reported by Novor Cloud.

[0303] As shown in Figure 12B, the distribution of minimum overlap (MO) scores for the light chains indicates that 263 proteins have the same MO score, with an MO score of 10, which is also the highest score. Therefore, only this dataset was considered when generating the FASTA file for the light chains. This resulted in 331 sequences used with the search engine pLink to improve the confidence in pairing distant regions of the heavy and light chains (68 candidates for the heavy chain and 263 candidates for the light chain). This approach aims to reduce the number of possible candidates by performing some additional orthogonal experiments. In this case, digestion was performed under non-reducing conditions to identify long-distance pairings among the selected candidates. This resulted in a smaller dataset that generated a FASTA file for searching for potential heavy and light chain candidates by using non-reducing digestion combined with pLink for long-distance region pairing. This method is used to improve the confidence in pairing distant regions (based on disulfide binding) and further reduce the number of possible candidate sequences.

[0304] For the non-reducing digestion analysis of pLink, sample preparation consists of two conditions. Condition 1 involves diluting 20 µg of the polyclonal antibody PD108 to 20 µL in 10 mM phosphate-buffered saline (PBS) containing 2 mM N-ethylmaleimide, followed by shaking incubation at 37°C for 2 hours. The sample is then degassed under low pressure in a SpeedVac™, resuspended in 100 mM tris buffer containing 25 µL of 8 M urea, pH adjusted to 6.5, and digested overnight at 37°C with LysC at a 1:20 ratio (protease to protein ratio). The next day, the sample is diluted to a final 2 M urea concentration with 100 mM tris buffer at pH 6.5 and digested with AspN at a 1:20 ratio (protease to protein ratio) at 37°C for 4 hours. Condition 2 involved 20 µg of PD108, completely dried under low pressure, then reconstituted in 100 mM Tris buffer containing 20 µL of 8M urea, adjusted to pH 6.5, and NEM was added to a final concentration of 2 mM. The mixture was then incubated at 37°C with shaking for 2 hours. LysC was added at a 1:20 ratio (protease to protein ratio) for overnight digestion at 37°C, followed by AspN digestion at a 1:20 ratio (protease to protein ratio) for 4 hours at 37°C. After protease digestion, samples from both conditions were completely dried under low pressure and reconstituted in 40 µL of 0.1% formic acid. Then, 2.5 µL of the digested sample was loaded onto Evotips according to the manufacturer's instructions and run on a Thermo Orbitrap Exploris 240 at a gradient of 44 minutes. Peptides containing disulfide bonds were identified using pLink v2. The software identification parameters were set as follows: disulfide bond (HCD-SS); enzyme: LysC-AspN, Try, or Try-GluC; peptide mass: 300-9000; peptide length: 3-90; fixed modification: Gln->pyro-Glu; and variable modifications: "N-acetylmethylene[C]", "oxidation[M]", "deamidation[N]", "deamidation[Q]", and "protein N-terminal acetylation". The protein database contained the target sequence and corresponding reverse sequence of the target pAb mixture. Table 9 shows some of the sequences identified in the non-reduction experiments. Peptides. Since peptides can correspond to several possible protein sequences, peptide assignment is limited to specific regions (FR1, CDR1, FR2, etc.), with parentheses indicating one of two peptides linked by disulfide bonds. In some cases, the information obtained from peptides is insufficient to identify a given region; therefore, ranges as shown in Table 6 are proposed based on the examples above. Peptides in the FR4 region were not reported because they were not identified. For the light chain peptides #1 and #2, CDR1[4] can pair with CDR3

[10] , thus limiting some possibilities for FR2 and FR3. Table 9 CQASQSISSYLAWYQQK (SEQ ID NO:152); DAATYYCQSY(SEQ ID NO:153);CQASQSISSYLAWYQQKPGQPPK (SEQ ID NO:154); QSVEESGGRLVTPGGSLTLTCTVSGFSLNNYPMAWVRQAPGK (SEQ ID NO:155); DTATYFCAR (SEQ ID NO:156); QSVEESGGRLVTPGGSLTLTCTVSGFSLNTYPMGWVRQAPGK (SEQ ID NO:157); QSVEESGGRLVTPGGSLTLTCTVSGFSLSSYGVSWVRQAPGK (SEQ ID NO:158); DTATYFCTR (SEQ ID NO:159); QSVEESGGRLVTPGGSLTLTCTVSGFSLSTYSMSWVRQAPGK (SEQ ID NO:160); QSVEESGGRLVTPGTPLTLTCTVSGFSLNNYPMAWVRQAPGK(SEQ ID NO:161); QSVEESGGRLVTPGTPLTLTCTVSGFSLNNYPMGWVRQAPGK (SEQ ID NO:162);QSVEESGGRLVTPGTPLTLTCTVSGFSLNNYPMGWVRQAPGK (SEQ ID NO:163); QSVEESGGRLVTPGTPLTLTCTVSGFSLNTYPMGWVRQAPGK (SEQ ID NO:164); QSVEESGGRLVTPGTPLTLTCTVSGFSLSSYAMGWVRQAPGK (SEQ ID NO:165); QSVEESGGRLVTPGTPLTLTCTVSGFSLSTYSMSWVRQAPGK (SEQID NO:166); DTATYFCAT (SEQ ID NO:167); QSVEESGGRLVTPGGSLTLTCTVSGFSLSTYAMGWVRQAPEK (SEQ ID NO:168); CQASQSVYNNDRLSWVQQK (SEQ ID NO:169)。

[0305] The light chain has one potential candidate sequence: Seq-a-4-(4-7)-d-(2,3,4,6,7,8)-10-2, which can be interpreted as having any of the eight possible FR1s, CDR1[4], FR2 having any of 4,5,6,7, any of the eight possible CDR2s, FR3 having any of 2,3,4,6,7,8, CDR3

[10] , and FR4[2] (since only peptide evidence for CDR3

[10] and FR4[2] was found). This yields a FASTA dataset of 8×1×4×8×6×1×1=1536 sequences. For the heavy chain, no CDR3 peptide was identified; however, one of the most promising candidates is Seq-2-4-6-d-4-fg, which provides a FASTA dataset of 5×6×3=90 sequences to help confirm this heavy chain. (See Table 8 for details).

[0306] Perform repeated searches on the Novor cloud platform using the following criteria:

[0307] Condition 1: Extract a smaller FASTA combinatorial dataset from the NR dataset experiment. The light chain sequence is Seq-a-4-(4-7)-d-(2,3,4,6,7,8)-10-2 (1536 candidates), and the heavy chain sequence is Seq-2-4-6-d-4-fg (90 candidates).

[0308] Condition 2: Use a larger experimental LC-MS dataset, which includes the initial dataset plus some additional LC-MS analyses obtained in “middle-down” mode (including initial different digestion conditions plus some selected additional samples such as AspN, LyC, trypsin, and LyC, chymotrypsin, and pepsin with cysteine ​​modified with ethyl bromide, run in middle-down proteomics mode to capture longer peptides).

[0309] The parameters on Novor Cloud were the same as those used in the first round of analysis. The search was performed in two steps: First, using the combined database of condition 1, the data obtained under standard Evosep conditions (standard proteomics in a bottom-up mode) were searched. Then, the search was repeated on the mid-to-bottom dataset. As mentioned earlier, a list of peptides found in both searches was extracted, retaining only non-redundant peptides.

[0310] Regarding heavy chains, Novor cloud was used to generate a list of 17 sequences from the 90 provided FASTA sequences. Minimum overlap scores were evaluated from amino acid positions 6 to 119. Table 10 ranks these 17 proteins according to their minimum overlap scores. All 17 selected proteins have a “d” value of “1” (in seq-abcdefg, therefore we have seq-2-4-6-1-4-fg). While the minimum overlap score helps exclude protein sequences with no overlap, it does not provide a detailed overview of the overall overlap score distribution. Figure 13 shows the three best-hitting sequences in Table 10 starting with seq-2-4-6-1-4; the sequence seq-2-4-6-1-4-3-2 has higher quality overlap with the included CDR3 and FR4 and is therefore retained as a potential candidate for heavy chains. Table 10

[0311] Regarding light chains, Novor cloud retained 1320 sequences from the 1536 FASTA sequences searched. Of these, 1099 had no zero-value gaps. The top 3 candidates had the same MO score of 16 and are shown in a superimposed format in Figure 14:

[0312] G0740 (Seq-1-4-5-5-8-10-2),

[0313] G1080 (Seq-1-4-6-3-8-10-2) and

[0314] G1143 (Seq-1-4-5-1-8-10-2).

[0315] All three sequences had the same minimum overlap score at amino acid 89 (Figure 14). However, plotting the overall overlap scores showed that Seq-1-4-5-5-8-10-2 (G0740) had a higher overlap score for covering the regions of CDR1, FR2, CDR2, and FR3 (indicated by the arrows in Figure 14). Therefore, this sequence was selected as the best candidate sequence and paired with the heavy chain using the experimental and bioinformatics strategies described in PCT / CA2022 / 051194.

[0316] Recombinant antibodies were prepared, and their affinity was tested using natural polyclonal antibodies. Figure 15 shows the ELISA curves of the recombinant form R1, the natural form PD108, and the negative control. The affinity of the recombinant antibody was found to be slightly higher than that of the original natural pAb, which may be due to the purity difference between the recombinant antibody and the natural polyclonal mixture.

[0317] Example 4: Using crosslinking to help long-distance sequence extension within single chains in a mixture

[0318] In this embodiment, the use of crosslinking to facilitate long-range sequence extension within a single chain was investigated in a mixture (an artificial mixture of rabbit monoclonal antibodies (internal standard names P17 and P18)). Equal volumes of antibody mixtures (40 µg each, 1 µg / µL) were prepared. BS3 crosslinking agent from ThermoFisher (catalog number A39266) was prepared to 50 mM in 25 mM sodium phosphate at pH 7.4 and added to the antibodies in excesses of 125 and 150 moles. The reaction was carried out at 25 °C for 1 hour, followed by neutralization with Tris at pH 8 (final Tris concentration 60 mM) for 30 minutes at room temperature. 20 µg of each condition was then diluted to 50 µL with HPLC-grade water and reduced with DTT (final concentration 30 mM) by heating at 95 °C for 15 minutes. Subsequently, IAA was added to a final concentration of 50 mM, and alkylation was performed in the dark at room temperature for 30 minutes. Then, acetone three times the sample volume was added at -20°C for 1 hour, followed by centrifugation at 23000 xg for 10 minutes at 4°C to precipitate the sample. The precipitate was dried under low pressure using SpeedVac™ and then resuspended in 4 µL of 4M urea at 37°C for 10 minutes. The sample was digested with 1 µg of pepsin, acidified with 2 µL of 1N HCl, and diluted to 50 µL with HPLC-grade water. The pepsin was then inactivated by heating the sample at 95°C for 3 minutes. The sample was dried under low pressure using SpeedVac™ and then reconstituted in 40 µL of 0.1% FA. 5 µg was loaded onto Evotips and pLink data were analyzed using an 88-minute gradient method on a Thermo Orbitrap™ Exploris 240.

[0319] As shown in Table 11, the highest number of sequences were observed at a 150x BS3 molar excess. In the mixture of two relatively similar antibodies, specific long-distance region pairings were observed on both the heavy and light chains, including CDR1 with CDR2 (peptides 1, 3, 14, and 18), and CDR1 with CDR3 (peptides #8 and #9). Furthermore, different frame regions also paired with other frame or CDR regions, with up to five distinct regions assembled from two linked peptides (e.g., peptide #3, which contains a single peptide from FR1, CDR1, and FR2, linked to another peptide from CDR2 and FR3). Interchain pairings were also observed (peptides 20 and 21). No crosslinking between different antibodies was observed. Table 11 ASGVPSRFKGSGSGTE (SEQ ID NO:170); IKCQASQSIY (SEQ ID NO:171); IYDASKL (SEQ ID NO:172); ETGVPSRFKGSGSGTRF (SEQ ID NO:173); NO:175); IYEASKL (SEQ ID NO:176); IKCQASQSIY (SEQ ID NO:177); KGSGSGTE (SEQ ID NO:178); KGSGSGTEFT (SEQ ID NO:179); TCKASGFD (SEQ ID NO:180); IDNO:182); FCGKD (SEQ ID NO:183); ISCQSSQSVNKNDL (SEQ ID NO:184); KGSGSGTRF(SEQ ID NO:185); ISCQSSQSVNKNDLS (SEQ ID NO:186); ISCQSSQSVNKNDLSW (SEQ ID NO:187); KGSGSGTRFTL (SEQ ID NO:188); ID NO:191); LIYDASKL (SEQ ID NO:192); YFCGKDLGL (SEQ ID NO:193); LIYEASKL (SEQ ID NO:194); KGSGSGTRFTLT (SEQ ID NO:195).

[0320] This experiment demonstrates that cross-linking followed by proteolytic digestion of a mixture of similar antibodies can pair distant regions within a single chain of the antibody-like mixture.

[0321] Example 5: Sequencing of human IgG receptor-binding domains isolated from plasma of individuals vaccinated against COVID-19

[0322] In this embodiment, sequence-specific antibodies that bind to antigens generated by human vaccination are enriched to evaluate the ability of a COVID-19 vaccine containing a receptor-binding domain (RBD) to generate antibodies against the SARS-CoV-2 virus RBD in a human host. These antigen-specific antibodies are purified and sequenced using a de novo proteomics-based approach. Following sequencing, recombinant forms are generated and detected by ELISA. The primary objective of this particular embodiment is to demonstrate that a complex mixture of intact Fab2 antibody protein fragments specific to the RBD antigen can be separated using a natural gel method, allowing for the assembly of distant regions based on hierarchical spectroscopy. Starting from the mixture of antigen-specific antibodies, the proposed method is a two-step approach comprising: 1) digesting polyclonal antibodies with several proteases (i.e., a bottom-up approach that allows for the generation of contigs) and 2) separating intact antibodies or Fab2 to assemble different contigs / distant regions.

[0323] Three samples were collected from three different healthy donors. This article presents data from one female donor (named "522"). The subject received her third dose of Moderna Spikevax. ® Blood is drawn approximately two months after receiving the COVID-19 vaccine.

[0324] IgG enriched with antigen:

[0325] Total IgG was enriched from 3 mL of human serum from subjects vaccinated against COVID-19. Briefly, 3 mL of precipitated G protein agarose resin (Genscript Cat#L00209) in a 20 mL gravity flow Biorad column (catalog number 7321010EDU) was equilibrated after washing twice with 15 mL of 10 mM phosphate-buffered saline (PBS). Serum enrichment was performed by centrifugation at 23000 RCF for 10 min at 4 °C to obtain precipitate fragments. 3 mL of serum was mixed with 9 mL of 10 mM PBS and filtered. The filtered serum was passed through the G protein agarose resin using gravity flow, and the eluent was then passed through the same G protein column twice to bind IgG. The G protein resin was washed three times with 10 mL of 10 mM PBS and then eluted with 12.5 mL of 0.1 M pH 2.5 glycine buffer. The eluted IgG was then concentrated, and the buffer was exchanged into 10 mM PBS using a 30 kDa Amicon™ filter (Sigma catalog number UFC803024). The total amount of captured IgG was found to be 22.2 mg.

[0326] Antigen enrichment was performed to enrich anti-RBD antibodies by incubating 0.4 mg of the SARS-CoV-2 spike protein receptor-binding domain (RBD) with 54 µL of streptavidin-coated agarose beads (Sigma, catalog number GE17-5113-01) at 4°C for 1 hour. First, the RBD was biotinylated by reconstitution in water with a 20-fold excess of 20 mM biotin (Sigma, catalog number A39259), followed by buffer exchange to 10 mM PBS using a 3 kDa Amicon™ filter (Sigma, catalog number UFC500396) to remove excess biotin. After conjugating the biotinylated RBD with streptavidin resin, the mixture was washed three times with 0.4 mL of 10 mM PBS to remove any unconjugated RBD. Then, 20 mg of the purified total IgG was added, and the mixture was incubated at 4°C for 1 hour. Non-specific conjugates were removed by washing twice with 0.4 mL of 10 mM PBS, followed by washing with 10 mM PBS containing 0.4 mL of 0.5% CHAPS, and then washing six times with 0.4 mL of 10 mM PBS. The anti-RBD antibody was eluted twice by incubating with 0.4 mL of 0.1 M glycine buffer (pH 2.5) at room temperature for 5 minutes. The washed antibodies were pooled and then neutralized with 0.2 mL of 1 M Tris (pH 8) to obtain 111 µg of anti-RBD antibody, internally named PD124.

[0327] In-solution digestion: A 25 µg sample of PD124 was in-solution digested by concentrating the sample to 100 µL under low pressure (i.e., Speedvac™) and then reducing it with dithiothreitol (DTT) at 95 °C for 15 min. The sample was then aliquoted into two portions: 5 / 7 was alkylated with iodoacetamide (IAA) at room temperature in the dark for 30 min to a final concentration of 50 mM; the other portion was treated with 2-bromoethylamine hydrobromide (BEA) at 25 °C with shaking for 4 h, adding 10 µL of 1 M Tris pH 8 every hour (total 40 µL) to maintain a neutral pH during the reaction. Precipitation was then carried out by adding three times the sample volume of acetone to the IAA-treated sample, incubating at -20 °C for 1 h, and then centrifuging at 23,000 xg for 10 min at 4 °C. Pour off the acetone, dry the precipitate under low pressure in a Speedvac™, then reconstitute with 10 µL of 4M urea and incubate with shaking at 37°C for 10 min to reconstitute the precipitate. Then, divide the sample into 5 tubes. Add 30 µL of 50 mM ammonium bicarbonate to 4 of these tubes and digest with trypsin, LysC, AspN, and chymotrypsin at a 1:20 protease:protein ratio. These samples are digested overnight at 37°C. For the fifth digestion, add 30 µL of HPLC-grade water and 2 µL of 1N HCl and pepsin at a protease:protein ratio of 1:20. Digest with pepsin at 37°C with shaking for 15 min, then inactivate at 95°C for 3 min. The digest is then dried under low pressure in a Speedvac™.

[0328] The BEA-treated sample was precipitated by adding trichloroacetic acid to a final volume of 20% and incubating overnight at 4°C. The next day, the sample was centrifuged at 17,000 RPM for 30 minutes at 4°C to remove the precipitate. The precipitate was then washed twice with 500 µL of 80% acetone, centrifuged at 23,000 RPM for 10 minutes after each wash. The precipitate was dried under low pressure using Speedvac™ and then shaken in 4 µL of 4 M urea at 37°C for 10 minutes. The precipitate was then divided into two tubes and digested with trypsin and LysC at a final concentration of 30 mM ammonium bicarbonate at a 1:20 protease:protein ratio. The digest was incubated overnight at 37°C. All proteases used for PD124 digestion were purchased from Promega.

[0329] Digested samples were dried to complete under low pressure in Speedvac™ and then resuspended in 40 µL of 0.1% formic acid. Following the manufacturer's instructions, 2.5 µg was loaded onto Evotips for each digestion and run on a 15 cm PepSep column on an Orbitrap™ 240 Exploris at a rate of 30 samples per day (44-minute method). Data for the precursor, ranging from 400 to 2000 m / z, were acquired at 60,000 resolution with standard AGC target and maximum injection time set to automatic. An intensity threshold of 2.5e4 was applied, and charge states 2–8 were included. Dynamic exclusion was set to 15 s. MS / MS lysis was induced using a fixed 30% HCD, and Orbitrap™ detected at 7500 resolution with a maximum injection time of 50 ms.

[0330] Natural gel electrophoresis. PD124 was separated on a BioRad pre-fabricated 7.5% polyacrylamide gel (catalog number 4561024), divided into two groups: IdeS-treated and untreated. For the IdeS-treated group, 50 μg of PD124 was incubated with 50 units of IdeS (from Promega) at 37°C for 1 hour, followed by low-pressure drying in a Speedvac™ to reduce the volume to 30 µL. NativePage 4x buffer (Thermo, catalog number BN20032) was added at a 4:1 ratio. For undigested samples, only 25 µg was loaded onto the gel after adding NativePage 4x buffer at a 4:1 ratio.

[0331] Run buffer was prepared by diluting 10X stock Tris / glycine buffer (BioRad, catalog number 1610771) to 1X with Milli-Q water. Gels were run for 180 minutes at 130V using a BioRad electrophoresis apparatus, stained with Coomassie Brilliant Blue (BioRad, catalog number 1610436) for 30 minutes, and destained overnight with BioRad destaining solution (catalog number 1610488).

[0332] The gel was cut into 12 strips (Figure 16). Strips / fractions 4, 7, 9, and 10 had higher protein abundance and were further divided into 2 or 3 fragments. The first, second, and third strips of each band (if any) were digested with trypsin, pepsin, and chymotrypsin, respectively, using a standard gel digestion protocol. The strips were dehydrated with 200 µL of 100 mM tetraethylammonium bicarbonate (TEAB) and ACN in a 1:1 ratio, allowed to stand for 5 minutes, then aspirated to remove the ACN. The strips were then further dehydrated with 200 µL of ACN for 30 seconds. After removing the ACN, the strips were air-dried at room temperature, then reconstituted in 100 mM TEAB containing 25 mM DTT and reduced by shaking at 56 °C for 30 minutes. After this, the strips were cooled to room temperature, the DTT was removed, and alkylation was performed in the dark at room temperature for 30 minutes with 55 mM IAA. To remove IAA, the band was washed twice with 0.4 mL of HPLC-grade water, then dehydrated for 5 minutes with 200 µL of 100 mM TEAB and ACN in a 1:1 ratio, aspirated, and further dehydrated for 30 seconds with 200 µL of ACN. Trypsin and chymotrypsin were diluted to 6 ng / µL in 100 mM TEAB, and 100 µL was added to a gel sheet for digestion. Pepsin was diluted to 20 ng / µL, acidified with 1N HCl, and 100 µL was added to a gel sheet for digestion. All digestions were incubated overnight at 37°C. The next day, the supernatant was collected in a new tube and dehydrated with 100 µL of 60% ACN in 40% 0.1% FA to extract additional peptides, followed by sonication for 30 minutes. The extracted supernatant was mixed with the digestion supernatant and dried under low pressure in a Speedvac™.

[0333] After gel separation, the IdeS-digested and undigested bands were cut with a blade and washed twice with 200 µL of HPLC-grade water. The darker bands were cut into multiple pieces and subjected to multiple enzymatic digestions.

[0334] The dried sample was resuspended in 40 µL of 0.1% FA, and then 100% of the sample was loaded onto Evotips according to the manufacturer's instructions and run on an Orbitrap™ 240 Exploris, using 30 samples per day on a 15 cm PepSep column (44-minute method). Data for the 400–2000 m / z precursor were acquired at 60,000 resolution with standard AGC target and maximum injection time set to automatic. An intensity threshold of 2.5e4 was applied, and charge states 2–8 were included. Dynamic exclusion was set to 15 s. MS / MS lysis was induced using fixed 30% HCD, and Orbitrap™ was detected at 7500 resolution with a maximum injection time of 50 ms. The IdeS-digested sample in Figure 16 shows clear bands, indicating the fractional separation of the polyclonal mixture of antibodies.

[0335] Data analysis. De novo sequencing was performed on the MS / MS spectra from the bottom-up data using Novor software. Candidate contigs covering each CDR region were assembled by combining overlapping de novo peptides. Assembly could proceed as long as the ambiguity of the overlap between the two peptides was low, for example, when the overlap involved multiple (>1) amino acid mutations from the germline sequence. Table 12 shows some examples of contigs obtained from this assembly. Table 12

[0336] With the aid of quantitative data obtained from natural gel separation experiments, contigs covering distant CDRs were assembled. For each unique peptide in a contig, the amount (peak area) of the peptide in each fraction was calculated using a label-free quantification method with MaxQuant software. The normalized quantification vector of the peptide was obtained by dividing its amount in each fraction by the total amount in all fractions. The normalized quantification vectors of the peptides for each contig in Table 12 are shown as curves in Figures 17A-C. It is clear from Figures 17A-C that HCDR1-c02, HCDR2-c01, and HCDR3-c01 have very similar intensity curves. This strongly indicates that these three contigs belong to the same antibody protein and should be paired together. The same conclusion can be reached by calculating the similarity score as follows: The similarity score between a pair of contigs is the average Pearson correlation coefficient between each pair of unique peptides in the two contigs. In the HCDR1 and HCDR2 contigs, HCDR1-c02 and HCDR2-c01 have the highest similarity scores with HCDR3-c01, respectively. This also leads to the conclusion that these three contigs should be paired together. As shown in Example 1, by using the isolation of complete antibodies combined with the digestion of the entire polyclonal mixture, distant regions can be assembled.

[0337] Finally, the three selected contigs HCDR1-c02, HCDR2-c01, and HCDR3-c01 were aligned with the variant germline gene, and gaps between them were filled by adding additional de novo peptides linking them. Figure 18 illustrates this process. The correct allocation of isomeric ectoleucine and isoleucine was determined by their w-ions using C-terminal chemistry in separate EThCD experiments, as described in PCT / CA2019 / 051870. After adding the constant regions, the following final sequence was obtained:

[0338] >PD124-R5-heavy (SEQ ID NO:203)

[0339] EVQLVESGGDLVQPGGSLRLSCAASGFTFSNYDMHWVRQVTGKGLEWVSGIGKDGDTYYLGSVKGRFAISRDNAKNSLYLQMNSLRAGDTALYYCARVGTTGYDLYGMDVWG QGTTVSSTSTKGPSVFPLAPSSKSTSGGTAALGCLVKDYFPEPVTVSWNSGALTSGVHTFPAVLQSSGLYSLSSVVTVPSSSLGTQTYICNVNHKPSNTKVDKKVEPKSCD KTHTCPPCPAPELLGGPSVFLFPPKPKDTLMISRTPEVTCVVVDVSHEDPEVKFNWYVDGVEVHNAKTKPREEQYNSTYRVVSVLTVLHQDWLNGKEYKCKVSNKALPAPIEK TISKAKGQPREPQVYTLPPSRDELTKNQVSLTCLVKGFYPSDIAVEWESNGQPENNYKTTPPVLDSDGSFFLYSKLTVDKSRWQQGNVFSCSVMHEALHNHYTQKSLSLSPGK (SEQ ID NO:203)

[0340] A similar process also yielded a light chain sequence:

[0341] >PD124-R5-light (SEQ ID NO:204)

[0342] SYELTQPPSVSVSPGQTARITCSGNVFPRQYAYWYQQKPGQAPVLLIYKDSERPSGIPERFSGSGSGTTVTLTITGVQAEDEADYYCQSGDSGGWVFGGGTKLTVL GQPKAAPSVTLFPPSSEELQANKATLVCLISDFYPGAVTVAWKADSSPVKAGVETTTPSKQSNNKYAASSYLSLTPEQWKSHRSYSCQVTHEGSTVEKTVAPTECS

[0343] The heavy and light chain sequences were paired by comparing quantitative changes in the separation experiments, in a manner similar to that used for contiguous groups of different CDRs (heavy chain and short chain pairing was based on procedures using experimental and bioinformatics strategies described in PCT / CA2022 / 051194). The antibody was recombinantly expressed in HEK293 cells and showed binding affinity similar to that of human anti-RBD pAb extracted from serum.

[0344] Although the invention has been described above with reference to specific embodiments thereof, modifications may be made thereto without departing from the spirit and nature of the invention as defined in the appended claims. In the claims, the word “comprising” is used as an open-ended term and is substantially equivalent to the word “including but not limited to”. Unless the context clearly specifies otherwise, the singular forms “a,” “an,” and “the” include their respective plural references.

[0345] References 1. https: / / www.marketdataforecast.com / market-reports / antibodies-market2. Wang et al.,Back to the future: recombinant polyclonal antibodytherapeutics. Curr Opin Chem En g. 2013 Nov;2(4):405-415. doi: 10.1016 / j.coche.2013.08.005.3. Cheung, W., Beausoleil, S., Zhang, X.et al.A proteomics approachfor the identification and cloning of monoclonal antibodies from serum.NatBiotechnol30, 447–452 (2012). https: / / doi.org / 10.1038 / nbt.21674. Wine et al., Molecular deconvolution of the monoclonal antibodiesthat comprise the polyclonal serum response, PNAS 110 (8) 2993-2998 (2013).5. Gilchuk et al., Proteo-Genomic Analysis Identifies Two Major Sitesof Vulnerability on Ebolavirus Glycoprotein for Neutralizing Antibodies inConvalescent Human Plasma. Front. Immunol., 15 July 2021, Volume 12 - 2021https: / / doi.org / 10.3389 / fimmu.2021.7067576. Nesvizhskii et al., Interpretation of shotgun proteomic data: theprotein inference problem. Mol Cell Proteomics. 2005 Oct;4(10):1419-40. doi:10.1074 / mcp.R500012-MCP200. Epub 2005 Jul 11.7. Guthals et al. De Novo MS / MS Sequencing of Native HumanAntibodies. J. Proteome Res. 2017, 16, 1, 45–54. October 25, 2016, https: / / doi.org / 10.1021 / acs.jproteome.6b006088. Fan et al. Using pLink to Analyze Cross-Linked Peptides. CurrProtoc Bioinformatics. 2015 Mar 9:49:8.21.1-8.21.19. doi: 10.1002 / 0471250953.bi0821s49.9. Yu, F., Li, N., and Yu, W. (2016). ECL: an exhaustive search toolfor the identification of cross-linked peptides using whole database. BMCBioinformatics, 17(1), 110. Rinner et al., Identification of cross-linked peptides from largesequence databases. Nat Methods. 2008 Apr; 5(4): 315–318. Published online2008 Mar 9. doi: 10.1038 / nmeth.1192.11. Leitner A, Walzthoeni T, Aebersold R "Lysine-specific chemicalcross-linking of protein complexes and identification of cross-linking sitesusing LC-MS / MS and the xQuest / xProphet software pipeline." Nat Protoc 2014;9(1):120-37.12. Chu F, Baker PR, Burlingame AL, Chalkley RJ. Finding chimeras: abioinformatics strategy for identification of cross-linked peptides. Mol CellProteomics. 2010;9:25–31. doi: 10.1074 / mcp.M800555-MCP200.13.Trnka MJ, Baker PR, Robinson PJ, Burlingame A, Chalkley RJ.Matching cross-linked peptide spectra: only as good as the worseidentification.Mol Cell Proteomics.2014;13(2):420–34. doi: 10.1074 / mcp.M113.034009.14. Hoopmann MR, Zelter A, Johnson RS, Riffle M, MacCoss MJ, DavisTN, Moritz RL. Kojak: efficient analysis of chemically cross-linked proteincomplexes.J Proteome Res.2015;14(5):2190–198. doi: 10.1021 / pr501321h.15. Netz et al. OpenPepXL: An Open-Source Tool for SensitiveIdentification of Cross-Linked Peptides in XL-MS. Mol Cell Proteomics. 2020Dec;19(12):2157-2168. doi: 10.1074 / mcp.TIR120.002186. Epub 2020 Oct 16.16. Pirklbauer et al. MS Annika: A New Cross-Linking Search Engine.J. Proteome Res. 2021, 20, 5, 2560–2569. Publication Date: April 14, 2021.https: / / doi.org / 10.1021 / acs.jproteome.0c0100017. Chenet al., A high-speed search engine pLink 2 with systematicevaluation for proteome-scale identification of cross-linked peptides.NatureCommunications, volume 10, Article number: 3404 (2019))18. Yang, B., Wu, YJ., Zhu, M.et al.Identification of cross-linkedpeptides from complex samples.Nat Methods9, 904–906 (2012). https: / / doi.org / 10.1038 / nmeth.2099.19. PCT application No. PCT / CA2022 / 051194, published as WO 2023 / 010219.

Claims

1. A method for determining the amino acid sequence of one or more antibodies or antibody chains present in an antibody mixture, the method comprising: (A) An antibody-derived peptide sequence library is generated by the following steps: (i) contacting multiple samples from the mixture with a reducing agent; (ii) Optionally contact the plurality of samples with an agent that prevents disulfide bond formation or modifies cysteine ​​residues into lysine analogs; (iii) contact the plurality of samples with one or more proteases and / or chemical proteolytic agents to obtain antibody-derived peptides, wherein each sample is contacted with a different protease, chemical proteolytic agent, or combination thereof to obtain a plurality of antibody-derived peptide digests; (iv) determine the amino acid sequence and intensity of short antibody-derived peptides present in the antibody-derived peptide digests by mass spectrometry using de novo sequencing methods, wherein the length of the short antibody-derived peptides is less than 50 amino acids; (v) assign the antibody-derived peptide sequences to specific complementarity-determining regions (CDR1, CDR2, or CDR3) or frame regions (FR1, FR2, FR3, or FR4); (vi) generate a candidate antibody chain sequence library on a computer by combining antibody-derived peptide sequences from different antibody regions identified in (v); (B) perform at least one of (a) to (e) to identify one or more antibodies or antibody chains present in the antibody mixture from the candidate antibody chain sequence library. The amino acid sequence is as follows: (a) (i) Optionally, the sample is contacted with an immunoglobulin (Ig) domain separating protease to obtain Fab, F(ab')2, and Fc fragments; optionally, the Fc fragment is removed (e.g., using protein A / G beads); (ii) The sample in the mixture is subjected to a separation step to separate the antibodies, antibody chains, or antibody fragments (Fab, F(ab')2, and Fc) present in the mixture into multiple fractions; (iii) Optionally, the sample from the antibody mixture is contacted with a reducing agent and / or a reagent that modifies cysteine ​​residues to prevent disulfide bond formation or to modify it to a lysine analog; (iv) The fractions are contacted with one or more proteases and / or chemical proteolytic agents to obtain a digested fraction containing antibody-derived peptides, wherein the peptides correspond to different regions of the antibody or antibody chain; (v) The digested fractions are subjected to mass spectrometry (MS) to obtain MS and / or MS / MS spectra; (vi) The MS and / or MS / MS spectra are analyzed using a proteomics search engine using the antibody-derived peptide sequence library obtained in (A). (vii) Assemble antibody-derived peptides from the same antibody or antibody chain based on their co-elution curves in different fractions; (b) (i) optionally incubate samples from antibody mixtures with reagents that modify free cysteine ​​residues; (ii) incubate samples under denaturing, non-reducing conditions; (iii) contact the samples with one or more proteases and / or chemical proteolytic agents to obtain digested antibody-derived peptides, wherein the digested antibody-derived peptides contain dimers of peptides from distant antibody regions linked by disulfide bonds; (iv) perform mass spectrometry (MS) analysis on the digested fractions to obtain MS and / or MS / MS spectra; (v) Analyze MS and / or MS / MS spectra using a proteomics search engine, the proteomics search engine being adapted to identify cross-linked peptides using the antibody-derived peptide sequences identified in (A); (c) (i) Incubating a sample from an antibody mixture with a protein cross-linking agent; (ii) Incubating the sample under denaturing and reducing conditions; (iii) Optionally contacting the sample with a reagent that modifies cysteine ​​residues to prevent disulfide bond formation or to convert cysteine ​​to a lysine analog; (iv) Contacting the sample with one or more proteases and / or chemical proteolytic agents to obtain digested antibody-derived peptides, wherein the digested antibody-derived peptides comprise cross-linked dimers of peptides from distant antibody regions; (v) Mass spectrometry (MS) is performed on the digested fractions to obtain MS and / or MS / MS spectra; (vi) MS and / or MS / MS spectra are analyzed using a proteomics search engine adapted to identify cross-linked peptides using the antibody-derived peptide sequences identified in (A); (d)(i) Generate a candidate antibody chain sequence library on a computer by combining antibody-derived peptide sequences from different antibody regions identified in (A)(v); (ii) Determine the amino acid sequence and intensity of long antibody-derived peptides present in multiple antibody-derived peptide digests in (A)(iii) by mass spectrometry, wherein the length of the long antibody-derived peptides is greater than 50 amino acids; (iii) Compare the long antibody-derived peptide sequences with the candidate antibody chain sequence library to identify antibody chain sequences present in the antibody mixture; (e)(i) Assess the overlap between antibody-derived peptides present in the sample, wherein a high level of overlap indicates that the antibody-derived peptides belong to the same antibody or antibody chain.

2. The method according to claim 1, wherein the separation step comprises chromatography or gel separation.

3. The method of claim 2, wherein the chromatography is hydrophobic interaction chromatography (HIC), and wherein the gel separation is natural gel, isoelectric focusing (IEF) gel, or 2D gel separation.

4. The method according to any one of claims 1 to 3, wherein the separation step is performed under non-reducing conditions.

5. The method according to any one of claims 1 to 4, wherein the protease is pepsin, trypsin, chymotrypsin, AspN, LysC, GluC, or any combination thereof.

6. The method according to any one of claims 1 to 5, comprising incubating a sample from an antibody mixture with a protein cross-linking agent.

7. The method of claim 6, wherein the protein crosslinking agent comprises bis(sulfosuccinimide) octanoate (BS3).

8. The method according to any one of claims 1 to 7, wherein the protein sequence database search engine is Mascot, Sequest, Novor Cloud, or Maxquant.

9. The method according to any one of claims 1 to 8, comprising contacting the sample with a reagent that modifies cysteine ​​residues into lysine analogs.

10. The method of claim 9, wherein the agent for modifying cysteine ​​residues to prevent disulfide bond formation comprises iodoacetamide, and / or the agent for modifying cysteine ​​residues into lysine analogs comprises 2-bromoethylamine hydrobromide (BEA).

11. The method according to any one of claims 1 to 10, comprising contacting the sample with a reagent that prevents the formation of disulfide bonds from cysteine ​​residues.

12. The method of claim 11, wherein the agent preventing the formation of disulfide bonds from cysteine ​​residues comprises N-ethylmaleimide.

13. The method according to any one of claims 1 to 12, wherein the Ig domain isolates the protease comprising IdeS and / or IdeZ.

14. The method according to any one of claims 1 to 13, wherein the short antibody-derived peptide has a length of 5 to 20 amino acids.

15. The method according to any one of claims 1 to 14, wherein the length of the long antibody-derived peptide is 25 to 100 amino acids.

16. The method of claim 15, wherein the length of the long antibody-derived peptide is 40 to 80 amino acids.

17. The method of any one of claims 1 to 16, wherein the level of overlap between the two antibody-derived peptides is determined by calculating an overlap score, and wherein the score of overlap between amino acids located at the amino and carboxyl termini of the antibody-derived peptide is lower than that between amino acids located at internal positions of the antibody-derived peptide.

18. The method according to any one of claims 1 to 17, wherein the method comprises performing at least (a).

19. The method according to any one of claims 1 to 18, wherein the method comprises performing at least (b).

20. The method according to any one of claims 1 to 19, wherein the method includes performing at least (c).

21. The method according to any one of claims 1 to 20, wherein the method includes performing at least (d).

22. The method according to any one of claims 1 to 21, wherein the method includes performing at least (e).

23. The method according to any one of claims 1 to 22, wherein the antibody mixture is a mixture of polyclonal antibodies or a mixture of monoclonal antibodies.

24. The method according to any one of claims 1 to 23, wherein the MS is liquid chromatography-MS (LC-MS) and the MS / MS is tandem MS / MS.

25. The method according to any one of claims 1 to 24, further comprising expressing one or more recombinant antibodies or antibody fragments, said recombinant antibodies or antibody fragments corresponding to one or more antibodies identified by the method defined in any one of claims 1 to 14.

26. The method of claim 25 further includes evaluating the binding of the recombinant antibody or antibody fragment to the target antigen.

27. The method of claim 26, wherein evaluating the binding of the antibody or antibody fragment comprises performing an immunoassay.

28. The method of claim 27, wherein the immunoassay is an enzyme-linked immunosorbent assay (ELISA).

Citation Information

Patent Citations

  • Systems and methods for antibody chain pairing

    WO2023010219A1