An antibody sequence identification method based on de novo sequencing and homologous tag graph theory

Through the binding method of de novo sequencing and homologous tag graph theory, the problems of insufficient peptide coverage and difficulty in isomer distinction in antibody sequence identification are solved, and high accuracy and stable antibody sequence assembly is achieved, which is suitable for a variety of antibody samples.

CN116779037BActive Publication Date: 2025-08-05BEIJING INST OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310815532.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2022-07-27
Filing Date
2023-07-04
Publication Date
2025-08-05
Estimated Expiration
2043-07-04

AI Technical Summary

Technical Problem

Existing de novo sequencing methods have insufficient peptide coverage and ambiguity in antibody sequence identification, especially when distinguishing leucine in isomer forms from isoleucine, resulting in difficult assembly tasks.

Method used

The method based on de novo sequencing and homologous tag graph theory was used to determine the constant region sequence of light and heavy chains through liquid phase-mass spectrometry analysis, combined with homologous search and amino acid probability distribution table, optimize the peptide assembly process, and use a variety of proteolytic methods and microwave-assisted methods to improve the overlap of peptides, and correct errors through homologous tags and kmer assembly techniques.

Benefits of technology

It improves the coverage and assembly accuracy of antibody sequences, can effectively distinguish isomers, improves the accuracy and stability of antibody sequence identification, and is suitable for a variety of antibody sample types.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116779037B_ABST
    Figure CN116779037B_ABST
Patent Text Reader

Abstract

The present invention relates to a method for antibody sequence identification based on de novo sequencing and homology tag graph theory, belonging to the field of bioanalysis technology. The method first involves cleaving the antibody to be identified, and subjecting the resulting peptide solution to liquid chromatography-mass spectrometry analysis. The light and heavy chain constant region peptide sequences are then determined through retrieval, and candidate variable region peptide sequences are determined through de novo sequencing. Finally, a dynamic windowing method, combined with a species-specific antibody homology library, is developed to effectively identify de novo sequencing errors and distinguish between isomeric forms of leucine and isoleucine, thereby improving assembly accuracy and stability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an antibody sequence identification method based on de novo sequencing and homology tag graph theory, and belongs to the technical field of biological analysis. Background Art

[0002] The amino acid sequence and post-translational modifications of antibodies are decisive factors affecting the specificity and efficacy of antibody drugs. During the R&D stage of monoclonal antibody drugs, DNA sequencing technology is mainly used to preliminarily identify the antibody gene sequence, and the corresponding antibody amino acid sequence can be obtained after codon decoding. However, for some antibody drugs used in clinical practice, they may be derived from immune hosts, commercialized antibodies, or hybridoma cell products. For many of these antibodies, their cDNA sequences are not available. In such scenarios, it is necessary to directly identify the amino acid sequence at the protein level to meet the antibody sequencing requirements in clinical trials.

[0003] In proteomics, there are two main types of mass spectrometry-based protein sequencing methods: database search and de novo sequencing. Database search methods can only identify protein sequences already in the database, while de novo sequencing methods do not rely on existing sequence information and directly deduce unknown protein sequences from secondary spectrum ion information. Due to the high variability of the variable regions of antibody sequences, antibody sequence data is often missing from existing protein sequence databases. This results in low sequence coverage when using database search methods to identify antibody sequences. To complete the identification of antibody sequences or protein sequences that do not exist in the database, de novo sequencing methods can only be used for identification.

[0004] In order to obtain as rich sequence fragments as possible, the current mainstream method is to combine multiple enzyme cutting and multiple mass spectrometry fragmentation methods to collect high-precision, high-quality mass spectrometry data, and then use de novo sequencing software to analyze and design sequence assembly algorithms to obtain the complete sequence of the antibody.

[0005] To address the lack of coverage and ambiguity in spectral interpretation caused by fragmentation during peptide sequence assembly, the Peaks Proteomics team proposed the integrated system ALPS in 2016, which automated the assembly of full-length monoclonal antibodies for the first time. This system integrates de novo sequenced peptides processed with three enzymes (Asp N, Chymotrypsin, and Trypsin), mass spectral intensities, positional confidence scores, and database error correction information into a weighted de Bruijn graph for protein sequence assembly. The ALPS system achieves 100% coverage and an assembly accuracy of 96.64%-100%, with the longest achievable assembly length being 441 amino acids. However, the sequence assembly performance of the ALPS system is susceptible to missing overlapping peptides, de novo sequencing errors, and sequence homology. In 2020, Yang Chao et al. developed a method for full protein sequence determination based on continuous enzymatic digestion with nonspecific proteases. This method constructs a continuous enzymatic digestion device and uses multiple nonspecific proteases to continuously digest proteins. By exploiting the nonspecificity of nonspecific protease digestion sites, varying digestion times, and the complementarity of peptides generated by different protease types, the diversity and overlap of peptides generated by protein digestion were increased. A protein sequence assembly algorithm was developed to assemble peptide sequences obtained by liquid chromatography-mass spectrometry (LC-MS / MS) and de novo sequencing. This method was applied to the full sequence determination of bovine serum albumin and the monoclonal antibody Herceptin. Without considering the presence of leucine and isoleucine, the sequencing accuracy for both the light chain of bovine serum albumin and Herceptin reached 100%, and the sequencing accuracy for the heavy chain of Herceptin reached 99.7%.

[0006] However, when peptide fragments with very similar masses appear in a peptide sequence, de novo sequencing methods often have difficulty distinguishing them. There are three common types of peptides with similar amino acid masses: 1) isomers I and L; 2) amino acid combinations with similar masses, such as AG = Q and GG = N; and 3) identical amino acid combinations but with different arrangements, such as AF and FA, or KR and RK. Most errors in de novo sequencing come from these three situations. Because the assembly process is based on the overlap between fragments, erroneous peptides reported by de novo sequencing methods pose a huge challenge to achieving full-length coverage of the assembly task. Selecting the correct peptide from ambiguous peptides and correcting errors at individual amino acid sites are the primary prerequisites for completing the assembly process. Summary of the Invention

[0007] In view of this, the object of the present invention is to provide an antibody sequence identification method based on de novo sequencing and homology tag graph theory.

[0008] To achieve the above object, the technical solution of the present invention is as follows:

[0009] A method for identifying antibody sequences based on de novo sequencing and homology tag graph theory, the method comprising the following steps:

[0010] (1) Perform amino acid level cleavage on the antibody to be identified to obtain a peptide solution;

[0011] (2) performing liquid phase separation and mass spectrometry analysis on the obtained peptide solution to obtain a mass spectrum file;

[0012] (3) performing pFind search on the obtained mass spectrometry file to determine the light chain constant region peptide sequence and the heavy chain constant region peptide sequence of the antibody to be identified; performing pNovo de novo sequencing on the obtained mass spectrometry file to determine the variable region candidate peptide sequence of the antibody to be identified;

[0013] (4) performing homology searches on the light chain constant region peptide sequence and the heavy chain constant region peptide sequence of the antibody to be identified, respectively, to obtain complete light chains and heavy chains with high homology to the light chain constant region and heavy chain constant region; performing statistics on the amino acid composition of each site in the variable region of the complete light chain and the complete heavy chain, respectively, to obtain a light chain variable region amino acid probability distribution table and a heavy chain variable region amino acid probability distribution table;

[0014] (5) selecting homologous tags from the light chain variable region amino acid probability distribution table and the heavy chain variable region amino acid probability distribution table, respectively, wherein the homologous tags consist of a sequence of 5 to 8 amino acids, and the amino acid sequence contains more than 2 amino acids with a probability greater than 60% and 3 amino acids with probabilities in the top 5; screening peptides containing the homologous tags from the variable region candidate peptide sequences of the antibody to be identified in step (3) to obtain high-confidence variable region peptide sequences;

[0015] (6) Divide the high-confidence variable region peptide sequence into multiple small peptide sequences (kmers) with a length of K of 5 to 8, count the number of occurrences and probability of each small peptide sequence, and obtain the assembly score of each small peptide. The small peptide sequence that meets the homology tag and has the highest score is used as the assembly starting point, and small peptides that meet the requirements of K-1 overlapping amino acids are searched along both ends. An amino acid is added to the end according to the assembly score, and the assembly is repeated to obtain the light chain variable region peptide sequence and the heavy chain variable region peptide sequence;

[0016] (7) The light chain constant region peptide sequence and the heavy chain constant region peptide sequence in step (3) are reassembled with the light chain variable region peptide sequence and the heavy chain variable region peptide sequence respectively to obtain the completed antibody light chain and heavy chain.

[0017] Furthermore, in step (1), the antibody to be identified is cleaved at the amino acid level using one or more of proteolysis, microwave hydrolysis, and microwave-assisted proteolysis, and there are more than 4 overlapping amino acids between the peptides obtained by different methods.

[0018] Furthermore, in step (1), the antibody to be identified is proteolyzed using specific proteases and non-specific proteases, respectively, and the total number of specific proteases and non-specific proteases is greater than or equal to 3.

[0019] Furthermore, in step (2), the peptide solution is subjected to liquid phase separation using mobile phase gradient elution, and a C18 reverse phase chromatographic column is used as the chromatographic column; during mass spectrometry analysis, the first 20 parent ions in the primary spectrum are selected for secondary spectrum analysis, and the parent ions are subjected to secondary fragmentation using high energy collision fragmentation mode (HCD).

[0020] Furthermore, in step (3), the database searched by pFind is the entire Swissprot library; during the search, the parent ion deviation and the fragment deviation are both ±20 ppm, and the amino acid sequence results with the highest coverage and close in length to the antibody light chain constant region sequence and the antibody heavy chain constant region sequence are selected from the search results as the light chain constant region peptide sequence and the heavy chain constant region peptide sequence of the antibody to be identified.

[0021] Furthermore, in step (3), the obtained mass spectrum file is subjected to pNovo de novo sequencing, and peptides with a parent ion mass deviation of less than 10 ppm are retained as candidate peptide sequences of the variable region of the antibody to be identified.

[0022] Furthermore, in step (4), the online antibody library abYsis is used for homology search.

[0023] Furthermore, in step (4), during homology search, the E value of the light chain constant region and the heavy chain constant region is less than 10 -5 The E-value refers to the probability of the sum of the scores assigned to each pair of amino acid residues in a random sequence of two amino acid residues of the same length, based on the scoring matrix. A smaller E-value indicates a lower probability of obtaining that total score by chance. A higher total score and a lower E-value indicate a higher homology between the two protein sequences.

[0024] Furthermore, in step (6), the assembly score of each small peptide segment is S = R × 10 P , where R is the number of occurrences of the small peptide segment and P is the amino acid probability of the small peptide segment.

[0025] Furthermore, in step (6), during assembly, if there is ambiguity at the same site, the amino acid with the highest probability in the probability distribution table is selected to be an amino acid with a probability more than three times that of the suboptimal amino acid.

[0026] Furthermore, in step (6), during assembly, the variable region C-terminus is extended by 4 to 7 amino acids to overlap with the first 4 to 7 amino acids at the N-terminus of the constant region.

[0027] Furthermore, the antibodies to be identified include mixtures of polyclonal antibodies with different specificities in serum, pure monoclonal antibodies, monoclonal antibodies secreted by hybridoma cells, or monoclonal antibodies secreted by plasma cells.

[0028] Beneficial effects

[0029] The present invention provides an antibody sequence identification method based on de novo sequencing and homology tag graph theory. First, the antibody to be identified is cut, and the resulting peptide solution is subjected to liquid chromatography-mass spectrometry analysis; then, the light chain and heavy chain constant region peptide sequences are determined by retrieval, and the variable region candidate peptide sequences are determined by de novo sequencing; then, by developing a dynamic window method combined with a species-specific antibody homology library, de novo sequencing errors are effectively identified, and isomeric forms of leucine and isoleucine can be distinguished, thereby improving assembly accuracy and stability.

[0030] During the analysis of antibody constant and variable regions determined by de novo sequencing, the antibodies were treated using three protein sequence cleavage methods (proteolysis, microwave hydrolysis, and microwave-assisted proteolysis). This ensured a high degree of overlap between antibody peptide sequences, improving coverage of the antibody protein sequence. Furthermore, the protease was selected, and both the precursor ion bias and fragment bias were set to 20 ppm to ensure accuracy at the amino acid level of the peptides.

[0031] Key parameters in the sequence assembly process include the amino acid probability threshold in the homology tag and the kmer size. A probability threshold greater than 60% in the homology tag is selected to extract highly conserved sequences for de novo peptide sequencing, thereby improving sequence assembly accuracy. kmer selections of 5, 6, 7, or 8 have complementary effects. When the kmer is small, the assembly results are susceptible to repetitive peptides, resulting in erroneous repeated amino acid fragments and reduced accuracy. However, smaller kmers can increase the length of the assembled sequence, thereby improving coverage. When the kmer is large, the assembly results are less susceptible to repetitive peptides, improving assembly accuracy. However, larger kmers filter out shorter sequences and reduce overlap between kmers, resulting in shorter assembled sequences and lower coverage.

[0032] The method described in the present invention uses a homology distribution probability table obtained through data mining to align peptides. Low-probability amino acids are replaced with high-probability ones, correcting ambiguous and isomeric amino acids while retaining peptides with the best matching scores, resulting in optimal de novo sequencing results. This method can improve the accuracy of assembled peptides and also enhance the accuracy of antibody sequence identification.

[0033] The methods described herein are applicable to a wide range of antibody sample types, including mixtures of polyclonal antibodies with varying specificities found in serum, pure monoclonal antibodies, monoclonal antibodies secreted by hybridoma cells, and monoclonal antibodies secreted by plasma cells. Antibody proteins from various sources can be cleaved using protein sequence cleavage methods at the amino acid level to form overlapping peptide segments, which can then be assembled to obtain complete antibody sequences. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] Figure 1 The present invention is a flowchart of the method.

[0035] Figure 2 It is a homologous label obtained from the probability distribution table by the method of the present invention.

[0036] Figure 3 This is a diagram of the peptide assembly method of the present invention. DETAILED DESCRIPTION

[0037] The present invention will be further described in detail below with reference to specific embodiments.

[0038] Antibodies (Ig) are large immunoglobulins (approximately 150 kDa), approximately 10 nm in size, and structurally resembling a Y. In humans and most mammals, an antibody unit is composed of four polypeptide chains: the two longer heavy chains (approximately 450-550 AA) and the two shorter light chains (approximately 214 AA). Within an antibody, the sequences of the two heavy chains and the two light chains are identical. Based on the variability of peptide chain composition, antibodies can be divided into variable and constant regions. Based on the length characteristics of the peptide chains, antibodies can be divided into four major regions: the light chain constant region (CL), the light chain variable region (VL), the heavy chain constant region (CH), and the heavy chain variable region (VH). The amino acid sequence composition of the variable regions of the light and heavy chains of antibodies varies significantly and is located near the N-terminus, occupying one-quarter and one-half of the length of the heavy and light chains, respectively. The amino acid sequence composition of the constant regions of the light and heavy chains of antibodies is relatively stable and located near the C-terminus, occupying three-quarters and one-half of the length of the heavy and light chains, respectively.

[0039] The variable regions of the light and heavy chains each contain three highly variable amino acid sequences, known as complementarity determining regions (CDRs). The variable regions also contain four more stable amino acid sequences, known as framework regions (FRs). The FRs and CDRs intersect to form the variable region. The VH and VL chains each contain three highly variable regions, known as complementarity determining regions (CDRs), namely CDR1, CDR2, and CDR3, with CDR3 being the most variable. The three CDRs of the VH chain are located at amino acids 29-31, 49-58, and 95-102, respectively, while the three CDRs of the VL chain are located at amino acids 28-35, 49-56, and 91-98, respectively. Constant regions: The C regions of the heavy and light chains are called CH and CL, respectively. The CL lengths of different types (κ or λ) of Ig are basically the same, but the CH lengths of different classes of Ig are different. For example, IgG, IgA, and IgD include CH1, CH2, and CH3, while IgM and IgE include CH1, CH2, CH3, and CH4.

[0040] Since the amino acid sequence composition of the constant region is relatively stable, this part of the sequence exists in the Swiss-prot protein sequence database. In addition, since the database search method is more accurate than the de novo sequencing method, combining these two characteristics, the database search method can be used to identify this part of the region. However, the CDR region in the variable region is highly variable, and this part of the sequence is usually not included in the protein sequence database, so it is not suitable for database search methods. Since the de novo sequencing method does not rely on the sequence database, it can complete the identification of the variable region sequence. However, the accuracy of the de novo sequencing method is low, and many ambiguous peptides contained in the sequencing results require reliable information to select reliable peptide sequences. Using the constant region sequence for homology search in the antibody database can obtain multiple homologous heavy chain and light chain sequences. The amino acid composition of each site in the variable region of these homologous sequences can be statistically analyzed to obtain an amino acid frequency distribution table of the variable region. Based on this frequency table, ambiguous peptides can be identified and the most reliable candidate variable region sequence can be selected. See the specific technical process for details. Figure 1 .

[0041] A method for identifying antibody sequences based on de novo sequencing and homology tag graph theory, the method comprising the following steps:

[0042] 1. Sample processing part (wet experiment)

[0043] 1. Antibody enzymatic hydrolysis

[0044] (1) Take 20 μg of Herceptin antibody as one portion, and take 5 portions separately. Each portion of the antibody is denatured with a 2% mass volume ratio of sodium deoxycholate aqueous solution, 200 mM Tris-HCl, and 10 mM triscarboxyethylphosphine at pH 8.0 and 95°C for 10 minutes, and then incubated at 35°C for 30 minutes for reduction. Finally, the sample is alkylated to a final concentration of 40 mM by adding iodoacetic acid, and incubated at room temperature at 25°C in the dark for 45 minutes. The antibody solution sample is stored at -4°C.

[0045] (2) Three different methods are used to cleave the antibody at the amino acid level, namely, proteolysis, microwave hydrolysis, and microwave-assisted proteolysis. The three methods can ensure that there is at least 4 amino acid overlaps between the peptides formed after protein cleavage, thereby improving the accuracy of sequencing and the coverage of antibody sequence assembly. The three amino acid level cleavage methods are first processed in step (1), and the subsequent specific implementation steps are as follows:

[0046] 1) Proteolysis method: Take 5 portions of antibody solution sample, each containing 3μg, and enzymatically hydrolyze the antibody sample at 37°C using one of the following proteases: aspartic protease, trypsin, chymotrypsin, lysine protease, and glutamic protease. The enzyme-to-antibody mass ratio is 1:50. The hydrolyzed mixture is dissolved in 100μL of 50mM ammonium bicarbonate solution. After standing for 4 hours, the 5 enzymatic hydrolysis product solutions are refrigerated at -4°C for later use. This method combines specific proteases (trypsin, aspartic protease, glutamic protease, lysine protease) with a non-specific protease (chymotrypsin) to ensure that the cleavage sites are inconsistent and overlapping peptides are formed, thereby improving the antibody sequence coverage.

[0047] 2) Microwave Hydrolysis: Place a 3 μg aliquot of the antibody solution in a glass vial and add HCl to a final concentration of 3 M. Place the vial on ice in a beaker and microwave at 200 W for 4 minutes (stop and refill with ice every minute to prevent denaturation of the antibody peptides). Store the microwave hydrolyzate at -4°C.

[0048] 3) Microwave-assisted enzymatic hydrolysis: Take 1 3μg sample of the antibody solution in a glass vial, add HCl to a final concentration of 3M, place the vial in a beaker filled with ice, and microwave-heat in a microwave oven at 200W power for 2 minutes (stop and add ice every 1 minute). The pH of the microwave hydrolyzate is adjusted to the optimal pH value of 8 for trypsin using 0.5M NaOH aqueous solution; then 0.3μg of trypsin is added. Microwave-assisted enzymatic hydrolysis of the antibody is carried out by microwave heating at 100W power for 6 minutes (stop and add ice every 1 minute). After the reaction is completed, 1M HCl is added to adjust the pH to 2 to terminate the enzymatic hydrolysis reaction. The product solution of the microwave-assisted enzymatic hydrolysis is refrigerated at -4°C for later use.

[0049] (3) To the product solutions of the cleavage methods at the three amino acid levels in step (2), 2 μL of formic acid was added respectively, and the mixture was centrifuged at 14,000 g for 20 min to remove sodium deoxycholate.

[0050] (4) After centrifugation, the supernatant was collected and desalted on a 30 μm Oasis HLB 96-well plate.

[0051] (5) The Oasis HLB adsorbent was first activated with 100 v% acetonitrile and then equilibrated with 10 v% formic acid aqueous solution.

[0052] (6) After the enzymatically hydrolyzed peptide fragments were bound to the adsorbent, they were eluted twice with a 10 v% formic acid aqueous solution and then eluted with 100 mL of a 50 v% acetonitrile and 5 v% formic acid solution.

[0053] (7) The eluted enzymatic peptide solution is vacuum dried and stored.

[0054] 2. Liquid phase separation and mass spectrometry analysis

[0055] (1) The enzymatically digested peptide samples were dissolved in 0.1 v% formic acid and analyzed online using an Agilent 1290 UHPLC coupled with an Orbitrap Q-Exactive HFX mass spectrometer. 1 μg of sample was loaded each time at a sample loading rate of 0.8 μL / min.

[0056] (2) The peptides were separated using a 15 cm C18 reverse phase column (100 μm inner diameter, 1.9 μm resin) with a 120 min elution gradient. Mobile phase A consisted of 0.1 v% trifluoroacetic acid and 2 v% acetonitrile, and mobile phase B consisted of 0.1 v% trifluoroacetic acid and 98 v% acetonitrile. A 120 min gradient (mobile phase B: 3 v% at 0 min, 5 v% at 5 min, 22 v% at 95 min, 30 v% at 105 min, 90 v% at 115 min, and 90 v% at 120 min) was used at a flow rate of 450 nL / min.

[0057] (3) Orbitrap Q-Exactive mass spectrometer parameters are as follows: spray voltage 2100 V, ion transfer tube temperature 300°C. The automatic gain control for the primary spectrum scan was 1 × 10^5, the m / z range scanned was 350–2000, and the resolution was 3 × 10^4. The first 20 parent ions in the primary spectrum were selected for secondary spectrum analysis. The parent ions were fragmented in the secondary mode using high-energy collision fragmentation (HCD) with a fragmentation energy of 30%.

[0058] 2. Database search to determine the constant region (hereinafter referred to as the dry test), specifically:

[0059] 1. The original mass spectrometry files of the three sequence cutting methods were searched using pFind. The searched sequence database was the entire Swissprot library.

[0060] 2. The main parameters of the database search are as follows: sequence database: Swiss-prot protein sequence database; protease type: 1) protease hydrolysis method: aspartic protease, trypsin, chymotrypsin, lysine protease, glutamic protease (select the protease corresponding to the wet experiment); 2) microwave hydrolysis method: no enzyme; 3) microwave-assisted protease hydrolysis method: no enzyme and trypsin (select the protease corresponding to the wet experiment); maximum number of missed cuts: 3; open search: no; parent ion deviation: ±20ppm; fragment deviation: ±20ppm; fixed modification: carbamidomethylation; variable modification: methionine oxidation; false discovery rate <1%; peptide mass range: 600-8000; peptide length: 6-40.

[0061] 3. Select the amino acid sequence results with the highest coverage and close to the length of the antibody light chain and heavy chain constant region sequences from the search results (the length of the antibody light chain constant region is approximately 107 amino acid residues, and the length of the antibody heavy chain constant region is approximately 330 amino acid residues), check their sequence names, and determine the species and the type and sequence composition of the light chain and heavy chain constant regions.

[0062] 3. De novo sequencing to obtain candidate variable region sequences, specifically:

[0063] 1. Use mass spectrometry data format conversion software to extract the mgf format files corresponding to the five enzymes from the mass spectrometry original files.

[0064] 2. Use the extracted mgf format mass spectrometry data as input for de novo sequencing in pNovo. De novo sequencing was performed for each of the three protein sequence cleavage methods, with all parameters remaining the same except for the enzyme cleavage type.

[0065] 3. The main parameters for pNovo de novo sequencing retrieval are as follows: fixed modification: carbamidomethylation; variable modification: methionine oxidation; parent ion deviation: ±20 ppm; fragment deviation: ±20 ppm; protease type: 1) protease hydrolysis method: aspartic protease, trypsin, chymotrypsin, lysine protease, glutamic protease (select the protease corresponding to the wet experiment); 2) microwave hydrolysis method: no enzyme; 3) microwave-assisted protease hydrolysis method: no enzyme and trypsin (select the protease corresponding to the wet experiment).

[0066] 4. The peptides obtained after de novo sequencing by proteolysis, microwave hydrolysis, and microwave-assisted proteolysis were merged, and the peptides with a de novo sequencing parent ion mass deviation of less than 10 ppm were retained as candidate variable region peptide sequences.

[0067] 4. Homology Search

[0068] 1. Use the determined light chain constant region sequence to perform homology search in the online antibody library abYsis and obtain a light chain constant region with an E value less than 10 -5 The E-value refers to the probability of the sum of the scores assigned to each pair of amino acid residues in a random sequence of two amino acid residues of the same length, based on the scoring matrix. A smaller E-value indicates a lower probability of obtaining that total score under random circumstances. When the total score is high and the E-value is low, the two protein sequences are more homologous.

[0069] 2. Count the amino acid composition of each site in the variable region of the complete light chain to obtain the amino acid probability distribution table of the light chain variable region.

[0070] 3. Repeat steps 1 and 2 for the heavy chain constant region sequence to obtain the amino acid probability distribution table of the heavy chain variable region.

[0071] 5. Amino Acid Frequency Distribution-Assisted Assembly

[0072] 1. Read the adjacent amino acids with a probability greater than 60% from the light chain variable region amino acid probability distribution table to obtain the homology tags of these sites (amino acid composition pattern: within a 5-amino acid window, there are 2 or more amino acids with a probability greater than 60%, and 3 amino acids in the top 3-5 probabilities.), such as Figure 2 shown.

[0073] 2. Use regular expressions to write homology tags, and search for candidate peptide fragments that meet this composition pattern in the retained results of de novo sequencing as high-confidence variable region sequences. Use multiple homology tags to further filter the de novo sequencing results and obtain a score table for each peptide segment. In this table, peptide segments are arranged in the order in which the homology tags appear, such as Figure 3 shown.

[0074] 3. Convert the high-confidence variable region sequence results into kmer with a default length K of 6 (the optional range of K is 5, 6, 7 or 8) and count the number of occurrences. Obtain the probability distribution of each kmer from the peptide score table using the formula S = R × 10 P Converted into an assembly score S (R represents the number of kmer occurrences, and P represents the amino acid probability value of the kmer). Select the kmer that meets the homology tag and has the highest score as the starting point of the assembly, and search for kmers that meet K-1 overlaps along both ends. Use the kmer score to divide the ambiguous amino acids and add an amino acid at the end. Repeat the above process to obtain a contig result for the starting point of the homology tag. Repeat the above assembly process at different homology tag starting points to obtain a series of contig results. If the high-confidence variable region sequence is ambiguous at the same site, use the probability distribution table of the 20 amino acids at the site for selection. The basis for selection is that the amino acid with the highest probability at the site is more than 3 times the probability of the suboptimal amino acid. Repeat the above process to complete the assembly of the variable region. On this basis, the C-terminus of the variable region is extended by 4 to 7 amino acids to facilitate subsequent assembly with the constant region.

[0075] 4. The light chain constant region sequence determined by searching the pFind database is reassembled with the light chain variable region sequence determined. The C-terminal 4 to 7 extended amino acids of the variable region must overlap with the N-terminal 4 to 7 amino acids of the constant region to obtain a complete antibody light chain.

[0076] 5. Repeat steps 1-4 above for the heavy chain constant region sequence to obtain a complete antibody heavy chain.

[0077] The above method was used to assemble the Herceptin antibody (214 amino acids in the light chain and 450 amino acids in the heavy chain). After alignment with the correct antibody sequence, the accuracy rates for the Herceptin antibody light chain and the Herceptin antibody heavy chain were 99.5% and 100%, respectively, with 100% sequence coverage for both the heavy and light chains.

[0078] One of the difficulties in de novo peptide mass spectrometry sequencing is distinguishing between isomeric forms of leucine and isoleucine residues. Currently, the most commonly used method is to use matrix-assisted laser desorption ionization-tandem time-of-flight mass spectrometry (MALDI-TOF / TOF) to generate w-type ions of varying masses to distinguish between these two amino acid residues. As an additional sequencing experimental method, this significantly increases sequencing costs and is easily affected by the intensity of the w-ion signal. In this protocol, the homology distribution probability table obtained through data mining methods can achieve the same effect of 100% accuracy in distinguishing isomeric leucine and isoleucine, without the need for additional complex experimental methods and high costs.

[0079] Using three methods for amino acid-level fragmentation of antibody sequences, approximately 170,000 spectrum-level antibody peptides were identified for the Herceptin antibody. After processing them into kmer-level segments and counting their occurrences, the average number of kmer occurrences was 20, significantly higher than the 10 kmer occurrences used in mainstream antibody sequencing methods. The amino acid-level confidence level for kmer-level segments with >20 occurrences was increased from 60% in mainstream methods to 90%.

[0080] Traditional mass spectrometry-based antibody sequencing methods are mostly based on multiple enzyme digestion. However, when multiple adjacent cleavage sites exist in an antibody sequence, sequence overlap cannot be guaranteed, resulting in incomplete sequence coverage. In the present invention, by introducing additional non-multiple enzyme digestion methods such as microwave hydrolysis and microwave-assisted enzymatic hydrolysis, the cleavage sites are not fixed, thereby forming diverse peptide segments. This significantly increases the overlap between peptide segments and enables stable and complete coverage of antibody sequences.

[0081] The present invention uses the amino acid probability distribution table obtained after homology alignment, and uses the probability distribution table to extract homology tags to assist in the preferred de novo sequencing peptide results. The constant region sequence identified by database search is used, and a limited homology search (limited species, sequence subtype) is performed in the abYsis antibody database. After obtaining the homologous sequence, the aligned homologous sequence is obtained through multiple sequence alignment, and the number of occurrences of 20 kinds of amino acids in each position is divided by the total number of sequences to obtain the amino acid probability distribution of the site. This process is repeated from the N-terminus to the C-terminus of the variable region to obtain the amino acid distribution probability table of the variable region. At present, the accuracy of peptide identification by de novo sequencing method is not high, and there are three forms of typical errors. Peptides are aligned using the homology probability table, and low-probability amino acids are replaced with high-probability amino acids to achieve amino acid error correction. At the same time, the peptide with the best matching score is retained to obtain the preferred de novo sequencing result. The above method can improve the accuracy of the peptides involved in assembly, and also improve the accuracy of antibody sequence identification.

[0082] Ambiguous amino acids and isomeric amino acids are corrected using amino acid probability distribution. Mass spectrometry distinguishes different amino acids based on the different masses of amino acids, but de novo sequencing methods cannot distinguish leucine and isoleucine isomers directly based on the original mass spectrum (amino acid mass information) because there is no reference sequence information. Currently, the different mass w-type ions generated by matrix-assisted laser desorption ionization-tandem time-of-flight mass spectrometry can distinguish leucine from isoleucine, but this method still has major limitations: 1) When the signal intensity of the w-type ions is weak or the w-type ions cannot be generated due to incomplete fragmentation, this method cannot distinguish leucine from isoleucine. 2) The use of w-type ions to distinguish leucine from isoleucine is limited by the length and sequence composition characteristics, and generally can only distinguish leucine from isoleucine from peptides with a length of 7 to 15 amino acids. In the present invention, a homology search is performed based on a large antibody sequence library to obtain the probability distribution of leucine and isoleucine, without the need for additional costly experiments, and the same distinction effect can be achieved.

[0083] When assembling a sequence, determining the starting point (N-terminus) and end point (C-terminus) of the sequence is one of the most difficult links. At present, the mainstream methods for determining the starting point of a protein sequence are the Edman chemical degradation method and the chemical isotope labeling method. The Edman chemical degradation method has high accuracy in determining the starting point of the sequence, but because the sequence is degraded one by one at the amino acid level, the required reaction time is long, the cost is high, and the maximum length of identification is only 30 amino acids. The N-terminus and C-terminus of the chemical isotope-labeled protein sequence are affected by the signal strength of the isotope-labeled peptide segment, and the internal peptide segment of the antibody sequence is also easily labeled, making it impossible to determine the starting point and end point of the sequence. In the present invention, the homologous tags collected after the homology search are used to determine the starting point and end point of the antibody variable region sequence. The sequence at the starting point is encoded by the V gene, and the sequence at the end point is 4 to 7 amino acids at the beginning of the constant region determined by a database search. This method does not require additional experiments. Mining homologous information based on antibody sequence data can achieve the same accuracy as the Edman chemical method, and greatly shortens the analysis time, from several hours to less than 10 minutes. Another difficulty in sequence assembly is determining the number of overlaps between peptides and how to distinguish the best overlapping peptides when ambiguity arises. In this paper, peptides sequenced de novo are processed into kmers of uniform length using a sliding window method. K-1 overlaps between different kmers are considered candidate assembly fragments. When multiple candidate assembly fragments exist, the homology probability score P and the frequency of occurrence R of the candidate assembly fragments are integrated into a comprehensive score S to distinguish ambiguous amino acids and improve assembly accuracy.

[0084] In summary, the invention includes but is not limited to the above embodiments. Any equivalent replacement or partial improvement made under the spirit and principle of the present invention shall be deemed to be within the scope of protection of the present invention.

Claims

1. A method for antibody sequence identification based on de novo sequencing and homology tag graph theory, characterized by: The method steps include: (1) Perform amino acid level cleavage on the antibody to be identified to obtain a peptide solution; (2) performing liquid phase separation and mass spectrometry analysis on the obtained peptide solution to obtain a mass spectrum file; (3) performing pFind search on the obtained mass spectrometry file to determine the light chain constant region peptide sequence and the heavy chain constant region peptide sequence of the antibody to be identified; performing pNovo de novo sequencing on the obtained mass spectrometry file to determine the candidate peptide sequences of the light chain and heavy chain variable regions of the antibody to be identified; (4) Perform homology searches on the peptide sequences of the light chain and heavy chain constant regions of the antibody to be identified, respectively, to obtain highly homologous complete light chains and heavy chains; statistically analyze the amino acid composition of each site in the variable region of the complete light chain and heavy chain, respectively, to obtain a probability distribution table of amino acids in the light chain variable region and a probability distribution table of amino acids in the heavy chain variable region; (5) obtaining a homology tag from a probability distribution table, and screening peptides containing the homology tag from the candidate peptide sequences of the variable region of the antibody to be identified in step (3) to obtain a highly reliable variable region peptide sequence; the homology tag is composed of a sequence of 5 to 8 amino acids, and the amino acid sequence contains more than 2 amino acids with a probability greater than 60% and 3 amino acids with a probability in the top 5; (6) Divide the high-confidence variable region peptide sequence into multiple small peptide sequences with a length of K of 5 to 8, count the number of occurrences and probability of each small peptide sequence, and obtain the assembly score of each small peptide. The small peptide sequence that meets the homology tag and has the highest score is used as the assembly starting point, and small peptides that meet the requirements of K-1 overlapping amino acids are searched along both ends. An amino acid is added to the end according to the assembly score, and the assembly is repeated to obtain the light chain and heavy chain variable region peptide sequences; (7) The light chain constant region peptide sequence and the heavy chain constant region peptide sequence in step (3) are reassembled with the light chain variable region peptide sequence and the heavy chain variable region peptide sequence respectively to obtain the completed antibody light chain and heavy chain.

2. The method for antibody sequence identification based on de novo sequencing and homology tag graph theory according to claim 1, characterized in that: In step (1), the antibody to be identified is cleaved at the amino acid level using one or more of the following methods: proteolysis, microwave hydrolysis, and microwave-assisted proteolysis. There are more than four overlapping amino acids between the peptides obtained by different methods.

3. The method for antibody sequence identification based on de novo sequencing and homology tag graph theory according to claim 2, characterized in that: In step (1), the antibody to be identified is proteolyzed using specific proteases and non-specific proteases, respectively, and the total number of specific proteases and non-specific proteases is greater than or equal to 3.

4. The method for antibody sequence identification based on de novo sequencing and homology tag graph theory according to claim 1, wherein: In step (2), the peptide solution is subjected to liquid phase separation by mobile phase gradient elution, and a C18 reverse phase column is used as the chromatographic column; during mass spectrometry analysis, the first 20 parent ions in the primary spectrum are selected for secondary spectrum analysis, and the parent ions are subjected to secondary fragmentation using a high energy collision fragmentation mode.

5. The method for antibody sequence identification based on de novo sequencing and homology tag graph theory according to claim 1, characterized in that: In step (3), the database searched by pFind is the entire Swissprot library; during the search, the parent ion deviation and fragment deviation are both ±20 ppm, and the amino acid sequence results with the highest coverage and close length to the antibody light chain constant region sequence and the antibody heavy chain constant region sequence are selected from the search results as the light chain constant region peptide sequence and the heavy chain constant region peptide sequence of the antibody to be identified; the obtained mass spectrometry file is subjected to pNovo de novo sequencing, and the peptides with parent ion mass deviation less than 10 ppm are retained as the candidate peptide sequences of the variable region of the antibody to be identified.

6. The method for antibody sequence identification based on de novo sequencing and homology tag graph theory according to claim 1, wherein: In step (4), the online antibody library abYsis was used for homology search; during homology search, the E value of the light chain constant region and the heavy chain constant region was less than 10 -5 Highly homologous complete light and heavy chains.

7. The method for antibody sequence identification based on de novo sequencing and homology tag graph theory according to claim 1, wherein: In step (6), the assembly score of each small peptide segment is S = R × 10 P , where R is the number of occurrences of the small peptide segment and P is the amino acid probability of the small peptide segment.

8. The method for antibody sequence identification based on de novo sequencing and homology tag graph theory according to claim 1, wherein: In step (6), during assembly, if there is ambiguity at the same site, the amino acid with the highest probability in the probability distribution table is selected to be an amino acid with a probability more than 3 times that of the suboptimal amino acid.

9. The method for antibody sequence identification based on de novo sequencing and homology tag graph theory according to claim 1, wherein: In step (6), during assembly, the C-terminus of the variable region is extended by 4 to 7 amino acids to overlap with the first 4 to 7 amino acids of the N-terminus of the constant region.

10. The method for antibody sequence identification based on de novo sequencing and homology tag graph theory according to any one of claims 1 to 9, characterized in that: The antibodies to be identified include a mixture of polyclonal antibodies with different specificities in serum, pure monoclonal antibodies, monoclonal antibodies secreted by hybridoma cells, or monoclonal antibodies secreted by plasma cells.

Citation Information

Patent Citations

  • Methods and reagents for creating monoclonal antibodies

    CN103842379A

  • Peptide fragment-based targeted proteome accurate quantification method

    CN113774074A