A method for analyzing and identifying cross-linked polypeptides

The enrichment and identification of cross-linked peptides by a carboxypeptidase Y-dependent method solves the problem of difficult disulfide bond identification in the existing technology, achieves efficient and accurate disulfide bond identification, reduces the complexity of cross-linked peptides, and improves identification efficiency.

CN115436495BActive Publication Date: 2025-09-26SHANGHAI INST OF ORGANIC CHEM CHINESE ACAD OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202110619926.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-06-03
Publication Date
2025-09-26
Estimated Expiration
2041-06-03

AI Technical Summary

Technical Problem

The lack of suitable enrichment methods in existing technologies makes it difficult to identify disulfide bonds in complex samples, especially at the omics level. In addition, existing algorithms are difficult to use, have small identification scales, and take a long time to search.

Method used

A carboxypeptidase Y-dependent method was used to enrich cross-linked peptides. By setting up an enzymatic hydrolysis time gradient experiment, disulfide bond sites were retained, and a linearized database was established. Conventional proteomics software was used for identification, and a strict target-decoy strategy was combined to control the false positive discovery rate.

Benefits of technology

It achieves efficient enrichment and identification of disulfide bonds, reduces the complexity of cross-linked polypeptides, makes disulfide bond mass spectrometry identification more common, and improves the sensitivity and accuracy of identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure SMS_1
    Figure SMS_1
  • Figure SMS_2
    Figure SMS_2
  • Figure SMS_3
    Figure SMS_3
Patent Text Reader

Abstract

The present invention discloses a method for analyzing and identifying cross-linked peptides. The method comprises enriching cross-linked peptide segments, such as disulfide bonds, isopeptide bonds, and peptides cross-linked by a cross-linking agent; data collection; establishing a linearized database; and identifying cross-linking sites using a conventional proteomics search engine. The present invention applies the CADI method to multiple systems. The feasibility of the method is first verified on a standard cross-linked peptide dataset, and the sensitivity and enrichment efficiency of the method are tested. The method is then used to identify disulfide bond sites in antibodies and multiple serum proteins, successfully identifying multiple known disulfide bond sites in these proteins.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of biotechnology, and in particular relates to a method for analyzing and identifying cross-linked polypeptides, in particular to a method for analyzing and identifying cross-linked polypeptides that are dependent on carboxypeptidase Y. Background Art

[0002] Currently, mass spectrometry identification of cross-linked peptides, such as disulfide bonds, remains challenging. The diverse nature of these cross-links, their low abundance, the lack of suitable enrichment methods, and the complexity of secondary mass spectrometry analysis are the main reasons for this difficulty. Although researchers have developed multiple algorithms for cross-linked peptides, these algorithms suffer from poor usability, typically small identification scales, and time-consuming searches. Furthermore, due to the lack of suitable enrichment methods, the successful identification of disulfide bond sites in complex samples has been unsuccessful. Therefore, identifying disulfide bonds in complex samples, especially at the omics level, remains a significant challenge. Summary of the Invention

[0003] The technical problem to be solved by the present invention is to overcome the defects of the prior art in lacking a suitable enrichment method and the corresponding difficulty in identifying disulfide bonds, and to provide a method for analyzing and identifying cross-linked polypeptides. The method includes a method for enriching cross-linked polypeptides and a simple identification of cross-linked polypeptides, particularly a carboxypeptidase Y-dependent disulfide bond analysis and identification method (CADI). The present invention utilizes carboxypeptidase Y to achieve the enrichment of disulfide bond cross-linked polypeptides for the first time, laying the foundation for the large-scale discovery of disulfide bonds and having important significance for the dynamic study of disulfide bonds in organisms.

[0004] The CADI method for disulfide bond enrichment relies on two conditions: first, that disulfide-bonded peptides are not completely digested by carboxypeptidase Y, preserving their disulfide bond sites intact. Second, that linear peptides in the sample are removed as completely as possible. Therefore, appropriate carboxypeptidase Y digestion conditions are crucial. This study employed a carboxypeptidase Y digestion time gradient experiment on both cross-linked and linear peptide samples to determine the optimal digestion time.

[0005] The present invention mainly solves the above technical problems through the following technical solutions.

[0006] One of the technical solutions of the present invention is: a method for enriching cross-linked polypeptides, the method comprising:

[0007] (1) Mixing and incubating the unreduced cross-linked protein sample to be identified with an endonuclease;

[0008] (2) mixing an exonuclease with the cross-linked polypeptide to be identified and incubating the mixture; the exonuclease includes carboxypeptidase Y and / or aminopeptidase.

[0009] The method for determining the type of protein endonuclease in step (1) can be conventional in the art. For example, the sequence of the cross-linked polypeptide can be determined before enzyme cleavage to determine the type of protein endonuclease required for enzyme cleavage.

[0010] Commonly used protein endo-cleavage enzymes include one or more of trypsin, chymotrypsin, Lys-C protease, Glu-C protease and Lys-N protease, such as trypsin and / or Lys-C protease;

[0011] When the endoproteinase is trypsin and Lys-C protease:

[0012] The relative amount of the protein endonuclease and the cross-linked polypeptide in step (1) is preferably 1:50 to 1:100 (w / w), and the enzyme cleavage system preferably further comprises 1M to 2M urea; or the incubation conditions are preferably 10 to 16 hours.

[0013] In step (2), the exonuclease is mixed with the cross-linked polypeptide to be identified and incubated to perform exo-cleavage of the cross-linked polypeptide, wherein the mixed incubation time is preferably 4 to 16 hours, for example, 12 hours.

[0014] The relative amount of the exonuclease in step (2) to the cross-linked polypeptide to be identified is preferably 1:10 to 1:100 (w / w), for example 1:50 (w / w).

[0015] The second technical solution of the present invention is: a method for analyzing and identifying cross-linked polypeptides, the method comprising the following steps:

[0016] 1) using LC-MS / MS to collect data on the product obtained by the method described in one of the technical solutions;

[0017] 2) Establishing a linearized database: The linearized database includes a positive sequence database and a bait database; wherein:

[0018] The forward sequence (FF) database includes linearized sequences, which are formed by reversing the sequence of one of the two cross-linked polypeptides after enzyme cleavage and then splicing the carboxyl ends of the two peptides together;

[0019] The decoy database includes a forward-reverse (FR) database, a reverse-forward (RF) database, and a reverse-reverse (RR) database. The FR database is the linearized sequence obtained by reversing the second peptide sequence of the cross-linked polypeptide in the FF database.

[0020] The RF database is a linearized sequence obtained by reversing the first peptide sequence of the cross-linked polypeptide in the FF database;

[0021] The RR database is a linearized sequence obtained by reversing the sequences of the two peptide segments of the cross-linked polypeptide in the FF database;

[0022] 3) Analyze the cross-linked peptide spectra using conventional proteomics search software and linearized database.

[0023] The cross-linked polypeptides described in the present invention include disulfide bond cross-linked polypeptides, isopeptide bond cross-linked polypeptides or polypeptides cross-linked by a cross-linking agent.

[0024] In a preferred embodiment of the present invention, the establishment of the linear database in step 2) comprises:

[0025] a) If it is a protein sequence, first perform simulated die-cutting according to the protein endonuclease used to obtain the polypeptide sequence of the corresponding protein. If it is a polypeptide sequence, proceed directly to step b);

[0026] b) screening the peptides containing cross-linking sites (wherein the disulfide cross-linking sites are cysteine) in the above polypeptides, and sequentially combining them in pairs to obtain cross-linked peptide sequences;

[0027] c) performing simulated enzymatic cleavage on the polypeptide sequences in the cross-linked polypeptide according to the enzymatic cleavage characteristics of carboxypeptidase Y, leaving 0-2 amino acids after the cross-linked amino acid site, i.e., each cross-linked polypeptide has 9 sequences containing different numbers of C-end AA; the amino acid abbreviation of the cross-linked site is replaced with the letter O outside the 20 common amino acid abbreviations. For example, to identify a disulfide bond site, the cross-linked cysteine ​​site is replaced with the letter C and the letter O, while the cysteine ​​at other positions is still represented by the letter C;

[0028] d) Constructing a forward sequence database: The two cross-linked polypeptides obtained in step c) are linearized by reversing the sequence of one of the peptide segments and then splicing the carboxyl ends of the two peptide segments together to form a linearized sequence; in software that can edit new amino acids, such as Mascot, a virtual amino acid J is added, and sequences containing different numbers of C-end AA are separated by the letter J to obtain a forward sequence database. In software that cannot edit amino acids, such as MaxQuant, sequences containing different numbers of C-end AA are separated by the letter U to obtain a forward sequence database;

[0029] e) Build a bait database.

[0030] In a preferred embodiment of the present invention, the data acquisition in step b) is performed using a Thermo Q-ExactiveHF mass spectrometer in DDA mode.

[0031] In a more preferred embodiment, the analytical column is prepared with 1.9 μm C18 filler and has a column length of 150 mm.

[0032] In a more preferred embodiment, mobile phase A is 98% H2O, 2% ACN, 0.1% FA; mobile phase B is 98% ACN, 2% H2O, 0.1% FA; and the liquid phase gradient is 60 min, with the concentration of mobile phase B increasing from 4% to 30% in 53 minutes.

[0033] In a more preferred embodiment, the scanning range of the primary mass spectrometer is: 350-1500 m / z, the resolution is 60000, and the AGC is 3e 6 .

[0034] In a more preferred embodiment, the secondary mass spectrometry fragmentation mode is HCD, the fragmentation energy is 27, the resolution is 15000, the maximum ion injection time is 150 ms, and the AGC is 2e 5 .

[0035] The analysis software used in step 3) may be linear protein search software such as Mascot, MaxQuant or Proteome Discovery.

[0036] In a preferred embodiment of the present invention, the analysis method used in step 3) uses software with existing scoring and FDR threshold screening results. The calculation method of FDR is conventional in the art.

[0037]

[0038] (NFR is the number of PSMs of Forward-Reverse cross-linked polypeptides, NRF is the number of PSMs of Reverse-Forward cross-linked polypeptides, NRR is the number of PSMs of Reverse-Reverse cross-linked polypeptides, and NFF is the number of PSMs of Forward-Forward cross-linked polypeptides).

[0039] In a specific embodiment of the present invention, in the Mascot results, after screening according to pep_expect≤0.05, pep_isbold=1, pep_rank=1, the FR and RF sites are re-analyzed, and the length of the reverse peptide segments of the FR and RF parts, i.e., the number of amino acids (length) and the number of matching b and y ions (ions) are calculated. According to the rule of length≤3 or the ratio of ions to length≤0.32, the FR or RF is converted to FF; then all the results are sorted from small to large by PEP, the FDR is calculated, and the results are screened according to the FDR threshold of 0.05;

[0040] Alternatively, peptides with a PEP less than 0.05 were screened out from the MaxQuant software results, and the length of the reverse peptides in the FR and RF parts, i.e., the number of amino acids (length) and the number of matching b and y ions (ions), were calculated. FR or RF was converted to FF according to the rule that length ≤ 3 or the ratio of ions to length ≤ 0.32; then all results were sorted in ascending order of PEP, and the FDR was calculated using the same formula, with an FDR threshold of 0.05.

[0041] On the basis of conforming to the common sense in this field, the above-mentioned preferred conditions can be arbitrarily combined to obtain the preferred embodiments of the present invention.

[0042] The reagents and raw materials used in the present invention are commercially available.

[0043] The positive progress effect of the present invention is:

[0044] A unique feature of the CADI method of the present invention for identifying cross-linked polypeptides is that it eliminates the need for developing specialized algorithms. Instead, conventional proteomics database search software can be used to identify disulfide bonds. In a preferred embodiment, after enzymatic digestion with carboxypeptidase Y, the complexity of the cross-linked polypeptides is reduced, leaving only a few amino acids at the carboxyl end of the sites. Based on these characteristics, the present invention establishes a linearized database of disulfide bond sites. The theoretical spectra of the linearized database can match those of actual cross-linked polypeptides. The use of conventional identification software further lowers the bar for disulfide bond identification, making mass spectrometric identification of disulfide bonds more common. Furthermore, the present invention employs a currently more stringent target-decoy strategy applicable to cross-linked polypeptides to calculate the FDR and control the false positive discovery rate.

[0045] The present invention applies the CADI method to multiple systems. The feasibility of the method was first verified on a standard cross-linked peptide dataset, and its sensitivity and enrichment efficiency were tested. The method was then used to identify disulfide bond sites in antibodies and multiple serum proteins, successfully identifying multiple known disulfide bond sites in these proteins. Finally, the present invention used the method to identify disulfide bonds in HeLa cells. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] Figure 1 Flowchart of the carboxypeptidase Y-dependent disulfide bond enrichment method.

[0047] Figure 2 Schematic diagram of the fragmentation patterns of disulfide-cross-linked peptides and linearized peptides.

[0048] Figure 3Schematic diagram of fragmentation ions of disulfide-cross-linked peptides and linearized peptides; (A) Fragmentation characteristics when cleavage occurs at the front segment of the cross-linked peptide, peptide α, bα1 = b1', yα5 = y1'; (B) Fragmentation characteristics when fragmentation occurs at the rear segment of the peptide, peptide β, y1' = bβ1, b11' = yβ5; (C) When the break occurs at the disulfide bond position.

[0049] Figure 4 The following table shows the carboxypeptidase Y digestion characteristics of disulfide-crosslinked peptides. After 16 hours of carboxypeptidase Y digestion of a standard crosslinked peptide, the carboxyl termini of the cysteine ​​residues that form the disulfide bond primarily retain 0, 1, and 2 amino acids. The horizontal axis shows the number of Cys carboxyl-terminal amino acids in the peptide without carboxypeptidase Y digestion.

[0050] Figure 5 Figure 1 shows the composition of the disulfide bond site database. (A) Schematic diagram of the target and decoy database construction. The complete database consists of FF, FR, RF, and RR. FF is formed by two peptides connected at their carboxyl ends. FR and RF are peptides with one sequence opposite to FF, while RR is formed by both sequences opposite to FF.

[0051] Figure 6 (A) Violin distribution plot of linear peptides in a database of Escherichia coli disulfide bond sites. (B) Cumulative score plot of pairwise KS-test comparisons between the four matching results in Figure A. The sample is standard cross-linked peptide data, digested with carboxypeptidase Y for 16 hours. The database is a database of Escherichia coli disulfide bond sites. Construction parameters are as follows: KR represents the endonuclease cleavage site, and 0-2 amino acids are retained after the disulfide bond site.

[0052] Figure 7 Specific matching results of the cross-linked peptide database; (A) Violin distribution plot of linear peptides in the disulfide bond site database; (B) KS-test comparison cumulative score plot of linear peptides in four cross-linked databases; the sample is standard cross-linked peptide data, digested by carboxypeptidase Y for 16 hours; the database is the disulfide bond site of Escherichia coli, and the construction parameters are as follows: KR is the endonuclease cleavage site, and 0-2 amino acids are retained after the disulfide bond site.

[0053] Figure 8 The cross-matching results of cross-linked peptides in the FF, RF, and FR databases.

[0054] Figure 9When fragmentation occurs only on one peptide segment of a cross-linked polypeptide, the same spectrum will match both the FF site and the FR\RF. (A) If fragmentation occurs on both disulfide-crosslinked peptide segments, the MS2 will only match FF. (B) If fragmentation occurs on only one of the two disulfide-crosslinked peptide segments, the MS2 can match both FF and FR, and the scores for both are the same.

[0055] Figure 10 Figure 3. Distribution of fragment ions on cross-linked peptide chains. (A) Cross-matching results of cross-linked peptides in the FF, RF, and FR databases. (B) Distribution of the ratio of fragment ions to peptide chain length for the three peptide chains A, B, and C in the Venn diagram of Figure A. (C) Distribution of fragment ions for the three peptide chains A, B, and C in the Venn diagram of Figure A.

[0056] Figure 11 Figure 2 shows the MS2 of a standard cross-linked peptide; (A) Fragment ion spectrum with a low Mascot score (13.29). (B) Fragment ion spectrum with a high Mascot score (96).

[0057] Figure 12 To protect the disulfide bond sites from being digested by carboxypeptidase Y; carboxypeptidase Y was added to the standard cross-linked peptide sample at a ratio of 1:10 (w / w), and the database was a Mascot linearized database retaining 0-2 amino acids at the carboxyl terminus.

[0058] Figure 13 The figure shows a linear peptide sample digested by carboxypeptidase Y; Figures A, B, and C show the changes in protein number, peptide number, and peptide intensity with carboxypeptidase Y digestion time, respectively; Figures D and E show the changes in tryptic peptide number and intensity with carboxypeptidase Y digestion time; and Figures F and G show the changes in semi-tryptic peptide number and intensity with carboxypeptidase Y digestion time.

[0059] Figure 14 This is the motif of a linear peptide after digestion with carboxypeptidase Y. Carboxypeptidase Y was added to the linear peptide sample at a ratio of 1:10 (w / w) and digested at 25 degrees for 0 h, 1 h, 6 h, 12 h, 20 h, and 24 h. The figure shows the motif of the 8 amino acids at the carboxyl terminus of the peptide.

[0060] Figure 15 Comparison of peptides with varying CPY sensitivities; Grand average hydropathicity (GRAVY) (A), isoelectric point (pI) (B), and length (C) of four groups of peptides with varying CPY sensitivities were compared. No significant differences were observed between the groups. GRAVY is a positive index with larger values ​​indicating greater hydrophobicity, while negative values ​​indicate greater hydrophilicity.

[0061] Figure 16 For a standard cross-linked peptide spike-in assay, various amounts of the standard peptide were mixed with an equal amount of linear peptide at a w:w ratio (1:10, 1:100, 1:1000, 1:10000). After carboxypeptidase Y digestion, 2 μg of the linear peptide was loaded. Disulfide bond identification using the Rituximab antibody.

[0062] Figure 17 (A) Identified disulfide bonds within the light and heavy chains of Rituximab. The pep_expect value is the optimal value for that site, and missing values ​​are gray. (B) Secondary spectrum of all disulfide bond sites identified in Rituximab.

[0063] Figure 18 The disulfide bond sites are known in standard proteins. After 6 hours of digestion with carboxypeptidase Y, 0-2 amino acids are retained at the carboxyl end of Cys in the CADI-Mascot database, with 4 repeats per group.

[0064] Figure 19 The unknown disulfide bond sites in the standard protein; the identification of the standard protein under different enzyme digestion conditions, the criteria for CADI-Mascot screening sites are pep_expect < 0.001 and PSM ≥ 2, and the criteria for pLink 2 screening sites are E_value < 0.001 and PSM ≥ 2.

[0065] Figure 20 The disulfide bond sites identified by RNase A were verified for synthetic peptides; A and B are disulfide bond sites not annotated by UniProt, and C is an unknown site; the spectrum on the left is the spectrum identified in the protein mixture, and the spectrum on the right is the spectrum identified after only carboxypeptidase Y digestion of the synthetic peptide.

[0066] Figure 21 These are the disulfide bond sites in HeLa cells identified by CADI-Mascot. DETAILED DESCRIPTION

[0067] The present invention is further illustrated by way of examples below, but the present invention is not limited to the scope of the examples. Experimental methods in the following examples where specific conditions are not specified were performed according to conventional methods and conditions, or selected according to the product specifications.

[0068] Example 1 Establishment of CADI method

[0069] 1. Standard cross-linking method for disulfide bonds between peptides

[0070] Ten standard peptides were artificially synthesized. Each of these peptides contains a cysteine, and the terminal end simulates the trypsin cleavage characteristics, ending with lysine or arginine. The number of amino acids left at the carboxyl end of the cysteine ​​of each peptide varies, ranging from 1 to 10 amino acids. The specific sequences are shown in Table 1. The peptides were dissolved in ddH2O, and equal amounts of peptides were mixed together. They were incubated at 37°C with shaking in a 100mM Tris-HCl solution (pH 8.5) containing 20% ​​DMSO for 17 hours. 10mM iodoacetamide (IAA) was added to alkylate the peptides that did not form disulfide bonds, and the mixture was incubated at 25°C in the dark for 15 minutes. The DMSO concentration was diluted to below 5% with 0.1% formaldehyde (FA) solution, and the FA was analyzed using Thermo Fisher Pierce TM C18 Tips were desalted, vacuum dried, and stored at -80°C.

[0071] Table 1 Standard peptide sequences

[0072]

[0073] 2. Confirm the carboxypeptidase Y enzyme cleavage characteristics

[0074] 1) Take 10 μg of successfully cross-linked standard peptide and add 10 μL of 100 mM ammonium acetate solution, pH 5, and 0.5 units of carboxypeptidase Y (185 units / mg, Thermo). Digest at 25°C for 16 hours, desalt, and spin dry. Obtain 1 μg of the treated sample and acquire data using the following LC-MS / MS method.

[0075] Data acquisition was performed on a Thermo Q-Exactive HF mass spectrometer in DDA mode. The analytical column was prepared with 1.9 μm C18 packing and had a length of 150 mm. Mobile phase A consisted of 98% H₂O, 2% ACN, and 0.1% FA; mobile phase B consisted of 98% ACN, 2% H₂O, and 0.1% FA. The liquid phase gradient lasted 60 min, with the concentration of mobile phase B increasing from 4% to 30% at 53 min. The primary mass spectrometer scan range was 350–1500 m / z, with a resolution of 60,000 and an AGC of 3 e. 6 The fragmentation mode of the secondary mass spectrometer was HCD, the fragmentation energy was 27, the resolution was 15000, the maximum ion injection time was 150 ms, and the AGC was 2e 5 .

[0076] 3) Method for establishing a standard peptide disulfide bond sequence database

[0077] The idea of ​​building a standard peptide database used in Mascot software is as follows:

[0078] a) All peptides were combined in pairs, and 55 pairs of cross-linked peptides were generated from the 10 peptides.

[0079] b) Simulate enzymatic cleavage of these peptide sequences based on the cleavage characteristics of carboxypeptidase Y. Consider retaining varying numbers of amino acids at the carboxyl end of the cysteine ​​residue; these amino acids are referred to as C-end amino acids. In experiments examining the cleavage characteristics of carboxypeptidase Y, the standard peptides were designed to contain 0-10 amino acids (C-end AA) following the cysteine ​​residue.

[0080] c) The letter C was changed to O for the cross-linked cysteine ​​site, while the cysteine ​​residues at other positions were still represented by C;

[0081] d) Constructing a forward-forward database (FF): The two peptides of the cross-linked peptide obtained in step c) are linearized by reversing the sequence of one of the peptides and then splicing the carboxyl termini of the two peptides together. In software that can edit new amino acids, such as Mascot, a virtual amino acid J can be added, and sequences containing different numbers of C-end AA are separated by the letter J to obtain a forward-forward database. In software that cannot edit amino acids, such as MaxQuant, sequences with different numbers of C-end AA are separated by the letter U to obtain a forward-forward database.

[0082] 4) CADI disulfide bond data analysis method

[0083] The raw mass spectrometry data were converted to mgf format using ProteoWizard msconvert and then searched using Mascot (version 2.5.0). First, the chemical composition of amino acid "O" was modified to C(3)H(6)NO(2)S, monoisotopic: 120.011924. The carboxypeptidase Y cleavage sites were set at the carboxyl and amino termini of amino acid J. The database was a disulfide bond site database of proteins from the corresponding species linearized using R language. The MS1 mass accuracy was set to 10 ppm, and the MS2 to 0.02 Da. Carboxypeptidase Y cleavage had no missed cleavage sites. The fixed modification was the carboxyl terminus minus the mass of one water molecule (H2O, monoisotopic: -18.010565). Variable modifications included Nethylmaleimide (+125.047678 Da, C) and oxidation (+15.994915 Da, M). The decoy database option was not checked. The search results were first filtered by significance threshold p less than 0.05 and ions score or expect cut-off less than 0.05, and the results were exported to csv format, selecting peptides with rank 1 and pep_isbold.

[0084] The following are the enzymatic cleavage characteristics of cross-linked peptides by carboxypeptidase Y:

[0085] Carboxypeptidase Y stops enzymatic cleavage due to the presence of steric hindrance of cross-linked polypeptides, but the position where it stops enzymatic cleavage is still unknown. In order to understand the enzymatic cleavage characteristics of carboxypeptidase Y on cross-linked polypeptides, the present invention carried out cross-linking and enzymatic hydrolysis experiments on 10 standard polypeptides. Each of the 10 polypeptides contains a cysteine ​​(Cys), and the carboxyl end of the cysteine ​​of each peptide segment has a different number of amino acids left. These polypeptides will produce 55 different cross-linking forms after cross-linking. The present invention qualitatively analyzes the enzymatic cleavage products of 55 cross-linked polypeptides that form disulfide bonds between the 10 standard polypeptides, and counts all enzymatic cleavage states of carboxypeptidase Y, that is, the number of amino acids at the disulfide bond cross-linking position, that is, the carboxyl end of cysteine, retains all possibilities. The present invention found that: in all cross-linking forms, mainly one amino acid is retained after the disulfide bond cross-linking position, and the situation of retaining 0 or 2 amino acids also occurs. These three situations account for about 90% ( Figure 4 ).

[0086] 3. Establishment of Decoy database

[0087] Due to the special situation of cross-linked peptides, although the present invention adopts a linear method to identify disulfide bond sites, the decoy database and FDR estimation method need to be changed for cross-linked peptides.1 and Dong Mengqiu 2 et al. proposed a similar target-decoy strategy for calculating FDR for cross-linked peptides. Based on this, the present invention constructed a decoy database suitable for the CADI method, which consists of a Forward-Reverse (FR) database, a Reverse-Forward (RF) database, and a Reverse-Reverse (RR) database. FR and RF are databases where the sequence of one disulfide bond site is opposite to the amino acid sequence in FF, while RR is databases where the sequence of both disulfide bond sites is opposite to the amino acid sequence in FF (see schematic diagram). Figure 5 Based on the characteristics of carboxypeptidase Y cleavage, only 0-2 C-end amino acids were retained in this database construction.

[0088] After obtaining mass spectrometry data for cross-linked peptides, we analyzed them using Mascot using the target and three decoy databases, respectively. Mascot software settings were the same as above. The search results were first filtered by a significance threshold of p < 0.05 and an ions score or expect cutoff of less than 0.05. The results were exported to CSV format, and peptides with rank 1 and pep_isbold were selected. The length of the reverse peptides in the FR and RF regions (i.e., the number of amino acids (length) and the number of matching b and y ions (ions)) were calculated. If length ≤ 3 or the ratio of ions to length ≤ 0.32, the FR or RF was considered FF. All results were then sorted by pep_expect in ascending order, and the FDR was calculated using the following formula. The FDR threshold was 0.05.

[0089]

[0090] Note: N FR is the number of PSMs of forward-reverse cross-linked peptides, N RF is the number of PSMs of Reverse-Forward cross-linked peptides, N RR is the number of PSMs of Reverse-Reverse cross-linked peptides, N FF is the number of PSMs in the Forward-Forward cross-linked polypeptides.

[0091] The following is the result of the Decoy database construction:

[0092] To verify the rationality of the database establishment, the present invention conducted a random matching test. When using E. coli sequences to search for standard cross-linked polypeptide data, the number of disulfide bond sites matched was very small, and their number showed a roughly 1:1:1:1 distribution in the four databases, which was basically consistent with the theory. However, it was also found that there were slightly more FR and RF matching results ( Figure 6 -A). Therefore, the Kolmogorov-Smirnov test (KS test) was performed on each matching data set to compare whether there were significant differences between the two data sets. In the KS results, the p-values ​​between the two data sets were all greater than 0.05, indicating that there were no significant differences between the four data sets. In other words, when the sequence of the target site was not among them, the result was indeed the result of random matching ( Figure 6 -B).

[0093] When the target sequence, i.e., the disulfide bond sites formed by the cross-linking of the standard cross-linked peptides, was added to the database, the correct target matches (FF) increased significantly from 21 spectra to 1,373 spectra, while the other three matching results, i.e., the matching results of the decoy database, did not change significantly ( Figure 7 -A), which shows that the database established by the present invention is indeed a cross-linked peptide segment specific database. Similarly, the present invention also performed a pairwise KS test analysis on these results. The p-values ​​of the target database matching result (FF) and the matching results of the other three bait databases are all less than 0.05, while the p-value between FR and RF is greater than 0.05. This result also shows that there is a significant difference between the matching results of the target database FF and the other databases ( Figure 7 -B).

[0094] In addition, the present invention found that in the current method using the three decoy databases, many RF and FR identification results appeared in the identification results, and these decoy matching scores were not very low. Further observation of these spectra revealed that most of them could match the correct FF site ( Figure 8 ), for example, only 5 spectra in the RF database were matched to RF alone, while the remaining 108 spectra were matched to both FF and RF. The same is true for FR. Analyzing these spectra, we found that the fragmentation information of the two peptide segments of the cross-linked peptides for which FF was the only matching result was relatively rich ( Figure 9- A). Most of the fragmentation of cross-linked peptides that match both FF and FR or RF occurs in the positive sequence of the peptide segment, or the remaining peptide segment after one of the peptide segments is removed by carboxypeptidase Y is very short, and naturally the fragmentation of this short peptide is also very rare. Take the cross-linked peptide segment of HSAILASPNPDCEK and LLYCPPETGLFLVR as an example ( Figure 9 -B), LLYCPPETGLFLVR is cleaved by carboxypeptidase Y, leaving no amino acids at the carboxyl terminus of the cysteine, and the entire peptide consists of only four amino acids. The difference in MS2 between FR and FF lies in this reverse peptide, and no fragment ions were matched to it, making the identification of this peptide inaccurate. Furthermore, 64 spectra matched simultaneously to three databases, and the majority of these matches were due to disulfide cross-links between two identical peptides.

[0095] Due to the existence of the above situation, it is hoped to establish some screening conditions to exclude these inaccurate matching results. Mascot software will give a score for each matching result. The better the matching result, the higher the score. In theory, the score of the correct cross-linked peptide should be the highest, but because the cross-linked peptide is composed of two chains, when one of them is in the positive sequence, it can also match many ions. At this time, the spectrum of this decoy will also have a higher score, so the decoy and target cannot be separated well by the score alone. You need to find a suitable value to distinguish them. Analysis Figure 10 -A matches only the spectrum of FF (part A), the spectrum of FF and FR (part B), and the spectrum of FF and RF (part C), and statistically compares the length ratios of the ions matched on each peptide segment of the cross-linked polypeptide to the length of this peptide segment. The results show that the ion length ratio on the forward peptide chain (i.e., the correct hit) is much higher than the ion length ratio on the decoy peptide chain. The minimum score calculated to include 95% of the peptide segments in part A is 0.32. It was found that when the ion length ratio is greater than 0.32, most of the forward peptide chains are included, and most of the decoy peptide chains are excluded ( Figure 10 -B). The present invention also counted the number of secondary ions (b, y ions) identified on the forward peptide chain and the reverse peptide chain. The results showed that the number of fragmentation ions on the forward peptide chain was significantly higher than that on the decoy chain. Similarly, the present invention selected the minimum number of ions of 3, which includes 95% of the A part of the peptide segment, as the threshold ( Figure 10-C). That is to say, when the number of fragmentation ions on this peptide chain is greater than or equal to 3 and its secondary ion length ratio is greater than 0.32, it is not a random match. The peptide chain that does not meet this condition is an incorrect match.

[0096] 4. Identification of standard cross-linked peptides using CADI-Mascot and CADI-MaxQuant

[0097] The present invention first applies the CADI method to a standard peptide dataset and compares it with the current pLink 2 software that identifies a larger number of disulfide bonds.

[0098] 1) Sample Preparation: Using cross-linked samples from the same batch, two groups were divided. One group was enriched with carboxypeptidase Y, while the other was not treated with carboxypeptidase Y. 1 μg of each treated sample was collected on a Thermo QE-HF mass spectrometer using the same data acquisition method as in step 2 below. Libraries for disulfide bond identification and analysis were constructed using the CADI-Mascot and CADI-MaxQuant methods for the carboxypeptidase Y-enriched samples, respectively.

[0099] 2) LC-MS / MS Data Acquisition Method: Data acquisition was performed on a Thermo Q-Exactive HF mass spectrometer in DDA mode. The analytical column was prepared with 1.9 μm C18 packing and had a length of 150 mm. Mobile phase A consisted of 98% H₂O, 2% ACN, and 0.1% FA; mobile phase B consisted of 98% ACN, 2% H₂O, and 0.1% FA. The liquid phase gradient lasted 60 min, with the concentration of mobile phase B increasing from 4% to 30% at 53 minutes. The primary mass spectrometer scan range was 350–1500 m / z, with a resolution of 60,000 and an AGC of 3e. 6 The secondary mass spectrometer fragmentation mode was HCD, the fragmentation energy was 27, the resolution was 15000, the maximum ion injection time was 150 ms, and the AGC was 2e 5 .

[0100] 3) Database Construction: The Mascot database construction approach is the same as above. Based on the results of the carboxypeptidase Y enzyme digestion characteristic experiment, only 0-2 C-end amino acids are considered in the database construction here. The MaxQuant database construction approach is the same as Mascot. However, since MaxQuant software cannot independently add new amino acids to the amino acid list or modify the amino acid composition, when constructing the MaxQuant database in the present invention, the disulfide bond database is constructed using minor encoded amino acids other than standard amino acids, namely selenocysteine ​​(U) and pyrrolysine (O). (Currently, only 25 selenoproteins have been identified, and human proteins do not contain pyrrolysine.)

[0101] 4) Analytical methods:

[0102] The Mascot search parameters used in this study were the same as above, and the results were analyzed as follows:

[0103] After filtering the results using pep_expect ≤ 0.05, pep_isbold = 1, and pep_rank = 1, the FR and RF sites were reanalyzed. The length of the reverse peptides (i.e., the number of amino acids) and the number of matching b and y ions (i.e., ions) in the FR and RF regions were calculated. FR or RF was converted to FF using the rule that length ≤ 3 or the ratio of ions to length ≤ 0.32. All results were then sorted by PEP in ascending order, and the FDR was calculated using the same formula, with an FDR threshold of 0.05.

[0104] MaxQuant version 1.6.0.1 was used for CADI data analysis. Carboxypeptidase Y was first added to the enzymes.xml file according to the specified file format. The enzyme cleavage sites were the carboxyl and amino termini of amino acid U, and no missed sites were set for carboxypeptidase Y. Fixed modifications included the subtraction of one water molecule from the carboxyl terminus, a chemical composition of H(-2)O(-1), a monoisotopic value of -18.010565 Da, and a mass difference of -117.1358027 Da from amino acid O, a chemical composition of H(-13)C(-9)N(-2)S, and a monoisotopic value of -117.1358027 Da. Variable modifications included oxidation (+15.994915 Da, M) and niethylmaleimide (+125.047678 Da, C). The match interval run was 5 min. The mass accuracy of the primary mass spectrometer was set to 10 ppm, and the secondary mass accuracy was set to 0.05 Da. The FDR at both the PSM and protein levels was set to 1. Peptides with a PEP less than 0.05 were selected from the results. The length of the reverse peptides in the FR and RF regions (i.e., the number of amino acids and the number of matching b and y ions) was calculated. The FR or RF regions were converted to FF using the rule that length ≤ 3 or the ratio of ions to length ≤ 0.32. All results were then sorted by PEP in ascending order and the FDR was calculated using the same formula, with an FDR threshold of 0.05.

[0105] pLink software version 2.3.7 was used. The Disulfide Bond (HCD-SS) module was selected for disulfide bond identification, with the linker set to SS. The restriction sites were set according to the endonuclease used, with three missed cleavage sites set. The mass accuracy for both primary and secondary mass spectrometry was 20 ppm. The variable modification was Nethylmaleimide (+125.047678 Da, C). The false detection rate (FDR) at the PSM level was set to 0.01, and the E-value was calculated. Sites with an E-value less than 0.01 and at least two PSMs were selected as reliable sites.

[0106] The following are the results of CADI-Mascot and CADI-MaxQuant identification of standard cross-linked peptides:

[0107] The CADI-Mascot method identified 53 of the 55 theoretical sites, while the CADI-MaxQuant method and pLink 2 only identified 51 sites (Table 2).

[0108] CADI-Mascot identified two more disulfide bond sites than pLink 2, indicating that the CADI method has more advantages than the traditional direct identification method.

[0109] Table 2 Number of disulfide bonds in standard peptides obtained by different software

[0110]

[0111] To further confirm the reliability of the identified sites, the present invention checked the secondary spectra of the identified peptides. The lowest score was LNEQASEOEVOAV, with a score of 13.29. Because this peptide produced fewer secondary ions, the score was low. However, fragmentation occurred on both cross-linked peptides, and each peptide had ions containing the mass of the other peptide ( Figure 11 -A), and similar ion features to the high-scoring spectrum (score 96) ( Figure 11 -B), therefore, the identification of this spectrum is reliable.

[0112] Example 2 Optimization of Carboxypeptidase Y-Dependent Disulfide Bond Enrichment Conditions

[0113] 1. Preparation of linear peptide samples

[0114] This step obtains linear peptide samples, so a conventional proteolysis process is used. The samples are derived from HEK293T cells, and the linear peptide samples involved in the following examples are all obtained using this method.

[0115] All protein extraction steps were performed on ice. HEK293T cells were washed three times with pre-chilled PBS, lysed with RIPA buffer, sonicated for 30 seconds, and centrifuged at 18,000 g for 10 minutes. The supernatant was collected and protein concentration was determined by BCA assay. 10 μg of protein was precipitated with 4 volumes of pre-chilled acetone and re-dissolved in 8 M urea (pH 8.5) before use in the next step. 1M TCEP (final concentration 5mM) was added for 20 minutes at room temperature to reduce disulfide bonds. 500mM IAA (final concentration 10mM) was added for 25 minutes at room temperature in the dark. The sample was diluted to 2M urea with 100mM Tris-HCl (pH 8.5), and trypsin (1:50, w / w) and a final concentration of 1mM CaCl2 were added. After incubation at 37°C for 16 hours, 90% formic acid (final concentration 5%) was added to terminate the digestion. The supernatant was desalted by centrifugation at 18,000g for 10 minutes. The sample was desalted using a C18 tip, dried under vacuum, and stored at -80°C.

[0116] 2. Carboxypeptidase Y time gradient experiment

[0117] The enzymatic activity of carboxypeptidase Y produced by Sigma-Aldrich is 72 units / mg. This method is used to optimize the enzymatic cleavage conditions of carboxypeptidase Y produced by Sigma.

[0118] 1) Dissolve 10 μg of linear peptide or cross-linked standard peptide sample from HEK293T cells in 100 mM ammonium acetate, pH 5, at a peptide concentration of 1 μg / μL. Add carboxypeptidase Y and incubate at 25°C for 16 h at a carboxypeptidase to peptide ratio of 1:10 (w / w). Five time points were set: 1 h, 6 h, 12 h, 20 h, and 24 h, with 0 h being the time when no carboxypeptidase Y was added. After digestion, remove salts, and collect 1 μg of the treated sample using the following LC-MS / MS parameters.

[0119] 2) Data acquisition was performed on a Thermo Q-Exactive HF mass spectrometer in DDA mode. The analytical column was prepared with 1.9 μm C18 packing and had a column length of 150 mm. Mobile phase A was 98% H2O, 2% ACN, and 0.1% FA; mobile phase B was 98% ACN, 2% H2O, and 0.1% FA. The acquisition parameters for the standard cross-linked peptide were the same as in Example 1. The acquisition parameters for the linear peptide were as follows: a 60-min liquid phase gradient, with a mobile phase B concentration gradient consistent with the standard cross-linked peptide gradient, increasing from 4% to 30% at 53 minutes. The primary mass spectrometer scan range was 350–1500 m / z, with a resolution of 60,000 and an AGC of 3e. 6The secondary mass spectrometer fragmentation mode was HCD, the fragmentation energy was 27, the resolution was 30000, the maximum ion injection time was 45ms, and the AGC was 1e 5 .

[0120] 3) The collected CADI disulfide bond data were analyzed by Masscot and MaxQuant, respectively, and the data of carboxypeptidase Y-digested linear peptides were analyzed by MaxQunat for label-free quantitative analysis.

[0121] Linear peptide data were analyzed and identified using MaxQuant software. The human database was downloaded from SwissProt in July 2019. Trypsin was used as the protease, and the carboxyl-terminal semi-cleavage (semispecific free C-terminus) was used. Carbamidoformylation of cysteine ​​(+57.021464 Da, C) was set as a fixed modification, while oxidation of methionine (+15.994915 Da, M) and amino-terminal acetylation of the protein (+42.010564) were set as variable modifications. The match between runs were 2 min. The mass accuracy of the primary mass spectrometer was set to 10 ppm, and the secondary mass accuracy was set to 0.05 Da. The FDR at both the PSM and protein levels was set to 0.01.

[0122] 4) Construction of CADI database: Based on the conclusion of Example 1, the CADI database used in Example 2 only needs to retain 0-2 C-end amino acids, and the rest of the construction method is the same as Example 1.

[0123] The following are the results of the carboxypeptidase Y time gradient experiment:

[0124] In the cross-linked polypeptide sample, the number of spectra that meet the characteristics of carboxypeptidase Y enzymatic digestion increased significantly after 1 hour of enzymatic digestion. After 12 hours of enzymatic digestion, the number of disulfide bond site spectra was the largest, and then it slowly decreased, but its number was also equivalent to 1 hour, indicating that as the enzymatic digestion time of carboxypeptidase Y increased, some peptides containing disulfide bonds did not resist the enzymatic digestion of carboxypeptidase Y, and their disulfide bond sites or complete sequences were completely enzymatically digested. The present invention counted the changes in the number of disulfide bond site identifications over time and found that the number of disulfide bond sites obtained after 6 hours of enzymatic digestion was the largest, with 51 disulfide bond sites identified, and then the number gradually decreased ( Figure 12 ).

[0125] Similarly, carboxypeptidase Y digestion experiments were also performed on linear peptide samples (peptide mixtures obtained from HEK293 cell whole protein lysate after trypsin digestion). As the digestion time increased, the number and intensity of peptides in the linear peptide samples gradually decreased, although the number of proteins identified by these peptides only decreased by 17% ( Figure 13-A), but the peptide intensity decreased by about 93%. After 12 hours of enzyme digestion, it decreased by one order of magnitude compared to the initial state. In addition, the number of peptides decreased by about 50% ( Figure 13 -B and C), the decrease in the number of proteins is not obvious due to the high sensitivity of mass spectrometry. Even if most of the peptides in a protein are cleaved by enzymes, as long as there are peptides belonging to it, the protein can still be identified. The present invention also analyzed tryptic peptides (the carboxyl end of the peptide is K or R) and semi-tryptic peptides (the amino end meets the characteristics of trypsin cleavage and the carboxyl end is any amino acid) and found that after 1 hour of enzymatic hydrolysis by carboxypeptidase Y, the number of tryptic peptides decreased ( Figure 13 -D and E), corresponding to which the amount of semi-tryptic peptide increased rapidly ( Figure 13 -F and G), indicating that carboxypeptidase Y began to play a role, and then the number and intensity of both gradually decreased. However, after 24 hours of enzymatic hydrolysis by carboxypeptidase Y, there were still more than 5,000 tryptic peptides, accounting for about 50% of the remaining polypeptides.

[0126] The present invention analyzed the motifs of the peptides in each time period and found that even after 24 hours of enzyme digestion, there were still a large number of peptides ending with KR that were not digested by enzyme, indicating that there were still a large number of peptides that were not digested by carboxypeptidase Y ( Figure 14 ), which is also related to Figure 13 The conclusion is consistent with the conclusion. This part of the protein was divided into peptides ending with K, R and peptides not ending with KR based on the last amino acid. The amino acid composition of the peptides was analyzed separately. It was found that carboxypeptidase Y had low enzymatic efficiency for proline and glycine, and low enzymatic activity for peptides with lysine at the carboxyl end and proline at the previous position. This enzymatic cleavage feature is consistent with the current reports on the enzymatic cleavage activity of carboxypeptidase Y. 3 In order to reduce the number of peptides with PK at the carboxyl terminus, we also tried to use Lys-N and trypsin together, but the effect was limited, and carboxypeptidase Y still could not completely enzymatically hydrolyze linear peptides.

[0127] In order to analyze the reasons why these peptides are not easily cleaved by enzymes, we counted some biochemical properties of peptides that are easily cleaved by carboxypeptidase Y and peptides that are not easily cleaved by carboxypeptidase Y. We found that there was no significant difference between the two parts in isoelectric point, length, and hydrophilicity. Figure 15 ).

[0128] The efficiency of CADI disulfide bond enrichment is closely related to the carboxypeptidase Y digestion conditions. These results confirm the basic strategy for disulfide bond enrichment using the CADI method. For simple samples such as peptides or pure protein mixtures, where background signal (i.e., linear peptides) is low, the carboxypeptidase Y digestion time can be reduced. For example, in this experiment, carboxypeptidase Y digestion alone only required 1 hour to clear approximately 88% of disulfide bonds in the standard peptide. For complex samples such as cell lysates, the carboxypeptidase Y digestion time should be increased, with 12-16 hours being the optimal time, depending on enzyme activity.

[0129] Example 3 CADI method enrichment efficiency detection

[0130] In order to test the enrichment efficiency of the CADI method, the present invention conducted a spike-in experiment using a standard cross-linked peptide dataset. The cross-linked peptides formed between 10 standard peptides were mixed in a linear cross-linked peptide sample at different ratios. The detection limits of the standard cross-linked peptides and linear peptides at different mixing ratios were calculated. The method flow chart is shown in the figure below. Figure 16 shown.

[0131] 1) Sample Preparation: Dissolve the linear peptide and standard cross-linked peptide samples in 100 mM ammonium acetate, pH 5, at a peptide concentration of 1 μg / μL. Add the standard cross-linked peptide to 10 μg of the linear peptide at ratios of 1:10, 1:100, 1:1000, and 1:1000, respectively. Digest the peptides at 25°C for 12 hours, desalt, and dry them by spin drying. Data were acquired on a Thermo Fisher Scientific QE-HF using 2 μg of each treated sample. Equal amounts of untreated cross-linked and linear peptides were loaded directly onto the samples and analyzed using pLink.

[0132] 2) LC-MS / MS Data Acquisition: The data acquisition method was as follows: a 60-min liquid phase gradient was used, with the concentration gradient of mobile phase B increasing from 4% to 30% at 53 minutes, consistent with the gradient of the standard cross-linked peptide. The primary mass spectrometer scan range was 350-1500 m / z, with a resolution of 60,000 and an AGC of 3e. 6 The secondary mass spectrometer fragmentation mode was HCD, the fragmentation energy was 27, the resolution was 30000, the maximum ion injection time was 45ms, and the AGC was 1e 5 .

[0133] 3) Analysis and identification: The collected data were analyzed using CADI-Mascot and pLink respectively. The CADI-Mascot database construction method and software parameters were the same as in Example 1. The data not treated with carboxypeptidase Y were identified using pLink2.3.7 for cross-linked polypeptides. The Disulfide Bond (HCD-SS) module was selected for disulfide bond identification. The linker was set to SS. The cleavage site was set according to the protein endonuclease used. Three missed cleavage sites were set. The mass accuracy of the primary and secondary mass spectra was 20ppm. The variable modification was Nethylmaleimide (+125.047678Da, C), the FDR of the PSM level was set to 0.01, and the E-value was calculated. Sites with an E-value less than 0.01 and at least two PSM numbers were selected as credible sites. The pLink software parameters were the same as in Example 1.

[0134] The following are the results of the enrichment efficiency of the CADI method:

[0135] When the mixing ratio of standard cross-linked peptides to linear peptides was 1:10 and 1:100, there was no significant difference in the number of disulfide bonds identified by pLink 2 and CADI-Mascot, with CADI-Mascot slightly exceeding pLink 2. However, when the cross-linked peptides only accounted for 0.01% of the sample, CADI-Mascot was superior to pLink 2 in identifying small libraries (Table 3).

[0136] Table 3 Identified disulfide bonds of standard cross-linked peptides in the spike-in experiment

[0137]

[0138] Note: Different amounts of standard peptides were mixed with equal amounts of linear peptides at a ratio of 1:10, 1:100, 1:1000, and 1:10000, w:w. After digestion with carboxypeptidase Y, 2 μg of the linear peptide was loaded. Each group had two replicates. The linear peptide sample was a mixture of peptides digested with trypsin from HEK293T cell lysate.

[0139] Example 4 CADI method sensitivity detection

[0140] Dissolve 10 μg of standard cross-linked peptide in 10 μL of 100 mM ammonium acetate (pH 5). Add carboxypeptidase Y at a 1:10 (w / w) ratio and digest at 25°C for 1 hour. Desalt and spin dry. Dilute the carboxypeptidase Y-digested standard cross-linked peptide sample 10-, 100-, 1000-, and 10,000-fold, respectively, and load the mass spectrometer at 0.1 μg, 0.01 μg, 0.001 μg, and 0.0001 μg.

[0141] The data acquisition method and CADI-Mascot database construction method are the same as in Example 2, the Mascot software parameters are the same as in Example 2, and the pLink software parameters are the same as in Example 1.

[0142] The following are the sensitivity test results of the CADI method:

[0143] The number of disulfide bond identifications decreased with decreasing sample load. The number of pLink identifications for the directly diluted sample decreased significantly at a 1000-fold dilution, from 38 sites to only 6 sites. However, the CADI-enriched sample still identified 33 sites at a 1000-fold dilution, and 6 sites at the highest dilution (Table 4). When E. coli sequences were added to the database, both methods continued to identify disulfide bond sites. Furthermore, both methods also identified non-target disulfide bond sites, namely those in E. coli.

[0144] Table 4 Number of disulfide bond sites in standard cross-linked polypeptides

[0145]

[0146] Note: pLink search samples: Standard peptide dataset (55 disulfide bond sites) were loaded at 0.1 μg, 0.01 μg, 0.001 μg, and 0.0001 μg, respectively; CADI-Mascot samples, peptides diluted after 6 hours of carboxypeptidase digestion were loaded, and the loading amount was the initial peptide amount, 0.1 μg, 0.01 μg, 0.001 μg, and 0.0001 μg, respectively; P10X is a standard cross-linked peptide disulfide bond site database, and the carboxyl terminus of cysteine ​​in the CADI-Mascot and CADI-MaxQuant databases retains 0-2 amino acids; each dilution gradient has 4 replicates.

[0147] Example 5 CADI method for identifying disulfide bonds in simple protein samples

[0148] 1. Identification and Analysis of Disulfide Bonds in Rituximab

[0149] Take 10 μg of Rituximab antibody, add 4 times the volume of pre-cooled acetone, precipitate for 30min and then dissolve in 8M urea solution (pH 6.5), add NEM with a final concentration of 2mM, react at 37 degrees for 2h to alkylate the free cysteine, then use different protein endonuclease combinations for enzymatic digestion, and the enzymatic digestion conditions are shown in Table 5. After the protein endonuclease is completed and desalted and spin-dried, the sample is redissolved in 10 μL of 100mM ammonium acetate solution, pH 5, 1 μg of carboxypeptidase Y is added, and the enzyme digestion is carried out at 25°C for 16h, desalted and spin-dried. The mass spectrometry acquisition method is the same as in Example 3. The data is analyzed using the CADI-Mascot method. After obtaining the antibody sequence, the library construction method, Mascot software parameters and analysis method are the same as in Example 1.

[0150] Table 5 Endoproteinase combinations and cleavage conditions for Rituximab

[0151]

[0152] The following are the results of the identification and analysis of disulfide bonds in Rituximab:

[0153] Disulfide bonds in antibodies are crucial for maintaining their structure and function. Rituximab is a chimeric monoclonal antibody targeting human CD20, primarily used to treat diseases such as non-Hodgkin lymphoma (NHL), chronic lymphocytic leukemia (CLL), and rheumatoid arthritis (RA). Rituximab has 15 disulfide bonds: eight in the heavy chain, four in the light chain, two in the hinge region, and one between the heavy and light chains. The disulfide bonds on both heavy and light chains are identical.

[0154] In order to obtain as many disulfide bond sites as possible, the present invention used Trypsin, Chymotrypsin single enzyme digestion and Trypsin, Chymotrypsin combined use on Rituximab, successfully obtained all the intrachain disulfide bond sites ( Figure 17-A). The number of disulfide bonds identified by digestion with either trypsin or chymotrypsin alone was lower than that identified by the combined digestion of both endoproteinases. This may be because single disulfide-bonded cross-linked peptides cannot be obtained using a single endoproteinase alone, and the cross-linked peptides obtained by dual digestion with both endoproteinases are of a length more suitable for mass spectrometry detection. The disulfide bond counts obtained by the combined digestion of trypsin and chymotrypsin include all intrachain disulfide bonds in Rituximab. The hinge region cysteines differ by only two amino acids, making single disulfide-bonded cross-linked peptides difficult to detect. The disulfide bonds between the heavy and light chains are not included in this database. Figure 17 -B shows the secondary spectra of all disulfide bond sites in Rituximab.

[0155] 2. Identification and analysis of disulfide bonds in standard protein samples

[0156] Four standard proteins, albumin, lysozyme, transferrin, and RNase A, were dissolved in 8 M urea solution (pH 6.5) at a protein concentration of 2 μg / μL. After equal volumes were mixed, 20 μg of the protein mixture was added to NEM at a final concentration of 2 mM and reacted at 37°C for 2 h to alkylate the free cysteines. Different combinations of protein endonucleases were then used for enzymatic digestion. The enzymatic digestion conditions are shown in Table 6.

[0157] Table 6 Standard protein endonuclease combinations and enzyme digestion conditions

[0158]

[0159] After the protein endonuclease was completed, the sample was desalted and dried, and the sample was divided into two groups. In one group, the treated sample was redissolved in 10 μL of 100 mM ammonium acetate solution, pH 5, and 1 μg of carboxypeptidase Y (Sigma) was added. The enzyme digestion was carried out at 25°C for 12 h, and the sample was desalted and dried. The disulfide bond identification was performed according to the following mass spectrometry acquisition parameters and CADI-Mascot analysis method. The other group of samples was directly subjected to mass spectrometry acquisition and disulfide bond identification was performed using pLink software.

[0160] The following are the results of the disulfide bond identification analysis of the standard protein sample:

[0161] The present invention selected Albumin, Lysozyme, Transferrin, and RNase A as standard proteins, and used pLink and CADI-Mascot to identify their disulfide bond sites. This experiment also used a combination of multiple protein endonucleases, including Trypsin, Lys-N, and Glu-C. Using CADI, the present invention identified a total of 14 known disulfide bond sites ( Figure 18 ), in addition, the present invention also identified 9, 3, 63, and 30 previously unreported disulfide bonds on Albumin, Lysozyme, Transferrin, and RNaseA ( Figure 19 ).

[0162] 3. Verification of disulfide bond sites in RNase A by in vitro peptide synthesis

[0163] Three peptides on the artificially synthesized RNase A protein that conform to the Lys-N, Glu-C, and Trypsin cleavage sites, with specific sequences shown in Table 7, cross-linked to form three disulfide bond sites, including one known site, RNase A C2-C7 (66-12), identified by CADI-Mascot, and two unknown sites, RNase A C1-C7 (52-12) and RNase A C1-C2 (52-66). The three peptides were cross-linked in 20% DMSO for 17 hours (cross-linking method as in Example 1). After cross-linking, carboxypeptidase Y was added (1:10, w / w, 6 hours) and then mass spectrometered. The secondary spectrum of the site was established after searching the CADI-Mascot library.

[0164] Table 7 Peptide sequences synthesized in vitro for verification of identified RNase A disulfide bonds

[0165]

[0166] The following are the verification results of the disulfide bond sites in RNase A:

[0167] The present invention synthesized three peptide segments, which can form three sites by cross-linking. The secondary spectra generated after cross-linking are similar to the spectra obtained in the protein mixture, indicating that all three sites exist ( Figure 20 ).

[0168] Example 6 Identification of disulfide bonds in HeLa cells using the CADI method

[0169] 1. Stepwise alkylation labeling of reducible cysteine

[0170] 50 μg of the protein sample to be labeled was dissolved in 8 M urea containing 2 mM NEM and pH 6.5, incubated at 37°C for 2 hours to block free cysteine, and the solution was transferred to a 10 kD ultrafiltration tube and centrifuged at 14,000 g for 30 minutes. The filtrate was removed and repeated three times. Unreacted NEM was filtered out, and then 5 mM TCEP was added for reduction for 30 minutes. 200 μL of 8 M urea (pH 8.5) was added and centrifuged at 14,000 g for 30 minutes. After removing unreacted TCEP, IAA was added to a final concentration of 10 mM. After incubation at 25°C in the dark for 30 minutes, the solution was entered into the proteolysis part. The solution was diluted to 2 M urea and trypsin (1:100, w / w) was added. After incubation at 37°C for 16 hours, 100 μL of 2 M urea was added and the solution was centrifuged at 14,000 g for 30 minutes. Repeated three times, all filtrates were collected. After desalting and drying, it can be stored at -80℃.

[0171] Data were acquired using a Q-Exactive HF mass spectrometer. Label-free quantification was performed using MaxQuant software (version 1.6.01). The human database was downloaded from SwissProt in July 2019. Trypsin was used as the protease, with two missed cleavage sites. Variable modifications included oxidation on methionine (+15.994915 Da, M), amino-terminal acetylation of the protein (+42.010564), carbamidoformylation on cysteine ​​(+57.021464 Da, C), and N-ethylmaleimidation on cysteine ​​(+125.047679, C). The match-between run was 2 min. The mass accuracy of the primary mass spectrometer was set to 10 ppm, and the secondary mass accuracy was set to 0.05 Da. The FDR was set to 0.01 at both the PSM and protein levels.

[0172] The following are the results of step-by-step alkylation labeling of reducible cysteine:

[0173] The present invention used a stepwise alkylation method to label a total of 2,488 reducible cysteines located on 1,142 proteins. Based on these results, the present invention constructed a reduced CADI database of human disulfide bond sites. The disulfide bonds in this database contain at least one reducible cysteine ​​(the FF database contains 32,724 sequences) and retain 0-2 C-terminal amino acids.

[0174] 2. CADI method to identify disulfide bond sites in HeLa cells

[0175] An appropriate number of HeLa cells were seeded in a 10 cm dish. Once confluent, the medium was aspirated and 3 mL of PBS was added twice, carefully washing away any remaining medium. 5 mL of PBS was added, and equal volumes of DMSO and 1 M tetramethyl azodicarbonamide (diamide) were added to the control and experimental groups, respectively, to a final diamide concentration of 1 mM. The cells were incubated at 37°C for 15 min, and the solution was carefully aspirated. The cells were washed three times with 5 mL of PBS on ice. Then, 2 mL of 20% trichloroacetic acid (TCA) was added and the cells were incubated at 4°C for 20 min. The cells were harvested by scraping into a 2 mL centrifuge tube and centrifuged at 20,000 g for 30 min at 4°C. The supernatant was removed, and the pellet was washed with 1 mL of 10% TCA solution and 1 mL of 5% TCA solution, respectively. Finally, the pellet was washed twice with pre-chilled acetone. After aspiration of the acetone, the remaining acetone was evaporated and the sample was re-dissolved in 8 M urea solution (pH 6.5) containing 2 mM NEM. Protein concentration was determined by the BCA method. Six mg of HeLa cell lysate was diluted to 2 M urea with 100 mM Tris-HCl (pH 6.5). Trypsin (1:100, w / w) was added and digested at 37°C for 16 h. The sample was then diluted to 1 M urea with 100 mM Tris-HCl (pH 6.5). Glu-C (1:100, w / w) was added and digested at 25°C for 10 h. The sample was then desalted and dried by spin drying. Peptides were separated using Strong Cation Exchange (SCX) separation.

[0176] The desalted peptides were dissolved in 10% formic acid. Peptide separation was performed on an Agilent 1260 HPLC instrument using a strong cation exchange column (Luna, 250 × 4.6 mm). Mobile phase A: 0.05% formic acid in 20% acetonitrile; mobile phase B: 0.05% formic acid, 0.5 M NaCl in 20% acetonitrile. The fractions were collected over a 60-minute period, with a 2-minute collection interval for a total of 30 fractions. The liquid phase gradient is shown in the table. The last 10 highly charged fractions were selected, and two adjacent fractions were combined to yield five fractions. After desalting and drying, the fractions can be stored at -80°C.

[0177] Table 8 Mobile phase gradient for strong cation exchange chromatography for peptide pre-separation

[0178]

[0179]

[0180] The sample after SCX separation was redissolved in 100 mM ammonium acetate solution, pH 5, and 30 μg of carboxypeptidase Y (Sigma) was added. The sample was digested at 25°C for 12 h, and then the salt was removed and dried by spin drying.

[0181] 2 μg of the treated samples were respectively taken and data were collected on a Thermo QE-HF mass spectrometer using the same method as in Example 5. The database used was the reduced human disulfide bond database constructed above, and the Mascot parameters and analysis method were the same as in Example 6.

[0182] The following are the results of the CADI method for identifying disulfide bond sites in HeLa cells:

[0183] The present invention used the CADI-Mascot method to identify disulfide bonds in HeLa cells. To improve disulfide bond identification efficiency, the present invention utilized the previously constructed reduced human disulfide bond database. Disulfide bonds in this database contain at least one reducible cysteine ​​(the FF database contains 32,724 sequences). Using strong cation exchange chromatography, fractions containing highly charged peptides were selected for CADI enrichment and identification. A total of 102 disulfide bonds were identified across three replicates (Figure 21).

[0184] Overview of the CADI Method

[0185] The present invention has developed a disulfide bond identification method (CADI) dependent on carboxypeptidase Y. Carboxypeptidase Y is a peptide chain exo-tide enzyme that can degrade the carboxyl end of the peptide chain one by one to release free amino acids. First, the present invention uses protein endonucleases such as Lys-C, trypsin, and Glu-C to hydrolyze non-reduced proteases into polypeptides, and strives to make each cross-linked polypeptide contain only one disulfide bond site. At this time, most of the sample is linear polypeptides, and the proportion of disulfide-crosslinked polypeptides is very low. Then carboxypeptidase Y is added. When carboxypeptidase Y is hydrolyzed to the vicinity of the disulfide bond, the steric hindrance formed by the disulfide cross-linking makes it impossible for the enzyme to effectively bind to the peptide segment, thereby preventing further enzymatic cleavage. Since linear polypeptides do not have this steric hindrance, they will be completely degraded within a certain period of time, thereby achieving the purpose of enriching disulfide bonds. After the enzymatic cleavage of carboxypeptidase Y, the amino acids after the carboxyl end of the cross-linking site of the two peptide segments in the cross-linked peptide are basically enzymatically degraded. At this time, if one of the peptide segments of the cross-linked peptide is reversed, it can be regarded as a quasi-linear peptide segment. After the database is reconstructed, the disulfide bond can be identified using the traditional linear peptide search software (see Figure 1 ).

[0186] Fragmentation pattern of linearized cross-linked peptides

[0187] In the HCD / CID fragmentation mode, both peptides of a disulfide cross-linked peptide can be fragmented, resulting in two groups of b and y ions. However, their spectra are more complex than those of the fragment ions produced by the common fragmentation of two linear peptides. This is because after the parent ion is fragmented, the fragment ions are connected together by cross-linking bonds, which changes the m / z of the corresponding fragment ions. Figure 2After enzymatic cleavage by carboxypeptidase Y, a pair of disulfide-crosslinked peptides can undergo fragmentation on both peptide α and peptide β, with the y ion containing the mass of the other peptide. Peptides linearized using the method of the present invention can also generate these ions. Simply reconstructing the database based on the fragmentation characteristics of the cross-linked peptides and matching the fragment ion masses of the reconstructed peptides with the actual mass allows identification of the disulfide-crosslinked peptides. Successful identification of cross-linked peptides requires that both the MS1 and MS2 spectra generated by the theoretical database match the actual spectra.

[0188] First, keep the parent ion masses of the two peptides consistent, even if the MS1 matches. After the two peptide sequences are spliced ​​into one peptide sequence, the molecular weight of one oxygen atom (O) will be reduced. In order to ensure that the parent ion mass remains unchanged after the cross-linked peptide is linearized, the linearized peptide needs to be modified. Add an OH to each of the two cysteines that form the disulfide bond, and then subtract the mass of a water molecule (H2O) from the carboxyl end of the linearized peptide. In this way, the parent ion masses of the two peptides are consistent, and the cysteines that form the disulfide bond and other cysteines can be distinguished ( Figure 2 ).

[0189] Secondly, the fragment ions of the two peptides must be consistent, even if the MS2 matches. To accurately identify the peptide, in addition to maintaining the consistency of the parent ion mass of the linearized peptide and the actual peptide, it is also necessary to ensure that the fragment ions of the two peptides match. When cleavage occurs in the front section of the cross-linked peptide (peptide α), the fragmentation pattern of the cross-linked peptide and the ordinary linear peptide of the same sequence is exactly the same. In this case, the cross-linked peptide can be correctly interpreted by Mascot without linearization. After the linearization modification of the present invention, the fragment ion types and masses of the two peptides can also be completely matched. Figure 3 -b in the cross-linked polypeptide in A α1 and y α5 The ions correspond to b1' and y1' of the linearized peptide in the database. When fragmentation occurs in the latter peptide, another set of ions (such as b β1 and y β5 ). This group of fragment ions has a different mass from the ion produced by the unmodified linearized peptide sequence, while the y1' produced by the modified linearized peptide is different from the b β1 Consistent quality, b 11 'with y β5 The quality was consistent (Figure 3-B).

[0190] References

[0191] Walzthoeni,T.et al.False discovery rate estimation for cross-linkedpeptides identified by mass spectrometry.Nature methods 9,901-903(2012).

[0192] Yang,B.et al.Identification of cross-linked peptides from complexsamples.Nature methods 9, 904-906(2012).

[0193] Klemm,P.in Proteins.(ed.J.M.Walker)255-259(Humana Press,Totowa,NJ;1984)。

Claims

1. A method for analyzing and identifying cross-linked polypeptides, characterized in that: The analysis and identification method comprises the following steps: 1) using LC-MS / MS to collect data on the product; the method for obtaining the product comprises: ① mixing an unreduced protein sample to be identified with a protein endoenzyme, and incubating the mixture to obtain a cross-linked polypeptide; ② mixing an exonuclease with the cross-linked polypeptide to be identified, and incubating the mixture; the exonuclease is carboxypeptidase Y; 2) Establishing a linearized database: The linearized database includes a positive sequence database and a decoy database; wherein: The forward sequence database includes linearized sequences, which are formed by reversing the sequence of one of the two cross-linked polypeptides after enzyme digestion and then splicing the carboxyl ends of the two peptides together; The bait database includes a forward-reverse database, a reverse-forward database and a reverse-reverse database, wherein the forward-reverse database is a linearized sequence obtained by reversing the sequence of the second peptide segment of the cross-linked polypeptide in the forward database; The reverse-forward database is a linearized sequence obtained by reversing the first peptide sequence of the cross-linked polypeptide in the forward database; The reverse-reverse database is a linearized sequence obtained by reversing the sequences of the two peptide segments of the cross-linked polypeptide in the forward sequence database; 3) Analyze the cross-linked peptide spectra using conventional proteomics search software and linearized databases; The cross-linked polypeptide is a disulfide bond cross-linked polypeptide.

2. The analysis and identification method according to claim 1, wherein The protein endonuclease described in step ① is selected according to the sequence of the cross-linked protein; it includes one or more of trypsin, chymotrypsin, Lys-C protease, Glu-C protease and Lys-N protease.

3. The analysis and identification method according to claim 2, wherein: The protein endonuclease is trypsin and / or Glu-C protease; When the endoproteinase is trypsin and Glu-C protease: The relative amount of the protein endonuclease and the cross-linked protein in step ① is 1:50 to 1:100 (w / w); And / or, the enzymatic digestion system in step ① further comprises 1M to 2M urea; And / or, the incubation condition described in step ① is 10 to 16 hours.

4. The analysis and identification method according to claim 1, wherein The mixing and incubation time in step ② is 4 to 16 hours; And / or, the relative amount of the exonuclease in step ② to the cross-linked polypeptide to be identified is 1:10 to 1:100 (w / w).

5. The analysis and identification method according to claim 4, wherein: The mixed incubation time in step ② is 12 hours; the relative amount of the exonuclease in step ② to the cross-linked polypeptide to be identified is 1:50 (w / w).

6. The analysis and identification method according to any one of claims 1 to 5, characterized in that: The step 2) of establishing a linearization database includes: a) If it is a protein sequence, first perform simulated die-cutting according to the protein endonuclease used to obtain the polypeptide sequence of the corresponding protein. If it is a polypeptide sequence, proceed directly to step b); b) screening the peptides containing cross-linking sites in the above polypeptides, and combining them in pairs in order to obtain cross-linked peptide sequences; c) performing simulated enzymatic cleavage on the polypeptide sequences in the cross-linked polypeptide according to the enzymatic cleavage characteristics of carboxypeptidase Y, leaving 0-2 amino acids after the cross-linked amino acid site, i.e., each cross-linked polypeptide has 9 sequences containing different numbers of C-end AA; and replacing the amino acid abbreviation of the cross-linked site with the letter O, which is not one of the 20 common amino acid abbreviations; d) Constructing a forward sequence database: The two cross-linked polypeptides obtained in step c) are linearized by reversing the sequence of one of the peptide segments and then splicing the carboxyl ends of the two peptide segments together to form a linearized sequence; in software that can edit new amino acids, a virtual amino acid J is added, and sequences containing different numbers of C-end AA are separated by the letter J to obtain a forward sequence database. In software that cannot edit amino acids, sequences containing different numbers of C-end AA are separated by the letter U to obtain a forward sequence database; e) Build a bait database.

7. The analysis and identification method according to claim 6, wherein: In step c), when the disulfide bond site is identified, the cross-linked cysteine ​​site is changed from the letter C to the letter O, and the cysteine ​​at other positions is still represented by the letter C; In step d), the software for editing new amino acids is Mascot; the software for non-editable amino acids is MaxQuant.

8. The analysis and identification method according to claim 6, wherein: The data acquisition in step 1) was performed using a Thermo Q-Exactive HF mass spectrometer in DDA mode.

9. The analysis and identification method according to claim 8, wherein: The analytical column was prepared with 1.9 μm C18 filler and the column length was 150 mm; and / or, mobile phase A is 98% H2O, 2% ACN, 0.1% FA; mobile phase B is 98% ACN, 2% H2O, 0.1% FA; and the liquid phase gradient is 60 min, with the concentration of mobile phase B increasing from 4% to 30% in 53 minutes; And / or, the scanning range of the primary mass spectrometer is: 350-1500 m / z, the resolution is 60000, and the AGC is 3e 6 ; And / or, the secondary mass spectrometry fragmentation mode is HCD, the fragmentation energy is 27, the resolution is 15000, the maximum ion injection time is 150ms, and the AGC is 2e 5 .

10. The analysis and identification method according to any one of claims 1 to 5, characterized in that: Step 3) The analysis software used is Mascot, MaxQuant or Proteome Discovery.

11. The analysis and identification method according to any one of claims 1 to 5, characterized in that: The analysis method of step 3) uses the existing scoring and FDR threshold of the software to filter the results.

12. The analysis and identification method according to claim 11, wherein: In the Mascot results, after screening according to pep_expect≤0.05, pep_isbold=1, pep_rank=1, the FR and RF sites were re-analyzed. The length of the reverse peptide segments of the FR and RF parts, i.e., the number of amino acids (length) and the number of matching b and y ions (ions), was calculated. According to the rule of length≤3 or the ratio of ions to length≤0.32, FR or RF was converted to FF. Then, all the results were sorted from small to large by PEP, and the FDR was calculated. The results were screened according to the FDR threshold of 0.

05. Alternatively, peptides with a PEP less than 0.05 were screened out from the MaxQuant software results, and the length of the reverse peptides in the FR and RF parts, i.e., the number of amino acids (length) and the number of matching b and y ions (ions), were calculated. FR or RF was converted to FF according to the rule that length ≤ 3 or the ratio of ions to length ≤ 0.32; then all results were sorted in ascending order of PEP, and the FDR was calculated using the same formula, with an FDR threshold of 0.05.