Direct-read precision de novo sequencing method for proteins and proteome and based on product ion type identification and full product-ion coverage
Through K/R NeuCode double labeling and mirroring enzyme cutting technology, the problems of low coverage and unknown type in protein de novo sequencing are solved, full coverage and type recognition of ions are achieved, and amino acid sequences are directly read, which improves the accuracy and coverage of sequencing, and is suitable for accurate sequencing of various proteomes.
Patent Information
- Application Number
- PCT/CN2024/142075
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-22
- Filing Date
- 2024-12-25
- Publication Date
- 2025-07-31
AI Technical Summary
In the existing protein de novo sequencing technology, the coverage of ions in the secondary spectrum is low and the type is unknown, resulting in low accuracy of amino acid sequence resolution, making it difficult to achieve accurate sequencing of proteomic research.
The method of K/R NeuCode double labeling combined with paired mirror enzyme cleavage was used to label proteins with mass loss lysine and arginine, and the protein was cut by Trypsin/LysargiNase or LysC/LysN enzyme to generate mirror peptides. Combined with high-resolution mass spectrometry analysis, the sub-ion type can be identified and full coverage, and the amino acid sequence can be read directly.
It achieves accurate identification and full coverage of ion types, improves the accuracy and coverage of protein de novo sequencing, and can directly read amino acid sequences, which are suitable for accurate sequencing of proteomes or complex proteomes without sequence reference.
Smart Images

Figure CN2024142075_31072025_PF_FP_ABST
Abstract
Description
A direct-reading method for accurate de novo protein and proteome sequencing based on identifiable product ion types and full product ion coverage Technical Field
[0001] The present invention relates to the technical principles and methods of protein de novo sequencing. Specifically, it is a method for producing NeuCode dual-labeled proteins with heavy stable isotope labels lysine (K) / arginine (R) through paired mirror enzyme digestion such as trypsin / lysine arginine N-terminal protease, thereby generating paired mirror peptides with quasi-isoheavy amino acid labels. The method can be used to achieve direct-reading protein accurate de novo sequencing with full coverage of identifiable sequences by daughter ion type. Background Art
[0002] As the executor of life activities and functions, proteins are directly involved in regulating various biological processes. The structure of proteins determines their functional diversity, and accurately analyzing the amino acid sequence information of proteins, that is, protein sequencing, is the prerequisite for exploring the higher-level structure and function of proteins. Early protein sequencing methods were mainly based on chemical degradation, such as the Edman degradation method (EDMAN PA method for the determination of amino acid sequence in peptides. Arch Biochem. 1949 Jul; 22 (3): 475. PMID: 18134557), which uses chemical reagents to hydrolyze the N-terminal amino acids of a single protein in sequence, and then separates the amino acids by high-performance liquid chromatography for identification. This method can analyze the sequence of nearly 20 amino acid residues at the N-terminus of proteins and peptides, but it is difficult to achieve full sequence determination of large molecular weight proteins. With the development of high-throughput and high-sensitivity mass spectrometry technology, mass spectrometry has become the preferred method for protein sequencing. Mass spectrometry-based protein sequencing encompasses two strategies. Top-down protein sequencing directly analyzes intact proteins using techniques such as electrospray ionization (ESI) or matrix-assisted laser desorption ionization (MALDI). This analytical approach can provide a comprehensive understanding of protein molecular weight and structure. However, for proteins with larger molecular weights, the fragment ions generated by fragmentation are often difficult to distinguish due to their similar masses, which reduces the accuracy of the analysis and hinders the determination of the protein's amino acid sequence. Therefore, a bottom-up analysis strategy, also known as shotgun protein sequencing, is more commonly used for amino acid sequence analysis. In this approach, a protein or proteome is digested to form a peptide mixture. This peptide mixture is then chromatographically separated and ionized, and then detected by tandem mass spectrometry to generate peptide fragmentation patterns for peptide sequence identification. Finally, the identified peptides are used to deduce the potential proteins. Analyzing the large amount of peptide fragmentation data generated offers various options. For proteins or species with existing protein sequences or database references, database searches can be used for identification. Commonly used database search software include MASCOT, MaxQuant, and pFind.
[0003] However, for species for which no protein databases exist, or for proteins whose amino acid sequences have changed due to gene mutations, database searches are not feasible. In these cases, de novo protein sequencing is necessary. A notable feature of de novo protein sequencing is its independence from databases. Therefore, de novo sequencing plays a vital role in identifying unknown proteins, such as antibodies and neoantigens, discovering new proteins and genetic events, and correcting misannotated proteins in databases. The development of new technologies for de novo protein sequencing holds significant scientific and practical value.
[0004] The principle of de novo protein sequencing based on mass spectrometry is to use the mass-to-charge ratio and intensity of fragment ions generated by peptide fragmentation to calculate the mass difference between peaks in the secondary spectrum, infer amino acid information and post-translational modifications, and thus obtain peptide sequence information. Finally, the sequenced fragments are spliced and assembled to obtain the complete protein sequence. Although a large number of algorithms and software have been developed for de novo protein sequencing, including the PEAKS series, the pNovo series, Pep Novo, and Novor, which have effectively improved the efficiency of de novo sequencing, their accuracy still needs to be improved. Thilo Muth et al. evaluated the accuracy of three commonly used sequencing software, Novor, PEAKS, and PepNovor, and found that only approximately 40% of the de novo sequencing results were consistent with the database search results, indicating that there is still significant room for improvement in the accuracy of de novo sequencing algorithms. Further analysis of the spectral data revealed that the main reasons for the low accuracy of de novo sequencing are the presence of a large amount of noise peak interference and the low coverage of peptide fragment ions in the tandem mass spectra. Especially for the latter, when the fragment ion coverage drops from 100% to 50%, the loss of fragment ions will cause the order of continuous amino acids to change, and the proportion of correctly sequenced peptides will drop directly from 80% to 20% (Muth, T., & Renard, BY (2018). Evaluating de novo sequencing in proteomics: already an accurate alternative to database-driven peptide identification?. Briefings in bioinformatics, 19(5), 954-970.; Yang, H., Chi, H., Zeng, WF, Zhou, WJ, & He, SM (2019). pNovo 3: precise de novo peptide sequencing using a learning-to-rank framework. Bioinformatics (Oxford, England), 35(14), i183-i190). The presence of noise and interfering ions in actual spectra can make it impossible to accurately identify the characteristic fragment ions of peptides in the spectra. Furthermore, incomplete fragmentation of the peptide precursor ions results in the loss of some fragment ions, which reduces the accuracy of spectral analysis. Consequently, when the accuracy of peptide de novo sequencing is low, the correctness of subsequent sequence assembly is also difficult to guarantee. Therefore, determining the depth of daughter ion coverage and daughter ion type in secondary spectra has become the key to de novo sequencing technology, but it is also the main challenge facing the current development of de novo sequencing technology.To this end, researchers have also proposed targeted solutions from multiple aspects.
[0005] In order to determine the type of fragment ions in the secondary spectrum, the researchers used heavy stable isotope labeling technology to chemically modify and label the N-terminus and C-terminus of the peptide at the same time, and used the double peak signal generated by the fragment ions of two identical sequence peptides with a slight mass difference that were separately labeled on the secondary spectrum to help identify the type of fragment ions and facilitate the identification of the peptide ( A.,Sedo,O., J., & Zdráhal, Z. (2014). A simplified method for peptide de novo sequencing using (18)O labeling. European journal of mass spectrometry (Chichester, England), 20 (3), 255-260.; Zhang S, Shan Y, Zhang S, Sui Z, Zhang L, Liang Z, Zhang Y. NIPTL-Novo: Non-isobaric peptide termini labeling assisted peptide de novo sequencing. J Proteomics. 2017 Feb 10; 154: 40-48). However, these strategies still have shortcomings. For example, after the same peptide segment is labeled, it may change its physicochemical properties, resulting in a change in the chromatographic retention time and unable to be eluted simultaneously. Even if it can be eluted simultaneously, in order to allow the parent ions to be fragmented simultaneously, it is necessary to increase the secondary fragmentation window, which will increase the number of co-fragmented ions and increase the complexity of the spectrum. However, the mass-defect lysine coding (K NeuCode) labeling strategy proposed by Richards AL et al. can circumvent these shortcomings. It uses two quasi-isotopic lysines: L-lysine- 13 C6 15 N2(K602) and L-lysine- 2H8 (K080) respectively labeled the same protein, and after LysC digestion, produced quasi-isobaric peptides with a C-terminal K602 or K080 mass loss difference of only 0.036Da. During liquid chromatography-tandem mass spectrometry analysis, the quasi-isobaric peptides can be eluted simultaneously and simultaneously selected and fragmented by HCD without changing the separation window. In high-resolution secondary mass spectrometry, the generated y ion peak is a NeuCode doublet with a mass difference of 36mDa. Based on the NeuCode doublet signal, the fragment ion is determined to be a y-type product ion, thereby improving the product ion identification in the secondary spectrum (Richards AL, Vincent CE, Guthals A, Rose CM, Westphall MS, Bandeira N, Coon JJ. Neutron-encoded signatures enable product ion annotation from tandem mass spectra. Mol Cell Proteomics. 2013 Dec; 12(12): 3812-23). This method can identify y ions in the spectrum, but it still has some shortcomings when used for de novo sequencing, including: (1) the peptide fragments produced by LysC digestion are long, which is not conducive to mass spectrometry identification, and the identification of NeuCode double peaks in the high mass-to-charge ratio region requires a higher secondary mass spectrometry resolution; (2) there are cases where spectral peaks are missing, and it is still difficult to obtain complete y ion information.
[0006] To obtain complete secondary fragment ion information, the Xu Ping team from the Beijing Proteome Research Center and the He Simin team from the Institute of Computing Technology, Chinese Academy of Sciences, used trypsin and its mirror protease, LysargiNase, to digest proteins to produce "mirror" peptides. The fragmentation of mirror peptides ending in K / R residues gave strong y-series daughter ion signals, while the fragmentation of mirror peptides starting with K / R residues gave strong b-series daughter ion signals. Therefore, in the mirror spectra corresponding to the mirror peptides, the b and y daughter ion peaks complement each other to achieve full coverage. A precise de novo sequencing method with full daughter ion coverage was established, and high-throughput sequence analysis was achieved using the pNovoM algorithm (Yang H, Li YC, Zhao MZ, Wu FL, Wang X, Xiao WD, Wang YH, Zhang JL, Wang FQ, Xu F, Zeng WF, Overall CM, He SM, Chi H, Xu P. Precision De Novo Peptide Sequencing Using Mirror Proteases of Ac-LysargiNase and Trypsin for Large-scale Proteomics. Mol Cell Proteomics. 2019Apr;18(4):773-785). However, this study did not solve the problem of product ion type identification, and paired mirror spectra accounted for only 48%, indicating low spectrum utilization.
[0007] In summary, there is no systematic and effective solution to the problem of low product ion coverage and unknown types in secondary spectra during de novo protein sequencing, and there is no direct de novo sequencing technology for reading protein amino acid sequences. It is urgent to develop a new direct-reading de novo protein sequencing method that can simultaneously take into account the identification of ion types and full coverage of product ions, so as to achieve a breakthrough in the underlying technology of proteomics research. Summary of the Invention
[0008] The key to de novo protein sequencing lies in the analysis of the secondary spectrum, but in the actual spectrum, incomplete ion fragmentation, noise peak interference and other reasons make the daughter ions missing and the daughter ion types difficult to distinguish, which directly leads to the difficulty of secondary spectrum analysis, low sequence reading accuracy, and difficulty in de novo sequencing. In order to solve the technical problems of missing daughter ions and unknown daughter ion types in the secondary spectrum, the present invention proposes to use paired mirror proteases to digest proteins labeled with mass-deficient lysine / arginine (K / R NeuCode) to produce paired mirror peptides with terminal quasi-equal-weight amino acid labels, so that the peptides produce NeuCode double-peak signals in the secondary mass spectrometry spectrum (MS2), realizing the underlying technical principle of direct-reading protein accurate de novo sequencing with identifiable daughter ion types and full sequence coverage. This sequencing technology not only simplifies the secondary mass spectrometry spectrum, but also can correctly use the mass difference of adjacent daughter ion peaks in the spectrum to directly read the amino acid sequence.
[0009] The purpose of the present invention is to provide a new method of direct-reading protein accurate de novo sequencing technology based on full coverage of identifiable sequences of daughter ion types.
[0010] The technical solution adopted in the present invention is:
[0011] Proteins were metabolically labeled using K / R NeuCode reagent. The labeled proteins were quantitatively mixed and then, after reduction and alkylation, divided equally into two aliquots. One aliquot was digested with Trypsin / LysC to produce peptides labeled at the C-terminus with K / R NeuCode, while the other aliquot was digested with its mirror image protease, LysargiNase / LysN, to produce peptides labeled at the N-terminus with K / R NeuCode. The two peptide samples were separated and analyzed separately by LC-MS / MS. In HCD fragmentation mode, the y ion fragments produced by fragmentation of the peptides digested with Trypsin / LysC appeared as NeuCode doublets in the secondary spectrum, while the b ion fragments produced by fragmentation of the peptides digested with LysargiNase / LysN also appeared as NeuCode doublets in the secondary spectrum. The paired mirror images were combined to complement missing product ion peaks in the single spectrum. This resulted in deep coverage of product ions in the secondary spectrum, and their types were clearly identifiable. The protein amino acid sequence was then directly determined based on the mass differences between adjacent product ion peaks of the same type.
[0012] The direct-reading protein accurate de novo sequencing method based on K / R NeuCode dual labeling combined with mirror enzyme cleavage described in the present invention comprises the following steps:
[0013] 1) using different pairs of mass-deficient amino acids to completely label the proteins of the organism to be sequenced to generate mass-deficient amino acid coding signals; the different pairs of mass-deficient amino acids are lysine and arginine;
[0014] 2) digesting the labeled protein with a protease and its mirror image protease, respectively, to obtain a mirror image peptide with a mass-defective amino acid coding signal; the protease and its mirror image protease are a protease pair capable of producing N-terminal and C-terminal mirror images;
[0015] 3) analyzing the mirror peptide generated in step 2) to obtain a secondary spectrum of the mirror peptide;
[0016] 4) Directly read out the peptide sequence based on the peptide secondary spectrum obtained in step 3).
[0017] The method of the present invention is preferably wherein the step 1) comprises: labeling the mass-deficient amino acid pair with a heavy stable isotope to produce a labeled mass-deficient lysine pair L-lysine- 13 C6 15 N2(K602) and L-lysine- 2 H8(K080), mass loss of arginine to L-arginine- 15 N4(R004) and L-arginine- 2 H4(R040).
[0018] In the method of the present invention, the protease and its mirror protease in step 2) are preferably a Trypsin / LysargiNase or LysC / LysN protease pair.
[0019] The method of the present invention, wherein the step 3) includes using a secondary high-resolution mass spectrometer to analyze the mirror peptide segment to obtain a secondary spectrum of the mirror peptide segment; and the step 4) includes determining a paired mirror spectrum according to the mass characteristics of the mirror peptide segment parent ion, identifying the mass-deficient amino acid coding signal, i.e., the NeuCode double peak signal, in the spectrum according to the mass difference, judging the y ion or b ion in the mirror spectrum, merging the mirror spectra, filling the missing daughter ions in the single spectrum, obtaining a secondary spectrum with known daughter ion type and full coverage of daughter ions, judging the type of amino acid according to the mass difference of adjacent daughter ion peaks, and directly reading the polypeptide sequence from the secondary spectrum.
[0020] The resolution of the secondary mass spectrometry in step 3) is set to 6,0000 (m / z=200) or above.
[0021] In one embodiment of the present invention, the labeling method of step 1) comprises:
[0022] The labeled mass-deficient amino acid pairs are added to the culture medium of the organism for metabolic labeling culture. After culture, the cells are collected, the protein is extracted, and the protein sample with the mass-deficient amino acid code (NeuCode) is obtained after quantitative mixing; wherein the labeled mass-deficient amino acid pair is L-lysine-13 C6 15 N2(K602) and L-lysine- 2 H8(K080), L-arginine- 15 N4(R004) and L-arginine- 2 H4(R040).
[0023] Quasi-isotopic amino acids are added to the culture medium of the organism for metabolic labeling culture. After culture, the cells are collected, the protein is extracted, and the protein sample with a mass-defect amino acid code (NeuCode) is obtained after quantitative mixing; wherein the quasi-isotopic amino acid is L-lysine- 13 C6 15 N2(K602) and L-lysine- 2 H8(K080), L-arginine- 15 N4(R004) and L-arginine- 2 H4(R040).
[0024] In a specific embodiment, the concentration of quasi-isotope lysine added is 30 μg / mL, and the concentration of quasi-isotope arginine added is 20 μg / mL; the protein extraction solution is PBS non-denaturing lysis solution or denaturing lysis solution containing 8 M urea.
[0025] In one embodiment of the present invention, the digestion conditions in step 2) include:
[0026] The labeled protein in step 1) was fully denatured and divided into two parts. One part was digested with Trypsin / LysC, and the other part was digested with LysargiNase / LysN. After the digestion was completed, the digestion reaction was terminated, and the digested sample was desalted to obtain a peptide sample for mass spectrometry detection.
[0027] More preferably, 5 mM tris(2-carbonylethyl) phosphine hydrochloride (TCEP) and 15 mM iodoacetamide (IAM) are used for reduction and alkylation treatment at 25°C in the dark for 15 minutes to fully denature the protein; the treated protein solution is divided into two equal parts, and the urea concentration in the solution can be diluted to less than 1 M using 50 mM ammonium bicarbonate. One part is added with Trypsin / LysC at a substrate / enzyme ratio of 50:1 (mass ratio), and the pH is adjusted to between 7.5 and 8.5 using 1 M ammonium bicarbonate solution, and digested at 37°C for 10-16 hours; the other part is added with LysargiNase / LysN at a substrate / enzyme ratio of 50:1 (mass ratio), and the pH is adjusted to between 7.5 and 8.5 using 1 M HEPES buffer, and digested at 37°C for 10-16 hours; then 1% FA is added to each enzymatic hydrolysis system to terminate the digestion reaction, and the enzymatically digested sample is desalted using C18 StageTips to obtain a peptide sample that can be used for mass spectrometry detection.
[0028] The method of the present invention, wherein the chromatographic separation and mass spectrometry analysis in step 3) comprises: dissolving the obtained peptide samples in a liquid chromatography tandem mass spectrometry loading buffer containing 0.1% formic acid and 1% acetonitrile, separating them by reversed-phase high performance liquid chromatography, detecting them by tandem mass spectrometry, calculating the labeling efficiency, and setting the secondary mass spectrometry resolution to 60,000 (m / z = 200) or above to identify the mass-defective amino acid coding signal (NeuCode doublet signal) in the secondary spectrum;
[0029] Wherein step 4) comprises:
[0030] 4a) Mirror Spectrum Pair Matching and Product Ion Identification: Paired mirror spectra were obtained by matching the mirror peptide precursor ion mass, retention time, and other characteristics. In the Trypsin / LysC digestion spectrum, the mass-defective amino acid coding signal (NeuCode doublet signal) generated by the K / R NeuCode double labeling was used to identify the y-type product ion peak. In the LysargiNase / LysN digestion spectrum, the mass-defective amino acid coding signal generated by the K / R NeuCode double labeling was used to identify the b-type product ion peak.
[0031] 4b) Mirror spectrum merging and direct sequence reading: The missing product ions in the single spectrum are supplemented based on the mass characteristics of the mirror peptide product ion fragments, the mass difference between adjacent product ion peaks is calculated, the amino acid type is determined, and the peptide sequence is directly and accurately read;
[0032] 4c) Sequence assembly: Based on multiple pairs of mirror-image enzyme digests, such as Trypsin and LysargiNase, and LysC and LysN, combined with K / RNeuCode dual labeling and KNeuCode single labeling, the sequencing peptides are accurately read directly. Overlapping sequences are then spliced and reconstructed to directly and accurately read the protein sequence.
[0033] The direct-reading de novo protein sequencing method of the present invention can be applied to the precise sequencing of any proteome without a sequence reference or proteins of unknown sequence.
[0034] According to another aspect of the present invention, the novel direct-reading, precise de novo sequencing method proposed in the present invention is also applicable to de novo sequencing of chemically labeled complex proteome samples. Therefore, the present invention also includes a direct-reading, precise de novo protein sequencing method comprising the following steps:
[0035] i) The biological protein to be sequenced is digested with a combination of Trypsin and LysC proteases, and the protein is purified by ionizing the protein with H2O and oxygen 18 heavy water (H2 18 O) were used as digestion buffer, and 16 O or 18 O-heavy stable isotope labeled enzymatic digestion peptides 16 O-Pep and 18 O-Pep;
[0036] ii) using isotopes to fully label the peptides from step i) with deuterium (D) or hydrogen (H) labeled formaldehyde to produce peptides with chemically quasi-isobaric labels 16 O-(CD3)2-Pep and 18 O-(CH3)2-Pep, the mass difference between the two peptides is 16.6 mDa;
[0037] iii) analyzing the labeled peptides produced in step ii) to obtain a secondary spectrum of the peptides;
[0038] iv) Based on the peptide secondary spectrum obtained in step iii), the protein sequence is directly read out according to the type of amino acid-encoded product ions and the mass difference between similar product ions.
[0039] In a preferred embodiment, the method of the present invention comprises the following steps:
[0040] (1) The protein to be sequenced is reduced and alkylated; the protein is evenly divided into two parts, one of which is digested with a combination of Trypsin and LysC, using water (H2O) as the digestion buffer to obtain trypsin-digested peptides, which are recorded as 16 O-Pep; the other was digested with Trypsin and LysC. 18 O Heavy water (H2 18O) replaces water (H2O) to obtain 18 O-labeled trypsin-digested peptides are denoted as 18 O-Pep;
[0041] (2) The obtained 18 O-Pep was labeled with formaldehyde (CH2O) to obtain 18 O-(CH3)2-Pep, 16 O-Pep was labeled with deuterated formaldehyde (CD2O) to obtain 16 O-(CD3)2-Pep;
[0042] (3) 18 O-(CH3)2-Pep and 16 The two peptides O-(CD3)2-Pep were mixed in equal amounts and fractionated to obtain peptide fractions;
[0043] (4) Direct-read sequencing of quasi-isobaric chemically labeled protein sequences.
[0044] Definition of terms
[0045] In the present invention, the term "heavy stable isotope" refers to the same chemical element in the periodic table with the same atomic number but different neutron number, which has basically the same chemical properties, stable mass number and is 1 to 2 atomic weight units heavier than ordinary elements. Common heavy stable isotopes include carbon-13 ( 13 C), nitrogen-15 ( 15 N), hydrogen-2( 2 H) and oxygen-18 ( 18 O). Amino acids synthesized using heavy stable isotopes are heavy stable isotope amino acids. For example, compared to common L-lysine (chemical formula is 12 C6 1 H 14 14 N2 16 O2), heavy stable isotope lysine L-lysine- 13 C6 1 H 14 15 N2 16 6 of O2(K602) 12 C is replaced by 13 C, 2 14 N is replaced by 15 N, which increases the mass number by 8, and the heavy stable isotope lysine L-lysine- 12 C6 1 H6 2 H8 14 N2 16 In O2 (K080), there are 81 H was replaced by 2 H, the mass number increased by 8; compared with the common L-arginine (chemical formula 12 C6 1 H 14 13 N4 16 O2), heavy stable isotope lysine L-arginine- 12 C6 1 H 14 15 N4 16 4 in O2(R004) 14 N is replaced by 15 N, mass number increased by 4, heavy stable isotope arginine L-arginine- 12 C6 1 H 10 2 H4 13 N4 16 4 in O2(R040) 1 H was replaced by 2 H, the mass number increased by 4.
[0046] In the present invention, the term "quasi-isotopic amino acids" refers to amino acids of the same heavy stable isotope type that contain different heavy stable isotopes but have the same increased mass number, such as L-lysine- 2 H8(K080) and L-lysine- 13 C6 15 N2(K602), compared with ordinary L-lysine, has a mass number increased by 8, and they are a pair of quasi-isotopic lysine. 15 N4(R004) and L-arginine- 2 H4(R040) has a mass number increased by 4 compared to ordinary L-arginine, and they are a pair of quasi-isotopes of arginine.
[0047] In the present invention, the term "mass-defective amino acid" refers to the different mass defects produced by the same amino acid labeled with different heavy stable isotopes compared to ordinary elements, for example 12 C / 13 The mass loss of C is +3.3 mDa. 1 H / 2 The mass loss caused by H is +6.3 mDa, 14 N / 15 The mass loss produced by N is -3.0 mDa, which makes the mass loss produced by amino acids containing different heavy stable isotopes different.2 The mass loss of H8(K080) is +50.4mDa, L-lysine- 13 C6 15 The mass loss of N2(K602) is +13.8mDa, L-arginine- 2 The mass loss of H4(R040) is +25.2mDa, L-arginine- 15 The mass loss of N4 (R004) is -12.0 mDa.
[0048] In the present invention, the term "mass defect amino acid pair" refers to a pair of quasi-isotopic amino acids with different mass defects, such as L-lysine- 2 H8(K080) and L-lysine- 13 C6 15 N2(K602), L-arginine- 2 H4(R040) and L-arginine- 15 N4(R004).
[0049] In the present invention, the term "mass defect amino acid coding signal (NeuCode)" refers to the fact that after the same peptide is labeled with quasi-isotopic amino acids with different mass defects, the same peptide has a slight mass difference, which can be distinguished in the high-resolution mass spectrometry and presents a double peak signal with a certain mass difference in the spectrum, namely the NeuCode signal, for example, L-lysine- 2 H8(K080) and L-lysine- 13 C6 15 The same peptide segment labeled with N2 (K602) differs by 36 mDa, and a double peak signal with 36 mDa can be generated in the mass spectrum. 2 H4(R040) and L-arginine- 15 N4 (R004) differs by 37 mDa and can produce a double peak signal with a peak of 37 mDa in the mass spectrum.
[0050] In the present invention, the term "KNeuCode single labeling" refers to the use of only L-lysine- 2 H8(K080) and L-lysine- 13 C6 15 The N2 (K602) pair of mass-deficient amino acids is metabolically labeled to produce a 36 mDa mass-deficient amino acid encoding signal.
[0051] In the present invention, the term "RNeuCode labeling" refers to the use of L-arginine- 2H4(R040) and L-arginine- 15 The N4 (R004) pair of mass-deficient amino acids is metabolically labeled to produce a 37 mDa mass-deficient amino acid encoding signal.
[0052] In the present invention, the term "K / R NeuCode dual labeling" refers to the simultaneous use of L-lysine- 2 H8(K080) and L-lysine- 13 C6 15 N2(K602), L-arginine- 2 H4(R040) and L-arginine- 15 The two pairs of mass-deficient amino acids in N4 (R004) were metabolically labeled to produce mass-deficient amino acid coding signals of 36 mDa or 37 mDa.
[0053] In the present invention, the term "protease and its mirror protease" refers to a protease pair that can produce N-terminal and C-terminal mirror images, such as trypsin and its mirror enzyme lysargiNase, Lys C and its mirror enzyme Lys N, Arg C and its possible mirror enzyme Arg N, Glu C and its possible mirror enzyme Asp N, etc.
[0054] In the present invention, the term "mirror peptide" refers to a peptide produced by enzymatic cleavage of the same protein using a protease and its mirror protease, which has the same amino acid composition and sequence except for the possible differences in the N-terminal and C-terminal amino acid residues. For example, the trypsin-cleaved peptide ABCDEF.K / R and the lysargiNase-cleaved peptide R / K.ABCDEF are mirror images of each other.
[0055] In the present invention, the term "mirror image" refers to the secondary spectrum corresponding to the mirror image peptide segment.
[0056] In this application, the terms "parent ion" and "daughter ion" refer to the charged ions produced by peptide ionization in a mass spectrometer. The parent ion has a mass equal to the molecular weight of the peptide. Collision-induced dissociation of the parent ion generates ions with different mass-to-charge ratios, known as daughter ions, by breaking peptide bonds. The amino acid sequence and composition of the peptide can be determined based on the mass-to-charge ratios of the daughter ions.
[0057] In this application, the term "peptide secondary spectrum" refers to the spectrum corresponding to the daughter ion fragments produced after the parent ion is fragmented by secondary mass spectrometry. The mass distribution of the daughter ion fragments in the secondary mass spectrum can be used to infer the amino acid sequence of the peptide and perform protein sequencing.
[0058] The present invention has the following beneficial effects:
[0059] 1. The present invention proposes a technology for direct reading of protein amino acid sequences based on the fully identifiable product ion type of the secondary mass spectrometry spectrum of mirror peptides obtained by paired mirror enzyme digestion of mass-deficient amino acid pairs labeled proteins. This technology is based on the underlying principles of de novo protein sequencing, proposes new principles, and combines a series of new technologies. It is a fundamental technological innovation in proteomics research.
[0060] 2. This invention expands mass-defect amino acid labeling from single mass-defect lysine labeling (KNeuCode single labeling) to dual mass-defect lysine and mass-defect arginine labeling (K / R NeuCode dual labeling). In addition to the existing mass-defect lysine pair K602 / K080, a new mass-defect arginine pair R004 / R040, with a mass difference of 37 mDa, is introduced. This allows for efficient KNeuCode and RNeuCode labeling of lysine and arginine residues in proteins. The two mass-defect amino acid pairs have similar mass differences (Figure 5 and Table 1), facilitating unbiased identification of mass-defect amino acid encoding signals (NeuCode doublets) generated by KNeuCode and RNeuCode labeling at the same resolution in MS / MS.
[0061] 3. By achieving dual labeling of protein lysine and arginine residues, the present invention not only allows for conventional LysC protease digestion and the LysN protease digestion proposed in the present invention, but also ArgC protease digestion and ArgN protease digestion. Furthermore, it adds the possibility of trypsin digestion, the most specific and efficient enzyme in the field of proteomics to date, as well as trypsin mirror enzyme LysargiNase protease digestion. In practice, this allows for direct-reading, accurate de novo protein sequencing at the proteome level. BRIEF DESCRIPTION OF THE DRAWINGS
[0062] Figure 1 Flowchart of mass spectrometry-based protein sequence determination.
[0063] Figure 2 Shotgun protease-cleaved peptide sequencing and protein sequencing peptide sequence splicing.
[0064] Figure 3 Principle of peptide amino acid sequence inference based on mass spectrometry fragmentation secondary spectrum.
[0065] Figure 4 Schematic diagram of the novel direct-reading protein precise de novo sequencing based on mirror image spectra and mass-defect amino acid coding (NeuCode) labeling.
[0066] Figure 5: A novel direct-reading, precise de novo protein sequencing method based on lysine and arginine mass-deficient amino acid coding (K / R NeuCode dual labeling) and Trypsin / LysargiNase paired mirror-image enzymatic peptide fragments.
[0067] Figure 6 Preparation of fully GST-labeled high-purity proteins from the K602 / R004 and K080 / R040 combinations.
[0068] Figure 7 Mirror image enzyme pairs of K602 / R004 and K080 / R040 combinations that fully labeled GST and were digested with Trypsin / Lysarginase, respectively.
[0069] Figure 8 De novo sequencing of the GST protease-cleaved peptide R.GSPGIPGST.R.
[0070] Figure 9: Efficiency of direct-reading accurate de novo sequencing of GST encoding lysine / arginine mass-deficient amino acids (K / R NeuCode dual-labeled) digested with Trypsin / LysargiNase, respectively.
[0071] Figure 10: A novel direct-reading, accurate de novo protein sequencing method based on lysine mass-defect amino acid coding (K NeuCode single labeling) and LysC / LysN paired mirror-image enzyme-cleaved peptides.
[0072] FIG11 Preparation of fully GST-tagged high-purity proteins of K602 and K080.
[0073] Figure 12 Mirror image enzyme pairs of K602 and K080, which completely label GST digested with LysC and LysN, respectively.
[0074] Figure 13 In-line de novo sequencing of the GST protease-cleaved peptide K.YIAWPLQGWQATFGGDHPP.K.
[0075] Figure 14. In-line accurate de novo sequencing efficiency of K602 and K080 after fully GST-tagged and digested with LysC and LysN, respectively.
[0076] Figure 15 Single protein GST direct-reading accurate de novo sequencing sequence assembly.
[0077] Figure 16: A novel direct-reading protein precise de novo sequencing of Escherichia coli proteins based on lysine and arginine mass-deficient amino acid coding (K / R NeuCode dual labeling) and Trypsin+LysC / LysargiNase+LysN paired combination mirror enzyme-cleaved peptides.
[0078] Figure 17 Peptide K.PHVNVGTIGHVDHG.K of the direct-read sequencing POCE47 (gene: tufA; protein: EFTU1) protein in the complex proteome sample of Escherichia coli.
[0079] Figure 18 Efficiency of direct-reading accurate de novo sequencing of P0CE47 protein in complex Escherichia coli proteome samples.
[0080] Figure 19 shows the depth of direct-read sequencing coverage of protein POCE47 in Escherichia coli. DETAILED DESCRIPTION
[0081] The principles and techniques proposed in the present invention are described in detail below through examples, but the present invention is not limited in any form.
[0082] Example 1. Principle of direct-reading protein accurate de novo sequencing technology combining mirror image strategy and quasi-isotopic amino acid labeling strategy to achieve full coverage of product ion type identifiable sequence
[0083] Protein sequencing is the foundation and prerequisite for proteomics research. Currently, the most commonly used strategy is shotgun protein sequencing (Figure 1). First, the extracted proteome is digested with a protease to produce peptide fragments. The peptide mixture is then separated by high-performance liquid chromatography (HPLC). After electrospray ionization, peptide precursor ions are formed, which are then selected for fragmentation in a secondary mass spectrometer (MS / MS). The mass-to-charge ratios and intensities of the daughter ions are mapped to the MS / MS spectrum as signal peaks. Each peptide has a corresponding MS / MS spectrum, and the sequence information of the peptide can be determined by analyzing the MS / MS spectrum. Although in shotgun protein sequencing, individual peptides may not be detected by mass spectrometry after protein digestion due to being too short, too long, or having low ionization efficiency, when digested with multiple proteases with different cleavage site specificities or nonspecific proteases, overlapping amino acid residues may occur between the resulting peptide sequences. In this case, using a multi-enzyme digestion strategy, sequence splicing and reconstruction can be performed, ultimately determining the sequence of the complete protein (Figure 2).
[0084] Therefore, in de novo protein sequencing based on mass spectrometry, the complete identification and accurate interpretation of product ions in the secondary spectrum have become key steps in de novo sequencing. The theoretical basis for de novo sequencing based on secondary spectra is that the peptide parent ion is selected and regularly induced to fragment under different fragmentation modes, producing secondary spectra of different types of product ions. According to the different naming of the fragmentation sites, fragmentation near the N-terminus of the peptide mainly produces three series of product ions: a, b, and c. Fragmentation near the C-terminus mainly produces three series of product ions: x, y, and z. In collision-induced dissociation (CID) or high-energy collision-induced dissociation (HCD) modes, peptide bonds are mainly broken, producing b and y series of product ions. The mass difference between two consecutive b ions or y ions corresponds to the mass of one amino acid residue (Steen H, Mann M. The ABC's (and XYZ's) of peptide sequencing. Nat Rev Mol Cell Biol. 2004; 5(9): 699-711.). Therefore, as shown in Figure 3, by analyzing the secondary spectrum, identifying the complete b-ion peak or y-ion peak in the secondary spectrum, and calculating the mass difference of adjacent product ions of the same type, the amino acid type can be determined, thereby determining the complete sequence of the peptide.
[0085] However, the complete and accurate identification of product ions in secondary spectra remains the rate-limiting step in de novo proteome sequencing. In actual spectra, incomplete ion fragmentation and interference from noise peaks can lead to missing product ions and difficulty in determining their types, ultimately hindering de novo sequencing and resulting in low accuracy.
[0086] Faced with the difficulties and challenges of protein de novo sequencing, the present invention starts from the underlying principles of protein de novo sequencing and proposes a direct-reading type of protein precise de novo sequencing technology principle with identifiable ion types, full coverage of daughter ions, and complete and spliced sequences (Figure 4). Specifically, in order to address the difficulties of incomplete ion fragmentation and low daughter ion coverage, the present invention adopts a mirror enzyme cleavage strategy to produce mirror peptides, thereby generating a secondary spectrum of mirror peptides, merging the spectra of mirror peptides, filling the missing daughter ions in the single spectrum, and generating a new spectrum with full coverage of daughter ions. In order to identify the type of daughter ions in the spectrum, a mass-deficient amino acid pair is used to label the protein. After mirror enzyme digestion, the mass-deficient amino acids are located at the C-terminus and N-terminus of the peptide, respectively. The parent ions of the peptides labeled with the mass-deficient amino acid pair with the same amino acid sequence can simultaneously enter the secondary mass spectrometer for fragmentation to generate a secondary spectrum, in which the y ion peak generated by the C-terminal labeled peptide is a double peak with a slight mass difference, and the b ion peak generated by the N-terminal labeled peptide is a double peak with a slight mass difference. The molecular ion type is distinguished based on the double peak signal (NeuCode signal). In the spectrum with full coverage of the identifiable sequence of daughter ion types, the amino acid sequence is accurately and directly read out through the mass difference information of adjacent daughter ions, and then the direct-read sequence is spliced and reconstructed to ultimately achieve accurate protein sequencing.
[0087] The principle of direct-reading de novo sequencing is based on the underlying technologies of proteomics research. As long as the secondary spectrum satisfies both identifiable product ion types and full sequence coverage, the amino acid sequence can be directly read by subtracting the mass of two adjacent ions of the same type. This sequencing technology is applicable to both simple systems, such as single proteins, and complex systems, such as proteomes.
[0088] Example 2. De novo sequencing of a single glutathione S-transferase (GST) protein based on dual labeling with lysine and arginine NeuCode and paired mirror enzyme digestion with trypsin and lysargiNase
[0089] As shown in Figure 5, R004 and R040, and K602 and K080, are two quasi-isobaric amino acids of arginine and lysine, respectively. The difference in mass defect between R004 and R040 is 37 mDa, and the difference in mass defect between K602 and K080 is 36 mDa. When two quasi-isobaric amino acids are used to label a protein or proteome, peptides with identical amino acid sequences are obtained by trypsin / lysarginase digestion, which are quasi-isobaric peptides. When these quasi-isobaric peptides enter mass spectrometry, their parent ions are indistinguishable within the separation window and can be simultaneously selected for secondary mass spectrometry fragmentation. The quasi-isobaric labeled lysine and arginine residues in the trypsin-digested peptides are located at the C-terminus of the peptides, resulting in a doublet of the y ion peak at 37 mDa or 36 mDa, while the b ion peak is a singlet. Similarly, after LysargiNase digestion, the quasi-equal weight labeled lysine and arginine residues are located at the N-terminus of the peptide, so that the b ion peak produced during the secondary mass spectrometry fragmentation is a bimodal form containing 37mDa or 36mDa, while the y ion peak is a single peak form. This type of mass difference is caused by the loss of atomic mass within the quasi-equal weight amino acids and is referred to as a mass loss amino acid encoding signal (NeuCode signal). The NeuCode signal can be used to accurately identify the daughter ion type in the secondary spectrum. In addition, Trypsin and LysargiNase are a pair of mirror-image proteases that specifically digest the C-terminus and N-terminus of lysine and arginine residues. The secondary spectrum of the mirror-image peptide generated by the mirror-image digestion effectively supplements the daughter ion peak missing in the single spectrum. In this embodiment, Trypsin and LysargiNase mirror-image digestion is combined with K / R NeuCode double labeling to simultaneously achieve de novo sequencing of the GST protein sequence with fully identifiable neutron ion types, further demonstrating the scientific nature and effectiveness of the direct-reading accurate de novo sequencing technology principle. The specific steps are as follows:
[0090] 1. Construction of GST protein E. coli expression strain
[0091] The recombinant expression vector pGEX-4T-GST containing the GST protein encoding gene sequence was introduced into Escherichia coli BL21(DE3)ΔlysAΔargA by chemical transformation to obtain the recombinant strain BL21(DE3)ΔlysAΔargA-GST containing the GST protein expression gene. This strain can stably express GST protein under isopropylthiogalactoside (IPTG) induction. The intact molecular weight (theoretical) of the expressed protein is 27376.76 Da. Single protein GST can be obtained by affinity purification.
[0092] 2. E. coli Lysine and Arginine NeuCode Dual Labeling (K / R NeuCode Dual Labeling) and GST Protein Expression
[0093] R004 and K602, R040 and K080 were combined and used for metabolic labeling and culturing Escherichia coli BL21 (DE3) ΔlysAΔargA-GST, respectively. The SILACE medium used for metabolic labeling refers to the formula described in patent CN102796682B, wherein lysine and arginine were replaced with the corresponding heavy stable isotope labeled amino acids, respectively. As shown in Figure 5 (A), H1 is a SILACE medium supplemented with 30 μg / mL K602 and 20 μg / mL R004, and H2 is a SILACE medium supplemented with 30 μg / mL K080 and 20 μg / mL R040. First, a fresh BL21 (DE3) ΔlysAΔargA-GST colony was picked and inoculated into LB liquid medium containing 50 μg / mL ampicillin. The colony was cultured at 37°C for 10 h at a speed of 200 rpm, and the cells were collected by centrifugation at 2000g for 3 min, washed twice with sterile water, and then quantified according to OD 600 The initial density was 0.01, and the cells were inoculated into 5 mL of H1 and H2 metabolic labeling medium, respectively, and cultured at 37°C for 6 to 7 generations until the OD 600 When the value is about 1.5, again with OD 600 The initial density was 0.01 and transferred to the corresponding 20 mL H1 and H2 metabolic labeling medium, and cultured at 37 ° C for 6 generations until the OD 600 When the value was about 0.8, 1 mM IPTG was added, and the culture was induced at 18°C at 200 rpm for 12 h. The cells were centrifuged at 3,700 g for 7 min at 4°C, and the bacteria were collected to obtain BL21 (DE3) ΔlysA ΔargA-GST bacteria expressing K602 and R004, K080 and R040 tagged GST proteins, respectively.
[0094] 3. Protein purification and labeling efficiency detection
[0095] PBS was used as the protein extraction solution, and the above-mentioned bacteria were lysed using a non-contact ultrasonic disruptor (QSONICA, model Q800R3 Sonicator). The cells were centrifuged at 21,000 g for 10 min at 4°C, and the supernatant was collected to obtain the total protein. The protein was purified using GST 4FF agarose resin ( Biotech, catalog number C600031-0025) were purified to obtain K602 / R004-labeled GST protein and K080 / R040-labeled GST protein. The purified protein was subjected to 10% SDS-PAGE gel electrophoresis and silver staining, and protein quantification was performed using Coomassie Brilliant Blue staining. In-gel digestion was performed with 10ng / μL trypsin for 14h to obtain peptides for labeling efficiency detection. The results are shown in Figure 6. The purified protein band is single with a molecular weight of approximately 27kDa. The GST peptide "SPILGYWK" was randomly selected, and the comparison of the parent ion mass-to-charge ratio after non-labeled and K / R NeuCode double labeling is shown in Figure 6(B), indicating that the protein is completely labeled.
[0096] 4. Trypsin / LysargiNase mirror enzyme digestion of K / R NeuCode dual-tagged protein GST
[0097] 10 μg of each K602 / R004-labeled and K080 / R040-labeled GST protein was mixed and reduced and alkylated with a final concentration of 5 mM tris(2-carbonylethyl)phosphine hydrochloride (TCEP) and a final concentration of 15 mM iodoacetamide (IAM) at 25°C for 15 minutes in the dark. The reduced and alkylated protein sample was divided into two equal parts, one digested with Trypsin and the other with Lysarginase, both at an enzyme / substrate ratio of 1:50 (mass ratio). Digestion was performed at 37°C for 14 hours to obtain peptide fragments.
[0098] 5. Chromatography-mass spectrometry tandem analysis
[0099] The peptide fragments obtained above were subjected to C 18 The sample was desalted using StageTips and dissolved in 25 μL of GC-MS loading buffer containing 1% acetonitrile and 0.1% formic acid. 5 μL of the sample was analyzed by GC-MS. The GC-MS instrument model and parameters were as follows: High Performance Liquid Chromatography (Thermo Scientific Easy-nLC 1000) Column: Analytical column (C 18 , 3μm, The chromatographic gradient was set as 7%-14% phase B for 6 min, 14%-28% phase B for 32 min, 28%-42% phase B for 14 min, 42%-95% phase B for 2 min, and 95%-95% phase B for 6 min. The flow rate was 600 nL / min. The mass spectrometer was a Thermo Scientific Q Exactive HF with a spray voltage of 2.1 kV, a capillary temperature of 275°C, an S-lens of 60%, a collision energy of 27% HCD, a primary resolution of 120,000 @ m / z 200, and a secondary resolution of 240,000 @ m / z 200, the maximum ion injection time (MIT) was 50 ms for the primary and 200 ms for the secondary, the dynamic gain control (AGC) was set to 3e6 for the primary and 5e5 for the secondary, the minimum AGC was 5e3, the parent ion scan range was set to m / z 150–1400, the product ion scan range started from m / z 120, the data-dependent scan (DDA) mode was selected, TopN was set to 10, the isolation window was set to 1.6 m / z, the dynamic exclusion time was set to 8 s, and the number of secondary microscans was set to 3.
[0100] The LC-MS base peaks of GST digested with trypsin and lysargiNase are shown in Figures 7(A, B), respectively. A total of 15 specific fully enzymatically cleaved peptides were identified after trypsin digestion of the K / R NeuCode-labeled GST protein, and 14 specific fully enzymatically cleaved peptides were identified after lysargiNase digestion. Furthermore, the clustered parent ion peaks corresponding to the peptides in the chromatographic peaks are doublets, with a mass difference of approximately 18 mDa and a signal intensity close to 1:1. This indicates that the two amino acid-labeled proteins are nearly equally mixed, and that the enzymatic cleavage specificity is not affected by the amino acid labeling, resulting in quasi-isometric peptides after enzymatic digestion. Figure 7 (C, D) shows the mirror-image spectra of the peptides "SDLVPR" and "KSDLVP" generated by trypsin and lysargiNase digestion, respectively. In the secondary spectrum of the trypsin-digested peptide, the y ion is a NeuCode doublet with a mass difference of 37 mDa, while the b ion is a singlet. In the LysargiNase-digested spectrum, the y ion is a singlet, while the b ion is a NeuCode doublet with a mass difference of 36 mDa. This indicates that the b and y ions can be effectively distinguished by the NeuCode doublet signal.
[0101] 6. Direct reading sequencing of GST protease cleavage peptides
[0102] First, the product ion types in the spectra are determined based on double peaks with only a slight mass difference in the paired mirror spectra. The mirror spectra are then combined to complete missing product ion peaks in the single spectrum. Ion types are identified and product ions are fully covered. The amino acid residue masses are then calculated based on the mass differences between adjacent product ions, confirming the amino acid types and enabling accurate direct readout of the complete sequence. As shown in Figure 8, for the mirror spectrum of a peptide precursor ion with a mass-to-charge ratio (m / z) of 466.7467 and two positive charges, the y ion is identified based on the double peak signal in the Trypsin spectrum, and the b ion is identified based on the double peak signal in the LysargiNase spectrum. By merging the mirror spectra, the complete product ion information for this precursor ion is obtained, allowing the direct readout of the peptide sequence with the precursor ion "m / z = 466.7467, z = 2": R.GSPGIPGST.R. Similarly, the complete amino acid (aa) sequence of a peptide derived from the mirror digestion of the GST protein by Trypsin and LysargiNase can be efficiently and accurately determined based on the mirror spectra. As shown in Figure 9, the theoretical enzyme digestion of Trypsin and LysargiNase produces a total of 17 fully digested peptides, of which 13 peptide sequence information was directly read in this example, with a reading efficiency of 100%, an accuracy of 100%, and an overall sequence coverage of 53.78% (128aa / 238aa), accounting for 80.0% (128aa / 160aa) of the theoretical sequence coverage (reading efficiency = number of directly read amino acids per directly read peptide / theoretical number of amino acids in the peptide; the theoretical sequenceable sequence refers to the peptide sequence with a length of 6-30 amino acids after specific full digestion by Trypsin or LysargiNase).
[0103] Example 3. De novo sequencing of a single protein of glutathione S-transferase (GST) based on lysine NeuCode single tagging (KNeuCode single tagging) combined with paired mirror enzyme digestion of LysC and LysN
[0104] In this example, a LysC / LysN mirror enzyme pair was used to digest a single protein GST labeled with KNeuCode, and direct-read sequencing was performed on it. As shown in Figure 10, K602 and K080 are a pair of quasi-isotopic amino acids of lysine. The molecular weight difference between the two is about 36mDa due to the internal atomic mass loss of the quasi-isotopic amino acids. After using these quasi-isotopic amino acids to label proteins or proteomes, equal amounts of mixed proteins were digested by LysC / LysN. After LysC or LysN digestion, the peptides with the same amino acid sequence were generated as quasi-isotopic peptides. The quasi-isotopic peptides can be simultaneously selected to enter the secondary mass spectrometer for fragmentation. The generated quasi-isotopic peptides with K602 and K080 are fragmented in the secondary mass spectrometer, generating a double peak signal with a mass difference of 36mDa. LysC specifically cleaves the carboxyl terminus (C-terminus) of lysine, so when the resulting LysC-cleaved peptide fragments are fragmented in the secondary mass spectrometry, the y ions produced are in the form of a double peak, while the b ions are in the form of a single peak. LysN specifically cleaves the amino terminus (N-terminus) of lysine, so when the resulting LysN-cleaved peptide fragments are fragmented in the secondary mass spectrometry, the b ions produced are in the form of a double peak, while the y ions are in the form of a single peak. Based on the mass-defect amino acid coding signal (NeuCode signal) in the spectrum, the type of daughter ions can be accurately determined, and the mirror spectra generated by the LysC and LysN mirror enzymes can be used to fill in the missing daughter ions in the single spectrum. At this point, the daughter ions are fully covered and the daughter ion types are distinguishable, which meets the prerequisites for direct-read sequencing and can be performed. The specific steps are as follows:
[0105] 1. Lysine NeuCode labeling and GST protein expression in Escherichia coli
[0106] The GST protein expression strain and labeling culture method used in this example were the same as those in the previous example. However, the lysine component of the labeling medium was replaced with a quasi-isotopic lysine. As shown in Figure 10(A), H1 is SILACE medium supplemented with 30 μg / mL K602, and H2 is SILACE medium supplemented with 30 μg / mL K080.
[0107] 2. Protein purification and labeling efficiency detection
[0108] The obtained cells were lysed using a non-contact ultrasonic disruptor (QSONICA, model Q800R3 Sonicator) using PBS as the protein extraction solution. The supernatant was collected to obtain the total protein after centrifugation at 21,000 g for 10 min at 4°C. The protein was purified using GST 4FF agarose resin ( Biotech, catalog number C600031-0025) were purified to obtain K602-labeled GST protein and K080-labeled GST protein. The purified protein was subjected to 10% SDS-PAGE gel electrophoresis and silver staining, and protein quantification was performed with Coomassie brilliant blue staining. In-gel digestion was performed with 10ng / μLLysC for 14h to obtain peptides for labeling efficiency testing. The results are shown in Figure 11. The purified protein band is single with a molecular weight of approximately 27kDa. The GST peptide "YEEHLYERDEGDK" was randomly selected, and the parent ion mass distribution after non-labeled and heavy-labeled amino acid labeling is shown in Figure 11(B), indicating that the protein is completely labeled.
[0109] 3. LysC / LysN mirror enzyme digestion of KNeuCode tagged protein GST
[0110] 10 μg of each K602-labeled and K080-labeled GST protein was mixed and reduced and alkylated with a final concentration of 5 mM tris(2-carbonylethyl)phosphine hydrochloride (TCEP) and a final concentration of 15 mM iodoacetamide (IAM) at 25°C in the dark for 15 minutes. The reduced and alkylated protein sample was divided into two equal parts, one digested with LysC and the other with LysN, both at an enzyme / substrate ratio of 1:50 (mass ratio). Digestion was performed at 37°C for 14 hours to obtain peptide fragments.
[0111] 4. Chromatography-mass spectrometry tandem analysis
[0112] The peptide fragments obtained above were subjected to C 18 The sample was desalted using StageTips and dissolved in 25 μL of GC-MS loading buffer containing 1% acetonitrile and 0.1% formic acid. 5 μL of the sample was analyzed by GC-MS. The GC-MS instrument model and parameters were as follows: High Performance Liquid Chromatography (Thermo Scientific Easy-nLC 1000) Column: Analytical column (C 18 , 3μm, The mobile phase was 75 μm × 20 cm), with mobile phases A and B containing 0.1% formic acid in acetonitrile, gradient 7% to 14% B in 6 min, 14% to 28% B in 32 min, 28% to 42% B in 14 min, 42% to 95% B in 2 min, and 95% to 95% B in 6 min. The flow rate was 600 nL / min. The mass spectrometer was a Thermo Scientific Q Exactive HF system with a spray voltage of 2.1 kV, a capillary temperature of 275°C, an S-lens of 60%, a collision energy of 27% HCD, a primary resolution of 120,000 at m / z 200, and a secondary resolution of 240,000 at m / z 200, the maximum ion injection time (MIT) was set to 50 ms for the first level and 200 ms for the second level, the dynamic gain control (AGC) was set to 3e6 for the first level and 5e5 for the second level, the minimum AGC was 5e3, the parent ion scan range was m / z 150-1400, the product ion scan range started from m / z 120, the data dependent scan (DDA) mode: TopN was set to 10, the isolation window was set to 1.6 m / z, the dynamic exclusion time was 8 s, and the number of second level micro scans was set to 3.
[0113] The LC-MS base peaks for GST digested with LysC and LysN are shown in Figures 12 (A, B), respectively. A total of 15 specific fully enzymatically cleaved peptides were identified after LysC digestion of the K NeuCode-labeled GST protein. Fifteen specific fully enzymatically cleaved peptides were also identified after LysN digestion. Furthermore, the clustered precursor ion peaks corresponding to the peptides in the chromatographic peaks are all doublets, with signal intensities approaching 1:1. This indicates that the proteins labeled with the two amino acids are nearly equally mixed, and that the enzymatic specificity is not affected by the amino acid labeling, resulting in quasi-isometric peptides after enzymatic digestion. Figure 12 (C, D) shows the mirror spectra of the mirror peptides "LPEMLK" and "KLPEML" obtained by LysC and LysN digestion, respectively. In the secondary spectrum of the LysC-digested peptide, the y ion is a NeuCode doublet with a mass difference of 36 mDa, while the b ion is a singlet. In the secondary spectrum of the LysN-digested peptide, the y ion is a singlet, while the b ion is a NeuCode doublet with a mass difference of 36 mDa. This indicates that the b and y ions can be effectively distinguished by the NeuCode doublet signal.
[0114] 5. Direct reading sequencing of GST protease cleavage peptides
[0115] First, the daughter ion type in the spectrum is determined based on the double peaks with only a slight mass difference in the paired mirror spectra, and the missing daughter ion peaks in the single spectrum are supplemented by the mirror spectra. The daughter ion type is identifiable and the daughter ions are fully covered, so the amino acid residue mass is calculated based on the mass difference of adjacent daughter ions, the amino acid type is determined, and the complete sequence is accurately read directly. As shown in Figure 13, it is a mirror spectrum of a mirror peptide with a mass-to-charge ratio (m / z) of 1168.0942 and a parent ion with two positive charges. The y ion is determined based on the double peak signal in its LysC spectrum, and the b ion is determined based on the double peak signal in the LysN spectrum. By merging the mirror spectra, the complete daughter ion information of the parent ion can be obtained, thereby directly reading the peptide sequence with the parent ion "m / z = 1168.0924, z = 2": K.YIAWPLQGWQATFGGDHPP.K. Similarly, the complete amino acid (aa) sequence of the peptide fragments obtained by mirror-image digestion of the GST protein by LysC and LysN can be efficiently and accurately directly read based on the mirror-image spectrum. As shown in Figure 14, LysC and LysN theoretically produce a total of 16 fully digested mirror-image peptide fragments. In this example, a total of 14 mirror-image peptide fragment sequence information was directly read, with an average reading efficiency of 96.3% and an accuracy rate of 100%. The overall sequence coverage reached 73.11% (174 aa / 238 aa), accounting for 90.2% of the theoretical sequence coverage (174 aa / 193 aa). (Reading efficiency = number of directly read amino acids per directly read peptide fragment / theoretical number of amino acids in the peptide fragment; the theoretical sequence that can be sequenced refers to the peptide fragment sequence with a length of 6-30 amino acids after the specific fully digested peptide fragments of Trypsin or LysargiNase.) These results demonstrate that accurate peptide sequencing can be achieved based on the direct-reading sequencing principle.
[0116] Example 4. Accurate de novo sequencing and sequence assembly of the amino acid code of a single GST protein based on two pairs of mirror enzymes
[0117] During de novo protein sequencing using the shotgun method, a protease is used to cleave long protein sequences into peptides with a length of 6 to 35 amino acids that are easily identified by mass spectrometry. When a single specific protease is used for enzymatic cleavage, the uneven distribution of cleavage sites and the influence of the cleavage efficiency of the cleavage sites on the surrounding amino acids can result in cleaved peptides that are too long or too short. In addition, different peptides have different amino acid compositions and inconsistent ionization efficiencies, making mass spectrometry identification difficult. Ultimately, only partial peptide sequences can be determined, and complete protein sequencing cannot be achieved. Based on the above Examples 2 and 3, two pairs of mirror enzymes, Trypsin / LysargiNase and LysC / LysN, were used to cleave a single protein, GST, to obtain a total of 20 peptides that were accurately sequenced directly. As shown in Figure 15, since the two groups of enzymes have different cleavage sites, the direct-read peptide amino acid sequences have overlapping parts. In this embodiment, the precise direct-read peptides identified by the two groups of enzymes Trypsin / LysargiNase and LysC / LysN were spliced to form a complete GST protein, with a sequence coverage of 76.05% (181aa / 238aa), accounting for 90.95% (181aa / 199aa) of the theoretical sequence coverage, indicating that the use of precise direct-read sequences can further achieve precise de novo protein sequencing sequence splicing and achieve deep coverage of de novo protein sequencing. (Theoretical enzyme-digested peptides: The theoretical enzyme-digested peptides are predicted using the deep learning algorithm DeepDigest, reference Yang, J., Gao, Z., Ren, X., Sheng, J., Xu, P., Chang, C., & Fu, Y. (2021). DeepDigest: Prediction of Protein Proteolytic Digestion with Deep Learning. Analytical chemistry, 93(15), 6094-6103., and it is stipulated that 6-30 amino acids are used as the length of the theoretical enzyme-digested detectable peptide, and the sequence coverage is calculated using the theoretical full enzyme-digested peptide as the theoretical sequence coverage of the measurable protein.) It can be further speculated that the use of more mirror enzyme pairs can more effectively improve the splicing efficiency of protein de novo sequencing.
[0118] Example 5. In-line accurate de novo protein sequencing in a complex system using Escherichia coli as an example
[0119] The novel direct-reading, precise de novo sequencing method proposed in this paper is also applicable to de novo sequencing of proteins in complex systems. Taking the E. coli proteome as an example, the experimental process shown in Figure 16 directly reads protein sequences in complex systems.
[0120] The specific steps are as follows:
[0121] 1. NeuCode dual labeling of lysine and arginine in Escherichia coli
[0122] The BL21(DE3)ΔlysAΔargA bacteria labeled with K602 / R004 and K080 / R040 were obtained by labeling and culturing according to the method described in Example 2. 600 The cells were mixed and protein extracted using PBS. The cells were disrupted by non-contact ultrasound. The lysate was centrifuged at 21,000 g for 10 min at 4°C. The supernatant was collected to obtain total protein. The protein was quantified using the BCA assay and then used for protease digestion.
[0123] 2. Mirror proteinase digestion
[0124] 120 μg of total protein was reduced and alkylated with TCEP at a final concentration of 5 mM and IAM at 15 mM in the dark at 25°C for 15 min. The protein was evenly divided into two parts. One part was digested with a combination of Trypsin and LysC (denoted as Trypsin+LysC). The pH of the digestion system was adjusted to between 7.5 and 8.5 with 50 mM sodium bicarbonate. Trypsin and LysC were added at a substrate-enzyme ratio of 50:1 (mass ratio) and digested at 37°C for 14 h. The other part was digested with 20 mM HEPES buffer was used to adjust the pH of the enzyme digestion system to between 7.5 and 8.5, and lysagiNase and LysN were used for sequential digestion (denoted as LysargiNase+LysN). LysargiNase was first added at a substrate enzyme dosage of 50:1 (mass ratio), and calcium chloride was added at a final concentration of 10mM as an activator of the metalloprotease LysargiNase. The enzyme digestion was carried out at 37°C for 4h, and then LysC was added at a substrate enzyme dosage of 50:1 (mass ratio), and zinc acetate was added at a final concentration of 1mM as an activator of the metalloprotease LysN. The enzyme digestion was carried out at 37°C for 10h. A mixture of K / R NeuCode C-terminal labeled and N-terminal labeled peptides was obtained respectively. The obtained peptides were subjected to C 18 After desalting and fractionation using StageTips, six fractions were obtained for each sample.
[0125] 3. Chromatography-mass spectrometry analysis
[0126] Dissolve the sample in 25 μL of 5% acetonitrile and 1% formic acid mass spectrometry loading buffer, and take 5 μL for chromatography-mass spectrometry analysis. The chromatography-mass spectrometry instrument model and parameter settings are as follows: High performance liquid chromatography (Thermo Scientific Easy-nLC 1000) Chromatographic column: analytical column (C 18 , 3μm, The HPLC-MS / MS instrument was used with a flow cytometer (75 μm × 20 cm). The mobile phases were: phase A: 0.1% formic acid in water; phase B: 0.1% formic acid in acetonitrile. The gradient was set at 5% to 12% phase B for 8 min, 12% to 30% phase B for 80 min, 30% to 42% phase B for 23 min, 42% to 95% phase B for 1 min, and 95% to 95% phase B for 8 min. The flow rate was 600 nL / min. The mass spectrometer was a Thermo Scientific Q Exactive HF instrument with a spray voltage of 2.1 kV, a capillary temperature of 275°C, an S-lens of 60%, a collision energy of 27% HCD, and a primary resolution of 120,000 @ m / z. 200, the secondary resolution was set to 240,000@m / z200, the maximum ion injection time (MIT) was set to 50 ms for primary and 80 ms for secondary, the dynamic gain control (AGC) was set to 3e6 for primary and 1e5 for secondary, the minimum AGC was 5e3, the parent ion scan range was set to m / z 300-1400, the product ion scan range started from m / z 120, the data dependent scan (DDA) mode: TopN was set to 10, the isolation window was set to 1.6 m / z, the dynamic exclusion time was set to 20 s, and the number of secondary microscans was set to 2.
[0127] 4. Direct-read sequencing of Escherichia coli proteins
[0128] A random sample of Escherichia coli proteins was randomly selected for de novo sequencing, using the direct-read sequencing of the E. coli protein P0CE47 (tufA, EFTU1) as an example. The figure shows the direct-read sequencing process for a peptide from the P0CE47 protein (number #1, m / z = 787.9197, z = 2). The mixed peptide sample, after mirror-image combined digestion, was first separated by HPLC. As shown in Figure (A), peptide #1, obtained by trypsin + LysC digestion, eluted between 16.24 and 16.78 min. At a primary mass spectrometer resolution of 120,000 (m / z = 200), its parent ion peak (m / z = 787.9197, z = 2) was a doublet at 18 mDa, with virtually no light-labeled parent ion signal in the spectrum. In the sample digested with LysargiNase+LysN, peptides similarly eluted between 16.70 and 17.20 min, with the parent ion peak (m / z = 787.9197, z = 2) as an 18 mDa doublet, and virtually no light-labeled parent ion signal. This suggests efficient quasi-isobaric amino acid labeling of the protein, and the peptides corresponding to the two parent ions are likely mirror-image peptides. Further analysis of the secondary spectra revealed the presence of a y ion in the Trypsin+LysC digestion spectrum, and a b ion in the LysargiNase+LysN digestion spectrum, both of which conform to the mirror-image peptide pattern. The amino acid sequence of peptide #1 was determined directly using direct-read sequencing: KPHVNVGTIGHVDHGK. The direct-read peak data are shown in Figure 17(C). Similarly, direct-read sequencing was performed on other peptides from this protein, based on this principle. A total of 19 pairs of mirror peptides of the P0CE47 protein were directly read, as shown in Figure 17, with an average reading efficiency of 98.4% and an accuracy of 100%. As shown in Figure 18, the direct-read peptide sequence coverage was 90.7% (267aa / 289aa) of the theoretically sequenable sequence.
[0129] The present invention proposes an underlying technical principle for de novo protein sequencing that combines a mirror image strategy and a quasi-isobaric isotope labeling strategy to simultaneously achieve full coverage of daughter ions and identifiable daughter ion types. Based on this method, the sequence information of amino acids in the secondary spectrum can be directly read, realizing innovation in the underlying technology of de novo protein sequencing and helping to promote the accurate de novo sequencing of the proteome.
[0130] Example 6. Accurate de novo sequencing of chemically labeled proteins from any proteome sample
[0131] The novel direct-reading, precise de novo sequencing method proposed in this paper is also applicable to de novo sequencing of chemically labeled complex proteome samples. Taking the E. coli proteome as an example, the specific steps are as follows:
[0132] 1. Preparation of E. coli proteome samples
[0133] Prepare E. coli culture according to the standard culture method of E. coli or a method similar to Example 2. Taking the standard culture method of E. coli as an example, first pick a fresh BL21 (DE3) ΔlysA ΔargA-GST colony and inoculate it into LB liquid medium containing 50 μg / mL ampicillin. Incubate at 37°C for 10 hours at 200 rpm to activate the bacteria and obtain a seed culture. Transfer an appropriate amount of seed culture to fresh LB liquid medium to obtain a starting culture concentration of 0.3 OD*mL and incubate at 37°C until OD 600 was 1-1.2, and BL21(DE3)ΔlysAΔargA bacteria were obtained.
[0134] The cells were disrupted by non-contact ultrasonication using PBS as the protein extraction medium. The lysate was centrifuged at 21,000 g for 10 minutes at 4°C, and the supernatant was collected to obtain total E. coli cellular protein. The protein was quantified using the BCA assay and then used for protease digestion.
[0135] 2. Trypsin digestion 16 O and 18 O mark
[0136] 120 μg of total protein was reduced and alkylated with TCEP and IAM at a final concentration of 5 mM and 15 mM, protected from light, at 25 ° C for 15 min. The protein was evenly divided into two parts, one of which was digested with a combination of Trypsin and LysC (denoted as Trypsin + LysC). The pH of the digestion system was adjusted to between 7.5 and 8.5 with 50 mM sodium bicarbonate. Trypsin and LysC were added at a substrate-enzyme ratio of 50:1 (mass ratio). Water (H2O) was used as the digestion buffer. The digestion was carried out at 37 ° C for 14 h to obtain trypsin-digested peptides, which were denoted as 16 O-Pep; the other was digested with a combination of Trypsin and LysC (denoted as Trypsin+LysC). The pH of the digestion system was adjusted to between 7.5 and 8.5 with 50 mM sodium bicarbonate, and Trypsin and LysC were added at a substrate-enzyme ratio of 50:1 (mass ratio). 18 O Heavy water (H2 18 O) to replace water (H2O), and then digested at 37℃ for 14h to obtain 18 O-labeled trypsin-digested peptides are denoted as 18 O-Pep, the molecular weight is increased by 4.0085Da compared to the trypsin digested peptides labeled with ordinary water.
[0137] 3. Hydrogen and deuterium dimethyl labeling
[0138] right 18O-Pep was labeled with 5 mM formaldehyde (CH2O) at room temperature for 30 minutes to obtain 18 O-(CH3)2-Pep, the molecular weight increased by 28.0313Da, and the total molecular weight increased by 32.0398Da; 16 O-Pep was labeled with 5 mM deuterated formaldehyde (CD2O) at room temperature for 30 minutes to obtain 16 O-(CD3)2-Pep, molecular weight increased by 32.0564Da. 18 O-(CH3)2-Pep and 16 The mass difference between the two peptides of O-(CD3)2-Pep is 16.6 mDa.
[0139] 4. Peptide fractionation and chromatography-mass spectrometry analysis
[0140] According to the same technique as in Example 2 18 O-(CH3)2-Pep and 16 The two peptide fragments O-(CD3)2-Pep were mixed in equal amounts and fractionated to obtain 6 fractions.
[0141] Dissolve the sample in 25 μL of 5% acetonitrile and 1% formic acid mass spectrometry loading buffer, and take 5 μL for chromatography-mass spectrometry analysis. The chromatography-mass spectrometry instrument model and parameter settings are as follows: High performance liquid chromatography (Thermo Scientific Easy-nLC 1000) Chromatographic column: analytical column (C 18 , 3μm, The chromatographic gradient was set at 5% to 12% phase B in 8 min, 12% to 30% phase B in 80 min, 30% to 42% phase B in 23 min, 42% to 95% phase B in 1 min, and 95% to 95% phase B in 8 min. The flow rate was 600 nL / min. The mass spectrometer was a Thermo Scientific Q Exactive HF with a spray voltage of 2.1 kV, a capillary temperature of 275°C, an S-lens of 60%, a collision energy of 27% HCD, a primary resolution of 120,000 @ m / z 200, and a secondary resolution of 240,000 @ m / z 200, maximum ion injection time (MIT) was set to 50 ms for primary and 80 ms for secondary, dynamic gain control (AGC) was set to 3e6 for primary and 1e5 for secondary, minimum AGC was 5e3, the parent ion scan range was set to m / z 300–1400, the product ion scan range started from m / z 120, data-dependent scan (DDA) mode: TopN was set to 10, the isolation window was set to 1.6 m / z, the dynamic exclusion time was set to 20 s, and the number of secondary microscans was set to 2.
[0142] 5. Direct-read sequencing of quasi-isobaric chemically labeled Escherichia coli protein sequences
[0143] Randomly selected E. coli protein-encoded proteins were sequenced de novo. Using the same E. coli protein POCE47 (tufA, EFTU1) as used in the previous examples, a peptide segment, KPHVNVGTIGHVDHGK, from the POCE47 protein was directly readable, consistent with Figure 17(C). Therefore, both chemical quasi-isobaric labeling and SILAC metabolic labeling using NeuCode amino acid labeling technology can achieve direct-reading, accurate de novo sequencing.
[0144] Those skilled in the art will understand that the above description and illustration of the preferred embodiments of the present invention do not limit the present invention, and should fall within the scope defined by the claims of the present invention as long as they do not depart from the spirit of the present invention.
Claims
1. A direct-reading method for accurate de novo sequencing of proteins that can identify sub-ion types and cover all sub-ions, comprising the following steps: 1) Completely label the proteins of the organism to be sequenced with different paired mass-deficient amino acids to generate mass-deficient amino acid coding signals; the different paired mass-deficient amino acids are lysine and arginine; 2) Digest the labeled proteins with protease and its mirror protease respectively to obtain mirror peptide segments with mass-deficient amino acid coding signals; the protease and its mirror protease are protease pairs that can produce N-terminal and C-terminal mirrors; 3) Analyze the mirror peptide segments generated in step 2) to obtain the secondary spectrum of the mirror peptide segments; 4) Directly read out the polypeptide sequence according to the peptide segment secondary spectrum obtained in step 3).
2. The method according to claim 1, wherein step 1) comprises: Label the mass deficit amino acid pair with heavy stable isotopes to produce labeled mass deficit lysine pair L-lysine- 13 C6 15 N2(K602) and L-lysine- 2 H8(K080), mass deficit arginine pair L-arginine- 15 N4(R004) and L-arginine- 2 H4(R040).
3. The method according to claim 2, wherein the protease and its mirror protease in step 2) are Trypsin / LysargiNase or LysC / LysN protease pairs.
4. The method according to claim 1, wherein step 3) includes using a secondary high-resolution mass spectrometer to analyze the mirror peptide segments to obtain the secondary spectrum of the mirror peptide segments; and step 4) includes determining paired mirror spectra according to the mass characteristics of the mirror peptide segment parent ions, identifying the mass-deficient amino acid coding signals (NeuCode doublets) according to the mass difference in the spectra, judging the y-ions or b-ions in the mirror spectra, merging the mirror spectra, filling in the missing sub-ions in the single spectrum, obtaining a secondary spectrum with known sub-ion types and full coverage of sub-ions, judging the types of amino acids according to the mass difference between adjacent sub-ion peaks in the spectra, and directly reading out the polypeptide sequence therefrom.
5. The method according to claim 4, wherein the resolution of the secondary mass spectrometry in step 3) is set to 60,000 (m / z = 200) or higher.
6. The method according to any one of claims 1-2, wherein the labeling method in step 1) includes: Metabolic labeling culture was carried out by adding labeled mass-deficient amino acid pairs to the biological medium respectively. After the culture, cells were collected, proteins were extracted, and protein samples with mass-deficient amino acid coding (NeuCode) were obtained after quantitative mixing; wherein the labeled mass-deficient amino acid pairs are L-lysine- 13 C6 15 N2 (K602) and L-lysine- 2 H8 (K080), L-argine- 15 N4 (R004) and L-argine- 2 H4 (R040).
7. The method according to any one of claims 1-3, wherein the digestion conditions in step 2) include: Fully denature the protein sample obtained in step 1), divide it into two parts, add Trypsin / LysC to one part for digestion, and add LysargiNase / LysN to the other part for digestion. After the digestion is completed, terminate the digestion reaction, desalt the digested sample, and obtain a peptide sample for mass spectrometry detection.
8. The method according to claim 6, wherein the addition concentration of the mass-deficient lysine pair is 30 μg / mL, and the addition concentration of the mass-deficient arginine pair is 20 μg / mL; the extraction solution for extracting proteins is PBS non-denaturing lysis solution or denaturing lysis solution containing 8 M urea.
9. The method according to claim 7, wherein 5 mM tris(2-carboxyethyl)phosphine hydrochloride (TCEP) and 15 mM iodoacetamide (IAM) are used for reduction and alkylation treatment in the dark at 25 °C for 15 min to fully denature the protein; the treated protein solution is divided into two equal parts, and the urea concentration in the solution is diluted to less than 1 M with 50 mM ammonium bicarbonate. One part is added with Trypsin / LysC / ArgC at a substrate / enzyme ratio of 50:1 (mass ratio), and the pH is adjusted to between 7.5 and 8.5 with 1 M ammonium bicarbonate solution, and digested at 37 °C for 10 - 16 h; the other part is added with LysargiNase / LysN / AgrN at a substrate / enzyme ratio of 50:1 (mass ratio), and the pH is adjusted to between 7.5 and 8.5 with 1 M HEPES buffer, and digested at 37 °C for 10 - 16 h; or digested with GluC or AspN respectively under their suitable conditions; then 1% FA is added to each enzymatic digestion system to terminate the digestion reaction, and the digested samples are desalted using C18 StageTips to obtain peptide samples for mass spectrometry detection.
10. The method according to any one of claims 1-9, wherein the chromatographic separation and mass spectrometry analysis in step 3) comprises: The obtained peptide samples are respectively dissolved with a liquid chromatography tandem mass spectrometry loading buffer containing 0.1% formic acid and 1% acetonitrile, separated by reversed-phase high-performance liquid chromatography, detected by tandem mass spectrometry, and the labeling efficiency is calculated. To identify the NeuCode doublet signal in the secondary spectrum, the secondary mass spectrometry resolution is set to 60000 (m / z = 200) or higher; Step 4) includes: 4a) Mirror spectrum pair matching and daughter ion identification: Paired mirror spectra are obtained by matching according to the characteristics of the mirror peptide parent ion mass and retention time. In the spectrum digested by Trypsin / LysC / ArgC, the y ion peak is determined by the K / R NeuCode doublet signal. In the spectrum digested by LysargiNase / LysN / ArgN, the b ion peak is determined by the K / R NeuCode doublet signal; 4b) Mirror spectrum merging and sequence direct reading: The missing daughter ions in the single spectrum are supplemented according to the mirror peptide fragment ion mass characteristics, the mass difference between adjacent daughter ion peaks is calculated, the amino acid type is determined, and the peptide sequence is directly and accurately read out; 4c) Sequence splicing: Based on multiple pairs of mirror enzymatic digestions, combined with K / R NeuCode double labeling and KNeuCode single labeling, the peptide segments are accurately sequenced directly, and the overlapping sequences are used for splicing and reconstruction to directly and accurately read out the protein sequence.
11. A method for direct reading of accurate de novo protein sequencing, comprising the following steps: i) Digest the proteins of the organism to be sequenced with a combination protease of Trypsin and LysC, using H2O and oxygen-18 heavy water (H2 18 O) as digestion buffers respectively, to obtain 16 O- or 18 O-heavy stable isotope-labeled digested peptides 16 O-Pep and 18 O-Pep; ii) Fully label the peptide segments in step i) with isotopes of deuterium (D) or hydrogen (H) of formaldehyde to produce peptide segments with chemically quasi-isobaric labels 16 O-(CD3)2-Pep and 18 O-(CH3)2-Pep, and the mass difference between the two peptide segments is 16.6 mDa; iii) Analyze the labeled peptide segments generated in step ii) to obtain peptide secondary spectra; iv) According to the peptide secondary spectra obtained in step iii), based on the amino acid coding daughter ion types, directly read out the protein sequence according to the mass difference between the same type of daughter ions.
12. The method according to claim 11, comprising the following steps: (1) Reduce and alkylate the proteins of the organism to be sequenced; divide the proteins evenly into two parts. Digest one part with a combination of Trypsin and LysC, using water (H2O) as the digestion buffer to obtain trypsin-digested peptide fragments, denoted as 16 O-Pep; Digest the other part with a combination of Trypsin and LysC, using 18 heavy water (H2 18 O) to replace water (H2O) to obtain 18 O-labeled trypsin-digested peptide fragments, denoted as 18 O-Pep; (2) For the 18 O-Pep is labeled with formaldehyde (CH2O) to obtain 18 O-(CH3)2-Pep, and for 16 O-Pep is labeled with deuterated formaldehyde (CD2O) to obtain 16 O-(CD3)2-Pep; (3) Mix equal amounts of 18 O-(CH3)2-Pep and 16 O-(CD3)2-Pep, two peptide segments, and fractionate them to obtain fractionated peptide segments; (4) Directly read and sequence the protein sequence of quasi-isobaric chemical labeling.
Citation Information
Patent Citations
Protein amino acid sequence de novo sequencing method based on unequal stable isotope labeling at two ends of polypeptide
CN105301119A
De novo sequencing method
CN107729719A
Novel protein methylation modification reverse enrichment method based on mirror image enzyme orthogonal principle and application
CN112285265A
Amino acid sequence determination method based on quasi-isobaric double labeling at two ends of polypeptide
CN112986570A
Direct-reading protein and proteome accurate de novo sequencing method based on identifiable daughter ion type and full daughter ion coverage
CN118033144A
Cited By
Method for screening yak milk and dzo milk differential protein markers based on proteomics and application in identification of yak milk and dzo milk
CN120741729A