Method and system for rapidly screening PRM targeted detection target peptide fragment based on scoring algorithm

By constructing a scoring model based on scoring algorithm, integrating fragment ion mass and quantitative reliability parameters, the problem of low screening efficiency of target peptides in PRM targeted detection is solved, and efficient and accurate screening of target peptides and PRM detection is achieved.

CN120472977APending Publication Date: 2025-08-12PUDU ZHONGHE (WUHAN) LIFE TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510550292.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-29
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

In the existing PRM targeted detection methods, the screening of target peptides depends on a single threshold or artificial experience, resulting in false positives, low quantitative accuracy and low efficiency in the screening results, and it is difficult to accurately screen out the appropriate target peptides in one go.

Method used

A scoring algorithm is used to integrate key parameters such as fragment ion mass, peptide specificity and quantitative reliability to build a data-driven scoring model. By calculating the score 1, score 2 and comprehensive score 3 of the candidate peptide, the target peptide segment is sorted from high to low based on the comprehensive score 3.

Benefits of technology

The screening efficiency and detection success rate of target peptides are significantly improved, and suitable target peptides can be screened out at one time. The success rate of PRM targeted detection can reach 100%, overcoming the subjectivity and low-throughput bottlenecks of traditional methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120472977A_ABST
    Figure CN120472977A_ABST
Patent Text Reader

Abstract

The invention relates to a method and system for rapidly screening PRM targeted detection target peptide fragments based on a scoring algorithm, and the method comprises the steps: firstly, according to a library searching result of original data of a biological sample non-targeted mass spectrometry detection proteome, taking a targeted peptide fragment corresponding to an identified target protein as a candidate peptide fragment; calculating a score 1, a score 2 and a comprehensive score 3 of each candidate peptide fragment, wherein # imgabs0 # imgabs1 # n is the total number of b and y fragment ions; and finally ranking the # imgabs2 # from high to low according to the comprehensive score 3, and selecting the candidate peptide fragments with the first three ranks as target peptide fragments. According to the method, the target peptide fragment corresponding to each target protein can be screened out from a plurality of target peptide fragments corresponding to a large number of target proteins at a time, the PRM target detection success rate based on the screened target peptide fragments can reach 100%, and the screening efficiency of the target peptide fragments is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to biological detection technology, and more specifically to a method and system for rapidly screening PRM targeted detection target peptides based on a scoring algorithm. Background Art

[0002] Parallel reaction monitoring (PRM) is a targeted proteomics technology based on high-resolution mass spectrometry. In the PRM mass spectrometry acquisition mode, according to the target peptide precursor ion (Precursor) parameters set for acquisition, the precursor ions of the target peptide are selectively detected in the quadrupole Q1 to obtain the primary mass spectrum MS1; these precursor ions are then fragmented in the collision cell HCD to generate corresponding fragment ions (Fragment), and the mass-to-charge ratio of the fragment ions is detected by a mass analyzer such as Orbitrap or TOF to obtain a high-resolution mass spectrum of all fragment ions, namely the secondary mass spectrum MS2; finally, the target peptide is quantitatively analyzed based on MS2. In the process of establishing the PRM detection method, the screening of target peptides is a key step, in which the target peptides must strictly meet the following dual standards:

[0003] ① Detectability: The target peptide must have good ionization efficiency and stable fragment ion generation ability.

[0004] ② Uniqueness: The target peptide can only be mapped to a unique target protein, avoiding interference from homologous proteins and affecting the accuracy of the final target protein quantification.

[0005] Currently, the screening of PRM target peptides mainly relies on simple software parameter filtering or screening based on manual experience. For example, Skyline software, which is commonly used for targeted proteomics analysis, uses a single threshold value such as peptide length, charge, intensity or number of spectra to preliminarily screen the target detection peptides, and uses the screened target peptides as target peptides for PRM targeted detection. However, after proteinase hydrolysis, the target will correspond to multiple target peptides, but not all target peptides can be used as a suitable target peptide for PRM targeted detection. For example, the target peptides screened by a single parameter of the existing software may have low signal intensity or insufficient fragment ion richness, resulting in false positives or low quantitative accuracy of the final PRM targeted detection target peptide.

[0006] However, the screening of targeted peptides based on manual experience is highly subjective, inefficient, and difficult to scale. Although existing tools such as TPP and PeptidePicker can be automated, they focus more on peptide identification rather than target optimization, and lack a systematic evaluation of fragment ion mass and quantitative accuracy. As a result, the screening results may contain targeted peptides with low response or susceptible to matrix interference, ultimately affecting the accuracy of PRM detection. It can be seen that the screening of target peptides determines the success of PRM targeted detection. Therefore, it is crucial to quickly screen out suitable target peptides from multiple targeted peptides corresponding to the target protein. However, the existing methods for screening target peptides are highly subjective or the parameters considered are not comprehensive enough. It is uncertain whether it is feasible to use the screened target peptides as target peptides for PRM targeted detection. It is difficult to accurately screen out target peptides at one time, resulting in low efficiency of PRM targeted detection of target proteins. Summary of the Invention

[0007] In response to the above defects or improvement needs of the prior art, the present invention provides a method and system for quickly screening target peptides for PRM targeted detection based on a scoring algorithm. The method aims to construct a data-driven scoring model by integrating key parameters such as fragment ion mass, peptide specificity and quantitative reliability, and screen target peptides according to the ranking of scores. The method can screen out the target peptides corresponding to each target protein from multiple targeted peptides corresponding to a large number of target proteins at one time, and the success rate of PRM targeted detection based on the screened target peptides can reach 100%, which significantly improves the screening efficiency of the target peptides, thereby solving the technical problem that the existing method for screening target peptides is difficult to accurately screen out the target peptides at one time and has low efficiency.

[0008] To achieve the above objectives, according to one aspect of the present invention, a method for rapidly screening target peptides for PRM targeted detection based on a scoring algorithm is provided, which comprises the following steps:

[0009] (1) Screening candidate peptides: Based on the search results of the original data of non-targeted mass spectrometry detection of biological samples, the targeted peptides corresponding to the identified target proteins are selected as candidate peptides;

[0010] (2) Calculate the candidate peptide scores: calculate the score 1, score 2 and comprehensive score 3 of each candidate peptide in step (1);

[0011]

[0012] n is the total number of b and y fragment ions;

[0013]

[0014] (3) Screening target peptides: Based on the comprehensive score 3 corresponding to each target protein calculated in step (2), the target peptides are screened according to the following method:

[0015] If the number of candidate peptides is less than 3, at least one candidate peptide with the highest comprehensive score 3 will be selected as the target peptide;

[0016] If the number of candidate peptides is ≥3, the candidate peptides are sorted from high to low according to the comprehensive score 3, and one or more peptides among the top three candidate peptides are selected as target peptides.

[0017] Preferably, in the method, if the number of candidate peptides is ≥3, the candidate peptides are sorted from high to low according to the comprehensive score 3, and the top two peptides among the candidate peptides are used as target peptides for PRM quantitative detection of the target protein.

[0018] Preferably, in the method, the comprehensive score 3 is calculated according to the normalized score 1 and score 2.

[0019] Preferably, in the method, the normalization processing is performed using Min-Max normalization method, Z-score normalization method or nonlinear normalization method.

[0020] Preferably, in the method, the library search result is obtained by performing library search analysis on the original mass spectrum data according to the library search software, and the original mass spectrum data includes the original data collected in DDA and DIA modes.

[0021] Preferably, in the method, the library search software includes one or more combinations of MaxQuant, FragPipe and DIA-NN.

[0022] According to another aspect of the present invention, a system for rapidly screening target peptides for PRM targeted detection based on a scoring algorithm is provided, which includes a data acquisition module, a judgment and calculation module, and a target peptide screening module;

[0023] The data acquisition module is used to obtain the search results based on the original data of the proteome detected by non-targeted mass spectrometry and submit them to the judgment and calculation module. The search results include the search result files of qualitative or quantitative data of proteins, peptides, b fragment ions and y fragment ions;

[0024] The judgment and calculation module is used to judge whether the target protein is identified and whether the identified target protein has a targeting peptide segment; if the target protein is identified and it has a targeting peptide segment, the targeting peptide segment is used as a candidate peptide segment, and the candidate peptide segment score 1, score 2 and comprehensive score 3 are calculated as follows;

[0025]

[0026] n is the total number of b and y fragment ions;

[0027] And submit the comprehensive score of 3 to the target peptide screening module;

[0028] The target peptide screening module screens target peptides according to the method of the present invention and automatically outputs the screening results.

[0029] Preferably, if the number of candidate peptides is ≥3, the system sorts them from high to low according to the comprehensive score 3, and uses the top two peptides among the candidate peptides as the target peptides for PRM quantitative detection of the target protein.

[0030] Preferably, in the system, the comprehensive score 3 is calculated according to the score 1 and the score 2 after normalization, and the normalization is performed using the Min-Max normalization method, the Z-score normalization method or the nonlinear normalization method.

[0031] Preferably, the raw data of the system are raw data collected in DDA mode, and the library search result files include proteinGroup.txt, peptides.txt and msms.txt files generated by MaxQuant software, and protein.tsv, peptide.tsv and library.tsv files output by FragPipe software;

[0032] The raw data is the raw data collected in the DIA mode, and the search result files include report.pg_matrix.tsv, report.pr_matrix.tsv, and report-lib.tsv files generated by DIA-NN software and FragPipe software.

[0033] In general, the above technical solutions conceived by the present invention can achieve the following beneficial effects compared with the prior art:

[0034] The method provided by the present invention, based on a scoring algorithm, rapidly screens target peptides for PRM-targeted detection. By integrating key parameters such as fragment ion mass, peptide specificity, and quantitative reliability, candidate peptides are first screened. Score 1, Score 2, and Comprehensive Score 3 are then calculated for each candidate peptide. Target peptides are then screened based on the Comprehensive Score 3. This method improves the detectability and quantitative reliability of targeted peptides through a data-driven scoring model, significantly increasing the screening efficiency and detection success rate of PRM-targeted peptides. It overcomes the subjectivity and low-throughput bottlenecks of traditional manual screening and provides efficient technical support for large-scale PRM-targeted detection of target proteins. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Figure 1 This is an overview of the method flow for screening PRM target peptides;

[0036] Figure 2 This is the fragment ion chromatogram of the target peptide corresponding to the protein GAPDH, from top to bottom are peptides VIISAPSADAPMFVMGVNHEK, GALQNIIPASTGAAK, VVDLMAHMASK;

[0037] Figure 3 This is the fragment ion chromatogram of the target peptide corresponding to the protein ACADM, from top to bottom are peptides EEIIPVAAEYDK, TGEYPVPLIR, and ANWYFLLAR;

[0038] Figure 4 This is the fragment ion chromatogram of the target peptide corresponding to the protein GNS, from top to bottom are peptides IQEPNTFPAILR, TQMDGMSLLPILR;

[0039] Figure 5 This is the fragment ion chromatogram of the target peptide corresponding to the protein CSNK2A2, from top to bottom are the peptides EPFFHGQDNYDQLVR, VLGTEELYGYLK, and HLVSPEALDLLDK;

[0040] Figure 6 This is the fragment ion chromatogram of the target peptide corresponding to the protein TCEA1, which includes peptides DTYVSSFPR and MTAEEMASDELK from top to bottom. DETAILED DESCRIPTION

[0041] In order to further illustrate the technical means and effects adopted by the present invention to achieve the predetermined purpose of the invention, the specific implementation methods, structures, features and effects of the present invention are described in detail below in conjunction with the accompanying drawings and preferred embodiments.

[0042] According to the different mass spectrometry scanning methods, it can be divided into three categories: data-dependent acquisition (DDA), data-independent acquisition (SWATH / DIA), and targeted analysis (SRM, PRM). Among them, PRM technology is based on a spectral library established based on samples containing target proteins and target peptides or synthetic standard peptides, and a specific mass spectrometry acquisition method for target proteins and target peptides is established to monitor the target peptides in the sample to achieve relative or absolute quantification. The target peptide used for PRM detection is a peptide unique to the target protein, namely the target peptide, but one target protein often corresponds to multiple target peptides. Therefore, before conducting PRM detection, it is necessary to screen suitable target peptides.

[0043] The present invention is based on non-targeted mass spectrometry detection (such as DDA or DIA mode) to extract target protein information. The target protein information includes whether the target protein is identified and whether there is a unique peptide segment. When the target protein is identified and there is a unique peptide segment, the targeted peptide segment corresponding to the identified target protein is used as a candidate peptide segment. In mass spectrometry analysis, fragment ions (FragmentIons) are produced by the fission of molecular ions (MolecularIon) during or after the ionization process. The molecular ions of the peptide segment are obtained in the first-stage mass spectrometry, and the ions of the target peptide segment are selected as parent ions, which collide with an inert gas to break the peptide bonds in the peptide chain and form a series of ions, an N-terminal fragment ion series (B series, i.e., b fragment ions) and a C-terminal fragment ion series (Y series, i.e., y fragment ions). Generally speaking, the longer the peptide segment, the richer the b and y fragment ions will be. The existing method usually uses the richer the richness of the b and y fragment ions to evaluate the higher the credibility of the peptide segment.

[0044] To more objectively evaluate the feasibility of PRM detection of target proteins, we propose to calculate scores based on the abundance and signal intensity of the b and y fragment ions identified by the target protein, and obtain a comprehensive score based on the consistent proportion of the two. The higher the comprehensive score, the higher the feasibility of the peptide in PRM detection. Specifically, the target peptide corresponding to the identified target protein is used as a candidate peptide, and the score 1, score 2 and comprehensive score 3 of each candidate peptide are calculated;

[0045]

[0046] n is the total number of b and y fragment ions;

[0047]

[0048] According to the calculated comprehensive score 3 corresponding to each target protein, the candidate peptides are sorted from high to low according to the comprehensive score 3, and one or more peptides among the top three candidate peptides are used as target peptides.

[0049] Furthermore, the score 1 and score 2 are the normalized scores 1 and 2. The comprehensive score 3 is calculated according to the normalized scores 1 and 2. The top three candidate peptides are sorted from high to low according to the comprehensive score 3, and one or more peptides are selected as target peptides. The target peptides corresponding to each target protein can be screened out at one time for PRM detection of quantitative target proteins. The PRM detection success rate of the target peptides screened based on this method can reach 100%.

[0050] Based on this, the present invention provides a method for rapidly screening PRM targeted detection target peptides based on a scoring algorithm, which comprises the following steps:

[0051] (1) Screening candidate peptides: Based on the results of non-targeted mass spectrometry detection of biological samples, the targeted peptides of the identified target proteins are selected as candidate peptides.

[0052] (2) Calculate the candidate peptide scores: calculate the score 1, score 2 and comprehensive score 3 of each candidate peptide in step (1);

[0053]

[0054] n is the total number of b and y fragment ions;

[0055]

[0056] (3) Screening target peptides: Based on the comprehensive scores 3 corresponding to each target protein calculated in step (2), the candidate peptides are sorted from high to low according to the comprehensive scores 3, and one or more peptides among the top three candidate peptides are used as target peptides; preferably, the top two candidate peptides are both used as target peptides.

[0057] In some embodiments, a non-targeted mass spectrometry protein group is detected on a biological sample, and a library search analysis is performed on the raw data of the non-targeted mass spectrometry detection to obtain a library search result file including qualitative or quantitative data of proteins, peptides, and fragment ions. The raw mass spectrometry data includes data collected in DDA and DIA modes, and the library search software types include one or more combinations of MaxQuant, FragPipe, and DIA-NN. For example, MaxQuant or FragPipe or DIA-NN software is used to perform a library search analysis on the raw mass spectrometry data, wherein for the DDA mode detection results, the library search result files include proteinGroup.txt, peptides.txt, and msms.txt files generated by MaxQuant software, and protein.tsv, peptide.tsv, and library.tsv files output by FragPipe software; for the DIA model detection results, the library search result files include report.pg_matrix.tsv, report.pr_matrix.tsv, and report-lib.tsv files generated by DIA-NN and FragPipe software, respectively.

[0058] Determine whether the target protein has been identified based on the search results. If the target protein has been identified, determine whether it has a targeted peptide. If the target protein has not been identified, the protein cannot be processed in the next step and should be excluded. For example, based on the DDA search results, if the target protein appears in the Protein IDs column of the proteinGroup.txt file generated by MaxQuant software, or in the Protein ID column of the protein.tsv file output by FragPipe software, then the target protein is considered to be identified. If the target protein does not appear in both MaxQuant and FragPipe software, then it is considered not to be identified.

[0059] Based on the DIA search results, if the target protein appears in the Protein.Group column of the report.pg_matrix.tsv file generated by DIA-NN or FragPipe software, then the target protein is considered to be identified. If the target protein does not appear in either DIA-NN or FragPipe software, then it is considered not to be identified.

[0060] All peptide sequences in the peptide search result file corresponding to the identified target protein are matched with the amino acid sequences of all proteins in the protein database FASTA file. If the peptide sequence matches only the amino acid sequence of the target protein, then this peptide can be determined to be a unique peptide of the target protein, that is, a targeted peptide, and can be used as a candidate peptide. If the peptide sequence matches the amino acid sequences of multiple target proteins, then this peptide can be determined to be not a targeted peptide, cannot be used as a candidate peptide, and should be excluded.

[0061] When the target protein is identified and there is a targeted peptide for the target protein, the targeted peptide and fragment ion information of each target protein are extracted. The scores of score 1, score 2 and comprehensive score 3 are calculated according to the abundance and signal intensity of the b and y fragment ions identified by the target protein, as follows:

[0062] ① Calculate the ratio of the total number of b and y fragment ions of each identified targeted peptide to the length of the targeted peptide to obtain a score of 1 (by fragment ratio). The higher the score 1, the higher the credibility of the targeted peptide segment. Preferably, the ratio of the b and y fragment ions is normalized so that it is within the numerical range of 0-1 to obtain Fragment_Ratio_norm, that is, the score 1 after normalization. The method of normalization described in the present invention is not limited in the present invention, and it mainly includes Min-Max normalization, Z-score normalization and nonlinear normalization, wherein nonlinear normalization includes logarithmic transformation, square root transformation, exponential transformation, etc. Min-Max normalization refers to scaling the data to between 0 and 1, and Z-score normalization refers to converting the data into a distribution with a mean of 0 and a standard deviation of 1.

[0063] ② Calculate the sum of the signal intensities of all b and y fragment ions of each identified targeted peptide to obtain a score of 2 (bytotal intensity). n is the total number of b and y fragment ions; a higher score of 2 indicates a higher feasibility of detecting the targeted peptide. The signal intensities of the b and y fragment ions are preferably normalized to a value between 0 and 1 to obtain Total_Intensity_norm, which is the normalized score of 2.

[0064] ③According to Calculate the combined score 3 (combined_score), preferably according to the normalized score 1 and score 2. The higher the combined score, the higher the feasibility of the targeted peptide in PRM detection.

[0065] If the number of candidate peptides is greater than 3, sort them from high to low according to the comprehensive score 3, and select one or more of the top three targeting peptides as target peptides. Preferably, multiple peptides from the top three candidate peptides are selected as target peptides, and more preferably, the top two candidate peptides are selected as target peptides.

[0066] In addition, the present invention also provides a system for rapid screening of PRM targeted detection target peptides based on a scoring algorithm, which includes a data acquisition module, a judgment and calculation module, and a target peptide screening module;

[0067] The data acquisition module is used to obtain the search result file based on the original data of the proteome detected by non-targeted mass spectrometry and submit it to the judgment and calculation module; the search result file includes the search result file of the qualitative or quantitative data of proteins, peptides, b fragment ions and y fragment ions.

[0068] The judgment and calculation module judges whether the target protein is identified based on the search result file. If the target protein is identified, it is judged whether the target protein has a targeting peptide segment; when the target protein is identified and the target protein contains a targeting peptide segment, the targeting peptide segment is used as a candidate peptide segment, and the score 1, score 2 and comprehensive score 3 of the candidate peptide segment are calculated according to the method described in the present invention, and the calculation results are submitted to the screening module, the score 1 is the ratio of the total number of b and y fragment ions of the targeting peptide segment to the length of the targeting peptide segment, and the score 2 is the sum of the signal intensities of the b and y fragment ions of the targeting peptide segment; the ratio of the b and y fragment ions is preferably normalized to obtain an A value; the sum of the signal intensities of the b and y fragment ions is preferably normalized to obtain a B value, and the comprehensive score 3 is preferably The normalized score.

[0069] The target peptide screening module screens target peptides according to the method of the present invention and automatically outputs the screening results. Specifically:

[0070] Based on the calculation results of comprehensive score 3, target peptides were screened as follows:

[0071] If the number of candidate peptides is less than 3, at least one candidate peptide with the highest comprehensive score 3 will be selected as the target peptide;

[0072] If the number of candidate peptides is ≥3, they are sorted from high to low according to the comprehensive score 3, and one or more peptides among the top three candidate peptides are selected as target peptides; preferably, the top two peptides among the candidate peptides are used as target peptides for PRM quantitative detection of the target protein.

[0073] Alternatively, a targeted peptide segment having a comprehensive score 3 greater than or equal to a threshold is used as the target peptide segment. In some embodiments, the threshold is 0.5.

[0074] In some embodiments, the raw data is raw data collected in DDA mode. For the DDA mode detection results, the data acquisition module of the system is used to obtain proteinGroup.txt, peptides.txt and msms.txt files generated by MaxQuant software, and protein.tsv, peptide.tsv and library.tsv files output by FragPipe software;

[0075] The raw data is the raw data collected by the DIA mode. For the DIA mode detection results, the data acquisition module of the system is used to obtain the report.pg_matrix.tsv, report.pr_matrix.tsv and report-lib.tsv files generated by the DIA-NN and FragPipe software.

[0076] The following are examples

[0077] Example 1 Automatic screening of suitable target peptides based on scoring algorithm

[0078] Before performing PRM targeted peptide screening, it is necessary to extract target protein information through non-targeted mass spectrometry detection (DDA or DIA mode), which includes whether the target protein is identified and whether there is a unique targeted peptide. The process of the method is as follows: Figure 1 The specific steps are as follows:

[0079] Sample pretreatment: The biological sample to be tested undergoes protein extraction, denaturation, disulfide bond reduction, alkylation, trypsin hydrolysis, and desalting in sequence. The peptides are vacuum-evacuated and stored at -20°C in preparation for mass spectrometry detection. Any sample that can be used for proteomics can be used in the present invention. Since the specific processing procedures for different types of biological samples vary, the steps for protein extraction, denaturation, etc. are described uniformly here. The biological samples to be tested include cells, tissues, body fluids (such as blood, urine, cerebrospinal fluid, etc.), microorganisms, and environmental samples (such as soil, feces, etc.).

[0080] Mass spectrometry detection: The enzymatic peptides were reconstituted in 0.1% formic acid aqueous solution and mass spectrometry detection was performed using a Thermo Fisher Scientific UltiMate 3000RSLCnano nanoliter liquid phase tandem Q Exactive HF mass spectrometer. To illustrate that this method is not affected by elution conditions or mass spectrometry detection conditions, two different elution and mass spectrometry detection conditions were used in this example. The mass spectrometry acquisition mode was 35 min DIA or 65 min DDA. The specific HPLC detection gradients are shown in Tables 1 and 2 below, and the DIA or DDA mass spectrometry acquisition parameters are shown in Tables 3 and 4 below.

[0081] Table 1 HPLC detection 35min gradient

[0082] Time (min) Flow rate (μL / min) Mobile phase A (%) Mobile phase B (%) 0 0.3 94 6 5 0.3 88 12 25 0.3 60 40 28 0.3 40 60 29 0.3 10 90 35 0.3 10 90

[0083] Table 2 HPLC detection 65min gradient

[0084] Time (min) Flow rate (μL / min) Mobile phase A (%) Mobile phase B (%) 0 0.3 94 6 5 0.3 89 11 18 0.3 85 15 40 0.3 75 25 54 0.3 53 47 55 0.3 10 90 65 0.3 10 90

[0085] Table 3 DIA mass spectrometry detection parameter settings

[0086]

[0087]

[0088] Table 4DDA mass spectrometry detection parameter settings

[0089]

[0090] The above two liquid phase gradient and mass spectrometry detections can be used for all types of biological samples.

[0091] Use MaxQuant, FragPipe, or DIA-NN software to perform library search analysis on the raw mass spectrometry data, generating library search results that include qualitative or quantitative data files for proteins, peptides, and fragment ions. For DDA results, these include the proteinGroup.txt, peptides.txt, and msms.txt files generated by MaxQuant, and the protein.tsv, peptide.tsv, and library.tsv files output by FragPipe. For DIA results, these include the report.pg_matrix.tsv, report.pr_matrix.tsv, and report-lib.tsv files generated by DIA-NN and FragPipe, respectively.

[0092] By inputting the protein search result file, if the target protein appears with the corresponding protein ID number in the protein search result file, it can be determined that the protein has been identified; if the target protein does not appear with the corresponding protein ID number in the protein search result file, it can be determined that the protein cannot be identified.

[0093] For example, for DDA results, if the target protein appears in the Protein IDs column of the proteinGroup.txt file generated by MaxQuant software, or in the Protein ID column of the protein.tsv file output by FragPipe software, then the target protein is considered to be identified. If the target protein does not appear in either MaxQuant or FragPipe software, then it is considered not to be identified.

[0094] For DIA detection results, if the target protein appears in the Protein.Group column of the report.pg_matrix.tsv file generated by DIA-NN or FragPipe software, then the target protein is considered to be identified. If the target protein does not appear in either DIA-NN or FragPipe software, then it is considered not to be identified.

[0095] Input a peptide or precursor search result file and match it with a protein database FASTA file to determine whether a unique peptide exists for the target protein. For example, after inputting a peptide search result file, a target protein may have multiple peptide sequences. These peptide sequences are then matched against the amino acid sequences of all proteins in the protein database FASTA file. If a peptide sequence matches only the amino acid sequence of the target protein, it can be determined to be a unique peptide for the target protein, i.e., a targeted peptide.

[0096] Based on the DDA test results, use the msms.txt file data generated by MaxQuant software or the library.tsv file data generated by FragPipe software to determine whether the peptide is unique based on the searched protein database.

[0097] Based on the DIA test results, the report-lib.tsv and report.pr_matrix.tsv file data are used to determine whether the peptide is unique based on the searched protein database.

[0098] Only when both are yes, that is, the target protein is identified and the target protein has a targeting peptide, can the next step of calculating the target peptide score of the target protein candidate be performed. In this example, 5 target proteins were randomly selected, and the specific protein list is as follows:

[0099] Table 5 Target protein information list

[0100] Protein Accession Number Gene name Has it been identified Whether there are unique peptides P04406 GAPDH yes yes P11310 ACADM yes yes P15586 GNS yes yes P19784 CSNK2A2 yes yes P23193 TCEA1 yes yes

[0101] After importing the target protein sequence .fasta file, DDA or DIA raw mass spectrometry data, and search results into Skyline software to build a spectral library, Skyline software then extracts peptide and fragment ion information for each target protein based on thresholds such as FDR and mass deviation, and filters the spectral library for the target protein's corresponding targeted peptides. For example, this is achieved in Skyline software by setting the maximum q-value to 0.01 (i.e., FDR) and the MS1 or MS2 tolerance (i.e., mass deviation, which varies depending on the mass spectrometer). A list of targeted peptide information corresponding to these five target proteins is shown in Table 6.

[0102] Table 6 Targeted peptide information list (selected results)

[0103]

[0104]

[0105] The targeting peptides in Table 6 were identified by skyline software. The results showed that 24 candidate targeting peptides were identified for these five proteins, and each protein corresponded to multiple targeting peptides.

[0106] After the candidate target peptide ions are collided and fragmented, secondary fragment ions are generated. The sequence information of the target peptide and protein is obtained by analyzing the mass-to-charge ratio (m / z) of these fragment ions. HCD collision produces b and y fragment ions. Although generally speaking, the richer the fragment ions, the higher the credibility of the peptide, in order to better evaluate the feasibility of the target peptides screened by the target protein for PRM detection, the present invention calculates the score value based on the richness and signal intensity of the b and y fragment ions identified by the target protein, as follows:

[0107] First, by calculating the ratio of the total number of b and y fragment ions of each identified targeted peptide segment to the length of the targeted peptide segment, a score 1 (by fragment ratio) is obtained. The higher the score 1, the higher the credibility of the targeted peptide segment. The normalization method mainly includes Min-Max normalization, Z-score normalization and nonlinear normalization, wherein Min-Max normalization refers to scaling the data to between 0 and 1, Z-score normalization refers to converting the data into a distribution with a mean of 0 and a standard deviation of 1, and nonlinear normalization includes logarithmic transformation, square root transformation, exponential transformation, etc. In this embodiment, Min-Max normalization is selected, that is, the normalization method is (X-Xmin) / (Xmax-Xmin).

[0108] Next, the total intensity score 2 (by total intensity) is calculated by summing the signal intensities of all b and y fragment ions for each identified targeted peptide. A higher score 2 indicates a higher feasibility of detecting the targeted peptide. The signal intensities of the b and y fragment ions are preferably normalized to a range of 0-1 to obtain Total_Intensity_norm. In this example, Min-Max normalization is used, i.e., the normalization method is (X-Xmin) / (Xmax-Xmin), the same below.

[0109] Finally, according to the normalized proportions of A and B after score 1 and score 2, the distance from the point (A, B) to the origin (0, 0) is calculated using A and B as the horizontal and vertical coordinates. The distance values were normalized to a range of 0-1 to obtain a combined score (combined_score). A higher combined score indicates a higher feasibility of the targeted peptide in PRM detection. In this example, Min-Max normalization was used, i.e., the normalization method was (X - Xmin) / (Xmax - Xmin).

[0110] According to the above method, the score 1 (by fragment ratio), score 2 (by total intensity), normalized A value (Fragment_Ratio_norm) and B value (Total_Intensity_norm) of different targeted peptides were obtained, and the comprehensive score (combined_score) was obtained. The corresponding results of DIANN are shown in Table 7.

[0111] Table 7 Scoring results of different targeted peptides (selected results)

[0112]

[0113]

[0114] In addition, the system of the present invention is equipped with a scoring algorithm program written according to this method. By inputting the msms.txt (MaxQuant software) or library.tsv (FragPipe software) or report-lib.tsv (DIA-NN software) file, an Excel file containing the score 1 (by fragment ratio), score 2 (by total intensity), normalized A value (Fragment_Ratio_norm) and B value (Total_Intensity_norm) of different targeted peptide segments can be output. The comprehensive score (combined_score) is an Excel file.

[0115] In order to ensure the accuracy and running speed of PRM targeted detection at the same time, it is generally sufficient to retain 2 to 3 target peptides for one target protein. In this embodiment, 3 target peptides are retained for each target protein, and 15 target peptides for the PRM targeted detection of these 5 target proteins are screened out at one time. In the present invention, if the number of identified targeted peptides corresponding to the target protein is ≤3, the identified targeted peptides are screened as target peptides; if the number of identified targeted peptides corresponding to the target protein is >3, screening can be performed based on the score of Fragment_Ratio_norm, Total_Intensity_norm or combined_score being greater than or equal to the threshold, such as setting the threshold to 0.5, or selecting the top 3 targeted peptides as target peptides according to the numerical value of Fragment_Ratio_norm, Total_Intensity_norm or combined_score.

[0116] In this example, the top three targeted peptides for each target protein were selected as target peptides for PRM targeted detection based on the ranking order of combined_score from high to low. The specific target peptide information screened out is shown below:

[0117] Table 8 List of target peptide information for PRM targeted detection screening

[0118]

[0119] If the target peptides are screened manually, the process will take about 1 hour and is highly subjective. However, according to this method, it only takes 5 minutes to screen the target peptides, reducing the interference of human subjectivity and making the screening results more objective and reliable.

[0120] Furthermore, this method was able to screen 15 suitable target peptides from the 24 identified target peptides for five target proteins. Existing methods, however, could not determine the feasibility of using these targeted peptides as target peptides for PRM-targeted detection. Compared to existing methods, the method provided by this invention demonstrates greater feasibility for PRM-targeted detection of target peptides, enabling single-step quantification of target proteins and improving the success rate of PRM-targeted detection.

[0121] Example 2 PRM mass spectrometry acquisition verifies the feasibility and stability of the target peptides screened by this method

[0122] To verify the feasibility and stability of the target peptides screened for PRM targeted detection in Example 1, this example used the PRM acquisition method derived from Skyline software based on the target peptide information screened in Table 8 for subsequent PRM targeted detection. After completing PRM mass spectrometry detection, the raw PRM detection data was imported into the Skyline software used to establish the acquisition method, and the retention time and intensity information of the fragment ion chromatogram peaks corresponding to the target peptides were extracted.

[0123] This example uses HeLa cell pellet samples as test samples. The specific experimental method is as follows:

[0124] 1. Add lysis buffer (1% SDC / 100 mM Tris-HCl, pH=8.5) to the sample, and ultrasonically disrupt it. After the ultrasonication, centrifuge the sample at 12000 g for 5 minutes.

[0125] 2. Take the protein supernatant and determine the protein concentration using the BCA assay. Take equal amounts of protein and use lysis buffer to bring all samples to the same volume.

[0126] 3. After adding TCEP and CAA, incubate at 37°C for 1 hour to complete the reduction and alkylation.

[0127] 4. Add 100mM Tris-HCl solution to dilute the Urea concentration to below 2M, add trypsin at a mass ratio of 1:50 between enzyme and protein, and incubate at 37°C with shaking overnight for enzymatic digestion.

[0128] 5. The next day, TFA was added to terminate the enzyme cleavage reaction. The sample was centrifuged at 12,000 g. The supernatant was taken and desalted using a homemade SDB desalting column. After vacuum drying, it was frozen at -20°C.

[0129] 6. Redissolve the enzymatic peptides in 0.1% formic acid aqueous solution and prepare for mass spectrometry detection.

[0130] The target proteins are GAPDH, ACADM, GNS, CSNK2A2 and TCEA1. The target peptides screened in Example 1 (see Table 8) are used to obtain the fragment ion chromatograms of GAPDH, ACADM, GNS, CSNK2A2 and TCEA1. Figures 2 to 6 As shown, samples 1, 2, and 3 are the results of PRM mass spectrometry detection repeated three times.

[0131] Depend on Figures 2 to 6 The PRM mass spectrometry detection results show that the targeted peptides screened by this method are the target peptides for PRM targeted detection, which can accurately detect the target peptides of proteins GAPDH, ACADM, GNS, CSNK2A2, and TCEA1. The success rate of PRM targeted detection of target proteins can reach 100%, and the repeatability of each PRM targeted detection is good, which can achieve accurate quantification of the target protein in one time.

[0132] The target protein is protein GNS, and the number of its candidate peptides is 3. All its corresponding target peptides are used as target peptides. Although the target peptide SDVLVEYQGEGR was not detected during PRM targeted detection, only one target peptide is needed to achieve quantitative detection of a target protein. When the top two candidate peptides screened are used as target peptides, PRM targeted detection can successfully detect the target peptide, and quantitative detection of the target protein can be achieved at one time.

[0133] The target protein is protein TCEA1, and the number of its candidate peptides is greater than 3. According to the ranking, the three target peptides were selected. Although the target peptide was NAAGALDLLK, the target peptide was not detected during the PRM targeted detection. However, when the top two candidate peptides were used as target peptides, the PRM targeted detection was able to successfully detect the target peptide, and the quantitative detection of the target protein could be achieved at one time.

[0134] Therefore, the target peptides screened by this method can achieve PRM targeted detection of the target protein in one go. For target proteins with ≥3 candidate peptides, the top two candidate peptides with the highest comprehensive scores are preferably used as target peptides.

[0135] The above description is merely a preferred embodiment of the present invention and does not constitute any form of limitation to the present invention. Although the present invention has been disclosed as above in terms of a preferred embodiment, it is not intended to limit the present invention. Any person skilled in the art can, without departing from the scope of the technical solution of the present invention, make some changes or modifications to equivalent embodiments using the technical contents disclosed above. However, any brief modifications, equivalent changes and modifications made to the above embodiments based on the technical essence of the present invention without departing from the content of the technical solution of the present invention are still within the scope of the technical solution of the present invention.

Claims

1. A method for rapid screening of PRM targeted detection target peptides based on a scoring algorithm, characterized in that: The following steps are involved: (1) Screening candidate peptides: Based on the search results of the original data of non-targeted mass spectrometry detection of biological samples, the targeted peptides corresponding to the identified target proteins are selected as candidate peptides; (2) Calculate the candidate peptide scores: calculate the score 1, score 2 and comprehensive score 3 of each candidate peptide in step (1); described described n is the total number of b and y fragment ions; described (3) Screening target peptides: Based on the comprehensive score 3 corresponding to each target protein calculated in step (2), the target peptides are screened according to the following method: If the number of candidate peptides is less than 3, at least one candidate peptide with the highest comprehensive score 3 will be selected as the target peptide; If the number of candidate peptides is ≥3, the candidate peptides are sorted from high to low according to the comprehensive score 3, and one or more peptides among the top three candidate peptides are selected as target peptides.

2. The method according to claim 1, wherein If the number of candidate peptides is ≥3, they are sorted from high to low according to the comprehensive score 3, and the top two peptides among the candidate peptides are used as the target peptides for PRM quantitative detection of the target protein.

3. The method according to claim 1 or 2, wherein: The comprehensive score 3 is calculated based on the normalized score 1 and score 2.

4. The method according to claim 3, wherein The normalization process is performed using a Min-Max normalization method, a Z-score normalization method, or a nonlinear normalization method.

5. The method according to claim 4, wherein The library search results are obtained by performing library search analysis on the original mass spectrum data using library search software, and the original mass spectrum data includes the original data collected in DDA and DIA modes.

6. The method according to claim 5, wherein The library search software includes one or more combinations of MaxQuant, FragPipe and DIA-NN.

7. A system for rapid screening of PRM targeted detection target peptides based on a scoring algorithm, characterized in that: It includes data acquisition module, judgment and calculation module and target peptide screening module; The data acquisition module is used to obtain the search results based on the original data of the proteome detected by non-targeted mass spectrometry and submit them to the judgment and calculation module. The search results include the search result files of qualitative or quantitative data of proteins, peptides, b fragment ions and y fragment ions; The judgment and calculation module is used to judge whether the target protein is identified and whether the identified target protein has a targeting peptide segment; If the target protein is identified and a targeting peptide exists, the targeting peptide is used as a candidate peptide, and the candidate peptide score 1, score 2, and comprehensive score 3 are calculated as follows; described described n is the total number of b and y fragment ions; described The calculated comprehensive score 3 is submitted to the target peptide screening module; The target peptide screening module screens target peptides according to the method of claim 1 and automatically outputs the screening results.

8. The system according to claim 7, wherein: If the number of candidate peptides is ≥3, they are sorted from high to low according to the comprehensive score 3, and the top two peptides among the candidate peptides are used as the target peptides for PRM quantitative detection of the target protein.

9. The system according to claim 8, wherein The comprehensive score 3 is calculated according to the normalized score 1 and score 2, and the normalization is performed using the Min-Max normalization method, the Z-score normalization method or the nonlinear normalization method.

10. The system according to claim 9, wherein: The raw data are the raw data collected in DDA mode, and the library search result files include proteinGroup.txt, peptides.txt and msms.txt files generated by MaxQuant software, and protein.tsv, peptide.tsv and library.tsv files output by FragPipe software; The raw data is the raw data collected in the DIA mode, and the search result files include report.pg_matrix.tsv, report.pr_matrix.tsv, and report-lib.tsv files generated by DIA-NN software and FragPipe software.

Citation Information

Patent Citations

  • Protein second-level mass spectrum identification method based on peak intensity recognition capability

    CN104076115A