Mirror peptide segment mass spectrum for identification method
By directly utilizing the information differences between mass spectra digested by mirror enzymes to calculate the matching score, mirror spectrum pairs are identified. This solves the problems of low efficiency and high cost caused by relying on prediction sequences in existing technologies, and achieves efficient and accurate identification of mirror spectrum pairs and complete spectrum coverage.
Patent Information
- Application Number
- CN202311757260.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-19
- Publication Date
- 2025-10-21
- Estimated Expiration
- 2043-12-19
AI Technical Summary
Existing mirror spectrum pair identification methods rely on the accuracy of peptide prediction sequences, resulting in low recognition efficiency and high cost, and are unable to effectively achieve complete coverage of fragment ion peaks and ion species identification in some spectra.
By obtaining two sets of mass spectra from paired mirror enzyme digestion, a set of spectrum pairs is generated. Candidate mirror spectrum pairs are screened, and the matching score is directly calculated using the information difference between the spectra to identify mirror spectrum pairs, thus avoiding reliance on predicted sequence results.
It improves the recognition efficiency and accuracy of mirror spectrum pairs, saves time and costs, and effectively achieves complete coverage of fragment ion peaks and identification of ion types in the spectrum.
Smart Images

Figure CN117746993B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of biotechnology, and in particular to a method for identifying mirror peptide mass spectra pairs. Background Art
[0002] In mass spectrometry-based proteomics research, de novo peptide sequencing has garnered widespread attention for its flexibility and efficiency. However, the accuracy of de novo sequencing is limited by the quality of mass spectrometry data, particularly low ion coverage. Mirror enzyme technology is a common approach to addressing this low ion coverage. Mirror enzyme cleavage involves the use of two enzymes to cleave the same amino acid at its C-terminus and N-terminus, respectively, to produce peptides that begin and end with the same amino acid and share the same intermediate sequence. These peptides exhibit mirror-image characteristics in the mass spectrometry, and are therefore called mirror peptides. The corresponding two spectra are called mirror spectrum pairs.
[0003] For the common separate spectra, the key is to identify which are mirror spectrum pairs. The current mainstream mirror spectrum pair identification method is to predict the sequence of the peptides and then match the spectrum pairs whose sequences are mirror images of each other. The accuracy of this type of method depends largely on the accuracy of the predicted sequencing results, and the time cost consumed by sequencing is also not negligible. Therefore, it is necessary to study the mirror spectrum pair identification algorithm that does not depend on the sequence. At the same time, the mirror spectrum pair identification results can help achieve complete coverage of the fragment ion peaks of some spectra and identify the ion species, thereby developing a more accurate peptide de novo sequencing method. Summary of the Invention
[0004] The present invention provides a mirror peptide mass spectrum pair identification method, which can accurately identify mirror peptide mass spectrum pairs without pre-sequencing, thereby improving identification efficiency.
[0005] To this end, the present invention provides the following technical solutions:
[0006] A method for identifying mirror peptide mass spectra pairs, the method comprising:
[0007] Obtain two sets of mass spectra obtained by pairwise mirror enzyme digestion to generate a spectrum pair set;
[0008] screening candidate mirror image spectrum pairs from the spectrum pair set to obtain a candidate mirror image spectrum pair list;
[0009] Directly using the information difference between the spectrogram pairs to calculate the matching score of each candidate mirror spectrogram pair in the candidate mirror spectrogram pair list;
[0010] The mirror image spectrum pair is determined according to the matching scores of the candidate mirror image spectrum pairs.
[0011] Optionally, the paired mirror enzymes include: any pair of mirror enzymes, or any combination of any pair of mirror enzymes.
[0012] Optionally, screening candidate mirror image spectrum pairs from the spectrum pair set to obtain a candidate mirror image spectrum pair list includes:
[0013] For each spectrum pair in the spectrum pair set, calculating the parent ion mass difference of the spectrum pair;
[0014] If the parent ion mass difference is within the error range set by the theoretical mass difference, the pair of spectra is taken as a candidate mirror image pair of spectra.
[0015] Optionally, the method further comprises: before calculating the matching score of each candidate mirror spectrum pair in the candidate mirror spectrum pair list, preprocessing the candidate mirror spectrum pairs.
[0016] Optionally, the preprocessing includes any one or more of the following: removing isotope peaks and converting them into single charges, removing dehydration and ammonia loss peaks, removing iminium ion peaks, removing noise peaks, normalizing spectral peak intensities, and generating complementary ion peaks.
[0017] Optionally, the information difference between the spectral pairs includes: difference in mass-to-charge ratio of fragment ion spectrum peaks and difference in parent ion mass;
[0018] The directly using the information difference between the spectrogram pairs to calculate the matching score of each candidate mirror spectrogram pair in the candidate mirror spectrogram pair list includes:
[0019] The matching scores of candidate mirror spectral pairs were calculated using the difference in mass-to-charge ratios of the fragment ion spectra peaks and the difference in the parent ion mass.
[0020] Optionally, calculating the matching score of the candidate mirror image spectrum pair using the difference in mass-to-charge ratio of the fragment ion spectrum peaks and the difference in the parent ion mass includes:
[0021] Evenly dividing the mass-to-charge ratio difference of the fragment ion spectrum peak into a plurality of small intervals, counting the sum of the fragment ion intensities and the number of fragment ion pairs falling within each small interval, and multiplying the sum of the fragment ion intensities and the number of fragment ion pairs in each interval as the mass difference statistic of the interval;
[0022] Calculating a statistical score for the candidate mirror image spectrum pair based on the mass difference statistic and the parent ion mass difference;
[0023] The matching score of the candidate mirror image spectrogram pair is determined according to the statistical score of the candidate mirror image spectrogram pair.
[0024] Optionally, calculating a statistical score of the candidate mirror image spectrum pair based on the mass difference statistic and the parent ion mass difference includes:
[0025] Determining one or more intervals where the theoretical mass difference of the fragment ions lies based on the mass difference of the parent ions;
[0026] Calculating the maximum value of the mass difference statistics in the interval where the theoretical mass difference of the fragment ions is located, and taking the ranking value of the maximum value in the mass difference statistics of all intervals as the statistical score of the candidate mirror spectrum pair; or
[0027] The minimum e-value of the interval where the theoretical mass difference of the fragment ion is located is calculated from the distribution of the mass difference statistics in all intervals as the statistical score of the candidate mirror image spectrum pair.
[0028] Optionally, the information difference between the spectral pairs further includes: retention time difference;
[0029] The method further comprises:
[0030] Before screening candidate mirror spectrum pairs from the spectrum pair set, for each spectrum pair in the spectrum pair set, calculating the retention time difference of the spectrum pair;
[0031] Remove candidate mirror spectrum pairs in the candidate mirror spectrum pair list whose retention time difference between the spectrum pairs is greater than a set time threshold.
[0032] Optionally, the information difference includes: difference in mass-to-charge ratio of fragment ion spectrum peaks, difference in parent ion mass, and difference in retention time;
[0033] The directly using the information difference between the spectrogram pairs to calculate the matching score of each candidate mirror spectrogram pair in the candidate mirror spectrogram pair list includes:
[0034] The fragment ion mass and intensity, parent ion mass, fragment ion spectrum peak mass-to-charge ratio difference, parent ion mass difference, and retention time difference of the candidate mirror image spectrum pair are input into a matching recognition model constructed manually or obtained based on machine learning training to obtain the matching score of the candidate mirror image spectrum pair.
[0035] Optionally, determining the mirror spectrum pair according to the matching scores of the candidate mirror spectrum pairs includes:
[0036] The candidate mirror spectrum pairs with matching scores greater than the set matching threshold are screened as mirror spectrum pairs.
[0037] A computer-readable storage medium stores a computer program, which, when executed by a processor, executes the steps of the aforementioned method for identifying mirror peptide mass spectra pairs.
[0038] The mirror peptide mass spectrum pair identification method provided by the present invention obtains two sets of mass spectra obtained by paired mirror enzyme digestion, generates a spectrum pair set, and then screens the spectrum pair set to select a list of candidate mirror spectrum pairs; then calculates a matching score for each candidate mirror spectrum pair in the candidate mirror spectrum pair list; and determines the mirror spectrum pair based on the matching scores of the candidate mirror spectrum pairs. Because the calculation of the candidate mirror spectrum pair matching scores is based only on information provided by the spectra themselves and does not rely on the prediction results of other software, it not only improves recognition efficiency and saves time and costs, but also effectively improves the accuracy of the recognition results. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 The present invention provides a flow chart of the method for identifying mirror peptide mass spectra pairs. DETAILED DESCRIPTION
[0040] To more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be derived from these drawings without inventive effort.
[0041] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. The embodiments cannot be described one by one here, but the embodiments of the present invention are not limited to the following embodiments.
[0042] The current mainstream mirror peptide mass spectrum pair identification method requires first predicting the peptides and then matching the spectrum pairs whose sequences are mirror images of each other. The accuracy of this method depends largely on the accuracy of the predicted sequencing results. The present invention provides a mirror peptide mass spectrum pair identification method that does not rely on the predicted sequencing results of the spectra and can accurately identify mirror spectrum pairs.
[0043] like Figure 1 FIG. 1 is a flow chart of the mirror peptide mass spectrum pair identification method provided by the present invention, comprising the following steps:
[0044] Step 101: Obtain two sets of mass spectra obtained by pairwise mirror enzyme digestion to generate a spectrum pair set.
[0045] A mass spectrum is a graph used to describe the mass-to-charge ratio and intensity information of the detected ions, and can also be simply referred to as a "spectrum".
[0046] The spectrum pair set refers to the set of all spectrum pairs consisting of spectra in two spectrum sets.
[0047] For the convenience of description, the two spectrum sets are referred to as spectrum set 1 and spectrum set 2, and the spectrum pair set represents the set of all spectrum pairs consisting of the spectrum in 1 and the spectrum in 2.
[0048] Step 102 : Screen candidate mirror spectrum pairs from the spectrum pair set to obtain a candidate mirror spectrum pair list.
[0049] The basic principle of screening candidate mirror image pairs is that the intermediate sequences of the peptides corresponding to each type of mirror image pair are the same, so the theoretical mass difference of their parent ions is determined.
[0050] To this end, in an embodiment of the present invention, each pair of spectrum pairs in the spectrum pair set, that is, each spectrum in spectrum set 1 and all spectra in spectrum set 2, can be judged in turn, and the parent ion mass difference of the spectrum pair can be calculated; if the parent ion mass difference is within the error range set by the theoretical mass difference, the pair of spectrum pairs is used as a candidate mirror image spectrum pair, thereby obtaining a candidate mirror image spectrum pair list.
[0051] Theoretical parent ion mass differences can have multiple values, depending on the mirror enzyme used in the sample (i.e., spectrum). For example, if trypsin and lysarginase cleave at the N-terminus and C-terminus of amino acids K and R, respectively, the theoretical parent ion mass differences for different mirror spectrum pairs are 0, ±28, ±128, and ±156, respectively. The error range, e, is user-specified; for example, e = 10 can be set.
[0052] Specifically, the spectrum pair currently being judged is spectrum t and spectrum l, and their parent ion masses are mass t and mass l , the theoretical mass difference of the parent ion type to be judged is Mdalton, the set error tolerance range is eppm, and the calculations are:
[0053] D1=[mass t (1-e·10 -6 )]-[mass l (1+e·10 -6 )]
[0054] D2=[mass t (1+e·10 -6 )]-[mass l (1-e·10 -6 )]
[0055] If M∈[D1,D2], then spectrum t and spectrum l are candidate mirror spectrum pairs of this type; otherwise they are not.
[0056] It should be noted that the type refers to the type of the spectrum pair. For example, in a non-limiting embodiment, there are 7 types of spectrum pairs in the spectrum pair set (see Table 1 below for details). For each pair of spectrum pairs, these 7 types must be judged in turn. When the type of the spectrum pair is judged to be type A, type A corresponds to a parent ion theoretical mass difference M(A). After calculating the above D1 and D2, if M(A)∈[D1,D2], then the spectrum pair is determined to be type A. Otherwise, the next type is judged until all 7 types are judged.
[0057] Step 103 , directly using the information differences between the spectrogram pairs to calculate the matching score of each candidate mirror spectrogram pair in the candidate mirror spectrogram pair list.
[0058] The information differences may include: fragment ion spectrum peak mass-to-charge ratio differences, parent ion mass differences, and further retention time differences. The fragment ion spectrum peak mass-to-charge ratio differences may be obtained by subtracting the fragment ion spectrum peak mass-to-charge ratios of the candidate mirror image spectrum pairs, the parent ion mass differences may be obtained by subtracting the parent ion masses of the candidate mirror image spectrum pairs, and the retention time differences may be obtained by subtracting the retention times of the candidate mirror image spectrum pairs.
[0059] The match calculation for a spectral pair is based solely on the information provided by the spectra themselves and does not rely on prediction results from other software. The basic principle is that the theoretical mass difference of the fragment ions for each type of mirror-image spectral pair is fixed, and the error range can be set as needed. Multiple theoretical mass differences between fragment ions are possible. Specifically, the match score is designed based on the peaks of each candidate mirror-image spectral pair, combined with the mass and intensity information of the fragment ions, as well as the mass and retention time of the parent ion.
[0060] In an embodiment of the present invention, the calculation of the matching degree can be implemented in a variety of ways, as long as it can reflect the confidence that the current spectrum pair is considered to be a mirror spectrum pair. For example, statistical scoring, machine learning scoring, etc. can be used, and this embodiment of the present invention does not limit this.
[0061] For example, the matching score of the candidate mirror image spectrum pair can be calculated using the mass-to-charge ratio difference of the fragment ion spectrum peak and the parent ion mass difference. The steps of calculating the matching score using a statistical scoring method may include: evenly dividing the mass-to-charge ratio difference of the fragment ion spectrum peak into multiple small intervals, counting the sum of the fragment ion intensities and the number of fragment ion pairs falling within each small interval, and multiplying the sum of the fragment ion intensities of each interval by the number of fragment ion pairs as the mass difference statistic of the interval; calculating the statistical score of the candidate mirror image spectrum pair based on the mass difference statistic and the parent ion mass difference; and determining the matching score of the candidate mirror image spectrum pair based on the statistical score of the candidate mirror image spectrum pair.
[0062] Among them, the step of calculating the statistical score may include: determining one or more intervals where the theoretical mass difference of the fragment ions is located based on the mass difference of the parent ions; calculating the maximum value of the mass difference statistic in the interval where the theoretical mass difference of the fragment ions is located, and taking the ranking value of the maximum value in the mass difference statistics of all intervals as the statistical score of the candidate mirror spectrum pair; or calculating the minimum e-value value of the interval where the theoretical mass difference of the fragment ions is located based on the distribution of the mass difference statistics of all intervals as the statistical score of the candidate mirror spectrum pair. The method of determining the matching score based on the statistical score may include: setting the statistical score to x, taking the matching score to 1 / x; or taking sigmoid(x); or taking e -x .
[0063] Calculating a match score using a machine learning scoring method can include inputting the precursor ion mass, fragment ion mass and intensity, fragment ion spectrum peak mass-to-charge ratio difference, and precursor ion mass difference of the candidate mirror image spectrum pair into a match recognition model trained based on machine learning to obtain a match score for the candidate mirror image spectrum pair. Furthermore, the input parameters of the match recognition model may further include retention time difference.
[0064] The matching degree recognition model can be constructed manually or obtained based on machine learning training, and the embodiment of the present invention does not limit its construction and training process.
[0065] Step 104: Determine a mirror image spectrum pair based on the matching scores of the candidate mirror image spectrum pairs.
[0066] Without loss of generality, it is assumed that the higher the matching score of a spectrum pair, the more reasonable it is to believe that it is a mirror spectrum pair. Therefore, the candidate mirror spectrum pairs with matching scores greater than a set threshold can be regarded as mirror spectrum pairs.
[0067] In another non-limiting embodiment, before calculating the matching score for each candidate mirror spectrum pair in the candidate mirror spectrum pair list in step 103, the candidate mirror spectrum pairs may be preprocessed. The preprocessing may include, but is not limited to, any one or more of the following: removing isotope peaks and converting them to single charges, removing dehydration and ammonia peaks, removing iminium ion peaks, removing noise peaks, normalizing peak intensities, and generating complementary ion peaks.
[0068] The method for removing noise peaks may include: global denoising and / or local denoising. Global denoising refers to filtering peaks with an intensity greater than a set value, or filtering a set number of the strongest peaks; local denoising refers to filtering peaks with an intensity greater than a set value, or filtering a set number of the strongest peaks within a window of xdalton.
[0069] The method of normalizing the spectrum peak intensity may include: maximum-minimum normalization, or zero-mean normalization, etc.
[0070] In another non-limiting embodiment, before screening candidate mirror spectrum pairs from the spectrum pair set in step 202, the retention time difference of each spectrum pair in the spectrum pair set is calculated. Candidate mirror spectrum pairs in the candidate mirror spectrum pair list where the retention time difference between the spectrum pairs exceeds a set time threshold are removed. This can further improve the efficiency and accuracy of subsequent determination of mirror spectrum pairs.
[0071] The mirror peptide mass spectrum pair identification method provided by the present invention obtains two sets of mass spectra obtained by paired mirror enzyme digestion to generate a spectrum pair set; first, a list of candidate mirror spectrum pairs is screened from the spectrum pair set; then, a matching score is calculated for each candidate mirror spectrum pair in the candidate mirror spectrum pair list; and the mirror spectrum pair is determined based on the matching scores of the candidate mirror spectrum pairs. Because the calculation of the candidate mirror spectrum pair matching scores is based solely on information provided by the spectra themselves and does not rely on the prediction results of other software, it not only improves recognition efficiency and saves time and costs, but also effectively increases the accuracy of the recognition results.
[0072] Compared to existing methods for identifying mirror-image peptide mass spectra pairs, the present invention's approach does not rely on predicted sequencing results from other software. A similar algorithm, pMerge, requires sequencing both sets of mass spectra using pNovo software before identifying mirror-image pairs. After obtaining the sequencing results, it then matches spectra pairs whose sequences are mirror images of each other. A matching score is calculated based on the sequencing results and spectra information, and the results for mirror-image pairs are filtered.
[0073] The mirror peptide mass spectrum pair identification method provided by the solution of the present invention is applicable to any pair of mirror enzymes, or a combination of paired mirror enzymes.
[0074] Taking the single pair of mirror enzymes Trypsin and LysargiNase as an example, suppose there is a protein sequence ...ABCKPEPTIDERDEF..., which can be hydrolyzed by Trypsin and LysargiNase, respectively, to produce mirror peptides PEPTIDER and KPEPTIDE, with a precursor mass difference of approximately 28 (RK). Under ideal conditions, if the peptide is completely hydrolyzed into fragment ions, then:
[0075] For Trypsin peptide, the b ions are P, PE, PEP, PEPT, PEPTI, PEPTID, PEPTIDE; the y ions are R, ER, DER, IDER, TIDER, PTIDER, EPTIDER;
[0076] For the LysargiNase peptide, the b ions are K, KP, KPE, KPEP, KPEPT, KPEPTI, KPEPTID; the y ions are E, DE, IDE, TIDE, PTIDE, EPTIDE, PEPTIDE, and the mass differences of its fragment ions tend to concentrate at -128 (-K, b ion) and 156 (R, y ion).
[0077] For other protein sequences, there are similar rules for the mass differences between precursor ions and fragment ions of peptides generated by Trypsin and LysargiNase digestion. Table 1 below shows the theoretical mass differences (approximate values) between precursor ions and fragment ions of different Trypsin and LysargiNase peptides.
[0078] Table 1
[0079]
[0080]
[0081] Suppose there is a trypsin spectrum t and a lysargiNase spectrum l.
[0082] First, judge whether the retention time difference meets the error. Secondly, calculate the mass difference of its precursor ions, and calculate the upper and lower limits of the mass difference D1 and D2 respectively according to the aforementioned definition. If D1 < 0 < D2, then this spectrum pair is a candidate spectrum pair of type A; if D1 < 28 < D2, then this spectrum pair is a candidate spectrum pair of type B; and so on. In this way, a list of candidate mirror spectrum pairs of all 7 types can be obtained.
[0083] For each pair of candidate spectrum pairs, a matching score needs to be calculated to characterize the credibility of this spectrum pair being a pair of mirror spectra. Suppose the spectrum pair t and l is a candidate spectrum pair of type B, and the mass differences of its fragment ions should theoretically concentrate at -128 and 156.
[0084] The following gives a specific implementation method of scoring:
[0085] First, take the square root of the peak intensity to make the overall scoring more average, and then assign weights to the intensities of different mass-to-charge ratios where x is the mass-to-charge ratio.
[0086] Next, calculate the mass differences of fragment ions pairwise, take an appropriate interval step size, divide the mass difference [-200, 200] into N sub-intervals, map each mass difference to an interval index, and merge the spectral peaks with the same index to obtain the weighted intensity sum sum_intensity i and the number of spectral peaks count i, calculate the quality difference statistics of each index s=[s1,…,s i ,…,s N ], where s i =sum_intensity i *count i Assuming that the non-negative statistic s obeys the maximum extreme value distribution, then in the high value region of s, s and F(s) = Pr(S>s) are in a log-linear relationship logF(s) = a1log s+a2+2. The parameters a1 and a2 are unknown, but can be estimated from the high value distribution of s. Once the estimated value of F(s) is obtained, the e-value can be calculated as escore(s) = N·F(s), thereby obtaining the statistical score of the spectrum pair, escore. Further, taking the inverse of escore can obtain the matching score final_score of the spectrum pair.
[0087] From the definition of final_score, we know that the higher the score, the more reliable the result. Therefore, we set a threshold α and filter all spectrum pairs with a matching score final_score ≥ α to obtain the identified mirror spectrum pairs.
[0088] The term “plurality” used in the embodiments of the present invention refers to two or more than two.
[0089] It should be noted that the terms "including" and "having" and any variations thereof in the specification and claims of the present invention and the above-mentioned drawings are intended to cover non-exclusive inclusions. For example, a process, method, system, product or apparatus comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or are inherent to these processes, methods, products or apparatuses.
[0090] The various embodiments in this specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments.
[0091] Accordingly, an embodiment of the present invention further provides a computer-readable storage medium having a computer program stored thereon, which is executed by a processor when the processor is running. Figure 1 The steps of the mirror peptide mass spectrum pair identification method.
[0092] The embodiments of the present invention are described in detail above. Specific implementation methods are used herein to illustrate the present invention. The description of the above embodiments is only used to help understand the method and system of the present invention. They are only embodiments of a part of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work should fall within the scope of protection of the present invention, and the content of this specification should not be understood as limiting the present invention. Therefore, any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A method for identifying mirror peptide mass spectra, characterized in that: The method comprises: Obtain two sets of mass spectra obtained by pairwise mirror enzyme digestion to generate a spectrum pair set; screening candidate mirror image spectrum pairs from the spectrum pair set to obtain a candidate mirror image spectrum pair list; Directly using the information difference between the spectrogram pairs to calculate the matching score of each candidate mirror spectrogram pair in the candidate mirror spectrogram pair list; Determining a mirror image spectrum pair according to the matching scores of the candidate mirror image spectrum pairs; The information differences between the spectral pairs include: differences in the mass-to-charge ratios of the fragment ion spectrum peaks and differences in the mass of the parent ions; The directly using the information difference between the spectrogram pairs to calculate the matching score of each candidate mirror spectrogram pair in the candidate mirror spectrogram pair list includes: The matching scores of candidate mirror image pairs are calculated using the difference in the mass-to-charge ratio of the fragment ion spectrum peaks and the difference in the parent ion mass; The calculation of the matching score of the candidate mirror image pair using the difference in mass-to-charge ratio of the fragment ion spectrum peaks and the difference in the mass of the parent ion comprises: Evenly dividing the mass-to-charge ratio difference of the fragment ion spectrum peak into a plurality of small intervals, counting the sum of the fragment ion intensities and the number of fragment ion pairs falling within each small interval, and multiplying the sum of the fragment ion intensities and the number of fragment ion pairs in each interval as the mass difference statistic of the interval; Calculating a statistical score for the candidate mirror image spectrum pair based on the mass difference statistic and the parent ion mass difference; Determining a matching score of the candidate mirror image spectrogram pair according to the statistical score of the candidate mirror image spectrogram pair; The information differences between the spectral pairs also include: retention time differences; Before screening candidate mirror spectrum pairs from the spectrum pair set, for each spectrum pair in the spectrum pair set, calculating the retention time difference of the spectrum pair; Remove candidate mirror spectrum pairs in the candidate mirror spectrum pair list whose retention time difference between the spectrum pairs is greater than a set time threshold.
2. The method for identifying mirror peptide mass spectra according to claim 1, wherein: The paired mirror enzymes include any pair of mirror enzymes or any combination of any pair of mirror enzymes.
3. The method for identifying mirror peptide mass spectra according to claim 1, wherein: The screening of candidate mirror image spectrum pairs from the spectrum pair set to obtain a list of candidate mirror image spectrum pairs comprises: For each spectrum pair in the spectrum pair set, calculating the parent ion mass difference of the spectrum pair; If the parent ion mass difference is within the error range set by the theoretical mass difference, the pair of spectra is taken as a candidate mirror image pair of spectra.
4. The method for identifying mirror peptide mass spectra according to claim 1, wherein: The method further comprises: Before calculating the matching score of each candidate mirror spectrum pair in the candidate mirror spectrum pair list, the candidate mirror spectrum pairs are preprocessed.
5. The method for identifying mirror peptide mass spectra pairs according to claim 4, wherein: The pretreatment includes any one or more of the following: removing isotope peaks and converting them into single charges, removing water loss and ammonia loss peaks, removing iminium ion peaks, removing noise peaks, normalizing spectral peak intensities, and generating complementary ion peaks.
6. The method for identifying mirror peptide mass spectra pairs according to claim 5, characterized in that: The calculating of the statistical score of the candidate mirror image spectrum pair according to the mass difference statistic and the parent ion mass difference comprises: Determining one or more intervals where the theoretical mass difference of the fragment ions lies based on the mass difference of the parent ions; Calculating the maximum value of the mass difference statistics in the interval where the theoretical mass difference of the fragment ions is located, and taking the ranking value of the maximum value in the mass difference statistics of all intervals as the statistical score of the candidate mirror spectrum pair; or The minimum e-value of the interval where the theoretical mass difference of the fragment ion is located is calculated from the distribution of the mass difference statistics in all intervals as the statistical score of the candidate mirror image spectrum pair.
7. The method for identifying mirror peptide mass spectra pairs according to claim 1, wherein: The information differences include: differences in fragment ion spectrum peak mass-to-charge ratios, parent ion mass differences, and retention time differences; The directly using the information difference between the spectrogram pairs to calculate the matching score of each candidate mirror spectrogram pair in the candidate mirror spectrogram pair list includes: The fragment ion mass and intensity, parent ion mass, fragment ion spectrum peak mass-to-charge ratio difference, parent ion mass difference, and retention time difference of the candidate mirror image spectrum pair are input into a matching recognition model constructed manually or obtained based on machine learning training to obtain the matching score of the candidate mirror image spectrum pair.
8. The method for identifying mirror peptide mass spectra pairs according to claim 1, wherein: Determining the mirror image spectrum pair according to the matching score of the candidate mirror image spectrum pair includes: The candidate mirror spectrum pairs with matching scores greater than the set matching threshold are screened as mirror spectrum pairs.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method for identifying mirror peptide mass spectrum pairs according to any one of claims 1 to 8 are performed.
Citation Information
Patent Citations
De novo sequencing method
CN107729719A