A method for identifying a breakable cross-linked peptide segment, a method for evaluating and screening the identification result
Through high-sensitivity spectral preprocessing and multi-level FDR control methods, the accuracy and sensitivity of identification of breakable cross-linked peptides are optimized, solving the problems of insufficient sensitivity, slow speed and high error rate in existing methods, and achieving more efficient cross-linked peptide identification.
Patent Information
- Application Number
- CN202111437089.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-30
- Publication Date
- 2025-10-21
- Estimated Expiration
- 2041-11-30
AI Technical Summary
Existing methods for identifying cleavable cross-linked peptides have the disadvantages of insufficient sensitivity, slow search speed, failure to target the fragmentation patterns of cleavable cross-linked peptides, failure to optimize the fragment ion types for peptide spectrum matching scoring, and inability to effectively control the false discovery rate of cross-linking sites and protein interaction levels.
A highly sensitive spectrum preprocessing algorithm was designed to extract the mass of the two cross-linked peptides, optimize the fragment ion types for peptide spectrum matching and scoring, and use a multi-level FDR control method, including spectrum, parent ion, peptide, site and protein interaction levels, combined with support vector machine for re-scoring and FDR estimation.
The accuracy and sensitivity of the identification method are improved, the search speed is significantly accelerated, and the false discovery rate at multiple levels is effectively controlled, reducing the error rate.
Smart Images

Figure CN114049915B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of bioinformatics, specifically, to the field of computational proteomics, and more specifically, to the identification of cross-linked peptides in cross-linking mass spectrometry technology, namely, a method for identifying cleavable cross-linked peptides and a method for evaluating and screening the identification results. Background Art
[0002] Cross-linking mass spectrometry combines cross-linking with mass spectrometry, offering advantages in analyzing protein structure and studying protein interactions, including rapid speed, high throughput, low protein sample purity requirements, and the ability to capture weak interactions. As a key step in cross-linking mass spectrometry, cross-linked peptide identification plays a crucial role in data analysis and quality control, and plays a crucial role in the promotion and application of cross-linking mass spectrometry.
[0003] The identification method of cross-linked peptides is closely related to the cross-linker used. Depending on whether they are easily broken in the mass spectrometer, cross-linkers can be divided into non-breakable cross-linkers and breakable cross-linkers, and the corresponding cross-linked peptides are called non-breakable cross-linked peptides and breakable cross-linked peptides. A typical representative of breakable cross-linkers is DSSO (others include DSBU and PIR). Because it easily breaks in the mass spectrometer, the two cross-linked peptides are separated from each other and form complete peptide ion characteristic peaks in the secondary mass spectrometer. By analyzing the characteristic peaks in the spectrum, the masses of the two peptides can be obtained. This simplifies the identification of cross-linked peptides to linear peptide identification, significantly reducing the time complexity of cross-linked peptide identification. Therefore, the identification of breakable cross-linked peptides has gradually become a research hotspot in the field. For the identification of breakable cross-linked peptides, the core issue is how to find the characteristic peaks from the secondary mass spectrometer and infer the masses of the two cross-linked single peptides.
[0004] There are currently two main strategies for identifying cleavable cross-linked peptides. One is to extract the masses of the two cross-linked peptides through spectral preprocessing, and then simplify the identification of cross-linked peptides to linear peptide identification. Software using this strategy includes MeroX and XlinkX. The other is to not perform the above simplification and directly identify cleavable cross-linked peptides as non-cleavable cross-linked peptides. Software using this strategy includes xiSEARCH. The following introduces these two strategies respectively:
[0005] The first strategy involves identifying the peptide ion characteristic peaks formed by the cleavage from the spectrum, then calculating the masses of the two cross-linked peptides, and finally simplifying the identification of the cross-linked peptides to the identification of two linear peptides. This strategy requires a highly sensitive spectrum preprocessing algorithm. If the peptide ion characteristic peaks identified during preprocessing are incorrect, the calculated peptide masses will also be incorrect, leading to incorrect identification of the cross-linked peptides. In particular, when cleavable cross-linked peptides are fragmented by step-energy HCD, the peptide ion characteristic peaks will further fragment to generate a large number of fragment ion characteristic peaks, which seriously interfere with the identification of the peptide ion characteristic peaks. Therefore, identification methods using this strategy require the design of a highly sensitive preprocessing algorithm. Existing work has shown that MeroX and XlinkX based on this strategy have problems with low sensitivity and accuracy.
[0006] The second strategy avoids this simplification and directly identifies cleavable cross-linked peptides as if they were non-cleavable. This strategy does not require the identification of peptide ion signatures from the spectrum, so its sensitivity is less affected by preprocessing algorithms. However, this strategy uses an algorithm for identifying non-cleavable cross-linked peptides, which suffers from a quadratic search space problem and a much greater time complexity than the first strategy. This results in very slow searches when searching large databases. Previous studies have tested xiSEARCH, a platform based on this strategy, and found that its search speed is significantly slower than MeroX and XlinkX.
[0007] Furthermore, whether using the first or second strategy, existing methods for identifying cleavable cross-linked peptides do not optimize the fragment ion types used for peptide spectrum matching and scoring based on the fragmentation patterns of cleavable cross-linked peptides. Finally, most methods for identifying cleavable cross-linked peptides can only control the false discovery rate (FDR) at the spectral or peptide level, but cannot control the FDR at the cross-linking site or protein interaction level. If only the FDR at low levels, such as the spectral and peptide levels, is controlled, the error rate at higher levels, such as the cross-linking site and protein interaction level, will increase significantly. Summary of the Invention
[0008] Therefore, the purpose of the present invention is to overcome the above-mentioned defects of the prior art and provide a new method for identifying cleavable cross-linked peptides and a method for evaluating the identification results.
[0009] According to a first aspect of the present invention, a method for identifying cleavable cross-linked peptides is provided. For each secondary spectrum of a cleavable cross-linked peptide, the method comprises the following steps: S1, extracting a characteristic doublet in the spectrum whose mass difference is the fixed mass difference corresponding to the cross-linker, taking the mass corresponding to the characteristic doublet as the α peptide mass, and obtaining the β peptide mass complementary to the α peptide mass based on the parent ion mass corresponding to the cleavable cross-linked peptide to obtain a candidate mass pair consisting of the α peptide mass and the β peptide mass; wherein the parent ion mass is obtained by calibration of the primary spectrum of the cleavable cross-linked peptide; S2, based on the candidate mass pair obtained in step S1, using the α peptide mass and the β peptide mass in the candidate mass pair as the basis, screening multiple candidate peptides that are equal to the α peptide mass and the β peptide mass from a protein database, and forming paired candidate cross-linked peptides with all candidate peptides of the α peptide mass and all candidate peptides of the β peptide mass to form a candidate cross-linked peptide set, matching and scoring the cross-linked peptides in the candidate cross-linked peptide set according to preset fragment ions, and selecting the cross-linked peptide with the highest score.
[0010] Preferably, the step S1 includes: S11, extracting all non-redundant characteristic double peaks in the spectrum whose mass difference is the fixed mass difference corresponding to the cross-linker, and taking the mass corresponding to each group of non-redundant characteristic double peaks as an α peptide mass; S12, for each group of non-redundant characteristic double peaks, obtaining the complementary β peptide mass of the current α peptide by the parent ion mass corresponding to the cross-linked peptide to be identified and the mass difference between the current α peptide, and querying whether there is a corresponding characteristic double peak in the spectrum based on the mass of the complementary β peptide, and forming an initial mass pair with the complementary β peptide mass with the corresponding characteristic double peak and the current α peptide mass; S13, filtering all initial mass pairs with a pre-trained filtering model to eliminate unreliable mass pairs with probability scores lower than a preset probability threshold to obtain candidate mass pairs. In some embodiments of the present invention, step S11 includes: S111, enumerating the charge of each peak in the spectrum and obtaining the mass corresponding to each peak under each charge based on the mass-to-charge ratio in the spectrum, and screening out all characteristic double peaks whose mass difference is the fixed mass difference corresponding to the cross-linker; S112, merging characteristic double peaks with equal mass to obtain non-redundant characteristic double peaks. In some embodiments of the present invention, in step S111, the charge of each peak is enumerated as follows: when the parent ion charge is ≥5, 1-3 are enumerated, otherwise 1-2 are enumerated; wherein, each time the peak charge is enumerated, each peak is regarded as a monoisotopic peak under that charge, and characteristic double peaks whose mass difference is the fixed mass difference corresponding to the cross-linker are screened out.
[0011] Preferably, the filtering model is obtained by training the logistic regression model in the following manner: obtaining a spectral dataset and marking the correct mass pairs of each spectrum in the dataset; extracting the mass pairs of each spectrum in the spectral dataset and screening out the erroneous mass pairs using steps S11 and S12; forming a training dataset with spectra containing correct mass pairs and erroneous mass pairs; and training the logistic regression model using the training dataset until convergence.
[0012] In some embodiments of the present invention, the preset probability threshold is 0.25.
[0013] Preferably, the step S2 comprises: S21, performing theoretical enzyme digestion on all proteins in the protein database to obtain a peptide set, establishing a peptide mass index according to the mass of the peptides so as to put peptides of the same mass into a set corresponding to the same index; S22, searching the peptide mass indexes according to the corresponding α peptide mass and β peptide mass of the candidate mass, respectively, and obtaining all candidate peptides corresponding to the α peptide mass and β peptide mass respectively; S23, matching the candidate peptides of the α peptide mass and β peptide mass with the spectra respectively, and roughly scoring them and retaining at least the top five peptides as the quality of each peptide. Preliminary candidate peptides, wherein only backbone fragment ions are matched during the rough matching, and the backbone fragment ions include conventional b / y ions of 1+ and 2+ and b / y ions with long or short arms of cross-linkers; S24, all preliminary candidate peptides of α peptide mass and all preliminary candidate peptides of β peptide mass are respectively formed into pairs of candidate cross-linked peptides to form a candidate cross-linked peptide set, and all cross-linked peptides in the candidate cross-linked peptide set are matched with the spectrum for fine matching to select the cross-linked peptide with the highest score, wherein the fine matching score is to match and score the cross-linked peptides using preset fragment ions. In some embodiments of the present invention, the preset fragment ions include the following 31 fragment ions: y 1+ 、Y 1+x 、b 1+ 、B 1+x 、a 1+ 、Y 2+x 、YB 1+x 、Y 1+x +H2O、y 3+x 、yb 1+ 、y 2+x 、YA 1+x 、b 1+ -H2O、y 1+ -NH3、y 2+ 、b 1+ -NH3、y 1+ -H2O, B 1+x -NH3、A 1+x 、Y 1+x -NH3, b 2+x 、B1+x -H2O, Y 1+x -H2O, Y 2+x -H2O, ya 1+ 、b 3+x 、yb 2+x 、ya 2+x 、Y 2+x -NH3, KL(N) 1+x 、b 2+x -NH3.
[0014] According to the second aspect of the present invention, a method for evaluating and screening the identification results of the method described in the first aspect of the present invention is provided, characterized in that the method includes: P1, obtaining a primary spectrum of the cleavable cross-linked peptide segment, and calibrating the parent ion mass corresponding to the cleavable cross-linked peptide segment based on the primary spectrum; P2, obtaining a secondary spectrum of the cross-linked peptide segment, and using the method described in any one of claims 1-8 to identify the cleavable cross-linked peptide segment for all secondary spectra, and obtaining the cleavable cross-linked peptide segment corresponding to each secondary spectrum and its matching detailed score; P3, using a support vector machine to re-score the cleavable cross-linked peptide segment obtained by identification of all secondary spectra, and re-sorting all secondary spectra according to the re-score of the cleavable cross-linked peptide segment obtained by identification, and cyclically re-scoring based on the re-sorting result until the final re-scoring result is obtained in a preset round; P4, based on the re-scoring result of step P3, performing a preset level of FDR estimation on the identification result, and retaining the cleavable cross-linked peptide segment identification result whose FDR is less than or equal to the preset error rate screening threshold.
[0015] Preferably, the step P3 includes repeatedly performing the following steps according to a preset round: P31 sorting all secondary spectra from high to low according to the re-scoring, wherein in the initial round all secondary spectra are sorted from high to low according to the detailed scores of the identified cleavable cross-linked peptides; P32, using the FDR estimation formula to perform FDR estimation on the identification results of each secondary spectrum, and taking the positive library spectra with FDR estimation results less than or equal to the preset error rate division threshold as positive samples, and all negative library spectra as negative samples to train the support vector machine classifier until convergence; P32, using the trained support vector machine classifier to re-score each secondary spectrum.
[0016] In some embodiments of the present invention, the preset number of rounds is 5 rounds.
[0017] In some embodiments of the present invention, the preset error rate division threshold is 1%.
[0018] In some embodiments of the present invention, step P4 includes executing the following steps in a preset hierarchical cycle and screening out identification results that are less than or equal to a preset error rate screening threshold at the end of the cycle: P41, taking the current level as the starting level, and re-scoring the identification results of the secondary spectra according to the support vector machine classifier, calculating the FDR of all secondary spectra, and retaining the secondary spectrum identification results whose FDR is less than or equal to the preset error rate screening threshold; P42, merging the secondary spectrum identification results obtained in step P41 into the cross-linked parent ion level, and using the number of secondary spectra merged into the parent ion level as the low-level voting number of the parent ion; P43, scoring all cross-linked parent ions, and the score is the highest score among the support vector machine classifier scores corresponding to all secondary spectra merged into the parent ion; P44, grouping in the parent ion level using its bottom-level voting number, and performing FDR estimation in each group.
[0019] Compared with the prior art, the advantages of the present invention are: higher accuracy and sensitivity, and faster speed. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] The embodiments of the present invention are further described below with reference to the accompanying drawings, in which:
[0021] Figure 1 Schematic diagram of the process of identifying cleavable cross-linked peptides according to an embodiment of the present invention;
[0022] Figure 2 Schematic diagram of a method for evaluating and screening identification results of the method for identifying cleavable cross-linked peptides according to an embodiment of the present invention. DETAILED DESCRIPTION
[0023] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below through specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0024] As described in the background art, the main problems with existing methods for identifying cleavable cross-linked peptides are:
[0025] 1) The method based on the first strategy has the problem of insufficient sensitivity, and the method based on the second strategy has the problem of slow search speed;
[0026] 2) Existing methods for identifying cleavable cross-linked peptides do not address the fragmentation patterns of cleavable cross-linked peptides and optimize the fragment ion types used in peptide spectrum matching and scoring;
[0027] 3) Existing methods for identifying cleavable cross-linked peptides all have high error rates, and most methods do not control FDR at multiple levels. If only the low-level FDR is controlled, the error rate at higher levels will increase significantly.
[0028] The purpose of the present invention is to solve the problems of low accuracy, low sensitivity and slow speed of the existing methods for identifying cleavable cross-linked peptides, and proposes a new method for identifying cleavable cross-linked peptides (called pLink-DSSO). The pLink-DSSO of the present invention calculates the mass of two cross-linked peptides through spectrum preprocessing, and simplifies the identification of cross-linked peptides to the identification of two linear peptides. By analyzing the significance of the fragment ion matching of the spectrum, 31 fragments were selected for matching and scoring. In order to control the FDR at multiple levels, a multi-level FDR control method was designed, which can control the FDR of the spectrum, parent ion, peptide, site and protein interaction. And the test results show that the accuracy, sensitivity and speed of pLink-DSSO are significantly better than other methods.
[0029] In order to solve the above problems, the present invention proposes a method for identifying cleavable cross-linked peptides from the following perspectives: 1) a highly sensitive spectrum preprocessing algorithm is designed to extract the masses of the two cross-linked peptides from the spectrum, and then simplify the identification of cross-linked peptides to linear peptide identification, thereby achieving higher sensitivity and faster search speed; 2) the fragmentation rules of cleavable cross-linked peptides under step energy HCD are analyzed, and the fragment ion types used for peptide spectrum matching scoring are optimized; 3) a multi-level FDR control method is designed, which can control FDR at five levels: spectrum, parent ion, peptide, site and protein interaction.
[0030] According to one embodiment of the present invention, a method for identifying cleavable cross-linked peptides is provided. Figure 1 As shown, steps S1 and S2 are performed on all secondary spectra of the cleavable cross-linked peptide segments. Each step is described in detail below in conjunction with the examples.
[0031] In step S1, a characteristic double peak with a mass difference corresponding to the fixed mass difference of the cross-linker is extracted from the spectrum, the mass corresponding to the characteristic double peak is taken as the mass of the α peptide segment, and the mass of the β peptide segment complementary to the mass of the α peptide segment is obtained based on the mass of the parent ion corresponding to the cleavable cross-linked peptide segment to obtain a candidate mass pair consisting of the mass of the α peptide segment and the mass of the β peptide segment; wherein the parent ion mass is obtained by calibration of the primary spectrum of the cleavable cross-linked peptide segment.
[0032] When the inventors were conducting research on the identification of cleavable cross-linked peptides, they first analyzed the spectral patterns of cleavable cross-linked peptides under step energy HCD fragmentation and found that the common DSSO / DSBU cleavable cross-linker has two cleavage sites. Each cleavage site will form two peptide ion characteristic peaks. Therefore, the DSSO / DSBU cross-linked peptide will form a total of four peptide ion characteristic peaks. These four peptide ion characteristic peaks can be divided into two groups, belonging to the two cross-linked peptides; each group contains two peaks, and the mass difference △m between the two peaks is equal to the mass difference between the long arm and the broken arm of the cross-linker. The two peaks that meet the mass difference of △m are called a group of characteristic double peaks. For example, for DSSO, △m≈32Da, and for DSBU, △m≈26Da. For example, a group of cross-linked peptides may form α after cleavage. L Peptides and β S Peptides, and α S Peptide and β L Peptide, where L and S represent the long and short arms of the cross-linker, respectively. L Peptides and α S The poor quality of the peptides and β S Peptides and β L The mass difference of the peptide segments is the mass difference of the long and short arms of the cross-linker, α L Peptides and α S Peptide or β S Peptides and β LThe spectral peaks corresponding to the peptides are called characteristic double peaks. If these two groups of peptide ion characteristic double peaks can be identified from the spectrum, the masses of the two cross-linked peptides can be calculated, and the identification of cross-linked peptides can be simplified to the identification of two linear peptides. Therefore, before designing the identification algorithm, it is first necessary to analyze the spectrum rules, count the frequencies of the two groups of characteristic double peaks, and design the corresponding identification algorithm based on the rules. The inventors found statistically on multiple data sets that for the spectra of DSSO / DSBU-breakable cross-linked peptides under step energy HCD fragmentation, more than 50% of the spectra have two groups of characteristic double peaks; more than 95% of the spectra have at least one group of characteristic double peaks; and only less than 5% of the spectra do not have any group of characteristic double peaks. Based on the above statistical rules, the peptide ion characteristic double peaks can be identified from the spectrum according to △m, and then the identification of cross-linked peptides can be simplified to the identification of two linear peptides. If two sets of characteristic doublets can be identified, the masses of the two cross-linked peptides can be directly calculated and filtered using the parent ion mass as a constraint. If only one set of characteristic doublets can be identified, the mass of the other set can be obtained by subtracting the mass of that set from the parent ion mass. In both cases, cross-linked peptide identification can be simplified to the identification of two linear peptides, significantly reducing the time complexity of cross-linked peptide identification. The primary challenge in identifying cleavable cross-linked peptides based on the first strategy lies in identifying the two sets of peptide ion characteristic doublets with high sensitivity during spectral preprocessing. Due to the use of stepped energy HCD fragmentation, the four peptide ion characteristic peaks of the cleavable cross-linked peptides are further fragmented to generate a large number of fragment ion peaks. The mass difference between the fragment ion peaks carrying the long and broken arms of the cross-linker is also Δm, which significantly interferes with the identification of the peptide ion characteristic doublet. If the peptide ion characteristic doublet is misidentified, the calculated masses of the two peptides will also be incorrect, leading to incorrect cross-linked peptide identification. Therefore, the design of a highly sensitive spectral preprocessing algorithm is necessary.
[0033] According to one embodiment of the present invention, the present invention performs spectrum preprocessing to obtain mass pairs in the following manner:
[0034] Step S11, extracting non-redundant characteristic double peaks, which further includes: step S111, searching for all characteristic double peaks with a mass difference of Δm in the secondary spectrum, in order not to miss possible characteristic peaks, enumerating the charge of each peak (wherein, when the parent ion charge is ≥5, enumerate 1+~3+, otherwise enumerate 1+~2+, since the information displayed in the mass spectrum is the mass-to-charge ratio, the mass corresponding to the peak is obtained by enumerating the peak charge), and assuming that each peak is a monoisotope peak under the charge, if their mass difference is Δm, it is considered to be a set of legal characteristic double peaks. According to one embodiment of the present invention, in order not to extract the characteristic double peaks of isotope peaks, when enumerating the peaks, check whether there is a peak at the isotope peak position on the left side of the peak. If so, it is considered that the peak is the isotope peak of the peak on its left, and no subsequent operations are performed. Step S112, after obtaining all characteristic double peaks with a mass difference of Δm, merge the characteristic double peaks with equal mass to obtain a non-redundant characteristic double peak set.
[0035] Step S12, extracting candidate mass pairs: for each set of non-redundant characteristic double peaks, assuming that it comes from peptide segment α in the cross-linked dipeptide (then a set of characteristic peaks contains α L Peptides and α S If a peptide is present, the mass of the complementary peptide β can be calculated based on the precursor ion mass. From this, the m / z of the characteristic peaks of the long and short arms of the crosslinker carried by peptide β can be calculated. If peaks are present at both m / z of peptide β's characteristic peaks, a valid mass pair is considered to have been found.
[0036] Step S13, filtering candidate mass pairs: all candidate mass pairs extracted in the previous step are filtered using a logistic regression model. For each candidate mass pair, 12 dimensional features are extracted. These 12 dimensional features can be divided into three categories, namely, parent ion charge, matching feature bimodal number, matching feature bimodal intensity. The annotation data set is then used to train a regression model, and the quality pair is filtered using a regression model, retaining the quality pair of a logistic regression probability score higher than 0.25. This step is in order to only filter the obviously untrustworthy quality pairs, and to try not to mistakenly delete the correct quality pair, so a lower probability threshold of 0.25 is selected. In actual application, according to the needs of actual credibility, other probability thresholds can be set. If there is a candidate mass pair to pass through filtering, then find the quality pair of higher credibility, otherwise, enter the rescue program of step S14.
[0037] Step S14, rescue procedure: If the characteristic double peaks of the correct α and β peptide segments do not exist at the same time, the first three steps cannot find the correct α and β mass pairs. In order to recover the spectrum in which only the characteristic double peak of α (or β) exists, and the two characteristic peaks of β (or α) do not exist at the same time, the present invention designs a rescue procedure. In the rescue procedure, it is simply assumed that all non-redundant characteristic double peaks are from peptide segment α, and the mass of peptide segment β is calculated using the parent ion mass, and it does not require the presence of any characteristic peaks in peptide segment β. The rescue module can recover the incomplete characteristic double peaks of the two cross-linked peptide segments, but more mass pairs are extracted, so the rescue procedure is only executed when steps S11-S13 do not extract mass pairs with higher credibility.
[0038] The present invention realizes the spectrum preprocessing of the secondary spectrum through steps S11-S14, and can extract the mass of the two cross-linked peptides with high accuracy and sensitivity. The inventors found through statistical analysis that for the spectra of DSSO / DSBU cleavable cross-linked peptides under step energy HCD fragmentation, more than 50% of the spectra have two groups of characteristic double peaks; more than 95% of the spectra have at least one group of characteristic double peaks; and only less than 5% of the spectra do not have any group of characteristic double peaks. Therefore, the first three steps of the preprocessing algorithm, namely steps S11, S12, and S13, first preprocess the spectra with two groups of characteristic double peaks with high accuracy, and step S14 preprocesses the spectra with only one group of characteristic double peaks with high sensitivity. The preprocessing algorithm as a whole can achieve high accuracy and sensitivity.
[0039] In step S2, based on the candidate mass pairs obtained in step S1, multiple candidate peptides with the same mass as the α peptide and the same mass as the β peptide in the candidate mass pairs are screened from the protein database, and all candidate peptides with the mass of the α peptide and the mass of the β peptide are paired to form a candidate cross-linked peptide set. The cross-linked peptides in the candidate cross-linked peptide set are matched and scored according to the preset fragment ions, and the cross-linked peptide with the highest score is selected.
[0040] Peptide spectrum matching and scoring is the core link in the identification of cross-linked peptides. Currently, almost all identification methods for cleavable cross-linked peptides do not consider the fragment ion patterns of cleavable cross-linked peptides under step-energy HCD fragmentation when scoring peptide spectrum matching. They only use conventional backbone fragment ions for scoring, and do not even consider fragment ions carrying long arms and broken arms of the cross-linker. This will result in some high-frequency fragment ions not being correctly matched, which may increase the risk of mismatching and thus affect the accuracy of the identification results. Since the spectrum generated by cleavable cross-linked peptides under step-energy HCD fragmentation is relatively complex, it contains fragment ions generated by the fragmentation of four complete peptide ion peaks, which is equivalent to a mixed spectrum of four peptides. Therefore, it is necessary to analyze the matching significance of the fragment ions generated by cleavable cross-linked peptides and select fragment ions with higher significance to participate in the matching scoring to improve the performance of the scoring algorithm.
[0041] According to one embodiment of the present invention, the inventors analyzed the significance of fragment ion matching of DSSO / DSBU cleavable cross-linked peptides under step energy HCD fragmentation, and selected 31 fragment ions as ion types for fine scoring. These 31 fragment ions are: y1+, Y1+x, b1+, B1+x, a1+, Y2+x, YB1+x, Y1+x+H2O, y3+x, yb1+, y2+x, YA1+x, b1+-H2O, y1+-NH3, y2+, b1+-NH3, y1+-H2O, B1+x-NH3, A1+x, Y1+x-NH3, b2+x, B1+x-H2O, Y1+x-H2O, Y2+x-H2O, ya1+, b3+x, yb2+x, ya2+x, Y2+x-NH3, KL(N)1+x, b2+x-NH3. Using these 31 fragment ions for matching and scoring can maximize the difference in scores between correct and incorrect peptides, that is, it can distinguish correct and incorrect peptides to the greatest extent.
[0042] According to one embodiment of the present invention, the process of matching and scoring to obtain peptide segments includes the following steps:
[0043] Step S21: Generate a peptide quality index: Perform theoretical enzymatic digestion on all proteins in the forward and reverse protein databases, then add modifications to the resulting peptides to generate a set of all peptides. A peptide quality index is created based on the peptide mass. The peptide quality index is essentially a hash table, where the key is the peptide mass and the value is the set of all peptides with that mass. Using the peptide quality index, a set of peptides with a given mass can be quickly retrieved.
[0044] Step S22: Retrieve peptide mass indexes: Obtain candidate single peptides. Using the candidate masses of the α-peptide and β-peptide segments extracted in step S1, retrieve the peptide mass indexes to obtain all candidate α-peptide segments and β-peptide segments, i.e., all candidate peptide segments with masses equal to the candidate mass of the α-peptide segment and all candidate peptide segments with masses equal to the candidate mass of the β-peptide segment.
[0045] Step S23, coarse scoring of candidate single peptides: The candidate α- and β-peptides are roughly scored by matching the secondary spectra. Only backbone fragment ions, i.e., conventional 1+ and 2+ b / y ions and b / y ions with long or short crosslinker arms, are matched during the coarse scoring. Preferably, only the top five candidate peptides are retained for both α- and β-peptides.
[0046] Step S24: Combine candidate single peptides into cross-linked dipeptides, and perform detailed scoring on the cross-linked dipeptides: The top five candidate α peptide segments and the top five candidate β peptide segments are combined in pairs, and then matched with the spectrum for detailed scoring. 31 fragment ions were considered in the detailed scoring, including y1+, Y1+x, b1+, B1+x, a1+, Y2+x, YB1+x, Y1+x+H2O, y3+x, yb1+, y2+x, YA1+x, b1+-H2O, y1+-NH3, y2+, b1+-NH3, y1+-H2O, B1+x-NH3, A1+x, Y1+x-NH3, b2+x, B1+x-H2O, Y1+x-H2O, Y2+x-H2O, ya1+, b3+x, yb2+x, ya2+x, Y2+x-NH3, KL(N)1+x, and b2+x-NH3. The candidate cross-linked dipeptide with the highest detailed score for each spectrum was retained as the identification result of the spectrum.
[0047] As mentioned in the background art, most methods for identifying cleavable cross-linked peptides can only control the false discovery rate (FDR) at the spectrum or peptide level, but cannot control the FDR at the cross-linking site or protein interaction level. Therefore, how to control the FDR at the cross-linking site and protein interaction level is also a problem that needs to be solved urgently. Currently, most methods for identifying cleavable cross-linked peptides can only control the FDR at the spectrum or peptide level, because correct spectrum or peptide results tend to cluster on a few correct cross-linking sites or protein interaction results; while the aggregation effect of incorrect spectrum or peptide results is weaker. Therefore, if the FDR is only controlled at the spectrum or peptide level, the error rate at the cross-linking site or protein interaction level will increase significantly. Therefore, it is necessary to design an algorithm to directly control the FDR at the cross-linking site and protein interaction level.
[0048] According to one embodiment of the present invention, based on the method described in the aforementioned embodiment, the present invention further proposes an evaluation and screening method for multi-level FDR control, by controlling the FDR at multiple levels among the five levels of spectrum, parent ion, peptide, site and protein interaction, thereby reducing the error rate at the site and protein interaction levels.
[0049] According to one embodiment of the present invention, Figure 2 As shown, the evaluation and screening method of the present invention comprises the following steps:
[0050] P1. Obtain the primary spectrum of the cleavable cross-linked peptide segment and calibrate the parent ion mass corresponding to the cleavable cross-linked peptide segment based on the primary spectrum.
[0051] P2. Obtain the secondary spectra of the cross-linked peptides, and use the cleavable cross-linked peptide identification method of the present invention to identify the cleavable cross-linked peptides in all the secondary spectra to obtain the cleavable cross-linked peptides corresponding to each secondary spectrum and their matching detailed scores; this includes extracting the candidate masses mα and mβ of the α peptide and β peptide according to the secondary spectra, then searching the database for candidate peptides of the α peptide and β peptide according to the peptide mass index, and performing detailed scoring on the cross-linked peptides composed of the candidate peptides.
[0052] P3. Use support vector machine to re-score the cleavable cross-linked peptides obtained by identification of all secondary spectra, and re-sort all secondary spectra according to the re-scoring of the cleavable cross-linked peptides obtained by identification. Based on the re-sorting results, re-scoring is performed cyclically until the final re-scoring result is obtained after the preset rounds.
[0053] According to one embodiment of the present invention, step P3 further includes the following steps:
[0054] First, all cross-linked spectra were sorted according to the scores of the cross-linked peptides identified.
[0055] The FDR estimation formula for cross-linking is then used to estimate the FDR of the identification results. When performing a database search, a target protein database (referred to as the positive library) is searched, plus a bait library of the same size as the positive library. The bait library is usually obtained by reversing the sequence of the proteins in the positive library (referred to as the reverse library). Therefore, for the identification results of cross-linked peptides, there are three situations: both peptides are from the positive library, referred to as TT, and their number is recorded as #TT; both peptides are from the reverse library, referred to as DD, and their number is recorded as #DD; one peptide is from the positive library and the other peptide is from the reverse library, referred to as TD, and its number is recorded as #TD; the calculation method of the cross-linked FDR estimation formula is FDR = (#TD–#DD) / #TT.
[0056] Then, the positive library spectra with FDR less than or equal to the error rate threshold (according to one embodiment of the present invention, the error rate threshold is 1%) are taken as positive samples, and all negative library spectra are taken as negative samples to train a support vector machine (SVM) classifier.
[0057] Finally, the SVM classifier was used to re-score all cross-linking spectra;
[0058] All cross-linking spectra were re-ranked using SVM rescoring, and the above operation was repeated five times.
[0059] P4. Based on the re-scoring results of step P3, the identification results are estimated at a preset level of FDR, and the identification results of cleavable cross-linked peptides with FDR less than or equal to the preset error rate screening threshold are retained.
[0060] According to one embodiment of the present invention, the present invention controls L=5 levels of FDR, performs multi-level FDR estimation, sets the error rate screening threshold to a=0.01 (the threshold is set according to actual needs, and the smaller the error rate screening threshold, the higher the credibility), and defines the current level as l=1, wherein step P4 further includes the following steps:
[0061] P41, estimate the FDR of all cross-linked spectra at the spectrum level based on SVM rescoring, and only retain cross-linked spectra with FDR ≤ a;
[0062] P42, merge the cross-linked spectrum level results into the cross-linked parent ion level, and the number of spectra merged into a certain cross-linked parent ion is called the low-level voting number of the parent ion;
[0063] P43. Score all cross-linked precursor ions using the highest SVM score of all spectra merged into that precursor ion. At the precursor ion level, group them using the low-level vote count, and use the cross-link FDR formula to estimate the FDR of the cross-linked precursor ions within each group.
[0064] P44. Let l = l + 1, and repeat the above steps until l = L. At the Lth layer, report the cross-linking identification results with FDR ≤ a.
[0065] To validate the effectiveness of the present invention, we used the method of the present invention (pLink-DSSO), MeroX, XlinkX, and xiSEARCH on different databases. Table 1 shows the number of cross-linked peptide pairs identified and their estimated accuracy for each method on four datasets. Table 2 shows the search time (in minutes) for each software on the four datasets. xiSEARCH was found to be too slow to search the E. coli-DSSO-15N dataset, yielding no results for over a week.
[0066] Table 1
[0067] Dataset MeroX XlinkX xiSEARCH pLink-DSSO Lumos-DSSO-E.coli 228(79.8%) 623(35.2%) 299(93.6%) 320(94.7%) QEHFX-DSSO-E.coli 164(72.6%) 396(37.9%) 270(89.6%) 280(93.6%) QEHFX-DSBU-E.coli 417(68.3%) 370(37.3%) 330(90.3%) 355(96.6%) E.coli-DSSO-15N 850(99.3%) 1746(95.0%) - 2367(98.1%)
[0068] Table 2
[0069] Dataset MeroX XlinkX xiSEARCH pLink-DSSO Lumos-DSSO-E.coli 18.3 7.6 147.0 2.1 QEHFX-DSSO-E.coli 18.5 8.3 132.1 2.6 QEHFX-DSBU-E.coli 16.7 7.1 366.1 3.1 E.coli-DSSO-15N 120.0 324.0 More than 1 week 96.0
[0070] As shown in Tables 1 and 2, the present invention achieves higher accuracy, higher sensitivity, and faster speeds across multiple datasets compared to existing solutions. For synthetic peptide datasets and 15N-labeled datasets, pLink-DSSO achieves significantly higher accuracy and sensitivity than MeroX and XlinkX, and is 1-6 times faster. While its sensitivity is comparable to xiSEARCH, pLink-DSSO offers significant accuracy advantages and is 40 to 100 times faster.
[0071] It should be noted that although the above describes the various steps in a specific order, it does not mean that the steps must be performed in the above specific order. In fact, some of these steps can be executed concurrently or even in a different order as long as the required functions can be achieved.
[0072] The present invention may be a system, a method and / or a computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for causing a processor to implement various aspects of the present invention.
[0073] Computer-readable storage media can be a tangible device that holds and stores the instructions used by an instruction execution device. Computer-readable storage media can, for example, include, but are not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination thereof. More specific examples (non-exhaustive list) of computer-readable storage media include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanical encoding device, a punch card or a raised structure in a groove on which instructions are stored, for example, and any suitable combination thereof.
[0074] While various embodiments of the present invention have been described above, the above descriptions are intended to be illustrative, non-exhaustive, and not limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is selected to best explain the principles of the embodiments, their practical applications, or technological improvements in the marketplace, or to enable others skilled in the art to understand the embodiments disclosed herein.
Claims
1. A method for identifying cleavable cross-linked peptides, characterized in that: For each secondary spectrum of the cleavable cross-linked peptide segment, the method comprises the following steps: S1. Extract a characteristic double peak in the spectrum whose mass difference is the fixed mass difference corresponding to the cross-linking agent, take the mass corresponding to the characteristic double peak as the mass of the α peptide segment, obtain the mass of the β peptide segment complementary to the mass of the α peptide segment based on the mass of the parent ion corresponding to the cleavable cross-linked peptide segment, and obtain a candidate mass pair consisting of the mass of the α peptide segment and the mass of the β peptide segment; wherein the parent ion mass is obtained by primary spectrum calibration of the cleavable cross-linked peptide segment; wherein, step S1 comprises: S11, extracting all non-redundant characteristic double peaks in the spectrum whose mass difference is the fixed mass difference corresponding to the cross-linker, and taking the mass corresponding to each group of non-redundant characteristic double peaks as an α peptide mass; wherein, the step S11 includes: S111, enumerating the charge of each peak in the spectrum and obtaining the mass corresponding to each peak under each charge according to the mass-to-charge ratio in the spectrum, and screening out all characteristic double peaks whose mass difference is the fixed mass difference corresponding to the cross-linker; S112, merging characteristic double peaks of equal mass to obtain non-redundant characteristic double peaks; S12. For each set of non-redundant characteristic doublets, the mass of the complementary β peptide of the current α peptide is obtained by taking the difference between the mass of the parent ion corresponding to the cross-linked peptide to be identified and the mass of the current α peptide, and querying whether there is a corresponding characteristic doublet in the spectrum based on the mass of the complementary β peptide, and forming an initial mass pair with the mass of the complementary β peptide with the corresponding characteristic doublet and the mass of the current α peptide; S13. Filter all initial quality pairs using a pre-trained filtering model to eliminate untrustworthy quality pairs with probability scores lower than a preset probability threshold, and then obtain candidate quality pairs; wherein the filtering model is obtained by training a logistic regression model in the following manner: Obtain a spectral dataset and annotate each spectrum in the dataset with the correct mass pair. Steps S11 and S12 are used to extract the mass pairs of each spectrum in the spectrum data set and filter out the erroneous mass pairs therein; The training data set is composed of spectra containing correct mass pairs and incorrect mass pairs; The logistic regression model is trained to convergence using the training dataset; S2. Based on the candidate mass pairs obtained in step S1, multiple candidate peptides with the same mass as the α peptide and the same mass as the β peptide in the candidate mass pairs are screened from the protein database, and all candidate peptides with the mass of the α peptide and all candidate peptides with the mass of the β peptide are paired to form a candidate cross-linked peptide set. The cross-linked peptides in the candidate cross-linked peptide set are matched and scored according to preset fragment ions, and the cross-linked peptide with the highest score is selected.
2. The method according to claim 1, characterized in that In step S111, the charge of each peak is enumerated in the following manner: When the parent ion charge is ≥5, enumerate 1-3, otherwise enumerate 1-2; Each time the spectral peak charges are enumerated, each peak is regarded as a monoisotopic peak under the charge, and characteristic double peaks with a mass difference corresponding to the fixed mass difference of the cross-linker are screened out.
3. The method according to claim 1, characterized in that The preset probability threshold is 0.
25.
4. The method according to claim 1, wherein The step S2 comprises: S21. Perform theoretical enzyme digestion on all proteins in the protein database to obtain a peptide set, and establish a peptide mass index based on the mass of the peptides to place peptides of the same mass into a set corresponding to the same index; S22, searching peptide mass indexes according to the α peptide mass and β peptide mass corresponding to the candidate mass pair, and obtaining all candidate peptides corresponding to the α peptide mass and β peptide mass respectively; S23. For the candidate peptides of α-peptide mass and β-peptide mass, the spectrum is matched and roughly scored respectively, and at least the top five peptides are retained as preliminary candidate peptides of each mass. In the matching rough score, only the backbone fragment ions are matched, and the backbone fragment ions include 1+ and 2+ conventional b / y ions and b / y ions with long or broken arms of the cross-linker; S24. All preliminary candidate peptides of α peptide mass and all preliminary candidate peptides of β peptide mass are respectively grouped into pairs of candidate cross-linked peptides to form a candidate cross-linked peptide set, and all cross-linked peptides in the candidate cross-linked peptide set are matched with the spectrum for fine scoring to select the cross-linked peptide with the highest score, wherein the matching fine scoring is to use preset fragment ions to match and score the cross-linked peptides.
5. The method according to claim 4, characterized in that The preset fragment ions include the following 30 fragment ions: y1+, Y1+x, b1+, B1+x, a1+, Y2+x, YB1+x, Y1+x+H2O, y3+x, yb1+, y2+x, YA1+x, b1+-H2O, y1+-NH3, y2+, b1+-NH3, y1+-H2 O, B1+x-NH3, A1+x, Y1+x-NH3, b2+x, B1+x-H2O, Y1+x-H2O, Y2+x-H2O, ya1+, b3+x, yb2+x, ya2+x, Y2+x-NH3, b2+x-NH3.
6. A method for evaluating and screening the identification results of any one of claims 1 to 5, characterized in that: The method comprises: P1. Obtain the primary spectrum of the cleavable cross-linked peptide segment and calibrate the parent ion mass corresponding to the cleavable cross-linked peptide segment based on the primary spectrum; P2. Obtaining secondary spectra of cross-linked peptides, and identifying cleavable cross-linked peptides in all secondary spectra using the method described in any one of claims 1 to 5, obtaining the cleavable cross-linked peptides corresponding to each secondary spectrum and their matching scores; P3. Use support vector machine to re-score all cleavable cross-linked peptides identified in the secondary spectra, and re-rank all the secondary spectra according to the re-scored cleavable cross-linked peptides identified, and re-rank based on the re-ranking results until the final re-scoring result is obtained after a preset round; P4. Based on the re-scoring results of step P3, the identification results are estimated at a preset level of FDR, and the identification results of cleavable cross-linked peptides with FDR less than or equal to the preset error rate screening threshold are retained.
7. The method according to claim 6, characterized in that Step P3 includes repeatedly performing the following steps according to a preset round: P31 sorts all secondary spectra from high to low according to the re-scoring. In the initial round, all secondary spectra are sorted from high to low according to the detailed scores of the identified cleavable cross-linked peptides. P32, use the FDR estimation formula to perform FDR estimation on the identification results of each secondary spectrum, and use the positive library spectra with FDR estimation results less than or equal to the preset error rate classification threshold as positive samples, and all negative library spectra as negative samples to train the support vector machine classifier until convergence; P32. Use the trained support vector machine classifier to re-score each secondary spectrum.
8. The method according to claim 7, characterized in that The preset number of rounds is 5 rounds.
9. The method according to claim 7, characterized in that The preset error rate division threshold is 1%.
10. The method according to claim 6, characterized in that Step P4 includes executing the following steps in a preset hierarchical cycle and screening out identification results that are less than or equal to a preset error rate screening threshold at the end of the cycle: P41. Using the current level as the starting level, re-score the identification results of the secondary spectra according to the support vector machine classifier, calculate the FDR of all secondary spectra, and retain the secondary spectra identification results with FDR less than or equal to the preset error rate screening threshold; P42, merging the secondary spectrum identification results obtained in step P41 into the cross-linked parent ion level, and using the number of secondary spectra merged into the parent ion level as the low-level vote number of the parent ion; P43. Score all cross-linked parent ions, and the score is the highest score among the support vector machine classifier scores corresponding to all secondary spectra merged into the parent ion; P44. Group the precursor ions using their underlying vote counts and perform FDR estimation in each group.
11. A computer-readable storage medium, characterized in that A computer program is stored thereon, and the computer program can be executed by a processor to implement the steps of any one of the methods of claims 1-5 or 6-10.
12. An electronic device, characterized in that: include: one or more processors; A storage device for storing one or more programs, which, when executed by the one or more processors, enables the electronic device to implement the steps of the method as described in any one of claims 1-5 or 6-10.
Citation Information
Patent Citations
Crosslinking dipeptide rapid identification method
CN106033501A