A neural network-based method for quantifying endogenous cross-linked peptides
By using a neural network-based method, peak area method and similar peak shape matching, the problem of low abundance of endogenous crosslinking that is difficult to detect was solved, and accurate quantification of endogenous crosslinking was achieved, improving detection accuracy and reliability.
Patent Information
- Application Number
- CN202510244196.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-03
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2045-03-03
AI Technical Summary
Existing technologies lack accurate quantitative methods for endogenous crosslinking, resulting in low abundance and difficulty in detection.
A neural network-based approach was adopted to obtain crosslinking identification results through crosslinking search software. Peak area integration and neural network matching of similar peak shapes were used, and cosine similarity was combined to filter out erroneous results, thereby achieving accurate quantification of endogenous crosslinking.
It achieves accurate quantification of endogenous crosslinking, solves the problem of low abundance making it difficult to detect, and improves the detection accuracy and reliability of endogenous crosslinking.
Smart Images

Figure CN120164529B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of bioinformatics technology, specifically to a method for quantifying endogenous cross-linked peptides based on neural networks. Background Technology
[0002] Protein cross-linking is a process that links two or more protein molecules together through covalent bonds. Based on the principle of covalent bond formation, protein cross-linking can be divided into exogenous cross-linking and endogenous cross-linking. Exogenous cross-linking refers to cross-linking formed under external conditions such as chemical reagents (called cross-linking agents) or physical methods (such as radiation), while endogenous cross-linking refers to cross-linking formed by enzyme catalysis within a living organism.
[0003] Endogenous protein cross-links, including disulfide bonds, isopeptide bonds, tryptophan cross-links, tyrosine cross-links, formaldehyde cross-links, methylglyoxal (MG) cross-links, and NOS bridges, have been identified, all of which have been shown to play important roles in biological processes. Many endogenous cross-links are closely related to disease. For example, it has been reported that in advanced liver fibrosis, TG2 enzymes are activated and catalyzed by extracellular calcium ions, mediating the formation of KQ cross-links. These cross-links are extremely difficult to hydrolyze, which may be a key inhibitory factor in the recovery from advanced fibrosis. Therefore, quantitatively analyzing the levels of endogenous cross-links and measuring changes in their content in the tissues of healthy individuals and diseased patients can help to more accurately assess the degree of disease progression and can serve as a potential biomarker for disease screening and early diagnosis.
[0004] MaxQuant (which enables high peptide identification rates, individualized ppb-range mass accuracies, and proteome-wide protein quantification) uses a brute-force search method to identify and quantify exogenous crosslinks such as BS3 and DSS, enabling researchers to better study protein-protein interactions and the three-dimensional structure of proteins.
[0005] pQuant (pQuant Improves Quantitation by Keeping out Interfering Signals and Evaluating the Accuracy of Calculated Ratios.) can detect interference signals during quantification, identifying a pair of isotope chromatograms with minimal interference for each peptide: one for the peptide labeled with the light isotope and the other for the peptide labeled with the heavy isotope, resulting in better quantification performance compared to MaxQuant on light and heavy isotope labeled data.
[0006] The patent "Label-Free Protein Quantification Method Based on Mass Spectrometry" establishes a multiple mapping table of isotope distribution and retention time of theoretical peptide sets and extracted ion current chromatograms, and optimizes the quantification results of single peptides through intelligent algorithms.
[0007] Existing technologies can accurately quantify single peptides and exogenous crosslinks, but there is no quantitative technology designed specifically for endogenous crosslinks. Because endogenous crosslinks are present in low abundance and are difficult to detect, a completely new quantitative protocol is needed to achieve accurate quantification of endogenous crosslinks. Summary of the Invention
[0008] In view of this, the present invention provides a neural network-based method for quantifying endogenous cross-linked peptides, which can achieve accurate quantification of endogenous cross-links and solve the problem that endogenous cross-links are low in abundance and difficult to detect.
[0009] To solve the above-mentioned technical problems, the present invention is implemented as follows.
[0010] A neural network-based method for quantifying endogenous cross-linked peptides, comprising:
[0011] Step 1: Use crosslinking search software to search the crosslinking experimental data collected by the mass spectrometer to obtain K crosslinking identification results; for each crosslinking identification result k, generate I isotope mass and intensity, which is called theoretical isotope distribution.
[0012] Step 2: From the M sets of crosslinking experimental data obtained by the mass spectrometer M repeated experiments, select one set as the basic experimental data, read the first-order spectrum, and take the first-order spectrum corresponding to the crosslinking identification result k as the center. In the adjacent N first-order spectra, obtain the isotope intensity array corresponding to I isotope masses, which is called the actual isotope distribution.
[0013] Step 3: In the actual isotope distribution, extract the chromatographic peak corresponding to the mass of each single isotope, perform area integration on the chromatographic peak, and obtain the quantitative intensity value of the basic experimental data corresponding to the crosslinking identification result k.
[0014] Step 4: For the crosslinking identification result k, in the primary spectrum of the crosslinking experimental data of group M-1 (excluding the basic experimental data), find the chromatographic peak with the most similar peak shape and perform area integration to obtain the quantitative intensity value of the other crosslinking experimental data of group M-1 corresponding to the crosslinking identification result k.
[0015] Step 5: Perform the operations from Step 2 to Step 4 on all K crosslinking identification results obtained in Step 1 to complete the quantitative processing of endogenous crosslinked peptides.
[0016] Preferably, the method further includes: screening based on the chromatographic peaks of K crosslinking identification results, and performing quantitative intensity value calculation operations in steps three and four for the screened crosslinking identification results; the screening is: calculating the cosine similarity score between the chromatographic peak of each crosslinking identification result and the theoretical crosslinked peptide isotope distribution and the theoretical single peptide isotope distribution, and filtering out crosslinking identification results with similarity scores lower than those of single peptides.
[0017] Preferably, in step three, the chromatographic peak corresponding to the mass of each single isotope extracted is:
[0018] Using the retention time of the crosslinking identification results as the center, extract data before and after the retention time for a set time period;
[0019] Using the retention time of the crosslinking identification result as the center of the chromatographic peak, record the maximum intensity value I at the center point. max Initialize two pointers, left and right, and simultaneously expand them from the center outwards. As the pointers move, record the strength value I of the current pointer. current Determine if the left and right pointers encounter the intensity value I. current Less than I threshold =0.1×I max The situation where the left and right pointers each encounter I twice. current threshold Stop the expansion; use the positions where the left and right pointers stop as the start and end points of the chromatographic peak.
[0020] Preferably, in step four, the process of finding the most similar chromatographic peak in the primary spectra of the M-1 group of crosslinking experimental data (excluding the basic experimental data) is implemented using a neural network.
[0021] Each of the M-1 sets of crosslinking experimental data, excluding the basic experimental data, is defined as crosslinking experimental data S;
[0022] Take the first-level spectrum number T corresponding to the crosslinking identification result k. k In the first-order spectrum of the crosslinking experimental data S, the first-order spectrum T is used as the basis. k Centered on a central point, obtain isotope intensity arrays corresponding to I isotope masses from xN adjacent primary spectra, where x is the magnification factor; based on the extracted isotope intensity arrays, extract the chromatographic peak corresponding to each single isotope mass, denoted as Y. j j = 1, 2, ...; the chromatographic peaks extracted from the basic experimental data are denoted as T;
[0023] Extract chromatographic peaks T and Y j The peak characteristics are input into a neural network to obtain chromatographic peaks T and Y. j The peak shape similarity was determined; the chromatographic peak with the highest similarity was selected and its area was integrated as the quantitative intensity value corresponding to the current crosslinking experimental data S.
[0024] Preferably, the peak features include peak position, peak height, and peak width.
[0025] Preferably, the peak characteristics include chromatographic features.
[0026] Preferably, the chromatographic features include the retention time of the peak features and / or the environmental matrix of the peak features; the environmental matrix of the peak features consists of N consecutive primary chromatographic data centered on the primary chromatographic spectrum corresponding to the chromatographic peak.
[0027] Preferably, the neural network is a transformer neural network.
[0028] Preferably, each group of cross-linking experimental data in group M is used as a basic experimental data, and steps two to five are performed to obtain complete quantitative results of endogenous cross-linked peptides.
[0029] Preferably, I = 5.
[0030] Beneficial effects:
[0031] (1) This invention provides a novel protein quantification scheme designed for endogenous crosslinking. It is a label-free protein quantification scheme that uses peak area integration and similar peak matching based on neural networks to obtain accurate quantification of endogenous crosslinking, thus solving the problem that the abundance of endogenous crosslinking is low and it is not easy to detect.
[0032] (2) This invention uses a transformer-based neural network to find similar peaks between multiple runs, which solves the problem that the abundance of endogenous crosslinking is low and traditional algorithms have difficulty in determining the signal between multiple runs.
[0033] (3) In this invention, the peak shape is fitted into a mass array with the same dimensions, and the neural network is trained to learn the peak shape, peak height, peak value, and also learn the retention time, environmental matrix and other features at the chromatographic level, so as to find the most similar peak among multiple matches.
[0034] (4) In a preferred embodiment, the present invention provides an error identification result filtering step, which uses a similarity matching algorithm to filter out erroneous endogenous crosslinking results generated during identification by comparing the similarity between the actual isotope distribution and the theoretical isotope distribution, so as to determine the reliability of the identification result. Attached Figure Description
[0035] Figure 1 This is a flowchart of the neural network-based method for quantifying endogenous cross-linked peptides according to the present invention.
[0036] Figure 2 This is a schematic diagram illustrating the use of isotope distribution to filter out incorrect identification results.
[0037] Figure 3 This is a schematic diagram of obtaining a quantitative value through integration.
[0038] Figure 4 This is a schematic diagram illustrating the calculation of peak similarity in a neural network. Detailed Implementation
[0039] The present invention will now be described in detail with reference to the accompanying drawings and embodiments.
[0040] Figure 1 The flowchart of the neural network-based endogenous cross-linked peptide quantification method of the present invention is shown, as follows: Figure 1 As shown, the method consists of the following steps:
[0041] Step 1: Use crosslinking search software to search the crosslinking experimental data collected by the mass spectrometer to obtain the crosslinking identification results and obtain the theoretical isotope distribution.
[0042] In this step, crosslinking search software was used to perform a search, resulting in K crosslinking identification results. For each crosslinking identification result, I isotopic masses and corresponding intensities of a peptide were extracted, termed the theoretical isotopic distribution, denoted as Q. k The following steps are performed for each crosslinking identification result k.
[0043] In this embodiment, I is preferably 5. For each crosslinking identification result k, I isotopic masses and corresponding intensities are extracted and stored in a two-dimensional array Q. k [i][j], i = 1 to 1, j = 1, 2. The subscript k indicates the data corresponding to the k-th crosslinking identification result.
[0044] The cross-linking identification results include the cross-linked peptide sequence, retention time, and may also include the protein number corresponding to the peptide.
[0045] The theoretical isotopic mass and intensity distribution for generating crosslinking identification results are based on the principle of calculating the mass of amino acid residues and linkers in the peptide. Since both amino acids and linkers are composed of several elements, such as carbon (C), hydrogen (H), nitrogen (N), oxygen (O), and sulfur (S), the entire peptide can be represented as a chemical formula.
[0046] These elements have different isotopic forms; for example, carbon exists in C12 and C13. By calculating the probability of each isotopic combination, the theoretical distribution of each isotopic peak can be determined. For example, a molecule composed of three carbon atoms may have the following isotopic combinations: 3 C12, 2 C12 and 1 C13, 1 C12 and 2 C13, and 3 C13. Since the abundance of carbon isotopes C12 is 98.93% and C13 is 1.07%, the probability of a molecule composed of three carbon atoms having 3 C12 is 0.9893%. 3 =0.9682, the probability of having 2 C12s and 1 C13 is 3 × 0.9893. 2 ×0.0107=0.0314, and so on.
[0047] Step 2: Perform M repeated experiments using a mass spectrometer to obtain M sets of crosslinking experimental data. Select one set of crosslinking experimental data as the basic experimental data S0, and extract the actual isotope distribution.
[0048] This step involves reading the primary spectrum from the basic experimental data S0. Based on the retention time t of the crosslinking identification result k... k Extract to retain time t k Given a set of N adjacent first-order spectra centered on the i isotopic masses, for each of the i isotopic masses, extract the intensity value corresponding to the isotopic mass from these N adjacent first-order spectra, resulting in an intensity array S corresponding to the isotopic mass in the N adjacent first-order spectra. m,k [n], n = 1 to N. The intensity arrays corresponding to the five theoretical isotopic masses of a peptide are calculated in this way. These intensity arrays are then concatenated into an I*N two-dimensional matrix, denoted as S. m,k [i][n], i = 1 ~ I, n = 1 ~ N, can be used to obtain a chromatographic curve corresponding to a peptide segment.
[0049] S m,k The subscript m in [i][n] represents the m-th crosslinking experimental data, k represents the k-th crosslinking identification result, i represents the mass of the i-th isotope, and n represents the mass of the isotope extracted from the n-th primary spectrum. n = 1 to N, i = 1 to I, m = 1 to M, k = 1 to K. In one example, N = 40, I = 5, M = 4, and K = 100.
[0050] The two-dimensional matrix S obtained in this stepm,k [i][n] is called the actual isotope distribution S. m,k The experimental data from M sets can yield the distribution S of M actual isotopes. 1,k ~S M,k .
[0051] In this step, when extracting the intensity value corresponding to the isotopic mass from the primary spectrum, an allowable range is set that includes the isotopic mass to be extracted. When the isotopic mass in the primary spectrum falls within this allowable range, a matching isotopic mass is considered to have been found, and the intensity value corresponding to the matching isotopic mass is extracted. In practice, a matching isotopic mass can be considered to have been found when the first three decimal places are consistent.
[0052] The obtained intensity array is standardized using Z-score, transforming the data into a distribution with a mean of 0 and a standard deviation of 1. The calculation formula is as follows:
[0053]
[0054] Where z is the original data, z' is the normalized data, μ is the mean of the data, and σ is the standard deviation of the data.
[0055] Step 3: For each individual isotope mass intensity array S in the actual isotope distribution m,k [i][n], extract chromatographic peaks.
[0056] This step involves using the intensity threshold method to determine the chromatographic peaks within the intensity array of single isotope mass quantities, i.e., determining the start and end points of the chromatographic peaks. Specifically, this includes the following steps:
[0057] Step S301: Using the retention time of the crosslinking identification result as the center, extract the chromatographic data 2n minutes before and after the retention time, where n is usually 1 or 2.
[0058] Step S302: Using the retention time of the crosslinking identification result as the center of the chromatographic peak, initialize the left and right pointers, and expand them outwards from the center point. Record the maximum intensity value I at the center point. max The left pointer moves to the left, and the right pointer moves to the right, recording the current pointer strength value I. current Determine if the left and right pointers encounter the intensity value I. current Less than I threshold =0.1×I max The situation where the left and right pointers each encounter I twice. current threshold Stop the expansion. Use the positions where the left and right pointers stop as the start and end points of the chromatographic peak.
[0059] In practice, other methods can also be used to obtain the chromatographic peak window. The extracted chromatographic peak is a part of the intensity array, S m,k [i][n], n = 1 to N', N' <N。
[0060] Step 4: Identification results of poor crosslinking in filtration.
[0061] For TG enzyme-mediated cross-linked peptides, cross-linking identification results may overlap with single-peptide identification results. These latter results are unreliable and therefore need to be filtered out. For the K cross-linking identification results obtained in step one, chromatographic peaks are extracted. This step calculates the cosine similarity score between the chromatographic peak of each cross-linking identification result and the theoretical isotopic distributions of the cross-linked peptide and the theoretical isotopic distributions of the single peptide. Based on these scores, the cross-linking identification results are reordered, and those with similarity scores lower than those of the single peptide are filtered out. For example... Figure 2 As shown. In practice, this filtration operation can be performed after obtaining the chromatographic peaks, and does not need to be limited to before calculating the area in step five.
[0062] The specific steps involved in this process include:
[0063] Step S401: Obtain the actual chromatographic curve of the peptide using mass spectrometry data, denoted as S. m,k [i][n], i=1~I, n=1~N;
[0064] Step S402: Based on the known amino acid sequence and elemental composition, calculate the theoretical isotopic mass distribution of the single peptide and cross-link corresponding to the same identification result (see Step 1), and denote them as D respectively. m,k [i][n], i=1~I, n=1~N; J m,k [i][n], i=1~I, n=1~N;
[0065] Step S403: Using a cosine similarity algorithm, the theoretical isotopic mass distributions of the single peptide and cross-linked peptides are compared with their actual distributions to determine their similarity. Taking a single peptide as an example, firstly, the two-dimensional arrays of the theoretical isotopic mass D and the actual chromatographic curve S of the single peptide are flattened into a one-dimensional long vector to obtain s and d. Then, the dot product and norm of the one-dimensional long vector are calculated. The similarity between the two one-dimensional long vectors is calculated using the cosine similarity formula, which can be used as the similarity between the theoretical isotopic mass distribution and the actual distribution of the single peptide.
[0066] s = flatten(S)
[0067] d = flatten(D)
[0068]
[0069] Step S404: According to the set similarity threshold, filter out cross-linking identification results with similarity scores lower than those of single peptides, i.e., filter out SIMs. SJ <SIM SD The part.
[0070] Step 5: Integrate the chromatographic peaks of the remaining crosslinking identification results after filtration using the peak area method to obtain the quantitative intensity value Q of the crosslinking identification result k. DL k,m .
[0071] In this step, numerical integration methods (such as the trapezoidal method or Simpson's method) are used to integrate the portion of the chromatographic peak between its starting and ending points, i.e., to calculate the peak area. Figure 3 As shown.
[0072]
[0073] Where I(t) represents the intensity value of the chromatographic curve at time t, t start and t end These are the starting and ending points of the chromatographic peak, respectively.
[0074] Step 6: For the crosslinking identification result k, find the chromatographic peak with the most similar peak shape in the primary spectrum corresponding to the experimental data of group M-1 outside the basic experimental data, integrate it, and use it as the quantitative intensity value of other experimental data.
[0075] If the basic experimental data is the first set of experimental data, and there are a total of 4 sets of experimental data, then for the crosslinking identification result k, step five obtains the quantitative intensity value Q based on the first set of experimental data. DL k,1 Then, based on the crosslinking identification result k obtained from the other three sets of experimental data, the quantitative intensity value is Q. DL k,2 Q DL k,3 Q DL k,4 .
[0076] Each group of crosslinking experimental data in the M-1 group (excluding the basic experimental data) is defined as crosslinking experimental data S. The method for finding the chromatographic peak with the most similar peak shape in crosslinking experimental data S is described in detail below. See [link to relevant documentation]. Figure 4 :
[0077] Step 601: Obtain the first-order spectrum number T of the crosslinking identification result k. k In the first-order spectrum of the crosslinking experimental data S, the spectrum T is used as the reference. kCentered on a central point, extract N×x primary spectra, where x is the magnification factor. Because there will be some error during identification across multiple runs, it is necessary to extract chromatographic curves over a larger area. Following the method in step two, use xN primary spectra to extract chromatographic curves, and then extract the chromatographic peak corresponding to the mass of each single isotope, denoted as Y. j , j = 1, 2, ...; the number of chromatographic peaks in each group of crosslinking experimental data S is variable. The chromatographic peaks extracted from the basic experimental data are denoted as T.
[0078] Step 602: Analyze the chromatographic peak T in the crosslinking identification result k and the chromatographic peak Y in other experimental data. j The peak characteristics of the chromatographic peaks were extracted.
[0079] The peak features include peak shape features and chromatographic layer features. Peak shape features can include peak position (i.e., the peak index in the array), peak height, and peak width (i.e., the width occupied by the peak in the array); the purpose of extracting these features is to further compare the peak shape similarity of different arrays. In addition to the features of the peak itself, it is also necessary to extract chromatographic layer features, which can include the retention time of the chromatographic peak at the chromatographic layer, and can also include the environmental matrix of the chromatographic peak at the chromatographic layer. The environmental matrix is the chromatographic peak Y j This corresponds to N consecutive primary chromatograms centered on the primary chromatogram. Selecting N chromatograms ensures consistency with the N chromatograms of chromatographic peak T.
[0080] Step 603: Train a neural network to compare chromatographic peak T with chromatographic peak Y j The peak shape similarity is used to obtain similar chromatographic peaks. In this embodiment, the neural network is preferably a transformer neural network.
[0081] The peak characteristics of chromatographic peak T and the chromatographic peak Y compared with it. j The peak features are input into a transformer neural network, which outputs a similarity score. All chromatographic peaks Y1 to Y2 in the crosslinking experimental data S are then used. P (P represents the total number of chromatographic peaks in the crosslinking experiment data S) are compared with chromatographic peak T to obtain their respective scores. The chromatographic peak with the highest score is the chromatographic peak similar to chromatographic peak T.
[0082] The transformer neural network first uses a position encoding function (sine, cosine, or a single-layer neural network with learned parameters) to encode T and T. yThe features are encoded into vectors, and these vectors are concatenated to form the input vector. A self-attention mechanism is then used to calculate the correlation of the input at each position across all positions. A multi-head mechanism allows the model to perform these calculations independently in different subspaces, and the results are then concatenated. In this step, the input vector is first mapped to query, key, and value matrices. The formulas for calculating the query, key, and value are as follows:
[0083] Q = XW Q
[0084] K = XW K
[0085] V = XW V
[0086] Where X is the input vector, W Q W K and W V These are the weight matrices for the query, key, and value, respectively. Then, the dot product of the query and key is calculated and scaled.
[0087]
[0088] Where, d k Given the dimension of the key, the softmax function is used to normalize the attention weights.
[0089] Through the above steps, the transformer can capture the dependencies between different positions in the input sequence. During training, the dependencies are compared between chromatographic peaks T and Y1~Y2. P By observing the peak shape characteristics, the neural network can learn how to determine the similarity between different arrays. Finally, by evaluating the similarity score output by the neural network, the similarity between chromatographic peak T and Y1~Y2 can be determined. P Which peak shape is more similar?
[0090] like Figure 4 As shown, among the three candidate chromatographic peaks A, B, and C, peak C has the highest similarity.
[0091] Step 604: For similar chromatographic peaks, perform area integration, and use the integration result as the quantitative intensity value corresponding to the crosslinking experimental data S. For each crosslinking experimental data point, similar chromatographic peaks are extracted and area integrated to obtain the quantitative intensity value.
[0092] This completes the process of obtaining the quantitative intensity values of M sets of crosslinking experimental data for the crosslinking identification result k.
[0093] Step 7: Perform steps 2 through 6 on all K crosslinking identification results obtained in step 1 to complete the quantitative processing of endogenous crosslinked peptides.
[0094] Repeat steps two through seven above, selecting another set of cross-linking experimental data as the base experimental data to obtain a set of quantitative intensity values. Each of the M sets of cross-linking experimental data serves as one base experimental data set, yielding M sets of quantitative intensity values. These M sets of quantitative intensity values can be cross-validated. For example, using cross-linking data A as the base for quantification, the quantitative value of peptide Pep1 in A is 100, and in B it is 10; then using cross-linking data B as the base for quantification, the quantitative value of peptide Pep1 in A is 98, and in B it is 14. The quantitative intensity values obtained from these two sets of data can be normalized to obtain more accurate results.
[0095] The specific embodiments described above only illustrate the design principles of the present invention. The shapes and names of the components in this description may differ and are not limited. Therefore, those skilled in the art can modify or make equivalent substitutions to the technical solutions described in the foregoing embodiments; and these modifications and substitutions do not depart from the inventive spirit and technical solutions of the present invention, and should all fall within the protection scope of the present invention.
Claims
1. A method for quantifying endogenous cross-linked peptides based on neural networks, characterized in that, include: Step 1: Use crosslinking search software to search the crosslinking experimental data collected by the mass spectrometer to obtain K crosslinking identification results; For each crosslinking identification result k, I isotopic mass and intensity are generated, which is called the theoretical isotopic distribution; Step 2: From the M sets of crosslinking experimental data obtained by the mass spectrometer M repeated experiments, select one set as the basic experimental data, read the first-order spectrum, and take the first-order spectrum corresponding to the crosslinking identification result k as the center. In the adjacent N first-order spectra, obtain the isotope intensity array corresponding to I isotope masses, which is called the actual isotope distribution. Step 3: In the actual isotope distribution, extract the chromatographic peak corresponding to the mass of each single isotope, perform area integration on the chromatographic peak, and obtain the quantitative intensity value of the basic experimental data corresponding to the crosslinking identification result k. Step 4: For the crosslinking identification result k, in the primary spectrum of the crosslinking experimental data of group M-1 (excluding the basic experimental data), find the chromatographic peak with the most similar peak shape and perform area integration to obtain the quantitative intensity value of the other crosslinking experimental data of group M-1 corresponding to the crosslinking identification result k. Step 5: Perform the operations from Step 2 to Step 4 on all K crosslinking identification results obtained in Step 1 to complete the quantitative processing of endogenous crosslinked peptides.
2. The method as described in claim 1, characterized in that, The method further includes: screening based on the chromatographic peaks of K crosslinking identification results, and performing quantitative intensity value calculation operations in steps three and four for the screened crosslinking identification results; the screening is: calculating the cosine similarity score between the chromatographic peak of each crosslinking identification result and the theoretical crosslinked peptide isotope distribution and the theoretical single peptide isotope distribution, and filtering out crosslinking identification results with similarity scores lower than those of single peptides.
3. The method as described in claim 1, characterized in that, In step three, the chromatographic peak corresponding to the mass of each single isotope is extracted as follows: Using the retention time of the crosslinking identification results as the center, extract data before and after the retention time for a set time period; Using the retention time of the crosslinking identification result as the center of the chromatographic peak, record the maximum intensity value I at the center point. max Initialize two pointers, left and right, and simultaneously expand them from the center outwards. As the pointers move, record the strength value I of the current pointer. current Determine if the left and right pointers encounter the intensity value I. current Less than I threshold =0.1×I max The situation where the left and right pointers each encounter I twice. current threshold Stop the expansion; use the positions where the left and right pointers stop as the start and end points of the chromatographic peak. 4. The method as described in claim 1, characterized in that, In step four, the process of finding the most similar chromatographic peak in the primary spectra of the M-1 group of crosslinking experimental data (excluding the basic experimental data) is implemented using a neural network. Each of the M-1 sets of crosslinking experimental data, excluding the basic experimental data, is defined as crosslinking experimental data S; Take the first-level spectrum number T corresponding to the crosslinking identification result k. k In the first-order spectrum of the crosslinking experimental data S, the first-order spectrum T is used as the basis. k Centered on a central point, obtain isotope intensity arrays corresponding to I isotope masses from xN adjacent primary spectra, where x is the magnification factor; based on the extracted isotope intensity arrays, extract the chromatographic peak corresponding to each single isotope mass, denoted as Y. j j = 1, 2, ...; the chromatographic peaks extracted from the basic experimental data are denoted as T; Extract chromatographic peaks T and Y j The peak characteristics are input into a neural network to obtain chromatographic peaks T and Y. j The peak shape similarity was determined; the chromatographic peak with the highest similarity was selected and its area was integrated as the quantitative intensity value corresponding to the current crosslinking experimental data S.
5. The method as described in claim 4, characterized in that, The peak characteristics include peak position, peak height, and peak width.
6. The method as described in claim 5, characterized in that, The peak characteristics include chromatographic features.
7. The method as described in claim 6, characterized in that, The chromatographic features include the retention time of the peak features and / or the environmental matrix of the peak features; the environmental matrix of the peak features consists of N consecutive primary chromatographic data centered on the primary chromatographic spectrum corresponding to the chromatographic peak.
8. The method as described in claim 4, characterized in that, The neural network used is a transformer neural network.
9. The method as described in claim 1, characterized in that, Each group of cross-linking experimental data in group M is used as a basic experimental data, and steps two through five are performed to obtain complete quantitative results of endogenous cross-linked peptides.
10. The method as described in claim 1, characterized in that, I=5。
Citation Information
Patent Citations
Cross-linked peptide fragment qualitative and quantitative analysis method
CN111208299A
Mass spectrometry-cleavable cross-linking agents to facilitate structural analysis of proteins and protein complexes, and method of using same
US20130144541A1