Neural network-based endogenous cross-linked peptide quantification method
Through the endogenous crosslinking peptide quantification method based on neural network, the peak area method and neural network match similar peak types are used to solve the problem of low endogenous crosslink abundance, and the accurate quantification of endogenous crosslinking is achieved, which improves the accuracy and reliability of signal matching.
Patent Information
- Application Number
- CN202510244196.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-03
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2045-03-03
AI Technical Summary
The prior art is difficult to achieve accurate quantities of endogenous crosslinking, especially because of its low abundance and is not easily detected.
The endogenous crosslinked peptide quantification method is adopted based on neural networks, and the mass spectrometer data is searched through crosslinking search software to generate theoretical isotope distribution, and the peak area method and the neural network match similar peak types are used to achieve accurate quantities of endogenous crosslinking.
The problem of low abundance of endogenous crosslinks was successfully solved, and the accurate amount of endogenous crosslinks was achieved, which improved the accuracy and reliability of signal matching.
Smart Images

Figure CN120164529A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of bioinformatics, and particularly relates to a method for quantifying endogenous cross-linked peptides based on neural networks. Background Art
[0002] Protein cross-linking is a process of covalently linking two or more protein molecules together. According to the formation principle of covalent bonds, protein cross-linking can be divided into exogenous cross-linking and endogenous cross-linking. Exogenous cross-linking refers to the cross-linking formed by external conditions such as chemical reagents (called cross-linking agents) or physical methods (such as radiation), and endogenous cross-linking refers to the cross-linking formed by the catalysis of enzymes in living organisms.
[0003] Currently, it has been found that endogenous protein cross-links include disulfide bonds, isopeptide bonds, tryptophan cross-links, tyrosine cross-links, formaldehyde cross-links, methylglyoxal (MG) cross-links, NOS bridges, etc., and they have all been proven to have important effects in biological processes. Many endogenous cross-links are closely related to diseases. For example, it has been reported that in advanced liver fibrosis, the TG2 enzyme will be activated by extracellular calcium ions and catalyze its activity, thereby mediating the formation of KQ cross-links, and this cross-link is extremely difficult to hydrolyze, which may be a key inhibitory factor for the recovery of advanced fibrosis. Therefore, quantitatively analyzing the content level of endogenous cross-links and measuring the content changes of endogenous cross-links in the tissues of healthy individuals and disease patients helps to more accurately evaluate the degree of disease development and can be used as a potential biomarker for disease screening and early diagnosis.
[0004] MaxQuant (MaxQuant enables high peptide identification rates, individualized ppb-range mass accuracies and proteome-wide protein quantification) uses a brute-force search method to achieve the identification and quantification of exogenous cross-links such as BS3 and DSS, enabling researchers to better study protein-protein interactions and the three-dimensional structure of proteins.
[0005] pQuant (pQuant Improves Quantitation by Keeping out Interfering Signals and Evaluating the Accuracy of Calculated Ratios.) can detect interfering signals during quantification, identify a pair of isotope chromatograms with the least interference for each peptide: one for the light-isotope-labeled peptide and the other for the heavy-isotope-labeled peptide, and has better quantification performance than MaxQuant on light and heavy-labeled data.
[0006] The patent "Label-Free Protein Quantification Method Based on Mass Spectrometry Spectra" established a multiple mapping relationship table of the theoretical peptide set, the isotope distribution, and the retention time of the extracted ion current chromatogram, and optimized the quantification results of single peptides through intelligent algorithms.
[0007] Existing technologies can already perform accurate quantification for single peptides and exogenous cross-links, but there is no quantification technology designed for endogenous cross-links. Since the abundance of endogenous cross-links is low and it is not easy to detect, a new set of quantification schemes needs to be designed to achieve accurate quantification of endogenous cross-links. Summary of the Invention
[0008] In view of this, the present invention provides a method for quantifying endogenous cross-linked peptides based on a neural network, which is used to achieve accurate quantification of endogenous cross-links and solve the problem that the abundance of endogenous cross-links is low and it is not easy to detect.
[0009] To solve the above technical problems, the present invention is implemented as follows.
[0010] A method for quantifying endogenous cross-linked peptides based on a neural network includes:
[0011] Step 1: Use cross-link search software to search the cross-link experiment data collected by a mass spectrometer to obtain K cross-link identification results; for each cross-link identification result k, generate I isotope masses and intensities, called the theoretical isotope distribution;
[0012] Step 2: Select one set from the M sets of cross-link experiment data obtained from M repeated experiments of the mass spectrometer as the basic experiment data, read the first-level spectrum, and take the first-level spectrum corresponding to the cross-link identification result k as the center to obtain an isotope intensity array corresponding to I isotope masses in the adjacent N first-level spectra, called the actual isotope distribution;
[0013] Step 3: In the actual isotope distribution, extract the chromatographic peak corresponding to each single isotope mass, perform area integration on the chromatographic peak, and obtain the quantification intensity value of the basic experiment data corresponding to the cross-link identification result k;
[0014] Step 4. For the cross-linking identification result k, in the first-level chromatogram of the M-1 group of cross-linking experiment data except for the basic experimental data, find the chromatographic peak with the most similar peak shape and perform area integration to obtain the quantitative intensity value of the other M-1 group of cross-linking experiment data corresponding to the cross-linking identification result k;
[0015] Step 5. Perform the operations of Step 2 to Step 4 for all the K cross-linking identification results obtained in Step 1 to complete the quantitative processing of the endogenous cross-linked peptides.
[0016] Preferably, the method further includes: screening according to the chromatographic peaks of the K cross-linking identification results, and performing the calculation operation of the quantitative intensity value in Step 3 and Step 4 for the screened cross-linking identification results; the screening is: calculating the cosine similarity score between the chromatographic peak of each cross-linking identification result and the theoretical isotope distribution of the cross-linked peptide segment and the theoretical isotope distribution of the single peptide, and filtering out the cross-linking identification results with similarity scores lower than that of the single peptide.
[0017] Preferably, in Step 3, the extraction of the chromatographic peak corresponding to each monoisotopic mass is as follows:
[0018] Taking the retention time of the cross-linking identification result as the center, extract the data with a set time length before and after this retention time;
[0019] Taking the retention time of the cross-linking identification result as the center of the chromatographic peak, record the maximum intensity value I at the center point max , initialize the left and right pointers, and synchronously expand from the center to both sides respectively. As the left and right pointers move, record the intensity value I of the current pointer current ; judge whether the left and right pointers encounter the intensity value I current less than I threshold = 0.1×I max ; if the left and right pointers each encounter I current <I threshold twice, stop expanding; take the stopping positions of the left pointer and the right pointer as the starting point and the ending point of the chromatographic peak.
[0020] Preferably, in Step 4, the finding of the chromatographic peak with the most similar peak shape in the first-level chromatogram of the M-1 group of cross-linking experiment data except for the basic experimental data is realized by a neural network:
[0021] Define each group of the M-1 group of cross-linking experiment data except for the basic experimental data as cross-linking experiment data S;
[0022] Take the first-level chromatogram number T corresponding to the cross-linking identification result k k , and in the first-level chromatogram of the cross-linking experiment data S, use the first-level chromatogram T kCentered around this, obtain an isotope intensity array corresponding to I isotope masses in the adjacent xN first-level spectra, where x is the magnification factor; according to the extracted isotope intensity array, extract the chromatographic peaks corresponding to each single isotope mass, denoted as Y j , j = 1, 2, …; the chromatographic peaks extracted from the basic experimental data are denoted as T;
[0023] Extract chromatographic peak T and chromatographic peak Y j of the peak characteristics, input them into the neural network, and obtain the peak shape similarity of chromatographic peak T and chromatographic peak Y j ; select the chromatographic peak with the highest similarity for area integration as the quantitative intensity value corresponding to the current cross-linking experimental data S.
[0024] Preferably, the peak characteristics include peak position, peak height, and peak width.
[0025] Preferably, the peak characteristics include chromatographic layer characteristics.
[0026] Preferably, the chromatographic layer characteristics include the retention time of the peak characteristics and / or the environmental matrix of the peak characteristics; the environmental matrix of the peak characteristics is the data of N consecutive first-level spectra centered on the first-level spectrum corresponding to the chromatographic peak.
[0027] Preferably, the neural network uses a transformer neural network.
[0028] Preferably, each of the M groups of cross-linking experimental data is used as a set of basic experimental data, and steps two to five are performed to obtain the complete quantitative results of endogenous cross-linked peptides.
[0029] Preferably, I = 5.
[0030] Beneficial effects:
[0031] (1) The present invention provides a brand-new protein quantification scheme designed for endogenous cross-linking, which is a label-free protein quantification scheme. By using the peak area method for integration and similarity peak shape matching based on a neural network, accurate quantification of endogenous cross-linking is obtained, solving the problem that the abundance of endogenous cross-linking is low and it is not easy to be detected.
[0032] (2) The present invention uses a neural network based on transformer to find similar peak shapes among multiple runs, solving the problem that the abundance of endogenous cross-linking is low and it is difficult for traditional algorithms to determine signals among multiple runs.
[0033] (3) The present invention fits the peak shape into a mass array with the same number of dimensions, trains the neural network to learn the peak shape, peak height, peak value, and also learn features such as retention time and environmental matrix at the chromatographic layer, so as to find the most similar peak among multiple matches.
[0034] (4) In a preferred embodiment, the present invention provides an error identification result filtering step, which uses a similarity matching algorithm to filter out the incorrect endogenous cross-linking results generated during identification by comparing the similarity between the actual isotope distribution and the theoretical isotope distribution, so as to judge the reliability of the identification results. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Figure 1 It is a flowchart of the method for quantitatively analyzing endogenous cross-linked peptides based on neural network according to the present invention.
[0036] Figure 2 It is a schematic diagram for filtering incorrect identification results using isotope distribution.
[0037] Figure 3 It is a schematic diagram for obtaining quantitative values by integration.
[0038] Figure 4 It is a schematic diagram for calculating the peak shape similarity by neural network calculation. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0039] The present invention will be described in detail below with reference to the accompanying drawings and examples.
[0040] Figure 1 The flowchart of the method for quantitatively analyzing endogenous cross-linked peptides based on neural network according to the present invention is shown as Figure 1 shown, and the method includes the following steps:
[0041] Step 1: Use cross-link search software to search the cross-linking experiment data collected by a mass spectrometer to obtain cross-linking identification results and obtain the theoretical isotope distribution.
[0042] In this step, when using the cross-link search software for searching, a total of K cross-linking identification results are obtained. For each cross-linking identification result, I isotope masses and the corresponding intensities of a peptide segment are extracted, which are called the theoretical isotope distribution and denoted as Q k . The following steps are executed for each cross-linking identification result k.
[0043] In this embodiment, I is preferably 5. For each cross-linking identification result k, I isotope masses and the corresponding intensities are extracted and stored in a two-dimensional array Q k [i][j], where i = 1 to I and j = 1, 2. The subscript k represents the data corresponding to the kth cross-linking identification result.
[0044] Among them, the cross-linking identification results include cross-linked peptide sequences, retention times, and may also include the protein numbers corresponding to the peptides.
[0045] Generate the theoretical isotope mass and intensity distribution for cross-linking identification results. The specific principle is to calculate the mass of amino acid residues in the peptide and the cross-linker. Since both amino acids and cross-linkers are composed of several elements, such as carbon (C), hydrogen (H), nitrogen (N), oxygen (O), and sulfur (S), the entire peptide can be represented as a chemical formula.
[0046] These elements all have different isotope forms. For example, carbon has C12 and C13. By calculating the probability of each isotope combination, the theoretical distribution of each isotope peak can be determined. For example, a molecule composed of three carbon atoms may have the following isotope combinations: 3 C12, 2 C12 and 1 C13, 1 C12 and 2 C13, 3 C13. Since the abundance of carbon isotope C12 is 98.93%, and the abundance of C13 is 1.07%, for a molecule composed of three carbon atoms, the probability of 3 C12 is 0.9893 3 = 0.9682, the probability of 2 C12 and 1 C13 is 3×0.9893 2 ×0.0107 = 0.0314, and so on.
[0047] Step 2: The mass spectrometer conducts M repeated experiments to obtain M sets of cross-linking experimental data. Select one set of cross-linking experimental data as the basic experimental data S0, and extract the actual isotope distribution.
[0048] In this step, the first-order spectrum is read from the basic experimental data S0. According to the retention time t of the cross-linking identification result k k , extract a set of N adjacent first-order spectra centered on the retention time t k . For each of the I isotope masses, extract the intensity value corresponding to the isotope mass from these N adjacent first-order spectra to obtain an intensity array S m,k [n] for the isotope mass corresponding to the adjacent N first-order spectra, where n = 1~N. Calculate the corresponding intensity arrays for the 5 theoretical isotope masses of a peptide in this way, and splice the intensity arrays into an I*N two-dimensional matrix, denoted as S m,k [i][n], where i = 1~I, n = 1~N, and the chromatographic curve corresponding to a peptide can be obtained.
[0049] S m,k [i][n] with subscript m representing the mth set of cross-linking experimental data, k representing the kth cross-linking identification result, i representing the ith isotope mass, and n representing the isotope mass extracted from the nth first-order spectrum. n = 1~N, i = 1~I, m = 1~M, k = 1~K. In one example, N = 40, I = 5, M = 4, K = 100.
[0050] The two-dimensional matrix S obtained in this stepm,k [i][n] is called the actual isotope distribution S m,k From M sets of experimental data, M actual isotope distributions S 1,k ~S M,k .
[0051] In this step, when extracting the intensity value corresponding to the isotope mass from the first-order spectrum, a permitted range containing the isotope mass to be extracted is set. When the isotope mass in the first-order spectrum falls within this permitted range, it is considered that the matching isotope mass is found, and the intensity value corresponding to this matching isotope mass is extracted. In practice, when the first three digits after the decimal point are the same, it can be considered that the matching isotope mass is found.
[0052] For the obtained intensity array, Z-score normalization is used to convert the data into a distribution with a mean of 0 and a standard deviation of 1. Its calculation formula is:
[0053]
[0054] Among them, z is the original data, z’ is the normalized data, μ is the mean of the data, and σ is the standard deviation of the data.
[0055] Step 3: For each single-isotope mass intensity array S m,k [i][n] in the actual isotope distribution, extract the chromatographic peak.
[0056] In this step, in the intensity array of the single-isotope mass, the intensity threshold method is used to determine the chromatographic peak, that is, to determine the starting point and ending point of the chromatographic peak. The specific steps are as follows:
[0057] Step S301: Centering on the retention time of the cross-linking identification result, extract the chromatographic data for 2n minutes before and after this retention time. Usually, n takes 1 or 2.
[0058] Step S302: Taking the retention time of the cross-linking identification result as the center of the chromatographic peak, initialize two pointers, the left and the right, and start expanding from the center point to both sides. Record the maximum intensity value I max at the center point. The left pointer moves to the left, and the right pointer moves to the right, and record the intensity value I current of the current pointer. Judge whether the left and right pointers encounter the intensity value I current less than I threshold = 0.1×I max ; if the left and right pointers each encounter I current <I threshold twice, stop expanding. Take the stopping positions of the left and right pointers as the starting point and ending point of the chromatographic peak.
[0059] In practice, other methods can also be used to obtain the window of the chromatographic peak. The extracted chromatographic peak is a part of the intensity array, S m,k [i][n], where n = 1 to N', and N' < N.
[0060] Step 4: Filter out the results of poor cross-linking identification.
[0061] For the cross-linked peptides mediated by TGase, the cross-linking identification results may overlap with the single peptide identification results, and this part of the identification results is not credible. Therefore, it is necessary to filter out the low-confidence identifications in this part of the cross-linking identification results. For the K cross-linking identification results obtained in Step 1, chromatographic peaks are extracted for each of them. In this step, the cosine similarity scores between the chromatographic peaks of each cross-linking identification result and the theoretical cross-linked peptide isotope distribution and the theoretical single peptide isotope distribution are calculated respectively. According to the scores, the cross-linking identification results are re-sorted, and the cross-linking identification results with similarity scores lower than that of the single peptide are filtered out, as Figure 2 shown. In practice, this filtering operation can be performed as long as after obtaining the chromatographic peaks, and there is no need to limit it before calculating the area in Step 5.
[0062] The specific process of this step includes:
[0063] Step S401: Obtain the actual chromatographic curve of the peptide from the mass spectrometry data, denoted as S m,k [i][n], where i = 1 to I and n = 1 to N;
[0064] Step S402: Based on the known amino acid sequence and elemental composition, calculate the theoretical isotope mass distributions of the single peptide and the cross-link corresponding to the same identification result (see Step 1), denoted as D m,k [i][n], where i = 1 to I and n = 1 to N; J m,k [i][n], where i = 1 to I and n = 1 to N;
[0065] Step S403: Use the cosine similarity algorithm to compare the theoretical isotope mass distributions of the single peptide and the cross-link with the actual distributions to determine the similarity between them. Taking the single peptide as an example, first, the two-dimensional arrays of the theoretical isotope mass D of the single peptide and the actual chromatographic curve S are flattened into a one-dimensional long vector to obtain s and d, and then the dot product and norm of the one-dimensional long vector are calculated. The cosine similarity formula (Cosine Similarity) is used to calculate the similarity between the two one-dimensional long vectors, which can be used as the similarity between the theoretical isotope mass distribution of the single peptide and the actual distribution.
[0066] s = flatten(S)
[0067] d = flatten(D)
[0068]
[0069] Step S404: Filter out the cross-linking identification results with similarity scores lower than that of the single peptide according to the set similarity threshold, that is, filter out the part of <SIM SJ <SIM SD .
[0070] Step Five: Integrate the chromatographic peaks of the remaining cross-linking identification results after filtration using the peak area method to obtain the quantitative intensity value Q of the cross-linking identification result k DL k,m .
[0071] In this step, use a numerical integration method (such as the trapezoidal method or Simpson's method) to integrate the middle part between the starting point and the ending point of the chromatographic peak, that is, calculate the area of the chromatographic peak, as shown Figure 3 .
[0072]
[0073] where I(t) represents the intensity value of the chromatographic curve at time t, and t start and t end are the starting point and the ending point of the chromatographic peak respectively
[0074] Step Six: For the cross-linking identification result k, find the chromatographic peak with the most similar peak shape in the first-order spectrogram corresponding to the M-1 group of experimental data outside the basic experimental data for integration, as the quantitative intensity value of other experimental data
[0075] If the basic experimental data is the first group of experimental data and there are a total of 4 groups of experimental data, then for the cross-linking identification result k, the quantitative intensity value Q obtained in Step Five is based on the first group of experimental data DL k,1 , then the quantitative intensity values of the cross-linking identification result k obtained from the other 3 groups of experimental data are Q DL k,2 , Q DL k,3 , Q DL k,4 .
[0076] Define each of the M-1 groups of cross-linking experimental data except the basic experimental data as cross-linking experimental data S. The following details the method of finding the chromatographic peak with the most similar peak shape in the cross-linking experimental data S. See Figure 4 :
[0077] Step 601: Obtain the first-order spectrogram number T of the cross-linking identification result k k , and in the first-order spectrogram of the cross-linking experimental data S, use this spectrogram T kCentered around it, extract N×x first-level spectrograms, where x is the magnification factor. Since there will be certain errors during identification in multiple runs, it is necessary to extract chromatographic curves in a larger range. According to the method in step two, use xN first-level spectrograms to extract chromatographic curves, and then extract the chromatographic peaks corresponding to each single-isotope mass, denoted as Y j , j = 1, 2, …; the number of chromatographic peaks in each set of cross-linking experimental data S is variable. The chromatographic peaks extracted from the basic experimental data are denoted as T
[0078] Step 602: For the chromatographic peak T of the cross-linking identification result k and the chromatographic peaks Y in other experimental data j , extract the peak characteristics of the chromatographic peaks
[0079] The peak characteristics include peak shape characteristics and chromatographic-level characteristics. The peak shape characteristics can include the peak position (i.e., the index of the peak value in the array), the peak height, and the peak width (i.e., the width occupied by the peak value in the array); the purpose of extracting these characteristics is to further compare the peak shape similarities of different arrays. In addition to the characteristics of the peak itself, it is also necessary to extract the chromatographic-level characteristics. The chromatographic-level characteristics can include the retention time of the chromatographic peak at the chromatographic level, and can also include the environmental matrix of the chromatographic peak at the chromatographic level. The environmental matrix is the N consecutive first-level spectrogram data centered around the first-level spectrogram corresponding to the chromatographic peak Y j Select N graphs to be consistent with the N graphs of the chromatographic peak T
[0080] Step 603: Train a neural network to compare the peak shape similarities between the chromatographic peak T and the chromatographic peak Y j to obtain similar chromatographic peaks. In this embodiment, the neural network is preferably a transformer neural network
[0081] Input the peak characteristics of the chromatographic peak T and the chromatographic peak Y j to be compared with it into the transformer neural network, and the transformer neural network outputs a similarity score. Compare all the chromatographic peaks Y1~Y P (P is the total number of chromatographic peaks in the cross-linking experimental data S) of the cross-linking experimental data S with the chromatographic peak T respectively to obtain their respective scores. The chromatographic peak with the highest score is the similar chromatographic peak to the chromatographic peak T
[0082] The transformer neural network first encodes the positions of T and T through a position encoding function (sine, cosine or a neural network with learnable parameters) yThe features are respectively encoded as vectors, and the encoded vectors are concatenated into an input vector. Then, the self-attention mechanism is used to calculate the correlation of the input at each position with all positions. The multi-head mechanism allows the model to independently perform these calculations in different subspaces and then concatenates the results together. In this step, the input vector is first mapped into query, key, and value matrices. The calculation formulas for query, key, and value are as follows:
[0083] Q = XW Q
[0084] K = XW K
[0085] V = XW V
[0086] where X is the input vector, and W Q , W K and W V are the weight matrices for query, key, and value respectively. Then, the dot product of the query and the key is calculated and scaled:
[0087]
[0088] where d k is the dimension of the key, and the softmax function is used to normalize the attention weights.
[0089] Through the above steps, the transformer can capture the dependencies between different positions in the input sequence. During the training process, by comparing the peak shape features of chromatographic peak T with those of Y1 to Y P , the neural network can learn how to judge the similarity between different arrays. Finally, by evaluating the similarity scores output by the neural network, it can be determined which peak shape among chromatographic peak T and Y1 to Y P is more similar.
[0090] As Figure 4 shown, for the three candidate chromatographic peaks A, B, and C, the similarity of chromatographic peak C is the highest.
[0091] Step 604: For the similar chromatographic peaks, perform area integration, and the integration result is used as the quantitative intensity value corresponding to the cross-linking experiment data S. Area integration is performed for each cross-linking experiment data after extracting the similar chromatographic peaks to obtain the quantitative intensity value.
[0092] So far, the process of obtaining the quantitative intensity values of M groups of cross-linking experiment data for the cross-linking identification result k is completed.
[0093] Step Seven: Perform the operations of Step Two to Step Six for each of the K cross-linking identification results obtained in Step One to complete the quantitative processing of the endogenous cross-linked peptides.
[0094] Repeat the above steps two to seven, select another set of cross-linking experimental data as the basic experimental data, and obtain a set of quantitative intensity values. Each set of the M sets of cross-linking experimental data can be used as the basic experimental data once, and M sets of quantitative intensity values can be obtained. These M sets of quantitative intensity values can be verified with each other. For example, when quantifying based on cross-linking data A, the quantitative value of peptide Pep1 in A is 100 and in B is 10; when quantifying based on cross-linking data B, the quantitative value of peptide Pep1 in A is 98 and in B is 14. The quantitative intensity values obtained separately from these two sets of data can be normalized to obtain more accurate results.
[0095] The above specific embodiments only describe the design principle of the present invention. The shapes and names of the components in this description can be different and are not limited. Therefore, those skilled in the art of the present invention can modify or make equivalent replacements to the technical solutions recorded in the foregoing embodiments; and these modifications and replacements do not depart from the purpose and technical solutions of the present invention, and should all fall within the protection scope of the present invention.
Claims
1. A method for quantifying endogenous cross-linked peptides based on a neural network, characterized in that: include: Step 1: Use cross-link search software to search the cross-link experimental data collected by the mass spectrometer to obtain K cross-link identification results; For each cross-link identification result k, I isotope mass and intensity are generated, which is called the theoretical isotope distribution; Step 2: From the M groups of cross-linking experimental data obtained by the mass spectrometer M times of repeated experiments, one group is selected as the basic experimental data, and the primary spectrum is read. With the primary spectrum corresponding to the cross-linking identification result k as the center, an isotope intensity array corresponding to I isotope mass is obtained in the adjacent N primary spectra, which is called the actual isotope distribution; Step 3: In the actual isotope distribution, extract the chromatographic peak corresponding to each monoisotopic mass, perform area integration on the chromatographic peak, and obtain the quantitative intensity value of the basic experimental data corresponding to the cross-linking identification result k; Step 4: for the cross-linking identification result k, in the primary spectrum of the M-1 group of cross-linking experimental data other than the basic experimental data, find the chromatographic peak with the most similar peak shape and perform area integration to obtain the quantitative intensity value of the other M-1 group of cross-linking experimental data corresponding to the cross-linking identification result k; Step 5: Perform the operations of steps 2 to 4 for the K cross-linking identification results obtained in step 1 to complete the quantitative processing of endogenous cross-linked peptides.
2. The method according to claim 1, characterized in that The method further comprises: screening according to the chromatographic peaks of K cross-linking identification results, and performing the quantitative intensity value calculation operation of step three and step four on the cross-linking identification results after screening; the screening comprises: calculating the cosine similarity score between the chromatographic peak of each cross-linking identification result and the theoretical cross-linking peptide segment isotope distribution and the theoretical single peptide isotope distribution, and filtering out the cross-linking identification results with a similarity score lower than that of a single peptide.
3. The method according to claim 1, characterized in that In step 3, the chromatographic peak corresponding to each monoisotopic mass is extracted as follows: Taking the retention time of the cross-linking identification result as the center, extract the data of the set time length before and after the retention time; The retention time of the cross-linking identification result is taken as the center of the chromatographic peak, and the maximum intensity value I of the center point is recorded. max Initialize the left and right pointers, expand them synchronously from the center to both sides, and record the current pointer strength value I as the left and right pointers move current ; Determine whether the left and right pointers meet the intensity value I current Less than I threshold =0.1×I max If the left and right pointers encounter I twice current threshold , stop expanding; the positions where the left and right pointers stop are used as the starting and ending points of the chromatographic peak. 4. The method according to claim 1, characterized in that In step 4, in the primary spectrum of the M-1 group cross-linking experimental data other than the basic experimental data, searching for the chromatographic peak with the most similar peak type is implemented by using a neural network: Each group of the M-1 groups of cross-linking experimental data except the basic experimental data is defined as cross-linking experimental data S; Take the first-level map number T corresponding to the cross-link identification result k k In the primary spectrum of the cross-linking experimental data S, the primary spectrum T k Centered at xN adjacent primary spectra, obtain the isotope intensity array corresponding to I isotope mass, x is the magnification factor; according to the extracted isotope intensity array, extract the chromatographic peak corresponding to each monoisotope mass, recorded as Y j , j = 1, 2, ...; the chromatographic peak extracted from the basic experimental data is recorded as T; Extract chromatographic peaks T and Y j The peak features are input into the neural network to obtain the chromatographic peaks T and Y j The peak similarity of the cross-linking experiment data S is obtained by selecting the chromatographic peak with the highest similarity for area integration.
5. The method according to claim 4, characterized in that The peak characteristics include peak position, peak height, and peak width.
6. The method according to claim 5, characterized in that The peak characteristics include chromatographic level characteristics.
7. The method according to claim 6, characterized in that The chromatographic level features include the retention time of peak features and / or the environment matrix of peak features; the environment matrix of peak features is N consecutive primary spectrum data centered on the primary spectrum corresponding to the chromatographic peak.
8. The method according to claim 4, characterized in that The neural network adopts a transformer neural network.
9. The method according to claim 1, characterized in that Each of the M groups of cross-linking experimental data is used as a basic experimental data, and steps 2 to 5 are performed to obtain complete quantitative results of endogenous cross-linked peptides.
10. The method according to claim 1, characterized in that I=5。
Citation Information
Patent Citations
Cross-linked peptide fragment qualitative and quantitative analysis method
CN111208299A
Method for realizing quantification of D-dimer in plasma based on immobilized metal ion affinity chromatography enrichment
CN113358760A
Mass spectrometry-cleavable cross-linking agents to facilitate structural analysis of proteins and protein complexes, and method of using same
US20130144541A1
Cross-linking compositions and related methods of isotope tagging of interacting proteins and analysis of protein interactions
US20140206091A1
Crosslinkers for mass spectrometric identification and quantitation of peptides
US20190084929A1