A method for integrating multiple liquid chromatography-mass spectrometry maps for polypeptide quantification analysis and application thereof

By integrating multiple liquid chromatography-mass spectrometry (LC-MS) chromatograms and utilizing the consistency of retention time and ion mass-to-charge ratio for the same component, inconsistent chromatographic peaks and fragment ions are eliminated, thereby improving the quantitative accuracy and repeatability of LC-MS technology and solving the problem of insufficient quantitative accuracy in existing technologies.

CN119827645BActive Publication Date: 2025-12-12HANGZHOU INST FOR ADVANCED STUDY UCAS
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202311323216.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-10-13
Publication Date
2025-12-12
Estimated Expiration
2043-10-13

AI Technical Summary

Technical Problem

Existing liquid chromatography-mass spectrometry (LC-MS) techniques have shortcomings in terms of quantitative accuracy and repeatability of multiple analysis results, especially in data-dependent acquisition modes, where interference signals lead to reduced quantitative accuracy, and existing algorithms do not fully utilize the repeatability of multiple analysis results.

Method used

Using a data-independent acquisition mode, multiple spectra are integrated through multiple liquid chromatography-mass spectrometry analyses. By utilizing the consistency of retention time, ion mass-to-charge ratio, and relative intensity of the same component, inconsistent chromatographic peaks and secondary fragment ions are eliminated, the relative content of peptides is calculated, and accurate quantitative chromatographic peaks and quantitative ions are screened out.

Benefits of technology

It improves quantitative accuracy, reduces quantitative error rate, and does not rely on additional experiments. The software automatically performs the screening process without the need for manual operation by professionals, and is suitable for different types of liquid chromatography-mass spectrometry datasets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119827645B_ABST
    Figure CN119827645B_ABST
Patent Text Reader

Abstract

The application discloses a polypeptide quantitative analysis method integrating multiple liquid chromatography-mass spectrometry (LC-MS) maps and application thereof, and belongs to the technical field of the cross of analytical chemistry and computer science. The application uses LC-MS technology to analyze samples, obtains multiple secondary fragment ion extraction ion flow chromatograms, and integrates multiple LC-MS maps, thereby changing the original mode of sequentially processing single maps and then merging results. The consistency of the retention time, ion mass-to-charge ratio and relative intensity of the same component in multiple analyses is utilized, inconsistent chromatographic peaks and secondary fragment ions in multiple analyses are excluded, the accuracy of quantitative chromatographic peaks and quantitative ions obtained through screening is improved, the relative content of polypeptides in samples is calculated by using the quantitative chromatographic peaks and quantitative ions obtained through screening, the quantitative accuracy of polypeptides and proteins is improved, and the application can be used for proteome non-target analysis of clinical samples and biological samples.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the technical field of the intersection of analytical chemistry and computer science, and specifically relates to a polypeptide quantitative analysis method integrating multiple liquid chromatography-mass spectrometry (LC-MS) maps and application thereof. BACKGROUND

[0002] Proteomics non-target analysis technology based on liquid chromatography-mass spectrometry (LC-MS) can simultaneously quantitate thousands of proteins, and is widely used to compare the differences in protein expression between samples, so as to mine important biomarkers. In the determination, proteins are usually enzymatically hydrolyzed into polypeptides, and the relative content of polypeptides is determined by LC-MS, and then the relative content of proteins is deduced. The quantitative accuracy of polypeptides determines the quantitative accuracy of proteins.

[0003] In quantitative analysis, data-independent acquisition (DIA) is a widely used mass spectrometry scanning mode, which can stably collect fragment ion spectra, make up for the defects of semi-random sampling of data-dependent acquisition (DDA) mode, and significantly improve the repeatability of multiple analysis results. DIA usually uses dozens of mass-to-charge ratio isolation windows, and each time the parent ions in an isolation window are fragmented and then the fragment ion spectrum is scanned. Since the mass-to-charge ratio range of the isolation window is significantly wider than that of DDA, combined with the complexity of the sample itself (containing a large number of polypeptides usually exceeding one hundred thousand), the obtained fragment ion spectrum is very complex. Many polypeptides with similar mass-to-charge ratios and similar sequences will produce fragment ions in the same isolation window, resulting in a large number of interference signals in the LC-MS spectrum. These interference signals will cause some target polypeptides to be incorrectly attributed to interference chromatographic peaks, and some fragment ions will not be used for polypeptide quantification due to interference, ultimately reducing the quantitative accuracy. Therefore, how to improve the quantitative accuracy under the premise of a large number of quantifications is an important problem to be solved in current research.

[0004] DIA-NN (V. Demichev, et al. DIA-NN: neural networks and interference correction enable deep proteome coverage in high throughput. Nat. Methods, 17, 41-44, 2020) and AvG (Vaca Jacome, A. S. et al. Avant-garde: an automated data-driven DIA data curation tool. Nat. Methods, 17, 1237-1244, 2020) are two algorithms with the best reported quantitative accuracy, DIA-NN uses deep neural networks to distinguish correct and false candidate chromatographic peaks to improve the accuracy of identification and quantification, and AvG uses genetic algorithms to optimize the screening of quantitative ions. Spectronaut is the most commonly used commercial software, and the specific algorithm principle of the current version is not publicly reported due to its commercial operation. However, the high repeatability of DIA multiple analysis results has not been fully valued and utilized, and the existing algorithms such as DIA-NN and AvG still process data in the mode of first processing each map one by one, and then merging the results of each map for subsequent processing, and the information contained in a single map is sometimes not enough to correctly screen chromatographic peaks and quantitative ions, and the accuracy of quantification still has room for further improvement.

[0005] The Chinese patent document with publication number CN115267033A discloses a macro proteomics analysis method based on mass spectrometry data, which includes: selecting protein data from a microbial protein database, and each time the selected protein data forms a first data set corresponding to each first mass spectrometry data; under the condition that the false discovery rate meets a first threshold, extracting protein data from each first data set to form a second data set; under the condition that the false discovery rate meets a second threshold, selecting protein data from the second data set based on each first mass spectrometry data, and constructing a first spectrum library; optimizing the first spectrum library to form a second spectrum library; based on the second spectrum library and second mass spectrometry data, performing qualitative and quantitative analysis on polypeptide segments and associated proteins contained in a second sample to obtain analysis results. However, this method focuses on identifying polypeptides and proteins, and does not compare the quantitative accuracy.

[0006] In summary, it is necessary to utilize the high repeatability of DIA multiple analysis results on the basis of existing technology to design a polypeptide quantification analysis method that integrates multiple liquid chromatography-mass spectrometry maps with higher quantitative accuracy. SUMMARY

[0007] The application provides a polypeptide quantitative analysis method integrating multiple liquid chromatography-mass spectrometry (LC-MS) maps.

[0008] The specific technical solutions are as follows:

[0009] A polypeptide quantitative analysis method integrating multiple liquid chromatography-mass spectrometry (LC-MS) maps comprises the following steps.

[0010] The data-independent acquisition mode is adopted, and liquid chromatography-mass spectrometry (LC-MS) technology is used to analyze each sample multiple times or analyze multiple groups of samples one or more times, ion flow chromatograms of secondary fragment ions in original LC-MS data are extracted, multiple LC-MS maps are obtained, the consistency of the retention time, ion mass-to-charge ratio and relative intensity of the same component in multiple analysis by liquid chromatography-mass spectrometry (LC-MS) technology is used, and the multiple LC-MS maps are integrated and processed, inconsistent chromatographic peaks and secondary fragment ions between multiple analyses are excluded, the accuracy of quantitative chromatographic peaks and quantitative ions obtained by screening is improved, and the relative content of the identified polypeptide in samples is calculated by using the quantitative chromatographic peaks and quantitative ions obtained by screening.

[0011] The sample is a peptide sample, which can be specifically selected from the enzymatic hydrolysate of one or more target proteins.

[0012] Specifically, the polypeptide quantitative analysis method integrating multiple liquid chromatography-mass spectrometry (LC-MS) maps comprises the following steps.

[0013] (1) The data-independent acquisition mode is adopted, and liquid chromatography-mass spectrometry (LC-MS) technology is used to analyze each sample multiple times or analyze multiple groups of samples one or more times, and original LC-MS data are obtained.

[0014] (2) According to the data in the spectrum library, the retention time of each target polypeptide in the sample is predicted, the scanning data of the target polypeptide within the allowed deviation range of the predicted retention time is extracted from each LC-MS original data, the extraction ion flow chromatogram of all secondary fragment ions of the target polypeptide within the allowed deviation range of the mass-to-charge ratio is obtained, and the chromatographic peak obtained by integration in the extraction ion flow chromatogram is used as a candidate chromatographic peak.

[0015] (3) For each candidate chromatographic peak, the consistency of the chromatographic peak type of the secondary fragment ions of the same component is used to screen and exclude the secondary fragment ions with inconsistent peak types in the candidate chromatographic peak, the measured peak area of the remaining secondary fragment ions is calculated, and the measured relative intensity is obtained, the relative intensity of the corresponding secondary fragment ion in the spectrum library is used as a reference, and the secondary fragment ions with large relative intensity deviation are excluded.

[0016] (4) According to the principle that multiple determinations are closer to the true value, more liquid chromatography-mass spectrometry maps are easier to find measured data close to the spectral library, and the spectral library matching score of the candidate chromatographic peak is calculated, and in all candidate chromatographic peaks corresponding to multiple liquid chromatography-mass spectrometry maps, the candidate chromatographic peak with the highest spectral library matching score is selected as the best matching peak (there is only one best matching peak);

[0017] (5) Taking the best matching peak as a reference, the chromatographic peak of the target polypeptide is found in the liquid chromatography-mass spectrometry map other than the liquid chromatography-mass spectrometry map where the best matching peak is located, specifically, the consistency of the retention time of the same component between multiple liquid chromatography analyses is utilized, the retention time of the best matching peak is aligned to the retention time of each liquid chromatography-mass spectrometry map, the allowed deviation of the retention time is reduced, and the candidate chromatographic peak with a measured retention time within the allowed deviation range is retained; then the consistency of the ion mass-to-charge ratio and its relative intensity between multiple liquid chromatography analyses is utilized, the similarity score between the retained candidate chromatographic peak and the best matching peak is calculated, and in the candidate chromatographic peak corresponding to each liquid chromatography-mass spectrometry map, the candidate chromatographic peak with the highest similarity score is selected as the measured chromatographic peak, and the measured chromatographic peak and the best matching peak form a chromatographic peak group, if the maximum similarity score in a liquid chromatography-mass spectrometry map does not meet the critical value requirement (the similarity score critical value range between the candidate chromatographic peak and the best matching peak is 0.4-0.8), it is considered that the target polypeptide is not detected in the liquid chromatography-mass spectrometry map; the consistency of the secondary fragment ions of the same component measured between multiple liquid chromatography analyses is utilized, the chromatographic peak with large differences in secondary fragment ions from other chromatographic peaks is excluded from the chromatographic peak group, and the retained chromatographic peak is used as a quantitative chromatographic peak for final quantitative calculation, and the different secondary fragment ions between the retained quantitative chromatographic peaks are excluded;

[0018] (6) The consistency of the ion relative intensity between multiple liquid chromatography analyses is utilized to finally screen the quantitative ions, specifically, the proportion of the measured area of each secondary fragment ion between the quantitative chromatographic peaks corresponding to each liquid chromatography analysis is calculated, if there is an outlier in the proportion, the corresponding secondary fragment ion is excluded, the correlation of the measured area of the retained secondary fragment ion between multiple quantitative chromatographic peaks is calculated, and the secondary fragment ion with low correlation is excluded, and the retained secondary fragment ion is used as a quantitative ion;

[0019] (7) According to the quantitative chromatographic peak and the quantitative ion screened in the chromatographic peak group, the chromatographic peak group score is calculated, and the weighted geometric mean or weighted arithmetic mean of the chromatographic peak group score and the highest spectral library matching score is obtained to obtain a polypeptide level score, and the polypeptide level score represents the reliability of the determination, and the weight coefficient ranges from 0.1 to 10;

[0020] (8) A polypeptide level score critical value is specified, and the target polypeptide with a polypeptide level score higher than the corresponding score critical value is the identified polypeptide, and the critical value ranges from 0.75 to 0.95;

[0021] (9) The area of each chromatographic peak in the chromatographic peak group is obtained by summing the area of the quantification ion, and the relative content of the identified polypeptide in the sample is calculated according to the area of each chromatographic peak in the chromatographic peak group of the sample.

[0022] Preferably, in step (2), the method for predicting the retention time of the target polypeptide is as follows: taking the measured retention time of the reference polypeptide in each liquid chromatography-mass spectrometry spectrum as the dependent variable, the relative retention time (iRT value) in the spectrum library as the independent variable, and the relative retention time of the target polypeptide as the sample point, the predicted retention time of the target polypeptide is obtained by linear interpolation or spline interpolation; the reference polypeptide is a polypeptide with high content added to the sample before liquid chromatography-mass spectrometry, or a polypeptide with high content contained in the sample, and the number of reference polypeptides is 5-200.

[0023] Preferably, in step (2), the allowed deviation range of the retention time is not more than 10 min, and the allowed deviation range of the mass-to-charge ratio is not more than 30 ppm.

[0024] Preferably, in step (3), the method for screening and excluding secondary fragment ions with inconsistent peak types in the candidate chromatographic peak is as follows: when the peak top of a chromatographic peak is outside the overall range of the chromatographic peak, the secondary fragment ion corresponding to the chromatographic peak is directly excluded; when the peak top of a chromatographic peak is within the overall range of the chromatographic peak, the sorting (from large to small) of the intensity of each secondary fragment ion at the overall peak top time point of the chromatographic peak within the range of the chromatographic peak is calculated, and the secondary fragment ion is excluded if the ratio of the sorting to the number of time points within the range of the chromatographic peak is greater than 0.8.

[0025] Preferably, in step (3), the method for excluding secondary fragment ions with large relative intensity deviation is as follows: first, the root mean square error (RMSE) between the measured relative intensity of the secondary fragment ion and the relative intensity of the corresponding secondary fragment ion in the spectrum library is calculated, and then the root mean square error (RMSE) between the measured relative intensity of the secondary fragment ion after being temporarily removed and the relative intensity of the corresponding secondary fragment ion in the spectrum library is calculated -1 If the RMSE of a secondary fragment ion is -1 < RMSE-0.05, the secondary fragment ion is excluded.

[0026] In step (4), the spectrum library matching score is obtained by weighted geometric mean or weighted arithmetic mean of multiple single scores; the single scores include: the similarity score of the peak area of each secondary fragment ion represented by the normalized root mean square error and the spectrum library intensity, the deviation score between the measured mass-to-charge ratio of each secondary fragment ion and the theoretical mass-to-charge ratio, the ratio score of the number of measured secondary fragment ions to the number of secondary fragment ions in the spectrum library, the peak top consistency score of the secondary fragment ions, the deviation score between the measured retention time and the predicted retention time, and the peak area sum score of the secondary fragment ions; the weight coefficient of each single score ranges from 0.01 to 100.

[0027] Specifically, a) the similarity score of each secondary fragment ion peak area to its library intensity is calculated according to formula (1):

[0028]

[0029] SRSS is the similarity score of the secondary fragment ion peak area to its library intensity, RMSE1 is the root mean square error between the normalized measured relative intensity of the secondary fragment ion and the corresponding relative intensity of the secondary fragment ion in the library; when RMSE1≥0.3, SRSS is set to 0.

[0030] b) the deviation score between the measured mass-to-charge ratio and the theoretical mass-to-charge ratio of each secondary fragment ion is calculated according to formula (2):

[0031]

[0032] MES is the deviation score between the measured mass-to-charge ratio and the theoretical mass-to-charge ratio of the secondary fragment ion, which is calculated by converting the deviation between the measured mass-to-charge ratio and the theoretical mass-to-charge ratio of the current retained secondary fragment ion to ppm, taking the absolute value, and then taking the median of the absolute values of the deviations of each secondary fragment ion, i.e. Δmz FP ; when Δmz FP ≥30, MES is set to 0.

[0033] c) the ratio score of the number of retained measured secondary fragment ions to the number of library secondary fragment ions is calculated according to formula (3):

[0034] FQS=n F-remain / n F (3)

[0035] FQS is the ratio score of the number of retained measured secondary fragment ions to the number of library secondary fragment ions; n F-remain is the number of fragment ions in the chromatographic peak that have not been excluded, and n F is the number of library ions.

[0036] d) the chromatographic peak top consistency score between secondary fragment ions is calculated according to formula (4):

[0037]

[0038] FPCS is the chromatographic peak top consistency score between secondary fragment ions; for each secondary fragment ion that has not been excluded, the intensity of the chromatographic peak top time point is calculated in the order of descending intensity in the ion intensity at each time point in the chromatographic peak range, and the average of each order is r apex,mean ; when r apex,mean ≥6, FPCS is set to 0.

[0039] e) the deviation score of the measured retention time from the predicted retention time is calculated according to formula (5):

[0040]

[0041] RTES is the score of the deviation of the measured retention time from the predicted retention time; RT error is the absolute value of the deviation of the measured retention time from the predicted retention time, TW 1 / 2 is the allowed deviation of the retention time in the extraction ion current of the secondary fragment ions in step (2).

[0042] f) the sum of the peak areas of the secondary fragment ions is calculated according to formula (6):

[0043]

[0044] FIS is the score of the sum of the peak areas of the secondary fragment ions; A sum is the sum of the peak areas of the secondary fragment ions at present not excluded; A max and A min are the maximum and minimum values of A sum of all the chromatographic peaks in the time range of the extraction ion current of the secondary fragment ions, respectively.

[0045] Preferably, in step (5), the calculation method of aligning the retention time of the best matching peak to the retention time of each LC-MS profile is as follows: taking the measured retention time of the reference polypeptide in each LC-MS profile as the dependent variable, taking the measured retention time of the reference polypeptide in the LC-MS profile where the best matching peak is located as the independent variable, taking the retention time of the best matching peak as the sample point, and using linear interpolation or spline interpolation method to interpolate, the aligned retention time is obtained. The purpose of alignment is to find the corresponding chromatographic peak of the component in the remaining LC-MS profiles (other than the LC-MS profile where the best matching peak is located).

[0046] Preferably, in step (5), the allowed deviation of the retention time is reduced to ±0.1-±3 min.

[0047] In step (5), the similarity score between the candidate chromatographic peak and the best matching peak is obtained by weighted geometric mean or weighted arithmetic mean of multiple single scores; the single scores include: the similarity score of the intensity of each secondary fragment ion peak after normalization to the intensity of the best matching peak, the deviation score between the measured mass-to-charge ratio and the theoretical mass-to-charge ratio of each secondary fragment ion, the score of the ratio of the number of measured secondary fragment ions to the number of ions of the best matching peak, and the score of the consistency of the peak top between the secondary fragment ions; the weight coefficient of each single score ranges from 0.01 to 100.

[0048] Specifically, in step (5), the similarity score of the peak area of each secondary fragment ion to the intensity of the best matching peak is calculated for the current retained secondary fragment ion according to formula (1), SRSS represents the similarity score of the peak area of the secondary fragment ion to the intensity of the best matching peak, RMSE1 takes the value of the root mean square error between the normalized measured relative intensity of the secondary fragment ion and the relative intensity of the corresponding secondary fragment ion in the best matching peak; the deviation score between the measured mass-to-charge ratio of each secondary fragment ion and the theoretical mass-to-charge ratio is calculated for the current retained secondary fragment ion according to formula (2); the score of the ratio of the number of retained measured secondary fragment ions to the number of ions in the best matching peak is calculated for the current retained secondary fragment ion according to formula (3), FQS represents the score of the ratio of the number of retained measured ions to the number of ions in the best matching peak, n F The value of the number of ions in the best matching peak; the consistency score of the peak top between the secondary fragment ions is calculated for the current retained secondary fragment ion according to formula (4).

[0049] Preferably, in step (6), the peak area of each secondary fragment ion retained in the quantitative chromatographic peak is taken as a vector, the cosine of the angle between each pair of vectors represents the correlation between the measured areas of different secondary fragment ions in multiple quantitative chromatographic peaks, and the cosine distance, i.e. 1 minus the cosine of the angle, is used for cluster analysis of the secondary fragment ions. If the cosine distance exceeds a threshold value, the corresponding secondary fragment ion is a secondary fragment ion with low correlation. The threshold value ranges from 0.02 to 0.1.

[0050] In step (7), the score of the chromatographic peak group is obtained by weighted geometric mean or weighted arithmetic mean of multiple single scores; the single scores include: the similarity score of the peak area of each secondary fragment ion to its library intensity represented by the normalized root mean square error, the deviation score between the measured mass-to-charge ratio of each secondary fragment ion and the theoretical mass-to-charge ratio, the score of the ratio of the number of retained measured secondary fragment ions to the number of library secondary fragment ions, the consistency score of the peak top between the secondary fragment ions, the deviation score of the measured retention time and the predicted retention time, the relative standard deviation score between the peak areas of multiple quantitative chromatographic peaks corresponding to the same sample, and the score of the ratio of the number of quantitative chromatographic peaks in the chromatographic peak group to the number of LC-MS maps; the weight coefficient of each single score ranges from 0.01 to 100.

[0051] When only one analysis is performed on multiple groups of samples respectively, the single score of the chromatographic peak group score does not include the relative standard deviation score between the peak areas of multiple quantitative chromatographic peaks corresponding to the same sample.

[0052] Specifically, in step (7), the peak area of each secondary fragment ion and the similarity score of its spectral library are calculated according to formula (1) for the current reserved secondary fragment ion; the deviation score between the measured mass-to-charge ratio and the theoretical mass-to-charge ratio of each secondary fragment ion is calculated according to formula (2) for the current reserved secondary fragment ion; the score of the ratio between the number of the reserved measured secondary fragment ions and the number of the spectral library secondary fragment ions is calculated according to formula (3) for the current reserved secondary fragment ion; the score of the consistency of the peak top between the secondary fragment ions is calculated according to formula (4) for the current reserved secondary fragment ion; and the deviation score of the measured retention time and the predicted retention time is calculated according to formula (5) for the current reserved secondary fragment ion.

[0053] The relative standard deviation score between the peak areas of the multiple quantitative chromatographic peaks corresponding to the same sample is calculated according to formula (7):

[0054]

[0055] The PARS is the relative standard deviation score between the peak areas of the multiple quantitative chromatographic peaks corresponding to the same sample. For each sample, the relative standard deviation RSD of the peak areas is calculated for the peaks corresponding to the sample in the chromatographic peak group. When the RSD is greater than 50%, it is set to 50%. The median of the RSDs of all samples is calculated to obtain the RSD CPG When a plurality of groups of samples are analyzed only once, respectively, the score is not used.

[0056] The score of the ratio between the number of the quantitative chromatographic peaks in the chromatographic peak group and the number of the liquid chromatography-mass spectrometry maps is calculated according to formula (8):

[0057]

[0058] The VPQS is the score of the ratio between the number of the quantitative chromatographic peaks in the chromatographic peak group and the number of the liquid chromatography-mass spectrometry maps; n runs-2 is the number of the quantitative chromatographic peaks in the chromatographic peak group, and n runs is the number of the liquid chromatography-mass spectrometry maps.

[0059] Preferably, in step (8), the method for specifying the polypeptide level score threshold value is that, after steps (2)-(7) are repeated for the false peptide spectral library, the polypeptide level scores of the target polypeptides and the false peptides are combined and sorted, the polypeptide level false discovery rate corresponding to each score threshold value is calculated, and the polypeptide level score threshold value is specified according to the false discovery rate requirement. The false discovery rate requirement ranges from 0.1% to 10%.

[0060] The application also provides an application of the polypeptide quantitative analysis method for integrating multiple liquid chromatography-mass spectrometry maps in protein analysis. The relative content of the protein between samples is calculated according to the relative content of the identified polypeptides between samples.

[0061] Specifically, the method for calculating the relative content of the protein between samples is as follows: in order to prevent a protein from being quantified by a polypeptide that does not actually exist due to a higher polypeptide level score caused by random factors when one protein corresponds to multiple polypeptides, the formula (9) is used to calculate the false discovery rate (FDR) of the polypeptide level when each protein is quantified peptide ),

[0062] FDR protein =1- (1-FDR peptide ) n (9)

[0063] Wherein n is the number of polypeptides corresponding to the protein, FDR protein is the set protein level false discovery rate requirement, the corresponding polypeptide level score threshold is obtained from FDR peptide , the median of the sample relative content (i.e. the ratio of the content) of the polypeptide with a chromatographic peak group score greater than the threshold is obtained to obtain the sample relative content of the protein, and the sum of the polypeptide level scores of these polypeptides is used as the quantification reliability score of each protein.

[0064] The application also provides an electronic device comprising at least a memory and a processor, wherein the memory stores a program, and the processor realizes the integrated polypeptide quantification analysis method of multiple liquid chromatography-mass spectrometry maps when executing the program stored in the memory.

[0065] Compared with the prior art, the application has the following beneficial effects:

[0066] (1) The application greatly improves the quantification accuracy and reduces the quantification error rate by changing the data processing mode without additional experiments under the premise that the liquid chromatography analysis process and data remain unchanged.

[0067] (2) The application takes the best matching peak actually measured as a reference, and the part of the quantification ion screening method does not depend on the spectrum library, and has good tolerance for the error of the intensity of a small number of fragment ions in the spectrum library.

[0068] (3) The improvement of the quantification accuracy of the method in the application is not at the expense of the reduction of the quantification number.

[0069] (4) The process of screening the quantification chromatographic peak and the quantification ion in the application is completely realized automatically by software without manual operation of professional personnel, and the algorithm parameters involved can be applied to different types of liquid chromatography data sets stably without special optimization of parameters. BRIEF DESCRIPTION OF DRAWINGS

[0070] Figure 1The error rate statistics chart of the ratio of polypeptides of three species measured in Example 1 by the method of the present application, the DIA-NN method and the AvG method, and the p value is obtained by chi-square test on the accurate and inaccurate quantification numbers of the software, indicating that there is a significant difference in the error rate.

[0071] Figure 2 The box plot of the correlation coefficient r between the ratio of polypeptide content and the ratio of theoretical content measured in Example 2 by the method of the present application, the DIA-NN method and Spectronaut. The middle horizontal line of the box plot represents the median, the upper and lower boundaries of the box represent the upper and lower quartiles, the cross represents the mean, the whisker represents 1.5 times the interquartile range, and the dot represents the outlier.

[0072] Figure 3 The box plot of the standard deviation (SD) of the log2-transformed (log2) ratio of polypeptides corresponding to the same protein in Example 3 by the method of the present application and Spectronaut. The middle horizontal line of the box plot represents the median, the upper and lower boundaries of the box represent the upper and lower quartiles, the cross represents the mean, the whisker represents 1.5 times the interquartile range, and the dot represents the outlier. DETAILED DESCRIPTION

[0073] The present application will be further illustrated below in conjunction with the embodiments and the accompanying drawings. It should be understood that these embodiments are only used to illustrate the present application, and are not used to limit the scope of the present application.

[0074] Example 1

[0075] In this embodiment, the polypeptide quantification analysis method for integrating multiple liquid chromatography-mass spectrometry maps is used to analyze the proteome digests, and the method comprises the following steps:

[0076] (1) Using data-independent acquisition (DIA), liquid chromatography-mass spectrometry technology is used to analyze each of the two samples A and B three times (64 isolation windows are used in liquid chromatography-mass spectrometry analysis, and the mass-to-charge ratio window width is different), to obtain 6 liquid chromatography-mass spectrometry maps as the original liquid chromatography-mass spectrometry data; wherein the two samples A and B are mixed from the proteome digests of 3 species (human, E. coli, and yeast), and the mass fractions of the proteome digests of human, E. coli, and yeast in sample A are 67%, 3%, and 30%, respectively, and the mass fractions of the proteome digests of human, E. coli, and yeast in sample B are 67%, 30%, and 3%, respectively.

[0077] The above steps have been completed by the literature (Navarro, P. et al. A multicenter study benchmarks software tools for label-free proteome quantification. Nat. Biotechnol., 34, 1130-1136, 2016) and given data, the dataset download link is: http: / / proteomecentral.proteomexchange.org / cgi / GetDataset?ID=PXD002952.

[0078] Liquid chromatography-mass spectrometry analysis was performed in data-dependent acquisition (DDA) mode, and the spectral library of the target polypeptide was constructed (the download link of the spectral library is:

[0079] https: / / ftp.pride.ebi.ac.uk / pride / data / archive / 2016 / 09 / PXD002952 / ecolihuman yeast_concat_mayu_IRR_cons_openswath_64var_curated.csv). According to each polypeptide in the spectral library, the following steps (2)-(7) were performed respectively.

[0080] (2) According to the data in the spectral library constructed in the above steps, the measured retention time of the reference polypeptide (11 polypeptides with high content in the sample added additionally before liquid chromatography-mass spectrometry analysis) in each liquid chromatography-mass spectrometry spectrum was taken as the dependent variable, the relative retention time (iRT value) in the spectral library was taken as the independent variable, and the iRT value of the target polypeptide was taken as the sample point. Linear interpolation was performed to obtain the predicted retention time of the polypeptide. According to the mass-to-charge ratio of the parent ion and fragment ion of each polypeptide in the spectral library, the allowed deviation range of the retention time was set to 10 min, and the allowed deviation range of the mass-to-charge ratio was set to 30 ppm. All extracted ion chromatograms of the secondary fragment ions within the above range were extracted from the original liquid chromatography-mass spectrometry data, and the chromatographic peaks obtained by integrating the extracted ion chromatograms were taken as candidate chromatographic peaks.

[0081] (3) For each candidate chromatographic peak, the consistency of the secondary fragment ion chromatographic peak type of the same component is used to preliminarily screen the secondary fragment ions, and the secondary fragment ions corresponding to the chromatographic peaks with inconsistent peak types are excluded. Specifically, when the peak top of a certain chromatographic peak is outside the overall range of the chromatographic peak, the secondary fragment ion corresponding to the chromatographic peak is directly excluded; when the peak top of a certain chromatographic peak is within the overall range of the chromatographic peak, the ranking (from large to small) of the intensity of the chromatographic peak at the overall peak top time point in the intensity of the ion at each time point within the chromatographic peak range is calculated for each secondary fragment ion, and the secondary fragment ion with a ratio greater than 0.8 is excluded. Among them, the starting time of each secondary fragment ion peak in the chromatographic peak is taken as the starting time of the overall chromatographic peak, and the ending time of each secondary fragment ion peak is taken as the ending time of the overall chromatographic peak. The range from the starting time to the ending time of the overall chromatographic peak is referred to as the overall range of the chromatographic peak.

[0082] The measured peak area of the retained secondary fragment ion is used to calculate the measured relative intensity, and the relative intensity of the corresponding secondary fragment ion in the spectral library is used as a reference to exclude secondary fragment ions with large relative intensity deviations; specifically, the root mean square error RMSE between the measured relative intensity of the secondary fragment ion and the relative intensity of the corresponding secondary fragment ion in the spectral library is first calculated, and then the root mean square error RMSE between the measured relative intensity of the secondary fragment ion after being temporarily removed and the relative intensity of the corresponding secondary fragment ion in the spectral library is calculated -1 If the RMSE -1 of a certain secondary fragment ion is less than RMSE-0.05, the secondary fragment ion is excluded.

[0083] (4) According to the principle that multiple determinations are closer to the true value, it is easier to find measured data similar to the spectral library in multiple liquid chromatography-mass spectrometry spectra, and based on this, the spectral library matching score of the candidate chromatographic peak is calculated. In all candidate chromatographic peaks corresponding to multiple liquid chromatography-mass spectrometry spectra, the candidate chromatographic peak with the highest spectral library matching score is selected as the best matching peak.

[0084] The spectral library matching score is obtained by weighted geometric mean of multiple single scores; the single scores include: the similarity score of the peak area of each secondary fragment ion to the spectral library intensity after normalization, the deviation score between the measured mass-to-charge ratio of each secondary fragment ion and the theoretical mass-to-charge ratio, the ratio score of the number of retained measured secondary fragment ions to the number of spectral library secondary fragment ions, the consistency score of the peak top of the secondary fragment ions, the deviation score of the measured retention time and the predicted retention time, and the sum score of the peak area of the secondary fragment ions; the weight coefficients of each single score are all 1;

[0085] Specifically, the similarity score of the peak area of each secondary fragment ion to its spectral library intensity is calculated according to formula (1); the deviation score between the measured mass-to-charge ratio and the theoretical mass-to-charge ratio of each secondary fragment ion is calculated according to formula (2); the score of the ratio of the number of the measured secondary fragment ions to the number of the spectral library secondary fragment ions is calculated according to formula (3); the score of the consistency of the peak top between the secondary fragment ions is calculated according to formula (4); the deviation score of the measured retention time and the predicted retention time is calculated according to formula (5); and the score of the sum of the peak areas of the secondary fragment ions is calculated according to formula (6).

[0086] (5) According to the best matching peak, find the corresponding chromatographic peak of the component in the remaining liquid chromatography-mass spectrometry (LC-MS) spectrum (other than the LC-MS spectrum where the best matching peak is located). Specifically, utilize the consistency of the retention time of the same component in multiple LC-MS analyses, take the retention time of the best matching peak as the center, narrow the allowed deviation of the retention time, exclude candidate chromatographic peaks with large retention time deviation, and retain candidate chromatographic peaks with measured retention time within the allowed deviation range. The calculation method of the retention time of the best matching peak aligned to the retention time of each LC-MS spectrum is as follows: take the measured retention time of the reference polypeptides (11 polypeptides with higher content added to the sample before LC-MS analysis) in each LC-MS spectrum as the dependent variable, take the measured retention time of the reference polypeptides in the LC-MS spectrum where the best matching peak is located as the independent variable, take the retention time of the best matching peak as the sample point, and interpolate by linear interpolation or spline interpolation to obtain the aligned retention time. The allowed deviation of the retention time is 0.27 min (obtained by leave-one-out cross-validation of the retention times of the above-mentioned 11 polypeptides).

[0087] Then, utilize the consistency of the ion mass-to-charge ratio and its relative intensity between multiple LC-MS analyses to calculate the similarity score between the retained candidate chromatographic peak and the best matching peak. In the candidate chromatographic peak corresponding to each LC-MS spectrum, select the candidate chromatographic peak with the largest similarity score as the measured chromatographic peak of the component selected in each determination.

[0088] The similarity score between the candidate chromatographic peak and the best matching peak is obtained by weighted geometric mean or weighted arithmetic mean of multiple single scores; the single scores include: the similarity score of the peak area of each secondary fragment ion to the intensity of the best matching peak, the deviation score between the measured mass-to-charge ratio and the theoretical mass-to-charge ratio of each secondary fragment ion, the ratio score of the number of the measured secondary fragment ions to the number of the ions of the best matching peak, and the consistency score of the peak top between the secondary fragment ions; and the weight coefficient of each single score ranges from 1. If the largest similarity score in a certain LC-MS spectrum does not reach the critical value requirement of 0.5 or does not reach the value obtained by subtracting 0.2 from the largest similarity score of the other LC-MS spectrum, it is considered that the component is not measured in this LC-MS spectrum.

[0089] Specifically, the similarity score of the area of each secondary fragment ion peak to the intensity of the best matching peak is calculated for the current retained secondary fragment ion according to formula (1), SRSS represents the similarity score of the area of the secondary fragment ion peak to the intensity of the best matching peak, RMSE1 takes the value of the root mean square error between the normalized measured relative intensity of the secondary fragment ion and the relative intensity of the corresponding secondary fragment ion in the best matching peak; the deviation score between the measured mass-to-charge ratio of each secondary fragment ion and the theoretical mass-to-charge ratio is calculated for the current retained secondary fragment ion according to formula (2); the score of the ratio of the number of retained measured secondary fragment ions to the number of ions in the best matching peak is calculated for the current retained secondary fragment ion according to formula (3), FQS represents the score of the ratio of the number of retained measured ions to the number of ions in the best matching peak, n F The score of the consistency of the peak top between the secondary fragment ions is calculated for the current retained secondary fragment ion according to formula (4).

[0090] The measured chromatographic peaks and the best matching peaks are combined into a chromatographic peak group, if the maximum similarity score in a certain liquid chromatography-mass spectrometry (LC-MS) spectrum does not reach the critical value requirement, it is considered that the target polypeptide is not detected in the LC-MS spectrum; the consistency of the secondary fragment ions of the same component measured in multiple LC-MS analyses is used to exclude the chromatographic peaks with large differences from the secondary fragment ions from the chromatographic peak group, and the retained chromatographic peaks are used as quantitative chromatographic peaks for final quantitative calculation, and the different secondary fragment ions between the retained quantitative chromatographic peaks are excluded;

[0091] (6) The consistency of the relative intensity of the ions between multiple LC-MS analyses is used to finally screen the quantitative ions, specifically, the proportion of the measured areas of each secondary fragment ion between the quantitative chromatographic peaks corresponding to each LC-MS analysis is calculated, the proportion is logarithmically converted (log2) with 2 as the base number, and then weighted and averaged with the intensity of each secondary fragment ion in the spectral library as the weight coefficient, if it exceeds the mean value ±1, it is an outlier, the secondary fragment ion corresponding to the outlier is excluded, the correlation of the measured areas of the retained secondary fragment ions in multiple quantitative chromatographic peaks is calculated, and the secondary fragment ions with low correlation are excluded, and the retained secondary fragment ions are used as quantitative ions;

[0092] Specifically, the peak areas of each secondary fragment ion retained in the quantitative chromatographic peak are used as vectors, the cosine of the angle between each two vectors represents the correlation between the measured areas of different secondary fragment ions in multiple quantitative chromatographic peaks, and the cosine distance, i.e. 1 minus the cosine of the angle, is used to perform cluster analysis on the secondary fragment ions, first, the two secondary fragment ions with the smallest cosine distance are retained, and other secondary fragment ions are retained according to whether the cosine distance between them and the already retained secondary fragment ions is within a critical value, and the unretained ions are excluded; the above critical value is the maximum of 0.05 and twice the distance between the already retained ions.

[0093] (7) According to the quantitative chromatographic peaks screened out from the chromatographic peak group and the quantitative ions, the chromatographic peak group score is calculated, and the polypeptide level score is obtained by weighted geometric mean of the chromatographic peak group score and the highest spectral library matching score, wherein the weight coefficients are both 1.

[0094] The chromatographic peak group score is obtained by weighted geometric mean or weighted arithmetic mean of multiple single scores; the single scores include: the score of the similarity between the normalized error root mean square of each secondary fragment ion peak area and the spectral library intensity, the deviation score between the measured mass-to-charge ratio and the theoretical mass-to-charge ratio of each secondary fragment ion, the score of the ratio between the number of retained measured secondary fragment ions and the number of spectral library secondary fragment ions, the chromatographic peak top consistency score between secondary fragment ions, the deviation score between the measured retention time and the predicted retention time, the relative standard deviation score between the peak areas of multiple quantitative chromatographic peaks corresponding to the same sample, and the score of the ratio between the number of quantitative chromatographic peaks in the chromatographic peak group and the number of liquid chromatography-mass spectrometry maps, and the weight coefficient of each single score ranges from 0.01 to 100;

[0095] Specifically, in step (7), the score of the similarity between the normalized error root mean square of each secondary fragment ion peak area and the spectral library intensity is calculated according to formula (1) for the current retained secondary fragment ion; the deviation score between the measured mass-to-charge ratio and the theoretical mass-to-charge ratio of each secondary fragment ion is calculated according to formula (2) for the current retained secondary fragment ion; the score of the ratio between the number of retained measured secondary fragment ions and the number of optimal matching peaks is calculated according to formula (3) for the current retained secondary fragment ion; the chromatographic peak top consistency score between secondary fragment ions is calculated according to formula (4) for the current retained secondary fragment ion; the deviation score between the measured retention time and the predicted retention time is calculated according to formula (5) for the current retained secondary fragment ion; the relative standard deviation score between the peak areas of multiple quantitative chromatographic peaks corresponding to the same sample is calculated according to formula (7); and the score of the ratio between the number of quantitative chromatographic peaks in the chromatographic peak group and the number of liquid chromatography-mass spectrometry maps is calculated according to formula (8).

[0096] (8) Repeat steps (2)-(7) for the decoy peptide library (download link ftp: / / massive.ucsd.edu / MSV000085540 / quant / LFQBench / Skyline / HYE110_6600_64Var_Skyline_PublishedPB.sky.zip, export the library from Skyline software), combine the peptide level scores of the target polypeptides and the decoy polypeptides, and sort them, and calculate the peptide level false discovery rate corresponding to each score threshold. According to the requirement of 1% of the peptide level false discovery rate, the polypeptides with peptide level scores higher than the corresponding score threshold (0.85088) are identified polypeptides, and the polypeptides of human, E. coli, and yeast are 15823, 8928, and 8450, respectively (according to the library data, the processing method of the software, and the same parent ion charge number of the same polypeptide in the library, count respectively, the same below).

[0097] (9) Calculate the relative content of the identified polypeptides in the sample according to the area of each peak in the sample peak group. In order to show the advantage of the method of the present application in quantitative accuracy, compare the quantitative results of DIA-NN (download link https: / / files.osf.io / v1 / resources / 6g3ux / providers / osfstorage / 5dc188e896af97000f1942c7?action=download&direct&version=1) and AvG (download link https: / / static-content.springer.com / esm / art%3A10.1038%2Fs41592-020-00986-4 / MediaObjects / 41592_2020_986_MOESM3_ESM.zip) two software. Take the relative error of the content ratio between samples A and B and the theoretical content ratio > 30% as the judgment basis of error (the same below). DIA-NN gets the content ratio between samples A and B, and the number of polypeptides accurately determined in human, E. coli, and yeast is 14677, 3369, and 2665, respectively, and the number of polypeptides incorrectly determined is 1652, 1820, and 1196, respectively, and the error rate is 10.1%, 35.1%, and 31.0%, respectively. Figure 1 ) Remove the polypeptides with low peptide level scores in the quantitative results of the method of the present application, and when the number of accurately determined polypeptides is the same as DIA-NN, the number of incorrectly determined polypeptides is 1527, 1271, and 672, respectively, and the error rate is 9.4%, 27.4%, and 20.1%, respectively. Figure 1The accuracy was significantly lower than that of DIA-NN, demonstrating the quantitative advantage of the method of this invention. AvG (using some functions of OpenSWATH software in its data processing) obtained the ratio of content between samples A and B. The number of accurately measured peptides in humans, E. coli, and yeast were 12096, 1611, and 1051, respectively, while the number of incorrectly measured peptides were 919, 593, and 432, respectively, with error rates of 7.1%, 26.9%, and 29.1%. Figure 1 By removing peptides with low peptide level scores from the quantitative results of the method of this invention, and ensuring that the number of accurate measurements is the same as that of AvG, the number of incorrect measurements were 925, 240, and 164, respectively, with error rates of 7.1%, 13.0%, and 13.5%. Figure 1 The result was significantly lower than that of AvG, which demonstrates the advantage of the present invention in terms of quantitative accuracy.

[0098] Furthermore, the relative content of proteins among samples is calculated based on the relative content of peptides identified in step (9). Specifically, the calculation method is as follows: to prevent the incorrect quantification of a protein when one protein corresponds to multiple peptides, where a peptide that does not actually exist receives a high peptide level score due to random factors, formula (9) is used to calculate the required peptide level false detection rate for each protein quantification. In this embodiment, FDR protein The value is 1%. Based on the correlation between the score threshold obtained in step (8) and the false discovery rate at the peptide level, from FDR... peptide The corresponding peptide level score threshold is obtained. The median of the relative content (i.e., the ratio of content) of peptides with peptide level scores greater than the threshold is used to obtain the relative content of proteins among samples. At the same time, the sum of the peptide level scores of these peptides is used as the quantitative reliability score of each protein.

[0099] To demonstrate the advantage of the method of this invention in protein quantification accuracy, the quantification results of DIA-NN were compared (AvG does not provide protein quantification results, so they are not compared here). Under the premise that at least two peptides for each protein are measured, DIA-NN yielded the following ratios of protein content between samples A and B: 1903, 488, and 427 proteins were accurately measured in humans, E. coli, and yeast, respectively; 47, 128, and 123 proteins were incorrectly measured, with error rates of 2.4%, 20.8%, and 22.4%, respectively. Under the premise that at least two peptides for each protein are measured, proteins with low reliability scores in the quantification results of the method of this invention were removed, so that the number of accurate measurements was the same as that of DIA-NN. The number of incorrect measurements was 30, 118, and 15, with error rates of 1.6%, 19.5%, and 3.4%, respectively, significantly lower than that of DIA-NN. This proves that the method of this invention still has an advantage in quantification accuracy at the protein level.

[0100] Example 2

[0101] This example analyzes proteome digests using the polypeptide quantification method described above, which includes the following steps:

[0102] (1) Using data-independent acquisition (DIA) and liquid chromatography-mass spectrometry technology, each of the Y1, Y2, Y3, and Y4 samples was analyzed three times (64 isolation windows were used in the DIA liquid chromatography-mass spectrometry analysis, the mass-to-charge ratio range was 400-1200 m / z, each isolation window was 13.5 m / z wide, and the window overlap was 1 m / z). Twelve liquid chromatography-mass spectrometry maps were obtained as the original liquid chromatography-mass spectrometry data. The Y1, Y2, Y3, and Y4 samples were each mixed from the proteome digests of two species (human and yeast). The human proteome digest was 800 ng in each of the four samples (serving as the background component for the determination). The yeast proteome digest was 200, 100, 50, and 25 ng in the Y1, Y2, Y3, and Y4 samples, respectively. That is, the theoretical ratio of the amount of polypeptides derived from yeast in the samples was Y2 / Y1 = 0.5, Y3 / Y1 = 0.25, and Y4 / Y1 = 0.125.

[0103] The measured data set can be downloaded from the PRIDE website (ID PXD039399). Liquid chromatography-mass spectrometry analysis of the yeast proteome digest was performed using data-dependent acquisition (DDA) mode, and the spectral library of the yeast target polypeptide was measured (which can be downloaded from the PRIDE website using the ID above). According to each polypeptide in the spectral library, the following steps (2)-(7) were performed, respectively.

[0104] (2) According to the spectral library data, the measured retention time of each liquid chromatogram of the reference polypeptide (11 polypeptides with high content contained in the sample: DSYVGDEAQSK, A GFAGDDAPR, ATAGDTHLGGEDFDNR, VATVSLPR, ELISNASDALDK, TTPSYVAFTDTER, VCENIPIVLCGNK, LGEHNIDVLEGNEQFINAAK, SYELPDGQVITIGNER, YFPTQALNFAFK, DSTLIMQLLR) is the dependent variable, and their iRT (relative retention time) values are the independent variables. The iRT value of the target polypeptide is the sample point, and the predicted retention time of the polypeptide is obtained after linear interpolation. According to the mass-to-charge ratio of the parent ion and fragment ion of each polypeptide in the spectral library, the allowed deviation range of the retention time is set to 5 min, and the allowed deviation range of the mass-to-charge ratio is set to 20 ppm. Extract the extracted ion chromatogram of all secondary fragment ions within the above range from the original liquid chromatography data, and the chromatographic peaks therein are used as candidate chromatographic peaks and are integrated.

[0105] Steps (3)-(7) in this example are the same as in Example 1, except that in step (5), the allowed deviation of the retention time is 0.92 min (obtained by leave-one-out cross-validation of the retention times of the above-mentioned 11 polypeptides).

[0106] (8) After repeating steps (2)-(7) on the decoy peptide spectral library (changing the order of amino acids in the target polypeptide in the spectral library, which can be downloaded from the PRIDE website using the above-mentioned ID), the polypeptide level scores of the target polypeptide and the decoy peptide are combined and sorted, and the polypeptide level false discovery rate corresponding to each score threshold is calculated. According to the requirement of 1% polypeptide level false discovery rate, the polypeptide with a polypeptide level score higher than the corresponding score threshold (0.84490) is identified as the polypeptide. A total of 8512 polypeptides of yeast were obtained (according to the spectral library data, the processing method of the software used for comparison, and the same mother ion charge number of the same polypeptide in the spectral library, respectively counted, the same below).

[0107] (9) The relative content of the identified peptides among the samples was calculated based on the area of ​​each peak in the chromatographic peak group. To demonstrate the advantage of the method of the present invention in terms of quantitative accuracy, the quantitative results of DIA-NN and Spectronaut software were compared respectively. The relative error of the ratio of the measured content among samples to the ratio of the theoretical content (Y2 / Y1, Y3 / Y1, Y4 / Y1) >30% was used as the criterion for judging measurement error (the same below). While examining the relative error of the ratio of the measured content to the ratio of the theoretical content, the correlation coefficient between the ratio of the measured content and the ratio of the theoretical content was used to examine the overall quantitative accuracy. In order to prevent the correlation coefficient used in subsequent comparisons from being distorted by missing values, peptides with missing values ​​in the ratios of content (Y2 / Y1, Y3 / Y1, Y4 / Y1) were excluded before comparison. DIA-NN yielded 6405, 4969, and 3764 accurately measured yeast polypeptides (Y2 / Y1, Y3 / Y1, Y4 / Y1), respectively, and 923, 2359, and 3564 incorrectly measured, respectively, with error rates of 12.6%, 32.2%, and 48.6%. The distribution of the correlation coefficient r between the measured ratios and the theoretical ratios was also shown below. Figure 2 As shown, the mean value was 0.9642, with 431 peptides having r < 0.95. After removing peptides with low level scores from the quantitative results of this invention, and making the number of measurements the same as DIA-NN, the accurate measurements were 6649, 5329, and 4023, respectively, and the incorrect measurements were 679, 1999, and 3305, respectively, with error rates of 9.3%, 27.3%, and 45.1%, significantly lower than DIA-NN. Simultaneously, the distribution of the correlation coefficient r between the measured content ratio and the theoretical content ratio is shown in the figure. Figure 2 As shown, the mean value was 0.9748, and only 288 peptides had r < 0.95, demonstrating the quantitative accuracy advantage of the method of this invention. Spectronaut obtained 4136, 3641, and 2951 accurately measured yeast peptides (Y2 / Y1, Y3 / Y1, Y4 / Y1), respectively, and 363, 858, and 1548 incorrectly measured, respectively, with error rates of 8.1%, 19.1%, and 34.4%. Simultaneously, the distribution of the correlation coefficient r between the measured content ratio and the theoretical content ratio is shown in the figure. Figure 2 As shown, the mean value was 0.9757, with 144 peptides having r < 0.95. After removing peptides with low level scores from the quantitative results of this invention, and making the number of measurements the same as Spectronaut, the accurate measurements were 4268, 3659, and 2893, respectively, and the incorrect measurements were 231, 840, and 1606, respectively, with error rates of 5.1%, 18.7%, and 35.7%, respectively. Overall, these error rates were lower than Spectronaut's. Furthermore, the distribution of the correlation coefficient r between the measured content ratio and the theoretical content ratio is shown in the figure.Figure 2 The mean is 0.9853, with only 62 polypeptides having r < 0.95, demonstrating the advantage of the method of the application in terms of quantitative accuracy.

[0108] The method for calculating the relative content of proteins between samples differs from that of Example 1 only in that FDR proteinThe ratio of the measured content to the theoretical content (Y2 / Y1, Y3 / Y1, Y4 / Y1) was compared. Before comparison, the proteins with missing values in the ratio of the content (Y2 / Y1, Y3 / Y1, Y4 / Y1) were excluded to prevent the correlation coefficient from being distorted by missing values. DIA-NN accurately determined the number of proteins (Y2 / Y1, Y3 / Y1, Y4 / Y1) in yeast as 1388, 983, and 679, respectively, and incorrectly determined the number of proteins as 125, 530, and 834, respectively, with error rates of 8.3%, 35.0%, and 55.1%, respectively. At the same time, the correlation coefficient r between the measured content ratio and the theoretical content ratio was 0.9679 on average, of which 121 proteins had r < 0.95. When the number of proteins with low reliability scores in the quantitative results of the present method was removed, the number of accurate determinations was 1429, 1163, and 870, respectively, and the number of incorrect determinations was 84, 350, and 643, respectively, with error rates of 5.6%, 23.1%, and 42.5%, respectively, which was significantly lower than DIA-NN. At the same time, the correlation coefficient r between the measured content ratio and the theoretical content ratio was 0.9757 on average, of which only 49 proteins had r < 0.95, which proved the advantage in quantitative accuracy. Spectronaut accurately determined the number of proteins (Y2 / Y1, Y3 / Y1, Y4 / Y1) in yeast as 978, 880, and 756, respectively, and incorrectly determined the number of proteins as 128, 226, and 350, respectively, with error rates of 11.6%, 20.4%, and 31.6%, respectively. At the same time, the correlation coefficient r between the measured content ratio and the theoretical content ratio was 0.9625 on average, of which 80 proteins had r < 0.95. When the number of proteins with low reliability scores in the quantitative results of the present method was removed, the number of accurate determinations was 1079, 918, and 724, respectively, and the number of incorrect determinations was 27, 188, and 382, respectively, with error rates of 2.4%, 17.0%, and 34.5%, respectively, which was lower than Spectronaut as a whole. At the same time, the correlation coefficient r between the measured content ratio and the theoretical content ratio was 0.9932 on average, of which only 12 proteins had r < 0.95, which proved that the present method still had an advantage in quantitative accuracy at the protein level.

[0109] Example 3

[0110] The present embodiment uses the polypeptide quantitative analysis method integrating multiple liquid chromatography-mass spectrometry maps to analyze proteome digests, which comprises the following steps:

[0111] (1) Using data-independent acquisition (DIA) mode, liquid chromatography-mass spectrometry technology is used to analyze multiple samples (33 isolation windows are used in DIA liquid chromatography-mass spectrometry analysis, and the mass-to-charge ratio window width is different), and each sample is analyzed only once (i.e. no repetition), 60 liquid chromatography-mass spectrometry maps are obtained as the original liquid chromatography-mass spectrometry data; wherein the samples are 60 cerebrospinal fluids collected in Sweden by literature (Bader, J. M. et al. Proteome profiling in cerebrospinal fluid reveals novel biomarkers of Alzheimer's disease. Mol. Syst. Biol., 16, 1-17, 2020), which are divided into Alzheimer's disease group (29 samples) and normal group (31 samples).

[0112] The above steps have been completed by the literature (Bader, J. M. et al. Proteome profiling in cerebrospinal fluid reveals novel biomarkers of Alzheimer's disease. Mol. Syst. Biol., 16, 1-17, 2020) and the data is given, and the data set download link is https: / / www.ebi.ac.uk / pride / archive / projects / PXD016278.

[0113] Liquid chromatography-mass spectrometry analysis is performed using data-dependent acquisition (DDA) mode, and the spectral library of the target polypeptide is measured (the download link of the spectral library is:

[0114] https: / / ftp.pride.ebi.ac.uk / pride / data / archive / 2020 / 05 / PXD016278 / final%20hy brid%20library_used%20for%20main%20study%20and%20CV%20experiment.zip), and according to each polypeptide in the spectral library, the following steps (2)-(7) are performed respectively;

[0115] (2) According to the data in the spectral library constructed in the above steps, taking the measured retention time of each polypeptide (8 polypeptides with high content contained in the sample: NLGKVGSK, ETCFAEEGKK, AAFTECCQAADK, AACLLPK, KVPQVSTPTLVEVSR, VPQVSTPTLVEVSR, LVAASQAALGL, HPYFYAPELLFFAKR) in each liquid chromatography-mass spectrometry spectrum as the dependent variable, the relative retention time (iRT value) as the independent variable, and the iRT value of the target polypeptide as the sample point, the predicted retention time of the polypeptide was obtained after linear interpolation. According to the mass-to-charge ratio of the parent ions and fragment ions of each polypeptide in the spectral library, the allowed deviation range of the retention time was set to ±10 min, and the allowed deviation range of the mass-to-charge ratio was set to 20 ppm. All the extracted ion chromatograms of the secondary fragment ions within the above range were extracted from the original liquid chromatography-mass spectrometry data, and the chromatographic peaks in the extracted ion chromatograms were taken as candidate chromatographic peaks and were integrated.

[0116] Steps (3) to (7) in this example are the same as in Example 1, except that in step (5), the allowed deviation of the retention time is 1.82 min (obtained by leave-one-out cross-validation of the retention times of the above-mentioned 8 polypeptides), and in step (7), the relative standard deviation between the peak areas of the multiple quantitative chromatographic peaks corresponding to the same sample is not included in the calculation of the overall calculated polypeptide level score.

[0117] (8) The relative contents of the identified polypeptides in the samples were calculated according to the areas of the peaks in the chromatographic peak group. Since the samples in this example are actual clinical samples, the true contents and proportions of the polypeptides and proteins are unknown, so it is not possible to directly judge whether the quantification is accurate or not. However, it is possible to judge whether the polypeptide quantification is accurate or not according to the principle that the relative contents of multiple polypeptides produced by the same protein are consistent between samples. Specifically, the ratios of the multiple polypeptides corresponding to the same protein between the Alzheimer's disease group and the normal group were measured, and the standard deviation (SD) was calculated after logarithmic conversion (log2) with 2 as the base number, to represent the consistency of the polypeptide quantification results. At the same time, the number of proteins for which the corresponding multiple polypeptides were both up-regulated (ratio greater than 1) and down-regulated (ratio less than 1) was counted, to represent the inconsistency of the quantification results. In addition, since the above-mentioned indicators for comparing the accuracy of quantification do not involve the false discovery rate of polypeptide levels, the decoy polypeptides do not need to be calculated in this example.

[0118] (9) In order to show the advantage of the method of the present application in terms of quantification accuracy, the quantification results of Spectronaut were compared (download link:

[0119] https: / / ftp.pride.ebi.ac.uk / pride / data / archive / 2020 / 05 / PXD016278 / full%20report_extra%20data.zip). The number of proteins measured by Spectronaut (only counting proteins corresponding to multiple peptides) was 967, among which 724 proteins corresponding to multiple peptides were both up-regulated and down-regulated, and the SD distribution of the peptides was as shown in Figure 3 The number of proteins measured by Spectronaut (only counting proteins corresponding to multiple peptides) was 967, among which 724 proteins corresponding to multiple peptides were both up-regulated and down-regulated, and the SD distribution of the peptides was as shown in

[0120] The above-described embodiments of the technical solutions of the present application are described in detail, and it should be understood that the above-described is only a specific embodiment of the present application and is not used to limit the present application. Any modification, supplement or similar replacement within the principle range of the present application shall be included in the protection scope of the present application.

Claims

1. A method for quantitative analysis of peptides integrating multiple liquid chromatography-mass spectra, characterized in that, include: A data-independent acquisition mode was adopted, and liquid chromatography-mass spectrometry (LC-MS) was used to analyze each sample multiple times, or to analyze multiple groups of samples once or multiple times. Ion chromatograms of secondary fragment ions were extracted from the raw LC-MS data to obtain multiple LC-MS chromatograms. The consistency of retention time, ion mass-to-charge ratio, and relative intensity of the same component among multiple analyses was analyzed using LC-MS. At the same time, the multiple LC-MS chromatograms were integrated and processed to exclude inconsistent chromatographic peaks and secondary fragment ions among multiple analyses. The relative content of identified peptides among samples was calculated using the screened quantitative chromatographic peaks and quantitative ions. The integration process includes the following steps: For each candidate chromatographic peak, select the candidate chromatographic peak with the highest library matching score from multiple liquid chromatography-mass spectrometry (LC-MS) chromatograms as the best matching peak; using the best matching peak as a reference, calculate the similarity score between the candidate chromatographic peak and the best matching peak; among the candidate chromatographic peaks corresponding to each LC-MS chromatogram, select the candidate chromatographic peak with the highest similarity score as the measured chromatographic peak; combine the measured chromatographic peak with the best matching peak to form a chromatographic peak group; exclude chromatographic peaks containing secondary fragment ions that differ significantly from other chromatographic peaks from the chromatographic peak group; retain the chromatographic peaks as quantitative chromatographic peaks for final quantitative calculation; and exclude different secondary fragment ions among the retained quantitative chromatographic peaks. The consistency of ion relative intensities across multiple liquid chromatography-mass spectrometry (LC-MS) analyses was used to screen for quantitative ions.

2. The method for quantitative analysis of peptides integrating multiple liquid chromatography-mass spectra according to claim 1, characterized in that, Specifically, the following steps are included: (1) Using a data-independent acquisition mode, liquid chromatography-mass spectrometry is used to analyze each sample multiple times, or multiple groups of samples are analyzed once or multiple times to obtain raw liquid chromatography-mass spectrometry data. (2) Based on the data in the spectral library, predict the retention time of each target peptide in the sample, extract the scan data of the target peptide with the predicted retention time within the allowable deviation range from each liquid chromatography raw data, and obtain the extracted ion chromatogram of all secondary fragment ions corresponding to the target peptide within the allowable deviation range of mass-to-charge ratio. The chromatographic peak obtained by integrating the extracted ion chromatogram is used as the candidate chromatographic peak. (3) For each candidate chromatographic peak, the consistency of the peak shape of the secondary fragment ions of the same component is used to screen and exclude secondary fragment ions with inconsistent peak shapes in the candidate chromatographic peaks. Then, the measured peak area of ​​the remaining secondary fragment ions is calculated and the measured relative intensity is obtained. The relative intensity of the corresponding secondary fragment ions in the spectral library is used as a reference to exclude secondary fragment ions with large relative intensity deviations. (4) Based on the principle that multiple measurements are closer to the true value, it is easier to find measured data that are similar to the spectral library in multiple liquid chromatography-mass spectra. Based on this, the spectral library matching score of the candidate chromatographic peak is calculated. Among all the candidate chromatographic peaks corresponding to multiple liquid chromatography-mass spectra, the candidate chromatographic peak with the highest spectral library matching score is selected as the best matching peak. (5) Using the best matching peak as a reference, find the chromatographic peak of the target polypeptide in the liquid chromatography-mass spectra other than the liquid chromatography-mass spectra where the best matching peak is located. Specifically, by utilizing the consistency of the retention time of the same component between multiple liquid chromatography analyses, and taking the retention time of the best matching peak aligned with the retention time of each liquid chromatography-mass spectra as the center, reduce the allowable deviation of the retention time and retain the candidate chromatographic peaks whose measured retention time is within the allowable deviation range; then, by utilizing the consistency of the ion mass-charge ratio and its relative intensity between multiple liquid chromatography analyses, calculate the similarity score between the retained candidate chromatographic peaks and the best matching peak. Among the candidate chromatographic peaks corresponding to each liquid chromatography-mass spectra, select the candidate chromatographic peak with the largest similarity score as the measured chromatographic peak, and form a chromatographic peak group with the measured chromatographic peak and the best matching peak. If the largest similarity score in a certain liquid chromatography-mass spectra does not reach the critical value requirement, it is considered that the target polypeptide was not detected in that liquid chromatography-mass spectra. By utilizing the consistency of secondary fragment ions measured for the same component in multiple liquid chromatography-mass analysis, chromatographic peaks containing secondary fragment ions that differ significantly from other chromatographic peaks are excluded from the chromatographic peak group. The retained chromatographic peaks are used as quantitative chromatographic peaks for the final quantitative calculation. At the same time, different secondary fragment ions among the retained quantitative chromatographic peaks are excluded. (6) Utilize the consistency of the relative ion intensities among multiple liquid chromatography analyses to finally screen quantitative ions. Specifically, calculate the ratio between the measured areas of each secondary fragment ion retained between the quantitative chromatographic peaks corresponding to each liquid chromatography analysis. If there are outliers in the ratio, the corresponding secondary fragment ions are excluded. Calculate the correlation between the measured areas of the retained secondary fragment ions in multiple quantitative chromatographic peaks, exclude secondary fragment ions with low correlation, and retain the secondary fragment ions as quantitative ions. (7) Based on the quantitative chromatographic peaks and quantitative ions screened in the chromatographic peak group, calculate the overall score of the chromatographic peak group, and then calculate the weighted geometric mean or weighted arithmetic of the chromatographic peak group score and the highest spectral library matching score to obtain the peptide level score. The peptide level score is used to characterize the reliability of the determination, and the weighting coefficient ranges from 0.1 to 10. (8) Specify a critical value for peptide level score. Target peptides with peptide level scores higher than the corresponding critical value are identified peptides. The critical value range is 0.75~0.

95. (9) The area of ​​each chromatographic peak in the chromatographic peak group is obtained by summing the quantitative ion area. The relative content of the identified polypeptide among the samples is calculated based on the area of ​​each chromatographic peak in the sample chromatographic peak group.

3. The method for quantitative analysis of peptides integrating multiple liquid chromatography-mass spectra according to claim 2, characterized in that, In step (2), the method for predicting the retention time of the target peptide is as follows: the measured retention time of the reference peptide in each liquid chromatography-mass spectra is used as the dependent variable, the relative retention time in the spectral library is used as the independent variable, and the relative retention time of the target peptide is used as the sample point. The predicted retention time of the target peptide is obtained by linear interpolation or spline interpolation. And / or, in step (2), the allowable deviation range of retention time is no more than 10 min, and the allowable deviation range of mass-to-charge ratio is no more than 30 ppm.

4. The method for quantitative analysis of peptides integrating multiple liquid chromatography-mass spectra according to claim 2, characterized in that, In step (3), the method for excluding secondary fragment ions with large relative intensity deviations is as follows: first, calculate the root mean square error (RMSE) between the measured relative intensity of the secondary fragment ion and the relative intensity of the corresponding secondary fragment ion in the spectral library; then, calculate the root mean square error (RMSE) between the measured relative intensity of the secondary fragment ion after it has been temporarily removed and the relative intensity of the corresponding secondary fragment ion in the spectral library. -1 If the RMSE of a certain secondary fragment ion -1 If RMSE < 0.05, then the secondary fragment ion is excluded.

5. The method for quantitative analysis of peptides integrating multiple liquid chromatography-mass spectra according to claim 2, characterized in that, In step (4), the library matching score is obtained by weighted geometric mean or weighted arithmetic mean of multiple individual scores; the individual scores include: the similarity score between the peak area of ​​each secondary fragment ion and its intensity in the spectral library, characterized by the normalized root mean square error; the deviation score between the measured mass-to-charge ratio and the theoretical mass-to-charge ratio of each secondary fragment ion; the ratio score between the number of measured secondary fragment ions retained and the number of secondary fragment ions in the spectral library; the consistency score of the chromatographic peak tops among secondary fragment ions; the deviation score between the measured retention time and the predicted retention time; and the sum score of the peak areas of secondary fragment ions; the weight coefficient of each individual score ranges from 0.01 to 100.

6. The method for quantitative analysis of peptides integrating multiple liquid chromatography-mass spectra according to claim 2, characterized in that, In step (5), the method for aligning the retention time of the best matching peak to the retention time of each liquid chromatography-mass spectra is as follows: the measured retention time of the reference peptide in each liquid chromatography-mass spectra is taken as the dependent variable, the measured retention time of the reference peptide in the liquid chromatography-mass spectra where the best matching peak is located is taken as the independent variable, and the retention time of the best matching peak is taken as the sample point. The retention time after alignment is obtained by interpolation using linear interpolation or spline interpolation. And / or, in step (5), the allowable deviation of the retention time is reduced to ±0.1~±3 min; And / or, in step (5), the similarity score between the candidate chromatographic peak and the best matching peak is obtained by weighted geometric mean or weighted arithmetic mean of multiple individual scores; the individual scores include: the similarity score between the peak area of ​​each secondary fragment ion characterized by the normalized root mean square error and the intensity of the best matching peak, the deviation score between the measured mass-to-charge ratio and the theoretical mass-to-charge ratio of each secondary fragment ion, the score of the ratio of the number of retained measured secondary fragment ions to the number of ions in the best matching peak, and the score of the consistency of the chromatographic peak tops among secondary fragment ions; the weight coefficient of each individual score ranges from 0.01 to 100.

7. The method for quantitative analysis of peptides integrating multiple liquid chromatography-mass spectra according to claim 2, characterized in that, In step (6), the area of ​​each secondary fragment ion peak retained in the quantitative chromatographic peak is used as a vector, and the cosine of the angle between each pair of vectors is used to characterize the "correlation between the measured areas of different secondary fragment ions in multiple quantitative chromatographic peaks". The secondary fragment ions are clustered using the cosine distance of the angle, i.e., 1 minus the cosine of the angle. If the cosine distance of the angle exceeds the critical value, the corresponding secondary fragment ion is a secondary fragment ion with low correlation. The critical value range is 0.02~0.

1.

8. The method for quantitative analysis of peptides integrating multiple liquid chromatography-mass spectra according to claim 2, characterized in that, In step (7), the score of the chromatographic peak group is obtained by weighted geometric mean or weighted arithmetic mean of multiple individual scores. The individual scores include: the similarity score between the peak area of ​​each secondary fragment ion characterized by the normalized root mean square error and its intensity in the spectral library; the deviation score between the measured mass-to-charge ratio and the theoretical mass-to-charge ratio of each secondary fragment ion; the ratio score between the number of measured secondary fragment ions retained and the number of secondary fragment ions in the spectral library; the consistency score of the chromatographic peak tops among secondary fragment ions; the deviation score between the measured retention time and the predicted retention time; the relative standard deviation score between the peak areas of multiple quantitative chromatographic peaks corresponding to the same sample; and the ratio score between the number of quantitative chromatographic peaks in the chromatographic peak group and the number of liquid chromatography-mass spectra. The weight coefficient of each individual score ranges from 0.01 to 100. When multiple samples are analyzed only once, the individual score for the chromatographic peak group does not include the relative standard deviation score between the peak areas of multiple quantitative chromatographic peaks corresponding to the same sample.

9. The method for quantitative analysis of peptides integrating multiple liquid chromatography-mass spectra according to claim 2, characterized in that, In step (8), the method for specifying the peptide level score threshold is to repeat steps (2)-(7) on the pseudopeptide library, merge and sort the peptide level scores of the target peptide and the pseudopeptide, calculate the peptide level false detection rate corresponding to each score threshold, and specify the peptide level score threshold according to the requirements of the peptide level false detection rate; the required range of the false detection rate is 0.1%~10%.

10. The application of the polypeptide quantitative analysis method integrating multiple liquid chromatography-mass spectra according to any one of claims 1-9 in protein analysis.

11. An electronic device, characterized in that, It includes at least a memory and a processor, wherein the memory stores a program, and the processor, when executing the program in the memory, implements the peptide quantitative analysis method integrating multiple liquid chromatography-mass spectra as described in any one of claims 1-9.

Citation Information

Patent Citations

  • Mass spectrum data-based macro proteomics analysis method and electronic equipment

    CN115267033A

  • Method for analyzing non-data-dependent acquisition mode mass spectral data and application thereof

    CN106290684A

  • Full-ion monitoring and quantifying method based on second-level mass spectrum

    CN106324115A