A qualitative annotation method for high-resolution mass spectrometry data adapted to different instruments
By using standardized separation and UHPLC/HRMS analysis, combined with local least squares method and similarity evaluation, the database adaptation problem of metabolite qualitative annotation scheme under different conditions was solved, thereby improving the accuracy and data sharing of metabolomics research.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-29
- Publication Date
- 2026-04-03
AI Technical Summary
Existing qualitative annotation and analysis schemes for metabolites rely on metabolite databases. Due to the limited number of metabolites and the fact that qualitative information is related to experimental conditions, it is difficult to achieve the sharing and universal use of qualitative data under different conditions.
Metabolites or lipids were obtained using standardized separation and extraction methods. Samples were analyzed using UHPLC/HRMS. MGF files were generated using liquid chromatography-mass spectrometry data conversion tools. Mass spectrometry feature drift was corrected using local least squares method. Similarity was calculated and peak table data was constructed. Qualitative annotation was performed using similarity evaluation strategies. A qualitative identification database adapted to different instruments was constructed.
It enables accurate qualitative annotation of metabolic mass spectrometry features under different conditions, improves the accuracy of metabolomics research and the universality of the database, and enhances the data sharing of non-targeted metabolomics analysis.
Smart Images

Figure CN116879477B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of analytical chemistry and relates to a qualitative annotation method for high-resolution mass spectrometry data adapted to different instruments. Background Technology
[0002] Metabolite identification is crucial in untargeted metabolomics research, serving as a bridge between chemical and biological discoveries. [1] The achievement of various research objectives in metabolomics analysis relies on the accurate identification of metabolites. Therefore, unreliable qualitative analysis is highly likely to produce controversial or even inaccurate analytical results or mechanistic interpretations. However, due to the vast physical and chemical diversity of compounds, including numerous metabolites with varying physical and chemical properties, and the high specificity and uniqueness of compound identification tools and databases, the field is quite limited. It has been reported that only 1.8% of mass spectrometry features can be effectively annotated in non-targeted metabolomics analyses. [2] This is clearly a very low percentage. With the rapid development of mass spectrometry instrument hardware and the progress in metabolomics analysis experiments and theories, there is a great chance that this can be improved.
[0003] UPLC / HRMS coupled instruments have long been a mainstay of metabolomics and its wide applications. In metabolomics studies based on HRMS analysis, retention time (t... R ), Level 1 mass spectrometry (MS) 1 ) and secondary mass spectrometry (MS) 2 Information applicable to the main content of qualitative annotation analysis of compounds. [3,4] According to the Metabolomics Standards Initiative (MSI), four levels of accuracy for characterizing metabolite identification were proposed in 2007: Level 1 for definitively identified compounds, Level 2 for hypothetically annotated compounds, Level 3 for hypothetically annotated compound categories, and Level 4 for unknown compounds. In these UPLC / HRMS-based non-targeted metabolomics analyses, Level 1 clearly aligns perfectly with the ultimate goal pursued by researchers in characterizing unknown chemical signatures. [5,6] However, it requires validation using chemical standards under exactly the same analytical conditions, and the included consistency attributes include t R MS 1 and MS 2 , and the CCS value measured by ion mobility mass spectrometry (IM-MS).
[0004] Over the past two decades, researchers have proposed numerous strategies and methods. [1,7] To construct an experimentally and computationally based qualitative HRMS database for metabolites, and to qualitatively characterize unknown mass spectrometry features in metabolomics analysis, such as the development of a database-assisted MS...2 Deconvolution tools have improved the ability to identify metabolites in multiple sclerosis. Algorithms for generating computational metabolite databases, such as those for therapeutic peptides and proteins (TPPs) with branching, polycyclic structures, and other structural variations, have been proposed, yielding definitive metabolite structural characterization results. Furthermore, computational simulation methods have been used to generate lipid identification databases, including structure and boundary definitions, rule-based feature fragment prediction, and rigorous validation. Beyond these methodological advancements, numerous free databases have become mainstream tools in this field, significantly contributing to the progress of metabolomics practice, including the Human Metabolomics Database (HMDB), Metlin, GNPS, and others. However, based on these databases, only hypothetical annotations can be obtained for compound characterization, as described in levels 2 or 3 above. [8,9] .
[0005] Current solutions for qualitative annotation analysis of metabolites rely heavily on metabolite databases. However, due to the limited number of metabolites and the fact that qualitative information is related to experimental conditions, it is difficult to achieve qualitative data sharing and universal utilization under different conditions. Summary of the Invention
[0006] To address the current solutions for qualitative annotation analysis of metabolites, which heavily rely on metabolite databases but are limited by the number of metabolites and whose qualitative information is dependent on experimental conditions, making it difficult to achieve qualitative data sharing and universal utilization under different conditions, this invention provides the following technical solution:
[0007] A qualitative annotation method for high-resolution mass spectrometry data adapted to different instruments includes the following steps:
[0008] a. Use standardized separation and extraction methods or exploratory high-coverage analysis methods to extract metabolites or lipids from blood, tissue, urine, or cell metabolomics analysis samples;
[0009] b. The extracted research samples were analyzed and detected using UHPLC / HRMS to obtain the raw data required for metabolomics analysis;
[0010] c. Using a liquid chromatography-mass spectrometry (LC-MS) data conversion tool, the raw data was converted into an mzXML file, and then a secondary mass spectrometry conversion tool was used to obtain the data containing t. R MS 1 With MS 2 The information is in the mgf file;
[0011] d. For a single sample, identify and extract the mass spectrometry features of compounds from the raw data. For multiple samples with extracted features, accurately match the distribution of the same substance in different samples to construct peak table data for qualitative annotation analysis of metabolites, or data containing tR MS 1 With MS 2 The mgf file containing mass spectrometry information;
[0012] e. Using the mass spectrometry features generated by the internal standard compound added to the sample, or the mass spectrometry features whose substance is known, as a reference, the local least squares method is used to correct the retention time of the compound's mass spectrometry features between different samples and the drift of the first-order mass spectrometry, thereby obtaining high-quality peak table data with one row representing one metabolic mass spectrometry feature and one column representing one sample, and after data correction.
[0013] f. For the peak table data or mgf file mentioned above, find the t corresponding to each feature by matching the peak table data with the mgf file. R MS 1 With MS 2 Information used for qualitative annotation analysis of metabolites;
[0014] g. Using the af step, obtain the t corresponding to the qualitative metabolic mass spectrum characteristics. R MS 1 With MS 2 After obtaining the information, calculate the MS of the above-mentioned undetermined annotation analysis. 2 Mass spectrometry and corresponding MS of metabolites in known databases 2 Similarity between mass spectrometers;
[0015] h. Calculate based on t respectively R and MS 1 And overall similarity evaluation strategies;
[0016] Based on the above method of comparing the similarity between the mass spectrometry features to be determined and the known metabolites in the database, the qualitative analysis results of the mass spectrometry features to be determined in the database are arranged in descending order of their similarity values, and the one with the highest similarity is the qualitative annotation analysis result of the metabolite.
[0017] Further: The similarity between the MS2 mass spectra of the above-mentioned unidentified annotation analysis and the MS2 mass spectra of metabolites in the known database is calculated using the following formula, where the range is [0,1], A and B represent the secondary mass spectrometry data from the unknown unidentified database and the known database, respectively, i represents the i-th fragment of the mass spectrometry, which are the mass spectrometry fragments contained in the datasets A and B and their union region, respectively, with a value of 1 or 0, corresponding to the presence or absence of the feature, and correspondingly, is the weight based on ion intensity information;
[0018]
[0019] In the formula The range is [0,1]. A and B represent secondary mass spectrometry data from unknown and known databases, respectively. i represents the i-th fragment of the mass spectrometry. i,A p i,B p i,A∪B Let w represent the mass spectrometry fragments contained in datasets A and B, and their union region, respectively. The values are 1 or 0, corresponding to the presence or absence of the feature. i Weights are based on ion strength information.
[0020] Furthermore: the respective calculations based on t R and MS 1 And the overall similarity evaluation strategy, using the following formulas (2)-(4):
[0021]
[0022]
[0023]
[0024] in, Both are in the range [0,1], t R,1 , t R,2 , Δt R Corresponding to the aforementioned qualitative and database metabolite retention times, and the maximum permissible time threshold, based on MS 1 The calculated information is consistent with this.
[0025] Further: The method for obtaining the mgf file, and determining the noise level of the UHPLC / HRMS data, includes the following steps:
[0026] a1: Integrate the mgf files generated for each sample, then arrange the intensities of different mass spectrometry features in descending order, and divide the intensity values into m predefined intervals. Calculate the intensity distribution frequency of the metabolic mass spectrometry features within each intensity interval.
[0027] a2: Statistical analysis of frequency distribution. If a significant difference is found between the above intensity frequencies of two adjacent regions, then the intensity level of that region is considered as the noise level.
[0028] a3: Considering the influence of instrument status, instrument disturbance, matrix effect in complex measurement systems, and factors of multi-component overlapping systems on the determination of noise level, the noise level is calculated within a certain range in the analysis.
[0029] This invention provides a qualitative annotation method for high-resolution mass spectrometry data adapted to different instruments. Based on the construction of a qualitative identification and analysis database including tR, MS1, and MS2, it defines a method for qualitative identification under varying conditions, improving the accuracy of metabolite characterization and having significant implications for metabolomics research and applications. This method enables the construction of a metabolite qualitative annotation database under fixed conditions to still obtain good qualitative characterization and analysis results under changing conditions.
[0030] This method systematically studies the variation of qualitative characteristics in non-targeted metabolomics qualitative identification analysis across different instruments, such as AB SCIEX. The Agilent 6546Q-TOF, along with the Thermo Fisher Orbitrap Q / Exactive, present a quantitative scoring method for qualitative annotation of metabolites, effectively characterizing their variations under different conditions to efficiently extend the application of standard compound databases from one condition to another. This clearly has significant implications for non-targeted metabolomics analysis. Attached Figure Description
[0031] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0032] Figure 1 This is a schematic diagram illustrating the analytical principle of the present invention;
[0033] Figure 2 This is an overall flowchart of the data analysis and verification scheme of the present invention;
[0034] Figure 3 These are the total ion chromatograms of the mixed standard solution and the actual serum sample used to verify the results of this invention. (a) Total ion chromatogram of the standard solution obtained by detection and analysis on an Agilent instrument; (b) Total ion chromatogram of the standard solution obtained by detection and analysis on a SCIEX instrument; (c) Total ion chromatogram of the standard solution obtained by detection and analysis on a Thermo instrument; (d) Total ion chromatogram of the actual serum extract solution obtained by detection and analysis on an Agilent instrument; (e) Total ion chromatogram of the actual serum extract solution obtained by detection and analysis on a SCIEX instrument; (f) Total ion chromatogram of the actual serum extract solution obtained by detection and analysis on a Thermo instrument.
[0035] Figure 4This is a graph showing the PCA distribution results for the three instruments;
[0036] Figure 5 (a) tR-m / z distribution of data collected by Agilent instrument, (b) tR-m / z distribution of data collected by SCIEX instrument, (c) tR-m / z distribution of data collected by Thermo instrument, (d) PCA analysis results of HC and HCC sample data collected by Agilent instrument, (e) PCA analysis results of HC and HCC sample data collected by SCIEX instrument, (f) PCA analysis results of HC and HCC sample data collected by Thermo instrument;
[0037] Figure 6 The results are the similarity analysis results and the qualitative analysis results of the actual serum samples. (a) Comparison of similarity analysis results, (b) Venn diagram of the qualitative results obtained by the three instruments Agilent, SCIEX and Thermo, (c) Qualitative score distribution results of Agilent instrument analysis results, (d) Qualitative score distribution results of SCIEX instrument analysis results, and (e) Qualitative score distribution results of Thermo instrument analysis results. Detailed Implementation
[0038] It should be noted that, unless otherwise specified, the embodiments and features in the embodiments of the present invention can be combined with each other. The present invention will be described in detail below with reference to the accompanying drawings and embodiments.
[0039] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The following description of at least one exemplary embodiment is merely illustrative and is in no way intended to limit the present invention or its application or use. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0040] It should be noted that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of exemplary embodiments according to the invention. As used herein, the singular form is intended to include the plural form as well, unless the context clearly indicates otherwise. Furthermore, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0041] Unless otherwise specifically stated, the relative arrangement, numerical expressions, and values of the components and steps described in these embodiments do not limit the scope of the invention. It should also be understood that, for ease of description, the dimensions of the various parts shown in the drawings are not drawn to actual scale. Techniques, methods, and devices known to those skilled in the art may not be discussed in detail, but where appropriate, such techniques, methods, and devices should be considered part of the specification. In all examples shown and discussed herein, any specific values should be interpreted as merely exemplary and not as limitations. Therefore, other examples of exemplary embodiments may have different values. It should be noted that similar reference numerals and letters in the following figures denote similar items; therefore, once an item is defined in one figure, it need not be further discussed in subsequent figures.
[0042] In the description of this invention, it should be understood that the orientation or positional relationship indicated by directional terms such as "front, back, up, down, left, right", "horizontal, vertical, horizontal" and "top, bottom" is generally based on the orientation or positional relationship shown in the accompanying drawings, and is only for the convenience of describing this invention and simplifying the description. Unless otherwise stated, these directional terms do not indicate or imply that the device or element referred to must have a specific orientation or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation on the scope of protection of this invention. The directional terms "inner" and "outer" refer to the inner and outer contours relative to the outline of each component itself.
[0043] For ease of description, spatial relative terms such as "above," "over," "on the upper surface of," "above," etc., are used herein to describe the spatial positional relationship of a device or feature as shown in the figures to other devices or features. It should be understood that spatial relative terms are intended to encompass different orientations in use or operation besides the orientation of the device as described in the figures. For example, if the device in the figures is inverted, a device described as "above" or "above" other devices or structures would subsequently be positioned as "below" or "under" other devices or structures. Thus, the exemplary term "above" can include both "above" and "below." The device may also be positioned in other different ways (rotated 90 degrees or in other orientations), and the spatial relative descriptions used herein will be interpreted accordingly.
[0044] Furthermore, it should be noted that the use of terms such as "first" and "second" to define components is merely for the purpose of distinguishing the corresponding components. Unless otherwise stated, the above terms have no special meaning and therefore should not be construed as limiting the scope of protection of this invention.
[0045] Figure 1This is a schematic diagram illustrating the analytical principle of the present invention;
[0046] A qualitative annotation method for high-resolution mass spectrometry data adapted to different instruments includes the following steps:
[0047] a. Use standardized separation and extraction methods or exploratory high-coverage analysis methods to effectively extract metabolites or lipids from metabolomics analysis samples such as blood, tissue, urine, or cells;
[0048] b. The extracted research samples were analyzed and detected using UHPLC / HRMS to obtain the raw data required for metabolomics analysis;
[0049] c. Using a liquid chromatography-mass spectrometry (LC-MS) data conversion tool, the raw data was converted into an mzXML file, and then a secondary mass spectrometry conversion tool was used to obtain the data containing t. R MS 1 With MS 2 The information is in the mgf file;
[0050] d. Using methods such as XCMS and MZMine, for a single sample, mass spectrometry features of compounds are identified and extracted from the raw data. For multiple samples with extracted features, the distribution of the same substance in different samples is accurately matched to construct peak table data for qualitative annotation analysis of metabolites, or data containing t R MS 1 With MS 2 The mgf file containing mass spectrometry information;
[0051] e. Using the mass spectrometry features generated by the internal standard compound added to the sample, or the mass spectrometry features whose substance is known, as a reference, the local least squares method is used to correct the retention time of the compound's mass spectrometry features between different samples and the drift of the first-order mass spectrometry, thereby obtaining high-quality peak table data with one row representing one metabolic mass spectrometry feature and one column representing one sample, and after data correction.
[0052] f. For the peak table data or mgf file mentioned above, find the t corresponding to each feature by matching the peak table data with the mgf file. R MS 1 With MS 2 Information used for qualitative annotation analysis of metabolites;
[0053] g. Using the above steps and methods, obtain the t corresponding to the qualitative metabolic mass spectrum characteristics. R MS 1 With MS 2 After obtaining the information, the MS of the above-mentioned undetermined annotation analysis is calculated using the following formula (1). 2 Mass spectrometry and corresponding MS of metabolites in known databases 2The similarity between mass spectrometers, where The range is [0,1]. A and B represent secondary mass spectrometry data from unknown and known databases, respectively. i represents the i-th fragment of the mass spectrometry. i,A p iB p i,A∪B Let w represent the mass spectrometry fragments contained in datasets A and B, and their union region, respectively. The values are 1 or 0, corresponding to the presence or absence of the feature, respectively. i Weights are based on ion strength information;
[0054]
[0055] h. Using the following formulas (2)-(4), calculate the values based on t respectively. R and MS 1 And overall similarity evaluation strategies, among which, Both are in the range [0, 1], t R,1 , t R,2 , Δt R Corresponding to the aforementioned qualitative and database metabolite retention times, and the maximum permissible time threshold, based on MS 1 The calculated information is consistent with this;
[0056]
[0057]
[0058]
[0059] Based on the above method of comparing the similarity between the mass spectrometry features to be determined and the known metabolites in the database, the qualitative analysis results of the mass spectrometry features to be determined in the database are arranged in descending order of their similarity values, and the one with the highest similarity is the qualitative annotation analysis result of the metabolite.
[0060] Further: the method for acquiring mgf files and determining the noise level of UHPLC / HRMS data includes the following steps:
[0061] a1: Integrate the mgf files generated for each sample, then arrange the intensities of different mass spectrometry features in descending order, and divide the intensity values into predefined m intervals, such as 10, and calculate the intensity distribution frequency of the metabolic mass spectrometry features in different intensity intervals respectively.
[0062] a2: Statistical analysis of frequency distribution. If a significant difference is found between the above intensity frequencies of two adjacent regions, then the intensity level of that region is considered as the noise level.
[0063] a3: Considering the influence of factors such as the state of the analytical instrument, instrument disturbance, matrix effect in complex measurement systems, and multi-component overlapping systems on the determination of noise levels, the noise level is calculated within a range in actual analysis, such as 3-5 times the noise level mentioned above.
[0064] Example data: healthy control samples and HCC liver cancer samples.
[0065] Prepare standard mixtures. Weigh 1-3 mg of the standard and dissolve it in a suitable solvent to prepare a stock solution of 1 μg / mL or 2 μg / mL. Some of the compounds had already been prepared as stock solutions and were directly diluted for use in preparing mixed standards. A total of 37 mixed standards were ultimately selected. Based on the differences in the responses of different metabolites in the mass spectrometer, mixed standard working solutions containing different concentrations of metabolites were prepared by diluting with 1 / 4 acetonitrile / water.
[0066] Serum sample preparation. Serum from 10 patients with hepatocellular carcinoma was used in the disease group, and serum from 10 healthy individuals was used in the control group. Serum samples were thawed in ice water, thoroughly vortexed, and 50 μL of serum was taken and 200 μL of cold methanol was added. The mixture was vortexed for 1 min. The samples were centrifuged at 14,000 g at 4 °C for 10 min. 200 μL of the supernatant was lyophilized in a 1.5 ml EP tube. The lyophilized powder was redissolved in 50 μL of a 1 / 4 volume ratio mixed standard working solution (i.e., matrix spiked by adding to the reconstitution solvent), vortexed for 30 s, centrifuged, and the supernatant was collected for UPLC / HRMS analysis.
[0067] QC sample preparation: Take 5 μL of the supernatant from each sample after reconstitution and centrifugation, for a total of 100 μL for 20 samples, and insert a QC syringe for every 6-8 samples.
[0068] As mentioned above, the typical UPLC / HRMS data acquisition platform used in the three metabolomics analyses is ABSCIEX. 5600+, Agilent 6546Q-TOF, and Thermo Fisher Orbitrap Q / Exactive were used for metabolite data acquisition.
[0069] The chromatographic analysis conditions were identical for all three instrument types. The column was a BEH C8 (100 mm × 2.1 mm × 1.7 μm), manufactured by Waters. The column temperature was 50 °C, and the flow rate was 0.35 mL / min. Mobile phases A and B were 0.1% formic acid aqueous solution and acetonitrile, respectively. The gradient started at 5% B, held for 1 min, increased to 100% B over 23 min, held for 4 min, and finally decreased to 5% over 28.1 min, held for 1.9 min. A blank sample equilibration system was used before analyzing actual samples. In UPLC analysis, 6-8 actual samples were analyzed at intervals for each QC sample, used for data quality assessment and calibration analysis as needed.
[0070] The first analytical platform had an ionization temperature of 500°C, a spray voltage of 5,500 V, atomizing gas and atomizer pressures of 35 psi and 55 psi respectively, a decomposition potential (DP) voltage of 100 V, a collision energy of 10 V, an m / z measurement range of 50-1,200, and a resolution of 60,000. MS measurements were performed. 2 At that time, the m / z range was 50-1,200. The top 10 fragments with the highest intensity were automatically selected for m / z scanning, and four ISCID voltages (high, medium, low, and mixed levels) were used for systematic comparison, with voltages of 15, 30, 45, and 15 / 30 / 45 eV, respectively.
[0071] The second analytical platform has the following parameters: capillary temperature 325℃, capillary voltage 3,500V, drying gas 7L / min, nebulizing gas 35psi, sheath gas temperature 350℃, sheath gas flow rate 11L / min, fragmentation voltage 110V, cone voltage 65V, and RF Vpp 750V. The TOF full scan range is 50-1200 m / z with a resolution of 60,000. The secondary scan range is 30-1200 m / z. The top 5 ions with the highest response are automatically selected for secondary fragmentation scanning. Each sample is sampled twice; the first sample obtains the top 5 response fragments, and the second sample obtains the top 5-10 response fragments. The four ISCID voltages are the same as described above.
[0072] The third analysis platform has a capillary temperature of 300°C, a capillary voltage of 3500V, an auxiliary gas temperature of 350°C, an auxiliary gas flow rate of 10 in arbitrary units, a sheath gas flow rate of 45 in arbitrary units, and an S-lens rflevel of 50.0. The Orbitrap full scan range is 73.4–1,100 m / z with a resolution of 120,000. The scan range is automatically determined based on the fragmentation. The top 10 ions with the highest response are automatically selected for secondary fragmentation scanning, with only one sample collected per needle. The four ISCID voltages are the same as described above.
[0073] The raw sample data obtained from the above measurements, after peak matching, showed that metabolite t R ~m / z distribution, and PCA analysis results between instruments, such as Figure 2-3 As shown.
[0074] Figure 2 This is an overall flowchart of the data analysis and verification scheme of the present invention;
[0075] Figure 3 These are the total ion chromatograms of the mixed standard solution and the actual serum sample used to verify the results of this invention. (a) Total ion chromatogram of the standard solution obtained by detection and analysis on an Agilent instrument; (b) Total ion chromatogram of the standard solution obtained by detection and analysis on a SCIEX instrument; (c) Total ion chromatogram of the standard solution obtained by detection and analysis on a Thermo instrument; (d) Total ion chromatogram of the actual serum extract solution obtained by detection and analysis on an Agilent instrument; (e) Total ion chromatogram of the actual serum extract solution obtained by detection and analysis on a SCIEX instrument; (f) Total ion chromatogram of the actual serum extract solution obtained by detection and analysis on a Thermo instrument.
[0076] The results indicate differences in instrument types, sample complexity, and metabolite distribution characteristics.
[0077] First, the mixed standard solutions were analyzed to assess the accuracy of metabolite identification. After data fusion from the standard samples, 2, 4, and 1 compounds, respectively, were not detected in the three instrument types. For instrument measurements, these standards were accurately identified by comparative analysis with previously added standard compounds in the database. The results demonstrate the effectiveness of this method in qualitative annotation for metabolite identification analysis. The overall similarity results also indicate a high correlation of qualitative characteristics between different instrument pairs.
[0078] Then, qualitative information on metabolites in actual serum samples is obtained, including t R MS 1 and MS 2 Used for the identification of metabolites and comparison of results, such as Figure 4 As shown. Using our established metabolite database, 30, 28, and 31 compounds were reliably annotated across the three instruments, respectively. For mass spectrometry features from actual serum, the number of compounds with similarity greater than 0.5 reached 535, 107, and 796, respectively, demonstrating good qualitative annotation results, such as... Figure 5 As shown.
[0079] Figure 5(a) tR-m / z distribution of data collected by Agilent instrument, (b) tR-m / z distribution of data collected by SCIEX instrument, (c) tR-m / z distribution of data collected by Thermo instrument, (d) PCA analysis results of HC and HCC sample data collected by Agilent instrument, (e) PCA analysis results of HC and HCC sample data collected by SCIEX instrument, (f) PCA analysis results of HC and HCC sample data collected by Thermo instrument;
[0080] A comparison and scoring analysis of common compounds simultaneously measured by the three instruments further demonstrated the method's good effectiveness in metabolite identification. The qualitative strategy proposed in this invention can improve data sharing and database versatility for metabolite identification in non-targeted metabolomics.
[0081] Figure 6 The results are the similarity analysis results and the qualitative analysis results of the actual serum samples. (a) Comparison of similarity analysis results, (b) Venn diagram of the qualitative results obtained by the three instruments Agilent, SCIEX and Thermo, (c) Qualitative score distribution results of Agilent instrument analysis results, (d) Qualitative score distribution results of SCIEX instrument analysis results, and (e) Qualitative score distribution results of Thermo instrument analysis results.
[0082] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
[0083] References:
[0084] 1. N. Dai Hai, N. Canh Hao, H. Mamitsuka, Recent advances and prospects of computational methods for metabolite identification: a review with emphasis on machine learning approaches, Briefings in Bioinformatics 20(6)(2019)2028-2043.
[0085] 2.RRda Silva,PCDorrestein,RAQuinn,Illuminating the metabolomics of dark matter,Proceedings of the National Academy of Sciences of the United States of America 112(41)(2015)12549-12550.
[0086] 3.ARFernie,A.Aharoni,L.Willmitzer,M.Stitt,T.Tohge,J.Kopka,...V.DeLuca,Recommendations forReportingMetaboliteData,Plant Cell 23(7)(2011)2477-2482.
[0087] 4.RMSalek,C.Steinbeck,MRViant,R.Goodacre,WBDunn,The role ofreporting standards for metabolite annotation and identification inmetabolomic studies,Gigascience 2(2013).
[0088] 5.MMKoek,B.Muilwijk,MJvan der Werf,T.Hankemeier,Microbial metabolomics with gas chromatography / mass spectrometry,Analytical Chemistry78(4)(2006)1272-1281.
[0089] 6.Y.Tikunov,A.Lommen,CHRde Vos,HAVerhoeven,RJBino,RDHall,AGBovy,A novel approach for nontargeted data analysis format metabolomics.Large-scale profiling of tomato fruit volatiles,Plant Physiology139(3)(2005)1125-1137.
[0090] 7.Y.Wang,S.Liu,Y.Hu,P.Li,J.-B.Wan,Current state ofthe art ofmassspectrometry-based metabolomics studies-a review focusing on wide coverage,high throughput and easy identification,RscAdvances 5(96)(2015)78728-78737.
[0091] 8.D.S.Wishart,Y.D.Feunang,A.Marcu,A.C.Guo,K.Liang,R.Vazquez-Fresno,...A.Scalbert,HMDB 4.0:the human metabolome database for 2018,NucleicAcids Research 46(D1)(2018)D608-D617.
[0092] 9.D.S.Wishart,T.Jewison,A.C.Guo,M.Wilson,C.Knox,Y.Liu,...A.Scalbert,HMDB 3.0-The Human Metabolome Database in 2013,Nucleic Acids Research 41(D1)(2013)D801-D807.
Claims
1. A qualitative annotation method for high-resolution mass spectrometry data adapted to different instruments, characterized in that: Includes the following steps: a. Use standardized separation and extraction methods or exploratory high-coverage analysis methods to extract metabolites or lipids from blood, tissue, urine, or cell metabolomics analysis samples; b. The extracted samples were analyzed using UHPLC / HRMS to obtain the raw data required for metabolomics analysis; c. Using a liquid chromatography-mass spectrometry (LC-MS) data conversion tool, the raw data was converted into an mzXML file, and then a secondary mass spectrometry conversion tool was used to obtain the data containing t. R MS 1 With MS 2 The information is in the mgf file; d. For a single sample, identify and extract the mass spectrometry features of compounds from the raw data. For multiple samples with extracted features, accurately match the distribution of the same substance in different samples to construct peak table data for qualitative annotation analysis of metabolites, or data containing t R MS 1 With MS 2 The mgf file containing mass spectrometry information; e. Using the mass spectrometry features generated by the internal standard compound added to the sample, or the mass spectrometry features whose substance is known, as a reference, the local least squares method is used to correct the retention time of the compound's mass spectrometry features between different samples and the drift of the first-order mass spectrometry, thereby obtaining high-quality peak table data with one row representing one metabolic mass spectrometry feature and one column representing one sample, and after data correction. f. For the peak table data or mgf file mentioned above, find the t corresponding to each feature by matching the peak table data with the mgf file. R MS 1 With MS 2 Information used for qualitative annotation analysis of metabolites; g. Using the af step, obtain the t corresponding to the qualitative metabolic mass spectrum characteristics. R MS 1 With MS 2 After obtaining the information, calculate the MS of the above-mentioned undetermined annotation analysis. 2 Mass spectrometry and corresponding MS of metabolites in known databases 2 Similarity between mass spectrometers; h. Calculate based on t respectively R and MS 1 And overall similarity evaluation strategies; Based on the above method of comparing the similarity between the mass spectrometry features to be determined and the known metabolites in the database, the qualitative analysis results of the mass spectrometry features to be determined in the database are arranged in descending order of their similarity values, and the one with the highest similarity is the qualitative annotation analysis result of the metabolite.
2. The qualitative annotation method for high-resolution mass spectrometry data adapted to different instruments as described in claim 1, characterized in that: The calculation of the MS of the above-mentioned undetermined annotation analysis 2 The similarity between the mass spectrometer and the corresponding MS2 mass spectra of metabolites in known databases is calculated using the following formula. (1) In the formula The range is [0,1]. A and B represent secondary mass spectrometry data from unknown and known databases, respectively, and i represents the i-th fragment of the mass spectrometer. These represent the mass spectrometry fragments contained in datasets A and B, and their union region, respectively. Their values are 1 or 0, corresponding to the presence or absence of the feature, respectively. Weights are based on ion strength information.
3. The qualitative annotation method for high-resolution mass spectrometry data adapted to different instruments as described in claim 1, characterized in that: The calculations are based on t. R and MS 1 And the overall similarity evaluation strategy, using the following formulas (2)-(4): (2) (3) (4) Among them, S tR Based on retention time t R The defined similarity calculation method, S MS 1 The similarity calculation method defined for the MS1 mass spectrometer, S MS 2 A similarity calculation method defined for MS2 secondary mass spectrometry; S tR S MS 1 S MS 2 All are in the range [0,1], where t R1 For the retention time of the metabolite being identified, t R2 For the retention time of metabolites in the database, t R The maximum allowable time threshold, based on MS 1 1 represents the qualitative first-order mass spectrometry, MS. 1 2 represents the first-order mass spectrometer in the database. MS 1 This is the maximum permissible threshold for qualitative analysis and the first-order mass spectrometry in the database.
4. The qualitative annotation method for high-resolution mass spectrometry data adapted to different instruments as described in claim 1, characterized in that: The method for obtaining mgf files, and determining the noise level of UHPLC / HRMS data, includes the following steps: a1: Integrate the mgf files generated for each sample, then arrange the intensities of different mass spectrometry features in descending order, and divide the intensity values into m predefined intervals. Calculate the intensity distribution frequency of the metabolic mass spectrometry features within each intensity interval. a2: Statistical analysis of frequency distribution. If a significant difference in intensity frequency is found between two adjacent regions, then the intensity level of that region is considered as the noise level. a3: Considering the influence of instrument status, instrument disturbance, matrix effect in complex measurement systems, and factors of multi-component overlapping systems on the determination of noise level, the noise level is calculated within a certain range in the analysis.
Citation Information
Patent Citations
Metabonomics relative quantitative analysis method based on UPLC / HRMS
CN111579665A
Metabolome deep annotation method
CN114594171A