Method for quantitative analysis of 4d metabolomics and lipidomics data
By collapsing 4D LC-IMS-MS data into pseudo-3D format and using feature detection, the method addresses the challenge of data complexity, facilitating efficient metabolite identification and illness diagnosis.
Patent Information
- Application Number
- PCT/US2025/031569
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-05-31
- Filing Date
- 2025-05-30
- Publication Date
- 2025-12-04
AI Technical Summary
The rapid increase in data complexity from LC-IMS-MS systems due to enhanced ion mobility separation has outpaced the development of software capable of efficiently processing and analyzing this data, limiting the evaluation of untargeted metabolomics and lipidomics research.
A method is disclosed to process 4D data from LC-IMS-MS systems by collapsing it into pseudo-3D data, detecting features, and using feature-related data to identify metabolites, employing mathematical operations such as summing, differencing, or combining spectral intensities to maintain data integrity while reducing computational burden.
This approach allows for efficient analysis of complex LC-IMS-MS data, enabling accurate metabolite identification and diagnosis of illnesses, while reducing computational requirements and maintaining data integrity.
Smart Images

Figure US2025031569_04122025_PF_FP_ABST
Abstract
Description
CT METHOD FOR QUANTITATIVE ANALYSIS OF 4D METABOLOMICS AND LIPIDOMICS DATA CROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims priority to U.S. Provisional Patent Application No. 63 / 654,628, filed May 31, 2024, the entire contents of which are incorporated herein by reference. BACKGROUND OF THE INVENTION
[0002] Metabolomics is a rapidly evolving field that deals with the high- throughput characterization of metabolites and is the study of the metabolite composition of cell types, tissues, organs, or organisms. Metabolomics involves an unbiased global survey of all the status and level of low-molecular-weight molecules in a biofluid, cell, tissue, organ, or organism. Such molecules are the final downstream products of genomic, transcriptomic, and / or proteomic activity, and comprise, for example, low-molecular weight metabolites, such as amino acids, sugars, fatty acids, lipids, and steroids. Such an analysis directly reflects the activity of the metabolic network on a global scale, that leads to the production of the detected metabolites and yields essential information about the underlying biological status of the system in question. Such information may be used to determine end points of physiology and pathophysiology.
[0003] Metabolomic analysis begins with obtaining a chemical or biological sample and separating the multitude of metabolites contained therein. Both liquid chromatography (LC) and gas chromatography (GC) are used for metabolite separation. In recent years, ion mobility spectrometry (IMS) has also been used, often being coupled with LC to improve separation of the metabolites. These separation methods are coupled to a detection device that detects molecular features of the metabolites, thereby enabling their identification. Both nuclear magnetic resonance (NMR) and mass spectrometry (MS) have been used to detect such molecular features, although MS has been more commonly used since it is more sensitive, has higher throughput, and can detect more molecules in a complex chemical or biological sample.
[0004] Data collection is merely the first step in metabolomic analysis. Data collection must be followed by analysis of the data obtained for each individual metabolite, and finally identification of the individual metabolites. Thus, metabolomics ultimatelyCT comprises an integration of instrumentation, chemistry, statistics, and computer science to solve a biological problem. Given the large number of metabolites in a biological, and the numerous data points associated with each metabolite, a metabolic experiment results in the production of enormous amounts of data. It is not humanly possible to process the entirety of the data manually and computational tools are necessary for further analysis. Thus, software has been developed that aids in component analysis, statistical testing, data visualization and identifying specific signatures in the data. However, the recent addition of IMS to LC has increased the amount of resulting data, and currently there is a paucity of software for processing LC-IMS-MS data. This lack of available software has limited evaluation of the quantitative benefit of ion-mobility platforms in untargeted metabolomics and lipidomics research. To fully benefit from the increased separation offered by IMS, what is needed is a method of handling the increased level of chromatographic and spectral data in an efficient and cost-effective manner. The present disclosure provides such methodology and also provides other benefits as well. BRIEF DESCRIPTION OF THE DISCLOSURE
[0005] One embodiment is a method of processing four-dimensional (4D) data from a liquid chromatography (LC)-ion mobility spectrometry (IMS)-mass spectrometry (MS) (LC-IMS-MS) system, comprising collapsing the 4D data to produce pseudo-3D data, detecting one or more features in the pseudo-3D data, using feature-related pseudo-3D data to identify 4D data related to the feature, and using the identified feature- related 4D data in combination with the feature-related pseudo-3D data to identify an associated metabolite, wherein the 4D data is produced by analysis of a sample using the LC-IMS-MS system. In certain aspects, the 4D data may comprise LC retention times (RT), IMS drift times, m / z ratios and related intensities. In certain aspects, the 4D data may be collapsed in the drift time dimension. Collapsing the 4D data may comprise mathematically combining MS spectra having m / z ratios and associated intensities associated with a single RT, thereby producing a combined spectrum, wherein if an m / z ratio is present in a single spectrum the intensity of the m / z ratio in the combined spectrum is the same as the intensity of the m / z ratio in the single spectrum, and wherein if two or more spectra share a common m / z ratio the intensity of the common m / z ratio in the combined spectrum is the result of mathematically combining the individual intensities from the two or more spectra. In certain aspects, mathematically combining may comprise an operation selected from the groupCT consisting of summing the intensities of the MS spectra, determining a difference between intensities of the MS spectra, determining a mean of the MS spectra, determining a median of the MS spectra, determining a product of the MS spectra, determining a quotient of the MS spectra, or determining a ratio of the MS spectra. In certain aspects, feature detection may comprise determining the boundaries, centers, and intensities of two-dimensional (2D) signals, which may comprise mass / charge ratios and RT. Feature detection may be performed using feature detection software. In certain aspects, using feature-related 3D data to identify 4D data may comprise using a feature-related RT and / or a feature-related m / z ratio to identify one or more drift times associated with the RT and / or the m / z ratio. In certain aspects, the feature-related drift time, the feature-related m / z ratio and the associated one or more drift times may be used to identify the metabolite.
[0006] One embodiment is a method of processing 4D data from a LC-IMS- MS system, comprising: collapsing the 4D data associated with at least one RT frame to produce pseudo-3D data, detecting one or more features in the pseudo-3D data, using feature- related pseudo-3D data to identify one or more bins comprising 4D data associated with the feature, and using the identified feature-related 4D data in combination with the feature- related pseudo-3D data to identify an associated metabolite, wherein the 4D data is produced by analysis of a sample using the LC-IMS-MS system. In certain aspects, the 4D data may comprise LC retention times (RT), IMS drift times, m / z ratios and related intensities. In certain aspects, the 4D data may be collapsed in the drift time dimension. Collapsing the data may comprise combining the MS spectra from two or more bins associated with the same RT, thereby producing a combined spectrum, wherein if an m / z ratio is only present in a single bin the intensity of the m / z ratio in the combined spectrum is the same as the intensity of the m / z ratio in the single bin, and wherein if two or more bins share a common m / z ratio, the intensity of the common m / z ratio in the combined spectrum is the result of mathematically combining the individual intensities from the two or more bins. In certain aspects, feature detection may comprise determining the boundaries, centers, and intensities of two-dimensional (2D) signals, which may comprise mass / charge ratios and RT. Feature detection may be performed using feature detection software. The step of using feature- related 3D data to identify 4D data may comprise a feature-related RT and / or a feature- related m / z ratio to identify one or more drift times associated with the RT and / or the m / z ratio. In certain aspects, the feature-related drift time, the feature-related m / z ratio and the associated one or more drift times may be used to identify the metabolite.CT
[0007] One embodiment is a method of identifying a metabolite in a sample, comprising obtaining 4D data produced by analyzing the sample using a LC-IMS-MS system, and processing the 4D data using a method of the disclosure, thereby identifying the metabolite.
[0008] In an aspect of the disclosure, the sample may be a biological sample or a chemical sample, and the biological sample may comprise blood, serum, plasma, urine, cerebrospinal fluid, tears, saliva, sputum, lymph fluid, dialysate, lavage fluid, or a fluid derived from an organ or tissue.
[0009] One embodiment is a method of diagnosing an illness in an individual, comprising obtaining 4D data produced by analyzing a sample from the individual using a LC-IMS-MS system, processing the 4D data using the method of the disclosure to determine the presence, absence or level of one or more metabolites in the biological sample, and using the presence, absence and / or level of the one or more metabolites to diagnose the illness.
[0010] One embodiment is a system for performing a method of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] FIGS.1A-1C show a schematic overview for 4D metabolomics and XCiMS pre-processing workflow. FIG. 1A illustrates a generalized scheme of LC-IM-MS analysis. FIG. 1B illustrates a scheme for four-dimensional LC-IM-MS data acquisition. FIG.1C illustrates a scheme of 4D data pre-processing by XCiMS software. Vender-neutral LC-IM-MS spectra were constructed from .mzML files or TIMS .d files. LC-IM-MS spectra were collapsed onto LC-MS dimensions through spectra binning. LC peak profiles from pseudo-LCMS spectra are determined by the 3D centWave algorithm. RT-IM regions-of- interest (RT-IM ROIs) defined by LC peak profiles were extracted from LC-IM-MS spectra. ROIs were collapsed to IM mobilograms / driftograms followed by IM peak detection using 2D centWave algorithm. As a result, 4D features with LC and IM peak profiles are obtained. XCiMS also includes post-processing functions of CCS calculation, peak correspondence, and fold-change analysis.
[0012] FIGS.2A & 2B illustrate the data-processing Workflow of XCiMS Software for LC-IM-MS and DI-IM-MS Data.CT
[0013] FIG.3 illustrates the data structure of the XCiMS workflow. XCiMS processes LC-IM-MS data consecutively at raw data layer, spectra layer and feature layer. At raw data layer, mzML files (DTIMS / TWIMS data) are read through MSnbase R interface, while Bruker TIMS data are read through opentimsr R interface. At spectra layer, LCIMMSpectraL object containing LC-IM-MS spectra of multiple data files are constructed from the raw data. pLCMSpectraL or pIMMSpectraL object including pseudo-LCMS or pseudo-IMMS spectra are constructed by collapsing MS spectra at each RT frame or IM bin to single MS spectrum. Optionally, 3D intensity maps (RT-IM, IM-MZ, and RT-MZ) and 2D extracted ion traces (EIC,EIM / EID) are extracted from spectra data. At feature layer, pLCMSpeaks or pIMMSpeaks objects are constructed by detecting chromatographic and ion-mobility peaks from pLCMSpectraL or pIMMSpectraL object. LCIMMSpeaks object are constructed by extracting RtImMapL object (RT-IM ROIs) of pseudo-LCMS peaks and detecting IM peaks on the extracted mobilograms / driftograms. pLCMSpeakgroup, pIMMSpeakgroup, and LCIMMSpeakgroup objects are optionally constructed by grouping pLCMSpeaks pIMMSpeaks or LCIMMSpeaks when sample group information is available.
[0014] FIGS. 4A-4E illustrates Multidimensional peak detection assisted by spectra binning in XCiMS. FIG.4A illustrates how 4D MS spectra are archived by aligned retention time frames and ion mobility bins. MS spectra can be collapsed to pseudo-LCMS spectra or pseudo-IMMS spectra via spectra binning. FIG.4B illustrates a spectra binning of three schematic MS spectra at one retention time frame i and consecutive ion mobility bins j, j+1, j+2. MS peaks are sorted by m / z values followed by merging peaks with m / z values under a ppm threshold (red rect.) while retaining peaks with unique m / z values. c) a RT-IM ROI of m / z 268.1049 (Adenosine) extracted from LC-IM-MS spectra acquired by a DTIMS- MS system. FIG. 4D shows convolved chromatograms (red solid line) versus discrete chromatograms extracted from c along each IMS bin (color-coded dashed lines). FIG. 4E shows convolved driftogram (red solid line) versus discrete driftograms of c along each RT frame (color-coded dashed lines).
[0015] FIG.5 illustrates alignment of RT frames and IM bins within a single TWIMS data file. The RT frame and IM bin coordinates of LC-IM-MS spectra are plotted at RT range of 200-240 seconds, and IM range of 4-6 ms.
[0016] FIGS. 6A-6D show RT frames and IM bins across three replicates of TWIMS data files. FIG.6A shows misaligned RT frames in seconds among three technical replicates. FIG.6B shows aligned IM bins in ms across three technical replicates. FIG. 6CCT shows a Venn diagram of RT frame values across three replicates. FIG. 6D shows a Venn diagram of IM bin values across three replicates.
[0017] FIGS. 7A-7O show the results of 4D metabolomics analysis of amino acid standards using XCiMS Software.10 ^^M mixture of amino acid standards were analyzed using HILIC chromatography coupled to a TWIMS-MS mass spectrometer. Feature detection and correspondence were performed using XCiMS software. FIGS. 7A- 7N shows LC and IM peak profiles of an amino acid standard with red solid lines of peak boundaries determined by XCiMS. FIG.7O shows the correlation between LC and IM peak areas of the amino acid standards quantified by XCiMS. Metabolite annotation was performed by matching accurate mass and retention time to an in-house library with a mass ppm of 20 and retention time window of 2 seconds.
[0018] FIGS. 8A-8I show the results of 4D metabolomics analysis of SRM1950 plasma extracts spiked with two equivalents of in-house chemical standard mixture. One (red) and five (cyan) equivalents of 63 metabolite standard mixture were spiked in SRM1950 metabolic extracts and measured by HILIC-TWIMS-MS analysis. 4D peak detection, peak correspondence and fold-change analysis were performed by XCiMS. FIG. 8A shows the number of standards spiked in originally, detected by XCiMS, and with statistically significant changes quantified by XCiMS. FIG.8B shows the number of average 4D features per sample, 4D feature groups, and statistically significant altered features detected by XCiMS. FIG.8C show the correlation between fold changes of LC and IM peaks of XCiMS-detected standards. LC and IM peak profiles of selected standards with statistically significant changes between two spike-in conditions: FIG. 8D- R5P: ribose-5- phosphate; FIG. 8E- GAP / DHAP: glyceraldehyde 3-phosphate / dihydroxyacetone phosphate; FIG. 8F- F6P / F1P: Fructose-6-phosphate / glucose-1-phosphate; FIG. 8G- G6P: glucose-6-phosphate; FIG. 8H- cGMP: guanosine 3',5'-cyclic monophosphate; FIG. 8I- Glutamate. Metabolite annotation was performed by matching accurate mass and retention time to an in-house library with a mass ppm of 20 and retention time window of 2 seconds.
[0019] FIGS. 9A-9H show Spike-in standards detected by XCiMS with statistically non-significant changes.
[0020] FIGS. 10A-10I show histograms of IM elution time (FIGS. 10A- 10C), IM peak widths (FIGS. 10D-10F), and IM peak resolution (FIGS. 10G-10I) of the most abundant 10,000 4D features detected by XCiMS in E. coli samples in TWIMS, DTIMS, and TIMS system respectively.CT
[0021] FIGS. 11A-11G illustrate improved metabolite annotation using accurate mass and CCS values. FIG.11A shows scattered plot of CCS versus m / z values of MSMLS IROA standards detected by XCiMS (N=145). FIG. 11B shows the correlation between XCiMS-calculated CCS values and unified CCS Compendium values of IROA MSMLS standards. (N=130) FIG. 11C shows histograms of percent CCS measurement errors of MSMLS IROA standards compared to unified CCS Compendium (N=130). FIG. 11D shows histograms of HMDB database hits when querying m / z values (blue) versus m / z and CCS values (red) of MSMLS IROA standards. (N=137) FIG.11E shows scattered plot of CCS versus m / z values of XCiMS-detected 4D features in SRM1950 plasma polar extract (N=3234). FIG. 11F shows the number of top abundant 4D features (black) with HMDB database hits when querying m / z values (blue) versus querying mz and CCS values (red) (N=1000). FIG. 11G shows histograms of HMDB database hits as querying m / z (blue) and m / z and CCS values (red) of 4D features of SRM1950 sample (N=97). Databases searches of MSMLS standards were performed with a 20-ppm tolerance in m / z and 3-percent tolerance in CCS. Databases searches of SRM19504D features were performed with a 30- ppm tolerance in m / z and 5-percent tolerance in CCS.
[0022] FIG. 12 shows CiMS-detected sub-features attributing to charge isomers of in-source reduced NADPH from NADP+ .
[0023] FIGS. 13A & 13B show experimental PASEF-MS / MS spectra of F6P (1 / K0 = 0.68) (FIG.13A) and G1P (1 / K0 =0.70) (FIG.13B).
[0024] FIGS.14A-14C show LC and IM peak profiles of structural isomers in MSMLS IROA library detected by XCiMS measured by TWIMS. FIG. 14 A- cAMP / A- 2’, 3’-cP, FIG.14B- AMP / dGMP, FIG.14C- ATP / dGTP.
[0025] FIGS. 15A-15C show: FIG. 15A- EICs of Cytidine, Tryptophan, [Cytidine+Tryptophan-H]- cluster ion; FIG. 15B- Averaged MS spectrum between RT 257 and 287 second; and, FIG.15C - EIMs of Tryptophan, Cytidine, and [Cytidine+Tryptophan- H]- ] cluster ion.
[0026] FIGS. 16A-16D show: FIG. 16A- Total number of precursor pseudo-3D features, 4D features, and 4D sub-features derived from 39 compounds in PolarMix sample; FIG.16B – the distribution of number of 4D sub-features and unique 4D features detected per 3D precursor of 39 reference compounds; FIG.16C – the total number of precursor pseudo-3D features, 4D features, and 4D sub-features detected by XCiMS inCT the PolarMix sample; and, FIG. 16D – the distribution of number of 4D sub-features and unique 4D features detected per 3D precursor.
[0027] FIG. 17 shows an example 4D analysis system in accordance with an embodiment of the present disclosure.
[0028] FIG.18 shows an example computing device in accordance with an embodiment of the present disclosure.
[0029] FIGS. 19A-19D show XCiMS-detected 4D features and 4D sub- features in 6 TWIMS, 2 DTIMS, and 3 TIMS raw data. FIG. 19A shows 4D features (dark blue) and sub-4D features (red) detected by XCiMS software under different XCiMS parameters. FIG.19B shows credentialed LC and IM peak profiles of NADPH [M-H]- charge isomers detected by XCiMS in E. coli measured by TWIMs. FIG. 19C shows credentialed LC and IM peak profiles of Adenosine [M+H]+charge isomers detected by XCiMS in E. coli measured by DTIMS. Red CCS values indicates a database match, while the black CCS values indicates no database match. FIG. 19D shows LC and IM peak profiles of structural isomers G1P and F6P unresolved by HLIC and resolved by TIMS in PolarMix (polar metabolite standard mixtures). DETAILED DESCRIPTION OF THE DISCLOSURE
[0030] Liquid chromatography (LC)- mass spectrometry (MS) systems having ion mobility spectroscopy (IMS) capabilities are increasingly being produced. The increased ability of the IMS component to separate previously non-resolved metabolites significantly increases the amount and complexity of data produced by such systems. However, the development of software capable of analyzing such data has not kept pace. Thus, the present disclosure relates to an improved method of analyzing the massive amounts of data obtained from an LC-IMS-MS system. In particular, the present disclosure discloses a method of taking complex, 4-dimensional (4D) data from an LC-IMS-MS system and reducing the data into a 3-dimensional (3D) representation without losing any of the original data. The 3D data can then be analyzed using conventional computer systems and related software. The results of this 3D data analysis are then used to delineate metabolites within the original 4D data. Thus, a method of the disclosure may generally be performed by obtaining 4D data resulting from introduction of a chemical or biological sample into an LC- IMS-MS system, collapsing the 4D data in one dimension to produce pseudo-3D data, analyzing the pseudo-3D data to identify features in the pseudo-3D data, using identifiedCT feature-related data to identify additional feature-related data in the original 4D data, and using the combined identified feature-related data to identify a metabolite in the chemical or biological sample.
[0031] Before the present invention is further described, it is to be understood that this invention is not limited to particular embodiments described, as such may, of course, vary. It is also to be understood that the terminology used herein is for the purpose of describing particular embodiments only, and is not intended to be limiting, since the scope of the present invention will be limited only by the claims.
[0032] It must be noted that as used herein and in the appended claims, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. For example, a metabolite refers to one or more metabolites. As such, the terms "a", "an", "one or more" and "at least one" can be used interchangeably.
[0033] Similarly, the terms "comprising", "including" and "having" can be used interchangeably. As used herein, the term “comprising” may be replaced with “consisting” or with “consisting essentially of” in particular aspect, as desired.
[0034] It is further noted that the claims may be drafted to exclude any optional element. As such, this statement is intended to serve as antecedent basis for use of such exclusive terminology as "solely," "only" and the like in connection with the recitation of claim elements or use of a "negative" limitation.
[0035] Various terms relating to aspects of the present disclosure are used throughout the specification and claims. Such terms are to be given their ordinary meaning in the art, unless otherwise indicated. Other specifically defined terms are to be construed in a manner consistent with the definitions provided herein.
[0036] Unless otherwise expressly stated, it is in no way intended that any method or aspect set forth herein be construed as requiring that its steps be performed in a specific order. Accordingly, where a method claim does not specifically state in the claims or descriptions that the steps are to be limited to a specific order, it is in no way intended that an order be inferred, in any respect. This holds for any possible non-expressed basis for interpretation, including matters of logic with respect to arrangement of steps or operational flow, plain meaning derived from grammatical organization or punctuation, or the number or type of aspects described in the specification.
[0037] It is appreciated that certain features of the invention, which are, for clarity, described in the context of separate embodiments, may also be provided inCT combination in a single embodiment. Conversely, various features of the invention, which are, for brevity, described in the context of a single embodiment, may also be provided separately or in any suitable sub-combination. All combinations of the embodiments are specifically embraced by the present invention and are disclosed herein just as if each and every combination was individually and explicitly disclosed. In addition, all sub- combinations are also specifically embraced by the present invention and are disclosed herein just as if each and every such sub-combination was individually and explicitly disclosed herein.
[0038] Publications discussed herein are provided solely for their disclosure prior to the filing date of the present application. Nothing herein is to be construed as an admission that the present invention is not entitled to antedate such publication by virtue of prior invention. Further, the dates of publication provided may be different from the actual publication dates, which may need to be independently confirmed.
[0039] As used herein, the term “about” means that the recited numerical value is approximate and small variations would not significantly affect the practice of the disclosed embodiments. Where a numerical value is used, unless indicated otherwise by the context, the term “about” means the numerical value can vary by ±10% and remain within the scope of the disclosed embodiments.
[0040] As used herein, dimensional data refers to specific measurements obtained from the LC-IMS-MS system. For example, the time it takes for a metabolite to traverse a LC column is known as the retention time (RT). RT is one dimension of chromatographic (e.g., LC-MS or LC-IMS-MS) analysis. Similarly, the mass to charge ratio (m / z) is a second dimension of analysis, while the intensity of peaks of the spectral (spectral intensity) is a third dimension. Thus, as used herein, 3D data refers to RT, m / z ratio and spectral intensity data obtained from a chromatographic system. Four-dimensional data includes additional data obtained from a chromatographic system comprising IMS. In an LC- IMS-MS system, as multiple metabolites co-elute off the LC column, their unique characteristics (e.g., collision cross sections) allows further gas phase separation based on the time it takes for an individual metabolite to traverse the IMS cell. This “drift time” is another dimension of measurement. Thus, as used herein, the phrase “4D data” refers to RT, drift time, m / z ratio, and spectral intensity obtained from a LC-IMS-MS system.
[0041] As used herein, the phrases “collapsing data”, “collapsing the data”, “data collapse”, and the like, refer to a transformation of dimensional data involvingCT combining several instances of data in one dimension into a single data point. Such transformation may involve mathematical transformation and may comprise converting the original dimensional data into, for example, a mean, a sum, a median, a difference, a product, etc., datapoint. To illustrate such collapse, in an array (e.g., a table) comprising 4D LC-IMS- MS data, collapse of the data in the IMS dimension may comprise summing the spectral intensities from metabolites having the same RT and m / z ratio and having consecutive drift times. Such transformation results in a dataset containing the original RT and m / z ratio(s), and a set of transformed intensities associate with the RT and m / z ratios(s). According to the present disclosure, such a collapsed dataset is referred to as pseudo-3D data.
[0042] As used herein, the term “feature” refers to the one or more points of data, such as a peak RT or a m / z ratio, that characterize a chemical constituent. For example, a feature obtained using a separation technique, such as chromatography (e.g., HPLC), may comprise a RT or a span of consecutive RTs. Likewise, a feature obtained using mass spectrometry may comprise a m / z ratio or a span of closely related consecutive m / z ratios. In another example, a feature obtained using a separation technique (e.g., chromatography) and mass spectrometry may comprise both a RT, or s span of consecutive RTs, and a m / z ratio or a span of closely related consecutive m / z ratios. A “relevant feature” is a feature that characterizes a unique metabolite. A “non-relevant feature” is a feature that characterizes non-unique metabolites, contaminants, and that are artifacts. It will be understood by those of skill in the art that feature-related data is associated with one or more species forming a characteristic group (e.g., a group of molecules sharing separation characteristics), which may appear as a single peak when such data is visualized. Feature / peak detection may be achieved using computer software, which is known by and available to a person of skill in the art. Such software allows the determination of the boundaries, centers and intensities of two-dimensional (2D) signals obtained from chromatographic systems, a process known as feature detection. Examples of such software include, but is not limited to, centWave, MarkerLynx, metAlign, XCMS and MZmine.
[0043] As used herein, “metabolite” and “biological metabolite”, may be used interchangeably and refer to the set of small molecules comprising substrates, intermediates, and products of metabolism. Metabolites of the disclosure may generally be less than about 2000 kDa in size, optionally less than about 1500 kDa in size. Metabolites include both endogenous metabolites (i.e., produced by an individual), exogenous metabolites (i.e., drugs, environmental toxins, etc.), as well as metabolites produced by theCT microbiome of the individual. Examples of endogenous metabolites include, but are not limited to, peptides, organic acids, fatty acids, triglycerides, cholesterol, phospholipids, sugars, vitamins, and co-factors. It is understood by those of skill in the art that analysis of a metabolite using a technique such as mass spectrometry results in the production of numerous species of metabolite ions. The different species result from such things as fragmentation of the metabolite, incorporation into the metabolite of isotopes such as carbon- 13 or nitrogen-15, and binding of an ion species to such things as salts or solvents, to form adducts. This may result in the production of several signals, each signal coming from a different species of the metabolite. To perform methods of the disclosure, preferably, one single ion species is chosen to represent each metabolite. As used herein, a unique metabolite refers to a select, monoisotopic, single ion species of an intact metabolite. As such, the term “unique metabolite” excludes fragmented species of metabolites, metabolites comprising naturally abundant stable isotopes, such as carbon-13 or nitrogen-15, and ion species of the metabolite other than the selected ion species, such as other adducts of the metabolite. A non-unique metabolite refers to a fragment of a metabolite, a non-monoisotopic form of the selected metabolite and additional adducts of the selected ion species of the metabolite.
[0044] As used herein, the term "biological sample" refers to a sample obtained from an individual, including a sample of biological tissue or fluid origin obtained in vivo or in vitro. Such samples may include, but are not limited to, blood, serum, plasma, urine, cerebrospinal fluid, tears, saliva, sputum, lymph fluids, dialysates, lavage fluids, and fluids derived from organs or tissue.
[0045] The terms “individual”, “subject”, and “patient” are well-recognized in the art and are herein used interchangeably to refer to any human or non-human animal. Examples include, but are not limited to, humans and other primates, including non-human primates such as chimpanzees and other apes and monkey species; farm animals such as cattle, sheep, pigs, seals, goats and horses; domestic mammals such as dogs and cats; laboratory animals including rodents such as mice, rats and guinea pigs; birds, including domestic, wild and game birds such as chickens, turkeys and other gallinaceous birds, ducks, geese, and the like. Unless explicitly stated, the terms individual, subject, and patient by themselves, do not denote a particular age, sex, race, and the like. Thus, individuals of any age, whether male or female, are covered by the present disclosure and include, but are not limited to the elderly, adults, children, babies, infants, and toddlers. Likewise, methods of the disclosure may be applied to any race, including, for example, Caucasian (white),CT African-American (black), Native American, Native Hawaiian, Hispanic, Latino, Asian, and European.
[0046] The present disclosure relates to a method of processing four- dimensional (4D) data from a LC-IMS-MS system, the 4D data comprising or consisting of RTs, drift times, m / z ratios, and spectral intensities resulting from analysis of a chemical or biological sample using the LC-IMS-MS system. In the disclosed method, the data produced by the LC-IMS-MS system is mathematically manipulated to reduce the computational burden without loss of the original data. The first step of the manipulation may comprise collapsing the 4D data in one dimension. For illustrative purposes, the following description will discuss collapse of the data in the IMS (drift time) dimension. To perform such a collapse, the drift times, m / z ratios, and spectral intensities associated with a single RT are used. It will be understood by those of skill in the art that, because two or more metabolites may share some physical characteristics (e.g., charge), an LC feature having a specific RT may comprise multiple metabolites. However, once the multiple metabolites enter an IMS cell, differences in their structures will cause each metabolite to traverse the IMS cell at a different rate, resulting in each of the multiple metabolites having a different drift time. Consequently, the multiple metabolites, each exiting the IMS cell at different times, will result in the production of several MS spectra. Each of these unique MS spectra, while associated with the same RT, will have a unique drift time. Thus, to perform the data collapse, MS spectra associated with a specific RT are compared to identify common m / z ratios. The intensities in each spectrum are then combined to produce a single spectrum. If an m / z ratio is present in only one of the original spectra, then the combined spectrum contains the m / z ratio at its original intensity. If two or more spectra share a common m / z ratio, then the intensities of the common m / z ratio from each spectrum where it appears are mathematically combined to produce a new, transformed intensity. Such mathematical combining may comprise summing the spectral intensities, determining a difference between spectral intensities, determining a mean of the spectral intensities, determining a median of the spectral intensities, determining a product of the spectral intensities, determining a quotient of the spectral intensities, or determining a ratio of the spectral intensities. This process is illustrated in FIG.4B. FIG. 4B shows the individual spectrums resulting from a single RT. Spectrum (i, j) (top spectrum) shares two m / z ratios with spectrum (i, j+1) (middle spectrum), and three m / z ratios with spectrum (i, j+2) (bottom spectrum). Likewise, spectrum (i, j+1) shares two m / z ratios with spectrum (i, j+2). In this illustration, when the spectra areCT combined, the intensities of the common m / z ratios are summed, yielding the spectrum shown in FIG.4B. The collapsed data is referred to as pseudo-3D data.
[0047] Once the data has been collapsed, it is subject to feature detection. Because the collapsed data is represented in a pseudo-3D format, feature detection may be carried out on the collapsed data using commonly available feature detection software, such as centWave. For example, such software may be used to detect a feature having a specific m / z ratio in the pseudo-3D data. Once a feature has been identified, the original 4D data may be reviewed to identify a drift time associated with the identified feature (e.g., m / z ratio). The m / z ratio along with the drift time and retention time may then be used to identify the associated metabolite.
[0048] Prior to identifying a metabolite, the 4D data (e.g., the drift time) may be subjected to further manipulation. For example, because the sampling rate of the IMS cell is high (i.e., drift rate may be measured in milliseconds), it is likely that an individual metabolite will exit the IMS cell over a small range of times and thus be associated with a small range of consecutive drift times. Consequently, when using the pseudo 3D data and specific RT to identify drift times in the 4D data, the specific RT time and mz / ratio may be associated with a number of consecutive drift times. As a result, metabolites exiting the IMS cell may not be fully resolved. Thus, to fully resolve individual features, it may be useful to subject spectral intensities associate with consecutive drift times to convolution. For example, intensities associated with consecutive drift times, which are themselves associated with consecutive RTs, may be arrayed. By summing intensities for all RTs associated with each drift time, closely related features may be further resolved. Such convolution may be used to produce chromatograms and / or driftograms. An example of conducting such convolution is illustrated in FIGS.4C-4E.
[0049] One embodiment of the disclosure is a method of processing four- dimensional (4D) data from a LC-IMS-MS system, the method comprising: a. collapsing 4D data from a LC-IMS-MS system into pseudo-3D data;b. identifying one or more features in the pseudo-3D data;c. using feature-related 3D data to identify corresponding feature-relateddata in the 4D data; and, d. using the identified feature-related- 4D data in combination with thefeature-related 3D data to identify the associated metabolite,CT wherein the 4D data is produced by analyzing chemical or biological sample using the LC-IMS-MS system.
[0050] In certain aspects, the 3D data may comprise RTs, m / z ratios, and spectral intensities resulting from analysis of a chemical or biological sample using the LC- IMS-MS system. In certain aspects, the 4D data may comprise RTs, drift times, m / z ratios, and spectral intensities resulting from analysis of a chemical or biological sample using the LC-IMS-MS system. In certain aspects, collapsing the 4D data may comprise collapsing the 4D data in the IMS (drift time) dimension. In such aspects, collapsing the data may comprise comparing MS spectra associated with a single, specific RT to identify common m / z ratios. In such aspects, the intensities in each spectrum may be combined to produce a single spectrum comprising pseudo-3D data. In certain aspects, the intensities of m / z ratios present in more than one spectrum are mathematically combined, which may comprise summing the spectral intensities, determining a difference between spectral intensities, determining a mean of the spectral intensities, determining a median of the spectral intensities, or determining a product of the spectral intensities. In certain aspects, collapsing the 4D data may comprise mathematically transforming spectral intensities associated with a common m / z ratio, a common RT and having a range of consecutive drift times. In these aspects, mathematically transforming may comprise summing the spectral intensities, determining a difference between spectral intensities, determining a mean of the spectral intensities, determining a median of the spectral intensities, or determining a product of the spectral intensities.
[0051] In certain aspects, the step of identifying one or more features in the pseudo-3D data may comprise determining the boundaries, centers, and intensities of two- dimensional (2D) signals. In certain aspects, the 2D signals may comprise a RT, a m / z ratio and / or a m / z ratio intensity. In certain aspects, the step of identifying one or more features in the pseudo-3D data may comprise using feature detection software.
[0052] In certain aspects, the step of using feature-related 3D data to identify corresponding feature-related data in the 4D data may comprise using an RT and / or an associated m / z ratio to identify one or more drift times associated with the RT and / or the m / z ratio. In certain aspects, drift times associated with the m / z ratio may be subjected to convolution. Such convolution may be performed using intensities associated with consecutive RTs, all of which are associated with a specific drift time. In certain aspects, the step of identifying the associated metabolite may comprise using a RT, the one or more driftCT times, the convolved drift times, and / or the m / z ratio to search a database for a metabolite sharing such features.
[0053] With regard to the collection of data in the present disclosure, it is illustrative to discuss the MS spectra acquired by the LC-IMS-MS analysis as being ordered by RT frames and IMS bins. As used herein, an RT frame refers to data (e.g., drift time, m / z ratio, associated with a particular instant in time, which may be referenced in relation to the RT. Thus, a single frame refers to data at a single point in time that may be defined in, for example, picoseconds, nanoseconds, microseconds, milliseconds, and the like. Each individual frame is associated with 4D data (i.e., RT, drift time, m / z, and spectral intensity) at that specific instant in time. Moreover, the data associated with a frame may be viewed as being in bins, each bin representing the MS spectral data at a single instant in time in relation to the IMS drift. As with the RT, bins may be defined in time units such as picoseconds, nanoseconds, microseconds, milliseconds, and the like. From such description, it will be understood that a frame comprises a plurality of bins, the number of bins in each frame being determined by the drift time sampling rate and the RT sampling rate. Moreover, each bin represents the m / z ratio and related spectral intensity data at a particular instant of drift time. Thus, in the context of RT frames and drift time bins, collapsing data may comprise summing spectral data from all bins within a single frame, thereby yielding a pseudo-3D dataset of frames, each frame comprising an RT, m / z ratios and spectral intensities.
[0054] One embodiment of the disclosure is a method of processing four- dimensional (4D) data from a LC-IMS-MS system, comprising: a. collapsing 4D data associated with at least one frame to produce pseudo-3D data; b. identifying one or more features in the pseudo-3D data;c. using feature-related 3D data to identify one or more bins comprisingfeature-associated 4D data; d. identify corresponding feature-related data in the 4D data; and,e. using the identified 4D data in combination with the feature-related 3Ddata to identify the associated metabolite, wherein the 4D data is produced by analyzing a chemical or biological sample using the LC-IMS-MS system.
[0055] In certain aspects, the 3D data may comprise RTs, m / z ratios, and spectral intensities resulting from analysis of a chemical or biological sample using the LC-CT IMS-MS system. In certain aspects, the 4D data may comprise RTs, drift times, m / z ratios, and spectral intensities resulting from analysis of a chemical or biological sample using the LC-IMS-MS system. In certain aspects, collapsing the 4D data may comprise collapsing the 4D data in the IMS (drift time) dimension. In such aspects, collapsing the data may comprise comparing MS spectra from two or more bins associated with a single, specific RT to identify common m / z ratios. In such aspects, the spectral intensities in each bin may be combined to produce a single spectrum comprising pseudo-3D data. In certain aspects, the intensities of m / z ratios present in more than one bin may be mathematically combined, which may comprise summing the spectral intensities, determining a difference between spectral intensities, determining a mean of the spectral intensities, determining a median of the spectral intensities, or determining a product of the spectral intensities. In certain aspects, collapsing the 4D data may comprise mathematically transforming spectral intensities from all bins associated with a single, specific RT. In certain aspects, mathematically transforming may comprise summing the spectral intensities, determining a difference between spectral intensities, determining a mean of the spectral intensities, determining a median of the spectral intensities, or determining a product of the spectral intensities.
[0056] In certain aspects, collapsing the 4D data may comprise mathematically transforming spectral intensities associated with at least one m / z ratio from one or more bins associated with the frame, wherein the spectral intensities are associated with the same m / z ratio. In certain aspects, collapsing the 4D data comprises mathematically transforming spectral intensities associated with at least one m / z ratio from all bins associated with the frame, wherein the spectral intensities are associated with the same m / z ratio. In certain aspects, collapsing the 4D data comprises mathematically transforming spectral intensities from all bins associated with the frame, wherein if more than one spectral intensity is associated with a common m / z ratio, the mathematical transformation comprises combining all spectral intensities associated with the common m / z ratio. In certain aspects, the mathematical transformation comprises summing the spectral intensities, determining a difference between spectral intensities, determining a mean of the spectral intensities, determining a median of the spectral intensities, or determining a product of the spectral intensities.
[0057] In certain aspects, collapsing the 4D data may comprise spectral binning using spectra a specific RT frame with n m / z peaks. In certain aspects, spectral binning may use an algorithm comprising four steps. First, m / z intensity peaks may be sortedCT by m / z values in ascending order. Next, the relative differences in mz between each pair of peaks in the spectra may be calculated according to Equation 1: mzi−mz
[0058] Equation 1: A+1 ii,i+1= mzi
[0059] The be used to generate a binary array Biof length n-1 by assigning 0 to Biif Ai,i+1is smaller than 0.1 ppm, and assigning 1 to Bi if Ai,i+1 is larger than 0.1 ppm. Next, a splitting array may be generated with j unique numbers by calculating the cumulative sum of array {1,Ci}, after which n peaks are split into j groups based on Ci and the centroid m / z and summed intensities is calculated for each group. The binning algorithm may be iteratively applied to each MS spectra across all RT frames.
[0060] In certain aspects, the step of identifying one or more features in the pseudo-3D data may comprise determining the boundaries, centers and intensities of two- dimensional (2D) signals. In certain aspects, the 2D signals may comprise a RT and a m / z ratio. In certain aspects, the step of identifying one or more features in the pseudo-3D data may comprise using feature detection software.
[0061] In certain aspects, the step of using feature-related 3D data to identify bins comprising feature-associated 4D data may comprise using an RT time and an associated m / z ratio to identify bins comprising a drift time associated with the RT time and the m / z ratio. In certain aspects, the step of identifying the associated metabolite may comprise using the RT, the drift time, and / or the m / z ratio to search a database of a metabolite sharing such features.
[0062] It should be understood that in these methods, the RT frame and IMS bin coordinates are cross-aligned such that the IMS bin coordinates of each RT frame are consistently identical across RT frames.
[0063] One embodiment of the disclosure is a method of identifying a metabolite in a chemical or biological sample, using a method of the disclosure. In certain aspects, the method may comprise: a. obtaining 4D data resulting from analysis of the chemical or biologicalsample using an LC-IMS-MS system; b. collapsing the 4D data into pseudo-3D data;CT c. identifying one or more features in the pseudo-3D data;d. using feature-related 3D data to identify corresponding feature-relateddata in the 4D data; and, e. using the identified 4D data in combination with the feature-related 3Ddata to identify the associated metabolite, thereby identifying the metabolite. In certain aspects, the method may comprise: a. obtaining 4D data resulting from analysis of the chemical or biologicalsample using an LC-IMS-MS system; b. collapsing 4D data associated with at least one frame to produce pseudo-3D data; c. identifying one or more features in the pseudo-3D data;d. using feature-related 3D data to identify one or more bins comprisingfeature-associated 4D data; e. identify corresponding feature-related data in the 4D data; and,f. using the identified 4D data in combination with the feature-related 3Ddata to identify the associated metabolite, thereby identifying the metabolite.
[0064] In these aspects, the 4D data may be obtained by doing, or having done, the analysis of the chemical or biological sample using the LC-IMS-MS system. In these aspects, the 4D data may be obtained at the same time, or immediately prior to, conducting the disclosed method. In these aspects, the 4D data may be obtained at a time prior to conducting the disclosed method. In these aspects, the 4D data may be obtained from a database or a printed record.
[0065] In these aspects, the biological sample may comprise blood, serum, plasma, urine, cerebrospinal fluid, tears, saliva, sputum, lymph fluids, dialysates, lavage fluids, and fluids derived from organs or tissue.
[0066] One embodiment of the disclosure is method of determining a diagnosis or a prognosis in an individual, comprising: using a method of the disclosure to determine the presence, absence or level of one or more metabolites in a biological sample from the individual; and, using the presence, absence and / or level of the one or more metabolites to determine the diagnosis or prognosis.CT In certain aspects, the method comprises determining a metabolic fingerprint from the sample, wherein the metabolic fingerprint is known to be associated with a specific diagnosis or prognosis. As used herein, a metabolic fingerprint refers to the presence, absence, and / or level of two or more metabolites in a sample from an individual having a disease, wherein the presence, absence, and / or level of the two or more metabolites are known to be characteristic of the disease.
[0067] In an example, one or more operations and / or methods described above are performed by the one or more systems of the 4D analysis system 1710 shown in FIG.17.
[0068] FIG. 18 and its associated descriptions provide a discussion of an example computing device that may be used with the various systems described herein. However, the illustrated computing device is an example and is not limiting as a vast number of electronic device configurations may be utilized for practicing various aspects of the disclosure.
[0069] FIG. 18 is a block diagram illustrating physical components (e.g., hardware) of a computing device 1800 with which aspects of the disclosure may be practiced. The computing device 1800 may be integrated or otherwise associated with any of the systems described above with respect to FIG.17. The components of the computing device 1800 described below may have computer executable instructions for executing the various methods and / or operations described herein.
[0070] In a basic configuration, the computing device 1800 includes at least one processing unit 1810 and a system memory 1820. Depending on the configuration and type of computing device, the system memory 1820 may comprise, but is not limited to, volatile storage (e.g., random access memory), non-volatile storage (e.g., read-only memory), flash memory, or any combination of such memories. The system memory 1820 may include an operating system 1830 and one or more program modules 1840 or components suitable for performing the various operations described above.
[0071] The operating system 1830 may be suitable for controlling the operation of the computing device 1800. The system memory 1820 may also include a 4D analysis system 1850 including, but not limited to, the various subsystems described herein.
[0072] The computing device 1800 may have additional features or functionality. For example, the computing device 1800 may also include additional data storage devices (removable and / or non-removable) such as, for example, magnetic disks,CT optical disks, or tape. Such additional storage is illustrated in FIG.18 by a removable storage device 1860 and a non-removable storage device 1870.
[0073] As stated above, a number of program modules 1840 and data files may be stored in the system memory 1820. While executing on the processing unit 18.10, the program modules 1840 may perform the various processes including, but not limited to, the aspects, as described herein.
[0074] Furthermore, examples of the disclosure may be practiced in an electrical circuit comprising discrete electronic elements, packaged or integrated electronic chips containing logic gates, a circuit utilizing a microprocessor, or on a single chip containing electronic elements or microprocessors. For example, examples of the disclosure may be practiced via a system-on-a-chip (SOC) where each or many of the components illustrated in FIG.18 may be integrated onto a single integrated circuit. Such an SOC device may include one or more processing units, graphics units, communications units, system virtualization units and various application functionality all of which are integrated (or “burned”) onto the chip substrate as a single integrated circuit.
[0075] When operating via an SOC, the functionality, described herein, with respect to the capability of client to switch protocols may be operated via application- specific logic integrated with other components of the computing device 18.00 on the single integrated circuit (chip). Examples of the disclosure may also be practiced using other technologies capable of performing logical operations such as, for example, AND, OR, and NOT, including but not limited to mechanical, optical, fluidic, and quantum technologies. In addition, examples of the disclosure may be practiced within a general-purpose computer or in any other circuits or systems.
[0076] The computing device 1800 may also have one or more input / output device(s) 1890. These include, but are not limited to, a keyboard, a trackpad, a mouse, a pen, a sound or voice input device, a touch, force and / or swipe input device, a display, speakers, a printer, etc. The aforementioned devices are examples and others may be used. The computing device 1800 may include one or more communication systems 1880 that allow or otherwise enable the computing device 1800 to communicate with remote computing devices 1895. Examples of suitable communication connections include, but are not limited to, radio frequency (RF) transmitter, receiver, and / or transceiver circuitry; universal serial bus (USB), parallel, and / or serial ports.CT
[0077] The term computer-readable media as used herein may include computer storage media. Computer storage media may include volatile and nonvolatile, removable and non-removable media implemented in any method or technology for storage of information, such as computer readable instructions, data structures, or program modules.
[0078] The system memory 1820, the removable storage device 1860, and the non-removable storage device 1870 are all computer storage media examples (e.g., memory storage). Computer storage media may include RAM, ROM, electrically erasable read-only memory (EEPROM), flash memory or other memory technology, CD-ROM, digital versatile disks (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other article of manufacture which can be used to store information and which can be accessed by the computing device 1800. Any such computer storage media may be part of the computing device 1800. Computer storage media does not include a carrier wave or other propagated or modulated data signal.
[0079] Communication media may be embodied by computer readable instructions, data structures, program modules, or other data in a modulated data signal, such as a carrier wave or other transport mechanism, and includes any information delivery media. The term “modulated data signal” may describe a signal that has one or more characteristics set or changed in such a manner as to encode information in the signal. By way of example, and not limitation, communication media may include wired media such as a wired network or direct-wired connection, and wireless media such as acoustic, radio frequency (RF), infrared, and other wireless media. EXAMPLES
[0080] Example 1. General disclosure of XCiMS workflow
[0081] This Example discusses the production of XCiMS software to process four dimensional LC-IMS-MS data within the XCMS ecosystem. XCiMS integrates spectral reconstruction, multidimensional feature detection, feature correspondence, CCS calculation, and statistical analysis into one R package. The software is compatible with data from Agilent drift tube ion mobility-mass spectrometry (DTIM-MS), Waters traveling wave ion mobility-mass spectrometry (TWIM-MS), and Bruker trapped ion mobility-mass spectrometry (TIM-MS) instruments. A detailed description of each step in the XCiMS workflow is included in the Examples and related Figures. Briefly, after converting the rawCT instrument files into format-free spectra, XCiMS collapses 4D data into pseudo 3D data. Data reduction is accomplished by using the integral of each mobiligram or driftogram as a single intensity value, thereby effectively converting LC-IMS-MS data into pseudo LC-MS data. Using the pseudo 3D (i.e., LC-MS) data, XCiMS then applies the centWave algorithm to detect pseudo 3D features, just as is conventionally done by XCMS when processing LC- MS data. The pseudo 3D features are used to define RT-IM regions of interest (RT-IM ROIs) within the original LC-IMS-MS data. Next, the centwave algorithm is applied again to detect IM peak profiles. Notably, however, IM peak detection is only performed within RT-IM ROIs. Limiting IM peak detection to RT-IM ROIs prevents the unnecessary processing of the IMS dimension when there is no corresponding chromatographic peak. Finally, XCiMS combines IM peak profiles with corresponding precursors (i.e., pseudo 3D features) and produces a table of 4D features.
[0082] Prior to XCiMS analysis, raw data from Agilent drift tube ion mobility-mass spectrometry (DTIM-MS) and Waters traveling wave ion mobility-mass spectrometry (TWIM-MS) instruments are converted into standard mzML format by using MSconvert. Raw data from Bruker trapped ion mobility-mass spectrometry (TIM-MS) instruments are saved as .d files and imported directly. The first step in the workflow is for XCiMS to construct vender agnostic, format-free LC-IM-MS 4D spectra by parsing .mzML files into an OnDiskMSnExp object with the MSnbase package or Bruker .d files into an OpenTIMS object with the opentimsr package followed by extracting MS spectra along all RT frames and IM bins. (FIG.1C) The second step is for XCiMS to construct pseudo LC- MS spectra by using a built-in spectral binning algorithm that reduces the IM dimension of the 4D spectra, thereby collapsing it into a pseudo 3D format. Third, XCiMS applies the centWave algorithm to detect features on the basis of LC peak profiles, mass-to-charge ratios (m / z), and intensity values from pseudo LC-MS spectra, as is conventionally done with 3D data. Fourth, using the retention times and m / z windows as defined by each pseudo 3D feature, XCiMS extracts 2D intensity maps (i.e., RT-IM regions-of-interest or RT-IM ROIs) from original LC-IM-MS spectra and convolves the ROIs to IM mobilograms or driftograms (FIG. 1C). Then, XCiMS again applies the centWave algorithm to detect IM peak profiles from the collection of mobilograms / driftograms. XCiMS combines IM peak profiles with corresponding precursors (pseudo 3D features) and produces in a table of 4D features. In addition to a 4D workflow, XCiMS also provides a 3D data processing workflow for direct infusion, or flow injection-based IM-MS analysis (FIGS.2A & 2B).CT
[0083] When groups of samples are analyzed together, XCiMS can perform feature correspondence along the LC and IM dimensions by using the density algorithm within the XCMS ecosystem, which organizes features into feature groups. In LC-MS data, feature groups are determined by calculating the density distribution of identified LC peaks within every m / z slice at all samples. Peaks with similar retention times are grouped. (XCMS reference) In LC-IMS-MS data, feature groups are determined by two-step groupings on 4D features based on their LC peaks and IM peaks. The two-step groupings are defined by using the same density algorithm with different standard deviations. When two sample classes are processed, XCiMS performs statistical fold-change analysis by calculating the mean intensities of features within each class and calculating p-values of the intensities from a Welch’s t-test. Additionally, collisional-cross-section (CCS) values of 4D features can be calculated by using XCiMS functions that are specific to the type of IMS data collected. External calibration coefficients are required for TWIMS and DTIMS data, but no coefficients are needed for TIMS data. While only one feature detection and correspondence algorithm are currently employed, it is noted that the modular framework of three data layers (raw data layer, spectra data layer, feature layer) established in XCiMS enables easy extension and adaption to other data-preprocessing algorithms from the open-source community. (FIG.3).
[0084] Compared to conventional processing of LC-MS data, processing of LC-IM-MS data requires significantly more computer memory. Using raw data files of three IMS systems, we benchmarked the memory and compute time of XCiMS when using a workstation equipped with an 8-core Xeon W-3225 CPU and 128GB RAM. The memory and compute time depend on the total number of RT frames, IM bins, and m / z values in the data file. For example, processing an LC-DTIM-MS data file with 674 RT frames, 500 IM bins, and over 99-million m / z values in 305,216 spectra took 2.97 GB memory and 1.5 hours, as shown in Table 1.
[0085] Table 1 Platform LC-DTIM-MS LC-TWIM-MS LC-TIM-MS i1CT Total unique IM 502 200 2,464 bins Total m / z values 99327184 294262253 73000578to mzML., . . s an LC- TIM-MS data file with 2,719 RT frames in the same RT range, 2,464 IM bins, and 73-million m / z values in 5-million MS spectra. The memory cost increases with larger numbers of m / z values. Despite a significant decrease in the number of MS spectra, RT frames, and IM bins in the LC-TWIM-MS data file, it still took 7.57 GB memory and 6.2 hours to process the data. This is due to the presence of 294-million m / z values, which is three times more than the number of values in DTIMS and TIMS data files. We point out that the memory usage and compute time for processing TIMS data are highly efficient. Rather than converting TIMS data to an mzML file of 39 GB, XCiMS applies the opentimsr R package to extract spectra directly from Bruker .d files, thus preventing an unnecessary amount of memory loading.
[0087] Example 2. Principles of multidimensional peak detection by XCiMS
[0088] MS spectra acquired by LC-IMS-MS analysis are ordered by RT frames and IM bins (FIG. 4A). We developed a simple and efficient binning algorithm to sort and merge m / z values in multiple MS spectra and ultimately collapse 4D LC-IM-MS spectra into pseudo LC-MS spectra. Binning of MS spectra at a given RT frame with n m / z peaks is achieved in four steps. First, peaks are sorted by m / z values in ascending order. Second, the relative difference in m / z between each pair of peaks is calculated according to equation 1 and used to generate a binary array B_i of length n-1 by assigning 0 to B_i if A_(i,i+1) is smaller than 0.1 ppm and assigning 1 to B_i if A_(i,i+1) is larger than 0.1 ppm. mz
[0089] Equation 1: Ai+1−mzii i+1=CT
[0090] Third, a splitting array C_i is generated with j unique numbers by calculating the cumulative sum of array〖{1,C〗_i}. Fourth, n peaks are split into j groups based on C_i and the centroid m / z and summed intensities is calculated for each group (FIG. 4B). To generate pseudo LC-MS spectra, the binning algorithm is iteratively applied to each MS spectra across all RT frames. The representation of the IM bin is drift time in milliseconds for DTIM-MS and TWIM-MS data, and inverse reduced mobility (1 / K0) in Vs / cm2 for TIM-MS data. Accordingly, we refer to ion traces along IM dimensions as driftograms for DTIM-MS or TWIM-MS data and mobilograms for TIMS-MS data. The RT frame and IM bin coordinates within one file are cross-aligned, meaning the IM bin coordinates at each RT frame are consistently identical across RT frame, and vice versa (FIG.5). This characteristic makes it possible to collapse the LC dimension of RT-IM ROIs and convolve driftograms or mobilograms simply by summing intensities at each IM bin. As an example, the RT-IM ROI for the [M+H]+ ion of adenosine (m / z of 268.1049) was extracted with 20 ppm from 4D spectra acquired by DTIM-MS system (FIG.4C). Each dot represents the summed intensity of all m / z values within the 20-ppm window of m / z 268.1049 in 4D LC-IM-MS spectra at the corresponding RT frame and IM drift time. Compared with extracted ion chromatograms (EIC) from individual IM bins and extracted ion driftograms (EID) from individual RT frames, the convolved ion trace provides significantly higher abundances (FIG.4D & 4E).
[0091] Example 3. Isomers uniquely resolved by IMS produce sub-features
[0092] For some analytes, one pseudo LC-MS peak will only produce a single LC-IMS-MS peak. Such a one-to-one correspondence between pseudo 3D peaks and 4D peaks is expected to occur in cases when IMS does not resolve any isomers of the analyte and the analyte does not form ion clusters in the instrument. For other analytes, on the other hand, one pseudo LC-MS peak will produce multiple LC-IMS-MS peaks. We expect multiple 4D peaks to be measured from a single pseudo 3D peak when IMS enables differentiation of either structural or charged isomers. Another reason that a pseudo 3D peaks might produce multiple 4D peaks is that the analyte forms ion clusters within the instrument that fall apart prior to reaching the detector. For simplicity, when a pseudo 3D peak leads to multiple 4D peaks, we say that the 3D feature has “sub-features”. In this work, we wished to use the XCiMS software to quantitate the number of 3D features in a conventional untargeted metabolomics experiment that produce multiple 4D features (i.e., have sub-features) when using different types of IMS technologies.CT
[0093] Sub-features have the same chromatographic profile but are resolved in the IMS dimension. The summation of all sub-features will have unique profiles in the IM-intensity domain.
[0094] The profile of sub-features in the intensity-IM dimension. The intensity patterns of RT-IM ROI falls in a bimodal distribution along IM coordinate, while only a single distribution is observed along RT coordinate. The distribution pattern is more intuitive by visualizing the convolved EIC and EID. We also noted that across multiple data files, the reproducibility of IM bin coordinates is perfect as 100 percent, while the numeric values of RT frames can deviate from sample to sample (FIGS.6A-6D).
[0095] Example 4. Evaluation of XCiMS using amino acid standards
[0096] To benchmark the performance of XCiMS, an authentic mixture of 18 amino acids (AAmix) was analyzed using hydrophilic interaction liquid chromatography (HILIC) coupled to a Waters Synapt XS mass spectrometer operated in HDMS mode. The AAmix was analyzed in triplicate. The output of XCiMS was compared to manual analysis of the raw data. Serine, glycine, and cysteine were not detected in negative-ion mode. Additionally, although D / L-alanine isomers were present in the mixture, they could not be resolved by HILIC or TWIM-MS. Therefore, 14 amino acids are expected to be readily detected by XCiMS in three technical replicates, yielding a total of 42 reference 4D features. Given that the elution profiles of gas phase ions in ion mobility spectrometry are distinct from LC, optimization of centWave parameters for IM peak detection is necessary. Forty- two (42) pseudo LC-MS features returned by XCiMS were identified as amino acids by matching RT and m / z values to our in-house library. Some of these pseudo LC-MS features corresponded to amino acid adducts, fragments, and dimers. The most abundant pseudo LC- MS feature for each amino acid were then used to tune centWave parameters such as including peakwidth, snthresh, and noise.
[0097] For convenience, the most abundant 3D precursor for each amino acid standard were picked in each replicate. Then, centWave parameters including peakwidth, snthresh, and noise and the extraction ppm of RT-IM ROI (ROI_ppm) were sequentially tuned to maximize detection of IM peaks from 42 reference IM driftograms. With optimized parameters, XCiMS detected LC and IM peaks of all 14 amino acids other than one IM peak of valine in one replicate (FIGS. 7A-7N). LC and IM peak areas of all reference 4D features were highly proportional (FIG.7O), suggesting an accurate and robust peak detection by XCiMS.CT
[0098] Example 5. Assessing quantitative performance from a spike-in experiment
[0099] To demonstrate and evaluate quantitative performance of XCiMS for multi-file processing, a spike-in experiment was designed by spiking a mixture of 63 metabolite standards in SRM1950 human plasma extracts at one and five equivalents respectively. The standard mixture covers different classes of metabolites including nucleotide derivatives, sugar phosphates, organic acids, and amino acids. This generates an artificially created fold change in metabolite pool sizes in the context of biological matrix. Samples were prepared in triplicates at each equivalent and analyzed by HILIC-TWIM-MS workflow. Among 63 metabolite standards spiked in the plasma extract, 56 standards were readily detected by XCiMS across all replicates. 48 of the standards was picked up by statistical fold change analysis (FIG.8A). Overall, XCiMS detected 10,8444D features per replicate, implying a complex background of plasma metabolome.3,570 feature groups were determined after feature correspondence, among which 1,099 groups revealed statistically significant fold changes (FIG. 8B). Fold changes of LC versus IM peak intensities among XCiMS-detected standards were strongly correlated (FIG. 8C). Deviation of fold changes from the five-to-one ratio is due to variable baseline pool sizes of metabolites in SRM1950 plasma extracts. Intracellular metabolites like glyceraldehyde 3- phosphate / dihydroxyacetone phosphate, glucose-6-phosphate, fructose-6- phosphate / glucose-1-phosphate, and ribose-5-phosphate showed fold changes closer to the spiking ratio of 5 (FIGS. 8D-8G), whereas metabolites existing in plasma such as cGMP and glutamate had smaller fold changes (FIGS. 8H-8I). Owing to a relatively large measurement variances within each condition or similar abundances between two conditions, 8 metabolites did not reveal significant changes and were excluded after fold-change analysis (FIGS. 9A-9H). The result suggests that data processing workflow of 4D metabolomics by XCiMS detects and quantitates features and real metabolite signals from complex chemical or biological samples with high sensitivity and accuracy.
[0100] Example 6. The use of XCiMS to optimize instrument parameters for 4D metabolomics
[0101] This example demonstrates the use of XCiMS to extend untargeted metabolomics to a four-dimensional scale. XCiMS was first used to assess the systematic effect of wave height, wave velocity, and two TOF acquisition modes on data quality as shown in Table 2.CT
[0102] Table 2. TWIMS / MS parameters for 4D untargeted metabolomics Parameter set 1 Wave height: 25V; wave velocity 1000-300 m / s; High resolution mode (TOF pusher frequency: 145 us) Parameter set 2 Wave hei ht: 25 V wave velocit 1000-300 m / s Resolution moden mode (Parameter set 1) and resolution mode (Parameter set 3). XCiMS detected 3,642 pseudo 3D features and 3,0944D features under high resolution mode. In contrast, 17,144 and 20,203 pseudo 3D and 4D features were detected using resolution mode, increased by 470% and 660%. It was reasoned that the significant increase in detected features at both LC and IM dimensions is due to an increase in TOF pusher frequency from 145 µs (high resolution mode) to 65 µs (resolution mode). A faster TOF pusher frequency means faster acquisition speed to provide more data points during an ion elution from TWIM-MS separation for centWave to identify. On the contrary, insufficient data sampling at slower TOF pusher frequency will cause information loss and deteriorate both pseudo LC and IM peak shapes, resulting in a significant drop in both pseudo-3D and 4D features as shown.
[0104] Two sets of wave heights and wave velocity were then evaluated under resolution mode (Parameter set 2 and Parameter set 3) using credentialing approach. Credentialing is a stable isotope-based approach developed to reflect biologically relevant features in complex untargeted metabolomics data. Briefly, the credentialed sample was prepared by mixing unlabeled and fully 13C-labeled E. coli extracts in 1:1 and 1:2 ratios. Two HILIC-TWIM-MS analyses using parameter sets 2 and 3 were performed followed by XCiMS processing to generate pseudo-3D and 4D features. Credentialed pseudo-3D and 4D features were determined by credential R package based on 12C and 13C counterparts that pass dual isotope-ratio filters. While the number of credentialed pseudo-3D features detected using TWIMS parameter sets 1 and 3 were similar (2,815 and 2,844), credentialed 4D features were increased from 1,917 using parameter set 1 to 3,255 using parameter set 3.
[0105] Example 7. XCiMS enables 4D untargeted metabolomics
[0106] XCiMS was used to profile 4D features from 11 different datasets that were acquired from TWIM-MS, DTIM-MS, and TIM-MS system (FIG. 19A). The 11 datasets represent a broad range of biological samples such as polar metabolite extracts of E. coli, polar and lipid extracts of SRM1950 human plasma, as well as several metaboliteCT standards including amino acid mixture (AAMix), Mass Spectrometry Metabolite Library Standard (MSMLS), in-house polar metabolite standards (PolarMix), and lipids standards (LipidMix). The resulting 4D features only include those with an intensity that is three times and above higher than the noise level as reported by centWave algorithm. Using E. coli extract as a reference sample, XCiMS detected 27,239, 16,141, and 21,1004D features from datasets acquired by HILIC-TWIM-MS, HILIC-DTIM-MS, and HILIC-TIM-MS platforms, respectively. The distributions of IM elution time, IM peak widths, and IM resolution based on XCiMS-detected 4D features in TWIM-MS, DTIM-MS, and TIM-MS system were summarized. IM drift time of 4D features for TWIMS and DTIMS data were Gaussian distributed with a center at around 5 ms, and 20 ms, whereas TIM-MS records inverse reduced mobility (1 / K0) and shows a bimodal distribution centered at around 0.85 and 1.0 (FIGS. 10A-10C). Distributions of IM peak widths and peak resolutions of the 4D features were also visualized by XCiMS based on peak detection results (FIGS. 10D-10I). This information should provide a general guideline of parameter selection for centWave to detect IM peaks from different ion mobility system. Notably, XCiMS reports 4D sub-features, defined as groups of 4D features that share one LC peak profile (precursor pseudo-3D feature) but are differentiated by IM peak profiles. This is unique information provided by 4D untargeted metabolomics. Among 11 datasets, the prevalence of 4D sub-features is different among three IM-MS system, ranging from 6-9 percent in TWIM-MS, 16-20 percent in DTIM-MS, and over 50 percent in TIM-MS system (FIGS.11A-11G).
[0107] The emergence of 4D sub-features breaks the one ion-one m / z relationship that has been the default in conventional metabolomics data, indicating a previously unforeseen complexity of ion composition introduced by electrospray ionization. One possibility of 4D sub-features may be ESI-induced isomerization, yielding charged isomers. As an example, pairs of credentialed 4D sub-features in unlabeled (M0 of 744.0838) and fully labeled (M21 of 765.1581) form were detected from a HILIC-TWIM-MS dataset (FIG.19B). Accurate mass searches suggest the identity of this ion is [M-H]- NADPH. CCS values of both sub-features were calculated and searched against the CCS Compendium. CCS of the sub-feature at right (249.2 A) matches database entry of NADPH, while the CCS of the first peak does not have a database hit. However, retention time of this NADPH (458 seconds) does not match to in-house RT library (495 seconds). Instead, this sub-feature shares the same RT with NADP+, suggesting the NADPH may originate from in-source reduction as previously reported (FIG.12). The left sub-feature may be a charged isomer, asCT there are multiple possible sites of deprotonation in the molecule. Another similar example is credentialed 4D features of protonated adenosine (M0 of 268.1049 and M10 of 278.1361, FIG. 19C). Interestingly, there are two CCS entries of adenosine [M+H]+ in CCS compendium. XCiMS-calculated CCS values matches both values respectively. The second possibility of 4D sub-features may relate to additional resolving power provided by ion mobility spectrometry that differentiates chromatographically unresolved structural isomers in the sample. For example, our HILIC method cannot resolve fructose-6-phosne (F6P) and glucose-1-phosphate (G1P) in PolarMix. XCiMS detected 4D sub-features of [M-H]- ion of F6P / G1P (FIG. 19D). The identity of two IM peaks were confirmed by matching mobilograms of F6P and G1P standards and MS / MS fragmentation of the IM-resolved ions respectively using PASEF data acquisition (FIG.19D, FIGS.13A-13B). However, because the separation mechanism of LC and IMS are different, each technique has unique selectivity for structural isomers. As a counter example, three pairs of structural isomers, cAMP / A- 2’,3’-cyclic phosphate, AMP / dGMP, and ATP / dGTP in MSMLS analyzed by HILIC- TWIM-MS showed superior chromatographic separation over IMS separation, suggesting that HILIC outperforms TWIMS in resolving these structural isomers (FIGS. 14A-14C). The third possible origin of 4D sub-features can be derived from instrument-induced artifacts. As an example, cytidine and tryptophan in the PolarMix sample analyzed by HILIC-TIM-MS co-elutes after chromatographic separation along with the formation of a [Cytidine+Tryptophan-H] cluster ion (FIGS. 15A-15C). Five major 4D sub-features of Tryptophan [M-H]- ion is highlighted. Only sub-feature 2 has a matched CCS value to the database value of 149.1 Å. Notably, the Tryptophan sub-feature 4 co-elutes with the unique IM peak of the cluster ion, suggesting a dissociation event on the cluster ion after TIMS separation.
[0108] To estimate the prevalence of 4D sub-features with some chemical context, the 4D feature list of PolarMix dataset acquired by HILIC-TIM-MS workflow was analyzed. The distribution of 4D sub-features derived from each of the 109 precursor pseudo-3D features derived from 39 standard compounds in PolarMix suggest a surprising prevalence of 4D sub-features. The distribution suggests most of the precursor features diverges into multiple 4D sub-features ranging from 2 to 10 (FIGS.16A & 16B). A similar trend was observed among the 1,599 precursor pseudo-3D features detected by XCiMS in biological datasets (FIGS. 16C & 16D). In summary, XCiMS may be used toCT uncover a new level of complexity in 4D untargeted metabolomics data driven by LC-IM- MS analysis.
[0109] Example 7. Improved metabolite annotation using accurate mass and CCS
[0110] XCiMS can perform calculations of collision cross section (CCS) from drift time measured by DTIMS, TWIMS, or from inverse ion mobility (1 / K0) measured by TIMS. CCS is a unique molecular descriptor that describes the rotationally averaged surface area of gas phase ion. To demonstrate the utility of CCS for metabolite annotation, an untargeted metabolomics workflow coupling data acquisition by HILIC-TWIM-MS and data processing by XCiMS were performed on mass spectrometry metabolite library standards (MSMLS). CCS were calculated by XCiMS with drift time and an instrument provided exponential calibration equation for TWIM-MS. When calculating CCS, all 4D features were assumed to be [M-H]- ion as proof-of-principle. Among 145 XCiMS-detected MSMLS standards, m / z values were positively correlated with the measured CCS values (FIG. 11A). The measured CCS values were searched against experimental CCS values in unified CCS Compendium. Strong correlation between the measured and databases CCS values suggests an accurate CCS measurement of standards (FIG. 11B). The histogram of relative CCS measurement error indicates a mean relative error of 1.2% (FIG. 11C). Most measured CCS values were within 5% error compared with CCS Compendium. The 4D features with m / z and CCS values of these MSMLS standards were treated as unknown and searched against Human Metabolome Database (HMDB), a larger metabolite database with experimental and predicted CCS values. The average number of database hits per entry is 10.3 when searching by accurate mass only. When searching with accurate mass and CCS, the average database hits is 5.9, representing a 40% reduction. The distribution of database hits of the MSMLS standards when searching by m / z and CCS values resulted in a global reduction in false positive database hits as compared to the distribution of database hits when searching by m / z only. (FIG.11D).
[0111] The utility of CCS for metabolite annotation was further assessed by annotating 4D features from an SRM1950 plasma dataset acquired by HILIC-TWIM-MS and processed by XCiMS with calculated CCS values. Among 32344D features, correlation patterns between m / z values and measured CCS were similar to MSMLS standard (FIG. 11E). The diverged clusters deviating from the diagonal may be different classes of compounds such as peptides and lipids. The m / z and CCS values of the 1000 most abundantCT 4D features was selected to search against HMDB database. Only about half of the features (551) have HMDB database hits when searched using accurate mass only (FIG.11F). When searching with accurate mass and CCS, the number of 4D features with database hits decreased to 97 by 80 percent. Among the 974D features, the average number of HMDB hits is 19.5 when searching by accurate mass only. In contrast, the average hits dropped to 3.1 when searching by accurate mass and CCS. The histogram of HMDB hits when searching with accurate mass and CCS values is compressed and shifted to the very left compared to searching with accurate mass only (FIG. 11G). The result suggests that CCS values as a unique molecular descriptor facilitates metabolite annotation with higher specificity.
[0112] This written description uses examples to disclose the invention, including the best mode, and also to enable any person skilled in the art to practice the invention, including making and using any devices or systems and performing any incorporated methods. The patentable scope of the invention is defined by the claims, and may include other examples that occur to those skilled in the art. Such other examples are intended to be within the scope of the claims if they have structural elements that do not differ from the literal language of the claims, or if they include equivalent structural elements with insubstantial differences from the literal languages of the claims.
Claims
CT What is claimed:
1. A method of processing four-dimensional (4D) data from a liquidchromatography (LC)-ion mobility spectrometry (IMS)-mass spectrometry (MS) (LC-IMS-MS) system, comprising: a. collapsing the 4D data to produce pseudo-3D data;b. detecting one or more features in the pseudo-3D data;c. using feature-related pseudo-3D data to identify 4D data related to thefeature; and, d. using the identified feature-related 4D data in combination with the feature-related pseudo-3D data to identify an associated metabolite, wherein the 4D data is produced by analysis of a sample using the LC-IMS- MS system.
2. The method of claim 1, wherein the 4D data comprises LC retention times (RT),IMS drift times, m / z ratios and related intensities.
3. The method of claim 1 or 2, wherein the 4D data is collapsed in the drift timedimension.
4. The method of any one of claims 1-3, wherein collapsing the 4D data comprisesmathematically combining MS spectra having m / z ratios and associated intensities associated with a single RT, thereby producing a combined spectrum; wherein if an m / z ratio is present in a single spectrum the intensity of the m / z ratio in the combined spectrum is the same as the intensity of the m / z ratio in the single spectrum; and, wherein if two or more spectra share a common m / z ratio the intensity of the common m / z ratio in the combined spectrum is the result of mathematically combining the individual intensities from the two or more spectra.
5. The method of claim 4, wherein mathematically combing comprises anoperation selected from the group consisting of summing the intensities of the MS spectra, determining a difference between intensities of the MS spectra, determining a mean of the MS spectra, determining a median of the MS spectra, determining a product of the MS spectra, determining a quotient of the MS spectra, or determining a ratio of the MS spectra.
6. The method of any one of claims 1-5, wherein feature detection comprisesdetermining the boundaries, centers, and intensities of two-dimensional (2D) signals.
7. The method of claim 6, wherein the 2D signals comprise mass / charge ratios andRT.
8. The method of any one of claims 1-7, wherein feature detection is performedusing feature detection software.
9. The method of any one of claims 1-8, wherein the step of using feature-related3D data to identify 4D data comprises a feature-related RT and / or a feature- related m / z ratio to identify one or more drift times associated with the RT and / or the m / z ratio.
10. The method of claim 9, wherein the feature-related drift time, the feature-relatedm / z ratio and the associated one or more drift times are used to identify the metabolite.
11. A method of processing 4D data from a LC-IMS-MS system, comprising:a. collapsing the 4D data associated with at least one RT frame to producepseudo-3D data; b. detecting one or more features in the pseudo-3D data;c. using feature-related pseudo-3D data to identify one or more binscomprising 4D data associated with the feature; and, d. using the identified feature-related 4D data in combination with the feature-related pseudo-3D data to identify an associated metabolite, wherein the 4D data is produced by analysis of a sample using the LC-IMS- MS system.
12. The method of claim 11, wherein the 4D data comprises LC retention times(RT), IMS drift times, m / z ratios and related intensities.
13. The method of claim 11 or 12, wherein the 4D data is collapsed in the drift timedimension.
14. The method of any one of claims 11-13, wherein collapsing the data comprisescombining the MS spectra from two or more bins associated with the same RT, thereby producing a combined spectrum;wherein if an m / z ratio is only present in a single bin the intensity of the m / z ratio in the combined spectrum is the same as the intensity of the m / z ratio in the single bin; and, wherein if two or more bins share a common m / z ratio, the intensity of the common m / z ratio in the combined spectrum is the result of mathematically combining the individual intensities from the two or more bins.
15. The method of any one of claims 11-14, wherein feature detection comprisesdetermining the boundaries, centers, and intensities of two-dimensional (2D) signals.
16. The method of claim 15, wherein the 2D signals comprise mass / charge ratiosand RT.
17. The method of any one of claims 11-16, wherein feature detection is performedusing feature detection software.
18. The method of any one of claims 11-17, wherein the step of using feature-related 3D data to identify 4D data comprises a feature-related RT and / or a feature-related m / z ratio to identify one or more drift times associated with the RT and / or the m / z ratio.
19. The method of claim 18, wherein the feature-related drift time, the feature-related m / z ratio and the associated one or more drift times are used to identify the metabolite.
20. A method of identifying a metabolite in a sample, comprising obtaining 4D dataproduced by analyzing the sample using a LC-IMS-MS system, and processing the 4D data using the method of any one of claims1-19, thereby identifying the metabolite.
21. The method of any one of claims 1-20, wherein the sample is a biologicalsample or a chemical sample.
22. The method of claim 21, wherein the biological sample comprises blood, serum,plasma, urine, cerebrospinal fluid, tears, saliva, sputum, lymph fluid, dialysate, lavage fluid, or a fluid derived from an organ or tissue.
23. A method of diagnosing an illness in an individual, comprising:a. obtaining 4D data produced by analyzing a sample from the individual usinga LC-IMS-MS system;b. processing the 4D data using the method of any one of claims 1-22 todetermine the presence, absence or level of one or more metabolites in the biological sample; and, c. using the presence, absence and / or level of the one or more metabolites todiagnose the illness.
24. A system for performing the method of any one of claims 1-22.
Citation Information
Patent Citations
Method of recording adc saturation
GB2518491B
Ion detection and parameter estimation of N-dimensional data
JP5542433B2
Data dependent MS / MS analysis
US10615014B2