Improved methods for scalable untargeted metabolomics workflows
The method addresses scalability issues in untargeted metabolomics by using a reference sample to filter out irrelevant features, enhancing the efficiency of metabolomics workflows for large sample sets.
Patent Information
- Application Number
- JP2025532598
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-12-05
- Filing Date
- 2023-12-04
- Publication Date
- 2026-01-21
AI Technical Summary
Current untargeted metabolomics workflows are not scalable for large sample sizes, leading to computational challenges and inefficiencies in processing thousands of samples due to experimental drift and nonlinear measurement errors.
A method involving a reference sample analysis to identify irrelevant features, followed by filtering and applying relevant features to individual samples, reducing computational burden and enabling analysis of larger sample sets by generating metabolic fingerprints.
Enables the analysis of larger sample sets by filtering out irrelevant features, thereby reducing computational requirements and improving the efficiency of metabolomics workflows.
Smart Images

Figure 2026502062000001_ABST
Abstract
Description
[Technical Field]
[0001] (CROSS-REFERENCE TO RELATED APPLICATIONS) This application claims priority to U.S. Provisional Patent Application No. 63 / 386,087, entitled "IMPROVED METHOD FOR SCALABLE UNTARGETED METABOLOMIC WORKFLOW," filed December 5, 2022, the entire disclosure of which is incorporated herein by reference. [Background technology]
[0002] The goal of precision medicine is to identify subgroups within the population for which strategies for the prevention, diagnosis, and treatment of disease conditions can be uniquely tailored. Large-scale, diverse, and longitudinal cohort studies are expected to be essential to accelerate progress toward this goal over the next decade. Indeed, research projects involving hundreds of thousands of participants are already underway in the FinnGen, UK Biobank, and All of Us cohorts. Given the important role of metabolism in health and diagnosis, the application of untargeted metabolomics to large cohorts promises to inform precision medicine practice in a manner that is highly complementary to other techniques, such as genomics. Accordingly, various metabolomics methods have been developed (see, e.g., U.S. Patent Nos. 7,884,318, 8,428,881, and 10,267,777, the entire disclosures of which are incorporated herein by reference). However, currently, standard data processing workflows used in untargeted metabolomics are not suitable for such large sample sizes.
[0003] When performing untargeted metabolomics using liquid chromatography / mass spectrometry (LC / MS), a typical biological sample yields thousands of signals (often referred to as "features") with unique m / z values and retention times. After detecting features from individual biological samples, it is necessary to determine which represent the same analyte across all LC / MS analyses. This process is known as alignment or correspondence determination. Ultimately, the intensities of all features from each aligned sample must be evaluated. Complicating matters is that experimental drift is compound-specific, resulting in nonlinear measurement errors of varying magnitude throughout the experiment. Given the number of features in a traditional untargeted metabolomics experiment, it is impractical to manually inspect every feature on a global scale to process and analyze the data.
[0004] Over the past two decades, various software platforms have been developed for automatically processing untargeted metabolomics data, including mzMine, XCMS, and MS-DIAL. While these informatics tools are highly effective and widely used for analyzing small cohorts, they are not designed to support projects involving thousands of samples. While feature detection methods can generally be applied to each data file individually, making analysis of large sample numbers feasible, albeit cumbersome, algorithms for correspondence determination are not easily scalable. Methods such as Obiwarp in XCMS are designed to process all data files simultaneously on a single computational workstation. As the number of samples evaluated increases, the amount of memory required for data processing increases, ultimately leading to the possibility of analysis being unable to be completed. While the specific number of samples that can be supported depends on the computational power and software available to the user, recent reports have shown that untargeted metabolomics programs reproducibly crash when processing more than 250 data files. While approaches such as SLAW have developed modified algorithms with low computational overhead, user-friendly tools for processing large sample sets remain limited.
[0005] Therefore, there is a need for improved metabolomics workflows that are scalable for the analysis of large numbers of samples (e.g., 1,000 or more) and correspondingly large data sets. The present disclosure addresses this challenge and provides other advantages. Summary of the Invention [Means for solving the problem]
[0006] One aspect of the present disclosure is a method for non-targeted identification of unique biological metabolites in individual samples within a sample set, wherein each individual sample comprises a chemical component, the method comprising: (a) analyzing reference features obtained from a reference sample to identify irrelevant reference features; (b) a filtering step of filtering the reference features by removing irrelevant reference features to generate a set of relevant reference features that characterize unique biological metabolites; (c) applying unrelated reference features and / or related reference features to the sample features obtained from the individual samples to identify the unique biological metabolite composition in the individual samples.
[0007] In one aspect, the disclosed method includes determining unique biological metabolites in all individual samples in a sample set by performing the analyzing, filtering, and applying steps on each sample in the sample set. In one aspect, the disclosed method includes performing the analyzing, filtering, and applying steps on a plurality of the individual samples in the sample set, and then re-analyzing a reference sample to obtain updated reference features and replacing the reference features with the updated reference features. The number of the plurality of samples is less than the total number of individual samples in the sample set. The plurality of samples includes at least two samples, optionally at least three samples, optionally at least five samples, optionally at least 10 samples, optionally at least 20 samples, optionally at least 50 samples, optionally at least 100 samples, or optionally at least 150 samples. The plurality of samples may comprise at least 1%, at least 2%, at least 3%, at least 4%, at least 5%, at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 75%, at least 90%, or at least 95% of the samples in the sample set. The reference sample comprises an aliquot from each of the plurality of individual samples. The reference feature is obtained by a technique comprising subjecting the reference sample to a separation technique or mass spectrometry. The separation technique comprises chromatography. The chromatography comprises gas chromatography or liquid chromatography. In one embodiment, the unrelated reference feature characterizes artifacts, background chemical components, and non-specific biological metabolites. Background chemical components include contaminants. Non-specific biological metabolites include adducts and / or fragment ions. The reference feature comprises reference separation data and reference mass data generated from the reference sample.
[0008] In one aspect, the applying step includes: (a) identifying sample features among the sample features that correspond to irrelevant reference features and removing sample features that correspond to irrelevant reference features to generate a set of relevant sample features; and (b) using the relevant sample features to determine unique biological metabolites in each sample. In one aspect, the applying step includes: (a) identifying sample features among the sample features that correspond to relevant reference features to generate a set of relevant sample features; and (b) using the relevant sample features to determine unique biological metabolites in each sample.
[0009] The disclosed methods may include determining the chemical components in all individual samples in the sample set by performing an applying step on each sample in the sample set, and may include identifying the chemical components of the individual samples by comparing relevant sample features to a library of information including features that characterize the chemical components.
[0010] One aspect of the present disclosure is a method for non-targeted identification of chemical components in individual samples within a sample set, comprising: (a) acquiring reference features comprising separation data and mass data generated from a reference sample; (b) generating a filtered reference feature comprising filtered separation data and mass data by filtering the reference feature by removing separation data and mass data associated with non-unique chemical components; (c) using the filtered reference features to characterize the presence and / or amount of unique chemical components in each sample in the sample set.
[0011] One aspect of the present disclosure is a method for non-targeted identification of chemical components in individual samples within a sample set, comprising: (a) generating a reference signature comprising separation data and mass data from a reference sample using a separation technique and mass spectrometry; (b) collecting and storing reference features; (c) analyzing the stored reference features to identify unrelated reference features including separation and mass data that characterize non-unique chemical components; (d) generating a filtered data set comprising filtered reference features comprising separation data and mass data characterizing unique chemical components by removing irrelevant reference features comprising separation data and mass data characterizing non-unique chemical components; (e) using the filtered reference features to characterize the presence and / or amount of unique chemical components in each sample in the sample set.
[0012] The step of using the filtered reference features comprises: (a) acquiring sample characteristics from individual samples; (b) creating a list of relevant sample features that characterize unique chemical components by identifying sample features that correspond to the filtered reference features, thereby determining the chemical components in the sample.
[0013] One aspect of the present disclosure is a system for implementing any of the methods described above, comprising: (a) a separation device for separating chemical components and generating separation-related signatures that characterize the chemical components; (b) a mass analyzer for performing mass spectrometry on a portion of the separated chemical components to generate mass spectrometry-related signatures that characterize the chemical components; (c) a first module that receives, collects, and / or stores separation-related and / or mass spectrometry-related characteristics; (d) a user interface connected to the first module that provides separation-related features and / or mass spectrometry-related features in a human-usable form.
[0014] In one aspect, the present invention includes a library of features characterizing chemical components generated using a separation device and a mass spectrometer, wherein the library of features characterizing chemical components includes separation features and mass spectrometer features characterizing identified chemical components. The separation device includes a device for performing chromatography. Chromatography includes liquid chromatography or gas chromatography. Liquid chromatography includes HPLC (high performance liquid chromatography) and UPLC (ultra-high performance liquid chromatography). The separation device includes an electrophoresis device. The separation device is connected to a mass spectrometer. Features relevant to separation include peak retention time, peak intensity, and / or peak width. Features relevant to mass spectrometer include mass or m / z value. [Brief explanation of the drawings]
[0015] [Figure 1] Figure 1 shows solid-phase extraction of plasma using a CAPTIVA 96-well plate (Agilent Technologies). One aliquot of plasma was subjected to solid-phase extraction using a CAPTIVA EMR 96-well plate. After elution of the polar fraction (steps 1–8), lipids still bound to the solid-phase extraction material were eluted into a second collection plate, completely dried, and reconstituted for RP separation before being subjected to HRMS analysis. The polar metabolite extract was subjected to HILIC separation and then subjected to HRMS analysis without further manipulation. Created with BioRender.com. [Figure 2A] Figures 2A and 2B show total ion chromatograms: Figure 2A shows representative total ion chromatograms of a blank sample (black), a test sample (purple), and a QC sample (green) of lipid metabolites. [Figure 2B] Figure 2B shows representative total ion chromatograms of the blank sample (black), test sample (purple), and QC sample (green) of polar metabolites. [Figure 3]Figure 3 shows the pipeline for processing metabolomics data. Polar and lipid metabolites were extracted from plasma samples onto 96-well plates and subjected to LC / MS analysis. Pooled samples were prepared for feature detection, MS / MS acquisition, and use as QC samples. Untargeted metabolite analysis was performed on all samples. After feature detection from the pooled samples, background features and degeneracy were filtered. The remaining features were subjected to metabolite identification using DecoID and Lipid Annotator, and the returned putative identifications were manually curated. Peak areas of these metabolites were extracted from the study samples using Skyline. Retention time shifts were manually corrected for each batch for polar metabolites and automatically adjusted for lipid compounds using indexed retention times. Peak areas were imputed and normalized to remove missing values and batch effects from the data. The final output includes metabolite information (name, m / z, retention time) and normalized metabolite intensities for each study and QC sample. [Figure 4A] Figures 4A and 4B show the data preprocessing for a single batch: Figure 4A: Overlay of total ion chromatograms (TIC) of QC samples injected every 12th sample. [Figure 4B] Figure 4B: Visualization of peak boundaries and apexes (shown as black lines) allows for confirmation of stable retention times after importing data files into Skyline. If two compounds, such as leucine and isoleucine, are not fully resolved at baseline, the peaks may not integrate consistently. Peak boundaries can be adjusted and imported for all data files. The baseline signal from the QTOF instrument will result in retention times reported by Skyline, even when no peaks are present (e.g., blank samples). [Figure 5A] 5A to 5C show the deviation in retention time of lipid metabolites. Figure 5A shows the deviation in retention time (RT) (in minutes) for the run order of LPC 0:0 / 18:2. [Figure 5B]Figure 5B shows the retention time (RT) deviation (in minutes) for the run order of TG 48:1. Each dot represents a sample. The dots are color-coded by analytical batch. [Figure 5C] Figure 5C shows a histogram of all retention time deviations across all samples. [Figure 6A] Figures 6A-6E show a comparison of individual and pooled samples. Figure 6A is a Venn diagram showing the breakdown of features detected in at least one individual sample (red) and at least one pooled sample (green). Searches of the Human Metabolome Database and KEGG revealed a high proportion of features not detected in the pooled sample. [Figure 6B] Figure 6B is a histogram showing the percentage of samples in which a feature was detected, showing the distribution of features detected only in individual samples. [Figure 6C] FIG. 6C is a histogram showing the percentage of samples in which the feature was detected, illustrating the distribution of features detected in both individual and pooled samples. [Figure 6D] FIG. 6D is a histogram showing the log10 transformed mean intensity of features detected in individual samples only. [Figure 6E] FIG. 6E is a histogram showing the log10 transformed mean intensities of features detected in both individual and pooled samples. [Figure 7A] Figures 7A-7C show a comparison of targeted extraction of peak areas with traditional global peak integration. The pie chart in Figure 7A shows the percentage of missing values (gray) among all measurements when targeted extraction of peak areas is performed in Skyline. In Figure 7A, the percentage of missing values is too small to be displayed in the pie chart. [Figure 7B] The pie chart in Figure 7B represents the percentage of missing values (gray) among all measurements when conventional peak integration is performed in XCMS. [Figure 7C] Figure 7C is a box plot showing the coefficient of variation (CV) for peak integration using Skyline targeted extraction (blue) or XCMS conventional processing (red). [Figure 8A] Figures 8A-8E show the correction of batch effects in metabolomics data. Figure 8A: Principal component analysis (PCA) of non-normalized lipid metabolic profiles shows a strong batch effect. Each dot represents a unique sample. Dots are color-coded according to the corresponding batch number. [Figure 8B] Figure 8B shows a comparison of 14 batch correction algorithms for lipid metabolic profiles. The normalized score is the change in coefficient of variation (CV) in the study samples (for non-normalized data) divided by the change in CV in the QC samples. A higher score indicates less technical variation. [Figure 8C] Figure 8C: PCA plot of random forest normalized lipid metabolic profiles shows reduced clustering by batch. [Figure 8D] FIG. 8D shows the intensity of CE 16:0 as a function of run order for both the non-normalized data (top) and the random forest corrected data (bottom). [Figure 8E] Figure 8E shows the violin plots of the CV distributions for all compounds in the QC samples for each batch correction algorithm evaluated. The corresponding data for polar metabolites are shown in Figure 9. [Figure 9A] Figures 9A-9E show the correction of batch effects in metabolomics data. Figure 9A shows principal component analysis (PCA) of unnormalized polar metabolic profiles showing clustering by batch. Each dot represents a unique sample. Dots are color-coded according to the corresponding batch number. [Figure 9B] Figure 9B shows a comparison of 14 batch correction algorithms for polar metabolic profiles. The normalized score is the change in coefficient of variation (CV) in the study samples (for non-normalized data) divided by the change in CV in the QC samples. A higher score indicates less technical variation. [Figure 9C]Figure 9C shows a PCA plot of the random forest-normalized polar metabolic profiles. Samples are still clustered by batch, but this is due to biological differences between batches in the study samples. QC samples are not clustered by batch after correction (see Figure 11). [Figure 9D] Figure 9D shows the intensity of gingerol as a function of run order for both the non-normalized data (top) and the random forest corrected data (bottom). [Figure 9E] Figure 9E shows violin plots showing the CV distribution of all compounds in the QC samples for each batch correction algorithm evaluated. The corresponding lipid metabolite data are shown in Figure 8. [Figure 10A] Figures 10A and 10B show that QC samples enable the differentiation of batch and biological effects. Metabolite intensities can vary as a function of analytical batch. Because different types of variation cannot be easily distinguished when batches are biased by biologically distinct sample groups, applying non-QC-based batch normalization can remove biological variation in addition to technical variation. However, using QC samples allows technical variation due to batch number (e.g., DGTS17:0) to be distinguished from biological differences between batches (e.g., inosine). The plot in Figure 10A shows the non-normalized metabolite intensities as a function of run order for QC samples (left) and study samples (right). [Figure 10B] Figure 10B shows the random forest-corrected metabolite intensities for QC and study samples of the same compounds, with dots color-coded by analytical batch. [Figure 11A] Figures 11A-11D show internal standard variation. Figure 11A is a violin plot showing the distribution of the coefficient of variation (CV) of the internal standard across all samples within a batch or across all batches (1-22) for non-normalized data. [Figure 11B] Figure 11B is a violin plot showing the distribution of the coefficient of variation (CV) of the internal standard across all samples within a batch or across all batches (1-22) in the random forest corrected data. [Figure 11C] FIG. 11C shows a principal component analysis (PCA) of the internal standard intensities across all samples in the non-normalized data. [Figure 11D] Figure 11D shows a principal component analysis (PCA) of the internal standard intensities across all samples in the random forest-corrected data. Each dot represents a sample. Samples are color-coded by batch. [Figure 11E] FIG. 11E is a violin plot showing the distribution of CVs of the internal standards across all samples for the 14 batch correction methods evaluated in this study. [Figure 12A] Figures 12A-D show that random forest normalization reduces batch effects in QC samples. Figure 12A: Principal component analysis (PCA) of non-normalized metabolic profiles of QC samples shows a strong batch effect in polar metabolites. [Figure 12B] Figure 12B: Principal component analysis (PCA) of the non-normalized metabolic profiles of QC samples shows a strong batch effect in lipid metabolites. [Figure 12C] FIG. 12C: After random forest correction, QC samples show no clustering by batch in polar metabolites. [Figure 12D] Figure 12D: After random forest correction, QC samples show no batch-specific clustering in lipid metabolites. Each dot represents a QC sample. Dots are color-coded by lot number. [Figure 13A] Figures 13A-13D show that metabolic profiles reflect geographic location. Figure 13A: Principal component analysis (PCA) of normalized metabolic profiles (polar and lipid metabolites) shows clustering based on field site in the United States (BU = Boston, NY = New York, PT = Pittsburgh) and Denmark (Odense). Each dot represents a unique sample. Dots are color-coded by geographic location. [Figure 13B] FIG. 13B shows lipid metabolites associated with geographic location (|FC|>2, p<0.05, one-way ANOVA). [Figure 13C]FIG. 13C shows polar metabolites associated with geographic location (|FC|>2, p<0.05, one-way ANOVA). [Figure 13D] Figure 13D shows the age distribution of samples from different field sites. Data shown are median ± interquartile range. NDHB, N,N-diethyl-4-hydroxybenzamide; CMPF, 3-carboxy-4-methyl-5-propyl-2-furanpropanoic acid. [Figure 14A] Figures 14A-14D show the analysis of unknown metabolites. Figure 14A is a heatmap showing the relative metabolite intensities across all samples of the 3421 unknown metabolites profiled in this study (rows = metabolites, columns = samples). [Figure 14B] FIG. 14B is a histogram showing the coefficient of variation (CV) of QC samples for all unknown metabolites after random forest batch correction. [Figure 14C] Figure 14C is a heat map of 29 unknown metabolites significantly associated with field site. [Figure 14D] Figure 14D is a box plot showing an example of an unknown metabolite that exhibits significantly lower intensity in Denmark (DK) than in the US. The color of the columns in Figures 14A and 14C indicates the field site of each sample (BU = Boston, NY = New York, PT = Pittsburgh, DK = Odense). [Figure 15] Figure 15 shows that metabolites associated with geographic location do not reflect the age of the subjects. Principal component analysis (PCA) of normalized metabolic profiles (polar and lipid metabolites) reveals no age-dependent pattern in metabolites associated with geographic location. Each dot represents a unique sample. The dots are color-coded according to the age of the subjects. DETAILED DESCRIPTION OF THE INVENTION
[0016] The present disclosure relates to improved methods for performing metabolomics workflows. More specifically, the present disclosure relates to metabolomics workflows that reduce the computational burden required to analyze the abundance of compounds present in a sample, thereby enabling the process to be scaled to analyze larger quantities of samples than can be analyzed with currently available technology. The disclosed process achieves this scalability by using a reference sample (standard sample) that represents the chemical complexity of the entire sample set. Analysis of the reference sample can detect and identify thousands of features in a single sample, most of which do not characterize unique biological metabolites. These features can then be filtered to remove irrelevant features that characterize irrelevant chemical components, resulting in a list of relevant features. The list of relevant features can then be used to identify relevant features that characterize unique biological metabolites in individual samples, thereby reducing the computational burden required to analyze the sample set. Thus, the disclosed method can generally be performed by analyzing a reference sample to obtain biologically relevant reference features and using the relevant reference features to identify relevant features in one or more individual samples within the sample set. The relevant features can then be quantified to generate metabolic fingerprints of biological metabolites in the individual samples. Several relevant features and / or metabolic fingerprints can be used to identify the individual biological metabolites present in an individual sample. In certain embodiments, the metabolic fingerprints of many or all of the individual samples in a sample set can be determined by sequentially or consecutively evaluating the relevant features in each sample in the sample set.
[0017] Before the present invention is further described, it is to be understood that this invention is not limited to particular embodiments described, as these may, of course, vary. Furthermore, it is to be understood that the scope of the present invention will be limited only by the claims and that the terminology used herein is for the purpose of describing particular embodiments, and is not intended to limit the scope of the invention.
[0018] It should be noted that, as used herein and in the appended claims, the singular terms "a," "an," and "the" include the plural of their referents unless the context clearly dictates otherwise. For example, "a" metabolite also includes one or more metabolites. Thus, "a," "an," "one or more," and "at least one" can be used interchangeably.
[0019] Similarly, the terms "comprising," "including," and "having" can be used interchangeably. As used herein, the term "comprising" can be replaced, where appropriate, with the terms "consisting of" or "consisting essentially of" in certain embodiments.
[0020] Further, it should be noted that the claims may be drafted to exclude any optional element. Accordingly, this specification is intended to serve as an antecedent basis for using exclusive terminology, such as "solely," "only," and the like, as well as for using negative limitations in connection with the recitation of claim elements.
[0021] Various terms relating to aspects of the present disclosure are used throughout the specification and claims. These terms are intended to have their ordinary meaning in the art unless otherwise indicated. Other specifically defined terms are to be interpreted in a manner consistent with the definitions set forth herein.
[0022] Unless expressly stated otherwise, the methods or aspects described herein are not intended to be construed as requiring that their steps be performed in a particular order. Thus, unless the claims and specification specifically state that the method steps are to be limited to a particular order, no order is intended to be inferred in any respect. This applies to all implicit bases of interpretation, including the logic of the arrangement of steps or operational flow, the apparent meaning derived from grammatical construction and punctuation, or the number and type of aspects described in the specification.
[0023] It will be understood that certain features of the invention, which are, for clarity, described in the context of separate embodiments, may also be provided in combination with each other in a single embodiment. Conversely, various features of the invention, which are, for brevity, described in the context of a single embodiment, may also be provided separately or in any suitable subcombination. All combinations of the embodiments of the invention are specifically embraced by the invention and are disclosed herein as if each such combination were individually and explicitly disclosed. In addition, all subcombinations are also specifically embraced by the invention and are disclosed herein as if each such subcombination were individually and explicitly disclosed.
[0024] As used herein, the term "about" means that the indicated numerical value is an approximation and that small variations will not significantly affect the practice of the embodiments of the present disclosure. When numerical values are used, unless the context indicates otherwise, the term "about" means that the numerical value may vary within a range of ±10% and remain within the scope of the embodiments of the present disclosure.
[0025] The publications referenced herein are provided solely for their publication prior to the filing date of the present application. Nothing herein is to be construed as an admission that the present invention is not entitled to antedate such publication by virtue of prior invention of such publication. Further, the stated publication dates may be different from the actual publication dates, which may need to be independently confirmed.
[0026] One aspect of the present disclosure is a method for non-targeted identification of unique biological metabolites in individual samples in a sample set, where the individual samples comprise chemical constituents, the method comprising: (a) an analyzing step of analyzing reference features obtained from reference samples to identify irrelevant reference features; (b) a filtering step of filtering the reference features by removing the irrelevant reference features to generate a set of relevant reference features that characterize the unique biological metabolites; and (c) an applying step of applying the irrelevant and / or relevant reference features to sample features obtained from the individual samples to identify the composition of the unique biological metabolites in the individual samples.
[0027] As used herein, non-targeted determination refers to characterizing one or more chemical components in a complex sample without prior knowledge thereof. The term "chemical component" refers to a molecule present in a sample. "Chemical component" includes any molecule present in a sample, including, but not limited to, metabolites such as endogenous metabolites, exogenous metabolites, native metabolites, and non-native metabolites, as well as contaminants.
[0028] As used herein, the terms "metabolite" and "biological metabolite" are used interchangeably and refer to a set of small molecules, including substrates, intermediates, and metabolites. Metabolites in the present disclosure generally have a size of less than about 2000 kDa, and optionally have a size of less than about 1500 kDa. Metabolites include both endogenous metabolites (i.e., produced by an individual) and exogenous metabolites (i.e., drugs, environmental toxins, etc.). Examples of endogenous metabolites include, but are not limited to, organic acids, fatty acids, triglycerides, cholesterol, phospholipids, sugars, vitamins, and cofactors. Those skilled in the art will appreciate that analysis of metabolites using techniques such as mass spectrometry results in the generation of various types of metabolite ions. These various types of metabolite ions arise from fragmentation of the metabolite, incorporation of isotopes such as carbon-13 and nitrogen-15 into the metabolite, and the binding of ionic species to salts, solvents, etc. to form adducts. As a result, multiple signals are generated from various types of metabolites. To practice the methods of the present disclosure, it is preferable to select a single ionic species representative of each metabolite. As used herein, a unique metabolite refers to a selected monoisotopic single ionic species of an intact metabolite. Thus, the term "unique metabolite" excludes fragmented species of a metabolite, metabolites containing naturally abundant stable isotopes such as carbon-13 or nitrogen-15, and ionic species of a metabolite other than the selected ionic species (e.g., other adducts of a metabolite). Non-specific metabolites refer to fragments of a metabolite, non-monoisotopic forms of a selected metabolite, and additional adducts of a metabolite of the selected ionic species.
[0029] The term "background" refers to features that characterize contaminants and artifacts. The term "contaminant" refers to chemicals that are present in a sample but originate from outside the sample (e.g., the instrument used to analyze the sample). For example, the tube holding the sample may contain a plasticizer that can leach into the sample, thereby causing the sample to contain plasticizer. The plasticizer in the sample is a contaminant. As used herein, "artifact" refers to a feature that is not a feature that characterizes an actual chemical, but rather is due to electronic noise in the mass analyzer or errors in downstream peak detection software. Methods for identifying contaminants and artifacts are known to those of skill in the art and are disclosed herein.
[0030] As used herein, the term "feature" refers to one or more data points, such as peak retention time or mass-to-charge (m / z) value, that characterize a chemical component. For example, a feature obtained using a separation technique such as chromatography (e.g., HPLC) includes a retention time (rt). Similarly, a feature obtained using mass spectrometry includes an m / z value. In another example, a feature obtained using a separation technique (e.g., chromatography) and mass spectrometry includes both a retention time (rt) and an m / z value. A "relevant feature" refers to a feature that characterizes a unique metabolite. An "irrelevant feature" refers to a feature that characterizes a non-unique metabolite, contaminant, or artifact.
[0031] Features of the present disclosure (e.g., reference features, sample features, etc.) are obtained from samples of the present disclosure (e.g., reference samples, individual samples) using various methods. In certain embodiments, features are obtained by subjecting the sample to a separation technique. Any separation technique that adequately separates the chemical components in the sample can be used. Examples of separation techniques include, but are not limited to, chromatography or electrophoresis (such as capillary electrophoresis). Suitable methods of chromatography include, but are not limited to, gas chromatography (GC), high performance liquid chromatography (HPLC), or ultra performance liquid chromatography (UPLC). In certain embodiments, features are obtained by subjecting the sample to mass spectrometry. In certain embodiments, features are obtained by subjecting the sample to chromatography, such as HPLC or UPLC, and mass spectrometry. Thus, in certain embodiments, features include retention time (rt) and m / z value.
[0032] As used herein, the term "biological sample" refers to a sample obtained from an individual, such as a sample of biological tissue or fluid origin obtained in vivo or in vitro. Such samples include, but are not limited to, blood, serum, plasma, urine, cerebrospinal fluid, tears, saliva, sputum, lymph, dialysate, lavage fluid, and fluids derived from organs and tissues. An "individual sample" refers to a biological sample obtained from a single individual. The term "sample set" refers to a defined plurality of individual samples.
[0033] The terms "individual," "subject," and "patient" are well known in the art and are used interchangeably herein to refer to any human or other animal. Examples include, but are not limited to, primates such as humans (including chimpanzees, other apes, and monkeys), livestock (including cattle, sheep, pigs, seals, goats, horses, etc.), domesticated mammals (including dogs, cats, etc.), laboratory animals (including rodents such as mice, rats, and guinea pigs), and birds (including poultry, wild birds, and game birds, e.g., chickens, turkeys, other geese, ducks, geese, etc.). The terms "individual," "subject," and "patient," by themselves, do not denote a particular age, sex, race, etc. Accordingly, the present disclosure is directed to individuals of any age, both sexes, including, but not limited to, elderly individuals, adults, children, infants, newborns, young children, etc. Similarly, the methods of the present disclosure may be applied to people of all races, including, for example, Caucasians (whites), African Americans (blacks), Native Americans, Native Hawaiians, Hispanics, Latinos, Asians, and Europeans.
[0034] The reference sample used in the disclosed methods must capture the complexity of the chemical components present in the sample set. Therefore, any suitable reference sample capable of adequately capturing such complexity can be used. For example, in the case of a sample set including individual samples obtained from blood (e.g., plasma), the reference sample can be a sample designed to represent "normal" human plasma. One example of such a reference sample is Standard Reference Material (SRM) 1950 (Metabolites in Human Plasma) available from MilliporeSigma (SKU# NIST1950). Alternatively, the reference sample can be created by combining aliquots from multiple individual samples in the sample set. While any number of aliquots may be combined, it will be understood that the more aliquots combined, the more accurate the results. In certain embodiments, the reference sample includes aliquots from at least about 10%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, or 100% of the individual samples in the sample set. In certain embodiments, the reference sample comprises aliquots from at least 2, at least about 5, at least about 10, at least about 25, at least about 50, at least about 100, at least about 150, at least about 200, at least about 250, at least about 500, at least about 750, at least about 1000, at least about 1500, or at least about 2000 individual samples in the sample set.
[0035] As described above, the reference features are obtained from a reference sample. In certain embodiments, the reference features are obtained from a sample at a time significantly prior to the time at which the analysis of the reference features is performed (e.g., days or weeks prior). Furthermore, the entity that obtains the reference features from the reference sample and the entity that performs the analysis can be, but are not necessarily, the same entity. For example, the reference features can be stored, such as in a feature matrix, and retrieved later for analysis. In certain embodiments, obtaining the reference features from the reference sample is part of a method of the present disclosure, and obtaining and analyzing the reference features are relatively simultaneous. In certain embodiments, the reference features are obtained by subjecting the reference sample to a separation technique and / or mass spectrometry. In certain embodiments, the reference features are obtained by subjecting the reference sample to a separation technique and / or mass spectrometry to generate reference features including retention times (rt) and m / z values.
[0036] Filtering reference features involves identifying and removing irrelevant reference features to generate a set of reference features (e.g., a feature matrix, feature table, feature list, etc.) that characterize unique biological metabolites. The composition of unique biological metabolites in each individual sample can then be determined by analyzing data obtained from the individual samples using the irrelevant and / or relevant reference features. Such methods are advantageous because they limit the computations for identifying relevant and irrelevant features to the reference sample, thereby reducing the computational burden on the entire set of individual samples. In certain embodiments, analyzing data from the individual samples using relevant reference features includes applying the relevant reference features to the sample features. In such embodiments, applying the relevant reference features to the sample features includes comparing the relevant reference features to the sample features, identifying sample features that correspond to the relevant reference features, and using the sample features that correspond to the relevant reference features to determine the composition of unique biological metabolites in each sample. As used herein, with respect to comparing features, "corresponding features" refers to features that yield identical data points (e.g., identical retention time (rt), identical m / z value, or identical retention time (rt) and m / z value) from two different samples. In certain embodiments, applying relevant reference features to sample features includes limiting the analysis of data obtained from individual samples to data that have the retention time (rt) and / or m / z value of the relevant reference feature.
[0037] In certain embodiments, applying irrelevant reference features to sample features includes ignoring (i.e., excluding from further analysis) the sample features that correspond to the irrelevant reference features. The remaining set of relevant sample features characterize the unique biological metabolites in the sample from which the sample features were obtained and can be used to identify the composition of the unique biological metabolites in an individual sample.
[0038] Thus far, this disclosure has described applying unrelated and related reference features to sample features from a single individual sample. However, as will be appreciated by those skilled in the art, because the reference sample reflects the complexity of the entire sample set, the above process of applying reference features to sample features may be repeated for each individual sample in the sample set. The result of this repetition is the determination of the unique metabolite composition in each individual sample in the sample set.
[0039] Those skilled in the art will understand that when performing an iterative analytical process, environmental factors (e.g., temperature, humidity, etc.) affecting the measurement equipment can cause fluctuations in measurements. Such fluctuations in measurements can alter the characteristics obtained for a particular metabolite, making it difficult to compare and match characteristics between samples analyzed at different times. The inventors surprisingly discovered that this problem can be solved by applying the concept of indexed retention time (iRT) in the separation step. Therefore, when the feature identification and / or application step is repeated, the reference characteristics can be updated by periodically subjecting a reference sample to separation techniques and mass spectrometry. Preferably, several (e.g., 2, 3, 4, 5, or 6) chemical components with high detection intensities and easily distinguishable from other chemical components in the separation step peak are selected as indicator compounds. The change in retention time of these compounds between previous and subsequent "runs" can be easily determined, and this difference can be applied to all retention times. In certain embodiments, after the analysis of multiple individual samples, the reference sample is reanalyzed to update the reference characteristics. In certain embodiments, the plurality of individual samples comprises at least 2, at least 3, at least 5, at least 10, at least 20, at least 50, at least 100, or at least 150 samples.
[0040] This written description uses examples to disclose the invention, including the best mode, and also to enable any person skilled in the art to practice the invention, including making and using any devices or systems and practicing any methods incorporated therein. The patentable scope of the invention is defined by the claims, and may include alternative embodiments that occur to those skilled in the art. Such alternative embodiments are within the scope of the claims if they include elements that do not deviate from the literal language of the claims, or if they include equivalent elements that do not deviate substantially from the literal language of the claims.
[0041] Example
[0042] sample
[0043] Blood samples were collected at participants' homes in dipotassium ethylenediaminetetraacetic acid (K2-EDTA) collection tubes and immediately stored in frozen gel packs. Samples were transported to the laboratory by courier or mail. Upon arrival, plasma was separated by centrifugation and stored at -80°C. A pooled sample was prepared from a portion of the plasma samples. This pooled sample served as a quality control (QC) sample and was used for peak list generation and metabolite identification (see Supporting Information). QC samples were prepared from 58 samples in the first analytical batch, thereby avoiding subjecting the samples to additional freeze-thaw cycles. Because some samples were unavailable at the start of the analysis (blood collection is still ongoing), it was not possible to create a pooled sample from all subject samples in this study. QC aliquots were stored at -80°C. Additionally, SPLASH Lipidomix (Avanti Polar Lipids), a deuterium-labeled lipid mixture designed for human plasma analysis, was used as an internal standard for QC samples for lipid metabolite analysis. A QC sample was injected after every 12th study sample.
[0044] Metabolite extraction, LC / MS, and LC / MS / MS analysis
[0045] To maximize detection range while minimizing sample preparation time, solid-phase extraction (SPE) was performed to isolate polar and lipid metabolites. Plasma samples were thawed on ice and vortexed one batch at a time. Each batch typically consisted of 92 study samples, two QC samples, and two blank samples. Aliquots of each sample were then transferred to a 96-well SPE plate. Polar and lipid metabolites were separated into separate fractions by a two-step extraction using acetonitrile and methanol or methyl tert-butyl ether and methanol (Figure 1). The use of SPE eliminates the need for centrifugation to remove the protein fraction of the sample. The lipid extract was dried under a stream of nitrogen and reconstituted prior to analysis by reversed-phase (RP) chromatography coupled to high-resolution mass spectrometry (HRMS) in positive ion mode. The polar metabolite extract was analyzed directly by hydrophilic interaction liquid chromatography (HILIC) coupled to HRMS in negative ion mode without a drying step. Samples were randomized for both LC and MS analysis. Additionally, LC / MS / MS data was acquired to aid in metabolite identification.
[0046] Positive and negative mode data for lipid and polar metabolite extracts were not collected, but would be beneficial if resources permitted. Blank samples were injected at the beginning and end of each worklist to detect / remove background peaks.
[0047] Representative total ion chromatograms of blank, test, and QC samples for both lipid and polar metabolite extracts are shown in Figures 2A and 2B. Although significant signals were observed for the blank samples, features whose intensities in the study samples were at least three-fold less than those in the blank samples were excluded from downstream analysis. Details of metabolite extraction and LC / MS analysis are provided in the Supporting Information.
[0048] Peak list generation
[0049] The polar metabolite peak list was created by combining the results of centwave7 peak detection (within XCMS), background subtraction, and addict selection (CAMERA) from six pooled samples from different sample subsets. Furthermore, a similar workflow was performed on three additional pooled samples within AcquireX software, and the unfiltered features from this analysis were combined with the XCMS results. The R and Python scripts used to perform the peak detection analysis are available on GitHub (https: / / github.com / e-stan / metabolomics_workflow) and include values for all parameters used. The lipid metabolite peak list was created directly based on the identification results from Lipid Annotator. Any workflow or software can be used to create the peak list.
[0050] Metabolite identification
[0051] Identification of polar metabolites was aided by matching accurate mass and MS / MS fragmentation data with in-house and online MS / MS libraries created from certified standards using DecoID software. In the online database search, the top hit for each feature with a dot-product similarity of 80 or greater was considered a putative identification. These results were further filtered by in-house retention time, predicted retention time from a ReTip-like method, and manual curation to remove noise peaks, interfering peaks, and false identifications. MSI identification levels for polar and lipid metabolites are listed in Tables S1 and S2, respectively. The code and scripts used to perform the automated portion of the metabolite identification workflow are available on GitHub. Lipid repeat MS / MS data were annotated with Lipid Annotator software (Agilent Technologies). Lipid identification results were provided as total compositions because there was insufficient information to estimate specific fatty acid compositions. Lipid identification followed the same manual curation applied to polar metabolite data. Any workflow or software can be used for compound identification.
[0052] Peak area extraction
[0053] Following peak list creation and metabolite identification, all data files were analyzed in batches in Skyline (version 20.1.0.155) to obtain peak areas. Using m / z values from the metabolite target list, peak areas were extracted by considering retention time or indexed retention time (iRT) (see Supporting Information). Because the data were acquired over several months, 14 different batch correction approaches were tested for peak area normalization (see Supporting Information). Additionally, a report containing the collection times of all samples was exported for use in batch correction.
[0054] Support information
[0055] Metabolite extraction
[0056] Participant plasma was collected in dipotassium ethylenediaminetetraacetic acid (K2-EDTA) tubes and stored at -80°C after collection before thawing on ice. A 50 μL aliquot was transferred to a solid-phase extraction (SPE) CAPTIVA-EMRLipid 96-well plate (Agilent Technologies), followed by the addition of 200 μL of acetonitrile:methanol (1:1) (v / v) containing 12.5 μM internal standard (consisting of uniformly labeled amino acids with 13C and 15N, MSK-A2-1.2, Cambridge Isotope Laboratories, Inc.). Ideally, isotopically labeled internal standards would be used for various compounds, but due to resource constraints, only an isotopically labeled amino acid mixture was added to the extraction solvent. While this is sufficient to assess technical variability, the internal standard was not used for quantification in this study. Samples were mixed at room temperature for 1 min on an orbital shaker (360 rpm) before a 10-min incubation period at 4°C. The samples were then mixed with 150 μL of 2:2:1 acetonitrile:methanol:water (v / v / v). The samples were then mixed on an orbital shaker (360 rpm) at room temperature for an additional 10 minutes. Using a positive pressure manifold (Biotage PRESSURE+96), the samples were eluted into a 96-deep-well collection plate. A second elution was performed with 100 μL of 2:2:1 acetonitrile:methanol:water (v / v / v). The polar eluate was stored at -80°C until lipid data collection was complete and prior to LC / MS analysis (2.5 months).
[0057] Lipid metabolites still bound to the SPE material were then eluted into a second elution plate using two elution steps, applying 2 x 500 μL of 1:1 methyl tert-butyl ether:methanol (v / v) onto the SPE cartridge and positive pressure manifold. The combined eluates were dried under a stream of nitrogen at room temperature (Biotage SPE Dry Evaporation System) and reconstituted in 200 μL of 1:1 2-propanol:methanol (v / v) before LC / MS analysis. The same SPE batch was used for all 2,000 samples.
[0058] LC / MS analysis of polar metabolites
[0059] LC / MS analysis of 4 μL aliquots of the polar metabolite extract was performed using an Agilent 1290 Infinity II liquid chromatography (LC) system coupled to an Agilent 6545 Quadrupole-Time-of-Flight (Q-TOF) mass spectrometer equipped with a dual Agilent Jet Stream electrospray ionization source. Polar metabolites were separated using a SeQuant® ZIC®-pHILIC column (100 × 2.1 mm, 5 μm, polymer, Merck-Millipore) containing a ZIC®-pHILIC guard column (2.1 mm × 20 mm, 5 μm). The use of an in-line filter before the guard column is recommended. The temperature of the column compartment was maintained at 40 °C, and the flow rate was set at 250 μL / min. The mobile phase consisted of A: 95% water, 5% acetonitrile, 20 mM ammonium bicarbonate, 0.1% ammonium hydroxide solution (25% ammonia in water), and 2.5 μM medronic acid; and B: 95% acetonitrile, 5% water, and 2.5 μM medronic acid. In this study and previous studies, medronic acid was used in mobile phase B (mainly acetonitrile). However, this method can also be performed by adding 5 μM medronic acid to aqueous mobile phase A and no medronic acid to mobile phase B. The following linear gradient was applied: 0–1 min, 90% B; 12 min, 35% B; 12.5–14.5 min, 25% B; 15 min, 90% B, followed by a 4-min re-equilibration phase at 400 μL / min and a 2-min re-equilibration phase at 250 μL / min. Polar metabolites were detected in negative ion mode at a scan rate of 1 spectrum / second. The source parameters were as follows: gas temperature 200 °C, drying gas flow rate 10 L / min, nebulizer pressure 44 psi, sheath gas temperature 300 °C, sheath gas flow rate 12 L / min, VCap 3000 V, nozzle voltage 2000 V, fragmentor 100 V, skimmer 65 V, and Oct 1 RF Vpp 750 V. The m / z range was 50–1700. Data were acquired with sequential reference mass correction at m / z 119.0363 and 966.0007. Samples were randomized before analysis. Additionally, a quality control (QC) sample was injected every 12th sample to monitor instrument signal stability.The HILIC column was equilibrated by injection of three blank samples and four QC samples before the start of the actual sample analysis. Before the start of each batch, the mass spectrometer was calibrated, and the TIC, internal standard intensity, and mass accuracy of the QC samples were compared to the previous batch before the start of the actual sample injections. Further details on best practices for system suitability testing are available elsewhere.
[0060] LC / MS analysis of lipid metabolites
[0061] LC / MS analysis of 4 μL aliquots of lipid extract was performed using an Agilent 1290 Infinity II liquid chromatography (LC) system coupled to an Agilent 6545 Quadrupole-Time-of-Flight (Q-TOF) mass spectrometer equipped with a dual Agilent Jet Stream electrospray ionization source. Lipids were separated using an Acquity UPLC® HSS T3 column (2.1 × 150 mm, 1.8 μm) containing an Acquity UPLC® HSS T3 VanGuard precolumn (2.1 × 5 mm, 1.8 μm) at a temperature of 60°C and a flow rate of 250 μL / min. The mobile phase consisted of A: 60% acetonitrile, 40% water, 0.1% formic acid, 10 mM ammonium formate, and 2.5 μM medronic acid; and B: 90% 2-propanol, 10% acetonitrile, 0.1% formic acid, and 10 mM ammonium formate (dissolved in 1 mL of water). The following linear gradient was used: 0–2 min, 30% B; 17 min, 75% B; 20 min, 85% B; 23–26 min, 100% B; and 26 min, 30% B, followed by a 5-min re-equilibration phase. Lipids were detected in positive ion mode at a scan rate of 2 spectra / s. The source parameters were as follows: gas temperature 25 °C, drying gas flow rate 11 L / min, nebulizer pressure 35 psi, sheath gas temperature 300 °C, sheath gas flow rate 12 L / min, VCap 3000 V, nozzle voltage 500 V, fragmentor 160 V, skimmer 65 V, and Oct 1 RF Vpp 750 V. The m / z range was 50–1700. Data were acquired in positive ion mode with sequential reference mass correction at m / z 121.0509 and 922.0890. Samples were randomized before analysis. Additionally, a QC sample was injected every 12th sample to monitor instrument signal stability. The RP column was equilibrated by injection of three blank samples and three QC samples before the start of the actual sample analysis. The same system suitability test as for the analysis of polar metabolites was used for the analysis of lipid metabolites.
[0062] Preparation of pooled samples
[0063] A pooled sample from the first batch was prepared by combining 330 μL of all samples (58 samples) containing at least 500 μL. Next, 160 aliquots of 110 μL each were frozen at -80°C to include two 50 μL pooled QC samples for each batch. This pooled QC sample served as a reference sample between batches as well as for feature detection and metabolite identification.
[0064] Metabolite identification was aided by pooled QC samples and eight additional pooled samples prepared by pooling 10 μL aliquots of 10 samples from eight random batches. The eight additional pooled samples covered four different geographic regions and a wide age range. Three pooled samples contained only samples from Denmark, while the other five samples were a mix of samples from three different regions of the United States. The mean age ranged from 69.1 to 87.9 years, with a mean standard deviation of 13.3 years. Overall, these pooled samples were from 52% male and 48% female participants. MS / MS data were acquired from nine pooled samples in negative ion mode (details below). Peak detection to generate DecoID peak lists was performed using XCMS for six pooled samples and AcquireX for three pooled samples.
[0065] MS / MS data acquisition
[0066] MS / MS spectra of polar and lipid metabolites were acquired using an iterative data-dependent acquisition (iDDA) approach in MassHunter acquisition software (version 10.1.48, Agilent Technologies) on an Agilent 6545 QTOF. The same source settings were used as for MS1 data acquisition. MS / MS spectra were acquired at a scan rate of 3 spectra / s using various intensity thresholds and collision energies of 10, 20, and 40 V to increase the identification rate. iDDA can be easily implemented by setting up sequential acquisition of blank and study samples in an iterative workflow. This eliminates the need to manually create an exclusion list after each run. Details on how to implement iDDA on the Agilent system are described in the application note.
[0067] Additionally, MS / MS data for polar metabolites were acquired using an Orbitrap ID-X Tribrid mass spectrometer (ThermoScientific). A Vanquish Horizon UHPLC system was connected to the mass spectrometer via an electrospray ionization source (spray voltage 2.8 kV) in negative ion mode under the same chromatographic conditions as above. The following source conditions were used: sheath gas flow rate 50 arbitrary units (Arb), auxiliary gas flow rate 10 Arb, sweep gas flow rate 1 Arb, ion transfer tube temperature 300 °C, and vaporizer temperature 200 °C. The RF lens value was 60%. Data were acquired in data-dependent acquisition (DDA) mode using the built-in deep scan option (AcquireX) in MS1 scans, with a mass range of 67–900 m / z and a resolution of 120 K. MS / MS scans were acquired at a resolution of 15 K from three different pooled samples.
[0068] Indexed Retention Time
[0069] To adjust for sample-specific retention time drift of lipid metabolites, indexed retention time (iRT) was performed. iRT used the retention times of selected "indexed" compounds to adjust the retention times of other compounds. iRT was performed for each lipid class, and two to three lipid metabolites from each class were selected as indexed compounds (see Table 3). For indexed compound selection, lipids that were chromatographically well separated from other species and had high intensities were selected.
[0070] The first step in this process is to calculate the iRT for each lipid in a class. The calculation relies on the retention time of each lipid measured from a single sample. As shown in Equation 1 below, iRT is the retention time (RT min ) and the retention time (RT max ) to determine the measured retention time (RT j ) to the corresponding iRT (iRT j ) was converted to
[0071]
number
[0072] Next, for each subsequent sample, the iRT of the indexed compound (C) in the sample and the observed retention time for the indexed compound were used to calculate a linear regression of iRT versus retention time by minimizing the error between the observed retention time and the fit retention time according to Equation 2 below.
[0073]
number
[0074] Although other studies have used Lowess regression to convert iRT to retention time, we found that linear regression could accurately estimate retention time from iRT values, suggesting that retention time drift within lipid classes is linear.
[0075] Finally, for all non-indexed compounds in a lipid class, the adjusted retention time (RT) was calculated using the linear regression and the calculated iRT values of the non-indexed compounds according to Equation 3 below. adj ) was calculated.
[0076]
number
[0077] This process was repeated for each lipid class and correction was performed for each sample. The iRT adjustment was implemented in Python and the associated Google Colab link is available on GitHub.
[0078] Batch correction evaluation
[0079] In this study, we evaluated 14 batch normalization methods: L1 ("constant sum"), L2 ("fixed length"), median, QC normalization, ComBat, combined ComBat and QC normalization, support vector regression (SVR), linear regression, random forest, stochastic exponential normalization (PQN), WaveICA, EigenMS, Gaussian (normal) quantile transformation, and uniform quantile transformation. For L1, L2, and median normalization, we calculated the L1 norm, L2 norm, or median of each sample's metabolic profile and normalized the sample's metabolite intensities by that value. For QC normalization, we normalized the metabolite intensities of a particular study sample by the average of the metabolite intensities of its neighboring QC samples. Specifically, the metabolite intensities of a study sample were multiplied by the average of the metabolite intensities across its neighboring QC samples and divided by the average of the QC samples. For the combination of ComBat and QC normalization, each batch of samples was QC normalized individually, followed by inter-batch correction using ComBat. For SVR, linear regression, and random forest, each model was fitted to each metabolite using the batch and run-order position of the QC sample relative to the deviation of the metabolite's intensity from the mean QC intensity. The fitting procedure for these models was performed using default parameters in Scikit-learn. After fitting, the deviation of all QC and study samples was estimated, and this deviation was subtracted from the intensity of each sample. For all correction methods except QC normalization, metabolite intensities were log2-transformed before normalization. For QC normalization, metabolite intensities were log2-transformed after normalization. The Python implementation of the normalization algorithm, the associated Google Colab link, and the code for performing the comparative analysis are available on GitHub.
[0080] To evaluate the performance of each method, a metric was calculated from the coefficient of variation (CV) of both the non-normalized and normalized data. First, the CV of each metabolite for the QC and study samples was calculated separately for the non-normalized data. These values were then recalculated for the normalized data. The CV value of the study sample in the normalized data was divided by the CV value of the study sample in the non-normalized data, and the CV value of the QC sample in the normalized data was divided by the CV value of the QC sample in the non-normalized data. These two quantities were then divided by each other (the change in the CV value of the study sample divided by the change in the CV value of the QC sample) to obtain a normalized score for each metabolite. The average normalized score across all metabolites was used to score the overall performance of the method. CV values were calculated using metabolite intensities that were not log2-transformed. As a secondary performance metric to evaluate the batch correction method for polar metabolite data, the intensity CV of the internal standard spiked into each sample was calculated both between and within batches. Because these CV values were calculated across all samples, not just QC samples, this approach allows for detecting potential "overfitting" of the batch correction method.
[0081] Analysis of unknown metabolites
[0082] Of the 6,036 features detected in the pooled samples, 3,421 high-quality features were retained after manual inspection of the extracted ion chromatograms. The peak areas of these high-quality features were then extracted using Skyline-daily (version 21.2.1.485), interpolated using the half-minimum method, and normalized using the random forest algorithm. Features with CVs greater than 10% across QC samples were excluded from downstream analysis. Furthermore, to eliminate data redundancy, features with Pearson correlation coefficients greater than 0.95 with other unknown features or identified metabolites across all samples were excluded from statistical analysis. The remaining features were tested for association with field site using one-way ANOVA with Bonferroni correction.
[0083] Results and Discussion
[0084] The strategy presented in this study for analyzing large cohorts using untargeted metabolomics is based on the observation that most features in experiments do not correspond to unique, biologically relevant metabolites. Instead of assessing each feature in every sample, a small number of pooled reference samples were used to annotate features of interest. This process reduces the data burden of untargeted metabolomics, allowing informatics tools typically applied to targeted studies to efficiently and quickly profile study samples without the need to apply computationally intensive analyses (e.g., peak detection, correspondence determination, peak grouping, metabolite identification, etc.) to each sample. A schematic of the workflow is shown in Figure 3. As an example, the disclosed workflow was used to analyze a subset of approximately 2,000 human plasma samples extracted from the Longevity Family Study (LLFS), conducted by the National Institute on Aging at the National Institutes of Health (NIH), Bethesda, Maryland, USA. A description of each step of the workflow follows. For more details, see the Experimental Section and Supporting Information.
[0085] Sample preparation and data collection
[0086] Prior to acquiring LC / MS data for the LLFS sample set, 2,005 plasma samples were organized into 22 batches (approximately 92 samples per batch). Longitudinal samples from the same participant were included in the same batch. Polar and lipid metabolites were extracted from plasma samples into 96-well plates using solid-phase extraction. From the first batch of the LLFS sample set, pooled samples were prepared for use as QC samples. All samples were analyzed by LC / MS. An additional pooled sample (see Supporting Information) was prepared and analyzed by LC / MS / MS for metabolite identification.
[0087] Data processing of pooled samples and subsequent data extraction
[0088] The data collected from the pooled samples were subjected to a standard processing workflow for untargeted metabolomics. Specifically, feature detection, grouping, background and degeneracy filtering, and MS / MS-based compound identification were performed to generate peak lists of putatively identified metabolites suitable for multi-omics integration. Peak lists of identified polar and lipid metabolites are shown in Tables 1 and 2, respectively. To demonstrate that this workflow is suitable for the analysis of unknown samples at the scale typically encountered in untargeted metabolomics studies, over 3,000 unidentified features were added to the peak lists.
[0089] After creating the peak list, Skyline was used to extract peak areas from the study and QC samples for each batch. Using the Skyline command line interface, documentation for each batch could be automatically generated. The retention times of polar metabolites separated by HILIC were stable across all samples within a batch (see Figure 4A and Figure 4B). Therefore, retention time boundaries were established by inspection of the QC samples within each batch, and these boundaries were applied to all samples by importing the peak boundaries for each sample. In the current work, a Python script was used to generate the peak boundary read file; however, this is no longer necessary when using the "Synchronize Integration" feature in Skyline 21.2, which was released after the immediate data processing run.
[0090] Retention time values were generally stable across all experiments. Lipid metabolites showed more significant drift than polar metabolites, but only five samples exhibited retention time deviations greater than 0.25 min across all lipid analysis samples. Among the compounds that exhibited retention time drift, lysophosphatidylcholine and lysophosphatidylethanolamine were most pronounced in certain samples within each batch. This is likely due to matrix effects (Figure 5A-C). Therefore, the concept of iRT, initially established in proteomics and recently integrated into lipomics, was applied. In summary, from each lipid class, two to three compounds that were highly intense and easily distinguishable from neighboring peaks were selected as indexed compounds. Importing all data files into Skyline in centroid mode consistently and correctly selected the peaks for the indexed compounds. Inaccurate peak assignments were corrected by inspection of retention time replicate comparison batches and manual adjustment of peak boundaries to ensure that the apexes of all indexed peaks were within the integration boundaries. The retention times of the indexed compounds were then used in a Python script to refine and generate peak boundaries for all compounds in all samples (details are provided in the Supporting Information). Upon importing these boundaries, correct integration was manually verified before exporting the peak areas. This method proved to be efficient at correcting RT shifts exceeding 1 min. After peak area extraction in Skyline, a total of 172 identified polar metabolites, 188 identified lipid metabolites, and 3,421 unidentified features were profiled from the polarity data in the 2,001 study samples and 197 QC samples of the LLFS sample set. Notably, four study samples were excluded from downstream analysis due to abnormally low signal abundance. Extracting peak areas using the above process requires 20–30 min for each batch of samples analyzed, even for an experienced researcher. The primary time investment is in defining and curating the peak list created from the analysis of pooled samples. For LLFS, defining the peak list takes several weeks and involves extensive MS / MS data acquisition, manual removal of artifact peaks and interferences, and review of metabolite identifications.The time required to generate this peak list depends on the study and the number of non-biological and redundant signals filtered out.
[0091] Comparison of metabolite coverage between pooled and individually measured samples
[0092] Limiting data processing to pooled samples significantly reduces the computational burden of untargeted metabolomics, but it carries the risk of overlooking low-abundance compounds present in only a few study samples. To determine the number of features potentially overlooked by limiting data processing to pooled samples, a subset of 58 study samples from the LLFS was evaluated. For this analysis, nine replicate injections of the pooled sample were created by combining small aliquots from the 58 study samples. Feature detection, background subtraction, and M+ / -H ion selection were then performed on each of the 58 study samples, the 9 pooled sample, and the 4 blank sample. Peak detection was performed using Centwave7, and features with peak areas less than 10,000 in a particular sample were classified as not detected in that sample. Background subtraction was performed by removing features whose intensity in the blank sample was less than one-third of that in the study or pooled sample. M+ / -H ions were selected using the CAMERA software package.
[0093] After applying the disclosed data processing workflow, a list of all features detected in each of the 58 study samples was generated. The list, containing a total of 5,894 features, included features detected only in a single study sample. A total of 3,241 features from the list (including both identified and unknown features) were detected in at least one replicate of the pooled sample (Figure 6A). To assess the biological relevance of features not detected in the pooled sample, the m / z values of all features were searched against endogenous metabolites in the Human Metabolome Database and the Kyoto Encyclopedia of Genes and Genomes. Overall, 40.4% of features detected in the pooled sample had at least one database hit. Of the features not detected in the pooled sample, only 32.2% had a database hit, suggesting that the undetected features were likely exogenous metabolites (e.g., drugs, environmental toxins, cosmetics). Furthermore, features not detected in the pooled sample were detected in an average of less than 10% of the individual samples (Figure 6B). In contrast, features detected in the pooled sample were detected in the majority of individual samples (Figure 6C), and features not detected in the pooled sample were an order of magnitude lower in abundance (Figures 6D and 6E).
[0094] These data indicate that using pooled samples reduces the number of detected features, but undetected features are above the detection limit only in a small subset of study samples. Previously, it has been suggested that features undetected in at least 70% of samples should be excluded from untargeted metabolomics analyses, regardless of the data processing workflow. Adopting this threshold reduces the number of undetected features to 52 (less than 1% of the total number of features) when only pooled samples are used for data processing. Another major issue with tracking undetectable features in pooled samples is the difficulty of normalizing for batch effects and technical variability. Therefore, even if features detected in a small number of samples are biologically interesting, more sensitive methods must be developed to evaluate them. Currently, these features are not well suited for large-scale studies, regardless of the data processing workflow applied.
[0095] Data Post-Processing
[0096] After peak area extraction, missing values must be removed from the data to facilitate downstream processing. In the pooled sample workflow, missing values were rare (less than 0.03% of all measurements) and primarily resulted from metabolites at concentrations below the instrument's detection limit, not from random metabolite dropout during peak detection. Thus, missing values were imputed using the half-minimum method. Comparing the frequency of missing values generated by targeted extraction and conventional XCMS-based processing, we found that using XCMS, 9.2% of the total peak area in a single batch of LLFS samples was missing. In contrast, when targeted extraction was applied to the same sample set, less than 0.001% of all measurements were missing (see Figures 7A–7C). Furthermore, the technical variability of output peak areas was five-fold lower than the peak intensities of targeted extraction when compared with XCMS (Figure 7C).
[0097] Given the scale of the LLFS study, raw data had to be collected over several months. When combining data extracted from each group of 92 study samples, strong batch effects were observed for identified lipids and polar metabolites, as seen in Figures 8A and 9A, respectively. Unfortunately, samples from human subjects for large-scale studies cannot always be randomized into analytical batches due to practical constraints such as long-term sample collection and funding schedules. As a result, differences between batches may be due to technical or biological variation (see Figures 10A and 10B). For example, for the LLFS sample set, the last six batches consisted primarily of samples collected in Denmark, while the other batches consisted primarily of samples collected in the United States.
[0098] To distinguish between biological and technical variation, using identical QC samples across all batches is essential for evaluating and guiding batch correction methods. Here, 14 normalization algorithms were tested for their ability to minimize technical variation while preserving biological variation (details of this comparison are provided in the Supporting Information). Analysis revealed that the performance of many methods was dataset-dependent (see Figures 8B and 9B). However, the random forest-based batch correction algorithm performed better than the other methods evaluated for both lipid and polar metabolite data. As a secondary validation, we also compared the internal standard variation in the polar metabolite data. Random forest-based correction was again observed to reduce both intra- and inter-batch variation (Figures 11A–11D). Comparing the performance of random forest with other evaluation methods, QC, ComBat, and QC+ComBat performed similarly to random forest in reducing internal standard CVs, with ComBat+QC showing the lowest variation (Figure 11E). Overall, random forest performed well on the data, but it may be useful to test various batch correction methods before selecting an algorithm to apply to a dataset. Such evaluations can be performed by applying the code written for the analysis, available on GitHub. After correction, metabolites that were previously affected by batch no longer exhibited batch-dependent intensity drift, and technical variation within the data was reduced (Figures 8C-8E, 9C-9E). While some batch-related clustering remained in the polar metabolic profiles, no such clustering occurred when examining only the QC samples (Figures 12A-12D). These findings indicate that biological differences between batches, rather than technical variation, are driving the observed patterns. After batch correction, metabolic profiles can be subjected to downstream statistical analysis to identify interesting biological patterns.
[0099] Metabolic profiles are clustered by geographic location
[0100] While untargeted metabolomics data collection for the entire LLFS cohort is currently underway, we hoped to demonstrate that the disclosed workflow generates metabolic profiles containing biologically relevant information. Therefore, we investigated differences in polar and lipid metabolite profiles that could be attributed to the geographic origin of the samples. LLFS samples were collected at four different field sites: Boston, Massachusetts, USA; Pittsburgh, Pennsylvania, USA; New York City, New York, USA; and Odense, Denmark. Given differences in dietary habits between the United States and Denmark, we anticipated that the plasma metabolic profiles of samples collected in the United States and Denmark would differ. Indeed, analysis of the unnormalized metabolic profiles of identified metabolites in these data revealed 72 metabolites with a maximum absolute change of more than 2-fold between field sites and a p-value of less than 0.05 (one-way ANOVA). However, given that field site samples were not uniformly distributed across sample batches, technical variation may have artificially separated metabolic profiles. Indeed, after performing random forest batch correction, only 45 metabolites met the same statistical threshold, demonstrating the importance of batch correction in interpreting metabolomics data. These 45 metabolites resulted in significant clustering between the US and Danish samples (Figure 13A). Several of the differentiating lipid metabolites contain multiple unsaturated bonds (e.g., CE22:5, DG36:4, LPC20:5, PC37:5, PC38:6, PC40:7, TG56:7, TG58:7, TG58:8, TG60:11, and TG60:12). This likely reflects differences in dietary omega-3 and omega-6 polyunsaturated fatty acids (e.g., DHA, EPA, and linoleic acid) content (Figure 13B). Scandinavian countries have been shown to have higher concentrations of circulating omega-3 fatty acids compared to the US and other countries with Western dietary habits, likely due to their higher per capita intake of seafood. Among polar metabolites, the most significant difference was inosine (Figure 13C). Inosine is a dietary metabolite known to be present in high concentrations in milk.The results suggest differences in dairy intake between individuals in the United States and Denmark, which have also been noted in previous studies.
[0101] While identified metabolites were the primary focus of the LLFS analysis, and metabolomics data were linked to corresponding gene and protein measurements, differences in unknowns between sample groups were also investigated. Filtering over 3,000 unknowns from the polar metabolite extracts profiled in this study (see Supporting Information) and performing statistical analysis revealed 29 unique unknowns associated with field site location (|FC|>2, p<0.05, one-way ANOVA), as shown in Figures 14A-14D. While differences between field sites were significant, participants from Denmark had a younger mean age than participants from the United States (p<0.0001), which may have contributed to the observed metabolic profiles (Figure 13D). Notably, however, principal component analysis of these samples color-coded by age revealed no age-dependent patterns within the United States or Denmark samples (Figure 15). This suggests that differences between field sites are not solely attributable to age. An additional potential confounding factor is that the time between blood collection and centrifugation varied across samples and field sites. Previous studies have shown that this delay can cause changes in metabolite levels. Once LLFS data collection is complete, we plan to conduct further statistical analyses of this dataset, taking into account covariates and genetic relationships among LLFS participants.
[0102] conclusion
[0103] A trend in omics science is the evaluation of large sample sizes approaching tens to hundreds of thousands of specimens. Larger sample sets increase statistical power, enabling studies to distinguish subpopulations within groups with unique therapeutic effects, a vision of precision medicine. However, a challenge with using extremely large sample cohorts is the burden of collecting and processing large amounts of data. The application of untargeted metabolomics to large sample cohorts has been limited because standard software programs are not compatible with population-based studies. This study describes a workflow for performing untargeted metabolomics on over 12,000 human plasma samples collected from the LLFS cohort.
[0104] To reduce the computational burden of performing global processing of all HRMS data files within a cohort, the disclosed method focuses on processing data from a small number of pooled samples created by combining aliquots from each individual sample in the study. By initially limiting data processing to the pooled sample, the disclosed method leverages standard informatics tools for untargeted metabolomics optimized for small-scale samples. This not only facilitates feature detection but also achieves significant data reduction through the removal of adducts, background signals, and other data-degrading factors. While analysis of pooled samples requires a significant time investment, it can be completed in parallel with data collection for the study samples. Furthermore, because peak areas for metabolites detected in the pooled sample can be extracted from the study samples at the time of data generation, the speed-limiting factor in the disclosed workflow is data collection rather than data processing. This contrasts with traditional processing workflows, which require all data to be collected before analysis and curation.
[0105] When preparing reference samples, care should be taken to ensure that pooled samples cover different sample groups and represent the biological diversity present in the study, such as age and gender. Although pooled samples are intended to capture the entire range of unique compounds in the entire study cohort, they may in fact miss unique compounds from individual samples. Data analysis showed that unique compounds are most frequently overlooked when present at low concentrations in a small number of participant samples. This is because they are diluted below the detection limit in the pooled sample. Analytical results indicated that signals missed in the analysis of pooled samples are likely to be derived from rare exogenous compounds (e.g., unique hygiene products or chemicals derived from specific environments). While rare exogenous compounds may certainly be biologically interesting, their analysis requires the development of novel peak detection, alignment, and annotation algorithms that can be scaled up to thousands of samples without compromising functionality or accuracy. It should be noted that even when using traditional informatics workflows for untargeted metabolomics, features detected in only a low percentage of samples are typically discarded due to the difficulty of correcting for technical drift.
[0106] When processing untargeted metabolomics data in a traditional workflow, retention times are adjusted using corresponding algorithms such as Obiwarp. Here, sample-specific retention time drift was corrected by applying iRT. This approach requires some manual intervention by the user to ensure accurate peak detection of indexed compounds and adjust for batch-to-batch variations in retention time. In this disclosure, we demonstrated that the iRT approach corrected for retention time drift in lipid metabolites, but this was not necessary for polar metabolites, because no intra-batch retention time drift was observed in these data. If retention time drift is observed for polar metabolites, the iRT approach can be easily extended by determining groups of compounds with correlated retention time drift across samples and using them as indexes.
[0107] In large-scale, longitudinal studies such as LLFS, completely randomizing the analysis order of samples by LC / MS may be impractical. Therefore, to accurately measure biological differences, it is necessary to correct for technical variation between batches of samples. Here, we evaluated 14 methods for batch correction and found that a method that accounts for batch effects using random forest normalization to estimate the drift of each reference chemical was the most effective. While this study used a limited number of internal standards, it should be noted that incorporating additional compounds covering a broader range of chemical classes would be beneficial to assess the potential for metabolite degradation and evaluate the overall quality of batch correction results.
[0108] In summary, the present disclosure enables the high-throughput application of untargeted metabolomics on a population scale, such as that observed in LLFS. As an added advantage, the disclosed workflow collects the separation and mass spectrometry characteristics of every individual sample within a sample set. Therefore, as new computational resources become available, the disclosed method can be used with more advanced technologies.
[0109] [Table 1-1]
[0110] [Table 1-2]
[0111] [Table 1-3]
[0112] [Table 1-4]
[0113] [Table 1-5]
[0114] Table 2-1
[0115] Table 2-2
[0116] Table 2-3
[0117] Table 2-4
[0118] Table 2-5
[0119] Table 2-6
[0120] Table 2-7
[0121] Table 2-8
[0122] Table 2-9
[0123] Table 3
Claims
1. 1. A method for non-targeted identification of unique biological metabolites in individual samples within a sample set, each individual sample comprising chemical constituents, the method comprising: (a) analyzing reference features obtained from a reference sample to identify irrelevant reference features; (b) filtering the reference features by removing the irrelevant reference features to generate a set of relevant reference features that characterize the unique biological metabolites; (c) applying the unrelated reference features and / or the related reference features to sample features obtained from the individual samples to identify the unique biological metabolite composition in the individual samples.
2. 10. The method of claim 1, The method includes determining the unique biological metabolites in all of the individual samples in the sample set by performing the analyzing, filtering, and applying steps on each sample in the sample set.
3. 3. The method of claim 2, after performing the analyzing, filtering, and applying steps on a plurality of the individual samples in the sample set, re-analyzing the reference sample to obtain updated reference features and replacing the reference features with the updated reference features; The method, wherein the number of samples in the plurality is less than the total number of the individual samples in the sample set.
4. 4. The method of claim 3, The method, wherein said plurality of samples comprises at least 2 samples, optionally at least 3 samples, optionally at least 5 samples, optionally at least 10 samples, optionally at least 20 samples, optionally at least 50 samples, optionally at least 100 samples, or optionally at least 150 samples.
5. The method according to any one of claims 1 to 4, The method, wherein the reference sample comprises an aliquot from each of a plurality of the individual samples.
6. The method according to any one of claims 1 to 5, The method, wherein the reference feature is obtained by a technique comprising subjecting the reference sample to a separation technique or mass spectrometry.
7. The method according to any one of claims 1 to 6, The method, wherein the irrelevant reference features characterize artifacts, background chemical constituents, and non-specific biological metabolites.
8. 8. The method of claim 7, The method, wherein the background chemical constituents include contaminants.
9. 9. The method of claim 7 or 8, The method, wherein the non-specific biological metabolites include adducts and / or fragment ions.
10. The method according to any one of claims 6 to 9, The method wherein said separation technique comprises chromatography.
11. 11. The method of claim 10, The method, wherein said chromatography comprises gas chromatography or liquid chromatography.
12. The method according to any one of claims 6 to 11, The method, wherein the reference features include reference separation data and reference mass data generated from the reference sample.
13. A method according to any one of claims 1 to 12, The applying step includes: (a) generating a set of relevant sample features by identifying sample features among the sample features that correspond to the irrelevant reference features and removing sample features that correspond to the irrelevant reference features; (b) using the relevant sample features to determine the unique biological metabolites in the individual samples.
14. A method according to any one of claims 1 to 12, The applying step includes: (a) identifying sample features among the sample features that correspond to the relevant reference features to generate the set of relevant sample features; (b) using the relevant sample features to determine the unique biological metabolites in the individual samples.
15. 15. The method according to any one of claims 1 to 14, determining said chemical constituents in all said individual samples in said sample set by performing said applying step for each sample in said sample set.
16. 16. The method according to any one of claims 1 to 15, identifying chemical components of the individual samples by comparing the relevant sample features to a library of information containing features that characterize chemical components.
17. 1. A method for non-targeted identification of chemical components in individual samples within a sample set, comprising: (a) acquiring reference features comprising separation data and mass data generated from a reference sample; (b) filtering the reference feature by removing separation data and mass data associated with non-unique chemical components to generate a filtered reference feature comprising the filtered separation data and the filtered mass data; (c) using the filtered reference features to characterize the presence and / or amount of unique chemical components in each sample in the sample set.
18. 18. The method of claim 17, The method, wherein the reference sample comprises an aliquot from each of a plurality of the individual samples.
19. 1. A method for non-targeted identification of chemical components in individual samples within a sample set, comprising: (a) generating a reference signature comprising separation data and mass data from a reference sample using a separation technique and mass spectrometry; (b) collecting and storing said reference features; (c) analyzing the stored reference features to identify unrelated reference features comprising separation and mass data that characterize non-unique chemical components; (d) removing the irrelevant reference features comprising separation data and mass data characterizing the non-unique chemical components to generate a filtered data set comprising filtered reference features comprising separation data and mass data characterizing unique chemical components; (e) using the filtered reference features to characterize the presence and / or amount of unique chemical components in each sample in the sample set.
20. 20. The method of claim 19, The method, wherein the reference sample comprises an aliquot from each of a plurality of the individual samples.
21. 21. The method of claim 19 or 20, The step of using the filtered reference features comprises: (a) acquiring sample characteristics from the individual samples; (b) creating a list of relevant sample features that characterize unique chemical components by identifying sample features that correspond to the filtered reference features, thereby determining the chemical components in the sample.
22. A system for carrying out the method according to any one of claims 1 to 21, comprising: (a) a separation device for separating chemical components and generating separation-related signatures that characterize the chemical components; (b) a mass spectrometer for performing mass spectrometry on a portion of the separated chemical components to generate mass spectrometry-related signatures that characterize the chemical components; (c) a first module that receives, collects, and / or stores the separation-related characteristics and / or the mass spectrometry-related characteristics; (d) a user interface connected to the first module, the user interface providing the separation-related features and / or the mass spectrometry-related features in a human-usable form.
23. 23. The system of claim 22, a library of features characterizing the chemical components produced using the separation device and the mass spectrometer; The system wherein the library of features characterizing the chemical components includes separation features and mass spectrometry features characterizing the identified chemical components.
24. 24. A system according to claim 22 or 23, comprising: The system wherein the separation device comprises a device for performing chromatography.
25. 25. The system of claim 24, The system, wherein the chromatography comprises liquid chromatography.
26. 26. The system of claim 25, The liquid chromatography system includes HPLC (High Performance Liquid Chromatography) and UPLC (Ultra High Performance Liquid Chromatography).
27. 24. A system according to claim 22 or 23, comprising: The system, wherein the separation device comprises an electrophoresis device.
28. A system according to any one of claims 22 to 24, The separation device is connected to the mass spectrometer.
29. A system according to any one of claims 22 to 28, The separation-related characteristics include peak retention time, peak intensity, and / or peak width.
30. A system according to any one of claims 22 to 29, The mass spectrometry relevant feature comprises a mass or m / z value.