Harnessing the law of large numbers to overcome drift
Patent Information
- Application Number
- CA3320502
- Authority / Receiving Office
- CA · CA
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-05
- Filing Date
- 2025-03-04
- Publication Date
- 2025-09-11
AI Technical Summary
Existing mass spectrometry techniques face challenges in correcting retention time drift, particularly for metabolites with difficult isomer profiles, leading to blurring or swapping of isomers and inefficient signal detection due to low signal-to-noise ratios, and current computational methods fail to correct drift individually for each unique compound.
A method involving a 2D grid system is applied to map and correct retention time drift by aligning data points of extracted ion chromatograms, iteratively defining templates to correct linear and non-linear drift, and identifying EIC files with retention time profiles that exhibit non-linear drift.
This approach enhances signal amplification and detection by accurately aligning retention times, improving signal-to-noise ratios and ensuring reliable compound identification by correcting drift independently for each m/z value, thereby enhancing the accuracy and reliability of compound detection in complex samples.
Abstract
Description
HARNESSING THE LAW OF LARGE NUMBERS TO OVERCOME DRIFTGOVERNMENTAL RIGHTS
[0001] This invention was made with government support under DE-SC0023160 awarded by the Department of Energy. The government has certain rights in the invention.CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims priority from Provisional Application number 63 / 561 ,536, filed March 5, 2024, the entire contents of which are hereby incorporated by reference.FIELD OF THE INVENTION
[0003] The present disclosure relates generally to compound detection. More specifically, the present disclosure relates to amplification and detection of compound signals.BACKGROUND OF THE INVENTION
[0004] Mass spectrometry techniques such as liquid chromatography-mass spectrometry (LC-MS) are chemical techniques that identify different compounds in a sample as unique mass features. In LC-MS, a liquid chromatography system may separate the different compounds by structural properties, while a mass spectrometer subsequently determines the mass and intensity of the ions that elute from the chromatography column. Modern high-resolution mass spectrometry can now detect and quantify ions with high mass precision (< 5 ppm mass error) but may also result in significant amounts of noise.
[0005] Detecting valid compound peaks within mass spectrometry data may present a number of challenges when the compound may only be present at low levels relative to noise. For example, samples from complex systems may include large numbers of different compounds, some of which may only be present in relatively low quantities. Atypical mass spectrometry file may contain as many as millions of data points, while as few as several hundred to thousands may correspond to true compound signals that are interspersed in vast amounts of noise.
[0006] One of the great challenges in harnessing mass spectrometry techniques such as liquid chromatography-mass spectrometry (LC-MS) for untargeted metabolomics is having to correct for retention time drift, especially regarding metabolites with difficult to resolve isomer profiles. For example, the retention time instrumentation can be off when a batch within the course of experiments / runs is processed. This results in drift of the intensity signal value in the retention time domain (e.g., a shift in signal to an earlier or later time).
[0007] The difficulty in correcting drift arises in part because the outline of each isomer is often faint and fluctuates across samples - i.e. low signal to noise. The issue of low signal to noise can be addressed by pooling samples together in order to create a merged extracted ion chromatogram from which boundaries can be called for each isomer of a mass. Because signal is non-random, while noise is random, pooling samples together increases signal to noise. However, this has the consequence - when there is drift - of blurring isomers together. Alternatively, when there is no drift, isomers may be swapped with one another.
[0008] Existing strategies generally analyze compound signals and the retention time drifts computationally by aligning all the signals for various metabolites at once, or in large chunks that improperly treat metabolites with disparate characteristics the same. In other words, current techniques attempt to match all compound signals, which may not behave in predictable fashion, simultaneously. However, this approach introduces warping effects because drift isn’t being corrected on an individual basis for each unique compound. Especially problematic is when a metabolite’s signal is close to another metabolite, with each one experiencing its own unique chromatography drift. Such methods end up discarding valid compound signals, as well as being computationally expensive and slow.
[0009] For instance, one existing strategy to combat drift takes the retention time profiles for all m / z values in query samples and align them to their correspondingfeatures in a reference sample using a warping function. However, this strategy to determine a global fit is fraught with difficulties as often, there is no effective global solution for drift. For example, the reference may not even have all the metabolites present in the query.
[0010] Thus, there is a need for improved systems and methods correcting drift in retention times of metabolite signals.SUMMARY OF THE INVENTION
[0011] One embodiment of the instant disclosure encompasses a method for overcoming drift in retention times of compound signals in mass spectrometry (MS). The method comprises the steps of (a) mapping data point values of extracted ion chromatograms (EICs) of a portion of each of a plurality of MS chromatograms onto a two-dimensional (2D) grid system comprising grid boxes of a plurality of scans comprising retention time increments and a plurality of signal intensity strata, wherein the portions of the plurality of chromatograms bear the same m / z value; (b) selecting a first reference scan as a scan comprising the highest intensity data points of the highest number of EICs among the plurality of EICs; (c) defining a first data point template as all grid boxes in the grid comprising at least one data point from at least one EIC associated with the reference scan identified in (b); (d) for each remaining EIC of the plurality of EICs, generating a query EIC by shifting the data points of the query EIC by a time increment sufficient to overlap the highest data point of the query EIC with the highest data point of the first template, wherein a query EIC exhibits linear drift in relation to the first template if coordinates of all the data points of the EIC lie in the first template; (e) correcting linear drift by adding coordinates of all shifted data points of each query EIC to the first data point template to thereby generate a corrected template; (f) iteratively repeating steps (b) to (e) using remaining EICs that were not added to the first data point template to define additional data point templates and correct the linear drift of the EICs based on the identified additional templates; and (g) identifying, among the EIC files identified in steps (a) to (f), EIC files comprising retention time profiles thatexhibit non-linear drift in retention time by: (i) merging all data point coordinates of all the EIC files in the identified templates or corrected templates to generate a master coordinates file; (ii) generating a master coordinate pairs file comprising all pairs of coordinates in the master coordinates file; (iii) generating a child coordinate pairs file comprising all pairs of coordinates within each of the individual EIC files; and (iv) identifying mutually exclusive pairs of coordinates, wherein mutually exclusive pairs of coordinates are pairs of coordinates that occur in the master coordinate pairs file but not in the child coordinate pairs file, wherein coordinates of data points in mutually exclusive pairs of coordinates identify EICs that comprise retention time profiles that exhibit nonlinear drift. Steps (b) to (e) can be iteratively repeated until none of the remaining EICs are added to any of the identified templates.
[0012] In some embodiments, the compound is metabolites and isomers of metabolites. The chromatography-MS can be liquid chromatography-MS (LC-MS).
[0013] In some embodiments, defining a data point template further comprises identifying one or more EICs that introduce non-linear blur among the identified EICs of a template, and excluding the EICs that introduce non-linear blur before the data values of the identified EICs are merged to construct the data point template. Each stratum can comprise an equal number of data points and the plurality of scans comprise equal retention time increments. In some embodiments, stringency is adjusted by changing the width of the scan, the number of points in each stratum, by determining how many points match or don’t match the template.
[0014] In some embodiments, identifying EIC files comprising retention time profiles that exhibit non-linear drift comprises: (a) constructing a template box pairs file and constructing a EIC box pairs file for each EIC file determined to match the template; (b) constructing a combined EIC box pairs file by combining EIC box pairs files constructed for each EIC file determined to match the template; (c) identifying mutually exclusive pairs of grid boxes, wherein mutually exclusive pairs of grid boxes are pairs of grid boxes that are present in the template box pairs file but are absent in the EIC box pairs file or combined EIC box pairs file; and (d) identifying EICs that comprise retention time profiles that exhibit non-linear drift, wherein a data point in a grid box of a box pairdetermined to be mutually exclusive identifies an EIC that comprise retention time profiles that exhibit non-linear drift. Each pair of grid boxes can be represented by a unique identifier to facilitate the comparative analysis.
[0015] In other embodiments, identifying EIC files comprising retention time profiles that exhibit non-linear drift comprises: (a) merging all coordinates of two or more data point templates or corrected templates to generate a master coordinates file; (b) generating a master coordinate pairs file comprising all possible pairs of coordinates in the master coordinates file; (c) identifying all possible pairs of coordinates within each of the individual EIC files represented in the two or more data point templates or corrected data point templates; (d) generating a child coordinate pairs file comprising all possible pairs of coordinates within each of the individual EIC files represented in the two or more data point templates or corrected data point templates; (e) identifying mutually exclusive pairs of data point coordinates, wherein mutually exclusive pairs of data point coordinates are pairs of coordinates that occur in the master template but not in each EIC file comprising at least one data point in each of the pairs of coordinates; (f) identifying mutually exclusive pairs of grid boxes, wherein mutually exclusive pairs of coordinates identify mutually exclusive pairs of grid boxes; and (g) identifying EICs that comprise retention time profiles that exhibit non-linear drift, wherein data points in mutually exclusive pairs of coordinates or mutually exclusive pairs of grid boxes identify EICs that comprise retention time profiles that exhibit non-linear drift.
[0016] In some embodiments, identifying mutually exclusive pairs of coordinates comprises: (a) generating scan gaps in the master coordinate file and generating scan gaps in each of the EICs comprising data points in the master coordinate file, wherein generating scan gaps comprises concatenating data points in each file by connecting a data point with the highest intensity to a data point comprising the next highest intensity and wherein each connection is defined as a scan gap; and (b) identifying blur scan gaps, wherein each blur scan gap is a scan gap that links two grid boxes in the merged data points coordinates file that are not linked in any of the individual EICs, and wherein coordinates of data points in the mutually exclusive pairs of grid boxes identify EICs that comprise retention time profiles that exhibit nonlinear drift.
[0017] In some embodiments, defining a data point template further comprises identifying one or more EICs that comprise retention time profiles that introduce nonlinear blur among the identified EICs of the template and excluding the EICs that introduce non-linear blur before the data values of the identified EICs are merged to construct the data point template.
[0018] Another embodiment of the instant disclosure encompasses a computing system for overcoming drift in retention times of metabolite signals in mass spectrometry (MS). The computing system comprises: (a) a communication interface that receives EICs of portions of each of a plurality of MS chromatograms, wherein the portions of the plurality of chromatograms bear the same m / z value; and (b) a processor that executes instructions stored in memory. The processor executes the instructions to: (i) map data point values of the EICs onto a 2D grid system comprising grid boxes of a plurality of scans comprising retention time increments and a plurality of signal intensity strata; (ii) select a first reference scan as a scan comprising the highest intensity data points of the highest number of EICs among the plurality of EICs; (iii) define a first data point template as all coordinates in the grid comprising at least one data point from at least one EIC identified in step (b)(ii); (iv) for each remaining EIC of the plurality of EICs, generate a query EIC by shifting the data points of the query EIC by a time increment sufficient to overlap the highest data point of the query EIC with the highest data point of the first template, wherein a query EIC exhibits linear drift in relation to the first template if coordinates of all the data points of the EIC lie in the first template; (v) correct linear drift by adding coordinates of all shifted data points of each query EIC to the first data point template to thereby generate a corrected template; (vi) iteratively repeat steps (b)(ii) to (b)(v) using EIC files that were not added to the first data point template to define additional data point templates and correct the EICs for linear drift based on the identified additional templates; and (vii) identify, among the EIC files identified in steps (b)(i) to (b)(vii), EIC files comprising retention time profiles that exhibit nonlinear drift in retention time by: (1 ) merging all data point coordinates of all the EIC files in the identified templates or corrected templates to generate a master coordinates file; (2) generating a master coordinate pairs file comprising all pairs of coordinates in themaster coordinates file; (3) generating a child coordinate pairs file comprising all pairs of coordinates within each of the individual EIC files; and (4) identifying mutually exclusive pairs of coordinates, wherein mutually exclusive pairs of coordinates are pairs of coordinates that occur in the master coordinate pairs file but not in the child coordinate pairs file, wherein coordinates of data points in mutually exclusive pairs of coordinates identify EICs that comprise retention time profiles that exhibit nonlinear drift. In some embodiments, each stratum comprises an equal number of data points and the plurality of scans comprise equal retention time increments. The processor can further execute the instructions to iteratively repeat steps (b) to (e) until none of the remaining EICs are added to any of the identified templates. In some embodiments, stringency are adjusted by changing the width of the scan, the number of points in each stratum, by determining how many points match or don’t match the template.
[0019] In some embodiments, the processor further executes the instructions to define a data point template by identifying one or more EICs that introduce non-linear blur among the identified EICs of a template, and excluding the EICs that introduce nonlinear blur before the data values of the identified EICs are merged to construct the data point template.
[0020] In some embodiments, the processor further executes the instructions to identify, among the EIC files identified in steps (b)(i) to (b)(vii), EIC files comprising retention time profiles that exhibit nonlinear drift in retention time by further: (a) constructing a template box pairs file and constructing a EIC box pairs file for each EIC file determined to match the template; (b) constructing a combined EIC box pairs file by combining EIC box pairs files constructed for each EIC file determined to match the template; (c) identifying mutually exclusive pairs of grid boxes, wherein mutually exclusive pairs of grid boxes are pairs of grid boxes that are present in the template box pairs file but are absent in the EIC box pairs file or combined EIC box pairs file; and (d) identifying EICs that comprise retention time profiles that exhibit non-linear drift, wherein a data point in a grid box of a box pair determined to be mutually exclusive identifies an EIC that comprise retention time profiles that exhibit non-linear drift. In some embodiments, theprocessor further executes the instructions to assign a unique identifier to each pair of grid boxes to facilitate the comparative analysis.
[0021] The processor can further execute the instructions to identify, among the EIC files identified in steps (b)(i) to (b)(vii), EIC files comprising retention time profiles that exhibit nonlinear drift in retention time by further: (a) merging all coordinates of two or more data point templates or corrected templates to generate a master coordinates file;
[0022] generating a master coordinate pairs file comprising all possible pairs of coordinates in the master coordinates file; (b) identifying all possible pairs of coordinates within each of the individual EIC files represented in the two or more data point templates or corrected data point templates; (c) generating a child coordinate pairs file comprising all possible pairs of coordinates within each of the individual EIC files represented in the two or more data point templates or corrected data point templates; (d) identifying mutually exclusive pairs of data point coordinates, wherein mutually exclusive pairs of data point coordinates are pairs of coordinates that occur in the master template but not in each EIC file comprising at least one data point in each of the pairs of coordinates; (e) identifying mutually exclusive pairs of grid boxes, wherein mutually exclusive pairs of coordinates identify mutually exclusive pairs of grid boxes; and (f) identifying EICs that comprise retention time profiles that exhibit non-linear drift, wherein data points in mutually exclusive pairs of coordinates or mutually exclusive pairs of grid boxes identify EICs that comprise retention time profiles that exhibit nonlinear drift.
[0023] In some embodiments, the processor further executes the instructions to identify, among the EIC files identified in steps (b)(i) to (b)(vii), EIC files comprising retention time profiles that exhibit nonlinear drift in retention time by further: (a) generating scan gaps in the master coordinate file and generating scan gaps in each of the EICs comprising data points in the master coordinate file, wherein generating scan gaps comprises concatenating data points in each file by connecting a data point with the highest intensity to a data point comprising the next highest intensity and wherein each connection is defined as a scan gap; and (b) identifying blur scan gaps, wherein each blur scan gap is a scan gap that links two grid boxes in the merged data pointscoordinates file that are not linked in any of the individual EICs, and wherein coordinates of data points in the mutually exclusive pairs of grid boxes identify EICs that comprise retention time profiles that exhibit nonlinear drift.
[0024] In some embodiments, the processor further executes the instructions to define a data point template by further identifying one or more EICs that comprise retention time profiles that introduce non-linear blur among the identified EICs of the template and excluding the EICs that introduce non-linear blur before the data values of the identified EICs are merged to construct the data point template.
[0025] Another embodiment of the instant disclosure encompasses a non-transitory, computer-readable storage medium, having embodied thereon a program executable by a processor to perform a method for identifying metabolite derivatives in chromatography-mass spectrometry (MS) data, the method comprising: (a) receiving or having received an MS data file from a chromatography-MS instrument, wherein the data file comprises data points obtained from a sample analyzed by the chromatography-mass spectrometry instrument, and wherein each data point in the file comprises a retention time value and a mass-to-charge (m / z) value; (b) mapping data point values of the EICs onto a 2D grid system comprising grid boxes of a plurality of scans comprising retention time increments and a plurality of signal intensity strata;
[0026] selecting a first reference scan as a scan comprising the highest intensity data points of the highest number of EICs among the plurality of EICs; (c) defining a first data point template as all grid boxes in the grid comprising at least one data point from at least one EIC associated with the reference scan identified in (b); (d) for each remaining EIC of the plurality of EICs, generating a query EIC by shifting the data points of the query EIC by a time increment sufficient to overlap the highest data point of the query EIC with the highest data point of the first template, wherein a query EIC exhibits linear drift in relation to the first template if coordinates of all the data points of the EIC lie in the first template; (e) correcting linear drift by adding coordinates of all shifted data points of each query EIC to the first data point template to thereby generate a corrected template; (f) iteratively repeating steps (b) to (e) using remaining EICs that were notadded to the first data point template to define additional data point templates and correct the linear drift of the EICs based on the identified additional templates; and
[0027] identifying, among the EIC files identified in steps (a) to (f), EIC files comprising retention time profiles that exhibit non-linear drift in retention time by: (i) merging all data point coordinates of all the EIC files in the identified templates or corrected templates to generate a master coordinates file; (ii) generating a master coordinate pairs file comprising all pairs of coordinates in the master coordinates file; (iii) generating a child coordinate pairs file comprising all pairs of coordinates within each of the individual EIC files; and (iv) identifying mutually exclusive pairs of coordinates, wherein mutually exclusive pairs of coordinates are pairs of coordinates that occur in the master coordinate pairs file but not in the child coordinate pairs file, wherein coordinates of data points in mutually exclusive pairs of coordinates identify EICs that comprise retention time profiles that exhibit nonlinear drift. Each stratum can comprise an equal number of data points and the plurality of scans comprise equal retention time increments. In some embodiments, the method further comprises iteratively repeating steps (b) to (e) until none of the remaining EICs are added to any of the identified templates. Stringency can be adjusted by changing the width of the scan, the number of points in each stratum, by determining how many points match or don’t match the template.
[0028] In some embodiments, the method further comprises defining a data point template by identifying one or more EICs that introduce non-linear blur among the identified EICs of a template, and excluding the EICs that introduce non-linear blur before the data values of the identified EICs are merged to construct the data point template.
[0029] In some embodiments, the method further comprises identifying, among the EIC files identified in steps (c) to (h), EIC files comprising retention time profiles that exhibit nonlinear drift in retention time by further: (a) constructing a template box pairs file and constructing a EIC box pairs file for each EIC file determined to match the template; (b) constructing a combined EIC box pairs file by combining EIC box pairs files constructed for each EIC file determined to match the template; (c) identifying mutually exclusive pairs of grid boxes, wherein mutually exclusive pairs of grid boxes are pairs of gridboxes that are present in the template box pairs file but are absent in the EIC box pairs file or combined EIC box pairs file; and (d) identifying EICs that comprise retention time profiles that exhibit non-linear drift, wherein a data point in a grid box of a box pair determined to be mutually exclusive identifies an EIC that comprise retention time profiles that exhibit non-linear drift. In some embodiments, the method further comprises assigning a unique identifier to each pair of grid boxes to facilitate the comparative analysis.
[0030] In some embodiments, the method comprises identifying, among the EIC files identified in steps (c) to (h), EIC files comprising retention time profiles that exhibit nonlinear drift in retention time by further: (a) merging all coordinates of two or more data point templates or corrected templates to generate a master coordinates file; (b) generating a master coordinate pairs file comprising all possible pairs of coordinates in the master coordinates file; (c) identifying all possible pairs of coordinates within each of the individual EIC files represented in the two or more data point templates or corrected data point templates; (d) generating a child coordinate pairs file comprising all possible pairs of coordinates within each of the individual EIC files represented in the two or more data point templates or corrected data point templates; (e) identifying mutually exclusive pairs of data point coordinates, wherein mutually exclusive pairs of data point coordinates are pairs of coordinates that occur in the master template but not in each EIC file comprising at least one data point in each of the pairs of coordinates; (f) identifying mutually exclusive pairs of grid boxes, wherein mutually exclusive pairs of coordinates identify mutually exclusive pairs of grid boxes; and (f) identifying EICs that comprise retention time profiles that exhibit non-linear drift, wherein data points in mutually exclusive pairs of coordinates or mutually exclusive pairs of grid boxes identify EICs that comprise retention time profiles that exhibit non-linear drift.
[0031] In some embodiments, the method comprises identifying, among the EIC files identified in steps (c) to (h), EIC files comprising retention time profiles that exhibit nonlinear drift in retention time by further: (a) generating scan gaps in the master coordinate file and generating scan gaps in each of the EICs comprising data points in the master coordinate file, wherein generating scan gaps comprises concatenating datapoints in each file by connecting a data point with the highest intensity to a data point comprising the next highest intensity and wherein each connection is defined as a scan gap; and (b) identifying blur scan gaps, wherein each blur scan gap is a scan gap that links two grid boxes in the merged data points coordinates file that are not linked in any of the individual EICs, and wherein coordinates of data points in the mutually exclusive pairs of grid boxes identify EICs that comprise retention time profiles that exhibit nonlinear drift.
[0032] In some embodiments, the method further comprises defining a data point template by further identifying one or more EICs that comprise retention time profiles that introduce non-linear blur among the identified EICs of the template and excluding the EICs that introduce non-linear blur before the data values of the identified EICs are merged to construct the data point template.BRIEF DESCRIPTION OF THE FIGURES
[0033] The following drawings form part of the present specification and are included to further demonstrate certain embodiments of the present disclosure. Certain embodiments can be better understood by reference to one or more of these drawings in combination with the detailed description of specific embodiments presented herein. The patent or application file contains at least one drawing executed in color. Copies of this patent or patent application publication with color drawing(s) will be provided by the Office upon request and payment of the necessary fee.
[0034] FIG. 1 illustrates an exemplary mass spectrometry dataset from a UPLC-MS run displaying three axes representing RT, m / z ratio, and intensity. Inset: An enlarged portion of the same human urine data file highlighting the sensitivity and detail in the small molecule features collected.
[0035] FIG. 2 depicts a hypothetical retention time profile from an EIC of a segment of a MS chromatogram for a specific m / z value, in accordance with example embodiments.
[0036] FIG. 3 depicts pooled retention time profile of EIC files of a particular m / z value, derived from real data, in accordance with example embodiments.
[0037] FIG. 4 is a diagrammatical representation of data points in a merged EIC, sowing multiple waveforms and drift, in accordance with example embodiments.
[0038] FIG. 5 depicts an example of merged retention time profiles from a specific m / z value mapped onto a 2D grid of scans and strata, in accordance with example embodiments. Strata are depicted by various colors.
[0039] FIG. 6 graphically depicts an example showing waveforms that can be represented in a data point template, in accordance with example embodiments.
[0040] FIG. 7 graphically depicts an example showing waveforms that can be represented in a data point template before excluding an EIC that introduces non-linear blur, in accordance with example embodiments.
[0041] FIG. 8 graphically depicts an example showing data points of a data point template (green) that corresponds to a waveform representing EICs obtained from the analysis of samples comprising a compound (X) and one isomer (Y) and data points of an EIC (grey dots) obtained from the analysis of a sample also comprising a compound and one isomer.
[0042] FIG. 9A graphically depicts an example showing data points of two data point templates (green and yellow) that correspond to two waveforms (S1 and S2) representing EICs obtained from the analysis of samples comprising a compound (X) and one isomer (Y) and data points of two EICs (grey and red dots) obtained from the analysis of a sample also comprising the two waveforms S1 and S2, but exhibiting nonlinear drift resulting in peaks X’ and Y’ for the two isomers.
[0043] FIG. 9B graphically depicts the data points of the two data point templates (green and yellow) of FIG. 9A when merged.
[0044] FIG. 9C graphically depicts the data points of two EICs (grey and red dots) of FIG. 9A obtained from the analysis of a sample also comprising a compound and one isomer.
[0045] FIG. 10A graphically depicts two retention time profiles comprising the same waveform variant but are linearly drifted by a retention time value, in accordance with example embodiments.
[0046] FIG. 10B graphically depicts the two retention time profiles of FIG. 10A after correcting the data point of one of the profiles for linear drift, in accordance with example embodiments.
[0047] FIG. 11A highlights the data points of a query EIC, represented by red dots, where the data points of the query EIC do not align with the template's black dots, indicating the presence of drift in the retention time profile relative to the template, in accordance with example embodiments.
[0048] FIG. 11B highlights the data points of a query EIC, represented by red dots, are shifted by a first time increment, aligning the first isomer of the retention time profile of the query EIC with the first isomer of the template, in accordance with example embodiments.
[0049] FIG. 11C highlights the data points of a query EIC, represented by red dots, are shifted by a second time increment that aligns the first isomer of the retention time profile of the query EIC with the second isomer of the template, in accordance with example embodiments.
[0050] FIG. 12A is a merged EIC of EIC files that have a specific m / z value, in accordance with example embodiments. The data is from approximately 3,000 samples.
[0051] FIG. 12B depicts data from all the EICs included in a first defined data point template, including those EICs where linear drift was identified and subsequently corrected (the right panel displays). The left panel, in contrast, contains data points from the remaining EICs that were not included in the first template.
[0052] FIG. 12C depicts data from all the EICs included in a second defined data point template, including those EICs where linear drift was identified and subsequently corrected (the right panel displays). The left panel, in contrast, contains data points from the remaining EICs that were not included in the second template.
[0053] FIG. 12D depicts data from all the EICs included in a third defined data point template, including those EICs where linear drift was identified and subsequently corrected (the right panel displays). The left panel, in contrast, contains data points from the remaining EICs that were not included in the third template.
[0054] FIG. 12E presents a comparative view between the initial state of the 3,000 samples (left panel), and the merged EICs after they have been refined and polished through methods of the instant disclosure that are shown in the right panels of FIGs 12B-D.
[0055] FIG. 13 graphically illustrates the concept of scan gaps and blur scan gaps.
[0056] FIG. 14 graphically illustrate retention time profiles that exhibit nonlinear drift and how the non-linear drift in each EIC can be corrected.
[0057] FIG. 15 is a flowchart illustrating an exemplary method 500 for defining a template in an embodiment of the methods of the instant disclosure.
[0058] FIG. 16 is a flowchart illustrating an exemplary method 600 for identifying and correcting linear drift in an embodiment of the methods of the instant disclosure.
[0059] FIG. 17 is a flowchart illustrating an exemplary method 700 for identifying nonlinear drift in an embodiment of the methods of the instant disclosure.
[0060] FIG. 18 illustrates an example of a computing system to correct for drift, in accordance with example embodiments.DETAILED DESCRIPTION
[0061] Embodiments of the present disclosure include systems and methods for correcting for retention time drift of MS mass features for the amplification and detection of MS signals. The inventors devised methods of harnessing the law of large numbers to independently align retention times of mass-features associated with a single m / z value. Importantly, methods of the instant disclosure address drift using raw data from each MS data file, before processing steps normally performed on MS data files, including processing of the mass spectra for peak detection and quantification, calibration, annotation, and noise reduction.
[0062] Methods of the instant disclosure allow unambiguous definition of individual signals which indicate “blurring” (i.e. signals that do not improve signal amplification), selectively remove the blurred signals without compromising general amplification, and allow samples containing the same waveforms to decompose into subsets of files (laterreferred to herein as “data point templates”). This latter aspect maximizes the nonrandom nature of even poorly retained compounds so that a signal can be extracted. Using these methods, even if the retention time correction is calculated incorrectly for the compounds associated with a single m / z value, these errors will be independently distributed and avoid issues with systemic biases. Additionally, the methods can remove a wide range of artifacts. Importantly, the strategy is executable in parallel, thereby allowing a balance of scalability and accuracy not available in currently existing strategies.I. Methods
[0063] One aspect of the instant disclosure encompasses methods for overcoming drift in retention times of compound signals obtained using mass spectrometry (MS) techniques coupled with separation techniques such as chromatography (chromatography-MS). As used herein, the term “overcoming” when referring to drift during retention time alignment in two or more mass spectrometry data files comprises identifying MS chromatograms that exhibit drift and correctly aligning the retention times of mass features in the two or more mass spectrometry data files to effectively address and correct the drift. The methods significantly enhance signal amplification in MS data analysis by leveraging precise data alignment to pool together small signals, thereby amplifying the overall signal of compounds. This process not only enhances the detectability of low-concentration compounds by improving the signal-to-noise ratio but also ensures that true signals are distinctly separated from background noise, facilitating more reliable identification and analysis. By pooling correctly aligned signals, the method mitigates the risk of merging or confusing signals from different compounds, maintaining data integrity and allowing for accurate interpretation of complex sample compositions, thus significantly improving the accuracy and reliability of compound identification in complex biological analyses.
[0064] As explained herein above, the methods can independently align retention times of the mass features associated with a single m / z value before performing processing steps normally performed on MS data, including processing of the mass spectra forpeak detection and quantification, calibration, annotation, and noise reduction. By addressing drift at this foundational level, the methods ensure that subsequent data processing is based on accurately aligned mass features, thereby significantly enhancing the reliability of compound identification and quantification. Moreover, this approach to overcoming drift is designed to operate independently for mass features associated with a single m / z value, ensuring that the correction is tailored and precise, avoiding the introduction of systemic biases often associated with broader correction strategies.
[0065] The methods comprise applying a two-dimensional (2D) grid system onto extracted ion chromatogram files (EIC files) of portions of each of a plurality of MS chromatograms, wherein the portions of the plurality of chromatograms bear the same m / z value. The 2D grid system comprises coordinates of a plurality of retention time increments and a plurality of signal intensity strata. The methods further comprise using the 2D grid system to sift through the plurality of EIC files to separate the EIC files into groups of files that comprise retention time profiles having characteristic waveforms, thereby defining data point templates. Individual EIC files comprising drift relative to the identified templates can then be identified and the drift corrected, thereby amplifying a compound’s signal for proper identification.(a) Mass spectrometry
[0066] MS is a powerful analytical tool that can be combined with separation methods such as chromatography (chromatography-MS) for effective detection and quantitation of various compounds in a sample. Chromatography-MS comprises a first separation of components of a sample by chromatography where compounds with different properties separate and are collected at different times. The time point at which a certain fraction of a compound elutes from the column is called the retention time (RT). The resulting fractions are then inserted online into the mass spectrometer. Here the second separation takes place. By ionizing the compounds and accelerating them to fly through a spectrometer, the compounds are separated by their mass-to-charge (m / z) ratio (also referred to herein as “mass” for the sake of simplicity). Accordingly, a plurality of m / zsignal intensities may be captured by a mass spectrometer in an output file, while intensity data records the abundance of a species of a given m / z relative to retention time. The measured data can then be depicted in a total ion chromatogram (TIC) 3D image comprising three axes. One axis depicts the RT, the second represents the m / z value, and the third axis represents the intensity or quantity of a peptide. The combined data of RT, m / z value, and intensity for each sample can be obtained as a chromatography-MS data file. FIG. 1 shows an example 3D image of UPLC-MS data from a biological sample comprising a plurality of metabolites, displaying three axes representing RT, m / z ratio, and intensity.
[0067] Methods of the instant disclosure comprise overcoming drift by focusing the analysis on a portion of a MS chromatogram bearing the same m / z value. In a typical analysis, a chromatography-MS data file can comprise up to 20,000 or more m / z values that represent the masses of plausible compounds. The retention time profile of each of these m / z values can be drawn from a TIC and can be depicted in an extracted ion chromatogram file (EIC file or EIC for short). Accordingly, an EIC is a chromatogram for a particular m / z value, such as for ion 55, 57, 83, or 105, each comprising a characteristic retention time profile. The retention time profile displays retention time, which is indicative of the time it takes for a compound to travel through the chromatographic and MS systems, and intensity, which correlates to the abundance of the ions.
[0068] This method can be particularly useful in the context of isomers, which are compounds with the same molecular formula but different structural arrangements. Despite having the same m / z value due to their identical molecular weights, isomers often exhibit different retention times in the chromatographic system due to their distinct physical and chemical properties. This allows the separation and identification of isomers within a sample. Therefore, a mass spectrometry EIC of a specific m / z value can provide a comprehensive profile that includes not only the intensity and retention time of a compound but also allows the differentiation and analysis of its isomers. For instance, it is possible to discern detailed information about compounds from the retention time profile in an EIC file, including the presence or absence of compoundisomers that share this particular m / z value, the number of compound isomers, and their relative abundance. FIG. 2 depicts hypothetical retention time profiles from EICs of a specific m / z value. This figure illustrates various retention time profile variants at this m / z value that could be present in various samples, each comprising a different number of isomers and / or intensity levels for each isomer identified by the Gaussian distribution of data values around various retention times. These retention time profile variants can be conceptualized as waveform variants in the EIC files, providing a visual representation of the diversity in the presence and concentration of compounds in various samples. The distinct waveforms correspond to the differing retention times and intensities of isomers, effectively depicting the variability in composition between samples captured at this specific m / z value. Each waveform contains valuable information about the composition of the sample, revealing both the number and quantity of isomers present.
[0069] For example, consider a biological sample where the compounds of interest are metabolites, which often exist in various isomeric forms. These isomers may have slight differences in structure but can significantly differ in their biological functions and pathways. In such a scenario, the retention time profile waveforms provide crucial insights. A complex waveform might indicate a high number of isomers, suggesting a diverse metabolic activity within the sample. On the other hand, the intensity of the waveform peaks can reveal the relative abundance of each isomer, offering clues about dominant metabolic processes. This level of detail is invaluable in fields like metabolomics, where understanding the subtle differences in metabolite profiles can be important for deciphering physiological states or disease conditions.
[0070] As previously discussed, a major challenge in utilizing chromatography-MS techniques for untargeted compound detection and quantification, such as in untargeted metabolomics, is the need to adjust for retention time drift. Existing approaches try to synchronize samples in the retention time domain using either all data points or extensive segments of the signal. However, these segments might not always act predictably, making such alignment methods susceptible to statistical biases and inaccuracies. This is particularly problematic for metabolites with isomer profiles thatare difficult to differentiate. The issue arises because the characteristics of each species are often subtle and exhibit variability across samples, typically presenting as low signal-to-noise ratios. In individual files, the signal might be too weak to distinguish even a single isomer. Combining samples to form a composite extracted ion chromatogram could enhance or amplify the detectability of a compound’s signal. However, retention time of isomers in an retention time profile of an EIC can vary from sample to sample due to inevitable drift, leading to blurring of isomer signals. Conversely, in the absence of drift, there's a risk of incorrectly interchanging isomers. FIG. 3 depicts pooled retention time profile of EIC files of a particular m / z value, derived from real data, showcasing the difficulty in interpreting such complex information.
[0071] Methods of the instant disclosure comprise independently correcting retention time drift in each of a plurality of EICs of compounds bearing the same m / z value before combining data values for each compound and / or a compound’s isomer(s) to amplify the isomers’ signals. Since mass and retention time are analyzed separately, this approach has enabled the development of retention time correction algorithms that independently adjust the signals for each m / z value (See Sections l(b) to (d) herein below).
[0072] In some embodiments, the methods comprise obtaining or having obtained one or more data files from a chromatography-mass spectrometry machine that has analyzed a sample of unknown compounds. The data in the data file can include a retention time and a mass-to-charge (m / z) signal intensity for each data point in the data file. As methods of the instant disclosure comprise independently correcting drift in a plurality EICs of compounds bearing the same m / z value, in some embodiments, a method of the instant disclosure can further comprise identifying or having identified an m / z value most likely to belong to a compound for which an EIC can be extracted. In some embodiments, the m / z value most likely to belong to a compound is identified using methods described in International Application No. PCT / US2022 / 028150, the disclosure of which is incorporated herein in its entirety.
[0073] The methods can be used to correct drift in hundreds, to thousands, to hundreds of thousands of EICs for a specific m / z value extracted from a plurality TICs. Theplurality of chromatograms can originate from a diverse range of samples, offering a broad scope for analysis and comparison. These chromatograms can be derived from a single sample, providing a detailed, focused examination of its components. Alternatively, chromatograms can be obtained from multiple samples collected independently, allowing for a comparative analysis across a wider range of variables. This versatility is particularly useful in experiments involving various treatments or conditions. For instance, samples from different treatment groups in a study can be analyzed to compare and contrast the effects of these treatments at a molecular level. Similarly, samples under different experimental conditions can provide insights into how these conditions influence the chemical composition of the samples.
[0074] Methods of the instant disclosure can be used to overcome drift in retention times in chromatography-MS runs for untargeted analyses of any compound, including chemical, pharmaceutical, biochemical, and metabolomic compounds. Non-limiting examples of chemical, biochemical, and metabolomic compounds include biochemical compounds such as amino acids, peptides and proteins, nucleotides and nucleosides, DNA and RNA fragments, lipids (fatty acids, triglycerides, phospholipids, steroids), carbohydrates (monosaccharides, disaccharides, polysaccharides), vitamins, hormones, and enzymes; metabolomic compounds such as metabolic intermediates (glycolysis intermediates, Krebs cycle intermediates), neurotransmitters, plant metabolites (alkaloids, terpenes, flavonoids), bacterial and fungal metabolites, endogenous metabolites (bile acids, urea cycle intermediates), and xenobiotics (drugs, toxins); environmental contaminants such as pesticides and herbicides, polycyclic aromatic hydrocarbons (PAHs), persistent organic pollutants (POPs), and heavy metals and metalloids (in their organic forms); pharmaceuticals such as therapeutic drugs, drug metabolites, antibiotics; polymeric materials such as polymers and copolymers, additives and plasticizers, and oligomers; isotopically labeled compounds such as stable isotope labeled metabolites, deuterated compounds; industrial chemicals such as dyes and pigments, surfactants and detergents, and explosives; food and beverage analysis such as food additives, flavor compounds, contaminants; forensic analysis such as illicit drugs, explosive residues, and trace evidence compounds; and geologicaland cosmochemical analysis such as elemental isotopes, and organic molecules in extraterrestrial samples.
[0075] In some embodiments, methods of the instant disclosure can be used to overcome drift in retention times of chromatography-MS runs for analyses of metabolites. While specific types of compounds (e.g., metabolites such as fructose and galactose) may be discussed herein, such discussion of specific embodiments is for illustrative purposes and should not be interpreted as limiting the present disclosure to the specific embodiments being illustrated and discussed.
[0076] There are multiple categories of chromatography and MS, wherein each combination of chromatography and MS methods can cater to different analytical needs, depending on the complexity of the sample, the required sensitivity, resolution, and speed of analysis. The choice of technique is often determined by the specific application and the nature of the compounds being analyzed.
[0077] Non-limiting examples of major categories of chromatographic techniques that can be coupled with MS include Liquid Chromatography-Mass Spectrometry (LC-MS), Gas Chromatography-Mass Spectrometry (GC-MS), Ion Chromatography-Mass Spectrometry (IC-MS), Supercritical Fluid Chromatography-Mass Spectrometry (SFC- MS), Capillary Electrophoresis-Mass Spectrometry (CE-MS). Liquid chromatography methods can include High-Performance Liquid Chromatography (HPLC) - the most common form, utilizing high pressure to pass the sample through a column filled with stationary phase; Ultra-Performance Liquid Chromatography (UPLC) - similar to HPLC but uses smaller particle sizes in the column for higher resolution and faster analysis; Ion Exchange Chromatography - separates ions based on their affinity to an ion exchanger in the column; Size Exclusion Chromatography (SEC) - separates molecules based on size, using a column with pores of a specific size; Normal Phase Chromatography - uses a polar stationary phase and a non-polar mobile phase;Reverse Phase Chromatography - uses a non-polar stationary phase and a polar mobile phase, opposite to normal phase; Chiral Chromatography - separates enantiomers based on their interaction with a chiral stationary phase; Affinity Chromatography - uses a stationary phase made of materials that specifically bind tothe analyte of interest. Gas chromatography methods include Capillary Gas Chromatography - uses very narrow capillary tubes with a liquid stationary phase; Packed Column Gas Chromatography - utilizes columns packed with solid stationary phase or solid support coated with liquid stationary phase; Gas-Solid Chromatography (GSC) - involves a solid stationary phase and is used primarily for separating gases or volatile compounds that don't interact with liquid stationary phases. Ion chromatography is typically used for the separation of ions and polar molecules. Supercritical Fluid Chromatography-Mass Spectrometry (SFC-MS) uses a supercritical fluid (like CO2) as the mobile phase, combining aspects of both GC and LC. It's particularly effective for analyzing compounds that are difficult to separate by traditional LC, such as chiral compounds. Capillary Electrophoresis: Although not technically a chromatographic technique, capillary electrophoresis separates ions based on their charge-to-size ratio in an electric field.
[0078] Non-limiting examples of MS include Quadrupole Mass Spectrometry (QMS) - uses quadrupole filters for mass analysis, suitable for a broad range of masses; Time- of-Flight Mass Spectrometry (TOF-MS); separates ions by their different flight times; Ion Trap Mass Spectrometry - traps ions using electromagnetic fields and then sequentially ejects them for mass analysis; Fourier Transform Ion Cyclotron Resonance (FT-ICR) - offers very high resolution and accuracy, using a magnetic field to trap ions; Orbitrap Mass Spectrometry - uses an electrostatic field to trap ions in an orbital motion around a central electrode; Triple Quadrupole Mass Spectrometry (QqQ) which incorporates three quadrupoles in series; commonly used for quantification due to its high sensitivity and specificity; Tandem Mass Spectrometry (MS / MS) which involves multiple stages of mass spectrometry, often with fragmentation of analyte ions between stages;Quadrupole Time-of-Flight Mass Spectrometry (Q-TOF) which combines quadrupole mass filtering with TOF mass analysis for high accuracy and resolution; and Magnetic Sector Mass Spectrometry which uses a magnetic field to deflect ions, with separation based on mass-to-charge ratio.
[0079] In some embodiments, methods of the instant disclosure can be used to overcome drift in retention times of LC-MS runs. In some embodiments, methods of theinstant disclosure can be used to overcome drift in retention times of LC-MS runs for analyses of metabolites.
[0080] In addition to chromatography, different separation techniques can also be used in conjunction with mass spectrometry. Non-limiting examples of suitable separation techniques other than chromatography include electrophoresis, ion mobility, Field-Flow Fractionation (FFF), Capillary Electrophoresis (CE), Matrix-Assisted Laser Desorption / lonization (MALDI), Thermal Desorption (TD), Laser Ablation (LA), Desorption Electrospray Ionization (DES I). Electrophoresis separates charged particles under an electric field, effectively used for biomolecules like DNA and proteins. Ion mobility spectrometry, on the other hand, separates ions based on their mobility in a gas phase under an electric field, useful for distinguishing isomers and conformers. FFF techniques separate particles, macromolecules, or colloids based on their size or density in a fluid under various field forces (such as centrifugal, thermal, or electrical fields). Coupling FFF with MS allows for the analysis of large and complex molecules like proteins and polymers. CE is similar to electrophoresis but conducted in capillary tubes. CE is effective for separating ionic species using an electric field. Its coupling with MS provides high resolution and efficiency, particularly useful for analyzing biomolecules like peptides and nucleotides. Although MALDI is more of an ionization technique than a separation technique, it is often mentioned in the context of MS coupling. It's particularly useful for the analysis of large biomolecules like proteins, DNA, and polymers. TD involves heating a sample to release volatile and semi-volatile compounds. When coupled with MS, it allows for the analysis of compounds in air, materials, and environmental samples. LA comprises using a laser to remove material from a solid sample. Coupling LA with MS enables the analysis of solid samples, particularly in fields like geology and material science, allowing for elemental and isotopic analysis. DESI allows for the direct analysis of samples (even from surfaces) under ambient conditions. It's useful for a wide range of applications, including biological tissues and environmental samples.(b) Data point templates
[0081] Prior to the development of methods of the instant disclosure, several processing steps were typically necessary to convert MS data into a format suitable for analysis, annotation, correction of retention time drift, and signal amplification. These steps, essential for each chromatography-MS run, included peak detection and quantification, calibration, and noise reduction. However, these traditional methods face challenges due to the often faint and fluctuating outlines of mass features across samples, characterized by a low signal-to-noise ratio such as in metabolomic studies. This problem is particularly acute when the signal of one mass feature is closely aligned with another, with each experiencing its own distinct chromatographic drift. As a result, these conventional methods can inadvertently discard valid compound signals and tend to be computationally demanding and time-consuming.
[0082] Contrary to previous strategies for amplifying signals, methods of the instant disclosure correct drift using raw data obtained from LC / MS run, bypassing the standard preprocessing steps such as mass spectra processing for peak detection and quantification, calibration, and noise reduction. This represents a significant departure from previous strategies, focusing on signal amplification through direct manipulation of unprocessed data.
[0083] Each EIC among the plurality of EICs can comprise a retention time profile which may correspond to one of several waveforms and waveform variants observed in the collection of EICs for a specific m / z value. This concept is illustrated in FIG. 3 and FIG. 4. The complexity of MS data exemplified in the merged retention time profiles of EICs of a particular m / z value shown in FIG. 3, particularly the presence of drift, can be diagrammatically represented in the first panel of FIG. 4, where multiple waveforms within the merged retention time profiles are depicted. The first panel of FIG. 4 shows how complex and chaotic MS data can be before adjusting for drift. However, this complexity can be resolved by recognizing that each retention time profile can align with one of these waveforms. By matching each retention time profile to the waveform it most closely resembles, a method central to this disclosure, drift can be accurately corrected. This process simplifies and organizes the data, as shown in the second panel of FIG. 4. Thus, the once entangled signals can be effectively separated intodistinct waveforms through this waveform matching technique, demonstrating the utility of the methods disclosed herein in deconvoluting complex MS data for further analysis.
[0084] To deconvolute signals in a merged EIC, methods of the instant disclosure can comprise defining a plurality of data point templates, wherein each data point template can correspond to one of the plurality of waveforms found in the plurality of retention time profiles in the EICs from that specific m / z value. To define the data point templates, the methods comprise mapping all the data point values of all the EICs from that specific m / z value onto a two-dimensional (2D) grid system. The 2D grid system comprises a plurality of cells, hereinafter referred to as 'grid boxes', each designated by an X and a Y coordinate. The X coordinate is an increment of retention time, while the Y coordinate is an increment of signal intensity. The retention time increments are referred to hereinafter as scans and the signal intensity increments are referred to herein after as strata. Accordingly, each grid box is defined by the intersection of a specific scan and a specific stratum. As each data point in an EIC represents a specific ion characterized by its intensity and the time at which it is detected (RT), mapping the data points of an EIC onto the grid system assigns each data point the coordinates of the grid box into which it falls. Additionally, each grid box can comprise multiple data points. The number of data points in a grid box can and will vary depending on the density of data points in an EIC file and the dimensions of the grid box as defined by the size of the scan and the size of the stratum.
[0085] The range of values for the retention time increments of the scans within the 2D grid system can and will vary based on the distribution of obtained data values, the characteristics of the retention time profile, the shape and number of waveforms, the number of data point values in the EICs, and the desired level of precision for defining templates. In some embodiments, scans can be categorized by the duration of RT increments. For instance, scans can comprise retention time increments of durations ranging from about 0.1 second or less to about 15 minutes or more, from about 1 second to about 10, 9, 8, 7, 6, 5, 4, 3, 2, 1 minutes or less, from about 1 second to about 50, 40, 30, 20, 10, 5 4, 3, or 2 seconds, or from about 30 seconds to about 1 , 2, 3, 4, 5, 6, 7, 8, 9, 10, 11 , 12, 13, 14, 15, 16, 17, 18, 19 or about 20 minutes or longer.Scans can also be categorized by the total number of scans desired in a merged EIC. For instance, a 2D grid system can also comprise 10, 15, 20, 25, 30, 35, 40, 45, or about 50 or more scans. Furthermore, scans can be defined by the number of data point values in each retention time increment of a scan. For instance, scans can be defined as a range of RT increments, wherein all RT increments can comprise the same number of data point values.
[0086] In some embodiments, the scans comprise retention time increments categorized by the duration of each RT increment. In some embodiments, the scans comprise retention time increments of equal durations.
[0087] The range of intensity values of the strata within the 2D grid system can and will vary based on the distribution of obtained data values, the characteristics of the retention time profile, the shape and number of waveforms, the number of data point values in the EICs, the type of mass spectrometer used, the nature of the sample, the ionization method, and the sensitivity of the detector, and the desired level of precision for defining templates. Strata can be categorized by increments of intensity. Strata can also be categorized by the total number of strata desired in a merged EIC. For instance, a 2D grid system can comprise 10, 15, 20, 25, 30, 35, 40, 45, or about 50 or more strata. Furthermore, strata can be defined by the number of data points. For instance, strata can be defined as a range of numbers of data points, wherein all strata can comprise the same or different numbers of data points. In some embodiments, the strata comprise an equal number of data points.
[0088] Various types of 2D grid systems suitable for use in a method of the instant disclosure can be as described herein further below. In some embodiments, methods of the instant disclosure comprise mapping the plurality of retention time profiles onto a 2D cartesian grid system.
[0089] FIG. 5 depicts an example of merged retention time profiles from a specific m / z value mapped onto a 2D cartesian grid of scans and strata. In the embodiment of the grid system shown in FIG. 5, each grid box comprises a scan of 0.5 minutes and strata defined by having an equal number of data points in each stratum.
[0090] The methods further comprise defining templates for one or more waveforms that can be present in a merged EIC. Each comprises a complete set of grid boxes, each grid box defined by its coordinates, that contain one or more data points from one or more EICs comprising a retention time profiles that corresponds to a particular waveform and variants of that waveform (see further below). This means that a template (or even just the collection of grid boxes of a template, i,e, the template grid boxes) encapsulates the entire pattern of a specific waveform and its variants as represented by the grid boxes that contain data points from EIC files determined to match a template. For instance, referring to FIG. 2, if all the waveforms shown are represented in a plurality of EICs associated with a single m / z value, methods of the instant disclosure can define a data point template for each of the waveforms of FIG. 2.
[0091] According to embodiments of the instant disclosure, a data point template is defined by first selecting a reference scan. According to embodiments of the instant disclosure, a reference scan is selected based on two criteria: the intensity values of the data points in the scan and by the number of EICs these data points are associated with. More specifically, a scan can be selected as a reference scan if it comprises data points that meet a predefined intensity threshold. Additionally, these data points must be associated with a threshold number of EICs. That means that a scan can be selected as a reference scan if it includes data points that not only reach a certain intensity level but also correspond to a threshold number of distinct EICs. The process of selecting a reference scan can comprise (a) identifying scans that contain data points with the threshold intensity value, and (b) selecting the scan that has data points belonging to the threshold number of EICs as the reference scan or vice versa. As such, the term “reference scan” refers to a scan in the 2D grid, wherein the scan comprises data points of a threshold intensity value and wherein the data points in the reference scan also belong to a threshold number of the EICs in a merged EIC.
[0092] According to embodiments of the instant disclosure, all retention time profiles in EICs identified through this reference scan are considered to exhibit the same waveform characteristics. In essence, these retention time profiles, when viewedtogether, constitute a distinct waveform that is representative of the data points collected at that specific reference scan.
[0093] After selecting a reference scan, methods of the instant disclosure comprise defining a template based on the retention time profiles of the EICs identified through this reference scan. The data point template is essentially a collection of grid boxes on the 2D grid system. The template encompasses all grid boxes that comprise all the data points from the EICs identified by the reference scan. Specifically, the data point template comprises every grid box that contains at least one data point from any of these identified EICs. This comprehensive inclusion ensures that the template accurately represents the waveform's profile by including all relevant intensity and retention time data points.
[0094] FIG. 6 graphically depicts an example showing waveforms that can be represented in a data point template defined by a selected reference scan. As seen in FIG. 6, when using methods of the instant disclosure to select a reference scan, retention time profiles of EICs identified by the selected reference scan can correspond to more than one waveform variant. These EICs can be obtained from the analysis of samples exhibiting variations in their isomer composition, despite being identified under the same reference scan, showcasing diversity in molecular structure or concentration levels. For instance, FIG. 6 depicts EICs identified by a selected reference scan, wherein two waveform variants are captured. A first waveform variant shown in red were obtained from samples comprising two isomers (A and C) and a second waveform variant shown in black was obtained from samples comprising three isomers (A, B, and C).
[0095] A threshold intensity value and a threshold number of EICs suitable for use in identifying reference scans in a method of the instant disclosure can and will vary based on the distribution of data values in a merged EIC, the characteristics of the retention time profile EICs, the shape and number of waveforms, the number of data point values in the EICs, and the desired level of precision for defining templates. For instance, a threshold intensity can be set to include the highest about 1 %, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, or about 50% or higherintensity values. This means choosing data points with intensity values in the top percentile ranges according to the specified percentages. In some embodiments, a method of the instant disclosure comprises defining a data point template by selecting a reference scan, wherein the reference scan comprises the highest intensity values present in a threshold number of EICs in a merged EIC.
[0096] Similarly, a threshold number of EICs that comprise data points having a threshold intensity value can be a percentage of the total number EICs of the merged EICs or of the remaining EICs (see definition of remaining EICs herein further below). For instance, a threshold number of EICs can be the top 1 %, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, or about 50% or higher. This approach selects a proportion of EICs based on their relative position in the overall distribution of intensity values. In some embodiments, a method of the instant disclosure comprises defining a data point template by selecting a reference scan, wherein the reference scan comprises a threshold intensity value present in the highest number of EICs.
[0097] In some embodiments, a method of the instant disclosure comprises defining a data point template for the most prevalent waveform among the retention time profiles in a merged EIC. In some embodiments, the most prevalent waveform is a waveform represented by retention time profiles of EICs identified when a reference scan is selected as a scan comprising the highest intensity values present in the highest number of EICs in a merged EIC.
[0098] In some embodiments, defining a data point template can further comprise identifying among the EIC files identified by a reference scan, one or more EICs comprising retention time profiles that exhibit non-linear drift, and excluding the EICs that introduce non-linear drift before a data point template is constructed for the waveform of the EICs identified by a reference scan. This is because, when using methods of the instant disclosure to select a reference scan, retention time profiles of some EICs identified by the selected reference scan can exhibit non-linear drift, despite being identified by the same reference scan. FIG. 7 graphically depicts an exampleshowing waveforms that can be represented in a data point template defined by a selected reference scan before excluding an EIC that introduces non-linear blur (C’).
[0099] In some embodiments, identifying EICs identified by a reference scan and that comprising retention time profiles exhibiting non-linear drift can be as described in Section l(d) herein below.
[0100] In other embodiments, identifying an EIC file identified by a reference scan and that comprises a retention time profile exhibiting non-linear drift can comprise (a) calculating an average intensity value of all data points at each scan in the template, and (b) identifying an EIC file as a file comprising a retention time profile exhibiting nonlinear drift if the EIC comprises one or more data points having intensity values significantly higher than the average intensity value at one or more scans. According to embodiments of the instant disclosure, coordinates of data points of the identified EIC can be removed from, or not added to the template. Alternatively, non-linear drift can be corrected in each identified EIC file. Correcting non-linear drift can be as described in Section l(d) herein below.
[0101] This concept can be illustrated in FIG. 8. In FIG. 8, green dots represent data points of a data point template that corresponds to a waveform representing EICs obtained from the analysis of samples comprising a compound (X) and one isomer (Y). Grey dots represent data points of an EIC obtained from the analysis of a sample also comprising a compound and one isomer. However, in the case of the EIC represented by the grey dots, the isomer (Y’) exhibits non-linear drift in relation to the Y isomer even though the EIC was identified by a reference scan that also identified the EICs represented by the green dots. As illustrated in FIG. 8, the grey EIC comprises a retention time profile exhibiting non-linear drift can be identified because the EIC comprises data points comprising intensity values at Y’ that are significantly higher than the average intensity values of data points at the scans of the data points. The EIC file with the grey dots can be removed from the template or the non-linear drift can be corrected before adding to the template. This method can also be used on data files when all templates are combined, before or after correcting linear drift as depicted in FIGs. 9A, 9B, and 9C. The EICs comprising retention time profiles exhibiting non-lineardrift can be removed from the template, and linear drift corrected to match the two removed EICs. Non-linear drift can then be corrected as explained in Section 1(d) herein below. Correcting linear drift can be as described in Section 1(c) herein below.
[0102] In some embodiments, methods of the instant disclosure can comprise iteratively defining multiple data point templates in a merged EIC of a particular m / z value. As such, a method of the instant disclosure can comprise defining a first data point template, then defining a second data point template among EICs remaining after the first template has been identified. The process of defining data point templates can be repeatedly applied to the remaining EICs after each round of defining a data point template until no additional templates can be reliably defined.
[0103] In some embodiments, iteratively defining multiple data point templates in a merged EIC of a particular m / z value comprises defining a first data point template followed by identifying and correcting linear drift among the remaining EICs, before iteratively defining additional templates among the remaining EICs. Correcting linear drift can be as described in Section l(c) herein below. In some embodiments, iteratively defining multiple data point templates in a merged EIC of a particular m / z value comprises defining a first data point template, then iteratively defining additional template before identifying and correcting linear drift among the remaining EICs.
[0104] There are several 2D grid systems used for organizing and dividing data points with x and y coordinates. The choice of a 2D grid system for organizing this data depends on the analysis requirements and the nature of the data. Non-limiting examples of some commonly used 2D grid systems include a cartesian grid system, hexagonal grid system, triangular grid system, polar grid system, isometric grid system, logarithmic grid system, quadtrees, and Voronoi diagrams. When selecting a grid system for MS data, it's important to consider the nature of the data and the goals of the analysis. Different grid systems can highlight different aspects of the data, such as the distribution of ions, relationships between different ions, or specific regions of interest within the data. The choice of grid system can significantly influence the insights gained from the data and the ease with which these insights can be communicated to others.
[0105] A cartesian grid system is the most straightforward method for MS data. The x-axis could represent the m / z values, while the y-axis could depict the RT. Each point on the grid corresponds to a detected ion, with its position reflecting its m / z value and RT. The range of the grid is determined by the range of m / z and RT values in the dataset. Adjusting the scale of the axes can change the granularity of the data visualization, allowing for a more detailed look at specific regions or a broader overview.
[0106] In a hexagonal grid system, data points are enclosed in hexagons. This could be useful for clustering similar ions, as the hexagonal arrangement allows for a more uniform distribution of points compared to square grids. The range of values is divided into hexagonal cells, each representing a range of m / z and RT values. A triangular grid system is similar to the hexagonal system, but with triangular cells. This might be useful for certain types of pattern recognition or data categorization. Each triangle in the grid encompasses a specific range of m / z and RT values.
[0107] A polar grid system could be used if there's a central point of interest in the data, such as a specific m / z value or RT. The angle could represent the m / z value while the radius could represent the RT. This approach might be less intuitive for MS data but could provide novel insights for certain datasets.
[0108] An isometric grid system can appear three-dimensional. However, this system is used for 2D data and could provide a different perspective on MS data. It might be less conventional but could be useful for emphasizing relationships between data points.
[0109] If the MS data spans a wide range of m / z values or RT, a logarithmic scale might be more appropriate. This allows for a more detailed visualization of data with large variances in values.
[0110] Quadtrees is a hierarchical system that could be used for MS data that requires dynamic range adjustments. Different levels of the tree can represent different resolutions of m / z and RT values, useful for zooming in on areas of interest.
[0111] Voronoi diagrams could be used to visualize the density or distribution of ions in the MS data. Each ion's m / z and RT values define a region on the grid, with boundaries formed based on proximity to neighboring points.
[0112] For each of these systems, the range of values that the grid is split into can vary. For instance, in a Cartesian grid, one might choose to have a linear scale or a logarithmic scale, depending on the distribution of the data points. The granularity of the grid (i.e. , how fine or coarse the division lines are) can also be adjusted based on the level of detail required for analysis.(c) Identifying and correcting linear drift
[0113] After defining one or more data point templates as described in Section l(b) herein above, methods of the instant disclosure can comprise identifying EICs comprising retention time profiles exhibiting linear drift when compared to a data point template and correcting the linear drift. A linear drift in chromatography-MS retention times is characterized by a consistent and proportional change in the RT over time or across a sequence of runs. The key feature of a linear drift is its predictability and uniformity. In other words, for each unit of time or for each additional chromatographic run, the retention time drifts by a constant amount. This can be represented as a straight line when plotting RT changes against time or sequence number. For example, if every subsequent run results in a retention time that is 0.1 minutes later than the previous run, this is a linear drift. Mathematically, it can be expressed as RT = a + b * n, where RT is the retention time, a is the initial retention time, b is the constant rate of drift, and n is the number of runs or time units. Linear drift and correction of linear drift is illustrated in FIG. 10A. FIG. 10A graphically depicts two retention time profiles comprising the same waveform variant but are linearly drifted by a retention time value. When two retention time profiles exhibit linear drift, the data point values of each retention time profile can mostly overlap when all the points of one of the retention time profiles match in RT when shifted by the same time increment as shown in FIG. 10B.
[0011] In some embodiments, methods of the instant disclosure comprise iteratively defining more than one data point template and then identifying individual EICs from the remaining EICs that do not match the defined data point templates and that further exhibit linear drift in comparison to any of the defined templates. Individual EICs that comprise retention time profiles comprising linear drift in relation to one of thedata point templates can then be identified among the remaining EICs that were not found to match a data point template, thus capturing all EICs that have linear drift in relation to the first scan template. In some embodiments, this identification and correction process is repeated for each EIC until no more EICs comprising retention time profiles showing linear drift relative to the template(s) are found.
[0115] In some embodiments, methods of the instant disclosure comprise initially defining a first data point template and then detecting individual EICs from the remaining set of EICs, EIC files that comprise retention time profiles that exhibit linear drift in comparison to this first template, followed by correcting the drift. This identification and correction process can then be repeated for each EIC until no more EICs comprising retention time profiles showing linear drift relative to the first template are found. Subsequently, a second data point template is identified among the EICs that remain unprocessed. The process of identifying and correcting drift in EICs comprising retention time profiles with linear drift is then repeated with respect to the second template. This iterative approach continues until no additional data point templates exhibiting linear drift can be identified among the EICs.
[0116] To identify the EICs, among remaining EICs, that comprise retention time profiles that have linear drift in relation to the first template, the methods comprise generating a query EIC for each remaining EIC of the plurality of EICs by shifting the data points of the query EIC by a time increment sufficient to overlap data point of the query EIC comprising a predefined threshold intensity with the reference scan of a template. A retention time profile is considered to exhibit linear drift relative to the data point template if, after the shift, all data points of the query EIC lie in a grid box of the template. In other words, a successful alignment occurs when the shifted data points of the query EIC lie in grid boxes that coincide with existing grid boxes in the template, indicating good alignment. Said another way, a retention time profile is considered to exhibit linear drift relative to the data point template if, after the shift, none of the data points of the query EIC lie in a grid box that is not part of the template. Conversely, a poor alignment is indicated by a shift leading to data points that lie in previously empty grid boxes on the 2D grid, resulting in a 'blur'. A template can be defined as describedin Section 1(b) herein above. Once an EIC file is identified as comprising a retention time profile exhibiting linear drift, the linear drift can be corrected by adding the data points of the EIC to a merged corrected EIC file comprising all data points of EICs that match the template after shifting the data points as described herein, thereby contributing to enhancement and “de-blurring” of signals in the merged EIC file. Accordingly, a merged corrected template file comprises all coordinates of data points of EIC files of an identified template and coordinates of data points of EICs comprising data point values corrected for linear drift.
[0117] In some embodiments, to identify the EICs, among remaining EICs, that comprise retention time profiles that have linear drift in relation to the first template, the methods comprise generating a query EIC for each remaining EIC of the plurality of EICs by shifting the data points of the query EIC by a time increment sufficient to overlap the data point of the query EIC comprising the highest intensity with the reference scan of the data point template. In some embodiments, the reference scan is selected as having the highest data points in a threshold number of EICs of the data point template.
[0118] This concept is illustrated in FIGs. 11A to 11C using a data point template of a waveform of an m / z region comprising two isomers, each observed as a peak. The data points in the template shown in the figure are shown with black dots on a 2D grid system, while data points of a query EIC are shown with red dots. It is important to note that reference to the isomers in the description of FIGs. 11A to 11C herein is for demonstration purposes only, as methods of the instant disclosure apply to raw data from MS data before peak detection. In FIG. 11 A, the data points of the query EIC, represented by red dots, do not align with the template's black dots, indicating the presence of drift (linear or non-linear) in the retention time profile relative to the template. In FIG. 11 B, the data points of the query EIC are shifted by a first time increment as described herein above, aligning the first isomer of the retention time profile of the query EIC with the first isomer of the template. The alignment of all the data points of the query EIC resulting from shifting the data points of the query EIC by the first time increment indicates that the query EIC comprises a retention time profileexhibiting a linear drift. Further, alignment of all the data points of the query EIC resulting from shifting the data points of the query EIC by the first time increment also indicates that shifting the data points of the EIC by the first time increment can correct the linear drift. FIG. 11C presents a different case where the data points of the query EIC are shifted by a second time increment that aligns the first isomer of the retention time profile of the query EIC with the second isomer of the template. This adjustment fails to align all the data points of the query EIC within the template and results in data points being present in grid boxes that are not part of the template. This misalignment, termed as 'blur', suggests that grid boxes comprising data points are created in the template where none existed previously, indicating an incorrect correction for linear drift. FIG. 11C depicts an instance where the misalignment leads to an isomer swap as the time increment used for the shift aligns the wrong isomers. This phenomenon shows an example of how the methods of the instant disclosure can detect and overcome isomer swaps, overcoming challenges present in current MS data drift correction techniques.
[0119] FIGs. 12A-E present a real-world example of using methods of the instant disclosure to identify waveforms and correct linear drift, utilizing data from approximately 3,000 samples. In the case of each individual EIC file among these samples, the signals of compounds might be too faint to detect, posing a challenge in identifying specific compounds. However, an attempt to overcome this by pooling the EIC files together to amplify the signals leads to a different problem: the details become indistinct or 'blur' as shown in FIG. 12A depicting data points in a merged EIC file before any corrections are made according to methods of the instant disclosure. In FIG. 12A a fundamental dilemma in traditional mass spectrometry data analysis is highlighted: the data requires deconvolution for peak identification, but paradoxically, identifying these peaks is a prerequisite for effective deconvolution. While the pooled data suggest the presence of peaks corresponding to isomers, the lack of clarity and definition in this data makes it challenging to accurately identify and distinguish these peaks and their associated isomers. In this example, methods of the instant disclosure are applied to create templates, and the templates are then used to identify and correct linear drift in retention time profiles of EICs that have a specific m / z value. Using an embodiment ofmethods of the instant disclosure, a first template is defined, and EICs comprising retention time profiles exhibiting linear drift in relation to the first template are identified and corrected. In FIG. 12B, the right panel displays data points of a first merged corrected EIC file comprising data from all the EICs included in the first template, including those EICs where linear drift was identified and subsequently corrected. The left panel, in contrast, contains data points from the remaining EICs that were not included in the first template. In accordance with some embodiments of the instant disclosure, the remaining EICs not included in the first template were used in turn to define a second template. This second template serves to identify and address linear drift in the EICs among the remaining EICs that comprise linear drift relative to the second template, as depicted in the second panel of FIG. 12C depicting data points in a second merged corrected EIC file comprising data from all the EICs included in the second template, including those EICs where linear drift was identified and subsequently corrected. Subsequently, the EICs not accounted for in the second template, shown in the first panel of FIG. 12C, are employed to define a third template, as illustrated in the second panel of FIG. 12D, which depicts data points in a third merged corrected EIC file comprising data from all the EICs included in the third template, including those EICs where linear drift was identified and subsequently corrected. The EICs that remain after this third template identification are displayed in the first panel of FIG. 12D. FIG. 12E presents a comparative view between the initial merged EIC of the 3,000 sample (left panel), and the merged EICs after they have been refined and polished through methods of the instant disclosure that are shown in the right panels of FIGs 12B-D. This figure effectively demonstrates how the methods described here successfully identify and correct linear drift in the data, leading to a successful deconvolution of the data and amplification of signals.(d) Identifying and correcting non-linear drift
[0120] After deconvoluting the linear drift by defining templates and correcting for linear drift, methods of the instant disclosure can comprise identifying EIC files comprising retention time profiles that exhibit nonlinear drift and correcting the non-linear drift. In some embodiments, EIC files comprising retention time profiles that exhibit nonlinear drift are identified among EICs that match a template. In some embodiments, EIC files comprising retention time profiles that exhibit nonlinear drift are identified among EICs that match a template and data points of EIC files after correcting for drift. EIC files of templates and EIC files corrected for linear drift can be identified as described in Sections 1(b) and (c) herein above.
[0121] Non-linear drift in chromatography-MS retention times is characterized by changes that do not follow a straight-line pattern over time or across a series of runs and can result from a variety of unpredictable factors such as changes in the column, temperature fluctuations, or inconsistencies in the mobile phase. In non-linear drifts, the change in RT is disproportionate and can vary in rate and direction. Non-linear drifts are unpredictable and can be irregular, making them more challenging to correct. For instance, if the retention time increases rapidly after the first few runs and then starts to decrease or fluctuate irregularly in later runs, this would be an example of non-linear drift. The mathematical representation of non-linear drift is not as straightforward as linear drift and may require complex modeling to describe accurately.
[0122] One concept used in identifying non-linear drift in methods according to the instant disclosure comprises identifying mutually exclusive pairs of coordinates. More specifically, a concept used in identifying non-linear drift in methods according to the instant disclosure comprises a comparative analysis of pairs of coordinates in template files and all pairs of coordinates in EIC files determined to match the template to identify mutually exclusive pairs of coordinates. Mutually exclusive pairs of coordinates are pairs of coordinates that occur in template files but do not occur in any of the individual EIC files. Pairs of coordinates can be pairs of data point coordinates or pairs of grid boxes. When the pairs of coordinates are pairs of data point coordinates, the pair of coordinates identify the pair of grid boxes in which each data point of the pair of data point coordinates is present.
[0123] In some embodiments, pairs of coordinates are pairs of grid boxes. In such embodiments, methods according to the instant disclosure comprise performing a comparative analysis of all pairs of template grid boxes and all pairs of EIC grid boxes.Said another way, methods used in identifying non-linear drift can comprise performing a comparative analysis of all pairs of grid boxes that can be obtained by pairing template grid boxes in all possible combinations and pairs of grid boxes obtained by pairing all EIC box pairs for each individual EIC file determined to match the template. EIC box pairs of an EIC file can be obtained by pairing all grid boxes of the EIC in all possible combinations. EIC box pairs can be noted in a template box pairs file. In some embodiments, pairs of grid boxes obtained by pairing all EIC box pairs for each individual EIC file can be combined into a collection of pairs of grid boxes, which can be noted in a combined EIC box pairs file.
[0124] In some embodiments, each pair of grid boxes can be represented by a unique identifier to facilitate the comparative analysis. These identifiers can be derived from the specific scan and stratum values that define each grid box, effectively allowing for precise comparison and analysis of the grid box pairs. Identifiers can also include information about the EIC file or template from which the box pairs are obtained. In some embodiments, methods of the instant disclosure comprise a comparative analysis of pairs of grid boxes present in a corrected template. In other words, in some embodiments, grid boxes also include grid boxes that comprise data point coordinates of data points corrected for linear drift.
[0125] The crux of the analysis lies in the comparative analysis between a template box pairs file and an EIC box pairs file or a combined EIC box pairs file. The objective here is to compare these two files, particularly focusing on identifying box pairs that are present in the template box pairs file but are absent in the EIC box pairs file or combined EIC box pairs file. The pairs of grid boxes that are present in the template box pairs file but are absent in the EIC box pairs file or combined EIC box pairs file are referred to herein as mutually exclusive pairs of grid boxes. Accordingly, methods of the instant disclosure can comprise identifying mutually exclusive pairs of grid boxes. Data points in mutually exclusive pairs of grid boxes identify EICs that comprise retention time profiles that exhibit non-linear drift. More specifically, data points in a grid box of an individual EIC, wherein the grid box is part of a box pairdetermined to be mutually exclusive identify EICs that comprise retention time profiles that exhibit non-linear drift.
[0126] In some embodiments, a method of identifying non-linear drift comprises constructing a template box pairs file and constructing a EIC box pairs file for each EIC file determined to match the template. In some embodiments, the method further comprises constructing a combined EIC box pairs file by combining EIC box pairs files constructed for each EIC file determined to match the template. In some embodiments, each pair of grid boxes can be represented by a unique identifier to facilitate the comparative analysis. In some embodiments, the methods further comprise identifying mutually exclusive pairs of grid boxes, wherein mutually exclusive pairs of grid boxes are pairs of grid boxes that are present in the template box pairs file but are absent in the EIC box pairs file or combined EIC box pairs file. In some embodiments, the methods can further comprise identifying EICs that comprise retention time profiles that exhibit non-linear drift, wherein a data point in a grid box of a box pair determined to be mutually exclusive identifies an EIC that comprise retention time profiles that exhibit non-linear drift.
[0127] In other embodiments, pairs of coordinates are pairs of data point coordinates. In such embodiments, methods according to the instant disclosure comprise a comparative analysis of all pairs of data point coordinates present in an identified data point template and all pairs of data point coordinates in each individual EIC file represented in the data point template. In some embodiments, methods of the instant disclosure comprise a comparative analysis of pairs of coordinates of all data point coordinates present in a corrected template. In other words, in some embodiments, data point coordinates include coordinates of data points corrected for linear drift.
[0128] In such embodiments, an initial step in identifying EIC files comprising retention time profiles exhibiting non-linear drift comprises merging all data point coordinates of two or more data point templates or corrected templates into one master coordinates file. This master coordinates file, therefore, encapsulates the complete set of data point coordinates drawn from the entire collection of individual EIC filesrepresented in the two or more data point templates or corrected data point templates. All possible pairs of coordinates in the master coordinates file are then identified to generate a “master coordinate pairs file”. Additionally, in some embodiments, all possible pairs of coordinates are identified between all data points within each of the individual EIC files represented in the two or more data point templates or corrected data point templates. In some embodiments, the pairs of coordinates identified within each of the individual EIC files represented in the two or more data point templates or corrected data point template can be consolidated into a single, aggregated file, referred to herein as a “child coordinate pairs file”. This child coordinate pairs file represents a comprehensive compilation of coordinate pairs identified across all individual EIC files.
[0129] In these embodiments, the crux of the analysis lies in the comparative analysis between the master coordinate pairs file and the child coordinate pairs file. The objective here is to compare these two files, particularly focusing on identifying pairs of coordinates that are present in the master coordinate pairs file but are absent in the child coordinate pairs file. The pairs of coordinates that link grid boxes that are not linked in any of the individual EICs are referred to herein as mutually exclusive pairs of coordinates. Accordingly, methods of the instant disclosure can comprise identifying mutually exclusive pairs of coordinates. In turn, mutually exclusive pairs of coordinates identify mutually exclusive pairs of grid boxes. Data points in mutually exclusive pairs of coordinates or mutually exclusive pairs of grid boxes identify EICs that comprise retention time profiles that exhibit non-linear drift.
[0130] Accordingly, in some embodiments, methods of the instant disclosure can comprise merging all coordinates of two or more data point templates or corrected templates to generate a master data point coordinates file. In some embodiments, the methods further comprise generating a master coordinate pairs file comprising all possible pairs of coordinates in the master coordinates file. In some embodiments, the methods can further comprise identifying all possible pairs of data point coordinates within each of the individual EIC files represented in the two or more data point templates or corrected data point templates. In some embodiments, methods of the instant disclosure comprise generating a child coordinate pairs file comprising allpossible pairs of coordinates within each of the individual EIC files represented in the two or more data point templates or corrected data point templates. In some embodiments, the methods further comprise identifying mutually exclusive pairs of data point coordinates. Mutually exclusive pairs of data point coordinates are pairs of coordinates that occur in the master template but not in each EIC file comprising at least one data point in each of the pairs of coordinates. The methods can further comprise identifying mutually exclusive pairs of grid boxes, wherein mutually exclusive pairs of coordinates identify mutually exclusive pairs of grid boxes. Data points in mutually exclusive pairs of coordinates or mutually exclusive pairs of grid boxes identify EICs that comprise retention time profiles that exhibit non-linear drift.
[0131] In some embodiments, methods of the instant disclosure can comprise (a) merging all coordinates of two or more data point templates or corrected templates to generate a master coordinates file; (b) generating a master coordinate pairs file comprising all possible pairs of coordinates in the master coordinates file; (c) identifying all possible pairs of coordinates within each of the individual EIC files represented in the two or more data point templates or corrected data point templates; (d) generating a child coordinate pairs file comprising all possible pairs of coordinates within each of the individual EIC files represented in the two or more data point templates or corrected data point templates; (e) identifying mutually exclusive pairs of grid coordinates; (f) identifying mutually exclusive pairs of grid boxes, wherein mutually exclusive pairs of coordinates identify mutually exclusive pairs of grid boxes; and (g) identifying EICs that comprise retention time profiles that exhibit non-linear drift, wherein data points in mutually exclusive pairs of coordinates or mutually exclusive pairs of grid boxes identify EICs that comprise retention time profiles that exhibit non-linear drift.
[0132] The task at hand, given the sheer volume and complexity of the data involved, necessitates substantial computational resources and time. This is because a dataset can comprise thousands of individual files, each potentially containing thousands of data points, data point coordinates, grid boxes, and an exponentially higher number of pairs of grid boxes. The computational challenge here is twofold: first, in the meticulous identification and pairing of grid boxes or data point coordinates withinboth the template file and the individual EIC files, and second, in the intricate comparison between the template file and the combined EIC files.
[0133] This immense data size means that processing, even with advanced computational methods, will be a resource-intensive task. The identification of unique box pairs, especially those present in the template box pairs file but absent in the combined EIC box pairs files, requires not only sophisticated algorithms but also significant processing power to handle the data efficiently. The exercise demands high- capacity memory, fast processing speeds, and possibly parallel computing capabilities to manage and analyze the data within a reasonable timeframe. Moreover, the intricacy of the task may necessitate iterative processes, further adding to the computational load. As such, this endeavor is not just a test of data analysis methodologies but also a challenge in terms of computational capacity and efficiency, underscoring the need for robust and powerful computing systems to successfully undertake and complete this extensive data analysis exercise.
[0134] To alleviate the computational problem, a method of the instant disclosure can further comprise using methods that can reduce the computational resources needed. One of these methods comprises identifying mutually exclusive pairs of grid boxes by generating scan gaps between pairs of data point coordinates in the template file and generating scan gaps between pairs of data point coordinates in each of the individual EIC files. Generating scan gaps comprises connecting a data point coordinate with the highest intensity in a file to the data point coordinate comprising the next highest intensity scan in that file and wherein each connection is defined as a scan gap. The method then comprises identifying blur scan gaps, wherein each blur scan gap is a scan gap that links two grid boxes in the merged data points coordinates file that are not linked in any of the individual EICs. In other words, blur scan gaps identify mutually exclusive pairs of grid boxes. Coordinates of data points in the mutually exclusive pairs of grid boxes identify EICs that comprise retention time profiles that exhibit nonlinear drift. Using this method, the task of identifying mutually exclusive pairs of coordinates is greatly simplified when compared to the task of identifying all possiblepairs of coordinates in files and comparing that large number of pairs of coordinates between template files and EIC files.
[0135] FIG. 13 illustrates the concept of scan gaps and blur scan gaps. FIG. 13 depicts an example of merged retention time profiles from a specific m / z value mapped onto a 2D grid of scans and strata and a diagrammatical representation of a portion of a 2D grid comprising four grid boxes. The grid boxes all comprise the same stratum (X) but four different scans. FIG. 13 illustrates an embodiment of the instant disclosure where scan gaps between data point coordinates are used to comparatively analyze mutually exclusive pairs of coordinates. In FIG. 13, coordinates of data points of two individual EICs (red and blue dots) and scan gaps generated between the coordinates of the individual EICs and scan gaps generated between the coordinates in the master coordinate pairs file are shown as colored arrows. Red and blue arrows are scan gaps that are present in the individual EICs as well as the master coordinate pairs file, whereas yellow and green arrows are scan gaps that can only be generated in the master coordinate pairs file. However, only the green arrows are pairs of coordinates that link grid boxes in the merged data point coordinates file that are not linked in any of the individual EICs. In other words, only the green arrows identify blur scan gaps.
[0136] As depicted in FIG. 14, when the retention time profiles that exhibit nonlinear drift are identified, the non-linear drift in each EIC can be corrected, and data points belonging to each compound and its isomers can be added to a corrected merged EIC file comprising all EIC files that were identified as described in Sections (b), (c), and (d) herein above. Methods of correcting non-linear drift can be performed using methods know in the art and include a combination of data processing techniques aimed at accurately defining the start, apex, and end of a peak representing a compound or its isomers, thereby identifying the data points belonging to them. Nonlimiting example of commonly used methods to identify the data points that belong to a compound or its isomer include baseline detection and subtraction, peak picking algorithms, deconvolution, thresholding, chromatographic peak integration, cluster analysis, and manual inspection and adjustment. Baseline detection and subtraction involves identifying the baseline noise level of the chromatogram and subtracting it fromthe signal to enhance the visibility of the peaks. Accurate baseline detection helps in distinguishing between actual signal peaks and background noise, ensuring that only genuine data points are attributed to each peak. Peak Picking Algorithms analyze the intensity and RT data to identify local maxima (peak apexes) and the inflection points that signify the start and end of a peak. Common algorithms include the Savitzky-Golay filter for smoothing the data while preserving peak shapes, and the Gaussian model to fit the data points around each peak, identifying the boundaries of the peak. Deconvolution techniques mathematically separate overlapping signals into individual components, allowing for the accurate assignment of data points to each peak based on their unique RT and intensity profiles. These techniques are particularly useful in resolving overlapping peaks, which is critical when two peaks are close to each other. Thresholding involves setting intensity thresholds to help distinguish peaks from noise. Data points with intensities above a certain threshold can be considered part of a peak, while those below the threshold are treated as background noise. This method requires careful selection of the threshold to avoid excluding relevant data points or including noise. Chromatographic peak integration involves calculating the area under the curve for each peak, which requires defining the start and end points of the peak. Integration tools often provide visual cues or markers that can be manually adjusted to accurately capture the boundaries of each peak, ensuring that all relevant data points are included. Cluster analysis groups data points based on their RT and intensity similarities. This method can help in segregating data points into distinct peaks, especially in complex EICs where peak boundaries might not be clear-cut. Manual inspection and adjustment by experienced analysts can often be necessary to confirm the peak boundaries and ensure that data points are correctly assigned. This approach can be particularly important in complex analyses where the automated methods may struggle to accurately define peak boundaries.
[0137] It will be recognized that, if any one of EICs associated with a data point template is found to exhibit non-linear drift, it can be reasonably assumed that all EICs related to this template likely demonstrate a similar pattern of non-linear drift. As a result, the correction of non-linear drift for all EICs associated with the second templatecan be conducted collectively. This method eliminates the need for individual analysis of each EIC within this group. Such a streamlined approach can lead to a significant reduction in the computational resources and time required for the drift correction process, enhancing the efficiency of the overall analysis.(e) Embodiments
[0138] Methods of the instant disclosure comprise methods for overcoming drift in retention times of MS mass features for the amplification and detection of MS signals.In some embodiments, the methods comprise obtaining or having obtained one or more data files from a chromatography-MS machine that has analyzed samples of unknown compounds. In some embodiments, the chromatography-MS is LC-MS. In some embodiments, methods of the instant disclosure can be used to overcome drift in retention times of chromatography-MS runs for analyses of metabolites. In some embodiments, methods of the instant disclosure can be used to overcome drift in retention times of LC-MS runs for analyses of metabolites.
[0139] In some embodiments, methods of the instant disclosure comprise defining a plurality of data point templates, wherein each data point template corresponds to one of a plurality of waveforms found in the plurality of retention time profiles in EICs from a specific m / z value. To define the data point templates, the methods comprise mapping all the data point values of all the EICs from that specific m / z value onto a 2D grid system. The 2D grid system comprises a plurality of grid boxes defined by the intersection of a specific scan and a specific stratum. In some embodiments, methods of the instant disclosure comprise mapping the plurality of retention time profiles onto a 2D cartesian grid system. In some embodiments, the scans comprise retention time increments of equal durations. In some embodiments, the strata comprise an equal number of data points.
[0010] In some embodiments, a method of the instant disclosure comprises defining a data point template for the most prevalent waveform among the retention time profiles in a merged EIC. In some embodiments, a data point template is definedby first selecting a reference scan. In some embodiments, a reference scan is selected as a scan comprising the highest intensity data points of the highest number of EICs among the plurality of EICs.
[0011] In some embodiments, defining a data point template further comprises identifying among the EIC files identified by a reference scan, one or more EICs comprising retention time profiles that exhibit non-linear drift, and excluding the EICs that introduce non-linear drift before the data point values of the identified EICs are merged to construct the data point template. According to embodiments of the instant disclosure, data points of the identified EIC are removed from, or not added to the template. Alternatively, non-linear drift can be corrected in each identified EIC file, and data points of the EIC file added to the template.
[0142] In some embodiments, methods of the instant disclosure comprise iteratively defining multiple data point templates in a merged EIC of a particular m / z value, then defining a second data point template among EICs remaining after the first template has been identified. In some embodiments, iteratively defining multiple data point templates in a merged EIC of a particular m / z value comprises defining a first data point template followed by identifying and correcting linear drift among the remaining EICs, before iteratively defining additional templates among the remaining EICs.
[0143] After defining one or more data point templates, methods of the instant disclosure comprise identifying EICs comprising retention time profiles exhibiting linear drift when compared to a data point template and correcting the linear drift. To identify the EICs that comprise retention time profiles that exhibit linear drift in relation to a data point template, the methods can comprise generating a query EIC for each remaining EIC of the plurality of EICs by shifting the data points of the query EIC by a time increment sufficient to overlap the highest data point of the query EIC with the highest data point of the data point template. A retention time profile is considered to exhibit linear drift relative to the data point template if, after the shift, coordinates of all the data points of the EIC lie in the data point template. In some embodiments, when EIC files comprising linear drift are identified, the linear drift is corrected by adding coordinates of all shifted data points to the data point template to generate a corrected template file.
[0144] After deconvoluting the linear drift by defining templates and correcting for linear drift, methods of the instant disclosure can comprise identifying EIC files comprising retention time profiles that exhibit nonlinear drift and correcting the nonlinear drift. In some embodiments, EIC files comprising retention time profiles that exhibit nonlinear drift are identified among EICs that match a template. In some embodiments, EIC files comprising retention time profiles that exhibit nonlinear drift are identified among EICs that match a template and data points of EIC files after correcting for drift.
[0145] In some embodiments, methods according to the instant disclosure comprise performing a comparative analysis of all pairs of template grid boxes and all pairs of EIC grid boxes. In some embodiments, a method the instant disclosure comprises (a) constructing a template box pairs file and constructing a EIC box pairs file for each EIC file determined to match the template; (b) constructing a combined EIC box pairs file by combining EIC box pairs files constructed for each EIC file determined to match the template; (c) identifying mutually exclusive pairs of grid boxes, wherein mutually exclusive pairs of grid boxes are pairs of grid boxes that are present in the template box pairs file but are absent in the EIC box pairs file or combined EIC box pairs file; and (d) identifying EICs that comprise retention time profiles that exhibit nonlinear drift, wherein a data point in a grid box of a box pair determined to be mutually exclusive identifies an EIC that comprise retention time profiles that exhibit non-linear drift. In some embodiments, each pair of grid boxes is represented by a unique identifier to facilitate the comparative analysis.
[0146] In other embodiments, methods of the instant disclosure comprise (a) merging all coordinates of two or more data point templates or corrected templates to generate a master coordinates file; (b) generating a master coordinate pairs file comprising all possible pairs of coordinates in the master coordinates file; (c) identifying all possible pairs of coordinates within each of the individual EIC files represented in the two or more data point templates or corrected data point templates; (d) generating a child coordinate pairs file comprising all possible pairs of coordinates within each of the individual EIC files represented in the two or more data point templates or correcteddata point templates; (e) identifying mutually exclusive pairs of grid coordinates; (f) identifying mutually exclusive pairs of grid boxes, wherein mutually exclusive pairs of coordinates identify mutually exclusive pairs of grid boxes; and (g) identifying EICs that comprise retention time profiles that exhibit non-linear drift, wherein data points in mutually exclusive pairs of coordinates or mutually exclusive pairs of grid boxes identify EICs that comprise retention time profiles that exhibit non-linear drift.
[0017] In some embodiments, identifying mutually exclusive pairs of coordinates further comprises alleviating the computational challenge in methods of the instant disclosure. In some embodiments, alleviating computational challenge comprises generating scan gaps between pairs of data point coordinates in the template file and generating scan gaps between pairs of data point coordinates in each of the individual EIC files. In some embodiments, identifying scan gaps between grid boxes comprises generating scan gaps between data point coordinates of individual EICs and generating scan gaps between data points in a merged template EIC file or merged corrected EIC file and identifying blur scan gaps, wherein each blur scan gap is a scan gap that links two mutually exclusive grid boxes in the merged data points coordinates file that are not linked in any of the individual EICs, wherein data points in the mutually exclusive grid boxes identify EICs that comprise retention time profiles that exhibit non-linear drift. Using this method, the task of identifying mutually exclusive pairs of coordinates is greatly simplified when compared to the task of identifying all possible pairs of coordinates in files and comparing that large number of pairs of coordinates between template files and EIC files.
[0018] FIGs. 15-17 are flowcharts illustrating various steps of methods of the instant disclosure to overcome drift in retention times of compound signals in MS. In some embodiments, compounds can be metabolites and isomers of metabolites. In some embodiments, the chromatography-MS is liquid chromatography-MS (LC-MS). FIG. 15 is a flowchart illustrating an exemplary method 500 for defining a template in an embodiment of the methods of the instant disclosure. In step 502, EIC files of portions of MS chromatograms bearing the same m / z value can be obtained. Method 500 can be repeated iteratively to identify a number of templates according to someembodiments of the instant disclosure. In step 504, data point values of the EICs can be mapped onto a two-dimensional (2D) grid system comprising coordinates of a plurality of scans comprising retention time increments and a plurality of signal intensity strata. In some embodiments, each stratum comprises an equal number of data points and the plurality of scans comprise equal retention time increments. In step 506, a first reference scan can be selected as a scan comprising the highest intensity data points of the highest number of EICs among the plurality of EICs. In step 508, a first data point template can be defined as all grid boxes in the grid comprising at least one data point from at least one EIC associated with the reference scan identified in Step 506.Defining a data point template can further comprise identifying one or more EICs that can introduce non-linear blur among the identified EICs of a template, and excluding the EICs that introduce non-linear blur before the data values of the identified EICs are merged to construct the data point template. FIG. 16 is a flowchart illustrating an exemplary method 600 for identifying and correcting linear drift in an embodiment of the methods of the instant disclosure. As shown in Step 602, identifying linear drift can comprise generating a query EIC for each remaining EIC of the plurality of EICs, by shifting the data points of the query EIC by a time increment sufficient to overlap the highest data point of the query EIC with the highest data point of the first template. A query EIC can be defined as an EIC that exhibits linear drift in relation to the first template if coordinates of all the data points of the EIC lie in the first template (Step 604). In step 606, linear drift can be corrected by adding coordinates of all shifted data points of each query EIC to the first data point template. In some embodiments, method 600 can be iteratively repeated to identify additional templates using remaining EICs that were not added to the first data point template and to correct the linear drift of the EICs based on the identified additional templates. FIG. 17 is a flowchart illustrating an exemplary method 700 for identifying non-linear drift in an embodiment of the methods of the instant disclosure. In Step 702, all data point coordinates of all EIC files in identified templates can be merged to generate a master coordinates file. In some embodiments, a master coordinates file can also include data point coordinates of EIC files corrected for linear drift. In Step 704, a master coordinate pairs file can begenerated comprising all pairs of coordinates in the master coordinates file. In Step 706, a child coordinate pairs file can be generated comprising all pairs of coordinates within each of the individual EIC files. In step 708, mutually exclusive pairs of coordinates can be identified. Mutually exclusive pairs of coordinates are pairs of coordinates that occur in the master coordinate pairs file but not in the child coordinate pairs file. Further, coordinates of data points in mutually exclusive pairs of coordinates identify EICs that comprise retention time profiles that exhibit nonlinear drift.
[0149] The methods 500, 600, 700, or any combination thereof of FIGs. 15-17 can be embodied as executable instructions in a non-transitory computer readable storage medium including but not limited to a CD, DVD, or non-volatile memory such as a hard drive. The instructions of the storage medium may be executed by a processor (or processors) to cause various hardware components of a computing device hosting or otherwise accessing the storage medium to effectuate the method. The steps identified in FIGs. 15-17 (and the order thereof) are exemplary and may include various alternatives, equivalents, or derivations thereof including but not limited to the order of execution of the same.
[0150] In some embodiments, a method for overcoming drift in retention times of compound signals in MS comprises: (a) mapping data point values of extracted ion chromatograms (EICs) of a portion of each of a plurality of MS chromatograms onto a two-dimensional (2D) grid system comprising grid boxes of a plurality of scans comprising retention time increments and a plurality of signal intensity strata, wherein the portions of the plurality of chromatograms bear the same m / z value; (b) selecting a first reference scan as a scan comprising the highest intensity data points of the highest number of EICs among the plurality of EICs; (c) defining a first data point template as all grid boxes in the grid comprising at least one data point from at least one EIC associated with the reference scan identified in (b); (d) for each remaining EIC of the plurality of EICs, generating a query EIC by shifting the data points of the query EIC by a time increment sufficient to overlap the highest data point of the query EIC with the highest data point of the first template, wherein a query EIC exhibits linear drift in relation to the first template if coordinates of all the data points of the EIC lie in the firsttemplate; (e) correcting linear drift by adding coordinates of all shifted data points of each query EIC to the first data point template to thereby generate a corrected template; (f) iteratively repeating steps (b) to (e) using remaining EICs that were not added to the first data point template to define additional data point templates and correct the linear drift of the EICs based on the identified additional templates; and (g) identifying, among the EIC files identified in steps (a) to (f), EIC files comprising retention time profiles that exhibit non-linear drift in retention time by: (i) merging all data point coordinates of all the EIC files in the identified templates or corrected templates to generate a master coordinates file; (ii) generating a master coordinate pairs file comprising all pairs of coordinates in the master coordinates file; (iii) generating a child coordinate pairs file comprising all pairs of coordinates within each of the individual EIC files; and (iv) identifying mutually exclusive pairs of coordinates, wherein mutually exclusive pairs of coordinates are pairs of coordinates that occur in the master coordinate pairs file but not in the child coordinate pairs file, wherein coordinates of data points in mutually exclusive pairs of coordinates identify EICs that comprise retention time profiles that exhibit nonlinear drift. In some embodiments, the compound is metabolites and isomers of metabolites. In some embodiments, the chromatography-MS is liquid chromatography- MS (LC-MS). In some embodiments, defining a data point template further comprises identifying one or more EICs that introduce non-linear blur among the identified EICs of a template, and excluding the EICs that introduce non-linear blur before the data values of the identified EICs are merged to construct the data point template. Each stratum can comprise an equal number of data points and the plurality of scans comprise equal retention time increments. In some embodiments, steps (b) to (e) are iteratively repeated until none of the remaining EICs are added to any of the identified templates.
[0151] In some embodiments, identifying EIC files comprising retention time profiles that exhibit non-linear drift comprises: (a) constructing a template box pairs file and constructing a EIC box pairs file for each EIC file determined to match the template; (b) constructing a combined EIC box pairs file by combining EIC box pairs files constructed for each EIC file determined to match the template; (c) identifying mutually exclusive pairs of grid boxes, wherein mutually exclusive pairs of grid boxes are pairs ofgrid boxes that are present in the template box pairs file but are absent in the EIC box pairs file or combined EIC box pairs file; and (d) identifying EICs that comprise retention time profiles that exhibit non-linear drift, wherein a data point in a grid box of a box pair determined to be mutually exclusive identifies an EIC that comprise retention time profiles that exhibit non-linear drift. In some embodiments, each pair of grid boxes is represented by a unique identifier to facilitate the comparative analysis.
[0152] In some embodiments, identifying EIC files comprising retention time profiles that exhibit non-linear drift comprises: (a) merging all coordinates of two or more data point templates or corrected templates to generate a master coordinates file; (b) generating a master coordinate pairs file comprising all possible pairs of coordinates in the master coordinates file; (c) identifying all possible pairs of coordinates within each of the individual EIC files represented in the two or more data point templates or corrected data point templates; (d) generating a child coordinate pairs file comprising all possible pairs of coordinates within each of the individual EIC files represented in the two or more data point templates or corrected data point templates; (e) identifying mutually exclusive pairs of data point coordinates, wherein mutually exclusive pairs of data point coordinates are pairs of coordinates that occur in the master template but not in each EIC file comprising at least one data point in each of the pairs of coordinates; (f) identifying mutually exclusive pairs of grid boxes, wherein mutually exclusive pairs of coordinates identify mutually exclusive pairs of grid boxes; and (g) identifying EICs that comprise retention time profiles that exhibit non-linear drift, wherein data points in mutually exclusive pairs of coordinates or mutually exclusive pairs of grid boxes identify EICs that comprise retention time profiles that exhibit non-linear drift.
[0153] In some embodiments, identifying mutually exclusive pairs of coordinates comprises: (a) generating scan gaps in the master coordinate file and generating scan gaps in each of the EICs comprising data points in the master coordinate file, wherein generating scan gaps comprises concatenating data points in each file by connecting a data point with the highest intensity to a data point comprising the next highest intensity and wherein each connection is defined as a scan gap; and (b) identifying blur scan gaps, wherein each blur scan gap is a scan gap that links two grid boxes in the mergeddata points coordinates file that are not linked in any of the individual EICs, and wherein coordinates of data points in the mutually exclusive pairs of grid boxes identify EICs that comprise retention time profiles that exhibit nonlinear drift.
[0154] Stringency can be adjusted by changing the width of the scan, the number of points in each stratum, by determining how many points match or don’t match the template. In some embodiments, defining a data point template further comprises identifying one or more EICs that comprise retention time profiles that introduce nonlinear blur among the identified EICs of the template and excluding the EICs that introduce non-linear blur before the data values of the identified EICs are merged to construct the data point template.II. Computing system
[0155] Another embodiment of the instant disclosure encompasses a computing system for overcoming drift in retention times of metabolite signals in MS. In some embodiments, the computing system is attached to and / or used in conjunction with a chromatography-MS system. The computing system can comprise several known components and circuitry, including a processor, a memory system, input and output devices and interfaces (e.g., an interconnection mechanism), as well as other components, such as transport circuitry (e.g., one or more busses), a video and audio data input / output (I / O) subsystem, special-purpose hardware, as well as other components and circuitry, as described below in more detail. Further, the computer system(s) can be a multi-processor computer system or may include multiple computers connected over a computer network.
[0156] A processor can include one or more general purpose computers, dedicated microprocessors, graphics processors, or other processing devices capable of communicating electronic information. Non-limiting examples of a processor include one or more application-specific integrated circuits (ASICs), graphical processing units (GPUs), field programmable gate arrays (FPGAs), graphics processing units (GPUs), digital signal processors (DSPs) and any other suitable specific or general purposeprocessors. The processor can be implemented as appropriate in hardware, firmware, or combinations thereof with computer-executable instructions and / or software. Computer-executable instructions and software can include computer-executable or machine-executable instructions written in any suitable programming language to perform the various functions described.
[0157] The memory can include more than one memory and can be distributed throughout the computing system. The memory can store program instructions that are loadable and executable on the processor(s) as well as data generated during the execution of these programs. Depending on the configuration and type of memory, the memory can be volatile (such as random access memory (RAM)) and / or non-volatile (such as read-only memory (ROM), flash memory, or other memory). In some embodiments, the memory can include multiple different types of memory, such as static random access memory (SRAM), dynamic random access memory (DRAM), or ROM.
[0158] In some embodiments, the computing system can also include additional storage, which can include removable storage and / or non-removable storage. The additional storage can include, but is not limited to, magnetic storage, optical disks, and / or solid-state storage. The disk drives and their associated computer-readable media can provide non-volatile storage of computer-readable instructions, data structures, program modules, and other data for the computing devices. The memory and the additional storage, both removable and non-removable, are examples of computer-readable storage media. For example, computer-readable storage media can include volatile or non-volatile, removable, or non-removable media implemented in any suitable method or technology for storage of information such as computer-readable instructions, data structures, program modules, or other data. As used herein, modules, engines, and components, can refer to programming modules executed by computing systems (e.g., processors) that are part of the architecture.
[0159] The processor generally manipulates the data within the integrated circuit memory element in accordance with the program instructions and then copies the manipulated data to the non-volatile recording medium after processing is completed. Avariety of mechanisms are known for managing data movement between the nonvolatile recording medium and the integrated circuit memory element, and the computing system that implements the methods, steps, systems control and system elements control described above is not limited thereto. The computing system is not limited to a particular memory system.
[0160] At least part of such a memory system described above can be used to store one or more data structures (e.g., look-up tables) or equations such as calibration curve equations. For example, at least part of the non-volatile recording medium can store at least part of a database that includes one or more of such data structures. Such a database can be any of a variety of types of databases, for example, a file system including one or more flat-file data structures where data is organized into data units separated by delimiters, a relational database where data is organized into data units stored in tables, an object-oriented database where data is organized into data units stored as objects, another type of database, or any combination thereof.
[0161] The computer implemented control system(s) can include one or more output devices. Non-limiting example output devices include a cathode ray tube (CRT) display, liquid crystal displays (LCD) and other video output devices, printers, communication devices such as a modem or network interface, storage devices such as disk or tape, and audio output devices such as a speaker.
[0162] The computing system also can include one or more input devices. Example input devices include a keyboard, keypad, track ball, mouse, pen and tablet, communication devices such as described above, and data input devices such as audio and video capture devices and sensors. The computing system is not limited to the particular input or output devices described herein.
[0163] It should be appreciated that one or more of any type of computing system can be used to implement various embodiments described herein. Aspects of the invention may be implemented in software, hardware or firmware, or any combination thereof. The computing system can include specially programmed, special purpose hardware, for example, an application-specific integrated circuit (ASIC). Such specialpurpose hardware can be configured to implement one or more of the methods, steps,simulations, algorithms, systems control, and system elements control described above as part of the computer implemented control system(s) described above or as an independent component.
[0164] The computing system and components thereof may be programmable using any of a variety of one or more suitable computer programming languages. Such languages may include procedural programming languages, for example, LabView, C, Pascal, Fortran and BASIC, object-oriented languages, for example, C++, Java and Eiffel and other languages, such as a scripting language or even assembly language.
[0165] The methods, steps, simulations, algorithms, systems control, and system elements control may be implemented using any of a variety of suitable programming languages, including procedural programming languages, object- oriented programming languages, other languages and combinations thereof, which may be executed by such a computer system. Such methods, steps, simulations, algorithms, systems control, and system elements control can be implemented as separate modules of a computer program, or can be implemented individually as separate computer programs. Such modules and programs can be executed on separate computers.
[0166] Such methods, steps, simulations, algorithms, systems control, and system elements control, either individually or in combination, may be implemented as a computer program product tangibly embodied as computer-readable signals on a computer-readable medium, for example, a non-volatile recording medium, an integrated circuit memory element, or a combination thereof. For each such method, step, simulation, algorithm, system control, or system element control, such a computer program product may comprise computer-readable signals tangibly embodied on the computer-readable medium that define instructions, for example, as part of one or more programs, that, as a result of being executed by a computer, instruct the computer to perform the method, step, simulation, algorithm, system control, or system element control.
[0167] For clarity of explanation, in some instances the present technology may be presented as including individual functional blocks including functional blockscomprising devices, device components, steps or routines in a method embodied in software, or combinations of hardware and software.
[0168] Any of the steps, operations, functions, or processes described herein may be performed or implemented by a combination of hardware and software services or services, alone or in combination with other devices. In some embodiments, a service can be software that resides in memory of a client device and / or one or more servers of a content management system and perform one or more functions when a processor executes the software associated with the service. In some embodiments, a service is a program, or a collection of programs that carry out a specific function. In some embodiments, a service can be considered a server. The memory can be a non- transitory computer-readable medium.
[0169] In some embodiments the computer-readable storage devices, mediums, and memories can include a cable or wireless signal containing a bit stream and the like. However, when mentioned, non-transitory computer-readable storage media expressly exclude media such as energy, carrier signals, electromagnetic waves, and signals per se.
[0170] Methods according to the above-described examples can be implemented using computer-executable instructions that are stored or otherwise available from computer readable media. Such instructions can comprise, for example, instructions and data which cause or otherwise configure a general purpose computer, special purpose computer, or special purpose processing device to perform a certain function or group of functions. Portions of computer resources used can be accessible over a network. The computer executable instructions may be, for example, binaries, intermediate format instructions such as assembly language, firmware, or source code. Examples of computer-readable media that may be used to store instructions, information used, and / or information created during methods according to described examples include magnetic or optical disks, solid state memory devices, flash memory, USB devices provided with non-volatile memory, networked storage devices, and so on.
[0171] Devices implementing methods according to these disclosures can comprise hardware, firmware and / or software, and can take any of a variety of formfactors. Typical examples of such form factors include servers, laptops, smart phones, small form factor personal computers, personal digital assistants, and so on. Functionality described herein also can be embodied in peripherals or add-in cards. Such functionality can also be implemented on a circuit board among different chips or different processes executing in a single device, by way of further example.
[0172] The instructions, media for conveying such instructions, computing resources for executing them, and other structures for supporting such computing resources are means for providing the functions described in these disclosures.
[0173] FIG. 19 shows an example of a computing system 1800 in which the components of the system are in communication with each other using connection 1805. Connection 1805 can be a physical connection via a bus, or a direct connection into processor 1810, such as in a chipset architecture. Connection 1805 can also be a virtual connection, networked connection, or logical connection.
[0174] In some embodiments computing system 1800 is a distributed system in which the functions described in this disclosure can be distributed within a datacenter, multiple datacenters, a peer network, etc. In some embodiments, one or more of the described system components represents many such components each performing some or all of the function for which the component is described. In some embodiments, the components can be physical or virtual devices.
[0175] Example system 1100 includes at least one processing unit (CPU or processor) 1810 and connection 1805 that couples various system components including system memory 1815, such as read only memory (ROM) and random access memory (RAM) to processor 1810.
[0176] Computing system 1800 can include a cache of high-speed memory connected directly with, in close proximity to, or integrated as part of processor 1810.
[0177] Processor 1810 can include any general purpose processor and a hardware service or software service, such as services 1832, 1834, and 1836 stored in storage device 1830, configured to control processor 1810 as well as a special-purpose processor where software instructions are incorporated into the actual processor design. Processor 1810 may essentially be a completely self-contained computingsystem, containing multiple cores or processors, a bus, memory controller, cache, etc. A multi-core processor may be symmetric or asymmetric.
[0178] To enable user interaction, computing system 1800 includes an input device 1845, which can represent any number of input mechanisms, such as a microphone for speech, a touch-sensitive screen for gesture or graphical input, keyboard, mouse, motion input, speech, etc. Computing system 1800 can also include output device 1835, which can be one or more of a number of output mechanisms known to those of skill in the art. In some instances, multimodal systems can enable a user to provide multiple types of input / output to communicate with computing system 1800. Computing system 1800 can include communications interface 1840, which can generally govern and manage the user input and system output. There is no restriction on operating on any particular hardware arrangement and therefore the basic features here may easily be substituted for improved hardware or firmware arrangements as they are developed.
[0179] Storage device 1830 can be a non-volatile memory device and can be a hard disk or other types of computer readable media which can store data that are accessible by a computer, such as magnetic cassettes, flash memory cards, solid state memory devices, digital versatile disks, cartridges, random access memories (RAMs), read only memory (ROM), and / or some combination of these devices.
[0180] The storage device 1830 can include software services, servers, services, etc., that when the code that defines such software is executed by the processor 1810, it causes the system to perform a function. In some embodiments, a hardware service that performs a particular function can include the software component stored in a computer-readable medium in connection with the necessary hardware components, such as processor 1810, connection 1805, output device 1835, etc., to carry out the function.
[0181] In some embodiments, a computing system of the instant disclosure comprises a communication interface and a processor. The communication interface can receive EICs of portions of each of a plurality of MS chromatograms, wherein the portions of the plurality of chromatograms bear the same m / z value. The processor canexecute instructions stored in memory. In some embodiments, the processor executes the instructions to (a) map data point values of the EICs onto a 2D grid system comprising grid boxes of a plurality of scans comprising retention time increments and a plurality of signal intensity strata; (b) select a first reference scan as a scan comprising the highest intensity data points of the highest number of EICs among the plurality of EICs; (c) define a first data point template as all coordinates in the grid comprising at least one data point from at least one EIC identified in step (b); (d) for each remaining EIC of the plurality of EICs, generate a query EIC by shifting the data points of the query EIC by a time increment sufficient to overlap the highest data point of the query EIC with the highest data point of the first template, wherein a query EIC exhibits linear drift in relation to the first template if coordinates of all the data points of the EIC lie in the first template; (e) correct linear drift by adding coordinates of all shifted data points of each query EIC to the first data point template to thereby generate a corrected template; (f) iteratively repeat steps (b) to (e) using EIC files that were not added to the first data point template to define additional data point templates and correct the EICs for linear drift based on the identified additional templates; and identify, among the EIC files identified in steps (a) to (f), EIC files comprising retention time profiles that exhibit nonlinear drift in retention time by: (i) merging all data point coordinates of all the EIC files in the identified templates or corrected templates to generate a master coordinates file; (ii) generating a master coordinate pairs file comprising all pairs of coordinates in the master coordinates file; (iii) generating a child coordinate pairs file comprising all pairs of coordinates within each of the individual EIC files; and (iv) identifying mutually exclusive pairs of coordinates, wherein mutually exclusive pairs of coordinates are pairs of coordinates that occur in the master coordinate pairs file but not in the child coordinate pairs file, wherein coordinates of data points in mutually exclusive pairs of coordinates identify EICs that comprise retention time profiles that exhibit nonlinear drift.
[0182] In some embodiments, the processor can further execute the instructions to define a data point template by identifying one or more EICs that can introduce nonlinear blur among the identified EICs of a template, and excluding the EICs that introduce non-linear blur before the data values of the identified EICs are merged toconstruct the data point template. Each stratum can comprise an equal number of data points and the plurality of scans comprise equal retention time increments.
[0183] In some embodiments, the processor can further execute the instructions to identify, among the EIC files identified in steps (a) to (f), EIC files comprising retention time profiles that exhibit nonlinear drift in retention time by further: (1 ) constructing a template box pairs file and constructing a EIC box pairs file for each EIC file determined to match the template; (2) constructing a combined EIC box pairs file by combining EIC box pairs files constructed for each EIC file determined to match the template; (3) identifying mutually exclusive pairs of grid boxes, wherein mutually exclusive pairs of grid boxes are pairs of grid boxes that are present in the template box pairs file but are absent in the EIC box pairs file or combined EIC box pairs file; and (4) identifying EICs that comprise retention time profiles that exhibit non-linear drift, wherein a data point in a grid box of a box pair determined to be mutually exclusive identifies an EIC that comprise retention time profiles that exhibit non-linear drift. In some embodiments, the processor further executes the instructions to assign a unique identifier to each pair of grid boxes to facilitate the comparative analysis.
[0184] In some embodiments, the processor further executes the instructions to identify, among the EIC files identified in steps (a) to (f), EIC files comprising retention time profiles that exhibit nonlinear drift in retention time by further: (1 ) merging all coordinates of two or more data point templates or corrected templates to generate a master coordinates file; (2) generating a master coordinate pairs file comprising all possible pairs of coordinates in the master coordinates file; (3) identifying all possible pairs of coordinates within each of the individual EIC files represented in the two or more data point templates or corrected data point templates; (4) generating a child coordinate pairs file comprising all possible pairs of coordinates within each of the individual EIC files represented in the two or more data point templates or corrected data point templates; (5) identifying mutually exclusive pairs of data point coordinates, wherein mutually exclusive pairs of data point coordinates are pairs of coordinates that occur in the master template but not in each EIC file comprising at least one data point in each of the pairs of coordinates; (6) identifying mutually exclusive pairs of grid boxes,wherein mutually exclusive pairs of coordinates identify mutually exclusive pairs of grid boxes; and (7) identifying EICs that comprise retention time profiles that exhibit nonlinear drift, wherein data points in mutually exclusive pairs of coordinates or mutually exclusive pairs of grid boxes identify EICs that comprise retention time profiles that exhibit non-linear drift.
[0185] In some embodiments, the processor further executes the instructions to identify, among the EIC files identified in steps (a) to (f), EIC files comprising retention time profiles that exhibit nonlinear drift in retention time by further: (1 ) generating scan gaps in the master coordinate file and generating scan gaps in each of the EICs comprising data points in the master coordinate file, wherein generating scan gaps comprises concatenating data points in each file by connecting a data point with the highest intensity to a data point comprising the next highest intensity and wherein each connection is defined as a scan gap; and (2) identifying blur scan gaps, wherein each blur scan gap is a scan gap that links two grid boxes in the merged data points coordinates file that are not linked in any of the individual EICs, and wherein coordinates of data points in the mutually exclusive pairs of grid boxes identify EICs that comprise retention time profiles that exhibit nonlinear drift.
[0186] In some embodiments, the processor further executes the instructions to iteratively repeat steps (a) to (d) until none of the remaining EICs can be added to any of the identified templates. In some embodiments, wherein stringency can be adjusted by changing the width of the scan, the number of points in each stratum, by determining how many points match or don’t match the template. In some embodiments, the processor further executes the instructions to define a data point template by further identifying one or more EICs that comprise retention time profiles that introduce nonlinear blur among the identified EICs of the template and excluding the EICs that introduce non-linear blur before the data values of the identified EICs are merged to construct the data point template.
[0187] An additional embodiment of the instant disclosure encompasses a non- transitory, computer-readable storage medium, having embodied thereon a program executable by a processor to perform a method for identifying metabolite derivatives inchromatography- MS data. The method comprises (a) receiving or having received an MS data file from a chromatography-MS instrument, wherein the data file comprises data points obtained from a sample analyzed by the chromatography-mass spectrometry instrument, and wherein each data point in the file comprises a retention time value and a mass-to-charge (m / z) value; (b) mapping data point values of the EICs onto a 2D grid system comprising grid boxes of a plurality of scans comprising retention time increments and a plurality of signal intensity strata; (c) selecting a first reference scan as a scan comprising the highest intensity data points of the highest number of EICs among the plurality of EICs; (d) defining a first data point template as all grid boxes in the grid comprising at least one data point from at least one EIC associated with the reference scan identified in (b); (e) for each remaining EIC of the plurality of EICs, generating a query EIC by shifting the data points of the query EIC by a time increment sufficient to overlap the highest data point of the query EIC with the highest data point of the first template, wherein a query EIC exhibits linear drift in relation to the first template if coordinates of all the data points of the EIC lie in the first template; (f) correcting linear drift by adding coordinates of all shifted data points of each query EIC to the first data point template to thereby generate a corrected template; (g) iteratively repeating steps (b) to (e) using remaining EICs that were not added to the first data point template to define additional data point templates and correct the linear drift of the EICs based on the identified additional templates; and (h) identifying, among the EIC files identified in steps (a) to (f), EIC files comprising retention time profiles that exhibit non-linear drift in retention time by: (i) merging all data point coordinates of all the EIC files in the identified templates or corrected templates to generate a master coordinates file; (ii) generating a master coordinate pairs file comprising all pairs of coordinates in the master coordinates file; (iii) generating a child coordinate pairs file comprising all pairs of coordinates within each of the individual EIC files; and (iv) identifying mutually exclusive pairs of coordinates, wherein mutually exclusive pairs of coordinates are pairs of coordinates that occur in the master coordinate pairs file but not in the child coordinate pairs file, wherein coordinates of data points in mutually exclusive pairs of coordinates identify EICs that comprise retention time profiles that exhibit nonlinear drift.In some embodiments, each stratum comprises an equal number of data points and the plurality of scans comprise equal retention time increments.
[0188] In some embodiments, the method further comprises defining a data point template by identifying one or more EICs that can introduce non-linear blur among the identified EICs of a template, and excluding the EICs that introduce non-linear blur before the data values of the identified EICs are merged to construct the data point template.
[0189] In some embodiments, the method further comprises identifying, among the EIC files identified in steps (c) to (h), EIC files comprising retention time profiles that exhibit nonlinear drift in retention time by further: (a) constructing a template box pairs file and constructing a EIC box pairs file for each EIC file determined to match the template; (b) constructing a combined EIC box pairs file by combining EIC box pairs files constructed for each EIC file determined to match the template; (c) identifying mutually exclusive pairs of grid boxes, wherein mutually exclusive pairs of grid boxes are pairs of grid boxes that are present in the template box pairs file but are absent in the EIC box pairs file or combined EIC box pairs file; and (d) identifying EICs that comprise retention time profiles that exhibit non-linear drift, wherein a data point in a grid box of a box pair determined to be mutually exclusive identifies an EIC that comprise retention time profiles that exhibit non-linear drift. In some embodiments, the method further comprises assigning a unique identifier to each pair of grid boxes to facilitate the comparative analysis.
[0190] In some embodiments, the method further comprises identifying, among the EIC files identified in steps (c) to (h), EIC files comprising retention time profiles that exhibit nonlinear drift in retention time by further: (i) merging all coordinates of two or more data point templates or corrected templates to generate a master coordinates file; (ii) generating a master coordinate pairs file comprising all possible pairs of coordinates in the master coordinates file; (iii) identifying all possible pairs of coordinates within each of the individual EIC files represented in the two or more data point templates or corrected data point templates; (iv) generating a child coordinate pairs file comprising all possible pairs of coordinates within each of the individual EIC files represented in thetwo or more data point templates or corrected data point templates; (v) identifying mutually exclusive pairs of data point coordinates, wherein mutually exclusive pairs of data point coordinates are pairs of coordinates that occur in the master template but not in each EIC file comprising at least one data point in each of the pairs of coordinates; (vi) identifying mutually exclusive pairs of grid boxes, wherein mutually exclusive pairs of coordinates identify mutually exclusive pairs of grid boxes; and (vii) identifying EICs that comprise retention time profiles that exhibit non-linear drift, wherein data points in mutually exclusive pairs of coordinates or mutually exclusive pairs of grid boxes identify EICs that comprise retention time profiles that exhibit non-linear drift.
[0191] In some embodiments, the method further comprises identifying, among the EIC files identified in steps (c) to (h), EIC files comprising retention time profiles that exhibit nonlinear drift in retention time by further: (i) generating scan gaps in the master coordinate file and generating scan gaps in each of the EICs comprising data points in the master coordinate file, wherein generating scan gaps comprises concatenating data points in each file by connecting a data point with the highest intensity to a data point comprising the next highest intensity and wherein each connection is defined as a scan gap; and (ii) identifying blur scan gaps, wherein each blur scan gap is a scan gap that links two grid boxes in the merged data points coordinates file that are not linked in any of the individual EICs, and wherein coordinates of data points in the mutually exclusive pairs of grid boxes identify EICs that comprise retention time profiles that exhibit nonlinear drift.
[0192] In some embodiments, the method further comprises iteratively repeating steps (b) to (e) until none of the remaining EICs can be added to any of the identified templates. In some embodiments, stringency can be adjusted by changing the width of the scan, the number of points in each stratum, by determining how many points match or don’t match the template. In some embodiments, the method further comprises defining a data point template by further identifying one or more EICs that comprise retention time profiles that introduce non-linear blur among the identified EICs of the template and excluding the EICs that introduce non-linear blur before the data values of the identified EICs are merged to construct the data point template.
[0193] Although a variety of examples and other information was used to explain aspects within the scope of the appended claims, no limitation of the claims should be implied based on particular features or arrangements in such examples, as one of ordinary skill would be able to use these examples to derive a wide variety of implementations. Further and although some subject matter may have been described in language specific to examples of structural features and / or method steps, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to these described features or acts. For example, such functionality can be distributed differently or performed in components other than those identified herein. Rather, the described features and steps are disclosed as examples of components of systems and methods within the scope of the appended claims.DEFINITIONS
[0194] Unless defined otherwise, all technical and scientific terms used herein have the meaning commonly understood by a person skilled in the art to which this invention belongs. The following references provide one of skill with a general definition of many of the terms used in this invention: Singleton et al., Dictionary of Microbiology and Molecular Biology (2nd ed. 1994); The Cambridge Dictionary of Science and Technology (Walker ed., 1988); The Glossary of Genetics, 5th Ed., R. Rieger et al. (eds.), Springer Verlag (1991 ); and Hale & Marham, The Harper Collins Dictionary of Biology (1991 ). As used herein, the following terms have the meanings ascribed to them unless specified otherwise.
[0195] When introducing elements of the present disclosure or the preferred aspects(s) thereof, the articles "a", "an", "the" and "said" are intended to mean that there are one or more of the elements. The terms "comprising", "including" and "having" are intended to be inclusive and mean that there may be additional elements other than the listed elements.
[0196] As used herein, the term “mass feature” when used in the context of mass chromatography refers to a concentration of data points that exhibit a Gaussian distribution, indicating the possible presence of a compound, an isomer, or a derivative.
[0197] As used herein, the term “waveform” refers to a retention time profile in an EIC for a specific m / z ratio, wherein the retention time profile can comprise one or more mass features. Waveforms are considered distinct if each comprises a different number of mass features, wherein each mass feature indicates the possible presence of a compound, an isomer, or a derivative.
[0198] A “waveform variant” refers to a variant of a waveform that shares the same mass features in number and type but shows variation in the intensity of one or more mass features.
[0199] As used herein, the term “remaining EIC” refers to any EIC in a merged EIC that either has not been identified to comprise a retention time profile of a specific waveform / tem plate or has not undergone any adjustments to correct for drift.
[0200] As used herein, the term “merged template EIC file” refers to a data file comprising all data points of EICs identified to match a template.
[0201] As used herein, the term “merged corrected EIC file” refers to a data file comprising all data points of EICs identified to match a template and data points of EIC files after correcting for drift.
[0202] As used herein, the term “template” refers to the grid boxes or pattern of that contain data points from EIC files determined to match a template.
[0203] As used herein, the term “corrected template” refers to the grid boxes or pattern of the grid boxes that contain data points from EIC files determined to match a template as well as data points of EIC files after correcting for drift.
[0204] As used herein, the term “template grid box file” refers to a file comprising all the grid boxes comprising at least one data point from one or more EICs comprising a retention time profiles that corresponds to a particular waveform and variants of that waveform.
[0205] As used herein, the term “EIC grid box file” refers to a file comprising all the grid boxes comprising at least one data point from an individual EIC determined to match a template.
[0206] As used herein, the term “template box pairs file” is a file noting all possible pairs of grid boxes that can be obtained when template grid boxes are paired in all possible combinations.
[0207] As used herein, the term “combined EIC box pairs file” is a file noting a combination of EIC box pairs for each EIC file determined to match a template. EIC box pairs of an EIC file are obtained by pairing all grid boxes of the EIC in all possible combinations.
[0208] As various changes could be made in the above-described cells and methods without departing from the scope of the invention, it is intended that all matter contained in the above description and in the examples given below, shall be interpreted as illustrative and not in a limiting sense.
Claims
CLAIMSWhat is claimed is:1 . A method for overcoming drift in retention times of compound signals in mass spectrometry (MS), the method comprising: a. mapping data point values of extracted ion chromatograms (EICs) of a portion of each of a plurality of MS chromatograms onto a two-dimensional (2D) grid system comprising grid boxes of a plurality of scans comprising retention time increments and a plurality of signal intensity strata, wherein the portions of the plurality of chromatograms bear the same m / z value; b. selecting a first reference scan as a scan comprising the highest intensity data points of the highest number of EICs among the plurality of EICs; c. defining a first data point template as all grid boxes in the grid comprising at least one data point from at least one EIC associated with the reference scan identified in (b); d. for each remaining EIC of the plurality of EICs, generating a query EIC by shifting the data points of the query EIC by a time increment sufficient to overlap the highest data point of the query EIC with the highest data point of the first template, wherein a query EIC exhibits linear drift in relation to the first template if coordinates of all the data points of the EIC lie in the first template; e. correcting linear drift by adding coordinates of all shifted data points of each query EIC to the first data point template to thereby generate a corrected template; f. iteratively repeating steps (b) to (e) using remaining EICs that were not added to the first data point template to define additional data point templates and correct the linear drift of the EICs based on the identified additional templates; and g. identifying, among the EIC files identified in steps (a) to (f), EIC files comprising retention time profiles that exhibit non-linear drift in retention time by:i. merging all data point coordinates of all the EIC files in the identified templates or corrected templates to generate a master coordinates file; ii. generating a master coordinate pairs file comprising all pairs of coordinates in the master coordinates file; iii. generating a child coordinate pairs file comprising all pairs of coordinates within each of the individual EIC files; and iv. identifying mutually exclusive pairs of coordinates, wherein mutually exclusive pairs of coordinates are pairs of coordinates that occur in the master coordinate pairs file but not in the child coordinate pairs file, wherein coordinates of data points in mutually exclusive pairs of coordinates identify EICs that comprise retention time profiles that exhibit nonlinear drift.
2. The method of claim 1 , wherein the compound is metabolites and isomers of metabolites.
3. The method of any one of either claim 1 or claim 2, wherein the chromatography-MS is liquid chromatography-MS (LC-MS).
4. The method of claim any one of the preceding claims, wherein defining a data point template further comprises identifying one or more EICs that introduce non-linear blur among the identified EICs of a template, and excluding the EICs that introduce non-linear blur before the data values of the identified EICs are merged to construct the data point template.
5. The method of any one of the preceding claims, wherein each stratum comprises an equal number of data points and the plurality of scans comprise equal retention time increments.
6. The method of any one of the preceding claims, wherein identifying EIC files comprising retention time profiles that exhibit non-linear drift comprises:a. constructing a template box pairs file and constructing a EIC box pairs file for each EIC file determined to match the template; b. constructing a combined EIC box pairs file by combining EIC box pairs files constructed for each EIC file determined to match the template; c. identifying mutually exclusive pairs of grid boxes, wherein mutually exclusive pairs of grid boxes are pairs of grid boxes that are present in the template box pairs file but are absent in the EIC box pairs file or combined EIC box pairs file; and d. identifying EICs that comprise retention time profiles that exhibit non-linear drift, wherein a data point in a grid box of a box pair determined to be mutually exclusive identifies an EIC that comprise retention time profiles that exhibit non-linear drift.
7. The method of claim 6, wherein each pair of grid boxes is represented by a unique identifier to facilitate the comparative analysis.
8. The method of one of claims 1-5, wherein identifying EIC files comprising retention time profiles that exhibit non-linear drift comprises: a. merging all coordinates of two or more data point templates or corrected templates to generate a master coordinates file; b. generating a master coordinate pairs file comprising all possible pairs of coordinates in the master coordinates file; c. identifying all possible pairs of coordinates within each of the individual EIC files represented in the two or more data point templates or corrected data point templates; d. generating a child coordinate pairs file comprising all possible pairs of coordinates within each of the individual EIC files represented in the two or more data point templates or corrected data point templates; e. identifying mutually exclusive pairs of data point coordinates, wherein mutually exclusive pairs of data point coordinates are pairs of coordinatesthat occur in the master template but not in each EIC file comprising at least one data point in each of the pairs of coordinates; f. identifying mutually exclusive pairs of grid boxes, wherein mutually exclusive pairs of coordinates identify mutually exclusive pairs of grid boxes; and g. identifying EICs that comprise retention time profiles that exhibit non-linear drift, wherein data points in mutually exclusive pairs of coordinates or mutually exclusive pairs of grid boxes identify EICs that comprise retention time profiles that exhibit non-linear drift.
9. The method of one of claims 1-5, wherein identifying mutually exclusive pairs of coordinates comprises: a. generating scan gaps in the master coordinate file and generating scan gaps in each of the EICs comprising data points in the master coordinate file, wherein generating scan gaps comprises concatenating data points in each file by connecting a data point with the highest intensity to a data point comprising the next highest intensity and wherein each connection is defined as a scan gap; and b. identifying blur scan gaps, wherein each blur scan gap is a scan gap that links two grid boxes in the merged data points coordinates file that are not linked in any of the individual EICs, and wherein coordinates of data points in the mutually exclusive pairs of grid boxes identify EICs that comprise retention time profiles that exhibit nonlinear drift.
10. The method of any one of the preceding claims, wherein steps (b) to (e) are iteratively repeated until none of the remaining EICs are added to any of the identified templates.11 . The method of any one of the preceding claims, wherein stringency are adjusted by changing the width of the scan, the number of points in each stratum, by determining how many points match or don’t match the template.
12. The method of any of the preceding claims, wherein defining a data point template further comprises identifying one or more EICs that comprise retention time profiles that introduce non-linear blur among the identified EICs of the template and excluding the EICs that introduce non-linear blur before the data values of the identified EICs are merged to construct the data point template.
13. A computing system for overcoming drift in retention times of metabolite signals in mass spectrometry (MS), the system comprising: a. a communication interface that receives EICs of portions of each of a plurality of MS chromatograms, wherein the portions of the plurality of chromatograms bear the same m / z value; and b. a processor that executes instructions stored in memory, wherein the processor executes the instructions to: i. map data point values of the EICs onto a 2D grid system comprising grid boxes of a plurality of scans comprising retention time increments and a plurality of signal intensity strata; ii. select a first reference scan as a scan comprising the highest intensity data points of the highest number of EICs among the plurality of EICs; iii. define a first data point template as all coordinates in the grid comprising at least one data point from at least one EIC identified in step (b)(ii); iv. for each remaining EIC of the plurality of EICs, generate a query EIC by shifting the data points of the query EIC by a time increment sufficient to overlap the highest data point of the query EIC with the highest data point of the first template, wherein a query EIC exhibits linear drift in relation to the first template if coordinates of all the data points of the EIC lie in the first template;v. correct linear drift by adding coordinates of all shifted data points of each query EIC to the first data point template to thereby generate a corrected template; vi. iteratively repeat steps (b)(ii) to (b)(v) using EIC files that were not added to the first data point template to define additional data point templates and correct the EICs for linear drift based on the identified additional templates; and vii. identify, among the EIC files identified in steps (b)(i) to (b)(vii), EIC files comprising retention time profiles that exhibit nonlinear drift in retention time by:1 . merging all data point coordinates of all the EIC files in the identified templates or corrected templates to generate a master coordinates file;2. generating a master coordinate pairs file comprising all pairs of coordinates in the master coordinates file;3. generating a child coordinate pairs file comprising all pairs of coordinates within each of the individual EIC files; and4. identifying mutually exclusive pairs of coordinates, wherein mutually exclusive pairs of coordinates are pairs of coordinates that occur in the master coordinate pairs file but not in the child coordinate pairs file, wherein coordinates of data points in mutually exclusive pairs of coordinates identify EICs that comprise retention time profiles that exhibit nonlinear drift.
14. The computing system of claim 13, wherein the processor further executes the instructions to define a data point template by identifying one or more EICs that introduce non-linear blur among the identified EICs of a template, and excluding the EICs that introduce non-linear blur before the data values of the identified EICs are merged to construct the data point template.
15. The computing system of claim 13 or claim 14, wherein each stratum comprises an equal number of data points and the plurality of scans comprise equal retention time increments.
16. The computing system of any one of the preceding claims, wherein the processor further executes the instructions to identify, among the EIC files identified in steps (b)(i) to (b)(vii), EIC files comprising retention time profiles that exhibit nonlinear drift in retention time by further: a. constructing a template box pairs file and constructing a EIC box pairs file for each EIC file determined to match the template; b. constructing a combined EIC box pairs file by combining EIC box pairs files constructed for each EIC file determined to match the template; c. identifying mutually exclusive pairs of grid boxes, wherein mutually exclusive pairs of grid boxes are pairs of grid boxes that are present in the template box pairs file but are absent in the EIC box pairs file or combined EIC box pairs file; and d. identifying EICs that comprise retention time profiles that exhibit non-linear drift, wherein a data point in a grid box of a box pair determined to be mutually exclusive identifies an EIC that comprise retention time profiles that exhibit non-linear drift.
17. The computing system of claim 16, wherein the processor further executes the instructions to assign a unique identifier to each pair of grid boxes to facilitate the comparative analysis.
18. The computing system of one of claims 13-15, wherein the processor further executes the instructions to identify, among the EIC files identified in steps (b)(i) to (b)(vii), EIC files comprising retention time profiles that exhibit nonlinear drift in retention time by further: a. merging all coordinates of two or more data point templates or corrected templates to generate a master coordinates file;b. generating a master coordinate pairs file comprising all possible pairs of coordinates in the master coordinates file; c. identifying all possible pairs of coordinates within each of the individual EIC files represented in the two or more data point templates or corrected data point templates; d. generating a child coordinate pairs file comprising all possible pairs of coordinates within each of the individual EIC files represented in the two or more data point templates or corrected data point templates; e. identifying mutually exclusive pairs of data point coordinates, wherein mutually exclusive pairs of data point coordinates are pairs of coordinates that occur in the master template but not in each EIC file comprising at least one data point in each of the pairs of coordinates; f. identifying mutually exclusive pairs of grid boxes, wherein mutually exclusive pairs of coordinates identify mutually exclusive pairs of grid boxes; and g. identifying EICs that comprise retention time profiles that exhibit non-linear drift, wherein data points in mutually exclusive pairs of coordinates or mutually exclusive pairs of grid boxes identify EICs that comprise retention time profiles that exhibit non-linear drift.
19. The computing system of one of claims 13-15, wherein the processor further executes the instructions to identify, among the EIC files identified in steps (b)(i) to (b)(vii), EIC files comprising retention time profiles that exhibit nonlinear drift in retention time by further: a. generating scan gaps in the master coordinate file and generating scan gaps in each of the EICs comprising data points in the master coordinate file, wherein generating scan gaps comprises concatenating data points in each file by connecting a data point with the highest intensity to a data point comprising the next highest intensity and wherein each connection is defined as a scan gap; andb. identifying blur scan gaps, wherein each blur scan gap is a scan gap that links two grid boxes in the merged data points coordinates file that are not linked in any of the individual EICs, and wherein coordinates of data points in the mutually exclusive pairs of grid boxes identify EICs that comprise retention time profiles that exhibit nonlinear drift.
20. The computing system of any one of the preceding claims, wherein the processor further executes the instructions to iteratively repeat steps (b) to (e) until none of the remaining EICs are added to any of the identified templates.
21. The computing system of any one of the preceding claims, wherein stringency are adjusted by changing the width of the scan, the number of points in each stratum, by determining how many points match or don’t match the template.
22. The computing system of any one of the preceding claims, wherein the processor further executes the instructions to define a data point template by further identifying one or more EICs that comprise retention time profiles that introduce non-linear blur among the identified EICs of the template and excluding the EICs that introduce non-linear blur before the data values of the identified EICs are merged to construct the data point template.
23. A non-transitory, computer-readable storage medium, having embodied thereon a program executable by a processor to perform a method for identifying metabolite derivatives in chromatography-mass spectrometry (MS) data, the method comprising: a. receiving or having received an MS data file from a chromatography-MS instrument, wherein the data file comprises data points obtained from a sample analyzed by the chromatography-mass spectrometry instrument, and wherein each data point in the file comprises a retention time value and a mass-to-charge (m / z) value;b. mapping data point values of the EICs onto a 2D grid system comprising grid boxes of a plurality of scans comprising retention time increments and a plurality of signal intensity strata; c. selecting a first reference scan as a scan comprising the highest intensity data points of the highest number of EICs among the plurality of EICs; d. defining a first data point template as all grid boxes in the grid comprising at least one data point from at least one EIC associated with the reference scan identified in (b); e. for each remaining EIC of the plurality of EICs, generating a query EIC by shifting the data points of the query EIC by a time increment sufficient to overlap the highest data point of the query EIC with the highest data point of the first template, wherein a query EIC exhibits linear drift in relation to the first template if coordinates of all the data points of the EIC lie in the first template; f. correcting linear drift by adding coordinates of all shifted data points of each query EIC to the first data point template to thereby generate a corrected template; g. iteratively repeating steps (b) to (e) using remaining EICs that were not added to the first data point template to define additional data point templates and correct the linear drift of the EICs based on the identified additional templates; and h. identifying, among the EIC files identified in steps (a) to (f), EIC files comprising retention time profiles that exhibit non-linear drift in retention time by: i. merging all data point coordinates of all the EIC files in the identified templates or corrected templates to generate a master coordinates file; ii. generating a master coordinate pairs file comprising all pairs of coordinates in the master coordinates file;iii. generating a child coordinate pairs file comprising all pairs of coordinates within each of the individual EIC files; and iv. identifying mutually exclusive pairs of coordinates, wherein mutually exclusive pairs of coordinates are pairs of coordinates that occur in the master coordinate pairs file but not in the child coordinate pairs file, wherein coordinates of data points in mutually exclusive pairs of coordinates identify EICs that comprise retention time profiles that exhibit nonlinear drift.
24. The non-transitory, computer-readable storage medium of claim 23, the method further comprises defining a data point template by identifying one or more EICs that introduce non-linear blur among the identified EICs of a template, and excluding the EICs that introduce non-linear blur before the data values of the identified EICs are merged to construct the data point template.
25. The non-transitory, computer-readable storage medium of claim 23 or claim 24, wherein each stratum comprises an equal number of data points and the plurality of scans comprise equal retention time increments.
26. The non-transitory, computer-readable storage medium of any the preceding claims, wherein the method further comprises identifying, among the EIC files identified in steps (c) to (h), EIC files comprising retention time profiles that exhibit nonlinear drift in retention time by further: a. constructing a template box pairs file and constructing a EIC box pairs file for each EIC file determined to match the template; b. constructing a combined EIC box pairs file by combining EIC box pairs files constructed for each EIC file determined to match the template; c. identifying mutually exclusive pairs of grid boxes, wherein mutually exclusive pairs of grid boxes are pairs of grid boxes that are present in the template box pairs file but are absent in the EIC box pairs file or combined EIC box pairs file; andd. identifying EICs that comprise retention time profiles that exhibit non-linear drift, wherein a data point in a grid box of a box pair determined to be mutually exclusive identifies an EIC that comprise retention time profiles that exhibit non-linear drift.
27. The non-transitory, computer-readable storage medium of claim 26, wherein the method further comprises assigning a unique identifier to each pair of grid boxes to facilitate the comparative analysis.
28. The non-transitory, computer-readable storage medium of claims 23-25, wherein the method comprises identifying, among the EIC files identified in steps (c) to (h), EIC files comprising retention time profiles that exhibit nonlinear drift in retention time by further: a. merging all coordinates of two or more data point templates or corrected templates to generate a master coordinates file; b. generating a master coordinate pairs file comprising all possible pairs of coordinates in the master coordinates file; c. identifying all possible pairs of coordinates within each of the individual EIC files represented in the two or more data point templates or corrected data point templates; d. generating a child coordinate pairs file comprising all possible pairs of coordinates within each of the individual EIC files represented in the two or more data point templates or corrected data point templates; e. identifying mutually exclusive pairs of data point coordinates, wherein mutually exclusive pairs of data point coordinates are pairs of coordinates that occur in the master template but not in each EIC file comprising at least one data point in each of the pairs of coordinates; f. identifying mutually exclusive pairs of grid boxes, wherein mutually exclusive pairs of coordinates identify mutually exclusive pairs of grid boxes; andg. identifying EICs that comprise retention time profiles that exhibit non-linear drift, wherein data points in mutually exclusive pairs of coordinates or mutually exclusive pairs of grid boxes identify EICs that comprise retention time profiles that exhibit non-linear drift.
29. The non-transitory, computer-readable storage medium of claims 23-25, wherein the method comprises identifying, among the EIC files identified in steps (c) to (h), EIC files comprising retention time profiles that exhibit nonlinear drift in retention time by further: a. generating scan gaps in the master coordinate file and generating scan gaps in each of the EICs comprising data points in the master coordinate file, wherein generating scan gaps comprises concatenating data points in each file by connecting a data point with the highest intensity to a data point comprising the next highest intensity and wherein each connection is defined as a scan gap; and b. identifying blur scan gaps, wherein each blur scan gap is a scan gap that links two grid boxes in the merged data points coordinates file that are not linked in any of the individual EICs, and wherein coordinates of data points in the mutually exclusive pairs of grid boxes identify EICs that comprise retention time profiles that exhibit nonlinear drift.
30. The non-transitory, computer-readable storage medium of any of the preceding claims, wherein the method further comprises iteratively repeating steps (b) to (e) until none of the remaining EICs are added to any of the identified templates.
31. The non-transitory, computer-readable storage medium of any of the preceding claims, wherein stringency are adjusted by changing the width of the scan, the number of points in each stratum, by determining how many points match or don’t match the template.
32. The non-transitory, computer-readable storage medium of any of the preceding claims, wherein the method further comprises defining a data point template by further identifying one or more EICs that comprise retention time profiles thatintroduce non-linear blur among the identified EICs of the template and excluding the EICs that introduce non-linear blur before the data values of the identified EICs are merged to construct the data point template.