Correction of retention time drift in mass spectrometry through separate alignment of mass features
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2026-02-06
- Publication Date
- 2026-08-13
Smart Images

Figure US2026014232_13082026_PF_FP_ABST
Abstract
Description
PROVPATENT Danforth Docket No.: DDPSC0170-401 -PCT Via EFS-webCORRECTION OF RETENTION TIME DRIFT IN MASS SPECTROMETRY THROUGH SEPARATE ALIGNMENT OF MASS FEATURESGOVERNMENTAL RIGHTS
[0001] This invention was made with government support under DE-SC0018277 awarded by the Department of Energy. The government has certain rights in the invention.CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims priority from Provisional Application number 63 / 754,719, filed February 6, 2025, the entire contents of which are hereby incorporated by reference.FIELD OF THE INVENTION
[0003] The present disclosure relates generally to compound detection. More specifically, the present disclosure relates to methods for independently aligning retention times of mass feature signals in mass spectrometry data.BACKGROUND OF THE INVENTION
[0004] Mass spectrometry techniques such as liquid chromatography-mass spectrometry (LC-MS) are chemical techniques that identify different compounds in a sample as unique mass features. In LC-MS, a liquid chromatography system can separate the different compounds by structural properties, while a mass spectrometer subsequently determines the mass and intensity of the ions that elute from the chromatography column. Modern high-resolution mass spectrometry can now detect and quantify ions with high mass precision (< 5 ppm mass error) but can also result in significant amounts of noise.
[0005] Detecting valid compound peaks within mass spectrometry data presents a number of challenges when the compound is only be present at low levels relative to noise. For example, samples from complex systems can include large numbers of different compounds, some of which can only be present in relatively low quantities. A typical mass spectrometry file canPROV PATENT Danforth Docket No.: DDPSC0170-401 -PCT Via EFS-webcontain as many as millions of data points, while as few as several hundred to thousands can correspond to true mass feature signals that are interspersed in vast amounts of noise.
[0006] One of the great challenges in harnessing mass spectrometry techniques such as liquid chromatography-mass spectrometry (LC-MS) for untargeted metabolomics is having to correct for retention time drift, for instance regarding metabolites with difficult to resolve isomer profiles. For example, the retention time instrumentation can be off when a batch within the course of experiments / runs is processed. This results in drift of the intensity signal value in the retention time domain (e.g., a shift in signal to an earlier or later time).
[0007] The difficulty in correcting drift arises in part because the outline of each isomer is often faint and fluctuates across samples - i.e. low signal to noise. The issue of low signal to noise can be addressed by pooling samples together in order to create a merged extracted ion chromatogram from which boundaries can be called for each isomer of a mass. Because signal is non-random, while noise is random, pooling samples together increases signal to noise. However, this has the consequence - when there is drift - of blurring isomers together. Alternatively, when there is no drift, isomers can be swapped with one another.
[0008] Existing strategies generally analyze mass feature signals and the retention time drifts computationally by aligning all the signals for various metabolites at once, or in large chunks that improperly treat metabolites with disparate characteristics the same. In other words, current techniques attempt to match all mass feature signals, which can not behave in predictable fashion, simultaneously. However, this approach introduces warping effects because drift isn’t being corrected on an individual basis for each compound. Especially problematic is when a metabolite’s signal is close to another metabolite, with each one experiencing its own chromatography drift. Such methods end up discarding valid mass feature signals, as well as being computationally expensive and slow.
[0009] For instance, one existing strategy to combat drift takes the retention time profiles for all m / z values in query samples and align them to their corresponding features in a reference sample using a warping function. However, this strategy to determine a global fit is fraught with difficulties as often, there is no effective global solution for drift. For example, the reference can not even have all the metabolites present in the query.PROV PATENT Danforth Docket No.: DDPSC0170-401 -PCT Via EFS-web
[0010] Thus, there is a need for improved systems and methods correcting drift in retention times of mass feature signals such as those of metabolites.SUMMARY OF THE INVENTION
[0011] One aspect of the instant disclosure encompasses a method for correcting drift in retention times of mass features in mass spectrometry (MS). The method comprises: for an m / z bin, generating or having generated an extracted ion chromatogram (EIC) in each file of a dataset comprising a plurality of files. The method further comprises locally aligning mass features in an m / z bin of interest by: selecting a first reference retention time across files (refScan) as a scan comprising the highest intensity data points of the highest number of EICs across the files of the dataset among the plurality of EICs extracted from each file in the dataset and establishing retention-time bounds centered on the first refScan within which data points are identified and attributed to the first mass feature in each file, thereby defining a first mass feature. The method further comprises assigning a per-file peak center (fileScan) of each remaining EIC of the plurality of EICs to a nearest selected refScan, wherein a remaining EIC is an EIC in which the fileScan does not coincide with the refScan of the first mass feature, and shifting the data points of each remaining EIC by a time increment sufficient to align the fileScan of the EIC with its assigned refScan. The method further comprises iteratively repeating the previous steps of selecting a first refScan through shifting data points for each nth+1 refScan to define an nth+1 mass feature, thereby assigning each remaining EIC’s nth+1 fileScan to that refScan, shifting to align, and recording bounds; wherein each iteration uses an unshifted data file to identify the nth+1 mass features, wherein the nth+1 selection excludes data points within the bounds of previously defined mass features, and wherein n is an integer having a value equal to or more than 1. The method further comprises recording in a shiftTable, for every refScan-file pair, the retention time of the fileScan, the retention time of the refScan to which the fileScan is assigned, a shift equal to fileScan retention time minus refScan retention time, a mass feature identifier, and optionally the m / z bin for the mass feature, a file identifier, and one or more quality flags. The method further comprises aligning each nthmass feature defined in steps of selecting a first refScan through the step of recording in a shiftTable, based on the shiftTable and using the original unshifted EIC of each data file,PROV PATENT Danforth Docket No.: DDPSC0170-401 -PCT Via EFS-webby shifting the data points attributed to that nthmass feature along the retention-time axis by a time increment sufficient to bring the recorded fileScan of that nthmass feature to the closest refScan in that file; thereby correcting the drift in retention times of mass features.
[0012] The method can optionally, for EICs in which two or more mass features have an ambiguous shift when locally aligning mass features in an m / z bin of interest, wherein an ambiguous shift occurs when two or more fileScans in a single EIC are closest to, or assigned to, the same refScan such that more than one shift value is recorded for that refScan-file situation, resolve the ambiguity by: identifying a chromolog for each of the two or more mass features by constructing from the retention times of fileScans recorded in the shiftTable, for a query mass feature that has an ambiguous shift, a shiftRow, wherein the shiftRow comprises the recorded retention time of the refScan of the query mass feature, the recorded retention times of each fileScan of the query mass feature in each file of the dataset, and a mass feature identifier; and constructing, across the same files, a corresponding shiftRow for one or more candidate chromologs, wherein a “candidate chromolog” is a mass feature in a different m / z bin whose refScan retention time matches, within a predefined retention-time tolerance, the query mass feature’s refScan retention time; and assigning a chromolog from the one or more candidate chromologs to the query mass feature by computing, from the shiftRows, an aggregate co-location score defined as the sum, over files in which both the query mass feature and a candidate chromolog have recorded fileScans, of the value of the difference between the candidate chromolog’s fileScan retention time and the query mass feature’s corresponding fileScan retention time in the same file; and selecting as the chromolog the candidate chromolog that minimizes said aggregate co-location score. The method further comprises aligning fileScans of the query mass feature, in each file where the query mass feature has an ambiguous shift, by first shifting, using the original unshifted EIC of that file for the chromolog’s m / z bin, the data points already attributed to the chromolog along the retention-time axis by a time increment sufficient to bring the chromolog’s recorded fileScan to the chromolog’s refScan; and then, in the same file, shifting the data points already attributed to the query mass feature by the same time increment applied to the chromolog, thereby bringing the query mass feature’s recorded fileScan into the retention-time bounds centered on the refScan used for both the chromolog and the query mass feature. The method canPROV PATENT Danforth Docket No.: DDPSC0170-401 -PCT Via EFS-webiteratively repeat resolving the ambiguity for each additional mass feature that has an ambiguous shift associated with the same refScan and can optionally update the shiftTable to record, for each file in which the chromolog is used to align the query mass feature, the corrected fileScan retention time for the query mass feature and the time increment applied; wherein a mass feature is defined as a single chromatographic peak identified within an m / z bin and retention-time peak.
[0013] In some embodiments, the quality flags comprise at least: a proxi mity-to-refS can flag indicating whether the fileScan retention time lies within predefined bounds around the refScan; a signal-to-noise flag indicating the apex signal-to-noise ratio at the fileScan; an ambiguity flag indicating whether multiple candidate apexes were present near the refScan; and a correspondence flag indicating presence or absence of a credible fileScan for the refScan in the file. In some embodiments, the quality flags comprise a peak shape flag indicating whether one or more peak-shape measures, including symmetry, tailing factor, or width-at-half-maximum, fall within predetermined ranges, an intensity threshold flag indicating whether the apex intensity at the fileScan meets predetermined minimum or maximum thresholds, or both. In some embodiments, the quality flags further comprise an overlap or coelution flag indicating evidence of overlapping mass features within the refScan bounds, an outlier flag indicating whether the fileScan retention time or apex intensity is outside a robust distribution for the mass feature across the plurality of files, a drift magnitude flag indicating categorical magnitude of the shift relative to a threshold, or any combination thereof. In some embodiments, the quality flags further comprise an overlap or co-elution flag indicating evidence of overlapping mass features within the refScan bounds, an outlier flag indicating whether the fileScan retention time or apex intensity is outside a robust distribution for the mass feature across the plurality of files, a drift magnitude flag indicating categorical magnitude of the shift relative to a threshold, or any combination thereof.
[0014] In some embodiments, the unshifted data file is annotated with each mass feature, peak center of each mass feature, bounds of each mass feature, data point values of each mass feature, or any combination thereof. In some embodiments, an aligned nthpeak center is the peak center of a distribution curve that encompasses data points of an nthmass feature,PROV PATENT Danforth Docket No.: DDPSC0170-401 -PCT Via EFS-webthereby defining the nthmass feature. In some embodiments, retention time bounds of each mass feature are determined using a smoothing algorithm.
[0015] In some embodiments, a mass feature is metabolites and isomers of metabolites. In some embodiments, a mass feature is amino acids, peptides, and proteins. In some embodiments, the MS is chromatography-MS.
[0016] In some embodiments, the chromatography-MS is liquid chromatography-MS (LC-MS).
[0017] Another aspect of the instant disclosure encompasses a method for correcting drift in retention times of mass features in mass spectrometry (MS). The method comprises: for each m / z bin, generating or having generated an extracted ion chromatogram (EIC) in each file of a dataset comprising a plurality of files; wherein a file is the full LC-MS data acquired from one sample, and an EIC is the intensity-versus-retention-time (RT) signal extracted for one m / z bin within a file; and referencing a shiftTable for a dataset comprising a plurality of files, wherein the shiftTable records, for every refScan-file pair, the retention time of the fileScan, the retention time of the refScan to which the fileScan is assigned, a mass feature identifier, and the m / z bin for the mass feature, and optionally a file identifier and one or more quality flags. The method further comprises, for EICs in which two or more mass features have an ambiguous shift when performing the step of referencing a shiftTable, wherein an ambiguous shift occurs when two or more fileScans in a single EIC are closest to, or assigned to, the same refScan such that more than one shift value is recorded for that refScan-file situation, resolving the ambiguity by: identifying a chromolog for a query mass feature by constructing from the retention times of fileScans recorded in the shiftTable, for a query mass feature that has an ambiguous shift, a shiftRow, wherein the shiftRow comprises the recorded retention time of the refScan of the query mass feature, the recorded retention times of each fileScan of the query mass feature in each file of the dataset, and a mass feature identifier; constructing, across the same files, a corresponding shiftRow for one or more candidate chromologs, wherein a “candidate chromolog” is a mass feature in a different m / z bin whose refScan retention time matches, within a predefined retention-time tolerance, the query mass feature’s refScan retention time; and assigning a chromolog from the one or more candidate chromologs to the query mass feature by computing, from the shiftRows, an aggregate co-location score defined as the sum, over files in which both the query mass feature and a candidate chromologPROV PATENT Danforth Docket No.: DDPSC0170-401 -PCT Via EFS-webhave recorded fileScans, of the value of the difference between the candidate chromolog’s fileScan retention time and the query mass feature’s corresponding fileScan retention time in the same file; and selecting as the chromolog the candidate chromolog that minimizes said aggregate co-location score. The method further comprises aligning fileScans of the query mass feature, in each file where the query mass feature has an ambiguous shift, by first shifting, using the original unshifted EIC of that file for the chromolog’s m / z bin, the data points already attributed to the chromolog along the retention-time axis by a time increment sufficient to bring the chromolog’s recorded fileScan to the chromolog’s refScan; and then, in the same file, shifting the data points already attributed to the query mass feature by the same time increment applied to the chromolog, thereby bringing the query mass feature’s recorded fileScan into retention-time bounds centered on the refScan used for both the chromolog and the query mass feature. The method can iteratively repeat steps of identifying a chromolog for a query mass feature through aligning fileScans of the query mass feature that has an ambiguous shift associated with the same refScan and can optionally update the shiftTable to record, for each file in which the chromolog is used to align the query mass feature, the corrected fileScan retention time for the query mass feature and the time increment applied; wherein a file is the full LC-MS data acquired from one sample; an extracted ion chromatogram (EIC) is the intensity-versus-retention-time signal extracted for one m / z bin within a file; a mass feature is defined as a single chromatographic peak identified within an m / z bin and retention-time peak; a refScan is a reference retention time across files selected to represent a mass feature and used as an alignment target; and a fileScan is the per-file peak center nearest a refScan recorded for that mass feature in that file.
[0018] An additional aspect of the instant disclosure encompasses a computing system for correcting drift in retention times of mass features in mass spectrometry (MS). The system comprises: a communication interface that receives EICs generated for an m / z bin in each file of a dataset comprising a plurality of files; and a processor that executes instructions stored in memory, wherein the processor executes the instructions to: locally align mass features in an m / z bin of interest by: selecting a first reference retention time across files (refScan) as a scan comprising the highest intensity data points of the highest number of EICs across the files of the dataset among the plurality of EICs extracted from each file in the dataset and establishingPROV PATENT Danforth Docket No.: DDPSC0170-401 -PCT Via EFS-webretention-time bounds centered on the first refScan within which data points are identified and attributed to the first mass feature in each file, thereby defining a first mass feature; assigning a per-file peak center (fileScan) of each remaining EIC of the plurality of EICs to a nearest refScan selected in the step of locally align mass features in an m / z bin of interest, wherein a remaining EIC is an EIC in which the fileScan does not coincide with the refScan of the first mass feature; shifting the data points of each remaining EIC by a time increment sufficient to align the fileScan of the EIC with its assigned refScan; iteratively repeating steps of selecting a first refScan through shifting the data points of each remaining EIC for each nth+1 refScan to define an nth+1 mass feature, thereby assigning each remaining EIC’s nth+1 fileScan to that refScan, shifting to align, and recording bounds; wherein each iteration uses an unshifted data file to identify the nth+1 mass features, wherein the nth+1 selection excludes data points within the bounds of previously defined mass features, and wherein n is an integer having a value equal to or more than 1 ; recording in a shiftTable, for every refScan-file pair, the retention time of the fileScan, the retention time of the refScan to which the fileScan is assigned, a shift equal to fileScan retention time minus refScan retention time, a mass feature identifier, and optionally the m / z bin for the mass feature, a file identifier, and one or more quality flags; and aligning each nthmass feature defined in steps of selecting a first refScan through recording in a shiftTable, based on the shiftTable and using the original unshifted EIC of each data file, by shifting the data points attributed to that nthmass feature along the retention-time axis by a time increment sufficient to bring the recorded fileScan of that nthmass feature to the closest refScan in that file; thereby correcting the drift in retention times of mass features. The system can optionally, for EICs in which two or more mass features have an ambiguous shift when performing the step of locally align mass features, wherein an ambiguous shift occurs when two or more fileScans in a single EIC are closest to, or assigned to, the same refScan such that more than one shift value is recorded forthat refScan-file situation, resolve the ambiguity by identifying a chromolog for each of the two or more mass features by constructing shiftRows, assigning a chromolog by minimizing an aggregate co-location score, and aligning fileScans by applying the chromolog’s per-file shift, iteratively repeating for each additional mass feature that has an ambiguous shift associated with the same refScan, and optionallyPROV PATENT Danforth Docket No.: DDPSC0170-401 -PCT Via EFS-webupdating the shiftTable; wherein a mass feature is a single chromatographic peak identified within an m / z bin and retention-time peak.
[0019] In some embodiments, the processor can further execute instructions to generate the EIC in each file of a dataset comprising a plurality of files.
[0020] Yet another aspect of the instant disclosure encompasses a computing system for correcting drift in retention times of mass features in mass spectrometry (MS). The system comprises: a communication interface that receives EICs generated for an m / z bin in each file of a dataset comprising a plurality of files; wherein a file is the full LC-MS data acquired from one sample, and an EIC is the intensity-versus-retention-time (RT) signal extracted for one m / z bin within a file; and a processor that executes instructions stored in memory, wherein the processor executes the instructions to: reference a shiftTable for a dataset comprising a plurality of files, wherein the shiftTable records, for every refScan-file pair, the retention time of the fileScan, the retention time of the refScan to which the fileScan is assigned, a mass feature identifier, and the m / z bin for the mass feature, and optionally a file identifier and one or more quality flags; and for EICs in which two or more mass features have an ambiguous shift when referencing a shiftTable, wherein an ambiguous shift occurs when two or more fileScans in a single EIC are closest to, or assigned to, the same refScan such that more than one shift value is recorded for that refScan-file situation, resolve the ambiguity by identifying a chromolog for a query mass feature by constructing shiftRows, assigning a chromolog by minimizing an aggregate co-location score, and aligning fileScans by applying the chromolog’s per-file shift, iteratively repeating for each additional mass feature that has an ambiguous shift associated with the same refScan, and optionally updating the shiftTable; wherein a file is the full LC-MS data acquired from one sample; an extracted ion chromatogram (EIC) is the intensity-versus-retention-time signal extracted for one m / z bin within a file; a mass feature is defined as a single chromatographic peak identified within an m / z bin and retention-time peak; a refScan is a reference retention time across files selected to represent a mass feature and used as an alignment target; and a fileScan is the per-file peak center nearest a refScan recorded for that mass feature in that file.
[0021] One aspect of the instant disclosure encompasses a non-transitory, computer-readable storage medium, having embodied thereon a program executable by a processor to perform aPROV PATENT Danforth Docket No.: DDPSC0170-401 -PCT Via EFS-webmethod for correcting drift in retention times of mass features in mass spectrometry (MS). The method comprises: for an m / z bin, generating or having generated an extracted ion chromatogram (EIC) in each file of a dataset comprising a plurality of files. The method further comprises locally aligning mass features in an m / z bin of interest by: selecting a first reference retention time across files (refScan) as a scan comprising the highest intensity data points of the highest number of EICs across the files of the dataset among the plurality of EICs extracted from each file in the dataset and establishing retention-time bounds centered on the first refScan within which data points are identified and attributed to the first mass feature in each file, thereby defining a first mass feature; assigning a per-file peak center (fileScan) of each remaining EIC of the plurality of EICs to a nearest refScan selected in step (b)(i), wherein a remaining EIC is an EIC in which the fileScan does not coincide with the refScan of the first mass feature; and shifting the data points of each remaining EIC by a time increment sufficient to align the fileScan of the EIC with its assigned refScan. The method further comprises iteratively repeating steps of selecting a refScan through shifting the data points for each nth+1 refScan to define an nth+1 mass feature, thereby assigning each remaining EIC’s nth+1 fileScan to that refScan, shifting to align, and recording bounds; wherein each iteration uses an unshifted data file to identify the nth+1 mass features, wherein the nth+1 selection excludes data points within the bounds of previously defined mass features, and wherein n is an integer having a value equal to or more than 1. The method further comprises recording in a shiftTable, for every refScan-file pair, the retention time of the fileScan, the retention time of the refScan to which the fileScan is assigned, a shift equal to fileScan retention time minus refScan retention time, a mass feature identifier, and optionally the m / z bin for the mass feature, a file identifier, and one or more quality flags. The method further comprises aligning each nthmass feature defined in steps of selecting a refScan through recording in a shiftTable, based on the shiftTable and using the original unshifted EIC of each data file, by shifting the data points attributed to that nthmass feature along the retention-time axis by a time increment sufficient to bring the recorded fileScan of that nthmass feature to the closest refScan in that file; thereby correcting the drift in retention times of mass features.
[0022] The method can optionally, for EICs in which two or more mass features have an ambiguous shift when performing the step of locally aligning mass features in an m / z bin ofPROV PATENT Danforth Docket No.: DDPSC0170-401 -PCT Via EFS-webinterest, wherein an ambiguous shift occurs when two or more fileScans in a single EIC are closest to, or assigned to, the same refScan such that more than one shift value is recorded for that refScan-file situation, resolve the ambiguity by: identifying a chromolog for each of the two or more mass features by constructing from the retention times of fileScans recorded in the shiftTable, for a query mass feature that has an ambiguous shift, a shiftRow, wherein the shiftRow comprises the recorded retention time of the refScan of the query mass feature, the recorded retention times of each fileScan of the query mass feature in each file of the dataset, and a mass feature identifier; constructing, across the same files, a corresponding shiftRow for one or more candidate chromologs, wherein a “candidate chromolog” is a mass feature in a different m / z bin whose refScan retention time matches, within a predefined retention-time tolerance, the query mass feature’s refScan retention time; and assigning a chromolog from the one or more candidate chromologs to the query mass feature by computing, from the shiftRows, an aggregate co-location score defined as the sum, over files in which both the query mass feature and a candidate chromolog have recorded fileScans, of the value of the difference between the candidate chromolog’s fileScan retention time and the query mass feature’s corresponding fileScan retention time in the same file; and selecting as the chromolog the candidate chromolog that minimizes said aggregate co-location score. The method further comprises aligning fileScans of the query mass feature, in each file where the query mass feature has an ambiguous shift, by first shifting, using the original unshifted EIC of that file for the chromolog’s m / z bin, the data points already attributed to the chromolog along the retention-time axis by a time increment sufficient to bring the chromolog’s recorded fileScan to the chromolog’s refScan; and then, in the same file, shifting the data points already attributed to the query mass feature by the same time increment applied to the chromolog, thereby bringing the query mass feature’s recorded fileScan into the retention-time bounds centered on the refScan used for both the chromolog and the query mass feature. The method can iteratively repeat steps of identifying a chromolog for each of the two or more mass features through aligning fileScans of the query mass feature for each additional mass feature that has an ambiguous shift associated with the same refScan and can optionally update the shiftTable to record, for each file in which the chromolog is used to align the query mass feature, the corrected fileScan retention time for the query mass feature and the timePROV PATENT Danforth Docket No.: DDPSC0170-401 -PCT Via EFS-webincrement applied; wherein a mass feature is defined as a single chromatographic peak identified within an m / z bin and retention-time peak.
[0023] Another aspect of the instant disclosure encompasses a non-transitory, computer-readable storage medium, having embodied thereon a program executable by a processor to perform a method for correcting drift in retention times of mass features in mass spectrometry (MS). The method comprises: for each m / z bin, generating or having generated an extracted ion chromatogram (EIC) in each file of a dataset comprising a plurality of files; wherein a file is the full LC-MS data acquired from one sample, and an EIC is the intensity-versus-retention-time (RT) signal extracted for one m / z bin within a file. The method further comprises referencing a shiftTable for a dataset comprising a plurality of files, wherein the shiftTable records, for every refScan-file pair, the retention time of a fileScan, the retention time of a refScan to which the fileScan is assigned, a mass feature identifier, and the m / z bin for the mass feature, and optionally a file identifier and one or more quality flags. The method further comprises, for EICs in which two or more mass features have an ambiguous shift when referencing a shiftTable, wherein an ambiguous shift occurs when two or more fileScans in a single EIC are closest to, or assigned to, the same refScan such that more than one shift value is recorded for that refScan-file situation, resolving the ambiguity by: identifying a chromolog for a query mass feature by constructing from the retention times of fileScans recorded in the shiftTable, for a query mass feature that has an ambiguous shift, a shiftRow, wherein the shiftRow comprises the recorded retention time of the refScan of the query mass feature, the recorded retention times of each fileScan of the query mass feature in each file of the dataset, and a mass feature identifier; constructing, across the same files, a corresponding shiftRow for one or more candidate chromologs, wherein a “candidate chromolog” is a mass feature in a different m / z bin whose refScan retention time matches, within a predefined retention-time tolerance, the query mass feature’s refScan retention time; and assigning a chromolog from the one or more candidate chromologs to the query mass feature by computing, from the shiftRows, an aggregate co-location score defined as the sum, over files in which both the query mass feature and a candidate chromolog have recorded fileScans, of the value of the difference between the candidate chromolog’s fileScan retention time and the query mass feature’s corresponding fileScan retention time in the same file; and selecting as thePROV PATENT Danforth Docket No.: DDPSC0170-401 -PCT Via EFS-webchromolog the candidate chromolog that minimizes said aggregate co-location score. The method further comprises aligning fileScans of the query mass feature, in each file where the query mass feature has an ambiguous shift, by first shifting, using the original unshifted EIC of that file for the chromolog’s m / z bin, the data points already attributed to the chromolog along the retention-time axis by a time increment sufficient to bring the chromolog’s recorded fileScan to the chromolog’s refScan; and then, in the same file, shifting the data points already attributed to the query mass feature by the same time increment applied to the chromolog, thereby bringing the query mass feature’s recorded fileScan into retention-time bounds centered on the refScan used for both the chromolog and the query mass feature. The method can iteratively repeat steps of identifying a chromolog for a query mass feature through aligning fileScans of the query mass feature for each additional mass feature that has an ambiguous shift associated with the same refScan and can optionally update the shiftTable to record, for each file in which the chromolog is used to align the query mass feature, the corrected fileScan retention time for the query mass feature and the time increment applied; wherein a file is the full LC-MS data acquired from one sample; an extracted ion chromatogram (EIC) is the intensity-versus-retention-time signal extracted for one m / z bin within a file; a mass feature is defined as a single chromatographic peak identified within an m / z bin and retention-time peak; a refScan is a reference retention time across files selected to represent a mass feature and used as an alignment target; and a fileScan is the per-file peak center nearest a refScan recorded for that mass feature in that fileBRIEF DESCRIPTION OF THE FIGURES
[0024] The following drawings form part of the present specification and are included to further demonstrate certain embodiments of the present disclosure. Certain embodiments can be better understood by reference to one or more of these drawings in combination with the detailed description of specific embodiments presented herein. The patent or application file contains at least one drawing executed in color. Copies of this patent or patent application publication with color drawing(s) will be provided by the Office upon request and payment of the necessary fee.PROV PATENT Danforth Docket No.: DDPSC0170-401 -PCT Via EFS-web
[0025] FIG. 1 illustrates an exemplary mass spectrometry dataset from a UPLC-MS run displaying three axes representing RT, m / z ratio, and intensity. Inset: An enlarged portion of the same human urine data file highlighting the sensitivity and detail in the small molecule features collected.
[0026] FIG. 2 depicts pooled retention time profile of EIC files of a particular m / z value, derived from real data, in accordance with some embodiments of the instant disclosure.
[0027] FIG. 3 depicts a diagrammatic representation of a hypothetical retention time profiles from pooled EICs of a segment of a MS chromatogram for a specific m / z value obtained from samples comprising two isomers of a hypothetical compound, in accordance with some embodiments of the instant disclosure. The top panel depicts the pooled EICs before aligning mass features. Bottom panel depicts the pooled EICs after independently correcting drift for the two isomers in accordance with some embodiments of the instant disclosure.
[0028] FIG. 4 depicts a diagrammatic representation of aligning a first peak center, in accordance with some embodiments of the instant disclosure. The top panel depicts the pooled EICs before aligning mass features. The bottom panel depicts the EICs after aligning the first peak center of the first isomer, in accordance with some embodiments of the instant disclosure.
[0029] FIG. 5 depicts the bottom panel shown in FIG. 4, with data points of the first defined mass feature diagrammatically annotated by hatching and data points of misaligned mass features diagrammatically represented by red outlines, in accordance with some embodiments of the instant disclosure.
[0030] FIG. 6 depicts the bottom panel shown in FIG. 4, with data points of the first defined mass feature removed, in accordance with some embodiments of the instant disclosure. Data points of the faint non-linearly shifted second mass feature of the dark grey chromatogram is diagrammatically represented by green outlines, in accordance with some embodiments of the instant disclosure
[0031] FIG. 7 depicts a diagrammatic representation of aligning a second peak center, in accordance with some embodiments of the instant disclosure. The top panel depicts the pooled EICs before aligning mass features and having the data points of the first mass feature removed in accordance with some embodiments of the instant disclosure. The bottom panelPROV PATENT Danforth Docket No.: DDPSC0170-401 -PCT Via EFS-webdepicts the EICs after aligning the peak centers of the second isomer, in accordance with some embodiments of the instant disclosure. Identified data points of the second isomer of the dark grey EIC are visually represented by green outlining.
[0032] FIG. 8 depicts a diagrammatic representation of aligning a second peak center, in accordance with some embodiments of the instant disclosure. The top panel depicts the pooled EICs before aligning mass features and having the data points of the first mass feature annotated by hatching in accordance with some embodiments of the instant disclosure. The bottom panel depicts the EICs after aligning the peak centers of the second isomer, in accordance with some embodiments of the instant disclosure. Identified data points of the second isomer of the dark grey EIC are visually represented by green outlining.
[0033] FIG. 9 depicts a diagrammatic representation of the hypothetical retention time profiles from pooled EICs of FIG. 3, in accordance with some embodiments of the instant disclosure. The top panel depicts the pooled EICs before aligning mass features. Bottom panel depicts the pooled EICs after correcting drift for the two isomers and correcting swapped isomers in accordance with some embodiments of the instant disclosure.
[0034] FIG. 10 depicts the pooled retention time profile of EIC files of FIG. 2, in accordance with some embodiments of the instant disclosure. In this figure, the peak centers of various mass features are noted by the vertical lines.
[0035] FIG. 11A is a zoomed-in view of the early portion of the pooled retention time profile of FIGs. 2 and 10
[0036] FIG. 11B is a zoomed-in view of the middle portion of the pooled retention time profile of FIGs. 2 and 10
[0037] FIG. 11C is a zoomed-in view of the late portion of the pooled retention time profile of FIGs. 2 and 10
[0038] FIG. 12A corresponds to the early portion of the pooled retention time profile of FIG. 11A noting the first peak centers by the vertical line.
[0039] FIG. 12B corresponds to the early portion of the pooled retention time profile of FIG. 11A after aligning first peak centers and defining a first mass feature, here shown between the red and green vertical lines.PROV PATENT Danforth Docket No.: DDPSC0170-401 -PCT Via EFS-web
[0040] FIG. 12C corresponds to the early portion of the pooled retention time profile of FIG. 11A after aligning first peak centers and defining a first mass feature and after removing the data points of the defined mass feature.
[0041] FIG. 13 is a flowchart illustrating an exemplary method 500 for defining a template in an embodiment of the Last Resort methods of the instant disclosure.
[0042] FIG. 14 is a diagrammatic illustration of ambiguous shift: upper panel shows positions of leucine (Leu) and isoleucine (lie) mass features with refScans identified; lower panel shows a file in which both He and Leu fileScans drift left and are closest to the Leu refScan, producing an ambiguous shift.
[0043] FIG. 15. Candidate chromologs overview (panel set 1): top panel (m / z bin A) shows two mass features (isomer 1, drifted; isomer 2, minimally drifted); middle panel (m / z bin B) shows candidate chromolog 1 with similar drift dispersion; bottom panel (m / z bin C) shows candidate chromolog 2 with little dispersion.
[0044] FIG. 16. Co-Iocation with candidate chromolog 1: arrows from isomer 1 fileScans (top panel) to candidate chromolog 1 fileScans (middle panel) indicate file-by-file co-location in RT, supporting candidate 1 as the chromolog.
[0045] FIG. 17. Non-co-location with candidate chromolog 2: arrows from isomer 1 fileScans (top panel) to candidate chromolog 2 fileScans (bottom panel) show poor file-by-file co-location in RT, disqualifying candidate 2 as the chromolog.
[0046] FIG. 18. ShiftRows and aggregate co-location scoring: diagram of shiftRows for the query mass feature and two candidate chromologs, and example computation of aggregate colocation scores used to select the chromolog.
[0047] FIG. 19. Chromolog alignment setup (right peak): upper panel shows Leu and He at their refScans; second panel shows files with ambiguous shift; third panel shows chromolog central tendency aligned with the right main peak; fourth panel shows drifted chromolog fileScans aligning with the right drifted peak.
[0048] FIG. 20. Chromolog-guided alignment (right peak): after applying the chromolog’s perfile shift, the chromolog fileScans are brought to its refScan and the query fileScans are shifted into the bounds centered on the shared refScan, resolving the right-peak ambiguity.PROV PATENT Danforth Docket No.: DDPSC0170-401 -PCT Via EFS-web
[0049] FIG. 21. Chromolog alignment setup (left peak): analogous to FIG. 19 for the leftmost peak; chromolog central tendency and drifted chromolog fileScans shown for files with ambiguous shift.
[0050] FIG. 22. Chromolog-guided alignment (left peak): analogous to FIG. 20 for the leftmost peak; the chromolog’s per-file shift is applied to bring the query fileScans into the bounds centered on the appropriate refScan.
[0051] FIG. 23. Real LC-MS data illustrating alignment of amino-acid-derived mass features (lle / Leu): upper panel shows He and Leu mass features across ~400 files; lower panel shows a zoomed view of the lle / Leu region highlighting convolution due to drift in a subset of files.
[0052] FIG. 24. Histograms for He and its chromolog: histogram view of He fileScans and of the lie chromolog fileScans in the He zone of interest.
[0053] FIG. 25. He chromolog tracking: top panel depicts the diagrammatic Leu / lle motifs; middle panel shows the histogram of the lie chromolog; bottom panel overlays real fileScans (He: black dots; He chromolog: blue dots), illustrating file-by-file tracking.
[0054] FIG. 26. Real-data alignment of amino-acid-derived mass features (lie): before / after view showing chromolog fileScans aligned to the chromolog refScan and the same shift applied to align He fileScans to the He refScan, yielding cleaner data.
[0055] FIG. 27. Histogram perspective (lie): before / after histograms corresponding to FIG. 26, confirming cleanup and resolution of ambiguous shifts.
[0056] FIG. 28. Histograms for Leu and its chromolog: histogram view of lie fileScans and of the He chromolog fileScans in the Leu zone of interest.
[0057] FIG. 29. Leu chromolog tracking: top panel depicts the diagrammatic Leu / lle motifs; middle panel shows the histogram of the Leu chromolog; bottom panel overlays real fileScans (Leu: black dots; Leu chromolog: purple dots), illustrating file-by-file tracking.
[0058] FIG. 30. Real-data alignment of amino-acid-derived mass features (Leu): before / after view showing chromolog fileScans aligned to the chromolog refScan and the same shift applied to align Leu fileScans to the Leu refScan, yielding cleaner data.
[0059] FIG. 31. Histogram perspective (Leu): before / after histograms corresponding to FIG. 30, confirming cleanup and resolution of ambiguous shifts.PROV PATENT Danforth Docket No.: DDPSC0170-401 -PCT Via EFS-web
[0060] FIG. 32. is a flowchart illustrating an exemplary method 600 for defining a template in an embodiment of the chromolog methods of the instant disclosure.
[0061] FIG. 33 illustrates an example of a computing system to correct for drift, in accordance with example embodiments.
[0062] DETAILED DESCRIPTION
[0063] The present disclosure describes systems and methods for correcting for retention time drift of MS mass features for the amplification and detection of MS signals. The inventors devised methods of addressing drift using raw data from each MS data file, before processing steps normally performed on the MS data files, including processing of the mass spectra for peak detection and quantification, calibration, annotation, and noise reduction. Critically, the methods of the instant disclosure also depart from previous methods by correcting drift on an individual basis for each compound by aligning each mass feature separately.
[0064] The methods disclosed herein establish linear equations for each mass feature separately from other mass features, performing iterative linear shifts even for non-linearly shifted features. This approach enhances MS data analysis by amplifying signals through data pooling from multiple EICs, increasing the signal-to-noise ratio. Additionally, it allows for the individual alignment of mass features, capturing weak signals that might otherwise be lost as background noise.
[0065] Importantly, these methods can be applied to proteomics datasets, operating on peptide- or protein-derived mass features as observed in chromatography-mass spectrometry data.I. Methods
[0066] One embodiment of the present disclosure encompasses methods for correcting drift in retention times of mass features obtained using mass spectrometry (MS) techniques coupled with separation techniques such as chromatography (chromatography-MS). As used herein, the term “correcting” when referring to drift during retention-time alignment in two or more files comprises correctly aligning the retention times of each mass feature across the files of aPROV PATENT Danforth Docket No.: DDPSC0170-401 -PCT Via EFS-webdataset. Each mass feature is aligned independently of other mass features by performing a series of shifts linear to each mass feature to effectively address and correct the drift. The methods enhance MS data analysis by aligning signals per mass feature so they can be pooled across files, thereby improving signal-to-noise ratio, separating true signal from background noise, and facilitating more reliable identification and analysis without merging or confusing signals from different mass features.
[0067] The methods can align retention times of mass features associated with a single m / z bin before processing steps normally performed on MS data, including peak detection and quantification, calibration, annotation, and noise reduction. By addressing drift at this foundational level, the methods ensure that subsequent data processing is based on accurately aligned mass features. Moreover, this approach to overcoming drift operates independently for mass features associated with a single m / z bin, ensuring that the correction is tailored and precise and avoiding systemic biases often associated with broader correction strategies. These methods are broadly applicable across MS applications, including metabolomics and proteomics, and can be deployed in research and clinical settings as well as other operational contexts where robust, per-feature retention-time alignment is required.
[0068] Two complementary components are employed to implement this approach: (i) Last Resort, which defines refScans and assigns, records, and aligns fileScans per mass feature within an m / z bin; and (ii) a chromolog step that, in files with an ambiguous shift, identifies a chromolog in a different m / z bin and uses its per-file shift to resolve the query mass feature’s assignment. These components are described in detail below.(a) Mass spectrometry
[0069] MS is a powerful analytical tool that can be combined with separation methods such as chromatography (chromatography-MS) for effective detection and quantitation of various chemical species in a sample.
[0070] Chromatography-MS comprises a first separation of components of a sample by chromatography where species with different properties separate and are collected at different times. The time point at which a fraction elutes from the column is called the retention time (RT). The resulting fractions are then introduced into the mass spectrometer, where ions arePROV PATENT Danforth Docket No.: DDPSC0170-401 -PCT Via EFS-webseparated by their mass-to-charge (m / z) ratio. Accordingly, a plurality of m / z signal intensities can be captured by a mass spectrometer in an output file, while intensity data record the abundance of a species at a given m / z relative to RT. The measured data can be depicted in a total ion chromatogram (TIC) image comprising three axes: RT, m / z, and intensity. The combined data of RT, m / z, and intensity for each sample are obtained in a chromatography-MS data file. FIG. 1 shows a representative three-dimensional depiction of UPLC-MS data from a biological sample, displaying axes for RT, m / z, and intensity.
[0071] Methods of the instant disclosure comprise overcoming or correcting drift in retention times of mass features in mass spectrometry by focusing analysis on portions of an MS chromatogram bearing the same m / z bin. Methods for defining, selecting, and implementing m / z bins, including bin widths, tolerances, aggregation rules, and discretization strategies, are well known to individuals of ordinary skill in the art of mass spectrometry and liquid chromatography-mass spectrometry data analysis. For instance, the m / z axis of a file can be partitioned into a plurality of m / z bins, each bin corresponding to a contiguous m / z interval having a fixed or variable width, and an extracted ion chromatogram is generated for a selected bin by extracting signal intensities associated with that interval across retention time. In some embodiments, an m / z bin corresponds to a discrete index, channel, or position within a discretized representation of the m / z axis, such that the bin is identified by its index rather than explicit numerical boundaries. In some embodiments, an m / z bin corresponds to one or more adjacent sampled m / z points produced by the mass spectrometer, optionally aggregated to improve signal-to-noise or computational efficiency. An m / z bin can include signal intensities that are summed, averaged, integrated, or otherwise aggregated within the bin for each retention time point or scan. Unless otherwise stated, references herein to an m / z bin are not limited to any particular binning strategy, tolerance scheme, mass accuracy, data acquisition mode, or instrument type.
[0072] A chromatography-MS data file can comprise a plurality of m / z bins spanning an m / z axis, wherein the number of m / z bins in a given file can range from thousands to tens of thousands or more, depending on the binning method, m / z range, and resolution of the acquired data. The retention-time profile for each m / z bin can be depicted in an extracted ion chromatogram (EIC). An EIC is the intensity-versus-RT signal for one m / z bin, each comprisingPROV PATENT Danforth Docket No.: DDPSC0170-401 -PCT Via EFS-weba characteristic retention-time profile. The retention-time profile displays RT, which is indicative of the time it takes for a mass feature to travel through the chromatographic and MS systems, and intensity, which correlates to the abundance of the ions. The m / z axis of a file can be partitioned into a plurality of m / z bins, each bin corresponding to a contiguous m / z interval having a fixed or variable width, and an EIC can be generated for a selected bin by extracting signal intensities associated with that interval across retention time. An EIC can comprise signal contributions from one or more mass features that fall within the same m / z bin and elute within overlapping or distinct retention-time ranges.
[0073] The method of the instant disclosure is particularly useful in the context of isomers, which are distinct chemical species that share the same nominal m / z but differ in structure. Despite having the same m / z bin, isomers often exhibit different RTs in the chromatographic system due to their distinct physical and chemical properties, allowing separation and identification as distinct mass features within the same EIC. Therefore, an EIC for an m / z bin can provide a profile that includes intensity and RT for one or more mass features (peaks), enabling differentiation and analysis of isomeric species. For instance, it is possible to discern detailed information from the retention-time profile in an EIC, including the presence or absence of isomeric mass features that share an m / z bin, the number of such mass features, and their relative abundance.
[0074] A major challenge in utilizing chromatography-MS for untargeted detection and quantification (e.g., in metabolomics or proteomics) is adjusting for retention-time drift. Existing approaches try to synchronize samples in the RT domain using either all data points or extensive segments of the signal, which can behave unpredictably and can introduce biases. This is particularly problematic for EICs with multiple mass features (e.g., isomeric peaks) that are difficult to differentiate at low signal-to-noise ratios. In individual files, signal can be too weak to distinguish a single mass feature. Pooling files to form composite views can enhance detectability but, when drift is present, leads to blurring and misassignment among mass features; conversely, in the absence of drift, there is a risk of incorrectly interchanging mass features.
[0075] Methods of the instant disclosure comprise independently correcting retention-time drift in each of a plurality of EICs bearing the same m / z bin for each mass feature separately fromPROV PATENT Danforth Docket No.: DDPSC0170-401 -PCT Via EFS-webother mass features, to amplify signals. Since retention time is analyzed per mass feature, this approach simplifies correction to a series of linear shifts applied mass-feature-by-mass-feature, iteratively addressing even non-linearly shifted mass features. In some embodiments, the method comprises defining, for each mass feature, a reference retention time across files (refScan), which is the alignment target selected from evidence across the files (e.g., the scan that is the highest in the most EICs), and, for each file, a per-file peak center (fileScan), which is the apex RT in that file’s EIC selected for mapping to the refScan. The shift for a given refScan-file pair is computed as fileScan retention time minus refScan retention time and is used to align the mass feature along the retention-time axis.
[0076] In some embodiments, the methods comprise obtaining one or more files from a chromatography-MS system that has analyzed a sample. The data in a file can include RT and m / z signal intensity for each data point. As the methods independently correct drift in a plurality of EICs bearing the same m / z bin, in some embodiments a method can further comprise identifying an m / z bin for which an EIC is extracted (e.g., using m / z selection techniques described elsewhere). In some embodiments, a method of the instant disclosure can further comprise identifying or having identified an m / z value most likely to belong to a compound for which an EIC can be extracted. In some embodiments, the m / z value most likely to belong to a compound is identified using methods described in International Application No.PCT / US2022 / 028150, the disclosure of which is incorporated herein in its entirety.
[0077] The methods can be used to correct drift in hundreds, to thousands, to hundreds of thousands of EICs for an m / z value extracted from a plurality of files. The plurality of chromatograms can originate from a diverse range of samples, offering a broad scope for analysis and comparison. These chromatograms can be derived from a single sample, providing a detailed, focused examination of its components. Alternatively, chromatograms can be obtained from multiple samples collected independently, even using different chromatography and MS techniques, allowing for a comparative analysis across a wider range of variables. This versatility is particularly useful in experiments involving various treatments or conditions, or where samples were collected and analyzed using various chromatography-MS techniques. For instance, samples from different treatment groups in a study can be analyzed to compare and contrast the effects of these treatments at a molecular level. Similarly,PROV PATENT Danforth Docket No.: DDPSC0170-401 -PCT Via EFS-websamples under different experimental conditions can provide insights into how these conditions influence the chemical composition of the samples.
[0078] Methods of the instant disclosure can be used to overcome drift in retention times in chromatography-MS runs for untargeted analyses of any mass feature signal, including those from chemical, pharmaceutical, biochemical, metabolomic and proteomic applications.Chromatography-MS can be applied across a broad range of fields and applications, including but not limited to clinical diagnostics and clinical testing (e.g., biomarker detection, therapeutic drug monitoring, toxicology screening, and disease profiling), proteomics (e.g., processing and comparison of peptide- or protein-derived mass spectrometry data, including retention-time alignment and signal normalization in proteomics workflows for use in protein identification, peptide mapping, post-translational modification analysis, and quantitative proteome profiling), metabolomics (e.g., metabolite identification, pathway analysis, and comparative metabolic profiling), lipidomics, glycomics, pharmaceutical research and development (e.g., drug discovery, pharmacokinetics, and stability studies), environmental analysis (e.g., detection of pollutants, contaminants, and trace chemicals), food and agricultural analysis (e.g., quality control, adulterant detection, and residue analysis), forensic science, chemical biology, systems biology, and industrial process monitoring. In these and other applications, chromatography-MS datasets can comprise large volumes of high-dimensional data generated across many samples, retention times, and mass-to-charge (m / z) values.
[0079] Non-limiting examples of chemical, biochemical, and metabolomic compounds with mass features that can be usefully analyzed using the disclosed methods and systems include biochemical compounds such as amino acids, peptides and proteins, nucleotides and nucleosides, DNA and RNA fragments, lipids (fatty acids, triglycerides, phospholipids, steroids), carbohydrates (monosaccharides, disaccharides, polysaccharides), vitamins, hormones, and enzymes; metabolomic compounds such as metabolic intermediates (glycolysis intermediates, Krebs cycle intermediates), neurotransmitters, plant metabolites (alkaloids, terpenes, flavonoids), bacterial and fungal metabolites, endogenous metabolites (bile acids, urea cycle intermediates), and xenobiotics (drugs, toxins); environmental contaminants such as pesticides and herbicides, polycyclic aromatic hydrocarbons (PAHs), persistent organic pollutants (POPs), and heavy metals and metalloids (in their organic forms);PROV PATENT Danforth Docket No.: DDPSC0170-401 -PCT Via EFS-webpharmaceuticals such as therapeutic drugs, drug metabolites, antibiotics; polymeric materials such as polymers and copolymers, additives and plasticizers, and oligomers; isotopically labeled compounds such as stable isotope labeled metabolites, deuterated compounds; industrial chemicals such as dyes and pigments, surfactants and detergents, and explosives; food and beverage analysis such as food additives, flavor compounds, contaminants; forensic analysis such as illicit drugs, explosive residues, and trace evidence compounds; and geological and cosmochemical analysis such as elemental isotopes, and organic molecules in extraterrestrial samples.
[0080] In some embodiments, methods of the instant disclosure can be used to overcome drift in retention times of chromatography-MS runs for analyses of metabolites in metabolomic experiments. While specific types of compounds (e.g., metabolites such as fructose and galactose, metabolites of GABA, or amino acid isomers) can be discussed herein, such discussion of embodiments is for illustrative purposes and should not be interpreted as limiting the present disclosure to the embodiments being illustrated and discussed.
[0081] In some embodiments, methods of the instant disclosure can be used to overcome drift in retention times of chromatography-MS runs for analyses of amino acids, peptides, and proteins in proteomic experiments. While specific types of amino acids, peptides, and proteins (e.g., isoleucine (He) and leucine (Leu)) can be discussed herein, such discussion of embodiments is for illustrative purposes and should not be interpreted as limiting the present disclosure to the embodiments being illustrated and discussed.
[0082] As used herein, when chromatography-mass spectrometry data are derived from peptides or proteins, the disclosed methods operate on the resulting mass features observed in the data, rather than on peptides or proteins as biochemical entities. In proteomic experiments, peptides and proteins are detected by mass spectrometry as one or more chromatographic peaks characterized by an m / z bin and a retention-time apex, and each such peak is treated as an independent mass feature for purposes of retention-time alignment. The methods described herein therefore correct retention-time drift of peptide- or protein-derived mass features using observed m / z and retention-time information, without requiring identification, sequencing, or biochemical interpretation of the peptides or proteins.PROV PATENT Danforth Docket No.: DDPSC0170-401 -PCT Via EFS-web
[0083] There are multiple categories of chromatography and MS, and different combinations cater to differing analytical needs. Non-limiting chromatographic techniques include LC-MS, GC-MS, IC-MS, SFC-MS, and CE-MS. Liquid chromatography methods include HPLC, UPLC, ion-exchange chromatography, size exclusion chromatography, normal phase, reverse phase, chiral chromatography, and affinity chromatography. Gas chromatography methods include capillary GC, packed-column GC, and gas-solid chromatography. Ion chromatography is typically used for ions and polar molecules. SFC-MS uses a supercritical fluid as the mobile phase. CE separates ions by charge-to-size ratio in an electric field.
[0084] Non-limiting examples of major categories of chromatographic techniques that can be coupled with MS include Liquid Chromatography-Mass Spectrometry (LC-MS), Gas Chromatography-Mass Spectrometry (GC-MS), Ion Chromatography-Mass Spectrometry (IC-MS), Supercritical Fluid Chromatography-Mass Spectrometry (SFC-MS), Capillary Electrophoresis-Mass Spectrometry (CE-MS). Liquid chromatography methods can include High-Performance Liquid Chromatography (HPLC) - the most common form, utilizing high pressure to pass the sample through a column filled with stationary phase; Ultra-Performance Liquid Chromatography (UPLC) - similar to HPLC but uses smaller particle sizes in the column for higher resolution and faster analysis; Ion Exchange Chromatography - separates ions based on their affinity to an ion exchanger in the column; Size Exclusion Chromatography (SEC) - separates molecules based on size, using a column with pores of a specific size;Normal Phase Chromatography - uses a polar stationary phase and a non-polar mobile phase; Reverse Phase Chromatography - uses a non-polar stationary phase and a polar mobile phase, opposite to normal phase; Chiral Chromatography - separates enantiomers based on their interaction with a chiral stationary phase; Affinity Chromatography - uses a stationary phase made of materials that specifically bind to the analyte of interest. Gas chromatography methods include Capillary Gas Chromatography - uses very narrow capillary tubes with a liquid stationary phase; Packed Column Gas Chromatography - utilizes columns packed with solid stationary phase or solid support coated with liquid stationary phase; Gas-Solid Chromatography (GSC) - involves a solid stationary phase and is used primarily for separating gases or volatile compounds that don't interact with liquid stationary phases. Ion chromatography is typically used for the separation of ions and polar molecules. SupercriticalPROV PATENT Danforth Docket No.: DDPSC0170-401 -PCT Via EFS-webFluid Chromatography-Mass Spectrometry (SFC-MS) uses a supercritical fluid (like C02) as the mobile phase, combining aspects of both GC and LC. It's particularly effective for analyzing compounds that are difficult to separate by traditional LC, such as chiral compounds. Capillary Electrophoresis: Although not technically a chromatographic technique, capillary electrophoresis separates ions based on their charge-to-size ratio in an electric field.
[0085] Non-limiting examples of MS include Quadrupole Mass Spectrometry (QMS) - uses quadrupole filters for mass analysis, suitable for a broad range of masses; Time-of-Flight Mass Spectrometry (TOF-MS); separates ions by their different flight times; Ion Trap Mass Spectrometry - traps ions using electromagnetic fields and then sequentially ejects them for mass analysis; Fourier Transform Ion Cyclotron Resonance (FT-ICR) - offers very high resolution and accuracy, using a magnetic field to trap ions; Orbitrap Mass Spectrometry -uses an electrostatic field to trap ions in an orbital motion around a central electrode; Triple Quadrupole Mass Spectrometry (QqQ) which incorporates three quadrupoles in series; commonly used for quantification due to its high sensitivity and specificity; Tandem Mass Spectrometry (MS / MS) which involves multiple stages of mass spectrometry, often with fragmentation of analyte ions between stages; Quadrupole Time-of-Flight Mass Spectrometry (Q-TOF) which combines quadrupole mass filtering with TOF mass analysis for high accuracy and resolution; and Magnetic Sector Mass Spectrometry which uses a magnetic field to deflect ions, with separation based on mass-to-charge ratio.
[0086] In some embodiments, methods of the instant disclosure can be used to correct drift in retention times of LC-MS runs. In some embodiments, methods of the instant disclosure can be used to overcome drift in retention times of LC-MS runs for analyses of metabolites. In some embodiments, methods of the instant disclosure can be used to overcome drift in retention times of LC-MS runs for analyses of amino acids, peptides, and proteins in proteomic applications.
[0087] In addition to chromatography, different separation techniques can also be used in conjunction with mass spectrometry. Non-limiting examples of suitable separation techniques other than chromatography include electrophoresis, ion mobility, Field-Flow Fractionation (FFF), Capillary Electrophoresis (CE), Matrix-Assisted Laser Desorption / lonization (MALDI), Thermal Desorption (TD), Laser Ablation (LA), Desorption Electrospray Ionization (DESI).PROV PATENT Danforth Docket No.: DDPSC0170-401 -PCT Via EFS-webElectrophoresis separates charged particles under an electric field, effectively used for biomolecules like DNA and proteins. Ion mobility spectrometry, on the other hand, separates ions based on their mobility in a gas phase under an electric field, useful for distinguishing isomers and conformers. FFF techniques separate particles, macromolecules, or colloids based on their size or density in a fluid under various field forces (such as centrifugal, thermal, or electrical fields). Coupling FFF with MS allows for the analysis of large and complex molecules like proteins and polymers. CE is similar to electrophoresis but conducted in capillary tubes. CE is effective for separating ionic species using an electric field. Its coupling with MS provides high resolution and efficiency, particularly useful for analyzing biomolecules like peptides and nucleotides. Although MALDI is more of an ionization technique than a separation technique, it is often mentioned in the context of MS coupling. It's particularly useful for the analysis of large biomolecules like proteins, DNA, and polymers. TD involves heating a sample to release volatile and semi-volatile compounds. When coupled with MS, it allows for the analysis of compounds in air, materials, and environmental samples. LA comprises using a laser to remove material from a solid sample. Coupling LA with MS enables the analysis of solid samples, particularly in fields like geology and material science, allowing for elemental and isotopic analysis. DESI allows for the direct analysis of samples (even from surfaces) under ambient conditions. It's useful for a wide range of applications, including biological tissues and environmental samples.(b) Locally aligning shifts that resort (Last Resort)
[0088] One embodiment of the instant disclosure encompasses a method for correcting drift in retention times of mass features in MS by locally aligning mass features in an m / z bin of interest. Prior to the development of methods of the instant disclosure, several processing steps were normally necessary to convert MS data into a format suitable for analysis, annotation, correction of retention time drift, and signal amplification. These steps, essential for each Chromatography-MS run, included peak detection and quantification, calibration, and noise reduction. However, these traditional methods face challenges due to the often faint and fluctuating outlines of mass features across samples, characterized by a low signal-to-noise ratio such as in metabolomic studies. These approaches can also introduce unintendedPROV PATENT Danforth Docket No.: DDPSC0170-401 -PCT Via EFS-webwarping effects because drift isn’t being corrected on an individual basis for each mass feature. This problem is particularly acute when the signal of one mass feature is closely aligned with another in the space of signals, with each experiencing its own distinct chromatographic drift. As a result, these conventional methods can inadvertently discard valid mass feature signals and tend to be computationally demanding and time-consuming. The complexity of MS data is depicted in the merged retention time profiles of EICs of a particular m / z bin shown in FIG. 2.
[0089] Contrary to previous strategies for amplifying signals, methods of the instant disclosure correct drift using raw data obtained from LC-MS runs, bypassing the standard preprocessing steps such as mass spectra processing for peak detection and quantification, calibration, and noise reduction. This represents a significant departure from previous strategies by focusing on signal amplification through direct manipulation of unprocessed data. Critically, the methods of the instant disclosure also depart from previous methods by correcting drift on an individual basis for each mass feature by aligning each mass feature separately.
[0090] Methods of the instant disclosure create systems of linear shifts local to each mass feature. In other words, methods of the instant disclosure iteratively perform a series of shifts linear to each mass feature, even if the mass feature is non-linearly shifted. This series of shifts provides several benefits to MS data analysis. Firstly, it amplifies the signals by pooling data from multiple EICs, thereby increasing the signal-to-noise ratio. Secondly, by aligning each mass feature individually, the methods can capture even weak signals of mass features that might otherwise be lost as background noise using other methods.
[0091] The methods comprise several steps. The methods first comprise generating or having generated an extracted ion chromatogram (EIC) in each file of a dataset comprising a plurality of files. In some embodiments, methods of the instant disclosure comprise generating an EIC for a given m / z bin in each file of a dataset comprising a plurality of files.
[0092] For a given m / z bin, data points in EICs across the files of the dataset can be divided into scans based on retention-time increments. The range of values for the retention time increments of the scans within the 2D grid system can and will vary based on the distribution of obtained data values, the characteristics of the retention time profile, the shape and number of waveforms, the number of data point values in the EICs, and the desired level of precision forPROV PATENT Danforth Docket No.: DDPSC0170-401 -PCT Via EFS-webdefining templates. In some embodiments, scans can be categorized by the duration of RT increments. For instance, scans can comprise retention time increments of durations ranging from about 0.1 second or less to about 15 minutes or more, from about 1 second to about 10, 9, 8, 7, 6, 5, 4, 3, 2, 1 minutes or less, from about 1 second to about 50, 40, 30, 20, 10, 54, 3, or 2 seconds, or from about 30 seconds to about 1 , 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or about 20 minutes or longer. Scans can also be categorized by the total number of scans desired in a merged EIC. For instance, a 2D grid system can also comprise 10, 15, 20, 25, 30, 35, 40, 45, or about 50 or more scans. Furthermore, scans can be defined by the number of data point values in each retention time increment of a scan. For instance, scans can be defined as a range of RT increments, wherein all RT increments can comprise the same number of data point values.
[0093] In some embodiments, the scans comprise retention time increments categorized by the duration of each RT increment. In some embodiments, the scans comprise retention time increments of equal durations.
[0094] The methods then comprise locally aligning mass features in an m / z bin of interest. Locally aligning mass features comprises selecting a first reference retention time across files (refScan) as the scan comprising the highest intensity data points of the highest number of EICs across the files of the dataset, thereby defining a first mass feature. A refScan establishes the central tendency for a mass feature. As used herein, “central tendency” refers to the across-files reference position of a mass feature’s peak in retention time.
[0095] In some embodiments, defining a mass feature further comprises establishing retentiontime bounds centered on the first refScan within which data points are identified and attributed to the first mass feature in each file. Any suitable method for establishing retention-time bounds of a mass feature can be used, including approaches based on peak-shape criteria (e.g., inflection points, width-at-half-maximum), smoothing / spline fitting, or signal-to-noise thresholds, as are known to individuals of skill in the art. For each remaining EIC in which a per-file peak center (fileScan) does not initially coincide with the refScan, the fileScan is assigned to the refScan and the data points of that EIC are shifted by a time increment sufficient to bring the fileScan to the refScan.PROV PATENT Danforth Docket No.: DDPSC0170-401 -PCT Via EFS-web
[0096] The method continues by iteratively defining an nth+ 1 mass feature: select an additional refScan using the same criterion, exclude data points within the bounds of previously defined mass features, assign the nth+ 1 fileScan of each remaining EIC to the nth+ 1 refScan, and shift the data points of each such EIC by a time increment sufficient to align the fileScan with its assigned nth+ 1 refScan, thereby establishing bounds for the nth+ 1 mass feature. In some embodiments, data points of previously defined mass features can be excluded by annotating each file of the plurality of EICs with each defined mass feature, peak center of each defined mass feature, bounds of each defined mass feature, data points of each defined mass feature, or any combination thereof in each file of each EIC, and excluding the annotated data points from this second step.
[0097] Subsequently, based on the recorded refScan-file pairings, the method aligns each defined mass feature by shifting, in the original unshifted EIC of each file, the data points already attributed to that mass feature along the retention-time axis by a time increment sufficient to bring the recorded fileScan to the closest refScan in that file (by retention-time difference), thereby correcting the drift in retention times of mass features.
[0098] In some embodiments, locally aligning mass features in an m / z bin of interest comprises (a) selecting a first reference retention time across files (refScan) as a scan comprising the highest intensity data points of the highest number of EICs across the files of the dataset among the plurality of EICs extracted from each file in the dataset and establishing retentiontime bounds centered on the first refScan within which data points are identified and attributed to the first mass feature in each file, thereby defining a first mass feature; (b) assigning a perfile peak center (fileScan) of each remaining EIC of the plurality of EICs to a nearest refScan selected in step (a), wherein a remaining EIC is an EIC in which the fileScan does not coincide with the refScan of the first mass feature; (c) shifting the data points of each remaining EIC by a time increment sufficient to align the fileScan of the EIC with its assigned refScan; (d) iteratively repeating steps (a) through (c) for each nth+1 refScan to define an nth+1 mass feature, thereby assigning each remaining EIC’s nth+1 fileScan to that refScan, shifting to align, and recording bounds; wherein each iteration uses an unshifted data file to identify the nth+1 mass features, wherein the nth+1 selection excludes data points within the bounds of previously defined mass features, and wherein n is an integer having a value equal to or morePROV PATENT Danforth Docket No.: DDPSC0170-401 -PCT Via EFS-webthan 1; (e) recording in a shiftTable, for every refScan-file pair, the retention time of the fileScan, the retention time of the refScan to which the fileScan is assigned, a shift equal to fileScan retention time minus refScan retention time, a mass feature identifier, and optionally the m / z bin for the mass feature, a file identifier, and one or more quality flags; and (f) aligning each nthmass feature defined in steps (a) through (e), based on the shiftTable of step (e) and using the original unshifted EIC of each data file, by shifting the data points attributed to that nthmass feature along the retention-time axis by a time increment sufficient to bring the recorded fileScan of that nthmass feature to the closest refScan in that file; thereby correcting the drift in retention times of mass features.
[0099] In some embodiments, the quality flags comprise at least: a proxi mity-to-refS can flag indicating whether the fileScan retention time lies within predefined bounds around the refScan; a signal-to-noise flag indicating the apex signal-to-noise ratio at the fileScan; an ambiguity flag indicating whether multiple candidate apexes were present near the refScan; and a correspondence flag indicating presence or absence of a credible fileScan for the refScan in the file. In some embodiments, the quality flags further comprise a peak shape flag indicating whether one or more peak-shape measures, including symmetry, tailing factor, or width-at-half-maximum, fall within predetermined ranges, an intensity threshold flag indicating whether the apex intensity at the fileScan meets predetermined minimum or maximum thresholds, or both. In some embodiments, the quality flags further comprise an overlap or coelution flag indicating evidence of overlapping mass features within the refScan bounds, an outlier flag indicating whether the fileScan retention time or apex intensity is outside a robust distribution for the mass feature across the plurality of files, a drift magnitude flag indicating categorical magnitude of the shift relative to a threshold, or any combination thereof. In some embodiments, the quality flags further comprise an overlap or co-elution flag indicating evidence of overlapping mass features within the refScan bounds, an outlier flag indicating whether the fileScan retention time or apex intensity is outside a robust distribution for the mass feature across the plurality of files, a drift magnitude flag indicating categorical magnitude of the shift relative to a threshold, or any combination thereof.A. Diagrammatic illustrationPROV PATENT Danforth Docket No.: DDPSC0170-401 -PCT Via EFS-web
[0100] Last Resort methods of the instant disclosure can be diagrammatically illustrated in FIGs. 3-9 depicting MS data obtained from a metabolomic analysis, wherein a mass feature represents a compound, an isomer thereof, or a derivative thereof. FIG. 3 diagrammatically depicts merged EICs of a particular m / z bin of MS data obtained from samples comprising two isomers of a hypothetical compound, each isomer depicted as a mass feature. The top panel of FIG. 3 represents the merged EICs before any adjustment of drift, showing how complex and chaotic MS data can be before adjusting for drift. The presence of linear drift is illustrated by the yellow chromatogram, and the presence of non-linear drift is represented by the second isomers of the light and dark grey chromatograms in the top panel of FIG. 3. The green chromatograms do not exhibit any drift. In individual files, peak centers of isomer A are noted by the black vertical lines, and peak centers of isomer B are noted by the red vertical lines. The bottom panel of FIG. 3 depicts the result of the deconvolution of the MS data in the merged EICs when using the methods of the instant disclosure. As seen in the bottom panel, data points of both isomers A and B are correctly aligned with their corresponding mass features, correcting both linear and non-linear drift. By separately aligning each peak center of an isomer in all files, a method central to the instant disclosure, simple linear shifts can be used to align mass features that comprise both linear and non-linear drift. This process simplifies and organizes the data, as shown in the bottom panel of FIG. 3. Thus, the once entangled signals can be effectively adjusted to pool all data of each mass feature, demonstrating the utility of the methods disclosed herein in deconvoluting complex MS data for further analysis.
[0101] Importantly, methods of the instant disclosure enable the detection of mass features with faint signals that might otherwise be obscured by background noise. An example of such faint signals is depicted by the second mass feature (isomer B) of the dark grey chromatogram in FIGs. 3-9. Using the methods disclosed herein, these faint signals are accurately identified and recorded in the file. These data points are then merged with the identified peak centers of the second mass feature when the methods are applied. Additionally, the faint signal contributes to the cumulative signal of a mass feature, thereby enhancing the overall quality and accuracy of the data.
[0102] To deconvolute signals in a merged EIC, the methods disclosed herein can begin by providing or having provided files that contain data point values of EICs from portions ofPROV PATENT Danforth Docket No.: DDPSC0170-401 -PCT Via EFS-webeach of a plurality of MS chromatograms that bear the same m / z bin. In this diagrammatic example, each EIC depicted in FIGs. 3-9 comprises two isomers. According to the methods disclosed herein, the data points of each file can be divided into a plurality of scans, each comprising retention-time increments denoted as scans A-l in FIGs. 4-8.
[0103] Methods of the instant disclosure further comprise defining a first mass feature (isomer A). Defining a first mass feature can comprise selecting a first refScan as a scan comprising the highest intensity data points of the highest number of EICs among the plurality of EICs. In the diagrammatic representation of the Last Resort method diagrammatically depicted FIGs. 3-9, the first refScan would be scan C. As shown in FIG. 4, the first refScan (scan C) comprises the first peak center of the first mass feature (isomer A). The fileScan of the linearly shifted first mass feature of the yellow chromatogram, before any alignment is performed, is depicted in scan D in the top panel of FIG. 4. The fileScan of the first mass feature in the yellow chromatogram can then be assigned to the first refScan. Other drifted mass features can also be observed in the top panel of FIG. 4, but as explained herein above, these mass features can be ignored for the moment.
[0104] The methods then comprise shifting the data points of each remaining EIC (the yellow, blue, and light grey EIC by a time increment sufficient to align the fileScan of the EIC with the first refScan (scan C). The outcome of this step is shown in the bottom panel of FIG.4, where the data is significantly refined for the first mass feature. As can be seen for the yellow EIC, the peak center of the first mass feature is now aligned with the first peak center, thereby substantially improving the clarity and accuracy of the data. However, this step also results in the incorrect alignment of isomer B of the blue and light grey chromatograms with isomer A. This outcome is expected at this intermediate stage and is not problematic, as it will be addressed in subsequent steps.
[0105] Methods of the instant disclosure can then comprise repeating the alignment steps for isomer B with the important distinction that in each iteration after the first alignment, the selection excludes data points within the bounds of previously defined mass features. In this exemplary situation, the data points of the first mass feature (shown as hatched data points in FIG. 5) are excluded, resulting in the remaining data points shown in FIG. 6. In the exemplary situation depicted in FIGs. 3-9, the second refScan corresponds to the second peakPROV PATENT Danforth Docket No.: DDPSC0170-401 -PCT Via EFS-webcenter of the second mass feature in scan G, as shown in FIGs. 7 and 8. As a result of this step, the data points of the second isomers are aligned, as illustrated in FIGs. 7 and 8.However, similar to the alignment of the first peak centers, in some EICs, this step also leads to the incorrect alignment of isomer A of the blue and light grey chromatograms with isomer B.
[0106] In some embodiments, the data points of the previously defined mass features can be excluded by simply removing the hatched data points from the files in this step (FIG. 6).As shown in the top panel of FIG. 7, none of the hatched data points are included in the alignment of the second mass feature (isomer B). In other embodiments, the data points of the previously defined mass features can be excluded by simply ignoring (but not removing) the hatched data points in the files as shown in the top panel FIG. 8.
[0107] This illustrates a key advantage of the methods disclosed herein. By identifying the peak center of each mass feature, including those with weak signals, even faint data point values that might otherwise be lost in the background can be captured and added to the merged data. In the example depicted in FIGs. 3-9, this method ensures that the faint data point signals of the second isomer of the dark grey EIC are identified, preventing their loss due to low signal. Identified data points of the second isomer of the dark grey EIC are visually represented in FIGs. 7 and 8 by green outlining.
[0108] As explained above, this step also inadvertently leads to the swapping of isomers or mass features in some EICs, particularly evident in the blue and light grey second mass features, which now have swapped isomers A and B. This misalignment is visually represented by the red outlines around the data points of the swapped second mass feature in the blue and light grey chromatograms as shown in FIGs. 7 and 8. As noted above, this outcome is expected because, across files, peak height ordering can differ, and the described logic can temporarily map fileScans to the wrong refScan under those conditions. The methods described in this disclosure include subsequent steps specifically designed to rectify these swaps, ensuring the correct alignment of the first and second isomers in the blue and light grey chromatograms.
[0109] As explained herein below, although the example depicted in FIGs. 3-9 comprises only two isomers, the process of identifying peak centers or mass features can be iteratively applied after each round of identifying the nth+ 1 peak center until no additional peakPROV PATENT Danforth Docket No.: DDPSC0170-401 -PCT Via EFS-webcenters or mass features can be detected. Of course, it should be noted that when a sample comprises a single isomer, subsequent steps of identifying peak centers and defining mass features are not necessary.
[0110] In some embodiments, aligning the peak centers can comprise recording, for every refScan-file pair, the retention time of the fileScan, the retention time of the refScan to which the fileScan is assigned, a shift equal to fileScan retention time minus refScan retention time, a mass feature identifier, and optionally the m / z bin for the mass feature, a file identifier, and one or more quality flags. A shift that decreases retention time (RT) is a negative shift, and a shift that increases RT is a positive shift. The sum of shifts can be calculated for each identified mass feature, peak center of each mass feature, bounds of each mass feature, data point values of each mass feature, or any combination thereof to generate a total shift value for each identified mass feature, peak center of each mass feature, bounds of each mass feature, data point values of each mass feature, or any combination thereof. The mass feature, peak center of each mass feature, bounds of each mass feature, data point values of each mass feature, or any combination thereof can be shifted by the total shift value to thereby align each mass feature with its refScan. In some embodiments, shifts can be noted in a shiftTable, wherein the table comprises: each identified mass feature; the RT of the peak center of each identified mass feature, the bounds of each identified mass feature, or both; values of shifts of each refScan of each identified mass feature at every iteration; and optionally the sum of all shifts for each mass feature. An example of such a table can be as shown in Table A. In some embodiments, the sum of the shifts can be recalculated at every iteration for each mass feature. In other embodiments, the sum of the shifts can be calculated at the end of all the shifts.PROV PATENT Danforth Docket No.: DDPSC0170-401 -PCT Via EFS-web
[0111] Once all the identified mass features, peak centers of each mass feature, bounds of each mass feature, data point values of each mass feature, or any combination thereof have been identified and data points annotated in the files, the method can further comprise aligning the fileScan of each identified mass feature with the closest refScan in the file to thereby alignPROV PATENT Danforth Docket No.: DDPSC0170-401 -PCT Via EFS-webthe retention times of the mass features. In one embodiment, the data in Table A can be used to determine the closest refScan to the starting retention time of the fileScan of each mass feature in a given file (column 1 ) and the original data to correct any mass features with swapping in order by shifting the data in that file at that mass feature until it aligns with the correct and closest refScan. In another embodiment, Table A and the series of shifts can be used to recover the refScan at which a scrambled mass feature was originally located and then calculate the shift to align it with the proper mass feature. Therefore, the swapping can be regarded as inconsequential, because the goal of the preceding steps is to discover, in each file, all relevant fileScans as refScans are identified, so that each fileScan can subsequently be matched or assigned to the closest refScan. This step corrects any swapping that can have occurred (e.g., blue and light grey EICs in FIGs. 3-9). As illustrated in FIG. 9, when all data points have been assigned to a mass feature, as shown in the top panel of FIG. 9, aligning each mass feature with its closest refScan corrects the swapping observed for the blue and light grey EICs, as depicted in the lower panel of FIG. 9.
[0112] FIGs. 2 and 10-13 present a real-world example of using the methods of the instant disclosure to align retention times of mass features of compound MS signals, utilizing data of all the isomers at the mass of GABA in human microbiome samples. FIG. 2 shows data from approximately 600 pooled files before the methods of the invention are applied. In the case of each individual EIC file among these samples, the signals of compounds might be too faint to detect, posing a challenge in identifying specific compounds. However, an attempt to overcome this by pooling the EIC files together to amplify the signals leads to a different problem: the details become indistinct or 'blur,' as shown in FIGs. 2 and 10, depicting data points in a merged EIC file before any corrections are made according to the methods of the instant disclosure. FIGs. 11A-11C are zoomed-in views of the first, second, and third portions, respectively, of the pooled data of FIG. 2, divided into the early (FIG. 11 A), middle (FIG. 11 B), and late (FIG. 11C) sections of the pooled EIC of FIG. 2.
[0113] In FIGs. 2 and 10 a fundamental dilemma in traditional mass spectrometry data analysis is highlighted: the data requires deconvolution for peak identification, but paradoxically, identifying these peaks is a prerequisite for effective deconvolution. While the pooled data suggests the presence of peaks corresponding to isomers, the lack of clarity andPROV PATENT Danforth Docket No.: DDPSC0170-401 -PCT Via EFS-webdefinition in this data makes it challenging to accurately identify and distinguish these peaks and their associated isomers. In this example, methods of the instant disclosure are applied to iteratively but independently align peak centers of mass features and identifying the mass features. The identified mass features can then be aligned with the correct GABA isomer to thereby align retention times of the mass features. Using an embodiment of methods of the instant disclosure, the data point values of each data file are divided into a plurality of scans comprising retention time increments. A first mass feature is aligned to correct the drift of the first mass feature. The first peak center is shown by the purple line in FIG. 12A which corresponds to the FIG. 11A portion of the merged EIC of FIGs. 2 and 10, before a method of the instant disclosure is applied. FIG. 12B shows the chromatogram after aligning first peak centers and defining a first mass feature, here shown between the red and green vertical lines. In an embodiment of the methods of the instant disclosure used in this example, the data points of the first mass feature are excluded as shown in FIG. 12C, and the method is repeated to find additional peak centers and mass features.
[0114] In a non-limiting proteomics application, a chromatography-MS dataset may comprise LC-MS runs of digested protein samples in which peptides are detected as multiple mass features across files. Each peptide-derived mass feature exhibits a retention-time profile that may drift across runs due to chromatographic variation. Using the methods of the instant disclosure, retention-time drift of each peptide-derived mass feature is corrected independently by selecting reference retention times across files and applying per-feature shifts, thereby enabling pooling and comparison of peptide-derived signals across datasets without requiring peptide identification, sequencing, or fragmentation analysis.B. Embodiments
[0115] Methods of the instant disclosure comprise a method for correcting drift in retention times of mass features in MS by locally aligning mass features in an m / z bin of interest. In some embodiments, the methods comprise obtaining or having obtained one or more data files from a chromatography-MS machine that has analyzed samples of unknown compounds. In some embodiments, the chromatography-MS is LC-MS. In some embodiments, methods of the instant disclosure can be used to overcome drift in retentionPROV PATENT Danforth Docket No.: DDPSC0170-401 -PCT Via EFS-webtimes of chromatography-MS runs for analyses of metabolites. In some embodiments, methods of the instant disclosure can be used to overcome drift in retention times of LC-MS runs for analyses of metabolites.
[0116] In some embodiments, methods of the instant disclosure comprise, for an m / z bin, generating or having generated an extracted ion chromatogram (EIC) in each file of a dataset comprising a plurality of files.
[0117] The methods of the instant disclosure comprise selecting a first reference retention time across files (refScan). A refScan is a scan comprising the highest intensity data points of the highest number of EICs across the files of the dataset among the plurality of EICs extracted from each file in the dataset. The range of values for the retention time increments of the scans can and will vary depending on the distribution of obtained data values, the characteristics of the retention time profile, the shape and number of waveforms, the number of data point values in the EICs, the desired level of precision for defining templates, or any combination thereof. In some embodiments, scans can be categorized by the duration of RT increments. For instance, scans can comprise retention time increments of durations ranging from about 0.1 second or less to about 15 minutes or more, from about 1 second to about 10, 9, 8, 7, 6, 5, 4, 3, 2, 1 minutes or less, from about 1 second to about 50, 40, 30, 20, 10, 54, 3, or 2 seconds, or from about 30 seconds to about 1 , 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or about 20 minutes or longer. Scans can also be categorized by the total number of scans desired in a merged EIC. For instance, an EIC can be divided into 10, 15, 20, 25, 30, 35, 40, 45, or about 50 or more scans. Furthermore, scans can be defined by the number of data point values in each retention time increment of a scan. For instance, scans can be defined as a range of RT increments, wherein all RT increments can comprise the same number of data point values.
[0118] In some embodiments, the methods further comprise establishing retention-time bounds centered on the first refScan within which data points are identified and attributed to the first mass feature in each file, thereby defining a first mass feature. In some embodiments, the retention time bounds of each mass feature is determined using a smoothing algorithm. The mass features can be metabolites and isomers of metabolites. The mass features can also be amino acids, peptides, and proteins.PROV PATENT Danforth Docket No.: DDPSC0170-401 -PCT Via EFS-web
[0119] Methods of the instant disclosure further comprise assigning a per-file peak center (fileScan) of each remaining EIC of the plurality of EICs to a nearest refScan selected in the previous step. A remaining EIC is an EIC in which the fileScan of a mass feature does not coincide with the refScan of the first mass feature. The data points of each remaining EIC are then shifted by a time increment sufficient to align the fileScan of the EIC with its assigned refScan. This step aligns the fileScans of the remaining EICs. In some embodiments, data points that fall within the established bounds are attributed to the first mass feature
[0120] The selection of refScans and aligning remaining EICs is then iteratively repeated for each nth+1 refScan to define an nth+1 mass feature, thereby assigning each remaining EIC’s nth+1 fileScan to that refScan, shifting to align, and recording bounds. It is noted that n is an integer having a value equal to or more than 1. Importantly, each iteration uses an unshifted data file to identify the nth+1 mass features and the nth+1 selection excludes data points within the bounds of previously defined mass features. In some embodiments, the retention time of the fileScan for every refScan-file pair is recorded in a shiftTable. In the shiftTable, the retention time of the refScan to which the fileScan is assigned can be recorded as well as a shift equal to fileScan retention time minus refScan retention time, a mass feature identifier, and optionally the m / z bin for the mass feature, a file identifier, and one or more quality flags. In some embodiments, the unshifted data file is annotated with each mass feature, peak center of each mass feature, bounds of each mass feature, data point values of each mass feature, or any combination thereof. Excluding data points of previously defined mass features can comprise removing identified data points or simply not including the data points. Identifying data points of a mass feature to be excluded is explained herein further below.
[0121] In some embodiments, the quality flags comprise at least: a proximity-to-refScan flag indicating whether the fileScan retention time lies within predefined bounds around the refScan; a signal-to-noise flag indicating the apex signal-to-noise ratio at the fileScan; an ambiguity flag indicating whether multiple candidate apexes were present near the refScan; and a correspondence flag indicating presence or absence of a credible fileScan for the refScan in the file. In some embodiments, the quality flags further comprise a peak shape flag indicating whether one or more peak-shape measures, including symmetry, tailing factor, orPROV PATENT Danforth Docket No.: DDPSC0170-401 -PCT Via EFS-webwidth-at-half-maximum, fall within predetermined ranges, an intensity threshold flag indicating whether the apex intensity at the fileScan meets predetermined minimum or maximum thresholds, or both. In some embodiments, the quality flags further comprise an overlap or coelution flag indicating evidence of overlapping mass features within the refScan bounds, an outlier flag indicating whether the fileScan retention time or apex intensity is outside a robust distribution for the mass feature across the plurality of files, a drift magnitude flag indicating categorical magnitude of the shift relative to a threshold, or any combination thereof. In some embodiments, the quality flags further comprise an overlap or co-elution flag indicating evidence of overlapping mass features within the refScan bounds, an outlier flag indicating whether the fileScan retention time or apex intensity is outside a robust distribution for the mass feature across the plurality of files, a drift magnitude flag indicating categorical magnitude of the shift relative to a threshold, or any combination thereof.
[0122] The methods of the instant disclosure then comprise aligning each nthmass feature defined in the previous steps, based on the shiftTable and using the original unshifted EIC of each data file, by shifting the data points attributed to that nthmass feature along the retention-time axis by a time increment sufficient to bring the recorded fileScan of that nthmass feature to the closest refScan in that file; thereby correcting the drift in retention times of mass features. In some embodiments, an aligned nthpeak center is the peak center of a distribution curve that encompasses data points of an nthmass feature, thereby defining the nthmass feature.
[0123] FIG. 13 is a flowchart illustrating various steps of the Last Resort method to overcome drift in retention times of mass features in MS (e.g., LC-MS). Step 502: for an m / z bin, generate or having generated an extracted ion chromatogram (EIC) in each file of a dataset comprising a plurality of files. Step 504: divide EIC data points into scans comprising retention-time increments. Step 506: select a first reference retention time across files (refScan) as the scan comprising the highest intensity data points of the highest number of EICs; thereby define a first mass feature by establishing retention-time bounds centered on the refScan within which data points are identified and attributed to the first mass feature in each file. Step 508: assign the per-file peak center (fileScan) of each remaining EIC to the nearest refScan selected in Step 506 and shift the data points of each such remaining EIC by a timePROV PATENT Danforth Docket No.: DDPSC0170-401 -PCT Via EFS-webincrement sufficient to align the fileScan with its assigned refScan. Step 510: iteratively repeat Steps 506-508 to select each nth+1 refScan, assign the nth+1 fileScan of each remaining EIC, shift to align, and establish bounds for the nth+1 mass feature; each iteration uses an unshifted data file and excludes data points within the bounds of previously defined mass features from subsequent selection. Step 512: record in a shiftTable, for every refScan-file pair, the retention time of the fileScan, the retention time of the refScan, and a shift equal to fileScan retention time minus refScan retention time, together with a mass feature identifier and the m / z bin for the mass feature, and optionally a file identifier and one or more quality flags. Step 514: align each nthmass feature defined in Steps 506-512, based on the shiftTable and using the original unshifted EIC of each data file, by shifting the data points already attributed to that nthmass feature along the retention-time axis by a time increment sufficient to bring the recorded fileScan to the closest refScan in that file (by retention-time difference), thereby correcting drift in retention times of mass features.
[0124] Method 500 of FIG. 13 can be embodied as executable instructions in a non-transitory computer readable storage medium including but not limited to a CD, DVD, or nonvolatile memory such as a hard drive. The instructions of the storage medium can be executed by a processor (or processors) to cause various hardware components of a computing device hosting or otherwise accessing the storage medium to effectuate the method. The steps identified in FIG. 13 (and the order thereof) are exemplary and can include various alternatives, equivalents, or derivations thereof including but not limited to the order of execution of the same.(c) Chromolog methods
[0125] Another embodiment of the instant disclosure encompasses a method for correcting drift in retention times of mass features in MS by resolving an ambiguous shift for EICs in which two or more mass features have an ambiguous shift. As used herein, the term “ambiguous shift” refers to a refScan-file situation in which two or more fileScans in the same EIC are closest to, and therefore assigned to, the same refScan such that more than one shift value is recorded for that refScan in that file. Despite the power of the techniques described inPROV PATENT Danforth Docket No.: DDPSC0170-401 -PCT Via EFS-webSection l(b), this phenomenon can occur in some situations such as where drift exceeds one-half of the isomer separation distance.
[0126] Ambiguity can arise, for instance, when within a single EIC one mass feature’s fileScan drifts in retention time to lie closer, by difference, to a different mass feature’s refScan than to its own, producing multiple candidate assignments for the same refScan in that file. In embodiments where chromolog methods are used after Last Resort methods, files exhibiting an ambiguous shift constitute a minority of the dataset; for each mass feature implicated in an ambiguous shift, a large proportion of the fileScans across the files have already been uniquely associated with their correct refScan. For instance, zwitterionic species such as certain amino acids, and in some cases peptides or proteins, observed retention-time behavior can vary across files under changing chromatographic conditions, for example under subtly changing chromatography conditions (e.g., transient pH variation), causing pronounced RT drift. For instance, leucine and isoleucine are isomers that can exhibit large, file-specific shifts in some LC-MS runs. In a subset of files, the fileScan for isoleucine can drift leftward sufficiently that it becomes closer to the refScan of leucine than to the correct refScan for isoleucine; conversely, leucine can drift toward the isoleucine refScan. In such files, both fileScans can be closest to the same refScan, creating more than one possible shift value for that refScan-file pair and thus an ambiguous shift. Ambiguous shifts can also appear when low signal-to-noise yields multiple nearby local maxima around a refScan, or when waveform reversals alter peak-height ordering between mass features across files, increasing the likelihood that two apexes contend for the same refScan in a given file.
[0127] Chromolog methods for correcting drift of the instant disclosure comprise generating or having generated an extracted EIC for each m / z bin in each file of a dataset comprising a plurality of files. EICs, generating EICs, and m / z bins can be as described in Sections l(a) and l(b) herein above.
[0128] The chromolog methods of the instant disclosure further comprise, for a dataset comprising a plurality of files, referencing, for every refScan-file pair, the recorded retention time of the fileScan, the retention time of the refScan to which the fileScan is assigned, a mass feature identifier, and the m / z bin for the mass feature, and optionally a file identifier and one or more quality flags. In some embodiments, the retention time of the fileScan, the retention timePROV PATENT Danforth Docket No.: DDPSC0170-401 -PCT Via EFS-webof the refScan to which the fileScan is assigned, the mass feature identifier, and the m / z bin for the mass feature, and optionally a file identifier and one or more quality flags are recorded in a shiftTable. In some embodiments, a shiftTable can be as described in Section l(b) herein above.
[0129] Chromolog methods further comprise, for EICs in which two or more mass features have an ambiguous shift, identifying a chromolog for a query mass feature and using the identified chromolog to align the mass features that exhibit ambiguous shift. As used herein, the term “chromolog” refers to a mass feature in a different m / z bin than a mass feature with ambiguous shift whose per-file fileScan retention times most closely co-locate with the query’s per-file fileScan retention times across the files, as determined by an aggregate colocation score. Said another way, a chromolog is a reference mass feature, at a different m / z bin than a query mass feature, that exhibits homologous chromatography to the query across files. Conceptually, it is the across-files counterpart whose per-file fileScans track the query’s fileScans closely in the same file, such that both mass features undergo similar drift patterns in those files. Because the chromolog is drawn from a different m / z bin of the same file, it provides an external anchor whose drift has the same ambiguity to guide correction of ambiguous assignments in the query.
[0130] From a co-elution perspective, a chromolog is the mass feature whose per-file fileScan retention times lie on top of, or as close as possible to, the query’s fileScan retention times in each file.
[0131] From a statistical perspective, the chromolog is defined by minimizing an aggregate co-location score between the query’s fileScan retention times and a candidate’s fileScan retention times. For a given query mass feature and one of its refScans, construct a per-file vector of fileScan retention times where the query is present. Do the same for each candidate mass feature at a different m / z bin. Compute an aggregate co-location score over the files where both features have recorded fileScans, defined as the sum of differences between per-file fileScan retention times in the same file. The mass feature that yields the lowest aggregate co-location score is designated the chromolog for that query refScan
[0132] From an operational perspective, identifying a chromolog proceeds as follows. Using the shiftTable which, for each refScan-file pair, records at least the refScan retentionPROV PATENT Danforth Docket No.: DDPSC0170-401 -PCT Via EFS-webtime and the fileScan retention time for a mass feature, extract the query’s fileScan retention times across files for the refScan of interest. For each candidate mass feature in a different m / z bin, extract its fileScan retention times across the same files. Evaluate the aggregate colocation score across the overlap of files where both features have recorded fileScans. The selected candidate becomes the chromolog.
[0133] Importantly, the chromolog step does not require knowledge of the causal mechanism underlying drift. The method solves directly for observable consequences: a chromolog is selected solely on the basis of per-file co-location of fileScan retention times and the resulting aggregate co-location score, and its per-file shift (fileScan - refScan) is then used to resolve an ambiguous shift in the query mass feature. Thus, even when the physical or chemical basis of drift is not known, the empirical co-location and shift transfer are sufficient to correct the ambiguity.
[0134] Additionally, there are few restrictions on the chromolog beyond being in a different m / z bin and matching the query’s refScan retention time (within a predefined tolerance). A chromolog’s EIC can itself contain multiple peaks; in practice, non chromolog peaks “move out of the way” under changing conditions, while the relevant chromolog mass feature remains the one whose per-file fileScans track closest to the query’s fileScans across files. This holds even when peak order reversals occur: the chromolog’s per-file tracking remains aligned with the query’s per-file drift, preserving the validity of the co-location score and the shift transfer.
[0135] From a drift-transfer perspective, the chromolog is the feature whose per-file shift vector (fileScan - refScan, per file) provides the alignment cue for the query. Because the chromolog and the query mass feature drift together in the same files, virtually applying the chromolog’s per-file shift in a given file to the query mass feature indicates where the query’s fileScan can land relative to its refScan, enabling correction of ambiguous assignments for the query mass feature.
[0136] From an eligibility perspective, a valid chromolog can satisfy three conditions. First, the valid chromolog is drawn from a different m / z bin than the query mass feature.Second, has sufficient per-file coverage to compute reliable co-location scores (for example, a minimum number of files with credible overlapping FileScans). Third, it passes basic co-elutionPROV PATENT Danforth Docket No.: DDPSC0170-401 -PCT Via EFS-webchecks, such as having similar central tendencies and not systematically offset from the query beyond predefined bounds.
[0137] From a robustness perspective, chromolog selection tolerates real-world complications by using the ShiftTable’s per-file evidence rather than a single average RT. Missing FileScans are handled by scoring only over shared files. Outliers can be mitigated by differences, trimming, or weighting by quality flags. Overlap / co-elution flags help exclude candidates whose per-file apexes are confounded by neighboring mass features within the RefScan’s bounds.
[0138] Accordingly, a chromolog is the different-mass feature, across-files co-elution partner of a query mass feature, defined as the candidate whose per-file FileScan RTs most closely track the query’s FileScan RTs across the files of the dataset according to a co-location score. It is used as the alignment surrogate in files where the query’s assignment is ambiguous: the chromolog’s per-file shift indicates how to virtually shift the query mass feature’s fileScans to identify the correct correction for drift.
[0139] According to chromolog methods of the instant disclosure, identifying a chromolog comprises (a) constructing, for the query mass feature and the refScan to which the fileScans are assigned, a shiftRow from the shiftTable comprising, for each file, the recorded retention time of that refScan and the recorded retention times of each fileScan closest to that refScan in that file (by retention-time difference); and constructing, across the same files, a corresponding shiftRow for each candidate chromolog, wherein a “candidate chromolog” is a mass feature in a different m / z bin whose refScan retention time matches within a predefined retention-time tolerance the query mass feature’s refScan retention time; and (b) assigning the chromolog for the query mass feature by computing, from the shiftRows, an aggregate colocation score defined as the sum, over files in which both the query mass feature and a candidate chromolog have recorded fileScans, of the value of the difference between the candidate chromolog’s fileScan retention time and the query mass feature’s corresponding fileScan retention time in the same file; and selecting as the chromolog the candidate chromolog that minimizes said aggregate co-location score.
[0010] When the chromolog is identified, the method comprises aligning, in each file where the query mass feature has an ambiguous shift, by first shifting, using the originalPROV PATENT Danforth Docket No.: DDPSC0170-401 -PCT Via EFS-webunshifted EIC of that file for the chromolog’s m / z bin, the data points already attributed to the chromolog by a time increment sufficient to bring the chromolog’s recorded fileScan to the chromolog’s refScan; and then, in the same file, shifting the data points already attributed to the query mass feature by the same time increment applied to the chromolog, thereby bringing the query mass feature’s recorded fileScan into the retention-time bounds centered on the refScan used for both the chromolog and the query mass feature.
[0011] Chromolog methods then comprise iteratively repeating the identification and alignment steps for each additional mass feature that has an ambiguous shift associated with the same refScan. Optionally, the chromolog methods can further comprise updating the shiftTable to record, for each file in which the chromolog is used to align the query mass feature, the corrected fileScan retention time for the query mass feature and the time increment applied.
[0142] For EICs in which two or more mass features have an ambiguous shift, the method can comprise resolving the ambiguity by identifying a chromolog for each of the two or more mass features and then aligning the mass feature’s fileScans by applying, in each such file, the increment corresponding to the chromolog’s shift (fileScan - refScan) to bring the mass feature’s recorded fileScan into the retention-time bounds centered on the refScan used for both the chromolog and the mass feature. This is performed iteratively for each ambiguous mass feature and its chromolog to handle chromatograms of arbitrary complexity.A. Diagrammatic illustration
[0143] Chromolog methods of the instant disclosure can be diagrammatically illustrated in FIGs. 14-32. FIG. 14 diagrammatically depicts an ambiguous shift. In FIG. 14, the upper panel depicts the prevalent positions of leucine and isoleucine mass features across most files, with the refScan for each mass feature identified. The lower panel depicts a file in which both the leucine and isoleucine fileScans have drifted left sufficiently that, by retention-time difference, both fileScans are closest to the leucine refScan, resulting in an ambiguous shift.
[0144] FIGs. 15-17 diagrammatically illustrate the concept of identifying a chromolog for a mass feature using per-file co-location of fileScan retention times across different m / z bins. In the top panels of FIGs. 15-17 (m / z bin A), two mass features are shown for the same m / zPROV PATENT Danforth Docket No.: DDPSC0170-401 -PCT Via EFS-webbin. The first mass feature (isomer 1 ) is zwitterionic and exhibits pronounced retention-time drift across files, diagrammatically represented by multiple staggered peaks (each peak corresponding to a fileScan in a different file). The second mass feature (isomer 2) is non-zwitterionic and shows minimal drift, represented by a tight cluster of peaks with limited dispersion in RT.
[0145] In the middle panels (m / z bin B) of FIGs. 15-17, a candidate mass feature for a chromolog (candiate chromolog 1) at a different m / z bin but at approximately the same refScan RT as isomer 1 is shown. This candidate chromolog also exhibits pronounced drift across files, represented by multiple staggered peaks and, using methods of the instant disclosure, will be identified as the chromolog for isomer 1 because its per-file fileScan RTs most closely co-locate with those of isomer 1 across the files as illustrated by the arrows in FIG. 16 from individual peaks (fileScans) of isomer 1 in the top panel (m / z bin A) to the corresponding peaks in the middle panel (candidate 1, m / z bin B). The arrows illustrate that, file-by-file, peaks co-locate closely in RT: when isomer 1’s fileScans drift across files, candidate 1’s fileScans drift in the same files with similar magnitudes and directions. As it will be shown herein further below, the aggregate co-location score (sum over files of the RT differences between paired fileScans) is low, indicating close co-location. This visually demonstrates that the middle-panel candidate can be a valid chromolog for isomer 1.
[0146] In the bottom panels of FIGs. 15-17, (m / z bin C), another candidate mass feature for a chromolog (candidate chromolog 2) at a different m / z bin and at approximately the same refScan RT is shown. This candidate chromolog has a much cleaner (less dispersed) set of peaks, indicating little drift across files. Despite the similar refScan RT, this candidate will not be the chromolog for isomer 1 because its per-file fileScan RTs do not co-locate with those of isomer 1 across the files as illustrated by the arrows in FIG. 17 from individual peaks (fileScans) of isomer 1 in the top panel (m / z bin A) to the corresponding peaks in the bottom panel (candidate 2, m / z bin C). The arrows illustrate that, file-by-file, the peaks do not co-locate in RT: when isomer 1’s fileScans drift, candidate 2’s fileScans do not follow the same drift pattern. Consequently, the aggregate co-location score (sum over files of the RT differences between paired fileScans) will be large, indicating poor co-location. This visual demonstratesPROV PATENT Danforth Docket No.: DDPSC0170-401 -PCT Via EFS-webthat the bottom-panel candidate is not a suitable chromolog for isomer 1 despite having a similar refScan RT.
[0147] Accordingly, methods of the instant disclosure comprise identifying a chromolog for each of the two or more mass features that have an ambiguous shift in a single EIC.Identifying a chromolog for a mass feature that has ambiguous shift comprises constructing from the retention times of fileScans recorded in the shiftTable, for a query mass feature that has an ambiguous shift, a shiftRow, wherein the shiftRow comprises the recorded retention time of the refScan of the query mass feature, the recorded retention times of each fileScan of the query mass feature in each file of the dataset, and a mass feature identifier; and constructing a corresponding shiftRow for each candidate chromolog, across the same files, wherein a “candidate chromolog” is a mass feature in a different m / z bin whose refScan retention time matches within a predefined retention-time tolerance the query mass feature’s refScan retention time. ShiftRows for the mass feature with an ambiguous shift and the two candidate chromologs (corresponding to FIGs. 15-17) are diagrammatically illustrated in FIG.18. As shown in FIG. 18, each shiftRow for each mass feature comprises the retention time (RT) of the refScan of that mass feature, and the recorded RTs of the fileScans for that mass feature across the files.
[0148] The method then assigns a chromolog to the query mass feature by computing, from the shiftRows, an aggregate co-location score defined as the sum, over files in which both the query mass feature and a candidate chromolog have recorded fileScans, of the value of the difference between the candidate chromolog’s fileScan retention time and the query mass feature’s corresponding fileScan retention time in the same file; and selecting as the chromolog the candidate chromolog that minimizes said aggregate co-location score. The calculation of the aggregate co-location score for the mass features shown in FIGs. 15-17 is diagrammatically presented in FIG. 18. In this example, equation 1 can be used to calculate the score for the correct chromolog candidate isEq 1 : abs(1 -1 ) + abs(2-2) + abs(3-3) = 0,while the score for the incorrect candidate can be calculated using equation 2Eq 2: abs(1 -2) + abs(2-2) + abs(3-2) = 2.PROV PATENT Danforth Docket No.: DDPSC0170-401 -PCT Via EFS-webIn each of the equations, each term such as abs(1 -1 ) represents the value of the difference in RT between the query mass feature’s fileScan and the corresponding candidate chromolog’s fileScan in the same file. The figure shows an example using three files, but the aggregate colocation score can be computed over all or most files of the dataset. Because 0 is smaller than 2, the candidate with the lower aggregate co-location score is selected as the chromolog.
[0149] Once a chromolog is identified, methods of the instant disclosure comprises aligning fileScans of the query mass feature, in each file where the query mass feature has an ambiguous shift, by first shifting, using the original unshifted EIC of that file for the chromolog’s m / z bin, the data points already attributed to the chromolog along the retention-time axis by a time increment sufficient to bring the chromolog’s recorded fileScan to the chromolog’s refScan; and then, in the same file, shifting the data points already attributed to the query mass feature by the same time increment applied to the chromolog, thereby bringing the query mass feature’s recorded fileScan into the retention-time bounds centered on the refScan used for both the chromolog and the query mass feature.
[0150] The alignment of fileScans of the query mass feature is shown diagrammatically in FIGs. 19-22 for the diagrammatic lie and Leu examples shown in FIG. 14. FIGs. 19-22 illustrate, in an exemplary situation, how the alignment step is performed using a chromolog to resolve an ambiguous shift. In FIG. 19, the upper panel shows the main isoleucine (He) and leucine (Leu) mass features at their refScans derived from files without drift (central tendency). In the second panel, a set of files with ambiguous shift is shown, where He and Leu have drifted to the left sufficiently that their fileScans contend for the same refScan (the leucine refScan). In the third panel, the chromolog’s central tendency peak (from a different m / z bin, blue circles) is shown aligned with the “right” main peak (the one the chromolog co-locates with across files, here, lie). In the fourth panel, the drifted chromolog fileScans are displayed; in the ambiguous files, the drifted chromolog aligns with the “right” drifted peak of the query mass feature, indicating the same per-file shift pattern. In FIG. 20, the same sets of peaks are shown after applying the chromolog’s per-file time increment to both the chromolog and the query mass feature in the ambiguous files: the drifted chromolog fileScans are brought to the chromolog’s refScan, and the query mass feature’s drifted fileScans are shifted by the same increment into the retention-time bounds centered on the shared refScan, thereby resolvingPROV PATENT Danforth Docket No.: DDPSC0170-401 -PCT Via EFS-webthe ambiguous assignment for the “right” peak. FIGs. 21 and 22 iterate the process for the “left” peak (Leu, purple circles): the chromolog that co-locates with the leftmost main peak is shown at central tendency and in its drifted ambiguous files, followed by the post-alignment view (FIG. 22) in which the chromolog’s per-file time increment is applied to both the chromolog and the query mass feature, bringing the query fileScans into the retention-time bounds centered on the appropriate refScan for the left peak. Together, these figures demonstrate that, file-by-file, the chromolog’s per-file shift (fileScan - refScan) serves as the alignment cue to place the query mass feature’s fileScan at the correct refScan when an ambiguous shift is present.
[0151] This approach was validated on real LC-MS data for isoleucine (He) and leucine (Leu). FIG. 23 shows data from approximately 400 files, with the first panel displaying the He and Leu mass feature data and the second panel providing a zoomed view of their region. Further zooming on the He region reveals files in which the fileScans have drifted left, producing convolution of peaks and data points between the two mass features and resulting in ambiguous shifts in a subset of files. This can also be shown in histogram format for the He zone of interest, where FIG. 24 presents histograms of the lie fileScans and the He chromolog fileScans. FIG. 25 shows, in the top panel, the diagrammatic representation of the Leu and He mass features from FIGs 19-22, and below it, the histogram of the He chromolog in the middle panel, and the real data points for isoleucine (black dots) and the chromolog of isoleucine (blue dots) in the third panel, illustrating that the chromolog fileScans track with the lie fileScans across files. FIG. 26 shows the alignment step in which the chromolog’s fileScans are brought to its refScan and the same time increment is applied to align the lie fileScans to the He refScan, resulting in visibly cleaner data. FIG. 27 presents the corresponding before / after view in histogram format, confirming the cleanup and resolution of ambiguous shifts.
[0152] The method comprises then iteratively repeating steps (c)(i) through (c)(iii) for each additional mass feature that has an ambiguous shift associated with the same refScan.Slides 24-28 show this for the Leu isomer in the example explained previously
[0153] In some embodiments, the method further comprises updating the shiftTable to record, for each file in which the chromolog is used to align the query mass feature, the corrected fileScan retention time for the query mass feature and the time increment applied.PROV PATENT Danforth Docket No.: DDPSC0170-401 -PCT Via EFS-webB. Embodiments
[0154] Another embodiment of the instant disclosure encompasses a method for correcting drift in retention times of mass features in MS by resolving an ambiguous shift for EICs in which two or more mass features have an ambiguous shift. The method comprises for each m / z bin, generating an extracted ion chromatogram (EIC) in each file of a dataset comprising a plurality of files; wherein a file is the full LC-MS data acquired from one sample, and an EIC is the intensity-versus-retention-time (RT) signal extracted for one specific m / z bin within a file. The method further comprises referencing a shiftTable for a dataset comprising a plurality of files, wherein the shiftTable records, for every refScan-file pair, the retention time of the fileScan, the retention time of the refScan to which the fileScan is assigned, a mass feature identifier, and the m / z bin for the mass feature, and optionally a file identifier and one or more quality flags. A shift table can be as described in Section l(b) herein above.
[0155] The method further comprises resolving the ambiguity for EICs in which two or more mass features have an ambiguous shift when performing the previous step, wherein an ambiguous shift occurs when two or more fileScans in a single EIC are closest to, or assigned to, the same refScan such that more than one shift value is recorded for that refScan-file situation. Resolving the ambiguity comprises identifying a chromolog for a query mass feature. Identifying a chromolog comprises (a) constructing from the retention times of fileScans recorded in the shiftTable, for a query mass feature that has an ambiguous shift, a shiftRow, wherein the shiftRow comprises the recorded retention time of the refScan of the query mass feature, the recorded retention times of each fileScan of the query mass feature in each file of the dataset, and a mass feature identifier; (b) constructing, across the same files, a corresponding shiftRow for one or more candidate chromologs, wherein a “candidate chromolog” is a mass feature in a different m / z bin whose refScan retention time matches, within a predefined retention-time tolerance, the query mass feature’s refScan retention time; and (c) assigning a chromolog from the one or more candidate chromologs to the query mass feature by computing, from the shiftRows, an aggregate co-location score defined as the sum, over files in which both the query mass feature and a candidate chromolog have recorded fileScans, of the value of the difference between the candidate chromolog’s fileScan retentionPROV PATENT Danforth Docket No.: DDPSC0170-401 -PCT Via EFS-webtime and the query mass feature’s corresponding fileScan retention time in the same file; and selecting as the chromolog the candidate chromolog that minimizes said aggregate co-location score.
[0156] Resolving the ambiguity further comprises aligning fileScans of the query mass feature, in each file where the query mass feature has an ambiguous shift, by first shifting, using the original unshifted EIC of that file for the chromolog’s m / z bin, the data points already attributed to the chromolog along the retention-time axis by a time increment sufficient to bring the chromolog’s recorded fileScan to the chromolog’s refScan; and then, in the same file, shifting the data points already attributed to the query mass feature by the same time increment applied to the chromolog, thereby bringing the query mass feature’s recorded fileScan into the retention-time bounds centered on the refScan used for both the chromolog and the query mass feature; and iteratively repeating the previous steps for each additional mass feature that has an ambiguous shift associated with the same refScan. Optionally, the method can further comprise updating the shiftTable to record, for each file in which the chromolog is used to align the query mass feature, the corrected fileScan retention time for the query mass feature and the time increment applied.
[0157] A file is the full LC-MS data acquired from one sample; an extracted ion chromatogram (EIC) is the intensity-versus-retention-time signal extracted for one m / z bin within a file; a mass feature is defined as a single chromatographic peak identified within an m / z bin and retention-time peak; a refScan is a reference retention time across files selected to represent a mass feature and used as an alignment target; and a fileScan is the per-file peak center nearest a refScan recorded for that mass feature in that file.
[0158] FIG. 32 is a flowchart illustrating various steps of the chromolog method to resolve ambiguous shifts of mass features in MS. Step 602: reference a shiftTable for a dataset comprising a plurality of files, wherein the shiftTable records, for every refScan-file pair, the retention time of the fileScan, the retention time of the refScan to which the fileScan is assigned, and a shift equal to fileScan retention time minus refScan retention time, together with a mass feature identifier and the m / z bin for the mass feature, and optionally a file identifier and one or more quality flags. Step 604: identify files in which a mass feature has an ambiguous shift, wherein an “ambiguous shift” occurs when two or more fileScans in a singlePROV PATENT Danforth Docket No.: DDPSC0170-401 -PCT Via EFS-webEIC are closest to, or assigned to, the same refScan such that more than one shift value is recorded for that refScan-file situation. Step 606: for a query mass feature, construct a shiftRow from the shiftTable comprising, for each file, the recorded retention time of the refScan to which the query’s fileScans are assigned and the recorded retention times of each fileScan closest to that refScan in that file (by retention-time difference). Step 608: construct, across the same files, a corresponding shiftRow for each candidate chromolog, wherein a “candidate chromolog” is a mass feature in a different m / z bin whose refScan retention time matches, within a predefined retention-time tolerance, the query mass feature’s refScan retention time. Step 610: assign the chromolog by computing, from the shiftRows, an aggregate co-location score defined as the sum, over files in which both the query mass feature and a candidate chromolog have recorded fileScans, of the value of the difference between the candidate chromolog’s fileScan retention time and the query mass feature’s corresponding fileScan retention time in the same file; select as the chromolog the candidate chromolog that minimizes said aggregate co-location score. Step 612: align, in each file where the query mass feature has an ambiguous shift, by first shifting, using the original unshifted EIC of that file for the chromolog’s m / z bin, the data points already attributed to the chromolog along the retention-time axis by a time increment sufficient to bring the chromolog’s recorded fileScan to the chromolog’s refScan; and then, in the same file, shifting the data points already attributed to the query mass feature by the same time increment applied to the chromolog, thereby bringing the query mass feature’s recorded fileScan into the retention-time bounds centered on the refScan used for both the chromolog and the query mass feature. Step 614: iteratively repeat Steps 606-612 for each additional mass feature that has an ambiguous shift associated with the same refScan. Step 616 (optional): update the shiftTable to record, for each file in which the chromolog is used to align the query mass feature, the corrected fileScan retention time for the query mass feature and the time increment applied.(d) Last Resort + chromolog
[0159] Yet another embodiment of the instant disclosure encompasses a method for correcting drift in retention times of mass features in mass spectrometry (MS), the method comprising locally aligning mass features in an m / z bin of interest and, for EICs in which two orPROV PATENT Danforth Docket No.: DDPSC0170-401 -PCT Via EFS-webmore mass features have an ambiguous shift, wherein an ambiguous shift occurs when two or more fileScans in a single EIC are closest to, or assigned to, the same refScan such that more than one shift value is recorded for that refScan-file situation, resolving the ambiguity. Locally aligning mass features can be as described in Section l(b) herein above. Resolving the ambiguity can be as described in Section l(c) herein above.
[0160] Accordingly, a method of the instant disclosure comprises, for an m / z bin, generating or having generated an extracted ion chromatogram (EIC) in each file of a dataset comprising a plurality of files. The method further comprises locally aligning mass features in an m / z bin of interest by: (i) selecting a first reference retention time across files (refScan) as a scan comprising the highest intensity data points of the highest number of EICs across the files of the dataset among the plurality of EICs extracted from each file in the dataset and establishing retention-time bounds centered on the first refScan within which data points are identified and attributed to the first mass feature in each file, thereby defining a first mass feature; (ii) assigning a per-file peak center (fileScan) of each remaining EIC of the plurality of EICs to a nearest refScan selected in step (b)(i), wherein a remaining EIC is an EIC in which the fileScan does not coincide with the refScan of the first mass feature; (iii) shifting the data points of each remaining EIC by a time increment sufficient to align the fileScan of the EIC with its assigned refScan; (v) iteratively repeating steps (i) through (iii) for each nth+1 refScan to define an nth+1 mass feature, thereby assigning each remaining EIC’s nth+1 fileScan to that refScan, shifting to align, and recording bounds; wherein each iteration uses an unshifted data file to identify the nth+1 mass features, wherein the nth+1 selection excludes data points within the bounds of previously defined mass features, and wherein n is an integer having a value equal to or more than 1 ; (v) recording in a shiftTable, for every refScan-file pair, the retention time of the fileScan, the retention time of the refScan to which the fileScan is assigned, a shift equal to fileScan retention time minus refScan retention time, a mass feature identifier, and optionally the m / z bin for the mass feature, a file identifier, and one or more quality flags; and (vi) aligning each nthmass feature defined in steps (i) through (v), based on the shiftTable of step (v) and using the original unshifted EIC of each data file, by shifting the data points attributed to that nthmass feature along the retention-time axis by a time increment sufficient toPROV PATENT Danforth Docket No.: DDPSC0170-401 -PCT Via EFS-webbring the recorded fileScan of that nthmass feature to the closest refScan in that file; thereby correcting the drift in retention times of mass features.
[0161] The method further comprises for EICs in which two or more mass features have an ambiguous shift when performing step (b), wherein an ambiguous shift occurs when two or more fileScans in a single EIC are closest to, or assigned to, the same refScan such that more than one shift value is recorded for that refScan-file situation, resolving the ambiguity.Resolving the ambiguity comprises (i) identifying a chromolog for each of the two or more mass features; aligning fileScans of the query mass feature; (iii) iteratively repeating steps (i) through (ii) of resolving the ambiguity for each additional mass feature that has an ambiguous shift associated with the same refScan, and (iv) optionally updating the shiftTable to record, for each file in which the chromolog is used to align the query mass feature, the corrected fileScan retention time for the query mass feature and the time increment applied.II. Computing system
[0162] The instant disclosure encompasses a computing system for correcting drift in retention times of mass features in mass spectrometry (MS). In some embodiments, the computing system is attached to and / or used in conjunction with a chromatography-MS system. The computing system can comprise several known components and circuitry, including a processor, a memory system, input and output devices and interfaces (e.g., an interconnection mechanism), as well as other components, such as transport circuitry (e.g., one or more busses), a video and audio data input / output (I / O) subsystem, special-purpose hardware, as well as other components and circuitry, as described below in more detail.Further, the computer system(s) can be a multi-processor computer system or can include multiple computers connected over a computer network.
[0163] A processor can include one or more general purpose computers, dedicated microprocessors, graphics processors, or other processing devices capable of communicating electronic information. Non-limiting examples of a processor include one or more applicationspecific integrated circuits (ASICs), graphical processing units (GPUs), field programmable gate arrays (FPGAs), graphics processing units (GPUs), digital signal processors (DSPs) and any other suitable specific or general purpose processors. The processor can be implementedPROV PATENT Danforth Docket No.: DDPSC0170-401 -PCT Via EFS-webas appropriate in hardware, firmware, or combinations thereof with computer-executable instructions and / or software. Computer-executable instructions and software can include computer-executable or machine-executable instructions written in any suitable programming language to perform the various functions described.
[0164] The memory can include more than one memory and can be distributed throughout the computing system. The memory can store program instructions that are loadable and executable on the processor(s) as well as data generated during the execution of these programs. Depending on the configuration and type of memory, the memory can be volatile (such as random access memory (RAM)) and / or non-volatile (such as read-only memory (ROM), flash memory, or other memory). In some embodiments, the memory can include multiple different types of memory, such as static random access memory (SRAM), dynamic random access memory (DRAM), or ROM.
[0165] In some embodiments, the computing system can also include additional storage, which can include removable storage and / or non-removable storage. The additional storage can include, but is not limited to, magnetic storage, optical disks, and / or solid-state storage. The disk drives and their associated computer-readable media can provide nonvolatile storage of computer-readable instructions, data structures, program modules, and other data for the computing devices. The memory and the additional storage, both removable and non-removable, are examples of computer-readable storage media. For example, computer-readable storage media can include volatile or non-volatile, removable, or nonremovable media implemented in any suitable method or technology for storage of information such as computer-readable instructions, data structures, program modules, or other data. As used herein, modules, engines, and components, can refer to programming modules executed by computing systems (e.g., processors) that are part of the architecture.
[0166] The processor generally manipulates the data within the integrated circuit memory element in accordance with the program instructions and then copies the manipulated data to the non-volatile recording medium after processing is completed. A variety of mechanisms are known for managing data movement between the non-volatile recording medium and the integrated circuit memory element, and the computing system thatPROV PATENT Danforth Docket No.: DDPSC0170-401 -PCT Via EFS-webimplements the methods, steps, systems control and system elements control described above is not limited thereto. The computing system is not limited to a particular memory system.
[0167] At least part of such a memory system described above can be used to store one or more data structures (e.g., look-up tables) or equations such as calibration curve equations. For example, at least part of the non-volatile recording medium can store at least part of a database that includes one or more of such data structures. Such a database can be any of a variety of types of databases, for example, a file system including one or more flat-file data structures where data is organized into data units separated by delimiters, a relational database where data is organized into data units stored in tables, an object-oriented database where data is organized into data units stored as objects, another type of database, or any combination thereof.
[0168] The computer implemented control system(s) can include one or more output devices. Non-limiting example output devices include a cathode ray tube (CRT) display, liquid crystal displays (LCD) and other video output devices, printers, communication devices such as a modem or network interface, storage devices such as disk or tape, and audio output devices such as a speaker.
[0169] The computing system also can include one or more input devices. Example input devices include a keyboard, keypad, track ball, mouse, pen and tablet, communication devices such as described above, and data input devices such as audio and video capture devices and sensors. The computing system is not limited to the particular input or output devices described herein.
[0170] It should be appreciated that one or more of any type of computing system can be used to implement various embodiments described herein. Embodiments of the invention can be implemented in software, hardware or firmware, or any combination thereof. The computing system can include specially programmed, special purpose hardware, for example, an application-specific integrated circuit (ASIC). Such special-purpose hardware can be configured to implement one or more of the methods, steps, simulations, algorithms, systems control, and system elements control described above as part of the computer implemented control system(s) described above or as an independent component.PROV PATENT Danforth Docket No.: DDPSC0170-401 -PCT Via EFS-web
[0171] The computing system and components thereof can be programmable using any of a variety of one or more suitable computer programming languages. Such languages can include procedural programming languages, for example, LabView, C, Pascal, Fortran and BASIC, object-oriented languages, for example, C++, Java and Eiffel and other languages, such as a scripting language or even assembly language.
[0172] The methods, steps, simulations, algorithms, systems control, and system elements control can be implemented using any of a variety of suitable programming languages, including procedural programming languages, object- oriented programming languages, other languages and combinations thereof, which can be executed by such a computer system. Such methods, steps, simulations, algorithms, systems control, and system elements control can be implemented as separate modules of a computer program, or can be implemented individually as separate computer programs. Such modules and programs can be executed on separate computers.
[0173] Such methods, steps, simulations, algorithms, systems control, and system elements control, either individually or in combination, can be implemented as a computer program product tangibly embodied as computer-readable signals on a computer-readable medium, for example, a non-volatile recording medium, an integrated circuit memory element, or a combination thereof. For each such method, step, simulation, algorithm, system control, or system element control, such a computer program product can comprise computer-readable signals tangibly embodied on the computer-readable medium that define instructions, for example, as part of one or more programs, that, as a result of being executed by a computer, instruct the computer to perform the method, step, simulation, algorithm, system control, or system element control.
[0174] For clarity of explanation, in some instances the present technology can be presented as including individual functional blocks including functional blocks comprising devices, device components, steps or routines in a method embodied in software, or combinations of hardware and software.
[0175] Any of the steps, operations, functions, or processes described herein can be performed or implemented by a combination of hardware and software services or services, alone or in combination with other devices. In some embodiments, a service can be softwarePROV PATENT Danforth Docket No.: DDPSC0170-401 -PCT Via EFS-webthat resides in memory of a client device and / or one or more servers of a content management system and perform one or more functions when a processor executes the software associated with the service. In some embodiments, a service is a program, or a collection of programs that carry out a specific function. In some embodiments, a service can be considered a server. The memory can be a non-transitory computer-readable medium.
[0176] In some embodiments the computer-readable storage devices, mediums, and memories can include a cable or wireless signal containing a bit stream and the like. However, when mentioned, non-transitory computer-readable storage media expressly exclude media such as energy, carrier signals, electromagnetic waves, and signals per se.
[0177] Methods according to the above-described examples can be implemented using computer-executable instructions that are stored or otherwise available from computer readable media. Such instructions can comprise, for example, instructions and data which cause or otherwise configure a general purpose computer, special purpose computer, or special purpose processing device to perform a certain function or group of functions. Portions of computer resources used can be accessible over a network. The computer executable instructions can be, for example, binaries, intermediate format instructions such as assembly language, firmware, or source code. Examples of computer-readable media that can be used to store instructions, information used, and / or information created during methods according to described examples include magnetic or optical disks, solid state memory devices, flash memory, USB devices provided with non-volatile memory, networked storage devices, and so on.
[0178] Devices implementing methods according to these disclosures can comprise hardware, firmware and / or software, and can take any of a variety of form factors. Typical examples of such form factors include servers, laptops, smart phones, small form factor personal computers, personal digital assistants, and so on. Functionality described herein also can be embodied in peripherals or add-in cards. Such functionality can also be implemented on a circuit board among different chips or different processes executing in a single device, by way of further example.PROV PATENT Danforth Docket No.: DDPSC0170-401 -PCT Via EFS-web
[0179] The instructions, media for conveying such instructions, computing resources for executing them, and other structures for supporting such computing resources are means for providing the functions described in these disclosures.
[0180] FIG. 33 shows an example of a computing system 1100 in which the components of the system are in communication with each other using connection 1805. Connection 1805 can be a physical connection via a bus, or a direct connection into processor 1810, such as in a chipset architecture. Connection 1105 can also be a virtual connection, networked connection, or logical connection.
[0181] In some embodiments computing system 1800 is a distributed system in which the functions described in this disclosure can be distributed within a datacenter, multiple datacenters, a peer network, etc. In some embodiments, one or more of the described system components represents many such components each performing some or all of the function for which the component is described. In some embodiments, the components can be physical or virtual devices.
[0182] Example system 1100 includes at least one processing unit (CPU or processor) 1110 and connection 1105 that couples various system components including system memory 1115, such as read only memory (ROM) and random access memory (RAM) to processor 1110.
[0183] Computing system 1100 can include a cache of high-speed memory connected directly with, in close proximity to, or integrated as part of processor 1110.
[0184] Processor 1110 can include any general purpose processor and a hardware service or software service, such as services 1132, 1134, and 1136 stored in storage device 1130, configured to control processor 1110 as well as a special-purpose processor where software instructions are incorporated into the actual processor design. Processor 1110 can essentially be a completely self-contained computing system, containing multiple cores or processors, a bus, memory controller, cache, etc. A multi-core processor can be symmetric or asymmetric.
[0185] To enable user interaction, computing system 1100 includes an input device 1145, which can represent any number of input mechanisms, such as a microphone for speech, a touch-sensitive screen for gesture or graphical input, keyboard, mouse, motionPROV PATENT Danforth Docket No.: DDPSC0170-401 -PCT Via EFS-webinput, speech, etc. Computing system 1100 can also include output device 1135, which can be one or more of a number of output mechanisms known to those of skill in the art. In some instances, multimodal systems can enable a user to provide multiple types of input / output to communicate with computing system 1100. Computing system 1100 can include communications interface 1140, which can generally govern and manage the user input and system output. There is no restriction on operating on any particular hardware arrangement and therefore the basic features here can easily be substituted for improved hardware or firmware arrangements as they are developed.
[0186] Storage device 1130 can be a non-volatile memory device and can be a hard disk or other types of computer readable media which can store data that are accessible by a computer, such as magnetic cassettes, flash memory cards, solid state memory devices, digital versatile disks, cartridges, random access memories (RAMs), read only memory (ROM), and / or some combination of these devices.
[0187] The storage device 1130 can include software services, servers, services, etc., that when the code that defines such software is executed by the processor 1110, it causes the system to perform a function. In some embodiments, a hardware service that performs a particular function can include the software component stored in a computer-readable medium in connection with the necessary hardware components, such as processor 1110, connection 1105, output device 1135, etc., to carry out the function.
[0188] One embodiment of the instant disclosure encompasses a computing system of the instant disclosure comprises a communication interface and a processor. The communication interface can receive EICs generated for an m / z bin in each file of a dataset comprising a plurality of files. The processor can execute instructions stored in memory, wherein the processor executes the instructions to overcoming drift in retention times of metabolite signals in mass spectrometry (MS). The processor first executes instructions to locally align mass features in an m / z bin of interest by: (1) selecting a first reference retention time across files (refScan) as a scan comprising the highest intensity data points of the highest number of EICs across the files of the dataset among the plurality of EICs extracted from each file in the dataset and establishing retention-time bounds centered on the first refScan within which data points are identified and attributed to the first mass feature in each file, therebyPROV PATENT Danforth Docket No.: DDPSC0170-401 -PCT Via EFS-webdefining a first mass feature; (2) assigning a per-file peak center (fileScan) of each remaining EIC of the plurality of EICs to a nearest refScan selected in step (b)(i), wherein a remaining EIC is an EIC in which the fileScan does not coincide with the refScan of the first mass feature; (3) shifting the data points of each remaining EIC by a time increment sufficient to align the fileScan of the EIC with its assigned refScan; (4) iteratively repeating steps (b)(i)(1) through (b)(i)(3) for each nth+1 refScan to define an nth+1 mass feature, thereby assigning each remaining EIC’s nth+1 fileScan to that refScan, shifting to align, and recording bounds; wherein each iteration uses an unshifted data file to identify the nth+1 mass features, wherein the nth+1 selection excludes data points within the bounds of previously defined mass features, and wherein n is an integer having a value equal to or more than 1 ; (5) recording in a shiftTable, for every refScan-file pair, the retention time of the fileScan, the retention time of the refScan to which the fileScan is assigned, a shift equal to fileScan retention time minus refScan retention time, a mass feature identifier, and optionally the m / z bin for the mass feature, a file identifier, and one or more quality flags; and (6) aligning each nthmass feature defined in steps (b)(i) through (b)(v), based on the shiftTable of step (b)(v) and using the original unshifted EIC of each data file, by shifting the data points attributed to that nthmass feature along the retentiontime axis by a time increment sufficient to bring the recorded fileScan of that nthmass feature to the closest refScan in that file; thereby correcting the drift in retention times of mass features.
[0189] After executing instructions to locally align mass features in an m / z bin of interest, the processor can optionally, for EICs in which two or more mass features have an ambiguous shift when locally align mass features, wherein an ambiguous shift occurs when two or more fileScans in a single EIC are closest to, or assigned to, the same refScan such that more than one shift value is recorded for that refScan-file situation. Resolving the ambiguity can comprise first identifying a chromolog for each of the two or more mass features by (a) constructing from the retention times of fileScans recorded in the shiftTable, for a query mass feature that has an ambiguous shift, a shiftRow, wherein the shiftRow comprises the recorded retention time of the refScan of the query mass feature, the recorded retention times of each fileScan of the query mass feature in each file of the dataset, and a mass feature identifier; (b) constructing, across the same files, a corresponding shiftRow for one or morePROV PATENT Danforth Docket No.: DDPSC0170-401 -PCT Via EFS-webcandidate chromologs, wherein a “candidate chromolog” is a mass feature in a different m / z bin whose refScan retention time matches, within a predefined retention-time tolerance, the query mass feature’s refScan retention time; and (c) assigning a chromolog from the one or more candidate chromologs to the query mass feature by computing, from the shiftRows, an aggregate co-location score defined as the sum, over files in which both the query mass feature and a candidate chromolog have recorded fileScans, of the value of the difference between the candidate chromolog’s fileScan retention time and the query mass feature’s corresponding fileScan retention time in the same file; and selecting as the chromolog the candidate chromolog that minimizes said aggregate co-location score.
[0190] Second, resolving the ambiguity can comprise aligning fileScans of the query mass feature, in each file where the query mass feature has an ambiguous shift, by first shifting, using the original unshifted EIC of that file for the chromolog’s m / z bin, the data points already attributed to the chromolog along the retention-time axis by a time increment sufficient to bring the chromolog’s recorded fileScan to the chromolog’s refScan; and then, in the same file, shifting the data points already attributed to the query mass feature by the same time increment applied to the chromolog, thereby bringing the query mass feature’s recorded fileScan into the retention-time bounds centered on the refScan used for both the chromolog and the query mass feature; iteratively repeating steps (c)(i) through (c)(ii) for each additional mass feature that has an ambiguous shift associated with the same refScan; and optionally updating the shiftTable to record, for each file in which the chromolog is used to align the query mass feature, the corrected fileScan retention time for the query mass feature and the time increment applied.
[0191] In some embodiments, the processor can further execute instructions to generate the EIC in each file of a dataset comprising a plurality of files.
[0192] Another embodiment of the instant disclosure encompasses a computing system for correcting drift in retention times of mass features in mass spectrometry (MS). The computing system comprises a communication interface that receives EICs generated for an m / z bin in each file of a dataset comprising a plurality of files; wherein a file is the full LC-MS data acquired from one sample, and an EIC is the intensity-versus-retention-time (RT) signal extracted for one m / z bin within a file.PROV PATENT Danforth Docket No.: DDPSC0170-401 -PCT Via EFS-web
[0193] The computing system also comprises a processor that executes instructions stored in memory, wherein the processor executes the instructions to: (i) reference a shiftTable for a dataset comprising a plurality of files, wherein the shiftTable records, for every refScan-file pair, the retention time of the fileScan, the retention time of the refScan to which the fileScan is assigned, a mass feature identifier, and the m / z bin for the mass feature, and optionally a file identifier and one or more quality flags; (ii) for EICs in which two or more mass features have an ambiguous shift when performing step (i), wherein an ambiguous shift occurs when two or more fileScans in a single EIC are closest to, or assigned to, the same refScan such that more than one shift value is recorded for that refScan-file situation.
[0194] Resolving the ambiguity can comprise identifying a chromolog for a query mass feature; aligning fileScans of the query mass feature, in each file where the query mass feature has an ambiguous shift, iteratively repeating steps (i) through (iii) for each additional mass feature that has an ambiguous shift associated with the same refScan; and optionally updating the shiftTable to record, for each file in which the chromolog is used to align the query mass feature, the corrected fileScan retention time for the query mass feature and the time increment applied.
[0195] Resolve the ambiguity can comprise: (a) constructing from the retention times of fileScans recorded in the shiftTable, for a query mass feature that has an ambiguous shift, a shiftRow, wherein the shiftRow comprises the recorded retention time of the refScan of the query mass feature, the recorded retention times of each fileScan of the query mass feature in each file of the dataset, and a mass feature identifier; (b) constructing, across the same files, a corresponding shiftRow for one or more candidate chromologs, wherein a “candidate chromolog” is a mass feature in a different m / z bin whose refScan retention time matches, within a predefined retention-time tolerance, the query mass feature’s refScan retention time; and c. assigning a chromolog from the one or more candidate chromologs to the query mass feature by computing, from the shiftRows, an aggregate co-location score defined as the sum, over files in which both the query mass feature and a candidate chromolog have recorded fileScans, of the value of the difference between the candidate chromolog’s fileScan retention time and the query mass feature’s corresponding fileScan retention time in the same file; andPROV PATENT Danforth Docket No.: DDPSC0170-401 -PCT Via EFS-webselecting as the chromolog the candidate chromolog that minimizes said aggregate co-location score.
[0196] An additional embodiment of the instant disclosure encompasses a non-transitory, computer-readable storage medium, having embodied thereon a program executable by a processor to perform a method for correcting drift in retention times of mass features in mass spectrometry (MS). The method for correcting drift first comprises for an m / z bin, generating or having generated an extracted ion chromatogram (EIC) in each file of a dataset comprising a plurality of files; and then locally aligning mass features in an m / z bin of interest. Locally aligning mass features in an m / z bin of interest comprises (i) selecting a first reference retention time across files (refScan) as a scan comprising the highest intensity data points of the highest number of EICs across the files of the dataset among the plurality of EICs extracted from each file in the dataset and establishing retention-time bounds centered on the first refScan within which data points are identified and attributed to the first mass feature in each file, thereby defining a first mass feature; (ii) selecting a first reference retention time across files (refScan) as a scan comprising the highest intensity data points of the highest number of EICs across the files of the dataset among the plurality of EICs extracted from each file in the dataset and establishing retention-time bounds centered on the first refScan within which data points are identified and attributed to the first mass feature in each file, thereby defining a first mass feature; (iii) shifting the data points of each remaining EIC by a time increment sufficient to align the fileScan of the EIC with its assigned refScan; (iv) iteratively repeating steps (b)(i) through (b)(iii) for each nth+1 refScan to define an nth+1 mass feature, thereby assigning each remaining EIC’s nth+1 fileScan to that refScan, shifting to align, and recording bounds; wherein each iteration uses an unshifted data file to identify the nth+1 mass features, wherein the nth+1 selection excludes data points within the bounds of previously defined mass features, and wherein n is an integer having a value equal to or more than 1; (v) recording in a shiftTable, for every refScan-file pair, the retention time of the fileScan, the retention time of the refScan to which the fileScan is assigned, a shift equal to fileScan retention time minus refScan retention time, a mass feature identifier, and optionally the m / z bin for the mass feature, a file identifier, and one or more quality flags; and (vi) aligning each nthmass feature defined in steps (i) through (v), based on the shiftTable of step (v) and usingPROV PATENT Danforth Docket No.: DDPSC0170-401 -PCT Via EFS-webthe original unshifted EIC of each data file, by shifting the data points attributed to that nthmass feature along the retention-time axis by a time increment sufficient to bring the recorded fileScan of that nthmass feature to the closest refScan in that file; thereby correcting the drift in retention times of mass features.
[0197] In some embodiments, the method can further comprise optionally, for EICs in which two or more mass features have an ambiguous shift when locally aligning mass features in an m / z bin of interest, wherein an ambiguous shift occurs when two or more fileScans in a single EIC are closest to, or assigned to, the same refScan such that more than one shift value is recorded for that refScan-file situation, resolving the ambiguity by identifying a chromolog for each of the two or more mass features. Identifying a chromolog for each of the two or more mass features comprises (1) constructing from the retention times of fileScans recorded in the shiftTable, for a query mass feature that has an ambiguous shift, a shiftRow, wherein the shiftRow comprises the recorded retention time of the refScan of the query mass feature, the recorded retention times of each fileScan of the query mass feature in each file of the dataset, and a mass feature identifier; (2) constructing, across the same files, a corresponding shiftRow for one or more candidate chromologs, wherein a “candidate chromolog” is a mass feature in a different m / z bin whose refScan retention time matches, within a predefined retention-time tolerance, the query mass feature’s refScan retention time; and (3) assigning a chromolog from the one or more candidate chromologs to the query mass feature by computing, from the shiftRows, an aggregate co-location score defined as the sum, over files in which both the query mass feature and a candidate chromolog have recorded fileScans, of the value of the difference between the candidate chromolog’s fileScan retention time and the query mass feature’s corresponding fileScan retention time in the same file; and selecting as the chromolog the candidate chromolog that minimizes said aggregate co-location score.
[0198] Resolving the ambiguity then comprises aligning fileScans of the query mass feature, in each file where the query mass feature has an ambiguous shift, by first shifting, using the original unshifted EIC of that file for the chromolog’s m / z bin, the data points already attributed to the chromolog along the retention-time axis by a time increment sufficient to bring the chromolog’s recorded fileScan to the chromolog’s refScan; and then, in the same file, shifting the data points already attributed to the query mass feature by the same timePROV PATENT Danforth Docket No.: DDPSC0170-401 -PCT Via EFS-webincrement applied to the chromolog, thereby bringing the query mass feature’s recorded fileScan into the retention-time bounds centered on the refScan used for both the chromolog and the query mass feature; iteratively repeating the identification of a chromolog through aligning fileScans for each additional mass feature that has an ambiguous shift associated with the same refScan.
[0199] Resolving the ambiguity can optionally further comprise updating the shiftTable to record, for each file in which the chromolog is used to align the query mass feature, the corrected fileScan retention time for the query mass feature and the time increment applied.
[0200] Yet another embodiment of the instant disclosure encompasses a non-transitory, computer-readable storage medium, having embodied thereon a program executable by a processor to perform a method for correcting drift in retention times of mass features in mass spectrometry (MS). The method comprises (i) reference a shiftTable for a dataset comprising a plurality of files, wherein the shiftTable records, for every refScan-file pair, the retention time of the fileScan, the retention time of the refScan to which the fileScan is assigned, a mass feature identifier, and the m / z bin for the mass feature, and optionally a file identifier and one or more quality flags; (ii) for EICs in which two or more mass features have an ambiguous shift when performing step (i), wherein an ambiguous shift occurs when two or more fileScans in a single EIC are closest to, or assigned to, the same refScan such that more than one shift value is recorded for that refScan-file situation.
[0201] Resolving the ambiguity can comprise identifying a chromolog for a query mass feature; aligning fileScans of the query mass feature, in each file where the query mass feature has an ambiguous shift, iteratively repeating steps (i) through (iii) for each additional mass feature that has an ambiguous shift associated with the same refScan; and optionally updating the shiftTable to record, for each file in which the chromolog is used to align the query mass feature, the corrected fileScan retention time for the query mass feature and the time increment applied.
[0202] Resolve the ambiguity can comprise: (a) constructing from the retention times of fileScans recorded in the shiftTable, for a query mass feature that has an ambiguous shift, a shiftRow, wherein the shiftRow comprises the recorded retention time of the refScan of the query mass feature, the recorded retention times of each fileScan of the query mass feature inPROV PATENT Danforth Docket No.: DDPSC0170-401 -PCT Via EFS-webeach file of the dataset, and a mass feature identifier; (b) constructing, across the same files, a corresponding shiftRow for one or more candidate chromologs, wherein a “candidate chromolog” is a mass feature in a different m / z bin whose refScan retention time matches, within a predefined retention-time tolerance, the query mass feature’s refScan retention time; and c. assigning a chromolog from the one or more candidate chromologs to the query mass feature by computing, from the shiftRows, an aggregate co-location score defined as the sum, over files in which both the query mass feature and a candidate chromolog have recorded fileScans, of the value of the difference between the candidate chromolog’s fileScan retention time and the query mass feature’s corresponding fileScan retention time in the same file; and selecting as the chromolog the candidate chromolog that minimizes said aggregate co-location score.DEFINITIONS
[0203] Unless defined otherwise, all technical and scientific terms used herein have the meaning commonly understood by a person skilled in the art to which this invention belongs. The following references provide one of skill with a general definition of many of the terms used in this invention: Singleton et al., Dictionary of Microbiology and Molecular Biology (2nd ed. 1994); The Cambridge Dictionary of Science and Technology (Walker ed., 1988); The Glossary of Genetics, 5th Ed., R. Rieger et al. (eds.), Springer Verlag (1991); and Hale & Marham, The Harper Collins Dictionary of Biology (1991). As used herein, the following terms have the meanings ascribed to them unless specified otherwise.
[0204] As used herein, the terms “file” and “data file” are used interchangeably and refer to a full LC-MS data acquired from one sample. One file contains signals across the full m / z-RT space for that sample.
[0205] As used herein, the term “extracted ion chromatogram” (EIC) refers to the intensity-versus-retention-time (RT) signal extracted, within a given file, for one m / z bin (within a defined tolerance). Each file can contain multiple EICs, one per m / z bin of interest.
[0206] As used herein, the term “merged EIC” refers to the aggregation, across a plurality of files, of the EICs corresponding to the same m / z bin.PROV PATENT Danforth Docket No.: DDPSC0170-401 -PCT Via EFS-web
[0207] As used herein, the term “m / z bin” refers to a defined portion of a mass-to-charge (m / z) axis within a mass spectrometry dataset, from which signal intensities are selected or aggregated for analysis. Methods for defining and implementing m / z bins, including fixed or variable binning, tolerance-based extraction, and instrument-dependent discretization, are well known to individuals of ordinary skill in the art and can be selected as appropriate for a given analytical or computational context. An m / z bin can be defined by one or more boundary m / z values, a center m / z value with an associated width or tolerance, by an index or position within a discretized m / z axis, or any combination thereof. In some embodiments, an m / z bin is defined as a mass window centered on a target m / z value, wherein the bin includes signal intensities having m / z values within a predefined tolerance of the target m / z. The tolerance can be specified as a mass difference (e.g., in Daltons), a relative mass difference (e.g., in parts per million), or any other criterion related to instrument resolution or mass accuracy.
[0208] As used herein, the term “compound” refers to a distinct chemical species (a specific molecular entity). In a chromatography-MS context, each isomer, stereoisomer, or dominant tautomer under the chromatographic conditions can be treated as a separate compound because it can exhibit different retention behavior and produce its own chromatographic peak(s).
[0209] As used herein, the term “mass feature” refers to a single resolved chromatographic peak identified within an m / z bin and a retention-time apex. Formally, a mass feature is the pair (m / z bin, RT peak). Within one EIC (for a given m / z bin in a file), each distinct peak is a distinct mass feature, and the same mass feature can be tracked across multiple files by its m / z bin and RT peak.
[0210] In proteomic datasets, a single peptide or protein species can give rise to multiple observed mass features, for example due to differing charge states or ion forms that appear at different m / z bins. For purposes of the present disclosure, each such observed chromatographic peak, defined by its m / z bin and retention-time apex, is treated as a separate mass feature and aligned independently according to the disclosed methods.
[0211] As used herein, the phrase “mass features corresponding to peptides or proteins” refers to chromatographic peaks observed in chromatography-mass spectrometry data thatPROV PATENT Danforth Docket No.: DDPSC0170-401 -PCT Via EFS-webarise from peptides or proteins, without implying identification, sequencing, or biochemical characterization of the underlying species.
[0212] As used herein, the term “m / z bin” refers to a defined mass-to-charge interval used to group signals in an LC-MS dataset for extraction and analysis. An m / z bin is specified by a central m / z value and a tolerance (e.g., ±A), or by a start-stop range, such that all signals whose measured m / z falls within that interval are included. Each EIC is generated for one m / z bin, and mass features (peaks) are identified within the EIC corresponding to that bin.
[0213] As used herein, the term “peak” refers to a resolved chromatographic signal maximum in an EIC, characterized by a retention-time apex and surrounding signal that can be bounded to define the extent of the peak.
[0214] As used herein, the terms “mass feature” and “peak” are related such that each “peak” within an EIC corresponds to one “mass feature” when identified by its m / z bin and retention-time apex; “mass feature” is used when referring to the peak as a tracked entity across files (the pair m / z bin, RT peak), while “peak” is used when referring to the local chromatographic signal within a single EIC.
[0215] As used herein, the term “peak center” refers to the per-file apex retention time (RT) of a mass feature in a given file’s EIC. The peak center is the observed maximum intensity point for that mass feature within that file.
[0216] As used herein, the term “compound” refers to a distinct chemical species (a specific molecular entity) that, under the LC-MS conditions described, is observed within one m / z bin and can produce one or more resolved chromatographic peaks (mass features) in that m / z bin (e.g., due to isomers, stereoisomers, or tautomers that separate by retention time). The terms “compound,” “mass feature,” and “peak” are related as follows: within a given m / z bin in an EIC, each resolved peak (characterized by a retention-time apex) corresponds to one mass feature, and each such mass feature typically represents one compound present in that m / z bin. Multiple mass features (peaks) can represent different forms of the compound observed in the same m / z bin (e.g., isomers), separated by retention time.
[0217] As used herein, the term “refScan” refers to the reference retention time selected across the files of a dataset to represent a mass feature’s peak and used as the alignment target for that mass feature.PROV PATENT Danforth Docket No.: DDPSC0170-401 -PCT Via EFS-web
[0218] As used herein, the term “fileScan” refers to the per-file apex retention time (peak center) selected in a given file’s EIC for mapping to a refScan, i.e. , the file’s peak center associated with that refScan during an iteration of the alignment process.
[0219] As used herein, the term “shiftTable” refers to a record that, for each refScan-file pair in an LC-MS dataset, stores at least: the retention time of the fileScan (the per-file peak center associated with the refScan), the retention time of the refScan (the alignment target selected across files), and a shift equal to fileScan retention time minus refScan retention time. The shiftTable further includes a mass feature identifier and the m / z bin for the mass feature, and optionally a file identifier and one or more quality flags. The shiftTable is constructed during alignment by, for each mass feature in an m / z bin, selecting one or more refScans and, for each file, locating the fileScan nearest each refScan and recording the foregoing values for that refScan-file pair.
[0220] As used herein, the term “chromolog” refers to a mass feature in a different m / z bin than a query mass feature whose per-file fileScan retention times most closely co-locate with the query’s per-file fileScan retention times across the files, as determined by an aggregate co-location score. The aggregate co-location score is defined as the sum, over files in which both the query mass feature and a candidate chromolog have recorded fileScans, of the value of the difference between the candidate’s fileScan retention time and the query’s corresponding fileScan retention time in the same file; the mass feature that minimizes said score is designated the chromolog. The term chromolog is employed as an alignment surrogate in files where the query mass feature has an ambiguous shift. In such files, a chromolog shift is determined as the difference between the chromolog’s fileScan retention time and the chromolog’s refScan retention time; applying the same time increment to the query mass feature indicates where the query’s fileScan can land relative to its refScan, enabling correction of drift for the query mass feature.
[0221] As used herein, the term “shiftRow” refers to a per-refScan, per-m ass-feature record constructed from the shiftTable that aggregates, across a plurality of files, at least: the recorded retention time of the refScan for the mass feature, the recorded retention time(s) of the fileScan(s) associated with that refScan in each file (i.e., the per-file peak centers closest to the refScan by retention-time difference), and a mass feature identifier. A shiftRow isPROV PATENT Danforth Docket No.: DDPSC0170-401 -PCT Via EFS-webconstructed for a query mass feature and, across the same files, for each candidate chromolog, and is used to compute an aggregate co-location score by pairing, file-by-file, the query’s fileScan retention times with those of a candidate chromolog.
[0222] As used herein, the term “ambiguous shift” refers to a refScan-file situation in which two or more fileScans in the same EIC are assigned to the same refScan such that more than one shift value is recorded for that refScan in that file.
[0223] As used herein, the term “candidate chromolog” refers to a mass feature in a different m / z bin than a query mass feature whose refScan retention time matches, within a predefined retention-time tolerance, the query mass feature’s refScan retention time, and for which a shiftRow can be constructed across the same files to enable computation of an aggregate co-location score against the query’s shiftRow.
[0224] As used herein, the term “query mass feature” refers to the mass feature under evaluation for ambiguity resolution or chromolog identification. The query mass feature is defined by its m / z bin and refScan retention time, and its per-file peak centers (fileScans) recorded in the shiftTable are used to construct a shiftRow and to compute aggregate colocation scores against candidate chromologs.
[0225] When introducing elements of the present disclosure or the preferred aspect(s) thereof, the articles "a", "an", "the" and "said" are intended to mean that there are one or more of the elements. The terms "comprising", "including" and "having" are intended to be inclusive and mean that there may be additional elements other than the listed elements.
[0226] A “genetically modified” cell refers to a cell in which the nuclear, organellar or extrachromosomal nucleic acid sequences of a cell has been modified, i.e. , the cell contains at least one nucleic acid sequence that has been engineered to contain an insertion of at least one nucleotide, a deletion of at least one nucleotide, and / or a substitution of at least one nucleotide.
[0227] As various changes could be made in the above-described cells and methods without departing from the scope of the invention, it is intended that all matter contained in the above description and in the examples given below, shall be interpreted as illustrative and not in a limiting sense.
Claims
1. PROV PATENT Danforth Docket No.: DDPSC0170-401 -PCT Via EFS-webCLAIMSWhat is claimed is:
1. A method for correcting drift in retention times of mass features in mass spectrometry (MS), comprising:a. for an m / z bin, generating or having generated an extracted ion chromatogram (EIC) in each file of a dataset comprising a plurality of files;b. locally aligning mass features in an m / z bin of interest by:i. selecting a first reference retention time across files (refScan) as a scan comprising the highest intensity data points of the highest number of EICs across the files of the dataset among the plurality of EICs extracted from each file in the dataset and establishing retention-time bounds centered on the first refScan within which data points are identified and attributed to the first mass feature in each file, thereby defining a first mass feature; ii. assigning a per-file peak center (fileScan) of each remaining EIC of the plurality of EICs to a nearest refScan selected in step (b)(i), wherein a remaining EIC is an EIC in which the fileScan does not coincide with the refScan of the first mass feature;iii. shifting the data points of each remaining EIC by a time increment sufficient to align the fileScan of the EIC with its assigned refScan;iv. iteratively repeating steps (b)(i) through (b)(iii) for each nth+1 refScan to define an nth+1 mass feature, thereby assigning each remaining EIC’s nth+1 fileScan to that refScan, shifting to align, and recording bounds; wherein each iteration uses an unshifted data file to identify the nth+1 mass features, wherein the nth+1 selection excludes data points within the bounds of previously defined mass features, and wherein n is an integer having a value equal to or more than 1 ;v. recording in a shiftTable, for every refScan-file pair, the retention time of the fileScan, the retention time of the refScan to which the fileScan is assigned, a shift equal to fileScan retention time minus refScan retentionPROV PATENT Danforth Docket No.: DDPSC0170-401 -PCT Via EFS-webtime, a mass feature identifier, and optionally the m / z bin for the mass feature, a file identifier, and one or more quality flags; andvi. aligning each nthmass feature defined in steps (b)(i) through (b)(v), based on the shiftTable of step (b)(v) and using the original unshifted EIC of each data file, by shifting the data points attributed to that nthmass feature along the retention-time axis by a time increment sufficient to bring the recorded fileScan of that nthmass feature to the closest refScan in that file; thereby correcting the drift in retention times of mass features; and c. optionally, for EICs in which two or more mass features have an ambiguous shift when performing step (b), wherein an ambiguous shift occurs when two or more fileScans in a single EIC are closest to, or assigned to, the same refScan such that more than one shift value is recorded for that refScan-file situation, resolving the ambiguity by:i. identifying a chromolog for each of the two or more mass features by:
1. constructing from the retention times of fileScans recorded in the shiftTable, for a query mass feature that has an ambiguous shift, a shiftRow, wherein the shiftRow comprises the recorded retention time of the refScan of the query mass feature, the recorded retention times of each fileScan of the query mass feature in each file of the dataset, and a mass feature identifier;2. constructing, across the same files, a corresponding shiftRow for one or more candidate chromologs, wherein a “candidate chromolog” is a mass feature in a different m / z bin whose refScan retention time matches, within a predefined retention-time tolerance, the query mass feature’s refScan retention time; and 3. assigning a chromolog from the one or more candidate chromologs to the query mass feature by computing, from the shiftRows, an aggregate co-location score defined as the sum, over files in which both the query mass feature and a candidate chromolog have recorded fileScans, of the value of the difference between thePROV PATENT Danforth Docket No.: DDPSC0170-401 -PCT Via EFS-webcandidate chromolog’s fileScan retention time and the query mass feature’s corresponding fileScan retention time in the same file; and selecting as the chromolog the candidate chromolog that minimizes said aggregate co-location score;ii. aligning fileScans of the query mass feature, in each file where the query mass feature has an ambiguous shift, by first shifting, using the original unshifted EIC of that file for the chromolog’s m / z bin, the data points already attributed to the chromolog along the retention-time axis by a time increment sufficient to bring the chromolog’s recorded fileScan to the chromolog’s refScan; and then, in the same file, shifting the data points already attributed to the query mass feature by the same time increment applied to the chromolog, thereby bringing the query mass feature’s recorded fileScan into the retention-time bounds centered on the refScan used for both the chromolog and the query mass feature;iii. iteratively repeating steps (c)(i) through (c)(ii) for each additional mass feature that has an ambiguous shift associated with the same refScan; andiv. optionally updating the shiftTable to record, for each file in which the chromolog is used to align the query mass feature, the corrected fileScan retention time for the query mass feature and the time increment applied; wherein a mass feature is defined as a single chromatographic peak identified within an m / z bin and retention-time peak.
2. The method of claim 1 , wherein the quality flags comprise at least: a proximity-to- refScan flag indicating whether the fileScan retention time lies within predefined bounds around the refScan; a signal-to-noise flag indicating the apex signal-to-noise ratio at the fileScan; an ambiguity flag indicating whether multiple candidate apexes were present near the refScan; and a correspondence flag indicating presence or absence of a credible fileScan for the refScan in the file.PROV PATENT Danforth Docket No.: DDPSC0170-401 -PCT Via EFS-web3. The method of claim 2, wherein the quality flags further comprise a peak shape flag indicating whether one or more peak-shape measures, including symmetry, tailing factor, or width-at-half-maximum, fall within predetermined ranges, an intensity threshold flag indicating whether the apex intensity at the fileScan meets predetermined minimum or maximum thresholds, or both.
4. The method of claim 2, wherein the quality flags further comprise an overlap or coelution flag indicating evidence of overlapping mass features within the refScan bounds, an outlier flag indicating whether the fileScan retention time or apex intensity is outside a robust distribution for the mass feature across the plurality of files, a drift magnitude flag indicating categorical magnitude of the shift relative to a threshold, or any combination thereof.
5. The method of claim 2, wherein the quality flags further comprise an overlap or coelution flag indicating evidence of overlapping mass features within the refScan bounds, an outlier flag indicating whether the fileScan retention time or apex intensity is outside a robust distribution for the mass feature across the plurality of files, a drift magnitude flag indicating categorical magnitude of the shift relative to a threshold, or any combination thereof.
6. The method of any one of the preceding claims, wherein the unshifted data file is annotated with each mass feature, peak center of each mass feature, bounds of each mass feature, data point values of each mass feature, or any combination thereof.
7. The method of any one of the preceding claims, wherein an aligned nthpeak center is the peak center of a distribution curve that encompasses data points of an nthmass feature, thereby defining the nthmass feature.
8. The method of any one of the preceding claims, wherein retention time bounds of each mass feature are determined using a smoothing algorithm.
9. The method of any one of the preceding claims, wherein a mass feature is metabolites and isomers of metabolites.PROV PATENT Danforth Docket No.: DDPSC0170-401 -PCT Via EFS-web10. The method of any one of the preceding claims, wherein the mass features are derived from amino acids, peptides, and proteins.
11. The method of any one of the preceding claims, wherein the MS is chromatography-MS.
12. The method of claim 11, wherein the chromatography-MS is liquid chromatography-MS (LC-MS).
13. A method for correcting drift in retention times of mass features in mass spectrometry (MS), the method comprising:a. for each m / z bin, generating or having generated an extracted ion chromatogram (EIC) in each file of a dataset comprising a plurality of files; wherein a file is the full LC-MS data acquired from one sample, and an EIC is the intensity-versus- retention-time (RT) signal extracted for one m / z bin within a file;b. referencing a shiftTable for a dataset comprising a plurality of files, wherein the shiftTable records, for every refScan-file pair, the retention time of the fileScan, the retention time of the refScan to which the fileScan is assigned, a mass feature identifier, and the m / z bin for the mass feature, and optionally a file identifier and one or more quality flags;c. for EICs in which two or more mass features have an ambiguous shift when performing step (b), wherein an ambiguous shift occurs when two or more fileScans in a single EIC are closest to, or assigned to, the same refScan such that more than one shift value is recorded for that refScan-file situation, resolving the ambiguity by:i. identifying a chromolog for a query mass feature by:
1. constructing from the retention times of fileScans recorded in the shiftTable, for a query mass feature that has an ambiguous shift, a shiftRow, wherein the shiftRow comprises the recorded retention time of the refScan of the query mass feature, the recorded retention times of each fileScan of the query mass feature in each file of the dataset, and a mass feature identifier;PROV PATENT Danforth Docket No.: DDPSC0170-401 -PCT Via EFS-web2. constructing, across the same files, a corresponding shiftRow for one or more candidate chromologs, wherein a “candidate chromolog” is a mass feature in a different m / z bin whose refScan retention time matches, within a predefined retention-time tolerance, the query mass feature’s refScan retention time; and 3. assigning a chromolog from the one or more candidate chromologs to the query mass feature by computing, from the shiftRows, an aggregate co-location score defined as the sum, over files in which both the query mass feature and a candidate chromolog have recorded fileScans, of the value of the difference between the candidate chromolog’s fileScan retention time and the query mass feature’s corresponding fileScan retention time in the same file; and selecting as the chromolog the candidate chromolog that minimizes said aggregate co-location score;ii. aligning fileScans of the query mass feature, in each file where the query mass feature has an ambiguous shift, by first shifting, using the original unshifted EIC of that file for the chromolog’s m / z bin, the data points already attributed to the chromolog along the retention-time axis by a time increment sufficient to bring the chromolog’s recorded fileScan to the chromolog’s refScan; and then, in the same file, shifting the data points already attributed to the query mass feature by the same time increment applied to the chromolog, thereby bringing the query mass feature’s recorded fileScan into retention-time bounds centered on the refScan used for both the chromolog and the query mass feature;iii. iteratively repeating steps (b) through (d) for each additional mass feature that has an ambiguous shift associated with the same refScan; and iv. optionally updating the shiftTable to record, for each file in which the chromolog is used to align the query mass feature, the corrected fileScan retention time for the query mass feature and the time increment applied;PROV PATENT Danforth Docket No.: DDPSC0170-401 -PCT Via EFS-webwherein a file is the full LC-MS data acquired from one sample; an extracted ion chromatogram (EIC) is the intensity-versus-retention-time signal extracted for one m / z bin within a file; a mass feature is defined as a single chromatographic peak identified within an m / z bin and retention-time peak; a refScan is a reference retention time across files selected to represent a mass feature and used as an alignment target; and a fileScan is the per-file peak center nearest a refScan recorded for that mass feature in that file.
14. The method of any one of claims 1-13, wherein the mass features comprise mass features derived from amino acids, peptides, or proteins.
15. The method of any one of claims 1-13, wherein the dataset comprises chromatographymass spectrometry data from a proteomics experiment.
16. The method of any one of claims 1-13, wherein the mass features comprise peptidederived chromatographic peaks detected across a plurality of files.
17. The method of any one of claims 1 -13, wherein a peptide or protein gives rise to a plurality of mass features that are independently aligned.
18. The method of any one of claims 1-13, wherein correcting retention-time drift is performed in the absence of identification or sequencing of peptides or proteins.
19. The method of any one of claims 1-13, wherein retention-time drift is corrected for peptide-derived mass features across samples.
20. A computing system for correcting drift in retention times of mass features in mass spectrometry (MS), the system comprising:a. a communication interface that receives EICs generated for an m / z bin in each file of a dataset comprising a plurality of files; andb. a processor that executes instructions stored in memory, wherein the processor executes the instructions to:i. locally align mass features in an m / z bin of interest by:
1. selecting a first reference retention time across files (refScan) as a scan comprising the highest intensity data points of the highest number of EICs across the files of the dataset among the pluralityPROV PATENT Danforth Docket No.: DDPSC0170-401 -PCT Via EFS-webof EICs extracted from each file in the dataset and establishing retention-time bounds centered on the first refScan within which data points are identified and attributed to the first mass feature in each file, thereby defining a first mass feature;2. assigning a per-file peak center (fileScan) of each remaining EIC of the plurality of EICs to a nearest refScan selected in step (b)(i), wherein a remaining EIC is an EIC in which the fileScan does not coincide with the refScan of the first mass feature;3. shifting the data points of each remaining EIC by a time increment sufficient to align the fileScan of the EIC with its assigned refScan; 4. iteratively repeating steps (b)(i)( 1 ) through (b)(i)(3) for each nth+1 refScan to define an nth+1 mass feature, thereby assigning each remaining EIC’s nth+1 fileScan to that refScan, shifting to align, and recording bounds; wherein each iteration uses an unshifted data file to identify the nth +1 mass features, wherein the nth+1 selection excludes data points within the bounds of previously defined mass features, and wherein n is an integer having a value equal to or more than 1 ;5. recording in a shiftTable, for every refScan-file pair, the retention time of the fileScan, the retention time of the refScan to which the fileScan is assigned, a shift equal to fileScan retention time minus refScan retention time, a mass feature identifier, and optionally the m / z bin for the mass feature, a file identifier, and one or more quality flags; and6. aligning each nthmass feature defined in steps (b)(i) through (b)(v), based on the shiftTable of step (b)(v) and using the original unshifted EIC of each data file, by shifting the data points attributed to that nthmass feature along the retention-time axis by a time increment sufficient to bring the recorded fileScan of that nthmassPROV PATENT Danforth Docket No.: DDPSC0170-401 -PCT Via EFS-webfeature to the closest refScan in that file; thereby correcting the drift in retention times of mass features; andii. optionally, for EICs in which two or more mass features have an ambiguous shift when performing step (i), wherein an ambiguous shift occurs when two or more fileScans in a single EIC are closest to, or assigned to, the same refScan such that more than one shift value is recorded for that refScan-file situation, resolve the ambiguity by:
1. identifying a chromolog for each of the two or more mass features by:a. constructing from the retention times of fileScans recorded in the shiftTable, for a query mass feature that has an ambiguous shift, a shiftRow, wherein the shiftRow comprises the recorded retention time of the refScan of the query mass feature, the recorded retention times of each fileScan of the query mass feature in each file of the dataset, and a mass feature identifier;b. constructing, across the same files, a corresponding shiftRow for one or more candidate chromologs, wherein a “candidate chromolog” is a mass feature in a different m / z bin whose refScan retention time matches, within a predefined retention-time tolerance, the query mass feature’s refScan retention time; andc. assigning a chromolog from the one or more candidate chromologs to the query mass feature by computing, from the shiftRows, an aggregate co-location score defined as the sum, over files in which both the query mass feature and a candidate chromolog have recorded fileScans, of the value of the difference between the candidate chromolog’s fileScan retention time and the query mass feature’s corresponding fileScan retention time in the same file; andPROV PATENT Danforth Docket No.: DDPSC0170-401 -PCT Via EFS-webselecting as the chromolog the candidate chromolog that minimizes said aggregate co-location score;2. aligning fileScans of the query mass feature, in each file where the query mass feature has an ambiguous shift, by first shifting, using the original unshifted EIC of that file for the chromolog’s m / z bin, the data points already attributed to the chromolog along the retention-time axis by a time increment sufficient to bring the chromolog’s recorded fileScan to the chromolog’s refScan; and then, in the same file, shifting the data points already attributed to the query mass feature by the same time increment applied to the chromolog, thereby bringing the query mass feature’s recorded fileScan into the retention-time bounds centered on the refScan used for both the chromolog and the query mass feature;3. iteratively repeating steps (c)(i) through (c)(ii) for each additional mass feature that has an ambiguous shift associated with the same refScan; and4. optionally updating the shiftTable to record, for each file in which the chromolog is used to align the query mass feature, the corrected fileScan retention time for the query mass feature and the time increment applied;wherein a mass feature is a single chromatographic peak identified within an m / z bin and retention-time peak.
21. The computing system of claim 20, wherein the processor further executes instructions to generate the EIC in each file of a dataset comprising a plurality of files.
22. A computing system for correcting drift in retention times of mass features in mass spectrometry (MS), the system comprising:a. a communication interface that receives EICs generated for an m / z bin in each file of a dataset comprising a plurality of files; wherein a file is the full LC-MSPROV PATENT Danforth Docket No.: DDPSC0170-401 -PCT Via EFS-webdata acquired from one sample, and an EIC is the intensity-versus-retention-time (RT) signal extracted for one m / z bin within a file; andb. a processor that executes instructions stored in memory, wherein the processor executes the instructions to:i. reference a shiftTable for a dataset comprising a plurality of files, wherein the shiftTable records, for every refScan-file pair, the retention time of the fileScan, the retention time of the refScan to which the fileScan is assigned, a mass feature identifier, and the m / z bin for the mass feature, and optionally a file identifier and one or more quality flags;ii. for EICs in which two or more mass features have an ambiguous shift when performing step (i), wherein an ambiguous shift occurs when two or more fileScans in a single EIC are closest to, or assigned to, the same refScan such that more than one shift value is recorded for that refScan- file situation, resolve the ambiguity by:
1. identifying a chromolog for a query mass feature by:a. constructing from the retention times of fileScans recorded in the shiftTable, for a query mass feature that has an ambiguous shift, a shiftRow, wherein the shiftRow comprises the recorded retention time of the refScan of the query mass feature, the recorded retention times of each fileScan of the query mass feature in each file of the dataset, and a mass feature identifier;b. constructing, across the same files, a corresponding shiftRow for one or more candidate chromologs, wherein a “candidate chromolog” is a mass feature in a different m / z bin whose refScan retention time matches, within a predefined retention-time tolerance, the query mass feature’s refScan retention time; andc. assigning a chromolog from the one or more candidate chromologs to the query mass feature by computing, fromPROV PATENT Danforth Docket No.: DDPSC0170-401 -PCT Via EFS-webthe shiftRows, an aggregate co-location score defined as the sum, over files in which both the query mass feature and a candidate chromolog have recorded fileScans, of the value of the difference between the candidate chromolog’s fileScan retention time and the query mass feature’s corresponding fileScan retention time in the same file; and selecting as the chromolog the candidate chromolog that minimizes said aggregate co-location score;2. aligning fileScans of the query mass feature, in each file where the query mass feature has an ambiguous shift, by first shifting, using the original unshifted EIC of that file for the chromolog’s m / z bin, the data points already attributed to the chromolog along the retention-time axis by a time increment sufficient to bring the chromolog’s recorded fileScan to the chromolog’s refScan; and then, in the same file, shifting the data points already attributed to the query mass feature by the same time increment applied to the chromolog, thereby bringing the query mass feature’s recorded fileScan into retention-time bounds centered on the refScan used for both the chromolog and the query mass feature;3. iteratively repeating steps (i) through (iii) for each additional mass feature that has an ambiguous shift associated with the same refScan; and4. optionally updating the shiftTable to record, for each file in which the chromolog is used to align the query mass feature, the corrected fileScan retention time for the query mass feature and the time increment applied;wherein a file is the full LC-MS data acquired from one sample; an extracted ion chromatogram (EIC) is the intensity-versus-retention-time signal extracted for one m / z bin within a file; a mass feature is defined as a single chromatographic peak identified within an m / z bin and retention-time peak; a refScan is a referencePROV PATENT Danforth Docket No.: DDPSC0170-401 -PCT Via EFS-webretention time across files selected to represent a mass feature and used as an alignment target; and a fileScan is the per-file peak center nearest a refScan recorded for that mass feature in that file.
23. A non-transitory, computer-readable storage medium, having embodied thereon a program executable by a processor to perform a method for correcting drift in retention times of mass features in mass spectrometry (MS), the method comprising:a. for an m / z bin, generating or having generated an extracted ion chromatogram (EIC) in each file of a dataset comprising a plurality of files;b. locally aligning mass features in an m / z bin of interest by:i. selecting a first reference retention time across files (refScan) as a scan comprising the highest intensity data points of the highest number of EICs across the files of the dataset among the plurality of EICs extracted from each file in the dataset and establishing retention-time bounds centered on the first refScan within which data points are identified and attributed to the first mass feature in each file, thereby defining a first mass feature;ii. assigning a per-file peak center (fileScan) of each remaining EIC of the plurality of EICs to a nearest refScan selected in step (b)(i), wherein a remaining EIC is an EIC in which the fileScan does not coincide with the refScan of the first mass feature;iii. shifting the data points of each remaining EIC by a time increment sufficient to align the fileScan of the EIC with its assigned refScan; iv. iteratively repeating steps (b)(i) through (b)(iii) for each nth+1 refScan to define an nth+1 mass feature, thereby assigning each remaining EIC’s nth+1 fileScan to that refScan, shifting to align, and recording bounds; wherein each iteration uses an unshifted data file to identify the nth +1 mass features, wherein the nth+1 selection excludes data points within the bounds of previously defined mass features, and wherein n is an integer having a value equal to or more than 1 ;PROV PATENT Danforth Docket No.: DDPSC0170-401 -PCT Via EFS-webv. recording in a shiftTable, for every refScan-file pair, the retention time of the fileScan, the retention time of the refScan to which the fileScan is assigned, a shift equal to fileScan retention time minus refScan retention time, a mass feature identifier, and optionally the m / z bin for the mass feature, a file identifier, and one or more quality flags; andvi. aligning each nthmass feature defined in steps (b)(i) through (b)(v), based on the shiftTable of step (b)(v) and using the original unshifted EIC of each data file, by shifting the data points attributed to that nthmass feature along the retention-time axis by a time increment sufficient to bring the recorded fileScan of that nthmass feature to the closest refScan in that file; thereby correcting the drift in retention times of mass features; andc. optionally, for EICs in which two or more mass features have an ambiguous shift when performing step (b), wherein an ambiguous shift occurs when two or more fileScans in a single EIC are closest to, or assigned to, the same refScan such that more than one shift value is recorded for that refScan-file situation, resolving the ambiguity by:i. identifying a chromolog for each of the two or more mass features by:
1. constructing from the retention times of fileScans recorded in the shiftTable, for a query mass feature that has an ambiguous shift, a shiftRow, wherein the shiftRow comprises the recorded retention time of the refScan of the query mass feature, the recorded retention times of each fileScan of the query mass feature in each file of the dataset, and a mass feature identifier;2. constructing, across the same files, a corresponding shiftRow for one or more candidate chromologs, wherein a “candidate chromolog” is a mass feature in a different m / z bin whose refScan retention time matches, within a predefined retention-time tolerance, the query mass feature’s refScan retention time; andPROV PATENT Danforth Docket No.: DDPSC0170-401 -PCT Via EFS-web3. assigning a chromolog from the one or more candidate chromologs to the query mass feature by computing, from the shiftRows, an aggregate co-location score defined as the sum, over files in which both the query mass feature and a candidate chromolog have recorded fileScans, of the value of the difference between the candidate chromolog’s fileScan retention time and the query mass feature’s corresponding fileScan retention time in the same file; and selecting as the chromolog the candidate chromolog that minimizes said aggregate co-location score;ii. aligning fileScans of the query mass feature, in each file where the query mass feature has an ambiguous shift, by first shifting, using the original unshifted EIC of that file for the chromolog’s m / z bin, the data points already attributed to the chromolog along the retention-time axis by a time increment sufficient to bring the chromolog’s recorded fileScan to the chromolog’s refScan; and then, in the same file, shifting the data points already attributed to the query mass feature by the same time increment applied to the chromolog, thereby bringing the query mass feature’s recorded fileScan into the retention-time bounds centered on the refScan used for both the chromolog and the query mass feature; iii. iteratively repeating steps (c)(i) through (c)(ii) for each additional mass feature that has an ambiguous shift associated with the same refScan; andiv. optionally updating the shiftTable to record, for each file in which the chromolog is used to align the query mass feature, the corrected fileScan retention time for the query mass feature and the time increment applied;wherein a mass feature is defined as a single chromatographic peak identified within an m / z bin and retention-time peak.PROV PATENT Danforth Docket No.: DDPSC0170-401 -PCT Via EFS-web24. A non-transitory, computer-readable storage medium, having embodied thereon a program executable by a processor to perform a method for correcting drift in retention times of mass features in mass spectrometry (MS), the method comprising:a. for each m / z bin, generating or having generated an extracted ion chromatogram (EIC) in each file of a dataset comprising a plurality of files; wherein a file is the full LC-MS data acquired from one sample, and an EIC is the intensity-versus- retention-time (RT) signal extracted for one m / z bin within a file;b. referencing a shiftTable for a dataset comprising a plurality of files, wherein the shiftTable records, for every refScan-file pair, the retention time of a fileScan, the retention time of a refScan to which the fileScan is assigned, a mass feature identifier, and the m / z bin for the mass feature, and optionally a file identifier and one or more quality flags;c. for EICs in which two or more mass features have an ambiguous shift when performing step (b), wherein an ambiguous shift occurs when two or more fileScans in a single EIC are closest to, or assigned to, the same refScan such that more than one shift value is recorded for that refScan-file situation, resolving the ambiguity by:i. identifying a chromolog for a query mass feature by:
1. constructing from the retention times of fileScans recorded in the shiftTable, for a query mass feature that has an ambiguous shift, a shiftRow, wherein the shiftRow comprises the recorded retention time of the refScan of the query mass feature, the recorded retention times of each fileScan of the query mass feature in each file of the dataset, and a mass feature identifier;2. constructing, across the same files, a corresponding shiftRow for one or more candidate chromologs, wherein a “candidate chromolog” is a mass feature in a different m / z bin whose refScan retention time matches, within a predefined retention-time tolerance, the query mass feature’s refScan retention time; andPROV PATENT Danforth Docket No.: DDPSC0170-401 -PCT Via EFS-web3. assigning a chromolog from the one or more candidate chromologs to the query mass feature by computing, from the shiftRows, an aggregate co-location score defined as the sum, over files in which both the query mass feature and a candidate chromolog have recorded fileScans, of the value of the difference between the candidate chromolog’s fileScan retention time and the query mass feature’s corresponding fileScan retention time in the same file; and selecting as the chromolog the candidate chromolog that minimizes said aggregate co-location score;ii. aligning fileScans of the query mass feature, in each file where the query mass feature has an ambiguous shift, by first shifting, using the original unshifted EIC of that file for the chromolog’s m / z bin, the data points already attributed to the chromolog along the retention-time axis by a time increment sufficient to bring the chromolog’s recorded fileScan to the chromolog’s refScan; and then, in the same file, shifting the data points already attributed to the query mass feature by the same time increment applied to the chromolog, thereby bringing the query mass feature’s recorded fileScan into retention-time bounds centered on the refScan used for both the chromolog and the query mass feature;iii. iteratively repeating steps (b) through (d) for each additional mass feature that has an ambiguous shift associated with the same refScan; and iv. optionally updating the shiftTable to record, for each file in which the chromolog is used to align the query mass feature, the corrected fileScan retention time for the query mass feature and the time increment applied; wherein a file is the full LC-MS data acquired from one sample; an extracted ion chromatogram (EIC) is the intensity-versus-retention-time signal extracted for one m / z bin within a file; a mass feature is defined as a single chromatographic peak identified within an m / z bin and retention-time peak; a refScan is a reference retention time across files selected to represent a mass feature and used as anPROV PATENT Danforth Docket No.: DDPSC0170-401 -PCT Via EFS-webalignment target; and a fileScan is the per-file peak center nearest a refScan recorded for that mass feature in that file.