System and method for chromatographic data review via assignment of statistical distances
The method automates chromatogram validation using Markov Chain Monte Carlo analysis to rank chromatograms based on statistical distances, reducing manual review time and ensuring accuracy in chromatographic data analysis.
Patent Information
- Application Number
- PCT/EP2025/071553
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-26
- Filing Date
- 2025-07-25
- Publication Date
- 2026-01-29
AI Technical Summary
Laboratory facilities face a significant burden in manually reviewing chromatograms to ensure accuracy, consuming up to 80% of user time, while existing methods lack efficient automated validation techniques.
A method and instrument that utilize Markov Chain Monte Carlo statistical analysis to rank chromatograms based on statistical distances from a consensus data distribution, reducing manual review by automatically identifying chromatograms outside acceptable tolerances.
This approach significantly reduces manual review time by automating the validation process, ensuring high accuracy in chromatographic data analysis by ranking and ordering chromatograms, allowing quick identification of outliers.
Smart Images

Figure EP2025071553_29012026_PF_FP_ABST
Abstract
Description
Attorney Docket No. M-4739-WO SYSTEM AND METHOD FOR CHROMATOGRAPHIC DATA REVIEW VIAASSIGNMENT OF DISTANCES CROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims priority to U.S. Provisional Application Serial No. 63 / 676,046entitled "SYSTEM AND METHOD FOR CHROMATOGRAPHIC DATA REVIEW VIAASSIGNMENT OF DISTANCES" filed July 26, 2024, which is incorporated herein by referencein its entirety. FIELD OF THE TECHNOLOGY
[0002] The present disclosure relates to methods, techniques, and processes for validation ofchromatographic data. BACKGROUND
[0003] Typical laboratory facilities are tasked with analyzing hundreds of samples each day. Toensure accuracy of all test results, the users of the analytical equipment have to continuously monitor the quality of the chromatograms that are produced. For example, users of the analyticalequipment can spend up to about 10% of their time checking that certain statistics lie withinacceptable ranges (usually in comparison to known standards). Up to 80% of the user’s time canbe spent manually reviewing each chromatogram to ensure the chromatogram accurately reflects the composition of the sample. There is a need to reduce the burden of manual chromatogramreview by laboratory users while maintaining high levels of accuracy in testing facilities.SUMMARY
[0004] These unmet needs are addressed by the present instrument and method ofchromatographic data validation. For example, present methods can be used to classify multiplechromatograms obtained from multiple samples having the same general composition. By applying Markov Chain Monte Carlo statistical analysis to the chromatograms, the chromatograms can be ranked and ordered based on a statistical “distance from consensus” measure. The ranking and ordering of chromatograms can improve manual review of the chromatograms by allowing the 1Attorney Docket No. M-4739-WO user to easily review the worst chromatograms until the chromatograms fall within the acceptable limits of the test.
[0005] In general, the present technology is directed to methods of interpretating collectedinstrument data. The method collects raw data from an analytical instrument (e.g., LC, MS, LC- MS), determines statistical distances from a data distribution center for chromatographic peak features and / or MRM transition data, and ranks each chromatogram based on the determined statistical distances. The ranked chromatograms can be sorted to facilitate review of numerouschromatograms to exclude data that is outside of an accepted tolerance for the analytical test. Thepresent technology can address the challenges associated with manual review of chromatographic data.
[0006] In one aspect, the present technology is directed to a chromatographic instrument forquantifying analytes. The chromatographic instrument includes a processing device for executing computer readable instructions for performing a method of quantifying analytes. The method includes collecting chromatograms of one or more analytes of one or more samples; determining statistical distances with respect to a consensus data distribution for chromatographic peak features or MRM transition data of the collected chromatograms; and ranking each chromatogram based on the determined statistical distances.
[0007] The above aspect can include one or more of the following embodiments. In anembodiment, the statistical distances comprise one or more of: retention time (RT) distance withrespect to a retention time consensus data distribution; full width, half maximum (FWHM) distance with respect to a FWHM consensus data distribution; peak area (PA) distance with respect to a peak area consensus data distribution; peak asymmetry (ASYM) distance with respect to a ASYM consensus data distribution; and peak height (PH) distance with respect to a PH consensus data distribution. In an embodiment, the chromatograms can be ranked based on a mathematicalcombination of a plurality of the statistical distances. In an embodiment, the statistical distancesare calculated using a Markov Chain Monte Carlo (MCMC) method. In an embodiment, raw chromatographic data of the collected chromatograms are collected from a plurality of analytes that are run simultaneously or sequentially. In some embodiments, raw chromatographic data includes retention times and relative abundances. In an embodiment, the one or more samples comprise endogenous or isotopically labeled analytes. In an embodiment, the chromatographicAttorney Docket No. M-4739-WO instrument is a liquid chromatographic instrument. In some embodiments, the chromatographic instrument comprises a mass spectrometer.
[0008] In some embodiments, peak detection and ranking for chromatograms is based onvariations in retention time and peak width. When analyzing MS data, ranking can be based onvariations of peak height and peak area of the fragments of the sample generated during the MS process. The statistical “distances” are defined in terms of central estimates of these parameters in an ideal or average sample. BRIEF DESCRIPTION OF DRAWINGS
[0009] The technology will be more fully understood from the following detailed description takenin conjunction with the accompanying drawings, in which:
[0010] FIG. 1A depicts a simulated distribution of test points.
[0011] FIG. 1B depicts a GOOD classification of test points using a known Gaussian distribution.
[0012] FIG. 1C depicts a GOOD classification of test points using all points.
[0013] FIG. 1D depicts a Bayesian Markov Chain Monte Carlo method to explore GOOD / BADstates of a collection of test points.
[0014] FIG. 2 depicts a graphical representation of collected data with analyte everolimus and itsisotopically labeled internal standard.
[0015] FIG. 3A depicts an analysis of the features of a chromatographic peak which are expectedto vary significantly from one MRM transition to another of a particular precursor, consistently across a batch of samples or with pre-established parameters.
[0016] FIG. 3B depicts an analysis of the features of a chromatographic peak which are notexpected to vary significantly from one MRM transition to another of a particular precursor.
[0017] FIG. 4 depicts an exemplary histogram of distances from the distribution center that can beused to rank chromatograms.Attorney Docket No. M-4739-WO
[0018] FIG. 5A shows an analysis of a first chromatogram identified in FIG. 4 as having one ofthe highest total distance scores.
[0019] FIG.5B shows an analysis of a second chromatogram, identified as Peak 1 in FIG. 4, whichhas one of the highest total distance scores.
[0020] FIG. 5C shows an analysis of a third chromatogram identified in FIG. 4 as having one ofthe highest total distance scores.
[0021] FIG. 5D shows an analysis of a fourth chromatogram identified in FIG. 4 as having one ofthe highest total distance scores.
[0022] FIG. 6A shows an analysis of a first chromatogram identified in FIG. 4 as having a hightotal distance score.
[0023] FIG.6B shows an analysis of a second chromatogram, identified as Peak 2 in FIG. 4, whichhas a high total distance score.
[0024] FIG. 6C shows an analysis of a third chromatogram identified in FIG. 4 as having a hightotal distance score.
[0025] FIG. 6D shows an analysis of a fourth chromatogram identified in FIG. 4 as having a hightotal distance score.
[0026] FIG. 7A shows an analysis of a first chromatogram, identified as Peak 3 in FIG. 4, whichhas a moderately high total distance score.
[0027] FIG.7B shows an analysis of a second chromatogram, identified as Peak 4 in FIG. 4, whichhas a moderately high total distance score.
[0028] FIG. 7C shows an analysis of a third chromatogram identified in FIG. 4 as having amoderately high total distance score.
[0029] FIG. 7D shows an analysis of a fourth chromatogram identified in FIG. 4 as having amoderately high total distance score.Attorney Docket No. M-4739-WO
[0030] FIG. 8A depicts the area distance versus retention time distance.
[0031] FIG. 8B depicts the area distance versus peak FWHM distance.
[0032] FIG. 8C depicts the peak FWHM distance versus retention time distance.
[0033] FIG. 8D depicts a total distance versus posterior probability (Pr) of a peak being in the ONset.
[0034] FIG. 9 is a repeat of FIG. 8D, with the different sample types indicated.
[0035] FIG. 10A shows a chromatogram for a first standard peak indicated in FIG. 9 (circled).
[0036] FIG. 10B shows a chromatogram for a second standard peak indicated in FIG. 9 (ellipse).
[0037] FIG. 10C shows an internal standard chromatogram for comparison with the first andsecond standard peaks.
[0038] FIG. 11A shows a quantifier chromatogram.
[0039] FIG. 11B shows a chromatogram for the QC peak in FIG. 9 (square).
[0040] FIG. 11C shows an internal standard chromatogram for comparison with the QC peak.DETAILED DESCRIPTION
[0041] In general, the present technology is directed to a method and instrument for quantifyinganalytes in a sample. The method collects raw data from an analytical instrument (e.g., LC, MS,LC-MS), determines statistical distances from a consensus data distribution for chromatographic peak features or MRM transition data, and ranks each chromatogram based on the determined statistical distances. The ranked chromatograms can be sorted to facilitate review of numerous chromatograms to exclude data that is outside of an accepted tolerance for the analytical test. The present technology can address the challenges associated with manual review of chromatographic data.Attorney Docket No. M-4739-WO
[0042] In an example, the method collects raw chromatographic data from an analyticalinstrument. Preferably, the analytical instrument is a mass spectrometer (“MS”) or a liquidchromatography (LC) instrument.
[0043] In an example, the analytical instrument (e.g., MS or LC) from which raw data is collectedalso includes a processor that uses the sample classification method of the present technology. Theinstrument may collect a batch of samples / analytes either simultaneously or in sequence, which isadvantageous given that the method is MRM-like, in that it may run multiple analyzes for multiple analytes. The analytical instrument may include a processing device for executing computer readable instructions for performing a method of quantifying analytes.
[0044] A method of quantifying analytes includes collecting chromatograms of one or moreanalytes of one or more samples. The chromatograms include raw (unprocessed) chromatographic data related to the analytes present in the sample. Once the raw chromatographic data is collected, peak detection is performed by using peak detection parameters.
[0045] In one embodiment, a ranking scheme is developed and used to rank the obtainedchromatograms. In particular, the present technology may reduce the burden of manual chromatogram review on the user by employing a ranking scheme for chromatographic peaks from the raw chromatographic data. If the ranking scheme, which places chromatographic peaks in order from worst to best, is reliable, the measurements requiring adjustment or rejection should almostalways appear above those that are immediately acceptable (i.e., semi-automated review).Additionally, components of statistical distance may be provided to aid interpretability of the ranking scheme. These components would reflect the degree of misfit of various aspects for the chromatographic peak measurement, e.g., consistency of retention time placement, peak width and peak area (including with respect to any ion ratio information). Given the ranking and separate components of distance, the user should quickly be able to assess the point at which further review is unnecessary.
[0046] The ranking scheme peak detection parameter may use one or more model parameters tooptimize the ranking scheme. Examples of one or more model parameters include batch center,sample center (for each sample), compound center (for each compound, relative to the sample center), variation of transition measurements from compound center, overall variance scale forAttorney Docket No. M-4739-WO measurement, precursor abundances (one per compound), transition efficiencies (one for each transition), and overall variance scale for measurement.
[0047] Using estimates of model parameters,) a distance for each data point from the ideal isconstructed. The estimates may be calculated at each iteration of an MCMC algorithm, and thesquares of the distances averaged over the MCMC run to produce the final distances.
[0048] A Good-Bad data model approach is used to assign peaks either to an ON set governed bysub-model types A and B, or OFF data governed by sub-model type C. Type A for measurementsof ON peaks, such as retention time or peak width, generally have no systematic variation betweenMRM transitions of a compound in particular sample. Type B for measurements of ON peaks is associated with abundance, such as peak height or area, that do have a systematic variation betweenMRM transitions of a compound in a particular sample (due to different efficiencies of thefragmentation of precursor to product). Type C model is used for all the attributes of OFF peaks. This may be a simple uniform model. Distances are defined in terms of central estimates of the parameters of the type A and type B models for peaks in the ON group.
[0049] Type A parameters that can be applied to data include: (1) batch center; (2) sample center(for each sample); (3) compound center (for each compound, relative to the sample center); (4)variation of transition measurements from the compound center; and (5) overall variance scale formeasurement.
[0050] Type B parameters that can be applied to data include: (1) precursor abundances (one percompound); (2) transition efficiencies (one for each transition); and (3) overall variance scale for measurement.
[0051] Using estimates of these parameters a distance for each data point from the ideal can beconstructed. The estimates may be calculated at each iteration of an MCMC algorithm, and the squares of the distances averaged over the MCM run to produce the final distances.
[0052] The analytes of the sample are not particularly limited in chemical structure. In an example,the analytes may be one or more endogenous peptides. Preferably, the analyte includes peptides from cellular samples that either contain disease state or are free of a disease state. The sample may contain internal standards from which feature information is inferred. In a preferred example, the sample includes isotopically labeled peptides.Attorney Docket No. M-4739-WO
[0053] The analytes of the sample can include biological analytes including peptides, nucleicacids, sugars, and lipids. Additional analytes of the samples include organic molecules, particularly organic compounds used for the treatment of physiological conditions and / or diseases.
[0054] Once a model is developed for a particular analyte, the method may be applied to othersimilar analytes. Purely as a quality check, experts may compare MRM chromatograms with theirown manual interpretations to validate method results.Ranking Chromatographic Peak Measurements – Mahalanobis Distance
[0055] In one embodiment, Mahalanobis distance can be used to determine if a chromatographicpeak is within the accepted tolerance of the analysis procedure (“GOOD”) or outside the accepted tolerance (BAD). Chromatograms with BAD data would be flagged for manual review based on a ranking score.
[0056] Mahalanobis distance is determined using a Gaussian sampling distribution of the data.Generally, the Mahalanobis distance is the distance of a test point from the center (e.g., the mean)of an elliptical or hyper-elliptical Gaussian distribution. The Mahalanobis distance can becalculated by determining the distance of a test point from the center and dividing by the width ofthe elliptical distribution along the direction of the test point from the center.Clustering Data With Outliers
[0057] FIG. 1A depicts a simulated distribution of test points. In FIG. 1A, 40 BAD points arerandomly drawn in a uniform 5 X 5 unit square, centered at (0,0). 60 GOOD points are also drawnrandomly within a circularly symmetric Gaussian distribution with unit standard deviation, centered at (0,0). In FIG. 1B, the test points in the circle are classified as GOOD by using the know Gaussian distribution and selecting test points having a p-value of greater than 5%.
[0058] However, in many instances, the Gaussian distribution will not be known until after thedata has been collected. In FIG. 1C, the Gaussian distribution is estimated from all of the datacollected. The test points are then classified as GOOD by using the estimated Gaussian distribution and selecting test points having a p-value of greater than 5%. This leads to the inadvertent capture of BAD.
[0059] The problems associated with using an estimated Gaussian distribution can be amelioratedby use of a Bayesian Markov Chain Monte Carlo method to explore GOOD / BAD states of aAttorney Docket No. M-4739-WOcollection of test points (FIG. 1D). Using posterior probability to classify the points (GOOD testpoints have a posterior probability of >95%), a selection of GOOD points can be achieved that iscommensurate with the classification achieved using a known Gaussian distribution (FIG. 1B).While Bayesian methods are described herein, it should be understood that other outlier detection methods can be used.Application Example – Everolimus Immunosuppressant Data
[0060] Everolimus and its 13C22H4 labelled internal standard were analyzed by LC-MS.Chromatography runs were made of analyte (Analyte*), calibrator 0, analyte (Analyte**) and solvent blank (Blank). The internal standard was subtracted from analyte measurements. Two features of the peaks in the chromatogram were studied. Peak width was used as a surrogate for what a user would inspect as a “peak shape.” In some embodiments, more than one measurement of peak width peak width (for example, at different heights) and measurements of peak asymmetry can be used to capture peak shape. For peak width measurements, the logarithm of peak widthwas used, so that an error-bar associated with peak width can be described in relative terms, e.g.,a percentage. The logarithm for retention time was also used. However, retention time can also be used directly, since error-bar on measurement is not expected to change (e.g., increase) with increasing retention time.Attorney Docket No. M-4739-WOTABLE 1
[0061] FIG. 2 depicts a graphical representation of the collected data. Table 1 shows calculatedMahalanobis distances on the basis of a Gaussian model for points weighted by posterior probability. The data show that, generally, distance decreases as posterior probability increases. Modifications to Improve Accuracy
[0062] The above-described model can be improved by additional modifications to the statisticalanalysis. For example, internal standard measurements may not be reliable. To improve accuracy, internal standards can be included in the full analysis. In another example, peak areas and / or ion ratios can be considered if multiple transitions have been acquired for a particular precursor.Attorney Docket No. M-4739-WOInconsistent ratios between the quantifier and the qualifier ion area can be a sign of interference(e.g., a BAD chromatogram).
[0063] Additionally, batches of data will not always be available. In one embodiment, the methodhas the ability to learn from previous data. In one embodiment, the analysis can be primed fromprevious results or from a training set.Model Structure
[0064] Data is split into “analytical groups.” An analytical group contains related compounds,often just a single compound of interest and its internal standard, across all samples in a batch. Two competing models can be used for each chromatographic peak in the group. ON: belongs to a “cluster” of peaks fitting well with requirements of RT alignment, peak width consistency, etc. OFF: could have come from anywhere in the measurement ranges, independently of other peaks. FIG. 3A depicts an analysis of chromatographic peaks.
[0065] Two types of measurements can be applied to mass spectroscopy data. “Positional”measurements are not expected to vary with MRM transition for a given compound, e.g., retention time, peak width. Peak width and peak asymmetry is part of what a user would consider about peak shape. “Quantity” expected to vary with MRM transition, e.g., peak area. FIG. 3B depicts an analysis of MRM transitions. Ranking of Chromatograms
[0066] Using the statistical methodology set forth herein, chromatograms can be ranked indescending order of a “distance from consensus” measure. After automatic ordering of thechromatograms by the software, a user will begin manually reviewing the ordered chromatograms.Once the chromatograms that are manually reviewed are considered to be within the accepted tolerance of the analysis procedure (“GOOD”), the manual review can be discontinued and the chromatograms that were not manually reviewed can be designated as GOOD. In some instances, greater than 50%, greater than 60%, greater than 70%, greater than 80%, or greater than 90% ofthe chromatograms can be designated as GOOD and excluded from manual review.
[0067] The statistical analytical methodology described herein can used for samplechromatograms as well as control chromatograms. The statistical analytical methodology can beAttorney Docket No. M-4739-WOapplied to solvent blank chromatograms. Analysis of solvent blank chromatograms can indicatewhether the reconstitution solvents are contaminated after passing through the column.
[0068] The statistical analytical methodology can be applied to double blank chromatograms.Analysis of double blank chromatograms can indicate whether the used matrices (e.g., plasma,serum, urine) are contaminated.
[0069] The statistical analytical methodology can be applied to single blank / QC-0 chromatograms(No Analyte / with Internal Standard). Analysis of single blank chromatograms can indicatewhether the used internal standard (e.g., a stable isotope labelled standard) is contaminated.Blanks are studied because consistent (low distance) peaks in blanks may be a sign of problems,indicating carry-over or contamination. On the other hand, high distance peaks in blanks may stillimpinge upon the region of interest for a genuine analyte and so may also be of concern.
[0070] The method of ranking chromatograms is performed by determining statistical distancesfrom a data distribution center for chromatographic peak features and / or MRM transition data. Forexample, statistical distances can be calculated based on one or more data points of the chromatogram. Data points that can be used to calculate statistical distances include, but are not limited to: retention time (RT) distance with respect to a RT consensus data distribution; full width, half maximum (FWHM) distance with respect to a FWHM consensus data distribution; peak area (PA) distance with respect to a PA data consensus data distribution; peak asymmetry (ASYM) distance with respect to a ASYM consensus data distribution; and peak height (PH) distance with respect to a PH consensus data distribution..
[0071] In one embodiment, the chromatograms are ranked based on the statistical distances thatare calculated. The “highest” ranked chromatogram is the chromatogram with the highest calculated statistical distance. If multiple statistical distances were determined, the rank of the chromatogram is based on a mathematical combination of a plurality of the statistical distances. For example, the combined distances can be computed using as the Pythagorean distance (i.e., D2= d12+ d22+…dn2where D is the total statistical distance and dnis one of the individual statisticaldistances). The chromatogram with the highest mathematical combination of the statisticaldistances is given the highest rank. For example, the ranking of the chromatograms can beachieved by summing one or more of the RT distance, the FWHM distance, and the PA distance.Attorney Docket No. M-4739-WO
[0072] After the chromatograms are ranked, the chromatograms can be ordered (e.g., by sorting)sequentially, starting with the highest ranked chromatogram. Placing the highest ranked chromatograms in order places all the BAD chromatograms together. The ordered chromatograms are then reviewed sequentially, starting with the highest ranked chromatogram. As the review proceeds, the chromatograms will have progressively lower rankings, and will therefore be closer to GOOD chromatograms. Once the reviewer reaches the GOOD chromatograms, the review process can stop. The remaining, unreviewed chromatograms can be considered valid based on the low statistical distance from the ideal center of the distribution.
[0073] FIG. 4 depicts an exemplary histogram of distances from the consensus data distributionthat can be used to rank chromatograms. In the example presented, an LC-MS chromatographicanalysis is performed on a sample having an internal standard and two associated analytes (Analyte 1 and Analyte 2). After ranking the chromatograms based on total distance (based on the sum of the squares of the determined RT distance, FWHM distance, and PA distance) a sample section ofthe chromatograms is identified for manual review (designated as “Peak to review” in FIG. 4).Table 2 lists the total distance scores associated with peaks identified in FIG. 4 as Peaks 1-4.TABLE 2
[0074] FIGS. 5A-5D show an analysis of four different chromatograms identified as having thehighest total distance scores (total distance greater than 10). FIG. 5B shows the chromatogramassociated with Peak 1 (identified in FIG.4) is wide and its area ratio is high compared with normalpeaks.
[0075] FIGS. 6A-6D shows an analysis of four different chromatograms identified as having hightotal distance scores (total distance less than 10, greater than 5). FIG.6B shows the Peak 2, whichAttorney Docket No. M-4739-WO also has a relatively high total distance score. FIG. 6B shows the chromatogram associated with Peak 2 (identified in FIG.4) is wide and also includes an impurity having an earlier retention time.
[0076] FIGS. 7A-7D shows an analysis of four different chromatograms identified as havingmoderately high total distance scores (total distance less than 5, greater than 1). Although Peak 3 (FIG.7A) and Peak 4 (FIG.7B) have consistent shapes, the area ratio between the peaks is a little high. MCMC Statistical Model of Distance Measurements
[0077] It is assumed that set of compounds analyzed in a particular batch of samples is brokendown into analytical groups. Usually, a group will contain a single internal standard and a common use-case has groups comprising a single analyte compound and its isotopically labelled version as the internal standard. The analysis of the analytical groups is taken to be independent.
[0078] There are K samples to analyze for which there are Ji transitions for the ith compound inthe analytical group.
[0079] Each measurement dimension can be considered independently, e.g., retention time, sothat xijk refers to the measured retention time of the jth transition for the ith com- pound in theanalytical group. The possible range of measurements is taken to have size ∆.
[0080] The position might vary with sample, so the central position µk for the com- pounds in asingle sample can be allowed to be distributed around some global central position ν.
[0081] There might be systematic differences between the different compounds in the analyticalgroup which can be modeled as a shift δµi, so that we expect xijk ≈ µk + δµi for I, j in the acquired set of transitions.
[0082] Ion areas and ratios are treated slightly differently as there are a set of transitionefficiencies associated with a compound which are scaled by a quantity of the compound in a particular sample.
[0083] The aim is to produce a system that will provide “distances” for each measurement fromsome central estimate and probabilities of “goodness”, i.e., of the measurement belonging to a consensus of good measurements. This consensus might come from the analysis of an individual batch or may also be influenced by historical / training data from previous acquisitions.Attorney Docket No. M-4739-WO
[0084] The analysis can be divided into two parts:1. Those measurements which are expected to be invariant (within statistical error) acrossMRM transitions of the same precursor ion but may have systematic differences from sample to sample or compound to compound, e.g., retention time and peak width. These are call positional measurements due to their similarity to retention time (position along the x-axis) in this respect. 2. Those measurements which are expected to be invariant (within statistical error) acrosssamples and compounds but have systematic differences across MRM transitions of the same precursor ion, e.g., peak area. These are called quantity measurements. Completing the square and marginalization
[0085] Given a quadratic expression, Aθ2−2Bθ+C, we may “complete the square” to giveIn the following, we frequently encounter Gaussian joint probabilities of the formMarginalizing out θ involves taking the integral of the joint probability over some prior range ∆. The range may be infinite or large enough to justify approximating the result by integrating over infinite range,Attorney Docket No. M-4739-WO
[0086] The central estimate θˆθ = A−1b has covariance. Actually, the inverse A−1 need notbe calculated if covariances are not required; as A is symmetric and positive definite (a benefit of having proper priors), we may use Cholesky decomposition to find the lower triangular matrix L such that LLT= A. It is easy to solve the triangular system of equations Lx = b to find x = L−1b, so that bTA−1b = bTL−TL−1b = x · x. The central estimate θˆθ = A−1b = L−Tx is the solution to the (upper) triangular system of equations LTθˆθ = x. Markov Chain Monte Carlo
[0087] The switch states controlling the OFF / ON status of each measurement are explored usingMarkov Chain Monte Carlo (MCMC) techniques, as are the various variances associated the prior positional and quantity centers.
[0088] The switch states may be sampled using Gibbs sampling, where a new state (which maybe the same as the old state) is simply sampled from the prior probability distribution for OFF / ON for the particular sample and compound combination. The ON state may be subdivided to select a particular chromatographic peak if more than one has been measured in a particular chromatogram.
[0089] For the variances, we require a prior probability distribution which is positive only in (0,∞). An effective technique for exploring these parameters is slice sampling for which convenientAttorney Docket No. M-4739-WO priors have an easily invertible cumulant. One way to achieve both these requirements is to use a logistic prior on the logarithm of the standard deviation, for example,
[0090] Here ξ is the mean, median and mode of the logistic distribution while the scaleparameter ζ may be set by choosing the values of particular quantiles or choosing a standard ^^ deviation equal to√^The logistic distribution has heavier tails than a normal distribution with the same standard deviation which might be advantageous in the context of MCMC exploration.
[0091] A simple MCMC implementation would sample a state for the entire system from thecombined prior and then allow the state to evolve through a series of transitions, each obeying detailed balance. One iteration of this simple method might involve sampling all the parameters, accepting new states if they meet or exceed a log-likelihood threshold, log L∗, set at the start of the iteration as log L∗ = log L + log Uniform (0, 1). (8)
[0092] For a number of “burn-in” iterations no statistics are collected from the samples to givetime for the state to evolve into the so-called “posterior bubble”. Thereafter, statistics on any quantity of interest may be accumulated until a sufficient number of samples has been acquired. In the present context, we are mainly interested in accumulating the posterior probabilities of the switch states and squared distances of the measurements from current central model. We may also acquire statistics relating to the model, perhaps to inform the setting of priors for subsequently analyzed data.Attorney Docket No. M-4739-WO Positional measurements
[0093] Positional measurements, such as retention time, peak width or peak asymmetry, wherethe expected value does not vary between different MRM transitions of the same precursor ion, may be transformed to a convenient axis. For retention time the most convenient axis is probably the original measurement axis, e.g., minutes, as the error-bar on the measurement is assumed not to vary with retention time. For peak width, on the other hand, the error-bar on a measurement is x assumed to be approximately proportional to the value of the measurement. In this case the ^^ logarithm is used as δ log x ≈ ^for small δ log x.Prior probability distributions
[0094] The width parameters of distributions are given relative to some single underlying scale κwhich may be marginalized away later. Global central positionSample Central PositionGlobal positional offset of a compoundAnalysisAttorney Docket No. M-4739-WO
[0095] Given the prior probability distributions and likelihood functions, we could allow theMCMC to sample all the parameters involved. However, with some effort, we may marginalize out some parameters in advance, thereby making the MCMC more efficient. Step 1: Marginalize out the µkNow perform the marginalisation, letting A^, =Njk'ijik N^yk,Taking the product over K samples, withStep 2: Marginalize out vExamining the sum in the second exponential above, we haveDefine,Attorney Docket No. M-4739-WOStep 3: Marginalize out the δμifor W) = Ylkwik at fixed i, and the fact thatAttorney Docket No. M-4739-WOStep 4: Marginalize of the variance scale κ2Attorney Docket No. M-4739-WODistance
[0096] Working backwards through equations 34 & 35 to provide central estimates of the δµithrough A−1b, equation 19 to provide a central estimate of the global position ν, and equation 14 to provide central estimates of the sample positions µk, we can revert to equation 13 to calculateAttorney Docket No. M-4739-WO its exponent. We also need an estimate of the variance scale κ2: the mode ^^=is always available but perhaps better is the reciprocal of ^^^which is simply^^. The squared distances are accumulated and averaged over the MCMC run.
[0097] The variance of the estimated parameter in each MCMC iteration may be calculated asfollows: Step 1: Marginalize out ν The marginalization of parameters may be done in any convenient order. In this case we seek covariances between the µk and δµi, so it is convenient to remove ν first. Also, as we seek only the covariance structure, we need only consider what happens in the exponent. Ignoring κ2for the moment, and denoting the exponent f (ν, µ, δµ),Step 2: Take the second derivatives of g (μ, δμ)Attorney Docket No. M-4739-WOStep 3: Marginalize out δμ and μ We may replace either δµ or µ in terms of the other as they are constrained by µ+δµ = x. We then proceed to marginalize out the remaining variable leaving,Step 4: Take second derivative of τ (x)Step 5: Final AdjustmentsAttorney Docket No. M-4739-WO Quantity measurements Prior probability distributions
[0098] We will assume a normal distribution for the compound quantities λi with mean Λi andvariance ρ2(scaled by κ2), Ion areas and ratios
[0099] It might be appropriate to have common ion ratios associated with an analyte compoundand its internal standard if the internal standard is an isotopically labelled version of the analyte compound, but in the general case each compound would be associated only with its own ion ratios. To accommodate either option we use a subscript l to denote a group of compounds expected to have the same underlying ion ratios among their transitions, and i indicates a single compound within the lthgroup (usually containing a single compound or a compound / internal standard pair).
[0100] The overall scale of ion areas is taken to be log-normally distributed with eachcompound having its own parameters for the distribution. We may, therefore, assume a normal prior in log quantity λikover range Λ. Working in terms of log area, aijkfor a particular transition j and compound i in a particular compound group l and sample k,The ϕlj ≤ 0 may be thought of as the logarithm of the efficiency of the jthtransition of the lthcompound group. Letting a′ijk = aijk–− ϕlj and rearranging the inner sum in the exponent,Attorney Docket No. M-4739-WODistance
[0101] To calculate distances (squared) of the area measurements from the centralestimate of the model for the current MCMC iteration, we need only estimate the λik, as current samples of the ϕlj are provided, along with κ2. The estimates of λik are made according to equation 53 asAttorney Docket No. M-4739-WO
[0102] As for the positional measurements, an addition of σ2 is made to account formeasurement error before scaling by the current estimate of κ2, so that the square of the distance in the current MCMC iteration isMCMC exploration of ON / OFF states
[0103] The analysis above applies to all those measurements assumed to be in theON group, but membership of that group remains to be explored. Prior probabilities of belonging to the ON group could be broken down in terms of sample type and compound type (analyte or internal standard) and transition type (Quantifier, Qualifier). For full generality, we will use individual pijkfor each transitions in the kthsample. Let M be the total number of chromatogram fs with N the number assigned to the ON group and M–− N assigned to the OFF group. The likelihood for the measurement of a particular field f in the OFF group isAttorney Docket No. M-4739-WOWe can associate switch states with each transition and explore the configuration of switches using MCMC.
[0104] EXAMPLE 1: RANKING CHROMATOGRAMS: THERAPEUTIC DRUGMONITORING (ABSOLUTE DETERMINATION)
[0105] In another example, the present technology quantifies and / or classifies absoluteamounts of one or more therapeutic drugs as a method of therapeutic drug monitoring. For example, everolimus is a commonly used immunosuppressive agent with a variety of activemechanisms such as high inter- and intra-individual variability. Therefore, an accurate,analytically sensitive quantitative method using the present technology may play a role inresearching the pharmacokinetic and pharmacodynamic effects of administration.
[0106] Here, the present technology applies a learning model of the present technologythat is optimized by determining value(s) for therapeutic drug monitoring, including bias values that may influence measurement of hematocrit. The method of the present technology uses an LC- MS / MS instrument for the analysis of dried blood spot analysis of everolimus. The everolimus sample is analyzed by using peak detection, quantifying consistency, and applying the learningmodel based on the methods disclosed herein in order determine a value (e.g., a concentration oramount) for everolimus in the sample. This information could be used for therapeutic drugmonitoring. Likewise, calculated bias values on medical decision levels showed that there was noclinical influence of hematocrit on the results. FIGS 8 and 9 show the result of processing a batchAttorney Docket No. M-4739-WO of everolimus samples using the technique described. The batch of 36 samples included 4 solvent blanks, 1 double blank and 2 single blanks.
[0107] FIG. 8A depicts the area distance versus retention time distance. FIG. 8B depictsthe area distance versus peak FWHM distance. FIG. 8C depicts the peak FWHM distance versusretention time distance. FIG. 8D depicts a total distance versus posterior probability (Pr) of a peakbeing in the ON set. Those peaks with Pr(ON) ≥ 0.95 are colored yellow, the remainder are coloredblue. FIG. 9 is a repeat of FIG. 8D, with the different sample types indicated.
[0108] The chromatograms for standard peaks indicated in FIG. 9 as shown in FIG. 10A,FIG.10B, and FIG.10C. A first standard peak chromatogram is shown in FIG.10A (circled data point in FIG. 9). A second standard peak chromatogram is shown in FIG. 10B (ellipse encircled data point in FIG.9). An internal standard chromatogram is shown in FIG.10C. The presence of a significant baseline appears to have affected the ratio of areas between the quantifier and qualifier peaks which is 1.82 compared with an overall estimate of 2.24, leading to high area distances of 38 and 30 for the quantifier and qualifier, respectively.
[0109] The chromatogram for the QC peak in FIG. 9 (square box) is shown in FIG. 11Bbetween a quantifier chromatogram (FIG. 11A) and an internal standard chromatogram (FIG.11C). In this case it is the peak width of the qualifier (0.0473 minutes) that is somewhat wider than the overall estimate (0.0449 minutes), leading to a FWHM distance of 12 for peak area.
[0110] Table 6 gives the posterior probabilities and distances associated with the 40chromatographic peaks with highest Total Distances, listed in descending order of Total Distance. This is not quite equivalent to ranking the peaks in ascending order of posterior probability, asshown in FIGS 8 and 9, because: (a) The posterior probabilities are simply the proportion of timesthe peak was in the ON group over the MCMC run of 200 iterations, so there statistical variations in the estimation of the posterior probabilities; and (b) When a peak is in the OFF group it loses connection with the various parameters describing the expected retention time, peak width and area. This can produce large contributions to the overall distance particularly from the area measurement which has the most variability.Attorney Docket No. M-4739-WO Table 6 Injection Name Sample Type Compound Type Ion Posterior Distance(RT) Distance(FWHM) Distance(AREA) Total DistanceAttorney Docket No. M-4739-WO
[0111] Although the present technology has been described with reference to preferredembodiments, it will be understood by those skilled in the art that various changes in form and detail may be made without departing from the scope of the present invention as set forth in the accompanying claims.
Claims
Attorney Docket No. M-4739-WO What is claimed is:
1. A chromatographic instrument for quantifying analytes comprising:a processing device for executing computer readable instructions for performing a method of quantifying analytes, the method comprising: collecting chromatograms of one or more analytes of one or more samples; determining statistical distances with respect to a consensus data distribution for chromatographic peak features or MRM transition data in the collectedchromatograms ; and ranking each chromatogram based on the determined statistical distances.
2. The instrument of claim 1, wherein the statistical distances comprise one or more of:retention time (RT) distance with respect to a RT consensus data distribution; full width, half maximum (FWHM) distance with respect to a FWHM consensus data distribution ; peak area (PA) distance with respect to a PA consensus data distribution; peak asymmetry (ASYM) distance with respect to a ASYM consensus data distribution; andpeak height (PH) distance with respect to a PH consensus data distribution.
3. The instrument of claim 2, wherein the collected chromatograms are ranked based on amathematical combination of a plurality of the statistical distances.
4. The instrument of claim 2 or 3, further comprising ordering the chromatogramssequentially, starting with the highest ranked chromatogram, wherein the highest ranked chromatogram has the highest statistical distance or the highest mathematical combination of the statistical distances.
5. The instrument of claim 1, wherein the statistical distances are calculated using a MarkovChain Monte Carlo (“MCMC”) method.
6. The instrument of claim 1, wherein raw chromatographic data of the collectedchromatograms are collected from a plurality of analytes that are run simultaneously or sequentially.Attorney Docket No. M-4739-WO7. The instrument of claim 6, wherein the raw chromatographic data includes retentiontimes and relative abundances.
8. The instrument of any one of claims 1 to 7, wherein the one or more samples comprise ofendogenous or isotopically labeled analytes.
9. The instrument of any one of claims 1 to 8, wherein the instrument is a liquidchromatography instrument.
10. The instrument of any one of claims 1 to 8, wherein the instrument is a massspectrometer.
11. The instrument of any one of claims 1 to 8, wherein the instrument is a liquidchromatography-mass spectrometer.
12. A method of quantifying sample data from a chromatographic instrument comprising:collecting chromatograms of one or more analytes of one or more samples; determining statistical distances from a consensus data distribution for chromatographic peak features or MRM transition data in the collected chromatograms; andranking each chromatogram based on the determined statistical distances.
13. The method of claim 12, wherein the statistical distances comprise one or more of:retention time (RT) distance with respect to a RT consensus data distribution; full width, half maximum (FWHM) distance with respect to a FWHM consensus data distribution; peak area (PA) distance with respect to a PA consensus data distribution; peak asymmetry (ASYM) distance with respect to a ASYM consensus data distribution; and peak height (PH) distance with respect to a PH consensus data distribution.
14. The method of claim 13, wherein the collected chromatograms are ranked based on amathematical combination of a plurality of the statistical distances.
15. The method of claim 13 or 14, further comprising ordering the chromatogramssequentially, starting with the highest ranked chromatogram, wherein the highest ranked chromatogram has the highest statistical distance or the highest mathematical combination of the statistical distances.Attorney Docket No. M-4739-WO16. The method of claim 15, further comprising reviewing the ordered chromatogramssequentially, starting with the highest ranked chromatogram.
17. The method of claim 12, wherein the statistical distances are calculated using a MarkovChain Monte Carlo (“MCMC”) method.
18. The method of claim 12, wherein raw chromatographic data of the collectedchromatograms is collected from a plurality of analytes that are run simultaneously orsequentially.
19. The method of claim 18, wherein the raw chromatographic data of the collectedchromatograms includes retention times and relative abundances.
20. The method of any one of claims 12 to 19, wherein the one or more samples comprise ofendogenous and isotopically labeled analytes.
21. The method of any one of claims 12 to 20, wherein the instrument is a liquidchromatography instrument.
22. The method of any one of claims 12 to 20, wherein the instrument is a mass spectrometer.
23. The method of any one of claims 12 to 20, wherein the instrument is a liquidchromatography-mass spectrometer.
Citation Information
Patent Citations
Sample quantification consistency and classification workflow
US20240102977A1