Method, apparatus, and computer-readable medium for adaptive normalization of analyte levels

Adaptive normalization techniques address the issues of artifact introduction in median normalization by iteratively removing biased analytes, ensuring accurate proteomic analysis by minimizing noise and preserving biological signals.

JP7748360B2Active Publication Date: 2025-10-02SOMALOGIC OPERATING CO INC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2022506418
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-07-31
Filing Date
2020-07-24
Publication Date
2025-10-02
Estimated Expiration
2040-07-24

AI Technical Summary

Technical Problem

Existing median normalization methods introduce artifacts and attenuate true biological signals due to violations of assumptions related to sample collection and processing, particularly in proteomic assays with chronic kidney disease, leading to systematic differences and assay noise.

Method used

Adaptive normalization techniques that iteratively remove analytes with significant differences from reference distributions, using statistical methods like Mahalanobis distance and maximum likelihood to calculate scale factors, minimizing bias and noise while preserving biological signals.

Benefits of technology

The adaptive normalization effectively removes assay bias and decorrelates noise, maintaining data integrity by excluding artifacts and preserving biological variability, thus enhancing the accuracy of proteomic analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007748360000011
    Figure 0007748360000011
  • Figure 0007748360000012
    Figure 0007748360000012
  • Figure 0007748360000013
    Figure 0007748360000013
Patent Text Reader

Abstract

The method, apparatus, and computer-readable medium for adaptive normalization of analyte levels in one or more samples includes receiving one or more analyte levels corresponding to one or more analyte levels detected in the one or more samples, each analyte level corresponding to a detected amount of the analyte in the one or more samples; iteratively applying a scale factor to the one or more analyte levels over a number of iterations until a change in the scale factor between successive iterations is below a predetermined change threshold or until an amount in one or more iterations exceeds a maximum iteration value, each iteration including determining a distance between each analyte level in the one or more analyte levels and a corresponding reference distribution for the analyte in a reference dataset; determining a scale factor based at least in part on analyte levels that are within a predetermined distance from their corresponding reference distributions; and normalizing the one or more analyte levels by applying the scale factor.
Need to check novelty before this filing date? Find Prior Art

Description

[Background technology]

[0001] This application claims priority to U.S. Provisional Application No. 62 / 880,791, filed July 31, 2019, which is incorporated herein by reference in its entirety.

[0002] Median normalization has been developed to remove certain assay artifacts from datasets prior to analysis. Such normalization can remove sample or assay bias that may result from sample-to-sample differences in overall protein concentration (e.g., due to hydration state), pipetting errors, changes in reagent concentrations, assay timing, and other sources of systematic variability within a single assay run. Furthermore, proteomic assays (e.g., aptamer-based proteomic assays) can generate correlated noise, and the normalization process has been observed to significantly reduce these artifacts.

[0003] Median normalization relies on the concept that true biological markers (related to underlying physiology) are relatively rare, so most protein measurements in highly multiplexed proteomic assays will not vary across populations of interest. Therefore, the majority of protein measurements within samples and across populations of interest can be assumed to be sampled from a common population distribution for that analyte with a well-defined center and scale. When these assumptions do not hold, median normalization can introduce artifacts into the data, attenuating true biological signals and introducing systematic differences in analytes that are not differentially expressed within the sample set.

[0004] Certain pre-analytical variables related to sample collection and processing have been observed to violate the median normalization assumption, as a large number of analytes may be affected by spinning samples or by lysing cells prior to separation from the bulk fluid. Furthermore, protein measurements in patients with chronic kidney disease have shown that the levels of hundreds of proteins are affected by this condition, resulting in elevated circulating protein concentrations in these individuals compared to those with adequately functioning kidneys. Therefore, improvements are needed in systems that adequately remove assay bias and decorrelate assay noise while preventing the introduction of artifacts in the data due to sample collection artifacts or excessive numbers of disease-related proteomic changes. [Brief explanation of the drawings]

[0005] [Figure 1] 1 shows a flowchart for determining a scale factor based at least in part on an analyte level that is within a predetermined distance from a corresponding reference distribution, according to an exemplary embodiment. [Figure 2] 2 shows an example of a sample 200 having multiple detected analytes, including 201A and 202A, according to an exemplary embodiment, including Reference Distribution 1 and Reference Distribution 2, respectively. [Figure 3] 10 illustrates a process for each iteration of a scale factor application process, according to an example embodiment. [Figure 4A] 1 illustrates an example of an adaptive normalization process for a set of sample data, according to an example embodiment. [Figure 4B] 1 illustrates an example of an adaptive normalization process for a set of sample data, according to an example embodiment. [Figure 4C] 1 illustrates an example of an adaptive normalization process for a set of sample data, according to an example embodiment. [Figure 4D] 1 illustrates an example of an adaptive normalization process for a set of sample data, according to an example embodiment. [Figure 4E] 1 illustrates an example of an adaptive normalization process for a set of sample data, according to an example embodiment. [Figure 4F] 1 illustrates an example of an adaptive normalization process for a set of sample data, according to an example embodiment. [Figure 5A] 10 illustrates another example of an adaptive normalization process requiring two or more iterations, according to an example embodiment. [Figure 5B] 10 illustrates another example of an adaptive normalization process requiring two or more iterations, according to an example embodiment. [Figure 5C] 10 illustrates another example of an adaptive normalization process requiring two or more iterations, according to an example embodiment. [Figure 5D] 10 illustrates another example of an adaptive normalization process requiring two or more iterations, according to an example embodiment. [Figure 5E] 10 illustrates another example of an adaptive normalization process requiring two or more iterations, according to an example embodiment. [Figure 6A] 1 shows the analyte levels for all samples after one iteration of the adaptive normalization process described herein. [Figure 6B] 1 shows the analyte levels for all samples after one iteration of the adaptive normalization process described herein. [Figure 7] 10 illustrates components for determining a value of the scale factor that maximizes the probability that analyte levels within a predetermined distance from their corresponding reference distributions are part of their corresponding reference distributions, according to an exemplary embodiment. [Figure 8A] 1 illustrates the application of adaptive normalization by maximum likelihood to the sample data of sample 4 shown in the figure. [Figure 8B] 1 illustrates the application of adaptive normalization by maximum likelihood to the sample data of sample 4 shown in the figure. [Figure 8C] 1 illustrates the application of adaptive normalization by maximum likelihood to the sample data of sample 4 shown in the figure. [Figure 9A] 9 illustrates the application of population adaptive normalization to the data shown in Figures 10A-10B according to an exemplary embodiment. Figure 9 illustrates another method for adaptive normalization of analyte levels in one or more samples according to an exemplary embodiment. [Figure 9B]9 illustrates the application of population adaptive normalization to the data shown in Figures 10A-10B according to an exemplary embodiment. Figure 9 illustrates another method for adaptive normalization of analyte levels in one or more samples according to an exemplary embodiment. [Figure 9C] 9 illustrates the application of population adaptive normalization to the data shown in Figures 10A-10B according to an exemplary embodiment. Figure 9 illustrates another method for adaptive normalization of analyte levels in one or more samples according to an exemplary embodiment. [Figure 9D] 9 illustrates the application of population adaptive normalization to the data shown in Figures 10A-10B according to an exemplary embodiment. Figure 9 illustrates another method for adaptive normalization of analyte levels in one or more samples according to an exemplary embodiment. [Figure 9E] 9 illustrates the application of population adaptive normalization to the data shown in Figures 10A-10B according to an exemplary embodiment. Figure 9 illustrates another method for adaptive normalization of analyte levels in one or more samples according to an exemplary embodiment. [Figure 9F] 9 illustrates the application of population adaptive normalization to the data shown in Figures 10A-10B according to an exemplary embodiment. Figure 9 illustrates another method for adaptive normalization of analyte levels in one or more samples according to an exemplary embodiment. [Figure 10] 1 illustrates a dedicated computing environment for adaptive analyte-level normalization, according to an exemplary embodiment. [Figure 11] The median coefficient of variation across all aptamer-based proteomics assay measurements for 38 technical replicates is shown. [Figure 12] 1 shows the Kolmogorov-Smirnov statistics for gender-specific biomarkers for the samples with respect to maximum allowed replicates. [Figure 13] The number of QC samples by sample ID for plasma and serum used in the analysis is shown. [Figure 14] 1 shows the agreement of QC sample scale factors using median normalization and ANML. [Figure 15]CV decomposition of control samples using median normalization and ANML. Lines indicate the empirical cumulative distribution function of the CV of each control sample across plates (inter) and within plates (intra) in total. [Figure 16] Median QC ratios using median normalization and ANML are shown. [Figure 17] QC ratios in the tails using median normalization and ANML are shown. [Figure 18] Scale factor concordance in inter-spin time samples using SSAN and ANML is shown. [Figure 19] Median specimen CVs across 18 donors at time-to-spin under various normalization schemes are shown. [Figure 20] 1 shows a plot of agreement between scale factors from Covance (plasma) using SSAN and ANML. [Figure 21] The distribution of all pairwise specimen correlations for Covance samples before and after ANML is shown. [Figure 22] A comparison of distributions obtained from data normalized by several methods is shown. [Figure 23] 10 shows metrics for the smoking logistic regression classifier model on the holdout test set using SSAN and ANML normalized data. [Figure 24] Empirical CDFs for c-Raf measurements in plasma and serum samples stained by collection site are shown. [Figure 25] Concordance plots of scale factors using standard median normalization versus adaptive median normalization in plasma (top) and serum (bottom) are shown. [Figure 26] 1 shows the CDF by site for samples unaffected by site differences for the standard normalization scheme and adaptive normalization. [Figure 27] Plasma sample median normalized scale factors by dilution and Covance collection site are shown. [Figure 28]To increase rigor in adaptive normalization, the distribution of median normalized scale factors is shown. [Figure 29] Typical behavior for a specimen showing significant differences in RFU as a function of time to spin is shown. [Figure 30] The median dilution normalized scale factor for time to spin is shown. [Figure 31] 10 summarizes the effect of adaptive normalization on the median normalized scale factor versus time to spin. [Figure 32] Shown are median standard normalized scale factors by dilution and disease state divided by GFR values. [Figure 33] Median normalization scale factors by dilution and disease state with standard median normalization (top) and adaptive normalization with cutoff are shown. [Figure 34] This is shown together with the CDF of the Pearson correlation between all samples and GFR (log / log) for the different normalization procedures. [Figure 35] 1 shows the distribution of protein-protein Pearson correlations for the CKD dataset for non-normalized data, standard median normalization, and adaptive normalization. DETAILED DESCRIPTION OF THE INVENTION

[0006] While the methods, devices, and computer-readable media are described herein as examples and embodiments, those skilled in the art will recognize that the methods, devices, and computer-readable media for adaptive analyte level normalization are not limited to the described embodiments or drawings. It should be understood that the drawings and description are not intended to be limited to the particular forms disclosed. Rather, the invention encompasses all modifications, equivalents, and alternatives falling within the spirit and scope of the appended claims. Any headings used herein are for organizational purposes only and are not intended to limit the scope of the description or the claims. As used herein, the word "can" is used in a permissive sense (i.e., having the possibility) rather than a mandatory sense (i.e., must mean). Similarly, "include," "including," "includes," "comprise," "comprises," "comprising," and the like mean including, but not being limited to, an element.

[0007] Applicants have developed novel methods, apparatus, and computer-readable media for adaptive normalization of analyte levels detected in a sample. The techniques disclosed and claimed herein adequately remove assay bias and decorrelate assay noise while avoiding the introduction of artifacts in the data due to sample collection artifacts or an excessive number of disease-associated proteomic changes.

[0008] The disclosed adaptive normalization techniques and systems remove affected specimens from the normalization procedure when collection bias exists within the subject population or when an excessive number of specimens are biologically affected in the population being studied, thereby preventing the introduction of bias into the data.

[0009] The directed aspect of adaptive normalization utilizes the definition of comparisons within sample sets where bias may be suspected. These include distinct sites within a multi-site sample collection that have been shown to exhibit large variations in specific protein distributions and important clinical variables within a study. Testable clinical variables are those clinical variables of interest in the analysis, but where other confounding factors may exist.

[0010] The adaptive aspect of adaptive normalization refers to the removal of those analytes from the normalization procedure that are found to be significantly different in the directed comparisons defined at the beginning of the normalization procedure. Because each collection of clinical samples is somewhat unique, the method adapts to learn which analytes are required for removal from normalization, and the set of removed analytes will differ for different studies.

[0011] Furthermore, by removing affected analytes from the median normalization, the present system and method minimize the introduction of normalization artifacts without correcting for affected analytes. Conversely, sample processing artifacts are amplified by such analyses, as are the underlying biology in the study. These effects are described in more detail in the Examples section.

[0012] The disclosed technique for adaptive normalization follows a recursive methodology, checking for significant differences between user-specified samples at a sample-by-sample level. The dataset is hybridization normalized and first calibrated to remove the initially detected assay noise and bias. This dataset is then passed to an adaptive normalization process (described in more detail below) using the following parameters: (1) Directed groups of interest; (2) the test statistic used to determine the difference between the indicated groups; (3) Multiple testing correction method (4) Cutoff of significance level of the test

[0013] The user-directed set of groups can be defined by the sample itself, by collection site, sample quality metrics, etc., or by clinical covariates such as glomerular filtration rate (GFR), cases / controls, events / non-events, etc. Many test statistics can be used to detect artifacts in collection, such as Student's T-test, ANOVA, Kruskal-Wallis, or serial correlation. Multiple testing corrections include Bonferroni, Holm, and Benjamini-Hochberg (BH), to name a few.

[0014] The adaptive normalization process begins with data that has already been hybridization normalized and calibrated. Univariate test statistics are calculated for each analyte level between the indicated groups. The data is then median normalized to a reference (Covance dataset), and those analyte levels with significant variation between the defined groups are removed from the set of measurements used to generate the normalization scale factor. Through this adaptation step, the system removes analyte levels that have the potential to introduce systematic bias between the defined groups. The resulting adaptively normalized data is then used to recalculate the test statistics, followed by a new adaptive measurement set used to normalize the data, and so on.

[0015] This process can be repeated over multiple iterations until one or more conditions are met. These conditions can include convergence, i.e., the analyte levels selected from successive iterations are the same, the degree of change in analyte levels between successive iterations is less than a certain threshold, the degree of change in scale factor between successive iterations is less than a certain threshold, or passing a certain number of iterations. The output of the adaptive normalization process can be a normalization file annotated with a list of excluded analytes / analyte levels, the test statistic, and the corresponding statistical value (i.e., adjusted p-value).

[0016] As further described in the Examples section, for datasets containing an extreme number of artifacts (either biological or acquisition-related), the system can filter out artifacts and noise that were not detected by previous median normalization schemes.

[0017] 1 illustrates a method for adaptive normalization of analyte levels in one or more samples, according to an exemplary embodiment. One or more analyte levels corresponding to one or more analytes detected in the one or more samples are received. Each analyte level corresponds to a detected amount of that analyte in the one or more samples.

[0018] Figure 2 shows an example of a sample 200 having multiple detected analytes, according to an exemplary embodiment. As shown in Figure 2, the larger circle 200 represents the sample, and each of the smaller circles represents an analyte level for a different analyte detected in the sample. For example, circles 201A and 202A correspond to two different analyte levels for two different analytes. Of course, the analyte amounts shown in Figure 2 are for illustrative purposes only, and the analyte levels and number of analytes detected in a particular sample may vary.

[0019] As shown in FIG. 2, sample 200 includes various analytes, such as analyte 201A and analyte 202A. Reference distribution 1 is a reference distribution corresponding to analyte 201A, and reference distribution 2 is a reference distribution corresponding to analyte 202A. The reference distributions can take any suitable format. For example, as shown in FIG. 2, each reference distribution may plot the analyte levels of the analyte detected in a reference population or reference sample. Of course, reference distributions can be plotted and / or stored in a variety of different ways. For example, a reference distribution can be plotted based on the respective counts of analyte levels or ranges of analyte levels. Furthermore, a reference distribution can be processed to extract mean, median, and standard deviation values, and these stored values ​​can be used in the distance determination process, as described below. Many variations are possible, and these examples are not intended to be limiting.

[0020] As shown in FIG. 2, the analyte level of each analyte in the sample (e.g., analytes 201A and 202A) is compared to a corresponding reference distribution (e.g., distributions 1 and 2) either directly or via statistical measures extracted from the reference distributions (e.g., mean, median, and / or standard deviation), and the statistical and / or mathematical distance between each analyte level in the sample and the corresponding reference distribution is determined.

[0021] The one or more samples in which the analyte level is detected can include a biological sample, such as a blood sample, a plasma sample, a serum sample, a cerebrospinal fluid sample, a cell lysate sample, and / or a urine sample. Additionally, the one or more analytes can include, for example, a protein analyte, a peptide analyte, a sugar analyte, and / or a lipid analyte.

[0022] The analyte level of each analyte can be determined in various ways. For example, the analyte level of each analyte can be determined based on applying a binding partner of the analyte to one or more samples. The binding of the binding partner to the analyte generates a measurable signal. The measurable signal can then be measured to obtain the analyte level. In this case, the binding partner can be an antibody or an aptamer. Additionally or alternatively, the analyte level of each analyte can be determined based on mass analysis of one or more samples.

[0023] Returning to FIG. 1, in step 102C, scale factors are iteratively applied to one or more analyte levels over multiple iterations until the change in scale factor between successive iterations is below a predetermined change threshold 102D or until the amount of one or more iterations exceeds a maximum iteration value (102F).

[0024] The scale factor is a dynamic variable that is recalculated for each iteration. By determining and measuring the change in the scale factor between subsequent iterations, the system can detect when further iterations do not improve the results and thereby terminate the process.

[0025] Additionally, a maximum iteration value can be utilized as a failsafe to ensure that the scale factor application process does not repeat indefinitely (in an infinite loop), e.g., 10 iterations, 20 iterations, 30 iterations, 40 iterations, 50 iterations, 100 iterations, or 200 iterations.

[0026] If desired, the maximum iterations value may be omitted and the scale factor may be iteratively applied to one or more analyte levels over multiple iterations, without regard to the number of iterations required, until the change in the scale factor between successive iterations is below a predetermined change threshold.

[0027] The predetermined change threshold can be user-configurable or set to some default value. For example, the predetermined change threshold can be set to a very low decimal value (e.g., 0.001) so that the scale factor is required to reach "convergence" where there is very little measurable change in the scale factor between iterations for the process to terminate.

[0028] The change in scale factor between subsequent iterations may be measured as a percentage change, where the predetermined change threshold may be, for example, a value between 0 and 40 percent (inclusive), a value between 0 and 20 percent (inclusive), a value between 0 and 10 percent (inclusive), a value between 0 and 5 percent (inclusive), a value between 0 and 2 percent (inclusive), a value between 0 and 1 percent (inclusive), and / or 0 percent.

[0029] In step 102A, a distance is determined between each analyte level in one or more analyte levels and the corresponding reference distribution for that analyte in the reference data set. This distance may be a statistical or mathematical distance that represents the extent to which a particular analyte level differs from the corresponding reference distribution for that same analyte. can be a measure of Reference distributions for various analyte levels can be pre-compiled, stored in a database, and accessed as needed during the distance determination process. The reference distributions can be based on reference samples or populations and can be verified to be free of contamination or artifacts by a manual review process or other suitable techniques.

[0030] Determining the distance between each analyte level in one or more analyte levels and the corresponding reference distribution of that analyte in the reference dataset can include determining the absolute value of the Mahalanobis distance between each analyte level and the corresponding reference distribution of that analyte in the reference dataset. The Mahalanobis distance is a measure of the distance between a point P and a distribution D, and the origin for calculating this measure can be the centroid (center of mass) of the distribution. The origin for calculating the Mahalanobis distance ("M-distance") can also be the mean or median of the distribution, as discussed further below, and the standard deviation of the distribution can be utilized. Of course, there are other methods for measuring the statistical or mathematical distance between an analyte level in a sample and the corresponding reference distribution that can be utilized. For example, determining the distance between each analyte level in one or more analyte levels and the corresponding reference distribution of that analyte in the reference dataset can include determining the amount of standard deviation between each analyte level and the mean or median of the corresponding reference distribution of that analyte in the reference dataset.

[0031] Returning to Figure 1, in step 102B, a scale factor is determined based at least in part on analyte levels that are within a predetermined distance from the corresponding reference distribution. This step includes a first substep of identifying all analyte levels in the sample that are within a predetermined distance threshold from the corresponding reference distribution. The predetermined distance used as a cutoff for identifying analyte levels to be used in the scale factor determination process can be set by a user, or can be set to some default value and / or customized to the type of sample and analyte involved.

[0032] Additionally, the predetermined distance threshold will depend on how the statistical distance between the analyte level and the corresponding reference distribution is determined. When using M-distance, the predetermined distance can be a value in the range of 0.5 to 6, a value in the range of 1 to 4, a value in the range of 1.5 to 3.5, a value in the range of 1.5 to 2.5, and / or a value in the range of 2.0 to 2.5. The particular predetermined distance used to filter analyte levels from use in the scale factor determination process can depend on the underlying data set and associated biological parameters. Certain types of samples have greater inherent variability than others, warranting a higher predetermined distance threshold, while others may warrant a lower predetermined distance threshold.

[0033] Returning to FIG. 1 , in step 102A, the distance between each analyte level and the corresponding reference distribution for that analyte is calculated. The corresponding reference distribution can be identified and stored in memory based on an identifier associated with the analyte, or can be identified based on an analyte identification process that detects each type of analyte. The distance can be calculated, for example, as the M-distance, as described above. Because the M-distance is calculated based on the mean, median, and / or standard deviation of the corresponding reference distribution, it is not necessary to store the entire reference distribution in memory. For example, the M-distance between each analyte level in a sample and the corresponding reference distribution is given by:

[0034]

number

[0035] where M is the Mahalanobis distance ("M-distance"), the value of the analyte level in the sample, and x p is the analyte level value of the sample, and μ ref is the mean of the reference distribution corresponding to that sample, and σ ref,p is the standard deviation of the reference distribution corresponding to that sample.

[0036] 3 shows a flowchart for determining a scale factor based at least in part on an analyte level that is within a predetermined distance from a corresponding reference distribution, according to an exemplary embodiment. In step 301, an analyte scale factor is determined for each analyte level that is within a predetermined distance from the corresponding reference distribution. The analyte scale factor is determined, at least in part, based on the analyte level and the mean or median of the corresponding reference distribution. For example, the analyte scale factor for each analyte can be based on the mean of the corresponding reference distribution.

[0037]

number

[0038] Here, SF analyte is the scale factor for each analyte within a given distance from the corresponding reference distribution, and μ ref、p is the mean of the reference distribution corresponding to that sample, and x p is the value of the analyte level in the sample. The analyte scale factor may also be based on the median of the corresponding reference distribution.

[0039]

number

[0040] Here, SF analyte is the scale factor for each analyte within a given distance from its corresponding reference distribution, ~x is the median of the reference distribution corresponding to that analyte, and x is the value of the analyte level in the sample.

[0041] In step 302, an overall scale factor for the sample is determined by calculating either the mean or median of the analyte scale factors corresponding to analyte levels that are within a predetermined distance from the corresponding reference distribution. Thus, the overall scale factor is given by one of the following:

[0042]

number

[0043] Here, SF analyte is the overall scale factor (herein referred to as the "scale factor") to be applied to the analyte level in the sample, and ~x SFanalyte is the mean of the specimen scale factor, and σ SFanalyte is the median specimen scale factor.

[0044] In step 302, a determination is made whether the distance between the analyte level and the reference distribution is greater than a predetermined distance threshold. If so, the analyte level is flagged as an outlier in step 303, and the analyte level is excluded from the scale factor determination process in step 304. Otherwise, if the distance between the analyte level and the reference distribution is less than or equal to the predetermined distance threshold, the analyte level is flagged as being within an acceptable distance in step 305, and the analyte level is used in the scale factor determination process in step 306.

[0045] The flagging of each analyte level may be encoded and tracked by a data structure for each iteration of the scale factor application process, for example, by a bit vector or other Boolean value that stores a 1 or 0 for each analyte level, where the 1 or 0 indicates whether the analyte level should be used in the scale factor determination process. The corresponding data structure may be refreshed / re-encoded during each new iteration of the scale factor application process.

[0046] If a scale factor determination process is performed in step 306, the data structure encoding the results of the distance threshold evaluation process in steps 301-302 can be used to filter the analyte levels in the sample to extract and / or identify only those analyte levels that are used in the scale factor determination process.

[0047] The origin for calculating the predetermined distance for each reference distribution is shown as the centroid of the distribution for clarity, but it should be understood that other origins may be utilized, such as the mean or median of the distribution, or the mean or median adjusted based on the standard deviation of the distribution.

[0048] Returning to Figure 1, in step 102D, a determination is made as to whether the change in scale factor between the determined scale factor and the previously determined scale factor (for the previous iteration) is less than or equal to a predetermined threshold. If the first iteration of the scaling process is being performed, this step can be omitted. This step compares the current scale factor with the previous scale factor from the previous iteration and determines whether the change between the previous scale factor and the current scale factor exceeds a predetermined threshold.

[0049] As discussed above, this predetermined threshold can be some user-defined threshold, such as a 1% change, and / or can require nearly identical scale factors (~0% change) for the scale factors to converge to a particular value.

[0050] If the change in the scale factor between the i and (i-1) iterations is less than or equal to a predetermined threshold, the adaptive normalization process terminates at step 102F. Otherwise, if the change in the scale factor between the i and (i-1) iterations is greater than a predetermined threshold, the process proceeds to step 102C, where one or more analyte levels in the sample are normalized by applying a scale factor. Note that all analyte levels in the sample are normalized using this scale factor, not just the analyte levels used to calculate the scale factor. Thus, the adaptive normalization process does not "correct" for collection site bias or differences in protein levels due to disease, but rather , positive Ensure that such large differential effects are not removed during normalization. , because when you do this removal, Introduces artifacts into the data, disrupting the desired protein signature Because.

[0051] After the normalization step in step 102C, a determination is made in optional step 102E as to whether repeating another iteration of the scaling process would exceed the maximum iteration value (i.e., whether i+1 > maximum iteration value). If so, the process ends in step 102F. Otherwise, the next iteration is initialized (i++) and the procedure returns to step 102A for distance determination in step 102B, scale factor determination, and normalization in step 102C (if the change in scale factor exceeds a predetermined threshold in 102D). Steps 102A-102D are repeated for each iteration until the process ends in step 102F (based on either the change in scale factor falling within a predetermined threshold or exceeding the maximum iteration value).

[0052] 4A-4F show examples of adaptive normalization processes for a set of sample data, according to an example embodiment.

[0053] Figure 4A illustrates a set of reference data summary statistics to be used for both the calculation of scale factors and for determining analyte-level distance to the reference distribution. The reference data summary statistics summarize appropriate statistical measures for the reference distributions corresponding to 25 different analytes.

[0054] Figure 4B shows a set of sample data corresponding to the analyte levels of 25 different analytes measured across 10 samples, each expressed as relative fluorescence units, although it is understood that other units of measurement may be utilized.

[0055] The adaptive normalization process can be iterated through each sample by first calculating the Mahalanobis distance (M-distance) between each analyte level and the corresponding reference distribution, determining whether each M-distance is within a predetermined distance, calculating a scale factor (both analyte-level and globally), normalizing the analyte levels, and then repeating the process until the change in scale factor falls below a predetermined threshold. As an example, Figures 4C-4F utilize the measurements for Sample 3 in Figure 4B. As shown in Figure 4C, the M-distance between each analyte level in Sample 3 and the corresponding reference distribution is calculated. This M-distance is given by the formula (discussed above):

[0056]

number

[0057] Also shown in the table of FIG. 4C is a Boolean variable, Within-Cutoff, which indicates whether the absolute value of the M-distance for each specimen is within a predetermined distance required for use in the scale factor determination process. In this case, the predetermined distance is set to 2. As shown in FIG. 4C, specimens 3, 6, 7, 11, 17, 18, 20, and 23 are greater than the cutoff distance of |2|; therefore, they will not be used in the following scale factor determination step.

[0058] To determine the overall scale factor, the scale factor for each of the remaining specimens (those with a true Within-Cutoff value) is determined as described above. Figure 4D shows the specimen scale factor for each specimen. The median of these specimen scale factors is then set as the overall scale factor. Of course, the average of these specimen scale factors can also be used as the overall scale factor. In this case, the scale factor is given by:

[0059]

number

[0060] Here, SF analyte 1, ...p is the analyte scale factor for each of the analytes used in the scale factor determination process.

[0061] The 25 analyte measurements for sample 3 are then multiplied by this scale factor, and the process is repeated. A new M-distance is calculated for this normalized data, as shown in Figure 4E, and the analytes that fall within the predetermined distance threshold are determined. Figure 4F further illustrates the analyte scale factor for this next iteration. Using the formula above for the overall scale factor, the overall scale factor for this iteration is determined to be equal to 1 (the median analyte scale factor).

[0062] Since the overall scale factor is determined to be 1, applying this scale factor will not cause any change to the data and the next scale factor will also be 1, so the process can end.

[0063] 5A-5E show another example of an adaptive normalization process requiring two or more iterations, according to an example embodiment, using data corresponding to sample 4 in FIGS. 4A-4B.

[0064] Figure 5A shows the M-distance values ​​and corresponding Boolean Within-Cutoff values ​​for each of the analytes in sample 4. As shown in Figure 5A, analytes 1, 4, 6, 8, 12, 17, 19, and 21-25 are excluded from the scale factor determination process.

[0065] Figure 5B shows the specimen scale factors for each of the remaining specimens. The overall scale factor for this iteration is taken as the median of these values, as before, and is equal to 0.9663.

[0066] This scale factor is applied to the analyte levels to produce the analyte levels shown in Figure 5C. Figure 5C also shows the M-distance and cutoff determination results for the second iteration of the normalization process. In this case, analytes 1, 4, 6, 10, 12, 17, 19, and 21-25 are excluded from the scale factor determination process.

[0067] Figure 5D shows the specimen scale factors for each of the remaining specimens. The overall scale factor for this iteration is taken as the median of these values, as described above, and is equal to 0.8903. Because this scale factor has not yet converged to a value of 1 (indicating no further change in the scale factor), the process is repeated until convergence is reached (or until the change in the scale factor is within some other predetermined threshold).

[0068] 5E illustrates the scale factors determined for each sample shown in FIGS. 4A-4B over eight iterations of the scale factor determination and adaptive normalization process. As shown in FIG. 5E, the scale factors for sample 4 do not converge until the fifth iteration of the process.

[0069] The analyte level data for each sample changes after each iteration (assuming the determined scale factor is not 1). For example, Figure 6A shows the analyte levels for all samples after one iteration of the adaptive normalization process described herein. Figures 6A-6B show the analyte levels for all samples after the adaptive normalization process is complete (in this example, after all scale factors have converged to 1).

[0070] 1, the scale factor determination step 102B can be performed in other ways. In particular, determining the scale factor based at least in part on the analyte level within a predetermined distance from the corresponding reference distribution can include determining a value of the scale factor that maximizes the probability that the analyte level within a predetermined distance from the corresponding reference distribution is part of the corresponding reference distribution.

[0071] 7 illustrates the requirements for determining the value of the scale factor that maximizes the probability that an analyte measurement in a given sample is derived from a reference distribution, where the probability that each analyte level is part of the corresponding reference distribution can be determined based at least in part on the scale factor, the analyte level, the standard deviation of the corresponding reference distribution, and the median of the corresponding reference distribution.

[0072] In step 704, a value of the scale factor is determined that maximizes the probability that all analyte levels within a predetermined distance from the corresponding reference distribution are part of the corresponding reference distribution. As shown in Figure 7, this probability function utilizes the standard deviations of the corresponding reference distributions 702 and analyte levels 703 to determine the value of the scale factor 7015 that maximizes this probability.

[0073] Adaptive normalization using this technique for scale factor determination is referred to herein as adaptive normalization by maximum likelihood (ANML). The main difference between ANML and the previous technique for adaptive normalization mentioned above (which operates on a single sample and is referred to herein as single sample adaptive normalization (SSAN)) is the scale factor determination step.

[0074] While the median was used to calculate the scale factor in SSAN, ANML utilizes information from the reference distribution to maximize the probability that the sample is drawn from the reference distribution.

[0075]

number

[0076] This formula relies on the assumption that the reference distribution follows lognormal probability. Such an assumption allows for a simple closed form for the scale factor, but is not necessary. As mentioned above, the overall scale factor for ANML is a weighted variance mean. The contribution to the scale factor of specimen measurements that exhibit large population variances is SF overallare weighted less than those resulting from smaller population variances.

[0077] 8A-8C illustrate the application of maximum likelihood adaptive normalization to the sample data for sample 4 shown in FIGS. 4A-4B, according to an exemplary embodiment. FIG. 4A shows the M-distance and With-Cutoff values ​​for each analyte in the first iteration. As shown in FIG. 8A, the unusable analytes from the first iteration for sample 4 are analytes 1, 4, 6, 8, 12, 17, 19, 21, 22, 23, 24, and 25. For the calculation of the scale factor, the log10 transformed reference data, standard deviation, and sample data are taken and the above formula is applied to determine the scale factor.

[0078]

number

[0079] Applying this exponent to a base of 10 determines the scale factor for this sample / iteration as follows:

[0080]

number

[0081] Similar to the SSAN procedure, this intermediate scale factor is applied to measurements from sample 4, and the process is repeated for successive iterations.

[0082] Figure 8B shows the scale factors determined by applying ANML to the data in Figures 4A-4B over multiple iterations. The difference in normalized sample measurements between the first iteration and after convergence is quite clear for samples requiring more than one iteration. These additional iterations demonstrate an advantage for data generated using aptamer-based proteomics assays, which is further described in the Examples section. As shown in Figure 8B, these scale factors differ from those determined by SSAN (Figure 5E). These differences are due to the weighted population variance for each analyte, which helps balance the scale factor calculation for analytes with large reference population variances.

[0083] Figure 8C shows the normalized analyte levels resulting from application of ANML to the data in Figures 4A-4B across multiple replicates. As shown in Figure 8C, the normalized analyte levels differ from those determined by SSAN (Figure 5B).

[0084] Another type of adaptive normalization that can be performed using the disclosed techniques is population adaptive normalization (PAN), which can be utilized when the one or more samples include multiple samples and the one or more analyte levels corresponding to the one or more analytes include multiple analyte levels corresponding to each analyte.

[0085] When adaptive normalization is performed using PAN, the distance between each analyte level in the one or more analyte levels and the corresponding reference distribution for that analyte in the reference dataset is determined by determining a Student's T-test, a Kolmogorov-Smirnov test, or a Cohen's D statistic between the multiple analyte levels corresponding to each analyte and the corresponding reference distribution for each analyte in the reference dataset.

[0086] For PAN, clinical data are treated as a group to test for specimens that differ significantly from population reference data. PAN can be used when a group of samples is identified because they have a subset of similar attributes, such as being collected from the same testing site under specific collection conditions, or when a group of samples may have a clinical distinction (disease state) that differs from the reference distribution.

[0087] The power of a population normalization scheme is the ability to compare many measurements of the same analyte against a reference distribution. The general procedure for normalization is similar to the adaptive normalization method described above, again starting with an initial comparison of each analyte measurement to a reference distribution.

[0088] As mentioned above, several statistical tests can be used to determine the statistical difference between the specimen measurements from the test data and a reference distribution, including Student's T-test, Kolmogorov-Smirnov test, etc.

[0089] In the following example, we utilize Cohen's D statistic for distance measurement, which is a measure of effect size between two distributions and is very similar to the M distance calculation discussed earlier.

[0090]

number

[0091] where D p is the Cohen's D statistic, and μ p is the reference distribution median for a particular sample, and ~x p is the clinical data (sample) median across all samples, and √(σ ref,p 2 +σ x,p 2 )) is the pooled standard deviation (or median absolute deviation). As shown above, Cohen's D is defined as the difference between the reference distribution median and the clinical data median over the pooled standard deviation (or median absolute deviation).

[0092] 9A-9F illustrate the application of population-adaptive normalization to the data shown in FIGS. 4A-4B, according to an exemplary embodiment. For the reference data shown in FIG. 4A and the clinical data shown in FIG. 4B, 25 Cohen's D statistics are calculated, one for each analyte. FIG. 9A shows the Cohen's D statistic for each analyte across all samples. This calculation uses log 10 This can be done in transformation space.

[0093] In an exemplary embodiment, the predetermined distance threshold used to determine whether an analyte should be included in the scale factor determination process is Cohen's D of |0.5|. Analytes outside this window are excluded from the scale factor calculation. As shown in Figure 9A, this results in analytes 1, 4, 5, 8, 17, 21, and 22 being excluded from the scale factor calculation.

[0094] Figure 9B shows the scale factors calculated for each analyte across samples. The difference between population adaptive normalization (PAN) and the normalization methods described above is that in PAN, each sample includes / excludes the same analyte during scale factor calculation. In PAN, the scale factors for all samples are determined based on the remaining analytes. In this example, the scale factor can be given by the median or average of the analyte scale factors for the remaining analytes. Similar to the adaptive normalization methods described above, the scale factor can be determined as the average or median of the individual analyte scale factors. If the median is used, the scale factor for the data shown in Figure 9B is 0.8876.

[0095] This scale factor is multiplied with the data values ​​shown in Figure 4B to generate normalized data values, as shown in Figure 9C. Figure 9D shows the results of the second iteration of the scale factor determination process, including Cohen's D values ​​for each specimen and within-cutoff values ​​for each specimen.

[0096] For this iteration, analytes 1, 4, 5, 8, 16, 17, 20, and 22 should be excluded from the scale factor determination process. In addition to the analytes excluded in the first iteration, the second iteration also excludes analyte 16 from the scale factor calculation. The above steps are then repeated to remove additional analytes from the scale factor calculation for each sample.

[0097] Convergence of adaptive normalization (change in scale factor below a predefined threshold) occurs when the analytes removed from the i-th iteration are the same as those in the (i-1)-th iteration and the scale factors for all samples have converged. In this example, convergence requires five iterations. Figure 9E shows the scale factors for each of the samples in each of the five iterations. Additionally, Figure 9F shows the normalized analyte-level data after convergence has occurred and all scale factors have been applied.

[0098] The systems and methods described herein implement an adaptive normalization process that performs outlier detection to identify any outlier analyte levels and exclude them from the scale factor determination while including the outliers in the scaling aspect of normalization. The features of calculating and applying scale factors are also described in more detail with respect to the previous figures. Furthermore, removal of outlier analyte levels at one or more analyte levels by performing outlier analysis may be performed as described with respect to Figures 1-3. The outlier analysis method described in those figures and in the corresponding sections herein is a distance-based outlier analysis that filters analyte levels based on a predetermined distance threshold from a corresponding reference distribution.

[0099] However, other forms of outlier analysis can also be used to identify outlier analyte levels. For example, density-based outlier analysis, such as local outlier factor ("LOF"), can be used. LOF is based on the local density of data points within a distribution. The locality of each point is given by its k nearest neighbors, and the distance is used to estimate the density. By comparing the local density of an object to the local density of its neighbors, regions of similar density can be identified, as well as points with a lower density than their neighbors. These are considered outliers.

[0100] Density-based outlier detection is performed by evaluating the distance from a given node to its K nearest neighbors ("K-NN"). The K-NN method calculates a Euclidean distance matrix for all clusters in a cluster system, and then evaluates the local reachability distance from the center of each cluster to its K nearest neighbors. Based on the local reachability distances in the distance matrix, a density is calculated for each cluster, and a local outlier factor ("LOF") is determined for each data point. Data points with large LOF values ​​are considered outlier candidates. In this case, the LOF can be calculated for each analyte level in a sample relative to its reference distribution.

[0101] Normalizing one or more analyte levels over multiple iterations can include performing additional iterations, as described above with respect to FIG. 1, until the change in the scale factor between successive iterations is below a predetermined change threshold or until the amount of one or more iterations exceeds a maximum iteration value.

[0102] 10 illustrates a dedicated computing environment for adaptive analyte-level normalization, according to an exemplary embodiment. The computing environment 1000 includes memory 1001, which is a non-transitory computer-readable medium and can be volatile memory (e.g., registers, cache, RAM), non-volatile memory (e.g., ROM, EEPROM, flash memory, etc.), or some combination of the two.

[0103] As shown in FIG. 10, memory 1001 stores distance determination software 1001A for determining the statistical / mathematical distance between analyte levels and their corresponding reference distributions, outlier detection software 1001B for identifying analyte levels that are outside of a predetermined distance threshold, scale factor determination software 1001C for determining analyte scale factors and overall scale factors, and normalization software 1001D for applying the adaptive normalization techniques described herein to a dataset.

[0104] The memory 1001 further includes a storage device 1001 that can be used to store reference data distributions, statistical measures related to the reference data, variables such as scale factors and Boolean data structures, and intermediate data values ​​or variables resulting from each iteration of the adaptive normalization process. Any software stored in memory 1001 may be stored as computer-readable instructions that, when executed by one or more processors 1002, cause the processors to perform the functions described herein.

[0105] The processor 1002 executes computer-executable instructions and may be a real or virtual processor. In a multiprocessing system, multiple processors or multi-core processors may be used to execute computer-executable instructions, increase processing power, and / or execute certain software in parallel.

[0106] The computing environment further includes a communications interface 503, such as a network interface, used to monitor network communications, communicate with devices, applications, or processes on a computer network or computing system, collect data from devices on the network, and perform actions on network communications within the computer network or data stored in a database of the computer network. The communications interface conveys information such as computer-executable instructions, audio or video information, or other data in a modulated data signal. A modulated data signal is a signal that has one or more of its characteristics set or changed in such a manner as to encode information in the signal. By way of example, and not limitation, communications media include wired or wireless techniques implemented with electrical, optical, RF, infrared, acoustic, or other carrier waves.

[0107] The computing environment 1000 further includes an input / output interface 1004 that allows a user (such as a system administrator) to provide input to the system and display or otherwise transmit information for display to the user. For example, the input / output interface 1004 can be used to configure settings and thresholds, load data sets, and display results.

[0108] An interconnection mechanism, such as a bus, controller, or network (shown in FIG. 10 as solid lines), interconnects the components of computing environment 1000. Input / output interface 1004 may be coupled to input / output devices. The input devices may be touch input devices such as a keyboard, mouse, pen, trackball, touch screen, or game controller, voice input devices, scanning devices, digital cameras, remote controls, or another device that provides input to the computing environment. The output devices may be a display, television, monitor, printer, speakers, or another device that provides output from computing environment 1000. The display may include a graphical user interface (GUI) that presents options to a user, such as a system administrator, for configuring the adaptive normalization process.

[0109] Computing environment 1000 may additionally utilize removable or non-removable storage devices, such as magnetic disks, magnetic tapes or cassettes, CD-ROMs, CD-RWs, DVDs, USB drives, or any other media that can be used to store information and that can be accessed within computing environment 1000. Computing environment 1000 can be a set-top box, a personal computer, a client device, a database or databases, or one or more servers, e.g., a farm of networked servers, a clustered server environment, or a cloud network of computing devices and / or distributed databases.

[0110] As used herein, the terms "nucleic acid ligand," "aptamer," "SOMAmer," and "clone" are used interchangeably to refer to a non-naturally occurring nucleic acid that has a desired effect on a target molecule. Desired effects include, but are not limited to, binding the target, catalytically altering the target, reacting with the target in a manner that modifies or alters the functional activity of the target, covalently binding to the target (as in suicide inhibitors), and facilitating a reaction between the target and another molecule. In one embodiment, the effect is specific binding affinity for a target molecule. Such target molecules are three-dimensional chemical structures other than polynucleotides that bind to the aptamer through a mechanism independent of Watson / Crick base pairing or triple helix formation. Additionally, aptamers are not nucleic acids with a known physiological function that are bound by a target molecule. Aptamers for a given target include nucleic acids identified from a candidate mixture of nucleic acids.

[0111] An aptamer is a ligand mixture of a target by (a) contacting a candidate mixture with a target (nucleic acids with increased affinity for the target relative to other nucleic acids in the candidate mixture can be partitioned from the remainder of the candidate mixture), (b) partitioning the increased affinity nucleic acids from the remainder of the candidate mixture, and (c) amplifying the increased affinity nucleic acids to produce a ligand-enriched mixture of nucleic acids, thereby identifying aptamers for the target molecule. While it is recognized that affinity interactions are a matter of degree, in this context, the "specific binding affinity" of an aptamer for its target means that the aptamer generally binds to its target with much higher affinity than it binds to other non-target components in the mixture or sample. An "aptamer," "SOMAmer," or "nucleic acid ligand" is a set of copies of one type or species of nucleic acid molecule having a specific nucleotide sequence. An aptamer can contain any suitable number of nucleotides. "Aptamer" refers to a set of two or more such molecules. Different aptamers can have the same or different numbers of nucleotides. Aptamers may be DNA or RNA and may be single-stranded, double-stranded, or contain double- or triple-stranded regions. In some embodiments, aptamers are prepared using the SELEX process described herein or known in the art. As used herein, SOMAmers or slow off-rate modified aptamers refer to aptamers with improved off-rate properties. SOMAmers can be generated using the improved SELEX method described in U.S. Patent No. 7,947,447, entitled "Method for Generating Aptamers with Improved Off-Rates," the disclosure of which is incorporated herein by reference in its entirety. Further details regarding aptamer-based proteomics assays are described in U.S. Patents 7,855,054, 7,964,356, 7 and 8,945,830, U.S. Patent Application No. 14 / 569,241, and PCT Application No. PCT / US2013 / 044792, the disclosures of which are incorporated herein by reference in their entireties.

[0112] [Improved accuracy] Figure 11 shows the median coefficient of variation across all aptamer-based proteomic assay measurements for 38 technical replicates. Applicant performed 38 technical replicates from 13 aptamer-based proteomic assay runs (quality control (QC) samples) and calculated the coefficient of variation (CV), defined as the standard deviation of the measurements across the mean / median of the measurements, for each analyte across the aptamer-based proteomic assay menu. Using ANML, Applicant normalized each sample while controlling the maximum number of replicates each sample was allowed to undergo under the normalization process. The median CV across replicates showed decreasing CV as the maximum number of allowable replicates increased, indicating increasing precision as the replicates were allowed to converge.

[0113] [Improved biomarker identification] FIG. 12 shows the Kolmogorov-Smirnov statistics for the sex-specific biomarkers for the samples with respect to the maximum allowed replicates.

[0114] Applicant investigated the discriminatory power of known sex-specific biomarkers in an aptamer-based proteomics assay menu. To quantify the degree of separation between these samples, Applicant calculated the Kolmogorov-Smirnoff (KS) test to quantify the distance between the empirical distribution functions of 569 female and 460 male samples. A KS distance of 1 indicates perfect separation of the distributions (good discrimination), while a value of 0 indicates perfect overlap of the distributions (poor discrimination). As in the previous example, Applicant limited the number of replicates each sample could perform before calculating the KS distance for the group. This data shows that the discrimination power of the biomarkers for male / female sex determination increases as the samples are allowed to converge in the iterative normalization process.

[0115] [Application of ANML to QC samples] 662 replicates (BI, Boulder) were performed with 2066 QC samples. These replicates included four different QC lots. Figure 13 shows the number of QC samples by sample ID for plasma and serum used in the analysis.

[0116] A new version of the normalized population reference was generated (to match ANML and generate estimates for the reference SD). The data was hybridization normalized and calibrated according to standard procedures for V4 normalization. At that point, it was median normalized (representing differences due to changes in the reference median) using ANML (representing differences due to both adaptive and maximum likelihood changes in normalization to the population reference) for both the original and new population references.

[0117] Normalization Scale Factor The first comparison is to examine the agreement of scale factors between different normalization standards / methods. If there are only small differences, good agreement of all other metrics is expected. Figure 1 shows the scale factors for QC samples in plasma and serum, which for QC_1710255 (for which we have by far the largest number of replicates) shows that for the most part there are no large differences (the dashed lines represent a 0.1 difference in scale factor; therefore, the differences are mostly less than 0.05).

[0118] Figure 14 shows the agreement of QC sample scale factors using median normalization and ANML. Solid lines indicate identity, dashed lines indicate 0.1 difference above / below identity.

[0119] CV (Coefficient of Variation) We then calculated the CV resolution for control samples in plasma and serum samples using median normalization and ANML. Figure 15 shows the CV resolution for control samples using median normalization and ANML. The lines show the empirical cumulative distribution functions of the CV for each control sample within a plate, between plates, and overall. There is little, if any, discernible difference between the two normalization strategies, indicating that ANML does not alter the reproducibility of control samples. [QC ratio (relative to reference)] After ANML, references are calculated for each QC lot, and these reference values ​​are used to compare with the median QC value for each run. Empirical cumulative distribution functions for QC samples in plasma and serum. Figure 16 shows the median QC ratios using median normalization and ANML. Each line represents an individual plate. These ratio distributions indicate that where we have a "good" distribution, the distribution did not change significantly when ANML was used. On the other hand, a pair of abnormal distributions (light blue plasma) are somewhat better under ANML. While the tails appear less affected, for both methods, be sure to plot the tail percentages, as well as their differences and ratios. Figure 17 shows the QC ratios in the tails using median normalization and ANML. Each dot represents an individual plate, and the yellow lines indicate the plate failure criteria: the dotted lines in the delta plot are + / - 0.5%, while the dotted lines in the ratio plot are 0.9, 1.1. Applicant observes no change in failure (the only plotted run that exceeded 15% in the tail remains there, and the unplotted outliers remain outliers), and furthermore, the variance in the tail is well below 0.5% for almost all runs.

[0120] [Applying ANML to the dataset] Applicants compared the effectiveness of ANML against SSAN on clinical (Covance) and experimental (Time to Spin) datasets using a consistent Mahalanobis distance cutoff of 2.0 for sample rejection during normalization.

[0121] [Time-to-spin] Time-to-spin experiments were conducted on 18 individuals. Each of the six K2EDTA-plasma collection tubes was left for 0, 0.5, 1.5, 3, 9, and 24 hours before processing. Thousands of samples show signal variation as a function of processing time. The same samples, using uncontrolled protocols or processing protocols that do not match the SomaLogic acquisition protocol, show behavior similar to clinical samples. Scale factors from SSAN were compared with ALMN. Figure 18 shows the agreement of scale factors for time-to-spin samples using SSAN and ANML. Each dot represents an individual sample. There is very good agreement between the two methods.

[0122] This dataset is unique in that it contains multiple measurements of the same individual, even under increasingly poor sample quality. While many analyte signals are affected by time-to-spin, there are thousands of signals that are similarly unaffected. The reproducibility of these measurements across increasing time-to-spin can be quantified across multiple normalization schemes (standard median normalization, single-sample adaptive median normalization, and maximum likelihood adaptive normalization). Applicants calculated the CV for each of the 18 donors across time-to-spin and separated the analytes by their sensitivity to time-to-spin. Figure 19 shows the median analyte CV across the 18 donors at time-to-spin under various normalization schemes. Each dot represents one individual, joined by a dashed line across the varying normalization. The expectation for an analyte that does not exhibit sensitivity to time-to-spin would be high reproducibility for each donor across the six conditions; therefore, the adaptive normalization strategy should reduce the CV.

[0123] ANML showed improved CV relative to both standard median normalization and SSAN, indicating that this normalization procedure increases reproducibility against deleterious sample processing artifacts. Conversely, analytes affected by time-to-spin were amplified across the six time-to-spin conditions (Figure 19). This is consistent with previous observations that adaptive normalization schemes enhance true biological effects. While sample processing artifacts are magnified in this case, in other cases, such as chronic kidney disease, where many analytes are affected, we would expect a similar magnification of effect sizes for affected analytes.

[0124] [Covance] Applicants then tested ANML on the Covance plasma samples used to derive the population reference. A comparison of the scale factors obtained using the single sample adaptation scheme is shown by dilution group in Figure 20. Figure 20 shows the agreement plot between the scale factors from Covance (plasma) using SSAN and ANML. Each dot represents an individual and the solid line represents identity. Very good agreement is again obtained between the two methods.

[0125] The goal of normalization is to remove correlation noise that occurs during aptamer-based proteomics assays. Figure 21 shows the distribution of all pairwise analyte correlations for Covance samples before and after ANML. The red curve shows the correlation structure of the calibration data, which exhibits a clear positive correlation bias with little to no negative correlation between analytes. After normalization, this distribution is recentered to distinct populations of analytes with positive and negative correlations.

[0126] Next, we examined how ANML compared with SSAN for generating and validating insights using Covance smoking status. Figure 22 shows a comparison of distributions obtained from data normalized by several methods. The distributions for tobacco users (dotted line) and non-users (solid line) for these two analytes are virtually identical between ANML and SSAN. The distribution of alkaline phosphatase, shown in Figure 22, is the best predictor of smoking status and shows good discrimination under ANML.

[0127] Applicants trained a logistic regression classifier to predict smoking status using a complexity of 10 samples under SAMN-normalized data and ANML-normalized data with an 80 / 20 / test split. A summary of performance measurements for each normalization is shown in Figure 23. Figure 23 shows measurements of the smoking logistic regression classifier model on the holdout test set using SSAN- and ANML-normalized data. Under ANML, there is no loss in performance for smoking prediction, and potentially a small gain.

[0128] Adaptive normalization by maximum likelihood normalizes single samples using knowledge of the underlying analyte distribution. The adaptive scheme prevents the influence of analytes with large pre-analytical variability from biasing signals from unaffected analytes. The high agreement of scale factors between ANML and single-sample normalization indicates that while small adjustments are made, they can affect reproducibility and model performance. Furthermore, data from control samples show no plate breakage or changes in reproducibility of QC and calibration samples.

[0129] [Application to PAN datasets] Analysis begins with hybridization-normalized, internally calibrated data. In all studies below, unless otherwise noted, the adaptive normalization method uses a Student's T-test with a BH multiple testing correction to detect differences in defined groups. Typically, normalization is repeated with different cutoff values ​​to examine behavior. In all cases, adaptive normalization is compared to a standard median normalization scheme.

[0130] [Covance] Covance collected plasma and serum samples from healthy individuals across five different collection sites (San Diego, Honolulu, Portland, Boise, and Austin / Dallas). Only one sample from the Texas site was assayed and therefore removed from this analysis. 167 Covance samples for each matrix were run on an aptamer-based proteomics assay (V3 assay; 5k menu). Here, the indicated groups are defined by the first four collection sites.

[0131] While the number of analytes removed in Covance plasma samples using adaptive normalization was ~2500, or half the analyte menu, measurements on Covance serum samples did not show a significant amount of site bias, with fewer than 200 analytes removed. Empirical cumulative distribution functions (CDFs) by collection site for analyte measurement c-RAF demonstrate the site bias observed for plasma measurements and the lack of such bias in serum. Figure 24 shows the empirical CDFs for c-Raf measurements in plasma and serum samples colored by collection site. Significant differences in the plasma sample distribution (left) are collapsed in the serum samples (right). Adaptive normalization only removes analytes within the assay that appear problematic by statistical testing; therefore, Covance plasma and serum normalization is sensitive to observed differences.

[0132] A central assumption with median normalization is that clinical outcomes (i.e., collection sites in this case) affect a relatively small number of analytes (e.g., <5%), avoiding introducing bias into the analyte signal. This assumption holds well for Covance serum measurements but is clearly not valid for Covance plasma measurements. A comparison of median normalized scale factors from our standard procedure with those of adaptive normalization reveals that for serum, adaptive normalization faithfully reproduces the scale factors for the standard scheme. However, for plasma, many analyte measurements have site-dependent bias introduced by using the standard normalization procedure. Figure 25 shows agreement plots of scale factors using standard median normalization and adaptive median normalization in plasma (top) and serum (bottom). In plasma, several thousand analytes exhibit significant site bias that is accounted for and corrected for using the adaptive scheme. In serum, fewer than 200 analytes exhibit significant site bias, resulting in little or no change in scale factor between the two normalization schemes. Individual points represent the scale factor for each sample colored by collection site. Black lines indicate identity.

[0133] For example, consider an analyte that does not signal differently across the four sites in plasma. Due to numerous other analytes with higher signaling in the Honolulu, Portland, and San Diego samples, measurements for these analytes after standard median normalization are inflated for the Boise site while simultaneously contracted for the remaining three sites, introducing distinct artifacts into the data. This is observed in Figure 25, where the plasma scale factor for the Boise sample appears below the diagonal, while the remaining ones appear above the diagonal. In Figure 26, to illustrate the bias that misapplication of standard median normalization can induce, CDFs by site for analytes that are not affected by site differences are shown for the standard normalization scheme and adaptive normalization. Adaptive normalization works well to prevent artifacts from being introduced into the data during normalization due to collection site bias. For analytes that exhibit strong site bias, adaptive normalization preserves the differences, while standard median normalization tends to attenuate these differences (see c-RAF in Figure 26). Median RFUs for all sites except Boise are higher for the adaptively normalized set compared to the standard.

[0134] The Covance results demonstrate two important features of the adaptive normalization algorithm. (1) For datasets without collection site or biological bias, adaptive normalization faithfully reproduces standard normalized median results, as shown for serum measurements. In situations where multiple sites or preanalytical variations or other clinical covariates affect many specimen measurements, adaptive normalization correctly normalizes the data by removing altered measurements during scale factor determination. When the scale factor is calculated, the entire sample is scaled.

[0135] In practice, artifacts in median normalization can be detected by looking for bias in the set of scale factors generated during normalization. For the standard normalized medians, there are significant differences in scale factor distributions among the four collection sites, with Portland and San Diego being more similar than Boyce and Honolulu. Figure 27 shows plasma sample median normalized scale factors by dilution and Covance collection site. The bias in scale factors by site is most evident for measurements at 1% and 40% mixtures. A simple ANOVA test on the distribution of scale factors by site yielded a 2.4 x 10 -7 and 4.3 × 10 -6 This shows a statistically significant difference for the 1% and 40% dilution measurements with a p-value of 0.001, while the measurement at 0.005% dilution shows no bias with a p-value of 0.45. ANOVA tests for scale factor bias between groups defined for adaptive normalization provide an important metric for assessing normalization without introducing bias.

[0136] This is illustrated in Figure 28, which shows the distribution of median normalized scale factors for increasing stringency in adaptive normalization, with q-value cutoffs from 0.0 (the standard normalized median) to 0.05, 0.25, and 0.5. At the 0.05 cutoff, 2557 (~50%) specimens were identified as exhibiting variability with collection site. Increasing the cutoff to 0.25 and 0.5 identified 3479 and 4133 specimens, respectively. However, the extent to which increasing the cutoff eliminated site-specific differences in median scale factors was negligible. Measurements at the 1% dilution no longer showed site-specific differences in scale factors, site bias at the 40% dilution was significantly reduced by 4 logs in q-values, and the 0.005% distribution remained unchanged and initially unbiased.

[0137] [Sample Processing / Time to Spin] Samples from 18 individuals, with multiple tubes per individual, were placed at room temperature for 0, 0.5, 1.5, 3, 9, and 24 hours before rotation and were measured using standard aptamer-based proteomic assays.

[0138] The signal of certain analytes is dramatically affected by sample processing artifacts. Specifically, for plasma samples, the duration of time the sample is left in place before spinning can increase the signal by more than 10-fold over rapidly processed samples. Figure 29 shows typical behavior for an analyte that exhibits significant differences in RFU as a function of time to spin.

[0139] Many of the analytes whose signal is seen to increase as time-to-spin increases have been identified as analytes dependent on platelet activation (data not shown). Using measurements for such analytes within the normalization median can introduce dramatic artifacts into the process and negatively alter the overall sample that is not affected by time-to-spin. Conversely, Figure 29 also shows sample analytes that are not sensitive to time-to-spin, whose measurements may be distorted by including them in the normalization procedure that are affected by time-to-spin. It is important to remove measurements that are, for any reason, anomalous from the normalization procedure to ensure the integrity of the remaining measurements.

[0140] The standard normalized median across this time-to-spin data set yields significant and systematic differences in the median normalized scale factor across time-to-spin groups. Figure 30 shows the median normalized scale factor by dilution for time-to-spin. Samples that were left for a longer period before spin yield higher RFU values ​​and lower median scale factors.

[0141] The 0.005% dilution scale factor is much less affected by time to spin than the 1% and 40% dilutions. This is likely due to two distinctly different reasons. First, the number of highly abundant circulating analytes also present in platelets is relatively small, and therefore plasma analytes in the 0.005% dilution are less affected by platelet activation. Furthermore, extreme treatment times can result in cell death and lysis in the sample, releasing very basic nuclear proteins (e.g., histones) and increasing nonspecific binding (NSB), as evidenced by the signal on the negative control.

[0142] Due to the large dilutions, the effect of NSB is not observed at 0.005% dilution. The median normalized scale factors for 1% and 40% dilutions show a very strong bias against spin time. Due to the significant increase in signal with increasing spin time, short time samples have scale factors higher than 1, and the signal is increased by median normalization. And samples with longer time to spin have scale factors lower than 1, and the signal is decreased. This observed bias in the normalized scale factor results in a bias in measurements for these analytes that are not affected by time to spin, similar to that exemplified above for the Covance samples.

[0143] Many analytes are affected by platelet activation in plasma samples. Therefore, these data represent an extreme test of the adaptive normalization method, as both the number of affected analytes and the magnitude of the effect size are very large. We tested whether our adaptive normalization procedure could eliminate this inherent correlation between the median normalized scale factor and time to spin.

[0144] Adaptive normalization was performed on plasma time-to-spin samples using Kruskal-Wallis to test for significance, with BH used to control for multiple comparisons. Bonferroni multiple comparison correction was also used with similar results (not shown). At cutoffs of p = 0.05, 1020, or 23%, samples were identified as showing significant change with time-to-spin. Increasing the cutoff to 0.25 and 0.5 increased the number of significant samples to 1344 and 1598, respectively. The effect of adaptive normalization on median normalization scale factor versus time-to-spin is summarized in Figure 31.

[0145] Analytes within the 0.005% dilution were unbiased with standard median normalization, and their values ​​were unaffected by adaptive normalization. At all cutoff levels, the variability in the scale factor due to spin time for the 1% dilution is eliminated, but at the 40% dilution, some residual bias still exists, albeit dramatically reduced. Evidence suggests that the residual bias may be due to NSB induced by platelet activation and / or cell lysis.

[0146] In summary, using a fairly stringent cutoff of 0.25 for adaptive normalization results in normalization across the sample set that reduces the bias observed in standard normalization schemes, but does not completely mitigate all artifacts, which may be due to NSB, a confounding factor here; adaptive normalization removes this signal on average, thereby resulting in residual bias in the scale factor but potentially eliminating bias in the analyte signal.

[0147] [CKD / GFR (CL-13-069)] A final example of the utility of PBAN involves a dataset from a single site that is presumably collected consistently but has a very large biological effect due to the underlying physiological condition of interest, chronic kidney disease (CKD).

[0148] A CKD study involving 357 plasma samples was performed using an aptamer-based proteomics assay (V3 assay; 1129-plex menu). Samples were collected from healthy individuals with a GFR >90 mls / min / 1.73 m. 2 Samples were collected with glomerular filtration rate (GFR) as a measure of renal function, ranging from 0.01 to 0.01. GFR was measured for each sample either pre- or post-blood collection using iohexol. Applicants did not distinguish between pre- and post-iohexol treatment analyses, but paired samples were excluded from the analysis.

[0149] A decrease in GFR results in an increase in signal across most specimens, making standard median normalization problematic. Because the adaptive variable is now continuous, the data were subdivided by GFR percentage (>90 for normal, 60-90 for mild, 40-60 for severe, and 0-40 for severe) and these groups were included in the adaptive normalization procedure. Using standard normalized medians, we observed significant differences in the median normalized scale factors by disease (GFR) status across all dilutions, indicating a strong inverse correlation between GFR and plasma protein concentration. Figure 32 shows the median standard normalized scale factors by dilution and disease status divided by GFR values. This effect is present in all three dilutions but is weakest in the 0.005% mixture, suggesting that some of the observed bias is due to NSB, as in the example above.

[0150] Using adaptive normalization with the indicated disease-related groups and a p=0.05 cutoff, 738 (out of 1211), or 61%, of specimen measurements were excluded from the normalized median. The number of specimens excluded from normalization increases to 1081 (89%) and 1147 (95%) at p=0.25 and p=0.5, respectively. As in two other studies, adaptive normalization used a conservative cutoff value of p=0.05 to remove the correlation of scale factors with disease severity at the 0.005% and 1% dilutions, but a residual, but significantly reduced, correlation remained within the 40% dilution. At p=0.5, we removed all GFR bias, but at the cost of excluding nearly 95% of all specimens from median normalization. Figure 33 shows median-normalized scale factors by dilution and disease status using standard median normalization (top) and adaptive normalization by cutoff.

[0151] When the assumptions of standard median normalization are invalid, artifacts are introduced into the data using standard median normalization. In this extreme case, where a large proportion of analyte measurements are correlated with GFR, standard median normalization attempts to make all measurements appear to be drawn from the same underlying distribution, thus removing analyte correlation with GFR and reducing the sensitivity of the analysis. As a result of "correcting" for higher signaling analytes in CKD, further distortions are introduced by shifting analyte signals that are not affected by biology. These distortions are observed as analytes having a positive correlation between protein levels and GFR, as opposed to the true biological signal.

[0152] Figure 34 shows this, along with the CDF of the Pearson correlation between all analytes and GFR (log / log) for various normalization procedures. Standard median normalization (HybCalMed) shifts the distribution toward zero, introducing false-positive correlations between analyte signals and GFR. Using adaptive normalization reduces this effect as a function of the chosen cutoff value.

[0153] In addition to preserving the true biological correlation between GFR and analyte levels, adaptive normalization also removes assay-induced protein-protein correlations arising from correlation noise in aptamer-based proteomics assays, as shown in Figure 31. The distribution of protein-protein Pearson correlations for the CKD dataset for non-normalized data, standard median normalization, and adaptive normalization is shown in Figure 35.

[0154] The unnormalized data show protein-protein correlations centered at ~0.2 and ranging from ~-0.3 to +0.75. In the normalized data, these correlations are much more concentrated at 0.0 and in the range of -0.5 to +0.5. Although many spurious correlations are removed by adaptive normalization, meaningful biological correlations are preserved, as we have already demonstrated that adaptive normalization preserves physiological correlations with protein levels and GFR.

[0155] [PBAN method analysis] The use of population-based adaptive normalization relies on metadata associated with the dataset. In practice, when clinical variables, outcomes, or collection protocols affect multiple analyte measurements, it moves normalization from the standard data workup process to the analytical tool. Applicants have considered tests with pre-analytical variation as well as extreme physiological variation, and this procedure performs well using bias in the scale factor as a measure of performance.

[0156] Aptamer-based proteomics assay data normalization, consisting of hybridization normalization, plate scaling, calibration, and standard median normalization, is likely sufficient for samples collected and performed in-house using fully compliant Somalogic sample collection and processing protocols. This normalization protocol does not apply to samples collected remotely, such as the four sites used in the Covance assay, as samples may exhibit significant site differences (presumably from comparable sample populations across sites). Each clinical sample set should be examined for bias in the median normalization scale factor as a quality control step. Metrics to explore such bias should include unambiguous sites, if known, as well as other clinical variables that may violate fundamental assumptions for standard normalization.

[0157] The Covance example demonstrates the power of the adaptive normalization method. For serum samples, little site-dependent bias was observed in the median standard normalization scale factor, and the adaptive normalization procedure essentially reproduced the median standard normalization results. However, for Covance plasma samples, extreme bias was observed in the median standard normalization scale factor. The adaptive normalization procedure results in normalized data without introducing artifacts into specimen measurements unaffected by collection differences. The power of the adaptive normalization procedure lies in its ability to normalize data from well-collected samples with few biomarkers, as well as data from studies with significant collection or biological effects. The method is easily adapted to include all specimens unaffected by the metric of interest while excluding only affected specimens. This makes the adaptive normalization technique highly suitable for application to most clinical studies.

[0158] In addition to preventing the introduction of normalization artifacts into aptamer-based proteomics assay data, the adaptive normalization method also removes spurious correlations due to correlation noise observed in raw aptamer-based proteomics assay data. This is well illustrated in the CKD dataset, where unnormalized correlations are centered around 0.0, yet important biological correlations with protein levels and GFR are well preserved. Finally, adaptive normalization works by removing analytes that are inconsistent across collection sites or strongly correlated with disease state from the normalization calculation, while such differences are preserved and even enhanced after normalization. This procedure does not "correct" for collection site bias or protein levels due to GFR. Rather, it ensures that such large differential effects are not removed during normalization, as they would introduce artifacts into the data and disrupt protein signatures. The reverse is also true: most differences are accentuated after adaptive normalization, while undifferentiated measurements are more consistent.

[0159] [Conclusion] Applicant has developed a robust normalization procedure (population-based adaptive normalization, a.k.a. PBAN) that replicates standard normalization for datasets using consistently collected samples with biological responses containing a small number of analytes (e.g., less than 5% of measurements). For collections with site-dependent bias (pre-analytical variation) or for studies of clinical populations where many analytes are affected, the adaptive normalization procedure prevents the introduction of artifacts due to unintended sample bias and does not attenuate the biological response. The analyses presented here support the use of adaptive normalization, using key clinical variables or collection site, or both, to guide normalization during normalization.

[0160] Each of the three normalization techniques described herein has its own advantages. The appropriate technique depends on the extent of available clinical and reference data. For example, ANML can be used when the distribution of specimen measurements relative to a reference population is known. Otherwise, SSAN can be used as an approximation to normalize individual samples. Furthermore, population-adaptive normalization techniques are useful for normalizing specific cohorts of samples.

[0161] The combination of adaptive and iterative processes ensures that sample measurements are recentered around the reference distribution without the potential influence of analyte measurements outside the reference distribution from bias scale factors.

[0162] Although the principles of the present invention have been described and illustrated with reference to the described embodiments, it will be recognized that the described embodiments can be modified in arrangement and detail without departing from such principles. Elements of the embodiments shown in software can be implemented in hardware and vice versa.

[0163] In view of the many possible embodiments to which the principles of our invention may be applied, we claim as our invention all such embodiments as may come within the scope and spirit of the following claims and equivalents thereto.

Claims

1. 1. A method performed by one or more computing devices for adaptive normalization of analyte levels in one or more samples, comprising: receiving, by at least one of the one or more computing devices, one or more analyte levels corresponding to one or more analytes detected in one or more samples, each analyte level corresponding to a detected amount of the analyte in the one or more samples; normalizing one or more analyte levels across multiple replicates, wherein normalizing is performed by, for each replicate, identifying any outlier analyte levels among the one or more analyte levels, calculating a scale factor for the replicate based at least in part on at least one non-outlier analyte level among the one or more analyte levels and not based on the outlier analyte level, and applying the scale factor for the replicate to all of the one or more analyte levels; Equipped with The method, wherein outlier analyte levels at the one or more analyte levels are identified based at least in part on an outlier analysis between each analyte level and a corresponding reference distribution for that analyte in a reference dataset.

2. The method of claim 1 , wherein the outlier analysis comprises a distance-based outlier analysis.

3. The method of claim 1 , wherein the outlier analysis comprises a density-based outlier analysis.

4. 4. The method of claim 1, wherein normalizing one or more analyte levels across multiple iterations comprises performing additional iterations until a change in scale factor between successive iterations is below a predetermined change threshold or until the amount of one or more iterations exceeds a maximum iterations value.

5. 1. A computer-implemented method for adaptive normalization of analyte levels in one or more samples, the method comprising: receiving one or more analyte levels corresponding to one or more analytes detected in one or more samples, each analyte level corresponding to a detected amount of the analyte in the one or more samples; repeating the iterative application of the scale factor to the one or more analyte levels over a number of iterations until a change in the scale factor between successive iterations is below a predetermined change threshold or until the amount of one or more iterations exceeds a maximum iteration value; Including, Each iteration in the plurality of iterations comprises: determining a distance between each analyte level in the one or more analyte levels and a corresponding reference distribution for the analyte in a reference dataset; determining a scale factor in the current iteration based at least in part on analyte levels that are within a predetermined distance from a corresponding reference distribution and not on analyte levels that are outside the predetermined distance of the corresponding reference distribution; normalizing one or more analyte levels by applying the scale factor in the current iteration to all of the one or more analyte levels; 20. A computer-implemented method comprising:

6. 6. The method of claim 5, wherein determining the distance between each analyte level in the one or more analyte levels and a corresponding reference distribution of the analyte in the reference dataset comprises determining an absolute value of a Mahalanobis distance between each analyte level and a corresponding reference distribution of the analyte in the reference dataset.

7. 6. The method of claim 5, wherein determining the distance between each analyte level in the one or more analyte levels and a corresponding reference distribution for the analyte in the reference dataset comprises determining the amount of standard deviation between each analyte level and a mean or median of a corresponding reference distribution for the analyte in the reference dataset.

8. The method according to any one of claims 5 to 7, wherein the predetermined distance comprises a value in the range of 0.5 to 6.

9. The method of any one of claims 5 to 8, wherein the predetermined distance comprises a value in the range of 1 to 4.

10. The method of any one of claims 5 to 9, wherein the predetermined distance comprises a value in the range of 1.5 to 3.

5.

11. The method of any one of claims 5 to 10, wherein the predetermined distance comprises a value in the range of 1.5 to 2.

5.

12. The method of any one of claims 5 to 11, wherein the predetermined distance comprises a value in the range of 2.0 to 2.

5.

13. The step of determining the scale factor in the current iteration comprises: determining an analyte scale factor for each analyte level within a predetermined distance from a corresponding reference distribution, the analyte scale factor being determined based at least in part on the analyte level and the mean or median of the corresponding reference distribution; determining the scale factor in the current iteration by calculating either the mean or median of analyte scale factors corresponding to analyte levels that are within a predetermined distance from their corresponding reference distributions; The method according to any one of claims 5 to 12, comprising:

14. A method described in any one of claims 5 to 12, wherein the step of determining the scale factor in the current iteration includes determining a value of the scale factor that maximizes the probability that an analyte level within a predetermined distance from a corresponding reference distribution is part of the corresponding reference distribution.

15. 15. The method of claim 14, wherein the probability that each analyte level is part of a corresponding reference distribution is determined based at least in part on the scale factor, the analyte level, the standard deviation of the corresponding reference distribution, and the median of the corresponding reference distribution.

16. 16. The method of claim 4, wherein the change in the scale factor between successive iterations is measured as a percentage change, and the predetermined change threshold comprises a value between 0 and 40 percent.

17. The method of claim 16 , wherein the predetermined change threshold comprises a value between 0% and 20%.

18. The method of any one of claims 16 to 17, wherein the predetermined change threshold comprises a value between 0% and 10%.

19. The method of any one of claims 16 to 18, wherein the predetermined variation threshold comprises a value between 0% and 5%.

20. The method of any one of claims 16 to 19, wherein the predetermined variation threshold comprises a value between 0% and 2%.

21. The method of any one of claims 16 to 20, wherein the predetermined change threshold comprises a value between 0 percent and 1 percent.

22. The method of any one of claims 16 to 21, wherein the predetermined change threshold comprises 0 percent.

23. 23. The method of any one of claims 4 to 22, wherein the maximum iteration value comprises one of 10 iterations, 20 iterations, 30 iterations, 40 iterations, 50 iterations, 100 iterations, or 200 iterations.

24. A method according to any one of claims 1 to 4, wherein the scale factor is determined based on one or more analyte scale factors, each corresponding to a respective analyte level, and each analyte scale factor is determined at least in part based on the corresponding analyte level and the mean or median of a reference distribution corresponding to that analyte.

25. The method of any one of claims 1 to 4, wherein the scale factor is calculated by maximizing the probability that non-outlier analyte levels are part of their corresponding reference distributions.

26. The method of any one of claims 1 to 25, wherein the one or more samples comprise a biological sample.

27. 27. The method of claim 26, wherein the biological sample comprises one or more of a blood sample, a plasma sample, a serum sample, a cerebrospinal fluid sample, a cell lysate sample, or a urine sample.

28. 28. The method of any one of claims 1 to 27, wherein the one or more analyte levels corresponding to the one or more analytes detected in the one or more samples comprise a plurality of analyte levels corresponding to a plurality of analytes detected in the one or more samples.

29. The method of any one of claims 1 to 28, wherein the one or more analytes comprise one or more of a protein analyte, a peptide analyte, a sugar analyte, or a lipid analyte.

30. each analyte level is determined based on applying a binding partner for said analyte to one or more samples; Binding of the binding partner to the analyte produces a measurable signal; The method of any one of claims 1 to 29, wherein the measurable signal yields the analyte level.

31. 31. The method of claim 30, wherein the binding partner is an antibody or an aptamer.

32. 32. The method of any one of claims 1 to 31, wherein each analyte level is determined based on mass analysis of one or more samples.

33. the one or more samples comprise a plurality of samples, and the one or more analyte levels corresponding to the one or more analytes comprise a plurality of analyte levels corresponding to each analyte; determining the distance between each analyte level within the one or more analyte levels and a corresponding reference distribution of the analyte within the reference dataset; 6. The method of claim 5, comprising determining a Student's T-test, a Kolmogorov-Smirnoff test, or a Cohen's D statistic between a plurality of analyte levels corresponding to each analyte and a corresponding reference distribution for each analyte in the reference dataset.

34. A computer program which, when executed by one or more processors, causes said one or more processors to carry out a method according to any one of claims 1 to 33.

35. Apparatus configured to carry out the method of any one of claims 1 to 33.

Citation Information

Patent Citations

  • Method and system for determining the ratio of different cell subsets

    JP2018512071A

  • A normalization method for sample assays

    WO2017083310A1

  • Systems and methods for identifying sequence information from single nucleic acid molecule measurements

    WO2019113024A1