Systems and methods for determining assay matrix background
The method improves matrix background estimation in affinity assays by using cumulative F-statistics to differentiate between assay and biological variability, enhancing detection accuracy and enabling cross-platform correlation analysis.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- SOMALOGIC OPERATING CO INC
- Filing Date
- 2026-01-20
- Publication Date
- 2026-07-30
AI Technical Summary
Existing methods for estimating matrix background in affinity assays suffer from limited accuracy due to non-specific binding interactions caused by endogenous proteins, leading to spurious signals that hinder the detection of analytes.
A computer-implemented method using cumulative F-statistics to determine matrix background by evaluating biological variability against assay variability, expanding sample subsets until the F-statistic exceeds a critical value, allowing for precise estimation of the matrix background.
Enhances the accuracy of matrix background estimation, enabling reliable detection of analytes by distinguishing between assay variability and biological variability, and facilitating cross-platform correlation predictions.
Smart Images

Figure US2026011808_30072026_PF_FP_ABST
Abstract
Description
[0001] Koiitch Romano Dascenzo Gates LLP Attorney Docket No. SML317PCT
[0002] SYSTEMS AND METHODS FOR DETERMINING ASSAY MATRIX BACKGROUND
[0003] FIELD
[0004] This disclosure relates to systems and methods for determining matrix background of an assay.
[0005] INTRODUCTION
[0006] The matrix effect is a well-known problem in affinity assays. The matrix effect occurs when endogenous proteins, which are highly abundant, create non-specific binding interactions that lead to spurious signals on an assay readout. The spurious signals act as a background, called a matrix background, that imposes a limit of detection on analyte levels in the assay. However, known methods for estimating values of the matrix background have various drawbacks, such as limited accuracy. Better solutions are needed for estimating the matrix background.
[0007] SUMMARY
[0008] The present disclosure provides systems, apparatuses, and methods relating to estimating a matrix background.
[0009] In some embodiments, a computer-implemented method for estimating a matrix background of a protein-binding reagent in an assay may comprise: determining a first F-statistic based on a first subset of samples of a plurality of samples, wherein each sample of the plurality of samples has a signal level associated with the protein-binding reagent; in response to determining that the first F-statistic is less than a first critical value, determining a second F-statistic based on a second subset of samples of the plurality of samples; and in response to determining that the second F-statistic is at least a second critical value, estimating the matrix background of the protein-binding reagent based on the signal levels of one or more samples of the second subset.
[0010] In some embodiments, a computer-implemented method for determining a matrix background of a target analyte in an assay may comprise: determining, for each of a plurality of sets of samples, whether a biological variability of the set is greater than an assay variability of the set, wherein the biological variability is associated with levels of a target analyte in the samples of the set, and the assay variability is associated withKoiitch Romano Dascenzo Gates LLP Attorney Docket No. SML317PCT
[0011] nonspecific binding in the assay; and determining the matrix background based on at least a first set and a second set of the plurality of sets, wherein the assay variability of the first set is greater than the biological variability of the first set, and the biological variability of the second set is greater than the assay variability of the second set; wherein the second set of samples comprises at least one sample having a higher level of the target analyte than any sample of the first set.
[0012] In some embodiments, a data processing system for estimating a matrix background of a protein-binding reagent in an assay may comprise: one or more processors; a memory; and a plurality of instructions stored in the memory and executable by the one or more processors to: determine a first F-statistic based on a first subset of samples of a plurality of samples, wherein each sample of the plurality of samples has a signal level associated with the protein-binding reagent; in response to determining that the first F-statistic is less than a first critical value, determining a second F-statistic based on a second subset of samples of the plurality of samples; and in response to determining that the second F-statistic is at least a second critical value, estimating the matrix background of the protein-binding reagent based on the signal levels of one or more samples of the second subset.
[0013] Features, functions, and advantages may be achieved independently in various embodiments of the present disclosure, or may be combined in yet other embodiments, further details of which can be seen with reference to the following description and drawings.
[0014] BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Fig. 1 is a flow chart depicting steps of an illustrative method for estimating matrix background in accordance with aspects of the present teachings.
[0016] Fig. 2 is a schematic diagram depicting cumulatively expanding subsets of samples in order of increasing signal strength in accordance with aspects of the present teachings.
[0017] Fig. 3 is a flow chart depicting steps of another illustrative method for estimating matrix background in accordance with aspects of the present teachings.Koiitch Romano Dascenzo Gates LLP Attorney Docket No. SML317PCT
[0018] Fig. 4 is a flow chart depicting steps of an illustrative method for predicting crossplatform correlation based on estimated matrix background in accordance with aspects of the present teachings.
[0019] Fig. 5 is a plot depicting F-statistic value against signal strength in an example application of methods described herein.
[0020] Fig. 6 is a histogram of signal strength of samples of the example application of Fig. 5.
[0021] Fig. 7 is a table summarizing aspects of the example application of Fig. 5.
[0022] Fig. 8 is an illustrative box plot depicting cross-platform correlation.
[0023] Fig. 9 is a schematic diagram depicting an illustrative data processing system in accordance with aspects of the present teachings.
[0024] DETAILED DESCRIPTION
[0025] Various aspects and examples of systems and methods for matrix background estimation are described below and illustrated in the associated drawings. Unless otherwise specified, a system or method for matrix background estimation in accordance with the present teachings, and / or its various components, may contain at least one of the structures, components, functionalities, and / or variations described, illustrated, and / or incorporated herein. Furthermore, unless specifically excluded, the process steps, structures, components, functionalities, and / or variations described, illustrated, and / or incorporated herein in connection with the present teachings may be included in other similar devices and methods, including being interchangeable between disclosed embodiments. The following description of various examples is merely illustrative in nature and is in no way intended to limit the disclosure, its application, or uses. Additionally, the advantages provided by the examples and embodiments described below are illustrative in nature and not all examples and embodiments provide the same advantages or the same degree of advantages.
[0026] The following definitions apply herein, unless otherwise indicated.
[0027] “Comprising,” “including,” and “having” (and conjugations thereof) are used interchangeably to mean including but not necessarily limited to, and are open-ended terms not intended to exclude additional, unrecited elements or method steps.Koiitch Romano Dascenzo Gates LLP Attorney Docket No. SML317PCT
[0028] Terms such as “first”, “second”, and “third” are used to distinguish or identify various members of a group, or the like, and are not intended to show serial or numerical limitation.
[0029] “AKA” means “also known as,” and may be used to indicate an alternative or corresponding term for a given element or elements.
[0030] “Processing logic” describes any suitable device(s) or hardware configured to process data by performing one or more logical and / or arithmetic operations (e.g., executing coded instructions). For example, processing logic may include one or more processors (e.g., central processing units (CPUs) and / or graphics processing units (GPUs)), microprocessors, clusters of processing cores, FPGAs (field-programmable gate arrays), artificial intelligence (Al) accelerators, digital signal processors (DSPs), and / or any other suitable combination of logic hardware.
[0031] “Providing,” in the context of a method, may include receiving, obtaining, purchasing, manufacturing, generating, processing, preprocessing, and / or the like, such that the object or material provided is in a state and configuration for other steps to be carried out.
[0032] In this disclosure, one or more publications, patents, and / or patent applications may be incorporated by reference. However, such material is only incorporated to the extent that no conflict exists between the incorporated material and the statements and drawings set forth herein. In the event of any such conflict, including any conflict in terminology, the present disclosure is controlling.
[0033] Overview
[0034] The matrix background of an assay is the spurious signal associated with nonspecific binding of proteins. The matrix background is thus a contributing factor to the “assay variability” of an assay readout, that is, the variability associated with the assay itself, rather than with the analyte being measured. Variability associated with the analyte being measured is referred to as “biological variability.” For measurements of a given analyte with a given assay, if the biological variability is smaller than the assay variability, the analyte is below the detection limit of the assay and cannot reliably be detected using that assay. On the other hand, if the biological variability is larger than the assay variance,Koiitch Romano Dascenzo Gates LLP Attorney Docket No. SML317PCT
[0035] then the analyte is said to be signaling above the matrix background, or above the detection limit, and can generally be detected using that assay.
[0036] The present disclosure describes systems and methods for determining whether biological variability is greater than assay variability, estimating matrix background, and using data about the samples having analytes that signal above the matrix background to estimate other suitable information. Examples of suitable information may include, without limitation, specificity of a given reagent and cross-platform correlation.
[0037] In general, a method of estimating matrix background (AKA limit of detection) for a set of samples in accordance with aspects of the present teachings includes evaluating whether biological variability is greater than assay variability based on a subset of the set of samples. In response to the result of the evaluation failing to be that biological variability is greater than assay variability, the method includes repeating the evaluation based on an expanded subset of the samples. The expanded subset of samples includes the samples used in the previous evaluation as well as additional samples not previously used. Accordingly, the expanded subset can be described as cumulative. The evaluation is repeated, cumulatively expanding the subset of samples at each evaluation, until either it is found that biological variability is greater than assay variability, or some other end condition is reached. Suitable end conditions may include, without limitation, reaching a maximum number of evaluations, using all samples of the set of samples, using a maximum number of samples of the set of samples, or determining that the repetitions should be stopped for any other suitable reason(s).
[0038] Evaluating whether biological variability is greater than assay variability may comprise any suitable calculation, estimate, statistical test, or other suitable process(es). In some examples, evaluating whether biological variability is greater than assay variability includes performing an analysis of variance (ANOVA) test, an F-test, a Levene’s test, a Bartlett’s test, a Brown-Forsythe test, a Welch’s test, and / or any other suitable test or assessment.
[0039] Aspects of the present teachings may be embodied as a computer-implemented method, computer system, or computer program product. Accordingly, such aspects may take the form of an entirely hardware embodiment, an entirely software embodiment (including firmware, resident software, micro-code, and the like), or an embodimentKoiitch Romano Dascenzo Gates LLP Attorney Docket No. SML317PCT
[0040] combining software and hardware aspects, all of which may generally be referred to herein as a “circuit,” “module,” or “system.” Furthermore, aspects of the present teachings may take the form of a computer program product embodied in a computer-readable medium (or media) having computer-readable program code / instructions embodied thereon. In the context of this disclosure, a computer-readable storage medium may include any suitable non-transitory, tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device.
[0041] The following sections describe selected aspects of illustrative examples of systems and methods in accordance with aspects of the present teachings. The examples in these sections are intended for illustration and should not be interpreted as limiting the scope of the present disclosure. Each section may include one or more distinct embodiments or examples, and / or contextual or related information, function, and / or structure.
[0042] Illustrative Cumulative F-Statistic
[0043] Generally speaking, an F-statistic is a statistical test that compares the variances of two quantities. For example, an F-statistic comparing the variance of a population of individuals, <7p0p, to the variance of the assay itself, assay-. 's:
[0044] <
[0045] (Eq. 1)
[0046]
[0047] The variances may be calculated according to any suitable method. For example, when the population variance upopis being estimated based on a sample of the population, the population variance is typically calculated as upop= F=i( i - x)2 / (n -1), where n is the size of the sample, xtis the value of the ith data point in the sample, and x is the mean of the values of the sample. As another example, when the assay variance is being estimated based on data associated with quality control replicates, the assay variance may be calculated as a^ssay= - y)2 / (Nqc-
[0048]
[0049] 1), where Nqcis the number of quality control replicates, y7is the value of the jth quality control replicate data point, and y is the mean value of the data associated with the quality control replicates.
[0050] As shown in Eq. 1 , the higher the value of the F-statistic, the larger the population variance is compared to the assay variance. The population variance is the varianceKoiitch Romano Dascenzo Gates LLP Attorney Docket No. SML317PCT
[0051] associated with differences among individuals in the population being studied; that is, it is a characterization of biological variability. An F-statistic above a critical value indicates that the null hypothesis - that assay variance and population variance are equal - should be rejected. An F-statistic below the critical value indicates that the null hypothesis cannot be rejected. If the null hypothesis cannot be rejected, then the possibility remains that the null hypothesis is true, i.e., that the assay variance and population variance are equal. This would mean that the biological variance is no greater than the matrix background, and thus that signals of interest in the biological population cannot be reliably uncovered.
[0052] A critical value of the F-statistic may be determined using any suitable method, including methods known in statistics and / or based on a user-defined effect size. For example, the critical value may be calculated based on an F-distribution. As another example, the critical value may be found on an F-table. An F-table contains critical values of the F-distribution. F-tables are generally available in publicly accessible databases, on public websites, and in computational software packages and scientific calculators. Typically, the rows in an F-table correspond to the number of degrees of freedom for the denominator of the F-statistic and the columns correspond to the number of degrees of freedom for the numerator of the F-statistic. For example, the entry of the table at row 5, column 8 corresponds to a critical F-statistic value for 5 degrees of freedom in the denominator and 8 degrees of freedom in the numerator.
[0053] More specifically, any given F-table provides critical F-statistic values with a particular significance level a (alpha), which represents the confidence with which the null hypothesis can be rejected using the critical values found in that table. Conventional values of a include 0.10, 0.05, 0.025 and 0.01. Accordingly, the critical value of an F-statistic may be found by selecting the F table corresponding to the desired significance level and identifying the value in the table corresponding to the appropriate numbers of degrees of freedom in the F-statistic denominator and numerator.
[0054] For an F-statistic that compares two variances, such as the F-statistic of Equation 1, the number of degrees of freedom for the numerator is generally k - 1, where k represents the number of data points used to calculate the variance in the numerator. Similarly, the number of degrees of freedom for the denominator is generally m- 1 , where m represents the number of data points used to calculate the variance in the denominator.Koiitch Romano Dascenzo Gates LLP Attorney Docket No. SML317PCT
[0055] For example, in Equation 1, if there are K samples from the population, the number of degrees of freedom for the numerator is K - 1. If M data points (e.g., quality control replicates) were used to calculate the assay variance, the number of degrees of freedom for the denom inator is M - 1.
[0056] Methods in accordance with aspects of the present teachings calculate an F-statistic via cumulative batching. Cumulative batching as described herein refers to calculating an initial value of the F-statistic using only a subset of population samples, and repeatedly updating the calculation of the F-statistic by expanding the subset of samples used for the calculation until the updated value of the F-statistic is at least the critical value. Based on the samples that were used to calculate the F-statistic value that was at or above the critical value, the matrix background can be estimated.
[0057] The critical value of the F-statistic of the present teachings may be determined by reference to an F-table, by calculation based on an F-distribution, and / or by any other suitable method. The number of degrees of freedom for the numerator is typically the number of samples in the current subset minus one (i.e. , N - 1 , where N is the number of samples in the current subset). Accordingly, the critical value of the F-statistic generally depends on the size of the subset and therefore generally changes as the subset is expanded. The number of degrees of freedom for the denominator is determined based on the way the assay variance is estimated, and thus may or may not depend on the number of samples in the current subset. In some examples, the number of degrees of freedom for the denominator is the number of quality control replicates minus 1 (i.e., Nqc- 1 , where Nqcis the number of quality control replicates), and thus does not change as the subset is expanded.
[0058] In some examples, an initial F-statistic is calculated using an initial subset of samples, where the initial subset contains the samples of the population having the weakest signals for a particular analyte. If the initial F-statistic is below the critical value corresponding to that subset, the initial subset is expanded by adding an additional batch of samples that have somewhat stronger signals for that analyte. The F-statistic is calculated for the expanded subset and compared to the critical value corresponding to the expanded subset. The process is repeated, with the subset of samples being expanded at each repetition to include increasingly stronger-signaling samples, until theKoiitch Romano Dascenzo Gates LLP Attorney Docket No. SML317PCT
[0059] F-statistic is at least the critical value. The matrix background can be estimated based on the signal strength of the samples that were included in the calculation for which the F-statistic was at least the critical value. For example, it may be inferred that the matrix background corresponds to the analyte signal strength of the last batch of samples added to the subset, or that the matrix background lies somewhere between the analyte signal strengths of the last batch added and the second-to-last batch added. The process may be repeated for all analytes, or all assay reagents, for which an estimate of the matrix background is desired.
[0060] Optionally, methods according to aspects of the present teachings may include applying one or more statistical correction methods. For example, a multiple-comparisons correction method such as the Bonferroni method, Holm method (AKA Holm-Bonferroni method), Hochberg’s method, false discovery rate (FDR) method, and / or any other suitable method(s), may be used.
[0061] In some cases, the estimated matrix background for one or more reagents may be used to predict cross-platform correlation. For example, the correlation between measurements of the same analytes on two different platforms may be strongest for analytes that signal above the matrix background in large percentages of samples.
[0062] Illustrative Method 100 for Estimating Matrix Background
[0063] With reference to Figs. 1-2, this section describes steps of an illustrative method 100 for estimating a matrix background based on a cumulative F-statistic in accordance with aspects of the present teachings. Method 100 is an example of the methods for determining matrix background described more generally above. Where appropriate, reference may be made to components and systems that may be used in carrying out each step. These references are for illustration, and are not intended to limit the possible ways of carrying out any particular step of the method. Reference numerals are used herein to refer to particular steps of method 100; however, as described below, the reference numerals do not necessarily reflect the order in which the steps are performed.
[0064] In this example, the dataset being studied comprises a plurality of samples from a population. Any suitable number of samples may be used, including tens, hundreds, thousands, tens of thousands, or more. In a typical scenario, each sample correspondsKoiitch Romano Dascenzo Gates LLP Attorney Docket No. SML317PCT
[0065] to a different individual in the population, though this may not be the case in every example.
[0066] Method 100 is an example method for estimating the matrix background (i.e., the limit of detection) for a particular reagent of an affinity assay. The matrix background of the reagent can be thought of as the matrix background of the analyte the reagent is configured to bind to. Typically, the strength of the assay readout signal for a given analyte varies from sample to sample. For purposes of this description, it is imagined that the samples are ordered from the sample having the weakest signal for the analyte in question to the sample having the strongest signal for that analyte, as depicted schematically in Fig. 2. This description uses the index / to refer to the position of a given sample in this order. For example, / = 1 refers to the sample having the weakest signal, i = 2 refers to the sample having the second-weakest signal, and so on. Method 100 does not necessarily include creating an ordered list of samples from weakest-signaling to strongest-signaling, although in some examples such a list is created.
[0067] In this example, the cumulative F-statistic for N samples is expressed by the following equation:
[0068] >
[0069]
[0070]
[0071] samples i=i through N is the variance in measured analyte level calculated using samples / = 1 through / = N, and o-^ssayis a variance associated with the assay itself. In general, the assay variance Uassay may be determined or estimated by any suitable method(s). In this example, the assay variance is determined based on the variance in measured analyte levels among quality control replicates, and the number of quality control replicates in the study is Nqc.
[0072] Accordingly, F(N) is the biological variance (or other suitable variance of interest) of samples / - 1 though / - N, divided by the variance associated with the assay. For example, F(10) denotes an F-statistic calculated using samples 1 through 10, i.e., the 10 weakest-signaling samples. F(12) denotes an F-statistic calculated using the 12 weakest-signaling samples, i.e., the same 10 samples used to calculate F(10) as well as the two next-weakest-signal ing samples. The F-statistic may be described as cumulativeKoiitch Romano Dascenzo Gates LLP Attorney Docket No. SML317PCT
[0073] because increasing N adds additional samples to the set of samples used to calculate the F-statistic. However, some examples of method 100 include decreasing N as well as (or instead of) increasing N, as described below.
[0074] Furthermore, in this example, the index j is used to refer to calculations in a set of repeated calculations. Accordingly, Nj refers to the value of N used in the jth calculation, and F(Nj) refers to the F-statistic calculated using samples 1 through Nj.
[0075] At step 102, method 100 includes calculating the F-statistic F(Nj). Put another way, step 102 includes calculating the F-statistic using the A / , weakest-signaling samples. The first time step 102 is performed, j = 1. The initial value Ni used to calculate F(Ni) at the first calculation may be selected in any suitable manner on any suitable basis.
[0076] Calculating F(Nj) may include calculating the variance of the measured analyte levels of the Nj weakest-signaling samples, calculating the variance of the measured analyte levels of the quality control replicates, and dividing the former by the latter. Because the variance of the quality control replicate levels does not change at each calculation, however, in some examples it is not calculated anew at each step. Accordingly, in such examples, calculating F(Nj) may include calculating the variance of the measured analyte levels of the Nj weakest-signaling samples and dividing that variance by the fixed value of the variance of the quality control replicate levels.
[0077] At step 104, method 100 includes determining whether the calculated value of F(Nj) is greater than, or equal to, the critical value of F(Nj). The critical value of F(Nj) may be determined in any suitable manner, e.g., including reference to an appropriately chosen F- table, or by computing an appropriate F-distribution. The number of degrees of freedom for the numerator of F(Nj) is Nj - 1, and the number of degrees of freedom for the denominator of F(Nj) is Nqc- 1 (i.e. , the number of quality control replicates minus 1 ).
[0078] In some examples, at step 102, the variance of the measured analyte levels of the A / 7weakest-signaling samples is calculated, but is not divided by the variance of the quality control replicate levels. Instead, step 104 includes determining whether the calculated variance of the measured analyte levels of the Nj weakest-signaling samples is greater than or equal to the product of the critical value of F(Nj) and the variance of the quality control replicate levels. For example, a specialized F-table may be created for a given study by multiplying the standard F-table of the desired significance level by the assayKoiitch Romano Dascenzo Gates LLP Attorney Docket No. SML317PCT
[0079] variance, and the values of this specialized F-table may be compared to the biological variance of the Nj samples calculated at step 102 to determine whether the F-statistic F(Nj) is critical. This may avoid the computational time and expense of always dividing the biological variance by the fixed assay variance each time step 102 is performed. However, in some cases it may be inconvenient to use a specialized F table, because the standard F tables are integrated into many software packages and may be familiar to users.
[0080] At step 106, in response to determining that the calculated value of F(Nj) is not greater than or equal to the critical value, method 100 optionally includes determining whether one or more predetermined end conditions are satisfied. In general, an end condition reflects a situation in which it may be desirable to terminate the method even though the F-statistic has not reached the critical value. For example, an end condition may be employed to prevent a computer implementation of the method from repeating the process infinitely. In some examples, an end condition is that a predetermined maximum number of calculations has been reached. Additionally, or alternatively, an end condition may be that all samples have been used (i.e. , that the current value of N is the total number of samples available), or that a suitable number or fraction of available samples has been used. In response to determining that an end condition has been satisfied, the method terminates at step 108. Step 106 is optional and may be omitted.
[0081] If an end condition has not been reached, or if step 106 is not performed, the method proceeds to step 110. At step 110, N is increased.
[0082] Next, the method returns to step 102 for another calculation, where the F-statistic is re-calculated using the increased value of N. The increased value of N corresponds to an expanded subset of samples relative to the previous value of N. Accordingly, increasing N adds a batch of samples to the cumulative subset of samples.
[0083] At step 104, it is determined whether the re-calculated F-statistic is at least the critical value corresponding to the expanded subset of samples. If not, the method returns to step 106, and so on. This process continues until either it is determined at step 106 that an end condition has been reached, in which case the method terminates at step 108, or it is determined at step 104 that the current F-statistic is greater than or equal to the corresponding critical value, in which case the method proceeds to step 112.Koiitch Romano Dascenzo Gates LLP Attorney Docket No. SML317PCT
[0084] At step 112, method 100 includes estimating the matrix background. In some examples, the matrix background is estimated at step 112 based on the samples that were used in the current F-statistic calculation, but not any previous F-statistic calculation. Put another way, the matrix background may be estimated based on the batch of samples which, when added to the cumulative F-statistic calculation, led to the F-statistic being at or above the critical value. Because increasing N adds increasingly strongly-signaling samples to the F-statistic calculation, it can be inferred that at least some of the samples of the most recently added batch have sufficiently strong signal to be detected despite the matrix background, whereas without the most recently added batch, the cumulative subset of samples lacked a sufficiently strong signal to be detected. Accordingly, the matrix background may be estimated based on the most recently added batch of samples. For example, it may be estimated that the matrix background is equal to the strongest signal in the most recently added batch of samples, or to an average signal value of the most recently added batch of samples, or to any other suitable estimate of signal value based on the most recently added batch of samples.
[0085] In some examples, the matrix background is estimated at step 112 based on the most recently added batch of samples as well as one or more previous batches. For example, based on the signal strength of the current cumulative subset of samples and the signal strengths of one or more previous cumulative subsets, and the value of the F-statistic corresponding to each subset, the signal strength corresponding to the critical F-value may be estimated using, e.g., regression and / or any other suitable estimation method(s). The estimated signal strength corresponding to the critical F-value may be inferred to be equal to the matrix background.
[0086] Any suitable number may be selected for the initial value of N to be used at step 102. If the initial value of N is so large that the initial F-statistic is above the critical value, then the only information that would be obtained about the limit of detection would be that it fell somewhere within the range of signal strengths spanned by the first N samples. A more precise estimate than this is generally desired. On the other hand, if the initial value of N is too small, then many repetitions may be required before the F-statistic reaches the critical value, which needlessly uses time and computational resources. Accordingly, the initial value of N may be selected to fall somewhere between these extremes, such thatKoiitch Romano Dascenzo Gates LLP Attorney Docket No. SML317PCT
[0087] the initial subset is unlikely to lead to a critical value of the F-statistic itself, but only a reasonable number of repetitions is likely to be needed in order to reach the critical value.
[0088] The value of N may be increased by any suitable amount at step 110. Increasing N by a small amount may lead to a more precise estimate of the matrix background, but risks requiring a large number of repetitions to arrive at the matrix background estimate. Increasing N by a large amount may reduce the number of repetitions needed, but may lead to a less precise estimate of the matrix background. In some examples, the amount by which N is adjusted at step 110 need not be the same everytime step 110 is performed. For example, N may be increased by a smaller amount as the F-statistic approaches the critical value.
[0089] Method 100 is a method for estimating the matrix background for a particular analyte. Method 100 may be performed one or more additional times (in series, parallel, or any other suitable arrangement) to estimate the matrix background(s) for one or more additional analytes.
[0090] The estimate of the matrix background obtained using method 100 (and / or any other suitable method, including method 150 described below), may be used for various purposes. In some examples in which an assay has been performed for a plurality of analytes, and the assay results show that one or more of the analytes has signaled below the matrix background, the assay results are updated as follows: for each analyte that signaled below the matrix background, the signal value for that analyte is replaced by the estimate of the matrix background. This method of processing the sub-background signal values may have several advantages, such as maintaining consistency across the assay results, simplifying certain downstream calculations by ensuring a floor (i.e., lowest value) of the assay results, avoiding skewing statistical analyses of the assay results with low values that lack any real meaning, and / or readily indicating to people studying the assay results that these analytes signaled at or below the limit of detection. The updated assay results may be used to determine biological information such as signaling pathways, disease progression, metabolic processes, immune function, and / or any other suitable information.
[0091] In some examples in which an assay has been performed for a plurality of analytes, a quality control method may include checking the signals measured for one or moreKoiitch Romano Dascenzo Gates LLP Attorney Docket No. SML317PCT
[0092] analytes that are expected to signal above matrix background. In response to determining that one or more of those analytes did not signal above the matrix background (e.g., the estimated matrix background obtained by method 100 or 150), the method may include taking appropriate step(s) to investigate and / or correct the discrepancy.
[0093] In some examples in which an assay has been performed for a plurality of analytes, a method of processing the assay results includes checking the signals measured for each of one or more of the analytes across some or all of the samples studied in the assay. In response to determining that the signal measured for a given analyte is below the estimated matrix background for all samples, or for a significant fraction of samples, the method includes excluding that analyte from subsequent steps of the analysis. Excluding analyte(s) that consistently signal below the estimated matrix background prevents that analyte(s) from distorting the analysis and may simplify the analysis, e.g., by removing one or more factors from a multiple comparisons correction. In some cases, simplifying the analysis by excluding the analyte(s) also improves the accuracy and / or statistical power of the analysis.
[0094] Illustrative Method 150 for Estimating Matrix Background
[0095] With reference to Fig. 3, this section describes steps of an illustrative method 150 for estimating a matrix background. Method 150 is an example of the methods for matrix background estimation described more generally above. Where appropriate, reference may be made to components and systems that may be used in carrying out each step. These references are for illustration, and are not intended to limit the possible ways of carrying out any particular step of the method.
[0096] Method 150 is an example method for estimating a matrix background for a particular analyte. Method 150 involves a plurality of samples from a population. A measurement of the analyte in question has been made for each sample. Method 150 can be used to estimate the matrix background of that analyte, which corresponds to the limit of detection for that analyte.
[0097] At step 152, method 150 includes determining whether biological variability is greater than assay variability for the analyte measurement in at least two respective subsets of the samples. Determining whether the biological variability is greater than the assay variability for each subset may comprise any suitable process(es), such asKoiitch Romano Dascenzo Gates LLP Attorney Docket No. SML317PCT
[0098] performing an ANOVA test using the respective subset. In some examples, determining whether the biological variability is greater than the assay variability comprises calculating an F-statistic using the respective subsets and determining whether the F-statistic is above a critical value for either subset.
[0099] The subsets of samples are nonidentical, but are not necessarily disjoint. In some examples, the subsets are cumulative, with a second subset containing all the samples of a first subset plus at least one additional sample, a third subset containing all the samples of the second subset plus at least one additional sample, and so on. In some examples, the cumulative subsets are ordered based on the measured analyte level of the samples within the subset. For example, the first subset may contain the samples having the lowest measured analyte level, the second subset may contain the first subset as well as one or more samples having higher measured analyte levels, the third subset may contain the second subset as well as one or more samples having yet higher measured analyte levels, and so on. Fig. 2 depicts an example of this arrangement in which the samples are individually ordered from lowest analyte level to highest analyte level. In other examples, the subsets of samples are generally grouped according to increasing analyte level, but are not necessarily ordered individually from lowest to highest level. For example, the first subset may correspond to a first range of analyte levels, the second subset may correspond to a second range of analyte levels that encompasses the first range and also a higher range, the third subset may correspond to a third range of analyte levels that encompasses the second range and also a yet higher range, and so on.
[0100] At step 154, method 150 includes estimating the matrix background for the analyte based on at least one subset for which the biological variability is greater than the assay variability (referred to herein as a high-significance subset) and at least one subset for which the biological variability is not greater than the assay variability (referred to herein as a low-significance subset).
[0101] In examples using cumulative subsets, the high-significance subset includes the low-significance subset as well as one or more samples not included in the low-significance subset. The matrix background may be estimated based on the sample(s) that are in the high-significance subset but not in the low-significance subset. ForKoiitch Romano Dascenzo Gates LLP Attorney Docket No. SML317PCT
[0102] example, if the high-significance subset is known to contain one or more samples having a higher analyte level than any samples of the low-significance subset, then the matrix background may be estimated based on the analyte levels of the group of sample(s) that are in the high-significance subset but not in the low-significance subset. For example, the matrix background may be inferred to lie within the range of analyte levels spanned by that group of samples, may be estimated to be the highest analyte level of the group, may be estimated to be an average of the analyte levels of the group, or may be determined based on the analyte levels of the group in any suitable way. The precision with which the matrix background can be estimated may be based at least in part on the range of analyte levels encompassed by this group of samples.
[0103] In some examples using cumulative subsets of samples ordered according to analyte level, step 152 of method 150 includes determining whether biological variability is greater than assay variability for successively expanding subsets. For example, step 152 may include performing the determination using a subset corresponding to the 100 lowest-level samples, an expanded subset corresponding to the 120 lowest-level samples, a further expanded subset corresponding to the 140 lowest-level samples, and so on. The process may continue with increasingly expanded subsets until a subset is identified for which the biological variability is greater than assay variability.
[0104] However, in other examples using cumulative subsets, the subset may be expanded or contracted as the determination is performed, so as to identify a suitable high-significance subset and a suitable low-significance subset. For example, step 152 may include performing the determination using a first subset corresponding to the 100 lowest-level samples and an expanded subset corresponding to the 200 lowest-level samples. If the biological variability is greater than the assay variability for the 200 lowest-level samples, but not for the 100 lowest-level samples, the determination is performed again using a subset in between the first subset and the expanded subset, such as a subset corresponding to the 150 lowest-level samples. This 150-sample subset may be described as “contracted” relative to the 200-sample subset and “expanded” relative to the 100-sample subset. If the biological variability is greater than the assay variability for the 150 samples, then the determination is performed again using a subset in between the first subset and the 150-sample subset, such as 120 samples. On the other hand, ifKoiitch Romano Dascenzo Gates LLP Attorney Docket No. SML317PCT
[0105] the biological variability is lower than the assay variability for the 150 samples, then the determination is performed again using a subset in between the 150 samples and the 200 samples, such as 180 samples. Thus, the determination is performed repeatedly, with the subsets being adaptively expanded or contracted at each determination based on whether the biological variability is greater than the assay variability. This process continues until a high-significance subset and a low-significance subset are found that differ by a number of samples small enough (or a range of analyte levels small enough) to enable the matrix background to be estimated with a desired precision.
[0106] Illustrative Method 200 for Predicting Cross-Platform Correlation
[0107] With reference to Fig. 4, this section describes an illustrative method 200 for predicting correlation between assay platforms, in accordance with aspects of the present teachings.
[0108] The inventor(s) have found that the percentage of samples in a population for which the measured signal corresponding to a given analyte is above the matrix background is, in at least some cases, a good predictor of correlation between measurements of that analyte on different assay platforms. For example, when assays are performed on a first platform and all or nearly all samples measure above matrix background for a particular analyte, it is likely that the correlation between the first platform and a second platform is high. On the other hand, when few or no samples measure above the matrix background for an analyte on the first platform, it is likely that the correlation between the first and second platforms is low. Accordingly, whether a given analyte signals above the matrix background in most or all samples on a first platform predicts how strongly correlated the measurement of that analyte on the first platform is to measurements of that analyte on other platforms.
[0109] At step 202, method 200 optionally includes determining the matrix background for an analyte on a first assay platform. In some examples, determining the matrix background includes aspects of methods described herein, such as method 100 and / or method 150. Step 202 may be omitted in situations in which a value of the matrix background has already been obtained. For example, the matrix background may have been estimated at a previous time or provided by another party, and so there is no need to determine the matrix background at step 202.Koiitch Romano Dascenzo Gates LLP Attorney Docket No. SML317PCT
[0110] At step 204, method 200 optionally includes measuring the analyte level in a plurality of samples using the first assay platform. The measured analyte levels are used as described below with reference to steps 206 and 208. Step 204 may be omitted in situations in which the analyte levels are obtained in another way. For example, analyte levels of samples that were obtained in the process of determining the matrix background at step 202 may be used at steps 206 and 208, in which case there is no need to measure the analyte levels of additional samples at step 204. As another example, measurements or estimates of analyte levels in a plurality of samples may be provided by another party rather than measured at step 204.
[0111] At step 206, method 200 includes determining information about the samples of the plurality of samples for which the respective measured analyte level is above the matrix background. In some examples, step 206 includes determining a percentage of samples for which the measured analyte level is above the matrix background. In some examples, determining this percentage includes determining a percentage of samples for which the measured analyte level is at or below the matrix background, and inferring that the complementary percentage is above the matrix background.
[0112] At step 208, method 200 includes predicting, based on the determined information about the samples for which the measured analyte level was above the matrix background, a correlation between measurements of the analyte made on the first platform and measurements of the analyte made on a second platform. The result of the prediction may comprise any suitable qualitative or quantitative characterization of the cross-platform correlation, including a Spearman correlation, Kendall correlation, Pearson correlation, and / or any other suitable correlation. Predicting the correlation based on the determined information (e.g., percentage of samples above matrix background) may comprise any suitable process, such as a regression analysis, machine learning model, neural network, and / or any other suitable process.
[0113] In some examples, the correlation is predicted based on the determined information along with other factors. For example, percentage of samples above matrix background may be one of a plurality of features of a regression model.
[0114] In general, a high correlation indicates that measurements of the analyte from the first platform and measurements of the analyte from the second platform can reliably beKoiitch Romano Dascenzo Gates LLP Attorney Docket No. SML317PCT
[0115] compared to one another. A low correlation indicates that such a comparison is not likely to be reliable.
[0116] In some examples, the cross-correlation predicted by method 200 (and / or any other suitable method) may be used to determine information about one or both of the platforms. For example, suppose a first platform may be utilized to perform an assay for a first plurality of analytes, and a second platform may be utilized to perform an assay for a second plurality of analytes. Method 200 may be utilized to predict a cross-platform correlation between the first and second platforms for a first analyte. If the first plurality of analytes and the second plurality of analytes both include the first analyte, then the result of method 200 should be a relatively high cross-correlation (e.g., a cross-correlation sufficiently high to be deemed meaningful according to an appropriate statistical evaluation). On the other hand, if the first analyte is not included in the first plurality of analytes, or is not included in the second plurality of analytes, or is included in neither the first nor the second plurality of analytes, then the cross-correlation predicted by method 200 should be low (e.g. sufficiently low to be deemed nonmeaningful according to an appropriate statistical evaluation).
[0117] Accordingly, in some examples, a method of identifying problem(s) in a platform and / or an assay performed thereon includes predicting a cross-correlation for a first analyte according to method 200 (and / or any other suitable method) between a first platform performing a first assay and a second platform performing a second assay. If it was believed that the first and second assays both included the first analyte, and the predicted cross-correlation is low, then the method of identifying problem(s) further includes investigating the first and / or second platforms and / or assays for problems. If it was believed that the first and second assays both included the first analyte, and the predicted cross-correlation is high, then the method of identifying problem(s) further includes determining that the method has not revealed any cause for concern. If it was believed that at least one of the first and second assays did not include the first analyte, and the predicted cross-correlation is low, then the method further includes determining that the method has not revealed any cause for concern. If it was believed that at least one of the first and second assays did not include the first analyte, and the predicted cross-correlation is high, then the method further includes investigating the first and / orKoiitch Romano Dascenzo Gates LLP Attorney Docket No. SML317PCT
[0118] second platforms and / or assays for problems. Investigating a platform or assay for problems may include comparing parameters or measured signals to benchmarks, repeating the assay, using an alternative method to validate the assay, reviewing instrument and software logs for errors, checking instrument settings, checking sample(s) for signs of contamination, and / or any other suitable steps. Additionally, or alternatively, in some examples, the predicted cross-correlation may be used to determine whether a first assay performed on a first platform included a first analyte and / or whether the first analyte signaled above the matrix background in the first assay. In some examples, such a method includes predicting a cross-correlation for the first analyte between the first platform and a second platform on which a second assay is performed. The second assay performed on the second platform is known to include the first analyte and is known to be configured such that the first analyte should signal above the matrix background. In response to the cross-correlation being low, the method includes inferring that the first assay on the first platform does not include the first analyte or did not measure a signal above the matrix background for the first analyte. In response to the cross-correlation being high, the method includes inferring that the first assay on the first platform does include the first analyte signaling above the matrix background. This method may optionally be repeated fora plurality of analytes in order to determine which analytes were included in the first assay and signaled above the matrix background. In some examples in which the method is repeated for two or more analytes, at least some of the analytes may be included in a third (etc.) assay on a third (etc.) platform, such that at least some repetitions of the method include predicting a cross-platform correlation between the first platform and the third (etc.) platform.
[0119] Illustrative Method for Validating Aptamers
[0120] This section describes an example method of validating new aptamers for potential use in assays and / or other applications.
[0121] In general, candidate aptamers may be discovered through in vitro methods, through computational methods (also known as in silico methods), and / or through any other suitable methods. As an example, the SELEX process (Systematic Evolution of Ligands by Exponential enrichment) is a method for the in vitro evolution of nucleic acid molecules for a certain desired activity. SELEX can be used to identify aptamers that areKoiitch Romano Dascenzo Gates LLP Attorney Docket No. SML317PCT
[0122] configured to specifically bind to a given target molecule with high affinity. The SELEX process provides a class of products which are referred to as nucleic acid ligands or aptamers, each having a unique sequence, and having the property of binding specifically to a desired target compound or molecule. Each SELEX-identified nucleic acid capture reagent is a specific ligand of a given target compound or molecule. The SELEX process is based on the unique insight that nucleic acids have sufficient capacity for forming a variety of two- and three-dimensional structures and sufficient chemical versatility available within their monomers to act as ligands (form specific binding pairs) with virtually any chemical compound, whether monomeric or polymeric. Molecules of any size or composition can serve as targets.
[0123] The SELEX method involves selection from a mixture of candidate oligonucleotides and stepwise iterations of binding, partitioning and amplification, using the same general selection scheme, to achieve virtually any desired criterion of binding affinity and selectivity. Starting from a mixture of nucleic acids, preferably comprising a segment of randomized sequence, the SELEX method includes steps of contacting the mixture with the target under conditions favorable for binding, partitioning unbound nucleic acids from those nucleic acids which have bound specifically to target molecules, dissociating the nucleic acid-target complexes, amplifying the nucleic acids dissociated from the nucleic acid-target complexes to yield a ligand-enriched mixture of nucleic acids, and then reiterating the steps of binding, partitioning, dissociating and amplifying through as many cycles as desired to yield highly specific high affinity nucleic acid ligands to the target molecule. In this manner, aptamers suitable for binding to virtually any desired target protein can be discovered.
[0124] In some cases, it is desirable to validate a candidate aptamer that has been identified through the SELEX process and / or any other suitable process to confirm that the actual binding performance of the aptamer is suitable. In some examples, a method of validating the identified candidate aptamer comprises: estimating the matrix background for the aptamer in a binding assay (e.g., using method 100, method 150, and / or any other suitable method); and comparing the measured signal associated with the aptamer to the estimated matrix background to determine whether the aptamer should be deemed validated. In some such examples, the method further includes obtaining theKoiitch Romano Dascenzo Gates LLP Attorney Docket No. SML317PCT
[0125] signal levels that are used for estimating the matrix background by measuring the signal levels (e.g., by performing an assay).
[0126] In some examples, comparing the aptamer to the estimated matrix background comprises determining a number or percentage of samples for which the aptamer signals above the estimated matrix background. For example, in some cases, the aptamer is deemed to be validated if the aptamer signals above the estimated matrix background for at least a threshold percentage of unique samples. In some examples, the threshold is 75%, but in other examples, the threshold may be any other suitable percentage (e.g., 70%, 65%, 60%, 55%, 50%, 45%, 40% or less). Alternatively, or additionally, the aptamer may be deemed non-valid (i.e., unsuitable for use) if the aptamer signals above the estimated matrix background, for analytes known to be present in the matrix, for no more than a threshold percentage of unique samples, such as 20% of unique samples. When an aptamer signals above matrix background for such a small percentage of samples, one can reasonably infer that the aptamer is unsuitable for use in detecting analytes.
[0127] In some examples, comparing the aptamer to the estimated matrix background comprises comparing the measured signals associated with the aptamer to an expected distribution of signals. For example, in some cases, the aptamer is deemed to be validated if the measured signals associated with the aptamer reflect a bimodal distribution that is expected to be present in the samples. For example, if the aptamer is believed to detect an analyte that can be present in male populations but generally not in female populations, and the samples of the assay include samples from male populations and from female populations, the aptamer may be deemed validated if the measured signals associated with the aptamer follow a bimodal distribution in which one subset of samples signals above matrix background (e.g., at least a threshold percentage of samples of that subset signal above matrix background) and another subset fails to signal above matrix background (e.g., no more than a threshold percentage of samples of that subset signal above matrix background).
[0128] In some examples of the validation method, the aptamer’s signal is measured in more than one assay so as to increase confidence in the conclusion that the aptamer is valid. Thus, the aptamer may be deemed to be validated in response to determining thatKoiitch Romano Dascenzo Gates LLP Attorney Docket No. SML317PCT
[0129] the measured signal is above the estimated matrix background for a threshold percentage of samples in a suitable number of assays and / or with a suitable statistical confidence.
[0130] In some examples, the validation method is repeated in a plurality of assays of different types, and based on the assay types for which the measured signal is above matrix background, information about the assay types, contexts, and / or use cases in which the aptamer is suitable for use may be determined.
[0131] In some examples of the validation method, determining whether the measured signal is above the estimated matrix background includes determining how far above the estimated matrix background the measured signal is, and the aptamer is validated only in response to determining that the measured signal is sufficiently high above the estimated matrix background. For example, in some cases the method includes validating the aptamer only in response to determining that the measured signal is at least three standard deviations above the estimated matrix background. This may help to ensure that the aptamer is validated only if it has a very high likelihood of performing adequately under realistic experimental conditions.
[0132] In some examples, a method of identifying suitable aptamers comprises performing SELEX and / or another suitable method to identify one or more candidate aptamers to be validated, and performing the validation method(s) described above in this section on one or more of the identified candidate aptamers to determine which (if any) of the identified candidate aptamers are actually suitable for use.
[0133] Illustrative Method for Evaluating an Assay
[0134] This section describes example methods of evaluating an assay using the matrix background (as estimated by method 100, method 150, and / or any other suitable method) as a feedback metric. The matrix background represents a baseline signal caused by nonspecific binding and / or interference from the sample matrix (e.g., serum or plasma). A lower matrix background indicates less nonspecific binding and better assay specificity; a higher matrix background indicates greater nonspecific binding and worse assay specificity. Accordingly, if adjusting an aspect of an assay reduces the matrix background, it can be inferred that the adjustment improved the assay’s accuracy, and if adjusting the aspect increased the matrix background, it can be inferred that the adjustment worsened the assay’s accuracy. The matrix background may therefore be used to validate an assay,Koiitch Romano Dascenzo Gates LLP Attorney Docket No. SML317PCT
[0135] evaluate changes to an assay, improve an assay, and / or to understand the effects of aspect(s) of an assay on accuracy.
[0136] In some examples, a method of evaluating an assay includes performing an assay using a first set of parameters and estimating a first matrix background for a first aptamer in the assay (e.g., using method 100, method 150, and / or any other suitable method). The method further includes adjusting one or more parameters of the first set of parameters, performing the assay using the adjusted first set of parameters, and estimating a second matrix background for the first aptamer in the assay. In response to the second matrix background being lower than the first matrix background, it is determined that the adjustment to the one or more parameters improved the assay. In response to the second matrix background being higher than the first matrix background, it is determined that the adjustment to the one or more parameters worsened the assay. The method may be repeated, adjusting different parameters and / or different combinations of parameters of the assay and determining whether those adjustments improved or worsened the assay. This method may be thought of as a method for identifying suitable parameter values for the assay.
[0137] In some examples, the method is performed two or more times adjusting the same parameter in a different way. For example, the method may include performing the assay, adjusting a first parameter of the assay by a first amount relative to its initial value, performing the assay with the adjusted parameter, adjusting the first parameter by a second amount relative to its initial value, and performing the assay with the newly adjusted parameter. In this manner, the method may yield further information about how adjusting the first parameter affects the performance of the assay. For example, it may be determined that adjusting the first parameter has a linear effect on the assay, or a nonlinear effect on the assay, or that the effect is linear for some regimes of the first parameter and nonlinear for others.
[0138] In some examples, a method of measuring presence and / or abundance of a target analyte includes evaluating an assay using one or more method steps described above in this section to determine suitable value(s) for one or more parameters for the assay; and performing the assay using the determined suitable value(s) for the parameter(s) so as to accurately measure the presence and / or abundance of the target molecule.Koiitch Romano Dascenzo Gates LLP Attorney Docket No. SML317PCT
[0139] Suitable parameters to be adjusted in the method(s) described in this section may include, without limitation: buffer composition (e.g., type and / or concentration of salts and additives, pH of the composition); type and / or concentration of blocking agent; incubation time; incubation temperature; number of washes and / or strength of wash buffers; protein concentration; aptamer concentration; matrix dilution; surface chemistry of assay plates, beads, and / or the like; concentration and / or other properties of detection labels.
[0140] Illustrative Example Application Using EDTA-Plasma Samples
[0141] This section describes an example application of methods according to aspects of the present teachings. The example of this section involves a set of 1020 samples of human EDTA-plasma collected by Covance, Inc. The 1020 samples were run on the SomaScan® 11 K Assay v5.0 platform, which is an aptamer-based proteomics assay platform. Twelve assay runs were performed for the 1020 population samples plus 36 quality control replicate samples. Respective signals corresponding to a plurality of aptamers were measured. As discussed below, the aptamers measured include aptamers configured to bind to human proteins as well as aptamers not configured to bind to human proteins (e.g., spuriomers, randomers, aptamers configured to bind to non-human proteins). The aptamers not configured to bind to human proteins may be considered a control group, referred to as non-human controls.
[0142] Cumulative F-statistics were calculated for each aptamer. Fig. 5 is a plot depicting cumulative F-statistic values calculated in accordance with aspects of the present teachings against signal strength in relative fluorescence units (RFU) for a particular example aptamer configured to bind to human protein. Each point on the plot represents an F-statistic value calculated using a respective cumulative subset of the 1020 samples. The samples are ordered from the sample having the weakest signal for the aptamer in question (sample #1 ) to the sample having the strongest signal for the aptamer in question (sample #1020).
[0143] The initial F-statistic was calculated using the first 100 samples, i.e., using N = 100 in Equation 2. Each subsequent F-statistic was calculated using an additional 20 samples. That is, the second F-statistic was calculated using N = 120, the third F-statistic was calculated using N = 140, and so on. The assay variance was taken to be the variance associated with the set of 36 quality control replicates, and thus was assumed to haveKoiitch Romano Dascenzo Gates LLP Attorney Docket No. SML317PCT
[0144] the same value irrespective of N. The number of degrees of freedom in the numerator was N - 1 and the number of degrees of freedom in the denominator was the number of quality control replicates minus 1 (i.e., 36 - 1 = 35). Critical values of the F-statistic corresponding to a 95% confidence level were used, adjusted by a Bonferroni multiplecomparison correction.
[0145] As shown in Fig. 5, the F-statistic calculated using N = 280 was above the critical value. Accordingly, the signal strength corresponding to the aptamer in question in the most recently added batch of 20 samples (e.g., the highest RFU value of the most recently added batch of samples) is considered a “critical RFU value”, and the matrix background is inferred to be approximately equal to the critical RFU value. In this example, the critical RFU value is 465 RFU, and so the matrix background is estimated to be 465 RFU.
[0146] For illustrative purposes, Fig. 5 includes F-statistics calculated for even greater subsets than the N = 280 subset. That is, in the example of Fig. 5, additional F-statistics were calculated even after arriving at a value of N for which the F-statistic F(N) is above the critical value. In general, there is no need to continue expanding the subset and recalculating the F-statistic based on the expanded subset after the critical value has been reached. However, in some examples, the F-statistic may be recalculated for expanded subsets after reaching the critical value so as to confirm the results, evaluate the method, and / or for any other suitable reason.
[0147] Fig. 6 is a histogram depicting a plurality of bins of signal strength for the aptamer, and the number of samples having a signal strength within each bin. The critical RFU value of 465 RFU is indicated in a vertical dashed line. In this example, approximately 27.45% of the 1020 samples measured above the critical RFU value for the aptamer in question.
[0148] The matrix background was estimated for additional aptamers of the SomaScan® 11 K Assay v5.0 menu for the same 1020 samples. For each aptamer, the percentage of samples that signaled above the matrix background corresponding to the aptamer was calculated. Each aptamer falls into one of the following three categories: no samples signaled above the matrix background, or all 1020 samples signaled above the matrix background, or some but not all samples (i.e., between 1 and 1019 samples) signaled above the matrix background.Koiitch Romano Dascenzo Gates LLP Attorney Docket No. SML317PCT
[0149] Fig. 7 is a table depicting percentage of aptamers falling into each of the above three categories, broken down into two categories of aptamer. The first category, listed in the first column of the table, corresponds to aptamers configured to bind to human proteins. As the table shows, 69.3% of aptamers configured to bind to human proteins signaled above the matrix background in some of the samples (labeled “Some Sig” in Fig.
[0150] 7), and 16.6% signaled above the matrix background in all 1020 samples (labeled “All Sig” in Fig. 7), with only 14.0% of these aptamers signaling above the matrix background in zero of the samples (labeled “No Sig” in Fig. 7).
[0151] In contrast, the second category of aptamers, corresponding to negative control aptamers and / or aptamers configured to bind to non-human proteins, were significantly less likely to signal above the matrix background. For example, as shown in the second column of the table of Fig. 7, 72.5% of aptamers in the second category did not signal above the matrix background in any of the 1020 samples, and no aptamer in the second category signaled above the matrix background in all samples. This result is consistent with the idea that biological variation is less significant than assay variation for non-human control aptamers. Accordingly, the percentage of samples signaling above the matrix background for a given aptamer can be understood as a predictor of the aptamer’s specificity.
[0152] Fig. 8 is a box plot depicting Spearman correlations between the SomaScan platform and an enzyme-linked immunosorbent assay (ELISA) platform. More specifically, Fig. 8 depicts distributions of Spearman correlation values for the group of aptamers for which 0% to 50% of samples signaled above the matrix background, the group for which 50% to 75% of samples signaled above the matrix background, and the group for which 75% to 100% of samples signaled above the matrix background.
[0153] As Fig. 8 shows, for the group of aptamers for which 0% to 50% of samples signaled above matrix background, the median value of the cross-platform Spearman correlation is relatively low. This result is consistent with the idea that the assay variation tends to be greater than the biological variation for this group of aptamers. The median value of the cross-platform Spearman correlation is much higher for the group of aptamers for which 75% to 100% of samples signaled above matrix background. This result is consistent with the idea that the biological variation tends to be detectable aboveKoiitch Romano Dascenzo Gates LLP Attorney Docket No. SML317PCT
[0154] the matrix background for these aptamers. More variation in cross-platform Spearman correlation is observed for the group of aptamers for which 50% to 75% of samples signaled above matrix background.
[0155] Overall, Fig. 8 shows that the percentage of samples above matrix background as measured on a first platform is predictive of the correlation with a second platform for a given aptamer.
[0156] Illustrative Data Processing System
[0157] With reference to Fig. 9, this example describes a data processing system 800 (also referred to as a computer, computing system, and / or computer system) in accordance with aspects of the present disclosure. In this example, data processing system 800 is an illustrative data processing system suitable for implementing aspects of the methods described herein. More specifically, in some examples, devices that are embodiments of data processing systems (e.g., smartphones, tablets, personal computers) may be used to compute cumulative F-statistic values, compare a value of an F-statistic to a critical value, predict cross-platform correlation, and / or perform any suitable method, method step, or function described herein.
[0158] In this illustrative example, data processing system 800 includes a system bus 802 (also referred to as communications framework). System bus 802 may provide communications between a processor unit 804 (also referred to as a processor or processors), a memory 806, a persistent storage 808, a communications unit 810, an input / output (I / O) unit 812, a codec 830, and / or a display 814. Memory 806, persistent storage 808, communications unit 810, input / output (I / O) unit 812, display 814, and codec 830 are examples of resources that may be accessible by processor unit 804 via system bus 802.
[0159] Processor unit 804 serves to run instructions that may be loaded into memory 806. Processor unit 804 may comprise a number of processors, a multi-processor core, and / or a particular type of processor or processors (e.g., a central processing unit (CPU), graphics processing unit (GPU), etc.), depending on the particular implementation. Further, processor unit 804 may be implemented using a number of heterogeneous processor systems in which a main processor is present with secondary processors on aKoiitch Romano Dascenzo Gates LLP Attorney Docket No. SML317PCT
[0160] single chip. As another illustrative example, processor unit 804 may be a symmetric multiprocessor system containing multiple processors of the same type.
[0161] Memory 806 and persistent storage 808 are examples of storage devices 816. A storage device may include any suitable hardware capable of storing information (e.g., digital information), such as data, program code in functional form, and / or other suitable information, either on a temporary basis or a permanent basis.
[0162] Storage devices 816 also may be referred to as computer-readable storage devices or computer-readable media. Memory 806 may include a volatile storage memory 840 and a non-volatile memory 842. In some examples, a basic input / output system (BIOS), containing the basic routines to transfer information between elements within the data processing system 800, such as during start-up, may be stored in non-volatile memory 842. Persistent storage 808 may take various forms, depending on the particular implementation.
[0163] Persistent storage 808 may contain one or more components or devices. For example, persistent storage 808 may include one or more devices such as a magnetic disk drive (also referred to as a hard disk drive or HDD), solid state disk (SSD), floppy disk drive, tape drive, Jaz drive, Zip drive, flash memory card, memory stick, and / or the like, or any combination of these. One or more of these devices may be removable and / or portable, e.g., a removable hard drive. Persistent storage 808 may include one or more storage media separately or in combination with other storage media, including an optical disk drive such as a compact disk ROM device (CD-ROM), CD recordable drive (CD-R Drive), CD rewritable drive (CD-RW Drive), and / or a digital versatile disk ROM drive (DVD-ROM). To facilitate connection of the persistent storage devices 808 to system bus 802, a removable or non-removable interface is typically used, such as interface 828.
[0164] Input / output (I / O) unit 812 allows for input and output of data with other devices that may be connected to data processing system 800 (i.e., input devices and output devices). For example, an input device may include one or more pointing and / or information-input devices such as a keyboard, a mouse, a trackball, stylus, touch pad or touch screen, microphone, joystick, game pad, satellite dish, scanner, TV tuner card, digital camera, digital video camera, web camera, and / or the like. These and other input devices may connect to processor unit 804 through system bus 802 via interface port(s).Koiitch Romano Dascenzo Gates LLP Attorney Docket No. SML317PCT
[0165] Suitable interface port(s) may include, for example, a serial port, a parallel port, a game port, and / or a universal serial bus (USB).
[0166] One or more output devices may use some of the same types of ports, and in some cases the same actual ports, as the input device(s). For example, a USB port may be used to provide input to data processing system 800 and to output information from data processing system 800 to an output device. One or more output adapters may be provided for certain output devices (e.g., monitors, speakers, and printers, among others) which require special adapters. Suitable output adapters may include, e.g. video and sound cards that provide a means of connection between the output device and system bus 802. Other devices and / or systems of devices may provide both input and output capabilities, such as remote computer(s) 860. Display 814 may include any suitable human-machine interface or other mechanism configured to display information to a user, e.g., a CRT, LED, or LCD monitor or screen, etc.
[0167] Communications unit 810 refers to any suitable hardware and / or software employed to provide for communications with other data processing systems or devices. While communication unit 810 is shown inside data processing system 800, it may in some examples be at least partially external to data processing system 800. Communications unit 810 may include internal and external technologies, e.g., modems (including regular telephone grade modems, cable modems, and DSL modems), ISDN adapters, and / or wired and wireless Ethernet cards, hubs, routers, etc. Data processing system 800 may operate in a networked environment, using logical connections to one or more remote computers 860. A remote computer(s) 860 may include a personal computer (PC), a server, a router, a network PC, a workstation, a microprocessor-based appliance, a peer device, a smart phone, a tablet, another network note, and / or the like. Remote computer(s) 860 typically include many of the elements described relative to data processing system 800. Remote computer(s) 860 may be logically connected to data processing system 800 through a network interface 862 which is connected to data processing system 800 via communications unit 810. Network interface 862 encompasses wired and / or wireless communication networks, such as local-area networks (LAN), wide-area networks (WAN), and cellular networks. LAN technologies may include Fiber Distributed Data Interface (FDDI), Copper Distributed Data InterfaceKoiitch Romano Dascenzo Gates LLP Attorney Docket No. SML317PCT
[0168] (CDDI), Ethernet, Token Ring, and / or the like. WAN technologies include point-to-point links, circuit switching networks (e.g., Integrated Services Digital networks (ISDN) and variations thereon), packet switching networks, and Digital Subscriber Lines (DSL).
[0169] Codec 830 may include an encoder, a decoder, or both, comprising hardware, software, or a combination of hardware and software. Codec 830 may include any suitable device and / or software configured to encode, compress, and / or encrypt a data stream or signal for transmission and storage, and to decode the data stream or signal by decoding, decompressing, and / or decrypting the data stream or signal (e.g., for playback or editing of a video). Although codec 830 is depicted as a separate component, codec 830 may be contained or implemented in memory, e.g., non-volatile memory 842.
[0170] Non-volatile memory 842 may include read only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), flash memory, and / or the like, or any combination of these. Volatile memory 840 may include random access memory (RAM), which may act as external cache memory. RAM may comprise static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), and / or the like, or any combination of these.
[0171] Instructions for the operating system, applications, and / or programs may be located in storage devices 816, which are in communication with processor unit 804 through system bus 802. In these illustrative examples, the instructions are in a functional form in persistent storage 808. These instructions may be loaded into memory 806 for execution by processor unit 804. Processes of one or more embodiments of the present disclosure may be performed by processor unit 804 using computer-implemented instructions, which may be located in a memory, such as memory 806.
[0172] These instructions are referred to as program instructions, program code, computer usable program code, or computer-readable program code executed by a processor in processor unit 804. The program code in the different embodiments may be embodied on different physical or computer-readable storage media, such as memory 806 or persistent storage 808. Program code 818 may be located in a functional form on computer-readable media 820 that is selectively removable and may be loaded onto or transferred to data processing system 800 for execution by processor unit 804. ProgramKoiitch Romano Dascenzo Gates LLP Attorney Docket No. SML317PCT
[0173] code 818 and computer-readable media 820 form computer program product 822 in these examples. In one example, computer-readable media 820 may comprise computer-readable storage media 824 or computer-readable signal media 826.
[0174] Computer-readable storage media 824 may include, for example, an optical or magnetic disk that is inserted or placed into a drive or other device that is part of persistent storage 808 for transfer onto a storage device, such as a hard drive, that is part of persistent storage 808. Computer-readable storage media 824 also may take the form of a persistent storage, such as a hard drive, a thumb drive, or a flash memory, that is connected to data processing system 800. In some instances, computer-readable storage media 824 may not be removable from data processing system 800.
[0175] In these examples, computer-readable storage media 824 is a non-transitory, physical or tangible storage device used to store program code 818 rather than a medium that propagates or transmits program code 818. Computer-readable storage media 824 is also referred to as a computer-readable tangible storage device or a computer-readable physical storage device. In other words, computer-readable storage media 824 is media that can be touched by a person.
[0176] Alternatively, program code 818 may be transferred to data processing system 800, e.g., remotely over a network, using computer-readable signal media 826. Computer-readable signal media 826 may be, for example, a propagated data signal containing program code 818. For example, computer-readable signal media 826 may be an electromagnetic signal, an optical signal, and / or any other suitable type of signal. These signals may be transmitted over communications links, such as wireless communications links, optical fiber cable, coaxial cable, a wire, and / or any other suitable type of communications link. In other words, the communications link and / or the connection may be physical or wireless in the illustrative examples.
[0177] In some illustrative embodiments, program code 818 may be downloaded over a network to persistent storage 808 from another device or data processing system through computer-readable signal media 826 for use within data processing system 800. For instance, program code stored in a computer-readable storage medium in a server data processing system may be downloaded over a network from the server to data processing system 800. The computer providing program code 818 may be a server computer, aKoiitch Romano Dascenzo Gates LLP Attorney Docket No. SML317PCT
[0178] client computer, or some other device capable of storing and transmitting program code 818.
[0179] In some examples, program code 818 may comprise an operating system (OS) 850. Operating system 850, which may be stored on persistent storage 808, controls and allocates resources of data processing system 800. One or more applications 852 take advantage of the operating system’s management of resources via program modules 854, and program data 856 stored on storage devices 816. OS 850 may include any suitable software system configured to manage and expose hardware resources of computer 800 for sharing and use by applications 852. In some examples, OS 850 provides application programming interfaces (APIs) that facilitate connection of different type of hardware and / or provide applications 852 access to hardware and OS services. In some examples, certain applications 852 may provide further services for use by other applications 852, e.g., as is the case with so-called “middleware.” Aspects of present disclosure may be implemented with respect to various operating systems or combinations of operating systems.
[0180] The different components illustrated for data processing system 800 are not meant to provide architectural limitations to the manner in which different embodiments may be implemented. One or more embodiments of the present disclosure may be implemented in a data processing system that includes fewer components or includes components in addition to and / or in place of those illustrated for computer 800. Other components shown in Fig. 9 can be varied from the examples depicted. Different embodiments may be implemented using any hardware device or system capable of running program code. As one example, data processing system 800 may include organic components integrated with inorganic components and / or may be comprised entirely of organic components (excluding a human being). For example, a storage device may be comprised of an organic semiconductor.
[0181] In some examples, processor unit 804 may take the form of a hardware unit having hardware circuits that are specifically manufactured or configured for a particular use, or to produce a particular outcome or progress. This type of hardware may perform operations without needing program code 818 to be loaded into a memory from a storage device to be configured to perform the operations. For example, processor unit 804 mayKoiitch Romano Dascenzo Gates LLP Attorney Docket No. SML317PCT
[0182] be a circuit system, an application specific integrated circuit (ASIC), a programmable logic device, or some other suitable type of hardware configured (e.g., preconfigured or reconfigured) to perform a number of operations. With a programmable logic device, for example, the device is configured to perform the number of operations and may be reconfigured at a later time. Examples of programmable logic devices include, a programmable logic array, a field programmable logic array, a field programmable gate array (FPGA), and other suitable hardware devices. With this type of implementation, executable instructions (e.g., program code 818) may be implemented as hardware, e.g., by specifying an FPGA configuration using a hardware description language (HDL) and then using a resulting binary file to (re)configure the FPGA.
[0183] In another example, data processing system 800 may be implemented as an FPGA-based (or in some cases ASIC-based), dedicated-purpose set of state machines (e.g., Finite State Machines (FSM)), which may allow critical tasks to be isolated and run on custom hardware. Whereas a processor such as a CPU can be described as a shared-use, general purpose state machine that executes instructions provided to it, FPGA-based state machine(s) are constructed for a special purpose, and may execute hardware-coded logic without sharing resources. Such systems are often utilized for safety-related and mission-critical tasks.
[0184] In still another illustrative example, processor unit 804 may be implemented using a combination of processors found in computers and hardware units. Processor unit 804 may have a number of hardware units and a number of processors that are configured to run program code 818. With this depicted example, some of the processes may be implemented in the number of hardware units, while other processes may be implemented in the number of processors.
[0185] In another example, system bus 802 may comprise one or more buses, such as a system bus or an input / output bus. Of course, the bus system may be implemented using any suitable type of architecture that provides for a transfer of data between different components or devices attached to the bus system. System bus 802 may include several types of bus structure(s) including memory bus or memory controller, a peripheral bus or external bus, and / or a local bus using any variety of available bus architectures (e.g., Industrial Standard Architecture (ISA), Micro-Channel Architecture (MSA), Extended ISAKoiitch Romano Dascenzo Gates LLP Attorney Docket No. SML317PCT
[0186] (EISA), Intelligent Drive Electronics (IDE), VESA Local Bus (VLB), Peripheral Component Interconnect (PCI), Card Bus, Universal Serial Bus (USB), Advanced Graphics Port (AGP), Personal Computer Memory Card International Association bus (PCMCIA), Firewire (IEEE 1394), and Small Computer Systems Interface (SCSI)).
[0187] Additionally, communications unit 810 may include a number of devices that transmit data, receive data, or both transmit and receive data. Communications unit 810 may be, for example, a modem or a network adapter, two network adapters, or some combination thereof. Further, a memory may be, for example, memory 806, or a cache, such as that found in an interface and memory controller hub that may be present in system bus 802.
[0188] The flowcharts and block diagrams described herein illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various illustrative embodiments. In this regard, each block in the flowcharts or block diagrams may represent a module, segment, or portion of code, which comprises one or more executable instructions for implementing the specified logical function or functions. It should also be noted that, in some alternative implementations, the functions noted in a block may occur out of the order noted in the drawings. For example, the functions of two blocks shown in succession may be executed substantially concurrently, or the functions of the blocks may sometimes be executed in the reverse order, depending upon the functionality involved.
[0189] Illustrative Combinations and Additional Examples
[0190] This section describes additional aspects and features of systems and methods for estimating matrix background, presented without limitation as a series of paragraphs, some or all of which may be alphanumerically designated for clarity and efficiency. Each of these paragraphs can be combined with one or more other paragraphs, and / or with disclosure from elsewhere in this application, in any suitable manner. Some of the paragraphs below expressly refer to and further limit other paragraphs, providing without limitation examples of some of the suitable combinations.
[0191] A1. A computer-implemented method for estimating a matrix background of a protein-binding reagent in an assay, the method comprising:Koiitch Romano Dascenzo Gates LLP Attorney Docket No. SML317PCT
[0192] determining a first F-statistic based on a first subset of samples of a plurality of samples, wherein each sample of the plurality of samples has a signal level associated with the protein-binding reagent;
[0193] in response to determining that the first F-statistic is less than a first critical value, determining a second F-statistic based on a second subset of samples of the plurality of samples; and
[0194] in response to determining that the second F-statistic is at least a second critical value, estimating the matrix background of the protein-binding reagent based on the signal levels of one or more samples of the second subset.
[0195] A2. The method of paragraph A1, wherein the second subset of samples includes the first subset of samples and a batch of one or more samples that are not included in the first subset of samples.
[0196] A3. The method of paragraph A2, wherein estimating the matrix background based on the signal levels of the one or more samples of the second subset comprises estimating the matrix background based on the signal levels of the batch of one or more samples that are not included in the first subset of samples.
[0197] A4. The method of paragraph A3, wherein each signal level of the batch of one or more samples is higher than a highest signal level of the first subset of samples.
[0198] A5. The method of any one of paragraphs A3-A4, wherein estimating the matrix background based on the signal levels of the batch of one or more samples comprises:
[0199] estimating the matrix background to be equal to a highest signal level of the signal levels of the batch of one or more samples.
[0200] A6. The method of any one of paragraphs A1 -A5, wherein the signal levels of the plurality of samples are associated with a first assay platform, the method further comprising predicting a correlation for the protein-binding reagent between the first assayKoiitch Romano Dascenzo Gates LLP Attorney Docket No. SML317PCT
[0201] platform and a second assay platform based on the estimated matrix background of the protein-binding reagent.
[0202] A7. The method of paragraph A6, wherein predicting the correlation for the protein-binding reagent based on the estimated matrix background of the protein-binding reagent comprises determining a percentage of samples of the plurality of samples having signal levels greater than the estimated matrix background, and predicting the correlation based on the determined percentage.
[0203] A8. The method of any one of paragraphs A1 -A7, wherein determining the first F-statistic comprises determining a first variance of the signal levels of the first subset of samples, determining a second variance of signals associated with nonspecific binding to the protein-binding reagent, and dividing the first variance by the second variance.
[0204] A9. The method of paragraph A8, wherein the signals associated with nonspecific binding to the protein-binding reagent comprise signals associated with quality control replicates.
[0205] A10. The method of any one of paragraphs A1-A9, further comprising obtaining the signal levels of the plurality of samples by measuring the signal levels using the assay.
[0206] A11. The method of any one of paragraphs A1 -A10, wherein the protein-binding reagent is selected from the following: an aptamer, an antibody.
[0207] B1. A computer-implemented method for determining a matrix background of a target analyte in an assay, the method comprising:
[0208] determining, for each of a plurality of sets of samples, whether a biological variability of the set is greater than an assay variability of the set, wherein the biological variability is associated with levels of a target analyte in the samples of the set, and the assay variability is associated with nonspecific binding in the assay; andKoiitch Romano Dascenzo Gates LLP Attorney Docket No. SML317PCT
[0209] determining the matrix background based on at least a first set and a second set of the plurality of sets, wherein the assay variability of the first set is greater than the biological variability of the first set, and the biological variability of the second set is greater than the assay variability of the second set;
[0210] wherein the second set of samples comprises at least one sample having a higher level of the target analyte than any sample of the first set.
[0211] B2. The method of paragraph B1 , wherein the second set of samples comprises all the samples of the first set of samples and the at least one sample having a higher level of the target analyte than any sample of the first set.
[0212] B3. The method of paragraph B2, wherein the plurality of sets of samples are a cumulative series of sets ordered from a smallest set to a largest set, with each set of samples including all of the samples of the previous set and at least one sample not included in the previous set; and
[0213] wherein determining whether the biological variability is greater than the assay variability for each of the plurality of sets of samples comprises:
[0214] performing respective comparisons of the biological variability and assay variability of each set of the cumulative series in order; and
[0215] stopping the comparisons in response to determining that the biological variability of the second set is greater than the assay variability of the second set.
[0216] B4. The method of any one of paragraphs B1 -B3, wherein determining whether the biological variability is greater than the assay variability for each of the plurality of sets of samples comprises:
[0217] for each of the plurality of sets of samples, performing an analysis of variance (ANOVA) test comparing a variance of levels of the target analyte in the samples of the set to a variance of signals of one or more technical replicates.
[0218] B5. The method of paragraph B4, wherein performing the ANOVA test for each set of samples comprises calculating an F-statistic for each set of samples.Koiitch Romano Dascenzo Gates LLP Attorney Docket No. SML317PCT
[0219] C1. A data processing system for estimating a matrix background of a proteinbinding reagent in an assay, the system comprising:
[0220] one or more processors;
[0221] a memory; and
[0222] a plurality of instructions stored in the memory and executable by the one or more processors to:
[0223] determine a first F-statistic based on a first subset of samples of a plurality of samples, wherein each sample of the plurality of samples has a signal level associated with the protein-binding reagent;
[0224] in response to determining that the first F-statistic is less than a first critical value, determining a second F-statistic based on a second subset of samples of the plurality of samples; and
[0225] in response to determining that the second F-statistic is at least a second critical value, estimating the matrix background of the protein-binding reagent based on the signal levels of one or more samples of the second subset.
[0226] C2. The data processing system of paragraph C1 , wherein the second subset of samples includes the first subset of samples and a batch of one or more samples that are not included in the first subset of samples.
[0227] C3. The data processing system of paragraph C2, wherein estimating the matrix background based on the signal levels of the one or more samples of the second subset comprises estimating the matrix background based on the signal levels of the batch of one or more samples.
[0228] C4. The data processing system of paragraph C3, wherein each signal level of the batch of one or more samples is higher than a highest signal level of the first subset of samples.Koiitch Romano Dascenzo Gates LLP Attorney Docket No. SML317PCT
[0229] C5. The data processing system of any one of paragraphs C1-C4, wherein determining the first F-statistic comprises determining a first variance of the signal levels of the first subset of samples, determining a second variance of signals associated with nonspecific binding to the aptamer, and dividing the first variance by the second variance.
[0230] D1. A data processing system for estimating a matrix background of a proteinbinding reagent in an assay , the system comprising:
[0231] one or more processors;
[0232] a memory; and
[0233] a plurality of instructions stored in the memory and executable by the one or more processors to:
[0234] calculate an F-statistic based on a subset of samples of a plurality of samples, wherein each sample of the plurality of samples has a signal level associated with the protein-binding reagent; and
[0235] if the calculated F-statistic is less than a critical value:
[0236] select a batch of samples of the plurality of samples to add to the subset of samples, wherein the samples of the batch are selected such that each sample of the batch has a higher signal level than any sample already in the subset;
[0237] add the batch of samples to the subset of samples; and
[0238] re-calculate the F-statistic based on the subset of samples;
[0239] if the calculated F-statistic is equal to or greater than the critical value:
[0240] estimate the matrix background of the protein-binding reagent in the assay to be equal to a highest signal level of the subset of samples;
[0241] determine a percentage of the plurality of samples having signal levels greater than the estimated matrix background of the protein-binding reagent in the assay; and
[0242] determine a specificity of the protein-binding reagent based on the determined percentage.
[0243] D2. The data processing system of paragraph D1 , wherein the signal levels of the plurality of samples are associated with a first assay platform, the method furtherKoiitch Romano Dascenzo Gates LLP Attorney Docket No. SML317PCT
[0244] comprising predicting a correlation for the protein-binding reagent between the first assay platform and a second assay platform based on the estimated matrix background of the protein-binding reagent.
[0245] D3. The data processing system of paragraph D2, wherein predicting the correlation is further based on the determined percentage of the plurality of samples having signal levels greater than the estimated matrix background.
[0246] D4. The data processing system of any one of paragraphs D1-D3, wherein calculating the F-statistic comprises determining a first variance of the signal levels of the subset of samples, determining a second variance of signals associated with nonspecific binding to the protein-binding reagent, and dividing the first variance by the second variance.
[0247] D5. The data processing system of paragraph D4, wherein the signals associated with nonspecific binding to the protein-binding reagent comprise signals associated with quality control replicates.
[0248] D6. The data processing system of any one of paragraphs D1-D5, further comprising obtaining the signal levels of the plurality of samples by measuring the signal levels using the assay.
[0249] D7. The data processing system of any one of paragraphs D1-D6, wherein the protein-binding reagent is selected from the following: an aptamer, an antibody.
[0250] E1. A data processing system for estimating a matrix background of a proteinbinding reagent in an assay, the system comprising:
[0251] one or more processors;
[0252] a memory; and
[0253] a plurality of instructions stored in the memory and executable by the one or more processors to:Koiitch Romano Dascenzo Gates LLP Attorney Docket No. SML317PCT
[0254] determine a first F-statistic based on a first subset of samples of a plurality of samples, wherein each sample of the plurality of samples has a signal level associated with the protein-binding reagent;
[0255] determine whether the first F-statistic is less than a first critical value; in response to determining that the first F-statistic is less than the first critical value, determine a second F-statistic based on a second subset of samples of the plurality of samples, wherein the second subset of samples includes the first subset of samples and a batch of one or more samples that are not included in the first subset of samples, and each signal level of the batch of one or more samples is higher than a highest signal level of the first subset of samples;
[0256] determine whether the second F-statistic is at least a second critical value; in response to determining that the second F-statistic is at least the second critical value, estimate the matrix background of the protein-binding reagent based on the signal levels of the batch of one or more samples that are not included in the first subset of samples; and
[0257] determine a percentage of the plurality of samples having signal levels greater than the matrix background.
[0258] E2. The data processing system of paragraph E1 , wherein the signal levels of the plurality of samples are associated with a first assay platform, and the plurality of instructions are further executable by the one or more processors to predict a correlation for the protein-binding reagent between the first assay platform and a second assay platform based on the estimated matrix background of the protein-binding reagent.
[0259] E3. The data processing system of paragraph E2, wherein predicting the correlation is further based on the determined percentage of the plurality of samples having signal levels greater than the estimated matrix background.
[0260] E4. The data processing system of any one of paragraphs E1-E3, wherein determining the first F-statistic comprises determining a first variance of the signal levels of the first subset of samples, determining a second variance of signals associated withKoiitch Romano Dascenzo Gates LLP Attorney Docket No. SML317PCT
[0261] nonspecific binding to the protein-binding reagent, and dividing the first variance by the second variance.
[0262] F1. A data processing system for estimating a detection limit of a proteinbinding reagent in an assay, the system comprising:
[0263] one or more processors;
[0264] a memory; and
[0265] a plurality of instructions stored in the memory and executable by the one or more processors to:
[0266] perform one or more iterations, each iteration comprising:
[0267] determining an F-statistic based on a subset of samples of a plurality of samples, wherein each sample of the plurality of samples has a signal level associated with the protein-binding reagent; and
[0268] if the F-statistic is less than a critical value associated with the current iteration, expanding the subset of samples by adding a batch of samples to the subset, wherein each sample of the batch has a higher signal level than any sample that was in the subset before the adding of the batch, and continuing to another iteration;
[0269] if the F-statistic is equal to or greater than the critical value associated with the current iteration, terminating the iterations; and in response to terminating the iterations:
[0270] estimate the detection limit of the protein-binding reagent in the assay based on a signal level of at least one sample of the batch of samples added to the subset on the last iteration before the terminating of the iterations; and
[0271] determine a percentage of the plurality of samples having signal levels greater than the detection limit.
[0272] F2. The data processing system of paragraph F1, wherein estimating the detection limit based on the signal level of at least one sample of the batch of samplesKoiitch Romano Dascenzo Gates LLP Attorney Docket No. SML317PCT
[0273] added to the subset on the last iteration includes estimating the detection limit to be equal to a highest signal level of the batch of samples added to the subset on the last iteration.
[0274] F3. The data processing system of paragraph F1, wherein estimating the detection limit based on the signal level of at least one sample of the batch of samples added to the subset on the last iteration includes estimating the detection limit to be equal to an average of one or more signal levels of the batch of samples added to the subset on the last iteration.
[0275] F4. The data processing system of any one of paragraphs F1-F3, wherein the plurality of instructions are further executable by the one or more processors to estimate a specificity of the protein-binding reagent based on the determined percentage of the plurality of samples having signal levels greater than the detection limit.
[0276] F5. The data processing system of any one of paragraphs F1-F4, wherein determining the F-statistic at each iteration comprises determining a first variance of the signal levels of the subset of samples, determining a second variance of signals associated with nonspecific binding to the protein-binding reagent, and dividing the first variance by the second variance.
[0277] F6. The data processing system of any one of paragraphs F1-F5, wherein the batch of samples added at each iteration comprises an equal number of samples at each iteration.
[0278] F7. The data processing system of any one of paragraphs F1-F6, wherein the signal levels of the plurality of samples are associated with a first assay platform, and the plurality of instructions are further executable by the one or more processors to predict a correlation for the protein-binding reagent between the first assay platform and a second assay platform based on the estimated detection limit of the protein-binding reagent in the assay.Koiitch Romano Dascenzo Gates LLP Attorney Docket No. SML317PCT
[0279] F8. The data processing system of paragraph F7, wherein predicting the correlation is further based on the percentage of the plurality of samples having signal levels greater than the detection limit.
[0280] G1. A computer-implemented method for estimating a matrix background of a protein-binding reagent in an assay, the method comprising:
[0281] determining a first F-statistic based on a first subset of samples of a plurality of samples, wherein each sample of the plurality of samples has a signal level associated with the protein-binding reagent;
[0282] determining whether the first F-statistic is less than a first critical value;
[0283] in response to determining that the first F-statistic is less than the first critical value, determining a second F-statistic based on a second subset of samples of the plurality of samples, wherein the second subset of samples includes the first subset of samples and a batch of one or more samples that are not included in the first subset of samples, and each signal level of the batch of one or more samples is higher than a highest signal level of the first subset of samples;
[0284] determining whether the second F-statistic is at least a second critical value; in response to determining that the second F-statistic is at least the second critical value, estimating the matrix background of the protein-binding reagent based on the signal levels of at least one sample of the batch of one or more samples that are not included in the first subset; and
[0285] determining a percentage of the plurality of samples having signal levels greater than the estimated matrix background.
[0286] G2. The method of paragraph G1, wherein estimating the matrix background based on the signal levels of at least one sample of the batch comprises:
[0287] estimating the matrix background to be equal to a highest signal level of the signal levels of the batch.
[0288] G3. The method of any one of paragraphs G1-G2, wherein the signal levels of the plurality of samples are associated with a first assay platform, the method furtherKoiitch Romano Dascenzo Gates LLP Attorney Docket No. SML317PCT
[0289] comprising predicting a correlation for the protein-binding reagent between the first assay platform and a second assay platform based on the estimated matrix background of the protein-binding reagent.
[0290] G4. The method of paragraph G3, wherein predicting the correlation is further based on the determined percentage of the plurality of samples having signal levels greater than the estimated matrix background.
[0291] G5. The method of any one of paragraphs G1 -G4, wherein determining the first F-statistic comprises determining a first variance of the signal levels of the first subset of samples, determining a second variance of signals associated with nonspecific binding to the protein-binding reagent, and dividing the first variance by the second variance.
[0292] G6. The method of paragraph G5, wherein the signals associated with nonspecific binding to the protein-binding reagent comprise signals associated with quality control replicates.
[0293] G7. The method of any one of paragraphs G1 -G6, further comprising obtaining the signal levels of the plurality of samples by measuring the signal levels using the assay.
[0294] G8. The method of any one of paragraphs G1-G7, wherein the protein-binding reagent is selected from the following: an aptamer, an antibody.
[0295] Conclusion
[0296] The disclosure set forth above may encompass multiple distinct examples with independent utility. Although each of these has been disclosed in its preferred form(s), the specific embodiments thereof as disclosed and illustrated herein are not to be considered in a limiting sense, because numerous variations are possible. To the extent that section headings are used within this disclosure, such headings are for organizational purposes only. The subject matter of the disclosure includes all novel and nonobvious combinations and subcombinations of the various elements, features, functions, and / orKoiitch Romano Dascenzo Gates LLP Attorney Docket No. SML317PCT
[0297] properties disclosed herein. The following claims particularly point out certain combinations and subcombinations regarded as novel and nonobvious. Other combinations and subcombinations of features, functions, elements, and / or properties may be claimed in applications claiming priority from this or a related application. Such claims, whether broader, narrower, equal, or different in scope to the original claims, also are regarded as included within the subject matter of the present disclosure.
Claims
Koiitch Romano Dascenzo Gates LLP Attorney Docket No. SML317PCTCLAIMS1. A computer-implemented method for estimating a matrix background of a protein-binding reagent in an assay, the method comprising:determining a first F-statistic based on a first subset of samples of a plurality of samples, wherein each sample of the plurality of samples has a signal level associated with the protein-binding reagent;determining whether the first F-statistic is less than a first critical value;in response to determining that the first F-statistic is less than the first critical value, determining a second F-statistic based on a second subset of samples of the plurality of samples, wherein the second subset of samples includes the first subset of samples and a batch of one or more samples that are not included in the first subset of samples, and each signal level of the batch of one or more samples is higher than a highest signal level of the first subset of samples;determining whether the second F-statistic is at least a second critical value; and in response to determining that the second F-statistic is at least the second critical value, estimating the matrix background of the protein-binding reagent based on the signal levels of at least one sample of the batch of one or more samples that are not included in the first subset.
2. The method of claim 1 , wherein estimating the matrix background based on the signal levels of at least one sample of the batch comprises:estimating the matrix background to be equal to a highest signal level of the signal levels of the batch.
3. The method of claim 1 , wherein the signal levels of the plurality of samples are associated with a first assay platform, the method further comprising predicting a correlation for the protein-binding reagent between the first assay platform and a second assay platform based on the estimated matrix background of the protein-binding reagent.
4. The method of claim 3, further comprising determining a percentage of the plurality of samples having signal levels greater than the estimated matrix background;Koiitch Romano Dascenzo Gates LLP Attorney Docket No. SML317PCTwherein predicting the correlation is further based on the determined percentage of the plurality of samples having signal levels greater than the estimated matrix background.
5. The method of claim 1, wherein determining the first F-statistic comprises determining a first variance of the signal levels of the first subset of samples, determining a second variance of signals associated with nonspecific binding to the protein-binding reagent, and dividing the first variance by the second variance.
6. The method of claim 5, wherein the signals associated with nonspecific binding to the protein-binding reagent comprise signals associated with quality control replicates.
7. The method of claim 1 , wherein the protein-binding reagent is selected from the following: an aptamer, an antibody.
8. The method of claim 1 , further comprising obtaining the signal levels of the plurality of samples by measuring the signal levels using the assay.
9. The method of claim 8, wherein measuring the signal levels using the assay comprises performing the assay with a first set of assay parameters, and wherein the estimated matrix background is an estimated first matrix background, the method further comprising:adjusting at least one parameter of the first set of assay parameters to yield a second set of assay parameters;performing the assay with the second set of assay parameters;estimating a second matrix background of the protein-binding reagent based on the assay performed with the second set of assay parameters; anddetermining whether the estimated first matrix background of the protein-binding reagent is lower than the estimated second matrix background of the protein-binding reagent.Koiitch Romano Dascenzo Gates LLP Attorney Docket No. SML317PCT10. The method of claim 9, further comprising, in response to determining that the estimated second matrix background of the protein-binding reagent is lower than the estimated first matrix background:inferring that the second set of assay parameters results in improved assay performance; anddetermining an abundance of a target analyte associated with the protein-binding reagent by performing an assay with the second set of assay parameters.
11. The method of claim 1 , further comprising:identifying the protein-binding reagent as a reagent of interest using at least one of: a SELEX method or a computational method;obtaining the signal levels of the plurality of samples associated with the proteinbinding reagent by measuring the signal levels using the assay; andin response to the signal levels not being greater than the estimated matrix background for at least a threshold percentage of samples of the plurality samples, inferring that the protein-binding reagent is unsuitable for use.
12. The method of claim 1 , further comprising:identifying the protein-binding reagent as a reagent of interest using at least one of: a SELEX method or a computational method;obtaining the signal levels of the plurality of samples associated with the proteinbinding reagent by measuring the signal levels using the assay; andin response to the signal levels being greater than the estimated matrix background for at least a threshold percentage of samples of the plurality of samples, inferring that the protein-binding reagent is suitable for use.
13. A data processing system for estimating a detection limit of a proteinbinding reagent in an assay, the system comprising:one or more processors;a memory; andKoiitch Romano Dascenzo Gates LLP Attorney Docket No. SML317PCTa plurality of instructions stored in the memory and executable by the one or more processors to:perform one or more iterations, each iteration comprising:determining an F-statistic based on a subset of samples of a plurality of samples, wherein each sample of the plurality of samples has a signal level associated with the protein-binding reagent; andif the F-statistic is less than a critical value associated with the current iteration, expanding the subset of samples by adding a batch of samples to the subset, wherein each sample of the batch has a higher signal level than any sample that was in the subset before the adding of the batch, and continuing to another iteration;if the F-statistic is equal to or greater than the critical value associated with the current iteration, terminating the iterations; and in response to terminating the iterations:estimate the detection limit of the protein-binding reagent in the assay based on a signal level of at least one sample of the batch of samples added to the subset on the last iteration before the terminating of the iterations; anddetermine a percentage of the plurality of samples having signal levels greater than the detection limit.
14. The data processing system of claim 13, wherein estimating the detection limit based on the signal level of at least one sample of the batch of samples added to the subset on the last iteration includes estimating the detection limit to be equal to a highest signal level of the batch of samples added to the subset on the last iteration.
15. The data processing system of claim 13, wherein estimating the detection limit based on the signal level of at least one sample of the batch of samples added to the subset on the last iteration includes estimating the detection limit to be equal to an average of one or more signal levels of the batch of samples added to the subset on the last iteration.Koiitch Romano Dascenzo Gates LLP Attorney Docket No. SML317PCT16. The data processing system of claim 13, wherein the plurality of instructions are further executable by the one or more processors to estimate a specificity of the protein-binding reagent based on the determined percentage of the plurality of samples having signal levels greater than the detection limit.
17. The data processing system of claim 13, wherein determining the F-statistic at each iteration comprises determining a first variance of the signal levels of the subset of samples, determining a second variance of signals associated with nonspecific binding to the protein-binding reagent, and dividing the first variance by the second variance.
18. The data processing system of claim 13, wherein the batch of samples added at each iteration comprises an equal number of samples at each iteration.
19. The data processing system of claim 13, wherein the signal levels of the plurality of samples are associated with a first assay platform, and the plurality of instructions are further executable by the one or more processors to predict a correlation for the protein-binding reagent between the first assay platform and a second assay platform based on the estimated detection limit of the protein-binding reagent in the assay.
20. The data processing system of claim 19, wherein predicting the correlation is further based on the percentage of the plurality of samples having signal levels greater than the detection limit.