Methods for normalization and quantification of sequencing data
By adding internal quantitative standards in NGS sample preparation stage and normalizing the sequencing data set, the differences between laboratories and samples in NGS quantitative analysis were solved, and the absolute quantification of multiple target species was achieved, improving the accuracy and consistency of quantitative analysis.
Patent Information
- Application Number
- CN201980041439.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2018-04-20
- Filing Date
- 2019-04-18
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2039-04-18
AI Technical Summary
There are interlaboratory and intersample differences in existing NGS technologies in quantitative analysis, making it difficult to accurately distinguish pathogens from cohabitants, and the existing quantitative methods are complex and inaccurate enough to provide absolute concentration measurements.
Using internal quantitative standards (IQS), a known quantity of IQS is added during the sample preparation stage. Normalize the sequencing data set and the sequencing readings of the quantitative standards to ensure the consistent detection limits and achieve absolute quantification of unknown nucleic acids.
A universal internal standard method is provided that enables simultaneous quantification of multiple target species in NGS sequencing, reducing inter-laboratory and inter-sample differences, and improving the accuracy and consistency of quantitative analysis.
Smart Images

Figure CN112492883B_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims the benefit of and priority to U.S. Provisional Patent Application No. 62 / 660,460, filed April 20, 2018, which is incorporated herein by reference in its entirety.
[0003] background
[0004] The ability to read the base sequence that comprises polynucleotides has had a significant impact on biological research. For most of the past 40 years, dideoxy DNA 'Sanger' sequencing has been used as the standard sequencing technology in many laboratories, culminating in the completion of the human genome sequence. However, because Sanger sequencing is performed on individual amplicons, the throughput of Sanger sequencing is limited, and large-scale Sanger sequencing projects are expensive and laborious.
[0005] The paradigm of DNA sequencing changed with the advent of 'next-generation' sequencing (NGS) technology, which processes hundreds of thousands to millions of DNA fragments in parallel, resulting in low cost per base of sequence generated and throughput on the scale of gigabases (Gb) to terabases (Tb) in a single sequencing run. This is noteworthy considering that the first human genome, famously co-published in Science and Nature in 2001, took 15 years to sequence and cost nearly $3 billion. In contrast, modern NGS sequencers can sequence 45 human genomes in a single day, at approximately $1,000 per genome. Consequently, NGS can now be used to define characteristics of entire genomes and delineate differences between them, allowing researchers to gain a deeper understanding of the full spectrum of genetic variation and define its role in phenotypic variation and the pathogenesis of complex traits.
[0006] However, the application of NGS for clinical and diagnostic applications can be limited by high intra- and inter-laboratory variability. These issues reduce the value of any results and have prevented the use of these sequencing methods in molecular diagnostics. For example, the complexity and variability of NGS library preparation and sequencing reaction preparation can lead to sample-to-sample and inter-laboratory variability, which can make it difficult to determine, for example, the prevalence of genetic variations or pathogenic organisms detected in a sample.
[0007] In another context, many pathogens targeted in diagnostic assays (e.g., FilmArray panel, BioFire, SLC, UT) can be found in the environment and as commensals at the site of sample collection. For example, in diseases such as pneumonia, the most frequently encountered bacterial pathogens may also be present as "normal flora" of the oropharyngeal passage, which is itself often the site of sample collection (sputum and tracheal aspirates or nasopharyngeal swabs (NPS)), or the route for collecting more invasive specimens such as bronchoalveolar lavage (BAL). In such cases, frequent contamination by normal flora or co-collection of normal flora is essentially unavoidable. In such situations, the diagnostic power of NGS may be limited by the fact that clinically relevant organisms cannot be easily distinguished from commensal flora or contamination due to the potential for NGS to detect the presence of organisms at both high and minimal concentrations (i.e., NGS has an almost unlimited dynamic range) without providing a substantial amount of inherent context to interpret the clinical relevance of the detection in the sequencing data (e.g., NGS can detect the presence of a pathogen (i.e., nucleic acid from a pathogen) and its relative abundance (%) to other detected nucleic acids or organisms without providing any indication of whether the detected pathogen is present at clinically relevant concentrations).
[0008] Traditional practice in microbiology laboratories is to perform semiquantitative or quantitative cultures to distinguish bacterial pathogenicity burden from nonclinically relevant commensal flora carriage. Different diagnostic titer guidelines exist for different sample types. A similar approach has been applied to NGS assays. By its nature, NGS provides semiquantitative data, and in the absence of confounding factors (such as sample preparation errors or differences in sequencing efficiency), the number of sequencing reads for a target can be correlated with the target's abundance. Several groups have exploited this relationship to obtain relative quantitative data for unknown nucleic acids in NGS. For example, the relative abundance of nucleic acids in a sample can be determined by performing a series of serial dilutions (illustratively, 10-fold dilutions) on one or more samples, sequencing the series of diluted samples, and then plotting the number of reads found in each. These groups have hypothesized that if the relationship between the number of reads in the serially diluted samples is linear (e.g., a 10-fold dilution results in an approximately 10-fold decrease in the number of sequencing reads, an approximately 100-fold dilution results in a 100-fold decrease in the number of sequencing reads, etc.), then the number of sequencing reads can be used to relatively quantify different targets present in the sample (e.g., relative quantification of high and low concentrations of a target). In some embodiments, the present invention relates to a method for determining the concentration of a nucleic acid sequence. For example, if the first sequencing nucleic acid has 10 sequencing reads and the second sequencing nucleic acid has 100 sequencing reads, it can be concluded that the concentration of the second nucleic acid is 10X that of the first nucleic acid. This can be used for example to detect gene duplication and / or determine the gene copy number in the genome. However, this method is only relative, and therefore, due to the absence of nucleic acids with known absolute concentrations that can be used as a reference, the concentration of the first or second nucleic acid cannot be determined. As a result, this method may not be very accurate. For example, in the case of large differences (for example, several orders of magnitude) between high-concentration detection and low-concentration detection, resolution may be lost at lower and / or high concentrations, resulting in a reduction in the multiple difference between high and low compared to the true difference. This method is also not very specific due to the following fact: it is only relative, it is sample / sequencing run specific, and it does not illustrate intra-laboratory differences and inter-laboratory differences.
[0009] Another common quantification method is to quantify nucleic acids in samples for NGS in separate reactions. For example, quantitative PCR (qPCR) can be used for absolute quantification, frequently employing the standard curve approach. In this approach, a standard curve generated by plotting the crossing point (Cp) values obtained from real-time PCR against a known quantity of a single reference template provides a regression line that can be used to infer the quantity of the same target gene in the sample of interest. Along with samples containing the specific gene target to be quantified, a serial dilution of the reference template (illustratively, 10-fold dilutions) is established. Various separate reactions are run, typically once for each level of the reference target and once for each target sample. Additionally, because assay-specific differences in PCR efficiency often affect quantification, separate standard curves with separate reference templates can be established for different gene targets.
[0010] However, the power of NGS lies in its massive parallelism - that is, samples from 10s to 100s to 1000s can be processed simultaneously and in parallel. In this situation, quantification of targets using qPCR can be challenging. Although quantification of targets from 100s to 1000s of separate nucleic acid reactions has been performed using qPCR (see, for example, High-Throughput Droplet Digital PCR System for Absolute Quantitation of DNA Copy Number, Hindson et al., Anal Chem. 2011 Nov 15;83(22):8604–8610), this approach is technically challenging and requires specialized equipment. In addition, qPCR methods generally assume or require that the assay has the same PCR efficiency in singleplex and multiplex reactions, which is not the case. In addition, all standard curve-based quantification methods published to date require the establishment of external reactions and the calculation of standard curves.
[0011] Another approach is to use assay-specific competitive templates to quantify nucleic acids in NGS (see, e.g., US 2015 / 0292001). This approach aims to provide reproducible measurement of nucleic acid copy number in a sample by relying on the proportional relationship between the native target sequence and a separate competitive internal amplification control (the control is specifically designed for that native target sequence). The competitive templates described in US 2015 / 0292001 utilize priming sites identical to those of the target native nucleic acid template, but employ designed (e.g., artificial) interprimer sequences to mimic the kinetics of the native target in the PCR reaction and thereby control for target-specific variation in PCR efficiency. However, as a result, this approach is assay- and target-specific (i.e., the competitive template is target- and sample-specific). To employ the methods described in US 2015 / 0292001, a new competitive internal amplification control must be designed for each new assay and / or template to be sequenced, which limits the general applicability of this approach. Additionally, the target typically needs to be sequenced both with and without the competing template in order to deconvolute the sequencing response of the target alone from the sequencing response of the target plus the competing template. This increases the level of complexity of the method and has the potential to introduce errors into the calculations.
[0012] Therefore, there is a need in the art for universal internal quantitative standards and related quantitative methods for NGS. Since nucleic acid purification from patient samples is integrated into the NGS workflow, the effect of sample-driven variability in nucleic acid extraction, as well as the effect of any sample-derived inhibitors on PCR, and therefore on quantification, cannot be easily estimated by external standard curves.
[0013] Overview
[0014] The present disclosure provides methods, systems, and kits for universal internal standards that can provide simultaneous quantification of multiple target species while also accounting for the effects of assay-specific and matrix-derived variations in sequencing results. Because the standard is universal, it is not specific to the target, sample, or assay, and as such, the standards, methods, and kits described herein can be used in any sequencing assay (e.g., an NGS assay). The present disclosure also teaches the use of process controls and / or limit of detection (LOD) controls for assay-specific correction.
[0015] In one aspect of the present disclosure, a method for normalizing read values for a sequencing assay is disclosed. The method includes providing a sample comprising one or more unknown nucleic acids to be sequenced; adding a known number of an internal quantitative standard (IQS) to the sample; preparing the sample comprising the internal quantitative standard for sequencing; sequencing to generate a sequencing dataset for the sample, wherein the sequencing dataset comprises sequencing reads observed from the unknown nucleic acid and the internal quantitative standard; counting the number of sequencing reads in the sequencing dataset that originate from the unknown nucleic acid and the internal quantitative standard; and normalizing the sequencing dataset, wherein the normalization (1) applies data acceptance / rejection criteria to the sequencing dataset based on the presence of a minimum number of sequencing reads for the internal quantitative standard for the sample (e.g., retaining the normalizing number of sequencing reads), and (2) ensures that substantially the same limit of detection (LOD) is applied to all unknown nucleic acids in the sequencing assay. In one embodiment, preparing the sample comprising the internal quantitative standard for sequencing may include: introducing sequencing primer binding sites and sample-specific identification sequences into the unknown nucleic acids and the internal quantitative standard in the sample. In another embodiment, preparing a sample for sequencing containing an internal quantitative standard can include one or more of the following: lysing cells in the sample, recovering nucleic acids from the lysate, nucleic acid purification, a first multiplex PCR using target-specific primers with overhangs, a second nucleic acid purification, a second multiplex PCR using sample-specific sequencing adapter primers, a third nucleic acid purification, and pooling multiple similarly prepared samples for sequencing. Because the standard is added at the beginning of sample preparation, the standard is used throughout all steps, and system losses and efficiencies are accounted for for all steps.
[0016] In one embodiment, each unknown nucleic acid in the sample may have about 0-10 13 copies / ml concentration – that is, in some cases, the unknown nucleic acid may not be present (0 copies / ml present), or the concentration of the unknown may be very high (e.g., up to 10 13 In typical samples, the concentration of unknown nucleic acids can range from about 10 3 -10 9 In one embodiment, the known amount of IQS added to the sample is about 10 4 -10 6 copies / ml (e.g., approximately 5x10 5The concentration of IQS can be within the range of 100 copies / ml. More or less IQS can be added to the sample based on or to achieve the limit of detection (LOD). In one embodiment, a single type of IQS can be added to the sample. In another embodiment, two or more types of IQS can be added to the sample at different input concentrations, allowing a standard curve to be generated using two or more reference points for quantification.
[0017] In one embodiment, there may preferably be a linear relationship between the amount of internal quantitative standard added to the sample and the number of sequencing reads for the internal quantitative standard. That is, the number of sequencing reads attributed to the internal quantitative standard may not equal the input quantity of the internal quantitative standard, but the number of reads should be linearly related to the input quantity. If the relationship is not linear, this can be interpreted as indicating a problem in one or more of the addition of the internal quantitative standard, sample preparation, or sequencing.
[0018] In one embodiment, normalization retains the normalized number of sequencing reads (NORM), where NORM = number of sequencing reads * (F / observed number of sequencing reads derived from the internal quantitative standard). 'F' is a fixed user-set normalization coefficient that defines the minimum expected number of IQS sequencing reads. In one embodiment, the value of F can be assay-specific. 'F' is a sequencing read value set to ensure a sequencing read depth sufficient for the limit of detection (LOD) of the sequencing assay. NORM may represent a subset of the sequencing data and is the number of sequencing reads saved for further analysis and quantification. For example, if 'F' and the observed number of sequencing reads derived from the internal quantitative standard are the same, NORM is equal to the recorded number of sequencing reads. On the other hand, if the number of sequencing reads of the internal quantitative standard in the dataset is greater than 'F', NORM reduces the data size to address the problem of over-reading and ensures that the same LOD is applied across all samples. If the number of sequencing reads of the internal quantitative standard in the dataset is less than 'F', the data from that sample may be rejected. In one embodiment, the data can be normalized using the relationship ALPHA, where the unknown nucleic acid reads and the internal quantification standard reads are each normalized by the same ratio ALPHA, where ALPHA = F / observed number of sequencing reads derived from the internal quantification standard.
[0019] In one embodiment, the sequencing read depth is in the range of about 1000 internal quantification standard sequencing reads to about 100,000 internal quantification standard sequencing reads, preferably about 2000 internal quantification standard sequencing reads to about 75,000 internal quantification standard sequencing reads, more preferably about 5000 internal quantification standard sequencing reads to about 50,000 internal quantification standard sequencing reads, or most preferably at least 5000 internal quantification standard sequencing reads. If sequencing data with less than 'F' is recorded for the internal quantification standard, data for samples associated with insufficient numbers of 'F' reads may be rejected.
[0020] In one embodiment, the method may include calculating an input quantity (IQT) of unknown nucleic acids in the sample after normalization, wherein since the input quantity of the internal quantification standard is known, the input quantity (IQT) of the unknown nucleic acids may be calculated by IQT = normalized number of unknown nucleic acid sequencing reads attributed to the unknown nucleic acid * (input quantity of the internal quantification standard / F). In one embodiment, the method may include calculating the input quantity of two or more unknown nucleic acids in the sample by calculating the input quantity (IQTi, IQTj, IQTk ... IQTn) of a plurality of unknown nucleic acids (if any) in the sample after normalization as IQTn = normalized number of unknown "n" nucleic acid sequencing reads attributed to the "nth" unknown nucleic acid * (input quantity of the internal quantification standard / F).
[0021] In one embodiment, two or more samples can be merged before being subjected to order-checking. In one embodiment, each sample can be prepared for order-checking after and before actual order-checking and merge samples. In one embodiment, 2-1000 samples can be merged after preparation and before order-checking, preferably 2-500 samples can be merged after preparation and before order-checking, more preferably 2-100 samples can be merged after preparation and before order-checking, more preferably 2-50 samples can be merged after preparation and before order-checking, or most preferably 2-32 samples can be merged after preparation and before order-checking. The number of samples that can be merged for order-checking is generally only subject to the ability limit of distinguishing the sample in the obtained data. For example, the sample in the storehouse can be distinguished by sample-specific order-checking adapter primers, and the primers are used to identify the order-checking data that comes from specific samples. In this embodiment, the distinction may be subject to the length and the diversity restriction of the order-checking adapter primers.
[0022] In one embodiment, each pooled sample has associated therewith a unique set of sample-specific identification sequences, such that sequencing data from each sample in the pool can be distinguished and separated. In one embodiment, each pooled sample has its own internal quantitative standard associated with its own unique set of sample-specific identification sequences, and wherein normalization is applied separately to each sample in the pool. In one embodiment, normalization separately (1) applies data acceptance / rejection criteria to each sample in the pool, and (2) ensures that substantially the same limit of detection (LOD) is applied to all unknown nucleic acids in each sample.
[0023] In one embodiment, the method can include calculating the input quantity (IQT) of unknown nucleic acids in each sample after the data attributed to each sample has been normalized. Because the input quantity of the internal quantification standard in each sample is known, the input quantity (IQT) of unknown nucleic acids can be calculated. After normalization as IQTn = normalized unknown "n" nucleic acid sequencing reads attributed to the "nth" unknown nucleic acid * (input quantity of internal quantification standard / F), the input quantities (IQTi, IQTj, IQTk ... IQTn) of the multiple unknown nucleic acids in the combined sample can be calculated.
[0024] In one embodiment, preparing a sample for sequencing can include lysing the sample to generate a lysate, recovering nucleic acid from the lysate, and optionally purifying the recovered nucleic acid, and introducing a primer binding site and a sample-specific identification sequence into a region of the nucleic acid to be sequenced. Attaching can include one of amplifying the nucleic acid to be sequenced in an amplification reaction using a target-specific primer having a dual-index sequencing overhang, the target-specific primer comprising a sequencing primer binding site and a sample-specific identification sequence, or fragmenting the nucleic acid to be sequenced and ligating to the fragmented nucleic acid sequencing-specific adapter, the adapter comprising a sequencing primer binding site and a sample-specific identification sequence.
[0025] In one embodiment, amplifying nucleic acids to be sequenced can include performing a first multiplex PCR reaction using target-specific primers having custom overhangs, performing a first nucleic acid purification, performing a second PCR reaction using dual-index sequencing adapter primers that anneal or ligate to the custom overhangs introduced in the first PCR, and performing a second nucleic acid purification. In one embodiment, the dual-index sequencing adapter primers are target-independent and comprise a sequencing primer binding site and a sample-specific identification sequence.
[0026] In one embodiment, amplification can be performed to limit or compress the upper limit of the dynamic range of nucleic acid concentrations in the sample to be sequenced. This can reduce sequencing and data analysis load and reduce the number of situations in which only very high concentrations of nucleic acids appear in the sequencing data set. In one embodiment, amplification can include a platform amplification that limits one or more of the target-specific primer concentration or cycle number in the first multiplex PCR reaction to nucleic acids present at a concentration greater than the desired dynamic range. As an example, compression can be performed to compress the dynamic range of nucleic acids to about 10 7 The platform concentration is less than 10 copies / ml. 7 Nucleic acids present at 10 copies / ml are present at a range of concentrations during the exponential amplification phase. In other embodiments, amplification can be performed to limit or compress other portions of the dynamic range. If, for example, only high concentration species are of interest in a sequencing assay, amplification can be limited so that only nucleic acids at higher concentrations (e.g., >10 5 copies / ml).
[0027] In any of the foregoing method embodiments, the sequencing assay can be a next generation sequencing assay.
[0028] In another aspect, a method for performing a quantitative next generation sequencing (NGS) assay is disclosed. The method includes providing a sample comprising one or more unknown nucleic acids to be sequenced, wherein the unknown nucleic acids have a relative abundance of about 0-10 13 A known amount of IQS is added to the sample, wherein the known amount of the IQS is about 10 4 -10 6 copies / ml (depending on the targeted dynamic range of the assay); preparing samples for sequencing that include an internal quantification standard; sequencing the unknown nucleic acid and the internal quantification standard in the sample to generate sequencing data; counting the number of sequencing reads originating from the unknown nucleic acid and the internal quantification standard in the sequencing data set; and normalizing the sequencing data set and calculating the input number of unknown nucleic acids (IQT) by IQT = normalized number of unknown nucleic acid sequencing reads originating from the unknown nucleic acid * (input number of internal quantification standard / F), where 'F' is the minimum expected number of sequencing reads of the fixed internal quantification standard.
[0029] In one embodiment, normalization can include retaining a normalized number of sequencing reads (NORM), where NORM = number of sequencing reads * (F / observed number of sequencing reads derived from an internal quantification standard). Normalization separately (1) applies data acceptance / rejection criteria to each sample in the assay, and (2) ensures that substantially the same limit of detection (LOD) is applied to all unknown nucleic acids in each sample.
[0030] In one embodiment, the unknown nucleic acid may have a 1 -10 12 copies / ml, or preferably about 10 1 -10 9 In one embodiment, 'F' is related to the LOD of the assay. For example, if for about 10 5 The IQS input concentration is 10 copies / ml, and 'F' is about 10 2 -10 3 , then the LOD for unknown nucleic acids in the sample is about 10 3 -10 2 copies / ml. If for the same IQS input concentration, 'F' is about 10 3 -10 4 (i.e., the read depth increases and the weight of each read increases accordingly), the LOD for unknown nucleic acids in the sample is approximately 10 2 -10 1 For a given IQS input concentration, the LOD can be raised or lowered by increasing or decreasing the read depth (i.e., by increasing or decreasing the degree of 'F'), and correspondingly increasing or decreasing the weight attributed to each individual sequencing read.
[0031] In one embodiment, a method for performing a quantitative next generation sequencing (NGS) assay may include calculating an input quantity (IQTi, IQTj, IQTk ... IQTn) of a plurality of unknown nucleic acids (if present) in a sample as IQTn = normalized unknown "n" nucleic acid sequencing reads attributed to the "nth" unknown nucleic acid * (input quantity of internal quantitation standard / F).
[0032] In one embodiment, a method for performing a quantitative next-generation sequencing (NGS) assay can include pooling two or more samples and subjecting them to sequencing simultaneously, wherein the two or more samples are pooled after preparation and prior to sequencing. In one embodiment, each pooled sample can have a unique set of sample-specific identification sequences associated with it, allowing sequencing data from each sample in the pool to be distinguished and separated. In one embodiment, each pooled sample can have its own internal quantification standard associated with its own unique set of sample-specific identification sequences, and quantification can be applied separately to each nucleic acid from each sample in the pool. As described above for multiple nucleic acids in a single sample, normalization and quantification can be applied to multiple nucleic acids in multiple samples. The method includes calculating the input quantity (IQTi, IQTj, IQTk…IQTn) of the multiple unknown nucleic acids in the pooled sample as IQTn = normalized number of unknown "n" nucleic acid sequencing reads attributed to the "nth" unknown nucleic acid * (input number of internal quantification standards / F).
[0033] In one embodiment, a method for performing a quantitative next-generation sequencing (NGS) assay may include providing a set of assay-specific positive controls to be sequenced, wherein the positive controls include a positive control corresponding to each of one or more unknown nucleic acids sequenced in the assay. In one embodiment, the method for performing a quantitative next-generation sequencing (NGS) assay may include applying an assay-specific correction factor to each of the one or more unknown nucleic acids based on the sequence read counts of the assay-specific positive controls. If the efficiency of PCR amplification is suboptimal or if the efficiency of the target and internal standard differs, the correction factor may depend on the number of positive control inputs within the reaction. Therefore, in one embodiment, the number of positive controls should be in the middle of the targeted dynamic range of the corresponding target.
[0034] In another aspect, a kit for normalizing and quantifying unknown nucleic acids in a next generation sequencing (NGS) assay is described. The kit can include an internal quantification standard (IQS), wherein the IQS is a nucleic acid configured to be added in known amounts to a sample containing unknown nucleic acids to be sequenced, and instructions for using the IQS for normalizing a sequencing data set and for calculating the input quantity of unknown nucleic acids. In one embodiment, the kit can include a set of IQS to be added at different known concentrations for generating a standard curve for quantifying unknown nucleic acids. In one embodiment, the IQS provided in the kit can be configured to be added in a known amount to about 10 4 -10 6A range of copies / ml can be added to the sample; however, the number of one or more IQSs can be increased or decreased to expand or compress the upper and lower limits of the detection range. In one embodiment, the sequencing data is normalized using an internal quantification standard to retain the normalized number of sequencing reads (NORM), where NORM = number of sequencing reads * (F / observed number of sequencing reads derived from the internal quantification standard), where 'F' is the minimum expected number of fixed internal quantification standard sequencing reads, and the input number (IQT) of the unknown nucleic acid is calculated using the internal quantification standard by IQT = normalized number of unknown nucleic acid sequencing reads derived from the unknown nucleic acid * (input number of internal quantification standard / F).
[0035] In one embodiment, the kit further comprises a sequencing-specific adapter for at least an internal quantitative standard comprising a sequencing adapter site and a sample-specific identification sequence. In one embodiment, the kit may further comprise a target-specific primer having a custom overhang configured for amplifying the internal quantitative standard and for annealing or ligating the sequencing-specific adapter.
[0036] In another aspect, a method for performing a comparator study is described. The method for performing a comparator study may include providing a first assay comprising single or multiplex amplification and detection of one or more target nucleic acids, the first assay having a limit of detection (LOD), and providing a second assay, different from the first assay, for confirming the detection and the LOD of the first assay. The second assay should have at least the same LOD as the first assay, but it may have a lower LOD. In one or more embodiments, the second assay may include preparing a sample for sequencing comprising at least one internal quantitative standard; sequencing to generate a sequencing data set for the sample, wherein the sequencing data set comprises sequencing reads observed from the target nucleic acid and the internal quantitative standard; counting the number of sequencing reads in the sequencing data set that originate from the target nucleic acid and the internal quantitative standard; and normalizing the sequencing data set, wherein the normalization (1) applies data acceptance / rejection criteria to the sequencing data set based on the presence of a minimum number of sequencing reads for the internal quantitative standard for the sample, and (2) ensures that substantially the same limit of detection (LOD) is applied to all unknown nucleic acids in the sequencing assay, and wherein the LOD of the second assay is substantially the same as the LOD of the first assay.
[0037] In one embodiment, the first assay can be a qualitative molecular diagnostic assay. In another embodiment, the first assay can be a semi-quantitative molecular diagnostic assay. In one embodiment, the first assay can further include adding one or more internal quantitative standards to the sample and performing a quantitative two-step amplification. In one embodiment, quantitative two-step amplification can include: amplifying the sample in a first-stage multiplex amplification mixture, the amplification mixture comprising a plurality of target primers, each target primer pair configured to amplify a different target that may be present in the sample, and at least one quantitative standard primer pair, the quantitative standard primer pair configured to amplify an internal quantitative standard nucleic acid, dividing the first-stage amplification mixture into a plurality of second-stage individual reactions, a first plurality of second-stage individual reactions each comprising at least one primer pair, the primer pair configured to further amplify one of the different targets that may be present in the sample, and a second plurality of second-stage individual reactions each comprising at least one primer pair configured to further amplify one of the internal quantitative standard nucleic acids, and subjecting the plurality of second-stage individual reactions to amplification conditions to generate one or more target amplicons and a plurality of quantitative standard amplicons, each quantitative standard amplicon having an associated quantitative standard crossing point (Cp), wherein each target nucleic acid has a Cp and each internal standard has a known concentration and a known quantitative standard Cp in the first assay.
[0038] In one embodiment, the method can further comprise: generating a standard curve from two or more quantification standard crossing points (Cp) in the first assay; and using the standard curve to quantify each of the one or more target nucleic acids. In one embodiment, each target nucleic acid can be quantified in the first assay using a standard curve generated using a least squares regression line fit to:
[0039] log 10 (Concentration) = (Cp - b) / a
[0040] where Cp is the intersection point measured for each target, b , intercept, represents the log of the target 10 The Cp value when (concentration) is zero, and a is the slope that represents how much Cp changes with a single unit change in concentration.
[0041] In one embodiment, the second assay can further include normalizing the sequencing dataset to retain the NORM sequencing reads, where NORM = number of sequencing reads * (F / observed number of sequencing reads derived from the internal quantification standard). In one embodiment, the nucleic acids in the second assay can be quantified by calculating the input quantity (IQT) for each target nucleic acid in the sample as IQT = normalized number of target nucleic acid sequencing reads derived from the target nucleic acid * (input number of internal quantification standard / F), where 'F' is the minimum expected number of fixed internal quantification standard sequencing reads. The same principle applies to calculating the input concentrations of multiple unknown nucleic acids in the sample. The input quantities (IQTi, IQTj, IQTk ... IQTn) of multiple target nucleic acids (if present) in the sample can be calculated as IQTn = normalized number of unknown "n" nucleic acid sequencing reads attributed to the "nth" unknown nucleic acid * (input number of internal quantification standard / F).
[0042] In one embodiment, the 'F' in the second assay may be related to the limit of detection (LOD) of the second assay. The LOD of the second assay is selected to be substantially the same as the LOD of the first assay. In one embodiment, 'F' is related to the LOD of the assay. For example, if for about 10 5 The IQS input concentration is 10 copies / ml, and 'F' is about 10 3 , then the LOD for unknown nucleic acids in the sample is about 10 2 copies / ml. If for the same IQS input concentration, 'F' is about 10 4 (i.e., read depth increases, and the weight of each read increases accordingly), the LOD for unknown nucleic acids in the sample is approximately 10 copies / ml. For a given IQS input concentration, the LOD can be increased or decreased by increasing or decreasing the read depth (i.e., by increasing or decreasing the degree of 'F'), and correspondingly increasing or decreasing the weight attributed to each individual sequencing read.
[0043] In one embodiment, the method for performing a comparator study can include merging two or more samples and subjecting them to sequencing simultaneously, wherein the two or more samples are merged after preparation and before sequencing. In one embodiment, each merged sample can have a set of unique sample-specific identification sequences associated therewith, so that the sequencing data from each sample in the library can be distinguished and separated. In one embodiment, each merged sample can have its own internal quantitative standard, which is associated with a set of unique sample-specific identification sequences of itself. As described above for multiple nucleic acids in a sample, normalization and quantification can be applied to the multiple nucleic acids in the multiple merged samples. The method includes calculating the input quantity (IQTi, IQTj, IQTk...IQTn) of the multiple unknown nucleic acids in the merged sample as IQTn=normalized unknown "n" nucleic acid sequencing reads attributed to the "nth" unknown nucleic acid*(input quantity / F of the internal quantitative standard).
[0044] In any of the foregoing embodiments, the sequencing assays, methods, and kits described herein do not include performing relative quantification. That is, the sequencing assays, methods, and kits described herein do include an internal quantification standard of known concentration, and thus do not necessarily rely on the relative number of sequencing reads or the relative abundance of detected nucleic acids to determine the relative concentration of unknowns.
[0045] In any of the foregoing embodiments, the sequencing assays, methods, and kits described herein do not include performing quantitation in a reaction separate from the sequencing assay. That is, the sequencing assays, methods, and kits described herein do include an internal quantitation standard of known concentration, thereby eliminating the need to perform another separate reaction (e.g., qPCR) to determine the input concentration of nucleic acid in the sequencing reaction.
[0046] In any of the foregoing embodiments, the sequencing assays, methods, and kits described herein do not include the use of assay- or template-specific quantification standards. That is, the sequencing assays, methods, and kits described herein do include universal internal standards that can provide simultaneous quantification of multiple target species in any sequencing assay (e.g., an NGS assay), rather than relying on standards designed for a specific assay or specific target.
[0047] In any of the foregoing embodiments, the sequencing assays, methods, and kits described herein do not include the use of a competitive template as a quantitative standard. Generally speaking, a competitive template is a specific type of assay or template-specific quantitative standard. However, a new competitive internal amplification control needs to be designed for each new assay and / or template to be sequenced. In contrast, the sequencing assays, methods, and kits described herein include a universal internal standard that can provide simultaneous quantification of multiple target species in any sequencing assay (e.g., an NGS assay).
[0048] This article describes:
[0049] A1. A method for normalizing read values of a sequencing assay, comprising:
[0050] Providing a sample comprising one or more unknown nucleic acids to be sequenced;
[0051] A known amount of internal quantitative standard was added to the sample;
[0052] preparing a sample comprising an internal quantitative standard for sequencing, wherein the preparing comprises introducing sequencing-specific adapter sites and sample-specific identification sequences into the unknown nucleic acids in the sample and the internal quantitative standard;
[0053] sequencing to generate a sequencing dataset for the sample, wherein the sequencing dataset includes sequencing reads observed from the unknown nucleic acid and an internal quantitative standard;
[0054] Counting the number of sequencing reads in a sequencing dataset that originate from unknown nucleic acids and internal quantification standards; and
[0055] The sequencing dataset is normalized, wherein the normalization (1) applies data acceptance / rejection criteria to the sequencing dataset based on the presence of a minimum number of sequencing reads with respect to an internal quantitative standard for a sample, and (2) ensures that substantially the same limit of detection (LOD) is applied to all unknown nucleic acids in the sequencing assay.
[0056] A2. The method of clause A1, wherein each unknown nucleic acid in the sample has about 0-10 13 copies / ml (e.g., approximately 10 2 -10 9 copies / ml) concentration.
[0057] A3. The method of at least one of clause A1 or clause A2, wherein the known amount of the internal quantitative standard added to the sample is about 10 3 -10 6 copies / ml (e.g., approximately 10 5 -10 6 copies / ml).
[0058] A4. The method of one or more of clauses A1-A3, wherein there is a linear relationship between the known amount of the internal quantification standard added to the sample and the number of sequencing reads for the internal quantification standard.
[0059] A5. The method of any one or more of clauses A1-A4, wherein the sequencing dataset is normalized to retain NORM sequencing reads, where NORM = number of sequencing reads * (F / observed number of sequencing reads derived from the internal quantification standard), and wherein 'F' is the minimum expected number of sequencing reads of the fixed internal quantification standard.
[0060] A6. The method of any one or more of clauses A1-A5, wherein 'F' is a number of sequencing reads set to ensure sufficient sequencing read depth for the LOD of the sequencing assay.
[0061] A7. The method of any one or more of clauses A1-A6, wherein the sequencing read depth is in the range of about 1000 internal quantification standard sequencing reads to about 100,000 internal quantification standard sequencing reads, preferably about 2000 internal quantification standard sequencing reads to about 75,000 internal quantification standard sequencing reads, more preferably about 5000 internal quantification standard sequencing reads to about 50,000 internal quantification standard sequencing reads, or most preferably at least 5000 internal quantification standard sequencing reads.
[0062] A8. The method of any one or more of clauses A1-A7, wherein the unknown nucleic acid reads and the internal quantification standard reads are each normalized by the same ratio ALPHA, where ALPHA = F / observed number of sequencing reads derived from the internal quantification standard.
[0063] A9. The method of any one or more of clauses A1-A8, further comprising calculating the input quantity (IQT) of the unknown nucleic acid in the sample after normalization, wherein since the input quantity of the internal quantitative standard is known, the input quantity (IQT) of the unknown nucleic acid can be calculated by IQT = normalized unknown nucleic acid sequencing read number attributed to the unknown nucleic acid * (input quantity of the internal quantitative standard / F).
[0064] A10. The method of any one or more of clauses A1-A9, further comprising: calculating the input quantity (IQTi, IQTj, IQTk…IQTn) of a plurality of unknown nucleic acids (if any) in the sample after normalization as IQTn = normalized number of unknown “n” nucleic acid sequencing reads attributed to the “nth” unknown nucleic acid * (input quantity of internal quantitative standard / F).
[0065] A11. The method of any one or more of clauses A1-A10, wherein preparing the sample comprises lysing the sample, recovering nucleic acid from the lysate, and optionally purifying the recovered nucleic acid, and attaching introduced primer binding sites and sample-specific identification sequences to regions of the nucleic acid to be sequenced.
[0066] A12. The method of any one or more of clauses A1-A11, wherein the attaching comprises one of:
[0067] Amplifying the nucleic acid to be sequenced in an amplification reaction using target-specific primers with dual-index sequencing overhangs, the target-specific primers comprising a sequencing primer binding site and a sample-specific identification sequence, or
[0068] The nucleic acid to be sequenced is fragmented and ligated to fragmented nucleic acid sequencing-specific adaptors, which contain sequencing primer binding sites and sample-specific identification sequences.
[0069] A13. The method of any one or more of clauses A1-A12, wherein amplifying the nucleic acid to be sequenced comprises:
[0070] Perform a first multiplex PCR reaction using target-specific primers with custom overhangs,
[0071] Perform the first nucleic acid purification,
[0072] performing a second PCR reaction using dual-index sequencing adapter primers that anneal or ligate to the overhangs introduced in the first PCR, wherein the dual-index sequencing adapter primers are target-independent and contain a sequencing primer binding site and a sample-specific identification sequence,
[0073] Perform a second nucleic acid purification.
[0074] A14. The method of any one or more of clauses A1-A13, further comprising limiting one or more of the target-specific primer concentration or the number of cycles in the first multiplex PCR reaction to a concentration greater than about 10 7 The platform amplification of nucleic acids present at a concentration of copies / ml and maintained in the exponential amplification phase is less than about 10 7 copies / ml of nucleic acid.
[0075] A15. The method of any one or more of clauses A1-A14, further comprising combining two or more samples and subjecting them to sequencing simultaneously, wherein the two or more samples are combined after preparation and prior to sequencing.
[0076] A16. The method of any one or more of clauses A1-A15, wherein 2-1000 samples are combined after preparation and before sequencing, preferably 2-500 samples are combined after preparation and before sequencing, more preferably 2-100 samples are combined after preparation and before sequencing, more preferably 2-50 samples are combined after preparation and before sequencing, or most preferably 2-32 samples are combined after preparation and before sequencing.
[0077] A17. The method of any one or more of clauses A1-A16, wherein each pooled sample has associated therewith a unique set of sample-specific identification sequences such that sequencing data from each sample in the pool can be distinguished and separated.
[0078] A18. The method of any one or more of clauses A1-A17, wherein each pooled sample has its own internal quantitative standard associated with its own set of unique sample-specific identification sequences, and wherein normalization is applied separately to each sample in the pool.
[0079] A19. The method of any one or more of clauses A1-A18, wherein the normalization separately (1) applies data acceptance / rejection criteria to each sample in the library, and (2) ensures that substantially the same limit of detection (LOD) is applied to all unknown nucleic acids in each sample.
[0080] A20. The method of any one or more of clauses A1-A19, further comprising: calculating the input quantity (IQTi, IQTj, IQTk…IQTn) of the plurality of unknown nucleic acids in the pooled sample after normalization as IQTn = normalized number of unknown “n” nucleic acid sequencing reads attributed to the “nth” unknown nucleic acid * (input quantity of internal quantitative standard / F).
[0081] A21. The method of any one or more of clauses A1-A20, wherein the sequencing assay is a next generation sequencing assay.
[0082] A22. The method of any one or more of clauses A1-A21, wherein the sequencing assay does not comprise performing relative quantification.
[0083] A23. The method of any one or more of clauses A1-A22, wherein the sequencing assay does not comprise performing quantification in a reaction separate from the sequencing assay.
[0084] A24. The method of one or more of clauses A1-A23, wherein the sequencing assay does not comprise the use of an assay or template-specific quantification standard.
[0085] A25. The method of any one or more of clauses A1-A24, wherein the sequencing assay does not comprise use of a competing template as a quantification standard.
[0086] B1. A method for performing a quantitative next generation sequencing (NGS) assay, comprising:
[0087] Providing a sample comprising one or more unknown nucleic acids to be sequenced;
[0088] A known amount of internal quantitative standard was added to the sample;
[0089] Prepare samples containing internal quantitative standards for sequencing;
[0090] Sequencing unknown nucleic acids in the sample and an internal quantitative standard to generate sequencing data;
[0091] Counting the number of sequencing reads in a sequencing dataset that originate from unknown nucleic acids and internal quantification standards; and
[0092] The sequencing data sets were normalized and the input number of unknown nucleic acids (IQT) was calculated by IQT = normalized unknown nucleic acid sequencing reads derived from unknown nucleic acids * (input number of internal quantification standard / F), where 'F' is the minimum expected number of fixed internal quantification standard sequencing reads.
[0093] B2. The method of clause B1, wherein each unknown nucleic acid in the sample has about 0-10 13 copies / ml (e.g., approximately 10 2 -10 9 copies / ml) concentration.
[0094] B3. The method of at least one of clauses B1 or B2, wherein the known amount of the internal quantitative standard added to the sample is about 10 4 -10 6 copies / ml (e.g., approximately 10 5 -10 6 copies / ml).
[0095] B4. The method of any one or more of clauses B1-B3, further comprising normalizing the sequencing dataset by NORM = number of sequencing reads * (F / observed number of sequencing reads derived from an internal quantification standard).
[0096] B5. The method of any one or more of clauses B1-B4, wherein the normalization separately (1) applies data acceptance / rejection criteria to each sample in the assay, and (2) ensures that substantially the same limit of detection (LOD) is applied to all unknown nucleic acids in each sample across multiple samples that are pooled together across multiple users, multiple days, etc.
[0097] B6. The method of any one or more of clauses B1-B5, wherein the unknown nucleic acid has about 10 2 -10 10 copies / ml, or preferably about 10 2 -109 The concentration of copies / ml.
[0098] B7. The method of any one or more of clauses B1-B6, wherein 'F' is related to the limit of detection (LOD) of the NGS assay, and wherein if 'F' is about 1x10 3 – 5x10 3 , then the LOD for unknown nucleic acids is about 10 2 – 10 3 copies / ml.
[0099] B8. The method of any one or more of clauses B1-B7, wherein 'F' is related to the limit of detection (LOD) of the NGS assay, and wherein if 'F' is about 1x10 4 – 5x10 4 , then the LOD for unknown nucleic acids is about 10 1 – 10 2 copies / ml.
[0100] B9. The method of any one or more of clauses B1-B8, further comprising: calculating the input quantity (IQTi, IQTj, IQTk…IQTn) of multiple unknown nucleic acids (if any) in the sample as IQTn = normalized unknown “n” nucleic acid sequencing reads attributed to the “nth” unknown nucleic acid * (input quantity of internal quantitative standard / F).
[0101] B10. The method of any one or more of clauses B1-B9, further comprising combining two or more samples and subjecting them to sequencing simultaneously, wherein the two or more samples are combined after preparation and prior to sequencing.
[0102] B11. The method of any one or more of clauses B1-B10, wherein each pooled sample has associated therewith a unique set of sample-specific identification sequences such that sequencing data from each sample in the pool can be distinguished and separated.
[0103] B12. The method of any one or more of clauses B1-B11, wherein each pooled sample has its own internal quantification standard associated with its own set of unique sample-specific identification sequences, and wherein quantification is applied separately to each nucleic acid from each sample in the library.
[0104] B13. The method of any one or more of clauses B1-B12, further comprising: calculating the input quantity (IQTi, IQTj, IQTk…IQTn) of the plurality of unknown nucleic acids in the combined sample as IQTn = normalized unknown “n” nucleic acid sequencing reads attributed to the “nth” unknown nucleic acid * (input quantity of internal quantitative standard / F).
[0105] B14. The method of any one or more of clauses B1-B13, further comprising providing a set of assay-specific positive controls to be sequenced in the sample, wherein the positive controls include a positive control corresponding to each of the one or more unknown nucleic acids sequenced in the assay.
[0106] B15. The method of any one or more of clauses B1-B14, further comprising applying an assay-specific correction factor to each of the one or more unknown nucleic acids based on sequencing of an assay-specific positive control.
[0107] B16. The method of any one or more of clauses B1-B15, wherein performing a quantitative NGS assay does not comprise performing relative quantification.
[0108] B17. The method of any one or more of clauses B1-B16, wherein performing the quantitative NGS assay does not comprise performing the quantification in a reaction separate from the sequencing assay.
[0109] B18. The method of any one or more of clauses B1-B17, wherein performing the quantitative NGS assay does not include use of assay or template-specific quantitation standards.
[0110] B19. The method of any one or more of clauses B1-B18, wherein performing the quantitative NGS assay does not include using a competing template as a quantitation standard.
[0111] C1. A kit for normalizing and quantifying unknown nucleic acids in a next-generation sequencing (NGS) assay, comprising:
[0112] an internal quantitative standard, wherein the internal quantitative standard is a nucleic acid configured to be added in known amounts to a sample comprising an unknown nucleic acid to be sequenced; and
[0113] Instructions for using an internal quantitative standard for normalizing sequencing data sets and for calculating the input quantity of unknown nucleic acids,
[0114] wherein the internal quantitative standard is configured to be used at about 10 4 -10 6 The range of copies / ml was added to the sample.
[0115] wherein sequencing data are normalized using the internal quantification standard by NORM = number of sequencing reads * (F / observed number of sequencing reads derived from the internal quantification standard), where 'F' is the minimum expected number of sequencing reads of the fixed internal quantification standard, and
[0116] The internal quantitative standard is used to calculate the input quantity (IQT) of unknown nucleic acids by IQT=normalized unknown nucleic acid sequencing read number derived from unknown nucleic acids*(input quantity of internal quantitative standard / F).
[0117] C2. The kit of clause C1, further comprising a sequencing-specific adaptor for at least an internal quantitation standard comprising a sequencing primer binding site and a sample-specific identification sequence.
[0118] C3. The kit of at least one of clause C1 or clause C2, further comprising target-specific primers with custom overhangs configured for amplifying an internal quantitation standard and for annealing or ligating sequencing-specific adapters.
[0119] C4. The kit of at least one of clauses C1 - C3, further comprising two or more internal quantitative standards, wherein the two or more internal quantitative standards are each configured to be added to the sample at a different known concentration for generating a standard curve for quantifying unknown nucleic acids.
[0120] D1. A method for performing a comparator study, comprising:
[0121] providing a first assay comprising multiplex amplification and detection of one or more target nucleic acids, the first assay having a limit of detection (LOD);
[0122] Providing a second assay, different from the first assay, for confirming the detection and LOD of the first assay, wherein the second assay comprises:
[0123] preparing a sample for sequencing comprising at least one internal quantitation standard;
[0124] sequencing to generate a sequencing dataset for the sample, wherein the sequencing dataset includes sequencing reads observed from the target nucleic acid and an internal quantitative standard;
[0125] counting the number of sequencing reads in a sequencing data set that originate from the target nucleic acid and the internal quantification standard; and
[0126] A sequencing data set is normalized, wherein the normalization (1) applies data acceptance / rejection criteria to the sequencing data set based on the presence of a minimum number of sequencing reads with respect to an internal quantification standard for a sample, and (2) ensures that substantially the same limit of detection (LOD) is applied to all unknown nucleic acids in a sequencing assay, and wherein the LOD of a second assay is substantially the same as the LOD of a first assay.
[0127] D2. The method of clause D1, wherein the first assay further comprises adding one or more internal quantitative standards to the sample, and performing a quantitative two-step amplification on the sample, the quantitative two-step amplification comprising:
[0128] amplifying the sample in a first stage multiplex amplification mixture comprising a plurality of target primers, each target primer configured to amplify a different target that may be present in the sample, and at least one quantification standard primer configured to amplify an internal quantification standard nucleic acid,
[0129] dividing the first stage amplification mixture into a plurality of second stage individual reactions, a first plurality of second stage individual reactions each comprising at least one primer configured to further amplify one of the different targets that may be present in the sample, and a second plurality of second stage individual reactions each comprising at least one primer configured to further amplify one of the internal quantitative standard nucleic acids, and
[0130] subjecting the plurality of second stage individual reactions to amplification conditions to generate one or more target amplicons and a plurality of quantification standard amplicons, each quantification standard amplicon having an associated quantification standard Cp,
[0131] Each target nucleic acid has a crossing point (Cp), and each internal standard has a known concentration and a known quantitative standard Cp in the first assay.
[0132] D3. The method of at least one of clause D1 or clause D2, further comprising
[0133] Generating a standard curve from the quantitative standard Cp; and
[0134] A standard curve is used to quantify each of the one or more target nucleic acids.
[0135] D4. The method of any one or more of clauses D1-D3, wherein each target nucleic acid is quantified using a standard curve generated using a least squares regression line fit to
[0136] log 10 (Concentration) = (Cp - b) / a
[0137] where Cp is the intersection point measured for each target,
[0138] b , intercept, represents the log of the target 10 The Cp value when (concentration) is zero, and
[0139] a is the slope that represents how much Cp changes with a single unit change in concentration.
[0140] D5. The method of any one or more of clauses D1-D4, further comprising normalizing the sequencing dataset in the second assay by NORM = number of sequencing reads * (F / observed number of sequencing reads derived from the internal quantification standard).
[0141] D6. The method of any one or more of clauses D1-D5, wherein the input quantity (IQT) of each target nucleic acid in the sample is calculated by IQT = normalized number of target nucleic acid sequencing reads derived from the target nucleic acid * (input quantity of internal quantification standard / F), where 'F' is the minimum expected number of fixed internal quantification standard sequencing reads.
[0142] D7. The method of any one or more of clauses D1-D6, wherein 'F' is related to the limit of detection (LOD) of a second assay, and wherein the LOD of the second assay is selected to be substantially the same as the LOD of the first assay.
[0143] D8. The method of any one or more of clauses D1-D7, wherein if 'F' is about 1x10 3 – 5x10 3 , the LOD for detecting the target nucleic acid in the second assay is about 10 2 – 10 3 copies / ml.
[0144] D9. The method of any one or more of clauses D1-D, wherein if 'F' is about 1x10 4 – 5x10 4 , the LOD for detecting the target nucleic acid in the second assay is about 10 1 – 10 2 copies / ml.
[0145] D10. The method of any one or more of clauses D1-D9, further comprising: calculating the input quantity (IQTi, IQTj, IQTk…IQTn) of the multiple target nucleic acids (if present) in the sample as IQTn = normalized unknown “n” nucleic acid sequencing reads attributed to the “nth” unknown nucleic acid * (input quantity of internal quantitative standard / F).
[0146] D11. The method of any one or more of clauses D1-D10, further comprising pooling the two or more samples and subjecting them to sequencing simultaneously in a second assay, wherein the two or more samples are pooled after preparation and prior to sequencing.
[0147] D12. The method of any one or more of clauses D1-D11, wherein each pooled sample has associated therewith a unique set of sample-specific identification sequences such that sequencing data from each sample in the pool can be distinguished and separated.
[0148] D13. The method of any one or more of clauses D1-D12, wherein each pooled sample has its own internal quantification standard associated with its own set of unique sample-specific identification sequences, and wherein quantification is applied separately to each nucleic acid from each sample in the library.
[0149] D14. The method of any one or more of clauses D1-D13, further comprising: calculating the input quantity (IQTi, IQTj, IQTk…IQTn) of the plurality of unknown nucleic acids in the combined sample as IQTn = normalized unknown “n” nucleic acid sequencing reads attributed to the “nth” unknown nucleic acid * (input quantity of internal quantitative standard / F).
[0150] D15. The method of any one or more of clauses D1-D14, wherein the quantitative standard nucleic acid and the target nucleic acid have similar amplification efficiency and sequencing efficiency.
[0151] D16. The method of any one or more of clauses D1-D15, wherein the second assay is a next generation sequencing assay.
[0152] D17. The method of any one of clauses D1-D16, wherein the second determination does not comprise performing relative quantification,
[0153] D18. The method of any one of clauses D1-D17, wherein the second assay does not comprise performing quantification in a reaction separate from the sequencing assay.
[0154] D19. The method of any of clauses D1-D18, wherein the second determination does not comprise the use of an assay or template-specific quantitation standard.
[0155] D20. The method of any of clauses D1-D19, wherein the second determination does not comprise use of a competing template as a quantitation standard.
[0156] In any of the foregoing embodiments of the method for performing a comparator study, the second assay can be a next generation sequencing assay.
[0157] This Summary is provided to introduce selected concepts in a simplified form that are further described below in the drawings and the detailed description. This Summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used as an aid in determining the scope of the claimed subject matter.
[0158] Additional features and advantages will be set forth in the following description, and in part will be apparent from the description, or may be learned through practice of the invention. These and other features will become more apparent from the following description and appended claims, or may be learned through practice of the invention as set forth below. BRIEF DESCRIPTION OF THE DRAWINGS
[0160] In order to describe the manner in which the above-recited and other advantages and features of the present invention can be obtained, a more particular description of the invention, briefly described above, will be rendered by reference to specific embodiments thereof that are illustrated in the accompanying drawings. Understanding that these drawings depict only typical embodiments of the invention and are not therefore to be considered limiting of its scope, the invention will be described and explained with additional specificity and detail through the use of the accompanying drawings, in which:
[0161] Figure 1 A flexible pouch according to one embodiment of the present invention is shown.
[0162] Figure 2 Together they form an exemplary embodiment according to the present invention for use with Figure 1 An exploded perspective view of an instrument for use with a pouch comprising Figure 1 Pouch.
[0163] Figure 3 According to an exemplary embodiment of the present invention, Figure 2 A partial cross-sectional view of the instrument, including Figure 2 Airbag components, wherein Figure 1 The pouch is shown in dotted lines.
[0164] Figure 4 Shown in Figure 2 An electric motor is used in one illustrative embodiment of an apparatus.
[0165] Figure 5 Shown are the Cp's across five dilutions of four different intended synthetic quantitation standards.
[0166] Figure 6A Similar to Figure 5 , but only data for the three quantification standards are shown. Figure 6BA single curve using data from all three quantitation standards is shown.
[0167] Figure 7 shows the presence of Acinetobacter baumannii ( A. baumannii ) along with curves generated from three quantitation standards. The x-axis is the amount of A. baumannii or quantitation standard included in the reaction, and the y-axis is Cp.
[0168] Figure 8A Shown are a composite standard curve from the quantitation standards and an external standard curve specific for A. baumannii without calibration. Figure 8B Shows the Figure 8A Same data, with assay-specific correction factors.
[0169] Figure 9 A hypothetical sequencing data set for three samples A, B, and C is shown.
[0170] Figure 10 Raw sequencing counts for the combined sample set are shown.
[0171] Figure 11 Shows Figure 10 Fragment counts / sample of the internal quantification standard (also referred to herein as QSM) in the pooled sample set.
[0172] Figure 12 Two alternative methods for binning and normalizing sequencing datasets are shown.
[0173] Figure 13 An example of a sequencing sample preparation workflow is shown.
[0174] Figure 14 The bacterial and / or viral load distribution in the sputum sample population is shown.
[0175] Figure 15 Shown are PCR amplification curves illustrating a method for preserving the dynamic range of a sequencing reaction, where the reaction is set up and stopped such that the high copy target is in a plateau phase and the low copy target is in an exponential amplification phase.
[0176] Figure 16 Shown are the batch positive controls (PC) quantifications across the assays.
[0177] Details
[0178] Exemplary embodiments are described below with reference to the accompanying drawings. Many different forms and embodiments are possible without departing from the spirit and teachings of the present disclosure, and therefore the disclosure should not be construed as limited to the exemplary embodiments set forth herein. On the contrary, these exemplary embodiments are provided so that the present disclosure will be thorough and complete and convey the scope of the present disclosure to those skilled in the art. In the accompanying drawings, the sizes and relative sizes of layers and regions may be exaggerated for clarity. Throughout the specification, the same reference numerals represent the same elements.
[0179] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as those generally understood by those of ordinary skill in the art to which this disclosure belongs. It will be further understood that terms, such as those defined in commonly used dictionaries, should be interpreted as having a meaning consistent with their meaning in the context of this application and the related art, and should not be interpreted in an idealized or overly formalized sense unless clearly defined as such in this article. The terms used in the description of the invention herein are only for the purpose of describing specific embodiments and are not intended to limit the present invention. Although only certain exemplary materials and methods are described herein, many methods and materials similar or equivalent to those described herein can be used in the practice of this disclosure.
[0180] All publications, patent applications, patents, or other references mentioned herein are incorporated by reference in their entirety. In the event of a conflict in terminology, the present specification controls.
[0181] Various aspects of the present disclosure, including devices, systems, methods, and the like, may be described with reference to one or more exemplary implementations. As used herein, the terms "exemplary" and "illustrative" mean "serving as an example, example, or illustration" and should not necessarily be construed as preferred or advantageous over other implementations disclosed herein. In addition, references to "implementations" or "embodiments" of the present disclosure or the invention include specific references to one or more embodiments thereof, and vice versa, and are intended to provide illustrative examples and not to limit the scope of the invention, which is indicated by the appended claims rather than the following description.
[0182] It should be noted that, as used in this specification and the appended claims, the singular forms "a," "an," and "the" include plural referents unless the content clearly dictates otherwise. Thus, for example, reference to "a tile" includes one, two, or more tiles. Similarly, reference to multiple referents should be construed to include a single referent and / or multiple referents unless the content and / or context clearly dictate otherwise. Thus, reference to a "tile" does not necessarily require a plurality of such tiles. Rather, it should be understood that conjugation is not relied upon; one or more tiles are contemplated herein.
[0183] As used throughout this application, the words "may" and "might" are used in a permissive sense (i.e., meaning having the potential to), rather than the mandatory sense (i.e., meaning must). Additionally, the terms "including," "having," "involving," "containing," "characterized by," variations thereof (e.g., "includes," "has," "involves," "contains," etc.), and similar terms as used herein, including in the claims, are intended to be inclusive and / or open-ended and have the same meaning as the word "comprising" and variations thereof (e.g., "comprise" and "comprises"), and do not exclude additional, unrecited illustrative elements or method steps.
[0184] As used herein, directions and / or any terms such as "top", "bottom", "left", "right", "up", "down", "upper", "lower", "inner", "outer", "interior", "exterior", "inner", "exterior", "proximal", "distal", "front", "back", etc. may be used only to indicate relative directions and / or orientations and may not otherwise be intended to limit the present disclosure, including the scope of the specification, invention and / or claims.
[0185] It should be understood that when an element is referred to as being “coupled,” “connected,” or “responsive” to another element or being “on another element,” it can be directly coupled, connected, or responsive to the other element or on another element, or intervening elements may be present. In contrast, when an element is referred to as being “directly coupled,” “directly connected,” or “directly responsive” to another element or being “directly on another element,” there are no intervening elements present.
[0186] Exemplary embodiments of the present inventive concepts are described herein with reference to cross-sectional illustrations that are schematic illustrations of idealized embodiments (and intermediate structures) of the exemplary embodiments. As such, variations from the illustrated shapes due to, for example, manufacturing techniques and / or tolerances are anticipated. Thus, exemplary embodiments of the present inventive concepts should not be construed as limited to the particular shapes of regions illustrated herein but are to include deviations in shapes that result, for example, from manufacturing. Accordingly, the regions illustrated in the figures are schematic in nature, and their shapes are not intended to illustrate the actual shape of a region of a device and are not intended to limit the scope of the exemplary embodiments.
[0187] It should be understood that although the terms "first," "second," etc., may be used herein to describe various elements, these elements should not be limited by these terms. These terms are merely used to distinguish one element from another. Thus, a "first" element may be referred to as a "second" element without departing from the teachings of this embodiment.
[0188] It should also be understood that the various implementations described herein can be utilized in combination with any other implementations described or disclosed without departing from the scope of the present disclosure. Thus, products, components, elements, devices, instruments, systems, methods, processes, compositions, and / or kits according to certain implementations of the present disclosure may include, incorporate, or otherwise incorporate properties, features, components, components, elements, steps, and / or the like described in other implementations (including systems, methods, instruments, and / or the like) disclosed herein without departing from the scope of the present disclosure. Thus, reference to specific features associated with one implementation should not be construed as limiting application to that implementation.
[0189] The headings used herein are for organizational purposes only and are not intended to limit the scope of the specification or claims. To facilitate understanding, identical reference numerals have been used, whenever possible, to denote identical elements common to the figures. Furthermore, identical numbers have been used for elements across the various figures, whenever possible. Furthermore, alternative configurations of a particular element may each include a separate letter appended to the element number.
[0190] The term "about" is used herein to mean approximately, in the region of approximately or roughly. When the term "about" is used in conjunction with a numerical range, it modifies the range by extending the boundaries above and below the number. Generally speaking, the term "about" is used herein to modify the numerical value above and below the value by 5%. When expressing such a range, another embodiment includes from a specific value and / or to another specific value. Similarly, when a value is expressed as an approximate value by using the antecedent "about", it should be understood that the specific value forms another embodiment. It should be further understood that the endpoints of each range are meaningful relative to the other endpoint and independently of the other endpoint.
[0191] As used herein, the word "or" means any one member of a particular list, and also includes any combination of members of that list.
[0192] "Sample" means an animal; a tissue or organ from an animal; a cell (in vivo, directly from a subject, or maintained in culture or from a cultured cell line); a cell lysate (or lysate fraction) or a cell extract; a solution containing one or more molecules derived from cells, cellular material, or viral material (e.g., polypeptides or nucleic acids); or a solution containing non-naturally occurring nucleic acids, illustratively cDNA or a next-generation sequencing library, which is assayed as described herein. A sample can also be any bodily fluid or excretion (e.g., but not limited to blood, urine, feces, saliva, tears, bile, or cerebrospinal fluid), which may or may not contain host or pathogen cells, cellular components, or nucleic acids.
[0193] As used herein, the phrase "nucleic acid" refers to a naturally occurring or synthetic oligonucleotide or polynucleotide, whether DNA or RNA or a DNA-RNA hybrid, single-stranded or double-stranded, sense or antisense, which is capable of hybridizing to a complementary nucleic acid via Watson-Crick base pairing. The nucleic acids of the present invention may also include nucleotide analogs (e.g., BrdU), modified or treated bases, and non-phosphodiester internucleoside linkages (e.g., peptide nucleic acid (PNA) or thiodiester linkages). In particular, nucleic acids may include, but are not limited to, DNA, cDNA, gDNA, ssDNA, dsDNA, RNA, including all RNA types such as miRNA, mtRNA, rRNA, including coding or non-coding regions, or any combination thereof.
[0194] "Probe," "primer," or "oligonucleotide" refers to a single-stranded nucleic acid molecule with a defined sequence that can base pair with a second nucleic acid molecule ("target") containing a complementary sequence. The stability of the resulting hybrid depends on the length, GC content, and the extent to which base pairing occurs. The extent of base pairing is affected by parameters such as the degree of complementarity between the probe and target molecules and the stringency of the hybridization conditions. Hybridization stringency is affected by parameters such as temperature, salt concentration, and the concentration of organic molecules such as formamide, and is determined by methods known to those skilled in the art. Probes, primers, and oligonucleotides can be detectably labeled radioactively, fluorescently, or non-radioactively by methods well known to those skilled in the art. dsDNA binding dyes can be used to detect dsDNA. It will be understood that a "primer" is specifically configured to be extended by a polymerase, while a "probe" or "oligonucleotide" may or may not be so configured. As probes, oligonucleotides can be used as part of a number of fluorescent PCR primers and probe-based chemistries known in the art, including those that share the use of fluorescence quenching and / or fluorescence resonance energy transfer (FRET) configurations, such as 5' nuclease probes (TaqMan® probes), dual hybridization probes (HybProbes®), or Eclipse® probes or molecular beacons, or Amplifluor® assays such as Scorpions®, LUX®, or QZyme® PCR primers, including those with natural or modified bases.
[0195] "dsDNA binding dye" means a dye that fluoresces differentially when bound to double-stranded DNA than when bound to single-stranded DNA or free in solution, typically by fluorescing more intensely. Although reference is made to dsDNA binding dyes, it should be understood that any suitable dye may be used herein, some non-limiting illustrative dyes of which are described in U.S. Patent No. 7,387,887, which is incorporated herein by reference. Other signal-generating substances can be used to detect nucleic acid amplification and melting, illustratively enzymes, antibodies, etc., as known in the art.
[0196] "Specifically hybridizes" means that the probe, primer, or oligonucleotide recognizes and physically interacts (ie, base pairs) with a substantially complementary nucleic acid (eg, a sample nucleic acid) under high stringency conditions, and does not substantially base pair with other nucleic acids.
[0197] "High stringency conditions" means conditions at about the melting temperature (Tm) minus 5°C (ie, 5° below the Tm of the nucleic acid). Functionally, high stringency conditions are used to identify nucleic acid sequences with at least 80% sequence identity.
[0198] While PCR is the amplification method used in the examples herein, it should be understood that any amplification method using primers is suitable. Such suitable procedures include any type of polymerase chain reaction (PCR) (single-step, two-step, or other); strand displacement amplification (SDA); nucleic acid sequence-based amplification (NASBA); cascade rolling circle amplification (CRCA); loop-mediated isothermal amplification of DNA (LAMP); isothermal and chimeric primer-initiated nucleic acid amplification (ICAN); target-based helicase-dependent amplification (HDA); transcription-mediated amplification (TMA), next-generation sequencing techniques, and the like. Therefore, when the term PCR is used, it should be understood to include other alternative amplification methods, including amino acid quantification methods. For amplification methods that do not contain discrete cycles, reaction times can be used in which the cycle or Cp is measured, and additional reaction times can be added in which additional PCR cycles are added in the embodiments described herein. It should be understood that the protocol may need to be adjusted accordingly.
[0199] As used herein, the term "crossing point" (Cp) (or alternatively, cycle threshold (Ct), quantification cycle (Cq), or synonymous terms used in the art) refers to the number of PCR cycles required to achieve a fluorescence signal above a certain threshold for a given PCR product (e.g., target or internal standard), as determined experimentally. The cycle at which each reaction rises above the threshold depends on the amount of target (i.e., reaction template) present at the start of the PCR reaction. The threshold is typically set at the point where a fluorescence signal of the product above background fluorescence is detectable; however, other thresholds may be employed. As an alternative to setting a somewhat arbitrary threshold, the Cp can be determined by calculating the reaction point at which the next, second, or nth-order derivative has its maximum, which identifies the cycle at which the curvature of the amplification curve is greatest. Illustrative derivation methods are taught in U.S. Patent No. 6,303,305, which is incorporated herein by reference in its entirety. However, where or how the threshold is set is generally unimportant, as long as the same threshold is used for all reactions being compared. As is known in the art, other points may also be used, and any such point may be substituted for Cp, Ct, or Cq in any of the methods discussed herein.
[0200] "Sample processing controls" means pathogens, microorganisms, cells (whether viable or not), nucleic acids, or any particles, natural or synthetic, that have the ability to simulate a pathogen, or a portion thereof, or nucleic acid, and its behavior during a sample workflow. Sample processing controls are often included in a device in known quantities to control some or all steps of the workflow followed by the sample, illustratively ensuring that the sample has been properly lysed, nucleic acids potentially infecting the target pathogen have been properly extracted and purified, and that proper amplification and detection of specific sequences of the target pathogen has occurred.
[0201] Illustratively, a microorganism (illustratively, Schizosaccharomyces pombe ( Schizosaccharomyces pombe ) (Schizosaccharomyces pombe ( S. pombe )) mimics the target microorganism to be detected and quantified as closely as possible. Sample processing control particles can replicate the structure of the pathogen to be detected (e.g., membrane and / or capsid and / or envelope), allowing them to simulate the behavior of the pathogen and its target nucleic acid along the workflow. The goal of the sample processing control is to ensure that target lysis and nucleic acid extraction yields are similar to those of the sample processing control, and that the purified nucleic acid is properly processed to ensure optimal amplification / detection. For qualitative results, the pathogen can be reported as positive or negative, or, if the run control fails, as undetermined. The sample processing control can be one of several run controls and should be positive, potentially within a specified range, to validate the run, as some inhibitory conditions can reduce the yield of extraction, purification, or PCR amplification / detection. The sample processing control can be used to monitor such inhibition, with yield reductions being similar between the sample processing control and the target pathogen. For qualitative results, failure to detect such inhibition can lead to false-negative results. For quantitative results, inhibition of one of the workflow steps can provide an underestimated quantitative result. Therefore, several illustrative embodiments of the present invention utilize at least one sample processing control (SPC) for at least two objectives:
[0202] 1) Control and verify workflow: the classic role of SPC as described above, and
[0203] 2) Helping to quantify target nucleic acids in test samples: A new role for SPC, which also serves as a quantitative standard.
[0204] Illustratively, the SPC follows some or all of the processes that the sample undergoes. Thus, the SPC can be added before or during the sample lysis step. Sample treatment controls can be selected based on the type of target pathogen. For example, a bacteriophage such as PhiX174 can be selected for viral assays, which are good candidates for simulating target viruses, or yeast such as Schizosaccharomyces pombe, which are used in a wide range of bacterial and yeast quantitative assays.
[0205] If a single pathogen is to be detected, two amplification assays, illustratively PCR assays (target pathogen and a sample treatment control used as a quantification standard) can be designed to achieve the same or similar thermodynamic characteristics and enable accurate quantification using a synthetic quantification standard (as in Example 5).
[0206] For quantification of multiple pathogens (i.e., multiplex amplification), it can be difficult to align the amplification protocol for sample treatment controls (illustratively, PCR design) with the protocol for the amplification assay (illustratively, the PCR assay for each pathogen) to achieve identical thermodynamic characteristics (illustratively due to sequence variability and amplicon length). Consequently, PCR efficiencies may vary for different target pathogens. To this end, a correction factor can be calculated for each pathogen, relating the obtained quantification to the quantification standard and input standard curve.
[0207] In an alternative to synthetic quantitative standards, calibration can be performed against known natural microorganisms with known concentrations or against other naturally occurring nucleic acid templates.
[0208] In another embodiment of the present invention, pathogens can also be reliably quantified in any amplification system having at least two different, illustratively three or four different sample treatment controls, provided that these sample treatment controls can be identified via known identification techniques, such as sequence-specific probes labeled with fluorescent, radioactive, chemiluminescent, enzymatic, etc., as known in the art.
[0209] Although various examples herein refer to human targets and human pathogens, these examples are illustrative only.The methods, kits, and devices described herein can be used to detect and sequence a wide variety of nucleic acid sequences from a wide variety of samples, including human, veterinary, industrial, and environmental.
[0210] Various embodiments disclosed herein use independent nucleic acid analysis pouches to determine the presence of various biological substances in a sample, illustratively antigens and nucleic acid sequences, illustratively in a single closed system. Such systems, including pouches and instruments for use with the pouches, are disclosed in more detail in U.S. Patent Nos. 8,394,608; and 8,895,295; and U.S. Patent Application No. 2014-0283945, which are incorporated herein by reference. However, it should be understood that such instruments and pouches are merely illustrative, and the nucleic acid preparation and amplification reactions discussed herein can be performed in any of a variety of open or closed system sample vessels as known in the art, including 96-well plates, other configurations of flat plates, arrays, turntables, and the like, using various nucleic acid purification and amplification systems as known in the art. Although the terms "sample well," "amplification well," "amplification vessel," and the like are used herein, these terms are intended to encompass wells, tubes, and various other reaction vessels used in these amplification systems. Such amplification systems can include a single multiplex step in an amplification vessel and can optionally include multiple second-stage individual or lower-order multiplex reactions in multiple individual reaction wells. In one embodiment, the pouch is used to assay for multiple pathogens. The pouch can include one or more blisters serving as sample wells, illustratively in a closed system. Illustratively, various steps can be performed in the optional single-use pouch, including nucleic acid preparation, primary large-volume multiplex PCR, dilution of the primary amplification product, and secondary PCR, followed by optional real-time detection or post-amplification analysis, such as melting curve analysis. Furthermore, it should be understood that while various steps can be performed in the pouch of the present invention, one or more steps may be omitted for certain applications, and the pouch configuration may be modified accordingly.
[0211] Figure 1 An illustrative pouch 510 is shown that can be used in various embodiments or that can be reconfigured for various embodiments. Pouch 510 is similar to that of U.S. Pat. No. 8,895,295. Figure 15 , where like items are numbered the same. Fitting 590 is provided with inlet channels 515a to 515l which also serve as reagent reservoirs or waste reservoirs. Illustratively, reagents can be freeze-dried in fitting 590 and rehydrated before use. Blisters 522, 544, 546, 548, 564 and 566 and their respective channels 514, 538, 543, 552, 553, 562 and 565 are similar to those of U.S. Pat. No. 8,895,295. Figure 15 Blisters of the same number. Figure 1 The second stage reaction zone 580 is similar to that of U.S. Patent Application No. 8,895,295, but the second stage reaction zones 582 of the high density array 581 are arranged in a slightly different pattern. Figure 1The more circular pattern of the high-density array 581 eliminates holes in the corners and can result in more uniform filling of the second-stage wells 582. As shown, the high-density array 581 is provided with 102 second-stage wells 582. The pouch 510 is suitable for use with a FilmArray® instrument (BioFire Diagnostics, LLC, Salt Lake City, UT). However, it should be understood that the pouch embodiment is illustrative only.
[0212] Although other containers may be used, illustratively, pouch 510 is formed from two layers of flexible plastic film or other flexible material, such as polyester, polyethylene terephthalate (PET), polycarbonate, polypropylene, polymethyl methacrylate, and mixtures thereof, which can be prepared by any method known in the art, including extrusion, plasma deposition, and lamination. Metal foil or plastic with an aluminum laminate structure may also be used. Other barrier materials are known in the art and can be sealed together to form blisters and channels. If a plastic film is used, the layers can illustratively be bonded together by heat sealing. Illustratively, such a material has low nucleic acid binding capacity.
[0213] For embodiments employing fluorescence monitoring, plastic films with sufficiently low absorbance and autofluorescence at the operating wavelength are preferred. Such materials can be identified by testing different plastics, varying plasticizer and compounding ratios, and films of varying thicknesses. For plastics with aluminum or other foil laminates, the portion of the pouch to be read by the fluorescence detection device can be left without the foil. For example, if fluorescence is detected in the second-stage well 582 of the second-stage reaction zone 580 of the pouch 510, one or both layers at the well 582 can be left without the foil. In the PCR example, a film laminate consisting of approximately 0.0048 inch (0.1219 mm) thick polyester (Mylar, DuPont, Wilmington, DE) and 0.001-0.003 inch (0.025-0.076 mm) thick polypropylene film performs well. Specifically, the pouch 510 is made of a transparent material that transmits approximately 80%-90% of incident light.
[0214] In an illustrative embodiment, by applying pressure on the blister and passage, illustratively pneumatic pressure, the material is moved between the blister. Accordingly, in the embodiment employing pressure, the pouch material is illustratively flexible enough to allow the pressure to have the desired effect. The term "flexible" is used herein to describe the physical characteristics of the material of the pouch. The term "flexible" is defined herein as being easy to deform by the pressure level used herein without breaking, destroying, cracking, etc. For example, thin plastic sheets, such as Saran™ packaging and Ziploc® bags, and thin metal foils, such as aluminum foil, are flexible. However, even in the embodiment employing pneumatic pressure, only certain areas of the blister and passage need to be flexible. In addition, only one side of the blister and passage needs to be flexible, as long as the blister and passage are easy to deform. Other areas of the pouch 510 can be made of rigid material, or can be reinforced with rigid material.
[0215] Illustratively, plastic film is used for pouch 510. Sheet metal, illustratively aluminum or another suitable material, can be milled or otherwise cut to create a mold with a raised surface pattern. When mounted on a pneumatic press (illustratively, A-5302-PDS, Janesville Tool Inc., Milton, WI), illustratively regulated at an operating temperature of 195°C, the press acts like a printing press, melting the mold only at the sealing surface of the plastic film where it contacts the film. Following the formation of pouch 510, various components, such as PCR primers (illustratively applied to the film and dried), antigen-binding substrates, magnetic beads, and zirconium silicate beads, can be sealed within various blisters. Reagents for sample processing can be applied to the film either collectively or separately prior to sealing. In one embodiment, nucleotide triphosphates (NTPs) are applied to the film separately from the polymerase and primers, essentially eliminating polymerase activity until the reaction is hydrated by an aqueous sample. If the aqueous sample is heated prior to hydration, this allows for true hot-start PCR and reduces or eliminates the need for expensive chemical hot-start components.
[0216] The pouch 510 can be used in a manner similar to that described in U.S. Patent No. 8,895,295. In one illustrative embodiment, 300 μl of a mixture containing the sample to be tested (100 μl) and lysis buffer (200 μl) is injected into an injection port (not shown) in the fitting 590 near the inlet channel 515a, and the sample mixture is drawn into the inlet channel 515a. Water is also injected into a second injection port (not shown) of the fitting 590 adjacent to the inlet channel 515l and distributed through channels (not shown) provided in the fitting 590, thereby hydrating up to 11 different reagents, each of which was previously provided in dry form at the inlet channels 515b to 515l. These reagents can illustratively include freeze-dried PCR reagents, DNA extraction reagents, wash solutions, immunoassay reagents or other chemical entities. Illustratively, the reagents are used for nucleic acid extraction, first stage multiplex PCR, dilution of multiplex reactions, and preparation of second stage PCR reagents and control reactions. In Figure 1 In the embodiment shown, all that needs to be injected is the sample solution in one injection port and water in the other. After injection, both injection ports can be sealed. For more information on various configurations of the pouch 510 and the accessory 590, see U.S. Patent No. 8,895,295, which has been incorporated by reference.
[0217] After injection, the sample moves from injection channel 515a via channel 514 to lysis bubble 522. Lysis bubble 522 is provided with beads or particles 534, such as ceramic beads, and is configured for vortexing via impact using a rotating blade or paddle provided in the FilmArray® instrument. Bead beating, which involves shaking or vortexing the sample in the presence of lysis particles, such as zirconium silicate (ZS) beads 534, is an effective method for forming lysis products. It should be understood that, as used herein, terms such as "lyse," "lysing," and "lysis product" are not limited to disrupting cells, but rather such terms include the destruction of non-cellular particles, such as viruses.
[0218] Figure 4 A bead beating motor 819 is shown which includes a blade 821 which may be mounted on a Figure 2800 is shown on a first side 811 of the support member 802 of the instrument 800. The blade 821 can extend through the slot 804 to contact the pouch 510. However, it should be understood that the motor 819 can be mounted on other structures of the instrument 800. In one illustrative embodiment, the motor 819 is a Mabuchi RC-280SA-2865 DC motor (Chiba, Japan) mounted on the support member 802. In one illustrative embodiment, the motor rotates at 5,000 to 25,000 rpm, more illustratively 10,000 to 20,000 rpm, and even more illustratively approximately 15,000 to 18,000 rpm. For the Mabuchi motor, 7.2 V has been found to provide sufficient rpm for lysis. However, it should be understood that the actual speed when the blade 821 impacts the pouch 510 may be slightly slower. Depending on the motor and blade used, other voltages and speeds may be used for lysis. Optionally, a controlled, small volume of air can be provided into the air bladder 822 adjacent to the lysis bubble 522. It has been found that, in some embodiments, partially filling the adjacent air bladder with one or more small volumes of air helps position and support the lysis bubble during the lysis process. Alternatively, other structures, illustratively a rigid or compliant gasket or other retaining structure around the lysis bubble 522, can be used to constrain the pouch 510 during lysis. It should also be understood that the motor 819 is illustrative only, and other devices can be used to grind, shake, or vortex the sample.
[0219] Once the cells have been fully lysed, the sample is moved to blister 546 by passage 538, blister 544 and passage 543, where the sample is mixed with nucleic acid binding substances, for example, silica-coated magnetic beads 533. The mixture is allowed to incubate for an appropriate length of time, illustratively about 10 seconds to 10 minutes. A retractable magnet positioned within the instrument adjacent to blister 546 captures magnetic beads 533 from the solution, forming a mass against the inner surface of blister 546. The liquid is then moved out of blister 546 and returned through blister 544 and into blister 522, which is now used as a waste container. One or more wash buffers from one or more of the injection channels 515c to 515e are provided to blister 546 via blister 544 and passage 543. Optionally, the magnet is retracted, and the magnetic beads 533 are cleaned by moving beads back and forth from blister 544 and 546 via passage 543. Once the magnetic beads 533 are washed, the magnetic beads 533 are recaptured in the blister 546 by activating the magnet and then the wash solution is moved to the blister 522. This process can be repeated as needed to wash the lysis buffer and sample fragments from the nucleic acid-binding magnetic beads 533.
[0220] After washing, the elution buffer stored in the injection channel 515f is moved to the bubble cap 548, and the magnet is retracted. The solution circulates between the bubbles 546 and 548 via the channel 552, breaking up the clumps of magnetic beads 533 in the bubble cap 546 and allowing the captured nucleic acids to separate from the beads and enter the solution. The magnet is activated again, capturing the magnetic beads 533 in the bubble cap 546, and the eluted nucleic acid solution is moved into the bubble cap 548.
[0221] The first-stage PCR master mix from injection channel 515g is mixed with the nucleic acid sample in bubble 548. Optionally, the mixture is mixed by forcing it between 548 and 564 via channel 553. After several mixing cycles, the solution is contained in bubble 564, where a pellet of first-stage PCR primers is provided, at least one set of primers for each target, and the first-stage multiplex PCR is performed. If an RNA target is present, a reverse transcription (RT) step may be performed before or concurrently with the first-stage multiplex PCR. The first-stage multiplex PCR temperature cycling in the FilmArray® instrument is illustratively performed for 15-30 cycles, although other levels of amplification may be desired depending on the requirements of the specific application. As is known in the art, the first-stage PCR master mix can be any of a variety of master mixes. In one illustrative embodiment, the first-stage PCR master mix can be any of the chemistries disclosed in US2015 / 0118715, which is incorporated herein by reference, for use with PCR protocols that take 20 seconds or less per cycle.
[0222] After the first-stage PCR has been performed for the desired number of cycles, the sample can be diluted, illustratively by forcing most of the sample back into bubble 548, leaving only a small amount in bubble 564, and adding the second-stage PCR master mix from injection channel 515i. Alternatively, dilution buffer from 515i can be moved to bubble 566 and then mixed with the amplified sample in bubble 564 by moving the fluid back and forth between bubbles 564 and 566. If desired, the dilution can be repeated several times using dilution buffer from injection channels 515j and 515k, or injection channel 515k can be reserved for sequencing or other post-PCR analysis, and the second-stage PCR master mix from injection channel 515h can then be added to some or all of the diluted amplified sample. It will be appreciated that the level of dilution can be adjusted by varying the number of dilution steps or by varying the percentage of sample discarded prior to mixing with the dilution buffer or the second-stage PCR master mix, which contains the components for amplification, illustratively a polymerase, dNTPs, and a suitable buffer, although other components may be suitable, particularly for non-PCR amplification methods. If desired, this mixture of sample and second stage PCR master mix can be preheated in blister 564 for second stage amplification before being moved to second stage wells 582. Such preheating can avoid the need for hot start components (antibodies, chemicals, or other) in the second stage PCR mix.
[0223] The illustrative second-stage PCR master mix is incomplete, lacking a primer pair, and each of the 102 second-stage wells 582 is preloaded with a specific PCR primer pair (or sometimes multiple primer pairs). If desired, the second-stage PCR master mix can lack other reaction components, and these components can also be preloaded in the second-stage wells 582. Each primer pair can be similar or identical to the first-stage PCR primer pair, or can be nested within a first-stage primer pair. The sample is moved from the blister 564 to the second-stage well 582 to complete the PCR reaction mixture. Once the high-density array 581 is filled, the individual second-stage reactions are sealed in their respective second-stage blisters by any number of means, as known in the art. Illustrative methods for filling and sealing the high-density array 581 without cross-contamination are discussed in U.S. Patent No. 8,895,295, which has been incorporated by reference. Illustratively, the various reactions in the wells 582 of the high-density array 581 are thermally cycled simultaneously, illustratively using one or more Peltier devices, although other devices for thermal cycling are known in the art.
[0224] In certain embodiments, the second-stage PCR master mix contains the dsDNA binding dye LCGreen® Plus (BioFire Diagnostics, LLC) to generate a signal indicative of amplification. However, it should be understood that this dye is illustrative only, and other signals can be used, including other dsDNA binding dyes and fluorescent, radioactive, chemiluminescent, enzymatic, etc. labeled probes, as known in the art. Alternatively, wells 582 of array 581 can be provided without a signal, with the results reported through subsequent processing.
[0225] When pneumatic pressure is used to move the material within the pouch 510, in one embodiment, an "air bag" may be employed. The air bag assembly 810, a portion of which is located in the Figure 2 AB and 3 show an airbag panel 824 that houses a plurality of inflatable airbags 822, 844, 846, 848, 864, and 866, each of which can be individually inflated (illustratively via a compressed gas source). Because airbag assembly 810 can withstand compressed gas and be used multiple times, it can be made of a tougher or thicker material than the pouches. Alternatively, airbags 822, 844, 846, 848, 864, and 866 can be formed from a series of panels secured together with gaskets, seals, valves, and pistons. Other arrangements are also within the scope of the present invention.
[0226] The success of the secondary PCR reaction depends on the template generated by the multiple first-stage reactions. Typically, PCR is performed using highly pure DNA. Methods such as phenol extraction or commercial DNA extraction kits provide highly pure DNA. Samples processed through pouch 510 may require adjustments to compensate for less pure preparations. PCR can be inhibited by components of the biological sample, which is a potential obstacle. Illustratively, hot-start PCR, higher concentrations of Taq polymerase, adjustments in MgCl2 concentration, adjustments in primer concentration, and the addition of adjuvants (e.g., DMSO, TMSO, or glycerol) can optionally be used to compensate for lower nucleic acid purity. Although purity issues may be more of a concern for first-stage amplification and single-stage PCR, it should be understood that similar adjustments can also be provided in second-stage amplification.
[0227] When the pouch 510 is placed within the instrument 800, the bladder assembly 810 is pressed against one face of the pouch 510 so that if a particular bladder is inflated, the pressure will force liquid out of the corresponding blister in the pouch 510. In addition to bladders corresponding to the many blisters of the pouch 510, the bladder assembly 810 can have additional pneumatic actuators, such as bladders or pneumatically driven pistons, corresponding to the various channels of the pouch 510. Figure 2 and 3Shown are illustrative multiple pistons or hard seals 838, 843, 852, 853 and 865, which correspond to the channels 538, 543, 553 and 565 of the pouch 510, and seals 871, 872, 873, 874, which minimize backflow into the fitting 590. When activated, the hard seals 838, 843, 852, 853 and 865 form a pinch valve to pinch off and close the corresponding channel. In order to confine the liquid in a specific blister of the pouch 510, the hard seal is activated on the channel back and forth between the bubbles so that the actuator acts as a pinch valve to close the channel. Illustratively, in order to mix two volumes of liquid in different bubbles, the pinch valve actuator that seals the connecting channel is activated, and the pneumatic bladder above the bubble is alternately pressurized, forcing the liquid to pass back and forth through the channel connecting the bubbles to mix the liquid therein. The pinch valve actuator can have various shapes and sizes and can be configured to pinch off more than one channel at a time. While pneumatic actuators are discussed herein, it should be understood that other ways of providing pressure to the pouch are contemplated, including various electromechanical actuators such as linear stepper motors, motor-driven cams, rigid paddles driven by pneumatic, hydraulic, or electromagnetic forces, rollers, rocker arms, and in some cases, plug springs. Additionally, in addition to applying pressure perpendicular to the axis of the channel, there are various methods of reversibly or irreversibly closing the channel. These include twisting the bag across the channel, heat sealing, rolling actuators, and various physical valves sealed in the channel, such as butterfly valves and ball valves. Additionally, a small Peltier device or other temperature regulator can be placed near the channel and set at a temperature sufficient to freeze the fluid, effectively forming a seal. Additionally, while Figure 1 The design is suitable for automated instruments featuring actuator elements positioned on each blister and channel, but it is also contemplated that the actuators can remain stationary and the pouch 510 can transition in one or two dimensions, so that a small number of actuators can be used for several processing stations, including sample disruption, nucleic acid capture, first and second stage PCR, and other applications of the pouch 510 such as immunoassays and immuno-PCR. Rollers acting on the channels and blisters may prove particularly useful in configurations in which the pouch 510 translates between stations. Thus, although pneumatic actuators are used in the presently disclosed embodiments, when the term "pneumatic actuator" is used herein, it should be understood that other actuators and other ways of providing pressure may be used, depending on the configuration of the pouch and instrument.
[0228] Other prior art instruments teach PCR within sealed flexible containers. See, for example, U.S. Patent Nos. 6,645,758, 6,780,617, and 9,586,208, which are incorporated herein by reference. However, including cell lysis within a sealed PCR container can improve ease of use and safety, particularly if the sample to be tested may contain biohazards. In the embodiment shown herein, waste from cell lysis and all other steps is retained within a sealed pouch. However, it should be understood that the pouch contents can be removed for further testing.
[0229] Figure 2 An illustrative instrument 800 is shown that can be used with the pouch 510. The instrument 800 includes a support member 802, which can form a wall of a housing or be mounted within the housing. The instrument 800 can also include a second support member (not shown), which is optionally movable relative to the support member 802 to allow for the insertion and withdrawal of the pouch 510. Illustratively, once the pouch 510 has been inserted into the instrument 800, a lid can cover the pouch 510. In another embodiment, both support members can be fixed, with the pouch 510 being held in place by other mechanical means or by pneumatic pressure.
[0230] In the illustrated embodiment, heaters 886 and 888 are mounted on support member 802. However, it should be understood that this arrangement is merely illustrative and that other arrangements are possible. The airbag panel 810 having airbags 822, 844, 846, 848, 864, 866, the hard seals 838, 843, 852, 853, and the seals 871, 872, 873, 874 forming the airbag assembly 808 may illustratively be mounted on a movable support structure that can be moved toward the pouch 510 such that the pneumatic actuator is placed in contact with the pouch 510. When the pouch 510 is inserted into the instrument 800 and the movable support member is moved toward the support member 802, the various blisters of the pouch 510 are positioned adjacent to the various bladders of the bladder assembly 810 and the various seals of the assembly 808 so that activation of the pneumatic actuator can force liquid from one or more blisters of the pouch 510 or can form a pinch valve with one or more channels of the pouch 510. Figure 3 The relationship between the blisters and channels of pouch 510 and the bladder and seal of assembly 808 is shown in more detail in FIG.
[0231] Each pneumatic actuator is connected to a compressed air source 895 via a valve 899. Figure 2Only a few hoses 878 are shown, but it should be understood that each pneumatic accessory is connected to a compressed gas source 895 via a hose 878. The compressed gas source 895 can be a compressor, or alternatively, the compressed gas source 895 can be a compressed gas cylinder, such as a carbon dioxide cylinder. Compressed gas cylinders are particularly useful if portability is desired. Other compressed gas sources are also within the scope of the present invention.
[0232] Assembly 808 is illustratively mounted on a movable support member, although it will be appreciated that other configurations are possible.
[0233] Several other components of the instrument 810 are also connected to a compressed gas source 895. A magnet 850 mounted on the second side 814 of the support member 802 is illustratively deployed and retracted using gas from the compressed gas source 895 via a hose 878, although other methods of moving the magnet 850 are known in the art. The magnet 850 is located in a recess 851 in the support member 802. It should be understood that the recess 851 can be a channel through the support member 802 so that the magnet 850 can contact the blister 546 of the pouch 510. However, depending on the material of the support member 802, it should be understood that the recess 851 does not need to extend all the way through the support member 802, as long as the magnet 850 is close enough to provide a sufficient magnetic field at the blister 546 when deployed and does not significantly affect any magnetic beads 533 present in the blister 546 when retracted. While reference is made to a retraction magnet 850, it should be understood that an electromagnet may be used and that the electromagnet may be activated and deactivated by controlling the current passing through the electromagnet. Thus, while this specification discusses a retraction or retraction magnet, it should be understood that these terms are broad enough to encompass other means of retracting the magnetic field. It should be understood that the pneumatic connection may be a pneumatic hose or a pneumatic air manifold, thereby reducing the number of hoses or valves required.
[0234] The various pneumatic pistons 868 of the pneumatic piston array 869 are also connected to a compressed gas source 895 via hoses 878. Although only two hoses 878 are shown connecting the pneumatic pistons 868 to the compressed gas source 895, it should be understood that the pneumatic pistons 868 are each connected to a compressed gas source 895. Twelve pneumatic pistons 868 are shown.
[0235] A pair of heating / cooling devices, illustratively Peltier heaters, are mounted on the second side 814 of the support 802. A first-stage heater 886 is positioned to heat and cool the contents of the blister 564 for the first-stage PCR. A second-stage heater 888 is positioned to heat and cool the contents of the second-stage blister 582 of the pouch 510 for the second-stage PCR. However, it should be understood that these heaters can also be used for other heating purposes, and other heaters can be used as appropriate for a particular application. Other configurations are also possible.
[0236] When fluorescence detection is required, an optical array 890 may be provided. Figure 2 As shown in FIG, the optical array 890 includes a light source 898, illustratively a filtered LED light source, filtered white light, or laser illumination, and a camera 896. The camera 896 illustratively has a plurality of photodetectors, each corresponding to a second-stage aperture 582 in the pouch 510. Alternatively, the camera 896 can capture an image containing all of the second-stage apertures 582, and the image can be separated into separate fields of view corresponding to each second-stage aperture 582. Depending on the configuration, the optical array 890 can be fixed, or the optical array 890 can be placed on a mover attached to one or more motors and moved to obtain a signal from each individual second-stage aperture 582. It should be understood that other arrangements are possible.
[0237] As shown, computer 894 controls valve 899 of compressed air source 895 and, therefore, all pneumatic devices of instrument 800. Computer 894 also controls heaters 886 and 888, as well as optical array 890. Each of these components is electrically connected, illustratively via cable 891, although other physical or wireless connections are also within the scope of the present invention. It should be understood that computer 894 can be housed within instrument 800 or can be external to instrument 800. In addition, computer 894 can include built-in circuit boards that control some or all of the components, can calculate amplification curves, melting curves, Cps, Cts, standard curves, and other relevant data, and can also include an external computer, such as a desktop or laptop PC, to receive and display data from the optical array. An interface can be provided, illustratively a keyboard interface, including keys for entering information and variables such as temperature, cycle time, etc. Illustratively, a display 892 is also provided. For example, display 892 can be an LED, LCD, or other such display.
[0238] Example 1 - High-density PCR
[0239] In one embodiment, a standard commercial immunofluorescence assay for common respiratory viruses is known to detect seven viruses: adenovirus, PIV1, PIV2, PIV3, RSV, influenza A, and influenza B. A more complete panel illustratively includes assays for other viruses, including: coronaviruses, human metapneumoviruses, rhinoviruses, and non-HRV enteroviruses. For highly variable viruses, such as adenovirus or HRV, it is desirable to use multiple primers to target all branches of the viral lineage (illustratively 4 outer and 4 inner primer sets, respectively). For other viruses such as coronaviruses, there are 4 different lineages (229E, NL63, OC43, HKU1) that do not change from season to season, but have diverged sufficiently to require separate primer sets. The FilmArray® Respiratory Panel (BioFire Diagnostics, LLC of Salt Lake City, UT) includes adenovirus, coronavirus HKU1, coronavirus NL63, coronavirus 229E, coronavirus OC43, human metapneumovirus, human rhinovirus / enterovirus, influenza A, influenza A / H1, influenza A / H3, influenza A / H1-2009, influenza B, parainfluenza 1, parainfluenza 2, parainfluenza 3, parainfluenza 4, and respiratory syncytial virus. In addition to these viruses, the FilmArray® Respiratory Panel also includes three bacteria: Bordetella pertussis ( Bordetella pertussis ), Chlamydia pneumoniae ( Chlamydophila pneumoniae ) and Mycoplasma pneumoniae ( Mycoplasma pneumoniae ). The high-density array 581 is capable of containing such panels in a single pouch 510. Other panels are available for the FilmArray®, each measuring at least 20 pathogens.
[0240] The illustrative second-stage PCR master mix contains the dsDNA binding dye LCGreen® Plus to generate a signal indicative of amplification. However, it should be understood that this dye is illustrative only, and other signals can be used, including other dsDNA binding dyes, as well as fluorescent, radioactive, chemiluminescent, enzymatic, etc. labeled probes, as known in the art.
[0241] The illustrative FilmArray instrument was programmed to call each second-stage reaction positive or negative based on post-PCR melting. The melting curve must produce a melting peak (first derivative maximum or negative first derivative maximum) within a predetermined temperature range to be called positive. It should be understood that this method of calling each second-stage reaction is illustrative only, and calls can be made using real-time amplification data or by other means, as is known in the art.
[0242] Example 2 - Design of Quantification Standards for Multiplex PCR
[0243] In systems such as FilmArray, in which a single multiplex PCR is performed in one reaction chamber, generating a standard curve using a 10-fold dilution of a single reference template is inconvenient because the individual levels cannot be easily distinguished. For example, if a single reference template is added to a single first-stage reaction chamber at concentrations of 10 copies, 100 copies, and 1000 copies, the final concentration of the reference template in that chamber will be 1110 copies, and absent some other labeling, the individual dilutions are indistinguishable. Furthermore, in a two-step multiplex PCR system, a standard curve generated only in the nested second-stage PCR may have limited value for quantitation because the single-plex standard template amplification reaction may not accurately reflect all upstream operations that the sample underwent, or may not amplify with similar efficiency, and therefore, may not reflect the entire process.
[0244] In this illustrative example, different nucleic acid templates (illustratively varying in sequence and / or length), illustratively synthesized quantitative standards, are used to represent different levels in a dilution series. In one illustrative embodiment, the assays used for all synthesized quantitative standards have similar amplification efficiencies and produce the same or similar Cp values at each given dilution point in the multiplex setup. Illustratively, all target assays are optimized for the same performance characteristics, including efficiency, although corrections can be applied to adjust for assay-specific variations in efficiency.
[0245] In one illustrative embodiment, the external and internal amplicon sizes used for the quantitative standards can represent the amplicon sizes used for the quantitative target assay. In addition, the sequence or GC content can be the same or similar between the priming regions. Illustratively, the sequences can be identical, except for at least one internal priming region, which should be different enough to avoid cross-reactivity between internal assays. If the sequences differ only by the internal primer binding region, the same PCR1 primers can be used to amplify all quantitative standards, thus minimizing potential differences in PCR1 assay performance. In addition, if tags are used, the sequences can be identical, and if tags are not used, even slight differences in sequence can be provided for detection, illustratively in a second stage singleplex reaction. However, it should be understood that these parameters are illustrative only, and other means for detecting and controlling amplification efficiency are also possible. It should be understood that the quantitative standards present in the multiplex reaction should be designed to match the reaction parameters, such as Mg 2+ , primer concentrations, Tm and cycling conditions. It will also be appreciated that it is desirable to minimize non-specific amplification of the synthesized templates in a multiplex PCR reaction.
[0246] Figure 5Shown are four kinds of expected internal quantitative standards Cp relative to concentration.In this illustrative embodiment, internal quantitative standard has synthetic sequence.In this illustrative embodiment, for second phase internal reaction, four kinds of quantitative standards all share a kind of common primer, and each has a kind of unique specific primer.They are also designed to share two kinds of external primers for first phase PCR.Therefore, each second phase hole for detecting quantitative standard will use common primer and the primer spotting for this quantitative standard, so that in each such hole, only a kind of quantitative standard should be increased.However, it should be understood that this is only illustrative embodiment, and other configurations are possible.Syn2, Syn3 and Syn4 each have similar amplification efficiency, and are selected for other research.Syn1 performance is different, and is omitted from further work.Therefore, in one embodiment, expectation has multiple quantitative standards, and it has similar amplification efficiency.
[0247] In this illustrative example, amplification is detected using the dsDNA binding dye LCGreen Plus. However, this is merely illustrative, and other dsDNA binding dyes, probes, signals, or other means of detecting amplification are within the scope of the present invention.
[0248] It should be understood that there are various methods for designing quantitative standards with similar amplification efficiencies. In one embodiment, the quantitative standards have identical sequences between internal primers and differ only in the internal primer binding sequence. In another embodiment, the quantitative standards all have substantially the same length and substantially the same GC content. In another embodiment, the sequences have different lengths, but also have different GC contents to compensate. Other methods for designing nucleic acids with similar amplification efficiencies are known in the art.
[0249] In one embodiment, illustratively, when quantitative standards are used in a two-step nested multiplex PCR reaction, the quantitative standards can all use the same outer primers in the first-stage PCR reaction, potentially even sharing the same region around the primers to avoid differences due to secondary structure formation. The quantitative standards can then be distinguished by using different inner primers in each second-stage PCR reaction, or each calibrator has a unique inner primer pair, or, as above, shares an inner primer and has a unique inner primer. The advantage of such an embodiment is that the quantitative standards each bind to their first-stage primers with the same kinetics, and the complexity of the first-stage multiplex PCR reaction can be minimized.
[0250] Although two-step PCR has been mentioned, identical principle can be used for single-step multiplex PCR.In this case, quantitative standard can have different forward or reverse primers or identical forward and reverse primers, and illustratively each has specific fluorescent probe or other identifiable label, and this paper considers such as chemiluminescence, bioluminescence, radioluminescence, electroluminescence, electrochemiluminescence, mechanical luminescence, crystal luminescence, thermoluminescence, sonoluminescence, phosphorescence and other forms of photoluminescence, enzymatic, radioactivity etc. This application is only limited to the number of available detection channels in any system or other method for distinguishing labels, as known in the art. Some labels may need amplification post-processing. In addition, it should be understood that the quantitative standard of label can be used for two-step PCR, wherein identical or different primer sequences can be used, and label is used for detecting in second phase PCR. In such embodiment, the quantitative standard of label optionally can be multiplexed in second phase PCR and distinguished by label.
[0251] Although synthetic quantitative standards are used in this embodiment, it should be understood that the sequences used for quantitative standards can be naturally occurring. For example, if yeast is used as SPC, yeast sequences can be used for one or more quantitative standard sequences. For the fission yeast Schizosaccharomyces pombe, the Tf2 type retrotransposon element / transposon exists in 13 copies, while the ribosomal RNA gene is repeated 47 times. In another embodiment, gene sequences existing in different copy numbers can be used. Illustratively, fungal pathogens have 50 to 200 copies of ribosomal RNA genes / nuclear genomes. These pathogens also have transposons / genomes ranging from 5 to 20 copies. Bacterial pathogens have 1 to 15 copies / genomes, but most have more than 5 copies. Other naturally occurring or synthetic templates can be used, such as bacterial phages for viruses and synthetic particles that can simulate membrane and / or capsid and / or envelope structures. Furthermore, although three quantitative standards are used in many of the examples herein, it should be understood that only two quantitative standards are required to define a linear standard curve, and more quantitative standards may be required in embodiments where a wide range of target concentrations is expected, or a non-linear standard curve is expected. Illustratively, the number of quantitative standards can be selected based on the dynamic range of the system and the requirements of the assay.
[0252] Alternatively, as shown in Example 5, it is also possible to use only one quantification standard in each experimental run (see the sample processing control (SPC) discussed in Example 5) and rely on an input standard curve for quantification, which was previously generated using a range of quantification standards, illustratively using at least 3 quantification standards (designated QS in Example 5) that can be included in the software used for the analysis.
[0253] Example 3 - Multiple Calibration
[0254] Figure 6A Similar to Figure 5 , but only data from three selected calibrator sequences are shown. Figure 6A Linearity and nearly identical amplification efficiencies were demonstrated for three illustrative quantification standards. Figure 6B A composite standard curve generated by combining three points for each of the three quantitative standards at a total of five dilutions is shown.
[0255] Now that an illustrative calibration curve has been generated, a standard curve using an assay-specific reference template for each target assay can be generated. Figure 7 A standard curve generated externally using a well-quantified synthetic reference template of Acinetobacter baumannii was compared to a composite internal standard curve generated using quantitative standards. Here, the three "Syn" templates were pre-mixed before addition to the reaction tubes; Syn4 was added at 10 3 copies were added, Syn2 was added with 10 4 copies are added, and Syn3 is added with 10 5 Copies / reaction were added. The Cp values from these templates were used to generate a composite internal standard curve for each reaction. An external standard curve was generated using a synthetic reference template. The Acinetobacter baumannii reference template was also tested at the same concentration as the "Syn" template. The composite internal standard curve and the Acinetobacter baumannii standard curve were very similar. Similar slopes showed that the efficiency was similar. It is expected that the unknown starting concentration of the Acinetobacter baumannii sample can be predicted using the internal standard curve. However, because the y-intercept transitions between the two curves, the quantification of Acinetobacter baumannii can benefit from a correction factor when using this internal standard curve.
[0256] Therefore, the concentration of the target organism can be calculated using a composite internal standard curve. Note that each internal standard is at a different known concentration and is amplified in the same process as the target organism. This method illustratively employs the cycle threshold (Ct) (or alternatively, the Cp value or other similar methods), which is the number of PCR cycles required to achieve a fluorescence signal above background fluorescence for both the target and the internal standard, as determined experimentally. Other points, such as the first, second, or nth derivative, can also be used, as illustratively taught in U.S. Patent No. 6,303,305, which is incorporated herein by reference in its entirety. As is known in the art, other points can also be used, and any such point can replace Cp or Ct in any of the methods discussed herein. Illustratively, in a two-step multiplex system, the Cp value is determined in a nested second-stage reaction. However, in other embodiments, it should be understood that the Cp value can be determined as appropriate for the amplification system. For example, the Cp value can be determined in a single multiplex reaction or in a subsequent second-stage reaction using oligonucleotide probes, each of which is specific for the quantification standard sequence and has a distinguishable fluorescence signal.
[0257] In an illustrative embodiment where a single internal standard is used, the Cp of the target organism can be used according to the following formula ( ), the concentration of the internal standard and Cp (concentration s , ), and the efficiency of target organisms ( efficiency t ) to calculate the concentration of target organisms.
[0258] concentration t = Concentration s * efficiency t (Cps - Cpt) [Equation 1],
[0259] The subscript s and t represent the internal quantitative standard and target organism, and
[0260] efficiency t = 1 + Efficiency as a percentage / 100 [Equation 2].
[0261] For example, the efficiency variable for a target with 100% amplification per PCR cycle would be equal to 2. Note that this assumes that the efficiency is predetermined and constant across the dynamic range. As discussed above, the efficiencies of the internal calibrators should all be similar, illustratively within 1%, 2%, 5%, or 10% of each other. Similarly, the efficiencies of the targets should each be similar to the efficiencies of the calibrators, illustratively within 2%, 5%, 10%, or 12% of the calibrators. It will be appreciated that for accurate quantification, efficiencies within a narrower range, such as 1%, 2%, or 5%, are desirable. However, for semi-quantitative or "binning" results (see below), greater variations in efficiency can be tolerated.
[0262] When two or more quantitative standards are used, a standard quantitative curve can be generated, illustratively using a least squares regression line fit to the data for each quantitative standard ( Ct , log 10 (concentration) ),like Figures 6A-6B As shown. Illustratively, the regression fit has the following form:
[0263] log 10 (concentration) =( Cp - b) / a [Equation 3],
[0264] in b is the intercept and represents the log 10 (concentration) The Cp value when is zero, and a is the slope, which represents how much Cp changes with a single unit change in template concentration (a function of efficiency). Given a calculated Cp value for an unknown target, this formula gives the log 10 The target concentration of the unit. As needed, other algorithms or equations can be applied to improve quantitative precision and accuracy. These can be included in the required adjustment for platform specificity, matrix specificity or determination specific deviation in extraction and / or amplification. These can also include the algorithm for the determination efficiency difference in the separate steps that can explain any multi-step amplification process. In some embodiments, the quantitative standard curve can be nonlinear, or can be linear only within a certain dynamic range. Illustratively, if there is a concentration-dependent variable slope, an S-shaped dose-response curve can be used. Other nonlinear curves are also within the scope of the present invention.
[0265] Based on the Cp values observed for the target and the regression equation for the standard curve generated using the internal quantitative standard, the above method can be used for target organisms with unknown concentrations. Ideally, all targets quantified using this method should have assays that have equivalent or similar PCR efficiencies to the internal quantitative standard assays. However, there may be some variation in the slope or intercept of the target assay standard curve. Given that the target assay may have different amplification characteristics than the internal standard, an assay-specific correction factor can be used to adjust for system assay-specific bias to improve the accuracy of the calculated concentration of the unknown target. Illustratively, when a linear quantitative curve is used, a Correction can be performed using a correction factor (which changes the slope) that indicates different assay-specific efficiencies, or b Corrections can be made for the lack of optimal PCR conditions for a particular target that contribute to a delay in target Cp. Both corrections can be used as appropriate. In another example, illustratively, when nested PCR is used, the observed Cp in PCR2 may be delayed due to variations in the efficiency of PCR1. b The differences in may be a result of the overall results of the PCR 1 assay. In this case, a correction factor can be calculated as a function of the Cp value or can be a constant depending on the desired quantitative accuracy.
[0266] For example, a set of control experiments can be run with known concentrations of the target organism. If multiple replicates at a single concentration are used, an assay-specific correction factor can be calculated as Log 10 The average difference between the known and calculated concentrations expressed in units. Illustratively, to obtain the corrected log concentration of the target organism, an assay-specific correction factor can be added to the log concentration of the target organism (as calculated above by the internal quantitative standard method). Each target sequence in a multiplex assay will illustratively have its own calibration (or no calibration at all if the composite standard curve is very similar).
[0267] An experiment was set up to compare the quantification of the target organism calculated by the assay-specific standard curve with the quantification calculated by the composite internal standard curve. In this experiment, 10-fold serial dilutions of known amounts of A. baumannii genomic nucleic acid were multiplexed with the internal quantitative standard in benchtop reactions. Figure 7 An external assay-specific standard curve was established as described. Figure 8A Results are shown using a composite standard curve from a quantification standard and an external standard curve specific for Acinetobacter baumannii. If the composite standard curve is used without correction, there is a clear systematic overquantification of the target organism (Acinetobacter baumannii) (~0.5 log copy units), whereas when a 0.5 log copy unit correction is applied, the corrected assay-specific standard curve gives a fairly accurate estimate of the Acinetobacter baumannii titer in the sample. Figure 8BIt is shown how the system can be corrected for quantitative bias by applying the average assay-specific correction factor generated as described above to the amounts calculated by the internal standard curve method. Similar corrections can be performed for each assay in a multiplex reaction.
[0268] In many embodiments, absolute quantification is not necessary, and semi-quantitative results may be sufficient. Results may be reported as absolute concentrations (with or without systematic error (illustratively 95% prediction intervals)), or may be binned into one of a number of ranges, illustratively reporting "high," "medium," or "low" concentrations, each covering one or more orders of magnitude. It will be understood that the number of bins may vary, as appropriate for the particular assay being used, and any number of bins may be used. Additionally, the binning range (orders of magnitude or other metric) for the semi-quantitative results may be adjusted, as appropriate for the particular embodiment.
[0269] Example 4 - Methods for normalizing and quantifying sequencing data
[0270] Next-generation sequencing (NGS) methods can be used in addition to or in place of the assay methods and systems described in Examples 1-3. NGS can be used, for example, to detect, identify, and quantify potentially pathogenic organisms in a sample. In another embodiment, NGS can be used as a so-called 'comparator' to confirm the performance of another assay, such as one of the assay methods and systems described in Examples 1-3. In a specific embodiment, NGS includes PCR followed by next-generation sequencing.
[0271] Modern sequencing instruments (e.g., NGS instruments) can generate large amounts of data that can be computationally cumbersome to analyze. Additionally, the sequencing process (e.g., sample type, sample preparation, amplification, and sequencing) and the data obtained can include many confounding factors that can make the data difficult to analyze and / or compare from experiment to experiment, laboratory to laboratory, and so on. For example, samples may not be consistently prepared for sequencing due to human (i.e., random) and systematic errors. In another case, samples cannot be consistently sequenced due to the presence of nucleic acids or multiple length ranges from multiple organisms at various concentrations. For example, sputum is a particularly abundant sample type that can have a range of approximately 0-10 13 The bioburden of the sample is approximately 10 organisms / ml, of which approximately 10 3 -10 9organisms / ml is quite typical. Other clinical samples (such as stool or blood culture) may have similar bioburdens. Similarly, clinical samples can include large amounts of host DNA (e.g., human DNA). In addition, sequencing libraries from multiple samples can be pooled together before sequencing, which can shorten the time required for sequencing because multiple samples can be analyzed simultaneously, but it can drastically increase the amount of data collected in a single sequencing run. In addition, due to the above factors, the individual samples in the sample library cannot be prepared consistently. The data from each pooled sample is usually separated in the dataset for individual alignment and analysis.
[0272] To address these and other similar issues, it was discovered that a known quantity of reference DNA sequence (referred to herein as quantitative standard material (QSM) or internal quantitative standard) could be added to the sample intended for sequencing prior to sample preparation (e.g., prior to nucleic acid extraction and amplification) and prior to sequencing; the internal quantitative standard is then used throughout all sample processing and sequencing steps. Because the internal quantitative standard is added at the beginning of sample processing, the number of sequencing reads for the sequencing standard (i.e., the number of molecules in the sequencing standard) accurately reflects all operational and system losses that the sample may have undergone, such as, but not limited to, sample preparation, amplification, nucleic acid recovery, purification, and sequencing. Similarly, because the internal quantitative standard is added separately to each sample before extraction, library preparation, and pooling, the standard added to each sample can accurately reflect all processing steps that each sample has undergone. In one embodiment, the internal quantitative standard added separately to each sample in the sample pool is identical. In one embodiment, each amplified sequence in the pool has its own sample-specific identification sequence (e.g., a DNA barcode), which allows nucleic acids from each sample to be tracked, separated, and normalized by sample.
[0273] The internal quantitative standard (i.e., reference DNA sequence) can be essentially any natural or synthetic nucleic acid sequence, provided that it can be distinguished from other nucleic acids in the sample that can be sequenced. For example, the internal quantitative standard can be a naturally occurring or modified plasmid, a naturally occurring or synthetic linear DNA fragment, or the like. Since the efficiency of sample preparation, DNA amplification, purification, and sequencing is affected by the size of the DNA, the internal quantitative standard should be within a size range close to that of the unknown sequence in the sample. For example, if the DNA fragment to be sequenced has a size within the range of 200-500 base pairs, the internal quantitative standard should also have a size within the range of 200-500 base pairs. In addition, the composition (e.g., relative GC content) of the internal quantitative standard should also be relatively similar to that of the DNA to be sequenced.
[0274] In one embodiment, a known quantity of an internal quantitative standard can be added to the sample to be sequenced. If multiple samples are to be sequenced, the internal quantitative standard can be added separately to each sample. In one embodiment, the known quantity of the internal quantitative standard is determined based on the assay. For example, the number of internal quantitative standards sufficient for an assay can depend on factors such as, but not limited to, the desired assay resolution, the efficiency of nucleic acid extraction, the concentration range of the nucleic acids to be sequenced, the prevalence of genetic mutations to be detected, or the desired sequencing read depth. In one embodiment, a known quantity of the internal quantitative standard can be added in an amount that is within the linear range of the assay, e.g., above the lower limit of detection (LOD) and below the maximum concentration expected for the assay. For example, in specific embodiments discussed herein, the linear range of DNA copies / ml that can be sequenced and distinguished in a sequencing assay is about 10 2 to about 10 7 -10 8 copies / ml, and in this case, the internal quantitative standard can be measured at approximately 10 4 -10 6 As described in more detail below, the amount of internal quantitation standard added and sequencing read depth correlated with the determined LOD.
[0275] In order to sequence the nucleic acids in a sample, RNA or DNA is extracted from the sample tissue / cells and fragmented. RNA is converted into cDNA by reverse transcription. Preparing a sample containing an internal quantitative standard for sequencing includes generating DNA fragments, which are converted into a 'library' by annealing or connecting sequencing adapters, which include specific sequences designed to interact with the sequencing platform and sample-specific identification sequences (e.g., DNA 'barcodes'), which can be used to identify sequencing data derived from a specific sample. If two or more samples are to be combined and sequenced in parallel, each sample has its own sample-specific identification sequence or 'barcode', which can be used to separate the data derived from each sample. This procedure is primarily compatible with Illumina sequencing technology. However, it should be noted that almost all of the principles discussed herein can be applied to NGS platforms developed by Life Technologies, Roche, Pacific Biosciences and others with minimal modifications.
[0276] The next step involves sequencing the DNA. The exact sequencing method depends on the sequencing platform, but all modern massively parallel sequencing technologies are somewhat similar and generate similar data.
[0277] After sequencing, the number of reads originating from the internal quantitative standard is then counted in each sample, and the NGS dataset for each sample can be normalized. As a preliminary check, there should be a linear relationship between the amount of internal quantitative standard spiked into the sample and the number of sequencing reads observed for the internal quantitative standard. Additionally, the minimum number of sequencing reads for the internal quantitative standard ('F') should be recorded. Normalization (1) applies data acceptance / rejection criteria to the sequencing dataset based on the presence of a minimum number of sequencing reads for the internal quantitative standard for a sample, and (2) ensures that essentially the same limit of detection (LOD) is applied to all unknown nucleic acids in the sequencing assay. If samples are pooled, normalization ensures that all samples are read to sufficient depth, that all samples retain the same number of QSM read pairs, and that essentially the same limit of detection (LOD) is applied to all unknown nucleic acids in all samples in the sequencing assay.
[0278] Assuming a linear relationship exists and that 'F' internal quantification standards / QSM reads are recorded, sequencing data can be processed according to Equation 4 to retain normalized sequencing read numbers:
[0279] NORM = Number of sequencing reads * (F / Observed number of sequencing reads derived from the internal quantification standard) [Equation 4]
[0280] 'F' is a fixed, user-defined minimum expected number of sequencing reads for the internal quantification standard that is specific to the assay. For example, 'F' is a number of sequencing reads set to ensure sufficient sequencing read depth to detect nucleic acids at or near the limit of detection (LOD) of the sequencing assay.
[0281] NORM may represent a subset of the sequencing data and is the number of sequencing reads saved for further analysis and quantification. As can be seen from Equation 4, if 'F' and the observed number of sequencing reads derived from the internal quantification standard are the same, the number in the brackets is reduced to '1' and NORM is then equal to the number of sequencing reads recorded. On the other hand, if the number of internal quantification standard sequencing reads in the dataset is greater than 'F', the number in the brackets is <1, and the NORM equation reduces the data size to account for the over-reading issue and ensures that the same LOD is applied across all samples. If the number of internal quantification standard sequencing reads in the dataset is less than 'F', the data from that sample may be rejected and the sample can be submitted for further sequencing until 'F' internal quantification standard reads are recorded.
[0282] The number of QSM reads retained, 'F', is a parameter selected during assay development and is related to read depth and the limit of detection (LOD) of the assay. This effect is shown in Table 1 below.
[0283]
[0284] For example, if the internal quantitation standard is 5x10 5 If the internal quantitative standard / QSM is spiked into the sample at a concentration of 5x10 copies / ml, then the number of QSM reads should be recorded to ensure that all unknown nucleic acids at or above the LOD are fully sequenced and detected. This is called sequencing read depth. For example, if the internal quantitative standard / QSM is spiked into the sample at a concentration of 5x10 5 If 10 copies / ml are added to the sample and 'F' is 500, the LOD of the assay for detecting the unknown nucleic acid is not less than 10 3 copies / ml (usually 10 3 -10 4 copies / ml); if F is 5000, the LOD of the assay is not less than 10 2 copies / ml (usually 10 2 -10 3 copies / ml); or if 'F is 50,000, the LOD of the assay is not less than 10 copies / ml (usually between 10-10 2 copies / ml). It can therefore be seen that the LOD of the assay can be lowered by increasing the read depth of the assay (i.e., increasing 'F'). However, setting 'F' arbitrarily high does not make sense because the sequencing burden also increases. For example, if the LOD of the assay is no less than 10 copies / ml ('F' = 50,000), and the most prevalent species is expressed at 10 9 If 10 copies / ml are present (which is not uncommon), 9 The most prevalent target entering at 10 copies / ml was sequenced 100 million times to achieve a 1 to 10 2 Only 1-10 (e.g., 4) reads of a target present in the range of 10 copies / ml may be impractical.
[0285] In the examples described herein, 'F' is typically set to 5000 QSM readings, as this is sufficient to reproducibly ensure that the 3 -10 70 copies / ml or more of the nucleic acid. For more stringent or less stringent detection, 'F' can be set differently. Therefore, in one embodiment, the sequencing read depth (i.e., 'F') can be as low as 100-1000 QSM reads, or in the range of about 1000 internal quantitative standard sequencing reads to about 100,000 internal quantitative standard sequencing reads, preferably about 2000 internal quantitative standard sequencing reads to about 75,000 internal quantitative standard sequencing reads, more preferably about 5000 internal quantitative standard sequencing reads to about 50,000 internal quantitative standard sequencing reads, or most preferably at least 5000 internal quantitative standard sequencing reads.
[0286] refer to Figure 9 The potency and 'F' values of the internal standards are further shown. Figure 9A hypothetical scenario for sequencing data sets for three samples, A, B, and C, is shown. Each sample is spiked with 10 copies of QSM, and each sample contains unknown amounts of nucleic acids X and Y. 20,000 sequencing reads were collected for each sample. In sample A, the sample clearly contains neither X nor Y, and all 20,000 reads were attributed to QSM. Of the 20,000 reads for sample B, 18,182 were attributed to QSM, and 1,818 were attributed to X, while no reads were attributed to Y. Sample C shows a different scenario. Of the 20,000 reads for sample C, 20 were attributed to QSM, 2 to X, and 19,978 to Y. Based on the initial read attribution, it can be seen that sample C has an extremely high concentration of Y and relatively few X. Nevertheless, sample C has fewer than 5,000 QSM reads, so the sample must be submitted for additional sequencing until at least 5,000 QSM reads are recorded. However, in the case of sample C, this does mean that Y will be sequenced almost 5 million times in order to achieve minimum QSM coverage. This may seem cumbersome compared to industry standards, which may set a minimum total read count for each sample (e.g., 20,000 reads, 200,000 reads, or 2 million reads) and then assume that if this total is reached, all detectable nucleic acids are sequenced. However, the method described herein ensures that the same read depth is achieved across all three samples A, B, and C, ensuring that sequencing has an equal chance of detecting rare nucleic acids in each sample. In the case of sample C, even with 2 million total reads recorded, the sample is still "read under-represented," which seems like a sufficient number. However, in this case, 2 million reads still yields an insufficient number of QSM reads, and the sample is still under-represented for the assay's LOD. However, when sequencing actual clinical samples with unknown amounts of nucleic acid, without internal quantitative standards, it may be impossible to know at what point sufficient read depth is achieved. That is, until sufficient QSM reads are attributed, unknown nucleic acids in a sample (e.g., X and Y in sample C) are not given an equal chance of being detected at the LOD. Furthermore, with 2 million or even 200,000 reads, X and Y should be detectable, but without QSM, it may not be possible to quantify them, and it may be impossible to determine when sufficient read depth has been reached to detect all rare nucleic acids in the sample. In samples A and B, more than 5,000 reads are attributed to QSM, thus meeting the 'F' requirement and passing. Therefore, based on the aforementioned hypothetical scenario, the internal quantification standards discussed herein are powerful tools.
[0287] Further return Figure 9, 'F' is the fixed, assay-defined minimum expected number of sequencing reads from the internal quantification standard, and NORM [Equation 4] generates a subset of sequencing reads that are saved for further analysis and quantification. If 'F' and the observed number of sequencing reads from the internal quantification standard are the same, then NORM is equal to the recorded number of sequencing reads. In contrast, if a sample is 'over-read' (e.g., Figure 9 For samples A and B in the same sequencing run, NORM reduces the data size to account for over-reads and ensures the same LOD is applied across all samples. For example, in sample A, 20,000 reads are attributed to QSM; this sample is 'over-read' for QSM, and NORM = 20,000 * (5000 / 20,000) = 5000. This means that only a quarter of the data in sample A is saved for further analysis. Similarly, in sample B, NORM = 20,000 * (5000 / 18182) = 5500. Sample C is 'under-read' for QSM, so it is not normalized or further analyzed and is submitted for additional sequencing.
[0288] Now refer to Figure 10 and 11 , showing the raw read numbers and QSM reads for a set of clinical samples. Figure 10 and Figure 11 The samples shown in are samples that were pooled for sequencing after spiking with QSM; they were then analyzed using sequencing library preparation and sequencing data collection. Figure 10 and 11 32 samples are shown that were pooled for sequencing. Each sample was spiked with its own QSM, and each sample had its own unique tracking sequence (e.g., DNA barcode) that was annealed, ligated, or otherwise bound to all unknown nucleic acids in the sample and to the QSM for each sample during sample preparation. These unique tracking sequences were bound to all nucleic acids in each different sample and throughout all steps of sample preparation and sequencing, and the unique tracking sequences allowed the data originating from the nucleic acids in each sample to be separated from the pooled sequencing data set for data analysis. Although in Figure 10 and 11In the embodiment shown in , 32 samples are merged and ordered together, but in fact more or less samples can be merged.For example, assuming that there are enough distinguishing abilities of barcode and enough reading depth for all samples, then 1000 or more samples can be merged for a sequence order.In one embodiment, can after preparing sample and before order-checking, merge 2-1000 sample, preferably can after preparing sample and before order-checking, merge 2-500 sample, more preferably can after preparing sample and before order-checking, merge 2-100 sample, more preferably can after preparing sample and before order-checking, merge 2-50 sample, or most preferably can after preparing sample and before order-checking, merge 2-32 sample.
[0289] exist Figure 10 and 11 In , 32 samples were combined and sequenced simultaneously. The QSM added to each sample and the unique tracking sequence associated with each sample allowed the sequences originating from each individual sample to be isolated and analyzed separately. Figure 10 It is interesting to note the wide range of sequencing reads associated with each sample. For example, sample 15 accounts for more than 1 / 3 of all sequencing reads in the dataset, while some samples (such as sample 13) have less than one-tenth of the reads of sample 15. Figure 11 QSM readings are shown. Three samples 12, 27 and 29 have less than 5000 QSM readings and fail the initial QSM inspection. Sample 23 has almost no enough QSM readings to pass through, and samples such as 6, 9, 10 and 14 are substantially over-read for QSM. Interestingly, even if sample 15 has a significant original number of reads, its QSM reading level is not abnormally high. However, due to normalization as described herein, all samples passed through are given equal weight. That is, samples such as 23 with only ~ 5000 QSM readings are not significantly trimmed during the normalization process, but data from samples such as 6, 9, 10 and 14 may be trimmed because a large amount of QSM readings in these samples indicate that the data from these samples are significantly over-presented in the data set. Therefore, it can be seen that the normalization process described herein ensures that all samples are given equal weight, and identical LOD is applied to all samples.
[0290] In one embodiment, unknown nucleic acid reads and internal quantitative standard reads can each be separately normalized by the same ratio ALPHA according to Equation 5:
[0291] ALPHA = F / observed number of sequencing reads derived from the internal quantification standard [Equation 5]
[0292] This is Figure 12The data for a sample includes both QSM and non-QSM reads. When these are normalized together (and scaled down together), the final number of QSM reads in the normalized dataset may not be exactly 'F' (e.g., 5000)—instead, for example, there may be ~4970-5030 QSM reads in the binned dataset. This is because the sampling process used to prepare the NORM dataset is randomized, and the underlying probability model for random sampling events predicts that the random seed used to prepare the NORM dataset may not have drawn exactly 'F' QSM reads from the entire dataset. Alternatively, the QSM and non-QSM data can be separated and binned, and normalized with ALPHA, which yields exactly 'F' QSM reads (e.g., 5000) in the QSM bin and an appropriate number of reads in the non-QSM bin. Specifically, pseudo-random stratified subsampling can be used to retain an appropriate number of reads originating from QSM in each sample's dataset. This normalization ensures that sequencing depth and LOD are consistent across samples and reduces the computational burden of subsequent analysis steps. To perform QSM alignment and normalization, QSM reads are separated from non-QSM reads by alignment according to well-established procedures known in the art. This separates the read pairs into a set derived from QSM and a set not derived from QSM. The counts of QSM read pairs are used to calculate a subsampling fraction for each sample, which is given by Equation 5. The subsampling fraction and the counts of non-QSM read pairs are used to calculate the number of non-QSM read pairs for subsampling of each sample. Pseudorandom read subsampling randomly selects a specified number of non-QSM read pairs based on a seed. The seed ensures that although the subsampling is random, it is also reproducible.
[0293] In one embodiment, a known amount of an internal quantitation standard added to a sample can be used to calculate the input quantity (IQT) of unknown nucleic acids in the sample after normalization by Equation 6.
[0294] IQT = Normalized unknown nucleic acid sequencing read counts attributed to unknown nucleic acids * (input number of internal quantitation standards / F) [Equation 6]
[0295] For example, returning a reference Figure 9 For sample B, Equation 6 can be used to calculate the IQT of unknown nucleic acid X as follows: After normalization, there are 500 reads attributed to X, and 'F' is set to 5000, so IQT = 500 * (10 / 5000) = 1. Many samples have more than one unknown nucleic acid. Therefore, the input quantities (IQTi, IQTj, IQTk…IQTn) of the multiple unknown nucleic acids in the sample can be calculated after normalization using Equation 7.
[0296] IQTn = Normalized number of unknown "n" nucleic acid sequencing reads attributed to the "nth" unknown nucleic acid * (input number of internal quantitation standards / F) [Equation 7]
[0297] After counting the normalized reads attributed to each of the different unknown nucleic acids (eg, by alignment, pseudo-alignment, or assembly methods), these normalized reads can be used to determine the input concentration of each of the unknown nucleic acids.
[0298] The same principle also applies to pooled samples. Each pooled sample has its own QSM and its own set of unique sample-specific identification sequences, making it possible to distinguish and separate the sequencing data from each sample in the pool. Therefore, normalization can be applied separately to the sequencing data originating from each sample in the pool. Therefore, each sample in the sample pool has its own (1) data acceptance / rejection criteria, and (2) essentially the same limit of detection (LOD) is applied to all unknown nucleic acids in each sample. This is important because, as Figure 10 and 11 As shown in , some samples may be very concentrated, and some may have only nucleic acids at or near the LOD, and some samples may be over-read or under-read for the QSM, while some samples may be read at or near the acceptance criteria for the QSM. For example, as shown for samples 6, 9, 10, and 14, the QSM of some samples may be over-represented in the combined dataset, while other samples such as 15, 17, 26, and 32 have QSM values at or near the acceptance criteria. Evaluating the QSM separately and applying normalization to each sample is powerful because (1) it ensures that each sample has sufficient read depth (some individual samples can be rejected without having to reject the entire dataset), (2) it ensures that very concentrated samples do not overwhelm low concentration samples, (3) it ensures that the same LOD is applied to all samples, and (4) it ensures that a comparable amount of data is objectively evaluated from each sample for normalization and quantification. Therefore, Equation 7 can be used to calculate the input quantity of multiple nucleic acids in multiple samples in a library.
[0299] In some embodiments, preparing a sample for sequencing can include specific procedures. In one embodiment, preparing a sample for sequencing can include sample lysis, recovering nucleic acids from the lysis products, and optionally purifying the recovered nucleic acids, and introducing sequencing adapter sites and sample-specific identification sequences into the region of the nucleic acid to be sequenced. In one embodiment, the sequencing adapter sites can include regions for attaching nucleic acids to a platform for sequencing (e.g., a flow cell) and sequencing primer binding sites. Figure 13An example of a sample preparation workflow is shown in FIG. In some embodiments, an internal quantitative standard / QSM can be added before or after lysis, e.g., depending on the likelihood that the lysis may be damaged or fragmented by the lysis protocol. For example, if sample lysis involves bead beating or sonication, the QSM can be added after lysis but before recovery and purification of the nucleic acid. If chemical lysis or mild mechanical lysis is employed, the QSM can be added before lysis.
[0300] In one illustrative embodiment, it is believed that an internal quantitative standard can be used to estimate the true input concentration of unknown nucleic acids in the original sample by accounting for losses during sample preparation and sequencing. That is, the internal quantitative standard is used throughout all steps of the analysis. Therefore, any losses are systematic, and by preparing the same losses due to sample extraction, sequencing, etc. for each sample throughout the process, the sequencing concentration can be used to back-calculate the true input concentration.
[0301] In one embodiment, introducing sequencing adapter sites and sample-specific identification sequences can include one of the following: (1) amplifying the nucleic acid to be sequenced in an amplification reaction using target-specific primers that contain sequencing primer binding sites and sample-specific identification sequences, or (2) fragmenting the nucleic acid to be sequenced and ligating to the fragmented nucleic acid sequencing-specific adapters that contain sequencing primer binding sites and sample-specific identification sequences.
[0302] In one embodiment, amplifying nucleic acids to be sequenced can include performing a first multiplex PCR reaction using target-specific primers having custom overhangs, performing a first nucleic acid purification, performing a second PCR reaction using sequencing adapter primers that anneal or ligate to the custom overhangs introduced in the first PCR, and performing a second nucleic acid purification. In one embodiment, the sequencing adapter primers are target-independent and comprise a sequencing primer adapter site and a sample-specific identification sequence.
[0303] like Figure 14 As shown in , the total bioburden in a typical sample can vary widely – between approximately 0 and 10 13 organisms / ml (usually about 10 2 -10 7 This can result in a very wide range of nucleic acid copies / ml input within the sequencing reaction. In one embodiment, the amplification reaction can be used to suppress the upper limit of the reaction so that the sample submitted for sequencing effectively has a nucleic acid number within a narrower relative dynamic range, for example about 10 2 –10 7 In one embodiment, the methods described herein may include limiting one or more of the target-specific primer concentration or cycle number in the first multiplex PCR reaction to a concentration greater than about 107 The platform amplification of nucleic acids present at a concentration of copies / ml and maintained in the exponential amplification phase is less than about 10 7 copies / ml of nucleic acid. Figure 15 For example, the upper LOD of the sequencing assay is about 10 for each nucleic acid. 2 – 10 7 The range of copies is likely linear. Above about 10 7 It is desirable to avoid unnecessary over-sequencing of nucleic acids at such concentrations that are outside the linear range of the assay, and such nucleic acids may be reported as having a concentration >10 7 The input concentration of nucleic acid copies / ml is about 10 2 – 10 7 All nucleic acids are kept within the linear range of the assay and are reported with their true input quantity by the corresponding normalized sequencing reads. In some embodiments, it may be desirable to suppress the lower end of the dynamic concentration range by limiting the number of amplification cycles. In one embodiment, the LOD of the assay can indicate a lower reporting range, and nucleic acids below the LOD can be reported as 'not detected'.
[0304] In one embodiment, the input quantity of nucleic acid in the sample can instead be reported semi-quantitatively or as a 'binning' range. In many embodiments, absolute quantification is not necessary, and semi-quantitative results may be sufficient. In addition, reporting a binning range may be simpler and may allow the report to take into account target-specific errors in the data. Results can be reported as absolute concentrations (with or without systematic error (illustratively, a 95% prediction interval)), or can be binned into one of a plurality of ranges, illustratively reporting a "high," "medium," or "low" concentration that each covers one or more orders of magnitude.
[0305] Likewise, the input quantity of nucleic acid in a sample can be reported in digital bins. Examples are shown in Table 2 below.
[0306] warehouse Semi-quantitative reporting level 10^3 <10^3.5 10^4 10^3.5 - ≤10^4.5 10^5 10^4.5 - ≤10^5.5 10^6 10^5.5 - ≤10^6.5 >10^7 ≥10^6.5
[0307] Table 2.
[0308] Each bin encompasses values spanning one log; a bin is defined as having a range of 0.5 log on either side of the bin label (i.e., a 10^5 bin contains quantities from 10^4.5 to 10^5.5, etc.). The additional 0.5 log on either side of the bin range accounts for measurement variability seen near the bin boundaries. The exceptions shown in Table 2 are the lowest bin, where any value below 10^3.5 can be reported as "not detected," and the highest bin, where any value equal to or greater than 10^6.5 can be reported as >10^7. However, it should be understood that these lower and upper reporting values are merely illustrative, and that different lower and upper reporting value ranges may be used depending on factors such as, but not limited to, the LOD and dynamic range of a given assay. Two standard deviations, or 0.5 log, from the bin boundaries is the theoretical range that captures 95% of the data points that do not center on the bin boundaries. It should be understood that the number of bins can vary, as appropriate for a specific assay, and any number of bins can be used. Additionally, the binning range (order of magnitude or other measure) for semi-quantitative results can be adjusted, as appropriate for a particular embodiment.
[0309] Example 5: Methods for performing comparator studies
[0310] When a diagnostic test provider (e.g., BioFire Diagnostics) develops a new commercial diagnostic test, the test provider is often required by the FDA to perform a comparator study to provide confirmation that the new commercial test delivers the results the provider claims it does. Typically, the comparator assay uses an independent and recognized gold standard technology (e.g., standard of care (SOC) culture method) to check the results of the new assay. However, current SOC culture methods may not provide highly accurate or consistent detection and quantification of bacterial targets in clinical samples. This is due to a number of factors, including the often-increased sensitivity of molecular assays over culture, the fact that culture can only detect viable organisms, and the challenges of establishing unique colony morphologies from complex, organism-rich sample types.
[0311] For molecular assays such as the FilmArray system as described herein, which uses PCR and fluorescence detection to detect the presence of organisms in samples, Sanger sequencing is an alternative comparator to recognized SOC culture methods. However, as molecular assays become more complex, the sequencing burden on traditional Sanger sequencing becomes larger. Similarly, in the case of quantitative requirements by commercial tests, comparator sequencing assays may also need to provide quantitative or semi-quantitative results.
[0312] Described herein is a high-throughput alternative to conventional PCR and bidirectional sequencing comparator methods previously used in the art. Preferably, the sequencing method employed in comparator studies is next-generation sequencing (NGS) or a similar massively parallel sequencing technology. For targets on diagnostic panels requiring semiquantitative bioburden reporting (e.g., bacteria, viruses, fungi, etc.), or antibiotic resistance markers associated with some bacteria, a preferred embodiment is the development of quantitative or semiquantitative molecular reference methods for use as comparators during clinical trials. The novel molecular reference methods described herein exploit the inherent digital and massively parallel nature of NGS. Although Illumina NGS is used in the specific embodiments discussed herein, the principles described herein can be applied to other sequencing platforms with minimal modification.
[0313] In one embodiment, a method for performing a comparator study is described. The method for performing a comparator study may include providing a first assay comprising multiplexed amplification and detection of one or more target nucleic acids, the first assay having a limit of detection (LOD), and providing a second assay, different from the first assay, for confirming the detection and the LOD of the first assay. The second assay may include preparing a sample for sequencing comprising at least one internal quantitative standard, sequencing to generate a sequencing dataset for the sample, wherein the sequencing dataset comprises sequencing reads observed from the target nucleic acid and the internal quantitative standard, counting the number of sequencing reads in the sequencing dataset that originate from the target nucleic acid and the internal quantitative standard, and normalizing the sequencing dataset, wherein the normalization (1) applies data acceptance / rejection criteria to the sequencing dataset based on the presence of a minimum number of sequencing reads for the internal quantitative standard for the sample, and (2) ensures that substantially the same limit of detection (LOD) is applied to all unknown nucleic acids in the sequencing assay. The LOD of the second assay is preferably substantially the same as the LOD of the first assay.
[0314] The results of the second assay can be compared to the results of the first assay.Preferably, the organisms and numbers of organisms detected in the first assay should be consistent with the detection and numbers reported by the second assay.
[0315] In one embodiment, a first assay may include adding one or more internal quantitative standards to a sample and performing a quantitative two-step amplification on the sample. In one embodiment, the quantitative two-step amplification in the first assay includes: amplifying the sample in a first-stage multiplex amplification mixture comprising a plurality of target primers, each configured to amplify a different target that may be present in the sample, and at least one quantitative standard primer configured to amplify an internal quantitative standard nucleic acid; dividing the first-stage amplification mixture into a plurality of second-stage individual reactions, each comprising at least one primer configured to further amplify one of the different targets that may be present in the sample, and a second plurality of second-stage individual reactions, each comprising at least one primer configured to further amplify one of the internal quantitative standard nucleic acids; and subjecting the plurality of second-stage individual reactions to amplification conditions to generate one or more target amplicons and a plurality of quantitative standard amplicons, each quantitative standard amplicon having an associated quantitative standard Cp. Each target nucleic acid has a crossover point (Cp), and each internal standard has a known concentration and a known quantitative standard Cp in the first assay.
[0316] In one embodiment, the input concentration of the unknown nucleic acid in the first assay may be determined by comparison with a quantitative standard of known input concentration. In one embodiment, the method may further comprise generating a standard curve from the quantitative standard Cp and using the standard curve to quantify each of the one or more target nucleic acids. In one embodiment, each target nucleic acid is quantified using a standard curve generated using a least squares regression line fit to Equation 3.
[0317] Log 10 (Concentration) = (Cp-b) / a [Equation 3]
[0318] In Equation 3, Cp is the intersection point measured for each target, and b, the intercept, represents the logarithm of the target. 10 is the Cp value when (concentration) is zero, and a is the slope representing how much Cp changes with a single unit change in concentration.
[0319] Referring again to the second assay, the data from the second assay can be normalized as described above in Example 4, and the input quantity of the unknown nucleic acid can be determined. Importantly, a read depth threshold checkpoint based on the expected number of predetermined QSM reads precedes target-specific analysis. Samples with insufficient read depth are flagged for additional sequencing, while samples with sufficient read depth are normalized. This ensures that the quantitative value of the reads is the same across each sample, and that target detection is independent of the total target load.
[0320] For different limits of detection, a predetermined number of QSM readings, 'F', can be selected to match the first assay. For example, if the QSM is 5x10 5 copies / ml were spiked into the sample, and 'F' was 5x10 3 , the LOD for detecting the target nucleic acid in the second assay is about 10 2 – 10 3 copies / ml. Similarly, if 'F' is 5x10 4 , the LOD for detecting the target nucleic acid in the second assay is about 10 1 – 10 2 In general, the greater the magnitude of 'F', the lower the LOD of the sequencing assay. This is because, as described in detail in Example 4, the greater the magnitude of 'F', the greater the sequencing depth of the assay. The QSM incorporation levels (e.g., 5x10 5 copies / ml), 'F' (e.g., 'F' = 5000), read counts, and the relationship between the semi-quantitative binned target report on representative normalized read counts at each reporting level.
[0321] Semi-quantitative reporting level Read Count <10^3.5 <32 10^3.5-10^4.5 32 - ≤420 10^4.5-10^5.5 420 - ≤5200 10^5.5-10^6.5 5200 - ≤62,000 >10^6.5 ≥62,000
[0322] Table 3.
[0323] Similarly, as described above, multiple samples can be combined for sequencing. Although merging is not necessary, it is where large-scale parallel DNA sequencing shows its true power. For example, the ability to merge and sequence in parallel significantly reduces the time and personnel required to perform comparative studies. As described in Example 4 above, samples can be combined, normalized, and quantified.
[0324] Example 6: Determination of specific correction factors
[0325] Ideally, next-generation sequencing (NGS) should report the exact same target quantification as any other quantitative molecular method (e.g., digital PCR). However, the reality is that variability in sampling and measurement can still exist among individual reference methods, and such randomness often prevents this ideal of consistency from being achieved. One possible solution is to report data with an appropriate error range or to bin the data. However, if the error is assay-specific and not consistent / systematic across all sample types, the error range or binning may not account for all the error in the data.
[0326] On the other hand, assay-specific correction factors can be used to further account for repeatable and systematic factors, such as differences in nucleic acid amplification efficiency, differences in nucleic acid purification efficiency, differences in sequencing library preparation, and differences in sequencing efficiency. Since such differences are repeatable and systematic for a given sample, analyte, and / or assay, the differences can be measured and used to generate assay-specific correction factors to correct target quantification. The NGS methods described herein can use assay-specific correction factors to remove systematic differences in target quantification performance observed across many (e.g., more than 50) experimental measurements for many targets.
[0327] The assay-specific correction factors described in this example were derived by analyzing assay-specific positive control (PC) sequences in over fifty batched positive controls. PC sequences can be added to the test samples, or they can be included as separate samples in each run. Prior to amplification, batched positive controls can be spiked with a mixture of PC sequences (e.g., modified gBlocks) at known concentrations (e.g., 50 copies of each PC gBlock / reaction) and quantitative standard material (e.g., 500 copies of QSM / reaction). If PCR amplification efficiency is suboptimal, or if the target and internal standard have different efficiencies, the correction factor can depend on the number of positive controls included in the reaction. Therefore, in one embodiment, the number of positive controls should be in the middle of the target dynamic range for the respective target. In one embodiment, PC sequences can be designed to be amplified using the same target-specific primers with overhangs as those used in the test assay, allowing sample-specific sequencing adapter primers to anneal to the amplified PC sequences in the same manner as in the test assay. Naturally, positive controls are designed to have their own unique identification sequence region between conserved primer binding site regions. In this way, sequencing data derived from positive controls can be identified and analyzed separately. PC sequences can contain assay-specific targets, and when each assay has the same performance (e.g., 10 4.7 copies / mL), quantification at equivalent levels is expected.
[0328] However, Figure 16Two assay populations were observed. For the purposes of this example, these two assay populations are referred to as good performers and poor performers. The assays designated as good performers (AM) are visible on the left side of the graph, while the assays designated as poor performers (OX) are visible on the right side of the graph. Sample N is an intermediate performer. The mean log10 quantification of the good performers is consistently close to the expected quantification, while the mean log10 quantification of the poor performers is not. That is, it can be assumed that a 'good performer' consistently quantifies close to its expected performance in the test assay, while it can be assumed that a 'poor performer' consistently quantifies at values significantly below its expected performance in the test assay. More specifically, a good performer is defined as an assay whose mean log10 quantification is within 0.5 log10 of the expected quantification, while a poor performer is defined as an assay whose mean log10 quantification is ≥ 0.5 log10 below the expected quantification. The mean log10 quantification of the good performers is 0.23 log10 below the expected quantification, while the mean log10 quantification of the poor performers is 1.0 log10 below the expected quantification. One assay (shown as N) was observed with intermediate performance.
[0329] The assay-specific correction factor described in this example is defined as the mean log10 quantification of each assay divided by the mean log10 quantification of the good performers (10 4.47 The difference between the two assays (copies / mL) was calculated, allowing for a correction factor of up to 1.0 log10. In this example, assay-specific correction factors were derived for a specific set of assays. However, one of ordinary skill will appreciate that the principles described in this example can be adapted and applied to other assays with only minimal assay-specific modifications.
[0330] Example 7: Kit for normalizing and quantifying unknown nucleic acids in next generation sequencing (NGS) assays
[0331] In one embodiment, a kit for normalizing and quantifying unknown nucleic acids in a next generation sequencing (NGS) assay is disclosed. The kit can include an internal quantification standard, wherein the internal quantification standard is a nucleic acid configured to be added in known amounts to a sample comprising an unknown nucleic acid to be sequenced, and instructions for using the internal quantification standard for normalizing and quantifying a sequencing data set and for calculating the input quantity of the unknown nucleic acid. In one embodiment, the kit can include a set of internal quantification standards to be added at different known concentrations for generating a standard curve for quantifying the unknown nucleic acid. In one embodiment, the internal quantification standard can be configured to be added at about 10 4 -10 6 In one embodiment, using an internal quantitative standard, sequencing data can be normalized by Equation 5.
[0332] NORM = Number of sequencing reads * (F / Observed number of sequencing reads derived from the internal quantification standard) [Equation 5]
[0333] Where 'F' is the minimum expected number of sequencing reads of a fixed internal quantification standard. In one embodiment, using an internal quantification standard, the input quantity (IQT) of the unknown nucleic acid can be calculated by Equation 6:
[0334] IQT = Normalized unknown nucleic acid sequencing read counts attributed to unknown nucleic acids * (input number of internal quantitation standards / F) [Equation 6]
[0335] In one embodiment, the kit can further include a sequencing-specific adapter for at least an internal quantitative standard, comprising a sequencing primer binding site and a sample-specific identification sequence. In one embodiment, the kit can further include a target-specific primer with a custom overhang configured to amplify the internal quantitative standard and to anneal or ligate the sequencing-specific adapter.
[0336] The present invention may be implemented in other specific forms without departing from its spirit or essential characteristics. The embodiments described are considered to be illustrative in all respects and not restrictive. Therefore, the scope of the present invention is indicated by the appended claims rather than the foregoing description. Although certain embodiments and details have been included herein and in the accompanying disclosure for the purpose of illustrating the present invention, it will be apparent to those skilled in the art that various changes may be made to the methods and apparatus disclosed herein without departing from the scope of the present invention as defined in the appended claims. All variations within the meaning and equivalent range of the claims are included within their scope.
Claims
1. A method for normalizing read values of a sequencing assay for non-diagnostic purposes, comprising: Providing a sample comprising one or more unknown nucleic acids to be sequenced; A known amount of internal quantitative standard was added to the sample; preparing a sample comprising an internal quantitative standard for sequencing, wherein the preparing comprises introducing sequencing-specific adapter sites and sample-specific identification sequences into the unknown nucleic acids in the sample and the internal quantitative standard; Set the number of sequencing reads F for the internal quantification standard, where 'F' is a fixed minimum number of sequencing reads expected for the internal quantification standard; sequencing to generate a sequencing dataset for the sample, wherein the sequencing dataset includes sequencing reads observed from the unknown nucleic acid and an internal quantitative standard; Count the number of sequencing reads originating from unknown nucleic acids and internal quantification standards in a sequencing dataset; determining whether a minimum number 'F' of sequencing reads for an internal quantification standard has been reached in the sequencing dataset; and The sequencing data set is normalized by NORM = number of sequencing reads * (F / observed number of sequencing reads derived from the internal quantification standard) to provide a normalized sequencing data set, and the input number of unknown nucleic acids (IQT) is calculated by IQT = normalized number of unknown nucleic acid sequencing reads derived from the unknown nucleic acid * (input number of internal quantification standard / F), Wherein if the number of internal quantification standard sequencing reads is less than 'F', the sequencing dataset is rejected, or subjected to additional sequencing until 'F' internal quantification standard reads are recorded.
2. The method of claim 1, wherein each unknown nucleic acid in the sample has 10 2 -10 7 The concentration of copies / ml.
3. The method of claim 1, wherein the known amount of the internal quantitative standard added to the sample is within 10 4 -10 6 copies / ml.
4. The method of claim 1, wherein the unknown nucleic acid reads and the internal quantification standard reads are each normalized by the same ratio ALPHA, wherein ALPHA = F / observed number of sequencing reads derived from the internal quantification standard.
5. The method of claim 1, wherein preparing the sample comprises lysing the sample, recovering nucleic acid from the lysate, and purifying the recovered nucleic acid, and introducing primer binding sites and sample-specific identification sequences into the region of the nucleic acid to be sequenced.
6. The method of claim 5, wherein said introducing a primer binding site and a sample-specific identification sequence into the region of the nucleic acid to be sequenced comprises one of the following: Amplifying the nucleic acid to be sequenced in an amplification reaction using target-specific primers with dual-index sequencing overhangs, the target-specific primers comprising a sequencing primer binding site and a sample-specific identification sequence, or The nucleic acid to be sequenced is fragmented and ligated to fragmented nucleic acid sequencing-specific adaptors, which contain sequencing primer binding sites and sample-specific identification sequences.
7. The method of claim 6, wherein amplifying the nucleic acid to be sequenced comprises: Perform a first multiplex PCR reaction using target-specific primers with custom overhangs, Perform the first nucleic acid purification, performing a second PCR reaction using dual-index sequencing adapter primers that anneal or ligate to the overhangs introduced in the first PCR, wherein the dual-index sequencing adapter primers are target-independent and contain a sequencing primer binding site and a sample-specific identification sequence, Perform a second nucleic acid purification.
8. The method of claim 1, further comprising pooling two or more samples and subjecting them to sequencing simultaneously, wherein the two or more samples are pooled after preparation and before sequencing.
9. The method of claim 1, wherein the sequencing assay is a next generation sequencing assay.
10. The method of claim 1, wherein the sequencing assay does not include performing relative quantification, performs quantification in a reaction separate from the sequencing assay, uses an assay or template-specific quantification standard, or uses a competing template as a quantification standard.
11. A method for performing a quantitative next generation sequencing (NGS) assay for non-diagnostic purposes, comprising: Providing a sample comprising one or more unknown nucleic acids to be sequenced; A known amount of internal quantitative standard was added to the sample; Prepare samples containing internal quantitative standards for sequencing; Set the number F of sequencing reads for the internal quantification standard, where 'F' is the minimum number of sequencing reads for the internal quantification standard; Sequencing unknown nucleic acids in the sample and an internal quantitative standard to generate sequencing data; Count the number of sequencing reads originating from unknown nucleic acids and internal quantification standards in a sequencing dataset; determining whether a minimum number 'F' of sequencing reads for an internal quantification standard has been reached in the sequencing dataset; and The sequencing data set was normalized by NORM = number of sequencing reads * (F / observed number of sequencing reads derived from the internal quantification standard) to provide a normalized sequencing data set, and the input number of unknown nucleic acids (IQT) was calculated by IQT = normalized number of unknown nucleic acid sequencing reads derived from the unknown nucleic acid * (input number of internal quantification standard / F); Wherein 'F' is set to ensure sufficient sequencing read depth of the sample to detect one or more unknown nucleic acids present in the sample at or above the limit of detection (LOD), wherein sufficient sequencing read depth of the sample is achieved when a minimum number of sequencing reads F of the internal quantification standard is detected.
12. The method of claim 11, wherein each unknown nucleic acid in the sample has 10 2 -10 7 The concentration of copies / ml.
13. The method of claim 11, wherein the known amount of the internal quantitative standard added to the sample is within 10 4 -10 6 copies / ml.
14. The method of claim 11, wherein normalization separately (1) applies data acceptance / rejection criteria to each sample in the assay and (2) ensures that the same limit of detection (LOD) is applied to all unknown nucleic acids in each sample.
15. The method of claim 11, further comprising: The input quantity of multiple unknown nucleic acids in a sample is calculated as IQTn = normalized unknown "n" nucleic acid sequencing reads attributed to the "nth" unknown nucleic acid * (input quantity of internal quantitation standard / F).
16. The method of claim 11, further comprising pooling two or more samples and subjecting them to sequencing simultaneously, wherein the two or more samples are pooled after preparation and before sequencing.
17. The method of claim 16, wherein each pooled sample has associated therewith a unique set of sample-specific identification sequences such that sequencing data from each sample in the pooled sample is distinguishable and separate.
18. The method of claim 17, wherein each pooled sample has its own internal quantitative standard associated with its own set of unique sample-specific identification sequences, and wherein quantification is applied separately to each nucleic acid from each sample in the pooled sample.
19. The method of claim 11, further comprising providing a set of assay-specific positive controls to be sequenced, wherein the positive controls include a positive control corresponding to each of the one or more unknown nucleic acids sequenced in the assay, and wherein the method further comprises applying an assay-specific correction factor to each of the one or more unknown nucleic acids based on the sequencing reads of the assay-specific positive controls.
Citation Information
Patent Citations
Loading vials
US20140283945A1
Extreme PCR
US20150118715A1
Methods for Standardized Sequencing of Nucleic Acids and Uses Thereof
US20150292001A1
Method for quantification of an analyte
US6303305B1
Containment cuvette for PCR and method of use
US6645758B1