Methods and compositions for mitigating index hops in DNA sequencing
By using a unique double indexing method during nucleic acid sample amplification, the index combination of each sample is ensured to be unique within the same group, solving the data contamination problem caused by index skipping, improving the accuracy and reliability of next-generation sequencing, and making it suitable for a variety of research fields.
Patent Information
- Application Number
- CN202380087725.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-12-02
- Filing Date
- 2023-12-04
- Publication Date
- 2025-12-05
AI Technical Summary
In next-generation sequencing, index skipping leads to data contamination and inaccurate results, especially in multiplexed libraries. Existing methods struggle to effectively eliminate or mitigate index skipping, impacting the accuracy of diagnosis and research.
A unique dual-indexing method is employed, in which indexed target-specific primers are used to hybridize with nucleic acids during nucleic acid sample amplification to form amplicones. The amplicones are then sequenced using next-generation sequencing, ensuring that the index combination for each sample is unique within the same group and reducing index misassignment.
It significantly reduces index skipping, improves the reliability and accuracy of sequencing data, especially when processing large numbers of samples, and reduces the risk of computer cross-contamination and misallocation. It is suitable for clinical sequencing, single-cell sequencing, and low allele fraction somatic variant analysis.
Smart Images

Figure CN121079433A_ABST
Abstract
Description
[0001] Related applications
[0002] This application claims priority to U.S. Provisional Patent Application No. 63 / 385,860, filed December 2, 2022, entitled “Dual Combination of Unique Indexing Method,” the entire contents of which are incorporated herein by reference. Background of the Invention
[0004] Targeted DNA sequencing allows users to selectively analyze specific regions of the genome. Instead of sequencing the entire genome, only the specific region of interest is sequenced. This approach is more focused and efficient because it allows those skilled in the art to gather information about specific genes or genomic regions without sequencing the entire genome. Targeted DNA sequencing typically involves using capture probes or primers designed to specifically bind to and capture the region of interest. These primers or probes are often complementary to the targeted DNA sequence, enabling selective amplification and sequencing. This is particularly useful when studying specific genes or genomic regions known to be associated with certain infections, pathogenicities, diseases, or traits. By targeting these regions, researchers can analyze variations, mutations, or structural changes that may be associated with a specific condition or characteristic. Compared to whole genome sequencing (WGS), targeted DNA sequencing is faster, cheaper, and requires fewer computational resources. It has become a valuable tool in a variety of research fields, including cancer genomics, genetic disease diagnosis, forensics, and personalized medicine.
[0005] Multiplex PCR (polymerase chain reaction) allows the simultaneous amplification of multiple target DNA sequences in a single reaction. It involves multiple primer pairs (each specific to a different target sequence), as well as the use of DNA polymerase and nucleotides. By incorporating multiple primer sets (each corresponding to a specific target region) into a single PCR reaction, multiple DNA fragments can be amplified simultaneously.
[0006] Multiplex PCR has a wide range of applications in many fields, including medical diagnostics and science, genetics, forensic medicine, microbiology and virus detection, food safety testing and food pathogens, drug resistance, pharmacogenetics, environmental testing, epigenetics, allergen testing, botany, ecology, evolutionary biology, zoology, and research.
[0007] By simultaneously amplifying multiple target genes or regions, multiplex PCR enables the detection of multiple disease-related variants in a single test. For pathogen identification, multiplex PCR can be employed for rapid detection and identification of pathogens such as bacteria, viruses, parasites, or fungi in clinical, food, or environmental samples. By targeting specific genomic regions unique to different pathogens, multiplex PCR can provide rapid and accurate diagnosis. For cancer mutation analysis, multiplex PCR can be employed to detect specific mutations or genetic alterations associated with cancer. By amplifying target genes known to have cancer-related mutations, multiplex PCR allows for efficient screening and profiling of tumor samples.
[0008] Next-generation sequencing (NGS) technologies can generate millions to billions of sequencing reads. As disclosed herein, the key to harnessing this increased capacity is to add a unique nucleic acid sequence (called an index or barcode) to each DNA fragment during library preparation during multiplex PCR. This allows for pooling and simultaneous sequencing of large numbers of libraries during a single sequencing run. However, the increased throughput that comes with multiplexing is accompanied by an increase in the level of complexity, as sequencing reads from pooled libraries need to be computationally identified and sorted in a process called demultiplexing prior to final data analysis.
[0009] DNA sequencing technology has significantly advanced our understanding of genetics and genomics. However, one of the challenges in high-throughput sequencing is "index jumping" (also known as "index switching," "index swapping," "barcode jumping," "index misassignment," or "sample bleeding"), in which sequences from one sample are erroneously assigned to another sample. This can lead to data contamination and inaccurate results in genomic studies. The present disclosure aims to mitigate index jumping and improve the reliability of DNA sequencing. Although the overall index jumping rate is relatively low, on certain devices that use patterned flow cells with exclusion amplification chemistry, the level of overall index jumping rate has been observed to be elevated compared to devices that do not use patterned flow cells. Even a relatively low index jumping rate can affect many diagnostic applications that use NGS systems, especially for the detection of microbial and viral pathogens, the screening for low copy number variants, and cfDNA and ctDNA. Therefore, there is a need for a method that eliminates or mitigates this risk.
[0010] The typical level of index jumping for certain sequencing platforms ranges from 0.1% to 2%, depending on the type, quality, and processing of the library, however, index jumping rates as high as 10% have been reported. The use of plus unique dual indexing combinations, in which the target sequence is labeled with unique index sequences at both the 5' and 3' ends, can significantly reduce the level of index jumping. However, the operational complexity of this plus unique dual indexing approach can significantly increase the chance of cross-contamination.
[0011] Large-scale sequencing of biological samples requires a highly streamlined workflow that maximizes the use of instrument output and eliminates the effects of lane-to-lane variability. To achieve maximum cost-effectiveness, sample multiplexing on the sequencer has become necessary. However, there have been many reports that certain amplification chemistries can result in severe data integrity issues due to the index jumping phenomenon. Some studies report that index jumping can be due to residual excess free primers or adapters in the sample that can cause library fragments to be falsely extended with oligonucleotides containing the wrong sample index. Since this phenomenon is a property of the flow cell chemistry itself, the only option to completely eliminate the effects of index jumping is to sequence one sample per lane, which unfortunately is not financially feasible in many cases.
[0012] As described above, while NGS-based targeted amplicon sequencing is a powerful approach, different errors and biases such as variation in sequencing depth between individual samples, sequencing error rates, and index hopping can play an important role within NGS data analysis. Currently, there are no standard requirements to provide detailed reports and explanations to correct such potential errors. Furthermore, the use of NGS platforms by sequencing companies, core facilities, diagnostic laboratories, research institutions, and other entities is increasing. Third-party provided NGS services often only provide sequencing data without including general information about NGS runs, individual sample de-multiplexing efficiency, and other related parameters.
[0013] Currently, a widely used approach to study large sample numbers is by analyzing pooled samples by combining DNA from multiple individuals into one sample for NGS library, thus precluding the opportunity to trace specific sequences back to individual samples. For example, in microbial and viral detection and screening, analysis of such pooled samples can lead to inaccurate results due to index hopping. In contrast, the present disclosure is an economically efficient NGS approach and is scalable to hundreds to thousands of individual samples, making it an ideal choice for handling any study of large sample numbers.
[0014] With the occurrence of index hopping, the misallocation of indices between multiplexed libraries and its rate can rise with more free adapters or primers present in the prepared NGS library. To address this issue, some methods differentiate between plus combinatorial dual indexing and plus unique dual indexing. Special kits can be used with unique dual indexing sequences (e.g., a set of 96 primer pairs) to address index hopping issues and de-multiplexing deficiencies. This is a choice for low sample numbers, but if one wants to individually index hundreds of samples in one sequencing run, it can be difficult to implement plus unique dual indexing due to high sample numbers and for cost reasons. For example, if shared flow cell lanes and reads are misallocated, computer cross-contamination between samples from different studies and altered or false results can occur. Even if samples are exclusively run on a single flow cell, index hopping can lead to barcode switching events between samples, resulting in misallocation of reads.
[0015] For certain library preparations run on NGS, two indices are often used to label individual samples (double indexing). Sequences are then de-multiplexed and sequencing data converted into FASTQ file format. The person skilled in the art assumes the assignment of sequences to samples to be correct when performing bioinformatics analysis with de-multiplexed data. It is extremely difficult to verify this, as the provided data set lacks all information about the de-multiplexing setup and, in addition, also lacks all information about the degree of sequencing errors within an index and index hopping. Thus, sequences can be wrongly assigned to samples and, in case of shared flow cells, even across sample groups. However, it is known that de-multiplexing errors occur and depend on multiple factors such as sequencing platform, library type used and index combination. The few existing studies that investigate index hopping in more detail give rates of 0.2-10%. The problem of sequencing errors within an index and index hopping can become particularly severe when many samples with different indices are sequenced together.
[0016] In general, there is no standard method or know-how about the used sequencing platform model and software tools to handle index hopping when analyzing sequence data, which can significantly influence the results and thereby the interpretation of the sequencing information, especially in diagnostic, detection and screening samples. In view of this lack of solution, it is a challenge to understand what has been done during sample processing and data analysis and it is not possible to compare results of different studies. To date, published NGS studies such as targeted amplicon sequencing are difficult to compare or evaluate due to the lack of this important information about data processing. Without effective tools to remove index hopping, it is possible to wrongly assign samples. Even methods with index exchange below 1% can generate false results in downstream data analysis, although reasonable laboratory practices are implemented to reduce residual adapters or primers.
[0017] Index hopping leads to incorrect demultiplexing, where reads are assigned to the wrong sample, which manifests as downstream read contamination in the data. Eliminating index hopping is critical for any sequencing study. While lower rates of index hopping can not affect the ability to trust variant calling in many germline DNA applications, it can lead to false results when looking for rare transcripts, fusion events in RNA-seq, or performing low-allele fraction somatic analysis or detection, or screening and detection of microbial and viral pathogens. While mutation callers such as the Data Analysis Software Tool filter many common artifacts that arise during the sequencing process, index swapped reads are unique in that they are high-quality reads, not errors, just assigned to the wrong sample. In cancer genomics studies, there are applications such as blood biopsies using circulating tumor DNA (ctDNA) where researchers are trying to detect mutations at allele fractions of 5% or lower in the ctDNA against a background of normal DNA. If this level of high sensitivity and confidence is required, then eliminating or mitigating the method is necessary because even low rates of sample cross-contamination can impede the accuracy and sensitivity of low-allele fraction variant calling. It should be noted that the index hopping phenomenon can occur in any case where multiplexed libraries are amplified together in the same sequencer and there are residual adaptors and active polymerase. Thus, when designing a sequencing experiment, it is important to remember that there is a risk of index hopping induced cross-contamination at any time that samples are amplified together in a pool, whether in tubes during library preparation or on a flow cell.
[0018] For example, Costello et al. (doi: 10.1186 / s12864-018-4703-0) demonstrated that index hopping led to the incorrect assignment of reads from 5 different gene fusion transcripts in cell line RNA-seq data (three cell lines were used) when four RNA-seq libraries were pooled for each cell line (a total of 12 libraries) and sequenced on a HiSeq 4000 lane. Only one cell line (K562) should carry the BCR-ABL1 translocation, however, due to index hopping, reads containing BCR-ABL1 were also found in the data files of the other two cell lines. Costello et al. also showed the variability of the index hopping rate between different pools and between different flow cells.
[0019] The present disclosure describes a method that can eliminate or significantly mitigate index hopping when analyzing a large number of samples in the same sequencing run. It allows filtering out swapped reads from pooled samples. The method of the present disclosure allows mitigating swapped reads caused by both swapping from multiplex PCR and sequencing chemistry. This is particularly critical in clinical sequencing environment, single cell sequencing, pathogen detection and screening, or low allele fraction somatic variant analysis, where even a low percentage of abnormal reads is not acceptable. SUMMARY
[0021] In some embodiments, the present disclosure describes a method of eliminating or mitigating index hopping during amplification of at least one nucleic acid sample, comprising the steps of: for each sample, hybridizing a plurality of indexed target-specific primers to nucleic acids from the sample in the presence of an indexed universal primer in a single reaction vessel to form a plurality of test reactions, wherein at least one indexed target-specific primer is configured to bind to at least one target nucleic acid sequence; subjecting each test reaction to amplification conditions to generate amplicons, wherein each amplicon contains a plurality of indexes, wherein one index (A) is located at the 5' end of the amplicon Watson (forward) strand, one index (B) is located at the 5' end of the amplicon Crick (reverse) strand, one index (X) is located on the Watson (forward) strand adjacent to the target-specific primer, one index (Y) is located on the Crick (reverse) strand adjacent to the target-specific primer; subjecting at least a portion of the amplicons to bead clean up to form enriched amplicons; and sequencing the enriched amplicons formed from each sample by next generation sequencing. In some embodiments, each index has a different sequence. In some embodiments, the amplicons are pooled and then immediately subjected to bead clean up.
[0022] In some embodiments, within a set of samples, the same X index (located on the Watson strand adjacent to the target-specific primer) can be shared among more than one sample in the same set, and the same Y index (located on the Crick strand adjacent to the target-specific primer) can be shared among more than one sample in the same set. However, the combination of X index and Y index needs to be unique for each sample within the same set. The same A index (located at the 5' end of the amplicon Watson strand) can be shared among more than one sample in the same set, and the same B index (located at the 5' end of the amplicon Crick strand) can be shared among more than one sample in the same set. However, the combination of A index and B index must be unique for each sample within the same set. The same X index and Y index can be used for samples in different sets of samples (e.g., Set 1, Set 2, Set 3, etc.) that are pooled in the same sequencing run. However, each combination of X index and Y index must be unique for each sample within each set. On the other hand, A index and B index are not shared among different sets and must be unique for each set.
[0023] In another embodiment, the present disclosure can be performed in two-step PCR during amplification of at least one nucleic acid sample, comprising the following steps: for the first step amplification, for each sample, hybridizing a plurality of indexed target-specific primers to nucleic acids from the sample in a reaction vessel to form test reactions, wherein at least one indexed target-specific primer is configured to bind to at least one target nucleic acid sequence; subjecting each test reaction to amplification conditions to generate amplicons, wherein each amplicon contains two indices, one index (X) located on the Watson strand proximal to the target-specific primer, and one index (Y) located on the Crick strand proximal to the target-specific primer; performing a second step amplification on at least a portion of the nucleic acid amplicon products from the first step amplification using indexed universal primers comprising two unique indices, or, alternatively, subjecting each test reaction to amplification conditions to generate amplicons, wherein each amplicon contains indices, wherein one index (A) is located at the 5' end of the amplicon Watson strand, one index (B) is located at the 5' end of the amplicon Crick strand, one index (X) is located on the Watson strand proximal to the target-specific primer, and one index (Y) is located on the Crick strand proximal to the target-specific primer; performing a bead clean on at least a portion of the amplicons to form enriched amplicons; and sequencing the enriched amplicons formed from each sample by next generation sequencing. In some embodiments, each final amplicon contains four (“quadruplexed”) combined indices. In some embodiments, each index has a different sequence. In some embodiments, the final amplicons are pooled and then immediately subjected to a bead clean.
[0024] In some embodiments, within a set of samples, the same X index can be shared within more than one sample in the same set, and the same Y index can be shared within more than one sample in the same set. However, the combination of the X index and the Y index needs to be unique within the same set. The same A index (located at the 5' end of the amplicon Watson strand) can be shared within more than one sample in the same set, and the same B index (located at the 5' end of the amplicon Crick strand) can be shared within more than one sample in the same set. The same X index and Y index can be used for different sample sets (e.g., S1, S2, S3, etc.) that are pooled in the same sequencing run. However, each combination of the X index and the Y index must be unique within each set. On the other hand, the A index and the B index cannot be shared between different sets and must be unique for each set.
[0025] In some embodiments, each indexed universal primer comprises: a universal priming portion at the 3' end; a barcode / indexing portion in the middle; and a universal priming portion at the 5' end. In some embodiments, each indexed target-specific primer comprises a universal priming portion, a barcode / indexing portion, and a specific sequence portion to a target nucleic acid sequence. In some embodiments, each sample is obtained from a subject, a human, an animal, a plant, a microorganism, a virus, or an environmental source. In some embodiments, the target-specific primers include primers configured to amplify at least 10, 20, 50, 100, 200, 1,000, or more targets.
[0026] In some embodiments, the method further comprises the step of pooling the enriched amplicons from each sample prior to sequencing.
[0027] In some embodiments, the disclosure describes a method of analyzing at least one sample from a human, food, animal, plant, and pathogen, comprising the steps of: for each sample, hybridizing a plurality of indexed target-specific primers to nucleic acids from the sample in the presence of an indexed universal primer in a single reaction vessel to form a test reaction, wherein at least one indexed target-specific primer is configured to bind to at least one target sequence, wherein each indexed universal primer comprises: a universal priming portion at the 3' end; a barcode / indexing portion in the middle; and a universal priming portion at the 5' end; and wherein each indexed target-specific primer comprises a universal priming portion, a barcode / indexing portion, and a target-specific sequence portion; subjecting each test reaction to amplification conditions to generate amplicons with four indices; pooling the amplicons from each sample; performing a bead clean on at least a portion of the pooled amplicons generated by each sample to form enriched amplicons; and sequencing the pooled and enriched amplicons formed by each sample by next generation sequencing.
[0028] In some embodiments, the disclosure describes a kit comprising a plurality of target-specific primers configured to bind to target sequences specific to a biological sample related to cancer, genetic disorders, forensic testing, allergens, microbe / virus species or pathogens, low frequency somatic variant detection, ancient DNA, gene expression, cell-free DNA, ctDNA, or any other nucleic acid biological application.
[0029] In some embodiments, the disclosure describes methods and compositions for amplifying one or more selective target regions in a nucleic acid sample. In some embodiments, the methods comprise the steps of:
[0030] (1) contacting a nucleic acid sample with indexed target-specific primers in a PCR reaction in the presence of indexed universal primers; and (2) allowing primer extension to generate target amplification products (amplicons) of different sizes, wherein each amplicon contains a four-fold (four) combination of indices. In some embodiments, four of the four indices in the amplicon comprise different sequences, or three of the four indices in the amplicon comprise different sequences. In some embodiments, the method comprises a step of determining the presence or absence of a target amplification product. In some embodiments, the method comprises a step of establishing the sequence of a target amplification product. In some embodiments, less than 50%, 40%, 30%, 20%, 10%, 5%, 0.5%, or 0.1% of the amplification products are primer dimers or pseudo products.
[0031] In some embodiments, the concentration of each indexed target-specific primer can be about 500, 250, 100, 80, 70, 50, 30, 10, 2, or 1 nM. In some embodiments, the GC content of the indexed target-specific primers can be different, and for example, it can be 40% to 70%, or 30% to 60% or 50% to 80%. In some embodiments, the melting temperature (Tm) of the indexed target-specific primers can be 55°C to 65°C or 40°C to 70°C, or 55°C to 68°C. In some embodiments, the length of the indexed target-specific primers can be 20 to 90 bases, 40 to 70 bases, 20 to 40 bases, or 25 to 50 bases. In some embodiments, the 5' region of the target-specific primer is a universal primer binding site that is not complementary to or specific for any region of a nucleic acid in the sample. In some embodiments, the length of the target amplicon is 50 to 500 bases, 90 to 350 bases, or 200 to 450 bases. m ) can be 55°C to 65°C or 40°C to 70°C, or 55°C to 68°C. In some embodiments, the length of the indexed target-specific primers can be 20 to 90 bases, 40 to 70 bases, 20 to 40 bases, or 25 to 50 bases. In some embodiments, the 5' region of the target-specific primer is a universal primer binding site that is not complementary to or specific for any region of a nucleic acid in the sample. In some embodiments, the length of the target amplicon is 50 to 500 bases, 90 to 350 bases, or 200 to 450 bases.
[0032] In various embodiments of any aspect of the disclosure, the primer extension method is based on the prior art polymerase chain reaction (PCR). In various embodiments, the annealing time can be greater than 0.5, 1, 2, 5, 8, 10, or 15 minutes. In various embodiments, the extension time can be greater than 0.5, 1, 2, 5, 8, 10, or 15 minutes.
[0033] In some embodiments, the methods disclosed herein quantify the copy number of a target sequence present in a sample.
[0034] In various embodiments of the disclosure, compatibility and incompatibility scores of selected primers are calculated based on different factors of target amplicon GC content, target amplicon melting temperature, target amplicon heterozygosity ratio, complementarity of candidate primers to target region; candidate primer size, target amplicon size and amplification efficiency, and off-target rate. Selected target-specific primers can hybridize to nucleic acid targets and selectively amplify target regions. In various embodiments, the test sample is from a subject, individual, food, plant, animal, soil, environment, or any nucleic acid subject suspected of having an infection or disease, or increased risk of infection or disease; and wherein one or more target nucleic acids comprise sequences at target regions associated with the infection or disease, or increased risk of infection or disease. In some embodiments, the test sample is from a subject, individual, animal, soil, environment, or nucleic acid subject not associated with any disease or infection. The characteristic profile of target regions can be used as an identity marker for the subject, individual, animal, or other sample in a similar manner to a fingerprint. In some embodiments, the information can be used for disease screening, detection, disease management, pathogen monitoring, food recall, disease outbreak, or epidemic.
[0035] In one embodiment, the methods disclosed herein can be used for screening, detecting, and identifying microbial and viral pathogens. In one embodiment, the methods disclosed herein can be used for screening, detecting, genotyping, serotyping, subtyping, and tracking (monitoring) infectious agents. In some embodiments, candidate primers are contacted with a nucleic acid sample, wherein forward and reverse strand indexed target-specific primers hybridize to target nucleic acid regions, if present in the sample, wherein the nucleic acid sample can have microbial and / or viral organisms, or is suspected of having microbial and / or viral organisms; a plurality of target nucleic acids are amplified in the presence of indexed universal primers to generate amplicons containing four combinatorial indices; the amplicons are sequenced by next-generation sequencing; and sequence data is analyzed by software analysis. In some embodiments, the detected infection can be clinically actionable. In some embodiments, the detected infection can be associated with drug resistance. In some embodiments, detection, identification, and quantification of microbial and viral species, strains, and subtypes can be associated with disease. In some embodiments, biological samples can be monitored for infectious agents or surveillance.
[0036] In one aspect, the methods and compositions disclosed herein are designed to detect, identify, and quantify target nucleic acids in a sample that can contain microorganisms and viral organisms such as sexually transmitted infections (STIs). In some embodiments, the disclosed methods comprise the following steps: (1) contacting nucleic acid targets in a sample with primers, wherein indexed forward strand and indexed reverse strand target-specific primers hybridize to different nucleic acid target regions in the presence of indexed universal primers in a test reaction; (2) amplifying the target nucleic acids under optimal amplification conditions to generate amplicons containing fourfold combinatorial indices; (3) sequencing the amplification products by NGS; and (4) analyzing and quantitatively measuring the generated sequence reads by mapping-and-counting methodology.
[0037] In one embodiment, the methods disclosed herein can be used to screen and analyze target regions of a genome for diseases such as cancer or genetic disorders. In one embodiment, the methods disclosed herein can be used to analyze a genome for forensic DNA analysis based on DNA profiles such as short tandem repeat (STR) regions. In some embodiments, the methods disclosed herein can be used for pharmacogenetics or drug resistance to detect genetic variations that affect how an individual responds to a drug. In some embodiments, candidate primers contact a nucleic acid sample; wherein forward strand and reverse strand indexed target-specific primers hybridize to target regions in the presence of indexed universal primers, a plurality of target nucleic acids are amplified to generate amplicons with fourfold combinatorial indices; the amplicons are sequenced by next generation sequencing; and sequence data is analyzed by software analysis. In some embodiments, the nucleic acid variations detected can be clinically actionable. In some embodiments, biological samples can be monitored for prognosis.
[0038] In some embodiments, the nucleic acid sample comprises genomic nucleic acids. In some embodiments, the sample comprises nucleic acid molecules obtained from food, vegetables, agricultural products, plants, soil, putrefaction, water, environment, food production facilities, or any nucleic acid subject. In some embodiments, the sample comprises nucleic acid molecules obtained from urine, tissue, saliva, biopsies, sputum, swabs, surgical resections, cervical swabs, tumor tissue, fine needle aspirations (FNA), scrapings, swabs, mucus, semen, other non-limiting clinically or laboratory obtained samples.
[0039] In another aspect, the disclosure relates to kits comprising indexed target-specific primers for amplifying target regions of interest in a sample.
[0040] In some embodiments, the disclosed method includes the steps of: performing multiplex barcoding amplification to generate amplicons containing four combined indices, and sequencing the resulting amplicons via NGS. In some embodiments, the samples are obtained from subjects with single or multiple co-infections. In some embodiments, the analytical sensitivity of the method is 10 copies of each microbial and viral species in the sample; high-multiplex PCR amplifies 20, 50, 100, 200, 500, 1,000 or more targets with minimal primer-primer interactions. In some embodiments, the method includes a single-reaction, single-step barcoding multiplex PCR step. In some embodiments, the method can analyze 5, 10, 20, 50, 100, 200, 500, 1,000, 2,000, 5,000, 10,000 or more samples with a single NGS sequencing run.
[0041] In some embodiments, this disclosure relates to methods, compositions, and kits for applying multiplex target amplification and target enrichment prior to downstream analyses such as next-generation sequencing. The methods rely on using multiple indexed target-specific primers in the presence of indexed universal primers to target-enrich and amplify DNA samples suspected of containing disease (cancer or genetic disorders), drug resistance, forensic information, or microbial and viral pathogens. The target-specific primers amplify target nucleic acids under optimal conditions in the presence of amplification reagents such as polymerases and dNTPs to amplify at least one or more nucleic acid targets of interest.
[0042] In some embodiments of the disclosed method, the primer design method selects candidate target-specific primers based on the following steps: (1) extracting genomic sequences; (2) designing a set of primers with appropriate GC content and T for the target sequence. m (2) Target-specific forward and reverse strand target-specific primers at different distances from each target region; (3) For each primer, search for target genome sequences for off-target matches; filter primers and retain those that pass the off-target threshold; (4) Search for the 3′ end portion of each primer for complementary matches with the sequence of the primer set; filter primers stepwise, wherein primers with the most complementary matches at the 3′ end are removed first; and (5) synthesize primers and run the entire wet laboratory experiment using next-generation sequencing; calibrate the performance of each primer and filter out primers with unsatisfactory performance. In some embodiments, steps 2 to 4 and steps 2 to 5 of the primer selection procedure are repeated until each target sequence is covered by at least one forward strand target-specific primer and one reverse strand target-specific primer from the primer set.
[0043] In various embodiments of any aspect of the disclosure, the methods and compositions are characterized by multiplexed barcoding amplification and target enrichment of target nucleic acid regions in a single reaction using indexed target-specific primers in the presence of indexed universal primers. In some embodiments, the disclosed methods comprise the following steps: (1) contacting indexed target-specific primers with target nucleic acid sequences in the presence of indexed universal primers and hybridizing to target nucleic acid sequences in the sample; (2) subjecting the test reaction to amplification under optimal amplification conditions and generating amplicons containing quadruplexed combinatorial indices; (3) pooling amplification products from each individual sample together; (4) performing a bead clean on a portion of the pooled amplification products to remove unconsumed primers and primer dimers and produce enriched amplification products; (5) performing standard normalization and quantification on a portion of the enriched amplification products; and (6) sequencing the amplicons by next generation sequencing.
[0044] In one embodiment, the indexed universal primer comprises: a) a universal priming portion at the 3' end; b) a barcode / index portion in the middle; and c) a universal priming portion at the 5' end. Figure 1 ) In one embodiment, each indexed target-specific primer comprises a universal priming portion, a barcode / index portion, a specific sequence portion to a target nucleic acid sequence. Figure 1
[0045] In some embodiments, the composition comprises a plurality of indexed target-specific primers, wherein at least one target-specific primer is at least 90% identical to any of the nucleic acid targets. In some embodiments, the composition comprises a plurality of target-specific primers that have at least 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity to the nucleic acid targets in the sample.
[0046] In some embodiments, the disclosure relates to a composition comprising a plurality of indexed target-specific primers, wherein the sequences complementary to the target nucleic acids of interest are between about 15 and 40 bases in length.
[0047] In some embodiments, the disclosure relates to a composition of pre-computationally designed indexed target primers that generate minimal cross-hybridization or primer-primer interactions with other target-specific primers in the composition. In some embodiments, the primers in the composition are designed to avoid non-specific priming that can lead to non-specific amplification. In some embodiments, amplification conditions such as annealing temperature, annealing duration, and primer concentration can be adjusted to minimize amplification artifacts such as primer dimers.
[0048] In some embodiments, the present disclosure relates to a method or composition comprising a plurality of indexed target-specific primers that have minimal cross-hybridization to non-specific sequences present in a sample. In some embodiments, such cross-hybridization to non-specific targets can be monitored and evaluated by downstream analysis such as next generation sequencing.
[0049] In some embodiments, the present disclosure relates to a method or composition comprising a plurality of indexed target primers that have minimal self-complementarity. In some embodiments, the composition comprises at least one target-specific primer that does not form a secondary structure such as a hairpin or loop. In some embodiments, the composition comprises a plurality of target-specific primers, a majority or possibly all of the target-specific primers do not form a secondary structure such as a hairpin and loop.
[0050] In some embodiments, the target nucleic acid is obtained from a subject. In some embodiments, the sample comprises proteins, cells, fluids, biological fluids, preservatives, and / or other substances. In certain embodiments, the sample is derived from urine, tissue, saliva, biopsies, sputum, swabs, surgical resections, cervical swabs, tumor tissue, fine needle aspirates (FNA), scrapings, swabs, mucus, semen, other non-limiting clinically or laboratory obtained samples.
[0051] In some embodiments, the target amplification products are sequenced by next generation sequencing by current state-of-the-art next generation sequencing technologies or platforms. In some embodiments, the disclosed methods are not limited to these next generation sequencing technology examples and can be applied to new sequencing innovations.
[0052] In certain embodiments, the foregoing methods can be performed at multiple time points. BRIEF DESCRIPTION OF DRAWINGS
[0054] The present disclosure can be better understood with reference to the following drawings. The components in the drawings are not necessarily to scale, emphasis instead being placed upon clearly illustrating the principles of the present disclosure. Furthermore, like reference numerals designate corresponding parts throughout the several views.
[0055] Figure 1 A- Figure 1 D depicts the following schematic: forward indexed target-specific primer comprising read 1 sequencing (universal) portion, X index portion, and target-specific sequence portion Figure 1 A); reverse indexed target-specific primer comprising read 2 sequencing (universal) portion, Y index portion, and target-specific sequence portion Figure 1 B); indexed universal primer comprising universal portion (P5 tail), A index portion, and read 1 sequencing (universal) portion Figure 1C); indexed universal primer comprising a universal portion (P7 tail), a B index portion, and a read 2 sequencing (universal) portion Figure 1 D).
[0056] Figure 2 A- Figure 2 D depicts a schematic of the process of amplification of a nucleic acid target by forward and reverse indexed target specific primers in a single tube and single amplification reaction in the presence of an indexed universal primer Figure 2 A- Figure 2 C). After amplification, each amplicon is labeled with four indices Figure 2 D).
[0057] Figure 3 Depicts amplicons after amplification with four indices and a schematic of the sequencing procedure. Four sequencing reads are generated in one sequencing run. Read 1 contains the X index and the target sequence. Read 2 contains the Y index and the target sequence. The A index and the B index are generated by independent reads.
[0058] Figure 4 Illustrates indexed primer layout for a 96 well plate.
[0059] Figure 5 Illustrates the same forward and reverse primer sets paired with a second index set: P5 primer with X index and P7 primer with Y index.
[0060] Figure 6 Illustrates the same forward and reverse primer sets paired with a third index set: P5 primer with X index and P7 primer with Y index.
[0061] Figure 7 Illustrates a single tube multiplexed barcoding amplification workflow.
[0062] Figure 8 A and Figure 8 B shows two tables listing read counts, Figure 8 A is a table listing read counts for three targets, where the numbers that are allowed to combine are highlighted in a gray background, and there are 57 unintended reads. Figure 8 B is a table showing 16 unintended reads after removal of reads with an index combination that is not allowed.
[0063] Figure 9 shows a table listing 5 sequencing runs of a total of 480 samples amplified by the HPV-STI assay (96 samples per sequencing run), where 45.7% to 66.6% of false positive targets per sequencing run were removed by the add-on quadruple indexing method.
[0064] Details
[0065] The present disclosure relates to a method of indexing where amplification and barcoding occur simultaneously in the same PCR reaction (final library product) where the amplicon contains four combinatorial indices. The four-indexed amplicon is further analyzed by systems such as next generation sequencing. The present invention has a universal method and can be applied to nucleic acid-based biological applications such as: (1) screening, detection, and identification of bacterial / viral / parasitic / fungal and microbiome organisms; (2) cancer / genetic disease nucleic acid sequence analysis; (3) pharmacogenetics and companion diagnostics for selecting the right treatment and prognosis; (4) drug resistance for applications such as antimicrobial therapy selection, monitoring, and epidemiology, and personalized medicine; (5) forensic DNA analysis for DNA profiling; (6) allergens; (7) low frequency somatic variant detection, ancient DNA, gene expression, cell-free DNA, ctDNA, and / or (8) or any nucleic acid-based biological application using four-index multiplexed barcoded amplification. The present invention provides methods, compositions, kits, systems, and devices that allow for such target enrichment.
[0066] Improvements in NGS technology have greatly increased sequencing speed and data output, leading to large sample throughput on current sequencing platforms. Key to utilizing this increased capacity is multiplexing, as disclosed methods, which can add unique sequences (called indices) to each DNA fragment during library preparation. This allows large numbers of libraries to be pooled and sequenced simultaneously during a single sequencing run. With multiplexing, there is the potential for index hopping, regardless of the library preparation method or sequencing system used. Index hopping can result in assigning sequencing reads to the wrong index during demultiplexing, resulting in misassignment. Libraries with higher levels of free adapters will have higher levels of index hopping. Since oligonucleotide / primer synthesis is inherently error prone, with error frequencies of 1-5% per base, it is desirable to maximize the difference between two indices. Given that indices are used to label samples and do not contribute to interrogation of target sequences in the sample, indices are therefore designed to have minimal length and maximal difference between them when used in the same multiplexed pool. For these reasons, indices are typically designed to be 6-10 bases in length, with 3-5 bases difference between them. Due to such limitations and other considerations such as avoiding low complexity, avoiding long extension polymers, maintaining balanced GC content, etc., only a limited number of indices can be designed.
[0067] To label different samples in a multiplexed pool, each sample must be labeled with a unique index in a process called single-indexing. For example, to multiplex 100 samples, 100 unique indexes are required. To multiplex 10,000 samples, 10,000 unique indexes are required. In addition to the burden of synthesizing and evaluating a large number of indexes, the operation of using so many indexes is prone to cross-contamination considering the need to open and close many containers with indexes. Furthermore, it is a challenge to streamline a functional, cost-effective, and practical workflow.
[0068] Unique dual indexing consists of two unique index or barcode sequences added to each DNA fragment. These indexes allow for multiplexing, enabling multiple samples to be sequenced together in a single sequencing run while maintaining the ability to identify and separate each individual sample data during analysis. Each DNA fragment from each sample is labeled with a unique combination of dual indexes. In this way, even if all samples are sequenced together in a pooled manner, the resulting sequencing data can be de-multiplexed using these unique dual indexes to assign each read to its original source, ensuring more accurate analysis and interpretation of sequencing data for each sample. Compared to other indexing strategies, unique dual indexing allows for an increase in the number of samples sequenced per run and a decrease in the cost per sample. With just one unique dual indexing plate, a user can pool 96 samples together. In addition to unique dual indexing, another strategy for dual-indexed sequencing is to use combinatorial dual indexing, which allows sequences to repeat across rows and columns of a well plate.
[0069] By using the combinatorial dual indexing approach, a smaller number of indexes can be used to label a larger number of samples. For example, 10 indexes can be used to label one end of a DNA fragment, and 10 indexes can be used to label the other end of the DNA fragment. A total of 20 indexes can be used to label 100 (10 x 10) samples. Similarly, 100 indexes can be used to label one end of a DNA fragment, and 100 indexes can be used to label the other end of the DNA fragment. A total of 200 combinatorial dual indexes can label 10,000 (100 x 100) samples. The operational complexity of the combinatorial dual indexing approach is significantly lower than the single-indexing or dual-unique indexing approach.
[0070] Unlike conventional combinatorial dual indexing, plus unique dual indexing has distinct, unrelated, and unique index sequences (96 unique A indexes and 96 unique B indexes for a 96-well plate) that mitigate misassigned reads. However, this indexing strategy is not optimal for enhanced multiplexing capacity and large sample size. On the other hand, in conventional combinatorial dual indexing, there are only 8 unique dual pairs in a 96-well plate, with most amplicons sharing common indexes at the A index and B index ends. Conventional combinatorial dual indexing is suitable for enhanced multiplexing capacity and larger sample size, but generates significantly higher contamination misassignment.
[0071] In practical conditions, the operation with many indexes is prone to cross-contamination because there are many containers with indexes that are opened and closed. Therefore, although the plus unique dual indexing method can theoretically reduce the risk of index jumping, the operation of this method can increase the risk of cross-contamination.
[0072] However, the conventional combinatorial dual indexing method increases the risk of index jumping because the samples share the same index at one end, while the index at the other end is different. As disclosed herein, a method was developed to solve this problem and mitigate the risk of index jumping by using combinatorial quadruple indexing (CQI), by which each end of a DNA fragment / amplicon is labeled with two indexes or a total of 4 combinatorial indexes at both ends Figure 2 and Figure 3 ).
[0073] Combinatorial quadruple indexing or "CQI" mitigates or eliminates the risk of index switching and has lower operational complexity than the plus combinatorial dual indexing method. In the CQI method, the targets are first labeled with plus-indexed forward primers and plus-indexed reverse primers at the early stage of the PCR reaction Figure 2 A and Figure 2 B). These early-stage products are then amplified by plus-indexed universal primers (P5 primer and P7 primer) to hybridize and amplify the early-stage amplicons, thereby generating plus-quadruple-indexed amplicons Figure 2 C and 2D). Each plus-indexed forward (F) primer contains a read 1 sequencing primer, an X index, and a target-specific primer Figure 1 A). Each plus-indexed reverse primer (R) contains a read 2 sequencing primer, a Y index, and a target-specific primer Figure 1 B). The plus-indexed universal primer (P5) contains a P5 tail, an A index, and a read 1 sequencing primer Figure 1 C). The plus-indexed universal primer (P7) contains a P7 tail, a B index, and a read 2 sequencing primer Figure 1 D). After amplification, each DNA fragment is labeled with a total of four combinatorial indexes Figure 2). The A-index and B-index are sequenced as independent reads. The X-index and Y-index are sequenced at the beginning of read 1 and read 2 Figure 3 ) The A-index and X-index in the forward primer are combined as two unique indices (forward indices). The B-index and Y-index in the reverse primer are combined as two unique indices (reverse indices). The final amplicon product is labeled with fourfold or four combined indices. It is worth noting that a single index switch cannot convert an allowed combination to another allowed combination. The forward indices and reverse indices collectively act as fourfold indices on the same amplicon, with the ability to significantly mitigate the risk of index hopping. The plus combined fourfold index method is based on plus two sets of combined indices, which mitigates or eliminates index hopping, uses a significantly smaller number of indices compared to plus unique duplex indexing, greatly enhances multiplexing capacity, and thus is suitable for barcoding a larger number of samples and achieving the same level of specificity as plus unique duplex indexing. The CQI is user-friendly, cost-effective, time-saving, and labor-saving in a laboratory environment. An example of using a 96-well plate plus fourfold index primer layout is shown in Figures 4-6
[0074] The embodiments, applications, descriptions, and contents disclosed herein are exemplary and explanatory, and are non-limiting and not in any way binding.
[0075] The present disclosure includes a one-step, single-tube four-index barcoding multiplex amplification step, which can be applied to a wide range of biological applications. In the PCR step, the amplification and four-index barcoding of nucleic acid targets occur simultaneously in the same reaction, and then NGS and data analysis are performed.
[0076] In one aspect, the methods and compositions disclosed herein are designed to analyze target nucleic acids in samples that are analyzed for diseases (cancer / genetic disorders), drug resistance, genetic profiles, forensic or microbial and viral organisms (bacteria, fungi, parasites or viruses), allergens, and other biological applications. In some embodiments, the disclosed methods include the following steps: (1) contacting a set of nucleic acid targets in a sample with primers, wherein indexed forward and reverse target-specific primers hybridize to the targets in a test reaction in the presence of indexed universal primers; (2) amplifying the target nucleic acids under optimal amplification conditions to determine the presence or absence of the target nucleic acids; (3) sequencing the four-indexed amplification products by NGS; and (4) analyzing the generated sequence reads.
[0077] It remains a challenge in the art to develop highly multiplexed amplification methods for nucleic acid samples with accurate and high copy number sensitivity. The present disclosure relates to an NGS-based assay that combines balanced target-specific multiplex amplification and sensitive copy number quantification, with balanced sequencing reads per target, while four-indexed barcoding of each target amplicon occurs simultaneously in the amplification step.
[0078] All scientific and technical terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the disclosure belongs unless specifically defined otherwise. The examples, materials, methods, figures, and tables are merely illustrative and are not intended to limit the disclosure.
[0079] As used herein, “amplification conditions” means conditions suitable for amplification using polymerase chain reaction. The polymerase chain reaction can be multiplex PCR. Amplification conditions include, but are not limited to, the examples provided in Examples 1-6 disclosed herein.
[0080] As used herein, “indexed universal primer” means a universal primer comprising a barcode / index sequence and at least one universal sequence. See, e.g., Figure 1 C and Figure 1 D.
[0081] As used herein, “bead cleanup” means the use of bead-based purification, wherein the beads are configured to bind one or more targets. As known to one of skill in the art, bead cleanup can use positive selection (i.e., the beads are configured to capture the target of interest) or negative selection (i.e., the beads are configured to not capture the target of interest). Various beads can be used as known in the art, such as streptavidin beads or magnetic beads.
[0082] As used herein, “compatibility score” means a score of a potential forward strand target-specific primer or reverse strand target-specific primer based on different factors of target amplicon GC content, target amplicon melting temperature, target amplicon heterozygosity ratio, complementarity of candidate primer to target region; candidate primer size, target amplicon size, primer-primer interaction and amplification efficiency, and off-target rate.
[0083] As used herein, “dsDNA” means double-stranded DNA.
[0084] As used herein, “environmental source” means any potential location in a natural and / or man-made environment from which a sample can be obtained. Environmental sources include, but are not limited to: water sources such as oceans, lakes, ponds, rivers, and streams; soil sources such as soil, sand, interior or exterior dust; gas sources such as air.
[0085] As used herein, “FNA” means fine needle aspirate.
[0086] As used herein, “forward strand” means one (single) strand of a dsDNA sample.
[0087] As used herein, “indexed forward strand target-specific primer” or “indexed forward primer” means a primer configured to bind to a target sequence on the forward strand, wherein the primer is configured to introduce an index in the resulting amplicon, see, e.g., Figure 1 A.
[0088] As used herein, “GC content” means guanine-cytosine content.
[0089] As used herein, “locus” means a specific physical location or site on a genome where a particular gene or genetic marker is located.
[0090] As used herein, “microbial and viral surveillance” means the systematic monitoring and tracking of bacterial and viral pathogens in a population or a specific geographic area. It involves the collection, analysis, and reporting of data related to the occurrence, distribution, and characteristics of bacterial and viral infections.
[0091] As used herein, “index” means a short, unique nucleic acid sequence added to individual DNA fragments prior to sequencing in order to distinguish and identify different samples or DNA fragments in a single sequencing run.
[0092] As used herein, “index hopping,” “index drop-out,” or “barcode hopping / switching / exchanging” means that an index (barcode) sequence originally assigned to a particular sample is erroneously assigned to other samples in a sample pool. The result of index hopping is the erroneous assignment of sample reads (DNA fragments) in the pooled samples.
[0093] As used herein, “multiplexed barcoded amplification” means multiplexed target amplification, wherein amplification and indexing / barcoding of each target occurs simultaneously in a PCR reaction.
[0094] As used herein, “NGS” means next generation sequencing.
[0095] As used herein, “PCR” means polymerase chain reaction.
[0096] As used herein, “reverse strand” means the second (single) strand of a dsDNA sample that is complementary to the forward strand.
[0097] As used herein, “indexed reverse strand target-specific primer” or “indexed reverse primer” means a primer configured to bind to a target sequence on the reverse strand, wherein the primer is configured to introduce an index in the resulting amplicon, see, e.g., Figure 1 B.
[0098] As used herein, "sample" means a specimen or preparation in the art to which the present invention pertains. Samples can be obtained from a variety of sources, such as subjects, food, plants, and environmental sources. Where the term is used in this specification with respect to a subject, for example, "sample" means "biological sample" or equivalent thereof. "Biological sample" means any preparation obtained from biological material (e.g., an individual, a fluid, a bodily fluid, a cell line, a cultured tissue, or a tissue section) as a source. Examples of "biological sample" include fluid (e.g., blood, saliva, dental plaque, serum, plasma, urine, synovial fluid, and cerebrospinal fluid) and tissue sources.
[0099] As used herein, "species" means a group of microorganisms that share similar genetic and phenotypic characteristics.
[0100] As used herein, "subject" means an animal, preferably a mammal, and most preferably a human.
[0101] As used herein, "target-specific primer" means a primer configured to bind a specific target.
[0102] As used herein in microbiology, "type" means a strain or specific type of microorganism. In virology, it means the classification of a virus based on its genetic and antigenic characteristics.
[0103] As used herein, "type-specific primer" means a primer configured to bind a target specific to a particular microorganism or viral genome.
[0104] As used herein, "universal sequence" means a sequence configured to be targeted by a universal sequence primer.
[0105] As used herein, "unique dual indexing" means short and different DNA sequences attached to individual DNA fragments prior to sequencing. They consist of two different index or barcode sequences tagged to each DNA fragment. These indexes allow for multiplexing, enabling multiple samples to be sequenced together in a single sequencing run while maintaining the ability to identify and separate each individual sample data during analysis.
[0106] As used herein, "unique dual indexing" means short and different DNA sequences attached to individual DNA fragments prior to sequencing. They consist of two different index or barcode sequences tagged to each DNA fragment. These indexes allow for multiplexing, enabling multiple samples to be sequenced together in a single sequencing run while maintaining the ability to identify and separate each individual sample data during analysis.
[0107] The present disclosure relates to the selective amplification of a set of target sequences by multiplexed barcoded amplification and further analysis by next-generation sequencing. The present disclosure has a universal method for a wide range of nucleic acid-based biological applications.
[0108] The disclosed methods provide a number of advantages over conventional methods, including but not limited to: (1) greatly enhanced multiplexing capacity by allowing barcoding of a large number of samples with a smaller number of indices; (2) at the same time, mitigating or eliminating index hopping rate to the same level as adding unique dual indices, (3) cost and time benefits due to the large coding capacity of the adding combinatorial index method, where N is the forward index and M is the reverse index, CQI can generate unique markers for N x M samples, which significantly saves costs in primer synthesis and has time and labor benefits in laboratory operation and workflow compared to adding unique dual indices; (4) user-friendly and cost-effective assay design and assay development, as only one set of internal combinatorial indices (indices close to gene-specific primers) and one set of external combinatorial indices (indices close to the end of the amplicon) are needed for a set of samples. See, e.g. Figure 7 .
[0109] Described herein is a method of multiplexed barcoded amplification and target enrichment for NGS analysis (e.g., Figure 7 ) In various embodiments of any aspect of the disclosure, the methods and compositions are characterized by multiplexed barcoded amplification and target enrichment of target nucleic acid regions of genomic material suitable for assessing cancer, genetic disorders, drug resistance, forensic, allergens, microbial and viral organisms, and other biological applications.
[0110] In some embodiments, the disclosure can be applied to the detection, identification, and typing of microbial and viral pathogens such as sexually transmitted infections and HPV. In some embodiments, the disclosed methods include the following steps: (1) contacting indexed target-specific primers with a set of target nucleic acid sequences in the presence of indexed universal primers and hybridizing to target nucleic acid sequences in the sample; (2) subjecting the test reaction to amplification under optimal amplification conditions to generate amplicons with quadruple combinatorial indices; (3) pooling amplification products from each individual or subject sample together; (4) performing bead cleanup on a portion of the pooled amplification products to remove possible unconsumed primers and primer dimers to yield enriched amplification products; (5) performing standard normalization and quantification on a portion of the enriched amplification products; and (6) sequencing the amplicons by next-generation sequencing. See, e.g. Figure 7 .
[0111] In some embodiments, the present disclosure can be applied to nucleic acid sequence analysis of target regions for cancer, genetic disorders, forensic, pharmacogenetics, and / or drug resistance. In some embodiments, the disclosed methods comprise the following steps: (1) contacting indexed target-specific primers with target nucleic acid sequences in a sample in the presence of indexed universal primers and hybridizing to the target nucleic acid sequences in the sample; (2) subjecting the test reaction to amplification under optimal amplification conditions, generating amplicons with four-combination indexing; (3) pooling amplification products from each individual or subject sample together; (4) performing a bead clean on a portion of the pooled amplification products to remove possible unspent primers and primer dimers, to produce enriched amplification products; (5) performing standard normalization and quantification on a portion of the enriched amplification products; and (6) sequencing the amplicons by next-generation sequencing. See, e.g., Figure 7 .
[0112] In one embodiment, the barcoded universal primers comprise: (a) a universal priming portion at the 3’ end; (b) a barcode / indexing portion in the middle; and (c) a universal priming portion at the 5’ end. Figure 1 C and Figure 1 D). In some embodiments, each indexed target-specific primer comprises a universal priming portion, a barcode / indexing portion, and a specific sequence portion for a target nucleic acid sequence.
[0113] In some embodiments, the disclosed methods utilize one round of multiplexed barcoding PCR in a single test reaction for each subject, which eliminates or minimizes index hopping and cross-contamination, and additional steps in the workflow. In contrast, conventional methods using more than one round of PCR are susceptible to DNA cross-contamination, longer workflow duration, and automation challenges.
[0114] In some embodiments, the disclosed methods comprise using four-combination indexing for each amplicon in a single PCR reaction, wherein the amplicons are barcoded by 1) indexed forward target-specific primers and indexed reverse target-specific primers, and 2) indexed universal primers, thereby eliminating and minimizing index hopping and cross-contamination. In some embodiments, each indexed universal primer comprises a universal priming portion at the 3’ end; an indexing / barcoding portion in the middle; and a universal priming portion at the 5’ end. In some embodiments, each indexed target-specific primer comprises a universal priming portion, a barcode / indexing portion, and a specific sequence portion for a target nucleic acid sequence. In some embodiments, each sample is obtained from a subject, a human, an animal, a plant, a microorganism, or an environmental source. In some embodiments, the target-specific primers include primers configured to amplify at least 10, 20, 50, 100, 200, 1,000, or more targets.
[0115] In some embodiments, the disclosed methods include screening, detecting, identifying, and quantifying nucleic acid targets using next generation sequencing. In certain embodiments, target nucleic acid sequences are amplified and sequenced to detect, identify, and type pathogens such as sexually transmitted infections. In some embodiments, the disclosed methods include using next generation sequencing to analyze target nucleic acid regions of genomic material of samples related to cancer / genetic disorders, drug resistance, forensic, allergens, and other biological applications.
[0116] In some embodiments, amplification conditions such as cycle number, annealing temperature, annealing duration, extension temperature, and extension duration are adjusted to optimal conditions for amplification. In some embodiments, cycle number, amplification conditions such as annealing temperature, annealing duration, extension temperature, and extension duration are adjusted to optimal conditions for amplification based on commercial DNA polymerase instructions.
[0117] In some embodiments, the nucleic acid sample includes genomic DNA or RNA. In another embodiment, the sample includes nucleic acid molecules obtained from fresh produce, food, imported food, food production facility, farm, fresh produce, animal farm, water, putrefaction, soil, and environment. In another embodiment, the sample includes nucleic acid molecules obtained from swabs or brushes. In some embodiments, the sample includes nucleic acid molecules obtained from saliva. In some embodiments, the sample includes nucleic acid molecules obtained from urine, tissue, saliva, biopsies, sputum, swabs, formalin-fixed paraffin-embedded material (FFPE), surgical resections, cervical swabs, tumor tissue, fine needle aspirates (FNA), scrapings, swabs, mucus, urine, semen, and other non-limiting clinically or laboratory obtained samples.
[0118] In some embodiments, the obtained nucleic acid sample can be from an animal such as a human or mammalian subject. In another embodiment, the obtained nucleic acid sample can be from a non-mammalian subject such as bacteria, parasites, viruses, fungi, and plants.
[0119] In some embodiments, the disclosure relates to target amplification of at least one target sequence from a biological sample from a normal or diseased subject. In some embodiments, the disclosure relates to specific and selective target amplification of at least one target sequence in a nucleic acid sample.
[0120] In some embodiments, the indexed target-specific primers include a plurality of primers designed to amplify microbial and viral target nucleic acid sequences. In some embodiments, the target-specific primers include a plurality of indexed target-specific primers designed to selectively amplify target nucleic acid sequences of genomic material of a sample related to cancer, genetic disorders, drug resistance, forensic, allergens, and / or other biological applications.
[0121] The range of amplification varies depending on the size of the fragment and the position of the primers on the nucleic acid fragment, and the size can vary within the range. In some embodiments, the target-specific primers include a plurality of primers selectively designed to amplify target nucleic acid sequences, wherein the length of the amplified target nucleic acid sequences differ from each other by no more than 90%, no more than 70%, no more than 50%, no more than 25%, or no more than 10%.
[0122] In some embodiments, the disclosed methods involve target enrichment by multiplexed barcoded target-specific PCR, which includes the steps of contacting a nucleic acid target with a plurality of indexed target-specific primers in the presence of indexed universal primers and PCR reagents such as DNA polymerase, dNTPs, and reaction buffer; primers hybridize to complementary target nucleic acid sequences and are extended under optimal temperature and time conditions for denaturation, annealing, and extension. In some embodiments, the amplification steps can be performed in any order. In some embodiments, amplification steps, purification steps, and clean-up steps can be added or removed after optimization for optimal multiplexed target amplification for downstream processes.
[0123] In some embodiments, the described methods use PCR and a DNA polymerase as one of the reaction components. In some embodiments, there is a wide selection of DNA polymerases characterized by different properties such as thermostability, fidelity, processivity, and hot-start. The method can use a DNA polymerase with one or more of these characteristics depending on the application. In some embodiments, the concentration of DNA polymerase for multiplexed PCR can be higher than for singleplex PCR.
[0124] In some embodiments, the methods disclosed herein use amplification of target nucleic acid sequences with multiplexed polymerase chain reaction, where more than one target sequence is amplified in a test reaction. In some embodiments, the amount of nucleic acid sample required for multiplexed amplification can be about 0.1 ng. In some embodiments, the amount of nucleic acid material can be about 1 ng, 5 ng, 10 ng, 50 ng, 100 ng, or 200 ng.
[0125] In some embodiments, the disclosed methods use amplification of target nucleic acid sequences using multiplexed polymerase chain reaction, where more than one target sequence is amplified in a test reaction. Polymerase chain reactions of the prior art are performed on a thermal cycler, and each PCR cycle includes denaturation, annealing, and extension. Each PCR cycle includes at least one denaturation step, one annealing step, and one extension step for nucleic acid extension. In some embodiments, annealing and extension can be combined. In some embodiments, the methods disclosed herein include 25 to 35 PCR cycles. Each cycle or each set of cycles can have different durations and temperatures, for example, the annealing step can have stepwise increases and decreases in temperature and duration, or the extension step can have stepwise increases and decreases in temperature and duration. In some embodiments, the duration can be decreased or increased in increments of 5 seconds, 10 seconds, 30 seconds, 1 minute, 2 minutes, 4 minutes, 8 minutes, or more. In some embodiments, the temperature can be decreased or increased in increments of 0.5, 1, 2, 4, 8, or 10 °C.
[0126] In some embodiments of the disclosure, the target-specific primers comprise nucleotide modifications at the 3’ end or 5’ end or throughout the sequence. In some embodiments, the length of the target-specific portion of the primer can be about 15 to 40 bases. In some embodiments, the Tm of each target-specific primer can be about 55 °C to about 72 °C. m may be about 55 °C to about 72 °C.
[0127] In some embodiments, the disclosure features target enrichment and multiplexed barcoded amplification methods for target-specific nucleic acid amplification of microbial and viral pathogens using indexed target-specific primers. In some embodiments, the disclosure features target enrichment and multiplexed barcoded amplification methods for target-specific nucleic acid amplification of genomic material related to cancer / genetic disorders, drug resistance, forensic, allergens, and other biological applications. In some embodiments, the selected indexed target-specific primers are contacted with and hybridize to target nucleic acid sequences that can be associated with disease. In one embodiment, the indexed target-specific primers hybridize to nucleic acid sequences of different sizes in the test reaction. In some embodiments, amplicon size selection can be used to sequence amplification products of a particular length range. In some embodiments, amplicons of 100 to 250 base pairs in length can be sequenced. In some embodiments, amplicons of 150 to 300 base pairs, or 120 to 350 base pairs, or 200 to 500 base pair range or larger length range can be sequenced.
[0128] In some embodiments, any method step may be removed or repeated. In some embodiments, a purification step may be added to produce optimal results. These procedures are non-limiting, and those skilled in the art can readily add, remove, or repeat these steps to obtain optimal results.
[0129] The ability to increase the number of indexed target-specific primers in multiplex PCR allows for the simultaneous amplification of large numbers of nucleic acid targets while reducing the amount of input DNA, labor, and time. This is particularly advantageous when the amount of initial input nucleic acid material is limited.
[0130] In some embodiments of the disclosed method, the primer design method selects candidate target-specific primers based on the following step-by-step procedure: (1) extracting genomic sequences around each target variant location; (2) for each variant in the target sequence, designing primers with appropriate GC content, T... m (2) Target-specific forward and reverse strand target-specific primers at different distances from each target variant; (3) For each primer, search for target genome sequences for off-target matches; filter primers and retain those that pass the off-target threshold; (4) Search for the 3′ end portion of each primer for complementary matches with the sequence of the primer set; filter primers stepwise, wherein primers with the most complementary matches at the 3′ end are removed first; (5) Synthesize primers and run the entire wet laboratory experiment including next-generation sequencing; calibrate the performance of each primer and filter out primers with unsatisfactory performance. In some embodiments, steps 2 to 4 and steps 2 to 5 of the primer selection procedure are repeated until each target variant is covered by at least one forward strand target-specific primer and one reverse strand target-specific primer from the primer set.
[0131] In some embodiments, this disclosure features a primer design method that eliminates incompatible primers that form pseudoproducts, such as primer dimers, that inhibit efficient amplification in high-multiplexity PCR. Such elimination systems remove or significantly minimize non-productive pseudoproducts such as primer dimers. In addition to improving downstream procedures such as high-throughput sequencing, the removal of incompatible and problematic primers significantly improves the overall performance and efficiency of high-multiplexity PCR. Pseudoproducts and primer dimers lead to significant failures in obtaining optimal sequence results, and a large portion of sequencing reads can be nonspecific and non-informative.
[0132] In some embodiments, the primer selection method is characterized by a primer compatibility score with respect to primer-primer interactions and specific target nucleic acid hybridization without non-specific priming or hybridization to off-target regions. A higher compatibility score of a candidate target-specific primer characterizes specific hybridization to the target nucleic acid with no or minimal interaction with other primers in the primer set. Primers that do not meet the compatibility score (i.e., above a minimum threshold) are eliminated. In various embodiments of the disclosed method, the compatibility score is calculated for at least 80%, 90%, 95%, 98%, 99%, or 99.5% of possible combinations of candidate primers in the set. The compatibility score in primer selection is calculated based on a number of parameters such as: target amplicon GC content, target amplicon melting temperature, target amplicon heterozygosity ratio, complementarity of candidate primers to the target region, candidate primer size, target amplicon size, and amplification efficiency. Due to the fact that determining the compatibility score involves several aspects, an average score is calculated based on multiple parameters, and this average can vary from application to application. The primer selection method will continue to eliminate low compatibility primers, and the elimination process is repeated until an optimal selection primer set is obtained that generates highly multiplexed target amplification PCR with no or minimal primer dimer.
[0133] In some embodiments, the primer selection method is characterized by a primer compatibility score with respect to primer-primer interactions and specific target nucleic acid hybridization without hybridization to off-target regions. Primers with a low compatibility score (i.e., above a minimum threshold) will be eliminated. However, if there are limitations in primer selection in certain applications, the minimum threshold can be increased to a higher second threshold level to facilitate primer selection for the primer set. In some embodiments, the selection process is repeated until candidate primers equal to or below the minimum threshold of the second level are selected.
[0134] In one embodiment, the disclosed methods feature multiplex amplification and target enrichment by utilizing indexed target-specific primers (in combination with indexed universal primers) that contact target nucleic acid sequences of genomic material of a sample related to cancer / genetic disorders, drug resistance, forensic, allergen, microorganism, fungi, parasite, virus, and other biological applications, where primer-dimers can be reduced or minimized by adjusting different parameters such as duration of annealing step, increase or decrease in temperature increment in combination with cycling arrays. In some embodiments, primer concentration can be reduced, and annealing temperature and duration can be increased, thereby enabling specific amplification while reducing or minimizing primer-dimers (primers have longer period to hybridize to target nucleic acids). In some embodiments, the concentration of primers can be 500 nM, 250 nM, 100 nM, 80 nM, 70 nM, 50 nM, 30 nM, 10 nM, 2 nM, 1 nM, or less than 1 nM. In some embodiments, the annealing temperature can be 1 minute, 3 minutes, 5 minutes, 8 minutes, 10 minutes, or longer. In some embodiments, amplification with longer annealing time uses 1 cycle, 2 cycles, 3 cycles, 5 cycles, 8 cycles, 10 cycles, or more cycles followed by standard annealing duration.
[0135] In one aspect, the disclosed methods include a step of amplifying a selective target nucleic acid sequence of a sample related to cancer / genetic disorders, drug resistance, forensic, allergen, microorganism, fungi, parasite, virus, and other biological applications. In some embodiments, the methods include a step of contacting a nucleic acid sample with indexed target-specific primers in a test reaction in the presence of indexed barcoded universal primers. In some embodiments, the methods include a step of determining the presence or absence of a target amplification product. In some embodiments, the methods include a step of determining the sequence of an amplified target product. In some embodiments, the methods identify microorganisms to the strain or sub-strain level. In some embodiments, less than 50%, 40%, 30%, 20%, 10%, 5%, 0.5%, or 0.1% of the amplified products are primer-dimers or pseudo-products. In one embodiment, there can be more than one set of target-specific primers, for example, there can be two sets of target-specific primers for two test reactions, three sets of target-specific primers for 3 test reactions, or five sets of target-specific primers for 5 test reactions or more. In some embodiments, the sample can also be divided into multiple parallel multiplex test reactions with multiple sets of target-specific primers for practical reasons such as limitations in primer design or selection.
[0136] In various embodiments, the concentration of each primer can be 500 nM, 250 nM, 100 nM, 80 nM, 70 nM, 50 nM, 30 nM, 10 nM, 2 nM, 1 nM, or less than 1 nM. In various embodiments, the primer concentration of each primer can be 1 mM to 1 nM, 1 nM to 80 nM, 1 nM to 100 nM, 10 nM to 50 nM, or 1 nM to 60 nM. In some embodiments, the GC content of the target-specific primer can be 40% to 70%, or 30% to 60% or 50% to 80% or 30% to 80%. In some embodiments, the primer GC content range can be less than 20%, 15%, 10%, or 5%. In some embodiments, the melting temperature (T m ) of the target-specific primer can be 55 °C to 65 °C or 40 °C to 72 °C, or 50 °C to 68 °C. In some embodiments, the melting temperature range of the primer can be less than 20 °C, 15 °C, 10 °C, 5 °C, 2 °C, or 1 °C. In some embodiments, the length of the target-specific primer can be 20 to 90 bases, 40 to 70 bases, 20 to 40 bases, or 25 to 50 bases. In some embodiments, the length range of the primer can be 60, 50, 40, 30, 20, 10, or 5 bases. In some embodiments, the 5' region of the target-specific primer is a universal priming site that is not complementary or specific to any region of the target nucleic acid.
[0137] In one aspect, the disclosure relates to a kit comprising a set of indexed target-specific primers; the primers are designed and selected based on the described criteria to have minimal primer-primer interactions or non-specific priming. In another embodiment, the kit can be formulated for detection, screening, diagnosis, prognosis, and treatment of disease. In another embodiment, the kit can be formulated for detection of resistance. In some embodiments, the kit can be used for screening, detection, identification, genotyping, subtyping, and monitoring of bacteria, fungi, parasites, and viruses. In some embodiments, the kit can be used for analysis of samples related to cancer / genetic disorders, forensic, resistance, pharmacogenetics, and other biological applications. In some embodiments, the kit can be used for detection of allergens.
[0138] In some embodiments, the disclosed methods comprise the following steps: (1) contacting a set of indexed target-specific primers to target nucleic acid sequences in the presence of indexed barcoded universal primers and hybridizing to the target nucleic acid sequences in each sample in a test reaction; (2) subjecting the test reaction to amplification under optimal amplification conditions to generate amplicons containing a four-combinatorial index; (3) pooling the amplification products from each individual sample together; (4) performing a bead clean on a portion of the pooled amplification products to remove possible primer dimers to yield enriched amplification products; (5) performing standard normalization and quantification on a portion of the enriched amplification products; and (6) sequencing the amplicons by next generation sequencing. The methods can further comprise additional steps, such as purification. In one aspect, highly multiplexed PCR is used in the disclosed methods. In some embodiments, 1 to 10 PCR cycles can be performed for the PCR; in some embodiments, 1 to 15 cycles or 1 to 20 cycles or 1 to 25 cycles or 1 to 30 cycles, 1 to 35 cycles or more can be performed.
[0139] In another embodiment, the disclosed methods can be used in a multiplexed manner when more than two targets are amplified, and are not limited to any number of multiplexing.
[0140] In some embodiments, the amplification products can be sequenced by next generation sequencing platforms. Next generation sequencing refers to non-sanger based massively parallel DNA nucleic acid sequencing technologies that can sequence millions to billions of DNA strands in parallel. Examples of current state-of-the-art next generation sequencing technologies and platforms are Illumina® platform (reversible dye terminator sequencing), pyrosequencing, Ion Torrent sequencing, SMRT sequencing, GeneReader sequencing technology, Element BioScience sequencing platform, and Oxford Nanopore® sequencing. The present disclosure is not limited to these next generation sequencing technology examples.
[0141] Example 1
[0142] Evaluation of the combinatorial four-indexed method to remove reads generated from index hopping on three synthetic DNA targets
[0143] Materials and Methods
[0144] Sample and assay design: Three target DNA templates were synthesized (IDT DNA, Coralville, IA). The following primers were synthesized for amplification and indexing: three indexed forward target-specific primers with three different indices, three indexed reverse target-specific primers with three different indices, three P5 primers (indexed universal primers) with three different indices, and three P7 primers (indexed universal primers) with three different indices.
[0145] Amplification: One-step multiplexed barcoding PCR was performed on the synthetic DNA templates by the indexed target-specific primers in the presence of P5 and P7 primers (indexed universal primers), DNA polymerase, dNTPs, and PCR buffer.
[0146] Next-generation sequencing: All amplicons were pooled into one tube and purified by SPRIselect beads (Beckman Coulter, Brea, CA). The concentration of the purified sample was measured on Qubit TM 3 and the concentration was normalized for sequencing. The Mid Output Sequencing Kit was used with the MiniSeq TM System to sequence the library.
[0147] Results: Figure 8 Column A lists the read counts for each target with different A / B index pairs. The numbers that are allowed to combine are highlighted with a gray background. There are 57 unintended reads. After removing the reads with unallowed index combinations, there are 16 unintended reads Figure 8 B). The correction process removed 72% of the unintended reads. The small number of unintended reads can be due to synthesis manufacturing errors, which can be addressed by using more diverse sequences in the indices. This result demonstrates that the combined quadruple indexing approach significantly mitigates index hopping and can be further reduced by adding more diverse index sequences. Figure 8 The numbers with a gray background in Column A are lower than their corresponding wells in Figure 8 Column B because the reads that perfectly match the target index are kept, while the reads that have the same A / B index pair but do not match the target index are removed.
[0148] Example 2
[0149] Index hopping evaluation of sexually transmitted infection clinical samples by multiplexed barcoding amplification and combined quadruple indexing approach
[0150] Materials and methods
[0151] Samples: The disclosed method was performed using 480 DNA samples that had been tested for human papillomavirus (HPV).
[0152] Assay description: The HPV-STI assay applied in this experiment utilized a combination of type-specific primers targeting 29 HPV (including HPV68a and HPV68b) and 13 STI (including Chlamydia trachomatis serotypes) and GAPDH internal control. Amplification and barcoding / indexing of each sample was performed simultaneously in a single tube and single PCR reaction. High-risk HPV included HPV16, 18, 31, 33, 35, 39, 45, 51, 52, 56, 58, 59, 66, 68a, 68b, 73, 26, 53, 82, and low-risk HPV included HPV6, 11, 40, 42, 43, 44, 55, 61, 81, 83. The 13 STI included Chlamydia trachomatis, Treponema pallidum, Mycoplasma genitalium, Trichomonas vaginalis, Neisseria gonorrhoeae, HSV-1, HSV-2, Mycoplasma hominis, Ureaplasma urealyticum, Ureaplasma parvum, Varicella zoster virus, Haemophilus ducreyi. Chlamydia serotypes included LI, L2, B, D, E, F, and G.
[0153] Multiplexed barcoding amplification: One-step multiplexed barcoding PCR was performed on 480 clinical samples (5 x 96 PCR plates) using 29 HPV, 13 STI, and one internal control indexed target-specific primers in the presence of indexed universal primers, sample DNA, DNA polymerase, dNTPs, and PCR buffer. Each generated amplicon contained a quadruplex index.
[0154] Next-generation sequencing: After amplification, the clinical samples from each plate were pooled into one tube. Then, a portion of the sample was purified with SPRI beads (Beckman Coulter, CA, USA) according to the manufacturer’s instructions. A portion of the purified sample was used to determine the concentration of the sample and to normalize the concentration for sequencing. The remaining sample was used for sequencing. TM 3The concentration of the purified sample was measured, and the concentration was normalized for sequencing. The remaining sample was used for sequencing. Mid Output Sequencing Kit, using MiSeqTM Libraries were sequenced by the system. One library was sequenced twice with two different concentrations. A total of five sequencing runs were performed for the five libraries, with each library consisting of 96 samples.
[0155] Results
[0156] Figure 9 Five libraries are listed, with column 2 listing the number of positive (read count > 0) targets before index jump reads were removed, column 3 listing the positive (read count > 0) targets after index jump removal, and column 4 listing the percentage of false positives per library due to index jumping. Libraries 1-5 generated 460 (66.6%), 276 (57.1%), 100 (45.7%), 90 (55.9%), and 467 (65.5%) false positive targets that were removed by the data analysis software as impermissible index combinations. Reads that perfectly matched the quadruplex target index were assigned to each designated sample, while reads with the same A / B index pair but not matching the target index were eliminated.
[0157] Results from the five libraries show that the addition of the quadruplex index method significantly mitigates false positives due to index jumping, with many samples that were reported as positive actually being negative.
Claims
1. A method of indexing at least one nucleic acid sample comprising the steps of: For each sample, hybridizing a plurality of indexed target-specific primers to the sample in the presence of an indexed universal primer in a single reaction vessel to form a plurality of test reactions, wherein at least one target-specific primer is configured to bind to at least one target nucleic acid sequence, and wherein the nucleic acid target is labeled with an indexed target-specific primer during primer extension in amplification; Simultaneously amplifying the indexed-labeled product using the indexed universal primer, wherein each test reaction is subjected to amplification conditions configured to generate an amplicon containing four combined indices, wherein one index (A) is located at the 5' end of the amplicon Watson (forward) strand, one index (B) is located at the 5' end of the amplicon Crick (reverse) strand, one index (X) is located on the Watson (forward) strand adjacent to the target-specific primer, and one index (Y) is located on the Crick (reverse) strand adjacent to the target-specific primer; and Performing next generation sequencing on at least a portion of the pooled amplicons.
2. The method of claim 1, wherein, Within a set of samples, the same X index (located on the Watson strand adjacent to the target-specific primer) is shared among more than one sample in the same set, the same Y index (located on the Crick strand adjacent to the target-specific primer) is shared among more than one sample in the same set, but wherein the combination of X index and Y index is unique for each sample within the same set, and wherein the same A index (located at the 5' end of the amplicon Watson strand) is shared among more than one sample in the same set, and the same B index (located at the 5' end of the amplicon Crick strand) is shared among more than one sample in the same set, but wherein the combination of A index and B index is unique for each sample within the same set.
3. The method of claim 1, wherein each indexed universal primer comprises: a universal priming portion located at the 3' end; an index portion in the middle; and a universal priming portion located at the 5' end.
4. The method of claim 1, wherein each indexed target-specific primer comprises a universal priming portion, an index portion, and a specific sequence portion for a target nucleic acid sequence.
5. The method of claim 1, wherein the generated amplicon has four indices.
6. The method of claim 1, wherein the amplification is performed in two steps, comprising a first step in which a target is amplified using an X-indexed primer and a Y-indexed primer to form a first amplicon; and amplifying the first amplicon using an A-indexed primer and a B-indexed primer and adding the A index and the B index to produce a second amplicon having four indices.
7. The method of claim 1, wherein each nucleic acid sample is obtained from a subject, a human, a plant, a food, one or more plants, or an environmental source.
8. The method of claim 1, wherein the indexed target-specific primers comprise at least two primers, wherein each primer is configured to amplify a different nucleic acid target.
9. The method of claim 1, further comprising the step of pooling the enriched amplicons from each sample, followed immediately by a bead clean-up.
10. The method of claim 1, further comprising the step of quantifying each nucleic acid target in each sample after sequencing the enriched amplicon.
11. The method of claim 1, wherein each index comprises a different sequence.
12. A method of indexing at least one nucleic acid sample comprising the steps of: for each sample, hybridizing a plurality of indexed target-specific primers to the sample in the presence of an indexed universal primer in a single reaction vessel to form a plurality of test reactions, wherein at least one target-specific primer is configured to bind to at least one target nucleic acid sequence, and wherein the nucleic acid target is labeled with the indexed target-specific primer during primer extension in amplification first; and simultaneously amplifying the indexed-labeled product using the indexed universal primer, wherein each test reaction is subjected to amplification conditions configured to generate an amplicon containing four combined indexes, wherein one index (A) is located at the 5' end of the amplicon Watson (forward) strand, one index (B) is located at the 5' end of the amplicon Crick (reverse) strand, one index (X) is located on the Watson (forward) strand adjacent to the target-specific primer, and one index (Y) is located on the Crick (reverse) strand adjacent to the target-specific primer; pooling the resulting amplicons; and performing next generation sequencing on at least a portion of the pooled amplicons.
13. A kit comprising a plurality of indexed target-specific primers configured to bind to target-specific nucleic acid sequences selected from the group consisting of: cancers, genetic disorders, allergens, forensic test targets, microbial and viral species, and any nucleic acid-based biological target.