Efficient methods and compositions for multiplex target amplification PCR

By using a combination of methylated primers and endonucleases in PCR amplification to digest and remove primer dimers, the problem of amplification artifacts in multiplex targeted amplification is solved, achieving efficient, specific, and uniform multiplex amplification, suitable for high-throughput analysis of various sample types.

CN114929896BActive Publication Date: 2026-03-31CHAPTER DIAGNOSTICS INC
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-10-26
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Existing technologies in multiplex targeted amplification PCR suffer from primer dimers and non-specific amplification products, leading to amplification artifacts, affecting amplification efficiency and the accuracy of downstream analysis. This is especially true when there is limited DNA sample, which limits the application and sensitivity of multiplex amplification.

Method used

By employing a combination of methylation primers and methylation-dependent restriction enzymes, methylation modification sites are introduced during amplification, primer dimers are digested and removed, and multiplex amplification is performed by connecting universal adaptors and barcoded primers, ensuring the specificity and uniformity of the amplification products.

Benefits of technology

It effectively removes primer dimers, improves the specificity and uniformity of multiplex amplification, reduces non-specific amplification, enhances amplification efficiency and the accuracy of downstream analysis, and is particularly suitable for high-throughput analysis of limited DNA samples.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure HDA0003707485470000011
    Figure HDA0003707485470000011
  • Figure HDA0003707485470000021
    Figure HDA0003707485470000021
  • Figure HDA0003707485470000031
    Figure HDA0003707485470000031
Patent Text Reader

Abstract

The present disclosure relates to methods of enzymatic treatment of double-stranded PCR amplification products for elimination or minimization of primer-dimers in multiplex PCR reactions and for efficient ligation of adaptors. The present disclosure relates to methods and compositions that allow for more efficient highly multiplexed target amplification by minimizing laboratory steps, eliminating primer-dimers, and increasing efficiency of adaptor ligation compared to conventional methods, compositions, and kits. The disclosed methods use multiple target-specific primers to specifically and selectively amplify targets in a subject's genome. The disclosed methods can be used for a number of downstream procedures and analyses, including DNA sequencing.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Related applications

[0002] This application claims priority to U.S. Provisional Patent Application 16 / 720,823, filed December 19, 2019, entitled “Efficient Methods and Compositions for Multiplex Target Amplification PCR,” the entire contents of which are incorporated herein by reference. Background Technology

[0003] Rapid and accurate identification of genetic variants that cause disease, drug response, or adverse drug effects is crucial for patient diagnosis, personalized treatment, companion diagnostics, and prognosis. Many cancers and congenital or hereditary conditions are complex diseases that may be associated with multiple genes and may involve heterozygous mutations. Furthermore, these mutations may be present in small amounts in a given sample.

[0004] Targeted gene sequencing is an effective method for analyzing specific mutations and genetic changes in a sample. It allows for deeper focus on a selected group of genes or gene regions known or suspected of being associated with a specific disease. Multiple genes can be analyzed in parallel across the genome, reducing time and cost. Targeted gene sequencing also produces more detailed datasets for the genes of interest, providing high sensitivity, specificity, and deep coverage, allowing for easier gene analysis and detection of rare gene variants.

[0005] Furthermore, since most known pathogenic mutations occur in the coding regions of genes, it is more practical to focus sequence analysis on a set of genes associated with a specific disease. Targeting, capturing, and sequencing genomic regions of interest is particularly valuable in clinical settings, enabling analysis of each target at greater depth and simultaneous analysis of large numbers of specific gene panels and high sample counts. Therefore, there remains a strong need to design applications that focus on specific genomic intervals or gene sets.

[0006] While whole-genome sequencing has become more cost-effective and practical for many indications, focused target-specific combinations continue to offer the following advantages: better coverage of target regions and stronger ability to detect multiple variant types (including CNVs and complex genomic rearrangements), significantly lower cost, higher throughput, simpler bioinformatics analysis, and more concentrated detection. This focused target-specific combination eliminates the need to handle minor / accidental discoveries that would otherwise be inevitable in whole-genome sequencing. Furthermore, targeted sequencing of specific regions of interest in a large number of samples is more cost-effective in characterizing disease states than sequencing the entire genome of a smaller number of individuals. For disease assessment, treatment, and prognosis, broader coverage of specific targets and deeper enrichment of targets allow for a wider dynamic range of allele frequencies and detection of a few sequences and low-frequency variants. Effective and specific target enrichment methods allow for more effective targeted sequencing. Important parameters for target enrichment are: (i) sensitivity; (ii) specificity; (iii) homogeneity; (iv) reproducibility; (v) cost; (vi) ease of use; and (vii) the amount of DNA required per experiment.

[0007] The human genome contains approximately 3 billion base pairs, about 21,000 coding genes, and over 220,000 exons. Exons comprise about 1–2% of the genome, and on average, each gene has 9 exons, with an average exon size of 170 nucleotides. Next-generation sequencing (NGS) is an important tool for analyzing the genome, offering higher sensitivity than Sanger sequencing and allowing the detection of mutations from samples containing only a few cells. It can be used to detect a variety of sequence variations in DNA and RNA, such as single and multinucleotide variants, insertions, deletions, and gene copy number variations. NGS can also be used to analyze gene expression levels by quantitatively measuring the levels of mRNA, microRNA, and factors that influence them, such as gene promoter methylation. Currently, there are three types of NGS used for sequencing DNA: (a) whole-genome sequencing (WGS); (b) whole-exome sequencing (WES); and (c) targeted sequencing. WGS covers and analyzes the entire genome content of an individual, while WES covers only the protein-coding regions of the genome. In contrast, targeted sequencing focuses on a group of genes or specific regions of the genome that are linked by common pathological mechanisms or known clinical phenotypes.

[0008] Notably, NGS is revolutionizing the molecular characterization of genetic diseases and cancers for the discovery of driver mutations and routine screening for genomic aberrations. Due to its wide range of applications, NGS has permeated many areas of life sciences and has significantly impacted medical genetics in both research and diagnostics. Isolating high-priority regions of the genome greatly enhances outcomes in clinical, diagnostic, and research settings. However, focusing on specific regions of interest using NGS requires enrichment of the relevant target region. Notably, target enrichment allows for increased coverage of both the target region and the specific region of interest, thereby facilitating sample multiplexing and simplifying sequence read data analysis. The fundamental advantages of target enrichment in genomic assays include enrichment factor, coverage or read depth, uniformity or homogeneity of coverage across the target region of interest, reproducibility, specificity (i.e., the on-target / off-target ratio of sequence reads), the amount of input DNA required, and the total cost per target base of useful sequence data.

[0009] PCR-based target enrichment methods are relatively rapid, requiring fewer steps and less input DNA, making them more suitable for samples containing small amounts of input DNA, such as FFPE, cfDNA, and ctDNA. The specificity of PCR target enrichment of regions of interest is significantly affected by the number of primers used in the reaction, and primer characteristics such as GC content and the presence of variations in the target region can interfere with optimal primer hybridization, leading to amplification failure of certain sequences, also known as allele loss. Amplicon size and coverage are important considerations for PCR-based target enrichment to produce uniform and consistent coverage.

[0010] Multiplex amplification of target nucleic acid sequences allows for high-volume application in a single polymerase chain reaction (PCR). The advantage of using multiple target-specific primers in a single PCR reaction is that it allows for the efficient multiplex amplification of selective targets, saving time, reducing costs and labor, and increasing throughput. However, increasing the number of oligonucleotide primers in the reaction can introduce primer cross-reactivity and the formation of amplification artifacts (e.g., primer dimers) or lead to nonspecific priming of nonspecific amplification products. Furthermore, due to this cross-reactivity, some nucleic acid targets may not amplify, resulting in target loss. These amplification artifacts can over-consume amplification components and reagents, such as dNTPs and DNA polymerase, affecting the overall efficiency and quality of the amplification reaction. These artifacts from highly multiplexed amplification can also affect downstream procedures, such as sample preparation for next-generation sequencing. In such cases, nonspecific amplification or artifactual amplification can be carried over to downstream steps, such as next-generation sequencing reads, producing overly dominant, non-informative sequencing reads.

[0011] Multiplex amplification allows for the amplification of multiple targets of interest in a single reaction, advantageously increasing the number of target regions that can be amplified from a limited amount of DNA in a single reaction, where hundreds to thousands of target regions can be amplified simultaneously for sequencing. Selective multiplex amplification has wide applications in clinical and research settings and can be used for mutation detection and analysis, single nucleotide polymorphism (SNP) detection in microorganisms and viruses, deletions and insertions, genotyping, copy number variation (CNV), epigenetic and methylation analysis, gene expression, and transcriptome analysis. These applications can be used for disease diagnosis, prognosis, and treatment.

[0012] However, as the number of nucleic acid target regions used for selective amplification increases, a proportionally larger number of primers needs to be introduced into the reaction. Higher primer numbers and concentrations in a single assay reaction can increase amplification artifacts, such as primer dimers, nonspecific amplification, and hyperamplifiers, and can lead to amplification failure due to interference between primers, with each primer potentially negatively impacting downstream steps. A common approach to avoid or minimize these amplification artifacts is to use commercial or in-house software packages to design primers for multiplex amplification assays to avoid or reduce primer dimer formation and nonspecific priming. This can be done by: (1) designing target-specific primers with rigorous design considerations to mitigate primer interactions; and (2) grouping primers into optimal subsets of non-overlapping pools to avoid artifacts.

[0013] However, such efforts have not completely solved the above problems. Therefore, there is a great need for methods or compositions for highly multiplex amplification of target-specific sequences without or with minimal amplification artifacts, such as primer dimers and nonspecific amplification products, as well as to eliminate or minimize primer packing for individual test reactions, which adds extra steps. Summary of the Invention

[0014] In some embodiments, this disclosure relates to methods, compositions, and kits for multiplex target amplification and target enrichment prior to downstream analysis, such as next-generation sequencing. In some embodiments, this disclosure relates to a method comprising the step of target enrichment amplification in a DNA or RNA sample using multiple target-specific primers, wherein said amplification is performed under optimal conditions in the presence of an amplification reagent comprising polymerase and dNTPs. In various embodiments, the method further includes the step of converting mRNA in the sample into cDNA. In some embodiments, the target-specific primers of this method comprise a target-specific sequence at the 3' end and an auxiliary sequence at the 5' end, wherein said auxiliary sequence is configured to allow digestion of PCR products around methylated sites with appropriate restriction enzymes.

[0015] In some embodiments, this disclosure relates to a method comprising the steps of: (1) hybridizing two or more target-specific primers to a nucleic acid target sequence in a test reaction, wherein the target-specific primers comprise a methylated universal helper portion having a methylation-dependent restriction enzyme recognition site and a target-specific portion configured to target a nucleic acid target sequence in a sample; (2) amplifying the test reaction under optimal amplification conditions to produce an amplification product containing an amplicon; and (3) digesting the amplification product with a methylation-dependent restriction enzyme to form a digestion product, wherein the digestion product contains each strand of the chain. (4) Size-selectively purify the digested product to remove digested primer dimers and unused primers to form a digested amplicon containing dsDNA; (5) Ligate a universal adaptor to the dsDNA from the digested amplicon to form a ligation product, wherein the ligation universal adaptor contains a universal sequence portion and sticky ends; (6) Amplify the ligation product with barcoded universal primers complementary to the sequence on the ligation universal adaptor to form a final amplification product; and (7) Quantify the final amplification product for next-generation sequencing.

[0016] In some embodiments, this disclosure relates to a method comprising the steps of: (1) hybridizing two or more target-specific primers to a nucleic acid target sequence in a test reaction, wherein the target-specific primers comprise a complementary universal helper portion at the 5' end and a target-specific portion of a nucleic acid target sequence configured to target a sample; (2) performing a first amplification of the test reaction under optimal amplification conditions using a universal helper primer to form an amplification product; (3) subjecting a portion of the amplification product to a second amplification using a methylated universal helper primer to form a second amplification product, wherein the methylated universal helper primer comprises a restriction enzyme recognition sequence; and (4) using a methylation-dependent restriction enzyme restriction endonuclease. (5) The digestion product is digested with an enzyme to form a digested product containing an amplicon, the amplicon having sticky ends at each end of the strand; (6) The digested product is size-selectively purified to remove digested primer dimers and unused primers to form a digested amplicon containing dsDNA; (7) A universal adaptor is ligated to the dsDNA from the digested amplicon to form a ligation product, wherein the ligation universal adaptor contains complementary sticky ends and a universal sequence portion; (8) The ligation product is amplified a third time using barcoded universal primers complementary to the sequence on the ligation universal adaptor; and (9) The final amplified product is quantified for next-generation sequencing.

[0017] In some embodiments, the disclosed method includes the steps of: (1) annealing at least 10, 20, 100, 500, 1,000, 2,500, 5,000, 10,000, 25,000, 50,000, 80,000, 100,000, 150,000 or more target nucleic acid sequences with target-specific primers, wherein the target-specific primers comprise both forward and reverse primers; and (2) using primer extension to generate target amplification products, wherein the target amplification products comprise amplicones of different sizes. In some embodiments, the disclosed method further includes the step of determining the presence or absence of at least one target amplification product. In some embodiments, the disclosed method further includes the step of determining the sequence of at least one target amplification product. In some embodiments, the disclosed method includes the step of using at least 10, 20, 100, 500, 1,000, 2,500, 5,000, 10,000, 25,000, 50,000, 80,000, 100,000, 150,000 or more forward and reverse primers, wherein each forward and reverse primer is directed to hybridize with a specific target nucleic acid sequence. In some embodiments, the disclosed method further includes the step of RNA analysis, which measures RNA expression in both control and subject samples by comparing control expression levels with subject sample expression levels.

[0018] In some embodiments of the disclosed method, amplification is performed using prior art polymerase chain reaction (PCR). In some embodiments, the PCR is performed with an annealing time greater than 0.1, 0.5, 1, 2, 5, 8, 10, or 15 minutes. In some embodiments, the PCR is performed with an extension time greater than 0.1, 0.5, 1, 2, 5, 8, 10, or 15 minutes.

[0019] In some embodiments, the disclosed method further includes the steps of: introducing a modification into the amplification product, wherein the modification enables digestion by a restriction endonuclease; and digesting the amplification product with a restriction endonuclease, wherein the restriction endonuclease is configured to cleave the DNA in the amplification product at a fixed distance away from the modification. In some embodiments, the modification is methylation. In some embodiments, the methylated amplification product is digested at a designated restriction site by a methylation-dependent endonuclease restriction enzyme. In some embodiments, the methylation-dependent endonuclease restriction enzyme is configured to digest double-stranded DNA. In some embodiments, the methylation-dependent endonuclease restriction enzyme is configured to digest both double-stranded and single-stranded DNA.

[0020] In some embodiments of this method, the concentration of the target-specific primer is 500, 250, 100, 80, 70, 50, 30, 10, 2, or 1 nM. In some embodiments, the GC content of the target-specific primer is 40% to 70%, or 30% to 60%, or 50% to 80%. In some embodiments, the GC content of the target-specific primer ranges from less than 20%, 15%, 10%, or 5%. In some embodiments, the melting temperature (T0) of the target-specific primer is... m The temperature range is 55°C to 65°C, 40°C to 70°C, or 55°C to 68°C. In some embodiments, the target-specific primers are 20 to 90 bases, 40 to 70 bases, 20 to 40 bases, or 25 to 50 bases in length. In some embodiments, the 5'-region of the modified methylated primer contains an auxiliary sequence that is not complementary to or non-specific to any nucleic acid region in the sample. In some embodiments, the target-specific primers specifically anneal to a target region, a gene, or a different region on a different exon of a gene. In some embodiments, the target-specific primers are configured to anneal to a target region and simultaneously amplify the target nucleic acid in the sample.

[0021] In some embodiments, this disclosure relates to a method for detecting cancer in a sample. In some embodiments, this disclosure relates to a method for detecting the presence or absence of a congenital or hereditary disease in a sample. In some embodiments, this disclosure relates to a method for detecting the ploidy status of a gestational fetus in a sample, wherein the method includes the step of measuring allele counts at polymorphic sites to determine the ploidy status. In some embodiments, the method further includes the step of selecting a treatment for a subject based on target sequences detected in the sample.

[0022] In some embodiments, the target sequence contains a clinically operable mutation. In some embodiments, the target sequence is associated with drug resistance or pharmacogenetic therapy (companion diagnostics). In some embodiments, the detection, identification, and / or quantification of genetic markers may be associated with organ transplantation or organ rejection.

[0023] In some embodiments, the sample is obtained from a healthy subject. In some embodiments, the sample is obtained from a pregnant subject, wherein the sample comprises maternal and fetal nucleic acids. In some embodiments, the sample is obtained from a subject suspected of having a disease or with an increased risk of having a disease, and wherein the target sequence comprises a disease-related mutation or variant. In some embodiments, the disease is cancer. In some embodiments, the sample comprises whole-genome DNA, mechanically or enzymatically fragmented DNA, cDNA, formalin-fixed paraffin-embedded tissue (FFPE), cell-free DNA (cfDNA), or circulating tumor DNA (ctDNA). In some embodiments, the sample comprises cell-free DNA from the plasma of a pregnant woman. In some embodiments, the nucleic acid target sequence in the sample comprises single nucleotide polymorphisms (SNPs), mutations, gene rearrangements and fusions, short tandem repeats, genes, exons, coding regions, and exomes. In some embodiments, the sample comprises mRNA. In some embodiments, the mRNA in the sample is reverse transcribed to produce DNA. In some embodiments, the nucleic acid is obtained from a single cell. In some embodiments, the sample comprises nucleic acids obtained from blood, serum, plasma, cerebrospinal fluid, urine, tissue, saliva, biopsy, sputum, swabs, surgical excisions, cervical swabs, tears, tumor tissue, fine needle aspiration (FNA), circulating cell-free DNA (cfDNA) and circulating tumor DNA (ctDNA), scrapings, swabs, mucus, urine, semen, hair, and other non-restrictive clinical or laboratory or forensic samples.

[0024] In one embodiment, at least one linker is a double-stranded oligonucleotide comprising a sticky end and a universal priming sequence, wherein the sticky end sequence of the linker is configured to be complementary to a sticky end of the digestion product produced by the methylation-dependent endonuclease restriction enzyme, wherein the universal priming sequence is used for downstream amplification and sequencing. In some embodiments, the disclosed method includes the step of ligating at least one linker to a target nucleic acid. In some embodiments, the linker comprises: (a) a universal priming sequence; (b) a barcode sequence (DNA marker sequence); and (c) a sticky end, wherein the sticky end sequence of the linker is complementary to a sticky end of the digestion product produced by the methylation-dependent endonuclease restriction enzyme. In some embodiments, the linker comprises: (a) a universal priming sequence; and (b) a sticky end, wherein the sticky end sequence is complementary to the sticky end of the digested amplicon according to the methylation-dependent endonuclease restriction enzyme used. In some embodiments, the linker comprises sticky ends complementary to both ends of the restriction enzyme digestion product. In another embodiment, the linker includes a universal priming sequence configured to allow amplification of a target sequence having a linker to be further amplified. In some embodiments, the ligation step uses more than one linker.

[0025] In some embodiments, this disclosure relates to a method comprising the steps of: contacting target-specific primers with nucleic acids in a sample, wherein the sample comprises both maternal and fetal DNA, and wherein the target-specific primers are configured to simultaneously hybridize with at least 10, 20, 100, 500, 1,000, 2,500, 5,000, 10,000, 25,000, 50,000, 80,000, 100,000, or 150,000 different target regions in the nucleic acids; amplifying multiple target nucleic acids to generate amplicons; performing next-generation sequencing on the amplicons to generate sequence data; and analyzing the sequence data using a software algorithm. In some embodiments, the method further comprises counting amplicons associated with control chromosomes and suspected chromosomes to determine the presence or absence of abnormal chromosome distributions. In some embodiments, the method further comprises counting amplicons to determine the presence of aneuploidy in fetal DNA. In some embodiments, the method further comprises counting amplicons and determining the presence or absence of deletions in nucleic acids. In some embodiments, the sample comprises both maternal and fetal DNA. In some implementations, the sample contains cell-free DNA from the subject's plasma.

[0026] In some embodiments, the method further includes determining the presence of aneuploidy or abnormal chromosome distribution by comparing the sample with a control sample containing control sequences. In some embodiments, the relative amount of the control sequences is known. In some embodiments, the relative amount of each target nucleic acid sequence is normalized by a reference genome. In some embodiments, the control or reference genome includes at least one chromosome abnormality or aneuploidy. In some embodiments, the aneuploidy is located on chromosome 13, chromosome 18, chromosome 21, chromosome X, or chromosome Y.

[0027] In some embodiments, this disclosure relates to a kit comprising two or more modified methylated target-specific primers configured to amplify a target sequence of interest in a sample. In some embodiments, this disclosure relates to a kit comprising two or more modified methylated universal primers configured to amplify a target sequence of interest in a sample. Attached Figure Description

[0028] The present disclosure can be better understood with reference to the following figures. The elements in the figures are not necessarily drawn to scale relative to each other, but rather the emphasis is on clearly illustrating the principles of the present disclosure. Furthermore, in several views, the same reference numerals denote corresponding parts.

[0029] Figure 1 A schematic diagram of target-specific primers for forward and reverse methylation is shown. Each primer consists of a target-specific sequence portion and a universal auxiliary sequence portion consisting of a methylated nucleotide C (mC), a restriction enzyme recognition site, and a cleavage site.

[0030] Figure 2 A schematic diagram of primers for forward and reverse universal auxiliary methylation used in the second method is shown. The universal auxiliary sequence primer consists of a methylated nucleotide C (mC), a restriction enzyme recognition site, and a cleavage site.

[0031] Figure 3 Four examples of modification-dependent endonucleases and their restriction sites are shown. The methylation-dependent endonuclease restriction enzymes digest modified (methylated) cytosine in double-stranded DNA at the restriction sites shown.

[0032] Figure 4 The amplicon amplified using methylated primers (Method 1 or 2) is shown, along with digestion by a methylation / modification-dependent endonuclease. Digested primers and excess / unused primers are removed by size-selective removal, and complementary sticky-end adaptors are ligated to the sticky-end amplicon.

[0033] Figure 5The amplicon is shown as amplicon linked with a barcoded universal primer amplifier prior to library preparation and sequencing.

[0034] Figure 6 A schematic diagram of the entire library preparation process is shown, in which: (1) double-stranded gDNA or RNA is amplified (converted to cDNA) by methylation of target-specific primers containing a universal auxiliary sequence (Method 1); or amplification is performed using target-specific primers containing a universal auxiliary sequence to form an amplification product, and then a portion of the amplification product is used for a next PCR using a methylated universal auxiliary primer (Method 2); (2) the amplification product from (1) is digested with a methylation-dependent endonuclease restriction enzyme, wherein the amplification product contains sticky ends at each end of the strand, and size-selective purification is performed to remove digested primer dimers and unused primers; (3) a universal adaptor containing complementary sticky ends is ligated to dsDNA, wherein the ligation adaptor contains a universal sequence portion; (4) the ligation product is amplified with a barcoded universal primer complementary to the sequence on the ligation adaptor to form a final amplification product; and (5) the final amplification product is prepared for next-generation sequencing.

[0035] Figure 7 It shows Figure 6 The library preparation workflows for the two publicly disclosed methods are described.

[0036] Figure 8 The results of bidirectional sequencing (electrophoresis) of multiplex PCR of the oncogene EGFR using methylated primers (method 1) are shown.

[0037] Figure 9 The results of bidirectional sequencing (electrophoresis) of multiplex PCR of the oncogene TP53 using methylated primers (method 1) are shown.

[0038] Figure 10 The bidirectional sequence results (electrophoresis diagram) of multiplex PCR with the oncogene KIT using methylated primers (Method 1) are shown.

[0039] Figure 11 Screenshots of Illumina sequence reads from a library generated using method 2 disclosed in this invention are shown, mapped onto human sequence reference hg19 for different target regions. Detailed Implementation

[0040] This disclosure relates to methods and compositions for amplifying and enriching specific target sequences. The following examples, applications, descriptions, and contents are exemplary and illustrative, and are in any way non-limiting and non-restrictive. This disclosure is characterized by a variety of applications, such as genotyping, detection of chromosomal abnormalities (e.g., fetal aneuploidy), analysis of gene mutations and polymorphisms (e.g., single nucleotide polymorphisms, SNPs), gene deletions, determination of paternity, analysis of genetic differences between populations, forensic analysis, measurement of disease predisposition, quantitative analysis of mRNA, and detection and identification of infectious agents (e.g., bacteria, parasites, and viruses). The methods disclosed herein can also be used for non-invasive prenatal testing, such as paternity testing or detection of fetal chromosomal abnormalities.

[0041] Next-generation sequencing has enabled many applications at extremely low cost; however, some applications, such as whole-genome sequencing and whole-transcriptome sequencing, while practical for research settings and discovery, remain impractical for disease diagnosis, treatment, and prognosis in clinical settings. Specific and uniform multiplex target sequencing offers numerous advantages in both clinical and research settings. To increase the output and power of bioassays (e.g., multiplex PCR and next-generation sequencing), the simultaneous amplification of multiple target genes using combinations of several target-specific primers allows for multiplex amplification of regions of interest. While the use of multiple primers reduces labor, cost, and time, the resulting non-specific amplification or amplification artifacts (such as primer-primer interactions (primer-dimers) and hyperamplifiers) can interfere with optimal amplification and further analysis, such as sequencing. These artifacts waste PCR reaction reagents and produce shorter fragments instead of the intended target sequence. Furthermore, these unwanted non-specific fragments tend to dominate the amplification reaction because they amplify more efficiently than the desired target sequence. These unwanted artifacts can also interfere with downstream procedures and applications involving a second PCR step, such as next-generation sequencing. These artifacts can consume a significant portion of the sequence read, resulting in non-informative results.

[0042] The current challenge of target enrichment (which selectively captures genomic regions from a DNA sample prior to sequencing) is achieving high specificity and uniformity, which would require fewer sequencing reads to generate sufficient coverage and sequence data for downstream analysis. In some applications, such as cancer or genetic diseases, deeper sequencing is needed to detect, identify, or validate somatic mutations in the genome with high specificity and uniformity.

[0043] Furthermore, minimizing primer dimers remains essential, allowing for the development of highly multiplex PCR, where multiplex amplification can simultaneously amplify large amounts of target nucleic acids in a single test reaction. Additionally, primer dimer removal allows for an increased number of primers used for multiplexing, higher primer concentrations for balanced amplification, and higher sensitivity. The ability to increase the number of target-specific targets in multiplex PCR allows for the simultaneous amplification of large numbers (thousands) of nucleic acid targets while reducing the amount of input DNA, labor, and time. This is particularly advantageous when the amount of initial input nucleic acid material is limited or when the sample is nucleic acid derived from a single cell.

[0044] To address the aforementioned needs, this paper describes a method for multiplex target enrichment that uses methylated primers to remove primer dimers, thereby increasing and amplifying the ability of multiple primers for further analysis, such as next-generation sequencing. This method can be used in two approaches. In the first approach, the target-specific primers contain methylated C(mC) in the universal helper portion, while in the second approach, two universal helper primers contain methylated C(mC) and are used in combination with target-specific primers containing portions complementary to the methylated universal helper primers.

[0045] Therefore, this disclosure can be performed in two variant methods. The first method is a method comprising the following steps: (1) contacting a methylated target-specific primer with a nucleic acid target sequence in a sample to hybridize the methylated target-specific primer with the target sequence in the sample; (2) amplifying the target nucleic acid sequence under optimal amplification conditions to form an amplification product; (3) digesting the amplification product with a methylation-dependent endonuclease restriction enzyme to form a digestion product, wherein the digestion product contains sticky ends at each end of the strand; (4) performing size-selective purification of the digestion product to remove digested primer dimers and unused primers to form a selected digestion product; (5) ligating a universal adaptor to dsDNA in the selected digestion product to form a ligation product, wherein the universal adaptor contains sticky ends complementary to the dsDNA and a universal sequence portion; (6) amplifying the ligation product using barcoded universal primers to form a final amplification product, wherein the barcoded universal primers are configured to be complementary to the sequence on the ligation adaptor; and (7) preparing the final amplification product for next-generation sequencing.

[0046] The second method includes the following steps: (1) contacting a target-specific primer with a nucleic acid target sequence in the sample to hybridize the target-specific primer with the target sequence, wherein the target-specific primer contains a universal auxiliary portion at its 5' end; (2) performing a first amplification of the target sequence with the universal auxiliary primer under optimal amplification conditions to form an amplification product; (3) performing a second amplification of a portion of the amplification product with a methylated universal auxiliary primer to form a second amplification product, wherein the methylated universal auxiliary primer contains a restriction enzyme recognition sequence; (4) digesting the second amplification product with a methylation-dependent endonuclease restriction enzyme to form a digestion product, wherein the digestion product contains sticky ends at each end of the chain; (5) performing size-selective purification of the digestion product to remove digested primer dimers and unused primers to form a selected digestion product; (6) ligating a universal adapter to the selected digestion product. DNA is used to form a ligation product, wherein the universal adapter comprises complementary sticky ends and a universal sequence portion; (6) the ligation product is amplified a third time using barcoded universal primers to form a final amplification product, wherein the barcoded universal primers comprise a sequence complementary to the sequence on the ligation adapter; and (7) the final amplification product is quantified for next-generation sequencing.

[0047] The disclosed methods may further include extracting DNA, such as FFPE or blood, plasma DNA or RNA (converted to ds cDNA), from a sample. The disclosed methods may further include purification known to those skilled in the art. The disclosed methods may comprise about 10, 20, 100, 500, 1,000, 2,500, 5,000, 10,000, 25,000, 50,000, 80,000, 100,000, or 150,000 or more target-specific primers and about 10, 20, 100, 500, 1,000, 2,500, 5,000, 10,000, 25,000, 50,000, 80,000, 100,000, or 150,000 or more target sequences.

[0048] Subjects can be mammals. In some cases, subjects are humans. Subjects can be healthy, diagnosed with a disease, or suspected of having a disease. In other cases, subjects can be non-mammals, such as bacteria, viruses, or fungi.

[0049] Target nucleic acids can be present in samples obtained from subjects. Samples may contain proteins, cells, fluids, biofluids, preservatives, blood, hair, biopsy material, and other materials containing nucleic acids. Nucleic acid samples may contain genomic DNA or RNA. Samples may also contain nucleic acid molecules obtained from FFPE or archived DNA samples. Samples may also contain mechanically or enzymatically cleaved or fragmented DNA. Samples may contain circulating cell-free DNA (cfDNA), such as material obtained from maternal subjects, or circulating tumor DNA (ctDNA) from subjects diagnosed with cancer or from subjects used for cancer screening purposes. Samples may include nucleic acid molecules obtained from blood, serum, plasma, cerebrospinal fluid, urine, tissue, saliva, biopsy, sputum, swabs, formalin-fixed paraffin-embedded material (FFPE), surgical excisions, cervical swabs, tears, tumor tissue, fine needle aspiration (FNA), circulating cell-free DNA (cfDNA) and circulating tumor DNA (ctDNA), scrapings, swabs, mucus, urine, semen, hair, laser capture microdissection, and other non-restricted clinical or laboratory-obtained samples. Samples may be epidemiological, bacterial, viral, fungal, agricultural, forensic, or pathogenic.

[0050] Multiple target-specific primers may contain a target-specific sequence and an auxiliary sequence at the 5'-, wherein the auxiliary sequence is configured to allow digestion of the methylation site using a restriction enzyme. Furthermore, methylation may be located directly on the auxiliary portion of the target-specific primer on a universal auxiliary primer. The target-specific primer may contain a methylated universal auxiliary sequence portion, which includes a restriction enzyme recognition site and a target-specific sequence portion (…). Figure 1 Amplification can utilize at least one forward target-specific primer and at least one reverse target-specific primer. The target-specific primer may contain a universal auxiliary sequence. This disclosure also contemplates a method in which a target-specific primer containing a universal auxiliary portion is used in the first amplification, and a universal auxiliary primer containing methylated nucleotides and complementary methylation of restriction enzyme recognition sites is used in the second amplification. Figure 2 Target-specific primers can be configured to hybridize with short tandem repeats (STRs) of the target sequence. Target-specific primers can contain nucleotide modifications at the 3' end, 5' end, or throughout the entire sequence. The length of the target-specific portion of a target-specific primer can be approximately 15 to 40 bases. The T-value for each target-specific primer... m It can be approximately 50°C to approximately 72°C or other temperature ranges.

[0051] Target-specific primers may target at least one target sequence, whereby one or more mutations in the target sequence indicate that the subject has a disease. The mutation may be a clinically operable mutation. The mutation may also be associated with drug resistance or companion diagnostic treatment. The mutation may include substitution, insertion, inversion, point mutation, deletion, mismatch, and translocation. In one embodiment, the mutation may include a change in copy number. In one embodiment, the mutation may include a germline or somatic mutation. In some embodiments, the mutation has an allele frequency of less than about 10%. In other embodiments, the mutation has an allele frequency of less than about 5%, 3%, 1%, 0.5%, 0.1%, or 0.01%.

[0052] Target-specific primers can target sequences associated with cancer-related diseases or one or more autoimmune, genetic, cardiovascular, developmental, metabolic, neurological, neuromuscular, neonatal diseases, or neonatal conditions. Target sequences may also be associated with organ transplantation or organ rejection.

[0053] Target-specific primers can target one or more genes associated with clinically relevant cancer genes that cover a wide range of cancers. Target-specific primers can be configured to amplify one or more clinically relevant cancer genes, including but not limited to: AIP, ALK, APC, ATM, BAP1, BARD1, BLM, BMPR1A, BRCA1, BRCA2, BRIP1, CDH1, CDK4, CDKN1B, CDKN2A, CHEK2, DICER1, EPCAM, FANCC, FH, FLCN, GALNT12, GREM1, HOXB13, MAX, MEN1, MET, MITF, MLH1, MRE11A, MSH2. MSH6, MUTYH, NBN, NF1, NF2, PALB2, PHOX2B, PMS2, POLD1, POLE, POT1, PRKAR1A, PTCH1, PTEN, RAD50, RAD51C, RAD51D, RB1, RE T, SDHA, SDHAF2, SDHB, SDHC, SDHD, SMAD4, SMARCA4, SMARCB1, SMARCE1, STK11, SUFU, TMEM127, TP53, TSC1, TSC2, VHL and XRCC2.

[0054] Target-specific primers can also target one or more genes associated with breast cancer. Target-specific primers can be configured to amplify one or more clinically relevant genes in breast cancer, including but not limited to: ATM, BARD1, BRCA1, BRCA2, BRIP1, CDH1, CHEK2, FANCC, MRE11A, MUTYH, NBN, NF1, PALB2, PTEN, RAD50, RAD51C, RAD51D, STK11, and TP53.

[0055] Target-specific primers can target one or more genes associated with ovarian cancer. Target-specific primers can be configured to amplify one or more clinically relevant genes in ovarian cancer, including but not limited to: ATM, BARD1, BRCA1, BRCA2, BRIP1, CDH1, CHEK2, DICER1, EPCAM, MLH1, MRE11A, MSH2, MSH6, MUTYH, NBN, NF1, PALB2, PMS2, PTEN, RAD50, RAD51C, RAD51D, SMARCA4, STK11, and TP53.

[0056] Target-specific primers can target one or more genes associated with colorectal cancer. Target-specific primers can be configured to amplify one or more clinically relevant genes for colorectal cancer, including but not limited to: APC, BMPR1A, CDH1, CHEK2, EPCAM, GREM1, MLH1, MSH2, MSH6, MUTYH, PMS2, POLD1, POLE, PTEN, SMAD4, STK11, and TP53.

[0057] Target-specific primers can target one or more genes associated with prostate cancer. Target-specific primers can be configured to amplify one or more clinically relevant genes for prostate cancer, including but not limited to: ATM, BRCA1, BRCA2, CHEK2, EPCAM, HOXB13, MLH1, MSH2, MSH6, NBN, PALB2, PMS2, RAD51D, and TP53.

[0058] Target-specific primers can be used to detect and identify fusion genes, such as aberrant gene fusions or transformed gene fusions in cancer (e.g., EML4-ALK or ROS1). Target-specific primers can be configured to amplify one or more fusion genes in cancer, including but not limited to: AKT3, ALK, ARHGAP26, AXL, BRAF, BRD3, BRD4, EGFR, ERG, ESR1, ETV1, ETV4, ETV5, ETV6, EWSR1, FGFR1, FGFR2, FGFR3, FGR, INSR, MAML2, MAST1, MAST2, MET, MSMB, MUSK, MYB, NOTCH1, NOTCH2, NRG1, NTRK1, NTRK2, NTRK3, NUMB1, NUTM1, PDGFRA, PDGFRB, PIK3CA, PKN1, PPARG, PRKCA, PRKCB, RAF1, RELA, RET, ROS1, RSPO2, RSPO3, TERT, TFE3, TFEB, THADA, and TMPRSS2.

[0059] Target-specific primers can be configured to selectively amplify target sequences carrying mutations associated with congenital or hereditary diseases. Mutations can be somatic or germline mutations. Mutations associated with congenital or hereditary diseases can include point mutations, insertions, deletions, inversions, substitutions, mismatches, translocations, and copy number variations. In some embodiments, at least one of the target-specific primers associated with hereditary diseases is at least 90% complementary to the target sequence.

[0060] Target-specific primers can target one or more genes associated with cardiovascular disease. These primers can be configured to amplify one or more clinically relevant genes associated with cardiovascular disease, including but not limited to ABCC9, ACTA2, ACTC1, ACTN2, AKAP9, ANK2, ANKRD1, BAG3, CACNA1C, CACNA2D1, CACNB2, CALM1, CASQ2, CAV3, CBS, COL3A1, COL5A1, COL5A2, CRYAB, CSRP3, DES, DMD, DSC2, DSG2, DSP, EMD, EYA4, FBN1, FBN2, FKTN, FLNA, FXN, GATA4, GATAD1, GLA, GPD1L, HCN4, JAG1, JPH2, JUP, KCND3, KCNE1, KCNE2, KCNE3, KCNH2, KCNJ2, KCNJ5, KCNJ8, KCNQ1, LA MA4, LAMP2, LDB3, LMNA, MED12, MYBPC3, MYH11, MYH6, MYH7, MYL2, MYL3, MYLK, MYOZ2, MYPN, NEXN , NKX2-5, NOTCH1, PKP2, PLN, PLOD1, PRKAG2, PRKG1, PTPN11, RAF1, RBM20, RYR2, SCN1B, SCN2B, SC N3B, SCN4B, SCN5A, SKI, SLC2A10, SMAD3, SMAD4, SNTA1, TAZ, TBX1, TBX20, TBX5, TCAP, TGFB2, TGFB3, TGFBR1, TGFBR2, TMEM43, TMPO, TNNC1, TNNI3, TNNT2, TPM1, TRDN, TRPM4, TTN, TTR, TXNRD, and CL.

[0061] This disclosure also relates to a method for target enrichment via multiplex PCR, comprising the steps of: contacting a target sequence with multiple target-specific primers in the presence of PCR reagents such as DNA polymerase, dNTPs, and reaction buffer; and providing optimal temperature and time conditions for denaturation, annealing, and extension, hybridizing primers to complement the target sequence and extend such target sequences. As determined by those skilled in the art, amplification, purification, and removal can be adjusted or removed as needed to optimize multiplex target amplification in downstream processes.

[0062] The methods disclosed herein are characterized by their broad applicability in clinical and research settings and can be used for mutation detection and analysis, single nucleotide polymorphism (SNP) detection, microbial and viral detection, deletion and insertion, genotyping, copy number variation (CNV), epigenetic and methylation analysis, gene expression, transcriptome analysis, and low-frequency allele mutations. The disclosed methods can also be used for disease detection, diagnosis, prognosis, and treatment. Furthermore, the disclosed methods can detect germline or somatic mutations in samples.

[0063] The disclosed method can be used with PCR and DNA polymerase. A wide selection of DNA polymerases is available, characterized by different properties such as thermal stability, high fidelity, sustained synthesis capability, and hot-start. Amplification conditions, such as cycle number, annealing temperature, annealing duration, extension temperature, and extension duration, can be adjusted to optimal amplification conditions, which can be based on the instructions provided with the selected commercial DNA polymerase. The concentration of DNA polymerase used for multiplex PCR can be higher than that used for singlex PCR.

[0064] The method disclosed herein uses a double-stranded linker configured to attach to a double-stranded nucleic acid fragment. This linker includes a universal sequence portion that is not complementary to the target sequence and sticky ends that are complementary to the sticky ends on the digested amplicon. The barcode sequence on the barcoded universal primer allows for tagging of the nucleic acid fragment for each subject and can distinguish the identity of each sample. Therefore, barcoding increases throughput by enabling sample pooling. The disclosed method can also use the linker for the purpose of universal amplification of large numbers of nucleic acid target sequences.

[0065] The disclosed method uses multiplex polymerase chain reaction (PCR) to amplify target sequences, wherein more than one target sequence is amplified in a single test reaction. The amount of nucleic acid in the sample required for multiplex amplification can be about 1 ng. Alternatively, the amount of nucleic acid material can be about 5 ng, 10 ng, 50 ng, 100 ng, or 200 ng or more. The multiplex PCR can be performed on a thermal cycler, and each cycle of multiplex PCR includes denaturation, annealing, and extension steps. Each cycle of multiplex PCR includes at least one denaturation step, one annealing step, and one extension step for extending the nucleic acid. The disclosed method can include 5 to 20 PCR cycles per round of amplification, but other numbers of cycles are also possible. For example, 1 to 10 cycles, 1 to 15 cycles, 1 to 20 cycles, 1 to 25 cycles, or 1 to 30 cycles or more can be performed. Each cycle or group of cycles can have different durations and temperatures. For example, the annealing step can have incremental increases and decreases in temperature and duration, or the extension step can have incremental increases and decreases in temperature and duration. The duration can decrease or increase in increments of approximately 5 seconds, 10 seconds, 30 seconds, 1 minute, 2 minutes, 4 minutes, 8 minutes, or more. The temperature can decrease or increase in increments of approximately 0.5 degrees Celsius, 1 degree Celsius, 2 degrees Celsius, 4 degrees Celsius, 8 degrees Celsius, 10 degrees Celsius, or more.

[0066] Amplicon size selection can be used to sequence amplicon products within a certain length range. For example, amplicon lengths of 100 to 250 base pairs, 150 to 300 base pairs, 120 to 350 base pairs, or 200 to 500 base pairs or longer can be sequenced.

[0067] Typically, a small number of primers in a primer set or primer library can cause amplification artifacts, such as primer dimers, in multiplex amplification reactions. However, by employing primer selection algorithms that can compute undesirable primer-primer interactions, target-specific primer selection can be performed in an efficient manner that minimizes primer-primer interactions to a negligible amount, thus allowing for the simultaneous amplification of large numbers of target sequences in a single test reaction. Furthermore, digestion of methylation sites on the common portion of target-specific primers allows for the removal of primer dimers via methylation-dependent restriction enzyme digestion and size-selective purification (e.g., SPRISelect beads). Additionally, methylated target-specific primers allow for an increased number of primers for multiplexing, higher concentrations of target-specific primers for balanced amplification, and higher sensitivity, regardless of primer dimers. The ability to increase the number of target-specific primers in multiplex PCR allows for the simultaneous amplification of large numbers (thousands) of nucleic acid target sequences while reducing the amount of input DNA, labor, and time. This is particularly advantageous when the amount of initial input nucleic acid material is limited, or when the sample contains nucleic acids from a single cell.

[0068] Primer dimers can be reduced or minimized by adjusting different parameters of the disclosed method, such as the duration of the annealing step, temperature increment, and / or the number of PCR cycles. In addition to reducing or minimizing primer dimers, primer concentrations can be decreased, and annealing temperatures and durations can be increased to allow for specific amplification (the primers have more time intervals to hybridize with the target nucleic acid). Target-specific primer concentrations can be approximately 500 nM, 250 nM, 100 nM, 80 nM, 70 nM, 50 nM, 30 nM, 10 nM, 2 nM, 1 nM, or less than 1 nM. Alternatively, the concentration of each target-specific primer can be 1 μM to 1 nM, 1 nM to 80 nM, 1 nM to 100 nM, 10 nM to 50 nM, or 1 nM to 60 nM. Annealing temperatures can be approximately 1 minute, 3 minutes, 5 minutes, 8 minutes, 10 minutes, or longer. Amplifications with increased annealing time can be performed using 1, 2, 3, 5, 8, 10 or more cycles, followed by the standard annealing duration.

[0069] The disclosed methods and kits may contain at least 10, 20, 100, 500, 1,000, 2,500, 5,000, 10,000, 25,000, 50,000, 80,000, 100,000, or 150,000 or more target-specific primers, each of which is directed to hybridize with a specific target sequence. More than one set of target-specific primers may be present; as an example, there may be two sets of target-specific primers for two test reactions, three sets for three test reactions, or five sets for five test reactions or more. For practical reasons, such as limitations in target-specific primer design or selection, samples may also be divided into multiple parallel multiplex tests with multiple sets of target-specific primers.

[0070] The GC content of target-specific primers can be between 40% and 70%, between 30% and 60%, between 50% and 80%, or between 30% and 80%. Alternatively, the GC content of target-specific primers can range from less than 20%, 15%, 10%, or 5%. The melting temperature (T) of the target-specific primers... mThe melting temperature range of the target-specific primer can be between 55°C and 65°C, between 40°C and 70°C, between 50°C and 68°C, or other such ranges as determined by those skilled in the art. The melting temperature range of the target-specific primer can vary. In some cases, this range can be less than 20°C, 15°C, 10°C, 5°C, 2°C, or 1°C. The length of the target-specific primer can also vary. In some cases, the length can be 20 to 90 bases, 40 to 70 bases, 20 to 40 bases, or 25 to 50 bases. The length range of the target-specific primer can also vary. For example, it can be 60, 50, 40, 30, or 20 bases. In some cases, the 5' region of the target-specific primer is an auxiliary or universal primer binding site or tag and is not complementary to or specific to any target sequence. In some cases, the target sequence is 50 to 500 bases, 90 to 350 bases, or 200 to 450 bases in length, but other lengths are also possible.

[0071] This disclosure also relates to kits containing two or more target-specific primers. In some cases, the kit contains multiple methylated target-specific primers. In other cases, the kit contains a combination of methylated universal primers and target-specific primers. Target-specific primers are designed and selected based on criteria described as having no or minimal primer-primer interactions or nonspecific priming. The kit can be formulated for the detection, diagnosis, prognosis, and treatment of diseases such as cancer or congenital or hereditary disorders. The kit can also be configured for the ploidy status of a pregnant fetus, for example by selecting target-specific primers that target sequences on chromosomes associated with trisomy in the fetus, such as chromosomes 13, 18, 21, X, and Y, other chromosomes, or some combination thereof. The kit may contain about 10, 20, 100, 500, 1,000, 2,500, 5,000, 10,000, 25,000, 50,000, 80,000, 100,000, or 150,000 or more target-specific primers.

[0072] The methods and kits disclosed herein may include a variety of target-specific primers that have little or no self-complementary structure and do not form secondary structures such as hairpins or loops. The methods and kits disclosed herein may further include a variety of target-specific primers that have minimal cross-hybridization with nonspecific sequences present in the sample.

[0073] The target-specific primers disclosed herein can be used to efficiently amplify short nucleic acid fragments, such as nucleic acids derived from FFPE samples, cell-free DNA (cfDNA), cell-free tumor DNA (ctDNA), and cell-free fetal DNA (cffDNA). These short nucleic acid fragments can be less than approximately 40, 50, 60, 70, 80, 90, 100, or 120 bases. The methods disclosed herein can also be used to detect and quantify a small percentage of mutations, such as the T790M mutation associated with drug resistance in lung cancer.

[0074] The methods disclosed herein produce amplified products that can be sequenced using next-generation sequencing platforms. Next-generation sequencing refers to non-Sanger massively parallel DNA sequencing technologies that can sequence tens of thousands, millions, or even billions of DNA strands in parallel. Examples of current state of existing next-generation sequencing technologies and platforms include the Illumina platform (reversible dye terminator sequencing), 454 pyrosequencing, Ion Torrent sequencing, PacBio SMRT sequencing, Qiagen GeneReader sequencing technology, and Oxford Nanopore sequencing. This disclosure is not limited to these examples of next-generation sequencing technologies. On the other hand, the methods disclosed herein can be used in multiplexes when amplifying more than two targets. The methods disclosed herein are not limited to any number of multiplexing operations.

[0075] Example

[0076] Example 1

[0077] Multiplex amplification using methylated primers to identify variants (Method 1)

[0078] Materials and methods

[0079] Human DNA was extracted using the Qiagen DNA Extraction Kit according to the manufacturer's instructions, and the amount of DNA was measured using NanoDrop (ThermoFisher, USA) and Qubit 3 (ThermoFisher, USA).

[0080] Three oncogenes were selected for this experiment. Methylation-modified forward and reverse primers were designed to detect variants of the EGFR, KIT, and TP53 oncogenes. Each primer consisted of a target-specific region and a methylation-modified auxiliary universal region.

[0081] Multiplex PCR was performed in a 20 μL reaction volume using six methylated primers in the presence of genomic DNA, DNA polymerase, dNTPs, and PCR buffer. The PCR conditions consisted of 15 cycles: an initiation at 98°C for 30 seconds, followed by 10 seconds at 98°C, 4 minutes at 63°C, 20 seconds at 72°C, and a final extension at 72°C for 2 minutes.

[0082] According to the manufacturer's instructions, the amplification products were digested with exonuclease I (NEB, USA) to remove redundant primers and purified with SPRIselect (Beckman Coulter, USA) beads to select large fragments.

[0083] According to the manufacturer's instructions, digest the PCR products in a separate reaction using MspJI or LpnPI (NEB, USA). Purify the digested products using SPRIselect (Beckman Coulter, USA) beads.

[0084] Following the manufacturer's instructions, the digested product was ligated to an adapter containing complementary sticky ends using Instant Sticky-EndLigase Master Mix (NEB, USA). The procedure was performed on an Applied Biosystems Veriti thermal cycler (ThermoFisher, USA).

[0085] The ligated DNA products were purified using SPRISelect (Beckman Coulter, USA) beads to remove excess ligation adaptors.

[0086] PCR was performed using barcoded universal primers to hybridize the ligation products to the universal initiation sites in a 20 μL reaction volume in the presence of DNA polymerase, dNTPs, and PCR buffer. The PCR conditions consisted of 21 cycles: an initiation at 98°C for 30 seconds, followed by cycles of 98°C for 10 seconds, 68°C for 30 seconds, and 72°C for 20 seconds, and a final extension at 72°C for 2 minutes.

[0087] The amplification products were digested with exonuclease I (NEB, USA) to remove redundant primers, purified using SPRIselect beads (Beckman Coulter, USA), and then measured on a Qubit 3 (ThermoFisher, USA). The final products were bidirectionally sequenced using Sanger sequencing.

[0088] Example 2

[0089] Cancer gene ensemble used to identify mutations and fusion genes in lung cancer from FFPE samples.

[0090] Materials and methods

[0091] Human genomic DNA was used in this experiment to analyze potential mutations that could affect treatment options.

[0092] DNA was extracted using the Qiagen DNA Extraction Kit according to the manufacturer's instructions, and the amount of DNA was measured using NanoDrop (ThermoFisher, USA) and Qubit 3 (ThermoFisher, USA).

[0093] Cancer Genes and Primer Design: Based on a literature search, 15 cancer-related genes were selected: AKT1, ALK, BRAF, CTNNB1, EGFR, ERBB2, HRAS, KIT, KRAS, MAP2K1, MET, NRAS, PDGFRA, PIK3CA, and TP53. To detect hotspot mutations in these genes, 61 pairs of forward and reverse primers were designed for multiplex amplification of the target nucleic acids.

[0094] Multiplex PCR was performed in a 20 μL reaction volume using 122 target-specific primers containing universal helper sequences in the presence of genomic DNA, DNA polymerase, dNTPs, and PCR buffer. The PCR conditions consisted of 10 cycles: an initiation at 98°C for 30 seconds, followed by 10 seconds at 98°C, 4 minutes at 63°C, 20 seconds at 72°C, and a final extension at 72°C for 2 minutes.

[0095] According to the manufacturer's instructions, the amplification product from the first amplification was treated with exonuclease I (NEB, USA) to remove redundant primers. The resulting first amplicon was then purified using SPRIselect beads (Beckman Coulter, USA).

[0096] In the presence of DNA polymerase, dNTPs, and PCR buffer, amplification was performed in a 20 μL reaction volume using a portion of the purified product from the first amplification and methylated universal auxiliary primers. The PCR conditions consisted of 15 cycles of initiation at 98°C for 30 seconds, followed by cycles of 98°C for 10 seconds, 69°C for 30 seconds, and 72°C for 30 seconds, and a final extension at 72°C for 2 minutes.

[0097] According to the manufacturer's instructions, the amplification products were digested with exonuclease I (NEB, USA) to remove redundant primers and purified with SPRIselect beads (Beckman Coulter, USA) to select large fragments.

[0098] According to the manufacturer's instructions, the amplified product is digested using MspJI or LpnPI (NEB, USA) and then purified by SPRIselect beads.

[0099] Following the manufacturer's instructions, the digested product was ligated to an adapter containing complementary sticky ends using Instant Sticky-EndLigase Master Mix (NEB, USA). The procedure was performed on an Applied Biosystems Veriti thermal cycler (ThermoFisher, USA).

[0100] The ligated DNA product was then purified using SPRIselect beads to remove excess ligands.

[0101] PCR was performed on the ligation product using barcoded universal primers. The primers hybridized to the universal initiation site of the ligation product in a 20 μL reaction volume in the presence of DNA polymerase, dNTPs, and PCR buffer. The PCR conditions consisted of 21 cycles: 98°C for 30 seconds, 98°C for 10 seconds, 68°C for 30 seconds, 72°C for 20 seconds, and a final extension at 72°C for 2 minutes.

[0102] The amplification products were digested with exonuclease I (NEB, USA) to remove redundant primers, purified with SPRIselect beads, and then measured on a Qubit 3 (ThermoFisher, USA).

[0103] Library sequencing was performed using the MiniSeq Mid Output Kit on a MiniSeq sequencing system (Illumina, CA, USA).

[0104] Analyze the mutations and variations in the sequence data generated from the above experiments.

[0105] The methods and various implementations thereof described herein are exemplary. Various other implementations of the methods described herein are possible.

Claims

1. A method of enriching a nucleic acid target sequence in a sample, comprising the steps of: hybridizing two or more target-specific primers to the target sequence in the sample in a test reaction, wherein the target-specific primers comprise a methylated universal helper portion having a methylation-dependent endonuclease restriction enzyme recognition site and a target-specific portion configured to target the nucleic acid target sequence in the sample; allowing the test reaction to proceed to amplification to produce an amplification product comprising amplicons; digesting the amplification product with a methylation-dependent endonuclease restriction enzyme to form a digestion product, wherein the digestion product comprises amplicons comprising sticky ends on each end of the strand; size selecting and purifying the digestion product to remove digested primer dimers and unused primers to produce digested amplicons comprising dsDNA; ligating universal adaptors to the dsDNA from the digested amplicons to form a ligation product, wherein the ligation universal adaptors comprise a universal sequence portion and a sticky end; and amplifying the ligation product with barcoded universal primers complementary to the sequence on the ligation universal adaptors to form a final amplification product comprising final amplicons.

2. The method of claim 1, further comprising at least one additional target-specific primer set in at least one additional test reaction.

3. The method of claim 1, wherein the sample comprises genomic DNA.

4. The method of claim 1, wherein the sample comprises RNA, and the method further comprises the step of subjecting the RNA to a reverse transcription reaction to produce double-stranded cDNA prior to first amplification.

5. The method of claim 1, wherein the method further comprises the step of subjecting the final amplicons to next-generation sequencing to produce sequence data.

6. The method of claim 5, further comprising the step of measuring allele counts at polymorphic sites in the sequence data.

7. The method of claim 1, wherein the ligation universal adaptors further comprise a barcode sequence.

8. The method of claim 1, wherein the sample comprises one or more of the following nucleic acids: a mixture of maternal cfDNA and cffDNA obtained from a pregnant subject, circulating cfDNA, and circulating ctDNA.

9. The method of claim 1, wherein the target-specific primers comprise at least one pair of forward target-specific primers and reverse target-specific primers.

10. The method of claim 1, wherein the target sequence comprises one or more mutations associated with cancer in a pregnant fetus.

11. The method of claim 1, wherein the target sequence comprises one or more mutations associated with infection in a pregnant fetus.

12. The method of claim 1, wherein the target sequence comprises one or more mutations associated with pharmacogenetic drug therapy in a pregnant fetus.

13. The method of claim 1, wherein the target sequence comprises one or more mutations associated with drug antibiotic resistance in a pregnant fetus.

14. The method of claim 1, wherein the target sequence comprises one or more mutations associated with an aneuploidy or trisomy in a pregnant fetus.

15. The method of claim 1, wherein the target sequence comprises one or more mutations associated with a disease in a pregnant fetus.

16. The method of claim 1, wherein the target sequence comprises one or more mutations associated with a condition in a pregnant fetus.

17. The method of claim 1, wherein the target sequence comprises one or more mutations associated with a companion diagnosis in a pregnant fetus.

18. The method of claim 1, wherein the target sequence comprises one or more mutations associated with a drug resistance in a pregnant fetus.

19. A method of enriching a nucleic acid target sequence in a sample, comprising the steps of: hybridizing two or more target-specific primers to a nucleic acid target sequence in a test reaction, wherein the target-specific primers comprise a complementary universal helper portion at the 5’ end and a target-specific portion configured to target the nucleic acid target sequence in the sample; performing a first amplification of the test reaction with a universal helper primer to form an amplification product; performing a second amplification of a portion of the amplification product with a methylated universal helper primer to form a second amplification product, wherein the methylated universal helper primer comprises a restriction enzyme recognition sequence; digesting the second amplification product with a methylation-dependent endonuclease restriction enzyme to form a digestion product comprising amplicons comprising sticky ends on each end of the strand; performing a size selection purification of the digestion product to remove digested primer dimers and unused primers to form a digested amplicon comprising dsDNA; ligating a universal adapter to the dsDNA from the digested amplicon to form a ligation product, wherein the ligation universal adapter comprises a complementary sticky end and a universal sequence portion; and performing a third amplification of the ligation product with a barcoded universal primer complementary to the sequence on the ligation universal adapter to form a final amplification product comprising a final amplicon.

20. The method of claim 19, further comprising at least one additional target-specific primer set in at least one additional test reaction.

21. The method of claim 19, wherein the sample comprises genomic DNA.

22. The method of claim 19, wherein the sample comprises RNA, and wherein the method further comprises performing a reverse transcription reaction on the RNA to produce double stranded cDNA prior to the first amplification.

23. The method of claim 19, wherein the method further comprises the step of performing next generation sequencing on the final amplicon to produce sequence data.

24. The method of claim 23, further comprising the step of measuring allele counts at polymorphic sites in the sequence data.

25. The method of claim 19, wherein the ligation universal adapter further comprises a barcode sequence.

26. The method of claim 19, wherein the sample comprises one or more of the following nucleic acids: a mixture of maternal cfDNA and cffDNA obtained from a pregnant subject, circulating cfDNA, and circulating ctDNA.

27. The method of claim 19, wherein the target-specific primers comprise at least one pair of forward target-specific primers and reverse target-specific primers.

28. The method of claim 19, wherein the target sequences comprise one or more mutations associated with cancer in a pregnant fetus.

29. The method of claim 19, wherein the target sequences comprise one or more mutations associated with infection in a pregnant fetus.

30. The method of claim 19, wherein the target sequences comprise one or more mutations associated with pharmacogenetic drug therapy in a pregnant fetus.

31. The method of claim 19, wherein the target sequences comprise one or more mutations associated with drug antibiotic resistance in a pregnant fetus.

32. The method of claim 19, wherein the target sequences comprise one or more mutations associated with an aneuploidy or trisomy in a pregnant fetus.

33. The method of claim 19, wherein the target sequences comprise one or more mutations associated with a disease in a pregnant fetus.

34. The method of claim 19, wherein the target sequences comprise one or more mutations associated with a condition in a pregnant fetus.

35. The method of claim 19, wherein the target sequences comprise one or more mutations associated with a companion diagnosis in a pregnant fetus.

36. The method of claim 19, wherein the target sequences comprise one or more mutations associated with resistance in a pregnant fetus.

37. A kit comprising: at least one pair of forward and reverse target-specific primers; a methylation-dependent endonuclease restriction enzyme; a ligation adaptor comprising a universal sequence portion and a sticky end portion; and a barcoded universal adaptor, wherein the target-specific primers comprise a methylated universal helper portion and a target-specific portion, the methylated universal helper portion comprising a methylation-dependent endonuclease restriction enzyme recognition site.

38. The kit of claim 37, wherein the target-specific primers comprise a universal helper portion and a target-specific portion, the universal helper portion comprising a methylation-dependent endonuclease restriction enzyme recognition site, and further comprising a complementary methylated universal helper primer, the complementary methylated universal helper primer comprising a restriction enzyme recognition sequence. ​

Citation Information

Patent Citations

  • Method for identification and quantification of nucleic acid expression, splice variant, translocation, copy number, or methylation changes

    CN107002145A

  • Methods for targeted nucleic acid sequence enrichment with applications to error corrected nucleic acid sequencing

    CN110520542A