Non-unique barcodes in genotyping assays

A method using non-unique barcodes with endogenous barcodes for ctDNA sequencing improves sensitivity and specificity, addressing the limitations of current genotyping technologies in detecting rare mutations in low DNA samples.

JP7781105B2Active Publication Date: 2025-12-05PERSONAL GENOME DIAGNOSTICS INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2023090732
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2016-11-15
Filing Date
2023-06-01
Publication Date
2025-12-05
Estimated Expiration
2037-11-14

AI Technical Summary

Technical Problem

Current genotyping technologies for analyzing circulating tumor DNA (ctDNA) face challenges due to low DNA percentages in blood samples, high error rates, and limited sensitivity and specificity, making it difficult to detect rare mutations.

Method used

A method using non-unique barcodes combined with endogenous barcodes for high-coverage sequencing of nucleic acids, allowing for the identification of tumor-specific somatic mutations and translocations with improved sensitivity and specificity.

Benefits of technology

The method achieves high sensitivity and specificity in detecting low-frequency mutations, particularly in small amounts of ctDNA, reducing sequencing errors and enhancing the detection of rare mutations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007781105000007
    Figure 0007781105000007
  • Figure 0007781105000008
    Figure 0007781105000008
  • Figure 0007781105000009
    Figure 0007781105000009
Patent Text Reader

Abstract

To provide ctDNA assays that interrogate many regions from a single sample with high precision and accuracy, while evaluating multiple forms of cancer-related genomic alterations including sequence mutations and structural alterations.SOLUTION: The disclosure provides simplified yet robust methods that achieve high sensitivity and specificity by analyzing cancer genes using a limited pool of non-unique barcodes in combination with endogenous barcodes. Samples are captured and sequenced using high-coverage next-generation sequencing to allow tumor-specific somatic mutations, amplifications and translocations to be identified.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of and priority to U.S. Provisional Patent Application No. 62 / 422,355, filed November 15, 2016, the entire contents of which are incorporated herein by reference.

[0002] FIELD OF THE INVENTION The present invention relates generally to barcoding methods for analyzing nucleic acids for tumor-specific biomarkers. [Background technology]

[0003] background Cancer kills more than 500,000 people each year in the United States alone. The success of current treatments depends on the type of cancer and the stage at which it is discovered. Many treatments involve expensive and painful surgery and chemotherapy, and often fail.

[0004] Early and accurate detection of mutations is essential for effective cancer treatment. One promising area in personalized cancer treatment is the analysis of circulating tumor DNA (ctDNA). ctDNA is released into the blood from tumor tissue, harbors tumor-specific genetic alterations, and can be analyzed by noninvasive liquid biopsy approaches to identify genetic alterations in cancer patients. Liquid biopsies offer considerable advantages because they eliminate the need for invasive procedures, allow early measurement of treatment response, and may enable detection of alterations in multiple metastatic lesions over the course of treatment.

[0005] However, examining ctDNA in blood has been problematic due to current limitations in genotyping technology. The percentage of ctDNA obtained from blood samples is often very low (<1.0%) and can be difficult to detect. Most methods for evaluating ctDNA examine single hotspot mutations or very subtle genetic changes. Conventional genotyping in cell-free DNA has an error rate of approximately 1%, making it difficult or impossible to identify mutations with a prevalence of <1% in samples using conventional molecular barcoding techniques. Current methods do not provide sufficient analytical sensitivity and specificity. Summary of the Invention

[0006] overview The present disclosure relates to a ctDNA assay that interrogates many genomic regions with high accuracy and precision from a single sample and evaluates multiple forms of cancer-related genomic alterations, including sequence mutations and structural changes. The present disclosure provides a simple yet robust method that achieves high sensitivity and specificity by analyzing cancer genes using a limited pool of non-unique barcodes combined with endogenous barcodes. Samples can be collected and sequenced using high-coverage next-generation sequencing to identify tumor-specific somatic mutations and translocations. Analysis for sequence mutations or rearrangements can be performed together or separately depending on the specific alteration of interest. The disclosed method provides improved sensitivity and specificity of sequencing for diagnostic, forensic, phylogenetic, and clinical purposes.

[0007] The disclosed method is particularly suitable for small amounts of sample DNA, such as in certain liquid biopsies. Liquid biopsies assess blood DNA for circulating tumor DNA. Circulating tumor DNA (ctDNA) can enter the bloodstream through tumor cell apoptosis, and when detected, it enables diagnosis, genotyping, and disease monitoring without the need for traditional invasive biopsy procedures. However, ctDNA levels are generally quite low, especially in early-stage tumors, making it difficult to rely on ctDNA for detection and analysis. The present invention addresses the problem of methods for identifying rare mutations in samples containing limited amounts of DNA template. The method of the present invention mitigates the impact of the error rate inherent in massively parallel sequencing devices. Without the disclosed method, the error rate inherent in these devices is generally too high to identify rare mutations in most samples.

[0008] The method can include extracting and isolating cell-free DNA from a plasma sample and assigning an exogenous barcode to each fragment to create a DNA library. The exogenous barcodes are from a limited pool of non-unique barcodes, for example, eight types of barcodes. Barcoded fragments are distinguished based on the combination of their exogenous barcodes and endogenous barcodes obtained from the genomic location of the fragment ends of each cell-free DNA molecule. The DNA library is sequenced in duplicate to align sequences with matching barcodes. The aligned sequences are aligned to a human genome reference, and variants present in the aligned sequences are identified as genuine mutations.

[0009] The present invention recognizes that a completely unique barcode sequence is not necessary. Instead, a predetermined set of non-unique sequences combined with endogenous barcodes can provide the same level of sensitivity and specificity as a unique barcode can with a biologically relevant amount of DNA. A limited pool of barcodes is more robust and easier to generate and use than a traditional unique set. This method can be used, for example, to assay a panel of well-characterized cancer genes. This method can also be used to evaluate subclonal mutations in tumor tissues.

[0010] An aspect of the present invention relates to a method for analyzing nucleic acid.Nucleic acid can be cell-free DNA, circulating tumor DNA or RNA.This method includes the following steps: obtain a sample containing nucleic acid fragments; introduce a set of non-unique barcodes into fragments to create a genome library; identify the end parts of fragments; sequence the fragments to create sequence reads; and align the sequence reads to identify mutations.

[0011] The obtaining step may include obtaining a plasma sample, extracting nucleic acid, and fragmenting the nucleic acid. The introducing step of the set of non-unique barcodes may include end repair, A-tailing, and adapter ligation. In some embodiments, the set of non-unique barcodes consists of eight sets of non-unique barcodes. The barcodes may include sequencing adapters.

[0012] The step of identifying the end portion can include hybrid capture or whole genome sequencing. The end portion of the DNA fragment can include an endogenous barcode. Hybrid capture can be performed to identify endogenous barcodes, for example, ABL1, AKT1, ALK, APC, AR, ATM, BCR, BRAF, CDH1, CDK4, CDK6, CDKN2A, CSF1R, CTNNB1, DNMT3A, EGFR, ERBB2, ERBB4, ESR1, EZH2, FBXW7, FGFR1, FGFR2, FGFR3, FLT3, GNA11, GNAQ, GNAS, HNF1A, HRAS, IDH1, IDH2, JAK2, JA The panel may include a well-characterized cancer gene panel including K3, KDR, KIT, KRAS, MAP2K1, MET, MLH1, MPL, MYC, NPM1, NRAS, NTRK1, PDGFRA, PDGFRB, PIK3CA, PIK3R1, PTEN, PTPN11, RARA, RB1, RET, ROS1, SMAD4, SMARCB1, SMO, SRC, STK11, TERT, TP53, and / or VHL.

[0013] The sequencing step can include single-end or paired-end sequencing. The sequencing step can include overlapping sequencing and determining a consensus sequence using overlapping sequence reads. The overlapping sequencing can be performed at a depth of 2x, 10x, 50x, or 100x, etc. The aligning step can include determining whether the loci of the barcoded fragments are identical across a predetermined percentage of overlapping sequence reads, such as 50%, 60%, 70%, 80%, 90%, or 99%.

[0014] In a related aspect, the invention includes a method for barcoding molecules, the method including obtaining a sample containing nucleic acid fragments, providing a plurality of sets of non-unique barcodes, and tagging the nucleic acid fragments with the barcodes to create a genomic library, wherein each nucleic acid fragment is tagged with the same barcode as another distinct nucleic acid fragment in the genomic library.

[0015] In an embodiment, the plurality of sets is limited to 20 or fewer unique barcodes. In another embodiment, the plurality of sets is limited to 10 or fewer unique barcodes.

[0016] The method may further include one or more of the following steps: identifying end portions of the fragments; redundantly sequencing the genomic library to generate multiple overlapping sequence reads for each nucleic acid fragment; aligning the overlapping sequence reads of similarly tagged nucleic acid fragments; and aligning the aligned sequence reads to a reference to determine a consensus sequence. [The present invention 1001] obtaining a sample containing nucleic acid fragments; introducing a set of non-unique barcodes into the fragments to create a genomic library; sequencing the fragments to generate sequence reads; aligning the sequence reads; identifying the genomic locations of the fragment ends; and Identifying mutations present in the plurality of molecules as determined by a combination of the non-unique barcodes and the genomic locations of the fragment ends. A method for analyzing nucleic acids, comprising: [The present invention 1002] 1001. The method of claim 1001, wherein said obtaining step comprises obtaining a plasma sample and extracting nucleic acid. [The present invention 1003] 1001. The method of claim 1001, wherein the step of introducing a set of non-unique barcodes comprises end repair, A-tailing, and adapter ligation. [The present invention 1004] 1001. The method of claim 1001, wherein said set of non-unique barcodes consists of eight sets of non-unique barcodes. [The present invention 1005] 1001. The method of claim 1001, wherein the step of identifying the genomic locations of the fragment ends comprises hybrid capture or whole genome sequencing. [The present invention 1006] 1001. The method of claim 1001, wherein the genomic locations of the fragment ends comprise endogenous barcodes. [The present invention 1007] The method of claim 1005, wherein the hybrid capture comprises a well-characterized cancer gene panel. [The present invention 1008] The oncogenes are ABL1, AKT1, ALK, APC, AR, ATM, BCR, BRAF, CDH1, CDK4, CDK6, CDKN2A, CSF1R, CTNNB1, DNMT3A, EGFR, ERBB2, ERBB4, ESR1, EZH2, FBXW7, FGFR1, FGFR2, FGFR3, FLT3, GNA11, GNAQ, GNAS, HNF1A, HRAS, IDH1, IDH2, J AK2, JAK3, KDR, KIT, KRAS, MAP2K1, MET, MLH1, MPL, MYC, NPM1, NRAS, NTRK1, PDGFRA, PDGFRB, PIK3CA, PIK3R1, PTEN, PTPN11, RARA, RB1, RET, ROS1, SMAD4, SMARCB1, SMO, SRC, STK11, TERT, TP53, and VHL. [The present invention 1009] 1001. The method of claim 1001, wherein the sequencing step comprises single-end or paired-end sequencing. [The present invention 1010] 1001. The method of claim 1001, wherein the sequencing step comprises overlapping sequencing. [The present invention 1011] The method of claim 1010, further comprising determining a consensus sequence using the overlapping sequence reads. [The present invention 1012] 10. The method of claim 10, wherein overlapping sequencing is performed at a depth of 10 times. [The present invention 1013] The method of the present invention 1001, wherein the mutations detected in a DNA molecule are based on using non-unique barcodes that are identical across a predetermined percentage of overlapping sequence reads for the DNA molecule, and genomic locations of fragment ends that are identical across a predetermined percentage of overlapping sequence reads for the DNA molecule. [The present invention 1014] 1013. The method of claim 10, wherein said predetermined percentage is 90%. [The present invention 1015] 1001. The method of claim 1001, wherein the nucleic acid comprises cell-free DNA, circulating tumor DNA, tumor-derived DNA, or RNA. [The present invention 1016] 1001. The method of claim 1001, wherein the barcode comprises a sequencing adapter. [The present invention 1017] obtaining a sample containing nucleic acid fragments; providing a plurality of sets of non-unique barcodes; and tagging the nucleic acid fragments with barcodes to generate a genomic library; Including, A method for barcoding molecules in which each nucleic acid fragment is tagged with the same barcode as another distinct nucleic acid fragment in a genomic library. [The present invention 1018] The method of claim 1017, wherein said plurality of sets consists of 20 or fewer unique barcodes. [The present invention 1019] The method of claim 1017, wherein said plurality of sets consists of 10 or fewer unique barcodes. [The present invention 1020] The method of claim 1017, further comprising the step of identifying the genomic locations of the fragment ends. [The present invention 1021] 1017. The method of claim 1017, further comprising overlapping sequencing the genomic library to generate multiple overlapping sequence reads for each nucleic acid fragment. [The present invention 1022] The method of claim 1021, further comprising a step of adjusting overlapping sequence reads of similarly tagged nucleic acid fragments. [The present invention 1023] 1023. The method of claim 1022, further comprising aligning said adjusted sequence reads to a reference to determine a consensus sequence. [Brief explanation of the drawings]

[0017] [Figure 1] A method for genotyping using non-unique barcodes in combination with endogenous barcodes is shown. [Figure 2] 1 illustrates a method for barcoding according to the present disclosure. [Figure 3] 1 shows a well-characterized cancer gene panel for use with the present invention. [Figure 4] 1 shows a well-characterized cancer gene panel for use with the present invention. [Figure 5] 1 shows a flow chart of the genotyping method. [Figure 6] The results of observed sequence variations and predicted mutant allele frequencies for all pan-cancer cell lines are shown. [Figure 7] The observed internal control breast cancer cell line and expected mutant allele frequency results are shown. DETAILED DESCRIPTION OF THE INVENTION

[0018] Detailed Description High-throughput sequencing of circulating tumor DNA (ctDNA) promises to individualize cancer diagnosis and treatment and eliminate the need for many invasive biopsy procedures. However, the low amount of cell-free DNA (cfDNA) in blood and the limitations of sequencing technology present challenges. The occurrence of sequencing artifacts limits the sensitivity of assays involving liquid biopsies of ctDNA. For example, Illumina sequencing has an error rate of up to 1%. Errors arise from errors in base assignment (base calling) during template preparation, library preparation, and sequencing. These errors are particularly problematic when searching for low-frequency mutations. The methods disclosed herein address these and other issues.

[0019] The method of the present invention provides high-throughput profiling of cancer gene panels with high sensitivity and specificity for gene variants.The method provides non-invasive genotyping and detection of ctDNA for both research and clinical purposes.The present invention utilizes non-unique barcodes along with the endogenous barcode of the target nucleic acid to provide high sensitivity and specificity in genotyping assays.The method is useful for small amounts of sample DNA, such as ctDNA.

[0020] The number of cfDNA input molecules (i.e., genome equivalents) is typically very low in plasma, making ctDNA recovery difficult. Library preparation and sequencing introduce errors that significantly hinder the investigation of rare mutations. The methods of the present invention achieve high detection limits (as low as 0.05-0.1%) in cfDNA and can discover mutations in many malignancies that would otherwise go undetected by conventional methods. These methods improve the sensitivity and specificity of detecting low-frequency alleles. The present invention recognizes that DNA can be identified with high levels of sensitivity and specificity using a combination of non-unique barcodes and the molecular ends of DNA molecules.

[0021] The method generally involves tagging cfDNA fragments with a pool of non-unique barcodes and performing paired-end sequencing to identify exogenous barcodes and fragment-specific endogenous barcodes. While most conventional barcoding methods are PCR-based, the method disclosed herein uses a capture-based approach that uses a limited, predetermined set of barcodes superimposed on endogenous barcodes. The capture-based approach involves creating a genome library and capturing specific regions. This approach is superior to PCR-based methods due to its scalability, flexibility, and improved coverage uniformity. The capture-based method can simultaneously interrogate thousands of genome locations with high sensitivity and specificity. In this method, each end of the fragment is sequenced to identify the endogenous barcode sequence at the fragment end combined with the exogenous barcode. The combination of the pool of exogenous barcodes with the mapping position of the DNA fragment provides all the complexity required to identify fragments with sufficient sensitivity and specificity.

[0022] For example, if 100 endogenous barcodes (which can be generated by random shearing, exonuclease digestion, or natural fragmentation that may be present with cell-free DNA) are present at either end of a fragment, 10,000 different molecules can be evaluated using paired-end sequencing. Thus, assigning a pool of eight non-unique barcodes, for example, would yield 80,000 combinations. Such an assay can identify mutations ranging from 0.1 to 0.05%. For assays requiring that level of sensitivity, the present disclosure demonstrates that a limited set of non-unique barcodes provides all the diversity required for such an assay. According to the present invention, a small pool of non-unique exogenous barcodes can be overlaid on endogenous terminal regions to achieve a level of sensitivity comparable to conventional, more complex barcoding schemes while providing a robust assay at significantly reduced cost and complexity. These figures are merely examples and can be increased or decreased as needed to suit a particular assay.

[0023] Sequencing can be performed at a depth of 2x, 10x, 50x, 100x, 1,000x, 10,000x, or 50,000x or greater. Duplicate sequence reads are compared and adjusted to distinguish somatic mutations from sequencing or other processing errors. If a mutation was present in the original DNA molecule, it should be present in all sequence reads for that locus, regardless of any subsequent sequencing errors. For example, a mutation can be called if a certain percentage of reads contain the putative mutation. The threshold percentage for making a mutation call can be 25%, 50%, 60%, 75%, 90%, 95%, and 99%, for example. The threshold can be set based on the number of sequence reads obtained and the specific requirements of the assay. Similarly, mutations that do not occur in the template DNA would not be expected to appear in a significant percentage of reads. These variants can then be dismissed as sequencing errors, replication errors, or other processing errors. A consensus sequence can be determined by comparing and aligning the sequence reads.

[0024] The methods of the present invention involve isolating nucleic acids from a sample. The nucleic acid can be cfDNA, including ctDNA. While this method is particularly useful for cfDNA, other types of nucleic acids, including RNA, can be used as well. The sample can include, for example, cell-free nucleic acids (including DNA or RNA) or nucleic acids isolated from tumor tissue samples such as biopsy tissue, formalin-fixed paraffin-embedded tissue (FFPE), frozen tissue, cell lines, DNA, and tumor explants. Samples provided as FFPE blocks or frozen tissue can be pathologically examined to determine tumor cellularity. Tumors can be excised macroscopically or microscopically to remove contaminating normal tissue. Samples can also be derived from a patient's lymphocytes, blood, saliva, cells obtained by oral swab, or other unaffected tissue. The cell-free nucleic acids can be fragments of DNA or ribonucleic acid (RNA) present in the patient's bloodstream. In a preferred embodiment, the circulating cell-free nucleic acids are one or more fragments of DNA obtained from the patient's plasma or serum.

[0025] Cell-free nucleic acids can be isolated, for example, using the QIAmp system from Qiagen (Venlo, Netherlands) using the Triton / Heat / Phenol protocol (THP) (Xue, et al., "Optimizing the Yield and Utility of Circulating Cell-Free DNA from Plasma and Serum", Clin. Chim. Acta., 2009; 404(2): 100-104), whole genome amplification by blunt-end ligation (BL-WGA) (Li, et al., "Whole Genome Amplification of Plasma-Circulating DNA Enables Expanded Screening for Allelic Imbalance in Plasma", J. Mol Diagn. 2006 Feb; 8(1): 22-30), or by other methods from Macherey-Nagel, GmbH & Co. KG (Duren, Germany). The nucleic acids can be isolated according to techniques known in the art, including the NucleoSpin system from Sigma-Aldrich (Germany). In an exemplary embodiment, a blood sample is drawn from a patient and plasma is isolated by centrifugation. Circulating cell-free nucleic acids can then be isolated by any of the above techniques.

[0026] In general, nucleic acids can be extracted, isolated, amplified, or analyzed by a variety of techniques, such as those described in Green and Sambrook, Molecular Cloning: A Laboratory Manual (Fourth Edition), Cold Spring Harbor Laboratory Press, Woodbury, NY 2,028 pages (2012); or those described in U.S. Patent Nos. 7,957,913; 7,776,616; 5,234,809; U.S. Patent Application Publication Nos. 2010 / 0285578; and 2002 / 0190663.

[0027] Nucleic acids obtained from biological samples can be fragmented to generate fragments suitable for analysis. Methods for fragmenting nucleic acids are known in the art. Template nucleic acids can be fragmented or sheared to the desired length using a variety of mechanical, chemical, and / or enzymatic methods. Nucleic acids can be sheared by sonication, brief exposure to DNase / RNase, a hydroshear device, one or more restriction enzymes, transposases, or cleavage enzymes, exposure to heat and magnesium, or by shearing. Nucleic acids can also be naturally fragmented, as can cell-free DNA. Biological samples can be lysed, homogenized, or fractionated, as needed, in the presence of detergents or surfactants. Suitable detergents include ionic detergents (e.g., sodium dodecyl sulfate or N-lauroyl sarcosine), or non-ionic detergents (e.g., polysorbate 80, sold under the trade name TWEEN by Uniqema Americas (Paterson, NJ), or C known as TRITON X-100). 14 H 22 O(C2H4) n ) can be used. The resulting fragments can be any size, for example, 10 bp, 50 bp, 100 bp, 500 bp, 1,000 bp, 5,000 bp, or larger. Shearing can be followed by end repair and A-tailing. Sequencing adapters can be ligated according to standard sequencing protocols.

[0028] Can use hybrid capture probe with selectable oligonucleotide to obtain nucleic acid of interest.For example, see Lapidus (US Patent No. 7,666,593), the whole content of which is incorporated herein by reference.The conventional method for making and using hybridization probe can be found in standard laboratory manuals: Genome Analysis: A Laboratory Manual Series (Vols.I-IV), Cold Spring Harbor Laboratory Press; PCR Primer: A Laboratory Manual, Cold Spring Harbor Laboratory Press; and Sambrook, J et al., (2001) Molecular Cloning: A Laboratory Manual, 2nd edition.(Vols.1-3), Cold Spring Harbor Laboratory Press, etc.

[0029] After the above-mentioned processing steps, nucleic acid can be sequenced.Sequencing can be by any method known in the art.DNA sequencing technology includes classical dideoxy sequencing reaction (Sanger method) using labeled terminator or primer and gel separation in slab or capillary, and next-generation sequencing, such as sequencing by synthesis using reversibly terminated labeled nucleotides, pyrosequencing, 454 sequencing, Illumina / Solexa sequencing, allele-specific hybridization to a library of labeled oligonucleotide probes, sequencing by synthesis and subsequent ligation using allele-specific hybridization to a library of labeled clones, real-time monitoring of the incorporation of labeled nucleotides during polymerization, polony sequencing, and SOLiD sequencing.Separated molecules can be sequenced by sequential or single extension reaction using polymerase or ligase, and by single or sequential differential hybridization using a library of probes.

[0030] Sequencing techniques that can be used include, for example, the use of sequencing-by-synthesis systems sold under the trade names GS JUNIOR, GS FLX+, and 454 SEQUENCING by 454 Life Sciences, a Roche company (Branford, CT), as described in Margulies, M. et al., Genome sequencing in micro-fabricated high-density picotiter reactors, Nature, 437:376-380 (2005); U.S. Patent Nos. 5,583,024; 5,674,713; and 5,700,673, the contents of which are incorporated herein by reference in their entireties.

[0031] Other examples of DNA sequencing technology include SOLiD technology by Applied Biosystems of Life Technologies Corporation (Carlsbad, CA), and ion semiconductor sequencing, for example, using the system sold by Ion Torrent of Life Technologies (South San Francisco, CA) under the trade name ION TORRENT.Ion semiconductor sequencing is described, for example, in Rothberg, et al., An integrated semiconductor device enabling non-optical genome sequencing, Nature 475:348-352 (2011); U.S. Patent Application Publication Nos. 2010 / 0304982; 2010 / 0301398; 2010 / 0300895; 2010 / 0300559; and 2009 / 0026082, the contents of each of which are incorporated herein by reference in their entirety.

[0032] Another example of a sequencing technology that can be used is Illumina sequencing. Illumina sequencing is based on the amplification of DNA on a solid surface using foldback PCR and anchor primers. Adapters are added to the 5' and 3' ends of naturally or experimentally fragmented DNA. The DNA fragments attached to the surface of the flow cell channel are extended and bridge-amplified. The fragments become double-stranded, and the double-stranded molecules are denatured. Multiple cycles of solid-phase amplification and subsequent denaturation can generate millions of clusters of approximately 1,000 copies of single-stranded DNA molecules of the same template in each channel of the flow cell. Continuous sequencing is performed using primers, DNA polymerase, and four fluorophore-labeled reversibly terminating nucleotides. After the nucleotide is incorporated, a laser is used to excite the fluorophore, and an image is taken to record the identity of the first base. The steps of removing the 3' terminator and fluorophore from each incorporated base, incorporating, detecting, and identifying are repeated. Sequencing by this technique is described in U.S. Patent Nos. 7,960,120; 7,835,871; 7,232,656; 7,598,035; 6,911,345; 6,833,246; 6,828,100; 6,306,597; 6,210,891; U.S. Patent Application Publication Nos. 2011 / 0009278; 2007 / 0114362; 2006 / 0292611; and 2006 / 0024681, each of which is incorporated by reference in its entirety.

[0033] One limitation of sequencing technology is the occurrence of sequencing artifacts.The general approach to reduce sequencing artifacts is molecular barcoding.Most barcode methods involve tagging DNA fragments with identifiers, and these identifiers can be tracked through assays, which allows distinguishing between somatic mutations and sequencing errors.

[0034] The term barcode encompasses both exogenous barcodes, which are introduced into sample DNA fragments, and endogenous barcodes, which are terminal sequences resulting from DNA fragmentation by biological or experimental shearing. Barcodes can contain any number of nucleotides, such as 2, 4, 8, 16, or more nucleotides.

[0035] Exogenous barcodes can be generated by methods known in the art. For example, exogenous barcodes can be created by adding random nucleotides to a short sequence assembled on a substrate. Exogenous barcodes can be enzymatically generated by polymerase extension on a degenerate synthetic template, or can be synthesized in a single unit with an adapter sequence. Synthesizing barcodes allows for greater control over their composition, but can be expensive. Therefore, using a limited pool of barcodes can make assays more cost-effective.

[0036] Barcodes can be completely random or genetically engineered with a predetermined sequence. Barcodes can have random or semi-random regions and other fixed regions. Barcodes can contain other regions such as priming sites, adapters, or other complementary regions that will facilitate further processing and analysis.

[0037] Exogenous barcodes can be attached to nucleic acid fragments by methods known in the art, such as PCR or enzymatic ligation. Exogenous barcodes can be attached to one or both ends of the fragment. Barcode molecules can be commercially available, such as from Integrated DNA Technologies (Coralville, IA). In certain embodiments, one or more barcodes are attached to each, any, or all of the fragments. Barcode sequences generally contain specific features that make the sequence useful in sequencing reactions. Methods for designing a set of barcode sequences are described, for example, in U.S. Patent No. 6,235,475, the entire contents of which are incorporated herein by reference. Attachment of barcode sequences to nucleic acid templates is described in U.S. Patent Application Publication Nos. 2008 / 0081330 and 2011 / 0301042, the entire contents of which are incorporated herein by reference. Methods for designing sets of barcode sequences and other methods for attaching barcode sequences are set forth in U.S. Patent Nos. 6,138,077; 6,352,828; 5,636,400; 6,172,214; 6,235,475; 7,393,665; 7,544,473; 5,846,719; 5,695,934; 5,604,097; 6,150,516; Reissue Patent No. 39,793; U.S. Patent No. 7,537,897; 6,172,218; and 5,863,722, the contents of each of which are incorporated herein by reference in their entirety. Barcodes for sequencing and copy number estimation are described in U.S. Patent Application Publication No. 2016 / 0046986, which is incorporated by reference in its entirety.

[0038] The present disclosure utilizes non-unique barcodes to confer high sensitivity and specificity in genotyping assays. In other contexts, such as the publications referenced above, barcodes are sometimes referred to as unique identifiers (UIDs). Herein, we do not use that term, as the exogenous barcodes of the present methods do not need to be unique. Traditional barcoding methods emphasize the need to generate thousands or millions of barcode sequences or combinations to ensure with high certainty that no two fragments receive the same barcode. The present disclosure demonstrates, contrary to conventional wisdom, that a smaller pool of non-unique barcodes overlaid on endogenous barcodes can have the same level of diversity as traditional schemes, while reducing complexity and increasing assay robustness.

[0039] The present invention recognizes that while some degree of barcoding is necessary to reduce background noise in sequencing assays, prior art barcoding methods overestimate the problem. Conventional methods involve generating thousands or even millions of barcode combinations. Generating these barcodes makes genotyping assays overly complex and less robust. The present disclosure demonstrates that the same level of specificity can be achieved with significantly less complexity.

[0040] When barcoded fragments are sequenced, multiple reads are generated. The reads can be about 50 to 200 bases in length. In some embodiments, shorter reads can be obtained, for example, reads less than about 50 or about 30 bases in length. Some sequencing techniques can generate reads hundreds or thousands of bases in length.

[0041] The set of sequence reads can be analyzed by any suitable method known in the art.For example, in some embodiments, the sequence reads are analyzed by the hardware or software that is provided as part of the sequencing device.In some embodiments, each sequence read is examined by visual inspection (for example, on a computer monitor).

[0042] Sequence assembly can be performed by methods known in the art, including reference-based assembly, de novo assembly, alignment-based assembly, or combinatorial methods. In some embodiments, sequence assembly uses the low-coverage sequence assembly software (LOCAS) tool described by Klein et al. in LOCAS-A low coverage sequence assembly tool for re-sequencing projects, PLoS One 6(8) article 23455 (2011), the entire contents of which are incorporated herein by reference. Sequence assembly is described in U.S. Patent Nos. 8,165,821; 7,809,509; 6,223,128; U.S. Patent Application Publication Nos. 2011 / 0257889; and 2009 / 0318310, the entire contents of which are incorporated herein by reference.

[0043] FIG. 1 shows a method 100 for analyzing nucleic acids according to the present disclosure. The method 100 includes step 113 of obtaining a sample containing nucleic acid fragments. Step 113 may include obtaining a plasma sample from a patient and extracting nucleic acid fragments. The nucleic acid may include cell-free DNA, circulating tumor DNA, tumor DNA, or RNA. The fragments may be end-repaired, A-tailed, and ligated with adapters. In step 119, a set of non-unique barcodes is introduced to generate a genomic library. In step 125, the fragments are sequenced to generate sequence reads, and the sequence reads are aligned. Sequencing may include sequencing each fragment in duplicate. In step 131, genomic locations of the fragment ends are identified. In step 137, mutations present in multiple molecules, as determined by a combination of the non-unique barcodes and the genomic locations of the fragment ends, are identified.

[0044] The method can include performing hybrid capture on a genomic library, which hybrid capture can include detecting ABL1, AKT1, ALK, APC, AR, ATM, BCR, BRAF, CDH1, CDK4, CDK6, CDKN2A, CSF1R, CTNNB1, DNMT3A, EGFR, ERBB2, ERBB4, ESR1, EZH2, FBXW7, FGFR1, FGFR2, FGFR3, FLT3, GNA11, GNAQ, GNAS, HNF1A, HRAS, IDH1, IDH2, JAK2, J May include a panel of well-characterized cancer genes such as AK3, KDR, KIT, KRAS, MAP2K1, MET, MLH1, MPL, MYC, NPM1, NRAS, NTRK1, PDGFRA, PDGFRB, PIK3CA, PIK3R1, PTEN, PTPN11, RARA, RB1, RET, ROS1, SMAD4, SMARCB1, SMO, SRC, STK11, TERT, TP53, and VHL.

[0045] Figure 2 shows a method 200 for molecular barcoding according to the present disclosure. Method 200 includes step 209 of obtaining a sample having nucleic acid fragments and step 215 of providing a plurality of sets of non-unique barcodes. In step 221, the nucleic acid fragments are tagged with barcodes to create a genomic library. Because there are a limited number of sets of non-unique barcodes (e.g., eight different sets), each nucleic acid fragment is tagged with the same barcode as at least one other type of nucleic acid fragment in the genomic library. Thus, the exogenous barcodes are not "unique." The genomic locations of the fragments can be identified by the endogenous barcodes resulting from the fragmentation of the nucleic acid.

[0046] In some embodiments, method 200 further comprises redundantly sequencing the genomic library to generate multiple overlapping sequence reads for each nucleic acid fragment. Method 200 may further comprise aligning the overlapping sequence reads of similarly tagged nucleic acid fragments. Method 200 may further comprise aligning the aligned sequence reads to a reference to determine a consensus sequence.

[0047] The disclosed approach is useful for any sequencing assay that requires high sensitivity and specificity. This method is particularly useful for sequencing small amounts of cfDNA isolated from plasma and examining them for somatic mutations. [Example]

[0048] Validation studies were conducted for use in research. The goal of the study was to demonstrate that next-generation library preparation combined with targeted gene capture using a panel is reproducible and accurate for sequencing on the Illumina HiSeq sequencing platform. The panel under test was a targeted panel of well-characterized cancer genes known as the PlasmaSelect™ panel, currently under development by PGDx (Baltimore, MD). Validation of this approach using a combination of cell line-derived and clinical plasma samples will enable the identification of tumor-specific sequence mutations, amplifications, and translocations in a set of genes relevant to clinical and biomedical cancer research. The scope of validation of this method includes the use of this assay in studies utilizing plasma samples from cancer patients for evaluation of the genes shown in Figures 3 and 4.

[0049] Method and process description 1. Sample preparation, library creation, and DNA capture DNA extraction and processing Sequencing analysis of targeted genes in cell-line- and plasma-derived cell-free DNA (cfDNA) has been performed to identify tumor-specific (somatic) alterations. Two technical challenges to implementing these approaches in the form of liquid biopsies include the limited amount of DNA obtained and the low mutant allele frequency associated with these alterations. It has been described that only a few thousand genome equivalents are obtained per milliliter of plasma, and mutant allele frequencies can range from <0.01% to >50% in total cfDNA (Bettegowda et al., 2014). The disclosed technology overcomes this issue and provides an optimized method for converting cell-free DNA into a genomic library, as well as improved digital sequencing approaches to enhance test sensitivity and the specificity of next-generation sequencing approaches. Utilizing digital sequencing technology in conjunction with a redundant sequencing error-correction approach effectively reduces the error rate introduced by next-generation sequencing and enables accurate identification of sequence mutations (see Figure 5, single bases and small insertions and deletions).

[0050] Library preparation and target capture Briefly, cell-free DNA was extracted from cell lines or plasma specimens and prepared into genomic libraries suitable for next-generation sequencing using oligonucleotide barcodes via end-repair, A-tailing, and adapter ligation. In-solution hybrid capture utilizing 120 base pair (bp) RNA oligonucleotides was performed for both the sequence variant panel (Figure 3) and the structural variant panel (Figure 4).

[0051] 2. Sequencing Enriched cell line- or plasma-derived captured DNA libraries were sequenced for each target base using paired-end Illumina HiSeq2500 sequencing chemistry to an average total target coverage of >20,000x for sequence variants or >5,000x for translocations. Sequence data were mapped to the reference human genome sequence to examine coding and intronic regions for somatic mutations.

[0052] 3. Bioinformatics The data are analyzed using sophisticated bioinformatics approaches, including novel genetic analysis methods and proprietary data analysis algorithms, to sensitively and specifically identify tumor-specific alterations and integrate sequence information, genomic data, and cancer genes and pathways to provide the most complete and informative data set to guide patient management. Briefly, these steps include: 1. Initial processing of next-generation sequencing data 2. Alignment of next-generation sequencing data to the human reference genome using ELAND and Novoalign 3. Analysis of next-generation sequencing data for sequence variation 4. Analysis of Next-Generation Sequencing Data for Focal Amplification 5. Analysis of next-generation sequencing data for translocations.

[0053] Study plan and sample set Sample type To evaluate assay performance, validation studies were performed using a combination of all cancer cell lines (Table 1), plasma from end-stage breast, colon, and lung cancer patients, and samples from healthy donors (Tables 1-4 and Figures 6 and 7). Clinical samples from both healthy donors and end-stage cancer patients were obtained retrospectively by ILSBio (Chestertown, MD). Cell line specimens were obtained from ATCC (Manassas, VA), and DNA was extracted, sheared, and purified to a fragment length profile consistent with cell-free DNA obtained from plasma. These samples were then evaluated using the PlasmaSelect™ 64 panel according to the relevant standard operating procedures (SOPs).

[0054] (Table 1) All cancer cell lines and sequence mutations TIFF0007781105000001.tif70128

[0055] Table 2. Sequence variation and amplification analysis performed for validation of the PlasmaSelect™ 64 method. TIFF0007781105000002.tif95148TIFF0007781105000003.tif221148 * Tumor purity of cell line samples was obtained by titrating tumor and normal DNA at the indicated ratios for a given DNA input to achieve the indicated tumor purity. Manufacturer guidelines were followed for reagents used in library preparation.

[0056] Table 3. Reconstitution analysis performed for validation of the PlasmaSelect™ 64 method. TIFF0007781105000004.tif213145 * Tumor purity of cell line samples was achieved by titrating tumor and normal DNA at the indicated ratios for a given DNA input to achieve the indicated tumor purity. Manufacturer guidelines were followed for reagents used in library preparation.

[0057] Table 4. Clinical plasma samples from 18 patients with breast, colon, and lung cancer. TIFF0007781105000005.tif105128

[0058] Test performance pass criteria: 1. Accuracy: Sequence variation Accuracy was assessed by comparing the results of the target capture panel and next-generation sequencing from proprietary cell lines with published, independently obtained Sanger sequencing results. A total of 19 positions known to be mutated in proprietary cell lines were included in the target panel, and these were assessed at 1%, 2%, 5%, 20%, 25%, and 100% tumor purity using 250 ng of DNA. Additionally, a combination cancer cell line containing 12 sequence mutations was assessed at 100% and 1% tumor purity using 250 ng of DNA. Finally, specificity was assessed by analyzing 18 plasma samples derived from healthy donors. None of these plasma samples would be expected to harbor any somatic alterations.

[0059] Performance Metrics Sensitivity 100.0% Specificity (in doubt) 99.9997% Specificity (healthy donors) 99.9996%

[0060] amplification Accuracy was assessed by comparing the results of the target capture panel and next-generation sequencing from proprietary cell lines with published, independently obtained SNP array results. Three amplifications were included in the target region of interest and were assessed at 20%, 25%, and 100% tumor purity using 250 ng of DNA. Specificity was further assessed by analysis of 18 plasma samples from healthy donors, none of which would be expected to harbor any somatic alterations.

[0061] Performance Metrics Sensitivity 100.0% Specificity (in doubtful cases) 91.7% Specificity (healthy donors) 100.0%

[0062] reorganization Accuracy was assessed by comparing results from various proprietary cell lines using target capture panels and next-generation sequencing at 1%, 2%, 20%, and 100% tumor purity combinations using 250 ng of DNA with published independently obtained results (Shibata et al., 2010 and Koivunen et al., 2008). Specificity was further assessed by analysis of 18 plasma samples from healthy donors, none of which would be expected to harbor any somatic alterations.

[0063] Performance Metrics Sensitivity 100.0% Specificity (in doubtful cases) 100.0% Specificity (healthy donors) 99.7%

[0064] 2. Analytical sensitivity (detection limit): Sequence variation Analytical sensitivity was assessed by comparing the results of the target capture panel and next-generation sequencing from proprietary cell lines with published, independently obtained Sanger sequencing results. A total of 19 positions known to be mutated in proprietary cell lines were included in the target panel and evaluated in duplicate at tumor purities of 0.1%, 0.2%, 0.5%, 1%, and 2% using 250 ng of DNA, and at tumor purities of 0.5%, 1%, 2%, 5%, and 10% using 25 ng of DNA. Additionally, combinatorially mutated cell lines containing 12 sequence mutations were evaluated at tumor purities of 0.1%, 0.2%, 0.5%, and 1% using 250 ng of DNA, and at tumor purities of 0.5%, 1.0%, 2.0%, and 5% using 25 ng of DNA.

[0065] Performance Metrics Analytical sensitivity: 99.4%

[0066] amplification Assay sensitivity was assessed by comparing the results of the target capture panel and next-generation sequencing from proprietary cell lines with published independently obtained SNP array results. Three amplifications were included in the target region of interest and were assessed in duplicate with 250 ng of DNA at 60%, 40%, and 20% tumor purity, and with 25 ng of DNA at 60%, 40%, and 20% tumor purity.

[0067] Performance Metrics Analytical sensitivity 97.2%

[0068] reorganization Assay sensitivity was assessed by comparing results from various proprietary cell lines with target capture panels and next-generation sequencing at combinations of 0.1%, 0.5%, and 1.0% tumor purity using 250 ng of DNA, and 0.5%, 1.0%, and 2.0% tumor purity using 25 ng of DNA, with independently published results (Shibata et al., 2010 and Koivunen et al., 2008).

[0069] Performance Metrics Analytical sensitivity: 94.4%

[0070] 3. Precision and robustness (intra- and inter-assay reproducibility): Sequence variation Accuracy and robustness were assessed by comparing the results obtained from the target capture panel and next-generation sequencing of proprietary cell lines with the independently obtained published Sanger sequencing results. A total of 19 positions known to be mutated in proprietary cell lines were included in the target panel, and these were assessed at 2% tumor purity using 150 ng of DNA both within and across sample preparations (different operators on different days).

[0071] Performance Metrics Intra-assay concordance 100.0% Inter-assay agreement: 100.0%

[0072] amplification Accuracy and robustness were assessed by comparing results from proprietary cell lines with target capture panels and next-generation sequencing to published, independently obtained SNP array results. Three amplifications were included in the target region of interest and were assessed at 20% tumor purity using 100 ng of DNA both within and across sample preparations (different operators on different days).

[0073] Performance Metrics Intra-assay concordance 94.7% Inter-assay agreement: 89.5%

[0074] reorganization Accuracy and robustness were assessed by comparing results from a range of proprietary cell lines with target capture panels and next-generation sequencing at 2% and 5% tumor purity using 25ng and 150ng of DNA both within and across sample preparations (different operators on different days) to published, independently obtained results (Shibata et al., 2010 and Koivunen et al., 2008).

[0075] Performance Metrics Intra-assay concordance 100.0% Inter-assay agreement: 100.0%

[0076] 4. Failure rate In total, there were 113 sequence panel (PS_Seq2) and 112 structural panel (PS_Str2) next-generation sequencing libraries generated using six libraries, with a processing failure rate of 2.7% (6 / 225).

[0077] 5. Comparison of blood collection tube types To evaluate the impact of blood collection tube type on the performance of the PlasmaSelect™ R64 approach, 4 x 10 ml blood samples were taken from nine cancer patients, 2 x 10 ml of blood was collected in K2EDTA tubes, and 2 x 10 ml was collected in Streck tubes and processed to plasma using either PGDx (K2EDTA) or according to the manufacturer's specifications (Streck). These data showed a very high concordance across reported results.

[0078] Performance Metrics Sequence mutation concordance rate 100.0% [MAF≧0.50%] Amplification concordance rate 98.8% Reorganization match rate 100.0%

[0079] 6. Stability: Manufacturer guidelines were followed for reagents used in sample library preparation, and all samples were collected following the same sample protocol and handling procedures.

[0080] Figure 6 shows the results of observed and predicted variant allele frequencies for sequence variations across all cancer cell lines. The calculated variant allele frequencies (MAFs) were compared to the predicted MAFs for the cases evaluated in a validation study of the method's accuracy, analytical sensitivity, precision, and robustness for the combined cancer cell lines (n = 12 predicted variations for each case).

[0081] Figure 7 shows the results of observed and expected variant allele frequencies for the internal control breast cancer cell line. The calculated variant allele frequencies (MAFs) were compared to the expected MAFs for cases evaluated in a validation study of the method for accuracy, analytical sensitivity, precision, and robustness of the combined cancer cell lines (n = 19 expected changes for each case).

[0082] Conclusions and Recommendations The PlasmaSelect™ assay has been validated to achieve high levels of sensitivity and specificity in the detection of sequence variants (SBS / indels), amplifications, and translocations in cell-free DNA obtained from the plasma of cancer patients for liquid biopsy analysis.

[0083] Performance Metrics (minimum sample input of 25ng): Table 5. Summary of PlasmaSelect™ 64 performance metrics TIFF0007781105000006.tif67139 * Per-base specificity provided by sequence variant analysis [99,359 bases evaluated]

[0084] INCORPORATION BY REFERENCE Any and all references and citations made throughout this disclosure to other documents, such as patents, patent applications, patent publications, journals, books, articles, web content, and the like, are incorporated herein by reference in their entirety for all purposes.

[0085] equivalent The present invention may be embodied in other specific forms without departing from its spirit or essential characteristics. The foregoing embodiments, therefore, are to be considered in all respects as illustrative rather than limiting of the invention described herein.

[0086] Sequence information SEQUENCE LISTING <110> PERSONAL GENOME DIAGNOSTICS INC. <120> NON-UNIQUE BARCODES IN A GENOTYPING ASSAY <150> US 62 / 422,355 <151> 2016-11-15 <160> 2 <170> PatentIn version 3.5 <210> 1 <211> 50 <212> DNA <213> Artificial Sequence <220> <223> Synthetic <400> 1 actgactgac tgactgactg actgactgac tgactgactg actgactgac 50 <210> 2 <211> 50 <212> DNA <213> Artificial Sequence <220> <223> Synthetic <400> 2 actgactgac tgactgactg actgactgac agactgactg actgactgac 50

Claims

1. A method for analyzing nucleic acids in a sample of nucleic acid fragments, comprising the steps of: introducing a limited pool of non-unique exogenous barcodes into the nucleic acid fragments, wherein the limited pool comprises eight sets of non-unique exogenous barcodes; preparing a genomic library; identifying the genomic locations of the ends of the nucleic acid fragments; sequencing the nucleic acid fragments to generate sequence reads, wherein the sequencing comprises overlapping sequencing, and wherein the overlapping sequencing is performed at a depth of 2x, 10x, 50x, or 100x; and aligning the sequence reads to identify variants; wherein the terminal portions of the nucleic acid fragments comprise endogenous barcodes.

2. 2. The method of claim 1, wherein the identified variant is selected from tumor-specific somatic mutations, amplifications, and translocations.

3. The method described in claim 1, wherein the step of identifying the genomic location of the terminal portion of the nucleic acid fragment comprises hybrid capture or whole genome sequencing.

4. 4. The method of claim 3, wherein the hybrid capture comprises a well-characterized cancer gene panel.

5. The oncogene is selected from the group consisting of ABL1, AKT1, ALK, APC, AR, ATM, BCR, BRAF, CDH1, CDK4, CDK6, CDKN2A, CSF1R, CTNNB1, DNMT3A, EGFR, ERBB2, ERBB4, ESR1, EZH2, FBXW7, FGFR1, FGFR2, FGFR3, FLT3, GNA11, GNAQ, GNAS, HNF1A, HRAS, IDH1, IDH2, 5. The method of claim 4, comprising JAK2, JAK3, KDR, KIT, KRAS, MAP2K1, MET, MLH1, MPL, MYC, NPM1, NRAS, NTRK1, PDGFRA, PDGFRB, PIK3CA, PIK3R1, PTEN, PTPN11, RARA, RB1, RET, ROS1, SMAD4, SMARCB1, SMO, SRC, STK11, TERT, TP53, and VHL.

6. 10. The method of claim 1, wherein the sequencing step comprises single-end or paired-end sequencing.

7. 10. The method of claim 1, wherein the nucleic acid comprises cell-free DNA, circulating tumor DNA, tumor-derived DNA, or RNA.

Citation Information

Patent Citations

  • Methods and systems for detecting genetic variants

    US20160046986A1