Methods for targeted nucleic acid sequence enrichment with application to error-corrected nucleic acid sequencing
Duplex sequencing with uniquely labeled strands corrects errors in nucleic acid sequencing by comparing complementary strands, addressing PCR stutter and allele dropout, enhancing accuracy and efficiency in forensic and clinical applications.
Patent Information
- Application Number
- JP2023057239
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2017-10-23
- Filing Date
- 2023-03-31
- Publication Date
- 2026-02-27
- Estimated Expiration
- 2038-03-23
AI Technical Summary
Current nucleic acid sequencing technologies face challenges in accurately genotyping short tandem repeat (STR) loci due to PCR stutter, allele dropout, and imbalanced amplification, especially in degraded DNA samples, limiting their effectiveness in forensic and clinical applications.
A method involving duplex sequencing with uniquely labeled strands in a double-stranded nucleic acid complex, allowing for error correction by comparing sequences of complementary strands to identify true variants and artifacts, thereby improving accuracy and reliability.
Enables highly accurate, cost-effective, and efficient sequencing of small nucleic acid samples, capable of detecting low-frequency mutations and improving genotyping accuracy in degraded or mixed DNA samples.
Smart Images

Figure 0007821756000012 
Figure 0007821756000013 
Figure 0007821756000014
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Patent Application No. 62 / 475,682, filed March 23, 2017, and U.S. Provisional Patent Application No. 62 / 575,958, filed October 23, 2017, the disclosures of which are incorporated herein by reference in their entireties.
[0002] Government Statement of Interest This invention was made with government support under Grant Nos. R01 CA160674 and R01 CA181308 awarded by the National Institutes of Health and Grant No. W911NF-15-2-0127 awarded by the U.S. Army Research Office. The government has certain rights in this invention. [Background technology]
[0003] Previous approaches to certain types of genetic analysis, such as forensic DNA analysis, rely on capillary electrophoresis (CE) separation of PCR amplification products (PCR-CE) to identify length polymorphisms in short tandem repeat sequences. This type of analysis has proven extremely valuable since its introduction around 1991. Since then, several publications have introduced standardized protocols, validated its use in laboratories worldwide, detailed its use in many different populations, and introduced more efficient approaches such as miniSTR.
[0004] Although this approach has proven highly successful, the technology suffers from a number of drawbacks that limit its usefulness. For example, current approaches to STR genotyping often generate background signals due to PCR stutter caused by polymerase slippage on the template DNA, ultimately resulting in a mixture of PCR amplification products of different lengths in the completed reaction. This issue is particularly significant in samples with more than one contributor (e.g., a mixture of DNA from different specific individuals containing a specific genetic makeup with different STR length variants) due to the difficulty of distinguishing between stutter and genuine alleles. Another problem arises when analyzing degraded DNA samples. Damaged DNA can exacerbate the degree of stutter and PCR errors. Fragment length variation often results in significantly fewer or even absent longer PCR fragments. As a result, capillary electropherogram profiles from degraded DNA often have lower discriminatory power.
[0005] The introduction of massively parallel sequencing (MPS, often known as next-generation DNA sequencing, or NGS) systems has the potential to address several challenging issues in forensic analysis. For example, these platforms offer the previously unparalleled ability to simultaneously analyze nuclear and mitochondrial DNA (mtDNA) STRs and single nucleotide polymorphisms (SNPs), dramatically increasing discriminatory power and offering the potential to determine ethnicity and even physical attributes (phenotypes). Furthermore, unlike PCR-CE, which simply reports the average genotype of a population of molecules, MPS technology digitally aggregates the complete nucleotide sequence of many individual DNA molecules, offering the unique ability to detect minor allele frequencies (MAFs) within heterogeneous DNA mixtures. Because forensic samples containing more than two contributors remain one of the most problematic issues in forensic science, the impact of MPS on the field of forensic science could be enormous.
[0006] The publication of the human genome highlighted the immense power of MPS platforms. However, until very recently, the full power of these platforms was limited to forensic applications due to the significantly shorter read lengths than short tandem repeat (STR) loci, preventing the ability to call genotypes based on length. Initially, pyrosequencers such as the MPS Roche 454 platform were the only platforms with sufficient read lengths to sequence core standard STR loci. However, as read lengths of competing technologies have increased, their utility for forensic applications has been realized. Overall, the overall result of all these studies is that, regardless of platform, STRs can be successfully classified, thereby generating genotypes comparable to CE analysis, even from damaged forensic samples.
[0007] While many studies have demonstrated agreement with traditional PCR-CE approaches and additional benefits, such as the detection of intra-STR SNPs (single nucleotide polymorphisms), they also highlight numerous issues with current technology. For example, the current MPS approach to STR genotyping relies on multiplex PCR to provide sufficient DNA for sequencing and PCR primer introduction. However, because multiplex PCR kits are designed for PCR-CE, they contain primers for amplification products of various sizes. This variation can lead to imbalanced inclusion, biasing toward amplification of small fragments and potentially resulting in allele dropout. Indeed, recent studies have shown that differences in PCR efficiency affect mixture components, especially at low MAFs.
[0008] Like PCR-CE, MPS is not susceptible to the occurrence of PCR stutter. Most MPS studies of STRs report the occurrence of artificial drop-in alleles. Recently, systematic MPS studies have reported that many stutter events manifest as shorter polymorphisms, differing from true alleles by four base pairs, most commonly at n-4, but also at n-8 and n-12 positions. The percentage of stutter typically occurred in approximately 1% of reads, but could reach 3% at some loci, indicating that MPS may exhibit a higher rate of stutter than PCR-CE.
[0009] Various approaches at the protocol development, chemistry / biochemistry, and data processing levels have been developed to mitigate the impact of PCR-based errors in MPS applications. Furthermore, techniques that can identify PCR copies resulting from individual DNA fragments before or during amplification based on unique random shear points or by exogenous tagging (i.e., the use of molecular barcodes, also known as molecular tags, unique molecular identifiers [UMIs], and single molecular identifiers [SMIs]) are commonly used. This approach has been used to improve the counting accuracy of DNA and RNA templates. Because all amplification products derived from a single starting molecule can be explicitly identified, sequence variations in identically tagged sequencing reads can be used to correct base errors that occur during PCR or sequencing. For example, Kinde et al. (Proc Natl Acad Sci USA 108, 9530-9535, 2011) introduced SafeSeqS, which uses single-strand molecular barcoding to group PCR copies that share barcode sequences and form a consensus, thereby reducing the error rate of sequencing. This approach results in an average limit of detection of point mutations of 0.5%, but its effectiveness for STR loci has not been extensively evaluated.
[0010] Another recently reported approach, MIPSTR, uses targeted capture of STR loci with single-molecule molecular inversion probes (smMIPs) that specifically anneal to sequences flanking the STR loci. After polymerase extension of the 3' end of the smMIP, the ends are ligated and subjected to PCR amplification and sequencing. The use of MIPs specific to the flanking regions of STR loci significantly improves target specificity and improves the accuracy of genotyping STR loci. However, similar to Safe-SeqS, the incorporation of single-stranded molecular barcodes cannot completely eliminate PCR artifacts arising from the first round of amplification that are carried over to derivative copies as "jackpot" events.
[0011] More accurate genotyping methods for STR loci, single nucleotide polymorphism (SNP) loci, and many other forms of mutations and genetic variants are desired for a variety of applications in forensic science, medicine, and the scientific industry. However, the challenge is how to most efficiently generate sequence information from as many relevant copies of the genetic material to be sequenced as possible with maximum reliability and at a reasonable cost. Various consensus sequencing methods (both molecular barcode-based and non-molecular barcode-based) have been used for error correction to better discriminate variants in mixtures (for a detailed discussion, see J. Salk et al., "Enhancing the accuracy of next-generation sequencing for detecting rare and subclonal mutations," Nature Reviews Genetics, 2018), but there are various tradeoffs in performance. We previously described duplex sequencing, an ultra-high-precision sequencing method that relies on genotyping and comparing the sequences of independent strands of double-stranded nucleic acid molecules for the purpose of error correction. The technology articulated herein describes methods for improving cost efficiency, recovery efficiency, and other performance criteria, as well as overall process speed, of duplex sequencing and related MPS sequencing methods. Summary of the Invention [Problem to be solved by the invention]
[0012] The present technology relates generally to methods for targeted nucleic acid sequence enrichment and the use of such enrichment for error-correcting nucleic acid sequencing applications. [Means for solving the problem]
[0013] In some embodiments, highly accurate, error-corrected, massively parallel sequencing of nucleic acid material is possible using a combination of uniquely labeled strands in a double-stranded nucleic acid complex, such that each strand is informationally related to its complementary strand but is differentiated following sequencing of each strand or from the amplification products derived from it; this information can be used for error correction of the determined sequence. Some aspects of the present technology provide methods and compositions for improving the cost, conversion of sequenced molecules, and time efficiency of generating labeled molecules for targeted, ultra-high-precision sequencing. In some embodiments, the provided methods and compositions enable the accurate analysis of very small amounts of nucleic acid material (e.g., from samples taken from crime scenes or small clinical samples or DNA floating freely in blood). In some embodiments, the provided methods and compositions enable the detection of mutations in samples of nucleic acid material present at frequencies of less than 1 in 100 cells or molecules (e.g., less than 1 in 1,000 cells or molecules, less than 1 in 10,000 cells or molecules, or less than 1 in 100,000 cells or molecules).
[0014] In some embodiments, the disclosure provides a method for amplifying nucleic acid material, the method comprising providing double-stranded nucleic acid material, the nucleic acid material comprising a single molecule identifier sequence on each strand of the nucleic acid material and an adapter sequence on at least one of the 5' end and the 3' end of each strand of the nucleic acid material, wherein a first adapter sequence is located on one of the 5' end or the 3' end of a first strand of the nucleic acid material and a second adapter sequence is located on an opposite end of a second strand of the nucleic acid material, the first strand and the second strand originating from the same double-stranded nucleic acid material; and amplifying the nucleic acid material. separating the amplified nucleic acid material into a first sample and a second sample, amplifying the first strand in the first sample through the use of a primer specific to a first adapter sequence to result in a first nucleic acid product, amplifying the second strand in the second sample through the use of a primer specific to a second adapter sequence to result in a second nucleic acid product, sequencing each of the first and second nucleic acid products, and comparing the sequence of the first nucleic acid product to the sequence of the second nucleic acid product. In some embodiments, the nucleic acid material comprises an adapter sequence at each of the 5' and 3' ends of each strand of the nucleic acid material.
[0015] In some embodiments, the present disclosure provides a method comprising: providing a double-stranded nucleic acid material comprising one or more double-stranded nucleic acid molecules, each double-stranded nucleic acid molecule comprising a single molecule identifier sequence on each strand and an adapter on at least one of the 5' and / or 3' ends of the nucleic acid molecule, wherein for each nucleic acid molecule, a first adapter sequence is associated with the first strand of the nucleic acid molecule and a second adapter sequence is associated with the second strand of the nucleic acid molecule; amplifying the nucleic acid material; separating the amplified nucleic acid material into a first sample and a second sample; amplifying the first strand in the first sample through the use of a primer specific to the first adapter sequence to result in a first nucleic acid product; amplifying the second strand in the second sample through the use of a primer specific to the second adapter sequence to result in a second nucleic acid product; sequencing each of the first and second nucleic acid products; and comparing the sequence of the first nucleic acid product to the sequence of the second nucleic acid product. In some embodiments, the nucleic acid material comprises an adapter sequence at each of the 5' and 3' ends of each strand of the nucleic acid material.
[0016] In some embodiments, the disclosure also provides a method for preparing a nucleic acid material comprising the steps of providing a double-stranded nucleic acid material, the nucleic acid material having been cleaved to result in strands of substantially similar length (about 1 to 1,000,000 bases, about 10 to 1,000 bases, or about 100 to 500 bases) upon cleavage with a targeting endonuclease (e.g., a CRISPR-associated (Cas) enzyme / guide RNA complex, e.g., Cas9 or Cpf1, a meganuclease, a transcription activator-like effector-based nuclease (TALEN), a zinc finger nuclease, an Argonaute nuclease), and the nucleic acid material comprising a single molecule identifier sequence on each strand of the nucleic acid material and an adaptor sequence on at least one of the 5' end and the 3' end of each strand of the nucleic acid material, wherein a first adaptor sequence is located on one of the 5' end or the 3' end of a first strand of the nucleic acid material. and a second adapter sequence located on an opposite end of the second strand of the nucleic acid material, the first strand and the second strand originating from the same double-stranded nucleic acid material; amplifying the nucleic acid material; separating the amplified nucleic acid material into a first sample and a second sample; amplifying the first strand in the first sample through the use of a primer specific to the first adapter sequence to produce a first nucleic acid product; amplifying the second strand in the second sample through the use of a primer specific to the second adapter sequence to produce a second nucleic acid product; sequencing each of the first and second nucleic acid products; and comparing the sequence of the first nucleic acid product with the sequence of the second nucleic acid product. In some embodiments, the nucleic acid material comprises an adapter sequence at each of the 5' and 3' ends of each strand of the nucleic acid material.
[0017] In some embodiments, sequencing each of the first and second nucleic acid products includes sequencing at least one of the first strands to determine a first-strand sequence read, sequencing at least one of the second strands to determine a second-strand sequence read, and comparing the first-strand sequence read with the second-strand sequence read to generate an error-corrected sequence read. In some embodiments, the error-corrected sequence read contains matching nucleotide bases between the first-strand sequence read and the second-strand sequence read. In some embodiments, variations occurring at specific positions in the error-corrected sequence read are identified as true variants. In some embodiments, variations occurring at specific positions in only one of the first-strand sequence read or the second-strand sequence read are identified as potential artifacts.
[0018] In some embodiments, the error-corrected sequence reads are used to identify or characterize cancer, cancer risk, cancer mutations, cancer metabolic states, mutator phenotypes, carcinogen exposure, toxin exposure, chronic inflammatory exposure, age, neurodegenerative diseases, pathogens, drug resistance variants, fetal molecules, forensically relevant molecules, immunologically relevant molecules, mutated T cell receptors, mutated B cell receptors, mutated immunoglobulin loci, genomic kataegis sites, genomic hypermutation sites, low frequency variants, subclonal variants, minority molecular populations, contaminant sources, nucleic acid synthesis errors, enzymatic modification errors, chemical modification errors, gene editing errors, gene therapy errors, nucleic acid information storage fragments, microbial quasispecies, viral quasispecies, organ transplants, organ transplant rejections, cancer recurrence, cancer recurrence after therapy, residual cancer after therapy, precancerous states, dysplastic states, microchimerism states, stem cell transplant states, cell therapy states, nucleic acid labels attached to another molecule, or combinations thereof, in the organism or subject from which the double-stranded target nucleic acid molecule is derived. In some embodiments, the error-corrected sequence reads are used to identify exposure to carcinogenic compounds or carcinogens. In some embodiments, the error-corrected sequence reads are used to identify exposure to mutagenic compounds or mutagens. In some embodiments, the nucleic acid material is derived from a forensic sample and the error-corrected sequence reads are used in forensic analysis.
[0019] In some embodiments, the single molecule identifier sequence comprises an endogenous shear point or an endogenous sequence that may be positionally related to a shear point. In some embodiments, the single molecule identifier sequence is at least one of a degenerate or semi-degenerate barcode sequence that uniquely labels a double-stranded nucleic acid molecule, one or more nucleic acid fragment ends of the nucleic acid material, or a combination thereof. In some embodiments, the adapter and / or adapter sequence comprises at least one nucleotide position that is at least partially non-complementary or contains at least one non-standard base. In some embodiments, the adapter comprises a single "U-shaped" oligonucleotide sequence formed by about five or more self-complementary nucleotides.
[0020] According to various embodiments, any of a variety of nucleic acid materials may be used. In some embodiments, the nucleic acid material may include at least one modification to a polynucleotide within a standard sugar-phosphate backbone. In some embodiments, the nucleic acid material may include at least one modification within any base of the nucleic acid material. For example, as a non-limiting example, in some embodiments, the nucleic acid material is or includes at least one of double-stranded DNA, double-stranded RNA, peptide nucleic acid (PNA), and locked nucleic acid (LNA).
[0021] In some embodiments, the providing step ligates the double-stranded nucleic acid material to at least one double-stranded degenerate barcode sequence to form a double-stranded nucleic acid molecule barcode complex, wherein the double-stranded degenerate barcode sequence comprises a single molecular identifier sequence in each strand.
[0022] In some embodiments, amplifying the nucleic acid material in the first sample comprises amplifying the first strand in the first sample through the use of a primer specific to a first adapter sequence of the first strand and a second primer specific to a non-adapter portion of the first strand to provide a first nucleic acid product. In some embodiments, amplifying the second strand in the second sample through the use of a primer specific to a second adapter sequence of the second strand and a second primer specific to a non-adapter portion of the second strand to provide a second nucleic acid product.
[0023] In some embodiments, amplifying the nucleic acid material in the first sample comprises using at least one single-stranded oligonucleotide that is at least partially complementary to a sequence present in the first adapter sequence and at least one single-stranded oligonucleotide that is at least partially complementary to the target sequence of interest to amplify the nucleic acid material derived from a single nucleic acid strand from the original double-stranded nucleic acid molecule such that the single molecule identifier sequence is at least partially maintained.
[0024] In some embodiments, this involves amplifying nucleic acid material derived from a single nucleic acid strand from the original double-stranded nucleic acid molecule using at least one single-stranded oligonucleotide that is at least partially complementary to a sequence present in the second adapter sequence and at least one single-stranded oligonucleotide that is at least partially complementary to the target sequence of interest, such that the single molecule identifier sequence is at least partially maintained.
[0025] In some embodiments, amplifying the nucleic acid material includes generating a plurality of amplification products derived from the first strand and a plurality of amplification products derived from the second strand.
[0026] In some embodiments, the provided methods further include, prior to the providing step, cleaving the nucleic acid material with one or more targeting endonucleases to form target nucleic acid fragments of substantially known length, and isolating the target nucleic acid fragments based on the substantially known length. In some embodiments, the provided methods further include, prior to the providing step, ligating an adaptor (e.g., an adaptor sequence) to the target nucleic acid (e.g., the target nucleic acid fragment).
[0027] In some embodiments, the nucleic acid material can be or can include one or more target nucleic acid fragments. In some embodiments, the one or more target nucleic acid fragments each include a genomic sequence of interest from one or more locations in a genome. In some embodiments, the one or more target nucleic acid fragments include target sequences from a substantially known region within the nucleic acid material. In some embodiments, isolating the target nucleic acid fragments based on a substantially known length includes concentrating the target nucleic acid fragments by gel electrophoresis, gel purification, liquid chromatography, size exclusion purification, filtration, or SPRI bead purification.
[0028] According to various embodiments, some of the methods provided may be useful for sequencing any of a variety of suboptimal (e.g., damaged or degraded) samples of nucleic acid material, for example, in some embodiments, at least a portion of the nucleic acid material is damaged. In some embodiments, the damage is oxidation, alkylation, deamination, methylation, hydrolysis, hydroxylation, nicking, intrastrand crosslinks, interstrand crosslinks, blunt-end strand breaks, sticky-end double-strand breaks, phosphorylation, dephosphorylation, sumoylation, glycosylation, deglycosylation, putresinylation, carboxylation, halogenation, formylation, single-strand gaps, heat damage, desiccation damage, UV exposure damage, gamma radiation damage, X-ray damage, ionizing radiation damage, non-ionizing radiation damage, heavy particle radiation damage, nuclear decay damage, beta radiation damage, alpha radiation damage, neutron radiation damage, proton radiation damage, cosmic radiation damage, high pH damage, low pH damage, reactive oxidative species damage, free radical damage, peroxide damage, hypochlorite damage, tissue fixation damage such as formalin or formaldehyde damage, reactive iron damage, low ionic conditions damage, high ionic conditions damage, unbuffered damage caused by conditions, damage caused by nucleases, damage caused by environmental exposure, damage caused by fire, damage caused by mechanical stress, damage caused by enzymatic degradation, damage caused by microorganisms, damage caused by preparative mechanical shearing, damage caused by preparative enzymatic fragmentation, damage caused naturally in vivo, damage caused during nucleic acid extraction, damage caused during sequencing library preparation, damage introduced by polymerases, damage introduced during nucleic acid repair, damage caused during nucleic acid end processing, damage caused during nucleic acid ligation, damage caused during sequencing, damage caused by mechanical handling of DNA, damage caused during passage through a nanopore, damage caused as part of organismal aging, damage caused as a result of exposure of an individual to chemicals, damage caused by mutagens, damage caused by carcinogens, damage caused by clastogens, damage caused by in vivo inflammatory damage due to oxygen exposure, damage caused by one or more strand breaks, and combinations thereof.
[0029] It is contemplated that nucleic acid material may be derived from a variety of sources. For example, in some embodiments, nucleic acid material (e.g., comprising one or more double-stranded nucleic acid molecules) is provided from a sample of a human subject, an animal, a plant, a fungus, a virus, a bacterium, a protozoan, or any other living organism. In other embodiments, the sample comprises at least partially artificially synthesized nucleic acid material. In some embodiments, the sample is body tissue, a biopsy, a skin sample, blood, serum, plasma, sweat, saliva, cerebrospinal fluid, mucus, uterine washing, vaginal swab, pap smear, nasal swab, oral swab, tissue scraping, hair, fingerprint, urine, stool, vitreous fluid, peritoneal washing, saliva, bronchial washing, oral washing, pleural washing, gastric washing, gastric juice, bile, pancreatic duct washing, bile duct washing, common bile duct washing, gallbladder fluid, synovial fluid, infected wound, non-infectious wound, archaeological sample, forensic sample, water sample, tissue sample, food sample, bioreactor sample, plant sample, bacterial sample, protozoan sample, fungal sample, animal sample, viral sample, sample from multiple organisms, nail scraping The nucleic acid material may be or comprise a sample, semen, prostatic fluid, vaginal fluid, vaginal swab, fallopian tube washing, cell-free nucleic acid, intracellular nucleic acid, metagenomics sample, washing or swab of an implanted foreign body, nasal wash, intestinal fluid, epithelial scraping, epithelial wash, tissue biopsy, autopsy sample, autopsy sample, organ sample, human identification sample, non-human identification sample, artificially produced nucleic acid sample, synthetic gene sample, banked or archived nucleic acid sample, tumor tissue, fetal sample, organ transplant sample, microbial culture sample, nuclear DNA sample, mitochondrial DNA sample, chloroplast DNA sample, apicoplast DNA sample, organelle sample, and any combination thereof. In some embodiments, the nucleic acid material is derived from more than one source.
[0030] As described herein, in some embodiments, it is advantageous to treat nucleic acid material to improve the efficiency, accuracy, and / or speed of the sequencing process. In some embodiments, the nucleic acid material comprises nucleic acid molecules of substantially uniform length and / or substantially known length. In some embodiments, the substantially uniform length and / or substantially known length is from about 1 to about 1,000,000 bases. For example, in some embodiments, the substantially uniform length and / or substantially known length can be at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 50, 60, 70, 80, 90, 100, 120, 150, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1200, 1500, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10,000, 15,000, 20,000, 30,000, 40,000, or 50,000 bases in length. In some embodiments, the substantially uniform and / or substantially known length can be at most 60,000, 70,000, 80,000, 90,000, 100,000, 120,000, 150,000, 200,000, 300,000, 400,000, 500,000, 600,000, 700,000, 800,000, 900,000, or 1,000,000 bases. As a specific, non-limiting example, in some embodiments, the substantially uniform and / or substantially known length is from about 100 to about 500 bases. In some embodiments, the nucleic acid material has been cleaved with one or more targeting endonucleases into nucleic acid molecules of substantially uniform and / or substantially known length. In some embodiments, the targeting endonucleases include at least one modification.
[0031] In some embodiments, the nucleic acid material includes nucleic acid molecules having lengths within one or more substantially known size ranges, which may be from 1 to about 1,000,000 bases, from about 10 to about 10,000 bases, from about 100 to about 1,000 bases, from about 100 to about 600 bases, from about 100 to about 500 bases, or some combination thereof.
[0032] In some embodiments, the targeting endonuclease is one or includes a restriction endonuclease (i.e., a restriction enzyme) that cleaves DNA at or near its recognition site (e.g., EcoRI, BamHI, XbaI, HindIII, AluI, AvaII, BsaJI, BstNI, DsaV, Fnu4HI, HaeIII, MaeIII, N1aIV, NSiI, MspJI, FspEI, NaeI, Bsu36I, NotI, HinF1, Sau3AI, PvuII, SmaI, HgaI, AluI, EcoRV, etc.). Lists of several restriction endonucleases are available in both printed and computer-readable format and are provided by many commercial suppliers (e.g., New England Biolabs, Ipswich, MA). Those skilled in the art will understand that any restriction endonuclease can be used in accordance with various embodiments of the present technology. In other embodiments, the targeting endonuclease is or includes at least one of a ribonucleoprotein complex, such as, for example, a CRISPR-associated (Cas) enzyme / guide RNA complex (e.g., Cas9 or Cpf1) or a Cas9-like enzyme. In other embodiments, the targeting endonuclease is or includes a homing endonuclease, a zinc finger nuclease, a TALEN, and / or a meganuclease (e.g., megaTAL nuclease), an Argonaute nuclease, or a combination thereof. In some embodiments, the targeting endonuclease includes Cas9 or CPF1, or a derivative thereof. In some embodiments, more than one targeting endonuclease may be used (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, or more). In some embodiments, the targeting endonuclease can be used to cleave more than one potential target region of a nucleic acid material (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, or more). In some embodiments, when more than one target region of nucleic acid material is present, each target region may be the same (or substantially the same) length.In some embodiments, when more than one target region of the nucleic acid material is present, at least two of the target regions of known length are different in length (e.g., a first target region that is 100 bp long and a second target region that is 1,000 bp long).
[0033] In some embodiments, specific modifications are made to a portion (e.g., adapter sequence) of a sample of nucleic acid material. As a particular example, in some embodiments, amplifying the nucleic acid material in a first sample further comprises, after the separating step and before amplifying the first sample, destroying or disrupting some or all of the second adapter sequence found on the nucleic acid material. As a further example, in some embodiments, amplifying the nucleic acid material in a second sample further comprises, after the separating step and before amplifying the second sample, destroying or disrupting the first adapter sequence found on the nucleic acid material. In some embodiments, the disruption or disintegration may be or may include at least one of enzymatic digestion, inclusion of at least one replication inhibitory molecule, enzymatic cleavage, enzymatic cleavage of one strand, enzymatic cleavage of both strands, incorporation of a modified nucleic acid followed by enzymatic treatment resulting in cleavage or one or both strands, incorporation of a replication-blocking nucleotide, incorporation of a chain terminator, incorporation of a photocleavable linker, incorporation of uracil, incorporation of a ribose base, incorporation of an 8-oxo-guanine adduct, use of a restriction endonuclease, use of a ribonucleoprotein endonuclease (e.g., a Cas enzyme such as Cas9 or CPF1), or other programmable endonuclease (e.g., a homing endonuclease, a zinc finger nuclease, a TALEN, a meganuclease (e.g., a megaTAL nuclease), an Argonaute nuclease, etc.), and any combination thereof. In some embodiments, in addition to or as an alternative to primer site disruption or disruption, methods such as affinity pull-down, size selection, or other known techniques for removing and / or de-amplifying undesired nucleic acid material from the sample are contemplated.
[0034] In some embodiments, at least one amplification step comprises at least one primer and / or adapter sequence that is or comprises at least one non-canonical nucleotide. By way of further example, in some embodiments, at least one adapter sequence is or comprises at least one non-standard nucleotide. In some embodiments, the non-standard nucleotide is selected from uracil, methylated nucleotides, RNA nucleotides, ribose nucleotides, 8-oxo-guanine, biotinylated nucleotides, desthiobiotin nucleotides, thiol-modified nucleotides, acrydite-modified nucleotides iso-dC, iso-dG, 2'-O-methyl nucleotides, inosine nucleotide-locked nucleic acids, peptide nucleic acids, 5-methyl-dC, 5-bromodeoxyuridine, 2,6-diaminopurine, 2-aminopurine nucleotides, abasic nucleotides, 5-nitroindole nucleotides, adenylated nucleotides, azido nucleotides, digoxigenin nucleotides, I-linkers, 5'-hexynyl-modified nucleotides, 5-octadiynyl-dU, photocleavable spacers, non-photocleavable spacers, click chemistry-enabled modified nucleotides, fluorescent dyes, biotin, furan, BrdU, fluoro-dU, loto-dU, and any combination thereof.
[0035] According to some embodiments, any of a variety of analytical steps can be used to improve one or more of the accuracy, speed, and efficiency of the provided processes. For example, in some embodiments, sequencing each of the first nucleic acid product and the second nucleic acid product comprises comparing the sequences of multiple strands in the first nucleic acid product to determine a consensus sequence of the first strand, and comparing the sequences of multiple strands in the second nucleic acid product to determine a consensus sequence of the second strand. In some embodiments, comparing the sequence of the first nucleic acid product with the sequence of the second nucleic acid product comprises comparing the consensus sequence of the first strand with the consensus sequence of the second strand to provide an error-corrected consensus sequence.
[0036] It is contemplated that any of a variety of methods for amplifying nucleic acid material may be used in accordance with various embodiments. For example, in some embodiments, at least one amplifying step includes polymerase chain reaction (PCR), rolling circle amplification (RCA), multiple displacement amplification (MDA), isothermal amplification, polony amplification in emulsion, bridge amplification on a surface, on a bead, or within a hydrogel, and any combination thereof. In some embodiments, amplifying the nucleic acid material includes the use of single-stranded oligonucleotides at least partially complementary to a region of the genomic sequence of interest and single-stranded oligonucleotides at least partially complementary to a region of an adapter sequence. In some embodiments, amplifying the nucleic acid material includes the use of single-stranded oligonucleotides at least partially complementary to a region of a first adapter sequence and a second adapter sequence (e.g., at least partially complementary to an adapter sequence at the 5' and / or 3' end of each strand of the nucleic acid material).
[0037] One aspect provided by some embodiments is the ability to generate high-quality sequencing information from very small amounts of nucleic acid material. In some embodiments, the provided methods and compositions can be used with starting nucleic acid material in amounts up to about 1 picogram (pg), 10 pg, 100 pg, 1 nanogram (ng), 10 ng, 100 ng, 200 ng, 300 ng, 400 ng, 500 ng, 600 ng, 700 ng, 800 ng, 900 ng, or 1000 ng. In some embodiments, the provided methods and compositions can be used with input amounts of nucleic acid material up to 1 molecular copy or genome equivalent, 10 molecular copies or genome equivalent, 100 molecular copies or genome equivalent, 1,000 molecular copies or genome equivalent, 10,000 molecular copies or genome equivalent, 100,000 molecular copies or genome equivalent, or 1,000,000 molecular copies or genome equivalent; for example, in some embodiments, up to 1,000 ng of nucleic acid material is initially provided to a particular sequencing process. For example, in some embodiments, up to 100 ng of nucleic acid material is initially provided to a particular sequencing process. For example, in some embodiments, up to 10 ng of nucleic acid material is initially provided to a particular sequencing process. For example, in some embodiments, up to 1 ng of nucleic acid material is initially provided to a particular sequencing process. For example, in some embodiments, up to 100 pg of nucleic acid material is initially provided to a particular sequencing process. For example, in some embodiments, up to 1 pg of nucleic acid material is initially provided to a particular sequencing process.
[0038] As used in this application, the terms "about" and "approximately" are used equivalently. All citations herein to publications, patents, or patent applications are incorporated herein by reference in their entirety. As used in this application, numbers with or without about / approximately are meant to encompass normal variations recognized by one of ordinary skill in the relevant art.
[0039] In various embodiments, enrichment of nucleic acid material, including enrichment of nucleic acid material in regions of interest, is provided at a faster rate (e.g., in fewer steps), at a lower cost (e.g., using fewer reagents), and results in increased data, which is desirable. Various aspects of the present technology have many applications in both preclinical and clinical trials and diagnostics, as well as other applications.
[0040] Specific details of several embodiments of the present technology are described below with reference to Figures 1A-24. While many of the embodiments are described herein with reference to duplex sequencing, other sequencing modalities capable of generating error-corrected and / or other sequencing reads in addition to those described herein are within the scope of the present technology. Furthermore, investigation of other nucleic acids is contemplated to benefit from the nucleic acid enrichment methods and reagents described herein. Furthermore, other embodiments of the present technology can have different configurations, components, or procedures than those described herein. Thus, those skilled in the art will accordingly understand that the technology can have other embodiments with additional elements and that the technology can have other embodiments without some of the features shown and described below with reference to Figures 1A-24. [Brief explanation of the drawings]
[0041] Many aspects of the present disclosure can be better understood with reference to the following drawings. The components in the drawings are not necessarily to scale, emphasis instead being placed upon clearly illustrating the principles of the present disclosure.
[0042] [Figure 1] (A) A nucleic acid adapter molecule for use in some embodiments of the present technology and a double-stranded adapter-nucleic acid complex resulting from ligation of the adapter molecule to a double-stranded nucleic acid fragment according to embodiments of the present technology. (B and C) Schematic diagrams of various duplex sequencing method steps according to embodiments of the present technology. [Figure 2]1 is a graph plotting positive predictive value as a function of variant allele frequency in a population of molecules for next generation sequencing (NGS), single-stranded tag-based error correction, and double-stranded sequencing error correction, according to certain embodiments of the present disclosure. [Figure 3] 3A shows a series of graphs depicting CODIS genotypes and number of paired sequencing reads in the absence of error correction (FIG. 3A) and standard DS analysis of three different loci (FIG. 3B), according to an embodiment of the present disclosure. [Figure 4] FIG. 1 is a conceptual diagram of the steps of the SPLiT-DS method according to one embodiment of the present technology. [Figure 5] FIG. 1 is a schematic diagram of the steps of the SPLiT-DS method, showing the steps of generating a double-stranded consensus sequence, according to one embodiment of the present technology. [Figure 6] FIG. 1 is a conceptual diagram of various SPLiT-DS method steps according to an embodiment of the present technology. [Figure 7] FIG. 10 is a schematic diagram of further SPLiT-DS method steps according to an embodiment of the present technology. [Figure 8A] FIG. 1 is a schematic diagram of the steps of the SPLiT-DS method incorporating a double-stranded primer site disruption scheme according to an additional embodiment of the present technology. [Figure 8B] 8B is a conceptual diagram of an example of steps in the SPLiT-DS method shown in FIG. 8A in accordance with an embodiment of the present technology. [Figure 8C] 8B is a conceptual diagram of an embodiment of a step of the SPLiT-DS method following the step of the method shown in FIG. 8A, in accordance with an additional aspect of the present technology. [Figure 8D] FIG. 1 is a schematic diagram of the steps of the SPLiT-DS method incorporating a double-stranded primer site disruption scheme according to another embodiment of the present technology. [Figure 9A] FIG. 1 is a schematic diagram of various embodiments of steps of the SPLiT-DS method incorporating a single-stranded primer site destruction scheme according to further aspects of the present technology. [Figure 9B] FIG. 1 is a schematic diagram of various embodiments of steps of the SPLiT-DS method incorporating a single-stranded primer site destruction scheme according to further aspects of the present technology. [Figure 10] FIG. 1 is a schematic diagram of the steps of the SPLiT-DS method using multiple targeted primers to generate a double-stranded consensus sequence of a longer nucleic acid molecule, according to yet another embodiment of the present technology. [Figure 11] (A) A graph plotting the relationship between nucleic acid insert size and the resulting family size after amplification, according to one embodiment of the present technology. (B) A schematic diagram showing sequencing data generated for different nucleic acid insert sizes, according to aspects of the present technology. (C) A schematic diagram showing method steps for generating targeted fragments cleaved to specific sizes by CRISPR / Cas9 to generate sequencing information, according to one embodiment of the present technology. [Figure 12]Schematic diagram of the steps of the CRISPR-DS method according to one embodiment of the present technology. Figure 12A shows the results from CRISPR / Cas9 digestion of TP53, with seven fragments containing all TP53 coding exons excised by targeted cleavage using gRNA. Dark gray represents the reference strand, and light gray represents the anti-reference strand. Figure 12B shows size selection using 0.5x SPRI beads; uncut genomic DNA binds to the beads, allowing recovery of the excised fragments in solution. Figure 12C shows a schematic diagram of double-stranded DNA molecules fragmented and ligated with double-stranded DS adapters containing 10 bp of random complementary nucleotides and 3'-dT overhangs. Figure 12D shows a schematic diagram of error correction by DS. Reads from the same DNA strand are compared to form a single-stranded consensus sequence (SSCS). Both strands of the same starting DNA molecule are then compared to each other to create a double-stranded consensus sequence (DSCS), and mutations found in both SSCS reads are counted as true mutations in the DSCS reads. Figures 12E-F provide a schematic comparison of the steps of CRISPR-DS and standard DS methods according to certain embodiments of the present technology. Figure 12E is a comparison of the library preparation steps of CRISPR-DS and standard DS. Each box represents 1 hour. Figure 12F shows a schematic of fragments produced using sonication that are shorter or longer than optimal (corresponding to missing or redundant information, respectively) compared to fragment products by CRISPR-DS that are optimal, consistent length, and have full coverage of sequencing reads. [Figure 13] Figure 13 shows data obtained from the SPLiT-DS procedure according to one embodiment of the present technology. Figure 13A is a representative gel showing insert fragment size before sequencing. Figures 13B and 13C are graphs showing the number of sequencing reads for CODIS genotypes without error correction (Figure 13B) and after analysis by SPLiT-DS (Figure 13C). [Figure 14] 14A and 14B are graphs showing CODIS genotypes versus the number of sequencing reads in the absence of error correction (FIG. 14A) and after analysis by SPLiT-DS (FIG. 14B) for highly damaged DNA according to an embodiment of the present technology. [Figure 15] 15A and 15B are visual representations of SPLiT-DS sequencing data for KRAS exon 2 generated from 10 ng (FIG. 15A) and 20 ng (FIG. 15B) of cfDNA according to an embodiment of the present technology. [Figure 16] Figure 16A is a schematic diagram of fragment lengths produced by sonication and CRISPR / Cas9 fragmentation according to an embodiment of the present technology. Figures 16B and 16C are histogram graphs showing fragment insert sizes for samples prepared with standard DS and CRISPR-DS protocols according to an embodiment of the present technology. The X-axis represents the percentage difference from the optimal fragment size, e.g., the fragment size that matches the length of the sequencing read after adjusting for molecular barcodes and clipping. The vertical column indicates the range of fragment sizes that fall within 10% of the optimal size, with the optimal size designated by the vertical hash line. [Figure 17A] 1 shows a CRISPR / Cas9 scheme for targeted enrichment of the coding region of human TP53 according to an embodiment of the present technology. TP53 tumor protein; Homo sapiens; NC_000017.11 Chr.17, Ref. GRCh38.p2. Gray text represents the coding region, and exon names are shown in the right margin and boxed together when they are in the same fragment. Gray-highlighted text represents the Cas9 cleavage site, with the PAM sequence double-underlined. Single underlined text represents the biotinylated probe, with the probe name shown in the left margin. [Figure 17B] 1 shows a CRISPR / Cas9 scheme for targeted enrichment of the coding region of human TP53 according to an embodiment of the present technology. TP53 tumor protein; Homo sapiens; NC_000017.11 Chr.17, Ref. GRCh38.p2. Gray text represents the coding region, and exon names are shown in the right margin and boxed together when they are in the same fragment. Gray-highlighted text represents the Cas9 cleavage site, with the PAM sequence double-underlined. Single underlined text represents the biotinylated probe, with the probe name shown in the left margin. [Figure 17C]1 shows a CRISPR / Cas9 scheme for targeted enrichment of the coding region of human TP53 according to an embodiment of the present technology. TP53 tumor protein; Homo sapiens; NC_000017.11 Chr.17, Ref. GRCh38.p2. Gray text represents the coding region, and exon names are shown in the right margin and boxed together when they are in the same fragment. Gray-highlighted text represents the Cas9 cleavage site, with the PAM sequence double-underlined. Single underlined text represents the biotinylated probe, with the probe name shown in the left margin. [Figure 18] 18A is a bar graph showing the percent of on-target, unprocessed sequencing reads (raw reads) (including TP53), showing the recovery percentage calculated by the percentage of genomes in the input DNA that produce double-stranded consensus sequence reads (FIG. 18B), and showing the median double-stranded consensus sequence depth across all targeted regions of various input amounts of DNA processed using standard DS and CRISPR-DS according to embodiments of the present technology (FIG. 18C). [Figure 19] 1 is a bar graph showing target enrichment achieved by CRISPR-DS with one capture step compared to two capture steps in three different blood DNA samples, according to one embodiment of the present technology. [Figure 20] 20A shows the results of pre-enrichment of high molecular weight DNA with BluePippin on a pulsed-field gel, and a bar graph (FIG. 20B) showing a comparison of the percentage of on-target raw reads and duplex consensus sequence depth for the same DNA sequenced before and after BluePippin pre-enrichment according to an embodiment of the present technology. [Figure 21] 21A is a schematic diagram of a synthetic double-stranded DNA molecule (FIG. 21A) and a chart of predicted fragment lengths after CRISPR / Cas9 digestion (FIG. 21B), as well as a resulting TapeStation gel image of actual DNA fragment lengths after CRISPR / Cas9 digestion of the synthetic double-stranded DNA molecule (FIG. 21C), demonstrating successful cleavage using CRISPR / Cas9 digestion according to an embodiment of the present technology. [Figure 22] Figure 22A shows a graph plotting the relationship between nucleic acid insert size and resulting family size after amplification of TP53 using CRISPR-DS and standard DS protocols according to one embodiment of the present technology. Dots represent the original barcode DNA molecules. With CRISPR-DS, all DNA molecules (light dots) have a predetermined size and generate a similar number of PCR copies (indicated by several "dot-like" clusters of light dots). With standard DS (dark dots), sonication shears the DNA into various fragment lengths (dark dots, more widely distributed on the plot than the light dots). The plot shows a higher number of short fragments than long fragments. Figures 22B-E show data on TP53 obtained from steps of the CRISPR-DS and standard DS methods according to an embodiment of the present technology. Figure 22B is a representative gel showing insert fragment size after adapter ligation and before sequencing. Figures 22C and 22D are electropherograms showing the peaks of the resulting nucleic acid libraries generated by CRISPR-DS (Figure 22C) and standard DS (Figure 22D) before sequencing. Figure 22E shows double-stranded consensus sequence reads of TP53 generated by the CRISPR-DS and standard DS protocols using the Integrative Genomics Viewer. Figure 22B shows a TapeStation gel with ladder and samples from CRISPR-DS (A1) and standard DS (B1). Band sizes correspond to the adaptor-containing CRISPR / Cas9 cleavage fragments. Figure 22E shows clear boundaries corresponding to the CRISPR / Cas9 breakpoints and a uniform distribution of depth across both intra- and inter-fragment locations. Standard DS shows a peak pattern generated by random shearing and hybridization capture of fragments, as well as uneven inclusion. [Figure 23] FIG. 1 is a schematic diagram of the steps of CRISPR-DS data processing according to one embodiment of the present technology. [Figure 24]
[0033] Figure 24A shows a chart (Figure 24A) and a graph (Figure 24B) illustrating the results of quantifying the degree of target enrichment after CRISPR / Cas9 digestion followed by size selection, according to one embodiment of the present technology. Figure 24A shows the DNA samples and the enrichment achieved for each. Figure 24B shows the percentage of raw reads that were "on target" compared to the amount of input DNA. DETAILED DESCRIPTION OF THE INVENTION
[0043] definition In order that this disclosure may be more readily understood, certain terms are first defined below. Additional definitions for these and other terms are found throughout the specification.
[0044] In this application, unless the context makes clear otherwise, the term "a" may be understood to mean "at least one." As used in this application, the term "or" may be understood to mean "and / or." As used in this application, the terms "comprising" and "including" may be understood to encompass the itemized component or step, whether presented by itself or with one or more additional components or steps. When ranges are provided herein, the endpoints are included. As used in this application, the term "comprise," and variations of the term such as "comprising" and "comprises," are not intended to exclude other additives, components, integers, or steps.
[0045] About: The term "about," when used herein with respect to a value, refers to a value that, in context, is similar to the referenced value. Generally, a person of ordinary skill in the art familiar with the context will understand the degree of relevant variation encompassed by "about" in that context. For example, in some embodiments, the term "about" can encompass a range of values within 25%, 20%, 19%, 18%, 17%, 16%, 15%, 14%, 13%, 12%, 11%, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, 1%, or less of the referenced value.
[0046] Analog: As used herein, the term "analog" refers to a substance that shares one or more particular structural features, elements, components, or moieties with a reference substance. Typically, an "analog" exhibits significant structural similarity with the reference substance, e.g., sharing a core or consensus structure, but also differs in certain individual respects. In some embodiments, an analog is a substance that can be produced from a reference substance, e.g., by chemical manipulation of the reference substance. In some embodiments, an analog is a substance that can be produced through the performance of a synthetic process that is substantially similar (e.g., shares multiple steps) to that which produces the reference substance. In some embodiments, an analog is produced, or can be produced, by the performance of a synthetic process that is different from that used to produce the reference substance.
[0047] Biological sample: As used herein, the term "biological sample" or "sample" generally refers to a sample obtained or derived from a biological source of interest (e.g., a tissue or organism or cell culture) as described herein. In some embodiments, the source of interest includes an organism, such as an animal or a human. In other embodiments, the source of interest includes a microorganism, such as a bacterium, a virus, a protozoan, or a fungus. In further embodiments, the source of interest may be a synthetic tissue, organism, cell culture, nucleic acid, or other material. In further embodiments, the source of interest may be a plant-based organism. In yet other embodiments, the sample may be an environmental sample, such as a water sample, soil sample, archaeological sample, or other sample collected from a non-living source. In other embodiments, the sample may be a multi-organism sample (e.g., a mixed biological sample). In some embodiments, the biological sample is a biological tissue or a biological fluid. In some embodiments, the biological sample can be or include bone marrow, blood, blood cells, ascites, fine needle biopsy sample, cell-containing body fluid, suspended nucleic acid, sputum, saliva, urine, cerebrospinal fluid, ascites, pleural fluid, feces, lymphatic fluid, gynecological fluid, skin swab, vaginal swab, Pap smear, oral swab, nasal swab, washings or lavage fluids such as ductal washings or bronchoalveolar lavage, vaginal fluid, aspirate, scraping, bone marrow sample, tissue biopsy sample, fetal tissue or fluid, excision sample, feces, other body fluids, secretions, and / or excretions, and / or cells therefrom, etc. In some embodiments, the biological sample is or includes cells obtained from an individual. In some embodiments, the obtained cells are or include cells of the individual from whom the sample was obtained. In certain embodiments, the biological sample is a liquid biopsy obtained from a subject. In some embodiments, the sample is a "primary sample" obtained directly from the source of interest by appropriate means. For example, in some embodiments, the primary biological sample is obtained by a method selected from the group consisting of biopsy (e.g., fine needle aspiration or tissue biopsy), surgery, collection of bodily fluids (e.g., blood, lymph, feces), and the like.In some embodiments, as will be clear from the context, the term "sample" refers to a preparation obtained by processing a primary sample (e.g., by removing one or more components and / or adding one or more agents), e.g., filtering using a semipermeable membrane. Such a "processed sample" can include, for example, nucleic acids or proteins extracted from the sample or obtained by subjecting the primary sample to techniques such as amplification or reverse transcription of mRNA, isolation and / or purification of specific components, etc.
[0048] Determining: Many methodologies described herein include a "determining" step. Those skilled in the art reading this specification will understand that such "determining" can be utilized or accomplished by utilizing any of a variety of techniques available to those skilled in the art, including, for example, the specific techniques explicitly mentioned herein. In some embodiments, determining involves physical manipulation of the sample. In some embodiments, determining involves reviewing and / or manipulating data or information, such as utilizing a computer or other processing unit adapted to perform the relevant analysis. In some embodiments, determining involves receiving relevant information and / or substances from a source. In some embodiments, determining involves comparing one or more characteristics of the sample or entity to a comparable reference.
[0049] Expression: As used herein, "expression" of a nucleic acid sequence refers to one or more of the following events: (1) production of an RNA template from a DNA sequence (e.g., by transcription), (2) processing of the RNA transcript (e.g., by splicing, editing, 5' capping, and / or 3' end formation), (3) translation of the RNA into a polypeptide or protein, and / or (4) post-translational modification of the polypeptide or protein.
[0050] gRNA: As used herein, "gRNA" or "guide RNA" refers to a short RNA molecule that contains a scaffold sequence suitable for targeting endonuclease (e.g., a Cas enzyme such as Cas9 or Cpfl, or another ribonucleoprotein with similar properties) binding to a substantially target-specific sequence that facilitates cleavage of a specific region of DNA or RNA.
[0051] Nucleic Acid: As used herein, in its broadest sense, refers to any compound and / or substance that is or can be incorporated into an oligonucleotide chain. In some embodiments, nucleic acids are compounds and / or substances that are or can be incorporated into an oligonucleotide chain via a phosphodiester bond. As is clear from the context, in some embodiments, "nucleic acid" refers to individual nucleic acid residues (e.g., nucleotides and / or nucleosides), and in some embodiments, "nucleic acid" refers to an oligonucleotide chain comprising individual nucleic acid residues. In some embodiments, "nucleic acid" is or comprises RNA, and in some embodiments, "nucleic acid" is or comprises DNA. In some embodiments, nucleic acids are, comprise, or consist of one or more naturally occurring nucleic acid residues. In some embodiments, nucleic acids are, comprise, or consist of one or more nucleic acid analogs. In some embodiments, nucleic acid analogs differ from nucleic acids in that they do not utilize a phosphodiester backbone. For example, in some embodiments, nucleic acids are, comprise, or consist of one or more "peptide nucleic acids," which are known in the art and have peptide bonds instead of phosphodiester bonds in the backbone and are considered within the scope of the present technology. Alternatively or additionally, in some embodiments, the nucleic acid has one or more phosphorothioate and / or 5'-N-phosphoramidite linkages rather than phosphodiester linkages. In some embodiments, the nucleic acid is, comprises, or consists of one or more naturally occurring nucleosides (e.g., adenosine, thymidine, guanosine, cytidine, uridine, deoxyadenosine, deoxythymidine, deoxyguanosine, and deoxycytidine).In some embodiments, the nucleic acid is, comprises, or consists of one or more nucleoside analogs (e.g., 2-aminoadenosine, 2-thiothymidine, inosine, pyrrolo-pyrimidine, 3-methyladenosine, 5-methylcytidine, C-5 propynyl-cytidine, C-5 propynyl-uridine, 2-aminoadenosine, C5-bromouridine, C5-fluorouridine, C5-iodouridine, C5-propynyl-uridine, C5-propynyl-cytidine, C5-methylcytidine, 2-aminoadenosine, 7-deazaadenosine, 7-deazaguanosine, 8-oxoadenosine, 8-oxoguanosine, O(6)-methylguanine, 2-thiocytidine, methylated bases, inserted bases, and combinations thereof). In some embodiments, the nucleic acid comprises one or more modified sugars (e.g., 2'-fluororibose, ribose, 2'-deoxyribose, arabinose, and hexose) compared to those of naturally occurring nucleic acids. In some embodiments, the nucleic acid has a nucleotide sequence that encodes a functional gene product such as RNA or a protein. In some embodiments, the nucleic acid comprises one or more introns. In some embodiments, the nucleic acid is prepared by one or more of isolation from a natural source, enzymatic synthesis by polymerization based on a complementary template (in vivo or in vitro), replication in a recombinant cell or system, and chemical synthesis. In some embodiments, the nucleic acid is at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 225, 250, 275, 300, 325, 350, 375, 400, 425, 450, 475, 500, 600, 700, 800, 900, 1000, 1500, 2000, 2500, 3000, 3500, 4000, 4500, 5000, or more residues in length. In some embodiments, the nucleic acid is partially or completely single-stranded, and in some embodiments, the nucleic acid is partially or completely double-stranded.In some embodiments, the nucleic acid has a nucleotide sequence that includes at least one element that encodes or is the complement of a sequence that encodes a polypeptide. In some embodiments, the nucleic acid has enzymatic activity. In some embodiments, the nucleic acid performs a mechanical function, for example, in a ribonucleoprotein complex or transfer RNA.
[0052] Reference: As used herein, describes a standard or control against which a comparison is made. For example, in some embodiments, an agent, animal, individual, population, sample, sequence, or value of interest is compared to a reference or control agent, animal, individual, population, sample, sequence, or value. In some embodiments, the reference or control is tested and / or determined substantially simultaneously with the test or determination of interest. In some embodiments, the reference or control is a historical reference or control, optionally embodied in a specific medium. Typically, as will be understood by those of skill in the art, a reference or control is determined or characterized under conditions or circumstances comparable to those under evaluation. Those of skill in the art will recognize when sufficient similarity exists to justify reliance on and / or comparison to a particular possible reference or control.
[0053] Single Molecular Identifier (SMI): As used herein, the term "single molecular identifier" or "SMI" (which may also be referred to as a "tag," "barcode," "molecular barcode," "unique molecular identifier," or "UMI," among other names) refers to any substance (e.g., a nucleotide sequence, a feature of a nucleic acid molecule) that can distinguish an individual molecule within a large, heterogeneous population of molecules. In some embodiments, an SMI can be or include an exogenously applied SMI. In some embodiments, an exogenously applied SMI can be or include a degenerate or semi-degenerate sequence. In some embodiments, a substantially degenerate SMI may be known as a random unique molecular identifier (R-UMI). In some embodiments, an SMI may include a code (e.g., a nucleic acid sequence) from within a pool of known codes. In some embodiments, a predefined SMI code is known as a defined unique molecular identifier (D-UMI). In some embodiments, an SMI can be or include an endogenous SMI. In some embodiments, endogenous SMIs may be or include information related to specific shear points of a target sequence or features related to the termini of individual molecules comprising the target sequence. In some embodiments, SMIs may be associated with sequence changes in nucleic acid molecules caused by random or semi-random damage, chemical modifications, enzymatic modifications, or other modifications to nucleic acid molecules. In some embodiments, the modification may be deamination of methylcytosine. In some embodiments, the modification may involve the site of a nucleic acid nick. In some embodiments, SMIs may include both exogenous and endogenous elements. In some embodiments, SMIs may include physically adjacent SMI elements. In some embodiments, SMI elements may be spatially distinct within a molecule. In some embodiments, SMIs may be non-nucleic acid. In some embodiments, SMIs may include two or more different types of SMI information. Various embodiments of SMIs are further disclosed in International Patent Publication No. WO 2017 / 100441, the entire contents of which are incorporated herein by reference.
[0054] Strand Defining Element (SDE): As used herein, the term "strand defining element" or "SDE" refers to any substance that allows for the identification of a particular strand of double-stranded nucleic acid material and therefore its distinction from other / complementary strands (e.g., any substance that renders the respective amplification products of two single-stranded nucleic acids resulting from a target double-stranded nucleic acid substantially distinguishable from one another after sequencing or other nucleic acid interrogation). In some embodiments, an SDE can be or include one or more segments of substantially non-complementary sequence within an adapter sequence. In certain embodiments, the segment of substantially non-complementary sequence within an adapter sequence can be provided by an adapter molecule comprising a Y-shape or "loop" shape. In other embodiments, the segment of substantially non-complementary sequence within an adapter sequence can form an unpaired "bubble" in the middle of adjacent complementary sequences within the adapter sequence. In other embodiments, an SDE can include a nucleic acid modification. In some embodiments, an SDE can include physical separation of paired strands into physically separated reaction compartments. In some embodiments, an SDE can include a chemical modification. In some embodiments, an SDE can include a modified nucleic acid. In some embodiments, SDEs may involve sequence changes in nucleic acid molecules caused by random or semi-random damage, chemical modifications, enzymatic modifications, or other modifications to nucleic acid molecules. In some embodiments, the modification may be deamination of methylcytosine. In some embodiments, the modification may involve the site of a nucleic acid nick. Various embodiments of SDEs are further disclosed in International Patent Publication No. WO2017 / 100441, the entire contents of which are incorporated herein by reference.
[0055] Subject: As used herein, the term "subject" refers to an organism, typically a mammal (e.g., a human in some embodiments, including prenatal human forms). In some embodiments, the subject is afflicted with an associated disease, disorder, or condition. In some embodiments, the subject is susceptible to a disease, disorder, or condition. In some embodiments, the subject exhibits one or more symptoms or characteristics of a disease, disorder, or condition. In some embodiments, the subject does not exhibit symptoms or characteristics of a disease, disorder, or condition. In some embodiments, a subject has one or more characteristics that characterize a susceptibility to or risk for a disease, disorder, or condition. In some embodiments, the subject is a patient. In some embodiments, the subject is an individual for whom and / or to whom diagnosis and / or treatment is being administered.
[0056] Substantially: As used herein, the term "substantially" refers to the qualitative condition of exhibiting the entire or nearly entire extent or degree of a desired characteristic or property. Those skilled in the art of biology will understand that biological and chemical phenomena, if they exist at all, rarely go to completion and / or rarely proceed perfectly, or rarely achieve or avoid absolute results. Thus, the term "substantially" is used herein to capture the potential lack of completeness inherent in many biological and chemical phenomena. <Detailed explanation>
[0057] Selected embodiments of adapters and reagents associated with duplex sequencing Duplex sequencing (DS) is a method for generating error-corrected DNA sequences from double-stranded nucleic acid molecules, originally described in International Patent Publication No. WO 2013 / 142389 and U.S. Patent No. 9,752,188, both of which are incorporated by reference in their entireties. As shown in Figures 1A-1C, and in certain embodiments of this technology, DS can be used to independently sequence both strands of an individual DNA molecule, such that derivative sequence reads are recognizable as originating from the same duplex nucleic acid parent molecule during MPS but are distinguishable from each other as distinct entities after sequencing. The sequence reads obtained from each strand are compared with the goal of obtaining an error-corrected sequence of the original double-stranded nucleic acid molecule, known as the duplex consensus sequence (DCS). The DS process can confirm whether one or both strands of the original double-stranded nucleic acid molecule are represented in the generated sequence data used to form the DCS.
[0058] In certain embodiments, methods of incorporating a DS may include ligating one or more sequencing adaptors to a target double-stranded nucleic acid molecule comprising a first strand target nucleic acid sequence and a second strand target nucleic acid sequence to produce a double-stranded target nucleic acid complex (e.g., Figure 1A).
[0059] In various embodiments, the resulting target nucleic acid complex can contain at least one SMI sequence, which may involve exogenously applied degenerate or semi-degenerate sequences, endogenous information related to the specific shear point of the target double-stranded nucleic acid molecule, or a combination thereof. The SMI can render the target nucleic acid molecule substantially distinguishable from multiple other molecules in the population to be sequenced. The substantially distinguishable characteristics of the SMI element can be carried independently by each of the single strands forming the double-stranded nucleic acid molecule, such that, after sequencing, the derivative amplification products of each strand can be recognized as originating from the same original, substantially unique double-stranded nucleic acid molecule. In other embodiments, the SMI can contain additional information and / or be used in other ways where molecular recognition functions are useful, such as those described in the publications mentioned above. In another embodiment, the SMI element can be incorporated after adapter ligation. In some embodiments, the SMI is essentially double-stranded. In other embodiments, it is essentially single-stranded. In other embodiments, it is essentially a combination of single-stranded and double-stranded.
[0060] In some embodiments, each double-stranded target nucleic acid sequence complex can further include an element (e.g., an SDE) that results in two single-stranded nucleic acid amplification products that form a target double-stranded nucleic acid molecule that is substantially distinguishable from one another after sequencing. In one embodiment, the SDE can include an asymmetric primer site contained within the sequencing adaptor, or in other arrangements, sequence asymmetry can be introduced into the adaptor molecule but not within the primer sequence, such that after amplification and sequencing, at least a portion of the nucleotide sequences of the first strand target nucleic acid sequence complex and the second strand of the target nucleic acid sequence complex differ from one another. In other embodiments, the SDE can include other biochemical asymmetries between the two strands that differ from the standard nucleotide sequences A, T, C, G, or U, but that are translated into at least one standard nucleotide sequence difference in the two amplified and sequenced molecules. In yet another embodiment, the SDE can be a means of physically separating the two strands prior to amplification, such that derivative amplification products from the first strand target nucleic acid sequence and the second strand target nucleic acid sequence are maintained in substantial physical separation to maintain their distinction. Other such configurations or methodologies that provide an SDE function that allows for differentiation between the first and second strands as described in the publications mentioned above, or other methods that serve the described functional purpose, may be utilized.
[0061] After generating a double-stranded target nucleic acid complex containing at least one SMI and at least one SDE, or if one or both of these elements are subsequently introduced, the complex can be subjected to DNA amplification, such as PCR or any other biochemical method of DNA amplification (e.g., rolling circle amplification, multiple displacement amplification, isothermal amplification, bridge amplification, or surface-bound amplification), to produce one or more copies of the first-strand target nucleic acid sequence and one or more copies of the second-strand target nucleic acid sequence (e.g., Figure 1B). One or more amplified copies of the first-strand target nucleic acid molecule and one or more amplified copies of the second target nucleic acid molecule can then be subjected to DNA sequencing, preferably using a "next-generation" massively parallel DNA sequencing platform (e.g., Figure 1B).
[0062] Sequence reads generated from either the first strand target nucleic acid molecule or the second strand target nucleic acid molecule derived from the original double-stranded target nucleic acid molecule are identified based on the shared substantially unique SMI and are distinguished from the opposite strand target nucleic acid molecule by SDE. In some embodiments, the SMI can be a sequence based on a mathematically based error correction code (e.g., a Hamming code), thereby allowing for certain amplification errors, sequencing errors, or SMI synthesis errors related to the sequence of the SMI sequence on the complementary strand of the original duplex (e.g., a double-stranded nucleic acid molecule). For example, in a double-stranded exogenous SMI where the SMI comprises 15 base pairs of a fully degenerate sequence of standard DNA bases, there are an estimated 4^15 = 1,073,741,824 SMI variants in the population of fully degenerate SMIs. If two SMIs are recovered from a population of 10,000 sampled SMIs by reading sequence data that differ by only one nucleotide in the SMI sequence, the probability of this occurring by random chance can be mathematically calculated, and it can be determined whether the single base pair difference likely reflects one of the types of errors described above, thereby determining that the SMI sequences are actually derived from the same original duplex molecule. In some embodiments where the SMIs are, at least in part, exogenously applied sequences, and the sequence variants are not completely degenerate from each other and are at least in part known sequences, the identity of the known sequences can be designed in some embodiments so that one or more errors of the types described above do not convert the identity of one known SMI sequence to that of another SMI sequence, reducing the likelihood that one SMI will be mistaken for another SMI. In some embodiments, this SMI design strategy involves a Hamming code approach or a derivative thereof. Once identified, one or more sequence reads generated from the first strand target nucleic acid molecule are compared with one or more sequence reads generated from the second strand target nucleic acid molecule to produce an error-corrected target nucleic acid molecule sequence (e.g., Figure 1C). For example, nucleotide positions where the bases of the target nucleic acid sequence in both the first and second strands match are considered to be true sequences, while nucleotide positions that do not match between the two strands are recognized as potential sites of technical error that can be ignored.Thus, an error-corrected sequence of the original double-stranded target nucleic acid molecule can be produced (shown in Figure 1C).
[0063] Alternatively, in some embodiments, a site of sequence mismatch between the two strands can be recognized as a potential site of mismatch originating from the biology of the original double-stranded target nucleic acid molecule. Alternatively, in some embodiments, a site of sequence mismatch between the two strands can be recognized as a potential site of mismatch originating from DNA synthesis of the original double-stranded target nucleic acid molecule. Alternatively, in some embodiments, a site of sequence mismatch between the two strands can be recognized as a potential site where damaged or modified nucleotide bases are present in one or both strands and converted into mismatches by an enzymatic process (e.g., DNA polymerase, DNA glycosylase, or another nucleic acid-modifying enzyme, or chemical process). In some embodiments, this latter finding can be used to infer the presence of nucleic acid damage or nucleotide modification prior to enzymatic or chemical processing.
[0064] Figure 2 is a graph plotting theoretical positive predictive value as a function of variant allele frequency in a molecular population for next-generation sequencing (NGS), single-stranded tag-based error correction, and double-stranded sequencing error correction, according to certain embodiments of the present disclosure. Referring to Figure 2, positive predictive value (e.g., the expected number of correct positive calls divided by the total number of positive calls) is plotted as a function of variant allele frequency in a molecular population for next-generation sequencing (NGS), single-stranded tag-based error correction, and DS error correction at a specified error rate. As can be seen from the overlap of the curves, when the frequency of detected variants is 2 per 10 or more, nearly all mutation calls are corrected using either method. However, the error rate of standard Illumina sequencing and single-stranded tag-based error correction result in a significant loss of positive predictive value at variant frequencies of approximately 1 per 100 and 1 per 1,000, respectively. The very low error rate provided by DS allows for reliable identification of variants below 1 per 100,000 (dotted line).
[0065] In some embodiments, and according to aspects of the present technology, sequence reads generated from the DS step discussed herein can be further filtered to remove sequencing reads from DNA-damaged molecules (e.g., during storage, transportation, or after tissue or blood extraction, library preparation, etc.). For example, DNA repair enzymes such as uracil DNA glycosylase (UDG), formamidopyrimidine DNA glycosylase (FPG), and 8-oxoguanine DNA glycosylase (OGG1) can be used to remove or correct DNA damage (e.g., in vitro DNA damage or in vivo damage). For example, these DNA repair enzymes are glycolases that remove damaged bases from DNA. For example, UDG removes uracil resulting from cytosine deamination (occurring through spontaneous hydrolysis of cytosine), and FPG removes 8-oxo-guanine (e.g., a common DNA damage caused by reactive oxygen species). FPG also has lyase activity, which creates a one-base gap at abasic sites. Such abasic sites generally result in PCR amplification failure, for example, because the polymerase fails to copy the template. Therefore, using such DNA damage repair / removal enzymes can effectively remove damaged DNA that does not harbor true mutations but would otherwise go undetected as errors after sequencing and double-strand sequence analysis. While rare base damage errors are often corrected by DNA splitting, complementary errors can theoretically occur at identical positions on both strands, so reducing damage that increases errors can reduce the likelihood of artifacts. Furthermore, during library preparation, certain fragments of DNA to be sequenced may be single-stranded from their source or processing steps (e.g., mechanical DNA shearing). These regions are typically converted to double-stranded DNA during the art-known "end repair" step, in which DNA polymerase and nucleoside substrates are added to the DNA sample and the 5' recessed end is extended.Mutant sites of DNA damage in the single-stranded portions of the DNA being copied (i.e., single-stranded 5' overhangs at one or both ends of the DNA duplex or internal single-stranded nicks or gaps) can lead to single-strand mutations, synthesis errors, or errors during fill-in reactions that can convert the nucleic acid damage site into double-stranded form, which can be mistaken in the final duplex consensus sequence as true mutations present in the original double-stranded nucleic acid molecule when in fact they are not. This scenario, known as "pseudoduplexing," can be reduced or prevented by using enzymes that destroy / repair such damage. In other embodiments, this occurrence can be reduced or eliminated by using strategies that destroy or prevent the single-stranded portions of the original duplex molecule (e.g., the use of specific enzymes that are not mechanical shearing or certain other enzymes that may leave nicks or gaps, as used to fragment the original double-stranded nucleic acid material). In other embodiments, the use of processes that eliminate the single-stranded portions of the original double-stranded nucleic acid (e.g., single-strand-specific nucleases such as S1 nuclease or mung bean nuclease) can be utilized for similar purposes.
[0066] In a further embodiment, the sequencing reads generated from the DS step discussed herein can be further filtered to eliminate spurious mutations by trimming the ends of reads most prone to false double-stranded artifacts. For example, DNA fragmentation can generate single-stranded portions at the ends of double-stranded molecules. These single-stranded portions can be filled during end repair (e.g., by Klenow or T4 polymerase). In some cases, the polymerase makes copy errors in these end-repaired regions, resulting in the generation of "pseudo-duplex molecules." These library preparation artifacts may erroneously appear to be true mutations when sequenced. These errors resulting from the end-repair mechanism can be eliminated or reduced from post-sequencing analysis by trimming the ends of sequencing reads to exclude any mutations that may have occurred in high-risk regions, thereby reducing the number of spurious mutations. In one embodiment, such trimming of sequencing reads can be achieved automatically (e.g., as a routine process step). In another embodiment, mutation frequencies can be assessed for fragment end regions, and if a threshold level of mutations is observed in the fragment end regions, trimming of the sequencing reads can be performed before generating double-stranded consensus sequence reads for the DNA fragments.
[0067] The advanced error correction provided by DS's strand comparison technology reduces sequencing errors of double-stranded nucleic acid molecules by several orders of magnitude compared to standard next-generation sequencing methods. This error reduction improves sequencing accuracy for almost all types of sequences, but is particularly suited to biochemically challenging sequences that are well known in the art for their error-prone nature. One non-limiting example of such a type of sequence is a homopolymer or other microsatellite / short tandem repeat. Another non-limiting example of an error-prone sequence that would benefit from DS error correction is a molecule damaged by, for example, heat, radiation, mechanical stress, or exposure to various chemicals that create error-prone chemical adducts during copying by one or more nucleotide polymerases. In a further embodiment, DS can also be used to accurately detect minor sequence variants among a population of double-stranded nucleic acid molecules. One non-limiting example of this application is the detection of a small number of DNA molecules derived from cancer among a large number of non-mutated molecules from non-cancerous tissues in a subject. Another non-limiting application of rare variant detection by DS is in forensic detection of DNA from one individual mixed at low concentration with DNA from another individual of a different genotype.
[0068] DS has been shown to be highly successful in removing artifacts from both mitochondrial and nuclear DNA amplification and sequencing. However, certain prior studies have focused on detecting somatic point mutations and small (e.g., less than 5 bp) insertions and deletions. DS holds great promise for the forensic community when addressing several challenges associated with forensic analysis (e.g., removal of PCR stutter, low levels of DNA, mixed samples, etc.). For example, referring to Figures 3A and 3B, DS demonstrated its ability to remove PCR stutter when compared to standard MPS. In this example, three representative CODIS loci from 10 ng of Promega 2800M standard reference material DNA were sequenced using conventional MPS (Figure 3A) and DS (Figure 3B) on an Illumina MiSeq platform with 300 bp paired-end reads, and the data were visualized using the STRait-Razor STR allele calling tool. Figure 3A shows three graphs depicting the number of sequencing reads for each of the three CODIS loci versus the number of reads in the absence of error correction (e.g., conventional MPS), indicating several stutter events (black arrows). In contrast, as shown in Figure 3B, DS eliminated the stutter events for the same three CODIS loci. Similar results are seen for all 13 original CODIS loci. Thus, various aspects of DS technology can overcome some of the limitations experienced by conventional methodologies for forensic analysis. Other aspects of forensic analysis, in addition to other uses of DS, may benefit from various aspects of conversion efficiency, or improvements in the percentage of input DNA that is converted into error-corrected sequence data. Forensic analysis may refer to applications related to human crime, natural disasters, mass casualty events, poaching, trafficking, or misuse of animals or other life forms, human or animal remains identification, assault identification, missing person identification, sexual assault identification, paleontological applications, and archaeological applications, among others.
[0069] Regarding the efficiency of the DS process, two types of efficiency are discussed in more detail: conversion efficiency and workflow efficiency. For the purposes of discussing DS efficiency, conversion efficiency can be defined as the percentage of unique nucleic acid molecules input into a sequencing library preparation reaction that produce at least one double-stranded consensus sequence read. Workflow efficiency can be related to the relative inefficiency of the time, relative number of steps, and / or financial cost of reagents / materials required to perform these steps to produce a duplex sequencing library and / or perform targeted enrichment of sequences of interest.
[0070] In some cases, limitations in either conversion efficiency or workflow efficiency, or both, may limit the usefulness of high-precision DS for some applications that are otherwise highly suitable. For example, low conversion efficiency can lead to a situation where the copy number of the target double-stranded nucleic acid is limited, potentially producing less sequence information than desired. Non-limiting examples of this concept include DNA from circulating tumor cells, cell-free DNA from tumors, or prenatal infants where DNA has been shed in bodily fluids such as plasma and mixed with excess DNA from other tissues. While DS is typically accurate enough to separate one mutant molecule from over 100,000 non-mutated molecules, for example, if only 10,000 molecules are available in a sample, and even with an ideal 100% efficiency in converting these into double-stranded consensus sequence reads, the lowest measurable mutation frequency would be 1 / (10,000 × 100%) = 1 / 10,000. For clinical diagnostics, it may be important to have maximum sensitivity for detecting low-level signals of cancer- or treatment-related mutations; therefore, relatively low conversion efficiency would be undesirable in this context. Similarly, in forensic applications, there is often very little DNA available for testing: when only nanogram or picogram quantities are recovered from crime scenes or natural disaster scenes, and DNA from multiple individuals is mixed, maximizing conversion efficiency is important to be able to detect the presence of DNA from all individuals in the mixture.
[0071] In some cases, workflow inefficiencies can be similarly challenging for certain nucleic acid testing applications. One non-limiting example of this is clinical microbiology testing. It is often desirable to rapidly detect the nature of one or more infectious organisms, for example, in microbial or polymicrobial bloodstream infections, where some organisms are resistant to specific antibiotics based on their unique genetic variations. However, the time required to culture the infectious organisms and empirically determine their antibiotic susceptibility is significantly longer than the time required to make a therapeutic decision about which antibiotic to use for treatment. DNA sequencing of DNA from blood (or other infected tissues or bodily fluids) can be even more rapid; for example, DS, among other high-precision sequencing methods, can very accurately detect therapeutically important minority variants in infected populations based on DNA signatures. Because workflow time to data generation can be critical for determining treatment options (e.g., as in the example used herein), applications that increase the speed at which data output is achieved would also be desirable.
[0072] Further disclosed herein are methods and compositions for the enrichment of targeted nucleic acid sequences, and the use of such enrichment for error-corrected nucleic acid sequencing applications that offer improvements in cost, conversion of sequenced molecules, and time efficiency in generating labeled molecules for targeted ultra-high precision sequencing.
[0073] SPLiT-DS In some embodiments, the provided methods provide a PCR-based targeted enrichment strategy that is compatible with the use of molecular barcodes for error correction. Figure 4 is a conceptual diagram of a sequencing enrichment strategy that utilizes steps of the Separation PCR of Linked Templates for Sequencing ("SPLiT-DS") method, according to one embodiment of the present technology. With reference to Figure 4, and in one embodiment, the SPLiT-DS approach can begin with labeling (e.g., tagging) fragmented double-stranded nucleic acid material (e.g., from a DNA sample) with molecular barcodes in a manner similar to that described above and with respect to standard DS library construction protocols (e.g., as shown in Figure 1B). In some embodiments, the double-stranded nucleic acid material may be fragmented (e.g., cell-free DNA, damaged DNA, etc.), while in other embodiments, various steps can include fragmenting the nucleic acid material using mechanical shearing, such as sonication, or other DNA cleavage methods as further described herein. Labeling aspects of the fragmented double-stranded nucleic acid material include end repair and 3'-dA-tailing, if required for a particular application, followed by ligation of a DS adapter containing an SMI to the double-stranded nucleic acid fragment (Figure 4, Step 1). In other embodiments, the SMI can be an endogenous or a combination of exogenous and endogenous sequences to uniquely associate information from both strands of the original nucleic acid molecule. After ligating the adapter molecule to the double-stranded nucleic acid material, the method can continue with amplification (e.g., PCR amplification, rolling circle amplification, multiple displacement amplification, isothermal amplification, bridge amplification, surface-bound amplification, etc.) (Figure 4, Step 2).
[0074] In certain embodiments, each strand of nucleic acid material can be amplified, for example, using primers specific to one or more adapter sequences, resulting in multiple copies of nucleic acid amplification products derived from each strand of the original double-stranded nucleic acid molecule, each amplification product retaining its originally associated SMI (Figure 4, step 2). After amplification and associated steps to remove reaction byproducts, the sample can be divided (preferably, but not necessarily, substantially equally) into two or more separate samples (e.g., into tubes, emulsion droplets, microchambers, isolated droplets on a surface, or other known containers collectively referred to as "tubes") (Figure 4, step 3). Alternatively, the amplification products can be divided in a manner that does not require them to be in solution, such as by binding to microbeads and then dividing the population of microbeads into two chambers, or by immobilizing the divided amplification products at two or more different physical locations on a surface. Either of these latter divided populations is functionally equivalent and is referred to herein as being in a separate "tube." In the example shown in FIG. 4, this step results in an average of half the copies of any given strand / barcode amplification product found in each tube. In other embodiments where the original sample is divided into more than two separate samples, such allocation of nucleic acid material results in a relatively equal reduction in the number of amplification products. Note that the random nature of how the amplification products are divided results in variance about this average. To account for this variance, a hypergeometric distribution (i.e., the probability of selecting k barcode copies without replacement) can be used as a model to determine the minimum number of amplification products (e.g., PCR copies) of an SMI (e.g., barcode) required to maximize the likelihood that each tube contains at least one copy from both strands. Without wishing to be bound by any particular theory, it is recommended that step 2 be performed using four or more PCR cycles (i.e., 2 4= 16 copies / barcode), there is an expected probability of greater than 99% that each barcode copy from each strand will appear at least once in each tube. In some embodiments, it may be preferable to divide the amplification products unevenly. If the nucleic acid material is divided into more than two tubes, additional amplification cycles can be used to generate additional copies to accommodate further divisions. After dividing the sample into two tubes, a primer specific to the adapter sequence and a primer specific to the target nucleic acid region of interest can be used to enrich for the target nucleic acid region (e.g., region, locus, etc. of interest) with multiplex PCR (Figure 4, step 3). In another embodiment, a linear amplification step can be added before the subsequent addition of a second primer, which allows for exponential amplification of the target region of interest.
[0075] In certain embodiments, multiplex target-specific PCR is performed so that the resulting PCR product in each tube is derived from only one of the two strands (e.g., the "top strand" or the "bottom strand"). As shown in Figure 4 (Step 3), in some embodiments, this is accomplished as follows: In the first tube (shown on the left), a primer at least partially complementary to "read 1" (e.g., Illumina P5) of the adapter sequence (Figure 4, Step 3, gray arrow) and a primer at least partially complementary to the nucleic acid region of interest and containing the "read 2" (i.e., Illumina P7, black arrow with gray tail) adapter sequence are used to specifically amplify (e.g., enrich) the "top strand" of the original nucleic acid molecule (Figure 4, Steps 3 and 4). In this first sample, the "bottom strand" does not amplify properly due to the nature of the SDE (e.g., in this case, the orientation of the unique adapter sequence relative to the target nucleic acid insert). Similarly, in a second tube (shown on the left), a primer (e.g., Illumina P5) at least partially complementary to the "read 2" of the adapter sequence (Figure 4, step 3, gray arrow) and a primer (i.e., Illumina P7, black arrow with gray tail) at least partially complementary to the nucleic acid region of interest and containing the "read 1" adapter sequence are used to specifically amplify (e.g., enrich) the "bottom strand" of the original nucleic acid molecule (Figure 4, steps 3 and 4). In this second sample, the "top strand" does not amplify properly. Following PCR or other amplification methods, multiple copies of the "top strand" are generated in the first tube, and multiple copies of the "bottom strand" are generated in the second tube. Because each of these resulting target-specific copies has both adapter sequences available at each end of the nucleic acid amplification product (e.g., Illumina P5 and Illumina P7 adapter sequences), these target-enriched products can be sequenced using standard MPS methods.
[0076] Figure 5 is a conceptual diagram of the steps of the SPLiT-DS method shown and discussed with respect to Figure 4, further illustrating the step of sequencing multiple copies of a PCR-enriched target region and generating a double consensus sequence according to an embodiment of the present technology. Following sequencing of multiple copies of the "top strand" from the first tube and multiple copies of the "bottom strand" from the second tube, the sequence data can be analyzed using a similar approach to DS, whereby sequence reads sharing the same molecular barcode (found in the first tube and the second tube, respectively) derived from the "top" or "bottom" strand of the original double-stranded target nucleic acid molecule are grouped separately. In some embodiments, the grouped sequence reads from the "top strand" are used to form a top-strand consensus sequence (e.g., a single-stranded consensus sequence (SSCS)), and the grouped sequence reads from the "bottom strand" are used to form a bottom-strand consensus sequence (e.g., a SSCS). Referring to Figure 5, the top and bottom SSCs can be compared to generate a double-stranded consensus sequence (DCS) with matching nucleotides between the two strands (e.g., a variant or mutation is considered true if it appears in sequencing reads from both strands) (see, e.g., Figure 1C).
[0077] As a specific example, in some embodiments, provided herein is a method for generating error-corrected sequence reads of double-stranded target nucleic acid material, comprising ligating the double-stranded target nucleic acid material to at least one adapter sequence to form an adapter-target nucleic acid material complex, wherein the at least one adapter sequence comprises: (a) a degenerate or semi-degenerate single molecule identifier (SMI) sequence that uniquely labels each molecule of the double-stranded target nucleic acid material; and (b) a first nucleotide adapter sequence that tags a first strand of the adapter-target nucleic acid material complex, and a second nucleotide adapter sequence that is at least partially non-complementary to the first nucleotide sequence that tags a second strand of the adapter-target nucleic acid material complex, such that each strand of the adapter-target nucleic acid material complex has a nucleotide sequence that is unambiguously identifiable with respect to its complementary strand. The method may then include amplifying each strand of the adaptor-target nucleic acid material complex to produce a plurality of first strand adaptor-target nucleic acid complex amplification products and a plurality of second strand adaptor-target nucleic acid complex amplification products, and separating the adaptor-target nucleic acid complex amplification products into a first sample and a second sample. The method may further include amplifying the first strand in the first sample using a first primer at least partially complementary to the first nucleotide adaptor sequence and a primer at least partially complementary to the target sequence of interest to provide a first nucleic acid product, and amplifying the second strand in the second sample using a second primer at least partially complementary to the second nucleotide adaptor sequence and a primer at least partially complementary to the target sequence of interest to provide a second nucleic acid product. The method may also include sequencing each of the first nucleic acid product and the second nucleic acid product to produce a plurality of first strand sequence reads and a plurality of second strand sequence reads, and confirming the presence of at least one first strand sequence read and at least one second strand sequence read.The method may further include generating error-corrected sequence reads of the double-stranded target nucleic acid material by comparing at least one first strand sequence read with at least one second strand sequence read and ignoring mismatched nucleotide positions, or alternatively, by removing compared first and second strand sequence reads having one or more nucleotide positions where the compared first and second strand sequence reads are non-complementary.
[0078] As a further specific example, in some embodiments, provided herein are methods of identifying DNA variants from a sample, comprising: ligating both strands of nucleic acid material (e.g., a double-stranded target DNA molecule) with at least one asymmetric adapter molecule to form an adapter-target nucleic acid material complex having a first nucleotide sequence associated with the top strand of the double-stranded target DNA molecule and a second nucleotide sequence that is at least partially non-complementary to the first nucleotide sequence and associated with the bottom strand of the double-stranded target DNA molecule; and amplifying each strand of the adapter-target nucleic acid material to obtain each strand generating a set of different, yet related, amplified adapter-target DNA products. The method may also include separating the adaptor-target DNA products into a first sample and a second sample; amplifying the top strand of the adaptor-target DNA product in the first sample using a primer specific to (e.g., at least partially complementary to) a first nucleotide sequence and a primer at least partially complementary to a target sequence of interest to result in a top strand adaptor-target nucleic acid complex amplification product; and amplifying the bottom strand in the second sample using a second primer specific to (e.g., at least partially complementary to) a second nucleotide sequence and a second primer to result in a bottom strand adaptor-target nucleic acid complex amplification product. The method may further include sequencing each of the top strand adaptor-target nucleic acid complex amplification products and the bottom strand adaptor-target nucleic acid complex amplification products, confirming the presence of at least one amplified sequence read from each strand of the adaptor-target DNA complex, and comparing the at least one amplified sequence read obtained from the top strand with the at least one amplified sequence read obtained from the bottom strand to form a consensus sequence read of the nucleic acid material (e.g., the double-stranded target DNA molecule) having only nucleotide bases that match in sequence on both strands of the nucleic acid material (e.g., the double-stranded target DNA molecule), such that variants occurring at specific positions in the consensus sequence read are identified as true DNA variants.
[0079] In some embodiments, provided herein are methods for generating error-correcting double-stranded consensus sequences from double-stranded nucleic acid material, the method comprising: tagging individual duplex DNA molecules with adapter molecules to form tagged DNA material, each adapter molecule comprising (a) a degenerate or semi-degenerate single molecule identifier (SMI) that uniquely labels the duplex DNA molecule, and (b) for each tagged DNA molecule, first and second non-complementary nucleotide adapter sequences that distinguish the original top strand from the original bottom strand of an individual DNA molecule within the tagged DNA material; and generating a set of replicas of the original top strand of the tagged DNA molecule and a set of replicas of the original bottom strand of the tagged DNA molecule to form amplified DNA material. The method may also include separating the amplified DNA material into a first sample and a second sample, generating additional copies of the original top strand in the first sample using a primer specific to a first nucleotide adapter sequence and a primer at least partially complementary to the target sequence of interest to produce a first nucleic acid product, and generating additional copies of the original bottom strand in the second sample using a primer specific to a second nucleotide adapter sequence (the same or different) and a primer at least partially complementary to the target sequence of interest to produce a second nucleic acid product. The method may further include generating a first single-stranded consensus sequence (SSCS) from the additional copies of the original top strand and a second single-stranded consensus sequence (SSCS) from the additional copies of the original bottom strand, comparing the first SSCS of the original top strand with the second SSCS of the original bottom strand, and generating an error-corrected double-stranded consensus sequence having only nucleotide bases where both the first SSCS of the original top strand and the second SSCS of the original bottom strand are complementary in sequence.
[0080] Single molecular identifier sequences (SMIs) According to various embodiments, the provided methods and compositions include one or more SMI sequences on each strand of nucleic acid material. The SMI can be carried independently by each single strand resulting from a double-stranded nucleic acid molecule, and after sequencing, the derivative amplification products of each strand can be recognized as originating from the same original, substantially unique double-stranded nucleic acid molecule. In some embodiments, the SMI can contain additional information and / or be used in other ways where molecular recognition functionality is useful, as will be understood by those skilled in the art. In some embodiments, the SMI element can be incorporated before, substantially simultaneously with, or after ligation of an adapter sequence to the nucleic acid material.
[0081] In some embodiments, the SMI sequence can include at least one degenerate or semi-degenerate nucleic acid. In other embodiments, the SMI sequence can be non-degenerate. In some embodiments, the SMI can be a sequence associated with or near the end of a fragment of a nucleic acid molecule (e.g., a randomly or semi-randomly sheared end of the ligated nucleic acid material). In some embodiments, the exogenous sequence can be considered in conjunction with a sequence corresponding to a randomly or semi-randomly sheared end of the ligated nucleic acid material (e.g., DNA) to obtain an SMI sequence that can distinguish single DNA molecules from each other. In some embodiments, the SMI sequence is part of an adapter sequence ligated to a double-stranded nucleic acid molecule. In certain embodiments, the adapter sequence containing the SMI sequence is double-stranded, and each strand of the double-stranded nucleic acid molecule contains an SMI after ligation to the adapter sequence. In another embodiment, the SMI sequence is single-stranded before or after ligation to the double-stranded nucleic acid molecule, and a complementary SMI sequence can be generated by extending the opposite strand with a DNA polymerase, resulting in a complementary double-stranded SMI sequence. In some embodiments, each SMI sequence includes from about 1 to about 30 nucleic acids (eg, 1, 2, 3, 4, 5, 8, 10, 12, 14, 16, 18, 20, or more degenerate or semi-degenerate nucleic acids).
[0082] In some embodiments, the SMI can ligate to one or both of the nucleic acid material and the adapter sequence, hi some embodiments, the SMI can ligate to at least one of a T overhang, an A overhang, a CG overhang, a dehydroxylated base, and a blunt end of the nucleic acid material.
[0083] In some embodiments, the sequence of the SMI can be considered (or designed accordingly) in conjunction with sequences corresponding to, for example, randomly or semi-randomly sheared ends of nucleic acid material (e.g., ligated nucleic acid material) to obtain an SMI sequence that can distinguish single nucleic acid molecules from each other.
[0084] In some embodiments, at least one SMI can be an endogenous SMI (e.g., an SMI associated with a shear point, e.g., using the shear point itself, or using a defined number of nucleotides in the nucleic acid material immediately adjacent to the shear point [e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10 nucleotides from the shear point]). In some embodiments, at least one SMI can be an exogenous SMI (e.g., an SMI comprising a sequence not found in the target nucleic acid material).
[0085] In some embodiments, the SMI may be or include an imaging moiety (e.g., a fluorescent or other optically detectable moiety). In some embodiments, such an SMI allows for detection and / or quantification without the need for an amplification step.
[0086] In some embodiments, the SMI element may comprise two or more distinct SMI elements located at different locations on the adaptor-target nucleic acid complex.
[0087] Various embodiments of the SMI are further disclosed in International Patent Publication No. WO2017 / 100441, which is incorporated herein by reference in its entirety.
[0088] Chain Definition Element (SDE): In some embodiments, each strand of the double-stranded nucleic acid material may further comprise an element that renders the amplification products of the two single-stranded nucleic acids forming the target double-stranded nucleic acid material substantially distinguishable from one another after sequencing. In some embodiments, the SDE may be or comprise an asymmetric primer site contained within the sequencing adaptor, or in other configurations, at least one position may be introduced into the adaptor sequence rather than within the primer sequence, such that at least one position in the nucleotide sequence of the second strand of the first-strand target nucleic acid sequence complex differs from one another after amplification and sequencing. In other embodiments, the SDE may differ from the standard nucleotide sequence A, T, C, G, or U, but may comprise another biochemical asymmetry between the two strands that translates into at least one standard nucleotide sequence difference in the two amplified and sequenced molecules. In yet other embodiments, the SDE may be or comprise a means for physically separating the two strands prior to amplification, such that derivative amplification products from the first-strand target nucleic acid sequence and the second-strand target nucleic acid sequence are maintained in substantial physical separation from one another to maintain distinction between the two derivative amplification products. Other such arrangements or methodologies for providing an SDE function that allows for differentiation between the first and second strands may be utilized.
[0089] In some embodiments, the SDE may form a loop (e.g., a hairpin loop). In some embodiments, the loop may include at least one endonuclease recognition site. In some embodiments, the target nucleic acid complex may include an endonuclease recognition site that promotes a cleavage event within the loop. In some embodiments, the loop may include a non-standard nucleotide sequence. In some embodiments, the included non-standard nucleotides may be recognizable by one or more enzymes that promote strand scission. In some embodiments, the included non-standard nucleotides may be targeted by one or more chemical processes that promote strand scission within the loop. In some embodiments, the loop may include a modified nucleic acid linker that may be targeted by one or more enzymatic, chemical, or physical processes that promote strand scission within the loop. In some embodiments, the modified linker is a photocleavable linker.
[0090] A variety of other molecular tools can function as SMIs and SDEs. In addition to shear points and DNA-based tags, single-molecule compartmentalization or other non-nucleic acid tagging methods that bring paired strands into physical proximity can serve strand-association functions. Similarly, asymmetric chemical labeling of adapter strands that allow for physical separation can serve as SDEs. A recently described modification of DS uses bisulfite conversion to convert naturally occurring strand asymmetry in the form of cytosine methylation into sequence differences that distinguish the two strands. While this incorporation limits the types of mutations that can be detected, the concept of exploiting natural asymmetry is noteworthy in the context of new sequencing technologies that can directly detect modified nucleotides. Various embodiments of SDEs are further disclosed in International Patent Publication No. WO 2017 / 100441, which is incorporated by reference in its entirety.
[0091] Adapters and adapter sequences Adapter molecules comprising SMIs (e.g., molecular barcodes), SDEs, primer sites, flow cell sequences, and / or other features in various configurations are contemplated for use in many embodiments disclosed herein. In some embodiments, the provided adapters can be or include one or more sequences (e.g., primer sites) that are complementary or at least partially complementary to PCR primers that have at least one of the following properties: 1) high target specificity, 2) amenable to multiplexing, and 3) exhibit robust, minimally biased amplification.
[0092] In some embodiments, adapter molecules can be "Y-shaped," "U-shaped," "hairpin"-shaped, have bubbles (e.g., a portion of the sequence that is non-complementary), or other features. In other embodiments, adapter molecules can include a "Y" shape, a "U" shape, a "hairpin" shape, or a bubble. Certain adapters can include modified or non-standard nucleotides, restriction sites, or other features for in vitro structural or functional manipulation. Adapter molecules can ligate to various nucleic acid materials with ends. For example, adapter molecules can be suitable for ligating to T-overhangs, A-overhangs, CG-overhangs, multi-nucleotide overhangs, dehydroxylated bases, and blunt ends of nucleic acid materials, and the end of the molecule is 5' of the target and is dephosphorylated or otherwise protected from conventional ligation. In other embodiments, adapter molecules can include a dephosphorylated or otherwise ligation-preventing modification on the 5' strand of the ligation site. In the latter two embodiments, such strategies can be useful to prevent dimerization of library fragments or adapter molecules.
[0093] An adapter sequence may refer to a single-stranded sequence, a double-stranded sequence, a complementary sequence, a non-complementary sequence, a partially complementary sequence, an asymmetric sequence, a primer binding sequence, a flow cell sequence, a ligation sequence, or other sequence provided by an adapter molecule. In certain embodiments, an adapter sequence may refer to a sequence used in amplification as a complement to an oligonucleotide.
[0094] In some embodiments, the provided methods and compositions include at least one adapter sequence (e.g., two adapter sequences, one for each of the 5' and 3' ends of the nucleic acid material). In some embodiments, the provided methods and compositions may include two or more (e.g., 3, 4, 5, 6, 7, 8, 9, 10 or more) adapter sequences. In some embodiments, at least two of the adapter sequences differ from each other (e.g., by sequence). In some embodiments, each adapter sequence differs from the other adapter sequences (e.g., by sequence). In some embodiments, at least one adapter sequence is at least partially non-complementary (e.g., non-complementary at at least one nucleotide) to at least a portion of at least one other adapter sequence.
[0095] In some embodiments, the adapter sequence comprises at least one non-standard nucleotide, which in some embodiments is an abasic site, uracil, tetrahydrofuran, 8-oxo-7,8-dihydro-2'-deoxyadenosine (8-oxo-A), 8-oxo-7,8-dihydro-2'-deoxyguanosine (8-oxo-G), deoxyinosine, 5'nitroindole, 5-hydroxydiethyl-2'-deoxycytidine, isocytosine, 5'-methylisocytosine, or isoguanosine, a methylated nucleotide, an RNA nucleotide, a ribose nucleotide, an 8-oxo-guanine, a photocleavable linker, a biotinylated nucleotide, a desthiobiotin nucleotide, a thiol-modified nucleotide, an acrydite-modified nucleotide, iso-dC, or iso-dC. The spacer is selected from dG, 2'-O-methyl nucleotides, inosine nucleotide-locked nucleic acids, peptide nucleic acids, 5-methyl dC, 5-bromodeoxyuridine, 2,6-diaminopurine, 2-aminopurine nucleotides, abasic nucleotides, 5-nitroindole nucleotides, adenylated nucleotides, azido nucleotides, digoxigenin nucleotides, I linkers, 5' hexynyl-modified nucleotides, 5-octadiynyl dU, photocleavable spacers, non-photocleavable spacers, click chemistry-compatible modified nucleotides, and any combination thereof.
[0096] In some embodiments, the adapter sequence includes a portion having magnetic properties (i.e., a magnetic portion). In some embodiments, the magnetic property is paramagnetic. In some embodiments, in which the adapter sequence includes a magnetic portion (e.g., nucleic acid material ligated to an adapter sequence including a magnetic portion), when a magnetic field is applied, the adapter sequence including the magnetic portion is substantially separated from the adapter portion that does not include the magnetic portion (e.g., nucleic acid material ligated to an adapter sequence that does not include the magnetic portion).
[0097] In some embodiments, at least one adapter sequence is located 5' to the SMI. In some embodiments, at least one adapter sequence is located 3' to the SMI.
[0098] In some embodiments, the adaptor sequence can be linked to at least one of the SMI and the nucleic acid material via one or more linker domains. In some embodiments, the linker domain can be composed of nucleotides. In some embodiments, the linker domain can include at least one modified nucleotide or non-nucleotide molecule (e.g., as described elsewhere in this disclosure). In some embodiments, the linker domain can be or include a loop.
[0099] In some embodiments, the adapter sequences at either or both ends of each strand of the double-stranded nucleic acid material may further comprise one or more elements that provide an SDE, hi some embodiments, the SDE may be or may include an asymmetric primer site contained within the adapter sequence.
[0100] In some embodiments, an adapter sequence may be or include at least one SDE and at least one ligation domain (i.e., at least one domain amendable to ligase activity, e.g., a domain suitable for ligation to nucleic acid material by ligase activity). In some embodiments, an adapter sequence may be or include, from 5' to 3', a primer binding site, an SDE, and a ligation domain.
[0101] Various methods for synthesizing DS adaptors have been previously described, for example, in U.S. Pat. No. 9,752,188 and International Patent Publication No. WO2017 / 100441, both of which are incorporated by reference herein in their entireties.
[0102] Primer In some embodiments, one or more PCR primers with at least one of the following characteristics are contemplated for use in various embodiments according to aspects of the present technology: 1) high target specificity, 2) multiplexable, and 3) robust and minimally biased amplification. Many prior studies and commercial products have designed primer mixtures that meet some of these criteria for traditional PCR-CE. However, it should be noted that these primer mixtures are not necessarily optimal for use in MPS. Indeed, developing highly multiplexed primer mixtures is a challenging and time-consuming process. Conveniently, both Illumina and Promega have recently developed multiplex-compatible primer mixtures for the Illumina platform, demonstrating robust and efficient amplification of a variety of standard and non-standard STR and SNP loci. Because these kits use PCR to amplify target regions prior to sequencing, the 5' end of each read in paired-end sequence data corresponds to the 5' end of the PCR primer used to amplify the DNA. In some embodiments, the provided methods and compositions include primers designed to ensure uniform amplification, which may require varying reaction concentrations, melting temperatures, secondary structures, and minimizing intra- and inter-primer interactions. Numerous techniques for highly multiplexed primer optimization for MPS applications have been described, particularly these techniques, often known as ampliseq methods, as are well described in the art.
[0103] amplification The provided methods and compositions, in various embodiments, employ or are the use of at least one amplification step in which nucleic acid material (or a portion thereof, e.g., a particular target region or locus) is amplified to form amplified nucleic acid material (e.g., a number of amplification products). In some embodiments, the provided methods include separating the amplified nucleic acid material, e.g., into first and second samples.
[0104] In some embodiments, amplifying the nucleic acid material in the first sample includes using at least one single-stranded oligonucleotide that is at least partially complementary to a sequence present in the first adapter sequence and at least one single-stranded oligonucleotide that is at least partially complementary to the target sequence of interest to amplify the nucleic acid material derived from a single nucleic acid strand from the original double-stranded nucleic acid material such that the SMI sequence is at least partially maintained.
[0105] In some embodiments, amplifying the nucleic acid material in the second sample includes using at least one single-stranded oligonucleotide that is at least partially complementary to a sequence present in the second adapter sequence and at least one single-stranded oligonucleotide that is at least partially complementary to the target sequence of interest to amplify the nucleic acid material derived from a single nucleic acid strand from the original double-stranded nucleic acid material such that the SMI sequence is at least partially maintained.
[0106] In some embodiments, the amplified nucleic acid material may be separated into three or more samples (e.g., 4, 5, 6, 7, 8, 9, 20, 20, 30, 40, 50, or more samples) prior to the second amplification step. In some embodiments, each sample contains approximately the same amount of amplified nucleic acid material as each other sample. In some embodiments, at least two samples contain substantially different amounts of amplified nucleic acid material.
[0107] In some embodiments, amplifying nucleic acid material of a first sample or a second sample includes amplifying samples in "tubes" (e.g., PCR tubes), emulsion droplets, microchambers, and other examples above or other known containers.
[0108] In some embodiments, at least one amplification step comprises at least one primer that is or comprises at least one non-standard nucleotide, hi some embodiments, the non-standard nucleotide is selected from uracil, methylated nucleotides, RNA nucleotides, ribose nucleotides, 8-oxo-guanine, biotinylated nucleotides, locked nucleic acids, peptide nucleic acids, high Tm nucleic acid variants, allele-discriminating nucleic acid variants, other nucleotide or linker variants described anywhere herein, and any combination thereof.
[0109] While any amplification reaction suitable for the application is contemplated to be compatible with some embodiments, as specific examples, in some embodiments the amplifying step may be or include polymerase chain reaction (PCR), rolling circle amplification (RCA), multiple displacement amplification (MDA), isothermal amplification, polony amplification in an emulsion, bridge amplification on a surface, on the surface of a bead or inside a hydrogel, and any combination thereof.
[0110] In some embodiments, specific modifications may be made to portions of a sample of nucleic acid material (e.g., adapter sequences). By way of example, in some embodiments, amplifying nucleic acid material in a first sample may further include, after the separating step and before amplifying the first sample, destroying or disrupting some or all of the second adapter sequences found on the nucleic acid material. As a further specific example, in some embodiments, amplifying nucleic acid material in a second sample may further include, after the separating step and before amplifying the second sample, destroying or disrupting at least some of the first adapter sequences found on the nucleic acid material. In some embodiments, the destruction or disruption may be or may include at least one of enzymatic digestion (e.g., by endonucleases and / or exonucleases), inclusion of at least one replication inhibitory molecule, enzymatic cleavage, enzymatic cleavage of one strand, enzymatic cleavage of both strands, incorporation of modified nucleic acids followed by enzymatic treatment resulting in cleavage or one or both strands, incorporation of replication-blocking nucleotides, incorporation of chain terminators, incorporation of photocleavable linkers, incorporation of uracil, incorporation of ribose bases, incorporation of 8-oxo-guanine adducts, use of sequence-specific restriction endonucleases, use of targeting endonucleases, and any combination thereof. In some embodiments, in addition to or as an alternative to primer site destruction or disruption, methods such as affinity pull-down, size selection, or other known techniques for removing and / or not amplifying undesired nucleic acid material from a sample are contemplated.
[0111] In some embodiments, the undesired first amplification product targeted for at least partial destruction results in a second amplification product following a second amplification with a targeting primer that ultimately contains two similar primer binding sites at each end of the molecule rather than two separate primer binding sites. In some embodiments, such a structure can be problematic for the performance or efficiency of the MPS DNA sequence.
[0112] In some embodiments, amplifying the nucleic acid material involves the use of a single-stranded oligonucleotide that is at least partially complementary to a target region or sequence of interest (e.g., a genomic sequence, a mitochondrial sequence, a plasmid sequence, a synthetically produced target nucleic acid, etc.) and at least one single-stranded oligonucleotide that is at least partially complementary to a region of an adapter sequence (e.g., a primer site). In some embodiments, amplifying the nucleic acid material involves the use of single-stranded oligonucleotides that are at least partially complementary to a region of an adapter sequence at the 5' and 3' ends of each strand of the nucleic acid material.
[0113] In general, robust amplification, such as PCR amplification, is highly dependent on reaction conditions. For example, multiplex PCR is sensitive to buffer composition, monovalent or divalent cation concentrations, detergent concentrations, crowding agent (i.e., PEG, glycerol, etc.) concentrations, primer concentrations, primer Tm, primer design, primer GC content, characteristics of primer-modified nucleotides, and cycling conditions (i.e., temperature and extension time, and temperature ramp rate). Optimizing buffer conditions can be a challenging and time-consuming process. In some embodiments, amplification reactions can use at least one of buffers, primer pool concentrations, and PCR conditions according to previously known amplification protocols. In some embodiments, new amplification protocols can be created and / or amplification reaction optimization can be used. As a specific example, in some embodiments, PCR optimization kits can be used, such as the Promega® PCR Optimization Kit, which includes a number of pre-prepared buffers partially optimized for various PCR applications, such as multiplex, real-time, GC-rich, and inhibitor-resistant amplification. These pre-prepared buffers include various Mg 2+and primer concentrations, as well as primer pool ratios, can be rapidly replenished. Additionally, in some embodiments, various cycling conditions (e.g., thermal cycling) may be evaluated and / or used. When assessing whether a particular embodiment is suitable for a particular desired application, one or more of the following may be evaluated: specificity, allele coverage of heterozygous loci, interlocus balance, and depth, among other aspects. Measurement of amplification success may include DNA sequencing of the products, evaluation of products by gel or capillary electrophoresis or HPLC or other size separation methods followed by visualization of fragments, melting curve analysis using double-stranded nucleic acid binding dyes or fluorescent probes, mass spectrometry, or other methods known in the art.
[0114] According to various embodiments, any of a variety of factors can affect the length of a particular amplification step (e.g., the number of cycles in a PCR reaction). For example, in some embodiments, the provided nucleic acid material may be damaged or otherwise suboptimal (e.g., degraded and / or contaminated). In such cases, a longer amplification step can help ensure that the desired product is acceptably amplified. In some embodiments, the amplification step may provide an average of 3-10 sequenced PCR copies from each starting DNA molecule, while in other embodiments, only a single copy of each of the top and bottom strands is necessary. Without wishing to be bound by any particular theory, using too many or too few PCR copies can reduce the efficiency of the assay and ultimately reduce its depth. In general, the number of nucleic acid (e.g., DNA) fragments used in an amplification (e.g., PCR) reaction is the primary adjustable variable that can determine the number of reads that share the same SMI / barcode sequence. Because SPLiT-DS uses an additional PCR step and does not require the use of hybridization-based targeted capture as in some previously described methods, the input amount requirements for double-stranded nucleic acids reported using previous methods may not be directly translatable to currently offered methods that may be more efficient.
[0115] Primer site destruction Figures 6-9B are conceptual diagrams of various SPLiT-DS method steps according to additional embodiments of the present technology. As described above and with reference to Figures 4-6, the method steps associated with SPLiT-DS result in amplified nucleic acid material having first and second strand amplification products (e.g., α, α', β, β', Figure 6) tagged with additional adapter sequences containing asymmetric primer sites (e.g., for Illumina P5 and P7 primers, Figure 6) after a first round of amplification that can be separated into SMI and multiple samples. Figure 7 illustrates a subsequent step in which nested PCR reactions can provide enriched amplification of the top and bottom strands of the original nucleic acid molecule in separate reaction samples (e.g., tubes). As shown in Figure 7, in addition to enrichment of the desired amplification products, some undesired amplification products and subsequent sequencing reads may be generated. Thus, and in some embodiments, efficiency may be reduced (e.g., the percentage of desired products for use in SPLiT-DS may be lower compared to products not useful in the SPLiT-DS protocol).
[0116] According to further aspects of the present technology, various aspects of conversion efficiency and workflow efficiency can be improved by employing one or more strategies to reduce and / or eliminate the amplification and sequencing of undesired amplification products. In some embodiments, primer site destruction or disruption (e.g., primer site destruction within adapter sequences) can be used as a method to enrich for specific nucleic acid products after the first round of amplification and separation of the amplified nucleic acid material into multiple samples (e.g., as shown in Figure 8A). In some embodiments, the provided methods can include the use of double-stranded primer site destruction. Several methods of primer site destruction are contemplated herein. Figures 8A-8D are conceptual diagrams of the steps of the SPLiT-DS method incorporating a double-stranded primer site destruction scheme. Double-stranded primer site destruction can be achieved by various means, including the introduction of primer site modifications into the targeting strand via modified primers used in the first amplification step (e.g., Figure 6). In some embodiments, primers in the first PCR can have modifications including uracil, methylation, RNA bases, 8-oxo-guanine, or other modifications that can be targeted in subsequent steps. In some embodiments, primer site destruction can be or can include, for example, digestion of sequences present in the adapter sequence with a restriction enzyme or other targeting endonuclease (Cas9, CPF1, etc.), where potential restriction sites have been determined to be unlikely to occur in the sequence of interest.
[0117] In certain embodiments, oligonucleotides complementary to the primer sequence to be destroyed can be added to a particular sample, followed by interrogation with a targeting endonuclease specific for double-stranded DNA. In another specific embodiment, hybridizing oligos bearing methyl groups can be used to recruit a methylation-specific restriction endonuclease to the complementary primer site. As shown in Figure 8A, double-stranded primer site destruction (e.g., primer site destruction on both copies of the non-target strand in the sample) can be used to destroy, disable, or remove the "P5" primer sequence from both the "top strand" and "bottom strand" copies in tube 1. Similarly, the "P7" primer sequence can be selectively destroyed, disabled, or removed from both the "top strand" and "bottom strand" copies in tube 2. Figure 8B is a conceptual diagram of one example for selectively destroying primer sequences in a sample. As shown in Figure 8B, a first sample can be treated with a first restriction endonuclease (e.g., MspJI) that selectively cleaves at a site found in a first primer sequence (e.g., Illumina "P5"), thereby destroying the first primer site in all nucleic acid material in the first sample. Similarly, a second sample can be treated with a second restriction endonuclease (e.g., FspEI) that selectively cleaves at a site found in a second primer sequence (e.g., Illumina "P7"), thereby destroying the second primer site in all nucleic acid material in the second sample.
[0118] 8A and 8C together, selectively amplifying (extending one or more linear cycles) the product in tube 1 using a target sequence primer (e.g., a gene-specific primer) having a "P7" primer and a "P5" primer site tail generates only a "bottom strand" species incorporating both the "P7" and "P5" primer sites (e.g., see FIG. 8C), while other nucleic acid species in tube 1 cannot be exponentially amplified or sequenced (e.g., lack the "P5" primer site). Similarly, selectively amplifying (extending one or more linear cycles) the product in tube 2 using a target sequence primer (e.g., a gene-specific primer) having a "P5" primer and a "P7" primer site tail generates only a "top strand" species incorporating both the "P5" and "P7" primer sites (e.g., see FIG. 8C), while other nucleic acid species in tube 2 cannot be exponentially amplified or sequenced (e.g., lack the "P5" primer site). It is understood that although unwanted linear products are not sequenced or exponentially amplified, they may consume primers and dNTPs and affect the efficiency of such reactions.
[0119] In some embodiments, methods involving primer site disruption may also use one or more biotinylated or other targeted primers. Figure 8D is a conceptual diagram of the steps of the SPLiT-DS method incorporating a double-stranded primer site disruption scheme according to another embodiment of the present technology. In the embodiment shown in Figure 8D, a target sequence primer with a "P5" primer site tail or a "P7" primer site tail is biotinylated. Referring to Figure 8D, following the extension step with the biotinylated targeted primer, streptavidin beads or hydrogel enrichment can be used to enrich for products with two primer sites, thereby eliminating the majority of nucleic acid species with only one primer site. In some such embodiments, such enrichment may improve PCR efficiency, facilitate multiplexing approaches, improve cluster amplification efficiency on an MPS DNA sequencer, and / or generate more useful sequencing data on an MPS DNA sequencer.
[0120] To further limit off-target enrichment of species captured by biotin / streptavidin enrichment, additional amplification with nested primers (e.g., a "P5" or "P7" primer and a second targeting primer nested internally with the opposite flow cell sequence) can be used to further enrich for on-target species and reduce undesired amplification products. In certain embodiments, for example, selective linear amplification using primers specific for the target sequence of interest can further enrich for desired species before the addition of paired nested primers for exponential amplification.
[0121] In some embodiments, single-stranded primer site destruction may be used. Figures 9A and 9B are conceptual diagrams of various embodiments of steps in the SPLiT-DS method incorporating a single-stranded primer site destruction scheme according to further aspects of the present technology. As a non-limiting example, as shown in Figure 9A, a modified primer (not shown) can be used during the first amplification step of SPLiT-DS to destroy the primer site on one strand of a double-stranded molecule (see, for example, Figure 6). The modified primer can include chemical modifications (e.g., uracil, methylation, RNA bases, 8-oxo-guanine, etc.) that can be targeted after the destruction or neutralization of the primer site on the affected strand. Subsequent amplification (one or more linear cycles of extension) of the desired target in tube 1 using the "P7" primer and a specially labeled (e.g., biotin, a different flow cell adapter tail, etc.) target sequence primer (e.g., a gene-specific primer) generates only "bottom strand" species incorporating both "P7" and the special label (e.g., biotin, a different primer site, etc.) (see, e.g., Figure 9A), while other nucleic acid species in tube 1 do not exponentially amplify. In a next step, undesired products are further selected by streptavidin bead enrichment (not shown) or by further amplification with the "P7" primer and a modified primer with a different primer site topological complement and a flow cell adapter tail bearing the "P5" primer site (Figure 9B). A final amplification reaction with the "P7" and "P5" primers yields enriched "bottom strand" products in the tube 1 sample (Figure 9B). A complementation step in the tube 2 sample can be performed to enrich for "top strand" products (Figure 9B). Without wishing to be bound by any particular theory, when the option of double-stranded primer site digestion is available, such option may be preferred over single-stranded digestion.
[0122] In further embodiments, one or more of the schemes described with respect to Figures 6-9B may be combined, or certain steps may be eliminated while still achieving certain efficiency improvements. For example, in one embodiment, a biotinylated targeting primer may be used during the extension step (e.g., following the method steps shown in Figure 6), and subsequent streptavidin probing may be used to recover the strand of interest. In this embodiment (e.g., without primer site disruption), species with two of the same primer sites (e.g., two "P5" primer sites, two "P7" primer sites) are also recovered.
[0123] Multiple PCRs per captured molecule In certain applications, targeted regions or sequences may be difficult to sequence because the nucleic acid breakpoints may be inaccessible to the target-specific primers, resulting in short fragments or completely missing regions. For example, randomly sheared DNA or circulating cell-free DNA (cfDNA), such as circulating tumor DNA or circulating fetal DNA, may have target sequences that are not retrievable (e.g., not detected and / or included in the sequencing readout). In some embodiments, the provided methods can overcome such challenges by targeting multiple regions within the target sequence (e.g., each primer targets a different region of the target sequence), such as by using multiple target primers complementary to staggered portions of the target sequence. To avoid the challenges associated with short fragments, and in one embodiment, DNA can be sheared into larger fragments than would typically be desirable for optimal sequencing. Figure 10 is a conceptual diagram of the steps of the SPLiT-DS method using multiple target primers to generate double-stranded consensus sequences of longer nucleic acid molecules, according to yet another embodiment of the present technology.
[0124] With reference to Figure 10, the provided methods can include the use of multiple amplification primers, e.g., multiple primers each targeting a region of the target sequence of interest (e.g., approximately 100 bp apart). According to various embodiments, such an approach can be performed in a single reaction (e.g., a tube) or, in other embodiments, in multiple reactions (e.g., tubes), e.g., to avoid nearby or adjacent primers from interacting with each other. In some embodiments, preventing interaction of multiple staggered primers within the same tube can be mitigated by performing extension with a strand-displacing polymerase so that primers priming downstream do not block primers priming further upstream. In some embodiments, extension is performed for several linear cycles with a first primer, followed by cleanup, extension of another set with a second primer, and so on. As shown in Figure 10, nested primer sets generate amplification products of different lengths that can then be sequenced. Read 1 of all amplification products produces identical sequence information, but paired-end sequence reads from each of amplification products A, B, and C combine with the read 1 sequence information to yield staggered sequencing information that provides a longer length of combined sequence than previously possible with MPS or standard DS protocols.
[0125] In some embodiments, analysis of multi-primer data is performed in a non-standard manner relative to other DS methods. As those skilled in the art will appreciate, multiplexed samples may contain products of various lengths with identical tags, making duplex assembly of multi-primer sequence reads impossible using SMI tags alone. To address this challenge, some embodiments include assembly of duplexes using tags that combine SMIs with the sequence (e.g., genomic) location of the target primer start site. In some embodiments, after duplex assembly, data can be evaluated for duplex reads of different lengths that share a common SMI. In some embodiments, individual duplex families can be grouped into a collective "multiple-read duplex family." Some such embodiments may be advantageous for certain applications and may facilitate the subassembly of DS target regions into longer single-molecule reads, increasing the effective genotyping length of target nucleic acid molecules on short-read sequencing platforms.
[0126] As known to those skilled in the art, the longest contiguous read currently obtainable on an Illumina NextSeq is approximately 300 bp, with a matched paired-end 150 bp read in the middle, provided the enzyme targeting and primers are carefully designed to produce fragments substantially close to this length. Thus, embodiments incorporating the multi-primer approach described herein may, in some embodiments, achieve longer whole-molecule DS sequences.
[0127] In some aspects, the provided methods, in some embodiments, allow for the use of multiple targeting primers in combination with SPLiT-DS to achieve, among other things, (i) contiguous sequences of a single long molecule, and, optionally, (ii) high specificity and / or (ii) accuracy of DS. The provided methods are likely to be useful, for example, in applications requiring long, accurate contiguous reads, de novo genome assembly, performing assays in uniquely difficult-to-map repetitive regions (i.e., regions of the genome with repetitive sequences), sequencing regions considered particularly challenging (e.g., HLA loci, cancer pseudogenes, microsatellites), cancer (e.g., drug-sensitizing mutations, resistance mutations), haplotype analysis (e.g., assessing the origin of mutations in circulating fetal DNA (e.g., maternal, paternal, or fetal origin)), assays for the co-occurrence of variants, such as metagenomics (e.g., antibiotic resistance), overcoming limitations of certain enzymes (e.g., limitations on how far apart certain regions need to be based on the location of Cas9 enzyme recognition sites), large-scale structural rearrangements, and / or indels.
[0128] Further embodiments for processing nucleic acid material In some embodiments, it is advantageous to treat nucleic acid material to improve the efficiency, accuracy, and / or speed of the sequencing process. According to further aspects of the present technology, for example, the efficiency of DS and / or SPLiT-DS can be enhanced by fragmenting targeted nucleic acids. Classically, fragmentation of nucleic acids (e.g., genomes, mitochondria, plasmids, etc.) is achieved either by physical shearing (e.g., sonication) or by somewhat non-sequence-specific enzymatic approaches that utilize enzyme cocktails to cleave DNA phosphodiester bonds. The result of either of these methods is a sample in which intact nucleic acid material (e.g., genomic DNA (gDNA)) has been reduced to a mixture of randomly or semi-randomly sized nucleic acid fragments. While effective, these approaches generate nucleic acid fragments of variable size, which can lead to amplification bias (e.g., short fragments tend to PCR amplify more readily than long fragments and clusters are more readily amplified during polony formation) and uneven sequencing depth. For example, Figure 11A is a graph plotting the relationship between nucleic acid insert size and the resulting family size after amplification. As shown in Figure 11A, shorter fragments tend to amplify preferentially, resulting in more copies of each of these shorter fragments being generated and sequenced, providing disproportionate levels of sequencing depth for these regions. Furthermore, for longer fragments, the portion of the DNA between the limits of the sequencing read (or between the ends of the paired-end sequencing read) cannot be interrogated, resulting in a "dark" image despite successful ligation, amplification, and capture (Figure 11B). Similarly, for short reads, and when using paired-end sequencing, identical sequences in the center of the molecule from both reads provide redundant information, making the method less cost-effective (Figure 11B). Random or semi-random nucleic acid fragmentation results in unpredictable breakpoints in the target molecule, generating fragments that may be less or more complementary to the hybrid capture bait strand, thereby reducing target capture efficiency. Random or semi-random fragmentation can also result in very small or very large fragments that disrupt the sequence of interest and / or are lost during other steps of library preparation, reducing data yield and efficiency.
[0129] Another problem with many random fragmentation methods, particularly mechanical or acoustic methods, is that they result in damage beyond double-strand breaks, which can cause portions of double-stranded DNA to cease to be double-stranded. For example, mechanical shearing can result in 3' or 5' overhangs at the ends of the molecule and single-stranded nicks in the center of the molecule. These single-stranded portions, suitable for adapter ligation, such as a cocktail of "end repair" enzymes, can be used to artificially re-make them double-stranded, potentially resulting in artifacts (such as those described above for "pseudoduplexes"). In many embodiments, it is optimal to maximize the amount of the desired double-stranded nucleic acid that remains in its natural double-stranded form during handling.
[0130] Thus, in some embodiments, provided methods and compositions utilize targeting endonucleases (e.g., ribonucleoprotein complexes (CRISPR-associated endonucleases such as Cas9, Cpfl), homing endonucleases, zinc finger nucleases, TALENs, Argonaute nucleases, and / or meganucleases (e.g., megaTAL nucleases, etc.), or combinations thereof) or other technologies (e.g., one or more restriction enzymes) that can cleave nucleic acid material to excise a target sequence of interest in fragments of optimal size for sequencing. In some embodiments, the targeting endonuclease has the ability to specifically and selectively excise a precise sequence region of interest. Figure 11C is a schematic diagram showing method steps for generating targeted fragments cleaved to specific sizes by CRISPR / Cas9 and generating sequencing information, according to one embodiment of the present technology. For example, by pre-selecting cleavage sites using a programmable endonuclease (e.g., a CRISPR-associated (Cas) enzyme / guide RNA complex) that results in fragments of a given, substantially uniform size (Figure 11C), bias and the presence of uninformative reads can be significantly reduced. Furthermore, due to the size difference between the excised fragments and the remaining uncut DNA, a size selection step (described further below) can be performed to remove large off-target regions and pre-enrich the sample before any further processing steps. The need for end-repair steps may likewise be reduced or eliminated, thus saving time and the risk of false duplication challenges and, in some cases, improving efficiency by reducing or eliminating the need for computational trimming of data near the ends of molecules.
[0131] Restriction endonucleases It is specifically contemplated that any of a variety of restriction endonucleases (i.e., enzymes) can be used to provide nucleic acid material of substantially uniform length. Generally, restriction enzymes are usually produced by specific bacteria / other prokaryotes and cut at, near, or between specific sequences in specific segments of DNA.
[0132] It will be apparent to those skilled in the art that restriction enzymes are selected to cleave at specific sites, or alternatively, at sites that are generated to create restriction sites for cleavage. In some embodiments, the restriction enzyme is a synthetic enzyme. In some embodiments, the restriction enzyme is not a synthetic enzyme. In some embodiments, the restriction enzymes used herein are modified to introduce one or more changes into their own genome. In some embodiments, the restriction enzyme generates a double-stranded break between defined sequences within a specific portion of DNA.
[0133] While any restriction enzyme (e.g., Type I, II, III, and / or IV) may be used according to some embodiments, the following represents a non-limiting list of restriction enzymes that can be used: AluI, ApoI, AspHI, BamHI, BfaI, BsaI, CfrI, DdeI, DpnI, DraI, EcoRI, EcoRII, EcoRV, HaeII, HaeIII, HgaI, HindII, HindIII, HinFI, KpnI, MamI, MseI, MstI, MstII, NcoI, NdeI, NotI, PacI, PstI, PvuI, PvuII, RcaI, RsaI, SacI, SacII, SalI, Sau3AI, ScaI, SmaI, SpeI, SphI, StuI, XbaI, XhoI, XhoII, XmaI, XmaII, and any combination thereof. Extensive, but not exhaustive, lists of suitable restriction enzymes can be found in public catalogs and on the internet (for example, available from New England Biolabs, Ipswich, MA, USA).
[0134] Targeting Endonucleases Targeting endonucleases (e.g., CRISPR-associated ribonucleoprotein complexes such as Cas9 or Cpf1, homing nucleases, zinc finger nucleases, TALENs, megaTAL nucleases, Argonaute nucleases, and / or derivatives thereof) can be used to selectively cleave and excise targeted portions of nucleic acid material for the purpose of enriching such targeted portions for sequencing applications. In some embodiments, targeting endonucleases can be modified, e.g., with amino acid substitutions to provide enhanced thermostability, salt tolerance, and / or pH tolerance. In other embodiments, targeting endonucleases may be biotinylated, fused to streptavidin, and / or incorporate other affinity-based (e.g., bait / prey) technologies. In certain embodiments, targeting endonucleases can have altered recognition site specificity (e.g., SpCas9 variants with altered PAM site specificity). CRISPR-based targeting endonucleases are further described herein to provide further detailed, non-limiting examples of the use of targeting endonucleases. It should be noted that the nomenclature surrounding such targeting nucleases is in flux. For purposes of this specification, the term "CRISPR-based" is used to generally refer to endonucleases that contain nucleic acid sequences that can be modified to redefine the nucleic acid sequence to be cleaved. While Cas9 and CPF1 are currently used examples of such targeting endonucleases, many more exist in various locations in nature, and the variety of such targeted, easily tunable nucleases is expected to rapidly increase over the next few years. Similarly, multiple engineered variants of these enzymes are becoming available to enhance or modify their properties. This specification expressly contemplates the use of substantially functionally similar targeting endonucleases not explicitly described herein or yet to be discovered, which achieve similar objectives to those disclosed herein.
[0135] CRISPR-DS A further aspect of the present technology relates to a method for enriching regions of interest using the programmable endonuclease CRISPR / Cas9. In particular, CRISPR / Cas9 (or other programmable endonucleases) can be used to selectively excise one or more sequence regions of interest, and the excised target regions can be designed to be one or more predetermined lengths, allowing for size selection prior to library preparation for sequencing applications such as DS and SPLiT-DS. These programmable endonucleases can be used alone or in combination with other forms of targeted nucleases, such as restriction endonucleases. This method, called CRISPR-DS, allows for a very high degree of on-target enrichment (reducing the need for a subsequent hybrid capture step), significantly reducing time and cost and increasing conversion efficiency. Figures 12A-12D are conceptual diagrams of the steps of the CRISPR-DS method according to one embodiment of the present technology. For example, CRISPR / Cas9 can be used to cleave one or more specific sites (e.g., PAM sites) within a target sequence (Figure 12A; in this example, the TP53 target region). Figure 12B shows one method for isolating the excised target portion using SPRI / Ampure beads and magnetic purification to remove high-molecular-weight DNA while leaving a given shorter fragment. In other embodiments, various size-selection methods, including but not limited to gel electrophoresis, gel purification, liquid chromatography, size-exclusion purification, and filtration purification, can be used to separate excised portions of a given length from undesired DNA fragments and other high-molecular-weight genomic DNA (if applicable). After size selection, the CRISPR-DS method includes steps consistent with those of the DS method (see, e.g., Figure 12E), including A-tailing (which leaves blunt ends after CRISPR / Cas9 excision), ligation of DS adapters (Figure 12C), duplex amplification (Figure 12D), sequencing of each strand, and capture and index amplification (e.g., PCR) prior to generation of a duplex consensus sequence (Figure 12D). In addition to the improved workflow efficiency evident in Figure 12E, CRISPR-DS provides optimal fragment lengths for highly efficient amplification and sequencing steps (Figure 12F).
[0136] In certain embodiments, CRISPR-DS solves multiple common problems associated with NGS, including, for example, inefficient target enrichment that can be optimized by CRISPR-based size selection, sequence errors that can be removed using DS methodology to generate error-corrected duplex consensus sequences, and uneven fragment sizes that are mitigated by pre-designed CRISPR / Cas9 fragmentation (Table 1).
[0137] JPEG0007821756000001.jpg58168Target description: description of target, Name: name, Sequence plus pam sige: sequence + pam site, Position start: start position, Position end: yield position, Zhang score: Zhang score upstream of exon 11: upstream of exon 11 downstream of exon 11: downstream of exon 11 upstream of exon 10: upstream of exon 10 downstream of exon 10: downstream of exon 10 upstream of exon 9-8: upstream of exon 9-8 downstream of exon 9-8: downstream of exon 9-8 downstream of exon 7: downstream of exon 7 upstream of exon 6-5: upstream of exon 6-5 downstream of exon 6-5: downstream of exon 6-5 upstream of exon 4-3: upstream of exon 4-3 downstream of exon 4-3: downstream of exon 4-3 downstream of exon 2: downstream of exon 2
[0138] In vitro digestion of DNA material with Cas9 nuclease relies on the formation of a ribonucleoprotein complex that recognizes and cleaves a given site (e.g., a PAM site, Figure 11C). This complex is formed by a guide RNA ("gRNA," e.g., crRNA + tracrRNA) and Cas9. For multiple cleavages, the gRNA can be complexed by pooling all crRNAs and then complexing with the tracrRNA, or by complexing each crRNA with the tracrRNA separately and then pooling. In some embodiments, the second option may be preferred because it eliminates competition between crRNAs.
[0139] As described herein, and as will be appreciated by those skilled in the art, CRISPR-DS may have applications for sensitive identification of mutations in situations where samples are limited to DNA, such as in forensic and early cancer detection applications.
[0140] In some embodiments, the nucleic acid material comprises nucleic acid molecules of substantially uniform length (in some embodiments, substantially uniform length is from about 1 to about 1,000,000 bases). For example, in some embodiments, the substantially uniform length can be at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 50, 60, 70, 80, 90, 100, 120, 150, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1200, 1500, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10,000, 15,000, 20,000, 30,000, 40,000, or 50,000 bases in length. In some embodiments, the substantially uniform length can be up to 60,000, 70,000, 80,000, 90,000, 100,000, 120,000, 150,000, 200,000, 300,000, 400,000, 500,000, 600,000, 700,000, 800,000, 900,000, or 1,000,000 bases. As a specific, non-limiting example, in some embodiments, the substantially uniform length is about 100 to about 500 bases. In some embodiments, a size selection step as described herein may be performed before any particular amplification step. In some embodiments, a size selection step as described herein may be performed after any particular amplification step. In some embodiments, a size selection step as described herein may be followed by additional steps, such as a digestion step and / or another size selection step.
[0141] In addition to the use of targeting endonucleases, any other application-appropriate method for achieving substantially uniform length nucleic acid molecules may be used. By way of non-limiting example, such methods may be or may include the use of one or more of agarose or other gels, affinity columns, HPLC, PAGE, filtration, SPRI / Ampure-type beads, or other suitable methods recognized by those of skill in the art.
[0142] In some embodiments, processing nucleic acid material to produce nucleic acid molecules of substantially uniform length (or mass) can be used to recover one or more desired target regions (e.g., target sequences of interest) from a sample. In some embodiments, processing nucleic acid material to produce nucleic acid molecules of substantially uniform length (or mass) can be used to exclude certain portions of a sample (e.g., nucleic acid material from undesired species or undesired subjects of the same species). In some embodiments, nucleic acid material can be present in a variety of sizes (e.g., not substantially uniform length or mass).
[0143] In some embodiments, more than one targeting endonuclease or other method for producing nucleic acid molecules of substantially uniform length may be used (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, or more). In some embodiments, a targeted nuclease can be used to cleave more than one potential target region of a nucleic acid material (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, or more). In some embodiments, when there is more than one target region of a nucleic acid material, each target region can be the same (or substantially the same) length. In some embodiments, when there is more than one target region of a nucleic acid material, at least two of the target regions of known length are different in length (e.g., a first target region that is 100 bp long and a second target region that is 1,000 bp long).
[0144] In some embodiments, multiple targeting endonucleases (e.g., programmable endonucleases) may be used in combination to fragment multiple regions of a target nucleic acid of interest. In some embodiments, one or more programmable targeting endonucleases may be used in combination with other targeting nucleases. In some embodiments, one or more targeting endonucleases may be used in combination with random or semi-random nucleases. In some embodiments, one or more targeting endonucleases may be used in combination with other random or semi-random nucleic acid fragmentation methods, such as mechanical or acoustic shearing. In some embodiments, it may be advantageous to perform cleavage in sequential steps, interleaved with one or more size selection steps. In some embodiments where targeted fragmentation is used in combination with random or semi-random fragmentation, the random or semi-random nature of the latter may be useful for achieving the purpose of SMI. In some embodiments where targeted fragmentation is used in combination with random or semi-random fragmentation, the random or semi-random nature of the latter may be useful for facilitating sequencing of regions of the nucleic acid that are not easily cleaved by targeted methods, such as long, highly repetitive regions.
[0145] More ways In some embodiments, the provided methods may include providing a nucleic acid material, cleaving the nucleic acid material with a targeting endonuclease (e.g., a ribonucleoprotein complex) such that a target region of a given length is separated from the remaining nucleic acid material, and analyzing the cleaved target region. In some embodiments, the provided methods may further include ligating at least one SMI and / or adapter sequence to at least one of the 5' or 3' ends of the cleaved target region of the given length. In some embodiments, the analyzing may be or may include quantification and / or sequencing.
[0146] In some embodiments, quantification may be or include spectrophotometric analysis, real-time PCR, and / or fluorescence-based quantification (e.g., using fluorescent dye tagging). In some embodiments, sequencing may be or include Sanger sequencing, shotgun sequencing, bridge PCR, nanopore sequencing, single-molecule real-time sequencing, ion torrent sequencing, pyrosequencing, digital sequencing (e.g., digital barcode-based sequencing), sequencing by ligation, polony-based sequencing, current-based sequencing (e.g., tunneling current), mass spectrometry sequencing, microfluidics-based sequencing, and any combination thereof.
[0147] In some embodiments, the targeting endonuclease is or includes at least one of a CRISPR-associated (Cas) enzyme (e.g., Cas9 or Cpfl), or other ribonucleoprotein complex, a homing endonuclease, a zinc finger nuclease, a transcription activator-like effector nuclease (TALEN), an Argonaute nuclease, and / or a megaTAL nuclease. In some embodiments, more than one targeting endonuclease may be used (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, or more). In some embodiments, the targeting nuclease can be used to cleave more than one potential target region of a given length (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, or more). In some embodiments, when more than one target region of a given length is present, each target region can be the same (or substantially the same) length. In some embodiments, when there is more than one target region of a given length, at least two of the target regions of a given length are different in length (e.g., a first target region that is 100 bp long and a second target region that is 1,000 bp long).
[0148] Further Aspects According to one aspect of the present disclosure, some embodiments provide high-quality sequencing information from very small amounts of nucleic acid material. In some embodiments, the provided methods and compositions can be used with starting nucleic acid material in amounts up to about 1 picogram (pg), 10 pg, 100 pg, 1 nanogram (ng), 10 ng, 100 ng, 200 ng, 300 ng, 400 ng, 500 ng, 600 ng, 700 ng, 800 ng, 900 ng, or 1000 ng. In some embodiments, the provided methods and compositions can be used with input amounts of nucleic acid material of up to 1 molecular copy or genome equivalent, 10 molecular copies or genome equivalent, 100 molecular copies or genome equivalent, 1,000 molecular copies or genome equivalent, 10,000 molecular copies or genome equivalent, 100,000 molecular copies or genome equivalent, or 1,000,000 molecular copies or genome equivalent. For example, in some embodiments, up to 1,000 ng of nucleic acid material is initially provided to a particular sequencing process. For example, in some embodiments, up to 100 ng of nucleic acid material is initially provided to a particular sequencing process. For example, in some embodiments, up to 10 ng of nucleic acid material is initially provided to a particular sequencing process. For example, in some embodiments, up to 1 ng of nucleic acid material is initially provided to a particular sequencing process. For example, in some embodiments, up to 100 pg of nucleic acid material is initially provided to a particular sequencing process. For example, in some embodiments, up to 1 pg of nucleic acid material is initially provided to a particular sequencing process.
[0149] In accordance with other aspects of the present technology, some methods provided may be useful for sequencing any of a variety of suboptimal (e.g., damaged or degraded) samples of nucleic acid material. For example, in some embodiments, at least a portion of the nucleic acid material is damaged. In some embodiments, the damage is oxidation, alkylation, deamination, methylation, hydrolysis, nicking, intrastrand crosslinks, interstrand crosslinks, blunt-end strand breaks, sticky-end double-strand breaks, phosphorylation, dephosphorylation, sumoylation, glycosylation, single-strand gaps, heat damage, desiccation damage, UV exposure damage, gamma radiation damage, X-ray damage, ionizing radiation damage, non-ionizing radiation damage, heavy particle radiation damage, nuclear decay damage, beta radiation damage, alpha radiation damage, neutron radiation damage, proton radiation damage, cosmic radiation damage, high pH damage, low pH damage, reactive oxidative species damage, free radical damage, peroxide damage, hypochlorite damage, damage from tissue fixation such as formalin or formaldehyde, reactive iron damage, low ionic conditions damage, high ionic conditions damage, unbuffered conditions damage, nuclease damage, environmental The damage may be or include at least one of damage due to exposure, damage due to fire, damage due to mechanical stress, damage due to enzymatic degradation, damage due to microorganisms, damage due to preparative mechanical shear, damage due to preparative enzymatic fragmentation, damage occurring naturally in vivo, damage occurring during nucleic acid extraction, damage occurring during sequencing library preparation, damage introduced by polymerases, damage introduced during nucleic acid repair, damage occurring during nucleic acid end processing, damage occurring during nucleic acid ligation, damage occurring during sequencing, damage occurring due to mechanical handling of DNA, damage occurring during passage through a nanopore, damage occurring as part of organismal aging, damage occurring as a result of exposure of an individual to chemicals, damage caused by mutagens, damage caused by carcinogens, damage caused by clastogens, damage caused by in vivo inflammatory damage due to oxygen exposure, damage due to one or more strand breaks, and combinations thereof.
[0150] Nucleic acid materials kinds According to various embodiments, any of a variety of nucleic acid materials may be used. In some embodiments, the nucleic acid material may include at least one modification to a polynucleotide within the standard sugar-phosphate backbone. In some embodiments, the nucleic acid material may include at least one modification within any base of the nucleic acid material. For example, as a non-limiting example, in some embodiments, the nucleic acid material is or includes at least one of double-stranded DNA, single-stranded DNA, double-stranded RNA, single-stranded RNA, peptide nucleic acid (PNA), and locked nucleic acid (LNA).
[0151] qualification According to various embodiments, the nucleic acid material may be subjected to one or more modifications before, substantially simultaneously with, or after a particular step, depending on the application for which the particular provided method or composition is to be used.
[0152] In some embodiments, the modification may be or may include repair of at least a portion of the nucleic acid material. While any method suitable for nucleic acid repair applications is contemplated as being compatible with some embodiments, certain exemplary methods and compositions are therefore described below and in the Examples.
[0153] As a non-limiting example, in some embodiments, DNA repair enzymes such as uracil DNA glycosylase (UDG), formamidopyrimidine DNA glycosylase (FPG), and 8-oxoguanine DNA glycosylase (OGG1) can be used to correct DNA damage (e.g., in vitro DNA damage). For example, these DNA repair enzymes are glycolases that remove damaged bases from DNA. For example, UDG removes uracil resulting from cytosine deamination (occurring through spontaneous hydrolysis of cytosine), while FPG removes 8-oxoguanine (e.g., a common DNA lesion caused by reactive oxygen species). FPG also has lyase activity that creates one-base gaps at abasic sites. Such abasic sites can result in PCR amplification failure, for example, because the polymerase fails to copy the template. Therefore, the use of such DNA damage repair enzymes can effectively remove damaged DNA that does not have true mutations but would otherwise go undetected as errors after sequencing and double-strand sequence analysis.
[0154] As mentioned above, in further embodiments, the sequencing reads generated from the processing steps discussed herein can be further filtered to eliminate spurious mutations by trimming the ends of reads that are most prone to artifacts. For example, DNA fragmentation can generate single-stranded portions at the ends of double-stranded molecules. These single-stranded portions can be filled during end repair (e.g., by Klenow). In some cases, polymerases make copy errors in these end-repaired regions, resulting in the creation of "pseudo-double-stranded molecules." These artifacts may appear to be true mutations when sequenced. These errors resulting from the end-repair mechanism can be eliminated from post-sequencing analysis by trimming the ends of the sequencing reads to exclude any mutations that may have occurred, thereby reducing the number of spurious mutations. In some embodiments, such trimming of sequencing reads can be achieved automatically (e.g., as a routine process step). In some embodiments, mutation frequencies can be assessed for fragment end regions, and if a threshold level of mutations is observed in the fragment end regions, trimming of the sequencing reads can be performed before generating double-stranded consensus sequence reads for the DNA fragments.
[0155] source It is contemplated that the nucleic acid material may be derived from any of a variety of sources. For example, in some embodiments, the nucleic acid material is provided from a sample from at least one subject (e.g., a human or animal subject) or other biological source. In some embodiments, the nucleic acid material is provided from a banked / deposited sample. In some embodiments, the sample is blood, serum, sweat, saliva, cerebrospinal fluid, mucus, uterine washings, vaginal swabs, nasal swabs, oral swabs, tissue scrapings, hair, fingerprints, urine, stool, vitreous fluid, peritoneal washings, saliva, bronchial washings, oral washings, pleural washings, gastric washings, gastric juice, bile, pancreatic duct washings, bile duct washings, common bile duct washings, gallbladder fluid, synovial fluid, infected wounds, non-infectious wounds, archaeological samples, forensic samples, water samples, tissue samples, food samples, bioliquids, or the like. The sample may be or include at least one of an actor sample, a plant sample, a nail scraping, semen, prostatic fluid, fallopian tube washings, cell-free nucleic acid, intracellular nucleic acid, a metagenomics sample, an implanted foreign body washing, a nasal wash, an intestinal fluid, an epithelial scraping, an epithelial wash, a tissue biopsy, an autopsy sample, an organ sample, a human identification sample, an artificially produced nucleic acid sample, a synthetic gene sample, a nucleic acid data repository sample, a tumor tissue, and any combination thereof. In other embodiments, the sample is or includes at least one of a microorganism, a plant-based organism, or a collected environmental sample (e.g., water, soil, archaeological, etc.).
[0156] Selected application examples As described herein, the provided methods and compositions may be used for any of a variety of purposes and / or in any of a variety of scenarios. The following are non-limiting examples of uses and / or scenarios that are for specific illustrative purposes only.
[0157] Forensic medicine Previous approaches to forensic DNA analysis have relied almost entirely on capillary electrophoretic separation of PCR amplification products to identify short tandem repeat sequence length polymorphisms. This type of analysis has proven extremely valuable since its introduction in 1991. Since then, several publications have introduced standardized protocols, validated its use in laboratories worldwide, detailed its use in many different populations, and introduced more efficient approaches such as miniSTR.
[0158] Although this approach has proven highly successful, the technology suffers from a number of drawbacks that limit its usefulness. For example, current approaches to STR genotyping often generate background signal due to PCR stutter caused by polymerase slippage on the template DNA. This issue is particularly significant in samples with two or more contributors, as it can be difficult to distinguish between stutter alleles and true alleles. Another problem arises when analyzing degraded DNA samples. Fragment length variation often results in significantly fewer or even absent longer PCR fragments. As a result, profiles from degraded DNA often have lower discriminatory power.
[0159] The introduction of MPS systems has the potential to address several challenging problems in forensic analysis. For example, these platforms offer the unparalleled ability to simultaneously analyze nuclear and mtDNA STRs and SNPs, dramatically increasing discriminatory power and offering the potential to determine ethnicity and even physical attributes. Furthermore, unlike PCR-CE, which simply reports the average genotype of a population of molecules, MPS technology digitally aggregates the complete nucleotide sequence of many individual DNA molecules, offering the unique ability to detect multiple nucleotide sequences within heterogeneous DNA mixtures. Because forensic samples containing more than two contributors remain one of the most problematic issues in forensic science, the impact of MPS on the field of forensic science could be enormous.
[0160] The publication of the human genome highlighted the immense power of the MPS platform. However, until very recently, the full power of these platforms was limited to forensic applications due to the read lengths significantly shorter than STR loci, making length-based genotyping impossible. Initially, pyrosequencers such as the Roche 454 platform were the only platforms with sufficient read lengths to sequence core STR loci. However, as the read lengths of competing technologies have increased, their utility for forensic applications has become apparent. Numerous studies have demonstrated the potential for MPS genotyping of STR loci. Overall, the general conclusion of all these studies is that, regardless of platform, STRs can be successfully classified, thereby generating genotypes comparable to CE analysis, even from damaged forensic samples.
[0161] While all these studies demonstrate agreement with traditional PCR-CE approaches and additional advantages, such as the detection of intra-STR SNPs, they also highlight numerous issues with current technology. For example, the current MPS approach to STR genotyping relies on multiplex PCR to provide sufficient DNA for sequencing and PCR primer introduction. However, because multiplex PCR kits are designed for PCR-CE, they contain primers for amplification products of various sizes. This variation can lead to biased amplification of small fragments, resulting in imbalanced inclusion and allele dropout. Indeed, recent studies have shown that differences in PCR efficiency affect mixture components, especially at low MAFs. To address this issue, several sequencing kits specifically designed for forensic applications are now commercially available, and validation studies are beginning to be reported. However, amplification bias remains evident due to the high level of multiplexing.
[0162] Like PCR-CE, MPS is not susceptible to the occurrence of PCR stutter. Most MPS studies of STRs report the occurrence of artificial drop-in alleles. Recently, systematic MPS studies have reported that many stutter events manifest as shorter polymorphisms, differing from true alleles by four base pairs, most commonly at n-4, but also at n-8 and n-12 positions. The percentage of stutter typically occurred in approximately 1% of reads, but could reach 3% at some loci, indicating that MPS may exhibit a higher rate of stutter than PCR-CE.
[0163] In contrast, in some embodiments, the provided methods and compositions allow for high-quality and efficient sequencing of low-quality and / or small-volume samples, as described above and in the Examples below. Thus, in some embodiments, the provided methods and / or compositions may be useful for detecting rare variants in DNA from one individual mixed at low concentrations with DNA from another individual of a different genotype.
[0164] Forensic DNA samples typically contain non-human DNA. Potential sources of this exogenous DNA include the source of the DNA (e.g., microorganisms in saliva or oral samples), the surface environment from which the sample was collected, and contamination from the laboratory (e.g., reagents, work areas, etc.). Another aspect provided by some embodiments is that certain provided methods and compositions enable differentiation of contaminating nucleic acid material from other sources (e.g., different species) and / or surface or environmental contaminants, allowing these substances (and / or their effects) to be removed from the final analysis without biasing the sequencing results.
[0165] In highly degraded DNA, locus-specific PCR does not work well due to DNA fragments that do not contain the necessary primer annealing sites, resulting in allele dropout. This situation limits the uniqueness of genotype calls, and the reliability of matches cannot be guaranteed, especially in mixed tests. However, in some embodiments, the provided methods and compositions allow the use of single nucleotide polymorphisms (SNPs) in addition to or as an alternative to STR markers.
[0166] Indeed, as data on human genetic variation continues to grow, SNPs are becoming increasingly important in forensic research. Thus, in some embodiments, the provided methods and compositions use primer design strategies to enable the creation of multiplex primer panels based on, for example, currently available sequencing kits, ensuring that reads span one or more SNP positions.
[0167] Patient stratification Patient stratification, which generally refers to dividing patients based on one or more non-treatment-related factors, is a topic of great interest in the medical community. Much of this interest can be attributed to the fact that certain therapeutic candidates have failed to receive FDA approval, in part, due to previously unrecognized differences between patients in trials. These differences can be or may include one or more genetic differences that cause one group of patients to metabolize the therapeutic drug differently or that result in or worsen side effects between one group of patients and one or more other groups of patients. In some cases, some or all of these differences can be detected as one or more different genetic profiles of patients that result in a different response to the therapeutic drug than other patients who do not exhibit the same genetic profile.
[0168] Thus, in some embodiments, the provided methods and compositions may be useful in determining which subjects in a particular patient population (e.g., patients suffering from a common disease, disorder, or condition) will respond to a particular treatment. For example, in some embodiments, the provided methods and / or compositions can be used to assess whether a particular subject possesses a genotype associated with a poor response to a treatment. In some embodiments, the provided methods and / or compositions can be used to assess whether a particular subject possesses a genotype associated with a positive response to a treatment.
[0169] Monitoring response to treatment (e.g., tumor mutations) The advent of next-generation sequencing (NGS) in genomic research has enabled the characterization of tumor mutational landscapes with unprecedented detail, resulting in the cataloging of diagnostic, prognostic, and clinically treatable mutations. Collectively, these mutations hold great promise for improving cancer outcomes through personalized medicine and potentially for early cancer detection and screening. Prior to this disclosure, a significant limitation in the field was the inability to detect these mutations when they are present at low frequencies. Clinical biopsies often consist largely of normal cells, making the detection of cancer cells based on DNA mutations a technical challenge even for modern NGS. Identifying tumor mutations among thousands of normal genomes is akin to finding a needle in a haystack, requiring a level of sequencing precision beyond previously known methods.
[0170] This problem is generally exacerbated in the case of liquid biopsies, where the challenge is not only to provide the extreme sensitivity necessary to detect tumor mutations but also to do so with the minimal amount of DNA typically present in these biopsies. The term "liquid biopsy" typically refers to blood, which has the ability to inform about cancer based on the presence of circulating tumor DNA (ctDNA). ctDNA, released into the bloodstream by cancer cells, shows great promise for monitoring, detecting, and predicting cancer, as well as enabling tumor genotyping and treatment selection. These applications could revolutionize the current management of cancer patients, but progress has been slower than previously anticipated. A major problem is that ctDNA typically represents a small fraction of all cell-free DNA (cfDNA) present in plasma. In metastatic cancers, its frequency can exceed 5%, while in localized cancers it ranges between 1% and 0.001%. Theoretically, by assaying a sufficient number of molecules, it should be possible to detect DNA subpopulations of any size. However, a fundamental limitation of previous methods is the high frequency of misscored bases. Errors often occur during cluster generation, sequencing cycles, low cluster resolution, and template degradation. As a result, approximately 0.1–1% of sequenced bases are incorrectly called. Further problems can arise from polymerase errors and amplification bias during PCR, which can lead to population distortion and the introduction of false mutant allele frequencies (MAFs). Taken together, known techniques, including conventional NGS, are unable to perform at the level required for low-frequency mutation detection.
[0171] Several approaches have been adopted to improve the accuracy of NGS. Removal of DNA damage using in vitro repair kits has been shown to reduce the number of incorrect variant calls in NGS. However, not all mutagenic lesions are recognized by these enzymes, and repair fidelity is not complete. Another approach that has gained significant traction is to utilize PCR overlaps resulting from individual DNA fragments to form a consensus. Referred to as "molecular barcoding," reads sharing unique random shear points or exogenously introduced random DNA sequences before or during PCR are grouped together, and the most common sequence is retained. Kinde et al. introduced this concept with SafeSeqS, which uses single-stranded molecular barcoding to reduce the sequencing error rate by grouping PCR copies that share barcode sequencing and forming a consensus. This approach resulted in an average detection limit of 0.5% and has been successful in detecting ctDNA in metastatic cancers, but only in approximately 40% of early-stage cancers. This detection limit can be significantly improved with digital droplet PCR (ddPCR), which can detect mutations as low as approximately 0.01% with MAF. However, mutations must be known in advance, which severely limits their application across multiple cancers. Furthermore, only 1–4 mutations can be tested at a time, making high-throughput screening impossible (Table 2). JPEG0007821756000002.jpg94159
[0172] Prior to this disclosure, the only technique with sensitivity comparable to ddPCR but that does not require a priori knowledge of tumor mutations is DS. DS extends the concept of molecular barcoding by using double-stranded molecular barcodes to exploit the fact that two DNA strands contain complementary information. We previously demonstrated that this approach yields unparalleled sensitivity of less than 0.005% of human nuclear DNA.
[0173] Due to the high accuracy of DS, SPLiT-DS, and CRISPR-DS, methods to improve the efficiency of DS, SPLiT-DS, and CRISPR-DS sequencing, as well as the translation and workflow of these sequencing platforms, hold promise in the field of oncology. As described herein, the provided methods and compositions enable an innovative approach to DS methodology that integrates double-stranded molecular tagging of DS with target sequence-specific amplification (e.g., PCR) to increase efficiency and scalability while maintaining error correction.
[0174] In addition to the need for highly accurate and efficient assays, clinical laboratory realities also demand assays that are fast, scalable, and reasonably cost-effective. Therefore, various embodiments of this technology (e.g., enrichment strategies for DS) that improve workflow efficiency for DS are highly desirable. As described herein, amplification-based enrichment and digestion / size-selection enrichment of specific target sequences for DS applications offer high target specificity, performance at low DNA input, scalability, and minimal cost (typically around $2-$3 / sample).
[0175] Some embodiments of the provided methods and compositions are particularly important in the field of cancer research in general, and ctDNA in particular, because the techniques developed herein have the potential to identify cancer mutations with unparalleled sensitivity while minimizing DNA input, preparation time, and cost. SPLiT-DS and CRISPR-DS, among other embodiments disclosed herein, may be useful in clinical applications where improved patient management and early cancer detection can significantly increase survival rates. [Example]
[0176] Example 1: SPLiT-DS SPLiT-DS is a PCR-based target enrichment strategy compatible with the use of molecular barcodes on each strand for duplex sequencing error correction (Figure 4A). In this exemplary embodiment, to begin SPLiT-DS analysis, one or more DNA samples are fragmented using one or more approaches (similar to previously described duplex sequencing library construction known in the art). Fragmentation is followed, most commonly by end repair and 3'-dA-tailing, followed by ligation of each DNA fragment with a T-tailed DS adapter containing a degenerate or semi-degenerate double-stranded barcode (Figure 4, step 1). Alternatively, other types of ligation overhangs, blunt-end ligation, or adapter ligation chemistries previously described in International Patent Publication No. WO 2017 / 100441 and U.S. Patent No. 9,752,188, can be used. Virtually all duel-matched DNA molecules are PCR-amplified using primers specific to the universal primer binding site of the single-stranded adapter tail, resulting in multiple barcode copies of DNA fragments ("barcode fragments") derived from each strand (Figure 4, step 2). After removing reaction byproducts, a given sample is split into two separate tubes (Figure 4, step 3) (i.e., the sample is split in half, with each tube containing approximately half the sample's contents). On average, half of the copies of a given barcode fragment are transferred to each tube, but randomness associated with splitting the sample can result in variance in the distribution of a given barcode fragment. To account for such variance, we use the hypergeometric distribution (i.e., the probability of selecting k barcode copies without replacement) as a model to determine the minimum number of PCR copies of a given barcode required to achieve a reasonably high probability that each tube contains at least one barcode fragment derived from each of the two (i.e., both) DNA strands of the original duplex. According to the hypergeometric model, more than four PCR cycles in step 1 (i.e., 2E4 = 16 copies / barcode) would likely provide a greater than 99% probability that each barcode fragment (from each strand) would be represented at least once in each tube.This assumes uniform, near-100% PCR amplification efficiency, which may not be realistic in all scenarios, but is a reasonable assumption for relatively low-input, high-quality DNA samples (e.g., 10 ng of human genomic DNA per 50 μL of PCR). After splitting the sample into two tubes, target loci are enriched by multiplex PCR using adapter sequences and primers specific to the loci of interest (Figure 4, step 4).
[0177] Multiplex locus-specific PCR is performed so that the PCR product obtained in each tube is derived from only one of the two original strands of a given DNA molecule sample. This is achieved by following the procedure described herein, using a sample split into two tubes (Tube 1 and Tube 2). In the first tube, PCR is performed using a primer (i.e., Illumina P5) specific for hybridizing to the "read 1" adapter sequence (Figure 4, Step 3, gray arrow) and a primer (i.e., Illumina P7) specific to the locus of interest and tailed with the sequence of the read 2 adapter sequence (Figure 4, Step 3, black arrow with gray tail). Alternatively, this tail can be shortened to not include the complete P7 sequence, or it can be added in a later PCR prior to sequencing. This step proposes that amplification products containing one P5 and one P7 sequence at each end will arise only from DNA derived from one strand of the original parent DNA molecule (i.e., the initial sample DNA). A similar reaction is repeated in a second tube, either sequentially or simultaneously, to amplify products originating from the opposite strand of the same genomic location compared to the amplification of the sample in the first tube. This is achieved by using locus-specific primers (i.e., anti-reference vs. reference sequences) tailed with the opposite universal primer sequence (i.e., P5 instead of P7) and an adapter primer to the opposite universal primer sequence (i.e., P7 instead of P5), which anneal in the opposite strand orientation as in tube 1. Data are analyzed using an approach similar to that used in conventional duplex sequencing analysis / library construction, whereby reads sharing a specific barcode derived from either the "original first strand" or the "original second strand" are grouped into a single-strand consensus sequence.
[0178] These single-stranded consensus sequences ("SSCSs") are then compared to the consensus calculated for the other original strand (e.g., the opposite strand as described herein). The identity of a nucleotide position is retained only if the resulting sequence at the same position is complementary to the two SSCSs from each of the original strands of the duplex. If the identity of a position does not match in the SSCSs, this is noted. For nucleotide positions that match between the paired SSCSs, the identity of this position is detailed in the final duplex consensus sequence (i.e., forming the DCS) (Figure 1C). Positions where the sequence identity between the two SSCSs does not match are flagged as potential error sites and are typically ignored by marking the position as unknown (i.e., "N"). Alternative strategies previously described in International Patent Publication No. WO 2017 / 100441 and U.S. Patent No. 9,752,188 include either disregarding the entire consensus read if a mismatch is found, or using a statistical approach to assign confidence to one variant versus the other, determining which is more likely to be the true variant based on how well a given SSCS is represented and how closely they match, given the prior probability of a particular type of error and the number of families that comprise it. Another approach is to retain uncertainty at the nucleotide position, for example, using IUPAC nomenclature (e.g., "K" representing a position that is either G or T). Additional information can be applied to the consensus sequence data file to reflect the relative likelihood of discrimination of one nucleotide versus another uncertain position, based, for example, on the prior probability of a particular type of sequencer, the relative number of reads supporting each variant at that position in a given sequence context or each pair of consensus families, or the raw red read quality score containing the SSCS family.
[0179] The duplex consensus calling approach is substantially similar to that described in International Patent Publication No. WO 2017 / 100441 and U.S. Patent No. 9,752,188, except that in SPLiT-DS, a single molecular identifier sequence at each end of the molecule is typically used to identify individual molecules (rather than one at each end), and sequence reads from one copy of the original strand may be found in one tube, while the complementary original strand may be found in the other tube. However, this need not be the case: as described herein, PCR reactions for duplex amplification libraries can be split into more than two tubes (e.g., four tubes with one specific primer pair in each tube), and the above process can be performed on both ends of the original molecule, creating two duplex consensus sequences per molecule. Similarly, initial PCR reactions can be split into multiple tubes (Figure 10) to generate multiple reads for duplex sequencing error correction and / or subassembly of long sequences with short read sequences.
[0180] It is often convenient to assign different indexes to the products in each tube to distinguish them after multiplex sequencing; however, this is not required. One advantage of SPLiT-DS is its ability to achieve targeted enrichment using PCR, which speeds up the workflow of previous versions of duplex sequencing or other approaches that rely on hybrid capture to enrich for regions of interest. At the same time, duplex adapters and tags can be used to achieve maximum precision, which is not achievable with traditional amplicon sequencing.
[0181] Example 2 Development of SPLiT-DS for CODIS STR loci This example is based on the insight that currently available methods for genotyping repetitive regions of DNA, such as short tandem repeats (STRs), would benefit from improved accuracy and sensitivity. This example extends and improves an established protocol for DS (which itself can eliminate "stutter"; Figure 3B) to create the "SPLiT-DS" assay / protocol. The current example demonstrates (1) the design and subsequent selection of primers for use in multiplex PCR, (2) methods for improving DNA library preparation, (3) evaluation of the accuracy, precision, sensitivity, and specificity of the provided techniques, e.g., using reduced DNA amounts, and (4) demonstration of significantly reduced stutter in the final error-corrected data.
[0182] Primer design and selection for multiplex PCR SPLiT-DS PCR primers are preferably designed to have the following properties: 1) high target specificity, 2) multiplexable, and 3) robust and minimally biased amplification. While many existing primer mixtures exist that meet these criteria for use in conventional PCR-capillary electrophoresis (PCR-CE), identical primer mixtures are unreliable for MPS. To this end, available data—mapping coordinates (i.e., the 5' end of each read in paired-end sequencing data corresponds to the 5' end of the PCR primer used to amplify the DNA)—from sequencing data obtained using commercially available kits that amplify target loci prior to sequencing were utilized to develop primers for use in this example. The insights described herein and data from previous examples are used to inform the design of initial primer sets for the extended CODIS core locus (CODIS20), as well as PentaD, PentaE, and SE3329 (collectively referred to as the CODIS loci for brevity, unless otherwise indicated). Because the mapping coordinates determined so far do not provide other information about the primers used in commercial (or other) kits, such as length, melting temperature, or concentration, primer development in this example focused on designs that maximize the probability of achieving uniform, robust, and specific amplification before multiplexing the reactions.
[0183] Results can be analyzed by direct sequencing (e.g., Illumina MiSeq platform) as opposed to gel analysis. To design optimal primer mixtures, each sample can be evaluated on several criteria. These criteria include: 1) specificity (i.e., the number of on-target reads divided by the number of off-target reads), 2) allele coverage of heterozygous loci (i.e., low-depth alleles divided by high-depth alleles; ideally 1.0), 3) interlocus balance (i.e., the lowest-depth locus divided by the highest-depth locus; ideally 1.0), and 4) depth variation (i.e., the average depth of each locus divided by the total average depth of all loci). Based on these criteria, at least one primer set can be selected for further analysis and development. Alternatively and / or additionally, primer design can involve the use of a web-based program, such as Primer3, for each STR marker.
[0184] Example 3: Improvement of library preparation method The library preparation protocol for SPLiT-DS follows known standard protocols, such as duplex sequencing protocols, until the first PCR step is complete. This example improves and extends this protocol by improving the locus-specific PCR unique to the SPLiT-DS technology provided herein, particularly the steps that occur after the initial duplex sequencing PCR step.
[0185] As a reference point, reactions are first run using known buffers, primer pool concentrations, and PCR conditions (e.g., as in standard DS protocols), but as applied to the SPLiT-DS approach, an initial duplex sequencing PCR is performed, and in some cases, other forms of targeted enrichment, such as hybrid capture, are performed. The effectiveness of these conditions in multiplex PCR is determined by directly sequencing the reactions on an Illumina MiSeq platform to monitor specificity, heterozygous locus allele coverage, inter-locus balance, and depth. This assay assesses PCR effectiveness (e.g., without error correction) and uses approximately 100,000–500,000 reads per condition, allowing for the analysis of at least 50 PCR conditions per sequencing run.
[0186] In this particular example, a successful analysis requires obtaining an average of 3-10 sequenced PCR copies (i.e., barcode families) from each starting DNA molecule. In other embodiments, a successful analysis may be defined as recovering one or more copies of each original DNA strand of a particular duplex molecule. It is believed that exceeding 3-10 copies may decrease assay efficiency in terms of sequencer resource usage, even without additional useful data. Too few average copies of each strand may not meet the defined criteria for a successful analysis, ultimately resulting in reduced depth. In some embodiments, defining a successful analysis as achieving a minimum number of sequenced copies of each strand may promote duplex sequencing with higher accuracy than duplex sequencing requiring fewer minimum copies per original strand.
[0187] Because SPLiT-DS is a unique approach compared to other currently available techniques, it cannot rely on known conditions for DNA input (e.g., conditions known from other assays), and therefore determines the DNA input amount used in PCR that occurs after splitting, as any change (e.g., reduction) in the amount of input until the first PCR step will inevitably affect the depth of post-processing.
[0188] After the DNA input range has been determined, use a qPCR-based assay to quantify the absolute amount of adapter-ligated target DNA (e.g., similar to step 3 in Figure 4).
[0189] Accuracy, precision, sensitivity, and specificity with decreasing DNA input The accuracy, precision, sensitivity, and specificity of commonly used standard reference material (SRM) DNA are used as a reference point for the improved technique described herein. SPLiT-DS is performed with decreasing amounts of input DNA (i.e., sensitivity) using serial dilutions (e.g., within the range of approximately 50 pg to approximately 10 ng) (e.g., to assess the accuracy and precision of the approach). At least six different libraries are prepared separately for each DNA input. After sequencing and error correction (using in-house software developed and designed specifically for the SPLiT-DS variant of duplex sequencing), STrait Razor is used to assess accuracy for: (i) genotyping the processed data, and / or (ii) determining the proportion of reads showing the "corrected" genotype at each CODIS locus (i.e., known from the standardized sample). Precision is assessed by determining: (i) allele coverage of heterozygous loci, (ii) interlocus balance, (iii) depth variation, and / or (iv) percent stutter (e.g., quantification of sample-to-sample variation).
[0190] Detection of contaminating DNA This example also focuses on improving currently available DNA evaluation methods for detecting contamination of a given sample with exogenous DNA (e.g., human forensic DNA contaminated with non-human DNA). SPLiT-DS analysis is performed on human DNA samples in the presence of contaminating DNA (mouse, canine, bovine, avian, Candida albicans, Escherichia coli, Staphylococcus aureus, etc.). The analysis includes sample DNA spiked with 10 ng of contaminating DNA, replicated three times at the following ratios: 50:50, 10:1, and 100:1 (contaminant:sample DNA, by weight), as well as a 100:0 control (i.e., no human DNA) and a 0:100 (unspiked human DNA). Each successfully generated library is sequenced and mapped to the reference genome and the corresponding contaminant in the human genome (GRCh38). This mapping is used to determine the percentage of reads that exhibit the correct genotype (e.g., consistent with the reference genome) at each locus and compare it to the control value. The alignment provides information about the range of contaminating DNA that will still allow SPLiT-DS to succeed (i.e., the level of contaminating DNA that can be present without adversely affecting the accuracy or robustness of SPLiT-DS).
[0191] Example 4: Validation of SPLiT-DS on single source samples. To validate SPLiT-DS as a viable high-precision genotyping method in a representative human population, DNA purified from cells obtained from the Personal Genome Project (PGP) is used (e.g., see PGP demographic details in Table 3). JPEG0007821756000003.jpg72120
[0192] The ability of SPLiT-DS to correctly genotype DNA single-source samples is evaluated. SPLiT-DS is performed in duplicate on DNA purified from cell lines of unrelated individuals from PGP. DNA from approximately 110 unique individuals is tested. SPLiT-DS is performed using the appropriate amount of DNA determined in the previous example (i.e., the minimum amount that reliably (e.g., greater than 80%) produces a sequencing library with an average post-processing depth of greater than 60x for each locus). Sequencing is performed using the in-house SPLiT-DS software described herein, and error correction is performed before genotyping the samples using STrait Razor.
[0193] As an interpretation guideline for genotyping of SPLiT-DS data, a modified "consensus" approach of two replicates is used, as follows:
[0194] No result: when at least one (e.g., one of two) replicates produces low inclusion (e.g., less than 60-fold);
[0195] Correct genotype: When all (e.g., 2 out of 2) replicates produce the expected genotype (i.e., match the genotype in the WGS data for a given sample).
[0196] Undefined genotype: when different genotypes are obtained in all replicates (e.g., two out of two) at a given locus, or when only one genotype differs from the WGS data.
[0197] Incorrect genotype: When all (2 out of 2) replicates show the same incorrect genotype.
[0198] The amount of stutter is quantified across all samples and loci by determining the stutter ratio for sequenced loci. The stutter ratio is calculated by dividing the number of reads for a given stutter allele by the number of reads for that allele in the actual sample. If more than one type of stutter event is observed, calculations for each stutter length are performed. To minimize bias in this analysis, stutter rates are calculated only for loci with an average depth of 60x or greater (80% power to detect one or more post-processing reads containing alternative stutter alleles, which occurs at 5% (1-Sample Binomial Test)). If consistent high-depth coverage is obtained for at least some loci, low-frequency stutter events can be examined and ratios calculated appropriately (e.g., adjusting power).
[0199] Another part of the analysis in this example involves the effect of STR length on various parameters, including comparing the results at a given locus with reference STR length (e.g., specificity, allele coverage at heterozygous loci, interlocus balance, and / or depth). Evaluation of these parameters may improve the interpretation of STR length-based polymorphisms (e.g., including the fact that the SPLiT-DS samples evaluated are generally from outbred populations and may, for example, harbor various STR length polymorphisms). In addition to evaluating the effect of STR length, stutter ratios are also determined. Finally, a calculation of the discriminatory power of each sample (based on loci correctly genotyped according to the guidelines described herein, using, for example, expected allele frequencies in the U.S. population) is performed.
[0200] The results of the analysis described in this example will determine the extent of use of SPLiT-DS (and the extent of any bias in the method), e.g., for various types of samples and / or genotyping of STRs.
[0201] Comparison and agreement study of capillary electrophoresis and MPS approaches For example, to demonstrate the superiority of SPLiT-DS as a sequencing method for forensic applications, a concordance study will be conducted against currently available methods. Currently, the "gold standard" for forensic STR genotyping is PCR-CE. SPLiT-DS results obtained according to the examples described herein will be compared to the same DNA sample genotyped using PCR-CE analysis and 1 ng of input DNA according to standard procedures. The two data sets (PCR-CE and SPLiT-DS, accompanied by appropriate controls / references (e.g., WGS PGP sample data)) allow the level of concordance between the two approaches to be determined. A concordance study will also be performed using a commercially available kit (e.g., the Illumina FORENSEQ DNA Signature Prep Kit), which uses targeted PCR amplification of 63 STRs and 95 discriminatory SNPs, including CODIS loci. The same samples used in the PCR-CE and SPLiT-DS concordance study will be used, and genotyping will be performed using the STRait-Razor. PCR stutter was also assessed for each approach (PCR-CE, commercial kit, SPLiT-DS), with stutter calculated if the true allele peak height was at least 600 RFU (probability threshold) but not more than 15,000 RFU. Positions two repeat units apart were not included to eliminate any additive effects of positive and negative stutter at repeat positions between heterozygous alleles. As described herein, the stutter percentage was calculated by dividing the peak height of the stutter peak by the peak height of the true allele. For samples analyzed by commercial kits, all alleles with 60 or more observed reads were called, and the stutter percentage was calculated as described herein. Comparisons were made between the stutter percentages for each locus tested. While stutter results between platforms are not directly comparable to each other, the data are believed to provide a reasonable estimate of the relative abundance of stutter in each method.
[0202] Example 5: Validation of SPLiT-DS on damaged DNA and DNA mixtures. Highly damaged / degraded DNA and mixtures confound currently available genotyping techniques. Therefore, this example demonstrates the ability of SPLiT-DS to correctly genotype samples containing damaged DNA and DNA mixtures, improving and extending currently available methodologies.
[0203] Validation of SPLiT-DS on damaged DNA from a single contributor SPLiT-DS is performed on sample DNA exposed to three forensically relevant categories: (i) chemical exposure, (ii) ultraviolet (UV) light, and (iii) elevated temperature (see Table 4 for a summary of exemplary exposure methods / conditions known to affect conventional STR analysis / used in previous studies). Due to the lack of SRM available for damaged DNA samples, the level of induced damage is standardized across biological replicates. DNA is first exposed to the environmental conditions and time points listed in Table 4, and then assessed using a commercially available kit used to determine DNA damage / degradation in a given sample (e.g., the KAPA Biosystems hgDNA Quantification and QC qPCR kit (Roche / KAPA Biosystems)). Only samples that exhibit comparable levels of damage (defined as within one standard deviation of the observed mean) for specific environmental conditions (as determined by the assays described herein) are used for this analysis.
[0204] Experiments evaluating SPLiT-DS with damaged / degraded DNA are performed in triplicate on Promega 2800M SRM DNA using the minimum amount of input DNA required to consistently (>50%) form libraries that can be sequenced using SPLiT-DS under the most stringent conditions possible for each category in Table 4 (determination of such amounts made as described herein). Conditions that do not produce consistent libraries will be considered to define the limit of sensitivity of SPLiT-DS to damaged / degraded DNA. No such libraries will be evaluated. JPEG0007821756000004.jpg101138
[0205] Samples are sequenced using 300-bp paired-end reads on an Illumina MiSeq platform, and data genotypes determined using STrait Razor are processed using the custom SPLiT-DS software described herein. Experimental conditions that result in inaccurate genotyping (as described in the previous example) are considered to define the limits of accuracy for SPLiT-DS for damaged / degraded DNA. Calculations are performed to determine specificity, allele coverage of heterozygous loci, and / or depth of each locus in damaged / degraded DNA, and results are compared to undamaged controls.
[0206] Because the relative performance of SPLiT-DS on high-quality DNA is not necessarily directly translatable to its performance on damaged DNA, comparisons are also performed using SPLiT-DS, standard PCR-CE, and MPS methods. These methods are performed using the 10 PGP samples genotyped in the previous examples, subjected to the most challenging conditions (determined by the results) for each category of damage for the successfully genotyped SPLiT-DS samples. Samples are genotyped by PCR-CE and conventional MPS using appropriate commercially available kits, as described in the previous examples. The relative performance of SPLiT-DS to PCR-CE and MPS is determined as described herein, including determining and comparing the relative amounts of stutter, allele dropout, intra-allelic balance, and genotyping success rates between approaches. SPLiT-DS may produce more sensitive and accurate results using smaller samples and / or samples with more damaged / degraded DNA than can be achieved with other methods.
[0207] Validation of SPLiT-DS in mixtures. This demonstrates the improved effectiveness (e.g., improved accuracy and sensitivity compared to available methods) of SPLiT-DS analysis of DNA mixtures consisting of two genetically unrelated individuals with a wide range of MAF ratios. For each mixture in Table 5, 10 pairs of pairs are selected from the PGP samples genotyped in the previous example. The specific PGP samples used in this example depend on the specific genotype determined in the previous example or by whole genome sequencing (available as part of the PGP). If possible, contributor pairs with at least two different repeat lengths at more than eight loci are selected. It is likely that more than 10 ng of DNA will be required from each sample. The exact amount will depend on how efficiently SPLiT-DS works at each locus, as determined in the previous example. JPEG0007821756000005.jpg67118
[0208] The input amount of DNA is adjusted so that any minor contributors are represented by at least 10 reads. Representation with at least 10 reads is considered to provide a greater than 95% chance of detecting both alleles at all CODIS loci. As shown in previous examples, the specific amount required to achieve 10 MAF reads depends on the sensitivity limit of SPLiT-DS.
[0209] To minimize inter-replicate variability, mixtures are constructed based on triplicate DNA quantification using the QUANTIFILER Duo DNA Quantification Kit (Thermo Fisher). Samples are sequenced on an Illumina MiSeq platform, data are processed using custom SPLiT-DS software, and genotypes are determined using STrait Razor. Evaluating the presence of stutter in these experiments helps evaluate the performance of SPLiT-DS on DNA mixtures. For each locus analyzed in each mixed sample, Wilson score intervals (in the form of binomial proportional confidence intervals) of known MAFs are calculated. The number of stutter events that differ by one repeat length from the known MAF of the mixture is counted. If the stutter read count is within the 95% Wilson score interval of one of the MAF alleles, the locus is considered a partial match. If both MAF alleles fail this test, the locus is considered a failed genotype call (if the MAF cannot be distinguished from the stutter, the homozygous allele automatically fails). As in the previous examples, comparative studies of SPLiT-DS against PCR-CE and MPS, as well as comparisons of the relative amounts of stutter, allele dropout, intra-allelic balance, and / or genotyping success rates, will also be performed and evaluated as described herein. The results of the two-person mixture experiment will then be used to perform a three-person mixture experiment (see, for example, Table 5) using the same sample selection criteria and analysis as the two-person mixture analysis.
[0210] SPLiT-DS is performed using single-source and two-person mixture mock casework samples, using DNA provided by the Washington State Patrol Forensic Laboratory Services Bureau from previously analyzed commercial forensic DNA proficiency tests. Genotyping using SPLiT-DS is compared to the online submitted consensus results for the samples.
[0211] Example 6: Improved performance of SPLiT-DS on damaged DNA samples Formalin fixation causes extreme DNA damage in the form of cytidine deamination, oxidative damage, and crosslinks. To demonstrate the functionality of SPLiT-DS compared to currently available methods, we performed an analysis of highly damaged DNA by sequencing formalin-fixed nuclear DNA at the D3S1358 locus on a Promega 2800M SRM (Figures 13B and 14A). Figures 13A-13C show data obtained from the SPLiT-DS procedure according to one embodiment of the present technology. Figure 13A is a representative gel showing insert fragment sizes before sequencing (lane 1 is the ladder, and lanes 2 and 3 are samples of PCR products from each tube; see, for example, step 4 in Figure 4). Figures 13B and 13C are graphs showing CODIS genotypes versus the number of sequencing reads in the absence of error correction (Figure 13B) and after analysis by SPLiT-DS (Figure 13C). Figure 13B shows a sample (D3S1358) in which the polymorphism was observed in the absence of error correction, with the stutter events indicated by black arrows. Figure 13C shows a sample (D3S1358-DCS) without detectable stutter events after analysis by SPLiT-DS. The x-axis of each of Figures 13B and 13C shows the CODIS genotype, and the y-axis shows the number of reads.
[0212] Figures 14A and 14B are graphs showing the number of sequencing reads for highly damaged DNA in the absence of CODIS genotype (Figure 14A) and after analysis with SPLiT-DS (Figure 14B) according to an embodiment of the present technology. The x-axis of each panel shows CODIS genotype, and the y-axis shows the number of reads. Figure 14A shows a damaged DNA sample not analyzed by SPLiT-DS (D3S1358), showing stutter events (black arrows) and a significant amount of obvious point mutations (not shown). Figure 14B shows a sample analyzed with SPLiT-DS error correction (D3S1358-DCS), showing no detectable stutter events. No obvious point mutations were observed.
[0213] SPLiT-DS results showed that all PCR- and sequencing-based artifacts present using standard sequencing methods were eliminated using SPLiT-DS in formalin-exposed DNA (Figures 13C and 14B). A reduction in efficiency (approximately 3-fold) was noted in these samples (see, e.g., Figure 14B vs. Figure 13C), although the presence of interstrand crosslinks common to formalin fixation may contribute to this reduction.
[0214] Example 7: Targeted genome fragmentation This example demonstrates targeted genome fragmentation as a method to improve the efficiency of genomic DNA (gDNA) sequencing. SPLiT-DS genome fragmentation is typically achieved by methods such as physical shearing or enzymatic digestion of DNA phosphodiester bonds. Such approaches produce samples in which intact gDNA is reduced to a mixture of randomly sized DNA fragments. While highly robust, the variable size of DNA fragments can lead to PCR amplification bias (shorter fragments are more abundant) and uneven sequencing depth (Figure 11A), as well as sequencing reads that do not overlap with the region of interest within the DNA fragment. Therefore, this example uses CRISPR / Cas9 to overcome these issues. Cleavage sites are designed to produce fragments of predetermined, uniform sizes. It is considered that a more homogeneous set of fragments is unlikely to overcome the presence of bias and / or uninformative reads, which can affect the efficiency of other techniques that do not use targeted fragmentation. It is also believed that targeted fragmentation likely facilitates pre-enrichment of a given sample prior to library preparation, as fragment separation from gDNA likely allows for removal of large off-target regions due to consistency / differences in fragment size.
[0215] Example 8: SPLiT-DS for cancer monitoring and diagnosis While the presence of circulating tumor DNA in blood has been recognized for decades, ultrasensitive methods are needed for the reliable development of cancer biomarkers (e.g., markers for diagnosing and / or tracking the presence / progression of disease). SPLiT-DS helps overcome a wide range of challenges, including the low amount of circulating tumor DNA in blood samples containing varying amounts of cell-free DNA. SPLiT-DS also improves and extends several highly sensitive and specific methods known in the art, such as BEAMing, SafeSeqS, TamSeq, and ddPCR, because it does not require prior knowledge of specific mutations. SPLiT-DS provides an approach that can detect cancer-associated mutations with the highest level of accuracy currently available, even at low DNA input and without prior knowledge of specific tumor mutations.
[0216] In this example, SPLiT-DS is used to evaluate sequences associated with circulating tumor cell DNA. Control samples of known mutations are used and run alongside samples from patients diagnosed with and / or suspected of having cancer.
[0217] SPLiT-DS and genomic DNA or cell-free DNA SPLiT-DS is used to develop assays for accurate sequencing of low-input gDNA (10-100 ng) and cfDNA (approximately 10 ng). Genomic DNA generally occurs in large fragments (>1 Kb), while cell-free DNA occurs almost exclusively as infrequent fragments of approximately 150 bp.
[0218] Rationale for low input (10-100ng) gDNA This example demonstrates the feasibility of SPLiT-DS for low DNA input and its suitability for multiplexing. While tissue can be obtained from cancer patient biopsies, it is preferable to use such samples sparingly to complete all required testing. Therefore, gDNA sequencing would benefit from improved platforms, such as those provided by SPLiT-DS, which require low input material.
[0219] Each target in SPLiT-DS is individually designed and optimized. The genes TP53, KRAS, and BRAF are analyzed as proof-of-principle. Notably, each gene has known target regions where cancer-associated mutations occur. TP53 has 10 coding exons (relatively small in size), all of which are targeted using SPLiT-DS. KRAS has known mutation hotspots at codons 12, 13, and 61 in exon 2, all of which are targeted. BRAF has a V600E mutation targeted in exon 15.
[0220] material and method The SPLiT-DS assay is performed on gDNA using DNA from de-identified tumors with known clonal mutations in TP53, KRAS, and BRAF, as well as leukocyte gDNA from cancer-free individuals, as outlined in Figures 4 and 5. Two different sets of experiments are performed to test efficiency and sensitivity, following any optimization / validation steps.
[0221] efficiency Efficiency is defined as the percentage of input DNA molecules converted into DCS reads. In this example, the efficiency is at least 30%, but we aim for greater than 50%. It is believed that 10 ng of input DNA is likely to achieve an average 1000-fold DCS depth across the entire locus of interest (10 ng = approximately 3200 genomes, therefore 3200 x 0.3 efficiency = approximately 1000 genomes sequenced). Efficiency depends in part on the performance of multiplex PCR. Using an in silico approach, PCR primers are designed for: i) high target specificity, ii) the ability to multiplex, and iii) the ability to perform robust, minimally biased amplification.
[0222] The CRISPR / Cas9 system is used to specifically produce approximately 500-550 bp fragments containing a particular region of interest (see Figure 11C). After completing the design of guide RNAs and PCR primers, a combinatorial approach is used to achieve: (i) target specificity (i.e., the percentage of target reads; greater than 70% is acceptable); and (ii) inter-locus depth balance (i.e., the locus with the lowest depth divided by the locus with the highest depth; greater than 0.5% is acceptable). The optimized guide and primer pools are then applied to 10 ng and 100 ng of identical gDNA. These pools are used for all subsequent experiments involving gDNA.
[0223] sensitivity TP53-mutated tumor gDNA is spiked into control, non-mutated leukocyte gDNA at ratios of 1:2, 1:10, 1:100, 1:1000, and 1:10,000. Using two additional tumor DNAs containing known clonal mutations in KRAS and BRAF, respectively, the same mixing experiment is performed on a total of 15 samples (5 dilutions for each of the three genes). These 15 samples are processed with SPLiT-DS as described herein using 10 ng and 100 ng of input DNA. The "expected" and "observed" MAFs are compared (maximum MAF is MAF). max = α, where N is the number of genomes and a is the efficiency of SPLiT-DS, e.g., at 30% efficiency, MAF max is 0.1% for 10 ng of DNA and 0.01% for 100 ng of DNA).
[0224] MAF based on binomial distribution max It is considered likely that the probability of detecting a given mutation present in the genotype reaches 63%. Since there are three spiked mutations in the experiment, it is statistically likely that at least one will be detected at 0.1% and 0.01%, and efficiency increases this probability above 30%.
[0225] In addition to the spiked mutations, SNPs are used to confirm sensitivity because the normal control DNA is from a different individual than the tumor DNA. SNPs are examined at identical dilutions (homozygous SNPs) and effective dilutions of 1:4, 1:20, 1:200, 1:2000, and 1:20,000 (heterozygous SNPs).
[0226] CRISPR / Cas9 efficiently cleaved all TP53 exons, facilitating enrichment through size selection and maximizing read utilization. CRISPR / Cas9 guides were designed to cleave TP53 exons (see Figure 12A). 10 ng of gDNA was digested and processed using SPLiT-DS (see Figures 12B and 12C) as described in previous examples with appropriate PCR primers to amplify exons 5-6 and 7 (Figures 12C and 12D). Both strands of DNA were properly sequenced with a high percentage of on-target reads, and DCS reads were generated after matching complementary random tags on each molecule (Figure 12D). Furthermore, the average depth achieved with a starting amount of 10 ng of DNA corresponded to an efficiency of 25% (i.e., from the original 3000 genomes, approximately 800x was sequenced on average), representing a 50-fold improvement over standard DS and an unparalleled improvement compared to traditional solution hybridization approaches.
[0227] Example 9: Development of SPLiT-DS for accurate sequencing of cfDNA This example demonstrates the use of SPLiT-DS for the detection of mutations in the following exemplary cancer-associated genes: TP53, KRAS, and BRAF in cfDNA.
[0228] material and method Cell-free DNA from commercially available plasma (Conversant Bio) is extracted using the QIAamp Circulating Nucleic Acid Kit. Three different synthetic 150-bp DNA molecules encoding known mutations in each of the three genes of interest are used. Each of these synthetic DNA molecules is spiked into cfDNA at ratios of 1:2, 1:10, 1:100, 1:1000, and 1:10,000. Two different sets of experiments are performed to optimize and validate the SPLiT-DS protocol parameters for cfDNA.
[0229] efficiency Because the cfDNA is already fragmented, no cleavage (e.g., CRISPR / Cas9) is required. Therefore, SPLiT-DS is performed as described in the previous example, with the addition of nested PCR. The resulting fragments are sequenced on a MiSeq v3 for 150 cycles, with approximately 10 samples multiplexed onto the cartridge for approximately 2.5 million reads each.
[0230] sensitivity Five mixed dilutions (1:2, 1:10, 1:100, 1:1000, 1:10,000) of cfDNA for TP53, KRAS, and BRAF mutations, respectively, were analyzed by SPLiT-DS using the optimized primers designed in this example, starting with 10 ng and 100 ng of DNA. The experiment was performed in parallel with SafeSeqS to compare the sensitivity between the two technologies (SafeSeqS is a known technology for accurate sequencing of ctDNA, and it uses single-strand correction to reduce NGS errors). SPLiT-DS likely outperforms SafeSeqS in detecting mutations at MAFs of 0.1% and 0.01%. SPLiT-DS likely can detect spike mutations with an estimated average sensitivity of 0.5% (Table 2), whereas SafeSeqS cannot detect spike mutations at such low frequencies.
[0231] Primers (for the nested PCR approach) were designed to amplify codons 12 and 13 of KRAS exon 2. 10 ng and 20 ng of cfDNA extracted from normal plasma (Conversant Bio) were processed in parallel. Figures 15A and 15B visually represent SPLiT-DS sequencing data for KRAS exon 2 generated from 10 ng (Figure 15A) and 20 ng (Figure 15B) of cfDNA using nested PCR according to one embodiment of the technology. In this example, target enrichment was performed using SPLiT-DS, and sequencing was performed on an Illumina MiSeq with 75 bp paired-end reads. SSCs for both the "A" and "B" strands prior to duplex formation and the final DCS read are shown. Arrows indicate the two locus-specific PCR primers (gray primers = nested PCR primers).
[0232] As shown in Figures 15A and 15B, we found that "Side A" and "Side B," corresponding to two different strands of DNA, were properly amplified, and their complementary strands formed highly accurate DCS reads. Although the depth achieved was modest (approximately 50 reads), it corresponded to an efficiency of approximately 1%, which is the current efficiency of standard DS. Thus, at baseline (i.e., without optimization), SPLiT-DS performed at the same efficiency as currently used approaches, but with only 10 ng of input DNA, a much smaller amount than previously available, demonstrating improved efficiency over other approaches available for sequencing cfDNA.
[0233] Example 10: SPLiT-DS for ctDNA-based pancreatic cancer detection and prognosis. This example demonstrates the improved detection of mutations in ctDNA from pancreatic ductal adenocarcinoma (PDAC) patients using SPLiT-DS (compared to currently available methods). SPLiT-DS improves the sensitivity of ddPCR at multiple target genes, including KRAS, TP53, and BRAF. The results of these assays likely demonstrate improved sensitivity, detecting a single mutation in 95% of PDAC patients with current approaches and detecting two mutations in over 50% of PDAC cases.
[0234] Furthermore, because most DNA in the circulation of human subjects (i.e., circulatory (e.g., cell-free) DNA) is of hematopoietic origin, leukocyte DNA will be compared in sequence and mutation to those found in cfDNA. These results are proposed to inform whether a particular background mutation originates from a leukocyte subclone with greater sensitivity and accuracy than other results.
[0235] Materials and Methods Fully de-identified cfDNA and matched leukocyte DNA samples from 40 patients with PDAC, 20 patients with chronic pancreatitis, and 20 age-matched normal controls will be evaluated. Blood samples will be processed within 2 hours of extraction, providing specimens containing 2-5 ml of plasma and 500 μl of buffy coat. Additionally, for PDAC patients, frozen tumor sections will be available to confirm tumor mutations. Blood will be procured preoperatively for all PDAC patients. All patients will be clinically followed, and detailed clinicopathological information, including time to recurrence and mortality, will be available. Patient samples will include 20 patients with localized cancer and 20 patients with metastatic cancer.
[0236] ctDNA is extracted using the QIAamp Circulating Nucleic Acid Kit, and gDNA is extracted using the QIAamp DNA Mini Kit. At least 10 ng of cfDNA (from collected plasma), 100 ng of gDNA, and all available ctDNA (up to 100 ng) are processed using the appropriate SPLiT-DS procedure as described herein for KRAS, BRAF, and TP53. Sequencing is performed using the Illumina 150-cycle MiSeq v3 Reagent Kit for ctDNA and the 600-cycle MiSeq v3 Reagent Kit for gDNA. Ten ctDNA samples are multiplexed with the 150-cycle kit, and 15 gDNA samples are multiplexed with the 600-cycle kit. Based on the experimental design, we believe that a sequencing depth of at least 1,000x with 10 ng of DNA and as much as 10,000x with 100 ng of DNA is likely to be achieved, with an expected efficiency of at least 30%. Data are analyzed after sequencing, DCS generation, and mutation identification.
[0237] Pancreatic cancer detection This example determines the sensitivity and specificity of SPLiT-DS for detecting KRAS, TP53, and BRAF mutations in cfDNA from patients with PDAC. To analyze sensitivity, mutations found in cfDNA are compared with tumor mutations (clonal and subclonal) identified by SPLiT-DS. Because SPLiT-DS results encompass nearly all PDAC cases with one mutation and over 50% of PDAC cases with two mutations, it is likely to detect at least one tumor mutation in the cfDNA of all metastatic cases and approximately 80% of localized cases, resulting in a combined sensitivity of approximately 90% for all PDAC cases.
[0238] Mutations found in cfDNA are compared with those found in matched white blood cells purified from the same patient. The white blood cells that match the mutations found in cfDNA are considered biological background and are ignored in the final cfDNA mutation count. After subtracting the shared mutations, cfDNA mutations from PDAC, pancreatitis, and controls are compared. Even if biological background mutations (e.g., age-related mutations) remain in the sample, cancer mutations are likely to occur at a higher frequency than biological background mutations. Using the area under the curve and an age-adjusted ROC model, the optimal mutation frequency threshold is determined to distinguish cancer from controls with maximum sensitivity and specificity.
[0239] Pancreatic cancer prognosis As demonstrated in previous examples, the improved sensitivity of SPLiT-DS suggests that ctDNA is likely detectable in nearly all (90%) PDAC patients, in contrast to previously available approaches. Instead of a binary variable (i.e., yes / no) for the presence of ctDNA, ctDNA MAF is analyzed as a quantitative variable, comparing MAF scores with clinical data (e.g., comparing MAF scores with prognosis). It is also determined whether mutated genes, codons, and / or mutation types correlate with recurrence or death. Multivariate COX models adjusted for confounding factors (including age and stage) are used to test the ability of these variables and their combinations to predict disease-free survival and overall survival. Kaplan-Meier curves are used to represent the predictive value of categorical variables.
[0240] Example 11: SPLiT-DS for the identification of resistance mutations in metastatic CRC Detecting early cancer and predicting recurrence using ctDNA In metastatic CRC (i.e., stage IV), which accounts for approximately 50% of cases at presentation, tumor genotyping is essential to guide treatment decisions: oncogenic mutations in KRAS, NRAS, and BRAF occur in approximately 50% of CRC patients and predict lack of response to the EGFR monoclonal antibodies cetuximab and panitumumab. Therefore, these genes are routinely assessed in both fixed and unfixed tissue biopsies, but currently available approaches often result in low-quality subclonal resolution and suffer from sampling bias. As a result, tumors with subclonal mutations may be missed, resulting in some patients receiving treatments that are sure to fail. Therefore, this example demonstrates that tumor genotyping by ctDNA using SPLiT-DS provides an assay with improved sensitivity over currently available techniques, which also improves diagnosis and treatment by detecting pre-existing resistance mutations in SPLiT-DS that condition patient eligibility for EGFR blockade therapy.
[0241] Detecting and predicting the presence and / or recurrence of CRC SPLiT-DS is used with a panel of five commonly mutated CRC genes to demonstrate detection of mutations in ctDNA without prior knowledge of specific tumor mutations. Results from this assay are likely to inform future CRC detection using much simpler tests (e.g., blood tests).
[0242] This example also demonstrates improvements in methods used to detect and / or predict recurrence. Currently, available technologies are limited by a lack of sufficient sensitivity and / or specificity, or, for technologies with sufficient sensitivity / specificity, they are prohibitively expensive. Thus, SPLiT-DS analysis of ctDNA demonstrates improved detection and prediction of CRC recurrence, providing improved accuracy (e.g., more than 100 times that of SafeSeqS, for example), and the ability to expand and evaluate multiple genes.
[0243] Materials and Methods This example uses patient samples from multiple biopsy types from over 300 patients who underwent surgical resection of their tumors. Available biopsies include tumor, plasma, and buffy coat. Patients from whom samples were obtained were followed longitudinally, with blood samples available at 6, 12, and 24 months after baseline resection. Detailed clinicopathological information, including recurrence, is available for all patients. All samples and coded medical information are fully anonymized. Samples from patients with metastatic disease were previously evaluated for KRAS and NRAS mutations to determine the likelihood of response to cetuximab or panitumumab. If no mutations were found, targeted therapy was administered. Resistance was documented via ongoing imaging studies.
[0244] Samples from 20 patients with metastatic cancer (stage IV) and 40 patients with localized cancer (stages I-III) will be evaluated. DNA will be purified from plasma (2-5 ml) and buffy coat samples obtained before surgery, as well as from frozen tumor samples. Patients classified as having metastatic cancer are those who tested negative for KRAS and NRAS mutations but did not respond to EGFR inhibitor therapy. At least 10 recurrent patients will also be included. ctDNA will be measured in blood collected 6, 12, and 24 months after surgery. As in previous examples, leukocyte DNA mutations will be used to identify potential biological background mutations that may be present in cfDNA.
[0245] Furthermore, APC is the most commonly mutated gene in CRC, and the SPLiT-DS panel used in this example includes the most commonly mutated region of APC, such as the mutation cluster region (299 bp) spanning from codon 1,286 to codon 1,585, which encompasses approximately 60% of CRC mutations in APC52. Additional top hits found in COSMIC total approximately 1,000 bp. NRAS codons 12, 13, and 61 are also included. Therefore, the panel used in this example includes APC (approximately 1,000 bp), TP53 (1,182 bp coding region), KRAS (codons 12, 13, and 61), BRAF (V600E), and NRAS (codons 12, 13, and 61), for a total size of approximately 2,700 bp. It is believed that the panel described in this example is likely to encompass all CRC samples, including a subset containing one mutation and two mutations.
[0246] Identification of resistance mutations in metastatic CRC SPLiT-DS is used to evaluate samples from metastatic CRC for clonal tumor mutations in cfDNA. All tumors are negative for KRAS and NRAS mutations but may harbor at least one clonal mutation (in APC or TP53) identified by the panel described in this example. SPLiT-DS is also used to determine whether the presence of very low-frequency (<0.1%) mutations in ctDNA that confer resistance to EGFR therapy is detectable. Samples from patients with metastatic disease are likely to be successfully sequenced at very high depths (approximately 10,000x). SPLiT-DS analysis also improves the detection of low-frequency KRAS, BRAF, and NRAF mutations in ctDNA from patients with metastatic disease who fail EGFR therapy but whose tumor DNA is negative for KRAS and NRAS by Sanger sequencing. SPLiT-DS is used to sequence tumor DNA at a similar high depth to determine the presence or absence of primary resistance mutations in ctDNA. Results are compared between ctDNA and DNA from intratumoral tissue.
[0247] Localized CRC detection SPLiT-DS is used to identify ctDNA in localized (stage I-III) cancer samples using the panel of five CRC genes described herein. Tumor DNA is also sequenced using SPLiT-DS. As described in previous examples, the presence of biological background mutations from white blood cells is also determined.
[0248] While certain currently available methods (e.g., CEA) offer an estimated 1.5–6 months of lead time compared to other methods for detecting recurrence, it is unclear whether such time impacts survival. Other techniques may improve lead time but require a priori knowledge of the tumor genotype. Therefore, we demonstrate the superior ability to sequence ctDNA using SPLiT-DS, improving lead time by several months, as described herein, without requiring prior knowledge of tumor genotype. In this example, we demonstrate the ability of SPLiT-DS to detect ctDNA at 6, 12, and 24 months after initial surgery in patients with localized CRC who experienced recurrence. Ten patients are selected based on having a recurrence whose tumor and baseline ctDNA harbor at least one mutation (ideally two) in the genes of the aforementioned panel. For each sample (individual), longitudinal clinical history (chemotherapy, CT scans, and other indicators of recurrence) is plotted against the total ctDNA levels for each mutation at baseline, 6, 12, and 24 months. Comparison of ctDNA and CEA in CEA levels and lead time to recurrence will also be evaluated.
[0249] Example 12: CRISPR-DS This example describes the creation of CRISPR-DS for highly accurate and sensitive sequencing. CRISPR-based technology was used to excise engineered target regions of predetermined, uniform length (Figure 12A). In this example, the CRISPR-compatible nuclease used was Cas9. This size control was used to facilitate size selection prior to library preparation (Figure 12B), followed by double-stranded barcoding (Figure 12C) to perform error removal (similar to the previously described DS method) (Figure 12D). Following barcoding, a single capture was performed (in contrast to other available methods), resulting in extremely high on-target enrichment and the production of fragments encompassing the entire sequencing read (Figures 12F and 16A). Fragmentation of hybridization capture is typically performed by sonication, which often produces sequencing reads that are too long and do not overlap the region of interest, and / or that are too short and overlap each other, rereading the same sequence (Figures 12F and 16A). Figures 16B and 16C are histogram graphs showing fragment insert sizes for samples prepared with standard DS and CRISPR-DS protocols according to embodiments of the present technology. The x-axis represents the percentage difference from the optimal fragment size, e.g., the fragment size that matches the length of the sequencing read after adjusting for molecular barcodes and clipping. The vertical column indicates the range of fragment sizes that fall within 10% of the optimal size, with the optimal size designated by the vertical hash line. As shown in Figures 16B and 16C, sonication resulted in significant variation in the amount of deviation from the optimal fragment size (Figure 16B), while CRISPR / Cas9 digestion yielded fragments with the majority of reads within the optimal fragment size (Figure 16C).
[0250] This example demonstrates how CRISPR-based fragmentation can prevent spurious mutations, including the fact that the enzyme Cas9 used in this example produces blunt ends and therefore does not require end repair.Therefore, the technology provided herein overcomes multiple common and widespread problems of NGS, including inefficient target enrichment, sequencing errors, and uneven fragment sizes.
[0251] Guide RNAs (gRNAs) were designed to excise the TP53 coding region and adjacent intronic regions (Figure 12A). Fragment size was set at approximately 500 bp. gRNAs were selected based on specificity scores and fragment lengths (Table 1, Figures 17A-17C). Test samples containing variable amounts of input DNA (10-250 ng) were digested with CRISPR / Cas9, followed by size selection with solid-phase reversible immobilization (SPRI) beads to remove undigested high-molecular-weight DNA and enrich for excised fragments containing the targeted region (Figure 12B). Subsequent library preparation was performed according to currently available standard protocols, but with only a single capture and minor modifications, as described herein. DNA was A-tailed, ligated with DS adapters, amplified, purified by bead washing, and captured by hybridization with a biotinylated 120 bp DNA probe targeting a TP53 exon (Table 6). The captured samples were amplified with index primers and sequenced on an Illumina MiSeq v3 600-cycle kit. Analysis was performed as per the standard protocol, but modified to include generation of a consensus sequence prior to alignment (Figure 23). JPEG0007821756000006.jpg73170
[0252] A side-by-side comparison of standard DS with one or two hybridization captures and CRISPR-DS with one hybridization capture is shown in Figures 18A-18C. Figures 18A-18C are a bar graph (Figure 18A) showing the percent of on-target raw sequencing reads (including TP53) and the median double-stranded consensus sequence depth across all target regions for various input amounts of DNA processed using standard DS and CRISPR-DS (Figure 18C). Figure 18A shows the percentage of on-target raw sequencing reads (including TP53) between Standard-DS with two captures and CRISPR-DS with one capture. Figure 18B shows the recovery rate calculated by the percentage of genomes in the input DNA that produced DCS reads. Figure 18C shows that the median DCS depth across all target regions was calculated for each input amount. Three input amounts (250 ng, 100 ng, and 25 ng) of the same DNA extracted from normal human bladder tissue were sequenced with the standard protocol (i.e., standard DS) and CRISPR-DS. In a single capture, CRISPR-DS achieved over 90% on-target raw reads (e.g., encompassing TP53) (Table 8, shown below), which represents a significant improvement over standard-DS (which achieved approximately 5% on-target raw reads in a single capture (Table 8, shown below). In a second capture, raw reads for CRISPR-DS increased minimally (Figure 19). Standard-DS produced approximately 1% recovery (e.g., percentage of input genome recovered as sequenced genome; also known as fractional genome equivalent recovery) across different inputs, while CRISPR-DS produced recoveries ranging from 6-12%. The recovery rate for CRISPR-DS translates to 25 ng of DNA producing a DCS depth (depth generated in DCS reads) comparable to that produced by 250 ng of DNA with standard-DS.A side-by-side comparison of the two methods demonstrated that CRISPR-DS can offer improvements in that over-representation of short fragments due to PCR amplification bias does not occur / affect results (i.e., even inclusion of the region of interest), distinct bands / peaks provide confirmation of correct library preparation prior to sequencing, and the distinct fragments created by targeted fragmentation fully spanned the desired target region with even inclusion (Figure 22E).
[0253] Materials and Methods sample Samples analyzed in this study included deidentified human genomic DNA from peripheral blood, bladder tissue with and without cancer, and ascites DNA. Patient information was available for ascites samples and was used to confirm the presence of tumor mutations. Fluid samples were obtained from the University of Washington Gynecological Tumor Tissue Bank, and specimen and clinical information were collected after informed consent under Protocol No. 27077, approved by the University of Washington Human Subjects Division Institutional Review Board. Deidentified frozen bladder samples were obtained from the University of Washington Genitourinary Cancer Biospecimen Repository and autopsy tissue that had not previously been fixed or frozen. DNA had been previously extracted with a QIAamp DNA Mini Kit (Qiagen, Inc., Valencia, CA, USA) and had never been denatured. DNA was quantified using a Qubit HS dsDNA Kit (ThermoFisher Scientific). DNA quality was assessed using a Genomic TapeStation (Agilent, Santa Clara, CA), and a DNA integrity number (DIN) was determined. DIN is a measure of genomic DNA quality ranging from 1 (highly degraded) to 10 (no degradation). The DIN for peripheral blood DNA and ascites DNA was greater than 7 (reflecting good quality DNA without degradation). Figure 19 is a bar graph showing target enrichment achieved by CRISPR-DS with one capture step compared to two capture steps on three different blood DNA samples.
[0254] Bladder samples were purposefully selected to contain different levels of DNA degradation. Bladder DNA samples B1 through B13 had DINs of 6.8 to 8.9 and were successfully analyzed by CRISPR-DS (Table 10, shown below). Samples B14 and B16 had DINs of 6 and 4, respectively, and were used to demonstrate the improvement made by pre-enriching high molecular weight DNA with the Bluepippin system (Figures 20A and 20B).
[0255] CRISPR guide design gRNAs for excising TP53 exons were designed to have the following characteristics: (1) the ability to produce approximately 500-bp fragments encompassing the TP53 coding region, and (2) the highest MIT website score ("MIT score"; CRISPR.mit.edu:8079 / ; Table 1 and Figures 17A-17C). For exon 7, guides were designed to produce smaller fragments to avoid the proximal polyA tract within the region of interest. A total of 12 gRNAs were designed to excise TP53 into seven distinct fragments (Figure 12A). All gRNAs had "MIT" scores above 60. The quality of excision was assessed by reviewing the alignment of the final DCS reads using Integrative Genomics Viewer. Successful guides produced typical inclusion patterns with sharp edges at the region boundaries and appropriate DCS depth (Figure 22E). If a guide "failed," a decrease in DCS depth was observed, along with the presence of long reads beyond the expected breakpoint. Such guides were redesigned as necessary. Guides were evaluated using synthetic GeneBlock DNA fragments (IDT, Coralville, IA) containing all gRNA sequences flanked by random DNA sequences (Table 7) (Figures 21A-21B). Using the CRISPR / Cas9 in vitro digestion protocol described herein, 3 ng of GeneBlock DNA was digested with each gRNA. After digestion, the reactions were analyzed on a TapeStation 4200 (Agilent Technologies, Santa Clara, CA, USA) (Figure 21C). The presence of defined fragment lengths confirmed proper gRNA assembly and the ability of the gRNA to cleave its target site.
[0256] JPEG0007821756000007.jpg156155
[0257] CRISPR / Cas9 in vitro digestion of genomic DNA crRNA and tracrRNA (IDT, Coralville, IA) were complexed with gRNA, and 30 nM gRNA was incubated with approximately 30 nM Cas9 nuclease (NEB, Ipswich, MA), 1x NEB Cas9 reaction buffer, and 23–27 µL of water at 25 °C for 10 min. Then, 10–250 ng of DNA was added to a final volume of 30 µL. The reaction was incubated overnight at 37 °C and then heat-shocked at 70 °C for 10 min for enzyme inactivation.
[0258] Select size. Size selection was used to select fragments of a predetermined length for target enrichment prior to library preparation. AMPure XP beads (Beckman Coulter, Brea, CA, USA) were used to remove off-target, undigested high-molecular-weight DNA. After heat inactivation, the reaction was combined with a 0.5x ratio of beads, briefly mixed, and then incubated for 3 minutes to allow high-molecular-weight DNA binding. The beads were then separated from the solution with a magnet, and the solution (containing DNA fragments of the desired length) was transferred to a new tube. Standard AMPure 1.8x ratio bead purification was performed and eluted in 50 μL of TE Low.
[0259] Library preparation A-tailing and ligation The fragmented DNA was A-tailed and ligated using the NEBNext Ultra II DNA Library Prep Kit (NEB, Ipswich, MA) according to the manufacturer's protocol. The NEB end repair and A-tailing (ERAT) reaction was incubated at 20°C for 30 minutes and 65°C for 30 minutes. Although end repair is not required for CRISPR-DS (Cas9 generates blunt ends), the ERAT reaction was used for efficient A-tailing. 2.5 μl of NEB Ligation Master Mix and 15 μM DS adapters were added and incubated at 20°C for 15 minutes. A commercially available adapter prototype (Figure 12C) was synthesized with the following differences from the adapters used in previous studies: (1) a 10-bp random double-stranded molecular tag was used instead of a 12-bp one; and (2) a simple 3'-dT overhang replaced the previous 5-bp conserved sequence at the 3' end and was used to ligate to 5'-dA-tailed DNA molecules. Upon ligation, the DNA was washed with 0.8x AMPure bead purification and eluted in 23 μL of nuclease-free water.
[0260] PCR The ligated DNA was amplified using the KAPA Real-Time Amplification kit and fluorescent standards (KAPA Biosystems, Woburn, MA, USA). A 50 μl reaction mixture was prepared containing KAPA HiFi HotStart Real-Time PCR Master Mix, 23 μl of pre-ligated and purified DNA, and DS primers MWS13 and MWS20 at a final concentration of 2 μM. The reaction was denatured at 98°C for 45 seconds and amplified for 6–8 cycles at 98°C for 15 seconds, 65°C for 30 seconds, and 72°C for 30 seconds, followed by a final extension at 72°C for 1 minute. The sample was amplified until fluorescent standard 3 (which produces a standardized number of DNA copies sufficient for capture across the entire sample, prevents overamplification, and indicates successful Cas9 cleavage and ligation) was reached, typically requiring 6–8 cycles depending on the amount of DNA input. The amplified fragment was purified using a 0.8x AMPure Bead wash and eluted in 40 μL of nuclease-free water. Compared to standard DS of the PCR step, CRISPR-DS offers improvements, including: (i) providing fragments of similar size (reducing amplification bias toward small fragments (Figure 22A)), (ii) generating more uniform inclusion of the region of interest (Figure 22E), and (iii) accurate assessment of successful library preparation (using predetermined fragment size characteristics) using a TapeStation 4200 (Agilent Technologies, Santa Clara, CA, USA). With standard DS, PCR products have a wide range of sizes due to sonication and appear as a broad smear that is difficult to compare across samples (Figure 22A). In contrast to other approaches, such as standard DS (which can produce results that are difficult to compare across samples), CRISPR-DS produces discrete peaks that clearly indicate successful cleavage and ligation and are suitable for quality control comparisons across samples (Figure 22B-D).
[0261] Capture and post-capture PCR Hybridization capture of TP53 exons was performed using TP53 xGen Lockdown Probes (IDT, Coralville, IA) according to previous studies, with the following modifications: Probes (from the IDT TP53 Lockdown probe set) were selected to encompass the entire TP53 coding region (exon 1 and portions of exon 11 are not coding regions) (Table 6). Each CRISPR / Cas9 excision fragment was encompassed by a minimum of two probes and a maximum of five probes (Figures 17A-17C). To generate capture probe pools, each probe for a given fragment was pooled in equimolar amounts, generating seven different pools (one for each fragment). The seven fragment pools were then mixed again in equimolar amounts (with the exception of exon 7 and exon 8-9 pools, which were represented at 40% and 90%, respectively). Reduction of capture probes for these exons was performed when overrepresentation of exons was observed in the sequence. The final capture pool was diluted to 0.75 pmol / μl. Hybridization capture was performed according to the standard IDT protocol with the following modifications: 75 μl (instead of 100 μl) of Dynabeads M-270 streptavidin beads were used, using blockers MWS60 and MSW61 specific for the DS adapter. Post-capture PCR was performed with the KAPA Hi-Fi HotStart PCR Kit (KAPA Biosystems, Woburn, MA, USA) using a final concentration of 0.8 μM MWS13 and indexed primer MWS21. Reactions were denatured at 98°C for 45 seconds, then amplified for 20 cycles of 98°C for 30 seconds, 60°C for 45 seconds, and 72°C for 45 seconds, followed by an extension at 72°C for 60 seconds. PCR products were purified with a 0.8x AMPure bead wash.
[0262] Sequencing Samples were quantified using the Qubit dsDNA HS Assay Kit, diluted, and pooled for sequencing. The sample pools were then visualized on an Agilent 4200 TapeStation to confirm library quality. TapeStation electropherograms showed sharp, distinct peaks corresponding to the fragment lengths of the designed CRISPR / Cas9 cleavage fragments (Figures 22B-22D). (This step can also be performed individually for each sample before pooling to verify the performance of each individual sample, if necessary / desired.) The final pool was quantified using the KAPA Library Quantification kit (KAPA Biosystems, Woburn, MA, USA). Libraries were sequenced on the MiSeq Illumina platform using the v3 600-cycle kit (Illumina, San Diego, CA, USA) according to the manufacturer's instructions. Each sample was assigned approximately 7–10% of the lanes (corresponding to approximately 2 million reads), and each sequencing run was spiked with approximately 1% PhiX control DNA.
[0263] Data Processing A custom bioinformatics pipeline was created to automate the analysis of raw FASTQ files into text files (Figure 23). This pipeline is similar to the method used for standard DS analysis, with the following modifications: (i) paired-read information is retained, and (ii) consensus generation is performed before alignment. Paired-end reads are used to analyze CRISPR-DS data, but this method offers an improvement over standard DS analysis because it provides quality control for fragment size and eliminates potential technical artifacts due to the presence of short fragments. Furthermore, while standard DS analysis generates consensus after all reads are mapped to the reference genome, CRISPR-DS analysis performs consensus generation as a first step, relying solely on sequencer reads. This modification likely improves consensus generation and reduces the time required for data processing. In CRISPR-DS, consensus generation was performed by a custom Python script called UnifiedConsensusMaker.py, which took all reads derived from the same tag, compared the bases called at each position, and produced a single-stranded consensus (SSCS) read. The SSCS reads for each complementary pair of tags were then compared position-by-position to generate a double-stranded consensus (DCS) read (Figure 12D). Two FASTQ files containing the resulting SSCS and DCS reads were generated (because DCS reads correspond to the original DNA molecule, the average DCS depth is an estimate of the number of genomes sequenced). The recovery rate (also known as fractional genome equivalent recovery) was calculated by dividing the average DCS depth (genomes sequenced) by the number of input genomes (1 ng of DNA corresponds to approximately 330 haploid genomes). On-target raw reads were calculated by counting the number of reads whose genomic coordinates were within the upstream and downstream CRISPR / Cas9 cut sites, with a 100-bp window added on either side. Paired-end DCS FASTQ files were aligned to the human reference genome v38 using bwa-mem v.0.7.419 with default parameters.Mapped reads were realigned using GATK Indel-Realigner, and low-quality bases were clipped from both ends using GATK Clip-Reads. Conservative clipping was performed: 30 bases from the 3' end and 7 bases from the 5' end. Furthermore, overlapping regions of read pairs, spanning approximately 80 bp in the TP53 design, were trimmed using fgbio ClipOverlappingReads. This algorithm clips from both ends of the paired reads until they match, maximizing the use of sequencing bases with high PHRED quality scores. A pileup file was created from the resulting files using SAMtools mpileup. The pileup file was then filtered using a custom Python script containing a BED file for the targeted genomic location. BED files can be easily created using the coordinates of the CRISPR / Cas9 gRNA. The filtered pileup file was then processed by a custom script, mut-position.1.33.py, which creates a tab-delimited text file containing mutation information called "mutpos." mutpos contains a summary of the DCS depth and mutations at each sequenced position (the software used in the CRISPR-DS analysis can be accessed via the hypertext transfer protocol secure: / / github.com / risqueslab / CRISPR-DS).
[0264] Standard DS Three amounts of DNA (25 ng, 100 ng, and 250 ng) from normal human bladder sample B9 were sequenced with standard DS in one and two capture runs and compared with CRISPR-DS results. Standard DS analysis was performed using the KAPA Hyperprep Kit (KAPA Biosystems, Woburn, MA, USA) for end repair and ligation, and the KAPA Hi-Fi HotStart PCR Kit (KAPA Biosystems, Woburn, MA, USA) for PCR amplification. Hybridization capture was performed using the xGen Lockdown probe encompassing TP53 exons 2–11 (the same probe was used for both standard DS and CRISPR-DS). Samples were sequenced on a HiSeq 2500 Illumina platform (approximately 10% of the time) to accommodate shorter fragment lengths.
[0265] CRISPR-DS target enrichment To characterize CRISPR-DS target enrichment, two separate analyses were performed.
[0266] Initial analyses included a comparison of one versus two captures (and comparison with standard DS results). Three DNA samples were processed for CRISPR-DS and split in half after one hybridization capture. As required by the original DS protocol, the first half was indexed and sequenced, and the second half was subjected to an additional capture. The percentage of raw reads that were "on-target" (i.e., encompassing TP53 exons) was compared between one versus two captures. Details of the comparison between standard DS and CRISPR-DS are shown in Table 8. JPEG0007821756000008.jpg77170
[0267] In the second analysis, we assessed the percentage of raw reads that were on-target without performing hybridization capture to determine the enrichment exclusively produced by size-selecting CRISPR excision fragments. Three different samples with different amounts of DNA (10 ng to 250 ng) were processed with the protocol described in the first analysis up to the first PCR (i.e., before hybridization capture). Figures 24A and 24B are a chart (Figure 24A) and a graph (Figure 24B) showing the results of quantifying the degree of target enrichment after CRISPR / Cas9 digestion and subsequent size selection, according to one embodiment of the present technology. Figure 24A shows the DNA samples and the enrichment achieved for each. Figure 24B shows the percentage of raw reads that were "on-target" relative to the amount of input DNA. The PCR products were then indexed and sequenced. The percentage of on-target raw reads was calculated and the enrichment fold was estimated (taking into account the targeted region size, in this case 3280 bp).
[0268] Preconcentration of high molecular weight DNA Selecting for high-molecular-weight DNA improves the performance of DNA degraded with CRISPR-DS. This selection was performed using the BluePippin system (Sage Science, Beverly, MA). Two bladder DNAs with DINs of 6 and 4 were run using a 0.75% gel cassette and high-pass settings to obtain fragments greater than 8 kb. Size selection was confirmed with TapeStation (Figure 20A). 250 ng of DNA before BluePippin and 250 ng of DNA after BluePippin were then processed in parallel by CRISPR-DS. The percentage of on-target raw reads and average DCS depth were quantified and compared (Figure 20B).
[0269] Example 13: CRISPR-DS in ovarian cancer samples To validate the ability of CRISPR-DS to detect low-frequency mutations, four ascites samples were collected and analyzed from women with ovarian cancer during debulking surgery. The presence of TP53 tumor mutations in these samples had previously been demonstrated by standard DS. 100 ng of DNA (30-100-fold less than that used for standard DS) was used for CRISPR-DS analysis, yielding DCS depth comparable to standard DS, and TP53 tumor mutations were successfully identified in all cases (Table 9). Recovery rates ranged from 6-12%, representing a 15- to 200-fold increase compared to standard DS with the same DNA. JPEG0007821756000009.jpg101154
[0270] Example 14: CRISPR-DS of bladder tissue samples This example describes the use of CRISPR-DS on a set of 13 DNA samples extracted from bladder tissue from various patients (Table 10). 250 ng of DNA from each sample was used in the assay, resulting in a median DCS depth of 6,143-fold, corresponding to a median recovery of 7.4%. Technical duplicates of two samples (B2 and B4) demonstrated reproducible performance. On-target DCS reads for all samples were greater than 98%, while the percentage of on-target raw reads (raw reads) ranged from 43% to 98%. Low target enrichment corresponded to samples with DNA Integrity Numbers (DIN) less than 7.
[0271] JPEG0007821756000010.jpg83148
[0272] To test the effect of DIN on assay performance, low-molecular-weight DNA was removed prior to CRISPR / Cas9 digestion. The pulsed-field feature of the BluePippin system was used to select high-molecular-weight DNA from two samples with "degraded DNA" (DIN 6 and 4). Pre-enrichment resulted in a 2-fold increase in on-target raw reads and a 5-fold increase in DCS depth (Figure 20B). To directly quantify the degree of enrichment conferred by simple CRISPR / Cas9 digestion followed by size selection, three samples were sequenced without capture. 10–250 ng of DNA was digested, size-selected, ligated, amplified, and sequenced. The percentage of "on-target" raw reads ranged from 0.2% to 5%, corresponding to approximately 2,000- to 50,000-fold enrichment (Table 11). Notably, low DNA inputs showed the highest enrichment, likely reflecting optimal removal of off-target, high-molecular-weight DNA fragments when present in low abundance. JPEG0007821756000011.jpg97121
[0273] CRISPR / Cas9 fragmentation followed by size selection successfully performed efficient target enrichment, eliminating the need for a second capture of small target regions. Furthermore, PCR bias was eliminated and uniform inclusion of the desired region was achieved, representing a significant improvement over currently available methods.
[0274] Equivalents and Scope The above detailed description of embodiments of the present technology is not intended to be exhaustive or to limit the technology to the precise form disclosed above. Specific embodiments and examples of the present technology have been described above for illustrative purposes, but those skilled in the art will recognize that various equivalent modifications are possible within the scope of the present technology. For example, while steps are presented in a given order, steps may be performed in a different order in alternative embodiments. The various embodiments described herein may be combined to provide further embodiments. All references cited herein are incorporated by reference as if fully set forth herein.
[0275] From the above, it will be understood that, although specific embodiments of the present technology have been described herein for illustrative purposes, well-known structures and functions have not been shown or described in detail to avoid unnecessarily obscuring the description of the embodiments of the present technology. Where the context permits, singular or plural terms may include the plural or singular terms, respectively. Furthermore, while advantages associated with specific embodiments of the technology have been described in the context of those embodiments, other embodiments may also exhibit such advantages, and not all embodiments necessarily exhibit such advantages to be within the scope of the present technology. Thus, the present disclosure and related technology may encompass other embodiments not explicitly shown or described herein.
[0276] Those skilled in the art will recognize, or be able to ascertain using no more than routine experimentation, many equivalents to the specific embodiments of the disclosed technology described herein. The scope of the technology is not intended to be limited to the above description, but rather is as set forth in the following claims.
Claims
1. providing nucleic acid constructs, each of the nucleic acid constructs comprising double-stranded DNA fragments, at least one single molecular identifier (SMI) sequence on each strand, a first adaptor attached to one end of the fragments, and a second adaptor attached to the other end of the fragments; wherein the first and second adaptors each comprise (i) a terminal portion comprising at least partially non-complementary 5' and 3' binding sequences, and (ii) a portion comprising a molecular barcode sequence located between the terminal portion and a nucleic acid fragment, and wherein at least one nucleic acid fragment of the construct comprises a target DNA sequence of interest; amplifying the nucleic acid construct to generate first-strand and second-strand amplification products, at least some of the first-strand and second-strand amplification products comprising a target DNA sequence of interest; separating the amplification products into a first sample and a second sample, each of the samples comprising a plurality of the first strand amplification products and a plurality of second strand amplification products; performing targeted amplification of the first sample, comprising exponentially amplifying only the first strand amplification products using at least one single-stranded oligonucleotide at least partially complementary to the 5' binding sequence and at least one single-stranded oligonucleotide at least partially complementary to a target DNA sequence of interest, resulting in first product nucleic acids such that the molecular barcode sequence of the first or second adapter is at least partially maintained; performing targeted amplification of the second sample, comprising exponentially amplifying only the second strand amplification products in the second sample using at least one single-stranded nucleotide at least partially complementary to the 3'-linked sequence and at least one single-stranded oligonucleotide at least partially complementary to a target DNA sequence of interest, resulting in second nucleic acid products such that the molecular barcode sequence at least partially maintained in the first nucleic acid products is at least partially maintained; sequencing each of the first nucleic acid product and the second nucleic acid product to obtain sequence reads each comprising an SMI sequence; and comparing a sequence read of the first nucleic acid product with a sequence read of the second nucleic acid product based at least in part on an SMI sequence for a target DNA sequence of interest derived from the same original nucleic acid construct; A method for determining the sequence of a target DNA sequence of interest.
2. For each double-stranded nucleic acid construct, providing a step of:
10. The method of claim 1, comprising ligating to the double-stranded DNA fragment a first adaptor comprising a first molecular barcode sequence and a second adaptor comprising a second molecular barcode sequence to form a double-stranded nucleic acid construct, wherein the molecular barcode sequence comprises an SMI sequence.
3. the molecular barcode sequence is single-stranded before or after ligating the first or second adaptor to the double-stranded nucleic acid molecule, and a complementary molecular barcode sequence is generated by extending the opposite strand with a DNA polymerase to generate a complementary double-stranded molecular barcode sequence; The method of claim 2.
4. The SMI sequence a) at least one of a molecular barcode sequence, the one or more DNA fragment ends, or a combination thereof, that uniquely labels double-stranded DNA fragments containing a target DNA sequence of interest; or b) comprising an intrinsic shear point or an intrinsic sequence spatially associated with said shear point; The method of claim 1.
5. 5. The method of any one of claims 1 to 4, wherein the double-stranded DNA fragments comprising the target DNA sequence of interest are provided by a sample originating from a subject or organism, said sample being or comprising body tissue, tissue biopsy, liquid biopsy, blood, serum, plasma, saliva, cerebrospinal fluid, lavage fluid, urine, stool, cell-free nucleic acid, intracellular nucleic acid, and any combination thereof.
6. Before the providing step, cleaving the DNA material with one or more targeting endonucleases so as to form double-stranded DNA fragments of substantially known length; and isolating the double-stranded DNA fragments based on the substantially known length; The method according to any one of claims 1 to 5, comprising: The method, wherein the one or more targeting endonucleases are selected from the group consisting of ribonucleoproteins, Cas enzymes, Cas9-like enzymes, meganucleases, transcription activator-like effector-based nucleases (TALENs), zinc finger nucleases, Argonaute nucleases, or combinations thereof.
7. 7. The method of claim 6, wherein cleaving the DNA material comprises cleaving the DNA material with one or more targeting endonucleases such that a plurality of double-stranded DNA fragments of interest of substantially known length are formed, each of the double-stranded DNA fragments comprises a target DNA sequence of interest from one or more different locations in the genome; The method.
8. the at least partially non-complementary terminal subsequences of the first and second adaptors comprise a Y-shape; The method according to any one of claims 1 to 7.
9. a) ligating double-stranded target DNA material to at least one adapter sequence to form an adapter-target nucleic acid material complex, wherein the at least one adapter sequence is: (i) a single molecule identifier (SMI) sequence that uniquely labels each adaptor-target nucleic acid material complex along with the sequence of the single or double stranded target DNA material; and (ii) a first nucleotide adapter sequence that tags a first strand of the adapter-target nucleic acid material complex, and a second nucleotide adapter sequence that tags a second strand of the adapter-target nucleic acid material complex, the second nucleotide adapter sequence being at least partially non-complementary to the first nucleotide adapter sequence; each strand of said adaptor-target nucleic acid material complex has a nucleotide sequence that is unambiguously identifiable relative to its complementary strand; b) amplifying each strand of the adaptor-target nucleic acid material complex to produce a plurality of first-strand amplification products and a plurality of second-strand amplification products, and separating the first and second amplification products into a first sample and a second sample, each sample comprising a plurality of first-strand amplification products and a plurality of second-strand amplification products; c) amplifying the first strand in the first sample using a first primer at least partially complementary to the first nucleotide adapter sequence and a primer at least partially complementary to the target sequence of interest to provide a first nucleic acid product, such that the SMI sequence is maintained in the first nucleic acid product; d) amplifying the second strand in the second sample using a second primer at least partially complementary to the second nucleotide adapter sequence and a primer at least partially complementary to the target sequence of interest, such that the SMI sequence maintained in the first nucleic acid product is also maintained in the second nucleic acid product, to provide a second nucleic acid product; e) sequencing each of the first nucleic acid product and the second nucleic acid product for a target DNA sequence of interest derived from the same original adaptor-target nucleic acid material complex to produce a plurality of first strand sequence reads and a plurality of second strand sequence reads, and confirming the presence of at least one first strand sequence read and at least one second strand sequence read; and, f) generating error-corrected sequence reads of the double-stranded target DNA material by comparing the at least one first strand sequence read with the at least one second strand sequence read and ignoring mismatched nucleotide positions, or alternatively, by removing compared first and second strand sequence reads having one or more nucleotide positions where the compared first and second strand sequence reads are non-complementary; Including, A method for generating error-corrected sequence reads of double-stranded target DNA material.
10. 10. The method of claim 9, wherein the sequence reads for which errors have been corrected in step (f) are used to identify or characterize cancer, cancer risk, or cancer recurrence.
11. the adapter sequence comprises a Y-shape; 11. The method according to claim 9 or 10.
Citation Information
Patent Citations
Methods of lowering the error rate of massively parallel DNA sequencing using duplex consensus sequencing
WO2013142389A1