Methods and compositions for nucleic acid library and template preparation for duplexed sequencing by expansion
Patent Information
- Application Number
- PCT/EP2024/087393
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-22
- Filing Date
- 2024-12-19
- Publication Date
- 2025-08-07
AI Technical Summary
Existing nucleic acid sequencing technologies face limitations in accuracy due to errors generated during sequencing of single-stranded nucleic acid templates, particularly in next-generation sequencing by synthesis and single-molecule nanopore sequencing. Duplex sequencing techniques, which determine sequences from both strands of a double-stranded nucleic acid, can improve accuracy but require efficient methods for generating and processing duplex nucleic acid templates.
The development of novel adapter compositions, including extendable Y adapters, cleavable hairpin adapters, and Y-hairpin hybrid adapters, enables the generation of duplex nucleic acid templates suitable for duplex sequencing by expansion. These adapters facilitate the synthesis of double-stranded end products and the subsequent joining of hairpin adapters, which are then amplified and used to generate Xpandomer copies for nanopore sequencing.
The proposed methods and compositions enhance the accuracy of nucleic acid sequencing by allowing for the reliable determination of both sense and anti-sense strands of a double-stranded nucleic acid template, improving the assembly of nucleic acid sequences and providing additional information for applications such as tumor DNA detection and epigenetic analysis.
Abstract
Description
METHODS AND COMPOSITIONS FOR NUCLEIC ACID LIBRARY AND TEMPLATE PREPARATION USING EXTENDABLE ADAPTERS FOR DUPLEXED SEQUENCING BY EXPANSIONFIELD OF THE INVENTION
[0001] The disclosure relates generally to methods and compositions for generating duplex nucleic acid template constructs that find use in duplex Sequencing by Expansion and improved reaction conditions for synthesizing Xpandomer copies of the duplex nucleic acid template constructs for nanopore sequencing. Also, the disclosure relates to novel adapter compositions for generating the duplex nucleic acid templates, and, in particular, extendable Y adapter, cleavable hairpin adapters, and Y-hairpin hybrid adapters.BACKGROUND OF THE INVENTION
[0002] Development of nucleic acid sequencing technologies has yielded countless advances in numerous areas. The ability to rapidly and reliably determine the sequence of DNA and RNA molecules has enabled numerous advances in molecular biology, evolutionary biology, medical diagnostics, and molecular medicine, among many other fields.
[0003] The accuracy of sequence data that can be reliably obtained when using certain next generation sequencing by synthesis or single molecules nanopore sequencing techniques, however, may be limited due to errors generated during sequencing of the single stranded nucleic acid template molecule. Thus, in many circumstances it can be advantageous to be able to reliably obtain further sequence data of the entire double stranded template molecule. To this end, paired-end, or duplex, sequencing techniques have been employed, e.g., particularly in the context of whole genome shotgun sequencing. Duplex sequencing can allow the determination of two “reads” of sequence from a double stranded nucleic acid target sequence, one from the “sense” strand and one from the “anti-sense” strand. The knowledge that the paired-end sequences are known to occur on a single duplex, and are therefore linked, or paired, in the genome, can greatly aid assembly of nucleic acid sequences into a consensus sequence, thus greatly improving the accuracy of the sequencing reads. The additional information obtained from paired-end sequencing can also benefit other applications, e.g., applications involving sequencing cell-free DNA such as detection of circulating tumor DNA and prenatal cell-free DNA screening, and various epigenetic detection methodologies.
[0004] Provided herein are novel and useful compositions and methods for carrying out paired-end, duplex sequencing. These compositions and methods provide advantages to a number of sequencing methods, e.g., nanopore-based, single molecule sequencing methods.
[0005] All of the subject matter discussed in the Background section is not necessarily prior art and should not be assumed to be prior art merely as a result of its discussion in the Background section. Along these lines, any recognition of problems in the prior art discussed in the Background section or associated with such subject matter should not be treated as prior art unless expressly stated to be prior art. Instead, the discussion of any subject matter in the Background section should be treated as part of the inventor’s approach to the particular problem, which in and of itself may also be inventive.BRIEF SUMMARY OF THE INVENTION
[0006] The present disclosure provides improved methods and compositions for generating duplex nucleic acid template constructs and their use in duplex sequencing methods, including e.g., Sequencing by Expansion.
[0007] In one aspect, the invention provides a method of producing a duplex nucleic acid template, the method including: providing a double stranded nucleic acid fragment including a first extendable Y adapter joined to a first end and a second extendable Y adapter joined to a second end, in which the first and the second extendable Y adapters includes a region of double stranded DNA, in which the region of double stranded DNA includes an internal break in one of the strands, and in which the internal break provides an extendable 3 ’ end within the double stranded region; contacting the double stranded nucleic acid fragment to a polymerase under nucleic acid synthesis conditions, in which the polymerase initiates template-dependent synthesis from the extendable 3 ’ ends in the first and second extendable Y adapters to produce a first and a second nucleic acid product, in which the first and second nucleic acid products include a double stranded end; and joining a third adapter to the double stranded ends of the first and the second nucleic acid synthesis products, in which the third adapter is a hairpin adapter, and in which the hairpin adapter includes a double stranded stem region and a single stranded loop region. In one embodiment, the strand of the double stranded region of the extendable Y adapter including a break further includes a 5’ single stranded tail. In another embodiment, the single stranded loop region of the hairpin adapter includes a cleavage site. In yet other embodiments, the single stranded loop regions of the hairpin adapter further includes a hybridization sequence complementary to an extension oligonucleotide. In yet other embodiments, the hairpin adapter further includes a first hybridization sequence and a second hybridization sequence, in which the first hybridization sequence is complementary to a first sequence in an invasive blockeroligonucleotide and the second hybridization sequence is complementary to a second sequence in the invasive blocker oligonucleotide. In some embodiments, the nucleic acid synthesis conditions include nucleotide analogs, in which the nucleotide analogs form weak hydrogen bonds with complementary nucleotides relative to native nucleotides. In further embodiments, the nucleotide analogs include N4-Me dCTP and 7-deaza dGTP. In some embodiments, any one of the first, second, or third adapters includes one or more features selected from the group consisting of a UMI, a SID, an extension oligonucleotide hybridization sequence, and a blocking oligonucleotide hybridization sequence. In some embodiments, the polymerase is a strand displacing DNA polymerase. In some embodiments, the duplex nucleic acid template is subjected treatment with a DNA glycosylase enzyme and an aminoxyalkyl uracil mimetic compound, in which the DNA glycosylase enzyme and the aminoxyalkyl uracil mimetic compound selectively converts epigenetically modified cytosine residues to uracil oxime mimetic residues.
[0008] In another aspect, the invention provides a method of sequencing a duplex nucleic acid template, the method including: providing a sample including any of the above duplex nucleic acid templates; amplifying the duplex nucleic acid template; generating a copy of the amplified duplex nucleic acid template, in which the copy of the amplified duplex nucleic acid template includes an Xpandomer, wherein the Xpandomer includes a sequence of reporter codes, wherein the sequence of reporter codes encodes the nucleic acid sequence of the duplex nucleic acid template; and determining the sequence of the reporter codes by passing the Xpandomer through a nanopore. In one embodiment, the step of generating the copy of the amplified duplex nucleic acid template includes a solid-state synthesis Xpandomer synthesis reaction. In some embodiments, the method further includes the step of cleaving the cleavage site of the hairpin adapter prior to the step of generating the copy of the amplified duplex nucleic acid template. In some embodiments, the method further includes the step of hybridizing an invasive blocker oligonucleotide to the cleaved amplified duplex nucleic acid template, in which the invasive blocker includes a first end capable of being extended and a second end including a moiety capable of being joined to a 5’ end of an Xpandomer. In some embodiments, the step of generating a copy of the amplified duplex nucleic acid template includes providing an Xpandomer synthesis reaction, in which the Xpandomer synthesis reaction includes the amplified duplex nucleic acid template, an extension oligonucleotide, a blocking oligonucleotide, a buffer / salt system, a polymerase cofactor, a polymerase enhancing moieties (PEMs), a DNA polymerase, XNTP substrates, a phosphate shield molecule, a solvent, and a crowding agent. In some embodiments, the Xpandomer synthesis reaction further includes one or more additives. In further embodiments, the one or more additives includes any of the additivesset forth in Table 2. In some embodiments, the DNA polymerase is a variant of wildtype DPO4 polymerase (SEQ ID NO: 1). In further embodiments, the variant of DPO4 polymerase is at least 85% identical to the variant of SEQ ID NO:2.
[0009] In another aspect, the invention provides an extendable Y adapter according to the embodiments described herein.
[0010] In another aspect, the invention provides a hairpin adapter according to the embodiments described herein.
[0011] In another aspect, the invention provides a Y-hairpin hybrid adapter including a double stranded stem region, a 3’ arm region, and a loop region, in which the 3’ arm region includes a single stranded 3’ end region, and in which the loop region includes in series, a first end, a single stranded region, and a second end, in which the first end is covalently joined to the double stranded stem region, in which the second end is not covalently joined to the double stranded stem region and in which the second end is hybridized to the 3’ arm region to provide an extendable 3’ end. In one embodiment, the loop region includes a cleavable linker interposed between the single stranded loop region and the double stranded stem region. In another embodiment, the Y-hairpin hybrid adapter further includes a biotin moiety interposed between the cleavable linker and the double stranded stem region. In another embodiment, the single stranded region of the loop region comprises a reverse polarity phosphoramidite.
[0012] In another aspect, the invention provides use of any of the Y-hairpin hybrid adapters described herein to produce a single stranded duplex template construct, in which the single stranded duplex template construct includes a parental sense strand and a daughter sense strand or a parental antisense strand and a daughter antisense strand.BRIEF DESCRIPTION OF THE DRAWINGS
[0013] FIGS. 1A, IB, 1C , ID, IE and IF are condensed schematics summarizing one embodiment of the methods of generating a duplex nucleic template construct of the present invention with an extendable Y adapter and use in duplexed Xpandomer synthesis.
[0014] FIGS. 2 A, 2B, 2C, 2D, 2E, 2F and 2G are condensed schematics summarizing another embodiment of the methods of generating a duplex nucleic template construct of the present invention with an extendable Y adapter and a cleavable hairpin adapter and use in duplexed Xpandomer synthesis.
[0015] FIGS. 3A, 3B, 3C , 3D, 3E, 3F, and 3G are are condensed schematics summarizing another embodiment of the methods of generating a duplex nucleic template construct of the present invention with an extendable Y adapter , a cleavable hairpin adapter, and an invasive blocker and use in duplexed Xpandomer synthesis.
[0016] FIGS. 4 A, 4B, 4C and 4D are condensed schematics summarizing another embodiment of the methods of generating a duplex nucleic template construct of the present invention with an extendable looped adapter.
[0017] FIG. 5 A and 5B are cartoons illustrating certain features of chemo-enzymatic conversion for detection of methylated cytosine nucleobases in a DNA target fragment.
[0018] FIG. 6 is a simplified illustration summarizing certain features of an exemplary xNTP substrate used for the synthesis of Xpandomer copies of library template constructs for nanopore sequence determination.DETAILED DESCRIPTION OF THE INVENTION
[0019] The present invention may be understood more readily by reference to the following detailed description of preferred embodiments of the invention and the Examples included herein. Unless otherwise explained, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs.
[0020] The practice of the present invention will employ, unless otherwise indicated, conventional techniques of molecular biology, microbiology, recombinant DNA, and so forth which are within the skill of the art. Such techniques are explained fully in the literature. See e g., Sambrook, Fritsch, and Maniatis, MOLECULAR CLONING: A LABORATORY MANUAL, Second Edition (1989), OLIGONUCLEOTIDE SYNTHESIS (M. J. Gait Ed., 1984), the series METHODS IN ENZYMOLOGY (Academic Press, Inc ), CURRENT PROTOCOLS IN MOLECULAR BIOLOGY (F. M. Ausubel, R. Brent, R. E. Kingston, D. D. Moore, J. G. Siedman, J. A. Smith, and K. Struhl, eds., 1987).
[0021] Reference throughout this specification to “one embodiment” or “an embodiment” and variations thereof means that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment. Thus, the appearances of the phrases “in one embodiment” or “in an embodiment” in various places throughout this specification are not necessarily all referring to the same embodiment. Furthermore, the particular features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.
[0022] As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” include plural referents, i.e., one or more, unless the content and context clearly dictates otherwise. It should also be noted that the conjunctive terms, “and” and “or” are generally employed in the broadest sense to include “and / or” unless the content and context clearly dictates inclusivity or exclusivity as the case may be. Thus, the use of the alternative(e.g., "or") should be understood to mean either one, both, or any combination thereof of the alternatives. In addition, the composition of “and” and “or” when recited herein as “and / or” is intended to encompass an embodiment that includes all of the associated items or ideas and one or more other alternative embodiments that include fewer than all of the associated items or ideas.
[0023] Unless the context requires otherwise, throughout the specification and claims that follow, the word “comprise” and synonyms and variants thereof such as “have” and “include”, as well as variations thereof such as “comprises” and “comprising” are to be construed in an open, inclusive sense, e.g., “including, but not limited to.” The term "consisting essentially of' limits the scope of a claim to the specified materials or steps, or to those that do not materially affect the basic and novel characteristics of the claimed invention.
[0024] The abbreviation, "e.g.," is derived from the Latin exempli gratia, and is used herein to indicate a non-limiting example. Thus, the abbreviation "e.g.," is synonymous with the term "for example." It is also to be understood that as used herein and in the appended claims, the singular forms “a,” “an,” and “the” include plural reference unless the context clearly dictates otherwise, the term “X and / or Y” means “X” or “Y” or both “X” and “Y”, and the letter “s” following a noun designates both the plural and singular forms of that noun. In addition, where features or aspects of the invention are described in terms of Markush groups, it is intended, and those skilled in the art will recognize, that the invention embraces and is also thereby described in terms of any individual member and any subgroup of members of the Markush group, and Applicants reserve the right to revise the application or claims to refer specifically to any individual member or any subgroup of members of the Markush group.
[0025] Any headings used within this document are only being utilized to expedite its review by the reader, and should not be construed as limiting the invention or claims in any manner. Thus, the headings and Abstract of the Disclosure provided herein are for convenience only and do not interpret the scope or meaning of the embodiments.
[0026] Where a range of values is provided herein, it is understood that each intervening value, to the tenth of the unit of the lower limit unless the context clearly dictates otherwise, between the upper and lower limit of that range and any other stated or intervening value in that stated range is encompassed within the invention. The upper and lower limits of these smaller ranges may independently be included in the smaller ranges is also encompassed within the invention, subject to any specifically excluded limit in the stated range. Where the stated range includes one or both of the limits, ranges excluding either or both of those included limits are also included in the invention.
[0027] For example, any concentration range, percentage range, ratio range, or integer range provided herein is to be understood to include the value of any integer within the recited range and, when appropriate, fractions thereof (such as one tenth and one hundredth of an integer), unless otherwise indicated. Also, any number range recited herein relating to any physical feature, such as polymer subunits, size or thickness, are to be understood to include any integer within the recited range, unless otherwise indicated. As used herein, the term "about" means ± 20% of the indicated range, value, or structure, unless otherwise indicated.Methods and Compositions for Duplex Sequencing by Expansion
[0028] Embodiments described herein may be applied any suitable sequencing platform, including next generation and nanopore sequencing, but are particularly useful for Sequencing by Expansion (SBX®). Sequencing by expansion is described in Applicant’s published PCT application, WO 2020 / 236526 Al, “Translocation control elements, reporter codes, and further means for translocation control for use in nanopore sequencing,” filed May 14, 2020, and issued patent, US 7,939,259 B2, “High throughput nucleic acid sequencing by expansion,” filed June 19, 2008, the entire contents of which are both incorporated herein by reference for all purposes.
[0029] In general terms, Sequencing by Expansion (SBX® ) uses biochemical polymerization to transcribe the sequence of a DNA template onto a measurable polymer called an “Xpandomer” using highly modified, non- natural nucleotide analog substrates referred to as “XNTPs”. The transcribed sequence is encoded along the Xpandomer backbone in high signal- to-noise reporter codes that are separated by ~10 nm and are designed for high-signal-to-noise, well-differentiated responses. The Xpandomer polymer thus preserves the original genetic information of the target nucleic acid template, while also increasing linear separation of the individual elements of the sequence data. These differences provide significant performance enhancements in sequence read efficiency and accuracy of Xpandomers relative to natural DNA.
[0030] Sequencing by Expansion is a single molecule sequencing method in that the nanopore detector reads electronic signals from a single Xpandomer. As such, SBX® sequencing reads provide information from a single, contiguous DNA template.
[0031] Sequencing by Expansion is discussed in greater detail further herein.
[0032] The present technology provides improved methods and compositions for nanoporebased, single molecule sequencing, particularly high accuracy duplex Sequencing by Expansion, which can be applied to both genomic and epigenomic, e.g., methylomic, sequence analysis. Described herein are methods and compositions for generating a library of paired-end, double stranded nucleic acid templates, each including at least a subset of sequence from a target nucleicacid sample. These methods and compositions allow for sequencing of nucleic acid templates in which both the sense and anti-sense strands (i.e., the two opposite strands of a double stranded nucleic acid target fragment) are copied onto the same Xpandomer molecule for nanopore sequence determination. The present invention also provides library preparation, template preparation, and analysis methods for Sequencing by Expansion, enabled by, e.g., novel adapter designs. In some embodiments, the sense and anti- sense strands of the nucleic acid target fragments in the library are separated, and operably joined (i.e., linked), by a known nucleic acid adapter segment.
[0033] The general method typically begins with a sample of double stranded nucleic acid fragments having defined ends, which could be blunt ends or ends with known overhang sequences (5' or 3' overhangs). These nucleic acid fragments can be of any size or size range and can include DNA, RNA, DNA-RNA hybrids (e.g., molecules produced by first-strand synthesis during cDNA preparation have one mRNA strand and one complementary DNA strand), genomic DNA, cDNA, mRNA, tRNA, etc. In some embodiments, the nucleotide sequence of the fragments is not known.
[0034] In some aspects, the invention provides for producing a library of paired-end nucleic acid template constructs from the sample of double stranded nucleic acid fragments for synthesis of Xpandomer copies for nanopore sequence determination. As discussed herein, the terms “paired-end” and “duplex” (or “duplexed”) may be used interchangeably as they relate to template construct for Xpandomer synthesis. Sequencing of the single, contiguous Xpandomer copies of the paired-end template provides duplexed reads of the original nucleic acid target fragments. The paired-end Xpandomer template constructs can be single nucleic acid chains that each have the following structure: adapter region 1, sense (i.e., forward) nucleic acid strand of the target fragment, adapter region 2, anti-sense (i.e., reverse) nucleic acid strand of the target fragment, adapter regions 3. In some embodiments, adapter region 2 forms a classic “hairpin” structure in which the stem portion of the hairpin adapter is double stranded and is ligated to one end of the double stranded nucleic acid target fragment. The loop portion of the hairpin adapter is single-stranded and operable joins (i.e., covalently links or couples) the sense and antisense strands of the double stranded nucleic acid target fragment. In some embodiments, adapter region 1 and adapter region 3 are derived from a classic “Y adapter” structure, in which the stem portion of the Y adapter is double stranded and is ligated to the opposite end of the double stranded nucleic acid target fragment. The arms of the Y adapter may be single stranded and provide a free 3’ end and a free 5’ end to the paired-end template construct. As described in further detail herein, in several embodiments, the hairpin and Y adapters structures of the present invention may include several novel features that enable synthesis of “daughter strand” copies ofthe paired-end template constructs useful for, e.g., epigenomic analyses. Certain, non-limiting, examples of alternative library formats and work-flows to enable high accuracy duplex and methylome Sequencing by Expansion are discussed below.
[0035] In one aspect, DNA from a biological sample is obtained or provided. The DNA obtained or provided from the biological sample may be genomic DNA, mitochondrial DNA, cell-free DNA (cfDNA), circulating tumor DNA (ctDNA), or a combination thereof.
[0036] DNA samples may be obtained from a patient or subject, from an environmental sample, or from an organism of interest. In embodiments, the DNA sample is extracted, purified, or derived from a cell or collection of cells, a body fluid, a tissue sample, an organ, and / or an organelle. In some embodiments, the sample DNA is whole genomic DNA.
[0037] In some instances, genomic DNA and mitochondrial DNA may be obtained separately from the same biological sample or source. Many different methods and technologies are available for the isolation of genomic DNA and mitochondrial DNA. In general, such methods involve disruption and lysis of the starting material followed by the removal of proteins and other contaminants and finally recovery of the DNA. Removal of proteins can be achieved, for example, by digestion with proteinase K, followed by salting-out, organic extraction, gradient separation, or binding of the DNA to a solid-phase support (either anion-exchange or silica technology). Mitochondrial DNA may be isolated similarly following initial isolation of mitochondria. DNA may be recovered by precipitation using ethanol or isopropanol. There are also commercial kits available for the isolation of nuclear or mitochondrial DNA. The choice of a method depends on many factors including, for example, the amount of sample, the required quantity and molecular weight of the DNA, the purity required for downstream applications, and the time and expense.
[0038] The methods of the present disclosure, in certain embodiments, utilize mild enzymatic and chemical reactions that avoid the substantial degradation associated with methods like bisulfite sequencing. Thus, the methods are useful in analysis of low-input samples, such as circulating cell-free DNA , circulating tumor DNA, and in single-cell analysis.
[0039] In some embodiments, the DNA sample is circulating cell-free DNA (cfDNA), which is DNA found in the blood and is not present within a cell. cfDNA can be isolated from blood or plasma using methods known in the art. Commercial kits are available for isolation of cfDNA including, for example, the Circulating DNA Kit (Qiagen). The DNA sample may result from an enrichment step, including, but is not limited to antibody immunoprecipitation, chromatin immunoprecipitation, restriction enzyme digestion-based enrichment, hybridization-based enrichment, or chemical labeling-based enrichment.
[0040] In some instances, the isolated DNA is fragmented into a plurality of shorter double stranded DNA target fragments. In general, fragmentation of DNA may be performed physically, or enzymatically.
[0041] For example, physical fragmentation may be performed by acoustic shearing, sonication, microwave irradiation, or hydrodynamic shear. Acoustic shearing and sonication are the main physical methods used to shear DNA. For example, the Covaris® instrument (Woburn, MA) is an acoustic device for breaking DNA into 100 bp - 5 kb. Another example is the Bioruptor® (Denville, NJ), a sonication device utilized for shearing chromatin, DNA and disrupting tissues. Small volumes of DNA can be sheared to 150 bp - 1 kb in length. The Hydroshear® from Digilab (Marlborough, MA) is another example and utilizes hydrodynamic forces to shear DNA. Nebulizers, such as those manufactured by Life Technologies (Grand Island, NY) can also be used to atomize liquid using compressed air, shearing DNA into 100 bp - 3 kb fragments in seconds. As nebulization may result in loss of sample, in some instances, it may not be a desirable fragmentation method for limited quantities samples. Sonication and acoustic shearing may be better fragmentation methods for smaller sample volumes because the entire amount of DNA from a sample may be retained more efficiently. Other physical fragmentation devices and methods that are known or developed can also be used.
[0042] Various enzymatic methods may also be used to fragment DNA. For example, DNA may be treated with DNase I, or a combination of maltose binding protein (MBP)-T7 Endo I and a non-specific nuclease such as Vibrio vulnificus nuclease (Vvn). The combination of nonspecific nuclease and T7 Endo synergistically work to produce non-specific nicks and counter nicks, generating fragments that disassociate 8 nucleotides or less from the nick site. In another example, DNA may be treated with NEBNext® dsDNA Fragmentase® (NEB, Ipswich, MA). NEBNext® dsDNA Fragmentase generates dsDNA breaks in a time-dependent manner to yield 50-1,000 bp DNA fragments depending on reaction time. NEBNext dsDNA Fragmentase contains two enzymes, one randomly generates nicks on dsDNA and the other recognizes the nicked site and cuts the opposite DNA strand across from the nick, producing dsDNA breaks. The resulting DNA fragments contain short overhangs, 5 '-phosphates, and 3 '-hydroxyl groups.
[0043] In some instances, the DNA sample is fragmented into specific size ranges of target fragments. For example, the DNA sample may be fragmented into fragments in the range of about 25-100 bp, about 25-150 bp, about 50-200 bp, about 25-200 bp, about 50-250 bp, about 25-250 bp, about 50-300 bp, about 25-300 bp, about 50-500 bp, about 25-500 bp, about 150-250 bp, about 100- 500 bp, about 200-800 bp, about 500-1300 bp, about 750-2500 bp, about 1000- 2800 bp, about 500-3000 bp, about 800-5000 bp, or any other size range within these ranges. Forexample, the DNA sample may be fragmented into fragments of about 50-250 bp. In some instances, the fragments may be larger or smaller by about 25 bp.
[0044] In certain embodiments, the fragments are treated to produce blunt ends that are compatible with ligation to a first adapter having a compatible blunt end. Any convenient method for producing blunt ends may be employed, including treatment with one or more enzyme having 5' and / or 3' single strand exonuclease activity (e.g., E. coli Exonuclease III) and / or performing a fill-in reaction to extend 3' recessed ends (e.g., with T4 DNA polymerase). No limitation in this regard is intended.
[0045] For epigenetic analysis, in certain embodiments, the DNA target fragments may be any DNA fragment, derived from a biological sample, having a sequence of interest that may or may not include epigenetic modifications or DNA damage to one or more nucleobases. In some aspects, the DNA target fragments may include cytosine modifications (i.e., 5-mC, 5-hmC, 5-fC, and / or 5-caC). The DNA target fragments can be a single DNA molecule in the sample, or may be the entire population of DNA molecules in a sample (or a subset thereof) having, e.g., a cytosine modification. The DNA target fragments can comprise a plurality of DNA sequences such that the methods described herein may be used to generate a library of DNA target fragments that can be analyzed individually (e.g., by determining the sequence of individual targets) or in a group (e.g., by multiplexed DNA sequencing methodologies).
[0046] In some embodiments, the methods described herein include the step of adding adapter DNA molecules to double stranded DNA target fragments. An adapter DNA, or DNA linker, is a short, chemically- synthesized, single- or double-stranded oligonucleotide that can be ligated to one or both ends of other DNA molecules. Double-stranded adapters can be synthesized so that each end of the adapter has a blunt end or a 5' or 3' overhang (i.e., sticky ends). DNA adapters are ligated to the DNA target fragments to provide sequences for, e.g., primer extension reactions and sequencing reactions with complimentary primers and / or for bioinformatic analysis (e.g., clustering of related sequences into families based on shared unique molecular identifier barcodes, UMIs).
[0047] Prior to ligation of adapters, the ends of the DNA fragments can be prepared for ligation. For example, by end repair and creating blunt ends with 5’ phosphate groups. Fragmented DNA may be rendered blunt-ended by a number of methods known to those skilled in the art. In a particular method, the ends of the fragmented DNA are “polished” with T4 DNA polymerase and Klenow polymerase, a procedure well known to skilled practitioners, and then phosphorylated with a polynucleotide kinase enzyme. A single ‘A’ deoxynucleotide is then added to both 3' ends of the DNA molecules using Taq polymerase orKlenow exo minus polymerase enzyme, producing a one-base 3' overhang that is complementary to the one-base 3' ‘T’ overhang on the double-stranded end of an adaptor.
[0048] In some instances, the adapters may include two oligonucleotides that are partially complementary such that they hybridize to form a region of double stranded sequence, but also retain a region of single stranded, non-hybridized sequence. The region of single stranded sequence may include “universal” oligonucleotide binding sequences, enabling all target fragments in a library to bind to the same oligonucleotide, which may be a capture oligonucleotide, to localize target fragments to a solid-support, an oligonucleotide primer for a primer extension reaction, a PCR primer, sequencing primer, or combinations thereof. In certain instances, the adapters may include two regions of single- stranded, non-hybridized sequence (i.e., a first, 5’ single stranded region and a second, 3’ single stranded region). This configuration is known in the art as a “Y” adapter. The first and second single stranded regions of a Y adapter are not complementary and may include different primer hybridization sequences and other features.
[0049] The portions of the two single stranded regions of the adapters typically include at least 10, or at least 15, or at least 20 consecutive nucleotides on each strand. The lower limit on the length of the single stranded regions will typically be determined by function, for example, the need to provide a suitable sequence for binding of a primer for primer extension, PCR and / or sequencing. Theoretically there is no upper limit on the length of the single stranded regions, except that in general it is advantageous to minimize the overall length of the adapter, for example, in order to facilitate separation of unbound adapters from adapter-ligated double stranded DNA target fragments following the ligation step. Therefore, it is preferred that the single stranded regions should be fewer than 50, or fewer than 40, or fewer than 30, or fewer than 25 consecutive nucleotides in length on each strand.
[0050] As used throughout, the term “hairpin adapter” refers to a nucleic acid sequence that has two complementary regions that hybridize to one another to form a double-stranded region with the two complementary regions being connected by a single-stranded loop. The hairpin adapters described herein can be of any length suitable for use in the provided methods. For example, the hairpin adapters can be at least 10, at least 20, at least 30, at least 40, or at least 50, nucleotides in length or longer. Optionally, the hairpin adapters are 15 to 40 base pairs in length.
[0051] The double stranded region of the adapter is a short double stranded region, typically comprising 5 or more consecutive base pairs, formed by annealing of the two partially complementary polynucleotide strands. Generally, it is advantageous for the double stranded region to be as short as possible without a loss of function. By “function” in this context ismeant that the double stranded regions forms a stable duplex under standard reaction conditions for the enzyme-catalyzed nucleic acid ligation reaction.
[0052] The precise nucleotide sequence of the adapters is generally not material to the invention and may be selected by the user such that the desired sequence elements are ultimately included in the common sequences of the library of adapter-ligated double stranded DNA target fragments. Additional sequence elements may be included, for example, to provide binding sites for primers which will ultimately be used in sequencing of complementary copy strands of the DNA target fragments. The adapters may further include “tag” sequences, unique molecular identifiers (UMI), and / or sample identifier sequences, which can be used to tag, track, and differentiate target fragments and complementary copies thereof derived from a particular source. The general features and use of such sequences is well known in the art.
[0053] The ends of the single stranded regions of the adapters may be biotinylated or bear another functionalities that enables it to be captured, or immobilized, on a surface, such as a solid support. Alternative functionalities other than biotin are known in the art, e.g., as described in Applicant’s published Patent Application no. WO2020 / 172479 entitled, “Methods and Devices for Solid-Phase Synthesis of Xpandomers for use in Single Molecule Sequencing”, which is herein incorporated by reference in its entirety.
[0054] “Ligation” of adapters to the 5' and 3 ' ends of each fragmented double stranded nucleic acid target fragment involves joining of the two polynucleotide strands of the adapter to the double-stranded target polynucleotide such that covalent linkages are formed between both strands of the two double-stranded molecules. Preferably such covalent linking takes place by formation of a phosphodiester linkage between the two polynucleotide strands but other means of covalent linkage (e.g., non-phosphodi ester backbone linkages) may be used. However, it is an essential requirement that the covalent linkages formed in the ligation reactions allow for read- through of a polymerase, such that the resultant construct can be copied in a primer extension reaction using primers which bind to sequences in the regions of the adapter-target construct that are derived from the adapter molecules.
[0055] In some instances, the adapters and DNA target fragments may be incubated with a ligase to covalently link the adapters and DNA target fragments. Ligase catalyzes the formation of a phosphodiester bond between juxtaposed 5' phosphate and 3 ' hydroxyl termini in duplex DNA or RNA. The enzyme will join blunt end and cohesive end termini as well as repair single stranded nicks in duplex DNA. An exemplary ligase is T4 ligase, which is the most frequently used enzyme for cloning. Another ligase that may be used is E. coli DNA ligase, which preferentially connects cohesive double-stranded DNA end but is also active on blunt ends DNAin the presence of Ficoll or polyethylene glycol. Another ligase that may be used is DNA ligase Ilia, which is known to function in mitochondria.
[0056] The products of the ligation reaction may be subjected to purification steps in order to remove unbound adapter molecules before the adapter-target constructs are processed further.
[0057] The ligation of adapters to both free ends of the double stranded DNA target fragments gives rise to a pool of adapter-ligated double stranded DNA target fragments with adapters at the 5’ and 3’ ends of the target.
[0058] There are several standard methods for separating the strands of an adapter-ligated double stranded DNA target fragment by denaturation, including thermal denaturation, or chemical denaturation in either 100 mM sodium hydroxide solution or formamide solution. The pH of a solution of single-stranded DNA fragments can be neutralized by adjusting with an appropriate solution of acid, or preferably by buffer-exchange through a size-exclusion chromatography column pre-equilibrated in a buffered solution, including SPRI and others for buffer exchange or size selection.
[0059] As used herein, the term “complementary” refers to nucleic acid sequences that are capable of forming Watson-Crick base-pairs. For example, a complementary sequence of a first sequence is a sequence which is capable of forming Watson-Crick base-pairs with the first sequence. The term “complementary” does not necessarily mean that a sequence is complementary to the full-length of its complementary strand, but the term can mean that the sequence is complementary to a portion thereof. Thus, in some embodiments, complementarity encompasses sequences that are complementary along the entire length of the sequence or a portion thereof. For example, two sequences can be complementary to each other along at least 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% of the length of the sequence. Here, the term “sequence” encompasses, but is not limited to, nucleic acid sequences, polynucleotides, oligonucleotides, probes, primers, primer-specific regions, and target-specific regions. Despite any mismatches, the two sequences should have the ability to selectively hybridize to one another under appropriate conditions.
[0060] One embodiment of a method and constructs of the present invention is set forth in FIG. 1. As shown in FIG. 1A, a double stranded target fragment 100 is derived from genomic DNA or some other nucleic acid source. The double stranded target fragment 100 includes parent (+) strand 100A (i.e., a sense strand) and parent (-) strand 100B (i.e., an antisense strand). In Step 1, extendable Y adapters 103 and 105 are joined to the ends of the double stranded target fragment. As shown in this embodiment, the extendable Y adapters have the same overall structure and each includes a region of double stranded DNA and three regions of single stranded DNA. The extendable Y adapters may include one or more of the following features: a free 3’end (103a and 105a) provided in the double stranded portion of the Y adapter that functions as an initiation site for polymerase-mediated DNA synthesis, as denoted by the arrows in the figure; a 5’ single stranded tail (103b and 105b) in the double stranded portion of the Y adapter, that may be of any suitable length; a designed UMI sequence (103c and 105c) in the double stranded portion of the Y adapter; a SID sequences (103d and 105d) in a single stranded portion of the Y adapter; primer hybridization sequences (103e and 105e) in a single stranded portion of the Y adapter; and a blocker hybridization sequence ( 103f and 105 f) in a single stranded portion of the Y adapter. In some embodiments, 5’ single stranded tails 103b and 105b may facilitate downstream DNA synthesis by a strand displacing DNA polymerase. In certain embodiments, the length of the 5’ single stranded tails may be from one to two, from one to three, from one to four, from one to five, or over 5 nucleotides in length. In certain embodiments, the extendable Y adapters may include additional sequences, or other features, that mediate downstream steps of the workflow. For example, the adapters may include nucleic acid sequences or other chemical moieties for immobilization of the template constructs on a solid support.
[0061] In step 1, the extended Y adapters are joined to the ends of the double stranded target fragment by DNA ligase-mediated. In this embodiment, proper alignment of the adapters and the target fragment may be facilitated by hybridization of a single 3 ’ T overhang in the double stranded portion of the Y adapter and a single 3 ’ A overhang in the double stranded target fragment. In this embodiment, ligation is referred to as “complete ligation” in that both strands of the target fragment at each end are ligated to a strand of the extendable Y adapter. The extendable adapter-ligated double stranded target fragment product 110 includes a single free 3’ end in each strand of the double stranded portion of the ligation product, denoted by arrows 103 a and 105a.
[0062] As shown in FIG. IB, in step 2, the extendable Y adapter-ligated double stranded target fragment construct 110 is contacted with a strand displacing polymerase which binds to the construct at the free 3’ ends, 103a and 103b. Any suitable strand displacing polymerase may be employed, including, but not limited, to Klenow fragment or BST large fragment. When placed under nucleic acid synthesis (i.e., extension reaction) conditions, the polymerase begins nucleic acid synthesis from the free 3 ' OH group provided by the extendable Y adapters using the opposite strand of the target fragment as a template, while simultaneously displacing the complementary strand of the target fragment in a 5' to 3' direction.
[0063] As the polymerases proceed along the target fragment strands in opposite directions (and on different strands of the target fragment), the two strands of the target fragment are separated, finally resulting in two nucleic acid extension products 113 and 115. Each of the nucleic acid extension products includes a strand of the original target fragment (e.g., parentstrands, 100a and 100b) and a newly synthesized complementary copy stand (e.g., daughter strands 117b and 117aa). Each of the nucleic acid extension products includes a fully ligated adapter at one end (i.e., there are no nicks or gaps between the fragment and the adapter) and either a blunt or overhang terminus at the opposite end. In the nucleic acid extension products 113 and 115 in FIG. IB, the end opposite the adapter ligated end (i.e., lacking an adapter) is expected to be blunt if the polymerase traversed the entire target fragment strand, while it will have a single-stranded 5' overhang in the parent strands if nucleic acid synthesis terminated before the polymerase reached the end of the target fragment parent strand. Further, if the polymerase employed has terminal transferase activity, there may be a 3 ' overhang instead of a blunt end if the polymerase traverses the entire target fragment. No limitation in this regard is intended. In the embodiment depicted in FIG. IB, the ends of the nucleic acid extension products 113 and 115 are A-tailed to produce single nucleotide 3’ overhangs.
[0064] In certain embodiments, the extension reaction may include modified nucleotides that weaken the duplex strength of the extension products, i.e., the strength of the hydrogen bonds between the parent template strand and the newly synthesized daughter strand. Additionally, in some embodiments, the modified nucleotides may confer resistance to methyl conversion chemistries. Exemplary, non-limiting modified nucleotides are N4-Me dCTP and 7-deaza dGTP. In embodiment illustrated in FIG. IB, daughter strands 117a and 117b incorporate modified dGTP and dCTP analogs, as denoted by the circles around “G” and “C” in these strands.
[0065] As shown in FIG. 1C, in step 3, the 5’ ends 113a and 115a of the parental target fragment strands of the extension products are activated by PNK-dependent phosphorylation. Hairpin adapters 119 are then contacted with the extension products and ligated to the double stranded ends to produce first duplex template 121 and second duplex template 123. Ligation of the hairpin adapters is facilitated by a single stranded T overhang in the hairpin adapters that is capable of base-pairing with the single A overhang in the daughter stand of the extension products. In each duplex template, the daughter strand is covalently coupled (i.e., covalently joined) to the parent strand by ligation to the intervening hairpin adapter.
[0066] As shown in FIG. ID, in step 4, the duplexed templates 121 and 123 are amplified by PCR in the presence of modified nucleotides, e.g., N4-Me dCTP and 7-deaza dGTP (as denoted by the circles around “C” and “G” in FIG. ID) to weaken the resulting duplex and, in certain embodiments, confer resistance to methyl conversion chemistries. PCR amplification is directed by primers 125 that hybridize to sequences 103 f and 105f in the single stranded portions of the Y adapters and primer 127 that hybridize to the complementary copies of sequences 103e and 105e included in the single stranded portions of the Y adapters. Products of the PCR amplificationreaction are duplexed template construct 131 that includes amplified parent antisense strand 131b and daughter sense strand 131a and duplexed template construct 133 that includes amplified parent sense strand 113a and daughter antisense strand 133b.
[0067] As shown in FIG. IE, the duplex template constructs 131 and 133 are prepared for the solid phase synthesis of Xpandomer copies of the duplexed templates. Here, extension oligonucleotide 140 for Xpandomer synthesis is joined to solid support 145 by cleavable linker moiety 147. The duplex template constructs are captured on the solid support by hybridization of adapter sequences 103e and 105e to extension oligonucleotide 140 joined to solid support 145. Invasive blocker 150 is then hybridized to the blocker hybridization sequences in the single stranded portions of the Y adapters. The extension oligonucleotides provide initiation sites for synthesis of Xpandomer copies of the duplex template constructs, while the invasive blockers terminate Xpandomer synthesis.
[0068] Xpandomer synthesis reactions are carried out in which Xpandomer copies of the duplex templates are synthesized from extension oligonucleotide primers 140 in the 5’ to 3’ direction as the parent and daughter strands of the duplex template are dissociated, or “unzipped”, which is facilitated by weak hydrogen bonding between the modified G and C nucleobases in the two strands. Full length Xpandomer copies 151 and 153 of duplexed templates 131 and 133 are illustrated in FIG. IF. From the 5’ to the 3’ direction, Xpandomer 151 includes copies of the following sequences of the duplexed template construct 131 : the antisense SID sequence 103d; the antisense UMI sequence 105c; the daughter sense strand 131a; the sense UMI sequence 103c; the hairpin adapter 119; the antisense UMI sequence 103c; the . From the 5’ to the 3’ direction, Xpandomer 153 includes copies of the following sequences of the duplexed template construct 133: the antisense SID sequence 103d; the antisense UMI sequence 103c; the daughter antisense strand 133b; the sense UMI sequence 105c; the hairpin adapter 119; the antisense UMI sequence 105c; the amplified parent sense strand 133a; and the sense UMI sequence 103c.
[0069] As Xpandomers 151 and 153 are individually passed through a nanopore to obtain sequence information, the resulting sequencing reads will include information from both strands of the duplexed template in a single read. This redundancy in sequence information offers significant advantages in the ability to assess the quality of the read data. Moreover, inclusion of UMI barcodes in the extendable Y adapters provides another level of quality assessment through the ability to bioinformatically link and compare the sequences of Xpandomer 151 and 153 that each include the sequence of a strand of the original nucleic acid target fragment.
[0070] Another embodiment of the methods and constructs of the present invention is set forth in FIG. 2. As shown in FIG. 2A, a double stranded target fragment 200 is derived fromgenomic DNA or some other nucleic acid source. The double stranded target fragment 200 includes parent (+) strand 200A (i.e., a sense strand) and parent (-) strand 200B (i.e., an antisense strand). In Step 1, extendable Y adapters 203 and 205 are joined to the ends of the double stranded target fragment. As shown in this embodiment, the extendable Y adapters each includes a region of double stranded DNA and three regions of single stranded DNA. The extendable Y adapters may include one or more of the following features: a free 3’ end (203a and 205a) provided in the double stranded portion of the Y adapter that functions as an initiation site for polymerase-mediated DNA synthesis, as denoted by the arrows in the figure; a 5’ single stranded tail (203b and 205b) in the double stranded portion of the Y adapter, that may be of any suitable length; a designed UMI sequence (203c and 205c) in the double stranded portion of the Y adapter; a SID sequences (203d and 205d) in a single stranded portion of the Y adapter; primer hybridization sequences (203e and 205e) in a single stranded portion of the Y adapter; and a blocker hybridization sequence (203f and 205 f) in a single stranded portion of the Y adapter. In certain embodiments, the extendable Y adapters may include additional sequences, or other features, that mediate downstream steps of the workflow. For example, the adapters may include nucleic acid sequences or other chemical moieties for immobilization of the template constructs on a solid support.
[0071] In step 1, the extendable Y adapters are joined to the ends of the double stranded target fragment by DNA ligase-mediated ligation, In this embodiment, proper alignment of the adapters and the target fragment may be facilitated by hybridization of a single 3’ T overhang in the double stranded portion of the Y adapter and a single 3’ A overhang in the double stranded target fragment. In this embodiment, ligation is referred to as “complete ligation” in that both strands of the target fragment at each end are ligated to a strand of the extendable Y adapter. The extendable adapter-ligated double stranded target fragment product 210 includes a single free 3’ end in each strand of the double stranded portion of the ligation product, denoted by arrows 203 a and 205a.
[0072] As shown in FIG. 2B, in step 2, the extendable adapter-ligated double stranded target fragment construct 210 is contacted with a strand displacing polymerase which binds to the construct at the free 3’ ends, 203a and 203b. Any suitable strand displacing polymerase may be employed, including, but not limited, to Klenow fragment and BST large fragment. When placed under nucleic acid synthesis (i.e., extension reaction) conditions, the polymerase begins nucleic acid synthesis from the free 3 ' OH group provided by the extendable Y adapters using the opposite strand of the target fragment as a template, while simultaneously displacing the complementary strand of the target fragment in a 5' to 3' direction.
[0073] As the polymerases proceed along the target fragment strands in opposite directions (and on different strands of the target fragment), the two strands of the target fragment are separated, finally resulting in two nucleic acid extension products 213 and 215. Each of the nucleic acid extension products includes a strand of the original target fragment (e.g., parent strands, 200a and 200b) and a newly synthesized complementary copy stand (e.g., daughter strands 217b and 217aa). Each of the nucleic acid extension products includes a fully ligated adapter at one end (i.e., there are no nicks or gaps between the fragment and the adapter) and either a blunt or overhang terminus at the opposite end. In the nucleic acid extension products 213 and 215 in FIG. 2B, the end opposite the adapter ligated end (i.e., lacking an adapter) is expected to be blunt if the polymerase traversed the entire target fragment strand, while it will have a single-stranded 5' overhang in the parent strands if nucleic acid synthesis terminated before the polymerase reached the end of the target fragment parent strand. Further, if the polymerase employed has terminal transferase activity, there may be a 3 ' overhang instead of a blunt end if the polymerase traverses the entire target fragment. No limitation in this regard is intended. In the embodiment depicted in FIG. 2B, the ends of the nucleic acid extension products 213 and 215 are A-tailed to produce single nucleotide 3’ overhangs.
[0074] In certain embodiments, the extension reaction may include modified nucleotides that weaken the duplex strength of the extension products, i.e., the strength of the hydrogen bonds between the parent template strand and the newly synthesized daughter strand. Additionally, in some embodiments, the modified nucleotides may confer resistance to methyl conversion chemistries. Exemplary, non-limiting modified nucleotides are N4-Me dCTP and 7-deaza dGTP. In the embodiment illustrated in FIG. 2B, daughter strands 217a and 217b incorporate modified dGTP and dCTP analogs, as denoted by the circles around “G” and “C” in these strands.
[0075] As shown in FIG. 2C, in step 3, the 5’ ends 213a and 215a of the parental target fragment component of the extension products are activated by PNK-dependent phosphorylation. In other embodiments, any suitable means of chemically “masking” and / or “unmasking” the 5’ or 3’ end of a nucleic acid strand may be employed to either block or promote ligation to another nucleic acid strand. Cleavable hairpin adapter 219 including cleavage site 219a is then contacted with the extension products under DNA ligation conditions and ligated to the double stranded ends to produce first duplex template 221 and second duplex template 223. Ligation of the hairpin adapters is facilitated by a single stranded T overhang in the hairpin adapters that is capable of base-pairing with the single A overhang in the daughter stand of the extension products. In each duplex template, the daughter strand is covalently coupled (i.e., covalently joined) to the parent strand by ligation to the intervening hairpinadapter. In this embodiment, cleavable hairpin adapter 219 includes extension oligonucleotide hybridization sequence 219b.
[0076] As shown in FIG. 2D, in step 4, duplex templates 221 and 223 are PCR amplified in the presence of modified nucleotides, e.g., N4-Me dCTP and 7-deaza dGTP (as denoted by the circles around “C” and “G” in FIG. 2D) to weaken the resulting duplex and, in certain embodiments, confer resistance to methyl conversion chemistries. PCR amplification is directed by primer 225 that hybridizes to sequences 203 f and 205f in the single stranded portions of the Y adapters and primer 227 that hybridizes to the complementary copies of sequences 203e and 205e in the single stranded portion of the Y adapter. Products of the PCR amplification reaction are duplex template construct 231 that includes amplified parent antisense strand 23 lb and daughter sense strand 231a covalently coupled by cleavable hairpin adapter 219 that includes cleavage site 219a and duplex template construct 233 that includes amplified parent sense strand 233a and daughter antisense strand 233b covalently coupled by cleavable hairpin adapter 219 that includes cleavage site 219a.
[0077] As shown in FIG. 2E, in step 5, duplex template constructs 231 and 233 are treated with a cleavage agent to generate a strand break in cleavable hairpin adapter 219 at cleavage site 219a. Cleavage of duplex template construct 231 yields single template constructs 241a and 241b. From the 5’ end to the 3’ end, single template construct 241a includes the following sequences: blocker hybridization sequence 205f, amplified parent antisense strand 231b, and extension oligonucleotide hybridization sequence 219b, derived from the hairpin adapter. From the 5’ end to the 3’ end, single template construct 241b includes the following sequences: amplified daughter sense strand 231a and extension oligonucleotide hybridization sequences 205e, derived from the Y adapter. From the 5’ end to the 3’ end, single template construct 243a includes the following sequences: amplified daughter antisense strand 233b and extension oligonucleotide hybridization sequences 203e, derived from the Y adapter. From the 5’ end to the 3’ end, single template construct 243b includes the following sequences: blocker hybridization sequence 203 f; amplified parent antisense strand 233a; and extension oligonucleotide hybridization sequence 219b, derived from the hairpin adapter.
[0078] As shown in FIG. 2F, step 6, the single template constructs 241a, 241b, 243 a, and 243b are prepared for the solid phase synthesis of Xpandomer copies of the templates. Here, extension oligonucleotide 250 for Xpandomer synthesis are joined to solid support 255 by cleavable linker moiety 257. The single template constructs are captured on the solid support by hybridization of adapter hybridization sequences 219b, 205e, 203e, and 219b to extension oligonucleotide 250 joined to solid support 255. Blocker oligonucleotides 260 and 265 are hybridized to 5’ terminal sequences in the single template constructs. The extensionoligonucleotides provide initiation sites for synthesis of Xpandomer copies of the single template constructs, while the blocker oligonucleotides terminate Xpandomer synthesis.
[0079] Xpandomer synthesis reactions are carried out in which Xpandomer copies of the single templates are synthesized from extension oligonucleotide primers 250 in the 5’ to 3’ direction. In this embodiment, the parent and daughter strands of the duplex constructs 231 and 233 have been physically dissociated by cleavage of the cleavable hairpin adapters. Therefore, the DNA polymerase does not necessarily require strand displacing activity and further is not required to synthesize an Xpandomer copy through an intervening hairpin structure.
[0080] Full length Xpandomer copies 251a, 251b, 253a, and 253a and 153 single template constructs 241a, 241b, 243a, and 243b are illustrated in FIG. 2G. From the 5’ to the 3’ direction, Xpandomer 251a includes copies of the following sequences of the single template construct 241a: the antisense SID sequence 203d; the antisense UMI sequence 203c; the amplified parent antisense strand 23 lb; and the sense UMI sequence 205c. From the 5’ to the 3’ direction, Xpandomer 251b includes copies of the following sequences of the single template construct 241b: the antisense SID sequence 205d; the antisense UMI sequence 205c; the amplified daughter sense strand 231a; and the sense UMI sequence 203c. From the 5’ to the 3’ direction, Xpandomer 253a includes copies of the following sequences of the single template construct 243a: the antisense SID sequence 203d; the antisense UMI sequence 205c; the amplified parent sense strand 233a; and the sense UMI sequence 203c. From the 5’ to the 3’ direction, Xpandomer 253b includes copies of the following sequences of the single template construct 243b: the antisense SID sequence 205d; the antisense UMI sequence 203c; the amplified daughter antisense strand 233b; and the sense UMI sequence 205c.
[0081] As Xpandomers 251a, 251b, 253a, and 253b are individually passed through a nanopore to obtain sequence information, the resulting sequencing reads will include information from both strands of the original double stranded target fragment. This redundancy in sequence information offers significant advantages in the ability to assess the quality of the read data. Moreover, inclusion of UMI barcodes in the extendable Y adapters provides another level of quality assessment through the ability to bioinformatically link and compare the sequences of Xpandomer 251a and 251b and 253a and 253b that each include the sequence of a strand of the original nucleic acid target fragment.
[0082] Another embodiment of the methods and constructs of the present invention is set forth in FIG. 3. As shown in FIG. 3A, a double stranded target fragment 300 is derived from genomic DNA or some other nucleic acid source. The double stranded target fragment 300 includes parent (+) strand 300A (i.e., a sense strand) and parent (-) strand 300B (i.e., an antisense strand). In Step 1, extendable Y adapters 303 and 305 are joined to the ends of the doublestranded target fragment. As shown in this embodiment, the extendable Y adapters each includes a region of double stranded DNA and three regions of single stranded DNA. The extendable Y adapters may include one or more of the following features: a free 3’ end (303a and 305a) provided in the double stranded portion of the Y adapter that functions as an initiation site for polymerase-mediated DNA synthesis, as denoted by the arrows in the figure; a 5’ single stranded tail (303b and 305b) in the double stranded portion of the Y adapter, that may be of any suitable length; a designed UMI sequence (303c and 305c) in the double stranded portion of the Y adapter; a SID sequences (303d and 305d) in a single stranded portion of the Y adapter; primer hybridization sequences (303e and 305e) in a single stranded portion of the Y adapter; and a blocker hybridization sequence (303f and 305f) in a single stranded portion of the Y adapter. In certain embodiments, the extendable Y adapters may include additional sequences, or other features, that mediate downstream steps of the workflow. For example, the adapters may include nucleic acid sequences or other chemical moieties for immobilization of the template constructs on a solid support.
[0083] In step 1, the extendable Y adapters are joined to the ends of the double stranded target fragment by ligation with a DNA ligase enzyme, In this embodiment, proper alignment of the adapters and the target fragment may be facilitated by hybridization of a single 3’ T overhang in the double stranded portion of the Y adapter and a single 3’ A overhang in the double stranded target fragment. In this embodiment, ligation is referred to as “complete ligation” in that both strands of the target fragment at each end are ligated to a strand of the extendable Y adapter. The extendable adapter-ligated double stranded target fragment product 310 includes a single free 3’ end in each strand of the double stranded portion of the ligation product, represented by arrows 303a and 305a.
[0084] As shown in FIG. 3B, in step 2, the extendable adapter-ligated double stranded target fragment construct 310 is contacted with a strand displacing polymerase which binds to the construct at the free 3’ ends, 303a and 303b. Any suitable strand displacing polymerase may be employed, including, but not limited, to Klenow fragment. When placed under nucleic acid synthesis (i.e., extension reaction) conditions, the polymerase begins nucleic acid synthesis from the free 3' OH group provided by the extendable Y adapters using the opposite strand of the target fragment as a template, while simultaneously displacing the complementary strand of the target fragment in a 5' to 3' direction.
[0085] As the polymerases proceed along the target fragment strands in opposite directions (and on different strands of the target fragment), the two strands of the target fragment are separated, finally resulting in two nucleic acid extension products 313 and 315. Each of the nucleic acid extension products includes a strand of the original target fragment (e.g., parentstrands, 300a and 300b) and a newly synthesized complementary copy stand (e.g., daughter strands 317b and 317aa). Each of the nucleic acid extension products includes a fully ligated adapter at one end (i.e., there are no nicks or gaps between the fragment and the adapter) and either a blunt or overhang terminus at the opposite end. In the nucleic acid extension products 313 and 315 in FIG. 3b, the end opposite the adapter ligated end (i.e., lacking an adapter) is expected to be blunt if the polymerase traversed the entire target fragment strand, while it will have a single-stranded 5' overhang in the parent strands if nucleic acid synthesis terminated before the polymerase reached the end of the target fragment parent strand. Further, if the polymerase employed has terminal transferase activity, there may be a 3 ' overhang instead of a blunt end if the polymerase traverses the entire target fragment. No limitation in this regard is intended. In the embodiment depicted in FIG. 3B, the blunt ends of the nucleic acid extension products 313 and 315 are A-tailed to produce single nucleotide 3’ overhangs.
[0086] In certain embodiments, the extension reaction may include modified nucleotides that weaken the duplex strength of the extension products, i.e., the strength of the hydrogen bonds between the parent template strand and the newly synthesized daughter strand. Additionally, in some embodiments, the modified nucleotides may confer resistance to methyl conversion chemistries. Exemplary, non-limiting modified nucleotides are N4-Me dCTP and 7-deaza dGTP. In the embodiment illustrated in FIG. 3B, daughter strands 317a and 317b incorporate modified dGTP and dCTP analogs, as denoted by the circles around “G” and “C” in these strands.
[0087] As shown in FIG. 3C, in step 3, the 5’ ends 313a and 315a of the parental target fragment strands of the extension products are activated by PNK-dependent phosphorylation. Cleavable hairpin adapter 319 including cleavage site 319a is then contacted with the extension products under DNA ligation conditions and ligated to the double stranded ends to produce first duplex template 321 and second duplex template 323. Ligation of the hairpin adapters is facilitated by a single stranded T overhang in the hairpin adapters that is capable of base-pairing with the single A overhang in the daughter stand of the extension products. In each duplex template, the daughter strand is covalently coupled (i.e., covalently joined) to the parent strand by ligation to the intervening hairpin adapter. In this embodiment, cleavable hairpin adapter 319 includes a primer capper hybridization sequence.
[0088] As shown in FIG. 3D, in step 4, duplex templates 321 and 323 are amplified by PCR in the presence of modified nucleotides, e.g., N4-Me dCTP and 7-deaza dGTP (as denoted by the circles around “C” and “G” in FIG. 3D) to weaken the resulting duplex and, in certain embodiments, confer resistance to methyl conversion chemistries. PCR amplification is directed by primer 325 that hybridizes to sequences 303f and in the single stranded portions of the Yadapters and primer 327 that hybridizes to the complementary copies of sequences 303e and 305e in the single stranded portions of the Y adapters. Products of the PCR amplification reaction are duplex template construct 331 that includes amplified parent antisense strand 33 lb and daughter sense strand 331a covalently coupled by cleavable hairpin adapter 319 that includes cleavage site 319a and duplex template construct 333 that includes amplified parent sense strand 333a and daughter antisense strand 333b covalently coupled by cleavable hairpin adapter 319 that includes cleavage site 319a.
[0089] As shown in FIG. 3E, in step 5a, duplex template constructs 331 and 333 are treated with a cleavage agent to generate a strand break in cleavable hairpin adapter 319 at cleavage site 319a. Cleavage of duplex template constructs 331 and 333 yields cleaved template constructs 341 and 343, respectively. Cleavage of the hairpin adapters produces single stranded terminal sequences 319b and 319, each derived from the hairpin adapter. In step 5b, primer / capper oligonucleotide 350 is hybridized to single stranded terminal sequences 319b and 319c in cleaved template constructs 341 and 343. The Primer / capper oligonucleotide is designed to include nucleotide sequences complementary to the sequences of terminal sequences 319b and 319c, derived from the hairpin adapter, joined by a non-hybridizing single stranded intervening sequence. The primer / capper sequence complementary to single stranded terminal sequence 319b provides terminal 3 ’ end 355 that is capable of being extended by a DNA polymerase and, as such, can serve as an initiation site for synthesis of a complementary copy of the cleaved template construct. The primer / capper sequence complementary to terminal single stranded sequence 319c includes a terminal joinable cap moiety 365. The terminal joinable cap moiety is capable of being covalently joined to a newly synthesized complementary copy strand of the cleaved template construct initiating from the 3 ’ single stranded end derived from the Y adapter.
[0090] As shown in FIG. 3F, step 6, cleaved template constructs 341 and 343 are prepared for the solid phase synthesis of Xpandomer copies of the templates. Here, extension oligonucleotide 360 for Xpandomer synthesis is joined to solid support 375 by cleavable linker moiety 365. The cleaved template constructs are captured on the solid support by hybridization of adapter hybridization sequences 305e and 303e to the extension oligonucleotides joined to the solid support. Invasive blocker oligonucleotides 370 are hybridized to 5’ terminal sequences in the cleaved template constructs. The extension oligonucleotides provide initiation sites for synthesis of Xpandomer copies of the cleaved template constructs, while the blocker oligonucleotides terminate Xpandomer synthesis.
[0091] Xpandomer synthesis reactions are carried out in which Xpandomer copies of the cleaved templates are synthesized from extension oligonucleotide primers 360 and extendable ends 355 of the primer / capper oligonucleotide 350 in the 5’ to 3’ direction . In this embodiment,extension of extension oligonucleotides 360 proceeds until the polymerase encounters the joinable end 365 of the capper / primer upon which the newly synthesized complementary copy strand is covalently joined to the primer / capper oligonucleotide. The concurrent extension from the extendable end 355 of the primer / capper oligonucleotide proceeds until the polymerase reaches the invasive blocker oligonucleotide 370 is encountered, whereupon synthesis is terminated. Thus, the resulting Xpandomer products will be contiguous linear polymers that include a copy of the parent and daughter strands of the original template in a single duplexed Xpandomer.
[0092] Full length Xpandomer copies 361 and 363 of cleaved template constructs 341 and 343 are illustrated in FIG. 3G. From the 5’ to the 3’ direction, Xpandomer 361 includes copies of the following sequences of the cleaved template construct 341 : the antisense SID sequence 303d, the antisense UMI sequence 305c, the amplified daughter sense strand 331a, the sense UMI sequence 303c, the primer / capper sequence 350, the antisense UMI sequence 303c, the amplified parent antisense strand 33 lb, and the sense UMI sequence 305c. From the 5’ to the 3’ direction, Xpandomer 363 includes copies of the following sequences of the cleaved template construct 343: the antisense SID sequence 305d; the antisense UMI sequence 3, the amplified daughter antisense strand 333b, and the sense UMI sequence 305c, the primer / capper sequence 350, the antisense UMI sequence 305c, the amplified parent sense strand 333a, and the sense UMI sequence 303c.
[0093] As Xpandomers 361 and 363 are individually passed through a nanopore to obtain sequence information, the resulting sequencing reads will include information from both strands of the original double stranded target fragment. This redundancy in sequence information offers significant advantages in the ability to assess the quality of the read data. Moreover, inclusion of UMI barcodes in the extendable Y adapters provides another level of quality assessment through the ability to bioinformatically link and compare the sequences of both Xpandomers^
[0094] Another embodiment of the methods and constructs of the present invention is set forth in FIG. 4. As shown in FIG. 4A, a double stranded target fragment 400 is derived from genomic DNA or some other nucleic acid source. The double stranded target fragment includes parent (+) strand 400A (i.e., a sense strand) and parent (-) strand 400B (i.e., an antisense strand). Also provided are Y-hairpin hybrid adapters 403 and 405, which in this embodiment have identical structures. As shown in this embodiment, extendable Y-hairpin hybrid adapters 403 and 405 each include the following features: double stranded stem region 403a and 405a, respectively; 3 ’ arm region 403b and 405b, respectively, which include a 3’ single stranded region; loop region 403c and 405c, respectively, in which the loop regions includes a single stranded region and includes a first end that is covalently joined to the double stranded stemregion and a second end that is not covalently joined to the double stranded stem region in which the second end includes a sequence complementary to a sequence in the 3’ arm region that enables the second end of the loop region to hybridize to the complementary region in the 3’ arm region and form a double stranded region that provides free 3 ’end, 403 f and 405f, respectively; a cleavable linker 403d and 405d, respectively, which is positioned at the first end of the single stranded loop region (e.g., the cleavable linker is interposed between the loop region and the double stranded stem region), biotin moiety 403e and 405e, respectively, which is joined to the double stranded stem portion by a flexible linker (e.g., the biotin moiety is interposed between the cleavable linker and the double stranded stem portion) and free 3’ ends 403f and 405f, respectively. In this embodiment, the single stranded portions of the loop regions include a “reverse polarity” phosphoramidite that inverts the polarity of the oligonucleotide during synthesis. As such, the second end of the loop region will provide a free 3’ end (i.e. free 3’ ends 403f and 405f). In this embodiment, the 3’ end of the 3’ arms of the adapters are protected with reversible protecting agent 403g and 405g. Reversible single stranded DNA protecting agents are known in the art. In some embodiment the cleavable linker may be a photocleavable linker. The Y-hairpin hybrid adapters may include additional sequences, or other features, that mediate downstream steps of the workflow. For example, the adapters may include a UMI sequence, and SID sequence, one or more oligonucleotide hybridization site, and the like.
[0095] In step 1, the Y-hairpin hybrid adapters are covalently joined to the ends of the double stranded target fragment by DNA ligase-mediated ligation. In this embodiment, proper alignment of the adapters and the target fragment may be facilitated by hybridization of a single 3’ T overhang in the double stranded portion of the adapter and a single 3’ A overhang in the double stranded target fragment. Adapter-ligated double stranded target fragment construct 410 includes free ends 403 f and 405 f, one in each strand of the ligation product.
[0096] As shown in FIG. 4B, and summarized in step 2, the adapter-ligated double stranded target fragment construct is contacted with a strand displacing polymerase which interacts with free 3’ ends 403f and 405f of the construct. Any suitable strand displacing polymerase may be employed, including, but not limited, to Klenow fragment or BST large fragment. When placed under nucleic acid synthesis (i.e., extension reaction) conditions, the polymerase begins nucleic acid synthesis from the free 3’ ends provided by the Y-hairpin hybrid adapters using the opposite strand of the construct as a template, while simultaneously displacing the complementary strand of the construct in a 5' to 3' direction.
[0097] The polymerases proceed along the construct template in opposite directions (and on different strands of the target fragment), starting with synthesis of a complementary copy of a strand of the stem portion of the adapter followed by synthesis of a complementary copy of atarget fragment strand. Accordingly, the two strands of the target fragment are separated, and extension proceeds until the polymerases reach the end of the stem portion of the opposite adapter. The final extension product is partially double stranded extension product 420. The partially double stranded extension product includes newly synthesized strand 420a that includes a newly synthesized daughter antisense copy hybridized to the sense strand of the target fragment (i.e., parental strand 400a) and a newly synthesized strand 420b that includes a newly synthesized daughter sense copy hybridized to the antisense strand of the target fragment (i.e., parental strand 400b). Newly synthesized strand 420a is joined to the parental antisense strand by single stranded loop 405c of the Y-hairpin hybrid adapter 405. Likewise, newly synthesized daughter strand 420b is joined to the parental sense strand of the original target fragment by single stranded loop 403 c of the Y-hairpin adapter 403.
[0098] In certain embodiments, the extension reaction may include modified nucleotides that weaken the duplex strength of the extension products, i.e., the strength of the hydrogen bonds between the parent template strand and the newly synthesized daughter strand. Additionally, in some embodiments, the modified nucleotides may confer resistance to methyl conversion chemistries. Exemplary, non-limiting modified nucleotides are N4-Me dCTP and 7-deaza dGTP. In embodiment illustrated in FIG. 4B, newly synthesized strand 420a and 420b incorporate modified dGTP and dCTP analogs, as denoted by the circles around “G” and “C” in these strands.
[0099] In step 3, partially double stranded extension product 420 is treated with denaturing conditions to produce single stranded products 430 and 435. The single stranded products are capture on a solid support by contacting the biotin moieties with streptavidin-coated beads 440 bound to solid support 445.
[0100] As shown in FIG. 4C, in step 4, in certain embodiments, 3’ single stranded end 405b of the adapter is deprotected. In this depiction, only single stranded product 430 is shown for the sake of simplicity. Protecting agent 405g is removed from the end of the single stranded arm of the Y adapter (illustrated here by the scissor symbol). In this embodiment, splint oligonucleotide 450 is provided that includes a 5’ sequence that is complementary to the sequence of single stranded loop 403 c of the adapter and a 3’ sequence that is complementary to a sequence in 3’ single stranded arm 405b of the adapter. Splint oligonucleotide 450 is contacted with single stranded product 430 under suitable nucleic acid hybridization conditions. The 5’ sequence of the splint oligonucleotide hybridizes to the single stranded loop of the adapter while the 3’ sequence of the splint oligonucleotide remains unhybridized. The nucleic acid hybridization conditions may be controlled to, e.g., apply a thermal bias that enables the splint oligonucleotide to preferentially hybridize to the single stranded loop of the adapter in this step.
[0101] In step 5, conditions are applied to enable 3’ single stranded arm portion 405b of the adapter to hybridize to the complementary 3’ sequence in splint oligonucleotide 450. This results in circularization of the region of the single stranded product that includes original parental sense strand 440a to produce partially circularized product 460 bound to solid support 445. The circularization of the product brings the free 3’ end of the single stranded arm of the adapter into proximity with photocleavable linker 403d.
[0102] As show in FIG. 4D, in step 6, partially circularized product 460 is treated with photocleaving conditions to cleave photocleavable linker 403 d. The releases the 5’ end of the circularized region of the product that is bound to the solid support and provides a new free 5 ’ end in the loop portion of the adapter at the cleavage site. Cleaved product is contacted with a DNA ligase enzyme under DNA ligation conditions to ligate the new free 5’ end to the proximal free 3’ end of 3’ single stranded arm 405b of the adapter at ligation site 465 to produce single stranded ligation product 470. Single stranded ligation product 470 includes parental sense strand the of the original double stranded target fragment covalently joined to the daughter sense strand copy of the original target fragment. Advantageously, both the parental and the daughter strands are in the same orientation (e.g., in this embodiment, they are both sense strands), and thus they will not hybridize to form a double stranded product.
[0103] In step 7, single stranded ligation product 470 may be treated with denaturing conditions to remove the split oligonucleotide and provide the single stranded ligation product as a single stranded duplex template for PCR amplification or other applications.Epigenetic Detection using Duplexed or Multiplexed Template Constructs
[0104] Advantageously, any of the duplexed template constructs for Xpandomer disclosed herein may be used for epigenetic analysis, e.g., for the detection of modified nucleobases in a DNA target fragment. Exemplary epigenetic modifications detectable by the disclosed methods include, but are not limited to, 5-methylcytosine (5-mC), 5-hydroxymethylcytosine (5-hmC), 5- carboxycytosine (5-caC), f5-ormylcytosine (5-fC), 8-oxo-7,8-dihyroguanine (oxoG), uracil, methyladenine (mA), and others.
[0105] In certain embodiments, “parent-daughter” duplex constructs may be synthesized under conditions in which the daughter strand copy of the parental template strand is synthesized with “native” dNTPS. In such embodiments, the parental template strand retains endogenous epigenetic information, while the daughter strand copy retains the genetic information of the parental target fragment, as it is synthesized in vitro with native dNPTs.
[0106] The differential base modifications between daughter strands relative to the paired parental strand can be identified using any suitable conversion chemistry, or methodology,known in the art. For example, bisulfite sequencing has been an accepted standard for mapping methylomes. Sodium bisulfite chemically modifies unmethylated cytosines, causing their deamination to uracils. However, 5mC and 5hmC are not converted (see, e.g., Li, Y. and Tollefsbol, Methods Mol Biol. 2011; 719: 11-21). Sequencing distinguishes cytosines from these modified forms as they are read as thymines and cytosines, respectively. Despite its widespread use, bisulfite sequencing has significant drawbacks. For example, it requires extreme temperatures and pH, which cause depyrimidination of DNA, resulting in DNA degradation. Furthermore, unmethylated cytosines are damaged disproportionately compared with 5mC or 5hmC, resulting in bisulfite libraries that have an unbalanced nucleotide composition. All these issues give rise to libraries with reduced mapping rates and skewed GC content representation.
[0107] An alternative epigenetic detection method is based on enzymatic conversion, are also known in the art, such as TET / APOBEC or EM-SEQ (see, e.g., Vaisvila, R. et al., Genome Res. 2021 Jul; 31(7): 1280-1289.). This method detects 5mC and 5hmC using two sets of enzymatic reactions. In the first reaction, TET2 and T4-BGT convert 5mC and 5hmC into products that cannot be deaminated by APOBEC3 A. In the second reaction, APOBEC3 A deaminates unmodified cytosines by converting them to uracils. Therefore, these three enzymes enable the identification of 5mC and 5hmC. As with the bisulfite method, these conversion chemistries lead to deamination of native cytosine to uracil, while certain modified forms of cytosine remain resistant. Thus, these enzymatic conversion options also reduce the complexity of the genome as native cytosine reads as uracil during a sequencing reaction.
[0108] In certain preferred embodiments, the epigenetic detection methods of the present invention using duplexed templates constructs employ an alternative chemo-enzymatic conversion strategy developed by the inventors, referred to herein as “abasic sequencing”, which is disclosed in PCT patent application no. PCT / EP2023 / 079149, filed October 19, 2023, the contents of which are disclosed herein in their entirety.
[0109] Fig. 5A illustrates a simplified, non-limiting example of the chemo-enzymatic conversion of 5-mC to a uracil oxime mimetic. Here, a DNA target molecule including a 5-mC residue is treated with a TET enzyme (I) to convert 5-mC to 5-caC, and a TDG glycosylase (II) to excise the 5-caC nucleobase and generate an abasic site. In this example, the DNA target is also treated with an aminoxyalkyl uracil mimetic (III), which chemically reacts with the abasic site to form a stable oxime mimetic adduct (IV). In this embodiment, aminoxyalkyl uracil mimetic (III) is l-[2-(aminooxy)ethyl] -uracil. Advantageously, the inventors have found that both the enzymatic conversion and excision of 5-mC with TET and TDG as well as the chemical conversion of the abasic nucleotide to the stable oxime adduct can be performed in a single reaction, i.e., a “one-pot” reaction. This one-pot reaction is also referred to herein as a “chemo-enzymatic nucleobase conversion reaction”. Importantly, the oxime mimetic adduct (IV) is capable of base-pairing with adenine and thus is read as uracil during DNA sequencing.
[0110] FIG. 5B illustrates a simplified, non-limiting example of how chemo-enzymatic conversion of 5-mC to the uracil oxime mimetic can be used in the detection of 5-mC in a duplexed DNA template of the present invention. Here, as discussed with reference to FIG. 5A, the 5-mC residues in the parental template strand are susceptible to steps (I) through (IV) that convert 5-mC to the uracil oxime mimetic. In contrast, the daughter strand copy of the duplex template is resistant to the conversion, as it incorporates native nucleotides, such that native G is incorporated into the daughter strand opposite positions of 5-mC in the parental template. For sequence comparison analysis, the duplexed template, including a converted parental strand covalently joined to an unconverted daughter strand by an intervening hairpin adapter serves as a template for the Sequencing by Expansion (SBX®) protocol (VI), as described further herein. The resulting sequencing reads of the daughter strand portion of the duplex template will indicate “G” at each of the positions of 5-mC in the daughter template, while the sequencing reads of the parent strand will indicate “A” at each of the positions of 5-mC in the parental template. Thus G: A mispairs in the sequence of the Xpandomer copy of the duplexed template reveal the positions of 5-mC in the original target fragment.
[0111] In other embodiments of the present invention, epigenetic detection methods may utilize the parent-parent duplexed template for Xpandomer synthesis. In some embodiments, prior to conversion, the parent-parent duplexed template may optionally be copied into a daughter-daughter duplexed template to provide a duplexed reference sequence encoding the genetic information of the parental template.Glycosylase-Mediated Excision of Modified Nucleobases
[0112] In one aspect, the methods of the present invention include the step of treating the duplex template constructs with a DNA glycosylase enzyme to specifically excise the modified base of interest. Many DNA glycosylases are known in the art, targeting a wide range of specifically modified nucleobases and DNA damage elements, including sequence mismatches and a large range of epigenetic modifications. Exemplary epigenetic modifications detectable by the described methods include, but are not limited to, 5 -methylcytosine (5-mC), 5- hydroxymethylcytosine (5-hmC), 5-carboxycytosine (5-caC), f5-ormylcytosine (5-fC), 8-oxo- 7,8-dihyroguanine (oxoG), uracil, methyladenine (mA), and others.
[0113] There are two main classes of DNA glycosylases: monofunctional and bifunctional. Monofunctional glycosylases have only glycosylase activity and cleave the V-glycosidic bond linking a damaged or modified nucleobase to the sugar-phosphate backbone of DNA. All DNAglycosylases cleave glycosidic bonds, but differ in their base substrate specificity and in their reaction mechanisms, Bifunctional glycosylases also possess apurinic or apyrimidinic site (AP) lyase activity that enables them to cut the phosphodiester bond of DNA at a base lesion, creating a single-strand break.
[0114] A non-limiting list of exemplary DNA glycosylases that are useful in the methods of the present invention are set forth in Table 1. In some instances, one or more of the DNA glycosylases listed in Table 1 may be used in the described methods to excise modified bases of interest from DNA target fragments. While select DNA glycosylases are specifically identified in this disclosure, it is understood that any suitable DNA glycosylase can be used in the performing the base excision step of the described methods.Table 1DNA Glycosylases
[0115] In one embodiment, the present methods utilize a DNA glycosylase that acts directly on 5-mC, i.e., a glycosylase that is capable of hydrolyzing the glycosidic bond between the 5-mC residue and the sugar-phosphate backbone. For example, a suitable DNA glycosylase thatdirectly excises 5-mC may be a member of the DEMETER (DME) family of DNA glycosylases, e.g., DME, ROS1, or DMEL. The DME gene of Arabidopsis encodes a 1,729 amino acid protein with a centrally located DNA glycosylase domain (amino acids 1167-1368) that includes a helix-hairpin-helix (HhH) motif. The HhH motif in DME catalyzes excision of 5-mC (see, e.g., Choi et al., 2002. Cell 110:33-42). In certain embodiments, the DME glycosylase may be a variant that comprises amino acids 1167-1368 but lacks certain other regions of the protein.
[0116] In some instances, a suitable DNA glycosylase that acts directly on 5-mC may be an orthologue of DME. As used herein, the term “orthologue” means one of two or more homologous gene sequences found in different species.
[0117] In instances where the DNA glycosylase is a bifunctional enzyme, the glycosylase (e.g., DME, or an orthologue thereof), may be mutated to inactivate lyase activity, while still retaining glycosylase activity. The reaction mechanism of bifunctional DNA glycosylases is well known in the art (see, e.g., Scharer and Jiricny. 2001. Bioessays 23: 270-281). In some cases, a conserved aspartic acid acquires a proton from a conserved lysine residue that attacks the Cl’ carbon of the deoxyribose ring, creating a covalent DNA-enzyme intermediate. Beta or gamma elimination reactions release the enzyme from the DNA and cleave one of the phosphodiester bonds. Mutant forms of DME in which the invariant aspartic acid at position 1304 or the lysine at position 1286 have been altered (e.g., variants D1304N or K1286Q) been shown to reduce DNA glycosylase activity while preserving enzyme structure and stability (see, e.g., Fromme et al. 2004 Nature 427: 652-656).
[0118] Other mutations that inactivate or optimize suitable features of the DNA glycosylase are also contemplated by the present invention. For example, the DNA glycosylase may be engineered to increase its stability and / or solubility. The DNA glycosylase may also be engineered to optimize for a desired substrate specificity.
[0119] In certain embodiments, thymine DNA glycosylase (TDG) may be used to excise its known targets, 5-carboxycytosine (5-caC) and 5 -formylcytosine (5-fC). In further embodiments, TDG may be used to identify 5-methylcytosine (5-mC) and 5-hydroxymethylcytosine (5-hmC), which are modified bases that it does not specifically recognize. For example, DNA target fragments may also be treated with a ten eleven translocation (TET) enzyme prior to treatment with TDG. The TET family proteins included three human proteins (TET1, TET2, and TET3) and are cytosine oxygenases that catalyze the conversion of 5-methylcytosine (5-mC) into 5- hydroxymethylcytosine (5-hmC). 5-hmC can be further oxidized into 5-formylcytosine (5-flC) and 5-carboxylcytosine (5-caC) by TET proteins (see, e.g., Parker, et. al. 2019. Biochemistry 58: 450-467). In another instance, a suitable TET enzyme may be any TET orthologue, e.g., ngTET, isolated from Naegleria (see, e.g., Hashimoto, et. al. 2014. Nature 506(7488): 391-395).Thus, in certain embodiments, TDG may be used to excise any existing 5-caC and 5-fC modified bases present in a DNA target fragment also treated with a TET enzyme.
[0120] Other comparable methods for altering the selective excision of modified bases are possible according to the present invention. For example, a similar method may be performed to detect the same bases discussed above using thymine DNA glycosylase (TDG) and uracil DNA glycosylase (UDG).Stabilization of Abasic Sites in DNA Target Fragments
[0121] Advantageously, according to the methods of the present invention, abasic sites generated in DNA target fragments may be protected from further degradation with a stabilizing agent. In certain embodiments, a suitable stabilizing agent may be a chemical that covalently binds to the abasic site to form a stable abasic adduct. Certain aldehyde-reactive compounds are known to react with the open-ring aldehyde form of the abasic site to create stable open structures, that are referred to herein abasic adducts. Abasic adducts are refractory to enzymatic activity (e.g., lyase-mediated degradation) or to degradation-inducing chemical conditions, such as high pH. Some exemplary, non-limiting, structural classes of aldehyde-reactive stabilizing agents are described below. Each class varies in reaction rates, stability, and size of the resulting protected adduct product. The chemical properties of each abasic adduct product provide different chemoenzymatic properties with regard to duration of stabilization and suitability as a template for extension by a DNA polymerase.
[0122] In one embodiment, suitable stabilizing agents may be from the group of O- hydroxylamines (compound Illa), which are a class of compounds known to react with the aldehydic group of the open-ring form of the abasic site to create very stable oxime structures that are refractory to P-elimination by enzymatic activity (e.g., AP or dRp lyases) or by high pH.Aminoxyalkyl Nucleobase Mimetics
[0123] As discussed herein, and with reference to FIG. 6, certain chemical stabilizing agents react with abasic sites in DNA to form a stable oxime adduct that prevents subsequent degradation of the phosphodiester backbone. As used herein, the term “oxime” refers to an organic compound belonging to the imines, with the general formula, RR’C=N-OH, where R is an organic side chain and R’ may be hydrogen, forming an aldoxime, or another organic group, forming a ketoxime. O-substituted oximes form a closely related family of compounds. One particularly useful class of stabilizing agents used to form oxime adducts are those with the generalized aminoxyalkyl structure, H2N-O-R, as disclosed herein. Advantageously, the inventors have discovered that certain oximes have the further capability to biologically mimicthe Watson-Crick base-pairing activity of natural nucleobases. Thus, they not only stabilize abasic sites, but also direct incorporation of specific nucleotides at opposing sites during daughter strand synthesis. Such aminoxyalkyl-based stabilizing reagents and their corresponding oxime adduct products may be referred to in certain embodiments herein alternatively as, “nucleobase mimetics”, “aminoxyalkyl nucleobase mimetics”, or “nucleobase oxime mimetics.”
[0124] In one embodiment, the uracil mimetic, l-[2-(amino)ethyl]-uracil, is used to stabilize abasic sites, as the aminoxyalkyl constituent of the mimetic compound reacts with the abasic site to form a stable oxime adduct. Advantageously, the heterocycle constituent of the compound is able to from Watson-Crick base pairs with adenine and will thus direct incorporation of dATP during daughter strand synthesis.
[0125] Other exemplary aminoxyalkyl nucleobase mimetics suitable for the methods of the present invention include l-[3-(aminoxy)propyl]-uracil, l-[4-(aminoxy)butyl] -uracil, l-[5- (aminoxy)pentyl]-uracil, commercialy available from, e.g., Enamine Ltd. In other embodiments, the present invention contemplates new aminoxyalkyl nucleobase mimetics in which certain chemical features are optimized for particular applications. For example, mimetics may include heterocycles other than uracil, such as thymine, cytosine, guanine, or adenine. In other embodiments, the mimetics may include alternative atomic distances between the oxime and the heterocycle, e.g., from two carbons to three, four, or five carbons. Certain exemplary aminoxyalkyl nucleobase mimetics include the following: l-[2-(aminoxy)ethyl]-2,4-diiodo-5- methyl benzene, compound (A); l-[2-(aminoxy)ethyl]-2,4-dibromo-5-methyl benzene, compound (B); l-[2-(aminoxy)ethyl]-2,4-dichloro-5-methyl benzene, compound (C); l-[2- (aminoxy)ethyl] -2, 4-difluoro-5 -methyl benzene, compound (D); l-[2-(aminoxy)ethyl]-thymine, compound (E); and further prophetic pseudo uridine analogs, compounds (F) and (G).Solid-Phase Synthesis
[0126] In certain embodiments, one or more steps of generating the duplexed template constructs and / or Xpandomer synthesis may be conducted on a solid support. As used herein, the terms "solid support", “solid-state”, "solid-phase", and "substrate" may be used interchangeably and refer to a material or group of materials having a rigid or semi-rigid surface or surfaces. In many embodiments, at least one surface of the solid support will be substantially flat, e.g., a surface of a polymeric microfluidic card or chip. In some embodiments it may be desirable to physically separate regions of a card or chip for different reactions with, for example, etched channels, trenches, wells, raised regions, pins, or the like. According to other embodiments, the solid support(s) will take the form of insoluble beads, resins, gels, membranes,microspheres, or other geometric configurations composed of, e.g., controlled pore glass (CPG) and / or polystyrene.
[0127] The invention encompasses solid-phase synthesis methods in which a capture moiety is immobilized on a solid support. In certain instances, the capture moiety includes a first end covalently bound to the solid support and a second end that provides a functional group capable of binding to the 5’ end of a single stranded target sequence. As used herein, the term "immobilized", refers to the association, attachment, or binding between a molecule (e.g., linker, adapter, or oligonucleotide) and a support in a manner that provides a stable association under the conditions of elongation, amplification, ligation, and other processes as described herein. Such binding can be covalent or non-covalent. Non-covalent binding includes electrostatic, hydrophilic and hydrophobic interactions. Covalent binding is the formation of covalent bonds that are characterized by sharing of pairs of electrons between atoms. Such covalent binding can be directly between the molecule and the support or can be formed by a cross linker or by inclusion of a specific reactive group on either the support or the molecule or both. Covalent attachment of a molecule can be achieved using a binding partner, such as avidin or streptavidin, immobilized to the support and the non-covalent binding of the biotinylated molecule to the avidin or streptavidin. Immobilization may also involve a combination of covalent and non- covalent interactions.
[0128] Any suitable covalent attachment means known in the art may be used for these purposes. The chosen attachment chemistry will depend on the nature of the solid support and any derivatization or functionalities applied thereto. The extension oligonucleotide may include a moiety, which may be a non-nucleotide chemical modification, to facilitate attachment. Certain exemplary embodiments of suitable surface chemistries include conventional streptavidin / biotin interaction chemistry and involve functionalization of a solid support, e.g., with a linker moiety that includes terminal a biotin moiety. In this embodiment, the 5’ end of single stranded DNA fragment (or oligonucleotide) is bound to the linker moiety. Attachment is mediated by a streptavidin moiety provided by the 5’ end of the single stranded DNA fragment. The linker moieties disclosed herein may be of sufficient length to connect the single stranded DNA fragment to the support such that the support does not significantly interfere with primer extension reaction.
[0129] Alternatively, immobilization of a capture moiety or oligonucleotide (e.g., an extension oligonucleotide) to a solid support may be accomplished by covalent linkage of the capture oligonucleotide to the solid support via a click reaction. In this embodiment, the covalent linkage may be mediated by a maleimide-PEG-alkyne linker that is crosslinked to the solid support. An alkyne moiety provided by the end of the linker distal to the substrate iscapable of reacting with an azide group provided by the 5’ end of the capture oligonucleotide. Methods of functionalizing a solid support with maleimide-linker polymers is provided in Applicant’s published Patent Application No. WO2020 / 172479, which is herein incorporated by reference in its entirety.
[0130] In certain instances, the linkage between the capture moiety and the solid support is cleavable, enabling primer extension products to be released from the support following synthesis. Cleavable linkers and methods of cleaving such linkers are known and can be employed in the provided methods using the knowledge of those of skill in the art. For example, the cleavable linker can be cleaved by an enzyme, a catalyst, a chemical compound, temperature, electromagnetic radiation or light. Optionally, the cleavable linker includes a moiety hydrolysable by beta-elimination, a moiety cleavable by acid hydrolysis, an enzymatically cleavable moiety, or a photo-cleavable moiety. In some embodiments, a suitable cleavable moiety is a photocleavable (PC) spacer or linker phosphoramidite available from Glen Research.Sequencing by Expansion
[0131] One nucleic acid sequencing methodology that may be implemented with the present invention is “Sequencing by Expansion” (SBX®), developed by Stratos Genomics (see, e.g., Kokoris et al., U.S. Pat. No. 7,939,259, "High Throughput Nucleic Acid Sequencing by Expansion", which is herein incorporated by reference in its entirety). As previously discussed, SBX® uses biochemical polymerization to transcribe the sequence of a DNA template, e.g., a duplexed template construct, onto a measurable polymer called an “Xpandomer”. SBX® is based on the polymerization of highly modified, non-natural nucleotide analogs, referred to as “XNTPs”. XNTPs are expandable, 5' triphosphate modified non-natural nucleotide analogs compatible with template dependent enzymatic polymerization. The XNTP has two distinct functional regions; namely, a selectively cleavable phosphoramidate bond, linking the 5’ a- phosphate to the nucleobase, and a symmetrically synthesized reporter tether (SSRT) that is attached within the nucleoside triphosphoramidate at positions that allow for controlled expansion by cleavage of the phosphoramidate bond. XNTPs are described in further details in Applicant’s U.S. patent no.s 10,301,345 and 10,774,105, which are herein incorporated by reference in their entireties. The SSRT includes linkers separated by the selectively cleavable phosphoramidate bond. Each linker attaches to one end of a reporter code. XNTP substrates incorporated into daughter strand products of template-dependent polymerization are in the “constrained” configuration. The constrained configuration of polymerized XNTPs is the precursor to the expanded configuration, as found in Xpandomer products.
[0132] The transition from the constrained configuration to an expanded configurationresults from cleavage of the selectively cleavable phosphoramidate bonds within the primary backbone of the daughter strand. In this embodiment, the SSRTs include one or more reporters or reporter codes, specific for the nucleobase to which they are linked, thereby encoding the sequence information of the template. In this manner, the SSRTs provide a means to expand the length of the Xpandomer and lower the linear density of the sequence information of the parent strand.
[0133] The SSRT (i.e., “tether”) of the XNTP includes several distinct functional elements, or features, such as polymerase enhancement regions, reporter codes, and translation control element (TCEs). These features are discussed in further details in Applicant’s published PCT application WO2020 / 236526, which is herein incorporated by reference in its entirety. Each of these features performs a unique function during translocation of the Xpandomer through a nanopore to produce a series of unique and reproducible electronic signal. The SSRT is designed for controlling the rate of Xpandomer translocation by the TCE through a combination of sterics and / or electrorepulsion, Different reporter codes are sized to block ion flow through a nanopore at different measurable levels. In certain embodiments, reference is made to the “reporter construct” of the XNTP, which includes, from a proximal end to a distal end, the TCE, a symmetrical Y brancher, and two symmetric reporter code, each joined to an end of the Y brancher distal to the TCE. The reporter construct is a feature of the larger SSR structure.
[0134] Specific SSRT polymeric sequences can be efficiently synthesized using phosphoramidite chemistry typically used for oligonucleotide synthesis. Reporter codes and other features can be designed by selecting a sequence of specific phosphoramidites from commercially available and / or proprietary libraries. Such libraries include, but are not limited to, polyethylene glycol with lengths of 1 to 12 or more ethylene glycol units and aliphatic polymers with lengths of 1 to 12 or more carbon units. In certain embodiments, the SSRTs include features referred to as “polymerase enhancement regions” at the ends of the SSRTs proximal to the nucleotide triphosphoramidate diester. Polymerase enhancement regions may include positively charged polyamine spacers (e.g., primary, secondary, tertiary, or quaternary amines) or triamine spacers (three secondary amines each separated by three carbons) that facilitate incorporation of XNTP structures by a nucleic acid polymerase. In certain embodiments, the polymerase enhancement region includes two repeat units spermine
[0135] As used throughout the present disclosure, the terms “linker A” and “linker B” refer to the regions of the SSRT that each include a polymerase enhancing region and one or more translocation deceleration features or regions, and, in certain embodiments, a spacer region that includes a polymer of, e.g., PEG6, which can be customized to modulate the length of theSSRT traversed in a nanopore.
[0136] In certain embodiments, an XNTP may be a compound having the generalized structure depicted in FIG. 6.
[0137] In one embodiment, R may be H, for example, when the compounds are used to sequence a DNA template.
[0138] In certain embodiments, nucleobase is adenine, cytosine, guanine, thymine, uracil or a nucleobase analog. As one of skill in the art will appreciate, adenine, cytosine, guanine, thymine, and uracil are naturally occurring nucleobases. As used herein, the term “nucleobase analog” refers to non-naturally occurring nucleobases that are capable of forming Watson and Crick base pair with a complementary nucleobase on an adjacent single-stranded nucleic acid template.
[0139] To obtain sequence information, an Xpandomer is translocated through a nanopore, from the cis reservoir to the trans reservoir. As the Xpandomer translocates, a reporter enters the stem until its translocation control element stops at the stem entrance. The reporter is held in the stem until the TCE is enabled to pass into and through the stem, whereupon translocation proceeds to the next reporter. Upon passage through the nanopore, each of the reporter codes of the linearized Xpandomer generates a distinct and reproducible electronic signal, specific for the nucleobase to which it is linked.
[0140] In certain embodiments, Xpandomers produced by the SBX chemistry may be analyzed using a nanopore-based sequencing chip. A nanopore-based sequencing chip can incorporate a large number of sensor cells configured as an array. For example, the chip may include an array of one million cells configured in 1000 rows by 1000 columns of cells. Each cell in the array may include a control circuit integrated on a silicon substrate. Such nanoporebased sequencing chips, devices, and systems are described, e.g., in Applicant’s published patent application no. WO2021 / 219795, which is herein incorporated by reference in its entirety.
[0141] Proprietary in-house bioinformatics pipelines are typically used to process sequencing reads. The methods disclosed herein leverage UMIs to enable pairing of related sequence reads. Read pairs may be quality filtered and trimmed of adapter and primer sequences. UMI sequences may be clustered together, defining UMI-families (all reads originating from a single DNA template).Xpandomer Synthesis Reaction
[0142] The Xpandomer synthesis reaction represent a critical step in SBX®, as it is responsible for accurately transcribing the sequence of the DNA template of interest into the sequence of the Xpandomer, which is the polymer directly read by the nanopore sensor.Through trial and error, the inventors have developed a complex reaction mixture for Xpandomer synthesis, which includes a DNA polymerase and many additives that enable incorporation of the bulky XNTP substrates by the polymerase into the very large Xpandomer structure.
[0143] In certain embodiments, a non-limiting Xpandomer synthesis reaction mixture may include the following reagents: a buffer / salt system, polymerase cofactors, polymerase enhancing moieties (PEMs), a DNA polymerase, XNTP substrates, a phosphate shield molecule, a solvent, a crowding agent, and optionally, additional additives. In some embodiments, the buffer / salt system may include TrisCi and NaCl; the polymerase cofactors may include MnCh formulated in MES; the PEMs may include molecules disclosed in Applicant’s published PCT applications, WO2019 / 135975 and W02020 / 263703, which are herein incorporated by reference in their entireties; the DNA polymerase may include a variant of DPO4 polymerase as disclosed in Applicant’s U.S. patent no. s 11,299,725, 11,708,566, 11,530,392 and U.S. provisional patent application no. 63 / 591,165, filed October 18, 2023, each of which is herein incorporated by reference in their entireties; the phosphate shield molecule may include hexametaphosphate (HMP); the solvent may include NMP and DMSO; the crowding agent may include PEG8k; and the additional additives may include imidazole and betaine.
[0144] In certain embodiments, the Xpandomer synthesis reaction comprises a variant of wildtype DPO4 polymerase, designated C7326, with the following amino acid substitutions: F37T_D39L_K56Y_A57S_I59M_E63R_M76W_K78E_E79P_Q82W_Q83G_S86E_K152A_I1 53 V_A155G D156S_M157K_D 179N_P184Q_G187P_N188Y I189F_E192QI248T_S272C_V 289W_E291S_D292R_L293W_D294N_I295S_V296Q_S297Y_G299W_R300S_T301W_K32 lQ_E324K_E325K_E327KA341-352 (SEQ ID NO:2). The amino acid sequence of wildtype DPO4 polymerase is set forth in SEQ ID NO: 1, while the amino acid sequence of variant C7326 is set forth in SEQ ID NO:2. In some embodiments, a variant of DPO4 polymerase suitable for the practice of the present invention may be a variant that is at least 85% identical to SEQ ID NO:2.
[0145] One of the challenges encountered by the DNA polymerase when replicating the duplex template constructs of the present invention are double stranded regions formed when the two complementary strands of the target fragment are hybridized. DPO4 polymerase does not possess robust strand displacement activity and, as such, it does not efficiently extend through double stranded regions in a template. Indeed, the inventors have observed that the DPO4 variants commonly used in Xpandomer synthesis will prematurely “jump” from one template strand to the other as the duplex template construct is copied, thereby synthesizing incomplete Xpandomer copies of the two template strands. This phenomenon is referred herein to as a polymerase “U-turn”.
[0146] In optimizing Xpandomer synthesis conditions using the duplex template constructs of the present invention, the inventors have tested numerous biological additives and other physical forces or manipulations to reduce the occurrence of polymerase U-turns and increase the percentage of full-length duplexed Xpandomers synthesized. The following classes of additives were observed to reduce the rate of polymerase U turns during replication of duplexed template constructs: 1) single stranded binding proteins (SSBs), which are proteins that bind to and help stabilize single stranded regions of DNA and prevent formation of more stable secondary structures; 2) proteins, or enzymes, known to participate in DNA recombination processes, or to otherwise manipulate regions of single stranded DNA; 3) non-protein additives known to beneficially impact Xpanodmer synthesis (e.g., PEMs or polyphosphate analogs); 4) the biochemistry conditions of the Xpandomer synthesis reaction, including order of addition (e.g., pre-treatment of the template with an additive such as SSB prior to the Xpandomer synthesis reaction) ; 4) stretching forces proposed to enhance solid-state replication of duplexed template constructs. For example, in certain embodiments a duplex template construct may be associated with a solid support through hybridization of a sequence in its 3’ end with an extension oligonucleotide that is covalently bound to the support. The 3’ end of the extension oligonucleotide provides an initiation site for Xpandomer synthesis by a DNA polymerase, as disclosed herein. A blocker oligonucleotide can be designed to hybridize to a sequence in the opposite 5’ end of the duplex template construct. In certain embodiments, the 3’ end of the blocker oligonucleotide can be joined to a moiety that is susceptible to, e.g., an applied external force that “stretches” apart double stranded regions by overcoming the strength of the hydrogen bonds between the two complementary strands of the duplex template construct.
[0147] A non-limiting list of SBX® synthesis enhancers according to the present invention is set forth in Table 2. It is to be emphasized that the following list of exemplary additives is intended to merely illustrate one of many suitable examples of the larger genus (i.e., class) of additives, and other forces recited in Table 2, contemplated by the present invention.Table 2SBX® synthesis enhancersDiagnostic and Prognostic Methods
[0148] In particular embodiments, the methods can be directed to diagnosing an individual with a condition that is characterized by a methylation level and / or pattern of methylation at particular loci in a test sample that are distinct from the methylation level and / or pattern of methylation for the same loci in a sample that is considered normal or for which the condition is considered to be absent. The methods can also be used for predicting the susceptibility of an individual to a condition that is characterized by a level and / or pattern of methylated loci that is distinct from the level and / or pattern of methylated loci exhibited in the absence of the condition.
[0149] With particular regards to cancer, changes in DNA methylation have been recognized as one of the most common molecular alterations in human neoplasia. Hyp ermethylation of CpG islands located in promoter regions of tumor suppressor genes is a well-established and common mechanism for gene inactivation in cancer (Esteller, Oncogene 21(35): 5427-40 (2002)). In contrast, a global hypomethylation of genomic DNA is observed in tumor cells; and a correlation between hypomethylation and increased gene expression has been reported for many oncogenes (Feinberg, Nature 301(5895): 89-92 (1983), Hanada, et al., Blood 82(6): 1820-8 (1993)). Cancer diagnosis or prognosis can be made in a method set forth herein based on the methylation state ofparticular sequence regions of a gene including, but not limited to, the coding sequence, the 5'- regulatory regions, or other regulatory regions that influence transcription efficiency.
[0150] A reference genomic DNA (for example, gDNA considered “normal”) and a test genomic DNA that are to be compared in a diagnostic or prognostic method, can be obtained from different individuals, from different tissues, and / or from different cell types. In particular embodiments, the genomic DNA samples to be compared can be from the same individual but from different tissues or different cell types, or from tissues or cell types that are differentially affected by a disease or condition. Similarly, the genomic DNA samples to be compared can be from the same tissue or the same cell type, wherein the cells or tissues are differentially affected by a disease or condition.
Claims
PATENT CLAIMSWhat is claimed is:
1. A method of producing a duplex nucleic acid template, the method comprising: providing a double stranded nucleic acid target fragment comprising a first extendable Y adapter joined to a first end and a second extendable Y adapter joined to a second end, wherein the first and the second extendable Y adapters comprises a region of double stranded DNA, wherein the region of double stranded DNA comprises an internal break in one of the strands, and wherein the internal break provides an extendable 3 ’ end within the double stranded region; contacting the double stranded nucleic acid fragment to a polymerase under nucleic acid synthesis conditions, wherein the polymerase initiates template-dependent synthesis from the extendable 3 ’ ends in the first and second extendable Y adapters to produce a first and a second nucleic acid product, wherein the first and second nucleic acid products comprises a double stranded end; and joining a third adapter to the double stranded ends of the first and the second nucleic acid synthesis products, wherein the third adapter is a hairpin adapter, and wherein the hairpin adapter comprises a double stranded stem region and a single stranded loop region.
2. The method of claim 1, wherein the strand of the double stranded region of the extendable Y adapter comprising a break further comprises a 5’ single stranded tail.
3. The method of claim 1 or 2, wherein the single stranded loop region of the hairpin adapter comprises a cleavage site.
4. The method of claim 3, wherein the single stranded loop region of the hairpin adapter further comprises a hybridization sequence complementary to an extension oligonucleotide.
5. The method of claim 3, wherein the hairpin adapter further comprises a first hybridization sequence and a second hybridization sequence, wherein the first hybridization sequence is complementary to a first sequence in an invasive blocker oligonucleotide and the second hybridization sequence is complementary to a second sequence in the invasive blocker oligonucleotide.
6. The method of any one of claims 1 to 5, wherein the nucleic acid synthesis conditions comprise nucleotide analogs, wherein the nucleotide analogs form weak hydrogen bonds with complementary nucleotides relative to native nucleotides.
7. The method of claim 6, wherein the nucleotide analogs comprise N4-Me dCTP and 7- deaza dGTP.
8. The method of any one of claims 1 to 7, wherein any one of the first, second, or third adapters comprises one or more features selected from the group consisting of a UMI, a SID, an extension oligonucleotide hybridization sequence, and a blocking oligonucleotide hybridization sequence.
9. The method of any one of claims 1 to 8, wherein the polymerase is a strand displacing DNA polymerase.
10. The method of any one of claims 1 to 9, wherein the duplex nucleic acid template is subjected treatment with a DNA glycosylase enzyme and an aminoxyalkyl uracil mimetic compound, wherein the DNA glycosylase enzyme and the aminoxyalkyl uracil mimetic compound selectively converts epigenetically modified cytosine residues to uracil oxime mimetic residues.
11. A method of sequencing a duplex nucleic acid template, the method comprising: providing a sample comprising the duplex nucleic acid template according to any one of claims 1-10; amplifying the duplex nucleic acid template; generating a copy of the amplified duplex nucleic acid template, wherein the copy of the amplified duplex nucleic acid template comprises an Xpandomer, wherein the Xpandomer comprises a sequence of reporter codes, wherein the sequence of reporter codes encodes the nucleic acid sequence of the duplex nucleic acid template; and determining the sequence of the reporter codes by passing the Xpandomer through a nanopore.
12. The method of claim 11, wherein the step of generating the copy of the amplified duplex nucleic acid template comprises a solid-state synthesis Xpandomer synthesis reaction.
13. The method of claim 12, wherein the method further comprises the step of cleaving the cleavage site of the hairpin adapter prior to the step of generating the copy of the amplified duplex nucleic acid template.
14. The method of claim 12, wherein the method further comprises the step of hybridizing an invasive blocker oligonucleotide to the cleaved amplified duplex nucleic acid template, wherein the invasive blocker comprises a first end capable of being extended and a second end comprising a moiety capable of being joined to a 5’ end of an Xpandomer.
15. The method of any one of claims 11 to 14, wherein the step of generating a copy of the amplified duplex nucleic acid template comprises providing an Xpandomer synthesis reaction, wherein the Xpandomer synthesis reaction comprises the amplified duplexnucleic acid template, an extension oligonucleotide, a blocking oligonucleotide, a buffer / salt system, a polymerase cofactor, a polymerase enhancing moieties (PEMs), a DNA polymerase, XNTP substrates, a phosphate shield molecule, a solvent, and a crowding agent.
16. The method of claim 15, wherein the Xpandomer synthesis reaction further comprises one or more additives.
17. The method of claim 16, wherein the one or more additives comprises any of the additives set forth in Table 2.
18. The method of claim 16, wherein the DNA polymerase is a variant of wildtype DPO4 polymerase (SEQ ID NO: 1).
19. The method of claim 18, wherein the variant of DPO4 polymerase is at least 85% identical to the variant of SEQ ID NO:2.
20. An extendable Y adapter according to claims 1 or 2.
21. A hairpin adapter according to any one of claims 1 to 5.
22. A Y-hairpin hybrid adapter comprising a double stranded stem region, a 3’ arm region, and a loop region, wherein the 3’ arm region comprises a single stranded 3’ end region, and wherein the loop region comprises a first end, a single stranded region, and a second end, wherein the first end is covalently joined to the double stranded stem region, wherein the second end is not covalently joined to the double stranded stem region and wherein the second end is hybridized to the 3 ’ arm region to provide an extendable 3 ’ end.
23. The Y-hairpin hybrid adapter of claim 22, wherein the loop region comprises a cleavable linker interposed between the single stranded loop region and the double stranded stem region.
24. The Y-hairpin hybrid adapter of claim 23, further comprising a biotin moiety interposed between the cleavable linker and the double stranded stem region.
25. The Y-hairpin hybrid adapter of any one of claims 22 to 24, wherein the single stranded region of the loop region comprises a reverse polarity phosphoramidite.
26. Use of the Y-hairpin hybrid adapter of any one of claims 22 to 25 to produce a single stranded duplex template construct, wherein the single stranded duplex template construct comprises a parental sense strand and a daughter sense strand or a parental antisense strand and a daughter antisense strand.
Citation Information
Patent Citations
Pinned-end double-chain connector element, kit and blunt-end library building method
CN116254320A
Methods and compositions for generating asymmetrically-tagged nucleic acid fragments
US20170362639A1
High throughput nucleic acid sequencing by expansion
WO2008157696A2
Linked ligation
WO2018104908A2
Linked paired strand sequencing
WO2021178893A2