Methods of preparing nucleic acids for sequencing and methylation analysis

By attaching methylated adaptors and tagged primers with binding partners, the methods improve nucleic acid sequencing and methylation analysis, addressing cumbersome protocols and low recovery rates, enabling efficient and interference-free combined sequencing.

WO2025221668A1PCT designated stage Publication Date: 2025-10-23AGILENT TECHNOLOGIES INC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/024549
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-19
Filing Date
2025-04-14
Publication Date
2025-10-23

AI Technical Summary

Technical Problem

Current methods for nucleic acid sequencing and methylation analysis face challenges such as cumbersome protocols, low on-target rates, low recovery rates from bisulfite conversion, and interference from tags like biotin, especially for low input and damaged DNA samples, which are critical for early cancer detection.

Method used

The methods involve attaching a methylated adaptor to nucleic acids, hybridizing a tagged primer, extending it to form a primer extension product, and using binding partners to separate and subject the nucleic acids to base conversion, enabling combined sequencing runs for efficient methylation analysis.

Benefits of technology

This approach enhances the compatibility of tagging for target enrichment and separation, allowing separate treatment of input nucleic acids, and facilitates genomic and methylation analysis with improved recovery rates and reduced interference, suitable for low input samples.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000073_0000
    Figure 00000073_0000
  • Figure 00000073_0001
    Figure 00000073_0001
  • Figure 00000074_0000
    Figure 00000074_0000
Patent Text Reader

Abstract

The present invention relates to preparation, sequencing and methylation analysis of nucleic acids. Input nucleic acids are attached, to methylated adaptors or tagged adaptors, or tagged primers are hybridized to the input nucleic acids, and tagged primer extension products are synthesized. Input nucleic acid can be separated from primer extension products or amplicons, and can be treated with a base conversion reagent with or without separation. Target sequences in tagged and untagged nucleic acid molecules can be captured for target enrichment.
Need to check novelty before this filing date? Find Prior Art

Description

METHODS OF PREPARING NUCLEIC ACIDS FOR SEQUENCING AND METHYLATION ANALYSISCROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims the benefit of U.S. Provisional Application No. 63 / 636,554, filed April 19, 2024, which is incorporated by reference herein in its entirety.FIELD OF THE INVENTION

[0002] The present invention relates to preparation of nucleic acids for methylation analysis. The present invention also relates to target enrichment of nucleic acids for methylation analysis.BACKGROUND

[0003] Sequencing of nucleic acids often provides extremely valuable information from a biological sample. Next-Generation Sequencing (NGS) methods and systems involve the parallel sequencing of a library of nucleic acids by a sequencing platform. Preparation of a sequencing library generally includes various steps such as amplification of the nucleic acids, attachment of adaptors, and / or other preparatory steps. Emerging polynucleotide sequencing platforms can enable direct detection and analysis of nucleic acid molecules without the need for amplification, though the nucleic acids often require an adaptor or other moiety to immobilize the nucleic acid for sequencing steps. An adaptor can be attached to one or both ends of nucleic acid molecules in order to add sites for primer binding, for immobilization of the nucleic acid on a surface such as a flowcell or a bead, and / to add other functional sequences to the fragments. Various kinds of adaptors are used in sequencing preparation kits to add these sites or sequencesto the nucleic acids from the sample. Adaptors can be attached in various ways, such as by ligation, primer extension, tagmentation, and other techniques.

[0004] In order to obtain a suitable signal from sequencing a single DNA fragment, many sequencing systems use clonal amplification to generate many identical copies of individual DNA molecules on a solid support. These copies are segregated in individual clusters or on beads. Sequencing reactions proceed on the identical copies of the fragment in parallel, thereby producing detectable signals from the clusters or beads, with signals simultaneously detected from an enormous number of distinct clusters or beads.

[0005] A sequencing library can be generated in a variety of ways, with different objectives regarding the nucleic acids to be used as inputs. For instance, PCR can be used with targetspecific primers to generate a library of amplicons covering regions of interest in the nucleic acid sample. Other methods of library preparation involve random fragmentation of the nucleic acid sample by enzymatic or physical shearing methods, followed by amplification using common adaptor sequences. Enrichment procedures are used to remove or separate sequences of interest from the rest of the sample.

[0006] The adaptors used in many target enrichment procedures are Y-shaped adaptors (or “Y-adaptors”) which can contain duplex-UMI (unique molecular identifier) sequences. Said adaptors are made full-length through subsequent PCR and typically include a sample identifier barcode. Alternatively, full-length adaptors are ligated directly (at lower efficiency than truncated adaptors). Many Y-adaptors and other adaptors do not include a tag which supports direct immobilization of the nucleic acid molecule on a surface, or analysis of specific strands from a nucleic acid molecule.

[0007] Nucleic acid target capture methods can allow specific genes, exons, and other genomic regions of interest to be enriched for targeted sequencing or other analysis. However, target capture-based sequencing methods can involve cumbersome lengthy protocols and costly processes, as well as a low on-target rate for a small capture panel (e.g., less than 500 probes). Moreover, current methods for nucleic acid target capture can be ill-suited for low input and damaged DNA because of a low recovery rate.

[0008] Bisulfite conversion of cytosine bases can be a useful technique to study the methylation pattern of nucleic acid molecules. This employs reactions that convert cytosines to uracils without converting methylated cytosines. However, bisulfite conversion can damage nucleic acids, such as by creating truncations for example. If nucleic acid molecules are treated with bisulfite, a substantial amount of the molecules can be damaged and be unable to be recovered in the subsequent amplification steps, and thereby resulting in a low recovery rate. Moreover, because bisulfite conversion can result in single stranded or fragmented DNA and reduced sequence complexity, base-converted DNA can be a difficult input for ligation of adaptors. Bisulfite-treated cell-free (cfDNA) or circulating tumor cell DNA (ctDNA), which typically is a small initial input, can present a bigger challenge given the low recovery rate (e.g., 5% or less for bisulfite treated cfDNA). A methylation-sensitive enzymatic treatment can also be used for base conversion. However, enzymatic treatment can still suffer from a loss of methylation status during the long and multi-step process, leading to a low recovery rate.

[0009] Another technique for preparing nucleic acids for methylation analysis is TET Assisted Pyridine borane Sequencing (TAPS). TAPS employs reactions that convert modified cytosines without converting unmodified cytosines. TET enzymes are used to oxidize 5-methylcytosine and 5-hydroxymethylcytosine to 5-carboxylcytosine (5caC) and / or 5-formylcytosine (5FC),which are converted to dihydrouracil (DHU) via reduction by a borane reducing agent. PCR then converts DHU to thymine. Additional information about TAPS can be found in Song et al.US20200370114A1.

[0010] Methylation analysis of cell-free DNA holds great potential for early cancer detection. In the plasma of early stage cancer patients, the tumor content is estimated to be less than 0.1%, often down to 0.01% or lower, and therefore requires a highly sensitive assay. For methylation analysis in cancer screening, various approaches are currently used, including whole genome bisulfite sequencing (WGBS), reduced representation bisulfite sequencing (RRBS) or affinitybased enrichment, and large targeted panels containing 10,000 or more of potential methylation markers. Targeted Methylation Sequencing (TMS) provides sensitive and specific analysis of methylation markers. Additional information about TMS can be found in Lin et al. US Pat. App. Publication 20210355485A1 and Lin et al. US Pat. App. Publication 20230193380A1.

[0011] There is a need for improved methods of preparing nucleic acid molecules for sequencing and methylation analysis. There is also a need for improved methods of separating and / or distinguishing nucleic acid molecules originally present in a sample from nucleic and molecules prepared by copying or amplifying the original molecules, including methods which are compatible with target enrichment. When a tag such as biotin is attached directly to a target nucleic acid prior to target enrichment, the tag will interfere with target enrichment protocols that employ the same tag, such as biotinylated probes. There remains a need for improved methods of tagging nucleic acids for sequencing or other analysis and / or for target enrichment.SUMMARY

[0012] The present disclosure provides methods of preparing nucleic acid molecules for sequencing and methylation analysis. As one aspect, the methods comprise attaching a methylated adaptor to an input nucleic acid, and hybridizing a tagged primer to the input nucleic acid. The tagged primer comprises a first tag comprising a binding partner of a first binding pair. The methods also comprise extending the tagged primer to produce a tagged primer extension (PE) product comprising a first segment complementary to at least a portion the input nucleic acid and a second segment complementary to at least a portion of the methylated adaptor, wherein the second segment is not methylated. The methods also comprise attaching the first tag to a reciprocal binding partner of the first binding pair. The methods can further include separating the tagged PE product from the input nucleic acid by binding the binding partner of the first tag to a reciprocal binding partner of the first binding pair. The methods can also comprise subjecting the input nucleic acid to base conversion after separation from the PE product to form a base-converted input nucleic acid. The methods can also comprise combining the base-converted nucleic acid and the PE product for sequencing in a combined sequencing run.

[0013] As another aspect, methods are provided for preparing nucleic acid molecules for methylation analysis. The methods comprise attaching a first tag to an input nucleic acid, wherein the first tag comprises a binding partner of a first binding pair; and performing an amplification of the input nucleic acid to produce amplicons of the input nucleic acid. After the amplification is performed, the input nucleic acid is separated from the amplicons by binding the binding partner of the first tag to a reciprocal binding partner of the first binding pair. The methods also comprise subjecting the input nucleic acid to base conversion after separation fromthe amplicons to form a base-converted nucleic acid. The methods can also comprise combining the base-converted nucleic acid and the amplicons for sequencing in a combined sequencing run.

[0014] In some embodiments, the foregoing methods also comprise attaching a second tag to the first tag to produce a nucleic acid complex, wherein the second tag comprises a reciprocal binding partner of the first binding pair. For example, the first tag can comprise an adaptor conjugated with digoxigenin (DIG), and the second tag can comprise anti-DIG antibody. In this example, DIG and anti-DIG antibody constitute the first binding pair.

[0015] In some embodiments, the second tag also comprises a binding partner of a second binding pair, and the present methods can comprise attaching the second tag to a reciprocal binding partner of the second binding pair, which itself may be attached to the solid support. For example, the second tag can comprise biotin attached to an anti-DIG antibody, and the reciprocal binding partner of the second binding pair can be a streptavidin coated bead; in this example, biotin and streptavidin constitute the second binding pair.

[0016] As yet another aspect, methods are provided for preparing nucleic acid molecules for methylation analysis. The methods comprise attaching a methylated adaptor to an input nucleic acid and hybridizing a primer to the input nucleic acid, wherein the primer is not methylated. The methods also comprise extending the primer to produce a primer extension product comprising a first segment complementary to at least a portion the input nucleic acid and a second segment complementary to at least a portion of the methylated adaptor, wherein the second segment is not methylated. The mixture of input nucleic acid and the primer extension product is contacted with one or more base conversion reagents that either (a) convert cytosine to uracil without conversion of 5-methylcytosine, or (b) converts 5-methylcytosine (5meC) to thymine without conversion of cytosine, thereby forming either converted input nucleic acid orconverted primer extension product. The converted input nucleic acid and / or the converted primer extension product are analyzed to determine methylation status of the input nucleic acid.

[0017] Any of the foregoing methods can further comprise one or more steps for capture or target enrichment of a desired nucleic acid molecule, such as a tagged or untagged primer extension product or a tagged or untagged input nucleic acid. For instance, an untagged input nucleic acid and a tagged PE product that both contain a target sequence can be hybridized with a probe, such as a capture probe or a bridge probe which is hybridized with an anchor probe. The capture probe or anchor probe can be attached to an enrichment tag such as biotin. The first tag on the tagged PE product is different from the enrichment tag and does not have the same reciprocal binding partner.BRIEF DESCRIPTION OF THE DRAWINGS

[0018] FIG. 1 illustrates synthesis of a tagged primer extension product by extending a tagged primer along an input nucleic acid template. The tagged primer comprises a first tag and a primer that binds an input nucleic acid or an adaptor attached thereto. The first tag comprises a binding partner of a first binding pair.

[0019] FIG. 2 illustrates attachment of a first tag to an input nucleic acid and production of untagged amplicons from the tagged input nucleic acid.

[0020] FIG. 3 illustrates a target enrichment procedure in which a tagged primer extension product having an attached first tag is hybridized with bridge probes, and the primer extension productbridge probe complex is hybridized with biotinylated universal (or anchor) probes. The first tag is not biotin and does not share a reciprocal binding partner with biotin.

[0021] FIG. 4 illustrates an embodiment of the present methods of preparing an input nucleic acid for methylation analysis. Methylated adaptors comprising methylated barcodes areattached to the input nucleic acid. The input nucleic acid and amplicons thereof are treated together and sequenced in a combined sequencing run.

[0022] FIG. 5 illustrates another embodiment of the present methods of preparing an input nucleic acid for methylation analysis. Methylated adaptors comprising unmethylated barcodes attached to the input nucleic acid. The input nucleic acid and amplicons thereof are treated together and sequenced in a combined sequencing run.

[0023] The present teachings are best understood from the following detailed description when read with the accompanying drawing figures. The features are not necessarily drawn to scale. Wherever practical, like reference numerals refer to like features.DETAILED DESCRIPTION

[0024] The present disclosure can enable compatibility of tagging for target enrichment and tagging for separation of an input nucleic acid from amplicons or primer extension products. In some embodiments, the present disclosure enables the separate treatment of input nucleic acids, followed by methylation analysis. In some embodiments, the present disclosure provides for genomic and methylation analysis with a combined sequencing run.

[0025] Before the various embodiments are described, it is to be understood that the teachings of this disclosure are not limited to the particular embodiments described, and as such can, of course, vary. The section headings used herein are for organizational purposes only and are not to be construed as limiting the subject matter described in any way.Tags and Binding Partners

[0026] An input nucleic acid can be attached to a first tag comprising an adaptor and a binding partner of a first binding pair, or an input nucleic acid can be hybridized to one or more primers comprising a first tag which comprises a binding partner.

[0027] In some embodiments, the present methods employ a first tag, which may be attached to an input nucleic acid or to a PE product. In some embodiments, the first tag comprises a binding partner selected from the group consisting of digoxigenin, 5-bromo-2’- deoxyuridine (BrdU), 2,4-dinitrophenyl (DNP), nitrilotriacetic acid or a nitrilotriacetate (NTA) such as nickel nitrilotriacetate (Ni-NTA), tris-Nitrilotriacetate (tris-NTA), a tyramine, a thiol, an amine (e.g., a primary amine), an aldehyde, an alkyne or an azide or other groups reacting by click chemistry, and mixtures thereof. Accordingly, examples of first binding pairs include DIG:anti-DIG antibody, BrdU: anti -BrdU antibody, DNP:anti-DNP antibody, Ni-NTA:poly- Histidine (His-tag), Tyramide:tyrosine residues, ThiokThiol (disulfide bonds), Amine:Aldehyde conjugation, alkyne:azide, or other click chemistry reactants. Examples of tyramines include Tyramine, N-Methyltyramine, N,N-Dimethyltyramine, and N,N,N-Trimethyltyramine. Examples of alkynes and azides binding via click chemistry include copper-catalyzed reaction of an azide and alkyne to form a triazole (Huisgen 1, 3 -dipolar cycloaddition) and strain-promoted azide alkyne cycloaddition (SPAAC).

[0028] Digoxigenin is advantageous as the binding partner included in the first tag, due to its very low non-specific binding and the availability of high-affinity anti-digoxigenin antibodies. Examples of anti -DIG antibodies include Perkin Elmer’s Anti -Digoxigenin biotin conjugate.Binding A Second Tag To Produce A Nucleic Acid Construct

[0029] In some embodiments, the present methods comprise attaching a second tag to a first tag. The second tag comprises a reciprocal binding partner of the first tag, in other words, the second tag comprises a binding partner which is a member of the first binding pair. For example, the first tag can comprise an antigen or hapten and the second tag can comprise an antibody that selectively binds that antigen or hapten. The attachment of the second tag to the nucleic acid constructs can facilitate separation or isolation of an input nucleic acid from amplicons or PE products. In some embodiments, the second tag comprises, in addition to the reciprocal binding partner of the first binding pair, a binding partner which is a member of a second binding part. The binding partners of the second binding pair do not bind with the binding partners of the first binding pair. In some embodiments, the binding partner of the second binding pair comprises a biotin moiety such as biotin, 5-bromo-2’-deoxyuridine (BrdU), 2,4-dinitrophenyl (DNP), nitrilotriacetic acid or a nitrilotriacetate (NTA) such as nickel nitrilotriacetate (Ni-NTA), tri s-Nitrilotri acetate (tris-NTA), a tyramine, a thiol, an amine (e.g., a primary amine), an aldehyde, an alkyne or an azide or other groups reacting by click chemistry, and mixtures thereof. The reciprocal binding partner of the second binding pair can comprise an avidin moiety such as avidin, streptavidin, neutravidin and captavidin, or antibodies that specifically bind the binding partner, His-tags, tyrosine residues, thiols, amines, aldehydes, alkynes, azides, etc. Accordingly, examples of second binding pairs include biotin: streptavidin, BrdU: anti -BrdU antibody, DNP:anti-DNP antibody, Ni-NTA:His-tag, Tyramide:tyrosine residues, ThiokThiol (disulfide bonds), Amine:Aldehyde conjugation, alkyne:azide, or other click chemistry reactants. The proteins avidin and streptavidin form exceptionally tight complexes with biotin. In general, when biotin is coupled to a second molecule through itscarboxyl side chain, the resulting conjugate is still tightly bound by avidin or streptavidin. The second molecule is said to be "biotinylated" when such conjugates are prepared.

[0030] In some embodiments, the present methods further comprise a step of attaching a first tag to a primer or to an adaptor wherein the first tag has a binding partner of a first binding pair. The first tag can be attached to the primer or to the adaptor by a covalent bond or by non- covalent binding the first tag has a binding partner of a first binding pair. In some embodiments, the methods also comprise attaching a second tag to the first tag to produce a dual tagged complex. The second tag comprises a reciprocal binding partner of the first binding pair, and the second tag can also comprise a binding partner of a second binding pair. The second tag can be attached to the first tag by a covalent bond or by non-covalent binding. The binding partners of the first binding pair do not bind with the binding partners of the second binding pair. In general, the first tag does not comprise a member of the second binding pair and the second tag can comprise binding partners from each of the first and second binding pairs.Primer Extension Products Comprising A First Tag

[0031] As one aspect, methods are provided for preparing nucleic acid molecules for methylation analysis. The methods comprise attaching a methylated adaptor to an input nucleic acid; hybridizing a tagged primer to the input nucleic acid, wherein the tagged primer comprises a first tag comprising a binding partner of a first binding pair; extending the tagged primer to produce a tagged primer extension (PE) product comprising a first segment complementary to at least a portion the input nucleic acid and a second segment complementary to at least a portion of the methylated adaptor, wherein the second segment is not methylated; and attaching the first tag to a reciprocal binding partner of the first binding pair. The methods can further comprise separating the tagged PE product from the input nucleic acid by binding the binding partner ofthe first tag to a reciprocal binding partner of the first binding pair, such as by binding a DIG- tagged PE product to an anti-DIG antibody on a solid support. The method can also comprise subjecting the input nucleic acid to base conversion after separation from the PE product to form a base-converted nucleic acid, and combining the base-converted nucleic acid and the amplicons for sequencing in a combined sequencing run.

[0032] In some embodiments, a target enrichment step is performed on a mixture comprising tagged PE product and untagged nucleic acid molecules, such as with biotinylated probes. The enriched tagged PE products can then be separated from the enriched input nucleic acid by use of the first tag, as described above.

[0033] FIG. 1 illustrates synthesis of a primer extension (PE) product by extension of a primer attached to first tag, thereby forming a tagged PE product. More particularly, FIG. 1 shows an input nucleic acid 102 ligated at its ends to methylated adaptors 104. A tagged primer 101 comprises a first tag 103 which is a binding partner of a first binding pair. For example, the first tag 103 can comprise biotin as the binding partner of a first binding pair (biotimstreptavidin or other avidin moiety). In some embodiments, the tagged primer 101 also comprises a sample specific index and / or molecular barcode to enable strand-pairing of sequence data or to facilitate other analysis. Tagged primers 101a, 101b hybridize with the input nucleic acid 102 or the adaptor 104, PE reactions are performed to form tagged primer extension products 105 comprising a first segment 106 complementary to at least a portion the input nucleic acid and a second segment 107 complementary to at least a portion of the methylated adaptor, wherein the second segment is not methylated. The binding partner of the first binding pair, can be attached to the tagged primer 101 in any suitable manner, such as aby conjugation. In FIG. 1, both of thefirst tags 103 attached to primers 101a and 101b are the same, but in some embodiments, the first tags 103 can be different and / or have different binding partners.

[0034] The tagged PE product 105 can be separated from the input nucleic acid 102 based on binding of the first tag 103 with its reciprocal binding partner, which can be attached to a solid support 108. In some embodiments, the first tag is biotin, and the reciprocal binding partner is an avidin moiety attached to a magnetic bead 108. In other embodiments, the first tag is DIG, and the reciprocal binding partner is an anti-DIG antibody attached to a magnetic bead 108.After separation, input nucleic acid 102 are present in a first subsample 110 or fraction, and the tagged PE product 105 are present in a second subsample 112 or fraction. The subsamples 110, 112 can undergo different treatments or handling in preparation for further analysis. For instance, the first subsample 110 can be subject to base conversion (e.g., bisulfite or enzymatic treatment, as discussed in more detail below). Optionally the second subsample 112 is treated by a different procedure to prepare it for sequencing; alternatively the second subsample is not subject to further treatment. After the different treatments, the subsamples 110, 112 can be combined and indexed all together 114, such as by amplification which adds sample barcodes, thereby making a single sequencing library 116. The single sequencing library is then sequenced (as discussed below) in a combined sequencing run. After sequencing run, the result of combined library can be analyzed using an analysis algorithm 118 that separates the sequencing reads into genomic and methylation (or epigenomic) categories. This multi-omic analysis strategy has the flexibility and capacity of combining molecules from whole genome and / or target enrichment assays into a single sequencing library.

[0035] As another aspect of the present disclosure, methods are provided for tagging input nucleic acids for separation and / or treatment after target enrichment, primer extension, and / oramplification. The methods enable selection and later isolation of input nucleic acid molecules (which may include epigenetic base modifications) as part of a sample preparation workflow for nucleic acid sequencing and methylation analysis. For example, input DNA molecules can be tagged with biotin-conjugated adaptors or DIG-conjugated adaptors. Amplicons or PE products can be synthesized, and the mixture can be enriched for desired target sequences. Following enrichment, the input nucleic acid with biotin- or DIG-conjugated adaptors could be separated from amplicons or PE products which do not have a biotin or DIG tag. The input nucleic acid could then be analyzed directly (e.g., on a nucleic acid sequencing platform capable of detecting base modifications) or processed through base conversion. Amplicons from the input nucleic acid could be processed through another treatment workflow and analyzed separately such as to determine genomic variants. This method would enable methylation and variant analysis from the same location and / or in a combined sequencing run.

[0036] The present disclosure provides methods of preparing nucleic acid molecules for methylation analysis comprising attaching a first tag to an input nucleic acid, wherein the first tag comprises a binding partner of a first binding pair. The methods also comprise performing an amplification of the input nucleic acid to produce amplicons of the input nucleic acid. After the amplification is performed, the input nucleic acid is separated from the amplicons by binding the binding partner of the first tag to a reciprocal binding partner of the first binding pair. The input nucleic acid is subject to base conversion after separation from the amplicons to form a base-converted nucleic acid. The base-converted nucleic acid and the amplicons are combined for sequencing in a combined sequencing run.

[0037] FIG. 2 illustrates an embodiment of the present methods in which a first tag is attached to an input nucleic acid, followed by separation of the tagged input nucleic acid fromamplicons of that input nucleic acid. More particularly, FIG. 2 shows an input nucleic acid construct 202 in which an input nucleic acid 201 is ligated at its ends to first tags 203. Input nucleic acid 201 includes base modifications 224 such as methylation of cytosine bases. First tag 203 comprises an adaptor 204 and a binding partner 205 of a first binding pair. In FIG. 2, first tag 203 comprises a Y-adaptor as adaptor 204 and DIG as the binding partner 205 of a first binding pair. In some embodiments, biotin is the binding partner 205 of the first pair.

[0038] In some embodiments, a first tag is attached to an input nucleic acid molecule without an adaptor. For instance, a nucleic acid can be directly tagged using biotin- or digoxigenin- 3’ end oligonucleotide labeling kits which might improve tagging efficiency over traditional adaptor ligation. Such labeling kits are commercially available. However, such an approach potentially limits inclusion of identifiers (sample index / UMI) to support pairing of sequences obtained from original DNA strands.

[0039] The input nucleic acid construct 202 is denatured to form a single-stranded input nucleic acid construct 202’, which is then amplified using first and second amplification primers 225, 226 to form a double-stranded product 228 of the input nucleic acid 201 hybridized to an amplicon 227. The double-stranded product 228 can be denatured, and the single-stranded input nucleic acid construct 202’, which still contains its base modifications 224, can be separated from amplicons 227 by attaching a second tag 216 comprises an anti -DIG antibody 217 conjugated to biotin 218 to the binding partner 205 (DIG) of the first tag 203. Amplicon 227 can be subjected to further amplification to form additional amplicons 229, before and / or after separation of the single-stranded input nucleic acid construct 202’ or the double-stranded product 228 from amplicons 227. The input nucleic acid construct 202’ can be further processed after amplification, such as for methylation analysis, while the amplicons 227 can be analyzed bysequencing or other techniques, such as for mutations in the target sequence, without being subjected to base conversion.Target Enrichment

[0040] In some embodiments, any of the present methods can further comprise target enrichment or capture. The capture and enrichment can be done by target probe hybridization. In some embodiments, the present methods comprise forming a complex between reciprocal binding partners of a second binding pair, such as a biotinylated target nucleic acid and solid- supported avidin or streptavidin. The target probe can include a capture probe or a bridge probe, and the bridge probe can also bind to an anchor probe. The target probe can comprise one or more binding moieties. The binding moiety can be a biotin. The binding moieties can be attached to a support. The support can be a bead. The bead can be a streptavidin coated bead.

[0041] FIG. 3 illustrates how a tagged nucleic acid molecule 311 can be further processed in a target enrichment procedure. The tagged nucleic acid molecule 311 can be a tagged primer extension product 105 as shown in FIG. 1 or a tagged input nucleic acid construct 202’ as shown in FIG. 2. In FIG. 3, tagged nucleic acid molecule 311 comprises an adaptor 307 and a DIG tag 309. The tagged nucleic acid molecule 311 is hybridized with one or more bridge probes 313, 315. Bridge probes 313, 315 are hybridized with universal probe 317 (also referred to as an anchor probe) conjugated with a binding partner 319 (for example, biotin) of a second binding pair (for example, the binding pair of biotin : streptavidin) wherein streptavidin 321 is the reciprocal binding partner. In the illustrated embodiment, the hybridizations proceed at the same time; in other embodiments, the hybridizations can be sequential. The binding partners of the first binding pair (DIG: anti -DIG antibody) do not bind with the binding partners of the second binding pair (biotin: streptavidin). In FIG. 3, target enrichment of the tagged nucleic acidmolecule 311 is performed by binding the binding partner 319 (biotin) to its reciprocal binding partner 321 (streptavidin) which is on a solid support 323 (e.g., a magnetic bead). The target probe can also capture untagged nucleic acid molecules in the same sample, and the first tag can be used for separation after forming a target enriched sample.

[0042] Although the present methods can employ a capture probe that directly hybridizes with a target sequence of a target nucleic acid, the use of bridge probes and anchor probes offer several advantages. Bridge probes can be used to hybridize a nucleic acid molecule and can further allow indirect association between an anchor probe and the nucleic acid molecule. The bridge probe can comprise a target specific region (TSR) that hybridizes to a target sequence. The bridge probe can also comprise an anchor-probe-landing sequence (ALS) that hybridizes to a bridgebinding-sequence of anchor probe. The bridge probe can comprise a linker connecting TSR and ALS. The TSR can be located in the 3 ’-portion of the bridge probe. The TSR can be located in the 5 ’-portion of the bridge probe. Additional description of the use of capture probes, bridge probes, and target probes is set forth below.Preparing Nucleic Acid Molecules for Methylation Analysis

[0043] As yet another aspect, methods are provided for preparing nucleic acid molecules for methylation analysis. The methods comprise attaching a methylated adaptor to an input nucleic acid and hybridizing a primer to the input nucleic acid, wherein the primer is not methylated. The methods also comprise extending the primer to produce a primer extension product comprising a first segment complementary to at least a portion the input nucleic acid and a second segment complementary to at least a portion of the methylated adaptor, wherein the second segment is not methylated. The mixture of input nucleic acid and the primer extension product is contacted with one or more base conversion reagents that either (a) convert cytosine touracil without conversion of 5 -methyl cytosine, or (b) converts 5-methylcytosine (5meC) to thymine without conversion of cytosine, thereby forming converted input nucleic acid and / or converted primer extension product. The converted input nucleic acid and / or the converted primer extension product are analyzed to determine methylation status of the input nucleic acid.

[0044] FIG. 4 illustrates an embodiment of the present methods of preparing an input nucleic acid for methylation analysis. More particularly, FIG. 4 shows an input nucleic acid construct 402 in which an input nucleic acid 401 is ligated at its ends to methylated adaptors 404. Input nucleic acid 401 includes base modifications 424 such as methylation of cytosine bases. The methylated adaptors 404 comprise barcodes 430 (e.g., sample identifiers and / or unique molecular identifiers) which include one or more 5-methylcytosines 432.

[0045] The input nucleic acid construct 402 is denatured to form a single-stranded input nucleic acid construct 402’, to which first and second primers 425, 426 are hybridized. Through primer extension reactions, first and second primers 425, 426 are extended to form primer extension products. In this embodiment, the PE reactions are performed with a PE reagent comprises 5- methyldeoxycytosine triphosphate (5mdCTP), so that the PE product will contain 5-meC rather than cytosine bases. In some embodiments, the primer extension products and input nucleic acid constructs 502’ can be subjected to target enrichment 540 to produce enriched input constructs 544 and enriched PE products 546. Enrichment is optional and may be performed if targeted sequencing is desired rather than whole genome sequencing. The method also comprises a base conversion step 542 using bisulfite or other reagent that converts cytosine bases to uracil including bases 543 on the barcode. The optionally-enriched input constructs 544 and optionally- enriched PE product 546 can be subjected to further amplification 545 to form a sequencing library 516. After sequencing run, the sequencing data from the library can be analyzed using ananalysis algorithm 418 that separates the sequencing reads into genomic and methylation (e.g., epigenomic) categories.

[0046] In FIG. 4, base conversion 442 converts 5meC bases to thymine, including 5meC bases 443 on the barcode. The base conversion step 442 can comprise contacting the nucleic acid molecules with one or more base conversion reagents comprising a ten eleven translocation (TET) enzyme that converts 5meC to 5-carboxylcytosine (5caC) and / or 5-formylcytosine (5fC), and a borane reducing agent that converts 5caC and / or 5fC to thymine.

[0047] FIG. 5 illustrates another embodiment of the present methods of preparing an input nucleic acid for methylation analysis. More particularly, FIG. 5 shows an input nucleic acid construct 502 in which an input nucleic acid 501 is ligated at its ends to methylated adaptors 504. Input nucleic acid 501 includes base modifications 524 such as methylation of cytosine bases as well as unmodified cytosines 533. The methylated adaptors 504 comprise unmethylated barcodes 530 (e.g., sample identifiers and / or unique molecular identifiers) which include one or more cytosines 532.

[0048] The input nucleic acid construct 502 is denatured to form a single-stranded input nucleic acid construct 502’, to which first and second primers 525, 526 are hybridized. Through primer extension reactions, first and second primers 525, 526 are extended to form primer extension products. In this embodiment, the PE reactions are performed with a PE reagent comprising 5-methyldeoxy cytosine triphosphate (5mdCTP), so that the PE product will contain 5meC rather than cytosine basis. In some embodiments, the primer extension products and input nucleic acid constructs 502’ can be subjected to target enrichment 540 to produce enriched input constructs 544 and enriched PE products 546. Enrichment is optional and may be performed if targeted sequencing is desired rather than whole genome sequencing. The method alsocomprises a base conversion step 542 using bisulfite or other reagent that converts cytosine bases to uracil including bases 543 on the barcode. The optionally-enriched input constructs 544 and optionally-enriched PE products 546 can be subjected to further amplification 545 to form a sequencing library 515. After sequencing run, the sequencing data from the library can be analyzed using an analysis algorithm 518 that separates the sequencing reads into genomic and methylation (e.g., epigenomic) categories.

[0049] In FIG. 5, base conversion 542 converts cytosine bases to uracil, which will be thymine in amplicons. The base conversion step 542 can comprise contacting the nucleic acid molecules with one or more base conversion reagents comprises bisulfite or an enzyme that converts cytosine to uracil.Probes for Target Enrichment

[0050] As mentioned above, the present methods can include target enrichment which may employ a capture probe, a bridge probe, and an anchor probe. The bridge probe can comprise DNA. The bridge probe can comprise of RNA. The bridge probe can comprise uracil and methylated cytosine. IN some embodiments, the bridge probe does not comprise uracil. The bridge probe can comprise about 400 nucleotides, about 300 nucleotides, about 200 nucleotides, about 120 nucleotides, about 100 nucleotides, about 90 nucleotides, about 80, about 70 nucleotides, about 50 nucleotides, about 40 nucleotides, about 30 nucleotides, about 20 nucleotides, or about 10 nucleotides. The bridge probe can comprise one or more molecular barcodes. The bridge probe can comprise one or more binding moieties. The binding moiety can be a biotin. The binding moieties can be attached to a support. The support can be a bead. The bead can be a streptavidin bead.

[0051] Multiple bridge probes can be used to anneal to multiple target sequences in a sample. The bridge probes can be designed to have similar melting temperatures. The melting temperatures for a set of bridge probes can be within about 15°C, within about 10°C, within about 5°C, or within about 2°C. The melting temperature for one or more bridge probes can be about 75°C, about 70°C, about 65°C, about 60°C, about 55°C, about 50°C, about 45°C, or about 40°C. The melting temperature for the bridge probe can be about 40°C to about 75°C, about 45°C to about 70°C, 45°C to about 60°C, or about 52°C to about 58°C.

[0052] Use of an anchor probe along with one or more bridge probe around a particular bridge probe can help to stabilize the hybridization of the particular bridge probe to its target sequence through synergistic effect. A hybridization temperature to form the multiple bridge probe assembly can be higher than the melting temperature of a single bridge probe. The higher temperature can result in a better capture specificity by reducing nonspecific hybridization that can occur at lower temperature. The hybridization temperature can be about 5°C, about 10°C, about 15°C, or about 20°C higher than the melting temperature of individual bridge probe. The hybridization temperature can be about 5°C to about 20°C higher than the melting temperature of a bridge probe, or about 5°C to about 20°C higher than an average melting temperature of a plurality of bridge probes.

[0053] The hybridization temperature for multiple bridge probes can be about 75°C, about 70°C, about 65°C, about 60°C, about 55°C, or about 50°C. The hybridization temperature for multiple bridge probes can be about 50°C to about 75°C, 55°C to about 75°C, 60°C to about 75°C, or 65°C to about 75°C.

[0054] The bridge probe can further comprise a label. The label can be fluorescent. The fluorescent label can be organic fluorescent dye, metal chelate, carbon nanotube, quantum dot,gold particle, or fluorescent mineral. The label can be radioactive. The label can be biotin. The bridge probe can bind to labeled nucleic acid binder molecule. The nucleic acid binder molecule can be antibody, antibiotic, histone, antibody, or nuclease.

[0055] The bridge probe can comprise a linker. In some embodiments, the linker comprises about 30 nucleotides, about 25 nucleotides, about 20 nucleotides, about 15 nucleotides, about 10 nucleotides, or about 5 nucleotides; any of those numbers can be combined to form a range for the number of nucleotides in a linker. The linker can comprise non-nucleic acid polymers (e.g., string of carbons). The linker non-nucleotide polymer can comprise about 30 units, about 25 units, about 20 units, about 15 units, about 10 units, or about 5 units; any of those numbers can be combined to form a range for the number of units in a linker.

[0056] The bridge probe can be blocked at the 3’ and / or 5’ end. The bridge probe can lack a 5’ phosphate. The bridge probe can lack a 3’ OH. The bridge probe can comprise a 3’ddC, 3’inverted dT, 3’C3 spacer, 3’ amino, or 3’ phosphorylation.

[0057] The anchor probe or universal anchor probe can comprise one or more bridge-binding- sequences (BBS) that hybridize to anchor-probe-landing sequence of the one or more bridge probes.

[0058] The anchor probe can comprise spacers in between the BBSs. The presence of the one or more spacers can improve the efficiency of the hybridization capture and increase the specificity of the capture.

[0059] The anchor probe can comprise a molecular barcode (MB). The anchor probe can comprise a BBS to which the one or more bridge probes can hybridize to. The anchor probe can comprise from 1 to 100 BBSs. The anchor probe can comprise an index for distinguishing samples. The molecular barcode or index can be 5’ of the adaptor sequence and 5’ of the BBS.

[0060] The anchor probe can comprise about 400 nucleotides, about 200 nucleotides, about 120 nucleotides, about 100 nucleotides, about 90 nucleotides, about 80 nucleotides, about 70 nucleotides, about 50 nucleotides, about 40 nucleotides, about 30 nucleotides, about 20 nucleotides, or about 10 nucleotides. The anchor probe can be about 20 to about 70 nucleotides.

[0061] The melting temperature of anchor probe to the bridge probe can be about 65°C, about 60°C, about 55°C, about 50°C, about 45°C, or about 45°C to about 70°C.

[0062] The anchor probe can comprise a label. The label can be fluorescent. The fluorescent label can be an organic fluorescent dye, metal chelate, carbon nanotube, quantum dot, gold particle, or fluorescent mineral. The label can be radioactive. The label can be biotin. The anchor probe can bind to labeled nucleic acid binder molecule. The nucleic acid binder molecule can be antibody, antibiotic, histone, antibody, or nuclease.

[0063] FIG. 3 demonstrates how the present methods can be used for target enrichment on a sample containing a tagged nucleic acid molecule 311 (such as a tagged PE product or a tagged input nucleic acid) tagged with a first tag. Preparation of nucleic acids for sequencing-by- synthesis or other analysis often employs target enrichment, and one or more target enrichment procedures can be included in any of the present methods. By enriching for one or more desired targets, sequencing or other analysis can be more focused with reduced effort and expense and / or with high coverage depth. Examples of target enrichment procedures include hybridization-based capture protocols such as SureSelect Hybrid Capture from Agilent and TruSeq Capture from Illumina. Other examples include PCR-based protocols such as HaloPlex from Agilent;AmpliSeq from ThermoFisher; TruSeq Amplicon from Illumina; and emulsion / digital PCR from Raindance.

[0064] Tn some embodiments, the present methods also comprise capture of tagged or untagged input nucleic acid molecules and / or tagged or untagged PE products and / or amplicons thereof comprising target sequences. The present methods allow efficient capture and enrichment of tagged nucleic acid molecules having a first tag attached thereto. Target enrichment can be performed after library construction by extension of a tagged primer comprising of a first tag and an adaptor to synthesize tagged PE products. Target enrichment can also be performed on input nucleic acids after attaching an adaptor (such as a methylated adaptor, or a first tag comprising the adaptor) to a 3' end of the input nucleic acid molecule. The present methods can be used to handle low input samples such cell-free DNA (cfDNA), and therefore can be suitable for methylation sequencing analysis of input nucleic acids.

[0065] In some embodiments, the present methods comprise target enrichment by indirect hybridization of the input nucleic acid and / or PE products with an anchor probe through hybridization of one or more bridge probes to the input nucleic acid and / or PE product. The one or more bridge probes can be designed to hybridize to particular target sequences in the nucleic acid molecules. An anchor probe in turn can be designed to hybridize to the one or more bridge probes, thereby creating an assembly of three or more hybridized nucleic acid molecules. The multi-structure hybridization assembly can act synergistic to provide more stability to the assembly. The hybridized input nucleic acid molecule and / or PE product can be subsequently treated with bisulfite or a base conversion enzyme for methylation sequencing.

[0066] Enrichment of a nucleic acid molecule containing a target sequence can be facilitated by interaction of the nucleic acid molecule and two or more probes that form a hybridization assembly. The multi-complex assembly can stabilize the hybridization interaction between the nucleic acid molecule and the enrichment such as bridge probes. A bridge probe can comprise atarget specific region that hybridizes to a target region of the input nucleic acid and anchorprobe-landing sequence (ALS) that hybridizes to bridge-binding-sequence (BBS) of an anchor probe. The hybridizations between the nucleic acid molecule and the bridge probe and between the bridge probe and the anchor probe can form a multi-complex assembly.

[0067] In some embodiments, the present methods comprise hybridizing a first target specific region of a first bridge probe to a first target sequence of a nucleic acid molecule with a sequence corresponding to the genome region, wherein a first anchor-probe-landing sequence of the first bridge probe is bound to a first bridge-binding-sequence of an anchor probe; and hybridizing a second target specific region of a second bridge probe to a second target sequence of the molecule with a sequence corresponding to the genome region, wherein a second anchor-probelanding sequence of the second bridge probe is bound to a second bridge-binding-sequence of the anchor probe. As described herein the anchor probe may comprise a binding moiety. The method generally comprises attaching adaptors to the 5’ end or the 3’ ends of nucleic acid molecules of the plurality of nucleic acid molecules, thereby generating a library of nucleic acid molecules comprising adaptors.

[0068] More than two bridge probes per input nucleic acid molecule can be used in the methods disclosed herein. For example, at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 25, 50, 75, 100, or more bridge probes can be used to bridge the input nucleic acid and the anchor probe. The target enrichment can further comprise hybridizing a second target specific region of a second bridge probe to a second target sequence of the input nucleic acid molecule, wherein a second anchor-probelanding sequence of the second bridge probe can be bound to a second bridge-binding-sequence of the anchor probe. In some cases, the target enrichment can be conducted after attachment of adaptors or other first tags to the input nucleic acid molecules.

[0069] The bridge probes can further comprise linkers that connect the target specific region and the anchor-probe-landing sequence. The adaptor anchor can comprise one or more spacers in between the bridge-binding-sequences. The presence of the one or more spacers can improve the efficiency of the hybridization capture and increase the specificity of the capture.

[0070] The nucleic acid molecules can be captured and enriched from low-input samples such as cell-free DNA (cfDNA) and circulating tumor DNA (ctDNA). The capture and enrichment can be done by the indirect association with anchor probe through hybridization with bridge probe. The bridge probe and / or anchor probe can comprise one or more binding moieties. The binding moiety can be a biotin. The binding moieties can be attached to a support. The support can be a bead. The bead can be a streptavidin bead.

[0071] The present methods of capture and enrichment can further include solid phase extraction of the nucleic acid molecules. The bridge probe or anchor probe can be bound to a solid support. The bridge probe, or anchor probe can comprise a label. The disclosed methods can further comprise capturing to the bridge probe, the anchor probe, or the hybridization complex comprising target nucleic acid molecule, bridge probe, and anchor probe by the label. The label can be biotin. The label can be a nucleic acid sequence, such as poly A or Poly T, or specific sequence. The nucleic acid sequence can be about 5 to 30 bases in length. The nucleic acid sequence can comprise DNA and / or RNA. The label can be at the 3’ end of the bridge probe, or anchor probe. The label can be a peptide, or modified nucleic acid that can be recognized by antibody such as 5-Bromouridine, and biotin. The label can be conjugated to the bridge probe, or anchor probe by reactions such as “click” chemistry. “Click” chemistry can allow for the conjugation of a reporter molecule like fluorescent dye to a biomolecule like DNA.Click Chemistry can be a reaction between and azide and alkyne that can yield a covalent product (e.g., 1,5-disubstituted 1,2, 3 -triazole). Copper can serve as a catalyst.

[0072] The label can be captured on a solid support. The solid support can be magnetic. The solid support can comprise a bead, flowcell, glass, plate, device comprising one or more microfluidic channels, or a column. The solid support can be a magnetic bead.

[0073] The solid support (e.g., bead) can comprise (e.g., by coated with) one or more capture moi eties that can bind the label. The capture moiety can be streptavidin, and the streptavidin can bind biotin. The capture moiety can be an antibody. The antibody can bind the label. The capture moiety can be a nucleic acid, e.g., a nucleic acid comprising DNA and / or RNA. The nucleic acid capture moiety can bind a sequence on, e.g., an anchor probe or bridge probe. In some cases, an anti-RNA / DNA hybrid antibody bound to a solid surface can be used as a capture moiety.

[0074] The label and the capture moiety can bind through one or more covalent or non-covalent bonds. Following capture of the bridge probe, anchor probe, or the hybridization complex on the solid support, the solid support can be washed to remove, e.g., unbound template from the sample. In some cases, no wash step is performed. The wash can be stringent or gentle. The capture probe or anchor probe that are hybridized to an input nucleic acid molecule can be eluted, e.g., by adding free biotin to the sample when the label is biotin and the capture moiety is streptavidin.

[0075] Cleanup can be performed using streptavidin beads after the input nucleic acid, bridge probe, and anchor probe hybridization, wherein the 3’ end of the anchor probe is biotinylated. The input nucleic acid complex hybridized to the bridge probes (and indirectly with the anchor probe) is bound to the bead. The input nucleic acid that has not hybridized to the bridge probecan be washed away. The 5’ end or the 3’ end of a first and or second bridge probe can be biotinylated. In this manner, streptavidin beads can be used to remove and separate the unhybridized input nucleic acid from input nucleic acid having the target sequence.Amplification

[0076] In some embodiments, the present method comprises amplifying a nucleic acid, before and / or after it is tagged with a first tag and / or a second tag. Nucleic acids are amplified by hybridizing primers to a primer site and using polymerase to add nucleotides to a synthesized strand using the nucleic acid as a template.

[0077] In some embodiments, an adaptor is located at a 5'-end of a target sequence in an input nucleic acid, and the adaptor provides a priming site for amplification of the target sequence. A nucleic acid can be amplified using a first amplification primer and a second amplification primer. In some embodiments, the first amplification primer has sequence specificity for a target sequence in the nucleic acid, and is capable of hybridizing to a portion of the target sequence (a nucleic acid of interest). The second amplification primer is capable of hybridizing to a priming site of the adaptor or to a target-specific priming site of the input nucleic acid. During the amplification step, the first amplification primer hybridizes to the target sequence and the second primer hybridizes to the sequence priming site on the adaptor. In some embodiments, the first amplification primer hybridizes at the 5'-end of the nucleic acid construct. The primers should be sufficiently large to provide adequate hybridization with the target sequence or other primer binding site.

[0078] Nucleic acid molecules may be amplified using any suitable method. In some embodiments, the input nucleic acid is amplified using polymerase chain reaction (PCR). Ingeneral, PCR comprises denaturation of polynucleotide strands (e.g., DNA melting), annealing of primers to the denatured polynucleotide strand, and extension of primers with a polymerase to synthesize the complementary polynucleotide. The process generally requires a DNA polymerase, forward and reverse primers, deoxynucleoside triphosphates, bivalent cations, and a buffer solution. In some embodiments, the input nucleic acid is amplified by linear amplification. In some embodiments, the nucleic acid molecule is amplified using Emulsion PCR, Bridge-PCR, or Rolling Circle amplification. The amplicons of the input nucleic acid molecule may be analyzed to determine the order of base pairs using a suitable sequencing method.Methylated Adaptors

[0079] In some embodiments, the present method comprises attaching one or more adaptors to an input nucleic acid, to form a nucleic acid construct. An adaptor can be attached to an input nucleic acid before or after amplification, and in some embodiments the adaptor is attached before amplification. The adaptor can be attached by any suitable technique, such as by ligation, use of a transposase, hybridization, and / or primer extension. In some embodiments, the input nucleic acid is ligated with an adaptor at one or both ends. In a ligation reaction, a covalent bond or linkage is formed between the termini of two or more nucleic acid molecules (such as an input and an adaptor). The nature of the bond or linkage may vary, and the ligation may be carried out enzymatically or chemically. Ligations are usually carried out enzymatically to form a phosphodiester linkage between a 5' carbon of a terminal nucleotide of one polynucleotide or oligonucleotide with 3' carbon of another polynucleotide or oligonucleotide. In someembodiments, the adaptor is a Y-adaptor. Other examples of adaptors including linear adaptors, circular adaptors, and bubble adaptors.

[0080] In some embodiments, the adaptor is a methylated adaptor. Methylated adaptors are adaptors comprising one or more methylcytosine bases. In some embodiments, the methylated adaptor includes at least 2, 3, 4, 5, 6, 7, 8, 9, 10, or more methylated cytosines. In some embodiments, all cysteines or substantially all cysteines in a methylated adaptor are methylated. In some embodiments, a methylated adaptor comprises an unmethylated barcode.

[0081] In some embodiments, the adaptor is an unmethylated adaptor, in that none or substantially none of the cysteines in the adaptor are methylated.

[0082] In some embodiments, a desired adaptor (e.g., a methylated adaptor or an unmethylated adaptor) is attached to an input nucleic acid. Generally an adaptor is attached to at least one strand of a double-stranded DNA molecule, and usually an adaptor can be a molecule that is at least partially double- stranded. An adaptor may be 40 to 150 bases in length, e.g., 50 to 120 bases. An adaptor can be joined to a 5' end and / or a 3' end of a nucleic acid molecule. A Y- adaptor is an adaptor that contains a double-stranded region and a single-stranded region in which the opposing sequences are not complementary. The end of the double-stranded region may be or can be joined to target molecules such as double-stranded fragments of genomic DNA, e.g., by via a transposase-catalyzed reaction. Each strand of a double-stranded DNA molecule that has been joined to a Y adaptor is asymmetrically tagged in that it has the sequence of one strand of the Y-adaptor at one end and the other strand of the Y-adaptor at the other end. Amplification of nucleic acid molecules that have been joined to Y-adaptors at both ends results in an asymmetrically tagged nucleic acid, i.e., a nucleic acid that has a 5' end containing one tag sequence and a 3' end that has another tag sequence.Input Nucleic Acid

[0083] The present methods can be used with samples comprising input nucleic acid molecules of various types, particularly where the sample comprises a mixture of input DNA and RNA molecules. Input DNA molecules include genomic DNA (gDNA), mitochondrial DNA, viral DNA, cDNA, cell-free DNA (cfDNA), circulating tumor DNA (ctDNA), cell-free fetal DNA (cffDNA), or synthetic DNA. The DNA can be double-stranded DNA, single-stranded DNA, fragmented DNA, or damaged DNA. The input nucleic molecules can include input RNA, input DNA, or a mixture of input RNA and input DNA. The input RNA molecules can be mRNA, pre-mRNA, tRNA, rRNA, microRNA, snRNA, piRNA, small non-coding RNA, polysomal RNA, intron RNA, pre-mRNA, viral RNA, or cell-free RNA. In some embodiments, the DNA comprises fragmented genomic DNA and the RNA comprises mRNA or pre-mRNA.

[0084] The input nucleic acid can be naturally occurring or synthetic. The input nucleic acid can have modified heterocyclic bases. The modification can be methylated purines or pyrimidines, acylated purines or pyrimidines, alkylated riboses, or other heterocycles. The input nucleic acid can have modified sugar moieties. The modified sugar moieties can include peptide nucleic acid. The input nucleic acid can comprise peptide nucleic acid. The input nucleic acid can comprise threose nucleic acid. The input nucleic acid can comprise locked nucleic acid. The input nucleic acid can comprise hexitol nucleic acid. The input nucleic acid can be flexible nucleic acid. The input nucleic acid can comprise glycerol nucleic acid.

[0085] Input nucleic acids can be prepared or treated in other ways that cooperate with the present methods. For instance, the input nucleic acids are prepared by a procedure that includes traditional end-repair, A-tailing, and ligation with Y-adaptors.

[0086] The input nucleic acid can be captured and enriched from low-input (e.g., 1 ng of nucleic acid materials) samples such as cell-free DNA (cfDNA), circulating tumor DNA (ctDNA), a single cell, or 10 or fewer cells. Examples of a single cell or other cells for which analysis may be desired include a neuron, a glial cell, a germ cell, a gamete, an embryonic stem cell, a pluripotent stem cell (including an induced pluripotent stem cell), an adult stem cell, a cell of the hematopoietic lineage, a differentiated somatic cell, a microbial cell, a cancer cell (including, for example a cancer stem cell), and a disease cell. In some embodiments, the input nucleic acid is captured from 10 or fewer cells (such as 1-10 cells, 2-10 cells, 5-10 cells, 1-2 cells, or 2-5 cells). The low-input samples can have 1 ng, 2 ng, 3 ng, 4 ng, 5 ng, 6 ng, 7 ng, 8 ng, 9 ng, 10 ng, or more of nucleic acid materials. The low-input samples can have less than 10 ng, 9 ng, 8 ng, 7 ng, 6 ng, 5 ng, 4 ng, 3 ng, 2 ng, 1 ng, or less of nucleic acid materials. The low-input samples can have from 200 pg to 10 ng of nucleic acid materials. The low-input samples can have less than 10 ng of nucleic acid materials. The low-input sample can less than 10 ng, 5 ng, 1 ng, 100 pg, 50 pg, 25 pg, or less of the nucleic acid materials. In some cases, the input samples can have 1 ng, 10 ng, 20 ng, 30 ng, 40 ng, 50 ng, or more of nucleic acid molecule. The input samples can have less than 50 ng, 40 ng, 30 ng, 20 ng, 10 ng, 1 ng, or less of nucleic acid materials.

[0087] The input nucleic acid can be damaged. The damaged nucleic acid can comprise altered or missing bases, and / or modified backbone. The input nucleic acid can be damaged by oxidation, radiation, or random mutation. The input nucleic acid can be damaged by bisulfite treatment.

[0088] Damaged dsDNA (with a nick) or ssDNA can be used as input nucleic acid for a library construction. For the damaged dsDNA, the dsDNA can be denatured so at least one undamagedstrand can be used as an input nucleic acid. The input nucleic acid can then be hybridized and attached to a capture probe and amplified using various primers.

[0089] The input nucleic acid can be derived from cell-free DNA (cfDNA) or circulating tumor DNA (ctDNA). The cfDNA can be fetal or tumor in source. The input nucleic acid can be derived from liquid biopsy, solid biopsy, or fixed tissue of a subject. The input nucleic acid can be cDNA and can be generated by reverse transcription. The input nucleic acid can be derived from fluid samples, including not limited to plasma, serum, sputum, saliva, urine, or sweat. The fluid samples can be bisulfite or enzyme treated to study the methylation pattern of the input nucleic acid and / or to determine the tissue origin of the input nucleic acid. The input nucleic acid can be derived from liver, esophagus, kidney, heart, lung, spleen, bladder, colon, or brain. The input nucleic acid can be treated with bisulfite to analyze methylation pattern of an organ from which the input nucleic acid is obtained. In some embodiments, the subject suffers from a methylation-related disease, such as autoimmune diseases, cardiovascular diseases, atherosclerosis, nervous disorders, and cancer.

[0090] The input nucleic acid can be derived from male or female subject. The subject can be an infant, a teenager, a young adult or an elderly person. The input nucleic acid can originate from human, rat, mouse, other animal, or specific plants, bacteria, algae, viruses, and the like. The input nucleic acid can originate from primates, such as chimpanzees or gorillas. Other animals include a rhesus macaque. The input nucleic acid can be from a mixture of genomes of different species including host-pathogen, bacterial populations, etc. In some embodiments, the input nucleic acid can be cDNA made from RNA expressed from genomes of one or more species.

[0091] The input nucleic acid can comprise a target sequence. The target sequence can be an exon, an intron, or a promoter. The target sequence can be previously known, partially known previously, or previously unknown. The target sequence can comprise a chromosome, chromosome arm, or a gene. The gene can be gene associated with a condition, e.g., cancer.Sequencing The Nucleic Acid Molecules

[0092] The present methods may be used as part of a high-throughput sequencing method such as a Next Generation Sequencing (NGS) method. In some embodiments, a high-throughput sequencing method comprises three steps: library preparation, immobilization, and sequencing. A nucleic acid sample generally is obtained, and adaptors are attached to one or both ends of the fragments or other nucleic acids to form a sequencing library. The sequencing library molecules are immobilized on a solid support, and sequencing reactions are performed to identify the nucleic acid sequence. The high-throughput sequencing method may employ Emulsion PCR, Bridge-PCR, or Rolling Circle amplification to provide colonies or copies of the nucleic acid molecules.

[0093] In some embodiments, the nucleic acid molecules are sequenced without amplification. The nucleic acid molecules can be sequenced by any suitable technique such as sequencing-by- synthesis or sequencing-by-hybridization. In some embodiments, the nucleic acid molecules are sequenced using a single-molecule sequencing platform, such as the methods discussed in Wbhrstein et al. US Patent 10,851,411 .

[0094] Emerging nucleic-acid sequencing platforms can enable direct detection and analysis of nucleic acid molecules without the need for traditional sample preparation and amplification steps. In one example of such emerging platforms, input nucleic acid moleculescan be tagged with a biotin molecule through a simple enzymatic addition step. The biotinylated nucleic acid molecule can then be bound to a flowcell containing avidin or streptavidin moieties. Subsequent analysis can then be performed, e.g., using fluorescently labelled probes.

[0095] If biotin is tagged to input nucleic acid molecules, as is done in some sample preparation procedures, it would interfere with certain probe-based capture and enrichment protocols that uses biotin-tagged oligonucleotide probes. However, addition of a biotin tag to nucleic acids after target enrichment is challenging due to the single stranded nature of the enriched nucleic acid molecules. In addition, residual components from the target enrichment protocol could also be processed, impacting subsequent analysis.

[0096] The present disclosure provides an approach for tagging input, or nucleic acid tagging PE products, that is compatible with existing target enrichment procedures and supports subsequent biotin-tag addition, enabling compatibility with direct nucleic acid sequence analysis platforms.

[0097] In some embodiments, the present methods comprise aligning sequence reads of the input nucleic acids and / or PE products and / or amplicons thereof. The sequence reads may be processed and grouped in any suitable way. In some embodiments, the sequence reads may be initially grouped by the fragment sequence and / or the identifier(s). In some implementations, initial processing of the sequence reads may include identification of molecular barcodes (including sample identifier sequences or sub-sample identifier sequences), and / or trimming reads to remove low quality or adaptor sequences. In addition, quality assessment metrics can be run to ensure that the dataset is of an acceptable quality. With sequencing platforms that require clonal amplification of input nucleic acid molecules, there is a concern a potential sequence variation is a PCR or amplification error rather than a true variation. An advantage fromidentifying and sequencing an input nucleic acid molecule without amplification is that it avoids such errors.

[0098] The amplified products generated using methods described herein can be further analyzed using various methods including southern blotting, polymerase chain reaction (PCR) (e g., real-time PCR (RT-PCR), digital PCR (dPCR), droplet digital PCR (ddPCR), quantitative PCR (Q-PCR), nCounter analysis (Nanostring technology), gel electrophoresis, DNA microarray, mass spectrometry (e.g., tandem mass spectrometry, matrix-assisted laser desorption ionization time of flight mass spectrometry (MALDI-TOF MS), chain termination sequencing (Sanger sequencing), or next generation sequencing. The input nucleic acids (or amplicons thereof) can also by analyzed by such methods.

[0099] The next generation sequencing can comprise 454 sequencing (ROCHE) (using pyrosequencing), sequencing using reversible terminator dyes (ILLUMINA sequencing), semiconductor sequencing (THERMOFISHER ION TORRENT), single molecule real time (SMRT) sequencing (PACIFIC BIOSCIENCES), nanopore sequencing (e g., using technology from OXFORD NANOPORE or GENIA), microdroplet single molecule sequencing using pyrophosphorolyis (BASE4), single molecule electronic detection sequencing, e.g., measuring tunnel current through nanoelectrodes as nucleic acid (DNA / RNA) passes through nanogaps and calculating the current difference (QUANTUM SEQUENCING from QUANTUM BIOSYSTEMS), GenapSys Gene Electronic Nano-Integrated Ultra-Sensitive (GENIUS) technology (GENAPYS), GENEREADER from QIAGEN, sequencing using sequential hybridization and ligation of partially random oligonucleotides with a central determined base (or pair of bases) identified by a specific fluorophore (SOLiD sequencing). The sequencing can be paired-end sequencing.

[0100] The performance of a panel or method for capturing targets or preparing a NGS library may be defined by a number of different metrics describing efficiency, accuracy, and precision. Such metrics can be obtained by sequencing the captured nucleic acid molecules or amplicons thereof. For example, coverage percentage region- wide (0.2X or 0.5X), coverage percentage base-wide, target coverage, depth of coverage, fold enrichment, percent mapped, percent on-target, AT or GC dropout rate, fold 80 base penalty, percent zero coverage targets, PF reads, percent selected bases, percent duplication, or other variables can be used to characterize a library.

[0101] The number of target sequences from a sample that can be sequenced using methods described herein can be about 5, 10, 15, 25, 50, 100, 1000, 10,000, 100,000, or 1,000,000, or about 5 to about 100, about 100 to about 1000, about 1000 to about 10,000, about 10,000 to about 100,000, or about 100,000 to about 1,000,000.

[0102] Nucleic acid libraries generated using methods described herein can be generated from more than one sample. Each library can have a different index associated with the sample. For example, a capture probe or an anchor probe can comprise an index that can be used to identify nucleic acids as coming from the same sample (e.g., a first set of capture probes or anchor probes comprising the same first index can be used to generate a first library from a first sample from a first subject, and a second set of capture probes or anchor probes comprising the same second index can be used to generate a second library from a second sample from a second subject, the first and second library can be pooled, sequenced, and an index can be used to discern from which sample a sequenced nucleic acid was derived). Amplified products generated using the methods described herein can be used to generate libraries from at least 2, 5, 10, 25, 50,100, 1000, or 10,000 samples, each library with a different index, and the libraries can be pooled and sequenced, e.g., using a next generation sequencing technology.

[0103] The sequencing can generate at least 100, 1000, 5000, 10,000, 100,000, 1,000,000, or 10,000,000 sequence reads. The sequencing can generate between about 100 sequence reads to about 1000 sequence reads, between about 1000 sequence reads to about 10,000 sequence reads, between about 10,000 sequence reads to about 100,000 sequence reads, between about 100,000 sequence reads and about 1,000,000 sequence reads, or between about 1,000,000 sequence reads and about 10,000,000 sequence reads.

[0104] The depth of sequencing can be about lx, 5x, lOx, 50x, lOOx, lOOOx, or 10,000x. The depth of sequencing can be between about lx and about lOx, between about lOx and about lOOx, between about lOOx and about lOOOx, or between about lOOOx and about lOOOOx.

[0105] The present disclosure provides methods in which separate fractions of a nucleic acid sample can be prepared for sequencing or treated with different procedures. The separated fractions can be combined for sequencing or other analysis. Because the present methods facilitate sequencing and analysis without amplification and enable the separation of different fractions, the separated input nucleic acids may be analyzed by sequencing, or may be bisulfite treated or enzymatically treated prior to sequencing to assess methylation. In some cases, a first fraction or subsample comprising amplicons or PE products may be analyzed by sequencing to assess mutations while a second fraction or subsample comprising input nucleic acid is bisulfide or enzymatically treated prior to sequencing to assess methylation. In some cases, a first fraction and a second fraction are both assessed by straightforward sequencing to access genomic alteration, however the samples may be sequenced at different depths. In some cases, an analysis of a fraction or subsample may be performed prior to performing a second target enrichmentstep. The results of the analysis of the initial fraction sample may be used to select a second panel for the second enrichment step.Treatment for Methylation Analysis

[0106] In some embodiments, the present methods further comprise treating the input nucleic acid (or a fraction or sub sample of the input nucleic acid molecules) or the PE products for methylation analysis. The methylation analysis can be done using one or more base conversion reagents, for example bisulfite treatment or TAPS. The base-converted nucleic acids can be used to study methylation of the nucleic acids. The bisulfite treatment can convert unmethylated cytosines to uracils. Methylation of a cytosine (e.g., 5’-methylctyosine) can prevent bisulfite from converting methylated cytosine to uracil. TAPS treatment can convert methylated cytosines to a uracil base, more particularly to dihydrouracil (DHU). Cytosine is not converted.

[0107] Base conversion reagents for DNA bisulfite modification are commercially available from, for example, MethylEasy™ (Human Genetic Signatures™) and CpGenome™ Modification Kit (Chemicon™). See also, WO04096825A1, which describes bisulfite modification methods and Olek et al. Nuc. Acids Res. 24:5064-6 (1994), which discloses methods of performing bisulfite treatment and subsequent amplification. Bisulfite treatment allows the methylation status of cytosines to be detected by a variety of methods.

[0108] The input nucleic acid molecules can be treated with bisulfite or other base conversion reagent either before or after target enrichment or hybridization capture using a capture probe or bridge probe / anchor probe. In some cases, the hybridized input nucleic acid molecules can be treated with bisulfite. Formation of double strand sequence (e.g., between a target sequent of an input nucleic acid and a target-specific region of a capture probe) can protect against conversionof cytosines in the hybridized region to uracils during bisulfite treatment. The double stranded sequence formed by the hybridization of the capture probe to the template or the bridge probe to the template and to an anchor probe can provide protection against bisulfite conversion of cytosines in the hybridized regions to uracils. Furthermore, since bisulfite treatment can convert non-methylated cytosine to uracil, the protection against conversion of cytosines to uracils at the target sequence can allow for the use of amplification primers designed to anneal to the non- bisulfite converted DNA. For the pre-bisulfite conversion capture, the probe can also be designed against the unconverted sequence. Probes and primers that anneal to unconverted cytosines can be more straightforward to design and provide better hybridization. In some cases, the enzymatic treatment can be performed for the methylation analysis. The enzyme can be methylationsensitive or methylation dependent enzymes. The enzymes can be restriction enzymes. The enzymes can be methylation-sensitive restriction endonucleases. In other cases, the methylation analysis can be done by using specific antibodies or proteins that specifically bind to methylation sites to enrich methylated nucleic acids.

[0109] In some embodiments of the present methods, the captured input nucleic acid constructs are treated with methylation-sensitive enzymes that convert cytosine to uracil. In another case, the methylated nucleic acids of the captured input nucleic acid molecules can be enriched by specifically binding to antibodies or proteins that target methylated CpG sites in the input nucleic acid molecule. The present methods can be compatible clinical samples with over a large range of nucleic material amount. In some embodiments, the present methods can be used sequence samples with nucleic acid molecules of less than 5 ng, less than 4 ng, less than 3 ng, less than 2 ng, or less than 1 ng.

[0110] The target specific sequence or target specific region (TSR) of a capture probe or a bridge probe can be designed based on the target sequence of the input nucleic acid molecule, and the target sequence of the input nucleic acid molecule can retain non-methylated cytosine after the bisulfite treatment.

[0111] In some embodiments, the bisulfite treatment or enzymatic treatment can be performed on input nucleic acids after amplification of the input nucleic acids, so that amplicons based on the original input nucleic acid sequence are obtained. An input nucleic acid can comprise one or more uracils after bisulfite treatment.

[0112] The present methods can include sequencing to determine methylation status (e.g., whole genome bisulfite sequencing (WGBS)). Methylation status can be useful as molecular markers for various diseases.

[0113] Methylation status includes information related to methylation of cysteine bases in a nucleic acid or a region thereof. Methylation status can comprise a methylation index of a CpG site, a methylation density of CpG sites in a region, a distribution of CpG sites over a contiguous region, a pattern or level of methylation for each individual CpG site within a region that contains more than one CpG site, and non-CpG methylation. Methylation status a substantial part of the genome can be considered equivalent to the methylome. Methylation generally refers to the addition of a methyl group to the 5' carbon of cytosine residues (i.e., 5-methylcytosines) among CpG dinucleotides. Methylation may occur in cytosines in other contexts, for example CHG and CHH, where H is adenine, cytosine or thymine. Cytosine methylation may also be in the form of 5-hydroxymethylcytosine.

[0114] Following base conversion of input nucleic acid and preparation for sequencing, the input nucleic acid is analyzed by sequencing that allows one to ascertain whether it containsmethylated cytosine bases. Aligners for sequencing data are available, and they output aligned reads along with methylation calls for each cytosine with sequence context information. For instance, sequencing reads can be aligned with Bismark (a bowtie2 wrapper for bisulfite sequencing alignment) or BS- Bolt (a bwa-mem wrapper for bisulfite sequencing alignment).Kits for Indirect Tagging of Nucleic Acids

[0115] As another aspect of the present invention, kits are provided which comprise first and second tagging reagents for making nucleic acid constructs as described herein. In some embodiments, a first tagging reagent comprises a tagged primer comprising first tag (according to any of the embodiments described herein) in a composition that comprises a solvent or other components. In some embodiments, a first tagging reagent comprises a first tag comprising an adaptor and a binding partner of a first binding pair. Likewise, a second tagging reagent comprises a second tag (according to any of the embodiments described herein) in a composition. The kits can comprise the first and second tagging reagents in one or more vessels, such as vials, tubes, etc.

[0116] In some embodiments, the present kits can comprise one or more tagged primers or first tags comprising functional sequences such as an adaptor configured to be attached to an end of the input nucleic acid molecule. The first tags also comprise a binding partner of a first binding pair. The adaptors can comprise one or more identifiers such as UMI sequences. The present kits can also comprise one or more second tags comprising a reciprocal binding partner of the first binding pair.

[0117] In some embodiments, the kits comprise one or more base conversion reagents(according to any of the embodiments described herein). In some embodiments, the adaptors inthe kit are methylated adaptors, which may have methylated barcodes or unmethylated barcodes. In some embodiments, the kits can comprise one or more polymerization regents for amplification or primer extension reactions. In some embodiments, the reagents include 5mdCTP.

[0118] In some embodiments, the present kits further comprise one or more bridge probes that comprises a target specific region which hybridizes to a target sequence of an input nucleic acid molecule; and an anchor probe that comprises a bridge-binding-sequence which hybridizes to an anchor-probe-landing sequence of the bridge probe. In some embodiments, the kit comprises two, three or more bridge probes.

[0119] In addition to above-mentioned components, the kits may further include instructions for using the components of the kit to practice the present methods, i.e., to prepare nucleic acids for sequencing. The instructions for practicing the present methods are generally recorded on a suitable recording medium. For example, the instructions may be printed on a substrate, such as paper or plastic, etc. As such, the instructions may be present in the kits as a package insert, in the labeling of the container of the kit or components thereof (i.e., associated with the packaging or subpackaging) etc. In other embodiments, the instructions are present as an electronic storage data file present on a suitable computer readable storage medium, e.g., CD- ROM, portable drive, or cloud-based storage, etc. In yet other embodiments, the actual instructions are not present in the kit, but means for obtaining the instructions from a remote source, e.g., via the internet, are provided. An example of this embodiment is a kit that includes a web address where the instructions can be viewed and / or from which the instructions can be downloaded. As with the instructions, this means for obtaining the instructions is recorded on a suitable substrate.TERMINOLOGY

[0120] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. Although any methods and materials similar or equivalent to those described herein can also be used in the practice or testing of the present teachings, some exemplary methods and materials are now described.

[0121] All patents and publications, including all sequences disclosed within such patents and publications, referred to herein are expressly incorporated by reference. The citation of any publication is for its disclosure prior to the filing date and should not be construed as an admission that the present claims are not entitled to antedate such publication by virtue of prior invention. Further, the dates of publication provided can be different from the actual publication dates which can need to be independently confirmed.

[0001] Numeric ranges are inclusive of the numbers defining the range. Unless otherwise indicated, nucleic acids are written left to right in 5' to 3' orientation; amino acid sequences are written left to right in amino to carboxy orientation, respectively.

[0002] The present technology may employ, unless otherwise indicated, techniques and descriptions of organic chemistry, polymer technology, molecular biology (including recombinant techniques), cell biology, biochemistry, and immunology, which are within the skill of the art. Such techniques include polymer array synthesis, hybridization, ligation, and detection of hybridization using a label.

[0003] As used herein, the singular forms “a”, “an”, and “the” include plural referents unless the context clearly dictates otherwise. For example, the term “a primer” refers to one or more primers, i.e., a single primer and multiple primers. A “plurality” contains at least 2 members. Incertain cases, a plurality may have at least 10, at least 100, at least 100, at least 10,000, at least 100,000, at least 106, at least 107, at least IO8or at least IO9or more members.

[0004] It is further noted that the claims can be drafted to exclude any optional element. As such, this statement is intended to serve as antecedent basis for use of such exclusive terminology as “solely,” “only” and the like in connection with the recitation of claim elements, or use of a “negative” limitation.

[0005] As used in the specification and appended claims, and in addition to their ordinary meanings, the terms "substantial" or "substantially" mean to within acceptable limits or degree to one having ordinary skill in the art. For example, "substantially inactive" means that one skilled in the art considers the level of activity to be negligible.

[0122] The term “target” as used herein refers to a nucleic acid of interest, or which is desired for sequencing and / or other analysis. One or more targets may be present within an input nucleic acid, or a construct or complex made from an input nucleic acid. A target may be singlestranded or double-stranded, and often is double-stranded DNA when attached to an adaptor to form a nucleic acid construct. Target as used herein can refer to a specific sequence or the complement thereof or to both. The term target encompasses any nucleic acid molecule of biological or synthetic origin whose sequence or other characteristic is of interest. The target sequence does not include identifiers, primer binding regions, or adaptors sequences which may be added to the input nucleic acid molecule to prepare an input nucleic acid construct for sequencing or other analysis. A target may be within a nucleic acid in vitro or in vivo within the genome of a cell, or within the cytoplasm of a cell (such as RNA), or with a biological fluid (such as blood, plasma, amniotic fluid, or other biological sample).

[0123] The term “input” refers to a nucleic acid molecule to be processed in accordance with the present methods. For example, an input nucleic acid molecule may be present in a nucleic acid sample. The input may include one or more target sequences of interest, or it may include other sequences from which a target is desired to be separated. In some embodiments, an input nucleic acid comprises one or more sequences complementary to sequences of one or more capture probes, bridge probes, or other types of probes.

[0124] The terms “amplifying” and “amplification” as used herein refer to synthesizing nucleic acid molecules that are complementary to one or both strands of an input nucleic acid. Amplifying a nucleic acid molecule may include denaturing a double-stranded input nucleic acid, annealing primers the input nucleic acid at a temperature that is below the melting temperatures of the primers, and enzymatically elongating from the primers to generate an amplification product. The terms “amplicon” or “amplification product” refer to the nucleic acid sequences which are produced from an amplifying process, including the nucleic acid molecules synthesized by amplifying the input nucleic acid or its complementary sequence, as well as the nucleic acid molecules synthesized from other amplicons. The denaturing, annealing and elongating steps each can be performed one or more times. Amplification generally does not change the target or input nucleic acid sequence unless errors arise during the amplification.

[0125] Amplification typically requires the presence of deoxyribonucleoside triphosphates, a DNA polymerase enzyme and an appropriate buffer and / or co-factors for optimal activity of the polymerase enzyme. Reverse transcription is a linear amplification reaction that employs a specialized DNA polymerase (reverse transcriptase) to copy RNA into cDNA (complementary DNA) using deoxyribonucleoside triphosphates.

[0126] The term “adaptor” generally refers to a nucleic acid molecule that is attached to an input nucleic acid molecule to add a desired structure or function. The term “tag” also generally refers to a moiety that can add a desired structure or function, though it is contemplated that a tag may be a nucleic acid molecule, a molecule other than a nucleic acid, or a combination thereof. For example, a “tag” as used herein can comprise an adaptor conjugated to a non-nucleic acid binding partner such as DIG. As another example, a “tag” as used herein can comprise an antibody conjugated to a biotin moiety. As another example, an adaptor can be attached to an input fragment or an amplicon thereof to add a binding site for a NGS platform. In some embodiments, an adaptor refers to molecules that are at least partially double-stranded. An adaptor or a tag may be any desired length, including but not limited to 40 to 150 bases in length, e.g., 50 to 120 bases, although adaptors and tags outside of this range are envisioned.

[0127] The terms “identifier” or “barcode” refers to a sequence of nucleotides used to identify the origin of a sequence. Identifiers may comprise sample indices or sample barcodes, where the same sequence is shared for all nucleic acids from a particular source, organism, or sample. Sample barcodes enable the mixing of nucleic acids from different samples in one sequencing run, as the different sample barcode sequences enable the correct assignment of sequencing reads to each sample. One, two, or more sample barcodes may be used. Identifiers also comprise molecular barcodes (MBCs) or unique molecular identifier (UMI) sequences, which function to identify copies of individual input nucleic acid molecules. UMIs may comprise random nucleotides, known nucleotides, or a mixture of random and known nucleotides. UMIs enable more accurate sequencing by allowing error correction of sequences and more accurate estimation of the original number of input nucleic acids. In some embodiments, a large number of UMIs is used (e.g., 100,000, 1 million, 1 billion, or morepossible sequences) such that each input nucleic acid has a unique molecular barcode. Molecular barcodes called degenerate base regions (DBR) are disclosed in US Patent 8,481,292 (Population Genetics Technologies Ltd.). The DBRs are random sequence tags that are attached to molecules that are present in the sample. DBRs and other molecular barcodes allow one to distinguish PCR errors during sample preparation from mutations and other variants that were present in the original input nucleic acid.

[0128] In other embodiments, a smaller number of molecular barcodes is used, and the beginning or ending positions (or both) of the sequence read are used together with the molecular barcode to identify copies arising from a unique input nucleic acid. Molecular barcodes may be combined with sample barcodes, on the same or different portions of the target nucleic acid. Molecular barcodes may be added to one end of a nucleic acid template (e.g., the 5’ end of the + strand, and the 3’ end of the - strand in a duplex), or to both ends of an input nucleic acid (e.g., to both the 5; and the 3’ ends of both the + and the - strands of the duplex).

[0129] The term “sample” as used herein relates to a material or mixture of materials containing one or more nucleic acids of interest. In some embodiments, the term refers to any plant, animal or viral material containing DNA, RNA, or other nucleic acid, such as, for example, tissue or fluid isolated from a patient (including without limitation plasma, serum, amniotic fluid, cerebrospinal fluid, lymph, tears, saliva and tissue sections), from preserved tissue (such as FFPE sections) or from in vitro cell culture constituents, as well as samples from the environment. Any sample containing nucleic acid, e.g., genomic DNA from tissue culture cells or from a sample of tissue, may be employed in the present technology.

[0130] The term “nucleic acid sample” as used herein denotes a sample containing nucleic acids. The nucleic acid samples may be complex in that they contain multiple differentmolecules that contain sequences. Nucleic acid samples from a mammal (e g., mouse or human) are types of complex samples. Complex samples may have more than 104, 105, 106or 107different nucleic acid molecules. Also, a complex sample may comprise only a few molecules, where the molecules collectively have more than 104, 105, 106or 107or more nucleotides. The term “complexity” generally refers the total number of different sequences in a population, such as in a population of fragments, adaptors, or adaptor-ligated fragments. For example, if a population has 4 different sequences, then that population has a complexity of 4. A population may have a complexity of at least 4, at least 8, at least 16, at least 100, at least 1,000, at least 10,000 or at least 100,000 or more, depending on the desired result.

[0131] The term “nucleotide” as used herein refers to a phosphate ester of a nucleoside, wherein the esterification site typically corresponds to the hydroxyl group attached to the C-5 position of the pentose sugar. In some cases nucleotides comprise nucleoside polyphosphates. However, the terms “added nucleotide,” “incorporated nucleotide,” “nucleotide added” and “nucleotide after incorporation” all refer to a nucleotide residue that is part of an oligonucleotide or polynucleotide chain.

[0132] The term “nucleotide” refers to naturally-occurring nucleotides including guanine, cytosine, adenine, thymine, uracil (G, C, A, T and U respectively), as well as modified pyrimidine and purine derivatives and other non-naturally occurring moieties that contain not only the known purine and pyrimidine bases, but also other heterocyclic bases that have been modified. Such modifications include methylated purines or pyrimidines, acylated purines or pyrimidines, alkylated riboses or other heterocycles. In addition, the term “nucleotide” includes those moieties that contain hapten or fluorescent labels and may contain not only conventional ribose and deoxyribose sugars, but other sugars as well. Modified nucleotides also includemodifications on the sugar moiety, e.g., wherein one or more of the hydroxyl groups are replaced with halogen atoms or aliphatic groups, are functionalized as ethers, amines, or the likes.

[0133] The term “nucleic acid” and “polynucleotide” are used interchangeably herein to describe a nucleotide-containing polymer of any length, e.g., greater than about 2 bases, greater than about 10 bases, greater than about 100 bases, greater than about 500 bases, greater than 1000 bases, up to about 10,000 or more bases composed of nucleotides, e.g., deoxyribonucleotides or ribonucleotides, and may be produced naturally, chemically, enzymatically or synthetically. The term includes polymers having PNA, LNA or UNA. DNA and RNA have a deoxyribose and ribose sugar backbone, respectively, whereas PNA's backbone is composed of repeating N-(2-aminoethyl)-glycine units linked by peptide bonds. In PNA various purine and pyrimidine bases are linked to the backbone by methylene carbonyl bonds. A locked nucleic acid (LNA), often referred to as inaccessible RNA, is a modified RNA nucleotide. The ribose moiety of an LNA nucleotide is modified with an extra bridge connecting the 2' oxygen and 4' carbon. The bridge "locks" the ribose in the 3'-endo (North) conformation, which is often found in the A-form duplexes. LNA nucleotides can be mixed with DNA or RNA residues in the oligonucleotide whenever desired. The term “unstructured nucleic acid”, or “UNA”, is a nucleic acid containing non-natural nucleotides that bind to each other with reduced stability. For example, an unstructured nucleic acid may contain a G’ residue and a C’ residue, where these residues correspond to non-naturally occurring forms, i.e., analogs, of G and C that base pair with each other with reduced stability, but retain an ability to base pair with naturally occurring C and G residues, respectively.

[0134] The terms “nucleoside”, “nucleotide”, “deoxynucleoside”, and “deoxynucleotide” are intended to include those moieties that contain not only the known purine and pyrimidinebases, but also other heterocyclic bases that have been modified. Such modifications include methylated purines or pyrimidines, acylated purines or pyrimidines, alkylated riboses or other heterocycles. In addition, the “nucleoside”, “nucleotide”, “deoxynucleoside”, and“deoxynucleotide” include those moieties that contain not only conventional ribose and deoxyribose sugars, but other sugars as well. Modified nucleosides, nucleotides, deoxynucleosides or deoxynucleotides also include modifications on the sugar moiety, e.g., wherein one or more of the hydroxyl groups are replaced with halogen atoms or aliphatic groups, or are functionalized as ethers, amines, or the like.

[0135] Natural nucleotides or nucleosides are defined herein as adenine (A), thymine (T), guanine (G), and cytosine (C). It is recognized that certain modifications of these nucleotides or nucleosides occur in nature. However, modifications of A, T, G, and C that occur in nature that affect hydrogen bonded base pairing are considered to be non-naturally occurring. For example, 2-aminoadenosine is found in nature, but is not a “naturally occurring” nucleotide or nucleoside as that term is used herein. Other non-limiting examples of modified nucleotides or nucleosides that occur in nature that do not affect base pairing and are considered to be naturally occurring are 5-methylcytosine, 3 -methyladenine, O(6)-methylguanine, and 8-oxoguanine, etc. Nucleotides include any nucleotide or nucleotide analog, whether naturally-occurring or synthetic. Exemplary nucleotides include phosphate esters of deoxyadenosine, deoxycytidine, deoxyguanosine, deoxythymidine, adenosine, cytidine, guanosine, and uridine. Other nucleotides include an adenine, cytosine, guanine, thymine base, a xanthine or hypoxanthine, -bromouracil, 2- aminopurine, deoxyinosine, or methylated cytosine, such as 5-methylcytosine, and N4- methoxydeoxycytosine. Also included are bases of polynucleotide mimetics, such as methylated nucleic acids, e.g., 2'-O-methRNA, peptide nucleic acids, modified peptide nucleic acids, lockednucleic acids and any other structural moiety that can act substantially like a nucleotide or base, for example, by exhibiting base-complementarity with one or more bases that occur in DNA or RNA and / or by being capable of base-complementary incorporation, and includes chainterminating analogs. A nucleotide corresponds to a specific nucleotide species if they share basecomplementarity with respect to at least one base.

[0136] In addition to purines and pyrimidines, modified nucleotides or analogs, as those terms are used herein, include any compound that can form a hydrogen bond with one or more naturally occurring nucleotides or with another nucleotide analog. Any compound that forms at least two hydrogen bonds with T or with a derivative of T is considered to be an analog of A or a modified A. Similarly, any compound that forms at least two hydrogen bonds with A or with a derivative of A is considered to be an analog of T or a modified T. Similarly, any compound that forms at least two hydrogen bonds with G or with a derivative of G is considered to be an analog of C or a modified C. Similarly, any compound that forms at least two hydrogen bonds with C or with a derivative of C is considered to be an analog of G or a modified G. It is recognized that under this scheme, some compounds will be considered for example to be both A analogs and G analogs (purine analogs) or both T analogs and C analogs (pyrimidine analogs).

[0137] As used herein, the term “nucleic acid construct” refers to a nucleic acid that is ligated or otherwise attached to another nucleic acid, such as an adaptor. For example, a nucleic acid construct may contain a nucleic acid molecule to be sequenced, a capture site for flowcell attachment, one or more identifier sequences such as SBC and UMI, and primer binding sites for a first and second primer.

[0138] As used herein, the term “capture site” refers to a nucleic acid sequence configured for attachment of a nucleic acid construct to a flowcell or other surface, for NGS sequencing or other analysis processing.

[0139] As used herein, the term “identifier” refers to a nucleic acid sequence that can be used to identify a particular nucleic acid construct. An “identifier” may be a “sample barcode” or “SBC” sequence for identifying a particular biological sample. An “identifier” may also refer to a “molecular barcode” for identification of unique molecules present in the sample. Also, an “identifier” may contain both an SBC and an UMI.

[0140] The term “antibody” is well understood by those in the field and is used interchangeably herein with “immunoglobulin” Those terms refer to a protein consisting of one or more polypeptides that specifically binds an antigen. One example of an antibody is the naturally occurring structural unit found in humans and other mammals which comprises a tetramer of two identical pairs of antibody chains, each pair having one light and one heavy chain. In each pair, the light and heavy chain variable regions are together responsible for binding to an antigen, and the constant regions are responsible for the antibody effector functions. The term antibody encompasses monoclonal antibodies, polyclonal antibodies, chimeric antibodies, humanized antibodies, human antibodies, murine antibodies, rabbit antibodies, camelid antibodies, and antibodies from other mammalian and non-mammalian species. The term antibody also encompasses single-chain antibodies, bi-specific hybrid antibodies, and fusion proteins comprising an antigen-binding portion of an antibody and a nonantibody protein. The term antibody also encompasses includes antigen-binding fragments of antibodies which retain specific binding to antigen, including, but not limited to, Fab, Fv, scFv, and Fd fragments.

[0141] The term “binding pair” as used herein refers to a pair of binding partners that exhibit specific binding between them. In some embodiments, a binding pair can selectively interact through covalent or non-covalent binding. In some embodiments, a binding pair can selectively interact by hybridization, ionic bonding, hydrogen bonding, van der Waals interactions, or any combination of these forces. In some embodiments, a binding partner can comprise, for example, biotin, avidin, streptavidin, digoxigenin, inosine, avidin, GST sequences, modified GST sequences, biotin ligase recognition (BiTag) sequences, S tags, SNAP -tags, enterokinase sites, thrombin sites, antibodies or antibody domains, antibody fragments, antigens, receptors, receptor domains, receptor fragments, or combinations thereof. Examples of binding pairs include biotimavidin, biotin: streptavidin, antibody: antigen, complementary nucleic acids, hapten / antibody, lectin / carbohydrate, apoprotein / cofactor and biotin / streptavidin, as well as others set forth above.

[0142] The term “specific binding” refers to the ability of a binding partner to preferentially bind to its reciprocal binding partner that is present in a homogeneous mixture of different molecules. In some embodiments, specific binding discriminates between a reciprocal binding partner and other molecules by at least 100-fold, 1000-fold, 10,000-fold, 100,000-fold, or more. In some embodiments, the affinity between binding partners of a binding pair when they are specifically bound in a complex is characterized by a KD (dissociation constant) of less than 10'6M, less than 10‘7M, less than 10'8M, less than 10‘9M, less than IO-10M, less than 10’11M, or less than about 10‘12M, or less.

[0143] As used herein, a “capture binding partner” refers to a binding partner that is configured to capture (e.g., isolate, purify, immobilize, extract) a nucleic acid tagged with its reciprocal binding partner. For example, streptavidin coated on a bead would be a capturebinding partner for an input nucleic acid complex having a tag comprising a biotin moiety. A capture binding partner and its reciprocal binding partner may comprise any suitable binding pair.Exemplary Embodiments

[0144] Embodiment 1. A method of preparing nucleic acid molecules for methylation analysis comprising: attaching a methylated adaptor to an input nucleic acid; hybridizing a tagged primer to the input nucleic acid, wherein the tagged primer comprises a first tag comprising a binding partner of a first binding pair; extending the tagged primer to produce a tagged primer extension product comprising a first segment complementary to at least a portion the input nucleic acid and a second segment complementary to at least a portion of the methylated adaptor, wherein the second segment is not methylated; separating the input nucleic acid from the tagged PE product by binding the binding partner of the first tag to a reciprocal binding partner of the first binding pair; subjecting the input nucleic acid to base conversion after separation from the tagged PE product to form a base-converted nucleic acid; and combining the base-converted nucleic acid and the amplicons for sequencing in a combined sequencing run.

[0145] Embodiment 2. The method of embodiment 1, further comprising hybridizing the input nucleic acid and the tagged primer extension product and construct to a target enrichment probe.

[0146] Embodiment 3. The method of embodiment 2, wherein the target enrichment probe is a bridge probe.

[0147] Embodiment 4. The method of embodiment 3, further comprising hybridizing the bridge probe with an anchor probe.

[0148] Embodiment 5. The method of embodiment 4, the anchor probe is attached to a binding partner of a second binding pair.

[0149] Embodiment 6. The method of embodiment 2, the target enrichment probe is attached to a binding partner of a second binding pair.

[0150] Embodiment 7. The method of any of embodiments 1 to 6, wherein the binding partner in the first tag is selected from biotin, digoxigenin, 5-bromo-2’ -deoxyuridine (BrdU), 2,4-dinitrophenyl (DNP), nitrilotriacetic acid or a nitrilotriacetate (NTA) such as nickel nitrilotriacetate (Ni-NTA), tri s-Nitrilotri acetate (tris-NTA), a tyramine, a thiol, an amine, an aldehyde, an alkyne or an azide.

[0151] Embodiment 8. The method of embodiment 7, wherein the reciprocal binding partner of the first binding pair is selected from an avidin moiety, an anti-DIG antibody, an anti- BrdU antibody, an anti-DNP antibody, poly-Histidine, a tyrosine, a thiol, an amine, an aldehyde, an alkyne, or an azide.

[0152] Embodiment 9. The method of any of embodiments 1 to 6, wherein the binding partner in the first tag is biotin.

[0153] Embodiment 10. The method of embodiment 9, wherein the reciprocal binding partner of the first binding pair is an avidin moiety.

[0154] Embodiment 11. The method of any of embodiments 1 to 10, further comprising contacting the tagged primer extension products with a second tag, wherein the second tag comprises a reciprocal binding partner of the first binding pair and the binding partner of a second binding pair.

[0155] Embodiment 12. The method of embodiment 11, wherein the binding partner of the second binding pair is selected from biotin, DIG, 5-bromo-2’ -deoxyuridine (BrdU), 2,4-dinitrophenyl (DNP), nitrilotri acetic acid or a nitrilotriacetate (NT A) such as nickel nitrilotriacetate (Ni-NTA), tris-Nitrilotriacetate (tris-NTA), a tyramine, a thiol, an amine, an aldehyde, an alkyne or an azide or other groups reacting by click chemistry, and mixtures thereof; with the proviso that the binding partner of the second binding pair is not a binding partner of the first binding pair.

[0156] Embodiment 13. The method of embodiment 12, wherein the reciprocal binding partner of the second pair is selected from an avidin moiety, an anti-DIG antibody, an anti-BrdU antibody, an anti-DNP antibody, poly-Histidine, a tyrosine residue, a thiol, an amine, an aldehyde, an alkyne, or an azide.

[0157] Embodiment 14. The method of embodiment 11, wherein the second tag comprises anti-DIG antibody as the reciprocal binding partner of the first binding pair, and biotin as the binding partner of the second binding pair.

[0158] Embodiment 15. The method of any of embodiments 11 to 14, further comprising attaching the second tag to a reciprocal binding pair of the second binding pair.

[0159] Embodiment 16. The method of embodiment 15, wherein the reciprocal binding pair of the second binding pair is attached to the solid support.

[0160] Embodiment 17. The method of embodiment 15 or 16, wherein the second tag comprises an antibody conjugated with a biotin moiety, and the reciprocal binding partner of the second binding pair is streptavidin coated on a solid support.

[0161] Embodiment 18. A method of preparing nucleic acid molecules for analysis comprising: attaching a first tag to an input nucleic acid, wherein the first tag comprises a binding partner of a first binding pair; performing an amplification of the input nucleic acid to produce amplicons of the input nucleic acid; after the amplification is performed, separating theinput nucleic acid from the amplicons by binding the binding partner of the first tag to a reciprocal binding partner of the first binding pair; subjecting the input nucleic acid to base conversion after separation from the amplicons to form a base-converted nucleic acid; and combining the base-converted nucleic acid and the amplicons for sequencing in a combined sequencing run.

[0162] Embodiment 19. The method of embodiment 18, wherein the binding partner in the first tag is selected from biotin, digoxigenin, 5-bromo-2’-deoxyuridine (BrdU), 2,4- dinitrophenyl (DNP), nitrilotriacetic acid or a nitrilotriacetate (NT A) such as nickel nitrilotriacetate (Ni-NTA), tris-Nitrilotriacetate (tris-NTA), a tyramine, a thiol, an amine, an aldehyde, an alkyne or an azide.

[0163] Embodiment 20. The method of embodiment 18, wherein the reciprocal binding partner of the first binding pair is selected from an avidin moiety, anti-DIG antibody, an anti- BrdU antibody, an anti-DNP antibody, poly-Histidine, a tyrosine, a thiol, an amine, an aldehyde, an alkyne, or an azide.

[0164] Embodiment 21. The method of any of embodiments 18 to 20, wherein the binding partner in the first tag is digoxigenin (DIG).

[0165] Embodiment 22. The method of embodiment 21, wherein the reciprocal binding partner of the first binding pair is an anti-DIG antibody.

[0166] Embodiment 23. The method of embodiment 18 or 19, further comprising attaching a second tag to the first tag to produce an input nucleic acid complex, wherein the second tag comprises a reciprocal binding partner of the first binding pair.

[0167] Embodiment 24. The method of embodiment 23, wherein the second tag further comprises a binding partner of a second binding pair.

[0168] Embodiment 25. The method of embodiment 24, wherein the second tag comprises an antibody as the reciprocal binding partner of the first binding pair and biotin as the binding partner of the second binding pair.

[0169] Embodiment 26. The method of embodiment 24, wherein the binding partner of the second binding pair is selected from biotin, DIG, 5-bromo-2’ -deoxyuridine (BrdU), 2,4- dinitrophenyl (DNP), nitrilotriacetic acid or a nitrilotriacetate (NTA) such as nickel nitrilotriacetate (Ni-NTA), tri s-Nitrilotri acetate (tris-NTA), a tyramine, a thiol, an amine, an aldehyde, an alkyne or an azide or other groups reacting by click chemistry, and mixtures thereof; with the proviso that the binding partner of the second binding pair is not a binding partner of the first binding pair.

[0170] Embodiment 27. The method of embodiment 26, wherein the reciprocal binding partner of the second binding pair is selected from an avidin moiety, an anti -DIG antibody, an anti-BrdU antibody, an anti-DNP antibody, poly-Histidine, a tyrosine residue, a thiol, an amine, an aldehyde, an alkyne, or an azide.

[0171] Embodiment 28. The method of any of embodiments 18 to 27, wherein the first tag is covalently bound to the input nucleic acid.

[0172] Embodiment 29. The method of any of embodiments 18 to 28, wherein the second tag is non-covalently bound to the first tag.

[0173] Embodiment 30. The method of any of embodiments 23 to 29, wherein the reciprocal binding pair of the second binding pair is attached to the solid support.

[0174] Embodiment 31. The method of embodiment 18, wherein the amplification is performed before attaching a second tag to the first tag.

[0175] Embodiment 32. The method of embodiment 18, wherein the amplification is performed after attaching a second tag to the first tag of the second binding pair.

[0176] Embodiment 33. The method of any of embodiments 18 to 32, further comprising hybridizing the input nucleic acid construct and / or the amplicons to a probe.

[0177] Embodiment 34. The method of embodiment 33, wherein the probe is a bridge probe.

[0178] Embodiment 35. The method of embodiment 34, further comprising hybridizing the bridge probe with an anchor probe.

[0179] Embodiment 36. The method of embodiment 35, the anchor probe is attached to an enrichment tag such as biotin.

[0180] Embodiment 37. A method of preparing nucleic acid molecules for methylation analysis comprising: attaching a methylated adaptor to an input nucleic acid; hybridizing a primer to the input nucleic acid, wherein the primer is not methylated; extending the primer to produce a primer extension product comprising a first segment complementary to at least a portion the input nucleic acid and a second segment complementary to at least a portion of the methylated adaptor, wherein the second segment is not methylated; contacting the mixture of input nucleic acid and the primer extension product with one or more base conversion reagent, thereby forming either converted input nucleic acid or converted primer extension product; and analyzing the converted input nucleic acid and / or the converted primer extension product to determine methylation status of the input nucleic acid.

[0181] Embodiment 38. The method of embodiment 37, wherein the one or more base conversion reagents comprises a ten eleven translocation (TET) enzyme that converts 5meC to 5-carboxyl cytosine (5caC) and / or 5-formylcytosine (5fC), and a borane reducing agent that converts 5caC and / or 5FC to thymine.

[0182] Embodiment 39. The method of embodiment 38, wherein the converted input nucleic acid is sequenced and analyzed for methylation status by Whole Genome Methylation Sequencing (WGMS).

[0183] Embodiment 40. The method of embodiment 39, further comprising enriching the input nucleic acid for a target to form an enriched input nucleic acid, and sequencing and analyzing by Targeted Methylation Sequencing (TMS).

[0184] Embodiment 41. The method of embodiment 40, wherein the primer extension product is not converted and is used for Whole Genome Sequence (WGS) or Targeted Sequencing (TS).

[0185] Embodiment 42. The method of embodiment 37, wherein the primer is extended with a sequencing primer extension reagent comprising 5medCTP and that does not contain dCTP, so that the primer extension product contains 5meC bases and does not contain cytosine bases.

[0186] Embodiment 43. The method of embodiment 42, wherein the one or more base conversion reagents comprise bisulfite or a cytosine deaminase, or other enzyme that converts cytosine to uracil.

[0187] Embodiment 44. The method of embodiment 37, wherein cytosine present in the input nucleic acid is converted to uracil, and the converted input nucleic acid is sequenced and analyzed for methylation status.

[0188] Embodiment 45. The method of embodiment 44, wherein the primer extension product is not converted and is used for Whole Genome Sequence (WGS) or Targeted Sequencing (TS).

[0189] Embodiment 46. The method of any of embodiments 37 to 45, further comprising hybridizing the input nucleic acid and / or the primer extension product to a target enrichment probe.

[0190] Embodiment 47. The method of embodiment 46, wherein the target enrichment probe is a bridge probe.

[0191] Embodiment 48. The method of embodiment 47, further comprising hybridizing the bridge probe with an anchor probe.

[0192] In view of this disclosure it is noted that the methods and kits can be implemented in keeping with the present teachings. Further, the various components, materials, structures and parameters are included by way of illustration and example only and not in any limiting sense.In view of this disclosure, the present teachings can be implemented in other applications and components, materials, structures and equipment to implement these applications can be determined, while remaining within the scope of the appended claims.

Claims

CLAIMSWe claim:

1. A method of preparing nucleic acid molecules for methylation analysis comprising: attaching a methylated adaptor to an input nucleic acid; hybridizing a tagged primer to the input nucleic acid, wherein the tagged primer comprises a first tag comprising a binding partner of a first binding pair; extending the tagged primer to produce a tagged primer extension product comprising a first segment complementary to at least a portion the input nucleic acid and a second segment complementary to at least a portion of the methylated adaptor, wherein the second segment is not methylated; separating the input nucleic acid from the tagged PE product by binding the binding partner of the first tag to a reciprocal binding partner of the first binding pair; subjecting the input nucleic acid to base conversion after separation from the tagged PE product to form a base-converted nucleic acid; and combining the base-converted nucleic acid and the amplicons for sequencing in a combined sequencing run.

2. The method of claim 1, further comprising hybridizing the input nucleic acid and the tagged primer extension product and construct to a target enrichment probe.

3. The method of claim 2, wherein the target enrichment probe is a bridge probe.

4. The method of claim 3, further comprising hybridizing the bridge probe with an anchor probe.

5. The method of claim 4, the anchor probe is attached to a binding partner of a second binding pair.

6. The method of claim 2, the target enrichment probe is attached to a binding partner of a second binding pair.

7. The method of claim 1, wherein the binding partner in the first tag is selected from biotin, digoxigenin, 5-bromo-2’-deoxyuridine (BrdU), 2,4-dinitrophenyl (DNP), nitrilotri acetic acid or a nitrilotriacetate (NTA) such as nickel nitrilotriacetate (Ni-NTA), trisNitrilotriacetate (tris-NTA), a tyramine, a thiol, an amine, an aldehyde, an alkyne or an azide.

8. The method of claim 7, wherein the reciprocal binding partner of the first binding pair is selected from an avidin moiety, an anti-DIG antibody, an anti-BrdU antibody, an anti- DNP antibody, poly-Histidine, a tyrosine, a thiol, an amine, an aldehyde, an alkyne, or an azide.

9. The method of claim 1, wherein the binding partner in the first tag is biotin.

10. The method of claim 9, wherein the reciprocal binding partner of the first binding pair is an avidin moiety.11 . The method of claim 1 , further comprising contacting the tagged primer extension products with a second tag, wherein the second tag comprises a reciprocal binding partner of the first binding pair and the binding partner of a second binding pair.

12. The method of claim 11, wherein the binding partner of the second binding pair is selected from biotin, DIG, 5-bromo-2’ -deoxyuridine (BrdU), 2,4-dinitrophenyl (DNP), nitrilotriacetic acid or a nitrilotriacetate (NTA) such as nickel nitrilotriacetate (Ni-NTA), trisNitrilotriacetate (tris-NTA), a tyramine, a thiol, an amine, an aldehyde, an alkyne or an azide or other groups reacting by click chemistry, and mixtures thereof; with the proviso that the binding partner of the second binding pair is not a binding partner of the first binding pair.

13. The method of claim 12, wherein the reciprocal binding partner of the second pair is selected from an avidin moiety, an anti-DIG antibody, an anti-BrdU antibody, an anti-DNP antibody, poly-Histidine, a tyrosine residue, a thiol, an amine, an aldehyde, an alkyne, or an azide.

14. The method of claim 11, wherein the second tag comprises anti-DIG antibody as the reciprocal binding partner of the first binding pair, and biotin as the binding partner of the second binding pair.

15. The method of claim 14, further comprising attaching the second tag to a reciprocal binding pair of the second binding pair.

16. The method of claim 15, wherein the reciprocal binding pair of the second binding pair is attached to the solid support.

17. The method of claim 16, wherein the second tag comprises an antibody conjugated with a biotin moiety, and the reciprocal binding partner of the second binding pair is streptavidin coated on a solid support.

18. A method of preparing nucleic acid molecules for analysis comprising: attaching a first tag to an input nucleic acid, wherein the first tag comprises a binding partner of a first binding pair; performing an amplification of the input nucleic acid to produce amplicons of the input nucleic acid; after the amplification is performed, separating the input nucleic acid from the amplicons by binding the binding partner of the first tag to a reciprocal binding partner of the first binding pair; subjecting the input nucleic acid to base conversion after separation from the amplicons to form a base-converted nucleic acid; and combining the base-converted nucleic acid and the amplicons for sequencing in a combined sequencing run.

19. The method of claim 18, wherein the binding partner in the first tag is selected from biotin, digoxigenin, 5-bromo-2’-deoxyuridine (BrdU), 2,4-dinitrophenyl (DNP),nitrilotri acetic acid or a nitrilotriacetate (NT A) such as nickel nitrilotriacetate (Ni-NTA), tris¬Nitrilotriacetate (tris-NTA), a tyramine, a thiol, an amine, an aldehyde, an alkyne or an azide.

20. The method of claim 18, wherein the reciprocal binding partner of the first binding pair is selected from an avidin moiety, anti-DIG antibody, an anti-BrdU antibody, an anti-DNP antibody, poly-Histidine, a tyrosine, a thiol, an amine, an aldehyde, an alkyne, or an azide.

21. The method of claim 18, wherein the binding partner in the first tag is digoxigenin (DIG).

22. The method of claim 21, wherein the reciprocal binding partner of the first binding pair is an anti-DIG antibody.

23. The method of claim 18, further comprising attaching a second tag to the first tag to produce an input nucleic acid complex, wherein the second tag comprises a reciprocal binding partner of the first binding pair.

24. The method of claim 23, wherein the second tag further comprises a binding partner of a second binding pair.

25. The method of claim 24, wherein the second tag comprises an antibody as the reciprocal binding partner of the first binding pair and biotin as the binding partner of the second binding pair.

26. The method of claim 24, wherein the binding partner of the second binding pair is selected from biotin, DIG, 5-bromo-2’ -deoxyuridine (BrdU), 2,4-dinitrophenyl (DNP), nitrilotriacetic acid or a nitrilotriacetate (NTA) such as nickel nitrilotriacetate (Ni-NTA), trisNitrilotriacetate (tris-NTA), a tyramine, a thiol, an amine, an aldehyde, an alkyne or an azide or other groups reacting by click chemistry, and mixtures thereof; with the proviso that the binding partner of the second binding pair is not a binding partner of the first binding pair.

27. The method of claim 26, wherein the reciprocal binding partner of the second binding pair is selected from an avidin moiety, an anti-DIG antibody, an anti-BrdU antibody, an anti-DNP antibody, poly-Histidine, a tyrosine residue, a thiol, an amine, an aldehyde, an alkyne, or an azide.

28. The method of claim 18, wherein the first tag is covalently bound to the input nucleic acid.

29. The method of claim 28, wherein the second tag is non-covalently bound to the first tag.

30. The method of claim 29, wherein the reciprocal binding pair of the second binding pair is attached to the solid support.

31. The method of claim 18, wherein the amplification is performed before attaching a second tag to the first tag.

32. The method of claim 18, wherein the amplification is performed after attaching a second tag to the first tag of the second binding pair.

33. The method of claim 18, further comprising hybridizing the input nucleic acid construct and / or the amplicons to a probe.

34. The method of claim 33, wherein the probe is a bridge probe.

35. The method of claim 34, further comprising hybridizing the bridge probe with an anchor probe.

36. The method of claim 35, the anchor probe is attached to an enrichment tag such as biotin.

37. A method of preparing nucleic acid molecules for methylation analysis comprising: attaching a methylated adaptor to an input nucleic acid;hybridizing a primer to the input nucleic acid, wherein the primer is not methylated; extending the primer to produce a primer extension product comprising a first segment complementary to at least a portion the input nucleic acid and a second segment complementary to at least a portion of the methylated adaptor, wherein the second segment is not methylated; contacting the mixture of input nucleic acid and the primer extension product with one or more base conversion reagent, thereby forming either converted input nucleic acid or converted primer extension product; and analyzing the converted input nucleic acid and / or the converted primer extension product to determine methylation status of the input nucleic acid.

38. The method of claim 37, wherein the one or more base conversion reagents comprises a ten eleven translocation (TET) enzyme that converts 5meC to 5-carboxylcytosine (5caC) and / or 5-formylcytosine (5fC), and a borane reducing agent that converts 5caC and / or 5FC to thymine.

39. The method of claim 38, wherein the converted input nucleic acid is sequenced and analyzed for methylation status by Whole Genome Methylation Sequencing (WGMS).

40. The method of claim 39, further comprising enriching the input nucleic acid for a target to form an enriched input nucleic acid, and sequencing and analyzing by Targeted Methylation Sequencing (TMS).41 . The method of claim 40, wherein the primer extension product is not converted and is used for Whole Genome Sequence (WGS) or Targeted Sequencing (TS).

42. The method of claim 37, wherein the primer is extended with a sequencing primer extension reagent comprising 5medCTP and that does not contain dCTP, so that the primer extension product contains 5meC bases and does not contain cytosine bases.

43. The method of claim 42, wherein the one or more base conversion reagents comprise bisulfite or a cytosine deaminase, or other enzyme that converts cytosine to uracil.

44. The method of claim 37, wherein cytosine present in the input nucleic acid is converted to uracil, and the converted input nucleic acid is sequenced and analyzed for methylation status.

45. The method of claim 44, wherein the primer extension product is not converted and is used for Whole Genome Sequence (WGS) or Targeted Sequencing (TS).

46. The method of claim 37, further comprising hybridizing the input nucleic acid and / or the primer extension product to a target enrichment probe.

47. The method of claim 46, wherein the target enrichment probe is a bridge probe.

48. The method of claim 47, further comprising hybridizing the bridge probe with an anchor probe.

Citation Information

Patent Citations

  • High throughput methylation detection method

    EP2725125A1

  • Preparation of templates for methylation analysis

    US20090148842A1

  • Method of estimating the amount of a methylated locus in a sample

    US20160047002A1

  • Methods for targeted nucleic acid library formation

    US20210355485A1

  • Systems and methods for targeted nucleic acid capture

    US20230193380A1