Nucleic acid preparation and analysis techniques

Adaptor sets with forked adaptors and strand-displacing polymerases generate tandem repeats of target nucleic acids, addressing the loss of contiguity in sequencing templates and enabling accurate sequencing and methylation analysis.

WO2026006746A9PCT designated stage Publication Date: 2026-04-09ILLUMINA INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-06-27
Publication Date
2026-04-09

AI Technical Summary

Technical Problem

Next-generation sequencing methods lose association between complementary strands of double-stranded nucleic acids during fragmentation and template preparation, leading to the loss of contiguity information and challenges in identifying contiguous fragments in the original genome.

Method used

The use of adaptor sets with forked adaptors and strand-displacing polymerases to generate tandem repeats of target nucleic acids, allowing for the extension and binding of complementary sequences to create sequencing templates with multiple copies of the same insert sequence, thereby maintaining contiguity information and facilitating error correction and methylation analysis.

Benefits of technology

The method preserves contiguity information and enables accurate sequencing by generating templates with tandem repeats, correcting random errors and identifying nucleobase damage, and supporting methylation analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025035718_09042026_PF_FP_ABST
    Figure US2025035718_09042026_PF_FP_ABST
Patent Text Reader

Abstract

Nucleic acid preparation and analysis techniques described. In an embodiment, techniques for generating tandem repeats include using adaptor sets having two types of adaptors with respective complementary regions present. When an adaptor of each type is present on end of a nucleic acid fragment, the complementary regions can bind to one another to generate tandem repeats of an insert, e.g., a fragment generated from a target nucleic acid.
Need to check novelty before this filing date? Find Prior Art

Description

NUCLEIC ACID PREPARATION AND ANALYSIS TECHNIQUESCROSS-REFERENCE TO RELATED APPLICATIONS

[0001] The present application claims priority to and the benefit of U.S. Provisional Application No. 63 / 665,758, filed June 28, 2024, the disclosure of which is incorporated by reference in its entirety herein for all purposes.REFERENCE TO ELECTRONIC SEQUENCE LISTING

[0002] The application contains a Sequence Listing which has been submitted electronically in .XML format and is hereby incorporated by reference in its entirety. Said .XML copy, created on June 25, 2025, is named “ ILUM0135PCT .xml” and is 39,130 bytes in size. The sequence listing contained in this .XML file is part of the specification and is hereby incorporated by reference herein in its entirety.BACKGROUND

[0003] The disclosed technology relates generally to nucleic acids such as oligonucleotides that are used in sequencing reactions.

[0004] The subject matter discussed in this section should not be assumed to be prior art merely as a result of its mention in this section. Similarly, a problem mentioned in this section or associated with the subject matter provided as background should not be assumed to have been previously recognized in the prior art. The subject matter in this section merely represents different approaches, which in and of themselves can also correspond to implementations of the claimed technology.

[0005] Sequencing methodology of next-generation sequencing (NGS) platforms typically makes use of nucleic acid fragment libraries. Short read sequencing methods may include generating short separate fragments from intact genomic DNA or RNA. These fragments are generated in several ways such as physical shearing, enzymatic digestion, or polymeraseextension from one or more primers. Template preparation then modifies and appends synthetic adaptors to these fragments to enable them to be sequenced. These sequencing templates almost always contain a single fragment from the original sample comprising the sequence of bases in the same order and juxtaposition as in the intact genome. Where a template is double-stranded, the complement of a sequence is associated by hybridization of the two strands. However, when a double-stranded template is denatured, the two complementary strands separate, and a template becomes a single strand comprising a single sequence fragment from the original sample. In this process, any association between the two complementary strands is lost. In addition, in this process of fragmentation and template preparation, any association between two or more fragments that were contiguous in the original unfragmented genome is also lost.BRIEF DESCRIPTION

[0006] In one embodiment, an adaptor set for sequence analysis is provided. The adaptor set includes a first forked adaptor having a double-stranded region and a forked region, the forked region comprising a first strand and a second strand noncomplementary to one another, the first strand comprising a first adaptor sequence and the second strand comprising a hybridization sequence and an extension binding sequence. The adaptor set also includes a second forked adaptor having a double-stranded region; and a forked region, the forked region comprising a first strand and a second strand noncomplementary to one another, the second strand comprising a complement of the hybridization sequence and wherein the second forked adaptor does not comprise the extension binding sequence or its complement, and the second strand comprising a second adaptor sequence different than the first adaptor sequence.

[0007] In one embodiment, a method of generating a tandem repeat of a target nucleic acid is provided. The method includes steps of providing an adaptor-linked fragment comprising a first adaptor of a first type at the first end of a target nucleic acid and a second adaptor of a second type at the second end of the target nucleic acid; using a strand-displacing polymerase to extend from a primer bound to a forked region of the first adaptor and not bound to a forked region of the second adaptor to displace a first strand of the adaptor-linked fragment;permitting binding of complementary sequences in the forked region of the first adaptor and the forked region of the second adaptor; and extending from at least one 3’ end in the forked region of the first adaptor or the forked region of the second adaptor to generate a product. The generated product includes a double-stranded region having a same sequence as the target nucleic acid and, extending from the double-stranded region, one of another double-stranded region having the same sequence as the target nucleic acid or a singled-stranded region complementary to one strand of the double-stranded region.

[0008] In one embodiment, a method of generating a sequencing library is provided. The method includes providing a pool of adaptor-linked fragments having different inserts generated from a target nucleic acid, wherein each adaptor-linked fragment comprises a double-stranded insert of the target nucleic acid flanked by forked adaptors. The pool of adaptor-linked fragments has a mix of fragments comprising: a first fragment set comprising a first adaptor type at both ends of the respective double-stranded insert; a second fragment set comprising a second adaptor type at both ends of the respective double-stranded insert; and a third fragment set comprising the first adaptor type at the first end and the second adaptor type at the second end of the respective double-stranded insert. The method also includes using a strand-displacing polymerase to extend from a primer bound to a forked region of the first adaptor type and not bound to a forked region of the second adaptor type to displace strands in the respective double-stranded insert from the first fragment set and the third fragment set, but not the second fragment set, to generate strand copies in the respective double-stranded insert; permitting binding of complementary sequences in the forked region of the first adaptor type and the forked region of the second adaptor type in the third fragment set; and extending from at least one 3’ end in the forked region of the first adaptor type or the forked region of the second adaptor type in the third fragment set to generate a product comprising a first copy of the double-stranded insert and a second copy of the double-stranded insert or a single strand of the double-stranded insert.

[0009] In one embodiment, a composition is provided that includes a pool of adaptor-linked fragments having different inserts generated from a target nucleic acid, wherein each adaptor-linked fragment comprises a double-stranded insert of the target nucleic acid flanked by forked adaptors. The pool of adaptor-linked fragments has a mix of fragments comprising: a first fragment set comprising a first adaptor type at both ends of the respective double-stranded insert; a second fragment set comprising a second adaptor type at both ends of the respective double-stranded insert; and a third fragment set comprising the first adaptor type at the first end and the second adaptor type at the second end of the respective double-stranded insert. The first adaptor type has a double-stranded region; a forked region, the forked region comprising a first strand and a second strand noncomplementary to one another, the second strand comprising a hybridization sequence and an extension binding sequence and the first strand comprising a first adaptor sequence; and a primer bound to the extension binding sequence. The second adaptor type has a double-stranded region; and a forked region, the forked region comprising a first strand and a second strand noncomplementary to one another, the second strand comprising a complement of the hybridization sequence and wherein the second forked adaptor does not comprise the extension binding sequence or its complement, and the first strand comprising a second adaptor sequence different than the first adaptor sequence.

[0010] In one embodiment, a method of generating a tandem repeat of a target nucleic acid is provided. The method includes steps of providing an adaptor-linked fragment comprising a first adaptor of a first type at the first end of a target nucleic acid and a second adaptor of a second type at the second end of the target nucleic acid; immobilizing the adaptor-linked fragment on a substrate; denaturing strands of the adaptor-linked fragment to generate a first strand and a second strand immobilized on the substrate; permitting hybridization of complementary sequences of the first adaptor on a 3 ’ end the first strand and the second adaptor on a 3’ end of the second strand; using a polymerase to extend from the 3’ end of the first strand, wherein the 3’ end of the second strand is blocked from extension by the polymerase to generate a partially double-stranded tandem repeat extension product; treating the partially double-stranded tandem repeat extension product with APOBEC; and generating sequence data from the treated partially double-stranded tandem repeat extension product.

[0011] In one embodiment, a method of generating a tandem repeat of a target nucleic acid is provided. The method includes steps of providing an adaptor-linked fragment comprising a first adaptor of a first type at the first end of a target nucleic acid and a second adaptor of a second type at the second end of the target nucleic acid; using a strand-displacing polymerase to extend from a primer bound to a sequence in the forked region of the first adaptor and the forked region of the second adaptor to displace strands of the adaptor-linked fragment; permitting binding of complementary sequences in the forked region of the first adaptor and the forked region of the second adaptor; and extending from 3’ ends in the forked region of the first adaptor and the forked region of the second adaptor to generate a product. The generated products isa double-stranded region having a same sequence as the target nucleic acid; and another double-stranded region having the same sequence as the target nucleic acid.

[0012] In one embodiment, an adaptor-linked fragment is provided that includes a doublestranded insert generated from a target nucleic acid flanked by adaptors at both ends, the adaptors comprising a first adaptor and a second adaptor each comprising: a double-stranded region adjacent to the double-stranded insert; and a forked region comprising a first strand and a second strand noncomplementary to one another, the first strand comprising a 3’ end and the second strand comprising a blocked 3’ end, wherein a portion of the first strand comprising the 3’ end is complementary to a complementary sequence of the second strand.

[0013] In one embodiment, a method is provided that includes contacting an adaptor-linked fragment as provided herein with a strand-displacing polymerase; extending from the 3 ’end of the first strand bound to the complementary sequence of the second strand to generate extension products using both strands of the double-stranded insert as template; and sequencing the extension products.

[0014] In one embodiment, a method is provided that includes denaturing an adaptor-linked fragment as provided herein to generate a top denatured strand comprising the first strand of the first adaptor and the second strand of the second adaptor; allowing the 3’ end of the first strand of the first to bind to the complementary sequence of the second strand of the second adaptor; extending from the 3’ end of the first adaptor to generate self-template extensionproducts; amplifying the self-template extension products; and sequencing the amplified selftemplate extension products.

[0015] In one embodiment, an adaptor-linked fragment is provided that includes a doublestranded insert generated from a target nucleic acid flanked by adaptors at both ends, the adaptors comprising a first adaptor and a second adaptor, wherein the first adaptor comprises a first double-stranded region and a second double-stranded region flanking a noncomplementary region; and wherein the second adaptor comprises a first double-stranded region and a second double-stranded region flanking a noncomplementary region and a hairpin loop adjacent to the second double-stranded region.

[0016] In one embodiment, a method is provided that includes denaturing an adaptor-linked fragment as provided herein to generate a single-stranded fragment; allowing the singlestranded fragment to hybridize to a substrate-linked primer complementary to a region of the single-stranded fragment generated from the first double-stranded region of the first adaptor; extending from the substrate-linked primer to copy the single-stranded fragment to generate a copied strand extending from the substrate-linked primer at a 5’ end, and wherein a 3’ end of the copied strand is complementary to the substrate-linked primer; forming a bridge with a second substrate-linked primer using the 3’ end; copying the bridge by extending from the second substrate-linked primer to generate a second copied strand to form a double-stranded bridge with the copied strand; nicking the hairpin loop to generate separated portions of the double-stranded bridge, the separated portions comprising a first strand extending from the substrate-linked primer and a second strand extending from the second substrate-linked primer; and generating sequencing data from the first strand and the second strand.

[0017] In one embodiment, an adaptor-linked fragment is provided that includes a doublestranded insert generated from a target nucleic acid flanked by adaptors at both ends, the adaptors comprising a first adaptor and a second adaptor, wherein the first adaptor comprises a double-stranded region comprising a first sequencing primer sequence and an index sequence; and wherein the second adaptor comprises a double-stranded region comprising a second sequencing primer sequence different than the first sequencing primer sequence andits complement and a hairpin loop, the hairpin loop comprising a nicking site or a cleavable region.

[0018] In one embodiment, a method for detecting methylated cytosine in a target double stranded DNA is provided. The method includes steps of treating a target double stranded DNA to reversibly covalently link two strands of the double stranded DNA to form linked DNA; denaturing the linked DNA so as to create single stranded regions within the linked DNA; converting 5 -methyl cytosine in the single stranded regions to thymine; reannealing the target DNA to reform linked double stranded DNA; reversing the reversable covalent link; and sequencing the DNA to identify regions that differ from a reference DNA.

[0019] In one embodiment, a kit for preparing DNA to detect methylated cytosine in a target DNA is provided. The kit includes components to reversibly covalently link the two strands of the double stranded DNA to from linked DNA; an enzyme with activity for converting 5- methylcytosine to thymine; and components to reverse the reversable covalent link.

[0020] In one embodiment, a complex is provided. The complex includes double stranded DNA where the two strands of the double stranded DNA are reversibly covalently linked; and a cytidine deaminase.

[0021] In one embodiment, a method for detecting methylated cytosine in a target doublestranded nucleic acid fragment is provided. The method includes steps of coupling hairpin adaptors to ends of a double-stranded nucleic acid fragment to form linked DNA; denaturing the linked DNA so as to create single- stranded regions of separate strands of the doublestranded nucleic acid fragment between the hairpin adaptors; selectively converting methylated cytosines in the single-stranded regions to thymine; reannealing the strands to reform the double-stranded nucleic acid fragment; sequencing the double-stranded nucleic acid fragment to generate sequence data; and identifying converted methylated cytosines in the sequence data based on sequence differences relative to a reference sequence.

[0022] In one embodiment, a method is provided that includes steps of providing fragments of a target nucleic acid with 3 ’ A-tails; contacting a substrate comprising immobilized duplexes with the fragments, wherein the immobilized duplexes are 3’ T-tailed; ligating a first fragment of the fragments to a first end of an individual immobilized duplex of the substrate- immobilized duplexes and a second fragment of the fragments to a second end of the individual immobilized duplex such that the individual immobilized duplex is positioned between the first fragment and the second fragment and such that the first fragment and the second fragment each have an available A-tailed 3’ end; ligating a first adaptor of a first type to the available A-tailed 3 ’ end of the first fragment and the second fragment; cleaving the immobilized duplex from the first fragment and the second fragment; and ligating a second adaptor of a second type to cleaved ends of the first fragment and the second fragment.

[0023] In one embodiment, a substrate is provided that includes immobilized double-stranded nucleic acids comprising 3’ T-tails, wherein the double-stranded nucleic acids are immobilized via a modified nucleotide comprising an affinity moiety and wherein an individual immobilized double-stranded nucleic acid comprises a first restriction enzyme recognition sequence at a first end and a second restriction enzyme recognition sequence at a second end.

[0024] In one embodiment, a method is provided that includes denaturing an adaptor-linked fragment including a hairpin to generate a single-stranded fragment; allowing the singlestranded fragment to hybridize to a substrate-linked primer of a first set of substrate-linked primers, the substrate-linked primer being complementary to a region of the single-stranded fragment generated from the first double-stranded region of the first adaptor; extending from the substrate-linked primer to copy the single-stranded fragment to generate a copied strand extending from the substrate-linked primer at a 5’ end, and wherein a 3’ end of the copied strand is complementary to the substrate-linked primer; forming a single-stranded bridge with a second substrate-linked primer of the first set using the 3’ end; copying the bridge by extending from the second substrate-linked primer to generate a second copied strand to form a first double-stranded bridge with the copied strand; nicking within the double- stranded bridge and removing non-immobilized nicked strands to generate a bridge comprisingpartially single-stranded bridge portions joined by a double-stranded portion having free ends; generating sequencing data from the by extending the free ends to form a second doublestranded bridge to generate read 1 sequencing data; deprotecting ends of a second set of substrate-linked primers having a different sequence than the first set of substrate-linked primers; denaturing the reformed bridge and allowing hybridization to primers of the second set of substrate-linked primers; extending from the deprotected ends of the second set of substrate-linked primers to form a third double-stranded bridge, the third double-stranded bridge extending between primers of the first set and second set; denaturing the third doublestranded bridge; and generating read 2 sequencing data from strands of the third doublestranded bridge using a sequencing primer.

[0025] In one embodiment, a method is provided that includes denaturing an adaptor-linked fragment to generate a single-stranded fragment; allowing the single-stranded fragment to hybridize to a substrate-linked primer of a first set of substrate-linked primers, the substrate- linked primer being complementary to a region of the single-stranded fragment generated from the first double-stranded region of the first adaptor; extending from the substrate-linked primer to copy the single- stranded fragment to generate a copied strand extending from the substrate- linked primer at a 5’ end, and wherein a 3’ end of the copied strand is complementary to the substrate-linked primer; forming a single-stranded bridge with a second substrate-linked primer of the first set using the 3’ end; copying the bridge by extending from the second substrate-linked primer to generate a second copied strand to form a first double-stranded bridge with the copied strand; nicking within the double-stranded bridge and removing nonimmobilized nicked strands to generate bridge comprising partially single-stranded bridge portions joined by a double-stranded portion having free ends; generating sequencing data from the by extending the free ends to form a second double-stranded bridge to generate read 1 sequencing data; cutting the second double-stranded bridge to generate single-strands immobilized by the substrate-linked primers of the first set; deprotecting ends of a second set of substrate-linked primers having a different sequence than the first set of substrate-linked primers; and allowing hybridization of the single-strands immobilized by the substrate-linkedprimers of the first set to primers of the second set of substrate-linked primers; extending from the deprotected ends of the second set of substrate-linked primers to form a third doublestranded bridge, the third double-stranded bridge extending between primers of the first set and second set; denaturing the third double-stranded bridge; and generating read 2 sequencing data from strands of the third double-stranded bridge using a sequencing primer.

[0026] In one embodiment, an adaptor for preparing a sequencing library includes a first oligonucleotide comprising a palindromic sequence and a second oligonucleotide, the first oligonucleotide and the second oligonucleotide together forming: a double-stranded region; and a forked region, wherein the first oligonucleotide and the second oligonucleotide are noncomplementaiy to one another in the forked region, wherein the palindromic sequence of the first oligonucleotide is partially within the forked region and partially within the doublestranded region.

[0027] In one embodiment, an adaptor-linked fragment includes a double-stranded insert generated from a target nucleic acid flanked by a first adaptor and a second adaptor, the first adaptor and the second adaptor having a same sequence and comprising: a first oligonucleotide comprising a palindromic sequence and a second oligonucleotide, the first oligonucleotide and the second oligonucleotide together forming: a double-stranded region; and a forked region, wherein the first oligonucleotide and the second oligonucleotide are noncomplementaiy to one another in the forked region, wherein the palindromic sequence of the first oligonucleotide is partially within the forked region and partially within the double- stranded region.

[0028] In one embodiment, a method of generating a tandem insert of a target nucleic acid includes providing an adaptor-linked fragment, wherein the first adaptor and the second adaptor are coupled to a surface of a substrate; separating the first oligonucleotide and the second oligonucleotide of the first adaptor and the second adaptor; permitting binding of complementary sequences in respective palindromic sequences of the first oligonucleotide of the first adaptor and the second adaptor; and extending from 3’ ends of the palindromic sequences to generate a product comprising: a double-stranded region having a same sequenceas the target nucleic acid; and another double-stranded region having the same sequence as the target nucleic acid.

[0029] In one embodiment, a method of generating a tandem insert of a target nucleic acid includes providing an adaptor-linked fragment, separating the first oligonucleotide and the second oligonucleotide of the first adaptor and the second adaptor without separating strands of the target nucleic acid; permitting binding of complementary sequences in respective palindromic sequences of the first oligonucleotide of the first adaptor and the second adaptor to form a circularized fragment; and extending from 3’ ends of the palindromic sequences using a strand-displacing polymerase to generate a product comprising: a double-stranded region having a same sequence as the target nucleic acid; and another double-stranded region having the same sequence as the target nucleic acid.

[0030] In one embodiment, a method of characterizing a target nucleic acid includes providing a substrate having a plurality of transposome complexes on a surface of the substrate, wherein the plurality of transposomes comprise a first type of transposome complex and a second type of transposome complex, wherein the first type of transposome complex comprises: a transposase; and a first oligonucleotide and a second oligonucleotide forming a doublestranded transposon end sequence, wherein a 5’ end of the first oligonucleotide is coupled to the surface and wherein a single-stranded region of the second oligonucleotide comprises a hybridization sequence; and wherein the second type of transposome complex comprises: a transposase; and a first oligonucleotide and a second oligonucleotide forming a doublestranded transposon end sequence, wherein a 5’ end of the first oligonucleotide is coupled to the surface and wherein a single-stranded region of the second oligonucleotide comprises a complement of the hybridization sequence. The method also includes loading a target nucleic acid onto the surface of the substrate under conditions sufficient to cause fragmentation of the target nucleic acid into a plurality of surface-associated double-stranded target fragments by the plurality of transposome complexes, wherein first oligonucleotides and second oligonucleotides of respective first and second types of transposome complexes are joined to ends of the surface-associated double-stranded target fragments as a result of thefragmentation; denaturing the surface-associated double-stranded target fragments into separate strands; treating the separated strands with a deaminase to convert methyl cytosine into thymine to generate treated strands; permitting binding of the hybridization sequence and its complement in the treated strands; and extending from 3’ ends of the hybridization sequence and its complement to generate a tandem repeat double-stranded product; generating clusters on the surface from the tandem repeat double-stranded product; and generating sequencing data from the clusters.

[0031] In one embodiment, a method of characterizing a target nucleic acid includes providing a substrate having a plurality of transposome complexes on a surface of the substrate, wherein the plurality of transposomes comprise a first type of transposome complex and a second type of transposome complex, wherein the first type of transposome complex comprises: a transposase; and a first oligonucleotide and a second oligonucleotide forming a doublestranded transposon end sequence, wherein a 5’ end of the first oligonucleotide is coupled to the surface and wherein a single-stranded region of the second oligonucleotide comprises a hybridization sequence; and wherein the second type of transposome complex comprises: a transposase; and a first oligonucleotide and a second oligonucleotide forming a doublestranded transposon end sequence, wherein a 5’ end of the first oligonucleotide is coupled to the surface and wherein a single-stranded region of the second oligonucleotide comprises a complement of the hybridization sequence. The method also includes loading a target nucleic acid onto the surface of the substrate under conditions sufficient to cause fragmentation of the target nucleic acid into a plurality of surface-associated double-stranded target fragments by the plurality of transposome complexes, wherein first oligonucleotides and second oligonucleotides of respective first and second types of transposome complexes are joined to ends of the surface-associated double-stranded target fragments as a result of the fragmentation; denaturing the surface-associated double-stranded target fragments into separate strands; permitting binding of the hybridization sequence and its complement in the treated strands; extending from 3’ ends of the hybridization sequence and its complement to generate a tandem repeat double-stranded product; denaturing the tandem repeat doublestranded product into single strands; treating the single strands with a deaminase to convertmethyl cytosine into thymine to generate treated strands; generating clusters on the surface from one or more strands of the treated strands; and generating sequencing data from the clusters.

[0032] In one embodiment, a method of characterizing a target nucleic acid includes providing a substrate having a plurality of transposome complexes on a surface of the substrate, wherein the plurality of transposomes comprise a first type of transposome complex and a second type of transposome complex, wherein the first type of transposome complex comprises: a transposase; and a first oligonucleotide and a second oligonucleotide forming a doublestranded transposon end sequence, wherein a 5’ end of the first oligonucleotide is coupled to the surface and wherein a single-stranded region of the second oligonucleotide comprises a hybridization sequence; and wherein the second type of transposome complex comprises: a transposase; and a first oligonucleotide and a second oligonucleotide forming a doublestranded transposon end sequence, wherein a 5’ end of the first oligonucleotide is coupled to the surface and wherein a single-stranded region of the second oligonucleotide comprises a complement of the hybridization sequence. The method also includes loading a target nucleic acid onto the surface of the substrate under conditions sufficient to cause fragmentation of the target nucleic acid into a plurality of surface-associated double-stranded target fragments by the plurality of transposome complexes, wherein first oligonucleotides and second oligonucleotides of respective first and second types of transposome complexes are joined to ends of the surface-associated double-stranded target fragments as a result of the fragmentation, wherein the target nucleic acid is treated with a deaminase to convert methyl cytosine into thymine; generating a double-stranded tandem insert product from the doublestranded target fragments; and generating clusters on the surface from one or more strands of the double-stranded tandem insert product; and generating sequencing data from the clusters.

[0033] In one embodiment, a method of characterizing a target nucleic acid includes providing a substrate having a plurality of transposome complexes on a surface of the substrate, wherein the plurality of transposomes comprise a first type of transposome complex and a second type of transposome complex, wherein the first type of transposome complex comprises: atransposase; and a first oligonucleotide and a second oligonucleotide forming a doublestranded transposon end sequence, wherein a 5’ end of the first oligonucleotide is coupled to the surface and wherein a single-stranded region of the second oligonucleotide comprises a hybridization sequence; and wherein the second type of transposome complex comprises: a transposase; and a first oligonucleotide and a second oligonucleotide forming a doublestranded transposon end sequence, wherein a 5’ end of the first oligonucleotide is coupled to the surface and wherein a single-stranded region of the second oligonucleotide comprises a complement of the hybridization sequence. The method also includes loading a target nucleic acid onto the surface of the substrate under conditions sufficient to cause fragmentation of the target nucleic acid into a plurality of surface-associated double-stranded target fragments by the plurality of transposome complexes, wherein first oligonucleotides and second oligonucleotides of respective first and second types of transposome complexes are joined to ends of the surface-associated double-stranded target fragments as a result of the fragmentation; generating a double-stranded tandem insert product from the double-stranded target fragments; treating the double-stranded tandem insert product with a deaminase to convert methyl cytosine into thymine; and generating clusters on the surface from one or more strands of the double-stranded tandem insert product; and generating sequencing data from the clusters.

[0034] The preceding description is presented to enable the making and use of the technology disclosed. Various modifications to the disclosed implementations will be apparent, and the general principles defined herein may be applied to other implementations and applications without departing from the spirit and scope of the technology disclosed. Thus, the technology disclosed is not intended to be limited to the implementations shown, but is to be accorded the widest scope consistent with the principles and features disclosed herein. The scope of the technology disclosed is defined by the appended claims.BRIEF DESCRIPTION OF THE DRAWINGS

[0035] These and other features, aspects, and advantages of the present invention will become better understood when the following detailed description is read with reference to theaccompanying drawings in which like characters represent like parts throughout the drawings, wherein:

[0036] FIG. l is a workflow of steps in generating immobilized tandem repeat nucleic acids products, in accordance with aspects of the present disclosure;

[0037] FIG. 2 shows example adaptors used to generate immobilized tandem repeat nucleic acids, in accordance with aspects of the present disclosure;

[0038] FIG. 3 shows example adaptor sets used to generate nucleic acid fragments having a tandem repeat of a target nucleic acid, in accordance with aspects of the present disclosure;

[0039] FIG. 4 shows example adaptor-linked nucleic acid fragments generated from the adaptor sets of FIG. 3, in accordance with aspects of the present disclosure;

[0040] FIG. 5 shows example adaptor-linked nucleic acid fragments generated from the adaptor sets of FIG. 3, in accordance with aspects of the present disclosure;

[0041] FIG. 6 is a schematic illustration of effects of a strand displacing polymerase extension step on the immobilized adaptor-linked nucleic acid fragments of FIG. 5, in accordance with aspects of the present disclosure;

[0042] FIG. 7 is a schematic illustration of effects of a hybridization step on the adaptor-linked nucleic acid fragments of FIG. 6, in accordance with aspects of the present disclosure;

[0043] FIG. 8 shows an example double-stranded nucleic acid having a tandem repeat of a target nucleic acid, in accordance with aspects of the present disclosure;

[0044] FIG. 9A shows an example partially double-stranded nucleic acid having a tandem repeat of a target nucleic acid, in accordance with aspects of the present disclosure;

[0045] FIG. 9B shows example adaptors used to generate the partially-stranded nucleic acid of FIG. 9A, in accordance with aspects of the present disclosure;

[0046] FIG. 10A shows example adaptors with an extension binding sequence present on both first and second adaptors, in accordance with aspects of the present disclosure;

[0047] FIG. 10B shows example adaptor-linked nucleic acid fragments generated from the adaptors of FIG. 10A, in accordance with aspects of the present disclosure;

[0048] FIG. 10C shows example adaptor-linked nucleic acid fragments generated from the adaptors of FIG. 10A, in accordance with aspects of the present disclosure;

[0049] FIG. 10D is a schematic illustration of effects of a strand displacing polymerase extension step on the immobilized adaptor-linked nucleic acid fragments of FIG. 10C, in accordance with aspects of the present disclosure;

[0050] FIG. 10E is a schematic illustration of effects of a hybridization step on the adaptor- linked nucleic acid fragments of FIG. 10D, in accordance with aspects of the present disclosure;

[0051] FIG. 10F shows an example double-stranded nucleic acid having a tandem repeat of a target nucleic acid generated from extension, in accordance with aspects of the present disclosure;

[0052] FIG. 10G shows an example double-stranded nucleic acid having a tandem repeat of a target nucleic acid generated from ligation, in accordance with aspects of the present disclosure;

[0053] FIG. 11 shows example adaptors used to generate experimental results;

[0054] FIG. 12 shows cluster density data of tandem repeat products generated using the adaptors of FIG. 11;

[0055] FIG. 13 shows analysis data of tandem repeat products generated using the adaptors of FIG. 11;

[0056] FIG. 14A shows an example adaptor with a palindromic sequencer that can be used in conjunction with a single adaptor-type or symmetric adaptor workflow, in accordance with aspects of the present disclosure;

[0057] FIG. 14B shows the adaptor of FIG. 14A with an example palindromic sequence;

[0058] FIG. 15 shows an example ligation product with symmetric adaptors with a palindromic sequence, in accordance with aspects of the present disclosure;

[0059] FIG. 16A shows noncomplementarity between forked portions of the palindromic sequence, in accordance with aspects of the present disclosure;

[0060] FIG. 16B shows noncomplementarity between forked portions of an example palindromic sequence, in accordance with aspects of the present disclosure;

[0061] FIG. 17A shows an example surface-linked ligation product with symmetric adaptors with a palindromic sequence, in accordance with aspects of the present disclosure;

[0062] FIG. 17B shows an example intermediate product with hybridized palindromic sequences, in accordance with aspects of the present disclosure;

[0063] FIG. 18 shows an example workflow to generate a tandem insert double-stranded product using symmetric adaptors with a palindromic sequence, in accordance with aspects of the present disclosure;

[0064] FIG. 19 shows self-circularization of a ligation product with symmetric adaptors with a palindromic sequence, in accordance with aspects of the present disclosure;

[0065] FIG. 20 shows downstream processing of a tandem insert double-stranded product, in accordance with aspects of the present disclosure;

[0066] FIG. 21 shows downstream processing of a tandem insert double-stranded product, in accordance with aspects of the present disclosure;

[0067] FIG. 22 shows downstream processing of a tandem insert double-stranded product, in accordance with aspects of the present disclosure;

[0068] FIG. 23 shows an example adaptor set and adaptor-linked fragment products to generate tandem repeats for methylation analysis, in accordance with aspects of the present disclosure;

[0069] FIG. 24 is a schematic illustration of effects of an extension step on the immobilized adaptor-linked nucleic acid fragments of FIG. 23, in accordance with aspects of the present disclosure;

[0070] FIG. 25 shows an example partially double-stranded nucleic acid extension product generated using an adaptor-linked nucleic acid fragment of FIG. 23 and having a tandem repeat of a target nucleic acid, in accordance with aspects of the present disclosure;

[0071] FIG. 26 shows an example partially double-stranded nucleic acid extension product with an example sequence including hydroxymethyl-C and methyl-C bases, in accordance with aspects of the present disclosure;

[0072] FIG. 27 shows effects of mutant APOBEC treatment on the extension product of FIG. 26;

[0073] FIG. 28 shows a sequencing preparation workflow using mutant APOBEC treatment on the extension product of FIG. 26;

[0074] FIG. 29 shows an example workflow for and treatment with P-glucosyltransferase and generation of a partially double-stranded nucleic acid extension product with an example sequence including hydroxymethyl-C and methyl-C bases, in accordance with aspects of the present disclosure;

[0075] FIG. 30 shows effects of mutant APOBEC treatment on the extension product of FIG. 29;

[0076] FIG. 31 shows a sequencing preparation workflow using mutant APOBEC treatment on the extension product of FIG. 26;

[0077] FIG. 32 shows an example adaptor-linked fragment including a 5 ’-5’ linker used to generate a forked region with two 3’ ends, in accordance with aspects of the present disclosure;

[0078] FIG. 33 shows self-annealing of the forked region of the adaptor-linked fragment of FIG. 23;

[0079] FIG. 34 shows same-strand and opposite-strand extension approaches to generating tandem insert extension products, in accordance with aspects of the present disclosure;

[0080] FIG. 35 shows an example same-strand extension reaction in which a strand of the forked region binds to a complement on the same strand and can be extended to generate extension products, in accordance with aspects of the present disclosure;

[0081] FIG. 36 shows the structure of the extension products of the reaction of FIG. 35;

[0082] FIG. 37 shows strand capture and extension of a sequencing reaction for the extension products of the reaction of FIG. 35;

[0083] FIG. 38 shows an example opposite-strand extension reaction in which a strand of the forked region binds to a complement on the opposite strand and can be extended using a stranddisplacing polymerase to generate extension products, in accordance with aspects of the present disclosure;

[0084] FIG. 39 shows the structure of the extension products of the reaction of FIG. 38;

[0085] FIG. 40 shows strand capture and extension of a sequencing reaction for the extension products of the reaction of FIG. 38;

[0086] FIG. 41 shows an example opposite-strand approach for 16QAM sequencing using methyl cytosine-containing adaptors, in accordance with aspects of the present disclosure;

[0087] FIG. 42 shows strand capture and extension of a sequencing reaction for the extension products of the reaction of FIG. 41;

[0088] FIG. 43 shows an example sequencing workflow using an adaptor-linked fragment generated using a hairpin adaptor, in accordance with aspects of the present disclosure;

[0089] FIG. 44 shows an example adaptor set for a sequencing workflow including a hairpin adaptor, in accordance with aspects of the present disclosure;

[0090] FIG. 45 shows an example adaptor-linked fragment for a sequencing workflow generated using the adaptor set of FIG. 44;

[0091] FIG. 46 shows an example sequencing workflow for sequencing the adaptor-linked fragment of FIG. 45;

[0092] FIG. 47 shows an example read and indexing scheme for the sequencing workflow of FIG. 46;

[0093] FIG. 48 shows an example adaptor set including a hairpin adaptor and example adaptor-linked fragment for a sequencing workflow generated using the adaptor set, in accordance with aspects of the present disclosure;

[0094] FIG. 49 shows an example sequencing workflow for sequencing the adaptor-linked fragment of FIG. 48;

[0095] FIG. 50 shows an example adaptor set including a hairpin adaptor and example adaptor-linked fragment for a sequencing workflow generated using the adaptor set an example sequencing workflow for sequencing the adaptor-linked fragment, in accordance with aspects of the present disclosure;

[0096] FIG. 51 shows an example adaptor set including a hairpin adaptor and example adaptor-linked fragment for a sequencing workflow generated using the adaptor set anexample sequencing workflow for sequencing the adaptor-linked fragment, in accordance with aspects of the present disclosure;

[0097] FIG. 52 shows an example read and indexing scheme for the sequencing workflow of FIG. 51;

[0098] FIG. 53 shows an example adaptor set for a sequencing workflow including a hairpin adaptor, in accordance with aspects of the present disclosure;

[0099] FIG. 54 shows an example adaptor-linked fragment for a sequencing workflow generated using the adaptor set of FIG. 53;

[0100] FIG. 55 shows an example sequencing workflow for sequencing the adaptor-linked fragment of FIG. 54;

[0101] FIG. 56 is an example sequencing workflow used in conjunction with hairpin or linked DNA, in accordance with aspects of the present disclosure;

[0102] FIG. 57 are example structures that can be used in conjunction with the workflow of FIG. 56;FIG. 58 shows example products of asymmetric adaptor ligation;

[0103] FIG. 59 shows an example workflow for more efficient asymmetric adaptor ligation, in accordance with aspects of the present disclosure;

[0104] FIG. 60 shows an example workflow for tandem insert product generation, in accordance with aspects of the present disclosure;

[0105] FIG. 61 shows novel transposomes that may be used in conjunction with the workflow of FIG. 60;

[0106] FIG. 62 shows an example workflow for deamination and tandem insert product generation that may use the transposomes of FIG. 61;

[0107] FIG. 63 shows sequencing results using the workflow of FIG. 62;

[0108] FIG. 64 shows an example workflow for tandem insert product generation and deamination that may use a transposome including a cleavable site, in accordance with aspects of the present disclosure;

[0109] FIG. 65 shows sequencing results using the workflow of FIG. 64 when either strand of the tandem library is clustered and the P5 primers are linearized;

[0110] FIG. 66 shows deamination using a dsDNA-compatible deaminase enzyme prior to the tandem reaction on the flow cell

[0111] FIG. 67 shows the tandem reaction on the flow cell followed by deamination using a dsDNA-compatible deaminase enzyme;

[0112] FIG 68 is a schematic illustration of resolvable base call data acquired from a9QAM sequencing reaction, in accordance with aspects of the present disclosure;

[0113] FIG 69 is a schematic illustration of resolvable base call data acquired from a 16QAM sequencing reaction, in accordance with aspects of the present disclosure;

[0114] FIG. 70 is a block diagram of a sequencing device configured to acquire sequencing data , according to an embodiment;

[0115] FIG. 70 shows example transposomes, according to an embodiment; and

[0116] FIG 71 shows a tagmentation reaction, according to an embodiment.DETAILED DESCRIPTION

[0117] The following discussion is presented to enable any person skilled in the art to make and use the technology disclosed, and is provided in the context of a particular application and its requirements. Various modifications to the disclosed implementations will be readily apparent to those skilled in the art, and the general principles defined herein may be applied toother implementations and applications without departing from the spirit and scope of the technology disclosed. Thus, the technology disclosed is not intended to be limited to the implementations shown, but is to be accorded the widest scope consistent with the principles and features disclosed herein.

[0118] Embodiments of the present disclosure are directed to sequencing techniques with improved nucleic acid templates. In an embodiment, the disclosed techniques include preparation of sequencing templates having multiple inserts, e.g., tandem repeated copies of a target nucleic acid present on a sequencing library fragment. This disclosure also relates to methods of use of such templates, including analysis of contiguity information. Further, sequencing templates comprising two copies of the same insert sequence (i.e., an insert sequence and a copy of an insert sequence) can be used to correct for random errors generated during sequencing or amplification or to identify nucleobase damage or other mutation that leads to non-canonical base pairing in a double-stranded nucleic acid. These sequencing templates comprising an insert sequence and a copy of the insert sequence can also be used for methylation analysis.

[0119] FIGS. 1-10 show embodiments of sequencing template preparation with repeated inserts (e.g., tandem repeats) that may be used in conjunction with the disclosed techniques. FIG. 1 illustrates an overview of a workflow to generate a nucleic acid sequencing library having fragments with tandem repeats. The workflow initiates with adaptor-linked fragments 12 generated from a target nucleic acid and immobilized on a surface 14 (e.g., a substrate, a flow cell surface). The adaptor-linked fragments 12 each include a double-stranded insert 20 generated from, in some cases, a target nucleic acid, whereby the double-stranded insert is flanked by adaptors. As illustrated, the adaptors are forked adaptors having noncompl ementary terminal regions such that a portion of each adaptor is single-stranded.

[0120] In the illustrated example, a pool of adaptor-linked fragments 12 having respective different inserts 20 shown as A / A’, B / B’, C / C’, D / D’, and E / E’, are bound to a surface, denatured, reannealed, and then extended to form concatenated sequencing templates. Templates that have a sequence from a first forked adaptor at both ends (e.g. a first adaptortype 22) or a sequence from a second forked adaptor (e.g. a second adaptor type 24) at both ends cannot anneal via their 3’ ends (e.g., templates C and E in FIG. 1) and thus cannot be extended. However, fragments with different adaptor types at respective ends (e.g., a first adaptor type 22 at one end and a second adaptor type 24 at the other end) are able to hybridize via complementary regions such that their respective 3’ ends can be extended. The doublestranded fragments (which are then denatured to single-stranded fragments) may be added (and immobilized) to the surface at a density that favors reannealing of the two fragments from a double-stranded fragments to produce a concatenated sequencing template comprising two copies of the same insert (e.g., A / A’ and A / A’, rather favoring annealing of two fragments from different double-stranded fragments.

[0121] In some embodiments, the adaptor-linked fragments 12 are prepared in solution by ligation of forked adaptors which can then be immobilized on the surface of a solid support via affinity tags (e.g., biotin and streptavidin binding or linkers). In some embodiments, the adaptor-linked fragments 12 are prepared in solution by tagmentation. A mixture of forked adaptors may be contacted with the double-stranded nucleic acid fragments to generate the adaptor-linked fragments 12. These double-stranded nucleic acid fragments may be prepared from DNA (such as genomic DNA or cDNA prepared from RNA) using well-known techniques in the art, such acoustics, nebulization, centrifugal force, needles, or hydrodynamics. Enzymatic means of preparing fragments are also well-known, such as DNase treatment, restriction endonuclease treatment, or tagmentation. When a mixture comprising a first forked adaptor 22 and a second forked adaptor 24 is combined with double-stranded nucleic acid fragments under conditions for ligating, the predicted ratio would be 50% of fragments would be tagged with a first forked adaptor 22 at one end and a second forked adaptor 24 at a second end, 25% of fragments would be tagged with a first forked adaptor 22 at both ends, and 25% of fragments would be tagged with a second forked adaptor 24 at both ends. Thus, as illustrated in FIG. 1, not all of the adaptor-linked fragments 12 have a heteroadaptor configuration with a first forked adaptor 22 at one end and a second forked adaptor 24. Instead, some adaptor-linked fragments 12 have a homadaptor configuration and have a same adaptor type at both ends.

[0122] In some embodiments, a method of generating one or more concatenated or tandem repeat nucleic acid sequencing templates comprises contacting a sample comprising doublestranded nucleic acid fragments each comprising an insert prepared from a target nucleic acid with a composition or kit comprising two forked adaptor types 22, 24, wherein one or both forked adaptors comprise a blocking oligonucleotide (see FIG. 2). In some embodiments, after contacting the sample with the two forked adaptor types, the method comprises ligating the forked adaptors to the double-stranded fragments to prepare adaptor-linked fragments 12 and immobilizing the adaptor-linked fragments 12 on a solid support. In some embodiments, adaptor-linked fragments 12 are prepared via tagmentation and subsequent immobilization to the surface 14. In some embodiments, adaptor-linked fragments 12 are applied to a solid support after ligation or linking (e.g., tagmentation) to forked adaptors. In some embodiments, both the 5’ ends of adaptor-linked fragments 12 comprise an affinity moiety (based on ligation of the first strand of a forked adaptor comprising an affinity moiety) that can bind to a binding moiety on the surface of a solid support. In some embodiments, binding of the affinity moiety to the binding moiety immobilizes adaptor-linked fragments 12 on the solid support, such that they will not be released from the support by temperature changes that can allow release of a blocking oligonucleotide bound to a hybridization sequence or its complement.

[0123] After immobilizing adaptor-linked fragments 12 on the surface of a solid support, certain techniques may include a step of denaturing (1) the immobilized tagged adaptor-linked fragments 12 to produce immobilized single- stranded fragments and (2) the blocking oligonucleotides present on the first adaptors 22 and / or the second adaptors 24 to unblock hybridization sequences and complements of hybridization sequences.

[0124] However, in embodiments disclosed herein, the denaturing step may be replaced by an isothermal strand-displacing extension to separate strands of the fragments 12 having certain adaptors at the ends that permit extension from an available primer. Thus, in embodiments, a denaturing temperature is not used as a first step. Isothermal amplification / extension methods may be used with the strand-displacing Phi 29 polymerase or Bst DNA polymerase large fragment. The extension temperature may be based on a reactiontemperature range of the polymerase, e.g., 40°C to 60°C. Thus, certain steps of the workflow do not include a denaturing step. In an embodiment, the temperature of the extension reaction does not exceed 70°C.

[0125] In some embodiments, a first single-stranded fragment comprises an insert, and a second single-stranded fragment comprises an insert that is the complement of the insert comprised in the first fragment. In some embodiments, a first single-stranded fragment comprises an insert, and a second fragment comprises an insert that is not the complement of the insert comprised in the first fragment. In some embodiments, hybridizing occurs between single-stranded fragments prepared from adaptor-linked fragments 12 comprising a first forked adaptor 22 ligated at one end of each fragment and a second forked adaptor 24 ligated at the other end of each fragment. In some embodiments, two immobilized single-stranded fragments do not hybridize to each other to form a bridge in the absence of binding of a hybridization sequence 30 in a first fragment to the complement of a hybridization sequence 32 in a second fragment. In some embodiments, hybridizing two immobilized single-stranded fragments to each other to form a bridge does not occur between single-stranded fragments prepared from double-stranded fragments comprising the same forked adaptor type (e.g., two adaptors 22 or two adaptors 24) ligated at both ends of each fragment.

[0126] In some embodiments, the surface of the solid support is washed after the denaturing, and the blocking oligonucleotides will be removed by the wash, while the singlestranded fragments remain immobilized due to the interaction between the 5’ affinity moiety on the fragments with the binding moiety of the surface of the solid support. In some embodiments, the immobilizing of double-stranded or single-stranded fragments is by binding of an affinity moiety from the first and / or second forked adaptor to one or more binding moieties on the surface of the solid support. In some embodiments, the affinity moiety is biotin, desthiobiotin, or dual biotin and the binding moiety is avidin or streptavidin.

[0127] Since the single-stranded fragments are prepared from adaptor-linked fragments 12 that were already immobilized on a single surface on a solid support, complementary singlestranded fragments from a double-stranded fragment are likely to be in close proximity. Thedenaturing of the blocking oligonucleotides means that the hybridization sequence and its complement are now available to bind each other. Next, the method comprises hybridizing two immobilized single-stranded fragments to each other to form a bridge by binding of the hybridization sequence in a first fragment to the complement of a hybridization sequence in a second fragment and extending from the 3’ ends of both single-stranded fragments to produce a double-stranded concatenated nucleic acid sequencing template wherein each strand of the template comprises inserts (or their complements) from both immobilized single-stranded fragments.

[0128] In some embodiments, a single- stranded fragment prepared from adaptor-linked fragments 12 with a first strand of a first forked adaptor at a first end and the second strand of a second forked adaptor can bind to another single-stranded fragment prepared from a doublestranded fragment ligated with a first strand of a first forked adaptor at a first end and the second strand of a second forked adaptor by association of the hybridization (HYB) sequence 30 in a first fragment to the complement (HYB’) of the hybridization sequence 32 in a second fragment. In some embodiments, one or more additional rounds of denaturing, hybridizing, and extending are performed. In this way, the method can proceed in making sequencing templates until single-stranded fragments do not have appropriate other single-stranded fragments with which to form bridges (and concatenated sequencing templates) via HYB / HYB’ binding. After elongation, a concatenated sequencing template can comprise two inserts that are copies of each other.

[0129] Embodiments of the disclosure include concatenated sequencing template techniques that have reduced template reannealing to improve the efficiency of generating desired end products. In embodiments, an adaptor includes a primer binding site from which a strand-displacing extension can occur. If this occurs on only one side of a desired asymmetrically-tagged fragment, a partially single- stranded product can be generated that is a concatenated template in which one portion of the insert sequence is single-stranded and the complementary sequence of the insert sequence is double-stranded. Thus, this product has different sensitivity to certain methylation analysis techniques in which active enzymesconvert unmodified cytosines to uracil in double-stranded DNA. The single-stranded insert sequence is not affected by these enzymes, while the double-stranded insert sequence is acted on by the enzyme. Thus, a single molecule can act as its own reference sequence for methylation analysis using the unaffected portion of the molecule resistant to enzyme conversion.

[0130] FIG. 2 shows the structure of the first adaptor 22 and the second adaptor 24, a pair of forked adaptors that may be used to prepare sequencing templates. The first adaptor 22 is a complex or assembly that includes, in an embodiment, three oligonucleotides. The first oligonucleotide has a single-stranded 5’ region and a double-stranded 3’ region. In some embodiments, the first oligonucleotide includes an adaptor sequence, such as a sequencing primer sequence. In the illustrated example, the adaptor sequence is a first read sequencing adaptor sequence (P5.Rl).The second oligonucleotide includes a single-stranded 3’ region and a double-stranded 5’ region. The first oligonucleotide and the second oligonucleotide are partially double-stranded. The single-stranded 3’ region of the second oligonucleotide of the first adaptor 22 includes a hybridization sequence 30, indicated as X’, which anneals to its complement 32, indicated as X in FIG. 2. In addition, the single-stranded 3’ region 51 includes an extension binder sequence 50, indicated as Y’, to which the extendable third oligonucleotide, indicated as Y, binds. The extension binder sequence 50 and its complement are not present on the second adaptor 24. Thus, the extension binder sequence 50 is unique to the first adaptor 22.

[0131] The second adaptor 24 may also be a complex or assembly having three oligonucleotides. The first oligonucleotide of the second adaptor 24 includes an adaptor sequence, such as a sequencing primer sequence. In the illustrated example, the adaptor sequence is a second read sequencing adaptor sequence (P7.R2). The adaptor sequence of the first adaptor 22 is different than the adaptor sequence of the second adaptor 24 in an embodiment. The first oligonucleotide 24 has a single-stranded 5’ region and a doublestranded 3’ region. The second oligonucleotide 24 includes a single-stranded 3’ region and a double-stranded 5’ region. The first oligonucleotide and the second oligonucleotide arepartially double-stranded. The single-stranded 3’ region of the second oligonucleotide of the second adaptor 24 includes a hybridization sequence 30 or its complement, indicated as X. If the hybridization sequence X is present on the first adaptor 22, then the complement X’ is present on the second adaptor 24, and vice versa. Thus, the first adaptor 22 and the second adaptor 24 are configured to partially hybridize to one another in the appropriate conditions via their respective X / X’ sequences. The hybridization between complementary portions of the first and second adaptors 22, 24 does not fully occupy the second oligonucleotide forked region of the first adaptor 22, indicated as reference number 51, leaving the extension binder sequence 50 (Y’) unoccupied and available binding by a primer. Thus, the 3’ forked region 53 of the strand of the second adaptor 24 may be shorter than the 3’ forked region of the first adaptor 22.

[0132] In order to block a hybridization sequence (X) and its complement (X’) from binding to each other at undesired times, a blocking third oligonucleotide of the second adaptor 24, indicated as X’B’, can be employed in some implementations. In some embodiments, blocking oligonucleotides comprise one or more modifications such that they are not targets of tagmentation. In other words, the blocking oligonucleotides may be designed to be resistant to transposases and thus avoid cleavage of the double-stranded nucleic acid formed by hybridization of a blocking oligonucleotide to a hybridization sequence or its complement. In some embodiments, a blocking oligonucleotide comprises a phosphorothioate backbone. In some embodiments, a blocking oligonucleotide comprises the complement of all or part of the sequence one wants to block from hybridizing. Thus, in some embodiments, a blocking oligonucleotide may be all or part of an X or X’ sequence. A blocking oligonucleotide may refer to an oligonucleotide that can be used to inhibit binding of two sequences to each other, until the blocking oligonucleotide bound to at least one of the two sequences is removed, e g., via denaturation. In some embodiments, a blocking oligonucleotide comprises a sequence that is fully or partially complementary to all or part of either the hybridization sequence or its complement. For example, a blocking oligonucleotide (X’B’) to block a HYB sequence (X in FIG. 2) may comprise all or part of a HYB’ sequence, and a blocking oligonucleotide (XB) to block a HYB’ sequence (X’ in FIG. 2) may comprise all or part of a HYB sequence. In thecase of the forked adaptors shown in FIG. 2, one or more blocking oligonucleotide can serve to block binding of a X sequence in one forked adaptor to a X’ sequence in the other forked adaptor. The blocking oligonucleotide may be fully or partially complementary to either an X or an X’ sequence. In some embodiments, the blocking oligonucleotide binds to the full X or X’ sequence. In some embodiments, the blocking oligonucleotide binds to a portion of the X or X’ sequence.

[0133] One or both forked adaptors may also comprise an affinity moiety on the 5’ end of the first strand of the forked adaptor. In some embodiments, such as that shown in FIG. 2, both the first strand or first oligonucleotide of the first forked adaptor 22 and the first strand or first oligonucleotide of the second forked adaptor 24 comprise an affinity moiety at the 5’ end of the strand. In some embodiments, the affinity moiety is biotin, desthiobiotin, or dual biotin. In some embodiments, the affinity moiety is a biotin (i.e., the first strand of one or both forked adaptors are biotinylated). In some embodiments, the affinity moiety binds to a binding moiety on a surface of a solid support, such as a planar surface, a shaped surface, a well, or a bead. In some embodiments, the binding moiety is avidin or streptavidin, which binds to an avidin or streptavidin on the surface of a solid support. A range of affinity moieties that can bind to binding moieties are known to those skilled in the art, and a user may choose any pair of an affinity / binding moiety of their choice. In some embodiments, the binding moiety serves to immobilize tagged fragments (prepared by ligation of forked adaptors to fragments) on a solid support. In some embodiments, single-stranded fragments ligated to at least one first strand of a forked adaptor will be immobilized on the solid support. In some embodiments, immobilized fragments can be washed and blocking oligonucleotides can be removed, without the fragments being released from the surface of the solid support. In some embodiments, the affinity element is connected via a linker attached to the first oligonucleotide. In some embodiments, this linker is a cleavable linker.

[0134] FIG. 3 different forked adaptors embodiments. A blocking oligonucleotide may be bound to the second strand of the second forked adaptor 24 (adaptor set 100). Alternatively, no blocking oligonucleotide may be used (adaptor set 102). However, in both embodiments,the extension binding sequence Y’ is present on the first adaptor 22, and the third oligonucleotide Y is bound to the first adaptor 22 in an initial configuration. It should be understood that the extension binding sequence may be sequence Y and third oligonucleotide may be sequence Y’. Further, in certain embodiments, the extension binding sequence may be present on the second adaptor 24 and not on the first adaptor 22. In practice, the extension binding sequence may be present on only one adaptor type in a mixed adaptor type composition. Where present, the extension binding sequence is positioned 5’ of the HYB / HBY’ sequence and in a forked portion of the adaptor. In an embodiment, the extension binding sequence is positioned on a strand that does not include the adaptor sequence. In an embodiment, the extension binding sequence is 4-50 nucleotides in length. The length and composition of the extension binding sequence and its complementary third oligonucleotide may be selected based on a desired melting temperature. In an embodiment, the melting temperature of the third oligonucleotide from the extension binding sequence is about a same melting temperature (e g., within 5 °C) as the blocking oligonucleotide for a particular adaptor set, e.g., adaptor set 100, where both are present.

[0135] In some embodiments, a forked adaptor is comprised in a mixture with another nonidentical forked adaptor to form an adaptor set. In some embodiments, a mixture comprises a first forked adaptor 22 and a second forked adaptor 24 that are different. In some embodiments, a composition or kit comprises two forked adaptors, wherein (a) the first forked adaptor 22 comprises a first strand comprising a first read sequencing primer sequence and a second strand comprising a complement of a hybridization sequence and (b) the second forked adaptor comprises a first strand comprising a second read sequencing primer sequence and a second strand comprising a hybridization sequence. In some embodiments, one or both forked adaptors 22, 24 comprise a blocking oligonucleotide.

[0136] FIG. 4 shows a group 102 different product types formed using an adaptor set 100 as illustrated in FIG. 3. The first forked adaptor 22 and the second forked adaptor 24 are combined with double-stranded nucleic acid fragments under conditions that coupled the adaptors to the insert fragment 110 ends. The predicted ratio would be a set 104 of about 50%of fragments would be tagged with a first forked adaptor 22 at one end and a second forked adaptor 24 at a second end, a set 106 of about 25% of fragments would be tagged with a first forked adaptor 22 at both ends, and a set 108 of about 25% of fragments would be tagged with a second forked adaptor 24 at both ends. A pool formed from the sets 104, 106, 108 can be immobilized on a substrate or solid support (see FIG. 1). FIG. 5 shows the different product types of FIG. 4 coupled to an example solid support via 5’ affinity molecules.

[0137] The resultant products from a reaction with strand-displacing polymerase and dNTPs added to the immobilized products of FIG. 5 under conditions to permit extension from the 3’ end of the bound oligonucleotide Y, which acts as a primer, are shown in FIG. 6 In the top portion, the target nucleic acid insert 110 (strand A / A’) is flanked by asymmetric adaptors 22, 24. The strand that includes the extension binding sequence Y’ to which the third oligonucleotide Y is bound is a template for extension from the oligonucleotide Y. Extension from the Primer Y displaces the complementary strand. In an embodiment, the blocking oligonucleotide X’B’ may be removed during or after extension from the primer Y using a temperature at which the longer extended strand is not denatured but that removes the blocking oligonucleotide.

[0138] With templates that have a first adaptor 22 at both ends, extension from the primer Y generates two double-stranded molecules with a partially single-stranded sequence X’ (after removal of the blocking oligonucleotide X’B’, where present). These double-stranded molecules are not complementary. Templates with the second adaptor 24 at both ends, which does not include a binding site for the primer Y, do not extend. In the illustrated example, the melting temperature of the blocking oligonucleotide does not denature the templates with the second adaptor 24 at both ends.

[0139] FIG. 7 shows formation of a tandem repeat or concatenated insert intermediate product 120 via hybridization of the available X / X’ complements for single-stranded Strand A and the mostly double-stranded Strand A7A template copy product to generate a partially double-stranded and partially single-stranded. The intermediate product has a first 3’ end 122 and a second 3’ end 124 that are part of a double-stranded region formed hybridizationsequence 130 and complement 132 annealing. In FIG. 8, addition of pol (polymerase) facilitates extension from the 3’ ends of both strands to generate a double-stranded product 136 suitable for tandem insert analysis. The 3’ ends includes the HYB’ or hybridization sequence complement 132 to the hybridization sequence 130. The Strand A temporary copy is removed via displacement. The generated double-stranded product shown in FIG. 8 may be more efficiently generated as a result of reduced reannealing of the fragment strands, which competes with polymerase extension. By using the primer Y to displace the original template copy, template reannealing is reduced.

[0140] In an embodiment, the adaptor that includes the extension binding sequence (e.g., the first adaptor 22) has a 3’ blocked forked end 124 that cannot be extended. As shown in FIG. 9A, blocking the 3’ forked end 124, e.g., via a terminal base (ddNTP) or other 3’ blocking moiety (inverted dT), prevents polymerase extension from a 3’ end that includes the HYB’ or hybridization sequence complement 132 to the hybridization sequence 130. These sequences anneal to form a double-stranded region with one blocked end and one available end. Extension from the free end of the hybridization sequence (shown as X) displaces the Strand A temporary copy, Blocked extension from the blocked end of the X’ or hybridization complement 132 prevents the available Strand A template from being copied. The resultant product 142 is partially double-stranded and partially single-stranded.

[0141] The product 142 may be used in methylation analysis, whereby the singled-stranded insert region 144 is vulnerable to enzyme conversion of unmethylated bases, while the doublestranded insert region 146 is not. Thus, within a single strand, there are copies of a same insert sequence, e.g., Strand A, such that the product 142 is a self-reference for methylation analysis. Certain embodiments replace bisulfite chemistry with treatment by TET 5-methylcytosine oxidase followed by apolipoprotein B mRNA editing enzyme, catalytic polypeptide like (APOBEC), a variant of the human cytosine deaminase. TET oxidizes 5-methylcytosine (5mC) and 5-hydroxymethylcytosine (5hmC) to 5- carboxylcytosine (5caC) while APOBEC deaminates unmodified cytosine, 5-methylcytosine, and 5- hydroxymethylcytosine to uracil. The 5mC and 5hmC converted to 5caC by TET are protectedfrom deamination by APOBEC and read as cytosine during sequencing while unmodified cytosine is deaminated by APOBEC and read as thymine during sequencing. Enzymatic deamination ( NEBNext Enzymatic Methyl-seq Kit (EM-seq™) may use APOBEC. By using these two enzymatic steps in combination, the modified bases 5mC and 5hmC can be detected without the harsh chemical treatment. APOBEC enzymes and treatment protocols may be as generally discussed in WO2023175037A2, which is incorporated by reference in its entirety herein. FIG. 9B shows an example adaptor set with a modified first adaptor 22 having the blocked 3’ forked end. As disclosed herein, a methylated cytosine may refer to one or more of 5-methylcytosine or 5-hydroxymethylcytosine.[00142J FIGS. 10A-10G show an adaptor and workflow arrangement that can avoid a kinetic trap that can potentially reduce the efficiency with which tandem inserts are created. In essence, it allows for a primer extension reaction with a strand displacing polymerase to occur that converts an individual template into two double stranded templates that have compatible overhanging ends. These ends can then hybridise and then two options are possible to create the tandem insert, extension or ligation.

[0143] An example adaptor set including first and second adaptors 22, 24 is shown in FIG. 10 A. The adaptors are as generally shown in FIG. 2. However, the second adaptor 24 includes an extension binding sequence 50. That is the extension binding sequence 50 is present on both first and second adaptors 22, 24. In this manner, the primer Y (e.g., denoted as the third oligo) can be used to initiate extension on both adaptor ends. The extension binding sequence 50 is positioned between the hybridization sequence 30 and the double-stranded region in the first adaptor 22 and between the complement of a hybridization sequence 32 and the doublestranded region in the second adaptor 24. The extension binding sequence 50 is positioned on the second strands having a 3’ forked end in both adaptor types, e.g., on the single-stranded 3’ region 51 and the single-stranded 3’ region 53.

[0144] FIG. 10B shows example adaptor-linked nucleic acid fragments 102 generated from the adaptors of FIG. 10A. Again, the proportions of inserts 110 having same-adaptor ends and different adaptor ends is generally as shown in FIG. 4. However, because of the presence ofthe extension binding sequence on the second adaptor type 24, the surface-bound (e.g., beadbound) adaptor-linked nucleic acid fragments of FIG. 10B, when contacted with a primer Y (and after removal of any blocking oligonucleotides XB or X’B’), will create extension products from all three different groups of the adaptor-linked nucleic acid fragments 102. This is in contrast to the adaptor-linked nucleic acid fragments 102 illustrated in FIG. 4, in which inserts 110 having two ends with the second adaptor type 24 do not form any extension products.

[0145] FIG. 10D is a schematic illustration of effects of a strand displacing polymerase extension step on the immobilized adaptor-linked nucleic acid fragments of FIG. 10C, and FIG. 10E is a schematic illustration of effects of a hybridization step on the adaptor-linked nucleic acid fragments of FIG. 10D in which the hybridization sequence 30 and the complement of a hybridization sequence 32 in the extension product are permitted to bind to one another. As illustrated, only the extension products formed from fragments having different adaptor type ends have an available complement for binding. The other extension products in the same adaptor-type fragments have only one of the hybridization sequence 30 or the complement of a hybridization sequence 32 and, therefore, no available complement to bind. After binding, gaps exist between 5’ ends of the Y primer and the 3’ ends of the extends strands.

[0146] FIG. 10F shows an example in which the gaps in the products of FIG. 10E are eliminated by continuing extension from the 3’ ends as displacing the strands initiated from the Y primer. FIG. 10G shows an example in which the gaps in the products of FIG. 10E are filled using ligation. In this example, the 5’ ligated ends of the Y primers are phosphorylated.

[0147] An experiment was conducted using two adaptor species (FIG. 11), which were an A14 and a B15 adaptor. Unless otherwise indicated, the reagents are Illumina reagents, e.g., Illumina DNA Prep reagents. Transposomes were assembled from each adaptor. A 1 pl mix of the two transposomes plus 9pl of bead-linked transposome storage buffer, 29pl H2O, andIOJJ.1 Tagmentation Buffer 1 (5x Tagmentation Buffer) were mixed with I pl human DNA (Promega) and incubated at 41 °C for 5 minutes to undergo a tagmentation reaction. The tagmentation generated an adaptor-linked fragment pool with mixed-composition fragments having either a same adaptor type on both ends or a different adaptor type on both ends. Tagmentation was stopped by adding lOpl of Stop Tagmentation 2, then cleaned up be performing a 1.8x SPRI reaction and then resuspended in lOpl Resuspension Buffer. An extension / ligation reaction was performed to fill-in the 9 base partially single stranded region from a Tn5 tagmentation reaction using Illumina tagmentation protocols. SPRI beads were used to clean up the extension / ligation products and to perform size selection. After resuspension in lOpl Resuspension Buffer, the library was added to lOpl MyOme Beads and incubated at room temperature for Ihr on a rocker platform to enable binding to the beads. In the reaction protocol, the beads served as the solid substrate to immobilize the adaptor-linked fragments. After immolibization, the beads were washed in Tagment Wash Buffer, then resuspended in lOOpl of Patterned Amplification Mix (compatible with HiSeq) with 50mM KC1 added. The amplification / extension occurred as an isothermal reaction in a thermocycler set to 5minutes at 40°C, 5 minutes at 50°C, and 5 minutes at 60°C with a stepped up temperature to remove the Hyb2 blocking oligonucleotide. Beads were then washed in Tagment Wash Buffer before undergoing PCR using enhanced PCR mix with P5-A14 primers and P7-B15 primers to generate a final sequencing library. The PCR program used was 98°C 5 minutes followed by 12 cycles of 98°C 45 seconds, 60°C 2 minutes , 68°C 2 minutes, and then 68°C for 5 minutes and held at 4°C.

[0148] FIG. 12 present data showing cluster density achieved on the flow cell lanes as well as numerical values for the percent pass filter (PF) achieved by each sample. The libraries were sequenced on an Illumina HiSeqX platform. One library was prepared with a 24 base Hyb2 non-extendable blocking oligonucleotide (Lane 3) and one library was prepared with a shorter 15 base Hyb2 non-extendable blocking oligonucleotide (Lane 4). Lane 5 was prepared without any Hyb2 non-extendable blocking oligonucleotide. Control 1 was a dual insert library prepared by mixing two libraries together and using PCR to combine them together. Control 2 was a tandem insert library prepared by conventional techniques. The resultindicated at the tandem insert library prepared with the isothermal workflow achieved equivalent performance to the control libraries, such as Control 2, which used conventional denaturation, annealing, and extension. FIG. 13 shows data analysis of the tandem inserts within each templated focusing on a percentage same position pairs metric, which is a measure of what percentage of tandem insert templates have the same sequence in both inserts within a given tandem template. Almost 70% of templates using the 24 nucleotide Hyb2 blocker were in the expected copy or tandem insert format.

[0149] As discussed above with respect to FIG. 4, a disadvantage of asymmetric adaptors, e.g., adaptors that have different sequences relative to one another, for tandem insert (e.g., tandem repeats of insert sequences) generation is that three different ligation products are generated after ligation with an insert. Only one of these ligation products is the desired outcome (product 104) with the two different adaptors at either end of the insert that can form a tandem insert. The tandem inserts are formed after a step of denaturation and hybridization at respective complementary portion of the different adaptors. The other two products (106, 108) cannot form tandems because they bear the same adaptor at both ends and, therefore, do not have complementary sequences. The maximum theoretical yield of the desired product is therefore 50% of the total yield obligated products. Thus, half of the input DNA is wasted.

[0150] Provided herein is a tandem insert generation workflow that can be performed with a single adaptor type. This is in contrast to workflow that provide a first adaptor type and a second adaptor type to the reaction to form desired ligation products having heteroadaptors. In certain embodiments, the workflows are used in conjunction with a forked adaptor 150 having a palindromic sequence, as shown in FIG. 14A and FIG. 14B.

[0151] Use of the single adaptor generates a composition having an insert (e.g., a ligated product) that can generate tandem inserts with a 100% theoretical yield. The adaptor 150 includes a first oligonucleotide 152 and a second oligonucleotide 154 that are partially doublestranded and partially single-stranded such that the adaptor 150 includes a single-stranded forked portion and a double-stranded stem portion. The first oligonucleotide 152 includes three sections, Seq 1, Seq 2 and Seq 3. Seq 1 and Seq 2 together form a palindromic sequence. Byway of example, Seq 1 could be 3’GATACATGT (SEQ ID NO: 38) and Seq 2 3’ATC which together form the sequence 3’-GATACATGTATC-5’ (SEQ ID NO: 35). This sequence is palindromic because the complement is also 3’-GATACATGTATC-5’ (FIG. 14B). Thus, the sequence and its complement can form a double-stranded structure (SEQ ID NO: 36). Other examples of palindromic sequences are possible and the adaptor arrangements is not restricted to the single adaptor represented herein. The palindromic sequence may be at least 10 nucleotides in an embodiment. The palindromic sequence may be 10-100 nucleotides in an embodiment. It should be understood that the adaptor 150 may include additional sequences that are not part of the palindromic sequence in an embodiment.[00152J The second oligonucleotide 154 includes at least two sections: Seq 2’ and Seq 3’ that are complementary in sequence to the Seq 2 and Seq 3 sections of the first oligonucleotide 152. The second oligonucleotide 154 may also, but not necessarily, contain a third section Seq 4 that is not complementary to Seq 1 of the first oligonucleotide 152. The 5’ end of Seq 4 may, but not necessarily, contain an anchoring moiety for attaching the adaptor and its ligated products to a surface. Seq 4 may contain sequences that function as recognition binding sites for endonucleases, for example, a Type Ils restriction enzyme that cuts several bases downstream of its recognition site. Seq 3 of the first oligonucleotide 152 is complementary to Seq 3’ of the second oligonucleotide 154. Seq 3’ may, but not necessarily, contain an additional 3’ overhang, for example a ‘T’ to facilitate efficient ligation of the adaptor to an insert.

[0153] In the adaptor 150, Seq 1 is on the forked portion of the first oligonucleotide 152, and Seq 2 is on the stem portion of the first oligonucleotide 152. Thus, the palindromic sequence is partially double-stranded and partially single stranded in the adaptor 150. As discussed herein, separation of the first oligonucleotide 152 and the second oligonucleotide 154 from one another allows Seq 2 to be available for hybridization. In an embodiment, the self-complementary palindromic sequence is inactive or unable to self-hybridize while Seq 2 is double-stranded. In an embodiment, Seq 1 is longer than Seq 2, e.g., more than 50% of the length of the palindromic sequence is found in Seq 1. In an embodiment, Seq 1 is a samelength or shorter than Seq 2. In an embodiment, Seq 2 is at least three nucleotides in length. In an embodiment, Seq 2 is 2-10, 5-15, or 10-20 nucleotides in length.

[0154] FIG. 15 shows an example ligation product 155 that includes an insert 156 positioned between flanking adaptors 150. Only one type of adaptor-linked product 155 results from a ligation with the adaptor 150. Therefore, the theoretical yield from a ligation reaction is 100% adapted templates that, in turn, can participate in a tandem insert generation process. This provides an advantage of less loss of starting material to unproductive ligation products that cannot form tandem inserts.

[0155] FIG. 16A and FIG. 16B show noncomplementarity between Seq 1 sequences (SEQ ID NO: 37 and SEQ ID NO: 38) of the first oligonucleotide 152. By design, the Seq 1 sequences are not complementary and thus an advantage of this adaptor composition is that adaptor dimers cannot form by hybridisation (FIG. 16A). This mitigates the requirement to have blocker oligos to prevent adaptor dimers as is the case with two-adaptor methods of tandem insert generation. Note that this concept could also apply to the two-adaptor method such that complementarity can be mediated by palindromic sequences in which part of the palindromic sequence is located in the stem or double-stranded portion of the adaptor. In both examples, denaturation reveals a portion of the complementary sequence to permit hybridization. A non-limiting exemplary sequence is illustrated in FIG. 16B.

[0156] FIG. 17A shows an example surface-linked product 155 that is linked to a surface via 5’ ends of the second oligonucleotide 154. The surface linkage may be as generally discussed herein. In certain embodiments, the surface may be pre-loaded with surface-linked adaptors 150 that are then ligated to respective insert ends to form the product 155. In an embodiment, the product 155 is formed and subsequently associated with the surface.

[0157] During tandem insert formation using a surface, a denaturation step is applied to separate the two strands of the template product 155, e.g., heat, chemical denaturations, etc. This provides an opportunity for the two strands to anneal by their 3’ ends as shown in FIG. 17B to generate an intermediate product 157 of tandem insert formation. As illustrated, thepalindromic sequence Seq2-Seql on one strand can anneal to the Seq2-Seql sequence on the other template because they are now available for binding. Previously, when part of the fork adaptor, these sequences could not anneal. The intermediate product 157 is double-stranded at the palindromic sequence and single-stranded in other regions.

[0158] FIG. 18 illustrates an exemplary template comprising the proposed novel adaptor composition undergoing tandem-insert generation on a surface. The surface-associated template product 155 is denatured to permit binding of the palindromic sequences on the respective first oligonucleotides. After hybridization, the 3’ ends are extended using the single-stranded regions as template to generate a double-stranded tandem insert product 158.

[0159] FIG. 19 illustrates an exemplary template comprising the proposed novel adaptor composition undergoing tandem-insert generation in solution via self-circularization. One or both of the 3 ’ ends may be extended using a strand-displacing polymerase to generate a tandem insert product.

[0160] Whether formed on a surface or in solution the resulting tandem insert template can be further processed, either by directly clustering and sequencing on a flow cell or optionally first amplifying using PCR or other amplification methods (FIG. 20).

[0161] In another embodiment the resulting tandem insert template can be further processed by ‘A-tailing and ligating of another forked adaptor to the ends (FIG. 21). In doing so, asymmetry of sequence cam be conferred on the ends of the top and bottom strands. UMIs could be included as part of these forked adaptors. The templates can be further processed, either by directly clustering and sequencing on a flow cell or optionally first amplifying using PCR or other amplification methods.

[0162] In yet another embodiment as shown in FIG. 22, the original single adaptor 150 is designed with a restriction enzyme site that leaves an overhang when the template is digested. The restriction enzyme can be selected from the subset of restriction enzymes that cleave outside of their recognition sites and downstream such that an overhang is created in the Seq3 sequence. The advantage of this method is that the ligation of the additional adaptor that confers asymmetric to the template strands can proceed with high efficiency because the adaptor comprises a complementary overhang.

[0163] Certain embodiments of the disclosure may be used in conjunction with workflows for generating a sequencing template comprising multiple inserts, as generally discussed in US20230407388A1, which is incorporated by reference herein in its entirety, for methylation analysis using APOBEC treatment with a mutant APOBEC as discussed in US20240182881A1, which is incorporated by reference herein in its entirety, to convert 5- methyl-cytosine to thymine without converting cytosine to uracil. Thus, methylated cytosines can be distinguished from unmethylated cytosines in the reaction products. APOBEC operates on single-stranded DNA. In an embodiment, the cytosine deaminase comprises an altered cytosine deaminase.

[0164] FIG. 23 shows an embodiment similar to that of FIG. 9A in which a 3’ end 161 of a first adaptor 160 is blocked (e.g., inverted dT, terminal ddNTP) and is not extendible via a polymerase in the blocked state. The block may or may not be removable. In contrast to the adaptors 22, 24 shown in FIG. 9A, the first adaptor 160 and a second adaptor 162 do not include any extension binding sequence that serves to initiate strand displacement. However, the first adaptor 160 and the second adaptor 162 include respective hybridization sequences that are complementary to one another, shown as X / X’ that are hybridized to blocking oligonucleotides at initial stages of the workflow. One or both forked adapters 160, 162 may also comprise an affinity moiety on the 5’ end of the first strand of the forked adapter.

[0165] As shown, a reaction mix includes target nucleic acid fragments 164 and the mixed adaptors 160, 162 yields When a mixture comprising a first forked adapter and a second forked adapter is combined double-stranded nucleic acid fragments under conditions for ligating / tagmentation, the predicted ratio of generated adaptor-linked fragments would be 50% of fragments would be tagged with the first forked adapter 160 at one end and the second forked adapter 162 at a second end, 25% of fragments would be tagged with a first forkedadapter 160 at both ends, and 25% of fragments would be tagged with a second forked adapter at both ends 162.

[0166] The affinity moieties of the adaptor-linked fragments are having are captured on a substrate as shown in FIG. 24. Because each strand can be coupled to an affinity moiety, denaturing the products yields captured single strands that are capable of hybridizing to one another via complementary portions X / X’ of the forked adaptors to initiate an extension. When linked to a substrate, the reaction products with mixed adaptors 160, 162 are capable of extension from a 3’ end of the hybridized complementary region as generally discussed with respect to FIG. 1, while the symmetrically tagged products having same adaptors at both ends are not capable of extension due to lack of hybridization. An extension reaction occurs using a polymerase and appropriate reagents (e.g., dNTPs) under extension conditions, such as a lower temperature suitable for polymerase activity (e.g., 37°C-65°C) to extend from the available 3’ end. The process of denaturation, reannealing, and extension can be repeated to ensure that up to 100% of templates are converted to tandem inserts

[0167] As shown in FIG. 25, using the adaptors of FIG. 23, there is only one available 3’ end for extension, because the first adaptor 160 has a blocked 3’ end 161. Thus, extension only occurs in one direction such that an extension product 170 has a single-stranded portion 172 including Strand A of the fragment 158 and a double-stranded portion 174 having both Strand A and Strand A’ of the double-stranded fragment 158. In this manner, mutant APOBEC can operate on the single-stranded portion 172 to selectively convert hydroxymethy-C and methyl-C to T, and not convert unmethylated C to T, while any hydroxymethy-C and methyl- C as well as unmethylated C in the double-stranded portion 174 is unaffected.

[0168] FIG. 26 shows an example sequence of the double-stranded fragment 158 and the corresponding sequences of the extension product 170. FIG. 27 shows conversion of 5mC in the extension product 170 to T only in the single-stranded portion 172 and not in the doublestranded portion 174. The converted template can be liberated from the surface by PCR, as shown in FIG. 28, or linker cleavage. When the tandem insert is subject to sequencing, each base in the original template give a two-base-call when both inserts are sequenced andcombined. In this embodiment, unmethylated Cs in the original template report as a (C,C) two-base call, and methylated and hydroxy-methylated Cs in the original template report as a (T,C) two-base call. This allows unmethylated and methylated / hydroxy-methylated Cs to be distinguished from one another in the original template based on the generated base call data.

[0169] In another embodiment, as shown in FIG. 29, a target nucleic acid, e.g., as part of an adaptor-linked fragment, can be treated with P-glucosyltransferase, which converts hydroxy-methyl C to beta-glucosyl-5-hydroxymethylcytosine. This is then subject to conversion to an extension product as discussed in FIGS. 23-25. The extension product, already subjected to P-glucosyltransferase treatment, can then be treated with wild type APOBEC, in contrast to selective mutant APOBEC as discussed in FIGS. 25-28.

[0170] Wild-type APOBEC converts both unmethylated Cs and methylated Cs to T. In this embodiment, because the hydroxy-methyl Cs have been converted to beta-glucosyl-5- hydroxymethylcytosine, these based are no longer converted to T by APOBEC. However, methylated and unmethylated Cs are still converted to T, as shown in FIG. 30.

[0171] The converted template can be liberated from the surface by PCR, as shown in FIG. 31, or linker cleavage. When the tandem insert is subject to sequencing, each base in the original template give a two-base-call when both inserts are sequenced and combined. In this embodiment, methylated and unmethylated Cs in the original template report as a (T,C) two- base call, and hydroxy-methylated Cs in the original template report as a (C,C) two-base call. This allows hydroxy-methylated Cs to be distinguished from methylated / unmethylated Cs in the original template based on the generated base call data.

[0172] FIGS. 32-42 show embodiments of tandem repeat nucleic acid preparation using modified adaptors and that can be conducted in solution. As discussed herein, a mixed or asymmetric adaptor set results in conversion efficiencies of less than 100%, because desired end products are not generated from fragments with same-type adaptors on both ends. For adaptor sets with two different adaptors, approximately 50% of intermediate products do notyield desired end products based on formation of A-A (first adaptor type) or B-B (second adaptor type) fragments, as illustrated in FIG. 4.

[0173] FIG. 32 shows an embodiment of an adaptor-linked fragment 200 prepared using symmetric adaptors 202 that flank an insert 204 (shown having a top strand 206 and complementary bottom strand 208 of the insert 204). That is, the adaptor 202 may be a same sequence or substantially identical (e.g., more than 95% identical) on both ends of the insert 204 such that a single adaptor type can be provided to the reaction mixture. The adaptor- linked fragment 200 may be prepared as generally discussed herein, e.g., via tagmentation or ligation (e.g., fragmenting the target nucleic acid, A-tailing, and adaptor ligation of the adaptors 202 to both ends).

[0174] The adaptor 202 includes a double-stranded stem region 209 that is proximate the insert 204, and a forked region 21 1. On a first strand of the forked region, first and second primer regions 210, 212 are separated by a 5’-5’ linker214. The 5’-5’ linker may be a synthetic composition (integrated DNA Technologies) having the first primer region 210 appended. The second strand also includes a primer complement region 218 that is a complement of the first primer region 210. The first strand fork is longer than the second strand fork based on a presence of two primer regions and the linker 214. The existing 5’ end of forked region 211 is modified using the linker to convert the 5’ end to a 3’ end 224. This arrangement yields 3’ ends on both strands of the forked region 211. However, as illustrated, a blocked 3’ end 220 is present on a second strand of the forked region 211. Thus, the available 3’ end for extension is a first strand 3’ end 224. It should be understood that the adaptor 200 may include any suitable adaptor sequences for compatibility with a desired sequencing platform. In addition, the adaptors 200 may include sample-specific index sequences in the stem 209, or on the first or second strands of the forked region 211. Because certain embodiments of the method are PCR-free, any incorporation of indexes or barcodes may be conducted prior to generating the adaptor-linked fragments 200.

[0175] As illustrated in FIG. 33, the primer region 210 and the primer complement region 218 of an individual adaptor 202 may anneal in an opposite strand annealing arrangement.However, a same strand annealing arrangement is also possible FIG. 34 is a schematic illustration of 3’ end extension of the adaptor-linked fragment 200 that may occur as a samestrand or opposite-strand reaction. In the same strand reaction, the primer region 210 of the top strand 206 anneals to the primer complement region 218 of the top strand 206, and the primer region 210 of the bottom strand 208 anneals to the primer complement region 218 of the bottom strand 208. In the opposite strand reaction, the primer region 210 of the top strand 206 anneals to the primer complement region 218 of the bottom strand 208, and the primer region 210 of the bottom strand 208 anneals to the primer complement region 218 of the top strand 206.[00176J FIG. 35 shows a process of generating extension products of the same strand reactions. The extension products are generated via annealing and polymerase extension in a PCR free arrangement. As discussed, the extension products may be generated in solution without immobilization to a surface. To promote intra-strand annealing, the adaptor-linked fragments 200 can be subjected to a denaturing step to separate top and bottom strands 206, 208 in conjunction with low concentrations of the adaptor-linked fragments 200 such that each individual strand is more likely to anneal to itself at complementary sequences rather than encountering a complementary sequence of a different strand.

[0177] The denaturing step may occur a denaturing temperature (e.g., greater than 80°C) such that the adaptor-linked fragment 200 is separated into component strands. After denaturing, an extension reaction occurs using a polymerase and appropriate reagents (e.g., dNTPs) under extension conditions, such as a lower temperature suitable for polymerase activity (e.g., 37°C-65°C) to extend from an available 3’ end. Each of the top strand 206 and the bottom strand 208 has one blocked 3’ end 220 and one extendable 3’ end 224. Thus, on each strand, extension can only occur from the 3’ end 224 of the first primer region 210. When the first primer region 210 is annealed to the primer complement region 218, the top strand 206 and the bottom strand 208 are self-templates, and polymerase extension continues to the linker 214. Extension using each of the top strand 206 or the bottom strand 208 as template yields a double-stranded extension product.

[0178] FIG. 36 shows extension products 250 of the reaction of FIG. 35 from a single strand, shown as the top strand 206 by way of example. It should be understood that the extension products also include those from the bottom strand 208, which generates complementary extension products. In the illustrated example, the extension products include an original strand insert sequence 252, indicated as A, which is one strand of the insert 204. In addition, as a result of 3’ extension, also present is a polymerase-generated complement sequence or template copy 254, indicated as A’. Because the original insert sequence 252 is directly obtained from fragmenting target nucleic acid, any nucleic acid modifications from natural nucleotides such as methylation will be present. Such modifications are not present in the complement sequence 254. Thus, within one single-stranded extension product is an original nucleic acid sequence and its complement as tandem inserts. The complement sequence 254 can serve as a reference or baseline to the original strand insert sequence 252.

[0179] As illustrated, enzyme treatment after generating the extension products 250, which operates on single-stranded nucleic acids, converts methyl-cytosine to thymine in the original insert sequence 252. In the complement sequence, the complement of the methyl-cytosine (mC) remains guanine at the corresponding position. Thus, embodiments of the present disclosure include treatment with a base modifying agent such as sodium bisulfite or a cytidine deaminase (e.g., APOBEC) and methylation sequence analysis to identify discrepancies between an expected complement sequence 254 and the original strand insert sequence 252. If a thymine is present in the original strand insert sequence 252 at a location corresponding to a guanine in the complement sequence 254, that position can be marked as likely being a methyl-cytosine.

[0180] FIG. 37 shows downstream PCR-free sequencing of the extension products 250 using Illumina flow cells having capture molecules that capture the first primer region 210, the second primer region 212, or their complements. Each individual extension product 250 is captured at both ends, and the extension product 250 is linked at the seed / capture stage. However, extension after capture forms a cluster with no links between copies. Under a9QAM sequencing readout configuration, Read 1 (e.g. a P5 primer) and Read 2 (e.g., a P7 primer) can be performed simultaneously.

[0181] FIG. 38 shows an opposite strand approach to generate extension products. The extension products are generated via opposite-strand annealing across an individual adaptor 200 and strand-displacing polymerase extension in a PCR free arrangement. As discussed, the extension products may be generated in solution without immobilization to a surface. The opposite-strand approach does not include a denaturing step, and accompanying temperature change, but instead relies on strand displacement. Thus, the opposite-strand approach may be isothermal.

[0182] Extension occurs from the 3’ end 224 of the first primer region 210 to displace the top strand 206 or the bottom strand 208, and extension continues to the linker 214.

[0183] FIG. 39 shows a single-stranded extension products 250 of the reaction of FIG. 29 from an opposite-strand approach, by way of example. It should be understood that the extension products also include those generated from both end adaptor opposite-strand binding to yield two different strands. In the illustrated example, the extension products include an original strand insert sequence 252, indicated as A, which is one strand of the insert 204 and corresponds to the top strand 206. In addition, as a result of 3’ extension, also present is a polymerase-generated complement sequence 260 of the bottom strand 208. As a result, the extension product shown includes two copies of A in reverse orientations to one another. Because the original insert sequence 252 is directly obtained from fragmenting target nucleic acid, any nucleic acid modifications from natural nucleotides such as methylation will be present. Such modifications are not present in the complement sequence 260. Thus, within one single-stranded extension product 250 is an original nucleic acid sequence and a same sequence, in reverse orientation. The complement sequence 260 can serve as a reference or baseline to the original strand insert sequence 252.

[0184] FIG. 40 shows downstream PCR-free sequencing of the extension products 250 using Illumina flow cells having capture molecules that capture the first primer region 210,the second primer region 212, or their complements. Each individual extension product 250 is captured at both ends, and the extension product 250 is linked at the seed / capture stage. However, extension after capture forms a cluster with no links between copies. Under a 16QAM sequencing readout configuration, Read 1 (e.g. a P5 primer) and Read 2 (e.g., a P7 primer) can be performed simultaneously.

[0185] FIG. 41 shows an example adaptor sequence using methylated bases in the stem 209 to permit 16QAM sequencing and associated extension products 250 before and after methylcytosine conversion to thymine. In the FIG. 42 sequencing substrate, the captured strand clusters can be read using two Read 1 primers and two Read 2 primers. The stem regions from the original strand have concerted methyl cytosines in the Read 1 or Read 2 primer region and, therefore, have a different binding sequence relative to the copied or extended stem regions.

[0186] Certain dual insert techniques may include preparation of adaptor-linked fragments 300 having a forked region 302 at one end and a loop 304 at another, as shown by way of example in the workflow of FIG. 43. In this manner, the adaptor-linked fragment 300 can be sequenced using a workflow that includes a nicking arrangement to facilitate sequencing of the entire inverted-repeat tandem-insert duplex. In the illustrated workflow, the fragment 300 is immobilized on a substrate via binding of substrate-linked lawn primers to a complementary sequence in the forked region. Extension and amplification steps yield a double-stranded bridge structure in which one strand has a sequence of the fragment 300 and in which complementary insert sequences A and A’ are one strand of the bridge, separated by the sequence of the loop region 304. The other strand includes copies of the insert sequences A and A’ as well. Nicking enzymes, specific for an alternative recognition site, are added to nick a recognition site within the loop sequence 304 to generate two start sites for simultaneous sequencing of the other strand of the original polynucleotide duplex. Following nicking and sequencing of the first strand (read 1), the free ends of the sequenced strands are blocked.

[0187] In the illustrated workflow, the technique for the Read 2 step may involve a doublestranded sequence by synthesis (SBS) sequence for simultaneous read (9QAM, with 9 detectable states, or 16QAM, with 16 detectable states using signal intensity) to add bases inthe extension against the strand not attached to the flowcell which is complex. In addition, acquiring index reads from dual indexes (e.g., i5 and i7) is complex and inefficient due to undesired hairpin formation of complementary insert strands. In the illustrated example, although two indexes, i5 and i7 are present, the incorporation of both is via a single adaptor in the forked region 302, Thus, this is not a true dual-indexing workflow. The illustrated nicking (BSPQ1 single-stranded and extendable cut site) may have a level of undesired background cutting that creates incorrect read start locations.

[0188] Provided herein are library preparation techniques that enable 9QAM or 16QAM sequencing for both reads with conventional clustering and sequencing workflows including standard indexing. The disclosed techniques also enable these reads to be performed with or without QAM based on the use of different primer sequences, depending on the desired sequencing encoding scheme. The disclosed techniques may involve generating sequencing libraries using a first adaptor 310 and a second adaptor 12, as shown in FIG. 44. The first adaptor 310 may be formed from annealing the oligonucleotides SEQ ID NO:32 and SEQ ID NO:33 by way of example. The second adaptor 312 may be formed by forming a hairpin loop (e.g. a contiguous loop structure with unpaired or non-Watson-Crick paired nucleotides) from the oligonucleotide SEQ ID NO:34. However, these sequences are shown by way of example. Generally, the first adaptor 310 may include double-stranded regions 320, 322 that flank a noncomplementary region 324. The first double-stranded region 320 includes a sequence that includes a capture region (e.g., P7) and its complement (e.g., P7’) to permit flow cell capture and an index sequence (e.g., i5) and it’s complement (e.g., i5’). The noncomplementary region 324 is shown as a bubble or partially-single-stranded region, which includes an A14 primer sequence and a Bl 5’ primer complement. The second double-stranded region 322 includes, for example, a mosaic end sequence (ME) and its complement (ME’).

[0189] The second adaptor 312 includes double-stranded regions 330, 332 that flank a noncomplementary region 334. A loop 336 terminates one end of the second adaptor 312 such that the first double-stranded region 330 is positioned between the loop 336 and the noncomplementary region 334. The first double- stranded region 330 includes a sequence thatincludes a capture region (e.g., P5) and its complement (e.g., P5’) to permit flow cell capture and an index sequence (e.g., i7) and it’s complement (e.g., i7’). The noncomplementary region 334 is shown as a bubble or partially-single-stranded region, which includes an A14 primer sequence and a Bl 5’ primer complement. The second double-stranded region 332 includes, for example, a mosaic end sequence (ME) and its complement (ME’).

[0190] The second double-stranded regions 322, 332 of the first and second adaptors 310, 312 includes a 3’ T overhang. To prepare the library, fragments 338 of target nucleic acid are repaired (if necessary) and A-tailed to permit asymmetric ligation of the first and second adaptors 310, 312 to generate an adaptor-linked fragment 340, as shown in FIG. 45. The double-stranded fragments 338 have positive X and negative X’ strands of the insert sequence. While an individual adaptor-linked fragment 340 is illustrated, it should be understood that the adaptor sets disclosed in the embodiments herein (e.g., the first and second adaptors 310, 312) can be used to generate a sequencing library of adaptor-linked fragments from fragments of a target nucleic acid. Thus, a universal or conserved set of adaptors can yield a plurality of adaptor-linked fragments (e.g., adaptor-linked fragment 340), each having a same first adaptor at one end and a same second adaptor at the other end. However, the adaptor-linked fragments may have different insert or target sequences. The various sequencing workflows disclosed herein use sequencing primers and flow cells that target conserved sequences in the adaptors.

[0191] FIG. 46 shows a workflow using the adaptor-linked fragment 340 of FIG. 45. The flowcell or substrate is pre-linearized at the uracil on the P5 end 350 prior to the library hydridization. This leaves a 3’ OH group on the P5 end 350 which cannot be extended until it is phosphorylated again during the resynthesis step of the paired end turn. Thus, P5 extension is blocked until read 1 is complete. The unwound adaptor-linked fragment 340 binds to P7 lawn primers on the substrate via the p7’ sequence at the 3’ end of the fragment 340. Via template hybridization and exclusion amplification (ExAmp), a double-stranded bridge is formed using only the P7 primers and not the P5 primers. That is, for a substrate having two different types of lawn primers, the initial double-stranded bridge is formed using only a firstprimer type and not the second primer type. Extension from the second primer type is blocked until later stages of the workflow.

[0192] Double-stranded cutting at I-Scel (or other restriction enzyme recognition site) followed by denaturation to remove the non-immobilized strands generates a first strand 354 of the double-stranded bridge and a second strand 356 of the double-stranded bridge that both are immobilized at their 5’ ends via the P7 immobilized primer. The workflow yields two different strands 354, 356, each with a respective complement of the insert and that include the previously internal P5’ and adjacent adaptor sequences, which serve as landing sites for a read 1 primer. After deprotection via phosphorylation of the P5 end 350, resynthesis of the immobilized P7 strand can occur from the P5 lawn primers from a single-stranded P7-P5 bridge. Generation of the complementary strand and linearization (e.g., via 8 oxoG) permits a read 2 step from a read 2 primer annealed to the complementary strand. The illustrated workflow enables a 9 QAM read 1 and read 2 utilizing simultaneous extension from the standard Illumina primer reagents (for A14-ME and B15-ME primers within the Illumina insert read primer mix HP21 on Illumina sequencing platforms). Other reagents may include the standard clustering and paired end turn reagents. It also allows though, if required, the two different primers to be hydridised onto each strand separately to enable 2*4QAM reads (e.g., two separate 4QAM reads) for both read 1 and read 2. In addition, an I-Scel or other suitable restriction enzyme is included in the workflow. It should be understood that the illustrated adaptor sequences and double-stranded cutting enzyme are by way of example.

[0193] An indexing scheme for the workflow of FIG. 46 is illustrated in FIG. 47. The indexing can use standard indexing primers with the index 1 read post read 1 (pre the P5 deprotection and resynthesis step) and index 2 read pre read 2 (post 8-oxo G linearisation). In the illustrated example, the ME’-B15’ and ME’-A14’ indexing primers can be used from the standard HP 14 primer mix from Illumina.

[0194] If the potential for 2 * 4QAM reads is not required, a library preparation, as illustrated in FIG. 48, may include a first adaptor 360 with only B15 and a second adapter 362 with only A14. End ligation to double-stranded target fragments 363 would generate adaptor-linked fragments 364 with the first adaptor 360 at one end and the second adaptor 362 at the other end. Such an approach would not use A14-ME and B15-ME primers for simultaneous reads (9QAM), but would use one primer at a time for the reads (e.g., A14-ME for read 1 and B15-ME for read 2), as shown in the workflow of FIG. 49.

[0195] Certain steps of the workflow of FIG. 49 may occur as discussed with respect to FIG. 46. For example, template hybridization and exclusion amplification to form a doublestranded bridge that is cleaved to generate separated substrate-linked strands, the substrate- linked P5 blocked and subsequently deprotected at a later point in the workflow, etc. Potential bias between primer hybridisation would also be removed using a kit in which the A14-ME and B15-ME primers are separated. In such an embodiment, the indexing primers in this version would be separated so only ME’-B15’ for index 1 and only ME’-A14’ for index 2 would be used.

[0196] The index in the loop adapter 312, 362 can also be removed as it is not on the end of the strand being clustered. Thus, the index will not be swapped during exclusion amplification clustering. Accordingly, having a dual index in this workflow is not required, and only having an index on the first adaptor 310, 360 would be sufficient. This could be two different indexes on the two oligos and read out separately with two primers if required.

[0197] The disclosed 9QAM-compatible techniques are also useful for improving accuracy of reads and for read-out of methylated converted reads. Methylation information is provided from both strands in every cluster.

[0198] FIG. 50 shows an example adaptor arrangement and associated workflow that includes a first adaptor 380 having an “almost fork” in the form of a noncomplementary region 382 and a simplified second adaptor 384. The noncomplementary region 382 includes a mismatch in which a first adaptor sequence Bl 5’ is present on a top strand while a second adaptor sequence A14 is present on the top strand. Together with the adjacent ME / ME’ regions, these sequences or their complements may form an annealing sequence for a sequence primer in the illustrated workflow. The first adaptor 380 includes an index sequence andP7 / P7’ sequences. The second adaptor 384 includes the P5 / P5’ sequences as well as a BspQl enzyme nicking sequence at one end of the second adaptor 384. The other end forms a hairpin loop 386, and the P5 / P5’ sequences are between the nicking sequence and the hairpin loop 386. As illustrated, the second adaptor 384 may not include any annealing sequences for a sequencing primer or any index sequence in one example.

[0199] Certain steps of the workflow may operate as generally discussed with respect to FIG. 46. However, rather than a cleavage within the hairpin loop 386, the cleavage occurs at a different location in the sequence of the second adaptor 384 and is a single-stranded cleavage. After cleavage and denaturing, read 1 data can be acquired via extension from the cleaved ends. In embodiments, an index read can also be acquired. After deprotecting the P5 substrate- linked primers, the P7-P5 bridge formed by the P7-linked strand can be extended from P5, linearized (e.g., 8-oxoG), denatured, and sequenced using a read 2 primer.

[0200] In certain embodiments, rather than the noncomplementary region 382, the first adaptor 380 may be fully double-stranded with only a single annealing sequence for a read 2 primer and its complement to prevent variabilities due to differential primer binding (e.g., because of secondary structures in the nucleic acid).

[0201] FIG. 51 shows an alternate arrangement in which the second adaptor 384 includes two separated or spaced apart cleavage sites, with a second cleavage site within the hairpin loop 386. Sequential cuts are used to separate read 1 and read 2 steps. In particular, after formation of the double-stranded bridge structure as generally discussed herein, a BspQl digestion and denaturing yields a partially single-stranded bridge structure. Read 1 can be performed via extension from a nicked end while retaining the association between the strands. After read 1, the cleavage site within the hairpin loop is used to generate a double-stranded cleavage. The P5 ends are deprotected, and amplification of the strands to form two bridges is performed, e.g., exclusion amplification or bridge amplification An index read may occur in an embodiment. After denaturing, read 2 primer is hybridized and extended to generate read 2 9QAM or 2* 4QAM sequencing can be performed dependent on the primers used.

[0202] FIG. 52 shows an example order of workflow steps for the illustrated workflows of FIG. 51. For embodiments in which the first adaptor includes two separate annealing sequences for a sequence primer in the noncomplementary region 382, read 1 may be a 9 QAM read via end extension. Read 2 may be a 9 QAM read or 2*4QAM reads using two separate primers. As illustrated, the primers are HP10 / HP11 (Illumina). It should be understood that, in any of the disclosed embodiments, the adaptor sequences are by way of example and dependent on the desired sequencing platform and the desired adaptor design features.

[0203] A different application for dual insert libraries is simultaneous PE reads which requires 16 QAM, where there are again two signals per cluster but they can be any mix of bases. In order to distinguish them one signal is half the intensity of the other, producing 16 separate signal clouds. In certain embodiments, the disclosed techniques may permit 16 QAM reads using straightforward clustering and sequencing methods. FIG. 53 shows an example adaptor set including a first adaptor 390 and a second adaptor 392. The first adaptor 390 is a forked adaptor, and the second adaptor includes a hairpin loop 394 having a linked affinity moiety 395 (e.g., biotin) and an internal cleavage site.

[0204] An example adaptor-linked fragment 396 formed using the first adaptor 390 and the second adaptor 392 is shown in FIG. 54. FIG. 55 is an example workflow for sequencing the adaptor-linked fragment 396. Post-hybridization of the fragment 396 to the flowcell, the hybridized fragment 396 undergoes a short or limited round of exclusion amplification (e.g., two cycles of bridge amplification or enough to form a smaller number of copies of the full fragment relative to conventional clustering) to form a double-stranded bridge with the P5 and P7 ends attached to the surface. Once the cut at I-Scel has exposed the P5’ and P7’ ends from the loop, the separated portions of the bridge can undergo exclusion amplification to fully develop two separate types of clusters and occupy the well with both inserts. Post a P5 (or P7) linearization, then the strands can be sequenced using a mix of A14-ME and B15-ME. One of these primers must be at a different (e.g. 50%) concentration for a 16 QAM read. This will give a read 1 with 16 QAM signal but simultaneously reading both ends of the original insert so giving a simultaneous paired end turn. Indexing could also be done on these libraries postthe 16 QAM read 1, it could be a 16 QAM index with a mix of ME’-B15 and ME’-A14’, one at 50%, to give the i5 and i7 read outs simultaneously or the two indexes could be read in separate reads. Both these library preps are novel and give benefits for enabling 9 QAM and 16 QAM readouts for dual insert libraries while maintaining fairly standard reagents and workflows for clustering and sequencing. In an embodiment, standard sequencing workflow reaction mixes may be modified with appropriate restriction enzymes, such as the I-Scel restriction enzyme to cut at the I-Scel site. The I-Scel site was chosen because the 18 base recognition site is sufficiently large to be less likely to cut in the incorrect place, as compared to for example BSPQI. However, in all of these designs, the restriction site design to be cut can be modified to match any different enzyme / site that was determined to be preferable with experimental design.

[0205] Certain embodiments of the disclosure provide hairpin or modified adaptor sequences that have benefits for sequencing libraries, such as sequencing libraries with tandem repeats of insert sequences (e.g., tandem inserts). As discussed herein, having tandem repeat inserts can provide benefits with respect to methylation analysis, in which a single fragment of a sequencing library can act as its own reference by including both treated or converted 5- methylcytosines (5mC) that are converted to a different base and untreated 5-methylcytosines to serve as a reference sequence. The disclosed embodiments may be used in conjunction with cytidine deaminase treatment or APOBEC treatment, such as APOBEC variant with a high level of selectivity for 5mC over C. In certain embodiments, the cytidine deaminases may be as disclosed in U.S. Patent Publication No. 20240182881 Al, which is hereby incorporated by reference in its entirety.

[0206] Wild type APOBEC3A deaminates cytosine (C), 5 methyl cytosine (5 mC), and 5- hydroxymethyl cytosine (5 hmC) efficiently in single-stranded DNA. An APOBEC3 A mutant containing a tyrosine to alanine point mutation in position 130 (Y130A) was found to preferentially deaminate 5 mC instead of C (5 mC was converted to T at a greater rate than C was converted to U), and an APOBEC3A mutant containing a tyrosine to leucine point mutation in position 130 (Y130L) was found to preferentially deaminate C instead of 5 mC (Cwas converted to U at a greater rate than 5 mC was converted to T). The deamination of 5 mC to T leads to C to T mutations which can be identified by standard sequencing methods. As a result, in one embodiment the treatment of DNA with an altered cytidine deaminase of the present disclosure preferentially converts 5 mC to thymidine.

[0207] In an embodiment, the altered cytosine deaminase is a member of the AID subfamily, the APOB EC 1 subfamily, the APOBEC2 subfamily, the AP0BEC3A subfamily, the AP0BEC3B subfamily, the APOBEC3C subfamily, the AP0BEC3D subfamily, the APOBEC3F subfamily, the AP0BEC3G subfamily, the AP0BEC3G subfamily, the AP0BEC3H subfamily, or the APOBEC4 subfamily, or an alteration thereof. In some aspects, the altered cytosine deaminase comprises an altered AP0BEC3A.

[0208] In an embodiment, the altered cytidine deaminase comprises an amino acid substitution mutation at a position functionally equivalent to (Tyr / Phe)130 in a wild-type AP0BEC3A protein and / or an amino acid substitution mutation at a position functionally equivalent to Tyrl32 in a wild-type AP0BEC3A protein. In some aspects, the altered cytidine deaminase comprises amino acid substitution mutations at positions functionally equivalent to (Tyr / Phe)130 and Tyrl32 in a wild-type AP0BEC3A protein.

[0209] In an embodiment, the substitution mutation at the position functionally equivalent to Tyrl30 comprises a mutation to alanine, glycine, phenylalanine, histidine, glutamine, methionine, asparagine, lysine, valine, aspartic acid, glutamic acid, serine, cysteine, proline, arginine, or threonine.

[0210] In an embodiment, the substitution mutation at the position functionally equivalent to Tyrl30 comprises a mutation to Ala, Vai, or Trp.

[0211] In an embodiment, the substitution mutation at the position functionally equivalent to Tyrl32 comprises a mutation to His, Arg, Gin, or Lys. In some aspects, the altered cytidine deaminase comprises an amino acid substitution mutation at a positionfunctionally equivalent to (Tyr / Phe)130 in a wild-type AP0BEC3A protein, wherein the substitution mutation is (Tyr / Phe)130Trp.

[0212] With this enzyme, treatment of DNA selectively converts methylated cytosine into T - leaving any other base untouched and generating a convenient 4-base genome. However, certain cytidine deaminase operate only on single stranded DNA. It would be more desirable to be able to act on double stranded DNA instead, because maintaining the DNA in dsDNA would allow retainment of information that would be instead lost - such as the symmetry of methylation at CpG sites. To address this problem, FIGS. 47-48 show adapters specifically functionalized to allow for retaining of dsDNA information.

[0213] APOBEC may have increased selectivity for 5-methylcytosine (5mC) over cytosine (C), allowing direct sequencing of methylated regions by 5mC to thymine (T) conversion. However, APOBEC is only active on ssDNA, complicating sequencing workflows and losing important information on strandness and the state of symmetrical methylation. The problem of APOBEC single-strand DNA exclusivity is solved by using various designs of circularized adapters. The hairpin-like structures will maintain the fragments connected, while allow for the formation of single-stranded DNA bubbles where APOBEC will convert 5mC into T. The strands will then quickly reanneal to dsDNA, maintaining the strand information intact.

[0214] Stem-loops, or hairpins include complementary regions located on the same strand that, when paired, ends in an unpaired or single-stranded loop linking the complementary portions. Stem-loops are typically GC-rich and have high melting temperature, allowing for exceptional stabilization of secondary structures. This concept can be used to “staple” dsDNA libraries together, perform chemistry on the libary insert, and sequence the insert after removal of the hairpin. In the illustrated examples, the hairpin-type arrangement can be used to improve methylation workflows requiring ssDNA, such as demethylation from APOBEC. In one iteration, a library ligated with hairpins (referred from here on as “stapled library”) would be treated with an agent causing formation of transient ssDNA within the stapled library. APOBEC would then be able to perform the deamination on the emerging ssDNA region.

[0215] FIG. 56 shows an example workflow with a linked DNA fragment 400 in which both ends are connected via a hairpin. The linked DNA fragment may be formed using dsDNA library preparation techniques and with hairpin adaptors including adaptor sequences (sequencing primer sequences, index sequences, etc.) such that both ends of a double-stranded fragment 402 generated from a target nucleic acid source are, via library preparation, coupled to hairpin adaptors. In an embodiment, the double-stranded fragment may be ligated to adaptors using conventional adaptors and subsequently modified to link the ends into a linked hairpin structure, e.g., via chemically-linked adaptors.

[0216] The linked DNA fragment 400 is treated with a DNA denaturing agent, such as NaOH, Betaine, DMSO, or heat, or any combination thereof or other denaturing technique, to open the duplex and contacted with an APOBEC enzyme to allow APOBEC conversion of any 5mC in the target fragment 402 are converted to T. The DNA is reannealed to dsDNA, with certain areas of mismatch 404 wherever a 5mC has been converted to T. After reannealing, the hairpins or links can be opened to permit downstream sequencing of a dsDNA library using conventional techniques. However, in certain cases, linked library preparation can be performed as discussed herein (see, e.g., FIG. 43) to retain coupling of top and bottom strands. Thus, only one hairpin loop in an asymmetric adaptor arrangement can be cleaved via a specific sequence not present on the other end in an embodiment.

[0217] FIG. 57 shows different design of cleavable hairpins, e.g., hairpin adaptors, that can be used in conjunction with the workflow of FIG. 58. A hairpin adaptor 410 can contain a cleavable section within a sequence in the structure that can be cleaved by a restriction enzyme specific for the sequence or CRISPr / Cas, so that upon treatment it would remove the stapling from the library. In another example, a hairpin adaptor 412 can contain a cleavable nucleotide within a sequence in the structure Similarly to the adaptor 410, the hairpin is cleaved from the stapled library. However, the cleavage could be much more specific when using a cleavable nucleotide such as 8-oxoG (enzymatic cleavage by Oggl), nitropiperonyl dC (cleaved with UV light), or an abasic site (cleaved with base). A hairpin adaptor 414 may contain a cleavable moiety within a nucleotide or a linker in the structure. In this iteration, the staple is not formedby high Tm regions - instead, the two strands are chemically linked with a synthetic functionalization, such as chemical formation of interstrand disulfides. Release of the two strands would simply require treatment with a reducing agent, and the ssDNA availability of the stapled library generated by use of mismatched sequences. In addition, this approach could also allow adding helicase recognition sites in the adapter region, meaning that the ssDNA could be generated enzymatically in a one-pot APOBEC / Helicase mix.

[0218] FIG. 59 shows an example workflow 450 of step-wise or sequential ligation. The method initiates with synthesis of a 3’ T-tailed, oligonucleotide duplex 452 with an internal biotin modification that acts as the blocking element. These blockers can then be pulled down onto streptavidin coated surfaces such as beads or wells. At each end of the duplex, identical enzymatic cleavage sites will be present for the release of library in subsequent steps.

[0219] Preparation of genomic DNA or target nucleic acid fragments 460 for library generation may involve use of standard, pre-existing mechanical shearing, end repair (blunting and 5’ phosphorylation) and 3’ A-tailing processes. This prepared material can then be added to the beads or wells, and through a ligation event attached to the biotinylated duplexes, further resulting in one end of the DNA insert being blocked. In the illustrated embodiment, a first target nucleic acid fragment 460a is ligated to one end of the duplex 452 and a second target nucleic acid fragment 460b is ligated to the other end of the duplex 452.

[0220] A second ligation event will attach the first adapter 464 to the free end of the DNA insert, whereby any excess adapter can be removed through supernatant removal and / or washing steps. At this instance the DNA can be removed from the blocking duplex and therefore the streptavidin surface, via the addition of a restriction enzyme (RE). Cleavage at the established sites in the blocking duplex that border the DNA, enable release of the insert into solution. After transferring material into new tubes, a third ligation event will affix the second adapter 468 to the unoccupied end of the DNA insert, producing the desired library with two different adapters.

[0221] The workflow 450 may include on-surface ligation facilitated by the inclusion in a reaction component of an affinity moiety and binder (e.g., biotin and streptavidin or similar). If performed in solution, multiple inserts and blocking duplexes could ligate and form contiguous strands. Either an ordered array or limited number of anchored blocking duplexes can be provided on the substrate or well surface to prevent unwanted attachment to neighboring DNA-blocker complexes. Blocking duplexes 452 can be sized and shaped to be long enough (e.g., at least 100 base pairs in length, at least 150 base pairs in length, at least 300 base pairs in length 100-1000 base pairs in an embodiment, 150-300 or 200-500 base pairs in an embodiment) that a single fragment 460 cannot ligate to either side, as this will result in no adapter ligation. The blocking duplexes 452 may be, in an embodiment, at least as long as an average target nucleic acid fragment 460 for a particular sequencing workflow in an embodiment. The workflow 450 may be suitable for smaller genomic DNA inserts or cell free DNA, as longer inserts may ligate in between two blocking duplexes, additionally resulting in no adapter ligation. The workflow also includes three separate ligation events to ensure formation of the correct product.

[0222] In an embodiment, a substrate is provided (e.g., surface, wells, beads) having blocking duplexes 452 immobilized thereon. The blocking duplexes 452 may be conserved or all have a same sequence such that a same cleavage enzyme or enzymes can be used for different target nucleic acid fragment ligations.

[0223] To achieve 9QAM readouts for methylation detection (see FIG. 68), certain techniques for preparing sequencing libraries as disclosed herein involve generating adaptor- linked fragments using asymmetric adapters via ligation (two non-identical adapters ligated on either end of the insert). A single ligation event involving the incorporation of two different adapters, results in only 50% formation of desired product; whereby further reaction inefficiencies facilitate the additional formation of unwanted adapter dimer, as shown in FIG. 58.

[0224] Certain embodiments of the present techniques ensure that ligation of the first adapter does not compromise ligation efficiency of the second adapter, and so are mutuallyexclusive events. Thus, all or most inserts will have two different adapters, such as a forked adaptor and a hairpin adaptor as illustrated. The additional risk of the adapters binding to themselves will also be eliminated, generating a purer library with better sequencing quality to reduce generation of undesired sequencing data via improved sample preparation. To achieve this, cleavable biotinylated duplexes are used to block one end of a double stranded DNA insert, subjecting the other for adapter A ligation. Upon removal of the blocker, the second end of the insert becomes exposed for ligation of adapter B and thus the following issues can be addressed: 1. The problem of inefficient asymmetric adapter ligation is solved by ensuring each adapter type can be added individually to the insert. 2. The problem of asymmetric adapter dimer generation is solved by performing separate ligation events that prevent adapter only interactions.

[0225] Certain embodiments disclosed herein provide workflows and associated reagents for preparing tandem repeat sequencing libraries (where top and bottom strands of a duplex DNA molecule are appended to one another) on a flow cell, whilst enabling deamination of methylated cytosine using a single- stranded deaminase enzyme (e.g., a deaminase enzyme with selective deaminase activity on single-stranded nucleic acid that is not active on doublestranded nucleic acid) for methylation detection. Tandem repeat libraries include oligonucleotides that contain sequence information from both original top and bottom strands, which can be used to detect errors in either strand. This approach leads to far greater overall sequencing accuracy than relying on the information of just one of the strands of a DNA duplex.

[0226] The disclosed techniques may be used in conjunction with a simplified NGS workflow that enables on-flow-cell 1. library prep and 2. clustering / sequencing, which eliminates standard library prep prior to sequencing that is performed using a separate device or as a separate workflow. Standard cluster generation and SBS sequencing is combined with cluster proximity information in DRAGEN algorithms to unlock long-distance and phasing information. That is, by combining library preparation and sequencing into a single workflow,fragment proximity information can be harnessed as additional information for sequence analysis.

[0227] DNA methylation can be a biomarker for diseases including cancer. A singlestranded deaminase enzyme that deaminates 5-methylcytosine in single stranded DNA, converting it to thymine, enables detection of methylation in the original sample from the sequencing data. In particular, workflows can use differential deamination activity for single vs. double-stranded oligonucleotides when the insert sequence itself is present in both doublestranded and single- stranded form on single molecule. By combining the increased sequencing accuracy of duplex sequencing with methylation detection and streamlined workflows, direct detection of methylation alongside high quality, long-distance sequencing information can be achieved. Methods for On Flow Cell Library Prep may be as described in PCT / US2022 / 082280, which is hereby incorporated by reference in its entirety.

[0228] FIG. 60 shows a method of tandem repeat product generation on a surface as generally discussed herein. The method includes ligation of adaptors to a double stranded DNA insert, followed by cycles of denaturing the double stranded molecule, annealing of the complementary hybridization (hyb) sequences at the 3’ end of each strand, and polymerase extension in which each strand acts as a template for extension of the other. Cycling is used as it is possible that the double stranded molecule will reform via annealing of the complementary insert sequences, rather than the complementary hyb sequences and therefore will not form a tandem. The tandem conversion efficiency can be increased by repeating the denaturation, annealing and extension steps. Accordingly, embodiments of the present disclosure include workflows with one or more cycles of denaturation, annealing, and extension. The cycling steps may be performed on the surface, e.g., of a magnetic bead, at a low relative concentration, so that both strands are held in proximity and are likely to anneal with each other via their complementary hyb sequences to form a self-tandem rather than with strands from other molecules which would result in a cross tandem. Similarly, by tagmenting the double stranded DNA molecules on a flow cell surface, both strands would be kept in proximity upon denaturation, which should increase the likelihood of self-tandem formation. This could becontrolled in certain embodiments by varying library loading concentration to the flow cell and transposome spacing on the flow cell surface.

[0229] A novel transposome design for library preparation with sequencing analysis using a tandem duplex workflow is illustrated by way of example in FIG. 61. Two different transposome dimer designs 470, 472 are illustrated by way of example. The first dimer 470 includes sequencing primers (e.g. 5’ P5-A14-ME 3’) on a first oligonucleotide and is attached to the flow cell (e.g. via a biotin-streptavidin interaction) and another or second oligonucleotide containing the hyb sequence (e.g. 5’ ME’-X 3’). The second transposome dimer 472 includes a first oligonucleotide containing the other sequencing primers (e.g. 5’ P7- B15-ME 3’) which is attached to the flow cell (e.g. via a biotin-streptavidin interaction) or intervening structure (e.g., a bead that in turn is associated with a flow cell surface or well) and another or second oligonucleotide containing the complement of the hyb sequence (e.g. 5’ ME’-X’ 3’). However, it should be understood that the illustrated transposomes or transposome complexes may include additional or different adaptor sequences as discussed herein and depending on the associated sequencing platform. Further, to facilitate bridge amplification, the transposome complexes may be provided in sets or pairs with a first sequencing primer sequence present on a first transposome complex and a second sequencing primer present on a second transposome complex. Tagmentation using this set generates surface-immobilized fragments that can be extended and captured at an extended end.

[0230] FIG. 62 uses the novel transposome design of FIG. 61 to generate libraries that can undergo the deamination reaction followed by the tandem insert generation reaction directly on or associated with the flow cell surface. The workflow begins with the loading of double stranded DNA to the flow cell. In certain cases, this technique may operate on relatively long fragments that are 500 or 1000 bases or above in length. In an embodiment, this technique may be used in conjunction with techniques that use read proximity as part of assembly or analysis and that operate on fragments in a range of lOkb or longer. The transposome complexes may be linked to the flow cell surface via a biotin-streptavidin interaction. The dsDNA undergoes tagmentation on the flow cell. The transposon oligonucleotides containingthe hyb sequences (illustrated as being the shorter strand with a 3’ noncomplementary end region) are not directly attached to the dsDNA by the tagmentation step, and a 9bp gap remains between the 3’ end of the dsDNA and the 5’ end of the transposon oligo. For this embodiment, the oligonucleotide containing the hyb sequence or its complement is then attached to the dsDNA fragment via a method such as gap-fill ligation on the surface. TheTn5 can be removed (e g. by heat or sodium dodecyl -sulphate) after tagmentation. Once both strands of the dsDNA are attached directly to both transposon oligos at either end, the denaturation step in the cycling reaction can occur. After denaturation, the DNA is single stranded, with the two strands from the original molecule being held in close proximity on the flow cell via the biotin-streptavi din- linked transposon oligo. Deamination of methylated cytosine to thymine can then take place using the single stranded SGx deaminase enzyme. Once deaminated, the strands can reanneal (via hyb sequences as desired or via their complementary insert sequences) and the reaction cycling can continue via extension. This generates a tandem library product where both strands are entirely complementary, meaning that clustering can go ahead, and all resulting strands from the clusters will contain the same sequences.

[0231] An advantage of carrying out deamination prior to the tandem reaction is the ability to detect methylation status of both original molecule strands from a single cluster, allowing direct detection of hemi-methylation. The resulting tandem libraries, when sequenced, will contain mismatches between reads which indicate the methylation status of each original strand. In FIG. 63, sequencing results are shown when either strand of the tandem library generated in FIG. 62 is clustered and the P5 primers are linearized (i.e. before the paired-end turn). In this case, a T:C mismatch (where the position contains a T in read 1, and the same position contains a C in read 2) indicates that the original A14-attached strand of the library contained a methylated cytosine at that position; whereas a G:A mismatch where the position contains a G in read 1, and the same position contains an A in read 2) indicates that the original B15-attached strand of the library contained a methylated cytosine at that position. Both mismatch types will be present in a single cluster. After the paired end turn, again both mismatch types will be present in a single cluster; however, they are inverted. Now, a T:C mismatch (where the position contains a T in read 3, and the same position contains a C inread 4) indicates that the original B15-attached strand of the library contained a methylated cytosine at that position; whereas a G:A mismatch where the position contains a G in read 3, and the same position contains an A in read 4) indicates that the original A14-attached strand of the library contained a methylated cytosine at that position.

[0232] FIG. 64 uses a similar custom transposome design to generate libraries which undergo the tandem insert product generation reaction from a sample of interest and followed by denaturation and deamination with the single stranded deaminase enzyme. The workflow begins with the loading of double stranded DNA (e.g., 500 bases and longer) to the flow cell. The dsDNA undergoes tagmentation on the flow cell (step 500), Tn5 removal (e g. by heat or sodium dodecyl-sulphate), and gap-fill ligation or another method to attach the 3’ end of the tagmented dsDNA and the 5’ end of the transposon oligos containing the hyb sequences (step 502). Once both strands of the dsDNA are attached directly to both transposon oligos at either end, the tandem reaction cycling can occur. This generates a double stranded tandem library product. This can be denatured to make the DNA single stranded, with the two strands from the original molecule being held in close proximity on the flow cell via the biotin-streptavidin- linked transposon oligo. Deamination of methylated cytosine to thymine can then take place using the single stranded deaminase enzyme. This generates a deaminated tandem library product where the strands are not completely complementary. Therefore, one strand of the tandem product is cleaved prior to clustering to avoid generation of a poly-clonal cluster. This can be achieved by including a cleavable site in the surface-bound transposon oligo of just one of the transposome designs (such as a uracil base, which could be cleaved by USER enzyme followed by denaturation). Accordingly, novel transposomes that include cleavable sites for use in tandem repeat product generation are encompassed in certain embodiments. This enables removal of one strand of the tandem product, leaving just a single strand to be amplified for cluster generation. Alternatively, the implementation of 9QAM as discussed herein would enable clustering of both strands (assuming relatively equal proportions within the well after clustering) as, although there will be polyclonality depending on which original strand the sequence was clustered from, it is possible to detect this as the cluster will move into a center cloud. This will indicate methylation at that position.

[0233] By deaminating after the tandem reaction (rather than before as shown in FIGS. 62- 63), a single cluster generated from one strand of tandem library (after cleavage of the other strand) will only provide the methylation status of one original strand, meaning that there is no direct detection of hemi-methylation in the original molecule. In FIG. 65, sequencing results are shown when either strand of the tandem library is clustered and the P5 primers are linearized (i.e. before the paired-end turn). Similarly to FIG. 63, a T:C mismatch between reads indicates that the original A14 strand of the library contained a methylated cytosine at that position; whereas a G:A mismatch indicates that the original B15 strand of the library contained a methylated cytosine at that position. However, unlike FIG. 63, only a single mismatch type (Either T:C or G:A) will be present in a single cluster from a single strand of the tandem library when comparing between reads 1 and 2 (before paired end turn). Therefore, clusters from both strands of the tandem library are used to determine the methylation status of both the A14 and B15 strand of the original molecule. For this reason, deamination prior to the tandem reaction (FIG. 62) provides more methylation information than the workflow shown in FIG 64

[0234] FIGS. 66-67 show alternative workflows for methylation detection with tandem duplex library preparation on flow cell which can deaminate double-stranded DNA instead of only single-stranded DNA. FIG. 66 shows deamination using a dsDNA-compatible deaminase enzyme prior to the tandem reaction on the flow cell. The deamination can be performed on the input DNA sample prior to the library prep as illustrated, or after tagmentation but before the tandem reaction. The tandem reaction can then occur as provided herein on the flow cell. As deamination occurs prior to tandem generation, both strands of the tandem library product are completely complementary and can therefore be clustered simultaneously to form a single cluster which will provide the methylation status of both original molecule strands, enabling direct detection of hemi-methylation (see FIG. 63).

[0235] FIG. 67 shows the tandem reaction on the flow cell followed by deamination using a dsDNA-compatible deaminase enzyme. The disadvantage of this method is that tandem generation prior to deamination results in a tandem library product where the strands are notcompletely complementary and therefore one must be cleaved prior to clustering. A single cluster generated from one strand of tandem product will only provide the methylation status of one original molecule strand, meaning that there is no direct detection of hemi-methylation in the original molecule (see FIG. 65).

[0236] Please note that for all diagrams, Illumina sequences A14, B 15, P5, P7 and ME have been used as well as the hyb sequence X; however, these libraries could use alternative sequences and may also include additional components such as indexes and pre-extension primer sites which have not been shown. However, the overall concepts described here would remain the same.Sample Preparation and Sequencing

[0237] Nucleic acid sequencing and certain sequencing library preparation steps suitable for use in conjunction with the disclosed embodiment may be performed as generally discussed in US20230407388A1, WO2023175021 Al and WO2023175013A1, which are hereby incorporated by reference in their entireties for all purposes. Following denaturation, a singlestranded library may be contacted in free solution onto a solid support comprising surface capture moieties (for example P5 and P7 lawn primers). Thus, embodiments of the present invention may be performed on a solid support, such as a flow cell. However, in alternative embodiments, seeding and clustering can be conducted off-flow cell using other types of solid support. The solid support may comprise a substrate. In one embodiment, the solid support comprises at least one first immobilized primer and at least one second immobilized primer. These immobilized primers may also be known as lawn primers. Thus, each well may comprise at least one first immobilized primer , and typically may comprise a plurality of first immobilized primers. In addition, each well may comprise at least one second immobilized primer, and typically may comprise a plurality of second immobilized primers. Thus, each well may comprise at least one first immobilized primer and at least one second immobilized primer, and typically may comprise a plurality of first immobilized primers and a plurality of second immobilized primers. The first immobilized primer may be attached via a 5 ’-end of its polynucleotide chain to the solid support. When extension occurs from the first immobilizedprimer 201 , the extension may be in a direction away from the solid support. The second immobilized primer may be attached via a 5 ’-end of its polynucleotide chain to the solid support. When extension occurs from second immobilized primer, the extension may be in a direction away from the solid support. The first immobilized primer may be different to the second immobilized primer and / or a complement of the second immobilized primer. The second immobilized primer may be different to the first immobilized primer and / or a complement of the first immobilized primer. The (or each of the) first or second immobilized primer(s) may comprise a sequence as defined in SEQ ID NO. 7 or 8, or a variant or complement thereof.[00238J By way of brief example, following attachment of the P5 and P7 primers to the solid support, the solid support may be contacted with the template to be amplified under conditions which permit hybridization (or annealing - such terms may be used interchangeably) between the template and the immobilized primers. The template is usually added in free solution under suitable hybridization conditions, which will be apparent to the skilled reader. Typically, hybridization conditions are, for example, 5xSSC at 40°C. However, other temperatures may be used during hybridization, for example about 50°C to about 75°C, about 55°C to about 70°C, or about 60°C to about 65°C. Solid-phase amplification can then proceed. The first step of the amplification is a primer extension step in which nucleotides are added to the 3' end of the immobilized primer using the template to produce a fully extended complementary strand. The template is then typically washed off the solid support. The complementary strand will include at its 3' end a primer-binding sequence (i.e. either P5’ or P7’) which is capable of bridging to the second primer molecule immobilized on the solid support and binding. The resulting structure is referred to herein as a sequence bridge. Further rounds of amplification (analogous to a standard PCR reaction) leads to the formation of clusters or colonies of template molecules bound to the solid support. This is called clustering. Thus, solid-phase amplification by either a method analogous to that of WO 98 / 44151 or that of WO 00 / 18957 (the contents of which are incorporated herein in their entirety by reference) will result in production of a clustered array comprised of colonies of "bridged" amplification products (or sequence bridges). This process is known as bridge amplification. Both strands of theamplification products will be immobilized on the solid support at or near the 5' end, this attachment being derived from the original attachment of the amplification primers. Typically, the amplification products within each colony will be derived from amplification of a single template molecule. Other amplification procedures may be used, and will be known to the skilled person. For example, amplification may be isothermal amplification using a strand displacement polymerase; or may be exclusion amplification as described in WO 2013 / 188582. Further information on amplification can be found in WO 2002 / 06456 and WO 2007 / 107710, the contents of which are incorporated herein in their entirety by reference.

[0239] Through such approaches, a cluster of template molecules is formed, comprising copies of a template strand and copies of the complement of the template strand. In some cases, to facilitate sequencing, one set of strands (either the original template strands or the complement strands thereof) may be removed from the solid support leaving either the original template strands or the complement strands. Suitable methods for removing such strands are described in more detail in application number WO 2007 / 010251 , the contents of which are incorporated herein by reference in their entirety. As described herein, the template provides information (e.g. identification of the genetic sequence, identification of epigenetic modifications) on the original target polynucleotide sequence. For example, a sequencing process (e.g. a sequencing-by-synthesis (referred to herein as SBS) or sequencing-by-ligation process) may reproduce information that was present in the original target polynucleotide sequence, by using complementary base pairing. In one embodiment, sequencing may be carried out using any suitable "sequencing-by- synthesis" technique, wherein nucleotides are added successively in cycles to the free 3' hydroxyl group, resulting in synthesis of a polynucleotide chain in the 5' to 3' direction. The nature of the nucleotide added may be determined after each addition. One particular sequencing method relies on the use of modified nucleotides that can act as reversible chain terminators. Such reversible chain terminators comprise removable 3' blocking groups. Once such a modified nucleotide has been incorporated into the growing polynucleotide chain complementary to the region of the template being sequenced there is no free 3'-OH group available to direct further sequence extension and therefore the polymerase cannot add further nucleotides. Once the nature of thebase incorporated into the growing chain has been determined, the 3' block may be removed to allow addition of the next successive nucleotide. By ordering the products derived using these modified nucleotides it is possible to deduce the DNA sequence of the DNA template. Such reactions can be done in a single experiment if each of the modified nucleotides has attached thereto a different label, known to correspond to the particular base, to facilitate discrimination between the bases added at each incorporation step. Suitable labels are described in PCT application PCT / GB2007 / 001770, the contents of which are incorporated herein by reference in their entirety. Alternatively, a separate reaction may be carried out containing each of the modified nucleotides added individually. The modified nucleotides may carry a label to facilitate their detection. Such a label may be configured to emit a signal, such as an electromagnetic signal, or a (visible) light signal.

[0240] In a particular embodiment, the label is a fluorescent label (e.g. a dye). Thus, such a label may be configured to emit an electromagnetic signal, or a (visible) light signal. One method for detecting the fluorescently labelled nucleotides comprises using laser light of a wavelength specific for the labelled nucleotides, or the use of other suitable sources of illumination. The fluorescence from the label on an incorporated nucleotide may be detected by a CCD camera or other suitable detection means. Suitable detection means are described in PCT / US2007 / 007991 , the contents of which are incorporated herein by reference in their entirety.

[0241] However, the detectable label need not be a fluorescent label. Any label can be used which allows the detection of the incorporation of the nucleotide into the DNA sequence. Each cycle may involve simultaneous delivery of four different nucleotide types to the array of template molecules. Alternatively, different nucleotide types can be added sequentially and an image of the array of template molecules can be obtained between each addition step. In some embodiments, each nucleotide type may have a (spectrally) distinct label. In other words, four channels may be used to detect four nucleobases (also known as 4- channel chemistry). For example, a first nucleotide type (e.g. A) may include a first label (e g. configured to emit a first wavelength, such as red light), a second nucleotide type (e.g. G) may include a secondlabel (e.g. configured to emit a second wavelength, such as blue light), a third nucleotide type (e g. T) may include a third label (e g. configured to emit a third wavelength, such as green light), and a fourth nucleotide type (e.g. C) may include a fourth label (e.g. configured to emit a fourth wavelength, such as yellow light). Four images can then be obtained, each using a detection channel that is selective for one of the four different labels. For example, the first nucleotide type (e.g. A) may be detected in a first channel (e.g. configured to detect the first wavelength, such as red light), the second nucleotide type (e.g. G) may be detected in a second channel (e.g. configured to detect the second wavelength, such as blue light), the third nucleotide type (e.g. T) may be detected in a third channel (e.g. configured to detect the third wavelength, such as green light), and the fourth nucleotide type (e.g. C) may be detected in a fourth channel (e.g. configured to detect the fourth wavelength, such as yellow light). Although specific pairings of bases to signal types (e.g. wavelengths) are described above, different signal types (e.g. wavelengths) and / or permutations may also be used. In some embodiments, detection of each nucleotide type may be conducted using fewer than four different labels. For example, sequencing-by-synthesis may be performed using methods and systems described in US 2013 / 0079232, which is incorporated herein by reference.

[0242] Thus, in some embodiments, two channels may be used to detect four nucleobases (also known as 2-channel chemistry). For example, a first nucleotide type (e.g. A) may include a first label (e.g. configured to emit a first wavelength, such as green light) and a second label (e g. configured to emit a second wavelength, such as red light), a second nucleotide type (e.g. G) may not include the first label and may not include the second label, a third nucleotide type (e.g. T) may include the first label (e.g. configured to emit the first wavelength, such as green light) and may not include the second label, and a fourth nucleotide type (e.g. C) may not include the first label and may include the second label (e.g. configured to emit the second wavelength, such as red light). Two images can then be obtained, using detection channels for the first label and the second label. For example, the first nucleotide type (e.g. A) may be detected in both a first channel (e.g. configured to detect the first wavelength, such as red light) and a second channel (e.g. configured to detect the second wavelength, such as green light), the second nucleotide type (e.g. G) may not be detected in the first channel and may not bedetected in the second channel, the third nucleotide type (e.g. T) may be detected in the first channel (e.g. configured to detect the first wavelength, such as red light) and may not be detected in the second channel, and the fourth nucleotide type (e.g. C) may not be detected in the first channel and may be detected in the second channel (e.g. configured to detect the second wavelength, such as green light). Although specific pairings of bases to signal types (e.g. wavelengths) and / or combinations of channels are described above, different signal types (e g. wavelengths) and / or permutations may also be used.

[0243] In some embodiments, one channel may be used to detect four nucleobases (also known as 1 -channel chemistry). For example, a first nucleotide type (e.g. A) may include a cleavable label (e.g. configured to emit a wavelength, such as green light), a second nucleotide type (e.g. G) may not include a label, a third nucleotide type (e.g. T) may include a non- cleavable label (e.g. configured to emit the wavelength, such as green light), and a fourth nucleotide type (e.g. C) may include a label-accepting site which does not include the label. A first image can then be obtained, and a subsequent treatment carried out to cleave the label attached to the first nucleotide type, and to attach the label to the label-accepting site on the fourth nucleotide type. A second image may then be obtained. For example, the first nucleotide type (e.g. A) may be detected in a channel (e.g. configured to detect the wavelength, such as green light) in the first image and not detected in the channel in the second image, the second nucleotide type (e.g. G) may not be detected in the channel in the first image and may not be detected in the channel in the second image, the third nucleotide type (e.g. T) may be detected in the channel (e.g. configured to detect the wavelength, such as green light) in the first image and may be detected in the channel (e.g. configured to detect the wavelength, such as green light) in the second image, and the fourth nucleotide type (e.g. C) may not be detected in the channel in the first image and may be detected in the channel in the second image (e.g. configured to detect the wavelength, such as green light). Although specific pairings of bases to signal types (e.g. wavelengths) and / or combinations of images are described above, different signal types (e.g. wavelengths), images and / or permutations may also be used.

[0244] In one embodiment, the sequencing process comprises a first sequencing read (referred to herein as Rl) and second sequencing read (referred to herein as R2). As described below, in each read at least two different polynucleotide strands may be sequenced simultaneously, generating a Rl.l and R1.2 read and a R2.1 and R2.2 read. The first sequencing read and the second sequencing read may also be conducted concurrently. In other words, the first sequencing read and the second sequencing read may be conducted at the same time. The first sequencing read may comprise the binding of a first sequencing primer (also known as a read 1 sequencing primer) to the first sequencing primer binding site. The second sequencing read may comprise the binding of a second sequencing primer (also known as a read 2 sequencing primer) to the second sequencing primer binding site. Alternative methods of sequencing include sequencing by ligation, for example as described in US 6,306,597 or WO 06 / 084132, the contents of which are incorporated herein by reference.

[0245] FIG. 70 is a schematic diagram of a sequencing device 600 that may be used in conjunction with the disclosed embodiments for acquiring sequencing data of fragments, such as adaptor-linked fragments as generally discussed herein. The sequence device 500 may be implemented according to any sequencing technique, such as those incorporating sequencing- by-synthesis methods described in U.S. Patent Publication Nos. 2007 / 0166705; 2006 / 0188901; 2006 / 0240439; 2006 / 0281109; 2005 / 0100900; U.S. Pat. No. 7,057,026; WO 05 / 065814; WO 06 / 064199; WO 07 / 010,251, the disclosures of which are incorporated herein by reference in their entireties. Alternatively, sequencing by ligation techniques may be used in the sequencing device 600. Such techniques use DNA ligase to incorporate oligonucleotides and identify the incorporation of such oligonucleotides and are described in U.S. Pat. No. 6,969,488; U.S. Pat. No. 6,172,218; and U.S. Pat. No. 6,306,597; the disclosures of which are incorporated herein by reference in their entireties. Some embodiments can utilize nanopore sequencing, whereby target nucleic acid strands, or nucleotides exonucleolytically removed from target nucleic acids, pass through a nanopore. As the target nucleic acids or nucleotides pass through the nanopore, each type of base can be identified by measuring fluctuations in the electrical conductance of the pore (U.S. Patent No. 7,001,792; Soni & Meller, Clin. Chem. 53, 1996-2001 (2007); Healy, Nanomed 2, 459-481 (2007); and Cockroft, et al. J. Am. Chem.Soc. 130, 818-820 (2008), the disclosures of which are incorporated herein by reference in their entireties). Yet other embodiments include detection of a proton released upon incorporation of a nucleotide into an extension product. For example, sequencing based on detection of released protons can use an electrical detector and associated techniques that are commercially available from Ion Torrent (Guilford, CT, a Life Technologies subsidiary) or sequencing methods and systems described in US 2009 / 0026082 Al; US 2009 / 0127589 Al; US 2010 / 0137143 Al; or US 2010 / 0282617 Al, each of which is incorporated herein by reference in its entirety. Particular embodiments can utilize methods involving the real-time monitoring of DNA polymerase activity. Nucleotide incorporations can be detected through fluorescence resonance energy transfer (FRET) interactions between a fluorophore-bearing polymerase and y-phosphate-labeled nucleotides, or with zeromode waveguides as described, for example, in Levene et al. Science 299, 682-686 (2003); Lundquist et al. Opt. Lett. 33, 1026-1028 (2008); Korlach et al. Proc. Natl. Acad. Set. USA 105, 1176-1181 (2008), the disclosures of which are incorporated herein by reference in their entireties. Other suitable alternative techniques include, for example, fluorescent in situ sequencing (FISSEQ), and Massively Parallel Signature Sequencing (MPSS). In particular embodiments, the sequencing device 600 may be a HiSeq, MiSeq, or HiScanSQ from Illumina (La Jolla, CA). In other embodiment, the sequencing device 500 may be configured to operate using a CMOS sensor with nanowells fabricated over photodiodes such that DNA deposition is aligned one-to-one with each photodiode.

[0246] The sequencing device 500 may be “one-channel” a detection device, in which only two of four nucleotides are labeled and detectable for any given image. For example, thymine may have a permanent fluorescent label, while adenine uses the same fluorescent label in a detachable form. Guanine may be permanently dark, and cytosine may be initially dark but capable of having a label added during the cycle. Accordingly, each cycle may involve an initial image and a second image in which dye is cleaved from any adenines and added to any cytosines such that only thymine and adenine are detectable in the initial image but only thymine and cytosine are detectable in the second image. Any base that is dark through both images in guanine and any base that is detectable through both images is thymine. A base thatis detectable in the first image but not the second is adenine, and a base that is not detectable in the first image but detectable in the second image is cytosine. By combining the information from the initial image and the second image, all four bases are able to be discriminated using one channel.

[0247] In the depicted embodiment, the sequencing device 500 includes a separate sample processing device 502 and an associated computer 504. However, as noted, these may be implemented as a single device. Further, the associated computer 504 may be local to or networked or otherwise in communication with the sample processing device 502. In the depicted embodiment, the biological sample may be loaded into the sample processing device 502 on a sample substrate 510, e.g., a flow cell or slide, that is imaged to generate sequence data. For example, reagents that interact with the biological sample fluoresce at particular wavelengths in response to an excitation beam generated by an imager 512 and thereby return radiation for imaging. For instance, the fluorescent components may be generated by fluorescently tagged nucleic acids that hybridize to complementary molecules of the components or to fluorescently tagged nucleotides that are incorporated into an oligonucleotide using a polymerase. As will be appreciated by those skilled in the art, the wavelength at which the dyes of the sample are excited and the wavelength at which they fluoresce will depend upon the absorption and emission spectra of the specific dyes. Such returned radiation may propagate back through the directing optics. This retrobeam may generally be directed toward detection optics of the imager 512.

[0248] The imager detection optics may be based upon any suitable technology, and may be, for example, a charged coupled device (CCD) sensor that generates pixilated image data based upon photons impacting locations in the device. However, it will be understood that any of a variety of other detectors may also be used including, but not limited to, a detector array configured for time delay integration (TDI) operation, a complementary metal oxide semiconductor (CMOS) detector, an avalanche photodiode (APD) detector, a Geiger-mode photon counter, or any other suitable detector. TDI mode detection can be coupled with line scanning as described in U. S. Patent No. 7,329,860, which is incorporated herein by reference.Other useful detectors are described, for example, in the references provided previously herein in the context of various nucleic acid sequencing methodologies.

[0249] The imager 512 may be under processor control, e.g., via a processor 514, and the sample receiving device 502 may also include I / O controls 516, an internal bus 518, nonvolatile memory 520, RAM 522 and any other memory structure such that the memory is capable of storing executable instructions, and other suitable hardware components that may be similar to those described with regard to FIG. 70. Further, the associated computer 504 may also include a processor 524, I / O controls 526, communications circuity 527, and a memory architecture including RAM 528 and non-volatile memory 530, such that the memory architecture is capable of storing executable instructions 532. The hardware components may be linked by an internal bus, which may also link to the display 534. In embodiments in which the sequencing device 500 is implemented as an all-in-one device, certain redundant hardware elements may be eliminated.

[0250] The processor 514, 524 may be programmed to assign individual sequencing reads to a sample based on the associated index sequence or sequences according to the techniques provided herein. In particular embodiments, based on the image data acquired by the imager 512, the sequencing device 500 may be configured to generate sequencing data that includes base calls for each base of a sequencing read. Further, based on the image data, even for sequencing reads that are performed in series, the individual reads may be linked to the same location via the image data and, therefore, to the same template strand. In this manner, index sequencing reads may be associated with a sequencing read of an insert sequence before being assigned to a sample of origin. The processor 514, 524 may also be programmed to perform downstream analysis on the sequences corresponding to the inserts for a particular sample subsequent to assignment of sequencing reads to the sample.

[0251] In some versions, the sequencing device 600 utilizes SBS to sequence nucleotide fragments into nucleotide reads and determine nucleobase calls for the nucleotide reads. By executing sequencing device system, sequencing device 600 may further store the nucleobase calls as part of base-call data that is formatted as a binary base call (BCL) file and send theBCL file to the local device and / or the server device(s). Sequencing device 600 may communicate the BCL file and / or other data to local device and / or client device via network or directly (i.e., bypassing network).

[0252] The sequencing devices may be configured in a 9QAM or 16QAM encoding scheme as generally discussed in WO2023175021 Al and WO2023175013A1, which are hereby incorporated by reference in their entireties for all purposes. In an embodiment, a 9QAM encoding scheme, shown by way of example in FIG. 68, can be used to accurately differentiate between two simultaneously received base calls. Plotting relative intensities of light signals obtained from Read 1.1 and Read 1.2 generates a constellation of 9 clouds. The four comer clouds represent high quality and accurate base calls, while off-corner clouds represent potential library prep I sequencing errors, which could be eliminated. A 9QAM encoding scheme can be used to simultaneously sequence genomic and epigenetic data; epigenetic conversion of the polynucleotide library strand by, for example, Bisulfite / EM-Seq or TAPS and subsequent sequencing enables mC and the canonical bases to be identified simultaneously. By plotting relative intensities of light signals obtained from Read 1.1 and Read 1.2 a constellation of 9 clouds is obtained. Each of these clouds allows sequence information to be identified from the two reads; in this particular encoding scheme, the top left corner of four clouds corresponds with base calls corresponding to A, the top right corner of four clouds corresponds with base calls corresponding to T, the bottom left comer of four clouds corresponds with base calls corresponding to G, and the bottom right corner of four clouds corresponds with base calls corresponding to C; however, other encoding schemes are possible and each of C, G, A and T may be mapped to different cloud permutations. By plotting the light intensities in this manner it is possible to determine an accurate base call from a library prep or sequencing error (and by library prep or sequencing error is meant here that there is a mismatch between read 1.1 and read 1.2, which may be indicative of asymmetry between the forward and reverse strands, for example, because of DNA damage to one strand).

[0253] The method described herein can also be used to simultaneously sequence genomic and epigenetic data. Following preparation of the polynucleotide library strand, an epigeneticconversion is applied. The modified library strand can then be sequenced as described above and the sequences of the duplex strands read simultaneously. A 9QAM system is used to decode the simultaneously-received read signals. Depending on which technology for epigenetic conversion is used, the C / C cloud may either represent a mC (Bisulfite / EM-Seq) or accurate C call (TAPS) and vice versa, the C / T cloud will represent the mC or accurate C calls respectively.

[0254] In a 16QAM encoding scheme, shown by way of example in FIG. 69, sixteen distributions (or bins) of intensity values from the combination of a brighter signal (i.e. a first signal as described herein) and a dimmer signal (i.e. a second signal as described herein). The intensity values shown in Figure 20 may be up to a scale or normalisation factor; the units of the intensity values may be arbitrary or relative (i.e., representing the ratio of the actual intensity to a reference intensity). The sum of the brighter signal generated by the first portions and the dimmer signal generated by the second portions results in a combined signal. The combined signal may be captured by a first optical channel and a second optical channel. Since the brighter signal may be A, T, C or G, and the dimmer signal may be A, T, C or G, there are sixteen possibilities for the combined signal, corresponding to sixteen distinguishable patterns when optically captured. That is, each of the sixteen possibilities corresponds to a bin. The computer system can map the combined signal generated into one of the sixteen bins, and thus determine the added nucleobase at the first portion and the added nucleobase at the second portion, respectively. For example, when the combined signal is mapped to bin 1612 for a base calling cycle, the computer processor base calls both the added nucleobase at the first portion and the added nucleobase at the second portion as C. When the combined signal is mapped to bin 1614 for the base calling cycle, the processor base calls the added nucleobase at the first portion as C and the added nucleobase at the second portion as T. When the combined signal is mapped to bin 1616 for the base calling cycle, the processor base calls the added nucleobase at the first portion as C and the added nucleobase at the second portion as G. When the combined signal is mapped to bin 1618 for the base calling cycle, the processor base calls the added nucleobase at the first portion as C and the added nucleobase at the second portion as A. When the combined signal is mapped to bin 1622 for the base calling cycle, the processorbase calls the added nucleobase at the first portion as T and the added nucleobase at the second portion as C. When the combined signal is mapped to bin 1624 for the base calling cycle, the processor base calls both the added nucleobase at the first portion and the added nucleobase at the second portion as T. When the combined signal is mapped to bin 1626 for the base calling cycle, the processor base calls the added nucleobase at the first portion as T and the added nucleobase at the second portion as G. When the combined signal is mapped to bin 1628 for the base calling cycle, the processor base calls the added nucleobase at the first portion as T and the added nucleobase at the second portion as A.

[0255] When the combined signal is mapped to bin 1632 for the base calling cycle, the processor base calls the added nucleobase at the first portion as G and the added nucleobase at the second portion as C. When the combined signal is mapped to bin 1634 for the base calling cycle, the processor base calls the added nucleobase at the first portion as G and the added nucleobase at the second portion as T. When the combined signal is mapped to bin 1636 for the base calling cycle, the processor base calls both the added nucleobase at the first portion and the added nucleobase at the second portion as G. When the combined signal is mapped to bin 1638 for the base calling cycle, the processor base calls the added nucleobase at the first portion as G and the added nucleobase at the second portion as A.

[0256] When the combined signal is mapped to bin 1642 for the base calling cycle, the processor base calls the added nucleobase at the first portion as A and the added nucleobase at the second portion as C. When the combined signal is mapped to bin 1644 for the base calling cycle, the processor base calls the added nucleobase at the first portion as A and the added nucleobase at the second portion as T. When the combined signal is mapped to bin 1646 for the base calling cycle, the processor base calls the added nucleobase at the first portion as A and the added nucleobase at the second portion as G. When the combined signal is mapped to bin 1648 for the base calling cycle, the processor base calls both the added nucleobase at the first portion and the added nucleobase at the second portion as A.Adaptor and Sequencing Reaction SequencesTable 1 provides a listing of certain sequences referenced herein.

[0257] The disclosed technique may be used in conjunction with universal or conserved sequences, such as primer sequences or adaptor sequences. In certain cases, these adaptor sequences may be incorporated onto fragments to generate adaptor-linked fragments in which a pool or population of adaptor-linked fragments has the same adaptors but different targets or inserts.

[0258] Example sequences used in adaptors A14-ME, ME, B15-ME, ME', A14, B15, and ME are provided below:

[0259] A14-ME: 5'-TCGTCGGCAGCGTCAGATGTGTATAAGAGACAG-3' (SEQ ID NO: 1)

[0260] B 15-ME: 5'-GTCTCGTGGGCTCGGAGATGTGTATAAGAGACAG-3' (SEQ ID NO: 2)

[0261] ME': 5'-phos-CTGTCTCTTATACACATCT-3' (SEQ ID NO: 3)

[0262] A14: 5'-TCGTCGGCAGCGTC-3' (SEQ ID NO: 4)

[0263] Bl 5: 5'-GTCTCGTGGGCTCGG-3' (SEQ ID NO: 5)

[0264] ME: AGATGTGTATAAGAGACAG (SEQ ID NO. : 6)

[0265] The primer region or primer binding region can include a region having the sequence of a universal Illumina® capture primer or a region specifically hybridizing with a universal Illumina® capture primer. Universal Illumina® capture primers include, e.g., P5 5’- AATGATACGGCGACCACCGA-3’ ((SEQ ID NO: 7)) or P7 (5’- CAAGCAGAAGACGGCATACGA-3’ (SEQ ID NO: 8)), or fragments thereof. A region specifically hybridizing with a universal Illumina® capture primer can include, e.g., the reverse complement sequence of the Illumina® capture primer P5 ("anti-P5": 5’- TCGGTGGTCGCCGTATCATT-3’ (SEQ ID NO: 9) or P7 ("anti-P7": 5’- TCGTATGCCGTCTTCTGCTTG-3’ (SEQ ID NO: 10)), or fragments thereof.

[0266] A conserved primer region can additionally or alternatively include a region having the sequence of an Illumina® sequencing primer, or fragment thereof, or a region specifically hybridizing with an Illumina® sequencing primer, or fragment thereof. Illumina® sequencing primers include, e.g, SBS3 (5’-ACACTCTTTCCCTACACGACGCTCTTCCGATCT-3’ (SEQ ID NO: 11)) or SBS8 (5’-CGGTCTCGGCATTCCTGCTGAACCGCTCTTCCGATCT-3’ (SEQ ID NO: 12)). A regionspecifically hybridizing with an Illumina® sequencing primer, or fragment thereof, can include, e.g., the reverse complement sequence of the Illumina® sequencing primer SB S3 ("anti-SBS3": 5’-AGATCGGAAGAGCGTCGTGTAGGGAAAGAGTGT-3’ (SEQ ID NO: 13)) or SBS8("anti-SBS8":5’-AGATCGGAAGAGCGGTTCAGCAGGAATGCCGAGACCG-3’ (SEQ ID NO: 14)), or fragments thereof. The incorporation of sequencing primer sequences in the adaptors may be either directly or via subsequent amplification, ligation, or other sequencing library preparation steps.

[0267] In an embodiment, the disclosed extension or amplification products may differ from one another based on different inserts but can have conserved adaptor sequences (e.g., a symmetric adaptor or different, conserved, adaptors on each end). In this manner, universal capture sequences can be used for sequencing libraries generated from extension products. In an embodiment, the sequencing may use Illumina® NGS primers. The following primers are shown by way of example.Read 1 5’ TCGTCGGCAGCGTCAGATGTGTATAAGAGACAG 3’ (SEQ ID NO: 15)Read 2 5’ GTCTCGTGGGCTCGGAGATGTGTATAAGAGACAG (SEQ ID NO: 16)Paired End Read 1 5' ACACTCTTTCCCTACACGACGCTCTTCCGATCT (SEQ ID NO: 17)Paired End Read 2 5' CGGTCTCGGCATTCCTGCTGAACCGCTCTTCCGATCT (SEQ ID NO: 18)Index 1 Read 5’ CAAGCAGAAGACGGCATACGAGAT[i7]GTCTCGTGGGCTCGG (SEQ ID NO: 19)Index 2 Read 5’ AATGATACGGCGACCACCGAGATCTACAC[i5]TCGTCGGCAGCGTC (SEQ ID NO: 20)It should be understood that the index read primers may be designed to include the particular index sequence associated with a particular sample in a sequencing reaction. Thus, the index primers may have a nucleotide region, shown as i5 or i7, that varies in sequence between different samples of a multiplexed sample. Other samples in the run can be prepared with primers that include their respective indexes. Accordingly, certain sequence reads may be obtained with universal primers while other sequence reads are obtained with primers or a mix of primers that are specific to indexes of one or more samples in a multiplexed reaction.

[0268] In an embodiment, unique molecular identifiers (UMIs) may be incorporated onto the extension products via ligation. UMIs are short sequences used to uniquely tag each molecule in a sample library to provide error correction and reduce sequencing bias.

[0269] It should be understood that the disclosed sequences are by way of example. In certain embodiments, other adaptor sequences compatible with other NGS platforms may be used.

[0270] Sequencing primers and adapter sequences that may be used for sequencing may include Illumina library preparation kits and sequencing platforms, e.g., Nextera, Illumina Prep, Ilumina PCR, AmpliSeq™, TruSight®, and TruSeq™, are as disclosed in Illumina Adapter Sequences Document #1000000002694 v!5, and is hereby incorporated by reference in its entirety. These sequencing primers and adapters may be modified in accordance with the present disclosure. Examples of said primers and adapters include the following: Read 1, Read 2, Index 1 Read, Index 2 Read, Index 1 (i7) Adapters, Index 2 (i5) Adapters, Index Adapters 1-27, TruSeq Universal Adapter, Index PCR Primers, Multiplexing Adapters, Multiplexing Read Sequencing Primers, Multiplexing Index Read Sequencing Primers, and PCR Primer Index SequencesNucleic Acids Comprising Multiple Insert Sequences

[0271] Described herein are nucleic acids that comprise multiple insert sequences, wherein each insert comprises at least a portion or contiguous region of one or more target nucleicacids. In some embodiments, a polynucleotide comprises two insert sequences. In some embodiments, a polynucleotide comprises three, four, or five insert sequences. A polynucleotide comprising more than one insert that can be used as a sequencing template may be referred to herein as a tandem repeat or tandem insert or sequencing template. Further, a tandem repeat may include a repeat of a top and / or bottom strand of an insert (e.g., target) sequence. That is, certain tandem repeats may include a top strand of an insert sequence as a first instance and a bottom strand of the insert sequence as the second instance or vice versa. It should be understood that the tandem inserts or tandem repeats may not necessarily be directly adjacent to one another in the polynucleotide. As referred to herein, a tandem repeat or tandem insert may include a first instance of a target sequence and a second instance of a target sequence that are separated by an intervening sequence, such as part an adaptor sequence. A tandem repeat may include the original template oligonucleotide and one or more copies of the original template oligonucleotide.

[0272] In some embodiments, polynucleotides comprise a hybridization sequence or the complement of a hybridization sequence. “Hybridization sequence” or “HYB,” as used herein, refers to a sequence that can hybridize to a complementary hybridization sequence. For example, hybridization of HYB in one fragment (such as a library product) to a HYB’ (the complement of a hybridization sequence) in another fragment can lead to a hybridization adduct or a bridge, wherein the two fragments anneal to each other via hybridization of HYB / HYB’. In some embodiments, HYB comprises sufficient nucleotides to attach two single-stranded fragments together when HYB hybridizes to HYB’. In some embodiments, a HYB or HYB’ comprises 10-30 nucleotides. In some embodiments, binding of the HYB in a first single-stranded nucleic acid fragment to the HYB’ in a second single-stranded nucleic acid fragment is sufficient to “bridge” the two fragments. The nucleotides comprised in a HYB or HYB’ may be naturally occurring or artificial or modified nucleotides. In some embodiments, HYB or HYB’ comprising artificial or modified nucleotides may require fewer nucleotides in these sequences to allow bridging between two single- stranded fragments.

[0273] In some embodiments, one or more nucleotide in the HYB or HYB’ is a locked nucleic acid or a bridged nucleic acid. As used herein, a “locked nucleic acid” or “LNA” refers to a modified nucleotide in which the ribose moiety is modified with an extra bridge connecting the 2’ oxygen and 4’ carbon. In some embodiments, LNAs confer heightened structural stability in the HYB or HYB’ sequence, thus increasing the hybridization melting temperature (Tm) of the HYB / HYB’ interaction. For example, HYB or HYB’ sequences comprising one or more LNAs may only comprise relatively short sequences (such as 10-20 nucleotides), yet still confer sufficiently strong binding to allow formation of bridges between a first single-stranded fragment comprising a HYB and a second single-stranded fragment comprising a HYB’ .

[0274] In some embodiments, the polynucleotide comprises two or more inserts. As described herein, these inserts may be copies of the same sequence from a target nucleic acid or separate sequences from a target nucleic acid. As used herein, a “chimeric template” refers to a template comprising different inserts.

[0275] In addition to more than one insert and a hybridization sequence (or its complement), the present polynucleotides may also comprise a variety of other types of inserts. For example, a polynucleotide may comprise one or more sequencing primer sequences. Such sequencing primer sequences may be used for binding primers to initiate sequencing when the polynucleotides are used as sequencing templates. In some embodiments, a polynucleotide comprises a first read sequencing primer sequence and / or a second read sequencing primer sequence. As used herein “first read sequencing primer sequence” and “second read sequencing primer sequences” refer to sequences that can bind to a primer that may be used in different sequencing reads. These terms do not limit to any specific sequence, and, for example, a first read sequencing primer sequence may be used to initiate a second sequencing read in a given experiment and a second read sequencing primer may be used to initiate a first sequencing read in a given experiment. Such primer sequences may vary based on the sequencing platform that a user plans to utilize, and such primer sequences would be well- known in the art, such as A14 (SEQ ID NO: 4) and B15 sequences (SEQ ID NO: 5).

[0276] In some embodiments, the first read sequencing primer sequence and the second read sequencing primer sequence are different. In some embodiments, the first read sequencing primer sequence and the second read sequencing primer sequence each comprise an A14 sequence or a Bl 5 sequence, or their complements. In some embodiments, the 3’ terminal polynucleotide comprises the complement of a P5 primer sequence (P5’) and the 5’ terminal polynucleotide comprises a P7 primer sequence (P7, SEQ ID NO: 8), or the 3’ terminal polynucleotide comprises the complement of a P7 primer sequence (P7’) and the 5’ terminal polynucleotide comprises a P5 primer sequence (P5, SEQ ID NO: 7).

[0277] In some embodiments, the 3’ terminal polynucleotide and / or the 5’ terminal polynucleotide each independently comprise at least one of an adaptor, a barcode sequence, a unique molecular identifier (UMI) sequence, an index sequence (e.g., a sample-specific index) a capture sequence, or a cleavage sequence. In other words, polynucleotides may comprise additional sequences of use in methods that a user wants to perform, such as sequencing.

[0278] Using methods described herein, one insert in a polynucleotide may be prepared from a fragment comprising a portion of a sense strand of a target nucleic acid and the other insert is prepared by elongation from a fragment comprising a portion of an antisense strand of a target nucleic acid. Using methods described herein, one insert may be prepared from a fragment comprising a portion of an antisense strand of a target nucleic acid and the other insert is prepared by elongation from a fragment comprising a portion of a strand of a target nucleic acid.

[0279] In some embodiments, a polynucleotide comprises two insert sequences that are copies of each other. In some embodiments, a polynucleotide comprises a 5’ terminal polynucleotide comprising (a) a first read sequencing primer sequence; (b) an insert sequence derived from a target nucleic acid, wherein the insert sequence is 3’ of the 5’ terminal polynucleotide; (c) a hybridization sequence 3’ of the insert sequence; (d) a copy of the insert sequence 3’ of the hybridization sequence; and (e) a 3’ terminal polynucleotide comprising the complement of a second read sequencing primer sequence. In some embodiments, this polynucleotide may be a sequencing template. While the two copies of the insert (i.e., the insertsequence and the copy of the insert sequence) may be expected to be identical, sequencing results may indicate that they are not. For example, the two copies of the insert may be different based on a mismatch mutation in the target nucleic acid or based on introduction of an error during PCR amplification.

[0280] In some embodiments, a polynucleotide comprises two insert sequences that are not copies of each other. In some embodiments, the two insert sequences may be different. In some embodiments, the two insert sequences comprised in a polynucleotide were prepared from different regions of a target nucleic acid. In some embodiments, a polynucleotide comprises (a) a 5’ terminal polynucleotide comprising a first read sequencing primer sequence; (b) a first insert sequence derived from a target nucleic acid, wherein the insert sequence is 3’ of the 5’ terminal polynucleotide; (c) a hybridization sequence 3’ of the insert sequence; (d) a second insert sequence 3’ of the hybridization sequence; and (e) a 3’ terminal polynucleotide comprising the complement of a second read sequencing primer sequence. As described herein for methods with immobilized transposomes, such templates with two different insert sequences can be used to determine contiguity data.

[0281] The two inserts comprised in a polynucleotide may be the same of different sizes. In some embodiments, inserts that are copies comprise the same number of nucleotides. In some embodiments, the insert sequences comprise 40 to 400 nucleotides, optionally wherein the insert sequences comprise 1000 or fewer nucleotides. In some embodiments, a paired sequencing read protocol may be performed for a larger insert, such as one comprising more than 500 nucleotides.

[0282] In some embodiments, a polynucleotide is immobilized on a solid support. In some embodiments, the polynucleotide is immobilized on the solid support via the 5’ terminal polynucleotide (such as in the embodiment shown in Figure 29). In some embodiments, a polynucleotide is immobilized to the solid support via binding of an affinity moiety on the 5’ terminal polynucleotide to a binding moiety on the surface of the solid support. In some embodiments, an affinity moiety is attached via a linker to the 5’ terminal polynucleotide. In some embodiments, the affinity moiety is biotin, desthiobiotin, or dual biotin.

[0283] In some embodiments, a composition comprises a polynucleotide hybridized to its complement. In some embodiments, a polynucleotide hybridized to its complement may be termed a double-stranded concatenated sequencing template. In some embodiments, a doublestranded concatenated sequencing template is immobilized to the surface of a solid support by both of its 5’ ends.

[0284] In some embodiments, a polynucleotide or a composition comprising a polynucleotide and its complement is immobilized on the surface of a solid support, wherein the affinity moiety is biotin, desthiobiotin, or dual biotin and the binding moiety is avidin or streptavidin. A wide range of different solid support may be used for immobilization. In some embodiments, the solid support is a bead, slide, wall of a vessel, a flow cell, or a nanowell comprised in a flow cell.

[0285] In some embodiments, a linker for attaching an affinity moiety to a polynucleotide is a cleavable linker. In some embodiments, a user can release a polynucleotide from a solid support at a desired time by cleaving this cleavable linker.

[0286] In some embodiments, the adaptor may be a forked adaptor, also known as a Y- adaptor. Forked adaptor-based technology can be utilized for generating polynucleotides, for example, as exemplified in the workflow for TruSeq™ sample preparation kits (Illumina, Inc.). Reagents from the workflow for TruSight® Oncology kits (Illumina, Inc.) may also be used to assemble forked adaptors as disclosed herein. In some embodiments, a forked adaptor comprises a HYB or HYB’ sequence.

[0287] A forked adaptor is a nucleic acid complex including a double-stranded region that is complementary to the other strand and a region that is not complementary to the other strand. In some embodiments, each forked adaptor comprises a first oligonucleotide and a second oligonucleotide that are partially hybridized to each other to form a double-stranded section and a single stranded section.Transposition Reactions

[0288] In some embodiments, a polynucleotide is prepared via a method comprising a transposition reaction. A transposition reaction is a reaction wherein one or more transposons are inserted into target nucleic acids at random sites or almost random sites. Components in a transposition reaction include a transposase (or other enzyme capable of fragmenting and tagging a nucleic acid as described herein, such as an integrase) and a transposon element that includes a double-stranded transposon end sequence that binds to the transposase (or other enzyme as described herein), and an adaptor sequence attached to one of the two transposon end sequences. One strand of the double-stranded transposon end sequence is transferred to one strand of the target nucleic acid and the complementary transposon end sequence strand is not (a non-transferred transposon sequence). The adaptor sequence can include one or more functional sequences or components (e.g., primer sequences, anchor sequences, universal sequences, spacer regions, or index tag sequences) as needed or desired.

[0289] Transposon based technology can be utilized for fragmenting DNA, for example, as exemplified in the workflow for NEXTERA™ FLEX DNA sample preparation kits (Illumina, Inc.), wherein target nucleic acids, such as genomic DNA, are treated with transposome complexes that simultaneously fragment and tag (“tagmentation”) the target, thereby creating a population of fragmented nucleic acid molecules tagged with unique adaptor sequences at the ends of the fragments. In some embodiments, bead-linked transposomes (BLTs) are used. In some embodiments, the reactions, transposomes in solution are used.

[0290] A “transposome complex” is comprised of at least one transposase (or other enzyme as described herein) and a transposon recognition sequence. In some such systems, the transposase binds to a transposon recognition sequence to form a functional complex that is capable of catalyzing a transposition reaction. In some aspects, the transposon recognition sequence is a double-stranded transposon end sequence. The transposase binds to a transposase recognition site in a target nucleic acid and insert sequences the transposon recognition sequence into a target nucleic acid. In some such insertion events, one strand of the transposon recognition sequence (or end sequence) is transferred into the target nucleic acid, resulting in a cleavage event. Exemplary transposition procedures and systems that can be readily adaptedfor use with the transposases. Exemplary transposases that can be used with certain embodiments provided herein include (or are encoded by): Tn5 transposase, Sleeping Beauty (SB) transposase, Vibrio harveyi, MuA transposase and a Mu transposase recognition site comprising R1 and R2 end sequences, Staphylococcus aureus Tn552, Tyl, Tn7 transposase, Tn / O and IS 10, Mariner transposase, Tel, P Element, Tn3, bacterial insertion sequences, retroviruses, and retrotransposon of yeast. More examples include IS5, TnlO, Tn903, IS911, and engineered versions of transposase family enzymes. The methods described herein could also include combinations of transposases, and not just a single transposase.

[0291] In some embodiments, the transposase is a Tn5, Tn7, MuA, or Vibrio harveyi transposase, or an active mutant thereof. In other embodiments, the transposase is a Tn5 transposase or a mutant thereof. In other embodiments, the transposase is a Tn5 transposase or a mutant thereof. In other embodiments, the transposase is a Tn5 transposase or an active mutant thereof. In some embodiments, the Tn5 transposase is a hyperactive Tn5 transposase, or an active mutant thereof. In some aspects, the Tn5 transposase is a Tn5 transposase as described in PCT Publ. No. WO2015 / 160895, which is incorporated herein by reference. In some aspects, the Tn5 transposase is a hyperactive Tn5 with mutations at positions 54, 56, 372, 212, 214, 251, and 338 relative to wild-type Tn5 transposase. In some aspects, the Tn5 transposase is a hyperactive Tn5 with the following mutations relative to wild-type Tn5 transposase: E54K, M56A, L372P, K212R, P214R, G251R, and A338V. In some embodiments, the Tn5 transposase is a fusion protein. In some embodiments, the Tn5 transposase fusion protein comprises a fused elongation factor Ts (Tsf) tag. In some embodiments, the Tn5 transposase is a hyperactive Tn5 transposase comprising mutations at amino acids 54, 56, and 372 relative to the wild type sequence. In some embodiments, the hyperactive Tn5 transposase is a fusion protein, optionally wherein the fused protein is elongation factor Ts (Tsf). In some embodiments, the recognition site is a Tn5-type transposase recognition site. In one embodiment, a transposase recognition site that forms a complex with a hyperactive Tn5 transposase is used (e.g., EZ-Tn5TM Transposase, Epicentre Biotechnologies, Madison, Wis.). In some embodiments, the Tn5 transposase is a wild-type Tn5 transposase.

[0292] In some embodiments, the transposome complex comprises a dimer of two molecules of a transposase. In some embodiments, the transposome complex is a homodimer, wherein two molecules of a transposase are each bound to first and second transposons of the same type (e.g., the sequences of the two transposons bound to each monomer are the same, forming a “homodimer”). In some embodiments, the compositions and methods described herein employ two populations of transposome complexes. In some embodiments, the transposases in each population are the same. In some embodiments, the transposome complexes in each population are homodimers, wherein the first population has a first adaptor sequence in each monomer and the second population has a different adaptor sequence in each monomer. Example transposomes are shown in FIG. 71. The example transposomes may be used in conjunction with certain embodiments discussed herein. However, it should be understood that the first adaptor and the second adaptor in the illustrated example may be exchanged with other first adaptors and second adaptors discussed herein to generate adaptor- linked fragments via transposition or tagmentation reactions. FIG. 72 shows an example tagmentation to generate adaptor-linked fragments as discussed herein.

[0293] In some embodiments, the transposase complex comprises a transposase (e.g., a Tn5 transposase) dimer comprising a first and a second monomer. In some aspects, each monomer comprises a first transposon, a second transposon, and an attachment polynucleotide, where the first transposon includes a transposon end sequence at its 3’ end (also referred to as a 3’ transposon end sequence) and an adaptor sequence at its 5’ end (also referred to as a 5’ adaptor sequence); the second transposon includes a transposon end sequence at its 5’ end (also referred to as a 5’ transposon end sequence) and an adaptor sequence at its 3’ end (also referred to as a 3’ adaptor sequence); and the attachment polynucleotide includes an attachment adaptor sequence hybridized to the 5’ adaptor sequence of the first transposon, a primer sequence, and a linker. In some embodiments, the 5’ transposon end sequence of the second transposon is at least partially complementary to the 3’ transposon end sequence of the first transposon. In some embodiments, the attachment adaptor sequence of the attachment polynucleotide is at least partially complementary to the 5’ adaptor sequence of the first transposon. In some embodiments, the linker of the attachment polynucleotide includes a binding element.

[0294] In some embodiments, a transposome complex comprises one or more adaptors as discussed herein. In an embodiment, a transposome complex composition includes a first adaptor type and a second adaptor type in a heterodimer arrangement or a homodimer arrangement. In some embodiments, the 3’ transposon end sequence comprises a mosaic end (ME) sequence and the 5’ transposon end sequence comprises an ME’ sequence, e.g., a stem region or double-stranded region of an adaptor.

[0295] In some embodiments, the first and second transposons as described herein are annealed to each other, and the first transposon is annealed to the attachment polynucleotide. The annealed polynucleotides are then loaded onto a transposase, such as a Tn5 transposase, thereby forming a transposome complex, which is then contacted with and bound to a solid support, such as a bead. In some embodiments, the annealed transposons are bound to a solid support such as a bead and a transposase is then complexed with the transposons, thereby creating a transposome that is bound to a solid support.

[0296] In some embodiments, the first transposon includes a 3’ transposon end sequence and the second transposon includes a 5’ transposon end sequence. In some embodiments, the 5’ transposon end sequence is at least partially complementary to the 3’ transposon end sequence. In some embodiments, the complementary transposon end sequences hybridize to form a double-stranded transposon end sequence that binds to the transposase (or other enzyme as described herein). In some embodiments, the transposon end sequence is a mosaic end (ME) sequence. Thus, in some embodiments, the 3’ transposon end sequence is an ME sequence and the 5’ transposon end sequence is an ME’ sequence.

[0297] In some embodiments, the transposome complex is immobilized to a solid support via the first or second transposon. In some embodiments, the transposome complex is immobilized on a bead. In some embodiments, the transposome complex is immobilized on a bead via the first or second transposon. The terms solid surface, solid support, or substrate may include, but are not limited to, glass and modified or functionalized glass, plastics (including acrylics, polystyrene and copolymers of styrene and other materials, polypropylene, polyethylene, polybutylene, polyurethanes, TEFLON, etc.), polysaccharides, polyhedralorganic silsesquioxane (POSS) materials, nylon or nitrocellulose, ceramics, resins, silica, or silica-based materials including silicon and modified silicon, carbon, metals, inorganic glasses, plastics, optical fiber bundles, beads, paramagnetic beads, and a variety of other polymers.

[0298] In some embodiments, the transposome complex is immobilized on the solid support via a binding element (and optional linker). In some embodiments, the solid support is a bead, a paramagnetic bead, a flowcell, a surface of a microfluidic device, a tube, a well of a plate, a slide, a patterned surface, or a microparticle. In some embodiments, the solid support comprises or is a bead. In one embodiment, the bead is a paramagnetic bead. In some embodiments, the solid support comprises a plurality of solid supports. In some embodiments, transposome complexes are immobilized on a plurality of solid supports. In some embodiments, the plurality of solid supports comprises a plurality of beads. In some embodiments, the plurality of transposome complexes are immobilized on the solid support at a density of at least 103, 104, 105, 106 complexes per mm2. In some embodiments, the solid support is a bead or a paramagnetic bead, and there are greater than 10,000, 20,000, 30,000, 40,000, 50,000, or 60,000 transposome complexes bound to each bead.

[0299] Suitable bead compositions include, but are not limited to, plastics, ceramics, glass, polystyrene, methylstyrene, acrylic polymers, paramagnetic materials, thoria sol, carbon graphite, titanium dioxide, latex or cross-linked dextran such as Sepharose, cellulose, nylon, cross-linked micelles and TEFLON, as well as any other materials outlined herein for solid supports. In certain embodiments, the microspheres are magnetic microspheres or beads, for example paramagnetic particles, spheres or beads. The beads need not be spherical; irregular particles may be used. Alternatively or additionally, the beads may be porous. The bead sizes ranee from nanometers, e g., 100 nm, to millimeters, e ., 1 mm, with beads from 0.2 micron to 200 microns being preferred, and from 0.5 to 5 micron being particularly preferred, although in some embodiments smaller or larger beads may be used. The bead may be coated with a binding partner, for example the bead may be streptavidin coated. In some embodiments, the beads are streptavidin coated paramagnetic beads, for example, Dynabeads MyOne streptavidin Cl beads (Thermo Scientific catalog # 65601), Streptavidin MagneSphereParamagnetic particles (Promega catalog #Z5481), Streptavidin Magnetic beads (NEB catalog # S1420S) and MaxBead Streptavidin (Abnova catalog # U0087). The solid support could also be a slide, for example a flowcell or other slide that has been modified such that the transposome complex can be immobilized thereon.

[0300] In some embodiments, the binding partner is present on the solid support or bead at a density of from 1000 to 6000 pmol / mg, or 2000 to 5000 pmol / mg, or 3000 to 5000 pmol / mg, or 3500 to 4500 pmol / mg.

[0301] In some embodiments, the solid surface is the inner surface of a sample tube. In some embodiments, the solid surface is a capture membrane. In one example, the capture membrane is a biotin-capture membrane (for example, available from Promega Corporation). In some embodiments, the capture membrane is filter paper. In some embodiments of the present disclosure, solid supports comprised of an inert substrate or matrix (e.g. glass slides, polymer beads etc.) which has been functionalized, for example by application of a layer or coating of an intermediate material comprising reactive groups which permit covalent attachment to molecules, such as polynucleotides. Examples of such supports include, but are not limited to, polyacrylamide hydrogels supported on an inert substrate such as glass, particularly polyacrylamide hydrogels as described in W02005 / 065814 and US2008 / 0280773, the contents of which are incorporated herein in their entirety by reference. The methods of tagmenting (fragmenting and tagging) DNA on a solid surface for the construction of a tagmented DNA library are described in WO2016 / 189331 and US2014 / 0093916A1, which are incorporated herein by reference in their entireties. In some embodiments, the transposome complex described herein is immobilized to a solid support via the binding element. In some such embodiments, the solid support comprises streptavidin as the binding partner and the binding element is biotin.

[0302] An affinity moiety can be used to bind, covalently or non-covalently, to a binding partner. The affinity moiety may be, for example, biotin, and the binding partner comprises or is avidin or streptavidin. In other embodiments, the binding element / binding partner combination comprises or is FITC / anti-FITC, digoxigenin / digoxigenin antibody, orhapten / antibody. Further suitable binding pairs include, but not limited to, desthiobiotinavidin, dithiobiotin-avidin, iminobiotin-avidin, biotin-avidin, dithiobiotin-succinilated avidin, iminobiotin-succinilated avidin, biotin-streptavidin, and biotin-succinilated avidin. In some embodiments, the binding element is a biotin and the binding partner is streptavidin.

[0303] In some embodiments, the binding element can bind to the binding partner via a chemical reaction or is bound covalently by reaction with the binding partner on the solid support, thereby covalently attaching the transposome complex to the solid support. In some aspects, the binding element / binding partner combination comprises or is amine / carboxylic acid (e.g., binding via standard peptide coupling reaction under conditions known to one of ordinary skill in the art, such as EDC or NHS-mediated coupling). The reaction of the two components joins the binding element and binding partner through an amide bond. Alternatively, the binding element and binding partner can be two click chemistry partners (e g., azide / alkyne, which react to form a triazole linkage).Kits

[0304] In some embodiments, an adaptor composition or kit comprises an adaptor or adaptor set as discussed herein. In addition, a kit may include reagents for reactions used in conjunction with the adaptor or adaptor set as discussed herein. Suitable reagents include dNTPs (e.g., natural or modified dNTPs) polymerases, buffers, stop buffers, wash buffers, resuspension buffers, separation beads for one or more separation and / or wash steps, nucleotide modification agents (e.g., sodium bisulfite, APOBEC or deaminase enzymes). In some embodiments, a kit may comprise solid support such as beads. Kits may include compartments or packaging to separate reaction components for different workflow steps.Nucleic Acids and Modified Nucleic Acids

[0305] The disclosed techniques may include workflow steps that operate on nucleic acids, such as individual nucleotides or oligonucleotides. As used herein, the terms “polynucleotide”, “oligonucleotide”, “nucleic acid”, and “nucleic acid sequence” may refer tosingle-stranded and double-stranded polymers of nucleotide monomers, including 2'- deoxyribonucleotides (DNA) and ribonucleotides (RNA) linked by internucleotide phosphodiester bond linkages, or intemucleotide analogs, and associated counter ions, e.g., H+, NH4+, trialkylammonium, tetraalkylammonium, Mg2, Na+and the like. A nucleic acid may be composed entirely of deoxyribonucleotides, entirely of ribonucleotides, or chimeric mixtures thereof. The nucleotide monomer units may comprise any of the nucleotides described herein, including, but not limited to, naturally occurring nucleotides and nucleotide analogs. Nucleic acids typically range in size from a few monomeric units, e.g. 5-40 when they are sometimes referred to in the art as oligonucleotides, to several thousands of monomeric nucleotide units. Nucleic acid sequence are shown in the 5’ to 3' orientation from left to right, unless otherwise apparent from the context or expressly indicated differently; and in such sequences, “A” denotes deoxyadenosine, “C” denotes deoxycytidine, “G” denotes deoxyguanosine, “T” denotes thymidine, and “U” denotes uridine.

[0306] The term “nucleotide analogs” refers to synthetic analogs having modified nucleotide base portions, modified pentose portions, and / or modified phosphate portions. Generally, modified phosphate portions comprise analogs of phosphate wherein the phosphorous atom is in the +5 oxidation state and one or more of the oxygen atoms is replaced with a non-oxygen moiety, e.g., sulfur. Exemplary phosphate analogs include but are not limited to phosphorothioate, phosphorodithioate, phosphoroselenoate, phosphorodi selenoate, phosphoroanilothioate, phosphoranilidate, phosphoramidate, boronophosphates, including associated counterions, e.g., H+, NEU+, Na+, if such counterions are present. Exemplary modified nucleotide base portions include but are not limited to 5-methylcytosine (5mC); C-5-propynyl analogs, including but not limited to, C- 5 propynyl-C and C-5 propynyl-U; 2,6-diaminopurine, also known as 2-amino adenine or 2- amino-dA); hypoxanthine, pseudouridine, 2-thiopyrimidine, isocytosine (isoC), 5-methyl isoC, and isoguanine (isoG; see, e.g., U.S. Pat. No. 5,432,272). Exemplary modified pentose portions include but are not limited to, locked nucleic acid (LNA) analogs including without limitation Bz-A-LNA, S-Me-Bz-C-LNA, dmf-G-LNA, and T-LNA , and 2'- or 3'- modifications where the 2'- or 3'-position is hydrogen, hydroxy, alkoxy (e.g., methoxy, ethoxy,allyloxy, isopropoxy, butoxy, isobutoxy and phenoxy), azido, amino, alkylamino, fluoro, chloro, or bromo. Modified internucleotide linkages include phosphate analogs, analogs having achiral and uncharged intersubunit linkages, and uncharged morpholino-based polymers having achiral intersubunit linkages (see, e.g., U.S. Pat. No. 5,034,506). Some internucleotide linkage analogs include morpholidate, acetal, and polyamide-linked heterocycles. In one class of nucleotide analogs, known as peptide nucleic acids, including pseudocomplementary peptide nucleic acids (“PNA”), a conventional sugar and internucleotide linkage has been replaced with a 2-aminoethylglycine amide backbone polymer. The term “Tmenhancing nucleotide analog” as used herein refers to a nucleotide analog that, when incorporated into a primer or extension product, increases the annealing temperature of that primer or extension product relative to a primer or extension product with the same sequence comprising conventional nucleotides (A, C, G, and / or T), but not the Tmenhancing nucleotide analog. Those in the art will appreciate that Tm can be determined experimentally using well-known methods or can be estimated using algorithms, thus one can readily determine whether a particular nucleotide analog will serve as a Tm enhancing nucleotide analog when used in a particular context, without undue experimentation. A wide range of nucleotide analogs are available as triphosphates, phoshoramidites, or CPG derivatives for use in enzymatic incorporation or chemical synthesis from, among other sources, Glen Research, Sterling, Md.; Link Technologies, Lanarkshire, Scotland, UK; and TriLink BioTechnologies, San Diego, Calif.

[0307] In one embodiment, a modified nucleotide may be a photocleavable nucleotide, such as 6-nitropiperonyl methyl group (NPM) on N4-dC and its corresponding hydroxymethylene analogs (NPOM) on N3-dT, N3-U and N’-dG. These groups may be photocleaved at appropriate photocleavage wavelengths (~365 nm).Nucleic Acid Manipulation and Modification

[0308] The disclosed techniques include one or more nucleic acid manipulation steps such as annealing or hybridization, washing, denaturing, amplifying, ligating, linearizing, nicking, cutting, end-blocking, phosphorylating, deaminating, coverting, and so on. The disclosedembodiments encompass suitable reagents (e.g., enzymes) and conditions under which the manipulation steps occur.

[0309] Hybridization, binding, or annealing of nucleic acids may refer to physical interaction of complementary (including partially complementary) polynucleotide strands by the formation of hydrogen bonds between complementary nucleotides when the strands are arranged antiparallel to each other. Hybridization and the strength of hybridization (e.g., the strength of the association between polynucleotides) is impacted by many factors well known in the art including the degree of complementarity between the polynucleotides, and the stringency of the conditions involved, which is affected by such conditions as the concentration of salts, the presence of other components (e.g., the presence or absence of polyethylene glycol), the molarity of the hybridizing strands and the G+C content of the polynucleotide strands, all of which results in a characteristic melting or denaturing temperature (Tm) of the formed hybrid. The terms “hybridization (hybridize)” and “binding,” when used in reference to nucleic acids, can be used interchangeably and can refer to the process by which single strands of nucleic acid sequences form double-helical segments through hydrogen bonding between complementary nucleotides. “Hybrid,” “duplex,” and “complex,” when used in reference to nucleic acids, can also be used interchangeably herein referring to a double-stranded nucleic acid molecule formed by hybridization (e.g., DNA- DNA, DNA-RNA, and RNA-RNA species). A variety of hybridization or washing conditions can be used in the methods disclosed herein. For example, hybridization complexes are immobilized on a solid support and washed under conditions sufficient to remove nonhybridized nucleic acids, i.e. non-hybridized probes and sample nucleic acids. In a particularly preferred embodiment immobilized complexes are washed under conditions sufficient to remove imperfectly hybridized complexes. Hybridization or washing conditions are well known in the art and can be found described in, for example, Sambrook et al., Molecular Cloning: A Laboratory Manual, Third Ed., Cold Spring Harbor Laboratory, New York (2001) and in Ansubel et al., Current Protocols in Molecular Biology, John Wiley and Sons, Baltimore, Md. (1999). Stringency of the hybridization or washing conditions include variations in temperature or buffer composition and can be varied according to the specificityof the reaction needed. A range of stringency includes, for example, high, moderate or low stringency conditions.

[0310] A cleavable site may refer to a moiety, such as a modified nucleotide, that allows selective cleavage to separate portions of a nucleic acid at the cleavable moiety. For example, a contiguous nucleic acid may be separated using a cleavable moiety. By way of non-limiting example, the cleavable site may comprise uracil bases, phosphorothioate groups, ribonucleotides, diol linkages, disulphide linkages, peptides etc. In one example, the cleavable site is a uracil. Uracil can be cleaved using a uracil glycosylase or USER enzyme mix (which is a cocktail of uracil glycosylase and endonuclease VIII). In another example, the cleavable site is 8-oxoguanine. 8- oxoguanine can be cleaved using a FPG glycosylase. Alternatively, the cleavable site is a restriction site.

[0311] A restriction site may refer to a sequence of nucleotides recognized by an endonuclease, such as a single-stranded endonuclease. A restriction site may also be referred to as a “recognition site” or “recognition sequence”, and such terms may be used interchangeably. In one embodiment, the endonuclease is a single strand restriction endonuclease, a nicking endonuclease or nicking enzyme or nickase (again, such terms may be used interchangeably). By any of these terms is meant an enzyme that can hydrolyze only one strand of the double-stranded polynucleotide (duplex), to produce DNA molecules that are “nicked”, rather than fully cleaved on both strands. Certain methods disclosed herein may include nicking enzymes, such as a nicking endonuclease or restriction endonucleases. Suitable nicking enzymes that may be used includeNb.BbvCI, Nb.Bsml, Nb.BsrDI, Nb.BtsI, Nt.Alwl, Nt.BsmAI, Nt.BspQI, Nt.BstNBI, BssSI, Nb.BpulOl and Nt.CviPI Exemplary restriction endonucleases include but are not limited to I-secI, EcoRI, EcoRII, BamHI, Hind III, TaqI, Notl. Other examples of restriction endonucleases can be found in New England Biolabs catalog (New England Biolabs, MA, USA).

[0312] Clustered Regularly Interspaced Short Palindromic Repeats (CRISPRs) are involved in interfering pathways that protect cells from phage and conjugative plasmids in many bacteria and archaea. In some embodiments, nucleic acid modification is performedusing a CRISPR Cas method. In some embodiments, a CRISPR Cas cut site is present in the sequence of interest. A CRISPR-Cas system may include a guide RNA (gRNA) sequence that includes an oligonucleotide sequence that is complementary or substantially complementary to a sequence within a target polynucleotide and a Cas protein. Cas proteins may have a variety of activities, such as nuclease activity. Thus, CRISPR-Cas systems provide a mechanism for targeting specific sequences (e.g., via gRNA) as well as certain enzymatic activities, such as cleaving on the sequences (e.g., via specific activity of Cas proteins).

[0313] Embodiments of the present disclosure relate to preparing nucleic acids for sequencing or other applications. In particular, embodiments of the proteins, methods, compositions, and kits provided herein relate to mapping of methylation status by using sequencing libraries and other methods. Certain methods of methylation analysis, including the workflows, enzymes, and relevant reagents, may be as discussed in WO2023175037A2, which is incorporated by reference in its entirety. Bisulfite sequencing (BS-seq) involves using bisulfite as the conversion agent. This process is described in Frommer et al. (Proc. Natl. Acad. Sci. U.S.A., 1992, 89, pp. 1827- 1831), which is incorporated herein by reference. This process converts unmodified cytosines in the target polynucleotide to uracil, as well as 5- formylcytosine and 5- carboxylcytosine to deaminated analogues, but does not convert 5- m ethyl cytosine and 5-hydroxymethylcytosine. Accordingly, BS-seq allows identification of the modified cytosines 5-mC and 5-hmC by reading them as C; whereas unmodified C, 5-fC and 5- caC are converted to nucleobases which are read as T / U. APOBEC-coupled epigenetic sequencing (ACE-seq) involves using a T4 bacteriophage P-glucosyltransferase as a further agent and APOBEC3A as the conversion agent. This process is described in Schutsky et al. (Nat. Biotechnol., 2018, 36, pp. 1083-1090), which is incorporated herein by reference. The T4 bacteriophage p-glucosyltransferase converts 5-hydroxymethylcytosine in the target polynucleotide to p-glucosyl-5- hydroxymethylcytosine, which prevents oxidation. Subsequent treatment with APOBEC3A converts unmodified cytosines in the target polynucleotide to uracil, as well as 5-methylcytosine to its deaminated analogue. 5- formylcytosine is also able to convert to its deaminated analogue, but reacts slower relative to unmodified cytosine and 5- methylcytosine. 5-carboxylcytosine is also able to convert to itsdeaminated analogue, but reacts far slower than unmodified cytosine and 5-methylcytosine, and slower than 5- formylcytosine. Accordingly, ACE-seq allows identification of the modified cytosine 5- hmC (as the protected glycosyl residue) by reading it as C; whereas unmodified C and 5-mC are converted to nucleobases which are read as T / ll; 5-fC is converted to a nucleobase which is read as T / ll to a limited extent; 5-caC is converted to a nucleobase which is read as T / ll to a more limited extent. Enzymatic Methyl sequencing (EM-seq) involves using T4 bacteriophage p- glucosyltransferase and a TET2 enzyme as the further agents and AP0BEC3A as the conversion agent. This process is described in Vaisvila et al. (Genome Res. 2021 , 31 , pp. 1280-1289), US 10,619,200 B2 and US 9,121 ,061 B2, which are incorporated herein by reference. The T4 bacteriophage p-glucosyltransferase converts 5- hydroxymethylcytosine in the target polynucleotide to p-glucosyl-5- hydroxymethylcytosine, which prevents oxidation. The TET2 enzyme causes oxidation of 5-methylcytosine in the target polynucleotide to 5-hydroxymethylcytosine, which in turn is converted to p-glucosyl-5- hydroxymethylcytosine by the T4 bacteriophage p- glucosyltransferase. The TET2 enzyme also causes oxidation of 5-formylcytosine in the target polynucleotide to 5-carboxylcytosine. Subsequent treatment with APOBEC3A converts unmodified cytosines in the target polynucleotide to uracil, as well as 5- carboxylcytosine (including residues that used to be 5- formylcytosine) to a limited extent. Accordingly, EM-seq allows identification of the modified cytosines 5-mC and 5-hmC (as protected glycosyl residues) by reading them as C; whereas unmodified C is converted to U; 5fC and 5-caC are converted to nucleobases which are read as T / U to a limited extent.

[0314] Modified APOBEC sequencing involves using a mutant APOBEC3A enzyme as the conversion agent. This process is described in US Provisional Application 63 / 328,444, which is incorporated herein by reference. TAPS TET-assisted pyridine borane sequencing (TAPS) involves using a TET1 enzyme as the further agent and pyridine borane as the conversion agent. This process is described in Liu et al. (Nature Biotechnology, 2019, 37, pp. 424-429), which is incorporated herein by reference. The TET1 enzyme causes oxidation of 5-methylcytosine, 5- hydroxymethylcytosine and 5-formylcytosine in the target polynucleotide to 5- carboxylcytosine. Subsequent treatment with pyridine borane converts 5-carboxylcytosine (including residues that used to be 5-methylcytosine, 5- hydroxymethylcytosine and 5-formylcytosine) to dihydrouracil, but does not convert unmodified cytosine. Accordingly, TAPS allows identification of the modified cytosines 5- mC, 5-hmC, 5-fC and 5-caC by reading them as T / ll; whereas unmodified cytosine is read as C. TET-assisted pyridine borane sequencing with p-glucosyltransferase blocking (TAPSP) involves using a T4 p-glucosyltransferase and a TET 1 enzyme as the further agents, and pyridine borane as the conversion agent. This process is described in Liu et al. (Nature Communications, 2021 , 12, 618), which is incorporated herein by reference. The T4 p- glucosyltransferase converts 5-hydroxymethylcytosine in the target polynucleotide to p- glucosyl-5-hydroxymethylcytosine, which prevents oxidation. The TET 1 enzyme causes oxidation of 5-methylcytosine and 5-formylcytosine in the target polynucleotide to 5- carb oxy 1 cytosine. Subsequent treatment with pyridine borane converts 5- carboxylcytosine (including residues that used to be 5-methylcytosine and 5- formylcytosine) to dihydrouracil, but does not convert unmodified cytosine or p-glucosyl- 5-hydroxymethylcytosine. Accordingly, TAPSp allows identification of the modified cytosines 5-mC, 5-fC and 5-caC by reading them as T / U; whereas unmodified cytosine and 5-hmC are read as C.

[0315] In certain embodiments, deaminase enzymes are provided that selectively act on certain modified cytosines of target nucleic acids and converts them to thymidine or modified thymidine analogues . The deaminase enzymes may act to selectively act on double-stranded or single-stranded nucleic acids. For example, a deaminase with selective action on singlestranded nucleic acids does not act on modified cytosines of double-stranded nucleic acids, and vice vera. These enzymes may include an altered cytidine deaminase comprising amino acid substitution mutations in a cytidine deaminase at positions functionally equivalent to (Tyr / Phe)130 and Tyrl32 in a wild-type APOBEC3A protein. In an embodiment, the deaminase enzyme may be as described in US20240182881A1, W02025072800A2, WO2025072783A1, WO2025072793A1, which are incorporated by reference in their entireties herein for all purposes.

[0316] Methods may include linearizing via abasic sites generated at non-natural / modified deoxyribonucleotides and cleaved by treatment with endonuclease, heat or alkali. For example, 8-oxo-guanine can be converted to an abasic site by exposure to FPG glycosylase. Deoxyinosine can be converted to an abasic site by exposure to AlkA glycosylase. The abasic sites thus generated may then be cleaved, typically by treatment with a suitable endonuclease (e g. EndoIV, AP lyase). A further example includes the use of a specific non methylated cytosine. If the remainder of the primer sequences, and the dNTP's used in cluster formation are methylated cytosine residues, then the non-methylated cytosines can be specifically converted to uracil residues by treatment with bisulfite. This allows the use of the same ‘USER’ treatment to linearise both strands of the cluster, as one of the primers may contain a uracil, and one may contain the cytosine that can be converted into a uracil (effectively a ‘protected uracil’ species). If the non-natural / modified nucleotide is to be incorporated into an amplification primer for use in solid-phase amplification, then the non-natural / modified nucleotide should be capable of being copied by the polymerase used for the amplification reaction.

[0317] In one embodiment, the molecules to be cleaved may be exposed to a mixture containing the appropriate glycosylase and one or more suitable endonucleases. In such mixtures the glycosylase and the endonuclease will typically be present in an activity ratio of at least about 2:1.

[0318] This method of cleavage has particular advantages in relation to the creation of templates for nucleic acid sequencing. In particular, cleavage of an abasic site generated by treatment with a reagent such as USER automatically releases a free 3' phosphate group on the cleaved strand which after phosphatase treatment can provide an initiation point for sequencing a region of the complementary strand. Moreover, if the initial double-stranded nucleic acid contains only one cleavable (e.g. uracil) base on one strand then a single “nick” can be generated at a unique position in this strand of the duplex. Since the cleavage reaction requires a residue, e.g. deoxyuridine, which does not occur naturally in DNA, but is otherwiseindependent of sequence context, if only one non-natural base is included there is no possibility of glycosylase-mediated cleavage occurring elsewhere at unwanted positions in the duplexTarget Nucleic Acid

[0319] Target nucleic acids used herein can be composed of DNA, RNA or analogs thereof. The source of the target nucleic acids can be genomic DNA, messenger RNA, or other nucleic acids from native sources. In some cases, the target nucleic acids that are derived from such sources can be amplified prior to use in a method or composition herein. In some embodiments, target nucleic acids can be obtained as fragments of one or more larger nucleic acids. Fragmentation can be carried out using any of a variety of techniques known in the art including, for example, nebulization, sonication, chemical cleavage, enzymatic cleavage, or physical shearing. Fragmentation may also result from use of a particular amplification technique that produces amplicons by copying only a portion of a larger nucleic acid. For example, PCR amplification produces fragments having a size defined by the length of the fragment between the flanking primers used for amplification.

[0320] A population of target nucleic acids, or amplicons thereof, can have an average strand length that is desired or appropriate for a particular application of the methods or compositions set forth herein. For example, the average strand length can be less than 100,000 nucleotides, 50,000 nucleotides, 10,000 nucleotides, 5,000 nucleotides, 1,000 nucleotides, 500 nucleotides, 100 nucleotides, or 50 nucleotides. Alternatively or additionally, the average strand length can be greater than 10 nucleotides, 50 nucleotides, 100 nucleotides, 500 nucleotides, 1,000 nucleotides, 5,000 nucleotides, 10,000 nucleotides, 50,000 nucleotides, or 100,000 nucleotides. The average strand length for population of target nucleic acids, or amplicons thereof, can be in a range between a maximum and minimum value set forth above. It will be understood that amplicons generated at an amplification site (or otherwise made or used herein) can have an average strand length that is in a range between an upper and lower limit selected from those exemplified above.

[0321] In some embodiments, the target nucleic acids have a relatively short average strand length, such as less than 200 nucleotides, less than 150 nucleotides, less than 100 nucleotides, less than 75 nucleotides, less than 50 nucleotides, or less than 36 nucleotides. Sequencing of target nucleic acids with relatively short average strand length are not limited by read-length, and increasing the number of reads could significantly increase sequencing output. Examples of sample types with relatively short average strand length are cell-free DNA (cfDNA) and exome sequencing sample.

[0322] In some embodiments, the target nucleic acids are cell-free DNA (cfDNA) from a maternal blood sample. In some embodiments, the cfDNA is extracted from a maternal plasma sample. In some embodiments, the cfDNA is for noninvasive prenatal testing (NIPT).

[0323] In some embodiments, the target nucleic acids are exomes. In some embodiments, exomes are prepared via targeted resequencing. In some embodiments, exomes are prepared by whole-genome enrichment. In some embodiments, exomes are prepared by hybridizationbased enrichment.

[0324] In some embodiments, the target nucleic acids are DNA and RNA. Separate libraries of RNA and DNA can be prepared to generate hybrid DNA / RNA polynucleotides. In some embodiments, polynucleotides comprise one or more insert comprising RNA and one or more insert comprising DNA. Such polynucleotides comprising RNA insert(s) and DNA insert(s) can be termed “hybrid polynucleotides” and allow multiple readouts to be generated from a single sequencing run. In some embodiments, polynucleotides comprising RNA and DNA inserts have a dual sample index to allow for self-normalizing. In some embodiments, the minimum of DNA or RNA in the starting libraries dictates the amount of hybrid polynucleotides generated.

[0325] Any of a variety of known amplification techniques can be used to increase the amount of template sequences present for use in a method set forth herein. Exemplary techniques include, but are not limited to, polymerase chain reaction (PCR), rolling circle amplification (RCA), multiple displacement amplification (MDA), or random primeamplification (RPA) of nucleic acid molecules having template sequences. It will be understood that amplification of target nucleic acids prior to use in a method or composition set forth herein is optional. As such, target nucleic acids will not be amplified prior to use in some embodiments of the methods and compositions set forth herein. Target nucleic acids can optionally be derived from synthetic libraries. Synthetic nucleic acids can have native DNA or RNA compositions or can be analogs thereof. Solid-phase amplification methods can also be used, including for example, cluster amplification, bridge amplification or other methods set forth below in the context of array -based methods.

[0326] In some embodiments, the polynucleotides disclosed herein can be sequenced using any suitable nucleic acid sequencing platform to determine the nucleic acid sequence of the target sequence. In some respects, sequences of interest are correlated with or associated with one or more congenital or inherited disorders, pathogenicity, antibiotic resistance, or genetic modifications. Sequencing may be used to determine the nucleic acid sequence of a short tandem repeat, single nucleotide polymorphism, gene, exon, coding region, exome, or portion thereof. As such, the methods and compositions described herein relate to methods useful in, but not limited to, cancer and disease diagnosis, prognosis and therapeutics, DNA fingerprinting applications (e.g., DNA databanking, criminal casework), metagenomic research and discovery, agrigenomic applications, and pathogen identification and monitoring.

[0327] In some embodiments, a sample used to prepare sequencing templates comprises double-stranded nucleic acid. This double-stranded nucleic acid may be referred to as target nucleic acid. In some embodiments, a double-stranded nucleic acid may be added to a solid support comprising immobilized transposomes. In some embodiments, a double-stranded nucleic acid may be fragmented and combined with a mixture of forked adaptors.

[0328] In some embodiments, a sample comprises multiple double-stranded nucleic acids.

[0329] A biological sample used in accordance with the present disclosure can be any type that comprises target nucleic acids. However, the sample need not be completely purified, and can comprise, for example, nucleic acid mixed with protein, other nucleic acid species, othercellular components, and / or any other contaminant. In some embodiments, the biological sample comprises a mixture of nucleic acid, protein, other nucleic acid species, other cellular components, and / or any other contaminant present in approximately the same proportion as found in vivo. For example, in some embodiments, the components are found in the same proportion as found in an intact cell. In some embodiments, the biological sample has a 260 / 280 absorbance ratio of less than or equal to 2.0, 1.9, 1.8, 1.7, 1.6, 1.5, 1.4, 1.3, 1.2, 1.1, 1.0, 0.9, 0.8, 0.7, or 0.60. In some embodiments, the biological sample has a 260 / 280 absorbance ratio of at least 2.0, 1.9, 1.8, 1.7, 1.6, 1.5, 1.4, 1.3, 1.2, 1.1, 1.0, 0.9, 0.8, 0.7, or 0.60. Because the methods provided herein allow nucleic acid to be bound to solid supports, other contaminants can be removed merely by washing the solid support after surface bound tagmentation occurs. The biological sample can comprise, for example, a crude cell lysate or whole cells. For example, a crude cell lysate that is applied to a solid support in a method set forth herein, need not have been subjected to one or more of the separation steps that are traditionally used to isolate nucleic acids from other cellular components. Exemplary separation steps are set forth in Maniatis et al., Molecular Cloning: A Laboratory Manual, 2d Edition, 1989, and Short Protocols in Molecular Biology, ed. Ausubel, et al, hereby incorporated by reference.

[0330] In some embodiments, the sample that is applied to the solid support has a 260 / 280 absorbance ratio that is less than or equal to 1.7.

[0331] In some embodiments, the sample is a biopsy sample. In some embodiments, the biopsy sample is a liquid or solid sample. In some embodiments, a biopsy sample from a cancer patient is used to evaluate sequences of interest to determine if the subject has certain mutations or variants in predictive genes.

[0332] In some embodiments, the sample comprises a target double-stranded DNA. In some embodiments, the DNA is genomic DNA. In some embodiments, the DNA is cell-free DNA (cfDNA). In some embodiments, the DNA is circulating tumor DNA (ctDNA).

[0333] In some embodiments, the DNA is double-stranded cDNA that is prepared from RNA. In some embodiments, the RNA is mRNA. In some embodiments, the RNA comprises coding, untranslated region (UTR), introns, and / or intergenic sequences.

[0334] The target nucleic acid can be derived from any in vivo or in vitro source, including from one or multiple cells, tissues, organs, or organisms, whether living or dead, or from any biological or environmental source (e.g., water, air, soil). For example, in some embodiments, the target nucleic acid comprises or consists of eukaryotic and / or prokaryotic dsDNA that originates or that is derived from humans, animals, plants, fungi, (e.g., molds or yeasts), bacteria, viruses, viroids, mycoplasma, or other microorganisms. In some embodiments, the target nucleic acid comprises or consists of genomic DNA, subgenomic DNA, chromosomal DNA (e.g., from an isolated chromosome or a portion of a chromosome, e.g., from one or more genes or loci from a chromosome), mitochondrial DNA, chloroplast DNA, plasmid or other episomal-derived DNA (or recombinant DNA contained therein), or double-stranded cDNA made by reverse transcription of RNA using an RNA-dependent DNA polymerase or reverse transcriptase to generate first-strand cDNA and then extending a primer annealed to the first-strand cDNA to generate dsDNA. In some embodiments, the target nucleic acid comprises multiple dsDNA molecules in or prepared from nucleic acid molecules (e.g., multiple dsDNA molecules in or prepared from genomic DNA or cDNA prepared from RNA in or from a biological (e.g., cell, tissue, organ, organism) or environmental (e.g., water, air, soil, saliva, sputum, urine, feces) source. In some embodiments, the target nucleic acid is from an in vitro source. For example, in some embodiments, the target nucleic acid comprises or consists of dsDNA that is prepared in vitro from single-stranded DNA (ssDNA) or from singlestranded or double-stranded RNA (e g., using methods that are well-known in the art, such as primer extension using a suitable DNA-dependent and / or RNA-dependent DNA polymerase (reverse transcriptase). In some embodiments, the target nucleic acid comprises or consists of dsDNA that is prepared from all or a portion of one or more double-stranded or single-stranded DNA or RNA molecules using any methods known in the art, including methods for: DNA or RNA amplification (e.g., PCR or reverse-transcriptase-PCR (RT-PCR), transcription- mediated amplification methods, with amplification of all or a portion of one or more nucleicacid molecules); molecular cloning of all or a portion of one or more nucleic acid molecules in a plasmid, fosmid, BAC or other vector that subsequently is replicated in a suitable host cell; or capture of one or more nucleic acid molecules by hybridization, such as by hybridization to DNA probes on an array or microarray. Target nucleic acids as provided herein may include, but are not limited to DNA, RNA, peptide nucleic acid, morpholino nucleic acid, locked nucleic acid, glycol nucleic acid, threose nucleic acid, mixtures thereof, and hybrids thereof. In an embodiment, genomic DNA fragments, or amplified copies thereof, are used as the target nucleic acid. In another embodiment, mitochondrial or chloroplast DNA is used. Still other embodiments are targeted to RNA or derivatives thereof such as mRNA or cDNA. In some embodiments, target nucleic acid can be from a single cell. In some embodiments, target nucleic acid can be from acellular body fluids, for example, plasma or sputum devoid of cells. In some embodiments, target nucleic acid can be from circulating tumor cells. In some embodiments, the biological sample can comprise, for example, blood, plasma, serum, lymph, mucus, sputum, urine, semen, cerebrospinal fluid, bronchial aspirate, feces, and macerated tissue, or a lysate thereof, or any other biological specimen comprising nucleic acid. In some embodiments, the sample is blood. In some embodiments, the sample is a cell lysate. In some embodiments, the cell lysate is a crude cell lysate. In some embodiments, the method further comprises lysing cells in the sample after applying the sample to a solid support to generate a cell lysate.

[0335] The disclosed adaptors and adaptor-linked fragments are non-naturally occurring molecules. In addition, other molecules generated from these molecules are also non-naturally occurring. Methods that incorporate these molecules operate on a biological sample to modify the biological sample to generate one or more unique and non-naturally occurring compositions. Further, sequencing data generated from these molecules is generated from non-naturally occurring sampled nucleic acids.

[0336] This written description uses examples to enable any person skilled in the art to practice the disclosed embodiments, including making and using any devices or systems and performing any incorporated methods. The patentable scope is defined by the claims, and mayinclude other examples that occur to those skilled in the art. Such other examples are intended to be within the scope of the claims if they have structural elements that do not differ from the literal language of the claims, or if they include equivalent structural elements with insubstantial differences from the literal languages of the claims.

Claims

CLAIMSWhat is claimed is:

1. An adaptor set for sequence analysis, comprising: a first forked adaptor, the first forked adaptor comprising: a double-stranded region; a forked region, the forked region comprising a first strand and a second strand noncomplementary to one another, the first strand comprising a first adaptor sequence and the second strand comprising a hybridization sequence and an extension binding sequence; and a primer bound to the extension binding sequence of the second strand; and a second forked adaptor, the second forked adaptor comprising: a double-stranded region; and a forked region, the forked region comprising a first strand and a second strand noncomplementary to one another, the second strand comprising a complement of the hybridization sequence, and the first strand comprising a second adaptor sequence different than the first adaptor sequence.

2. The adaptor set of claim 1, comprising a blocking oligonucleotide bound to the complement of the hybridization sequence such that the blocking oligonucleotide comprises all or part of the hybridization sequence.

3. The adaptor set of claim 2, wherein a 3’ end of the blocking oligonucleotide is blocked to polymerase extension.

4. The adaptor set of claim 1, wherein the second forked adaptor comprises the extension binding sequence on the second strand, wherein the extension binding sequence is adjacent to the double-stranded region.

5. The adaptor set of claim 1, wherein the target nucleic acid is 40 to 400 nucleotides in length.

6. The adaptor set of claim 1, wherein the double stranded region of one or both of the first forked adaptor or the second forked adaptor comprises a transposon end sequence.

7. The adaptor set of claim 1, wherein a 5’ end of the first strand of one or both of the first forked adaptor or the second forked adaptor comprises an affinity moiety.

8. The adaptor set of claim 1, wherein the extension binding sequence is 20-50 nucleotides in length.

9. The adaptor set of claim 1, wherein the second strand of the forked region of the first forked adaptor is longer than the second strand of the forked region of the second forked adaptor based on a presence of the extension binding sequence in the first forked adaptor.

10. The adaptor set of claim 1, wherein a 3’ end of the second strand of the forked region of the first forked adaptor is blocked to polymerase extension.

11. A method of generating a tandem repeat of a target nucleic acid, comprising: providing an adaptor-linked fragment comprising a first adaptor of a first type at the first end of a target nucleic acid and a second adaptor of a second type at the second end of the target nucleic acid;using a strand-displacing polymerase to extend from a primer bound to a forked region of the first adaptor and not bound to a forked region of the second adaptor to displace a first strand of the adaptor-linked fragment; permitting binding of complementary sequences in the forked region of the first adaptor and the forked region of the second adaptor; and extending from at least one 3’ end in the forked region of the first adaptor or the forked region of the second adaptor to generate a product comprising: a double-stranded region having a same sequence as the target nucleic acid; and another double-stranded region having the same sequence as the target nucleic acid or a singled-stranded region complementary to one strand of the double-stranded region.

12. The method of claim 11, wherein the forked region of the first adaptor comprises a first strand and a second strand noncomplementary to one another, the first strand comprising a first adaptor sequence and the second strand comprising a hybridization sequence and an extension binding sequence and; and wherein the primer is bound to the extension binding sequence.

13. The method of claim 12, wherein the forked region of the second adaptor comprises a first strand and a second strand noncomplementary to one another, the second strand comprising a complement of the hybridization sequence and wherein the second forked adaptor does not comprise the extension binding sequence or its complement, and the first strand comprising a second adaptor sequence different than the first adaptor sequence.

14. The method of claim 13, wherein the complementary sequences comprise the hybridization sequence and the complement of the hybridization sequence.

15. The method of claim 11, comprising increasing a temperature to remove a blocking oligonucleotide bound to a complementary sequence in the forked region of the second adaptor and not bound to the forked region of the first adaptor to permit binding of the complementary sequences.

16. The method of claim 15, wherein increasing the temperature occurs after or during using the strand-displacing polymerase to extend from the primer.

17. The method of claim 15, wherein increasing the temperature does not denature the double-stranded target nucleic acid.

18. The method of claim 11, comprising immobilizing the adaptor-linked fragment on a substrate using a first affinity moiety present on the first adaptor and a second affinity moiety present on the second adaptor.

19. The method of claim 11, comprising treating the product using a nucleotide modifying agent and sequencing the treated product.

20. The method of claim 19, wherein the nucleotide modifying agent is an APOBEC enzyme that only modifies single-stranded nucleotides to selectively modify methylated cytosines in the single-stranded region to thymine and wherein methylated cytosines in the double-stranded region of the product are not modified.

21. The method of claim 19, comprising generating a methylation analysis using the sequencing, wherein the sequencing generates sequence data for the treated product, wherein a sequence of a strand of the double-stranded region that includes the single- stranded regionhas tandem copies of the sequence complementary to one strand of the double-stranded region.

22. A method of generating a sequencing library, comprising: providing a pool of adaptor-linked fragments having different inserts generated from a target nucleic acid, wherein each adaptor-linked fragment comprises a double-stranded insert of the target nucleic acid flanked by forked adaptors, and wherein the pool of adaptor-linked fragments has a mix of fragments comprising: a first fragment set comprising a first adaptor type at both ends of the respective double-stranded insert; a second fragment set comprising a second adaptor type at both ends of the respective double-stranded insert; and a third fragment set comprising the first adaptor type at the first end and the second adaptor type at the second end of the respective double-stranded insert; using a strand-displacing polymerase to extend from a primer bound to a forked region of the first adaptor type and not bound to a forked region of the second adaptor type to displace strands in the respective double-stranded insert from the first fragment set and the third fragment set, but not the second fragment set, to generate strand copies in the respective double-stranded insert; permitting binding of complementary sequences in the forked region of the first adaptor type and the forked region of the second adaptor type in the third fragment set; and extending from at least one 3’ end in the forked region of the first adaptor type or the forked region of the second adaptor type in the third fragment set to generate a product comprising a first copy of the double-stranded insert and a second copy of the double-stranded insert or a single strand of the double-stranded insert.

23. A composition, comprising:a pool of adaptor-linked fragments having different inserts generated from a target nucleic acid, wherein each adaptor-linked fragment comprises a doublestranded insert of the target nucleic acid flanked by forked adaptors, and wherein the pool of adaptor-linked fragments has a mix of fragments comprising: a first fragment set comprising a first adaptor type at both ends of the respective double-stranded insert; a second fragment set comprising a second adaptor type at both ends of the respective double-stranded insert; and a third fragment set comprising the first adaptor type at the first end and the second adaptor type at the second end of the respective double-stranded insert, the first adaptor type comprising: a double-stranded region; a forked region, the forked region comprising a first strand and a second strand noncomplementary to one another, the first strand comprising a first adaptor sequence and the second strand comprising a hybridization sequence and an extension binding sequence; and a primer bound to the extension binding sequence; and the second adaptor type comprising: a double-stranded region; and a forked region, the forked region comprising a first strand and a second strand noncomplementary to one another, the second strand comprising a complement of the hybridization sequence and wherein the second forked adaptor does not comprise the extension binding sequence or its complement, and the second strand comprising a second adaptor sequence different than the first adaptor sequence.

24. The composition of claim 23, wherein a 5’ end of the forked region of the first adaptor type and a 5’ end of the forked region of the second adaptor type both comprise an affinity moiety.

25. A method of generating a tandem repeat of a target nucleic acid, comprising: providing an adaptor-linked fragment comprising a first adaptor of a first type at the first end of a target nucleic acid and a second adaptor of a second type at the second end of the target nucleic acid; using a strand-displacing polymerase to extend from a primer bound to a sequence in the forked region of the first adaptor and the forked region of the second adaptor to displace strands of the adaptor-linked fragment; permitting binding of complementary sequences in the forked region of the first adaptor and the forked region of the second adaptor; and extending from 3’ ends in the forked region of the first adaptor and the forked region of the second adaptor to generate a product comprising: a double-stranded region having a same sequence as the target nucleic acid; and another double-stranded region having the same sequence as the target nucleic acid.

26. A method of generating a tandem repeat of a target nucleic acid, comprising: providing an adaptor-linked fragment comprising a first adaptor of a first type at the first end of a target nucleic acid and a second adaptor of a second type at the second end of the target nucleic acid; immobilizing the adaptor-linked fragment on a substrate; denaturing strands of the adaptor-linked fragment to generate a first strand and a second strand immobilized on the substrate permitting hybridization of complementary sequences of the first adaptor on a 3 ’ end the first strand and the second adaptor on a 3 ’ end of the second strand;using a polymerase to extend from the 3 ’ end of the first strand, wherein the 3 ’ end of the second strand is blocked from extension by the polymerase to generate a partially double-stranded tandem repeat extension product; treating the partially double-stranded tandem repeat extension product with APOB EC; and generating sequence data from the treated partially double-stranded tandem repeat extension product.

27. The method of claim 26, wherein the adaptor-linked fragment is treated with P- glucosyltransferase before generating the partially double-stranded tandem repeat extension product.

28. An adaptor-linked fragment, comprising: a double-stranded insert generated from a target nucleic acid flanked by adaptors at both ends, the adaptors comprising a first adaptor and a second adaptor each comprising: a double-stranded region adjacent to the double-stranded insert; and a forked region comprising a first strand and a second strand noncomplementary to one another, the first strand comprising a 3’ end and the second strand comprising a blocked 3’ end, wherein a portion of the first strand comprising the 3’ end is complementary to a portion of the second strand.

29. The adaptor-linked fragment of claim 28, wherein the first strand of the forked region comprises a 5 ’-5’ linker positioned between a first primer region and a second primer region, wherein the first primer region is complementary to the complementary sequence.

30. The adaptor-linked fragment of claim 28, wherein the blocked 3’ end comprises an inverted dT, dideoxy-C, or a 3’ phosphate.

31. The adaptor-linked fragment of claim 28, wherein the double-stranded region comprises a plurality of methyl -cytosines.

32. The adaptor-linked fragment of claim 28, wherein the first adaptor and the second adaptor have a same sequence.

33. A method, comprising: contacting the adaptor-linked fragment of claim 28 with a strand-displacing polymerase; extending from the 3 ’end of the first strand bound to the complementary sequence of the second strand to generate extension products using both strands of the doublestranded insert as template; and sequencing the extension products.

34. The method of claim 33, wherein sequencing the extension products comprises: simultaneously generating data from extension from a set of first primers having at least two different sequences and bound to a portion of the extension products and a set of second primers having at least two different sequences and bound to another portion of the extension products; and resolving the data into base call data indicative of sequences of the extension products.

35. The method of claim 34, wherein sequencing the extension products comprises: simultaneously generating data from extension from the set of first primers bound to a portion of the extension products and the set of second primers bound to another portion of the extension products; and resolving the data into base call data indicative of sequences of the extension products.

36. The method of claim 34, wherein the sequencing comprises 16QAM sequencing.

37. The method of claim 33, wherein the strand-displaying polymerase extension is isothermal.

38. The method of claim 33, wherein the extension continues to a 5’ -5’ linker.

39. The method of claim 33, comprising treating the extension products with a modifying agent that converts methylated cytosimes to thymines.

40. The method of claim 33, wherein there is no amplification step prior to the sequencing.

41. The method of claim 33, wherein the sequencing comprises paired-end sequencing.

42. The method of claim 33, wherein the insert sequence and the copy of the insert sequence are both flanked by adaptor sequences.

43. The method of claim 33, wherein the insert sequence and the copy of the insert sequence are separated by a 5 ’-5’ linker.

44. The method of claim 33, wherein an individual extension product has two 3’ ends.

45. A method, comprising: denaturing the adaptor-linked fragment of claim 28 to generate a top denatured strand comprising the first strand of the first adaptor and the second strand of the second adaptor; allowing the 3’ end of the first strand of the first to bind to the complementary sequence of the second strand of the second adaptor;extending from the 3’ end of the first adaptor to generate self-template extension products; amplifying the self-template extension products; and sequencing the amplified self-template extension products.

46. The method of claim 45, wherein sequencing the amplified self-template extension products comprises: simultaneously generating data from extension from a first primer bound to a portion of the amplified extension products and a second primer bound to another portion of the amplified extension products; and resolving the data into base call data indicative of sequences of the self-template extension products.

47. The method of claim 46, wherein the first primer is a Read 1 primer and the second primer is a Read 2 primer.

48. The method of claim 46, wherein the sequencing comprises 9QAM sequencing.

49. The method of claim 45, comprising treating the self-template extension products with a modifying agent that converts methylated cytosines to thymines.

50. The method of claim 45, wherein the amplifying comprises clustering.

51. The method of claim 45, wherein the sequencing comprises paired-end sequencing.

52. The method of claim 45, wherein a single-stranded extension product of the selftemplate extension products comprises an insert sequence and a complement of the insert sequence.

53. The method of claim 52, wherein the insert sequence is a top strand of the doublestranded insert.

54. The method of claim 45, wherein denaturing the adaptor-linked fragment of claim 28 generates a bottom denatured strand comprising the second strand of the first adaptor and the first strand of the second adaptor; allowing the 3’ end of the first strand of the second adaptor to bind to the complementary sequencing of the second strand of the first adaptor; and extending from the 3’ end of the second adaptor to generate the self-template extension products.

55. An adaptor-linked fragment, comprising: a double-stranded insert generated from a target nucleic acid flanked by adaptors at both ends, the adaptors comprising a first adaptor and a second adaptor, wherein the first adaptor comprises a first double-stranded region and a second double- stranded region flanking a noncompl ementary region; and wherein the second adaptor comprises: a first double-stranded region and a second double-stranded region flanking a noncomplementary region; and a hairpin loop adjacent to the second double-stranded region.

56. The adaptor-linked fragment of claim 55, wherein the hairpin loop comprises a nicking site or a double-stranded endonuclease cutting site.

57. The adaptor-linked fragment of claim 55, wherein the hairpin loop forms one end of the adaptor-linked fragment.

58. The adaptor-linked fragment of claim 55, wherein a top strand of the noncomplementary region of the first adaptor and a bottom strand of the noncomplementary region of the second adaptor have a same first sequence.

59. The adaptor-linked fragment of claim 58, wherein a bottom strand of the noncomplementary region of the first adaptor and a top strand of the noncomplementary region of the second adaptor have a same second sequence.

60. The adaptor-linked fragment of claim 59, wherein the first sequence and the second sequence comprise sequencing primer sequences or their complements.

61. The adaptor-linked fragment of claim 55, comprising a first index sequence in the first double-stranded region of the first adaptor and a second index sequence in the second double-stranded region of the second adaptor.

62. A sequencing library comprising the adaptor-linked fragment of claim 55.

63. A method comprising: denaturing the adaptor-linked fragment of claim 55 to generate a single-stranded fragment; allowing the single-stranded fragment to hybridize to a substrate-linked primer complementary to a region of the single-stranded fragment generated from the first double-stranded region of the first adaptor; extending from the substrate-linked primer to copy the single-stranded fragment to generate a copied strand extending from the substrate-linked primer at a 5’ end, and wherein a 3’ end of the copied strand is complementary to the substrate-linked primer; forming a bridge with a second substrate-linked primer using the 3’ end; copying the bridge by extending from the second substrate-linked primer to generate a second copied strand to form a double-stranded bridge with the copied strand;cutting the hairpin loop to generate separated portions of the double-stranded bridge, the separated portions comprising a first strand extending from the substrate-linked primer and a second strand extending from the second substrate-linked primer; and generating sequencing data from the first strand and the second strand.

64. The method of claim 63, wherein generating the sequencing data from the first strand and the second strand comprises using a first sequencing primer to sequence the first strand and a second sequencing primer to sequence the second strand, the first sequencing primer being different than the second sequencing primer.

65. The method of claim 63, wherein the first sequencing primer and the second sequencing primer generate the sequencing data at the same time.

66. The method of claim 63, wherein generating the sequencing data from the first strand and the second strand comprises copying the first strand and the second strand to generate clusters and sequencing strands of the clusters.

67. The method of claim 63, wherein the first strand and the second strand both comprise an insert sequence of the target nucleic acid.

68. The method of claim 67, wherein the insert sequence is the same.

69. The method of claim 63, wherein the substrate-linked primer is a primer having a first sequence and is immobilized on a substrate, wherein the substrate comprises a plurality of primers having the first sequence distributed across the substrate.

70. The method of claim 69, wherein the substrate comprises a plurality of primers having a second sequence, wherein the second adaptor comprises the second sequence.

71. The method of claim 69, wherein 3’ ends of the plurality of primers are blocked from extension while the sequencing data from the first strand and the second strand is generated.

72. The method of claim 69, wherein 3’ ends of the plurality of primers are unblocked to copy the first strand and the second strand.

73. An adaptor-linked fragment, comprising: a double-stranded insert generated from a target nucleic acid flanked by adaptors at both ends, the adaptors comprising a first adaptor and a second adaptor, wherein the first adaptor comprises a double-stranded region and a forked region, the forked region comprising a first index sequence and its complement and, on a first strand, a first primer sequence complementary to a first substrate-linked primer of a first type and, on a second strand, a second primer sequence corresponding to a sequence of a second substrate- linked primer of a second type and; and wherein the second adaptor comprises a double-stranded region comprising a second index sequence and its complement and a hairpin loop, the hairpin loop comprising: a nicking site or a cleavable region; the second primer sequence corresponding to the sequence of the second substrate-linked primer of the second type at a first end of the hairpin loop; and the first primer sequence complementary to the first substrate-linked primer of the first type at a second end of the hairpin loop.

74. The adaptor-linked fragment of claim 73, wherein the hairpin loop is linked to an affinity binder.

75. The adaptor-linked fragment of claim 73, wherein the first primer sequence from the first adaptor is positioned proximate to a 3’ end of the adaptor-linked fragment and wherein the second primer sequence from the first adaptor is positioned proximate to a 5’ end of the adaptor-linked fragment and such that the second primer sequence and the first primer sequence from the second adaptor are on opposing sides of the cleavage or nicking site in the hairpin loop.

76. A method comprising: denaturing the adaptor-linked fragment of claim 73 to generate a single-stranded fragment; allowing the first primer sequence at a 3’ end of the single-stranded fragment to hybridize to the first substrate-linked primer of the first type; extending from the first substrate-linked primer to copy the single-stranded fragment to generate a copied strand extending from the first substrate-linked primer at a 5’ end, and wherein a 3’ end of the copied strand is complementary to the second substrate-linked primer of the second type; forming a single-stranded bridge with the second substrate-linked primer using the 3’ end; copying the bridge by extending from the second substrate-linked primer to generate a second copied strand to form a double-stranded bridge with the copied strand; cutting the double-stranded bridge and removing non-immobilized strands to generate immobilized single strands; clustering the immobilized strands using substrate-linked primers of the first type and the second type; linearizing the clusters; and generating sequencing data from the linearized clusters.

77. The method of claim 76, wherein generating sequencing data from the linearized clusters comprises generating index sequencing data of the first index sequence and thesecond index sequence from linearized strands immobilized to a first or second substrate- linked primer.

78. The method of claim 76, wherein copying the bridge by extending from the second substrate-linked primer comprising running a short extension cycle or fewer than three cycles of bridge amplification.

79. A method comprising: denaturing the adaptor-linked fragment of claim 55 to generate a single-stranded fragment; allowing the single-stranded fragment to hybridize to a substrate-linked primer of a first set of substrate-linked primers, the substrate-linked primer being complementary to a region of the single-stranded fragment generated from the first double-stranded region of the first adaptor; extending from the substrate-linked primer to copy the single-stranded fragment to generate a copied strand extending from the substrate-linked primer at a 5’ end, and wherein a 3’ end of the copied strand is complementary to the substrate-linked primer; forming a single-stranded bridge with a second substrate-linked primer of the first set using the 3’ end; copying the bridge by extending from the second substrate-linked primer to generate a second copied strand to form a first double-stranded bridge with the copied strand; nicking within the double-stranded bridge and removing non-immobilized nicked strands to generate a bridge comprising partially single-stranded bridge portions joined by a double-stranded portion having free ends; generating sequencing data from the by extending the free ends to form a second double-stranded bridge to generate read 1 sequencing data; deprotecting ends of a second set of substrate-linked primers having a different sequence than the first set of substrate-linked primers; anddenaturing the reformed bridge and allowing hybridization to primers of the second set of substrate-linked primers; extending from the deprotected ends of the second set of substrate-linked primers to form a third double-stranded bridge, the third double-stranded bridge extending between primers of the first set and second set; denaturing the third double-stranded bridge; and generating read 2 sequencing data from strands of the third double-stranded bridge using a sequencing primer.

80. A method comprising: denaturing the adaptor-linked fragment of claim 55 to generate a single-stranded fragment; allowing the single-stranded fragment to hybridize to a substrate-linked primer of a first set of substrate-linked primers, the substrate-linked primer being complementary to a region of the single-stranded fragment generated from the first double-stranded region of the first adaptor; extending from the substrate-linked primer to copy the single-stranded fragment to generate a copied strand extending from the substrate-linked primer at a 5’ end, and wherein a 3’ end of the copied strand is complementary to the substrate-linked primer; forming a single-stranded bridge with a second substrate-linked primer of the first set using the 3’ end; copying the bridge by extending from the second substrate-linked primer to generate a second copied strand to form a first double-stranded bridge with the copied strand; nicking within the double-stranded bridge and removing non-immobilized nicked strands to generate bridge comprising partially single-stranded bridge portions joined by a double-stranded portion having free ends; generating sequencing data from the by extending the free ends to form a second double-stranded bridge to generate read 1 sequencing data;cutting the second double-stranded bridge to generate single-strands immobilized by the substrate-linked primers of the first set; deprotecting ends of a second set of substrate-linked primers having a different sequence than the first set of substrate-linked primers; and allowing hybridization of the single-strands immobilized by the substrate-linked primers of the first set to primers of the second set of substrate-linked primers; extending from the deprotected ends of the second set of substrate-linked primers to form a third double-stranded bridge, the third double-stranded bridge extending between primers of the first set and second set; denaturing the third double-stranded bridge; and generating read 2 sequencing data from strands of the third double-stranded bridge using a sequencing primer.

81. A method for detecting methylated cytosine in a target double stranded DNA, the method comprising: treating a target double stranded DNA to reversibly covalently link two strands of the double stranded DNA to form linked DNA; denaturing the linked DNA so as to create single stranded regions within the linked DNA; converting methylated cytosines in the single stranded regions to thymine; reannealing the target DNA to reform linked double stranded DNA; reversing the reversable covalent link; and sequencing the DNA to identify regions that differ from a reference DNA.

82. The method according to claim 81, wherein the methylated cytosine is converted to thymine by a cytidine deaminase.

83. The method according to claim 82, wherein the reversable covalent link is a hairpin.

84. The method according to claim 83, wherein the hairpin is formed by extending the 5’ end of each strand to loop around and attach to the 3 ’ end of the other strand to create the hairpin.

85. The method according to claim 81, wherein the reversable covalent link comprises a cleavage site.

86. The method according to claim 85, wherein the cleavage site is comprised in the nucleotide sequence of the linked DNA structure.

87. The method according to claim 85, wherein the cleavage site is a restriction enzyme recognition site.

88. The method according to claim 85, wherein the cleavage site is a Cas nuclease site.

89. The method according to claim 85, wherein the cleavage site is a cleavable nucleotide.

90. The method according to claim 89, wherein the cleavable nucleotide is cleavable by an enzyme.

91. The method according to claim 90, wherein the cleavable nucleotide is 8-oxoG.

92. The method according to claim 89, wherein the cleavable nucleotide is cleavable by UV light.

93. The method according to claim 89, wherein the cleavable nucleotide is nitropiperonyl dC.

94. The method according to claim 89, wherein the cleavable nucleotide is cleaved by treatment with a base.

95. The method according to claim 94, wherein the cleavage site is an abasic site.

96. The method according to claim 89, wherein the cleavage site comprises a cleavable moiety.

97. The method according to claim 96, wherein the linked DNA is formed by a chemical link with a synthetic functionalization.

98. The method according to claim 97, wherein the chemical link is an interstrand disulfide.

99. The method according to claim 85, wherein the cleavage site is a helicase recognition sequence.

100. The method according to any of claims 81-99, wherein the sequencing is SBS sequencing.

101. The method according to any of claims 81-100, wherein the reference DNA is the original target DNA.

102. A kit for preparing DNA to detect methylated cytosine in a target DNA, the kit comprising:components to reversibly covalently link the two strands of the double stranded DNA to from linked DNA; an enzyme with activity for converting 5-methylcytosine to thymine; and components to reverse the reversable covalent link.

103. The kit of claim 102, further comprising a DNA denaturing agent.

104. The kit of claim 103, wherein the enzyme with activity for converting 5- methylcytosine to thymine is a cytidine deaminase.

105. The kit of claim 102, wherein the components to reversibly covalently link the two strands of the double stranded DNA comprise 8-oxoG.

106. The kit of claim 102, wherein the components to reversibly covalently link the two strands of the double stranded DNA comprise nitropiperonyl dC.

107. The kit of claim 102, wherein the components to reversibly covalently link the two strands of the double stranded DNA introduce interstrand disulfides.

108. A complex comprising: double stranded DNA where the two strands of the double stranded DNA are reversibly covalently linked; and a cytidine deaminase.

109. The complex of claim 108, wherein at least a portion of the double stranded DNA is denatured forming a region with two single DNA strands.

110. A method for detecting methylated cytosine in a target double-stranded nucleic acid fragment, the method comprising:coupling hairpin adaptors to ends of a double-stranded nucleic acid fragment to form linked DNA; denaturing the linked DNA so as to create single-stranded regions of separate strands of the double-stranded nucleic acid fragment between the hairpin adaptors; selectively converting methylated cytosines in the single-stranded regions to thymine; reannealing the strands to reform the double-stranded nucleic acid fragment; sequencing the double-stranded nucleic acid fragment to generate sequence data; and identifying converted methylated cytosines in the sequence data based on sequence differences relative to a reference sequence.

111. The method of claim 110, wherein selectively converting the methylated cytosines does not convert unmethylated cytosines to thymine.

112. The method of claim 110, wherein the coupling is via ligation.

113. The method of claim 110, wherein the coupling is via tagmentation.

114. The method of claim 110, wherein selective converting methylated cytosines to thymine is via a cytidine deaminase.

115. The method of claim 110, comprising cleaving one or both hairpin loops of the hairpin adaptors before or as part of the sequencing.

116. The method of claim 110, comprising cleaving only one hairpin loop of the hairpin adaptors before or as part of the sequencing.

117. The method of claim 110, wherein the hairpin adaptors comprise adaptor sequences.

118. The method of claim 110, wherein denaturing the linked DNA comprises partial denaturing.

119. A method comprising: providing fragments of a target nucleic acid with 3’ A-tails; contacting a substrate comprising immobilized duplexes with the fragments, wherein the immobilized duplexes are 3’ T-tailed; ligating a first fragment of the fragments to a first end of an individual immobilized duplex of the substrate-immobilized duplexes and a second fragment of the fragments to a second end of the individual immobilized duplex such that the individual immobilized duplex is positioned between the first fragment and the second fragment and such that the first fragment and the second fragment each have an available A- tailed 3’ end; ligating a first adaptor of a first type to the available A-tailed 3’ end of the first fragment and the second fragment; and cleaving the immobilized duplex from the first fragment and the second fragment; ligating a second adaptor of a second type to cleaved ends of the first fragment and the second fragment.

120. The method of claim 119, comprising washing excess of the first adaptor from the substrate.

121. The method of claim 119, wherein the cleaving is via a restriction endonuclease.

122. The method of claim 119, wherein the cleaving generates a blunt end or overhand end of the first fragment and the second fragment.

123. The method of claim 119, wherein the first adaptor and the second adaptor are different from one another.

124. The method of claim 119, wherein the substrate comprises a bead.

125. The method of claim 119, wherein the immobilized duplex comprises a doublestranded nucleic acid having a first restriction enzyme recognition site at a first end and a second restriction enzyme recognition site at a second end.

126. The method of claim 119, wherein first restriction enzyme recognition site end and the second restriction enzyme recognition site are the same.

127. The method of claim 119, wherein the immobilized duplex is at least 150 base pairs in length.

128. The method of claim 119, wherein the immobilized duplexes all have a same sequence.

129. A substrate comprising: immobilized double-stranded nucleic acids comprising 3’ T-tails, wherein the doublestranded nucleic acids are immobilized via a modified nucleotide comprising an affinity moiety and wherein an individual immobilized double-stranded nucleic acid comprises a first restriction enzyme recognition sequence at a first end and a second restriction enzyme recognition sequence at a second end.

130. An adaptor for preparing a sequencing library, comprising: a first oligonucleotide comprising a palindromic sequence and a second oligonucleotide, the first oligonucleotide and the second oligonucleotide together forming:a double-stranded region; and a forked region, wherein the first oligonucleotide and the second oligonucleotide are noncomplementary to one another in the forked region, wherein the palindromic sequence of the first oligonucleotide is partially within the forked region and partially within the double-stranded region.

131. The adaptor of claim 130, wherein a 3’ end of the first oligonucleotide is in the forked region.

132. The adaptor of claim 130, wherein all of a region of the first oligonucleotide within the forked region is part of the palindromic sequence.

133. The adaptor of claim 130, wherein the palindromic sequence comprises a first portion that is within the forked region and a second portion within the double-stranded region, and wherein the first portion is longer than the second portion.

134. The adaptor of claim 133, wherein the first portion is at least twice a length of the second portion.

135. The adaptor of claim 133, wherein the first portion cannot bind to a complementary part of the palindromic sequence of a different adaptor having a same sequence as the adaptor when the double-stranded region is present.

136. The adaptor of claim 130, wherein the adaptor comprises a restriction enzyme recognition site.

137. The adaptor of claim 130, wherein the palindromic sequence is at least 10 nucleotides in length.

138. An adaptor-linked fragment, comprising: a double-stranded insert generated from a target nucleic acid flanked by a first adaptor and a second adaptor, the first adaptor and the second adaptor having a same sequence and comprising: a first oligonucleotide comprising a palindromic sequence and a second oligonucleotide, the first oligonucleotide and the second oligonucleotide together forming: a double-stranded region; and a forked region, wherein the first oligonucleotide and the second oligonucleotide are noncomplementary to one another in the forked region, wherein the palindromic sequence of the first oligonucleotide is partially within the forked region and partially within the double-stranded region.

139. The adaptor-linked fragment of claim 138, wherein a 3’ end of the first oligonucleotide is in the forked region.

140. The adaptor-linked fragment of claim 138, wherein all of a region of the first oligonucleotide within the forked region is part of the palindromic sequence.

141. The adaptor-linked fragment of claim 138, wherein the palindromic sequence comprises a first portion that is within the forked region and a second portion within the double-stranded region, and wherein the first portion is longer than the second portion.

142. The adaptor-linked fragment of claim 141, wherein the first portion is at least twice a length of the second portion.

143. The adaptor-linked fragment of claim 138, wherein the palindromic sequence of the first adaptor cannot bind to the palindromic sequence of the second adaptor when the doublestranded region is present in the first adaptor and / or the second adaptor.

144. The adaptor-linked fragment of claim 138, wherein the adaptor comprises a restriction enzyme recognition site.

145. The adaptor-linked fragment of claim 138, wherein the palindromic sequence is at least 10 nucleotides in length.

146. A method of generating a tandem insert of a target nucleic acid, comprising: providing the adaptor-linked fragment of claim 136, wherein the first adaptor and the second adaptor are coupled to a surface of a substrate; separating the first oligonucleotide and the second oligonucleotide of the first adaptor and the second adaptor; permitting binding of complementary sequences in respective palindromic sequences of the first oligonucleotide of the first adaptor and the second adaptor; and extending from 3’ ends of the palindromic sequences to generate a product comprising: a double-stranded region having a same sequence as the target nucleic acid; and another double-stranded region having the same sequence as the target nucleic acid.

147. The method of claim 146, wherein the substrate comprises a bead or a flow cell.

148. The method of claim 146, wherein the first adaptor and the second adaptor are coupled to the surface via 5’ ends of the second oligonucleotide.

149. The method of claim 146, comprising: cutting the product at a restriction enzyme recognition site in the first adaptor and the second adaptor to generate a first cut end and a second cut end; and ligating a third adaptor to the first cut end and a fourth adaptor to the second cut end, wherein the third adaptor and the fourth adaptor are different from one another.

150. A method of generating a tandem insert of a target nucleic acid, comprising: providing the adaptor-linked fragment of claim 136; separating the first oligonucleotide and the second oligonucleotide of the first adaptor and the second adaptor without separating strands of the target nucleic acid; permitting binding of complementary sequences in respective palindromic sequences of the first oligonucleotide of the first adaptor and the second adaptor to form a circularized fragment; and extending from 3’ ends of the palindromic sequences using a strand-displacing polymerase to generate a product comprising: a double-stranded region having a same sequence as the target nucleic acid; and another double-stranded region having the same sequence as the target nucleic acid.

151. A method of characterizing a target nucleic acid, comprising: providing a substrate having a plurality of transposome complexes on a surface of the substrate, wherein the plurality of transposomes comprise a first type of transposome complex and a second type of transposome complex, wherein the first type of transposome complex comprises: a transposase; and a first oligonucleotide and a second oligonucleotide forming a doublestranded transposon end sequence, wherein a 5’ end of the first oligonucleotide iscoupled to the surface and wherein a single-stranded region of the second oligonucleotide comprises a hybridization sequence; and wherein the second type of transposome complex comprises: a transposase; and a first oligonucleotide and a second oligonucleotide forming a doublestranded transposon end sequence, wherein a 5’ end of the first oligonucleotide is coupled to the surface and wherein a single-stranded region of the second oligonucleotide comprises a complement of the hybridization sequence; loading a target nucleic acid onto the surface of the substrate under conditions sufficient to cause fragmentation of the target nucleic acid into a plurality of surface- associated double-stranded target fragments by the plurality of transposome complexes, wherein first oligonucleotides and second oligonucleotides of respective first and second types of transposome complexes are joined to ends of the surface- associated double-stranded target fragments as a result of the fragmentation; denaturing the surface-associated double-stranded target fragments into separate strands; treating the separated strands with a deaminase to convert methyl cytosine into thymine to generate treated strands; permitting binding of the hybridization sequence and its complement in the treated strands; and extending from 3’ ends of the hybridization sequence and its complement to generate a tandem repeat double-stranded product; generating clusters on the surface from the tandem repeat double-stranded product; and generating sequencing data from the clusters.

152. The method of claim 151, comprising using gap fill ligation to fill a gap between 5’ ends of the second oligonucleotide of the first type of transposome complexand the second type of transposome complex and respective 3’ ends of the surface- associated double-stranded target fragments.

153. The method of claim 151, comprising deactivating transposases of the plurality of transposome complexes after generating the surface-associated double-stranded target fragments.

154. The method of claim 151, wherein the deaminase has specificity for single-stranded nucleic acid.

155. The method of claim 151, wherein the surface comprises a plurality of sequencing primers of a first type and a second type.

156. The method of claim 155, wherein the first oligonucleotide of the first type of transposome complex comprises the sequencing primer of the first type or its complement.

157. The method of claim 156, wherein the first oligonucleotide of the second type of transposome complex comprises the sequencing primer of the second type or its complement.

158. The method of claim 156, wherein generating the clusters comprises ridge amplification using the plurality of sequencing primers to generate double-stranded bridges.

159. The method of claim 151, using the sequence data to characterize the methylated cytosines in the target nucleic acid.

160. A method of characterizing a target nucleic acid, comprising:providing a substrate having a plurality of transposome complexes on a surface of the substrate, wherein the plurality of transposomes comprise a first type of transposome complex and a second type of transposome complex, wherein the first type of transposome complex comprises: a transposase; and a first oligonucleotide and a second oligonucleotide forming a doublestranded transposon end sequence, wherein a 5’ end of the first oligonucleotide is coupled to the surface and wherein a single-stranded region of the second oligonucleotide comprises a hybridization sequence; and wherein the second type of transposome complex comprises: a transposase; and a first oligonucleotide and a second oligonucleotide forming a doublestranded transposon end sequence, wherein a 5’ end of the first oligonucleotide is coupled to the surface and wherein a single-stranded region of the second oligonucleotide comprises a complement of the hybridization sequence; loading a target nucleic acid onto the surface of the substrate under conditions sufficient to cause fragmentation of the target nucleic acid into a plurality of surface- associated double-stranded target fragments by the plurality of transposome complexes, wherein first oligonucleotides and second oligonucleotides of respective first and second types of transposome complexes are joined to ends of the surface- associated double-stranded target fragments as a result of the fragmentation; denaturing the surface-associated double-stranded target fragments into separate strands; permitting binding of the hybridization sequence and its complement in the treated strands; extending from 3’ ends of the hybridization sequence and its complement to generate a tandem repeat double-stranded product; denaturing the tandem repeat double-stranded product into single strands;treating the single strands with a deaminase to convert methyl cytosine into thymine to generate treated strands; generating clusters on the surface from one or more strands of the treated strands; and generating sequencing data from the clusters.

161. The method of claim 160, comprising using gap fill ligation to fill a gap between 5’ ends of the second oligonucleotide of the first type of transposome complex and the second type of transposome complex and respective 3’ ends of the surface-associated doublestranded target fragments.

162. The method of claim 160, comprising deactivating transposases of the plurality of transposome complexes after generating the surface-associated double-stranded target fragments.

163. The method of claim 160, wherein the deaminase has specificity for single-stranded nucleic acid.

164. The method of claim 160, wherein the surface comprises a plurality of sequencing primers of a first type and a second type.

165. The method of claim 164, wherein the first oligonucleotide of the first type of transposome complex comprises the sequencing primer of the first type or its complement.

166. The method of claim 165, wherein the first oligonucleotide of the second type of transposome complex comprises the sequencing primer of the second type or its complement.

167. The method of claim 166, wherein generating the clusters comprises ridge amplification using the plurality of sequencing primers to generate double-stranded bridges.

168. The method of claim 160, using the sequence data to characterize the methylated cytosines in the target nucleic acid.

169. The method of claim 160, wherein the first oligonucleotide of the first type of transposome complex or the second type of transposome complex comprises a cleavable base, and comprising cleaving at the cleavable base before generating the clusters.

170. A method of characterizing a target nucleic acid, comprising: providing a substrate having a plurality of transposome complexes on a surface of the substrate, wherein the plurality of transposomes comprise a first type of transposome complex and a second type of transposome complex, wherein the first type of transposome complex comprises: a transposase; and a first oligonucleotide and a second oligonucleotide forming a doublestranded transposon end sequence, wherein a 5’ end of the first oligonucleotide is coupled to the surface and wherein a single-stranded region of the second oligonucleotide comprises a hybridization sequence; and wherein the second type of transposome complex comprises: a transposase; and a first oligonucleotide and a second oligonucleotide forming a doublestranded transposon end sequence, wherein a 5’ end of the first oligonucleotide is coupled to the surface and wherein a single-stranded region of the second oligonucleotide comprises a complement of the hybridization sequence; loading a target nucleic acid onto the surface of the substrate under conditions sufficient to cause fragmentation of the target nucleic acid into a plurality of surface- associated double-stranded target fragments by the plurality of transposome complexes, wherein first oligonucleotides and second oligonucleotides of respective first and second types of transposome complexes are joined to ends of the surface-associated double-stranded target fragments as a result of the fragmentation, wherein the target nucleic acid is treated with a deaminase to convert methyl cytosine into thymine; generating a double-stranded tandem insert product from the double-stranded target fragments; generating clusters on the surface from one or more strands of the doublestranded tandem insert product; and generating sequencing data from the clusters.

171. The method of claim 170, comprising using gap fill ligation to fill a gap between 5’ ends of the second oligonucleotide of the first type of transposome complex and the second type of transposome complex and respective 3’ ends of the surface-associated doublestranded target fragments.

172. The method of claim 170, comprising deactivating transposases of the plurality of transposome complexes after generating the surface-associated double-stranded target fragments.

173. The method of claim 170, wherein the deaminase has specificity for double-stranded nucleic acid.

174. The method of claim 170, wherein the surface comprises a plurality of sequencing primers of a first type and a second type.

175. A method of characterizing a target nucleic acid, comprising: providing a substrate having a plurality of transposome complexes on a surface of the substrate, wherein the plurality of transposomes comprise a first type of transposome complex and a second type of transposome complex, wherein the first type of transposome complex comprises: a transposase; anda first oligonucleotide and a second oligonucleotide forming a doublestranded transposon end sequence, wherein a 5’ end of the first oligonucleotide is coupled to the surface and wherein a single-stranded region of the second oligonucleotide comprises a hybridization sequence; and wherein the second type of transposome complex comprises: a transposase; and a first oligonucleotide and a second oligonucleotide forming a doublestranded transposon end sequence, wherein a 5’ end of the first oligonucleotide is coupled to the surface and wherein a single-stranded region of the second oligonucleotide comprises a complement of the hybridization sequence; loading a target nucleic acid onto the surface of the substrate under conditions sufficient to cause fragmentation of the target nucleic acid into a plurality of surface- associated double-stranded target fragments by the plurality of transposome complexes, wherein first oligonucleotides and second oligonucleotides of respective first and second types of transposome complexes are joined to ends of the surface- associated double-stranded target fragments as a result of the fragmentation; generating a double-stranded tandem insert product from the double-stranded target fragments; treating the double-stranded tandem insert product with a deaminase to convert methyl cytosine into thymine; generating clusters on the surface from one or more strands of the doublestranded tandem insert product after the treating; and generating sequencing data from the clusters.

176. The method of claim 175, comprising using gap fill ligation to fill a gap between 5’ ends of the second oligonucleotide of the first type of transposome complex and the second type of transposome complex and respective 3’ ends of the surface-associated doublestranded target fragments.

177. The method of claim 175, comprising deactivating transposases of the plurality of transposome complexes after generating the surface-associated double-stranded target fragments.

178. The method of claim 175, wherein the deaminase has specificity for double-stranded nucleic acid.

179. The method of claim 175 wherein the surface comprises a plurality of sequencing primers of a first type and a second type.