Methods of improving sequencing accuracy

The method uses transposome complexes with adaptors to differentiate between forward and reverse strands in nucleic acid sequencing, addressing symmetry issues and modified cytosines, enhancing sequencing precision and diagnostic accuracy.

WO2026097017A2PCT designated stage Publication Date: 2026-05-07ILLUMINA INC
View PDF 33 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
ILLUMINA INC
Filing Date
2025-11-03
Publication Date
2026-05-07

AI Technical Summary

Technical Problem

Existing nucleic acid sequencing methods struggle with accuracy due to symmetry breakage in complementary DNA strands caused by DNA damage and the presence of modified cytosines, particularly 5-methylcytosine, which affects sequencing precision and clinical diagnostics.

Method used

A method involving transposome complexes with adaptors to distinguish between forward and reverse strands of double-stranded polynucleotides, incorporating strand identifiers through ligation and extension reactions, followed by sequencing to differentiate and detect errors or modified cytosines.

Benefits of technology

Enhances sequencing accuracy by distinguishing between forward and reverse strands, allowing for precise detection of modified cytosines and identifying sequencing errors, thereby improving diagnostic capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IMGF000030_0001
    Figure IMGF000030_0001
  • Figure IMGF000030_0002
    Figure IMGF000030_0002
  • Figure IMGF000032_0001
    Figure IMGF000032_0001
Patent Text Reader

Abstract

Methods of introducing a strand identifier to distinguish the descendants of a forward strand from the descendants of a reverse strand of a double-stranded target polynucleotide, methods of preparing polynucleotides for sequencing, and methods of sequencing polynucleotide sequences are described.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Methods of improving sequencing accuracy

[0002] Related Applications

[0003] This application claims priority to U. S. Prov. Application No. 63 / 715,069 filed on November 1, 2024, which is incorporated by reference in its entirety.

[0004] Reference to an Electronic Sequence Listing

[0005] The contents of the electronic sequence listing (14419400032. xml; Size: 17,939 bytes; and Date of Creation: October 31, 2025) is herein incorporated by reference in its entirety.

[0006] Field of the Disclosure

[0007] The present disclosure relates to methods of introducing a strand identifier to distinguish the descendants of a forward strand from the descendants of a reverse strand of a double-stranded target polynucleotide, methods of preparing polynucleotides for sequencing, and methods of sequencing polynucleotide sequences.

[0008] Background of the Disclosure

[0009] The common expectation is that the complementary sequences of a double-stranded DNA molecule should carry exactly the same information, and as such, sequencing one strand of the molecule should be sufficient. In practice, however, this notion is not accurate.

[0010] The most common occasion where the symmetry of information between complementary strands may break is due to DNA damage. Different bases of DNA have different susceptibilities to different forms of damage. This results in the creation of mismatched base pairs.

[0011] In addition, modified cytosines, including 5-methylcytosine (5mC), play fundamental roles in human development and disease. Its genome-wide distribution differs between tissue types, and between healthy and diseased states. In recent years, 5mC has also gained prominence as a tool for clinical diagnostics: its distribution in cell-free DNA (cfDNA) -obtained from a liquid biopsy - can be used for the tissue-specific prediction of early-stage cancer.

[0012] There remains a need to develop more accurate nucleic acid sequencing methods, as well as methods for identifying modified cytosines.

[0013] Summary of the Disclosure

[0014] According to an aspect of the present disclosure, there is provided a method of introducing a strand identifier to distinguish the descendants of a forward strand from the descendants of a reverse strand of a double-stranded target polynucleotide, the method comprising: providing one or more first transposome complexes, wherein the first transposome complex comprises a first adaptor, one or more second transposome complexes, wherein the second transposome complex comprises a second adaptor, and a double-stranded target polynucleotide comprising a forward target strand and a reverse target strand; ligating the first adaptor to a 5’ end of the reverse target strand using the first transposome complex and ligating the second adaptor to a 5’ end of the forward target strand using the second transposome complex, to form a partially adapted double-stranded target polynucleotide; removing a transposase enzyme from each of the first and second transposome complexes; and conducting an extension reaction of the partially adapted double-stranded template polynucleotide to form a fully adapted double-stranded target polynucleotide, wherein the extension reaction leads to the incorporation of a third adaptor at the 3’ end of the forward target strand and a fourth adaptor at the 3’ end of the reverse target strand, wherein the third adaptor comprises a forward strand descendant identifier and the fourth adaptor comprises a reverse strand descendant identifier.

[0015] In one aspect, the method further comprises providing a solid support comprising first immobilised primers and second immobilised primers.

[0016] In one aspect, the one or more first transposome complexes are immobilised at a 5’ end to the solid support, and the one or more second transposome complexes immobilised at a 3’ end or a 5’ end to the solid support.

[0017] In one aspect, the first transposome complex comprises a transferred strand and a non-transferred strand, wherein the transferred strand comprises the first adaptor.

[0018] In one aspect, the second transposome complex comprises a transferred strand and a non-transferred strand, wherein the transferred strand comprises the second adaptor.

[0019] In one aspect, the first transposome complex is immobilised to the solid support through the 5’ end of the transferred strand and the second transposome complex is immobilised to the solid support through the 5’ end of the transferred strand.

[0020] In one aspect, the first transposome complex is immobilised to the solid support through the 5’ end of the transferred strand and the second transposome complex is optionally immobilised to the solid support through 3’ end of the non-transferred strand.

[0021] In one aspect, the first adaptor comprises at least a first amplification domain, a reverse template strand identifier and a transposase recognition sequence.

[0022] In one aspect, the second adaptor comprises at least a second amplification domain, a forward template strand identifier and a transposase recognition sequence.

[0023] In one aspect, the method comprises removing the non-transferred strand from the first transposase complex and removing the non-transferred strand from the second transposase complex before conducting an extension reaction.

[0024] In one aspect, the third adaptor is covalently attached to a 3’ end of the forward strand and a fourth adaptor is covalently attached to a 3’ end of the reverse strand using a non-strand displacing polymerase and a ligase. In one aspect, the third adaptor comprises a transposase recognition sequence, a forward strand descendant identifier and at least one amplification domain.

[0025] In one aspect, the fourth adaptor comprises a transposase recognition sequence, a reverse strand descendant identifier and at least one amplification domain.

[0026] In one aspect, the forward strand descendant identifier is non-complementary to the reverse template strand identifier.

[0027] In one aspect, the reverse strand descendant identifier is non-complementary to the forward template strand identifier.

[0028] In one aspect, the forward strand descendant identifier differs from the complement of the forward template strand identifier by one or more bases.

[0029] In one aspect, the reverse strand descendant identifier differs from the complement of the reverse template strand identifier by one or more bases.

[0030] In one aspect, the reverse strand descendant identifier and / or the forward strand descendant identifier comprise at least one modified base, wherein the modified base can be converted to a different base following the extension reaction.

[0031] In one aspect, the reverse strand identifier and / or the forward strand template identifier comprise at least one modified base.

[0032] In one aspect, the modified base is 8-oxo-guanine.

[0033] In one aspect, the modified base is a methylated cytosine, wherein the method further comprises the step of applying a conversion agent following the extension reaction wherein the conversion agent is configured to convert the methylated cytosine to thymine or a nucleobase that is read as thymine / uracil.

[0034] In one aspect, the non-transferred strand of the first transposome complex comprises the fourth adaptor, wherein the fourth adaptor comprises a transposase recognition sequence, a reverse strand descendant identifier and at least one amplification domain, wherein the amplification domain is a complement of the first amplification domain or a fourth amplification domain.

[0035] In one aspect, the non-transferred strand of the second transposome complex comprises the third adaptor, wherein the third adaptor comprises a transposase recognition sequence, a forward strand descendant identifier and at least one amplification domain, wherein the amplification domain is a complement of the second amplification domain or a third amplification domain.

[0036] In one aspect, the method further comprises denaturing the forward strand of the fully adapted target polynucleotide from the reverse strand of the fully adapted target polynucleotide. In one aspect, the method further comprises a step of amplifying the forward and reverse strand of the fully adapted target polynucleotide.

[0037] In one aspect, the forward strand and amplified descendants thereof, and reverse strand and amplified descendants thereof, together form a cluster on the solid support.

[0038] In one aspect, the method comprises a step of preparing the forward strand and amplified descendants thereof, and reverse strand and amplified descendants thereof for sequencing, wherein the method comprises removing the forward strand and amplified reverse complement descendants thereof, or removing the reverse strand and forward complement amplified descendants thereof.

[0039] In one aspect, the solid support is a flow cell, optionally a non-patterned flow cell. According to a further aspect of the present disclosure, there is provided a method of distinguishing the descendants of a forward target strand from the descendants of a reverse target strand of a double-stranded target polynucleotide, the method comprising:

[0040] introducing at least one strand identifier into the descendants of a forward target strand and / or introducing at least one strand identifier into the descendants of a reverse target strand of a double-stranded target polynucleotide using the methods of introducing a strand identifier to distinguish the descendants of a forward strand from the descendants of a reverse strand of a double-stranded target polynucleotide as described herein; and

[0041] sequencing nucleobases in the amplified forward strands and the amplified reverse strands, and determining a difference in sequence output within the amplified forward strand read and / or a difference in sequence output within the amplified reverse strand read; and determining a proportion of sequences derived from the forward target strand and the proportion of sequences derived from the reverse target strand based on the difference in sequence output within the amplified forward strand read and / or the difference in sequence output within the amplified reverse strand read.

[0042] According to a further aspect of the present disclosure, there is provided a method of detecting an error in nucleic acid amplification of a target polynucleotide, the method comprising introducing at least one strand identifier into the descendants of a forward target strand and / or introducing at least one strand identifier into the descendants of a reverse target strand of a double-stranded target polynucleotide using the methods of introducing a strand identifier to distinguish the descendants of a forward strand from the descendants of a reverse strand of a double-stranded target polynucleotide as described herein; and

[0043] sequencing nucleobases in the amplified forward strands and the amplified reverse strands, and determining if there is a difference in sequence output within the amplified forward strand read and / or a difference in sequence output within the amplified reverse strand read; wherein the presence of a difference is indicative of an error during amplification of a target polynucleotide.

[0044] In one aspect, the method comprises conducting a first sequence read using a first sequencing primer that substantially hybridises to a 3’ adaptor of the amplified strand comprising the complement of the forward template strand identifier and conducting a second sequencing read using a second sequencing primer that substantially hybridises to a 3’ adaptor of the amplified strand comprising the complement of the reverse strand descendant identifier and determining a difference in sequence output.

[0045] In one aspect, the method comprises conducting a third sequence read using a first sequencing primer that substantially hybridises to a 3’ adaptor of the amplified strand comprising the complement of the reverse template strand identifier and conducting a fourth sequencing read using a fourth sequencing primer that substantially hybridises to a 3’ adaptor of the amplified strand comprising the complement of the forward strand descendant identifier and determining a difference in sequence output.

[0046] In one aspect, the first and second sequence reads and / or the third and fourth sequence reads are conducted sequentially or concurrently.

[0047] In one aspect, the method comprises conducting paired-end re-synthesis before conducting the third and fourth sequence read.

[0048] In one aspect, the method comprises: conducting a first sequence read using a first sequencing primer wherein the first sequencing primer substantially hybridises to a 3’ adaptor on a forward strand, sequencing nucleobases in the amplified forward strands, and determining a difference in sequence, wherein a forward strand derived from the forward target strand will have a different sequence to a reverse complement strand derived from the reverse target strand.

[0049] In one aspect, the method comprises: conducting a second sequence read using a second sequencing primer wherein the second sequencing primer substantially hybridises to a 3’ adaptor on a reverse strand, sequencing nucleobases in the amplified reverse strands, and determining a difference in sequence, wherein a reverse strand derived from the reverse target strand will have a different sequence to a forward complement strand derived from the forward target strand.

[0050] In one aspect, the first and second sequence reads are conducted sequentially or concurrently.

[0051] In one aspect, the method comprises conducting paired-end re-synthesis before conducting the second sequence read.

[0052] According to a further aspect of the present disclosure, there is provided a solid support comprising a first immobilised primer, a second immobilised primer, and optionally a third immobilised primer and a fourth immobilised primer, and a first and second transposome complex.

[0053] In one aspect, the first transposome complex comprises a transferred and nontransferred strand.

[0054] In one aspect, the second transposome complex comprises a transferred and nontransferred strand.

[0055] In one aspect, each transferred strand comprises at least one amplification domain, a target strand identifier and a transposome recognition sequence.

[0056] In one aspect, each non-transferred strand comprises at least one amplification domain, a descendant strand identifier and a transposome recognition sequence.

[0057] According to a further aspect of the present disclosure, there is provided a method of preparing polynucleotide sequences for sequencing, comprising:

[0058] providing a solid support comprising first immobilised primers and second immobilised primers, and a double-stranded target polynucleotide comprising a forward target strand and a reverse target strand, seeding both the forward target strand and the reverse target strand onto the solid support, amplifying the forward target strand and the reverse target strand to produce amplified forward strands and amplified reverse complement strands each covalently attached to a first immobilised primer, and amplified reverse strands and amplified forward complement strands each covalently attached to a second immobilised primer, and preparing the amplified forward strands and amplified reverse complement strands for concurrent sequencing, or preparing the amplified reverse strands and amplified forward complement strands for concurrent sequencing.

[0059] In one aspect, the step of seeding both the forward target strand and the reverse target strand onto the solid support is conducted using a tagmentation reaction.

[0060] In one aspect, the step of seeding both the forward target strand and the reverse target strand onto the solid support is conducted using a transposase.

[0061] In one aspect, both the forward target strand and the reverse target strand are seeded into a same well of the solid support.

[0062] In one aspect, a 5’-end of the forward target strand is covalently attached to a 3’-end of a first transferred strand.

[0063] In one aspect, a 5’-end of the first transferred strand is covalently attached to a 3’-end of a first amplification domain.

[0064] In one aspect, the first transferred strand is immobilised onto the solid support.

[0065] In one aspect, a 5’-end of the reverse target strand is covalently attached to a 3’-end of a second transferred strand. In one aspect, a 5’-end of the second transferred strand is covalently attached to a 3’-end of a second amplification domain.

[0066] In one aspect, the second transferred strand is immobilised onto the solid support. In one aspect, the step of amplifying the forward target strand and the reverse target strand is conducted using bridge amplification and / or exclusion amplification.

[0067] In one aspect, the second non-transferred strand is immobilised onto the solid support. In one aspect, the step of amplifying the forward target strand and the reverse target strand is conducted using exclusion amplification.

[0068] In one aspect, the step of preparing the amplified forward strands and amplified reverse complement strands for concurrent sequencing comprises simultaneously contacting first sequencing primer binding sites located after a 3’-end of the amplified forward strands with first primers and second sequencing primer binding sites located after a 3’-end of the amplified reverse complement strands with second primers; or wherein the step of preparing the amplified reverse strands and amplified forward complement strands for concurrent sequencing comprises simultaneously contacting third sequencing primer binding sites located after a 3’-end of the amplified reverse strands with third primers and fourth sequencing primer binding sites located after a 3’-end of the amplified forward complement strands with fourth primers.

[0069] In one aspect, the step of preparing the amplified forward strands and amplified reverse complement strands for concurrent sequencing comprises simultaneously nicking amplified forward complement strands hybridised to the amplified forward strands and nicking amplified reverse strands hybridised to the amplified reverse complement strands; or wherein the step of preparing the amplified reverse strands and amplified forward complement strands for concurrent sequencing comprises simultaneously nicking amplified reverse complement strands hybridised to the amplified reverse strands and nicking amplified forward strands hybridised to the amplified forward complement strands.

[0070] In one aspect, the method further comprises a step of treating the forward target strand and the reverse target strand with a conversion agent, prior to the step of amplifying the forward target strand and the reverse target strand,

[0071] wherein the conversion reagent is configured to convert a modified cytosine to thymine or a nucleobase which is read as thymine / uracil, and / or wherein the conversion reagent is configured to convert an unmodified cytosine to uracil or a nucleobase which is read as thymine / uracil.

[0072] In one aspect, the step of treating the forward target strand and the reverse target strand with a conversion agent is conducted after the step of seeding both the forward target strand and the reverse target strand onto the solid support. In one aspect, the conversion agent comprises an enzyme.

[0073] In one aspect, the enzyme is a cytidine deaminase.

[0074] In one aspect, the cytidine deaminase is a wild-type cytidine deaminase or a mutant cytidine deaminase; optionally a mutant cytidine deaminase.

[0075] In one aspect, the amplified forward strands and amplified reverse complement strands, and the amplified reverse strands and amplified forward complement strands, together form a cluster on the solid support.

[0076] In one aspect, the method comprises a step of preparing the amplified forward strands and amplified reverse complement strands for concurrent sequencing, and wherein the amplified reverse strands and the amplified forward complement strands are removed prior to concurrent sequencing.

[0077] In one aspect, the amplified forward strands and amplified reverse complement strands form a duoclonal cluster on the solid support.

[0078] In one aspect, the method comprises preparing the amplified reverse strands and amplified forward complement strands for concurrent sequencing, and wherein the amplified forward strands and the amplified reverse complement strands are removed prior to concurrent sequencing.

[0079] In one aspect, the amplified reverse strands and amplified forward complement strands form a duoclonal cluster on the solid support.

[0080] In one aspect, the population of amplified forward strands prepared for concurrent sequencing is substantially equal to the population of amplified reverse complement strands prepared for concurrent sequencing; and / or wherein the population of amplified reverse strands prepared for concurrent sequencing is substantially equal to the population of amplified forward complement strands prepared for concurrent sequencing.

[0081] In one aspect, the solid support is a flow cell; optionally a non-patterned flow cell. According to a further aspect of the present disclosure, there is provided a method of sequencing polynucleotide sequences, comprising:

[0082] preparing polynucleotides for sequencing using a method as described herein, and concurrently sequencing nucleobases in the amplified forward strands and amplified reverse complement strands, or concurrently sequencing the amplified reverse strands and amplified forward complement strands.

[0083] In one aspect, the method further comprises a step of identifying differences when comparing sequence outputs from the amplified forward strands and amplified reverse complement strands, or when comparing sequence outputs from the amplified reverse strands and amplified forward complement strands. In one aspect, the step of identifying differences comprises associating the differences with the presence of errors.

[0084] In one aspect, the step of identifying differences comprises associating the differences with the presence of modified cytosines.

[0085] In one aspect, the method further comprises a step of conducting paired-end reads. According to a further aspect of the present disclosure, there is provided a kit comprising instructions for preparing polynucleotide sequences for sequencing as described herein, and / or for sequencing polynucleotide sequences as described herein.

[0086] According to a further aspect of the present disclosure, there is provided a data processing device comprising means for carrying out a method as described herein.

[0087] In one aspect, the data processing device is a polynucleotide sequencer.

[0088] According to a further aspect of the present disclosure, there is provided a computer program product comprising instructions which, when the program is executed by a processor, cause the processor to carry out a method as described herein.

[0089] According to a further aspect of the present disclosure, there is provided a computer-readable storage medium comprising instructions which, when executed by a processor, cause the processor to carry out a method as described herein.

[0090] According to a further aspect of the present disclosure, there is provided a computer-readable data carrier having stored thereon a computer program product as described herein.

[0091] According to a further aspect of the present disclosure, there is provided a data carrier signal carrying a computer program product as described herein.

[0092] Description of the Drawings

[0093] Features of examples of the present disclosure will become apparent by reference to the following detailed description and drawings, in which like reference numerals correspond to similar, though perhaps not identical, components. For the sake of brevity, reference numerals or features having a previously described function may or may not be described in connection with other drawings in which they appear.

[0094] Figure 1 shows a symmetric (Figure 1a) and asymmetric (Figure 1b) attachment of a first and second transposome complex. Each transposome complex is a dimer. Each monomer of the dimer comprises a transferred strand comprising a first or second adaptor, where each adaptors comprise amplification domains, transposome recognition sequences and a forward or reverse strand identifier. Each monomer in a single dimer will comprise the same amplification domain and strand identifier - but these sequences may be different between a first and second transposome complex. This is shown in both Figures 1a and 1b -the black strands represent adaptors of the same sequence. The grey strands represent adaptors of different sequences. For example - the black adaptors could be the P5-A14-ME adaptors and the grey ones, could be P7-B15-ME, or vice versa.

[0095] Figure 2 shows the addition of dsDNA to a flow cell lane containing a symmetric composition of transposomes (Figure 2a), the product of tagmentation of the transposomes into the DNA (Figure 2b) and the tagmentation products after the transposase enzyme (such as Tn5) have been removed (Figure 2c).

[0096] Figure 3 shows a progenitor sequence-able bridged template comprising immobilised 5’ adaptor sequences.

[0097] Figure 4 shows a method of adding dsDNA to a flow cell comprising an asymmetric composition of transposomes.

[0098] Figure 5 shows a progenitor sequence-able bridged template immobilised through a 5’ adaptor on the bottom (reverse) strand.

[0099] Figure 6 shows the outcome of clustering when the bridged descendants of the top and bottom strand are in equal or unequal proportions.

[0100] Figure 7 shows the desired post-tagmentation / pre-clustering adaptor sequences that enable top (forward) and bottom (reverse) strand identification.

[0101] Figure 8 shows the descendant strands formed following re-synthesis of the forward (top) (Figure 8a) or reverse (bottom) (Figure 8b) strand, which can be obtained from either a symmetric or asymmetric surface transposome composition.

[0102] Figure 9a shows linearization of the clusters obtained in Figure 8, and then how they would present after the paired-end turn (Fig. 9b).

[0103] Figure 10 shows a method of sequencing of the linearized top (forward) and bottom (reverse) strands of Figure 8.

[0104] Figure 11 shows an alternative method to sequence the linear forward and reverse strands of Figure 8.

[0105] Figure 12 shows a method of preparing a fully-adapted double-stranded target polynucleotide comprising strand identifiers using a symmetric transposome composition.

[0106] Figure 13 shows a method of preparing clusters comprising descendant strand identifiers using a symmetric transposome composition.

[0107] Figure 14 shows another method of preparing clusters comprising descendant strand identifiers using a symmetric transposome composition.

[0108] Figure 15 shows another method of preparing clusters comprising descendant strand identifiers using a symmetric transposome composition.

[0109] Figure 16 shows another method of preparing clusters comprising descendant strand identifiers using a symmetric transposome composition. Figure 17 shows a method of preparing a fully-adapted double-stranded target polynucleotide comprising strand identifiers using an asymmetric transposome composition.

[0110] Figure 18 shows another method of preparing clusters comprising descendant strand identifiers using an asymmetric transposome composition.

[0111] Figure 19 shows another method of preparing clusters comprising descendant strand identifiers using an asymmetric transposome composition.

[0112] Figure 20 shows another method of preparing clusters comprising descendant strand identifiers using an asymmetric transposome composition.

[0113] Figure 21 shows another method of preparing a fully-adapted double-stranded target polynucleotide comprising strand identifiers using an asymmetric transposome composition.

[0114] Figure 22 shows shows another method of preparing clusters comprising descendant strand identifiers using an asymmetric transposome composition.

[0115] Figure 23 shows another method of preparing a fully-adapted double-stranded target polynucleotide comprising strand identifiers using a symmetric transposome composition.

[0116] Figure 24 shows another method of preparing a fully-adapted double-stranded target polynucleotide comprising strand identifiers using a symmetric transposome composition.

[0117] Figure 25 shows an example workflow of detecting errors in the native library using double-stranded seeding. After double-stranded seeding, replication and linearization, concurrent sequencing leads to a 9 QAM output. Four clouds of the 9 QAM output can be used to sequence the forward and reverse complement strands (or reverse and forward complement strands) conventionally, associating the nucleobase callouts with C, G, A or T. Five other clouds of the 9 QAM output can be used to detect differences between the forward and reverse complement strands (or reverse and forward complement strands), thus allowing the detection of errors in the native library.

[0118] Figure 26 shows a further workflow for detecting errors in the native library using double-stranded seeding. After double-stranded seeding, amplification using bridge amplification and / or exclusion amplification results in a cluster of amplified strands in the nanowell. Linearization and a sequencing read can then be conducted to sequence the forward and reverse complement strands (or reverse and forward complement strands). In this case, a paired end turnaround is also conducted, allowing a further sequencing read to be conducted to sequence the reverse and forward complement strands (or forward and reverse complement strands) that were removed in the linearization step.

[0119] Figure 27 shows an example workflow of detecting modified cytosines in the native library using double-stranded seeding. After treating the library with a conversion agent, double-stranded seeding, replication and linearization, concurrent sequencing leads to a 9 QAM output. Four clouds of the 9 QAM output can be used to sequence the forward and reverse complement strands (or reverse and forward complement strands) conventionally, associating the nucleobase callouts with C, G, A or T. Two other clouds can be used to detect differences between the forward and reverse complement strands (or reverse and forward complement strands) and associate them with the presence of modified cytosines (here, 5-methylcytosine) in the native library. Three additional clouds can be used to detect errors in the native library.

[0120] Figure 28 shows a further workflow for detecting modified cytosines in the native library using double-stranded seeding. After double-stranded seeding, treatment with a conversion agent on the flow cell leads to the generation of a mismatched base pair. Amplification using bridge amplification and / or exclusion amplification results in a cluster of amplified strands in the nanowell. Linearization and a sequencing read can then be conducted to sequence the forward and reverse complement strands (or reverse and forward complement strands). In this case, a paired end turnaround is also conducted, allowing a further sequencing read to be conducted to sequence the reverse and forward complement strands (or forward and reverse complement strands) that were removed in the linearization step.

[0121] Figure 29 shows an example workflow for detecting modified cytosines using “symmetric” double-stranded seeding. After tagmentation, forward and reverse target strands are seeded onto a solid support including P5 Tsm (first transposome complex) and P7 Tsm (second transposome complex), which each are immobilised to the solid support via streptavidin-biotin interactions. After removal of the transposases, an extension reaction is conducted to fully adapt the forward and reverse target strands. Methods can be used to convert the modified cytosines (here, 5-methylcytosine) to thymine while the forward and reverse target strands are hybridised to each other. After conversion, the strands are ready for amplification and then concurrent sequencing.

[0122] Figure 30 shows another example workflow for detecting modified cytosines using “symmetric” double-stranded seeding. After tagmentation, forward and reverse target strands are seeded onto a solid support including P5 Tsm (first transposome complex) and P7 Tsm (second transposome complex), which each are immobilised to the solid support via streptavidin-biotin interactions. After removal of the transposases, an extension reaction is conducted to fully adapt the forward and reverse target strands. The forward and reverse target strands can then temporarily be dehybridised from each other, allowing a deaminase to convert modified cytosines (here, 5-methylcytosine) to thymine. After rehybridisation, the strands are ready for amplification and then concurrent sequencing.

[0123] Figure 31 shows an example workflow for detecting modified cytosines using “asymmetric” double-stranded seeding. After tagmentation, forward and reverse target strands are seeded onto a solid support including P5 Tsm (first transposome complex) and P7 Tsm (second transposome complex), which each are immobilised to the solid support via streptavidin-biotin interactions. After removal of the transposases, an extension reaction is conducted to fully adapt the forward and reverse target strands. Methods can be used to convert the modified cytosines (here, 5-methylcytosine) to thymine while the forward and reverse target strands are hybridised to each other. After conversion, the strands are ready for amplification and then concurrent sequencing.

[0124] Detailed Description

[0125] All patents, patent applications, and other publications referred to herein, including all sequences disclosed within these references, are expressly incorporated herein by reference, to the same extent as if each individual publication, patent or patent application was specifically and individually indicated to be incorporated by reference. All documents cited are, in relevant part, incorporated herein by reference in their entireties for the purposes indicated by the context of their citation herein. However, the citation of any document is not to be construed as an admission that it is prior art with respect to the present disclosure.

[0126] The aspects of the present disclosure can be used in sequencing, in particular concurrent sequencing. Methodologies applicable to the present disclosure have been described in WO 08 / 041002, WO 07 / 052006, WO 98 / 44151, WO 00 / 18957, WO 02 / 06456, WO 07 / 107710, WO05 / 068656, US 13 / 661,524 and US 2012 / 0316086, the contents of which are herein incorporated by reference. Further information can be found in US 20060024681, US 20060292611, WO 06 / 110855, WO 06 / 135342, WO 03 / 074734, W007 / 010252, WO 07 / 091077, WO 00 / 179553, WO 98 / 44152 and WO 2022 / 087150, the contents of which are herein incorporated by reference.

[0127] As used herein, the term “variant” refers to a variant polypeptide sequence or part of the polypeptide sequence that retains desired function of the full non-variant sequence. For example, a desired function of the immobilised primer retains the ability to bind (i.e. hybridise) to a target sequence.

[0128] As used in any aspect described herein, a “variant” has at least 25%, 26%, 27%, 28%, 29%, 30%, 31%, 32%, 33%, 34%, 35%, 36%, 37%, 38%, 39%, 40%, 41%, 42%, 43%, 44%, 45%, 46%, 47%, 48%, 49%, 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or at least 99% overall sequence identity to the non-variant nucleic acid sequence. The sequence identity of a variant can be determined using any number of sequence alignment programs known in the art. As an example, Emboss Stretcher from the EMBL-EBI may be used: https: / / www.ebi.ac. uk / jdispatchebpsa / emboss stretcher (using default parameters: pair output format, Matrix = BLOSUM62, Gap open = 12, Gap extend = 2 for proteins; pair output format, Matrix = DNAfull, Gap open = 16, Gap extend = 4 for nucleotides).

[0129] As used herein, the term “fragment” refers to a functionally active series of consecutive nucleic acids from a longer nucleic acid sequence. The fragment may be at least 99%, at least 95%, at least 90%, at least 80%, at least 70%, at least 60%, at least 50%, at least 40% or at least 30% the length of the longer nucleic acid sequence. In particular embodiments, a fragment as used herein also retains the ability to bind (i.e. hybridise) to a target sequence.

[0130] Sequencing generally comprises four fundamental steps: 1) library preparation to form a plurality of target polynucleotides for identification; 2) cluster generation to form an array of amplified polynucleotides; 3) sequencing the cluster array of amplified polynucleotides; and 4) data analysis to identify characteristics of the target polynucleotides from the amplified polynucleotide sequences. These steps are described in greater detail below.

[0131] Library preparation

[0132] Library preparation is the first step in any high-throughput sequencing platform. During library preparation, nucleic acid sequences, for example a genomic DNA sample, or a cDNA or RNA sample, is converted into a sequencing library, which can then be sequenced. By way of example with a DNA sample, the first step in library preparation is random fragmentation of the DNA sample. Sample DNA is first fragmented and the fragments of a specific size (typically 200-500 bp, but can be larger) are ligated, sub-cloned or “inserted” in-between two oligo adaptors (adaptor sequences). The original sample DNA fragments are referred to as “inserts”.

[0133] As will be understood by the skilled person, a double-stranded nucleic acid will typically be formed from two complementary polynucleotide strands comprised of deoxyribonucleotides or ribonucleotides joined by phosphodiester bonds, but may additionally include one or more ribonucleotides and / or non-nucleotide chemical moieties and / or non-naturally occurring nucleotides and / or non-naturally occurring backbone linkages. In particular, the doublestranded nucleic acid may include non-nucleotide chemical moieties, e.g. linkers or spacers, at the 5' end of one or both strands. By way of non-limiting example, the double-stranded nucleic acid may include methylated nucleotides, uracil bases, phosphorothioate groups, peptide conjugates etc. Such non-DNA or non-natural modifications may be included in order to confer some desirable property to the nucleic acid, for example to enable covalent, non-covalent or metal-coordination attachment to a solid support, or to act as spacers to position the site of cleavage an optimal distance from the solid support. A single stranded nucleic acid consists of one such polynucleotide strand. Where a polynucleotide strand is only partially hybridised to a complementary strand - for example, a long polynucleotide strand hybridised to a short nucleotide primer - it may still be referred to herein as a single stranded nucleic acid. A sequence comprising at least a primer-binding sequence (in one embodiment, a primer-binding sequence and a sequencing primer binding site; and in a further embodiment, a combination of a primer-binding sequence, an index sequence and a sequencing primer binding site) may be referred to herein as an adaptor sequence, and an insert is flanked by a 5’ adaptor sequence and a 3’ adaptor sequence. The primer-binding sequence may also comprise a sequencing primer for the index read.

[0134] As used herein, an “adaptor” refers to a sequence that comprises a short sequencespecific oligonucleotide that is ligated to the 5' and 3' ends of each DNA (or RNA) fragment in a sequencing library as part of library preparation. The adaptor sequence may further comprise non-peptide linkers.

[0135] In a further embodiment, the P5’ (e.g. SEQ ID NO: 12 or SEQ ID NO: 15) and P7’ (e.g. SEQ ID NO: 13) primer-binding sequences are complementary to short primer sequences (or lawn primers) present on the surface of a flow cell. Binding of P5’ and P7’ to their complements (P5 and P7) on - for example - the surface of the flow cell, permits nucleic acid amplification. As used herein denotes the complementary strand.

[0136] The primer-binding sequences in the adaptor which permit hybridisation to amplification primers (e.g. lawn primers) will typically be around 20-40 nucleotides in length, although the present disclosure is not limited to sequences of this length. The precise identity of the amplification primers (e.g. lawn primers), and hence the cognate sequences in the adaptors, are generally not material to the present disclosure, as long as the primer-binding sequences are able to interact with the amplification primers in order to direct PCR amplification. The sequence of the amplification primers may be specific for a particular target nucleic acid that it is desired to amplify, but in other embodiments these sequences may be "universal" primer sequences which enable amplification of any target nucleic acid of known or unknown sequence which has been modified to enable amplification with the universal primers. The criteria for design of PCR primers are generally well known to those of ordinary skill in the art.

[0137] The index sequences (also known as a barcode or tag sequence) are unique short DNA (or RNA) sequences that are added to each DNA (or RNA) fragment during library preparation. The unique sequences allow many libraries to be pooled together and sequenced simultaneously. Sequencing reads from pooled libraries are identified and sorted computationally, based on their barcodes, before final data analysis. Library multiplexing is also a useful technique when working with small genomes or targeting genomic regions of interest. Multiplexing with barcodes can exponentially increase the number of samples analysed in a single run, without drastically increasing run cost or run time. Examples of tag sequences are found in WO 05 / 068656, whose contents are incorporated herein by reference in their entirety. The tag can be read at the end of the first read, or equally at the end of the second read, for example using a sequencing primer complementary to the strand marked P7. The present disclosure is not limited by the number of reads per cluster, for example two reads per cluster: three or more reads per cluster are obtainable simply by dehybridising a first extended sequencing primer, and rehybridising a second primer before or after a cluster repopulation / strand resynthesis step. Methods of preparing suitable samples for indexing are described in, for example WO 2008 / 093098, which is incorporated herein by reference. Single or dual indexing may also be used. With single indexing, up to 48 unique 6-base indexes can be used to generate up to 48 uniquely tagged libraries. With dual indexing, up to 24 unique 8-base Index 1 sequences and up to 16 unique 8-base Index 2 sequences can be used in combination to generate up to 384 uniquely tagged libraries. Pairs of indexes can also be used such that every i5 index and every i7 index are used only one time. With these unique dual indexes, it is possible to identify and filter indexed hopped reads, providing even higher confidence in multiplexed samples.

[0138] The sequencing primer binding sites are sequencing and / or index primer binding sites and indicate the starting point of the sequencing read. During the sequencing process, a sequencing primer anneals (i.e. hybridises) to at least a portion of the sequencing primer binding site on the template strand. The polymerase enzyme binds to this site and incorporates complementary nucleotides base by base into the growing opposite strand.

[0139] In some embodiments of the present disclosure, double-stranded sample polynucleotides are directly seeded onto a solid support, where the solid support comprises transposases. In such cases, the double-stranded sample polynucleotides are fragmented and inserted in-between a first amplification domain and a second amplification domain on the solid support by tagmentation, thus generating a library of (fragmented) target polynucleotides directly on the solid support, rather than in solution phase.

[0140] Non-limiting ways of sample preparation are illustrated below.

[0141] Sample Preparation

[0142] As described herein, a polynucleotide sample is introduced into the flow cell where tagmentation takes place. Several example methods disclosed herein involve the preparation of the polynucleotide sample. One example method includes: at an active temperature of a protease in an inorganic salt free lysis buffer including a serine protease, less than 2 mM ethylenediaminetetraacetic acid, a chaotropic detergent, and water, exposing a sample selected from the group consisting of whole blood, a tissue sample, blood spots, and saliva to the inorganic salt free lysis buffer, thereby extracting deoxyribonucleic acids from the whole blood sample and generating a crude lysate; inactivating the protease in the crude lysate using heat or an inhibitor; and adding a chelator of the chaotropic detergent to the crude lysate to generate a complexed crude lysate.

[0143] In this example method, the sample is first exposed to the inorganic salt free lysis buffer at the temperature at which the protease in the inorganic salt free lysis buffer is active. During exposure to the lysis buffer gDNA is extracted from the white blood cells in the sample. In one example of the method, the volume of the sample ranges from about 20 pL to about 50 pL.

[0144] The inorganic salt free lysis buffer includes, and in some examples consists of, a protease (e.g., a serine protease, such as thermolabile proteinase K, trypsin, etc.), less than 2 mM ethylenediaminetetraacetic acid (EDTA), a chaotropic detergent, and water.

[0145] The protease (e.g., serine protease, such as proteinase K, thermolabile proteinase K or trypsin) is included in the inorganic salt free lysis buffer to digest proteins that may deleteriously affect the gDNA during extraction and to protect the extracted gDNA from nucleases. The temperature at which the protease is active will depend upon the protease that is used. Proteinase K is active at temperatures ranging from about 25°C to about 65°C, and may be fully inactivated at about 80°C. Proteinase K can alternatively be inactivated when exposed to a protease inhibitor, such as tetrapeptidyl chloromethyl ketone (TCK). Thermolabile proteinase K is active at low temperatures ranging from about 20°C to about 40°C, and may be fully inactivated at about 55°C. Trypsin is active at temperatures ranging from about 4°C to about 65°C, and can be inactivated when exposed to a protease inhibitor, such as 4-(2-aminoethy)-benzenesulfonyl fluoride. This type of protease inhibitor does not contain DNA contaminants, and thus can remain in the lysate during library preparation. The ability to inactivate the protease, e.g., proteinase K, thermolabile proteinase K, or the trypsin, helps to minimise damage to the enzymes used in downstream processes that take place at higher temperatures, e.g., subsequent library preparation processes that take place at higher temperatures. The inorganic salt free lysis buffer includes 1 unit to 5 units of the protease, i.e., a concentration ranging from about 5 units / mL to about 20 units / mL. In one example, the inorganic salt free lysis buffer includes 3.6 units of the thermolabile proteinase K. In another example, the inorganic salt free lysis buffer includes about 0.8 mg / mL of the proteinase K.

[0146] The inorganic salt free lysis buffer includes less than 2 mM of ethylenediaminetetraacetic acid (EDTA). For the DNA extraction from the whole blood sample, the EDTA may be used as a chelating agent to chelate metal ions that are present in enzymes, which work as co-factors to increase the enzyme’s catalytic activity. While EDTA readily deactivates DNase enzymes (which digest DNA), some of the metal ions chelated by EDTA may be desirable for the enzyme activity that is to take place in the subsequent library preparation technique. The relatively low EDTA concentration in the inorganic salt free lysis buffer enables the desirable metal ion chelation during the lysis procedure, without inhibiting the activity of these enzymes during library preparation. In one example, the concentration of EDTA ranges from about 0.5 mM to 1.5 mM. In another example, the concentration of EDTA is about 1 mM.

[0147] The inorganic salt free lysis buffer includes a chaotropic detergent. The chaotropic detergent may be useful in DNA extraction, but may also act as an inhibitor for enzymatic reactions occurring during library prep. One example of a chaotropic detergent is sodium dodecyl sulfate (SDS), which inhibits the transposase Tn5 used during the tagmentation reaction. Another example of a chaotropic detergent is guanidine hydrochloride. In an example, the chaotropic detergent is present in the inorganic salt free lysis buffer in an amount of about 0.1% w / v to about 0.5% w / v. As one specific example, the chaotropic detergent is present in the inorganic salt free lysis buffer in an amount of about 0.2% w / v.

[0148] The balance of the inorganic salt free lysis buffer is water, such as deionised water or another form of purified water.

[0149] In some examples, the inorganic salt free lysis buffer also includes suitable lysis buffer additives, such as a neutral buffer, e.g., Tris-HCI (pH 8), and a non-ionic detergent, e.g., TWEEN™ 20 (a polyoxyethylene sorbitol ester commercially available from Croda, Inc.). The neutral buffer may be present in a concentration ranging from about 25 mM to about 100 mM, and the non-ionic detergent may be present in an amount ranging from about 0.5% active w / v to about 0.75% active w / v. In one example, the inorganic salt free lysis buffer includes about 50 mM Tris-HCI (pH 8) and about 0.5% active w / v TWEEN™ 20.

[0150] As described herein, the lysis buffer is free of inorganic salt(s). While inorganic salt(s), including sodium chloride, may be beneficial during DNA extraction, it / they can also function as an inhibitor of one or more enzyme(s) used in library preparation processes. The exclusion of the inorganic salt contributes to the fact that the complexed crude lysate disclosed herein does not have to undergo purification before being used in library preparation.

[0151] The exposure of the whole blood sample to the inorganic salt free lysis buffer at the lysis temperature generates a crude lysate. The protease in the crude lysate is then inactivated using heat or an inhibitor. The inactivating agent used will depend upon the protease present.

[0152] In one example, the temperature of the crude lysate (and the contents contained therein) can be raised to about 55°C, which inactivates thermolabile proteinase K. In another example, the temperature of the crude lysate (and the contents contained therein) can be raised to about 80°C, which inactivates proteinase K. It is to be understood that the inactivation temperature may vary depending upon the protease that is used.

[0153] In another example, an inhibitor may be added to the crude lysate. As one specific example, when the inorganic salt free lysis buffer includes trypsin, inactivating the trypsin in the crude lysate involves exposing the crude lysate to the inhibitor, and the inhibitor is 4-(2- aminoethyl)-benzenesulfonyl fluoride. As another specific example, when the inorganic salt free lysis buffer includes proteinase K, inactivating the proteinase K in the crude lysate involves exposing the crude lysate to the inhibitor, and the inhibitor is tetrapeptidyl chloromethyl ketone (TPK). In an example, the ratio of inhibitorserine protease ranges from 1:1 to 50:1. In one example, the inhibitor (e.g., TPK) that is introduced into the flow cell is present in an example of the tagmentation buffer in an amount ranging from about 0.04 mg / mL to about 0.1 mg / mL. Any inhibitor that is selected should be free of DNA contaminants so that the crude lysate containing the inhibitor can be used in library preparation without first being exposed to purification.

[0154] After or as the protease, e.g., proteinase K, thermolabile proteinase K, or the trypsin, in the crude lysate is inactivated, the chelator of the chaotropic detergent is added to the crude lysate to generate a complexed crude lysate. The chelator that is used depends upon the chaotropic detergent that is used. As examples, the chelator of the chaotropic detergent is a cyclodextrin selected from the group consisting of alphacyclodextrin, betacyclodextrin, and methyl betacyclodextrin. As mentioned above, cyclodextrins are cup-shaped structures that have a hydrophilic exterior and a hydrophobic core. These compounds have the ability to form complexes with hydrophilic molecules, such as SDS or another chaotropic detergent in the lysis buffer. In particular, the long hydrophobic tail of SDS is drawn into the cup of the cyclodextrin, which effectively locks the SDS following lysis and prevents it from denaturing enzymes, including those used in downstream library preparation. As such, the complexed crude lysate can be used directly in library preparation without first being exposed to purification.

[0155] In some examples, the chelator is in dry form, the dry form of the chelator is part of an encapsulated complex including a coating material surrounding the dry form of the chelator. The chelator may be lyophilised or otherwise dried to form microspheres, which are then encapsulated by the coating material. In these examples, adding the chelator to the crude lysate involves adding the encapsulated complex to the crude lysate, then triggering release of the dry form of the chelator from the encapsulated complex, and rehydrating the dry form of the chelator. Rehydration of the chelator will precede any activity of the chelator. The release mechanism used will depend upon the type of coating material that surrounds the dry form of the chelator.

[0156] In one example, the coating material is a temperature sensitive wax; and triggering release of the dry form of the chelator involves heating the crude lysate to above a melting temperature of the temperature sensitive wax. Some examples of suitable temperature sensitive waxes include paraffin wax or soy wax, each of which melts at a temperature ranging from about 40°C to about 50°C. Other examples of suitable temperature sensitive waxes include beeswax, microcrystalline polyethylene wax, or carnauba wax, each of which melts at a temperature ranging from about 60°C to about 80°C. In these examples, the crude lysate is heated to the desired temperature to melt the wax and release the dry form of the chelator. The chelator is rehydrated and is able to sequester the chaotropic detergent to form the complexed crude lysate. Rehydration may be accomplished with water or another aqueous rehydration solution.

[0157] In another example, the coating material is a temperature sensitive polymer; and triggering release of the dry form of the chelator involves heating the crude lysate above solubilisation temperature of the temperature sensitive polymer. Examples of suitable temperature sensitive polymers include poly(acrylamide-co-acrylonitrile), agarose, or gelatin, each of which solubilises at a temperature greater than 30°C. Other examples of suitable temperature sensitive polymers include poly(N-isopropylacrylamide), methylcellulose, or poloxamer, each of which solubilises at a temperature less than 30°C. In these examples, the crude lysate is heated to the desired temperature to solubilise the polymer and release the dry form of the chelator. The chelator is rehydrated (e.g., with water or an aqueous solution) and is able to sequester the chaotropic detergent to form the complexed crude lysate.

[0158] In still another example, the coating material is a pH sensitive polymer; and triggering release of the dry form of the chelator involves adjusting a pH of the crude lysate. Examples of suitable pH sensitive polymers include EUDRAGIT® L100 (methacrylic acid copolymer, type B from Evonik, Inc.) or KOLLICOAT® MAE-100 (methacrylic acid copolymer, type A from BASF Corp.), each of which is soluble at pH greater than 6. Other examples of suitable pH sensitive polymers include EUDRAGIT® RL / RS 100 (ammonio methacrylate from Evonik, Inc.) or carboxymethyl cellulose, each of which is soluble at pH greater than 8. Still other examples of suitable pH sensitive polymers include EUDRAGIT® E (amino dimethyl methacryate copolymer from Evonik, Inc.) or chitosan, each of which is soluble at pH less than 3. In these examples, the pH of the crude lysate is adjusted to be lower than the pH at which the polymer is sensitive (e.g., if solubility is pH X or lower) or to be higher than the pH at which the polymer is sensitive (e.g., if solubility is pH X or higher), and the polymer solubilises. Solubilisation of the polymer releases the dry form of the chelator. The chelator is rehydrated as described herein, and is able to sequester the chaotropic detergent to form the complexed crude lysate.

[0159] In yet a further example, the coating material is a time sensitive polymer; and triggering release of the dry form of the chelator involves allowing the encapsulated complex to incubate in the crude lysate for a predetermined amount of time. Examples of suitable time sensitive polymers include hydroxypropyl methylcellulose (a commercially available example of which is METHOCEL™ from DuPont), ethylcellulose (a commercially available example of which is ETHOCEL™ from DuPont), cellulose acetate (a commercially available example of which is OPADRY® CA from Colorcon), and poly(lactic-co-glycolic acid). In these examples, the crude lysate having the encapsulated complex therein is allowed to incubate for a predetermined time, which depends upon the type of polymer, the thickness of the polymer coating, and the temperature and makeup of the reaction taking place (e.g., pH). Over time, the polymer degrades and releases the dry form of the chelator. The chelator is rehydrated and is able to sequester the chaotropic detergent to form the complexed crude lysate.

[0160] The encapsulated complex may be formed by generating a dispersion of the coating material, spray coating the dispersion in a fluidised bed onto the dry form of the chelator, and allowing the complex to dry. Other suitable coating techniques, such a pan coating, spray drying, etc., may be used that effectively form a film of the coating material over the dry form of the chelator.

[0161] The chelator (whether encapsulated or not) may be mixed with a tagmentation buffer, such as a combination of Tris- acetate, magnesium acetate, and water. The use of the tagmentation buffer may be particularly desirable when the complexed crude lysate is to be used in a tagmentation library preparation process. It is to be understood that the tagmentation buffer may initiate the release of the dry form of the chelator, and thus incorporation of the encapsulated chelator and the tagmentation buffer may depend upon the coating of the encapsulated chelator, its dissolution characteristics, and the desired timing for achieving complexation.

[0162] In some examples, the inhibitor of the protease and the chelator (whether encapsulated or not) of the chaotropic agent may also be mixed with the tagmentation buffer. In this example, inactivation of the protease and generation of the complexed crude lysate may take place at the same time.

[0163] When incorporated into the tagmentation buffer, the amount of the chelator may range from about 2% w / v to about 5% w / v.

[0164] The complexed crude lysate may then be exposed to a library preparation technique. The extracted DNA sample in the complexed crude lysate may be exposed to tagmentation, or any other library preparation technique that fragments the longer piece(s) of genetic material and incorporates the desired adapters to the ends of the fragments.

[0165] In one example, the complexed crude lysate disclosed herein may be introduced into any example of the flow cell disclosed herein, where on flow cell tagmentation may be performed as described herein. As such, some examples of the method involve introducing the complexed crude lysate to a flow cell including first and second transposome complexes and first immobilised primers and second immobilised primers immobilised to a surface within the flow cell. In this example, the extracted deoxyribonucleic acids (in the complexed crude lysate) are exposed to tagmentation by: introducing a tagmentation buffer to the flow cell as the complexed crude lysate is introduced, and bringing the flow cell to a temperature of 30°C or higher. In this example, the transposase may be removed and amplification may be performed by: introducing a washing solution into the flow cell; heating the flow cell, containing the washing solution, to about 60°C; and then introducing an extension amplification mix into the flow cell.

[0166] In one example, the sample preparation method disclosed herein, i.e., the exposing, the inactivating, and the adding, are performed in a lysis tube or chamber off board the flow cell.

[0167] When the lysis tube is used, the complexed crude lysate may be dumped from the lysis tube into a sample input port of the sequencing instrument. Then, the fluid delivery system of the sequencing instrument delivers the complexed crude lysate to the flow cell for tagmentation and subsequent processing.

[0168] The chamber may be a polynucleotide sample preparation chamber that is a component of the sequencing instrument. Thus, when the chamber is used, the complexed crude lysate may be transported from the chamber to the flow cell for tagmentation and subsequent processing.

[0169] Another example sample preparation method includes: at a temperature ranging from about 18°C to about 25°C, exposing a cell sample to zinc oxide nanomaterials, thereby initiating detergent free extraction of deoxyribonucleic acids or ribonucleic acids from the cell sample and generating a crude lysate; and exposing the extracted deoxyribonucleic acids or ribonucleic acids in the crude lysate to tagmentation. In an example, the zinc oxide nanomaterials are added to the cell sample. In another example, the zinc oxide nanomaterials are dispersed in water and the dispersion is added to the cell sample. In the latter example, the cell sample may be in the form of a cell-pellet that is added to a zinc oxide nanomaterial dispersion.

[0170] The cell sample may be from any eukaryotic or prokaryotic cells, as well as from fungi cells. Some examples of the cell samples include whole blood, bone marrow aspirate, serum, plasma, tissue, blood spots, saliva, and other bodily fluids.

[0171] The zinc oxide nanomaterials may be in the form of nanoparticles, nanotubes, nanowires, or the like. The nanoparticles may have an average particle size ranging from about 1 nm to about 1000 nm. The nanotubes or nanowires may have an average diameter ranging from about 1 nm to about 1000 nm. In one example, the zinc oxide nanoparticles have an average particle size ranging from about 10 nm to about 950 nm. In one example, the zinc oxide nanoparticles have an average particle size of about 300 nm. The average particle size may represent the mean value for a distribution of particles, which may be associated with the basis of the distribution calculation (number, surface, or volume). As such, in some examples, the average particle size may represent a volume mean diameter, a number mean diameter, or a surface mean diameter. The average particle size may be determined using a particle size analyser, or using X-ray diffraction (XRD), scanning electron microscopy (SEM), transmission electron microscopy (TEM), etc. The use of the zinc oxide nanomaterials enables this DNA or RNA extraction to be detergent free.

[0172] At the outset of this example method, the cell sample and the zinc oxide nanomaterials are mixed together at a 1.5:1 to 1:1.5 volume:volume ratio or volume:weight ratio. In one example, the cell sample and the zinc oxide nanomaterials are mixed together at a 1:1 volume:volume ratio or volume:weight ratio. In one example of the method, the volume of each of the cell sample and the zinc oxide nanoparticles ranges from about 20 pL to about 100 pL. The cell sample and zinc oxide nanomaterials are mixed together and are allowed to incubate for a time ranging from about 2 minutes to about 30 minutes. During incubation, the cells are ruptured, which leads to DNA or RNA extraction. The mechanism by which rupture occurs may be chemical, where excessive Zn2+ released from the zinc oxide nanomaterials adsorbs on the cell membrane surface and then penetrates the cell wall causing rupture, and / or biological, where the zinc oxide nanomaterials trigger a reactive oxygen species (ROS) reaction that leads to apoptotic cell death and collapse of the cellular structure.

[0173] The DNA (e.g., gDNA) or RNA extraction using the zinc oxide nanomaterials can take place at room temperature, e.g., from about 18°C to about 25°C, and thus additional heating and heating equipment is not utilised.

[0174] In one example, the crude lysate is generated in a sample collection device (e.g., a lysis tube, MITRA® tips, dried blood spot cards, etc.), and the zinc oxide nanomaterials are dried on a surface of the sample collection device. In this example, the addition of the cell sample activates lysis.

[0175] In some examples, the zinc oxide nanomaterials are mixed with a protease. The protease (e.g., serine protease, such as thermolabile proteinase K) may be included to digest proteins that may deleteriously affect the DNA or RNA during extraction and to protect the extracted DNA or RNA from nucleases. The protease may be added in an amount ranging from 1 unit to 5 units, i.e., a concentration ranging from about 5 units / mL to about 20 units / mL. The temperature at which the protease is active will depend upon the protease that is used. Thermolabile proteinase K is active at low temperatures ranging from about 20°C to about 40°C, and may be fully inactivated at about 55°C. The ability to inactivate the protease, e.g., thermolabile proteinase K, helps to minimise damage to the enzymes used in downstream processes that take place at higher temperatures, e.g., subsequent library preparation processes that take place at higher temperatures. When the cell sample is a blood sample, hemoglobin interference may take place. To avoid hemoglobin interference, the zinc oxide nanomaterials may be mixed with hemoglobin capture agents, such as HemogloBind™ (from Biotech Support Group) and EasySep™ Direct RapidSpheres™ (from StemCell Technologies). Each of these hemoglobin capture agents may be used as directed by the manufacturer.

[0176] As noted herein, the zinc oxide nanomaterial can induce the generation of a reactive oxygen species. This species can potentially damage the extracted DNA or RNA, and thus it may be desirable to add an ROS-scavenger to reduce the likelihood of extracted DNA or RNA damage. The amount of ROS-scavenger may depend on how much of the zinc is present. In an example, the ROS-scavenger is added in an amount ranging from about 2 mM to about 50 mM.

[0177] The exposure of the cell sample to the zinc oxide nanomaterials generates a crude lysate. The crude lysate may then be exposed to a library preparation technique.

[0178] In some examples of this method, the zinc oxide nanomaterials are combined with the protease; and prior to exposing the extracted deoxyribonucleic acids in the crude lysate to tagmentation, the method further comprises inactivating the protease using heat. For example, the temperature of the crude lysate (and the contents contained therein) can be raised to about 55°C, which inactivates thermolabile proteinase K. It is to be understood that the inactivation temperature may vary depending upon the protease that is used.

[0179] For library preparation, the extracted DNA or RNA sample in the crude lysate may be exposed to tagmentation, or to any other library preparation technique that fragments the longer piece(s) of genetic material and incorporates the desired adapters to the ends of the fragments.

[0180] In one example, the crude lysate generated via this example of the method may be introduced into any example of the flow cell disclosed herein, where on flow cell tagmentation may be performed as described herein. As such, some examples of the method involve introducing the crude lysate to a flow cell including first and second transposome complexes and first immobilised primers and second immobilised primers primers immobilised to a surface within the flow cell. In this example, the extracted deoxyribonucleic acids or ribonucleic acids (in the crude lysate) are exposed to tagmentation by: introducing a tagmentation buffer to the flow cell as the crude lysate is introduced, and bringing the flow cell to a temperature of 30°C or higher. In this example, the tagmentation buffer includes water, an optional co-solvent (e.g., dimethylformamide), and a buffer salt (e.g., tris acetate salt, pH 7.6). Some examples of this tagmentation buffer also include the metal co-factor (e.g., magnesium acetate). Because the crude lysate includes Zn2+, this example of the tagmentation buffer may alternatively be free of the metal co-factor (e.g., magnesium acetate) for the transposase of the transposome complex. In this particular example, the zinc oxide materials may be used in excess. After tagmentation, the transposase may be removed and amplification performed by heating the flow cell to about 60°C; flowing a washing solution through the flow cell; and introducing an amplification mix into the flow cell.

[0181] In one example, the DNA sample preparation method utilizing the zinc oxide nanomaterials takes place in a chamber off board the flow cell. The chamber may be a DNA sample preparation chamber that is a component of the sequencing instrument. Thus, when the chamber is used, the crude lysate may be transported from the chamber to the flow cell for on flow cell tagmentation and subsequent processing.

[0182] Cluster generation and amplification

[0183] In some embodiments of the present disclosure, a double-stranded sample polynucleotide is contacted in free solution onto a solid support comprising surface capture moieties (for example P5 and P7 lawn primers).

[0184] Thus, embodiments of the present disclosure may be performed on a solid support, such as a flow cell. However, in alternative embodiments, seeding and clustering can be conducted off-flow cell using other types of solid support.

[0185] The solid support may comprise a substrate. The substrate comprises at least one well or depression (e.g. a nanowell), and typically comprises a plurality of wells or depressions (e.g. a plurality of nanowells). Seeding of the forward target strand and the reverse target strand may be conducted such that both the forward target strand and the reverse target strand are seeded into a same well of the solid support.

[0186] Where flow cells are used, the flow cell may be a patterned flow cell, where the wells or depressions are spatially arranged in an ordered pattern with defined sizes, defined shapes and ordered spacings. In other embodiments, the flow cell may be a non-patterned flow cell, where the wells or depressions have varied sizes, undefined shapes and irregular spacing.

[0187] The solid support typically comprises at least one first immobilised primer and at least one second immobilised primer, for example a plurality of first immobilised primers and a plurality of second immobilised primers.

[0188] Thus, each well may comprise at least one first immobilised primer, and typically may comprise a plurality of first immobilised primers. In addition, each well may comprise at least one second immobilised primer, and typically may comprise a plurality of second immobilised primers. Thus, each well may comprise at least one first immobilised primer and at least one second immobilised primer, and typically may comprise a plurality of first immobilised primers and a plurality of second immobilised primers. The first immobilised primer may be attached via a 5’-end of its polynucleotide chain to the solid support. When extension occurs from first immobilised primer, the extension may be in a direction away from the solid support.

[0189] The second immobilised primer may be attached via a 5’-end of its polynucleotide chain to the solid support. When extension occurs from second immobilised primer, the extension may be in a direction away from the solid support.

[0190] The first immobilised primer may be different (in sequence) to the second immobilised primer and / or a complement of the second immobilised primer. The second immobilised primer may be different (in sequence) to the first immobilised primer and / or a complement of the first immobilised primer.

[0191] The (or each of the) first immobilised primer(s) may comprise a sequence as defined in any one of SEQ ID NO: 1, SEQ ID NO: 3, SEQ ID NO: 4 and SEQ ID NO: 14, or a variant or fragment thereof. The second immobilised primer(s) may comprise a sequence as defined in any one of SEQ ID NO: 2, SEQ ID NO: 5 and SEQ ID NO: 6, ora variant or fragment thereof. However, it will be appreciated by the skilled person that the sequences used for the first immobilised primer(s) can also be swapped around with the sequences used for the second immobilised primer(s), and vice versa.

[0192] In some embodiments of the present disclosure, the solid support may comprise at least one third immobilised primer and at least one fourth immobilised primer, for example a plurality of third immobilised primers and a plurality of fourth immobilised primers.

[0193] Thus, each well may further comprise at least one third immobilised primer, and typically may comprise a plurality of third immobilised primers. In addition, each well may comprise at least one fourth immobilised primer, and typically may comprise a plurality of fourth immobilised primers. Thus, each well may comprise at least one first immobilised primer, at least one second immobilised primer, at least one third immobilised primer and at least one fourth immobilised primer; and typically may comprise a plurality of first immobilised primers, a plurality of second immobilised primers, a plurality of third immobilised primers and a plurality of fourth immobilised primers.

[0194] The third immobilised primer may be attached via a 5’-end of its polynucleotide chain to the solid support. When extension occurs from third immobilised primer, the extension may be in a direction away from the solid support.

[0195] The fourth immobilised primer may be attached via a 5’-end of its polynucleotide chain to the solid support. When extension occurs from fourth immobilised primer, the extension may be in a direction away from the solid support.

[0196] The third immobilised primer may be different (in sequence) to the first immobilised primer, a complement of the first immobilised primer, the second immobilised primer, a complement of the second immobilised primer, the fourth immobilised primer, and / or a complement of the fourth immobilised primer. The fourth immobilised primer may be different (in sequence) to the first immobilised primer, a complement of the first immobilised primer, the second immobilised primer, a complement of the second immobilised primer, the third immobilised primer, and / or a complement of the third immobilised primer.

[0197] The solid support may be adapted to allow both the forward and reverse strands of a double-stranded library fragment to become attached to the solid support. Suitable methods for allowing such double-stranded seeding to occur by using methods as described in WO 2023 / 122755 (the contents of which are incorporated herein in their entirety by reference), such as by tagmentation reactions.

[0198] Thus, in some embodiments, the solid supports (e.g. flow cells) as described herein comprise a transposase. In particular, the transposase may be immobilised to a surface of the solid support. The transposase may comprise a first transposase and a second transposase.

[0199] A transposase is an enzyme that is capable of forming a functional complex with a transposon end-containing composition (e.g., transposons, transposon ends, transposon end compositions) and catalysing insertion or transposition of the transposon end-containing composition into the double-stranded polynucleotide sample with which it is incubated, for example, in the in vitro transposition reaction (i.e., tagmentation). Any transposase that is capable of inserting a transposon end with sufficient efficiency to 5’-tag and fragment the DNA sample for its intended purpose can be used. Non-limiting examples of transposases include Tn5 transposase, Tn7 transposase, MuA transposase and Sleeping Beauty (SB) transposase. A transposase as presented herein can also include integrases from retrotransposons and retroviruses.

[0200] As used herein, a transposon end is a double-stranded nucleic acid strand that exhibits only the nucleotide sequences (the “transposon end sequences”) that are necessary to form the complex with the transposase that is functional in tagmentation. The double-stranded nucleic acid strand of the transposon end can include any nucleic acid or nucleic acid analogue suitable for forming the functional complex with the transposase. For example, the transposon end can include natural DNA or DNA analogs (with modified bases and / or backbones), and can include nicks in one or both strands.

[0201] The transposase typically forms a complex with the “transferred” strand of a transposon end and a shorter polynucleotide sequence that is hybridised to the “transferred” strand (the “non-transferred” strand of a transposon end). Such a transposase complex may be formed between a transposase and a double stranded nucleic acid including a transposase integration recognition site. For example, the transposome complex can be a transposase enzyme pre-incubated with double-stranded transposon DNA under conditions that support non-covalent complex formation. Double-stranded transposon DNA can include, for example, Tn5 DNA, a portion of Tn5 DNA, a transposon end composition, a mixture of transposon end compositions or other double-stranded DNAs capable of interacting with a transposase, such as the hyperactive Tn5 transposase.

[0202] Thus, a transposome complex may comprise the transposase, the transferred strand and the non-transferred strand. As used herein, the term “transferred strand” may refer to a sequence that includes a transferred portion of a transposon end. Similarly, the term “nontransferred strand” may refer to a sequence that includes the non-transferred portion of a transposon end. The 3’-end of a transferred strand is joined or transferred to a double stranded fragment during tagmentation. The non-transferred strand is not joined or transferred to the double stranded fragment during tagmentation. In an example, the transferred and nontransferred strands include at least partially complementary portions that are hybridised together.

[0203] A first transposase may be part of a complex (a first transposome complex), the first transposome complex comprising the first transposase, the first transferred strand and the first non-transferred strand (wherein the first non-transferred strand is hybridised to the first transferred strand).

[0204] The first transferred strand is typically attached (e.g. covalently) to a first amplification domain. For example, a 5’-end of the first transferred strand may be attached (e.g. covalently) to a 3’-end of the first amplification domain. The attachment of the first transferred strand to the first amplification domain may be direct or indirect (but overall may be covalently attached). For indirect attachment (e.g. covalent attachment), the first transferred strand may be spaced from the first amplification domain by a first sequencing primer sequence and / or a first index sequence. For example, from a 5’ to 3’ direction, the order of sequences may be:

[0205] first amplification domain, first sequencing primer sequence, first transferred strand;

[0206] first amplification domain, first index sequence, first transferred strand; first amplification domain, first sequencing primer sequence, first index sequence, first transferred strand; or

[0207] first amplification domain, first index sequence, first sequencing primer sequence, first transferred strand.

[0208] In embodiments, the first amplification domain has the same sequence as the first immobilised primer. The first amplification domain may be different (in sequence) to the second amplification domain and / or a complement of the second amplification domain. In some embodiments, the first transferred strand is immobilised onto the solid support. The first transferred strand may be covalently immobilised onto the solid support, or may be non-covalently immobilised onto the solid support (e.g. via streptavidin-biotin interactions).

[0209] Seeding causes the forward target strand to become attached (e.g. covalently) to the first transferred strand. For example, a 5’-end of the forward target strand may be attached (e.g. covalently) to a 3’-end of the first transferred strand. The attachment (e.g. covalent attachment) may be direct.

[0210] Similarly, a second transposase may be part of a complex (a second transposome complex), the second transposome complex comprising the second transposase, the second transferred strand and the second non-transferred strand (wherein the second non-transferred strand is hybridised to the second transferred strand).

[0211] The second transferred strand is typically attached (e.g. covalently) to a second amplification domain. For example, a 5’-end of the second transferred strand may be attached (e.g. covalently) to a 3’-end of the second amplification domain. The attachment of the second transferred strand to the second amplification domain may be direct or indirect (but overall may be covalently attached). For indirect attachment (e.g. covalent attachment), the second transferred strand may be spaced from the second amplification domain by a second sequencing primer sequence and / or a second index sequence. For example, from a 5’ to 3’ direction, the order of sequences may be:

[0212] second amplification domain, second sequencing primer sequence, second transferred strand;

[0213] second amplification domain, second index sequence, second transferred strand;

[0214] second amplification domain, second sequencing primer sequence, second index sequence, second transferred strand; or

[0215] second amplification domain, second index sequence, second sequencing primer sequence, second transferred strand.

[0216] In embodiments, the second amplification domain has the same sequence as the second immobilised primer. The second amplification domain may be different (in sequence) to the first amplification domain and / or a complement of the first amplification domain.

[0217] In some embodiments (“symmetric” double-stranded seeding), the second transferred strand is immobilised onto the solid support. The second transferred strand may be covalently immobilised onto the solid support, or may be non-covalently immobilised onto the solid support (e.g. via streptavidin-biotin interactions).

[0218] In other embodiments (“asymmetric” double-stranded seeding), the second nontransferred strand is immobilised onto the solid support. The second non-transferred strand may be covalently immobilised onto the solid support, or may be non-covalently immobilised onto the solid support (e.g. via streptavidin-biotin interactions).

[0219] Seeding causes the reverse target strand to become attached (e.g. covalently) to the second transferred strand. For example, a 5’-end of the reverse target strand may be attached (e.g. covalently) to a 3’-end of the second transferred strand. The attachment (e.g. covalent attachment) may be direct.

[0220] In embodiments where the detection of modified cytosines is desired, following doublestranded polynucleotide seeding, the forward target strand and the reverse target strand may be treated with a conversion agent, wherein the conversion reagent is configured to convert a modified cytosine to thymine or a nucleobase which is read as thymine / uracil, and / or wherein the conversion reagent is configured to convert an unmodified cytosine to uracil or a nucleobase which is read as thymine / uracil. Examples of suitable conversion agents are described in WO 2023 / 175037, the contents of which are incorporated herein by reference.

[0221] As used herein, the term “modified cytosine” may refer to any one or more of 5-methylcytosine (5-mC), 5-hydroxymethylcytosine (5-hmC), 5-formylcytosine (5-fC) and 5-carboxylcytosine (5-caC):

[0222]

[0223] 5-methylcytosine 5-hydroxymethylcytosine 5-formylcytosine 5-carboxylcytosine (5-mC) (5-hmC) (5-fC) (5-caC) wherein the wavy line indicates an attachment point of the modified cytosine to the polynucleotide.

[0224] As used herein, the term “ refers to cytosine (C):

[0225]

[0226] cytosine

[0227] (C)

[0228] wherein the wavy line indicates an attachment point of the unmodified cytosine to the polynucleotide.

[0229] As used herein, the term “conversion reagent configured to convert a modified cytosine to thymine or a nucleobase which is read as thymine / uracil” may refer to a reagent which converts one or more modified cytosines (e.g. 5-methylcytosine, 5-hydroxymethylcytosine, 5-formylcytosine and 5-carboxylcytosine) to thymine (i.e. would base pair with adenine), or to an equivalent nucleobase which would base pair with adenine. The conversion may comprise a deamination reaction converting the modified cytosine to thymine or nucleobase which is read as thymine / uracil.

[0230] As used herein, the term “conversion reagent configured to convert an unmodified cytosine to uracil or a nucleobase which is read as thymine / uracil” may refer to a reagent which converts one or more unmodified cytosines to uracil (i.e. would base pair with adenine), or to an equivalent nucleobase which would base pair with adenine. The conversion may comprise a deamination reaction converting the unmodified cytosine to uracil or nucleobase which is read as thymine / uracil.

[0231] In some embodiments, the conversion reagent configured to convert a modified cytosine to thymine or a nucleobase which is read as thymine / uracil may further be configured to be selective for converting one or more modified cytosines (e.g. 5-methylcytosine, 5-hydroxymethylcytosine, 5-formylcytosine and 5-carboxylcytosine) over converting unmodified cytosine. The selectivity may be measured by comparing reaction parameters (e.g. deamination reaction parameters) of the conversion of a particular modified cytosine to thymine or equivalent nucleobase which is read as thymine / uracil, with corresponding reaction parameters (e.g. deamination reaction parameters) of the conversion of unmodified cytosine to uracil or nucleobase which is read as thymine / uracil. For example, reaction parameters such as rate of reaction or yield may be compared. In the case of rate of reaction, a rate of a reaction (e.g. deamination) of the particular modified cytosine to thymine or nucleobase which is read as thymine / uracil may be greater (e.g. at least 2 times greater, at least 5 times greater, at least 10 times greater, at least 20 times greater, at least 50 times greater, or at least 100 times greater) than a corresponding rate of a reaction (e.g. deamination) of the unmodified cytosine to uracil or nucleobase which is read as thymine / uracil. In the case of yield, a yield of a reaction (e.g. deamination) of the particular modified cytosine to thymine or nucleobase which is read as thymine / uracil may be greater (e.g. at least 2 times greater, at least 5 times greater, at least 10 times greater, at least 20 times greater, at least 50 times greater, or at least 100 times greater) than a corresponding yield of a reaction (e.g. deamination) of the unmodified cytosine to uracil or nucleobase which is read as thymine / uracil.

[0232] In some embodiments, the conversion reagent configured to convert an unmodified cytosine to uracil or a nucleobase which is read as thymine / uracil may further be configured to be selective for converting unmodified cytosine over converting one or more modified cytosines (e.g. 5-methylcytosine, 5-hydroxymethylcytosine, 5-formylcytosine and 5-carboxylcytosine). The selectivity may be measured by comparing reaction parameters (e.g. deamination reaction parameters) of the conversion of unmodified cytosine to uracil or nucleobase which is read as thymine / uracil, with corresponding reaction parameters (e.g. deamination reaction parameters) of the conversion of a particular modified cytosine to thymine or nucleobase which is read as thymine / uracil. For example, reaction parameters such as rate of reaction or yield may be compared. In the case of rate of reaction, a rate of a reaction (e.g. deamination) of the unmodified cytosine to uracil or nucleobase which is read as thymine / uracil may be greater (e.g. at least 2 times greater, at least 5 times greater, at least 10 times greater, at least 20 times greater, at least 50 times greater, or at least 100 times greater) than a rate of a reaction (e.g. deamination) of the particular modified cytosine to uracil or the nucleobase which is read as thymine / uracil. In the case of yield, a yield of a reaction (e.g. deamination) the unmodified cytosine to uracil or nucleobase which is read as thymine / uracil may be greater (e.g. at least 2 times greater, at least 5 times greater, at least 10 times greater, at least 20 times greater, at least 50 times greater, or at least 100 times greater) than a corresponding yield of a reaction (e.g. deamination) of the particular modified cytosine to uracil or the nucleobase which is read as thymine / uracil.

[0233] In some embodiments, the conversion agent may comprise an enzyme.

[0234] In one embodiment, the enzyme may comprise a cytidine deaminase.

[0235] As used herein, the term “cytidine deaminase” may refer to an enzyme which is able to catalyse the following reaction:

[0236]

[0237] wherein R is hydrogen, methyl, hydroxymethyl, formyl or carboxyl, and wherein the wavy line indicates an attachment point to a polynucleotide.

[0238] In one embodiment, the cytidine deaminase is a wild-type cytidine deaminase or a mutant cytidine deaminase. In one example, the cytidine deaminase is a mutant cytidine deaminase.

[0239] In some embodiments, the cytidine deaminase is a member of the APOBEC protein family. In one embodiment, the cytidine deaminase is a member of the AID subfamily, the APOBEC1 subfamily, the APOBEC2 subfamily, the APOBEC3 subfamily (e.g. the APOBEC3A subfamily, the APOBEC3B subfamily, the APOBEC3C subfamily, the APOBEC3D subfamily, the APOBEC3F subfamily, the APOBEC3G subfamily, or the APOBEC3H subfamily), or the APOBEC4 subfamily. In one example, the cytidine deaminase is a member of the APOBEC3A subfamily.

[0240] In general, cytidine deaminases are able to catalyse the deamination of all modified cytosines (particularly 5-methylcytosine, 5-hydroxymethylcytosine and 5-formylcytosine) to their equivalent deaminated versions (i.e. nucleobases which are read as thymine / uracil), as well as catalysing the deamination of unmodified cytosines to uracil. Nevertheless, rates of reaction may differ depending on the type of modified cytosine; for example, wild-type APOBEC3A catalyses the deamination of unmodified cytosine and 5-methylcytosine relatively efficiently, whereas deamination of 5-hydroxymethylcytosine is ~5000-fold slower relative to unmodified cytosine, deamination of 5-formylcytosine is ~3700-fold slower relative to unmodified cytosine, and deamination of 5-carboxylcytosine is >20000-fold slower relative to unmodified cytosine. Where distinction between modified cytosines and unmodified cytosines is desired (or even between different types of modified cytosines), treatment with further agents as described herein prior to treatment with the cytidine deaminase may provide such distinction. In particular, the cytidine deaminase may be combined with ten-eleven translocation (TET) methylcytosine dioxygenases and / or p-glucosyltransferases as described herein. Alternatively, or in addition, particular cytidine deaminases (e.g. mutant cytidine deaminases) may be chosen which have higher affinities for modified cytosines as substrates over unmodified cytosines, or vice versa.

[0241] In one embodiment, the mutant cytidine deaminase may comprise amino acid substitution mutations at positions functionally equivalent to (Tyr / Phe)130 and Tyr132 in a wild-type APOBEC3A protein. Such mutant cytidine deaminases are described in further detail in WO 2023 / 196572, which is incorporated herein by reference. Other suitable mutant (altered) cytidine deaminases include those described in WO 2025 / 072783, WO 2025 / 072793 and WO 2025 / 072800, which are incorporated by reference in their entirety with regard to altered cytidine deaminases. Further, the mutant cytidine deaminase may be a fusion protein, such as those disclosed in WO / 2024 / 069581, the contents of which are incorporated in its entirety, including the fusion protein of a helicase and an altered cytidine deaminase. By "functionally equivalent" it is meant that the mutant cytidine deaminase has the amino acid substitution at the amino acid position in a reference (wild-type) cytidine deaminase that has the same functional role in both the reference (wild-type) cytidine deaminase and the mutant cytidine deaminase.

[0242] In one embodiment, the (Tyr / Phe)130 may be Tyr130, and the wild-type APOBEC3A protein may be SEQ ID NO: 16 (UniProt P31941).

[0243] In some embodiments, the mutant cytidine deaminase may convert 5-methylcytosine to thymine by deamination at a greater rate than conversion rate of cytosine to uracil by deamination; wherein the rate may be at least 100-fold greater.

[0244] In one embodiment, the substitution mutation at the position functionally equivalent to Tyr130 may comprise Ala, Vai or Trp.

[0245] In one embodiment, the substitution mutation at the position functionally equivalent to Tyr132 may comprise a mutation to His, Arg, Gin or Lys. When a step of treating the forward target strand and the reverse target strand with a conversion agent is conducted, it can be advantageous to include additional reagents that allow the dehybridisation of the forward target strand from the reverse target strand (at least during the conversion process). This allows the conversion agent to access the nucleobases to be converted. Non-limiting examples of such additional reagents include single-stranded binding proteins, helicases, and zinc ions. Other methods may involve the incorporation of internal biotin into the forward target strand and the reverse target strand, thus allowing the strands to be pulled apart by use of streptavidin.

[0246] Following double-stranded polynucleotide seeding (or following the treatment of the forward target strand and the reverse target strand with a conversion agent), the forward target strand and the reverse target strand can then be amplified on the solid support. In cases where “symmetric” double-stranded seeding has occurred, both bridge amplification (e.g. by a method analogous to that of WO 98 / 44151 or that of WO 00 / 18957) and exclusion amplification (such as in the presence of a recombinase, e.g. by a method analogous to that of WO 2013 / 188582) may be used to generate copies of the forward target strand and the reverse target strand. In cases where “asymmetric” double-stranded seeding has occurred, exclusion amplification may be more suitable because the recombinase can permit strand invasion of the second immobilised primer. In more general terms, amplification where “asymmetric” double-stranded seeding has occurred may be conducted in the presence of a recombinase.

[0247] Through such approaches, a cluster of template molecules is formed, comprising amplified forward strands, reverse complement strands, reverse strands and forward complement strands.

[0248] As used herein, the term “cluster” may refer to a clonal group of template polynucleotides (e.g. DNA or RNA) bound within a single well of a solid support (e.g. flow cell). As such, a cluster may refer to the population of polynucleotide molecules within a well that are then sequenced. A “cluster” may contain a sufficient number of copies of template polynucleotides such that the cluster is able to output a signal (e.g. a light signal) that allows sequencing reads to be performed on the cluster. A “cluster” may comprise, for example, about 500 to about 2000 copies, about 600 to about 1800 copies, about 700 to about 1600 copies, about 800 to 1400 copies, about 900 to 1200 copies, or about 1000 copies of template polynucleotides.

[0249] By “duoclonal” cluster is meant that the population of polynucleotide sequences that are then sequenced (as the next step) are substantially of two types - e.g. a first sequence and a second sequence. As such, a “duoclonal” cluster may refer to the population of single first sequences and single second sequences within a well that are then sequenced. A “duoclonal” cluster may contain a sufficient number of copies of a single first sequence and copies of a single second sequence such that the cluster is able to output a signal (e.g. a light signal) that allows sequencing reads to be performed on the “monoclonal” cluster. A “duoclonal” cluster may comprise, for example, about 500 to about 2000 combined copies, about 600 to about 1800 combined copies, about 700 to about 1600 combined copies, about 800 to 1400 combined copies, about 900 to 1200 combined copies, or about 1000 combined copies of single first sequences and single second sequences. The copies of single first sequences and single second sequences together may comprise at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, or about 95%, 98%, 99% or 100% of all polynucleotides within a single well of the flow cell, and thus providing a substantially duoclonal “cluster”.

[0250] Non-limiting ways of seeding and tagmenting double-stranded sample polynucleotides are illustrated below.

[0251] Transposome Complexes and Seeding

[0252] Examples of some of the transposome complexes are disclosed herein immobilised within a flow cell depression. While a single set of transposome complexes may be shown in the figures, it is to be understood that several sets of transposome complexes may be attached within a single depression. The transposome complexes may also be immobilised in a spatially separated manner.

[0253] Examples of the transposome complexes are separate entities, but will form dimers in solution and when attached to the flow cell surface. It is to be understood that for simplicity, single transposome complexes are shown in some of the figures. The dimer form of the transposome examples is also shown in some of the figures.

[0254] As mentioned herein, the transposome complexes include a transposase enzyme non-covalently bound to a transposon end. Each transposon end is a double-stranded nucleic acid strand, one strand of which is part of a transferred strand and the other strand of which is part of a non-transferred strand. In other words, the transposon end includes a portion of the transferred strand that is hybridised to a portion of the non-transferred strand.

[0255] In “symmetric” double-stranded seeding, the first transferred strand includes a 5’ end functional group that is capable of covalently or non-covalently attaching, directly or indirectly, to surface functional groups of a polymeric hydrogel in the flow cell depression, a first amplification domain, and a first sequencing primer sequence that is attached to one strand of the first transposon end. The strand of the first transposon end is positioned at the 3’ end of the first transferred strand. Similar to the first transferred strand, the second transferred strand includes a 5’ end functional group that is capable of covalently attaching to surface functional groups of a polymeric hydrogel in the flow cell depression, a second amplification domain, and a second sequencing primer sequence that is attached to one strand of the second transposon end. The strand of the second transposon end is positioned at the 3’ end of the second transferred strand.

[0256] The 5’ end functional groups may be any functional group that is capable of covalently or non-covalently attaching, directly or indirectly, to surface functional groups of a polymeric hydrogel, and thus will depend upon the surface functional groups of the polymer hydrogel. In one example, the polymeric hydrogel includes azide or tetrazine surface groups, and the 5’ end functional groups include a terminal alkyne (e.g., hexynyl) or an internal alkyne, where the alkyne is partof a cyclic compound (e.g., bicyclo[6.1.0]nonyne (BCN)). In another example, the polymeric hydrogel is functionalised with biotin surface groups (referred to herein as biotinylated polymeric hydrogel), and the 5’ end functional groups also include biotin. In these examples, additional streptavidin or avidin is added to indirectly attach the biotin groups to one another.

[0257] The first and second amplification domains have different sequences from each other, but have the same sequence, respectively, as first immobilised primers and second immobilised primers attached to the polymer hydrogel. The first amplification domain and the first immobilised primer together with the second amplification domain and the second immobilised primer enable the amplification of the fragmented target polynucleotides generated during tagmentation.

[0258] Examples of suitable sequences for the first amplification domain and / or first immobilised primer, and for the second amplification domain and / or second immobilised primer include P5 and P7 primer sequences; P15 and P7 primer sequences; or any combination of the PA primer sequences, the PB primer sequences, the PC primer sequences, and the PD primer sequences set forth herein. Examples of P5 and P7 primer sequences are used on the surface of commercial flow cells sold by Illumina Inc. for sequencing, for example, on HiSeq™, HiSeqX™, MiSeq™, MiSeqDX™, MiniSeq™, NextSeq™, NextSeqDX™, NovaSeq™, iSEQ™, Genome Analyzer™, and other instrument platforms.

[0259] The P5 primer sequence may be any one of the following:

[0260] P5 #1: 5’ > 3’ (SEQ ID NO: 3)

[0261] AATGATACGGCGACCACCGAGAnCTACAC

[0262] where “n” is uracil in the sequence.

[0263] P5 #2: 5’ - 3’ (SEQ ID NO: 4)

[0264] AATGATACGGCGACCACCGAGAnCTACAC

[0265] where “n” is alkene-thymidine (i.e., alkene-dT) in the sequence.

[0266] The P7 primer sequence may be any of the following:

[0267] P7 #1: 5’ > 3’ (SEQ ID NO: 5) CAAGCAGAAGACGGCATACGAnAT

[0268] where “n” is 8-oxoguanine in the sequence.

[0269] P7 #2: 5’ - 3’ (SEQ ID NO: 6)

[0270] CAAGC AGAAGACGGC AT AC nAGAT

[0271] where “n” is 8-oxoguanine in the sequence.

[0272] The P15 primer sequence is:

[0273] P15: 5’ > 3’ (SEQ ID NO: 7)

[0274] AATGATACGGCGACCACCGAGAnCTACAC

[0275] where “n” is allyl-T.

[0276] The other primer sequences (PA-PD) mentioned above include:

[0277] PA: 5’ - 3’ (SEQ ID NO: 8)

[0278] GCTGGCACGTCCGAACGCTTCGTTAATCCGTTGAG PB: 5’ - 3’ (SEQ ID NO: 9)

[0279] CGTCGTCTGCCATGGCGCTTCGGTGGATATGAACT PC: 5’ > 3’ (SEQ ID NO: 10)

[0280] ACGGCCGCTAATATCAACGCGTCGAATCCGCAACT PD: 5’ > 3’ (SEQ ID NO: 11)

[0281] GCCGCGTTACGTTAGCCGGACTATTCGATGCAGC

[0282] While not shown in the example sequences for PA-PD, it is to be understood that any of these sequences may include a cleavage site, such as uracil, 8-oxoguanine, allyl-T, diols, etc. at any point in the strand.

[0283] The sequences for the first amplification domain and / or first immobilised primer, and for the second amplification domain and / or second immobilised primer, may be selected to have orthogonal cleavage sites (i.e., one cleavage site is not susceptible to the cleaving agent used for the other cleavage site), so that after amplification, amplified forward and reverse complement strands can be cleaved, or amplified reverse and forward complement strands can be cleaved, leaving the other of the amplified reverse and forward complement strands, or amplified forward and reverse complement strands, for sequencing.

[0284] The first immobilised and second immobilised primers may also include a polyT sequence at the 5’ end of the primer sequence. In some examples, the polyT region includes from 2 T bases to 20 T bases. As specific examples, the polyT region may include 3, 4, 5, 6, 7, or 10 T bases.

[0285] The first and second sequencing primer sequences have different sequences from each other that respectively bind to sequencing primers introduced into the flow cell after tagmentation and amplification. As examples, the first sequencing primer sequence may bind a sequencing primer that primes synthesis of a new strand that is complementary to amplified forward strand and reverse complement strand fragments / fragment amplicons, and the second sequencing primer sequence may bind a sequencing primer that primes synthesis of a new strand that is complementary to amplified reverse strand and forward complement strand fragments / fragment amplicons.

[0286] The first and second transposon ends of each transposome complex include the first and second transferred strands respectively hybridised to the first and second non-transferred strands. As such, the first transferred strand and the first non-transferred strand are complementary, and the second transferred strand and the second non-transferred strand are complementary. The double-stranded first and second transposon ends are respectively capable of complexing with the first and second transposases. As examples, the first transferred strand and the first non-transferred strand, and the second transferred strand and the second non-transferred strand, of the first and second transposon ends may be the related but non-identical 19-base pair (bp) outer end (e.g., first and second transferred strands) and inner end (e.g., first and second non-transferred strands) sequences that serve as the substrate for the activity of the T n5 transposase, or the mosaic ends recognised by a wild-type or mutant Tn5 transposase, or the R1 end (e.g., first and second transferred strands) and the R2 end (e.g., first and second non-transferred strands) recognised by the MuA transposase.

[0287] The polynucleotide sample may be introduced with a tagmentation buffer, which may include water, an optional co-solvent (e.g., dimethylformamide), a metal co-factor for the transposase (e.g., magnesium acetate), and a buffer salt (e.g., Tris(hydroxymethyl) aminomethane (Tris or TRIS) acetate salt, pH 7.6). In an example, the optional co-solvent may be present in an amount up to about 11%, the metal co-factor may be present in a concentration ranging from about 3 mM to about 5.5 mM, and the buffer salt may be present in a concentration ranging from about 7 mM to about 12 mM.

[0288] When the polynucleotide sample is introduced into the flow cell including the first and second transposome complexes, in an embodiment (“symmetric” double-stranded seeding), the polynucleotide sample is fragmented and the 5’ ends of both the forward target strand and reverse target strand of the duplex fragments are ligated to respective 3’ ends of the first and second transferred strands of the transposome complexes. Fragmentation and ligation may take place at a temperature at or above 30°C. In one example, the temperature may range from 30°C to about 55°C. In another example, the temperature may range from 35°C to about 45°C. The 3’ ends of the duplex fragments are not ligated to the 5’ ends of the non-transferred strands. As such, a gap exists between the 3’ end of the forward target strand and the 5’ end of the second non-transferred strand, and a gap exists between the 3’ end of the reverse target strand and the 5’ end of the first non-transferred strand. In one example, each gap is nine (9) base pairs long. The first and second transposases are then removed from the complexes, which are now covalently attached to the forward and reverse target strands via the first and second transferred strands. Transposase removal may be accomplished, for example, using sodium dodecyl sulfate (SDS) or proteinase, or by heating the flow cell to about 60°C. As such, after tagmentation, some example methods involve introducing a washing solution into the flow cell; heating the flow cell, containing the washing solution, to about 60°C; and then introducing an extension amplification mix into the flow cell. Prior to introducing the extension amplification mix, the temperature may be reduced to about 38°C. It has been found that heating of the flow cell after tagmentation may be sufficient to denature the transposase, without including additional buffers or reagents.

[0289] An example extension amplification mix includes a recombinase, a polymerase, and accessory proteins. An example washing solution is an aqueous solution including a buffer agent (e.g., Tris), a salt (e.g., sodium chloride, sodium citrate, etc.), a surfactant (e.g., TWEEN polysorbates), and / or a chelating agent (e.g., EDTA). In one example, the washing solution includes water, the salt at a concentration ranging from about 25 mM to about 50 mM, the surfactant in an amount ranging from about 0.01 wt% to about 0.1 wt%, and optionally the chelating agent. The washing solution may have a relatively high pH, e.g., ranging from about 7 to about 10.

[0290] When the transposase removal is accomplished using sodium dodecyl sulfate (SDS) or another chaotropic detergent, the presence of this reagent can act as an inhibitor for subsequent enzymatic reactions involving, e.g., recombinases, ligases, polymerases, and the like. As such, the presence of the chaotropic detergent may deleteriously affect subsequent library preparation reactions and / or amplification reactions. In this example method, a chelator of the chaotropic detergent may be added to sequester the chaotropic detergent. The chelator that is used depends upon the chaotropic detergent that is used. In an example, the chaotropic detergent is sodium dodecyl sulfate, and the chelator of the chaotropic detergent is a cyclodextrin selected from the group consisting of alphacyclodextrin, betacyclodextrin, and methyl betacyclodextrin. An aqueous solution containing from about 1 wt% to about 4 wt% of the cyclodextrin may be used. In another example, the aqueous solution may include from about 1 wt% to about 2 wt% of the cyclodextrin. Cyclodextrins are cup-shaped structures that have a hydrophilic exterior and a hydrophobic core. These compounds have the ability to form complexes with hydrophilic molecules, such as SDS or another chaotropic detergent that is used to remove the transposase. In particular, the long hydrophobic tail of SDS is drawn into the cup of the cyclodextrin, which effectively locks the SDS following transposase removal and prevents it from denaturing enzymes, including those used in downstream library preparation and / or amplification. Whether heat or the chaotropic detergent is used, the removal of the transposases liberates the partially adapted polynucleotide targets, which include the first and second transferred strands and the forward and reverse target strands respectively attached thereto. Following this purification step to remove the transposases, the first and second nontransferred strands are dehybridised and additional sequences (adaptors) are added to the 3’ ends of the partially adapted fragments by an extension reaction using exclusion amplification reagents (e.g., the ExAMP reagents available from Illumina Inc.). It is to be understood that in examples where SDS or another chaotropic detergent has been chelated with a cyclodextrin, the extension reaction may take place without having to perform one or more wash cycles to remove the SDS or other chaotropic detergent from the flow cell. The extension reaction involves the addition of nucleotides in a template dependent fashion from the 3’ ends of the forward and reverse target strands using the respective first and second transferred strands as the template. As such, the forward target strand is extended along sections to generate complementary sections attached to the forward target strand; and the reverse target strand is extended along sections to generate complementary sections attached to the reverse target strand.

[0291] The introduction of the chaotropic detergent and the chelator may also be used in methods where the transposome complexes are attached to another substrate surface, such as a bead (e.g., a magnetic bead). This example method may more generally involve performing a tagmentation reaction with a polynucleotide sample and a surface bound transposome complex, thereby generating a plurality of partially adapted polynucleotide fragments; introducing a chaotropic detergent to remove a transposase of the transposome complex; introducing a chelator of the chaotropic detergent to sequester the chaotropic detergent; and performing an extension reaction to add additional adapters to the partially adapted polynucleotide target fragments. This method may be performed off of the flow cell, and then the beads (with fully adapted polynucleotide target fragments bound thereto) may be added to a flow cell including capture sites (for attaching the beads) and first and second immobilised primers. The beads can attach to the capture sites and then the fully adapted polynucleotide target fragments may be released from the beads and exposed to amplification using the primers on the flow cell surface.

[0292] In one example of cluster generation, the fully adapted polynucleotide target fragments are denatured from one another, and loop over to hybridise to an adjacent, complementary immobilised primer, and a polymerase copies the copied templates to form double stranded bridges, which are denatured to form two single stranded strands. These two strands loop over and hybridise to adjacent, complementary immobilised primers and are extended again to form two new double stranded loops. The process is repeated on each template copy by cycles of isothermal denaturation and amplification to create dense clonal clusters (bridge amplification). Each cluster of double stranded bridges is denatured. In an example, the amplified reverse and forward complement strands are removed by a cleaving agent suitable for the cleavage site of the first immobilised primers or the second immobilised primers to which the amplified reverse and forward complement strands are attached (e.g., specific base cleavage), leaving amplified forward and reverse complement strands / amplicons. In another example, the amplified forward and reverse complement strands are removed by a cleaving agent suitable for the cleavage site of the first immobilised primers or the second immobilised primers to which the amplified forward and reverse complement strands are attached (e.g., specific base cleavage), leaving amplified reverse and forward complement strands / amplicons. Clustering results in the formation of several template strands immobilised in the depressions.

[0293] In “asymmetric” double-stranded seeding, the transposome complexes are configured for asymmetric attachment to the flow cell surface. As such, one of the complexes, e.g., second transposome complex, includes a 3’ end group for attachment to the flow cell surface, and the other of the complexes, e.g., first transposome complex, includes a 5’ end group for attachment to the flow cell surface. Thus, the second non-transferred strand of the second transposome complex includes the 3’ end group, and the first transferred strand of the first transposome complex includes the 5’ end group. The 3’ end group and the 5’ end group may be any functional group that is capable of covalently or non-covalently attaching, directly or indirectly, to surface functional groups of the polymeric hydrogel, and thus will depend upon the surface functional groups of the polymer hydrogel. In one example, the polymeric hydrogel includes azide or tetrazine surface groups, and the 3’ end group and the 5’ end group each include a terminal alkyne (e.g., hexynyl) or an internal alkyne, where the alkyne is part of a cyclic compound (e.g., bicyclo[6.1.0]nonyne (BCN)). In another example, the biotinylated polymeric hydrogel is present in the depressions of the flow cell, and each of the 3’ end group and the 5’ end group is biotin. In these examples, additional streptavidin or avidin is added to indirectly attach the biotin groups to one another.

[0294] The transposome complexes may be introduced to the flow cell, which includes the depressions separated by the interstitial regions and an example of the polymeric hydrogel in the depressions.

[0295] The transposome complexes may also be included in a carrier liquid in a concentration ranging from about 0.1 pM to about 1 pM. The carrier liquid of this example of the transposome complex fluid may be water. When the polymeric hydrogel is used, a buffer and / or salt may be added to the carrier liquid for grafting the transposome complexes to suitable functional groups of the polymeric hydrogel. The buffer has a pH ranging from 5 to 12. Any of the neutral buffers and / or salts set forth herein may be added to this example of the transposome complex fluid.

[0296] For grafting, the transposome complex fluid is introduced into the flow cell. The transposome complex fluid may be introduced using flowthrough deposition. Grafting may be performed at a temperature ranging from about 35°C to about 45°C for a time ranging from about 30 minutes to about 120 minutes. During grafting, the transposome complexes attach to at least some of the azide or tetrazine groups of the polymeric hydrogel and have no affinity for the interstitial regions 38 or edge portions of the flow cell.

[0297] When the biotinylated polymeric hydrogel is used, the attachment of the transposome complexes to the flow cell surface may take place using any of the examples described herein. Briefly, it is to be understood that avidin or streptavidin may be used to attach the 3’ end group and the 5’ end group (which in this example are biotinylated) with the biotinylated polymeric hydrogel as described herein. As examples, the avidin or streptavidin may be pre-attached to the 3’ and 5’ biotinylated ends before the complexes are introduced into the flow cell and allowed to incubate; or the avidin or streptavidin may be introduced into the flow cell with the 3’ and 5’ biotinylated transposome complexes and allowed to incubate; or the avidin or streptavidin may be introduced into the flow cell and allowed to incubate with the polymeric hydrogel, and then the 3’ and 5’ biotinylated transposome complexes may be added and allowed to incubate; or the avidin or streptavidin and biotin are pre-attached to one another and then introduced into the flow cell where the biotin attaches to the surface of the polymeric hydrogel through some of the RAgroups (directly or through a linker), and then the biotinylated transposome complexes may be introduced into the flow cell containing the streptavidin-biotin bound pair as part of the polymeric hydrogel.

[0298] The flow cell also includes the first immobilised primers and the second immobilised primers. Any of first immobilised primers and the second immobilised primers and suitable attachment mechanisms of first immobilised primers and the second immobilised primers to the polymeric hydrogel described herein may be used for first immobilised primers and the second immobilised primers.

[0299] In this example, the plurality of second transposome complexes (which may be referred to herein as the 3’ attached transposome complexes) and the plurality of first transposome complexes (which may be referred to herein as the 5’ attached transposome complexes) are present in the flow cell at a ratio ranging from greater than 1: 1 to 5: 1. In another example, the 3’ biotinylated transposome complexes and the 5’ biotinylated transposome complexes are introduced into the flow cell at a ratio of 4:1.

[0300] The polynucleotide sample may be introduced with a tagmentation buffer, which may include water, an optional co-solvent (e.g., dimethylformamide), a metal co-factor for the transposase (e.g., magnesium acetate), and a buffer salt (e.g., Tris acetate salt, pH 7.6). In an example, the optional co-solvent may be present in an amount up to about 11%, the metal co-factor may be present in a concentration ranging from about 3 mM to about 5.5 mM, and the buffer salt may be present in a concentration ranging from about 7 mM to about 12 mM.

[0301] When the polynucleotide sample is introduced into the flow cell including the 3’ and 5’ attached transposome complexes (“asymmetric” double-stranded seeding), the polynucleotide sample is fragmented and the 5’ ends of both the forward target strand and reverse target strand of the duplex fragments are ligated to respective 3’ ends of the first and second transferred strands of the transposome complexes. Fragmentation and ligation may take place at a temperature at or above 30°C. In one example, the temperature may range from 30°C to about 55°C. In another example, the temperature may range from 35°C to about 45°C. The 3’ ends of the duplex fragments are not ligated to the 5’ ends of the non-transferred strands. As such, a gap exists between the 3’ end of the forward target strand and the 5’ end of the second non-transferred strand, and a gap exists between the 3’ end of the reverse target strand and the 5’ end of the first non-transferred strand. In one example, each gap is nine (9) base pairs long.

[0302] The first and second transposases are then removed from the complexes, which are now attached to the forward and reverse target strands via the first and second transferred strands. Transposase removal may be accomplished by heating the flow cell. In this example, a washing solution is introduced into the flow cell; the flow cell containing the washing solution is heated to about 60°C; and then the extension amplification mix is introduced into the flow cell. Prior to the introduction of the extension amplification mix, the temperature may be lowered to about 38°C.

[0303] The extension mix is then added to form the fully extended fragment with adapters at both ends. The extension of the forward and reverse target strands involves the addition of nucleotides in a template dependent fashion from the 3’ ends of the forward and reverse target strands. When at least some of the first non-transferred strands and first transferred strands, and the second non-transferred strands and second transferred strands, remain hybridised after transposase removal, the extension reaction may displace the non-transferred strands and allow the transferred strands to be copied, thus forming fully extended (fully adapted) fragments. When at least some of first non-transferred strands and first transferred strands, and the second non-transferred strands and second transferred strands, dehybridise during transposase removal, because the non-transferred strands are relatively short, their melting temperature may be low (e.g., from about 40°C to about 50°C); as such, the transposase removal temperature may be sufficient to dehybridise at least some of the transposon ends. In contrast, the longer forward target strands and reverse target strands may remain hybridised. It is to be understood that the longer forward target strands and reverse target strands may dehybridise at regions (e.g., AT rich regions) that have a lower melting temperature, but the overall insert size includes regions with higher melting temperature that do not dehybridise. These regions keep the longer forward target strands and reverse target strands from falling apart. The insert size may be shifted if the transposase removal temperature is increased above 60°C, which can lead to dehybridisation of some of the forward target strands and reverse target strands.

[0304] In another example, transposase removal is accomplished using sodium dodecyl sulfate (SDS) or proteinase, followed by the introduction of the wash solution, dehybridisation of the non-transferred strands, and the extension amplification reaction.

[0305] Flow Cells and Seeding

[0306] The transposome complexes disclosed herein, whether modified or non-modified, are immobilised on a flow cell surface. In one example, the transposome complexes are attached to the polymeric hydrogel within the depression of the flow cell. In another example, the transposome complexes are asymmetrically attached to the polymeric hydrogel within the depression of the flow cell. While a single depression is shown may be referred to herein and in the figures, it is to be understood that a plurality of depressions may be formed across the substrate surface in an array as described herein.

[0307] Each example of the flow cell disclosed herein includes a substrate having depressions separated by interstitial regions, and a polymeric hydrogel within the depressions. This type of substrate may be referred to as a patterned substrate, and the flow cell may include two patterned substrates bonded together, or may include a single patterned substrate bonded to a lid.

[0308] A flow channel is defined between the substrates or the substrate and the lid. The depth of the flow channel can be as small as a monolayer thick when microcontact, aerosol, or inkjet printing is used to deposit a material to bond the substrates or the substrate and the lid. In other examples, the depth of the flow channel can be about 1 pm, about 10 pm, about 50 pm, about 100 pm, or more. In an example, the depth may range from about 10 pm to about 100 pm. In another example, the depth may range from about 10 pm to about 30 pm. In still another example, the depth is about 5 pm or less. It is to be understood that the depth of the flow channel may be greater than, less than or between the values specified above.

[0309] Each flow channel is in fluid communication with an inlet and an outlet. The inlet and outlet may be positioned at opposed ends of the flow cell. The inlets and outlets of the respective flow channels may alternatively be positioned anywhere along the length and width of the flow channel that enables desirable fluid flow. The inlet allows fluids to be introduced into the flow channel, and the outlet allows fluid to be extracted from the flow channel. Each of the inlets and outlets is fluidly connected to a fluidic control system (including, e.g., reservoirs, pumps, valves, waste containers, and the like) which controls fluid introduction and expulsion.

[0310] The substrate may be a single layer base support having the depressions defined therein, or a multi-layer structure including a base support with another layer positioned thereon and having the depressions defined therein.

[0311] Examples of suitable single layer base supports include epoxy siloxane, glass, modified or functionalised glass, plastics (including acrylics, polystyrene and copolymers of styrene and other materials, polypropylene, polyethylene, polybutylene, polyurethanes, polytetrafluoroethylene (such as TEFLON® from Chemours), cyclic olefins / cyclo-olefin polymers (COP) (such as ZEONOR® from Zeon), polyimides, etc.), nylon (polyamides), ceram ics / ceramic oxides, silica, fused silica, or silica-based materials, aluminum silicate, silicon and modified silicon (e.g., boron doped p+ silicon), silicon nitride (SisN^, silicon oxide (SiO2), tantalum pentoxide (Ta2Os) or other tantalum oxide(s) (TaOx), hafnium oxide (HfO2), carbon, metals, inorganic glasses, or the like.

[0312] Examples of the multi-layered structure include any example of the base support and at least one other layer on the base support.

[0313] The other layer may be an inorganic oxide, such as tantalum oxide (e.g., Ta2Os), aluminium oxide (e.g., AI2O3), silicon oxide (e.g., SiCh), hafnium oxide (e.g., HfCh), etc. Inorganic oxides could be selectively applied, e.g., via vapor deposition, aerosol printing, or inkjet printing, to form the depressions and interstitial regions.

[0314] The other layer may alternatively be a polymeric resin that can be applied to the base support and then patterned. Some examples of suitable polymeric resins include a polyhedral oligomeric silsesquioxane based resin, an epoxy resin not based on a polyhedral oligomeric silsesquioxane, a poly(ethylene glycol) resin, a polyether resin (e.g., ring opened epoxies), an acrylic resin, an acrylate resin, a methacrylate resin, an amorphous fluoropolymer resin (e.g., CYTOP® from Bellex), and combinations thereof. Suitable deposition techniques include chemical vapor deposition, dip coating, dunk coating, spin coating, spray coating, puddle dispensing, ultrasonic spray coating, doctor blade coating, aerosol printing, screen printing, microcontact printing, etc. Suitable patterning techniques include photolithography, nanoimprint lithography (NIL), stamping techniques, embossing techniques, molding techniques, microetching techniques, etc.

[0315] In an example, the substrate may be a circular sheet, a panel, a wafer, a die etc. having a diameter ranging from about 2 mm to about 300 mm, e.g., from about 200 mm to about 300 mm, or may be a rectangular sheet, panel, wafer, die etc. having its largest dimension up to about 10 feet (~ 3 meters). While example dimensions have been provided, it is to be understood that a substrate with any suitable dimensions may be used.

[0316] The depressions may be formed in the substrate using any suitable patterning techniques, such as photolithography, nanoimprint lithography (NIL), stamping techniques, embossing techniques, molding techniques, microetching techniques, etc.

[0317] Many different layouts of the depressions and interstitial regions may be envisaged, including regular, repeating, and non-regular patterns. In an example, the plurality depressions and interstitial regions are disposed to create a hexagonal grid for close packing and improved density. Other layouts may include, for example, rectangular layouts, triangular layouts, and so forth. In some examples, the layout or pattern can be an x-y format in rows and columns. In some other examples, the layout or pattern can be a repeating arrangement of the depressions and the interstitial regions. In still other examples, the layout or pattern can be a random arrangement of the plurality of depressions and the interstitial regions.

[0318] The layout or pattern may be characterised with respect to the density (number) of the plurality depressions within a defined area. For example, the depressions may be present at a density of approximately 2 million per mm2. The density may be tuned to different densities including, for example, a density of about 100 per mm2, about 1,000 per mm2, about 0.1 million per mm2, about 1 million per mm2, about 2 million per mm2, about 5 million per mm2, about 10 million per mm2, about 50 million per mm2, or more, or less. It is to be further understood that the density can be between one of the lower values and one of the upper values selected from the ranges above, or that other densities (outside of the given ranges) may be used. As examples, a high density array may be characterised as having the plurality of depressions separated by less than about 100 nm, a medium density array may be characterised as having the plurality of depressions separated by about 400 nm to about 1 pm, and a low density array may be characterised as having the plurality of depressions separated by greater than about 1 pm.

[0319] The layout or pattern of the plurality of depressions may also or alternatively be characterised in terms of the average pitch, or the spacing from the center of one depression to the center of an adjacent depression (center-to-center spacing) or from the right edge of one depressions to the left edge of an adjacent depression (edge-to-edge spacing). The pattern can be regular, such that the coefficient of variation around the average pitch is small, or the pattern can be non-regular in which case the coefficient of variation can be relatively large. In either case, the average pitch can be, for example, about 50 nm, about 0.15 pm, about 0.5 pm, about 1 pm, about 5 pm, about 10 pm, about 100 pm, or more or less. The average pitch for a particular pattern of can be between one of the lower values and one of the upper values selected from the ranges above. In an example, the depressions have a pitch (center-to-center spacing) of about 1.5 pm. While example average pitch values have been provided, it is to be understood that other average pitch values may be used.

[0320] The size of each depression may be characterised by its volume, opening area, depth, and / or diameter (when the depression is circular) and / or length and width. For example, the volume can range from about 1x10"3m3to about 100 pm3, e.g., about 1x10-2pm3, about 0.1 pm3, about 1 pm3, about 10 pm3, or more, or less. For another example, the opening area can range from about 1xio-3pm2to about 100 pm2, e.g., about 1X10-2pm2, about 0.1 pm2, about 1 pm2, at least about 10 pm2, or more, or less. For still another example, the depth can range from about 0.1 pm to about 100 pm, e.g., about 0.5 pm, about 1 pm, about 10 pm, or more, or less. For another example, the depth can range from about 0.1 pm to about 100 pm, e.g., about 0.5 pm, about 1 pm, about 10 pm, or more, or less. For yet another example, the diameter or each of the length and width can range from about 0.1 pm to about 100 pm, e.g., about 0.5 pm, about 1 pm, about 10 pm, or more, or less.

[0321] The substrate may include edge regions that define interstitial like regions that extend the length of the flow channels and separate one flow channel from an adjacent flow channel. The edge regions provide bonding regions where two substrates can be attached to one another or where one substrate can be attached to the lid.

[0322] The depressions provide a designated area for the polymeric hydrogel. In the examples disclosed herein, the polymeric hydrogel includes an acrylamide copolymer. In this example, the acrylamide copolymer has a structure:

[0323]

[0324] wherein:

[0325] RAis an azide or a tetrazine or any other functional group that can attach to an alkyne;

[0326] RBis H or optionally substituted alkyl;

[0327] Rc, RD, and REare each independently selected from the group consisting of H and optionally substituted alkyl; each of the -(CH2)P- can be optionally substituted;

[0328] p is an integer in the range of 1 to 50;

[0329] n is an integer in the range of 1 to 50,000; and

[0330] m is an integer in the range of 1 to 100,000.

[0331] One specific example of the acrylamide copolymer represented by structure (I) is poly(N-(5-azidoacetamidylpentyl)acrylamide-co-acrylamide, PAZAM.

[0332] One of ordinary skill in the art will recognise that the arrangement of the recurring “n” and “m” features in structure (I) are representative, and the monomeric subunits may be present in any order in the polymer structure (e.g., random, block, patterned, or a combination thereof).

[0333] The molecular weight of the acrylamide copolymer may range from about 5 kDa to about 1500 kDa or from about 10 kDa to about 1000 kDa, or may be, in a specific example, about 312 kDa. The molecular weight may be a weight average molecular weight, or a number average molecular weight. The molecular weight may be measured using gel permeation chromatography.

[0334] In some examples, the acrylamide copolymer is a linear polymer. In some other examples, the acrylamide copolymer is a lightly cross-linked polymer.

[0335] In some examples, the gel material may be a variation of structure (I). In one example,

[0336]

[0337] the acrylamide unit may be replaced with N, N-dimethylacrylamide ( ). In another

[0338]

[0339] example, the acrylamide unit in structure (I) may be replaced with,

[0340] where RD, RE, and RFare each H or a Ci-Ce alkyl, and RGand RHare each a Ci-Ce alkyl (instead of H as is the case with the acrylamide). In this example, q may be an integer in the range of 1 to 100,000. In another example, the N, N-dimethylacrylamide may be used in addition to the acrylamide unit. In this example, structure (I) may include

[0341]

[0342] in addition to the recurring “n” and “m” features, where RD, RE, and RFare each H or a Ci-Ce alkyl, and RGand RHare each a Ci-Ce alkyl. In this example, q may be an integer in the range of 1 to 100,000.

[0343] As another example of the polymeric hydrogel, the recurring “n” feature in structure (I) may be replaced with a monomer including a heterocyclic azido group having structure (II):

[0344]

[0345] wherein R1is H or a Ci-Ce alkyl; R2 is H or a Ci-Ce alkyl; L is a linker including a linear chain with 2 to 20 atoms selected from the group consisting of carbon, oxygen, and nitrogen and 10 optional substituents on the carbon and any nitrogen atoms in the chain; E is a linear chain including 1 to 4 atoms selected from the group consisting of carbon, oxygen and nitrogen, and optional substituents on the carbon and any nitrogen atoms in the chain; A is an N substituted amide with an H or a C1-C4 alkyl attached to the N; and Z is a nitrogen containing heterocycle. Examples of Z include 5 to 10 carbon-containing ring members present as a single cyclic structure or a fused structure. Some specific examples of Z include pyrrolidinyl, pyridinyl, or pyrimidinyl.

[0346] As still another example, the gel material may include a recurring unit of each of structure (III) and (IV):

[0347]

[0348] wherein each of R1a, R2a, R1band R2bis independently selected from hydrogen, an optionally substituted alkyl or optionally substituted phenyl; each of R3aand R3bis independently selected from hydrogen, an optionally substituted alkyl, an optionally substituted phenyl, or an optionally substituted C7-C14 aralkyl; and each L1and L2is independently selected from an optionally substituted alkylene linker or an optionally substituted heteroalkylene linker

[0349] In some examples, the polymeric hydrogel is biotinylated. In these examples, biotin is attached to the surface of the polymeric hydrogel through some of the RAgroups (i.e., the azide, tetrazine, or other functional group that can attach to an alkyne). The biotin is attached to a linker, such as bicyclo[6.1.0]nonyne (BCN), which can covalently attach to some of the RAgroups. The combination of the biotin and the linker is one example of a biotin-containing linker.

[0350] In other examples, streptavidin and biotin are attached to one another, and the biotin is attached to the surface of the polymeric hydrogel through some of the RAgroups (i.e., the azide, tetrazine, or other functional group that can attach to an alkyne). The biotin may be attached to a linker, such as bicyclo[6.1.0]nonyne (BCN), which can covalently attach to some of the RAgroups. In this example, the biotin-containing linker includes the linker, the biotin, and the streptavidin.

[0351] In still other examples, the bicyclo[6.1.0]nonyne (BCN) linkers can be attached to some of the RAgroups of the polymeric hydrogel, and the biotin is not attached. In these examples, biotin can added after the polymeric hydrogel is applied to the depressions.

[0352] To introduce the polymeric hydrogel into the depressions, a mixture of the polymeric hydrogel may be generated and then applied to the substrate. In one example, the polymeric hydrogel may be present in a mixture (e.g., with water or with ethanol and water). The mixture may then be applied to the substrate surface using spin coating, or dipping or dip coating, or flow of the material under positive or negative pressure, or another suitable technique. These types of techniques blanketly deposit the polymeric hydrogel in the depressions and on the interstitial regions. Other selective deposition techniques (e.g., involving a mask, controlled printing techniques, etc.) may be used to specifically deposit the polymeric hydrogel in the depressions and not on the interstitial regions.

[0353] In some examples, the substrate may be activated, and then the mixture (including the polymeric hydrogel) may be applied thereto. In one example, a silane or silane derivative (e.g., norbornene silane) may be deposited on the surface of the substrate using vapor deposition, spin coating, or other deposition methods. In another example, the substrate may be exposed to plasma ashing to generate surface-activating agent(s) (e.g., -OH groups) that can adhere to the polymeric hydrogel.

[0354] The applied mixture may be exposed to a curing process to cure the polymeric hydrogel. In an example, curing may take place at a temperature ranging from room temperature (e.g., about 25°C) to about 95°C for a time ranging from about 1 millisecond to about several days.

[0355] Polishing may then be performed in order to remove the polymeric hydrogel from the interstitial regions, while leaving the polymeric hydrogel on the surface in the depressions at least substantially intact. The polishing process may be performed with a chemical slurry (including, e.g., an abrasive, a buffer, a chelating agent, a surfactant, and / or a dispersant) which can remove the polymeric hydrogel from the interstitial regions without deleteriously affecting the underlying substrate at those regions. Alternatively, polishing may be performed with a solution that does not include the abrasive particles.

[0356] The chemical slurry may be used in a chemical mechanical polishing system to polish the surface of the interstitial regions. The polishing head(s) / pad(s) or other polishing tool(s) is / are capable of polishing the polymeric hydrogel that may be present over the interstitial regions while leaving the polymeric hydrogel in the depressions at least substantially intact. As an example, the polishing head may be a Strasbaugh ViPRR II polishing head.

[0357] Cleaning and drying processes may be performed after polishing. The cleaning process may utilise a water bath and sonication. The water bath may be maintained at a relatively low temperature ranging from about 22°C to about 30°C. The drying process may involve spin drying, or drying via another suitable technique.

[0358] Each example of the flow cell also includes the first immobilised primers and the second immobilised primers, and the first and second transposome complexes. The first immobilised primers and the second immobilised primers include any two of the primer sequences set forth herein for the first and second amplification domains. The 5’ terminal end of first immobilised primers and the second immobilised primers will vary depending upon the chemistry of the polymeric hydrogel. Different examples of how first immobilised primers and the second immobilised primers and the first and second transposome complexes can be attached within the flow cell will be described.

[0359] In some embodiments, both of the first immobilised primers and the second immobilised primers and both of the first and second transposome complexes are attached to the polymeric hydrogel in the depression. In one example, the 5’ ends of the first immobilised primers and the second immobilised primers and of the first and second transposome complexes include functional groups that can covalently attach to the azide and / or tetrazine functional groups of the polymeric hydrogel. As examples, the 5’ end functional groups may be a terminal alkyne (e.g., hexynyl) or an internal alkyne, where the alkyne is part of a cyclic compound (e.g., bicyclo[6.1.0]nonyne (BCN)).

[0360] The first immobilised primers and the second immobilised primers may be included in a carrier liquid in a concentration ranging from about 0.5 pM to about 100 pM. In one example, the primer concentration ranges from about 5 pM to about 25 pM.

[0361] The carrier liquid of the primer fluid may be water. A buffer and / or salt may be added to the carrier liquid for grafting the first immobilised primers and the second immobilised primers to suitable functional groups of the polymeric hydrogel. The buffer has a pH ranging from 5 to 12, and the buffer used will depend upon the alkyne at the 5’ end of the first immobilised primers and the second immobilised primers. A neutral buffer and / or salt may be added to the primer fluid for grafting BCN terminated primers, while an alkaline buffer may be added to the primer fluid for copper-assisted grafting methods (e.g., the click reaction). Any of the primer fluids used in copper-assisted grafting methods may also include a copper catalyst. Example of neutral buffers include Tris(hydroxymethyl) aminomethane (Tris or TRIS) buffers, such as Tris-HCI or Tris-EDTA, or a carbonate buffer (e.g., 0.25 M to 1 M). Sodium sulfate (e.g., 1 M to 2 M) is a suitable salt that may be used. Examples of alkaline buffers include Tris(hydroxymethyl) aminomethane (CHES), 3-(Cyclohexylamino)-1-propanesulphonic acid (CAPS), and alkaline buffer solution (from Sigma-Aldrich).

[0362] For grafting, the primer fluid is introduced into the flow cell. The primer fluid may be introduced using flow through deposition. Grafting may be performed at a temperature ranging from about 55°C to about 65°C for a time ranging from about 20 minutes to about 60 minutes. In one example, grafting is performed at 60°C for about 30 minutes or 60 minutes. It is to be understood that a lower temperature and a longer time or a higher temperature and a shorter time may also be used. Some primer grafting techniques, such as those involving BCN grafting to tetrazine units, may be performed at room temperature (e.g., 18°C to about 25°C). During grafting, the first immobilised primers and the second immobilised primers attach to at least some of the azide or tetrazine groups of the polymeric hydrogel and have no affinity for the interstitial regions or edge portions of the flow cell. The first and second transposome complexes may also be included in a carrier liquid in a concentration ranging from about 0.1 pM to about 1 pM.

[0363] The carrier liquid of the transposome complex fluid may be water. A buffer and / or salt may be added to the carrier liquid for grafting the first and second transposome complexes to suitable functional groups of the polymeric hydrogel. The buffer has a pH ranging from 5 to 12. Any of the neutral buffers and / or salts set forth herein may be added to the transposome complex fluid.

[0364] For grafting, the transposome complex fluid is introduced into the flow cell. The transposome complex fluid may be introduced using flowthrough deposition. Grafting may be performed at a temperature ranging from about 35°C to about 45°C for a time ranging from about 30 minutes to about 120 minutes. In one example, grafting is performed at 37°C for about 90 minutes or 120 minutes. During grafting, the first and second transposome complexes attach to at least some of the azide or tetrazine groups of the polymeric hydrogel and have no affinity for the interstitial regions or edge portions of the flow cell.

[0365] The separate transposome complexes form dimers in solution, and these dimers attach to the polymeric hydrogel.

[0366] Transposome complex grafting may take place during flow cell manufacturing. Alternatively, the transposome complex fluid may be contained in a reagent cartridge, and transposome complex grafting may take place on the sequencing instrument prior to sample preparation.

[0367] In some embodiments, the attachment mechanisms utilise the non-covalent interaction between biotin and streptavidin.

[0368] Thus, in some embodiments, the polymeric hydrogel is biotinylated. The biotinylated polymeric hydrogel contains biotin-containing linkers grafted to the surface of the polymeric hydrogel. The biotinylated polymer hydrogel may be prepared by grafting bicyclononyne-biotin (BCN-biotin) to some of terminal azide or tetrazine groups of the polymeric hydrogel. The biotin-containing linkers may be grafted to the polymeric hydrogel prior to the polymeric hydrogel being added to the flow cell surface, or after the polymeric hydrogel has been added to the flow cell surface.

[0369] In one example method, the polymeric hydrogel is applied to the depressions, and then the biotin-containing linkers grafted to the surface of the polymeric hydrogel. In this example, the biotin-containing linkers may be added to a carrier fluid (e.g., water including a neutral buffer and / or salt as described herein), and the fluid may be introduced to the flow cell and allowed to incubate. Grafting may be performed at a temperature ranging from about 50°C to about 70°C for a time ranging from about 30 minutes to about 120 minutes. In one example, biotin-containing linker grafting is performed at 60°C for about 120 minutes. During grafting, the biotin-containing linkers attach to at least some of the azide or tetrazine groups of the polymeric hydrogel and have no affinity for the interstitial regions or edge portions of the flow cell. Once grafted, the biotin of the biotin-containing linkers is at the surface, and thus is available for interaction with subsequently introduced streptavidin.

[0370] The first immobilised primers and the second immobilised primers are grafted to some other of the terminal azide or tetrazine groups of the biotinylated polymer hydrogel.

[0371] It is to be understood that the first and second transposome complexes used in this example include biotin as the 5’ end functional groups. Thus, the transferred strands of these example transposome complexes include a 5’ biotinylated end.

[0372] In one example, these biotinylated transposome complexes are introduced to the flow cell, which includes the depressions separated by the interstitial regions and the biotinylated polymeric hydrogel in the depressions. This example method also includes introducing streptavidin to the flow cell, and incubating the transposome complexes and the streptavidin in the flow cell, whereby the streptavidin respectively attaches the 5’ biotinylated end of at least some of the transposome complexes to the biotinylated polymeric hydrogel. Incubation of the transposome complexes and the streptavidin in the flow cell may take place at room temperature (e.g., from about 18°C to about 25°C) for a time ranging from about 15 minutes to about 60 minutes. In one example, this incubation takes place for about 30 minutes at about 22°C.

[0373] In this example, the transposome complexes and the streptavidin may be added at a weight ratio ranging from about 1:1 to about 1:2 so that an ample amount of streptavidin is present to graft the transposome complexes to the biotin-containing linkers.

[0374] The streptavidin interacts with the biotin at the surface of the biotinylated polymeric hydrogel and with the 5’ end biotin functional groups.

[0375] In other embodiments, the first immobilised primers and the second immobilised primers are also biotinylated. Instead of the 5’ end of the first immobilised primers and the second immobilised primers containing an alkyne, the 5’ end contains biotin. These first immobilised primers and second immobilised primers are referred to as biotinylated primers.

[0376] In this example, biotinylated transposome complexes (i.e., complexes with 5’ end biotin functional groups) and the biotinylated primers (i.e., the first immobilised primers and the second immobilised primers with 5’ end biotin) respectively attach to biotin-containing linkers of the biotinylated polymeric hydrogel. In this example, streptavidin is introduced into the flow cell with the biotinylated transposome complexes in order to achieve the desired biotin-streptavidin-biotin interaction, and streptavidin is introduced into the flow cell with the biotinylated primers in order to achieve the desired biotin-streptavidin-biotin interaction. Thus, in this example, the first and second biotinylated immobilised primers are immobilised to the biotinylated polymeric hydrogel through streptavidin; the first biotinylated transposomes (with biotin 5’ end functional groups) are immobilised to the biotinylated polymeric hydrogel through streptavidin; and the second biotinylated transposome complexes (with biotin 5’ end functional groups) are immobilised to the biotinylated polymeric hydrogel through streptavidin.

[0377] Because the biotin-streptavidin interactions are reversible, the flow cell may be reusable. For example, after DNA sample tagmentation, fragment amplification, clustering, and sequencing, a biotin streptavidin cleavage composition may be introduced into the flow cell. Examples of the biotin streptavidin cleavage composition include about 95% formamide and about 10mM ethylenediaminetetraacetic acid (EDTA), or from about 10% by volume to about 50% by volume of a formamide reagent and a balance of a salt buffer. At suitable reaction temperatures, these cleavage compositions disrupt the biotin-streptavidin interactions, thus releasing whatever is attached (e.g., target polynucleotides, nascent strands, etc.) to the biotin. Alternatively a hot wash with biotin and / or desthiobiotin causes the newly added biotin and / or desthiobiotin to compete with the already bound biotin.

[0378] Thus, the example described is a reusable flow cell, comprising a substrate including depressions separated by interstitial regions; a biotinylated polymeric hydrogel in the depressions; first and second biotinylated primers immobilised within the depressions; first biotinylated transposome complexes (with biotin 5’ end functional groups) immobilised within the depressions, the first biotinylated transposome complexes including a first amplification domain; and second biotinylated transposome complexes (with biotin 5’ end functional groups) immobilised within the depressions, the second biotinylated transposome complexes including a second amplification domain.

[0379] In another example, the streptavidin may be introduced into the flow cell before the biotinylated transposome complexes are introduced into the flow cell. The streptavidin may be introduced into the flow cell at room temperature and binding of the streptavidin to the biotin of the polymeric hydrogel may take place in less than 30 minutes. The biotinylated transposome complexes may then be introduced into the flow cell and allowed to incubate to attach to the pre-assembled streptavidin.

[0380] In still another example, the streptavidin may be mixed with the biotinylated transposome complexes before they are introduced into the flow cell. The streptavidin and biotinylated transposome complexes may be allowed to incubate at room temperature for about 30 minutes or less. This will pre-assemble the streptavidin onto the biotinylated transposome complexes. The streptavidin-biotinylated transposome complexes may then be introduced into the flow cell and allowed to incubate to attach to the biotin of the polymeric hydrogel. In still another example, streptavidin and biotin are pre-attached to one another to form a streptavidin-biotin bound pair, and then the biotin of the streptavidin-biotin bound pair is attached to the surface of the polymeric hydrogel through some of the RAgroups (i.e., the azide, tetrazine, or other functional group that can attach to an alkyne). The biotin of the streptavidin-biotin bound pair may be attached to a linker, such as bicyclo[6.1.0]nonyne (BCN), which can covalently attach to some of the RAgroups. In this example, biotinylated transposome complexes may be introduced into the flow cell containing the streptavidin-biotin bound pair as part of the polymeric hydrogel.

[0381] In still other examples, the transposome complexes may be fully assembled with biotin-streptavidin-biotin, where biotin is the linker attached to streptavidin attached to another biotin. In this example, the polymeric hydrogel is used, which includes surface groups or surface bound linkers, such as bicyclo[6.1.0]nonyne (BCN), that can attach to the biotin of the fully assembled transposome complexes.

[0382] Other examples of the flow cell include the transposome complexes attached to the substrate surface in spatially separated arrangements. These examples of the flow cell include a substrate having depressions separated by interstitial regions; first and immobilised second primers immobilised within the depressions; first transposome complexes including a first amplification domain; second transposome complexes including a second amplification domain; wherein one of: i) the first and second transposome complexes are respectively immobilised at different regions of the depressions; or ii) the first and second transposome complexes are respectively immobilised within the depressions and on the interstitial regions; or iii) the first and second transposome complexes are respectively immobilised on different areas of the interstitial regions.

[0383] In one embodiment, the polymeric hydrogel includes two regions A, B. One region A has the first immobilised primers and the first transposome complexes attached thereto, and the other region B has the second immobilised primers and the second transposome complexes attached thereto.

[0384] In one embodiment, the polymeric hydrogel is chemically the same throughout the regions A, B, and suitable techniques may be used to immobilise the first immobilised primers, the second immobilised primers, first transposome complexes and second transposome complexes to the respective regions A, B. For example, the polymeric hydrogel across both regions A, B may include azides and / or tetrazines, which can attach to alkynes (e.g., BCN). One example of a suitable technique may include the use of a photoresist. In this example, the photoresist is developed to mask one region A while the surface chemistry (e.g., the first immobilised primers, the second immobilised primers, first transposome complexes and second transposome complexes) is added to the other region B. The photoresist is then removed, and the surface chemistry to the region A under conditions that will not deleteriously affect the region B or its surface chemistry. Another example of a suitable technique may include the use of a sacrificial layer, such as aluminium. In this example, the sacrificial layer is applied to the substrate so that it masks a portion (e.g., one half) of the depression. The other portion (e.g., other half) of the depression remains exposed. The exposed portion of the depression may be activated, e.g., by depositing a suitable silane (e.g., depositing norbornene silane using CVD). In this example, the polymeric hydrogel is applied over the sacrificial layer and over the exposed and activated portion of the depression. The sacrificial layer is then lifted off, which removes the overlying polymeric hydrogel and exposes the previously covered portion of the depression. The polymeric hydrogel on the other portion of the depression remains intact. The first immobilised primers, the second immobilised primers, first transposome complexes and second transposome complexes may then be grafted to the polymeric hydrogel. In one example, first immobilised primer grafting involves azide reduction of the polymeric hydrogel in region A, followed by application of an NHS-tetrazine agent, followed by grafting of BCN -terminated first transposome complexes and BCN -terminated first immobilised primers. The polymeric hydrogel in region A may be exposed to a capping reagent to deactivate any unreacted tetrazine groups. Because the substrate has no affinity for the first immobilised primers and first transposome complexes, the exposed portion of the depression is not affected. The substrate may then be activated, e.g., by exposure to a solution of norbornene silane. The polymeric hydrogel in region B may then be applied under conditions (e.g., under high ionic strength (e.g., in the presence of 10x PBS, NaCI, KCI, etc.)) that will not deleteriously affect the region A or its surface chemistry. The second immobilised primers and second transposome complexes may then be grafted to the polymeric hydrogel in region B. Still another example of a suitable technique may include pre-grafting the first immobilised primers, the second immobilised primers, the first transposome complexes and the second transposome complexes, and sequentially and selectively applying the pre-grafted polymeric materials to the regions A, B. It is to be understood that while several example methods have been provided, any other methods may be used to immobilise the first immobilised primers, the second immobilised primers, the first transposome complexes and the second transposome complexes to the respective regions A, B.

[0385] In still another example, the regions A, B of the polymeric hydrogel may be chemically different (i.e., orthogonal so that one region A can graft first immobilised primers and the other region B can graft second immobilised primers). In one example, the biotinylated polymeric hydrogel may be applied as region A and the azide or tetrazine terminated polymeric hydrogel may be applied as region B for respective immobilised primer and transposome complex attachment. In another example, the polymeric hydrogel in region A has tetrazine terminal groups and the polymeric hydrogel in region B has alkyne (non-BCN) terminal groups. In this example, first immobilised primers and first transposome complexes terminated with transcyclooctene (TCO), and second immobilised primers and second immobilised complexes terminated with picolyl azide, can simultaneously and respectively be grafted to the regions A and B of the polymeric hydrogel. In still another example, the polymeric hydrogel in region A has tetrazine terminal groups and the polymeric hydrogel in region B has dibenzocyclooctyne terminal groups. In this example, first immobilised primers and first transposome complexes terminated with trans-cyclooctene (TCO), and second immobilised primers and second immobilised complexes terminated with picolyl azide, can simultaneously and respectively be grafted to the regions A and B of the polymeric hydrogel. While several examples have been provided, it is to be understood that any combination of orthogonal chemistries may be used at the regions A and B in order to respectively graft the first immobilised primers and the second immobilised primers.

[0386] In a further embodiment, the polymeric hydrogel includes two regions A, B. One region A has the first immobilised primers and the first transposome complexes attached thereto, and the other region B has the second immobilised primers attached thereto. In this example, the second transposome complexes are attached to the interstitial regions. The methods already described above may be used to attach the first immobilised primers and the first transposome complexes to the region A and to attach the second immobilised primers to the region B. In this example, the interstitial regions may have material selectively added thereto that is capable of attaching the second transposome complexes thereto. Selective deposition techniques or salt exclusion methods may be used to selectively apply the material on the interstitial regions.

[0387] In a further embodiment, the polymeric hydrogel includes two regions A, B. One region A has the first immobilised primers attached thereto, and the other region B has the second immobilised primers attached thereto. In this example, the first and second transposome complexes are attached to areas of the interstitial regions. The methods already described above may be used to attach the first immobilised primers to the region A and to attach the second immobilised primers to the region B. The linkers of the transposome complexes may be selected so that they attach to the material applied to the interstitial regions or to the substrate (if the material is not used).

[0388] In still another example, the spatially separated arrangement of the transposome complexes is achieved using a multi-depth depression. In addition to enabling spatial separation, the use of this particular depression may also lead to increased insert sizes. The multi-depth depression includes a shallow portion and a deep portion. The depths of the shallow portion and the deep portion are within the ranges set forth herein for the depression, with the caveat that the deep portion has a greater depth than the shallow portion.

[0389] The polymeric hydrogel is applied on the surfaces within the multi-depth depression. It is to be understood that the polymeric hydrogel may be conformally coated along all of the surfaces (including the bottom surfaces and the sidewalls) within the multi-depth depression, as long as the interstitial regions remain free of the polymeric hydrogel. Any example of the polymeric hydrogel described herein may be used in this example.

[0390] A first immobilised primer set (including a cleavable first immobilised primer and an uncleavable second immobilised primer) is immobilised on a surface in the shallow portion; first transposome complexes are immobilised on the surface in the shallow portion, the first transposome complexes including the first amplification domain as described herein; a second primer set (an uncleavable first immobilised primer and a cleavable second immobilised primer) immobilised on a surface in the deep portion; and second transposome complexes are immobilised on the surface in the deep portion, the second transposome complexes including a second amplification domain as described herein. In these examples, the first amplification domain has the same sequence, e.g., as the uncleavable first immobilised primer and the cleavable first immobilised primer; and the second amplification domain has the same sequence, e.g., as the cleavable second immobilised primer and the uncleavable second immobilised primer.

[0391] The immobilised primers and the transposome complexes include 5’ functional groups that enable them to be attached to the polymeric hydrogel. Any of the examples set forth herein may be used, such as a terminal alkyne (e.g., hexynyl) or an internal alkyne, where the alkyne is part of a cyclic compound (e.g., bicyclo[6.1.0]nonyne (BCN)).

[0392] Sequencing

[0393] As described herein, the amplified polynucleotides provide information (e.g. identification of the genetic sequence, identification of epigenetic modifications) on the original target polynucleotide sequence. For example, a sequencing process (e.g. a sequencing-by-synthesis or sequencing-by-ligation process) may reproduce information that was present in the original target polynucleotide sequence, by using complementary base pairing.

[0394] In one embodiment, sequencing may be carried out using any suitable "sequencing-by-synthesis" technique, wherein nucleotides are added successively in cycles to the free 3' hydroxyl group, resulting in synthesis of a polynucleotide chain in the 5' to 3' direction. The nature of the nucleotide added is typically determined after each addition. One particular sequencing method relies on the use of modified nucleotides that can act as reversible chain terminators. Such reversible chain terminators comprise removable 3' blocking groups. Once such a modified nucleotide has been incorporated into the growing polynucleotide chain complementary to the region of the template being sequenced there is no free 3'-OH group available to direct further sequence extension and therefore the polymerase cannot add further nucleotides. Once the nature of the base incorporated into the growing chain has been determined, the 3' block may be removed to allow addition of the next successive nucleotide. By ordering the products derived using these modified nucleotides it is possible to deduce the sequence of the polynucleotide template. Such reactions can be done in a single experiment if each of the modified nucleotides has attached thereto a different label, known to correspond to the particular base, to facilitate discrimination between the bases added at each incorporation step. Suitable labels are described in PCT application PCT / GB2007 / 001770, the contents of which are incorporated herein by reference in their entirety. Alternatively, a separate reaction may be carried out containing each of the modified nucleotides added individually.

[0395] The modified nucleotides may carry a label to facilitate their detection. Such a label may be configured to emit a signal, for example an electromagnetic signal (e.g. a (visible) light signal).

[0396] In a particular embodiment, the label is a fluorescent label (e.g. a dye). Thus, such a label may be configured to emit an electromagnetic signal (e.g. a (visible) light signal). One method for detecting the fluorescently labelled nucleotides comprises using laser light of a wavelength specific for the labelled nucleotides, or the use of other suitable sources of illumination. The fluorescence from the label on an incorporated nucleotide may be detected by a CCD camera or other suitable detection means. Suitable detection means are described in PCT / US2007 / 007991, the contents of which are incorporated herein by reference in their entirety.

[0397] However, the detectable label need not be a fluorescent label. Any label can be used which allows the detection of the incorporation of the nucleotide into the polynucleotide sequence.

[0398] Each cycle may involve simultaneous delivery of four different nucleotide types to the array of template molecules. Alternatively, different nucleotide types can be added sequentially and an image of the array of template molecules can be obtained between each addition step.

[0399] In some embodiments, each nucleotide type may have a (spectrally) distinct label. In other words, four channels may be used to detect four nucleobases (also known as 4-channel chemistry). For example, a first nucleotide type (e.g. A) may include a first label (e.g. configured to emit a first wavelength, such as red light), a second nucleotide type (e.g. G) may include a second label (e.g. configured to emit a second wavelength, such as blue light), a third nucleotide type (e.g. T) may include a third label (e.g. configured to emit a third wavelength, such as green light), and a fourth nucleotide type (e.g. C) may include a fourth label (e.g. configured to emit a fourth wavelength, such as yellow light). Four images can then be obtained, each using a detection channel that is selective for one of the four different labels. For example, the first nucleotide type (e.g. A) may be detected in a first channel (e.g. configured to detect the first wavelength, such as red light), the second nucleotide type (e.g. G) may be detected in a second channel (e.g. configured to detect the second wavelength, such as blue light), the third nucleotide type (e.g. T) may be detected in a third channel (e.g. configured to detect the third wavelength, such as green light), and the fourth nucleotide type (e.g. C) may be detected in a fourth channel (e.g. configured to detect the fourth wavelength, such as yellow light). Although specific pairings of bases to signal types (e.g. wavelengths) are described above, different signal types (e.g. wavelengths) and / or permutations may also be used.

[0400] In some embodiments, detection of each nucleotide type may be conducted using fewer than four different labels. For example, sequencing-by-synthesis may be performed using methods and systems described in US 2013 / 0079232, which is incorporated herein by reference.

[0401] Thus, in some embodiments, two channels may be used to detect four nucleobases (also known as 2-channel chemistry). For example, a first nucleotide type (e.g. A) may include a first label (e.g. configured to emit a first wavelength, such as green light) and a second label (e.g. configured to emit a second wavelength, such as red light), a second nucleotide type (e.g. G) may not include the first label and may not include the second label, a third nucleotide type (e.g. T) may include the first label (e.g. configured to emit the first wavelength, such as green light) and may not include the second label, and a fourth nucleotide type (e.g. C) may not include the first label and may include the second label (e.g. configured to emit the second wavelength, such as red light). Two images can then be obtained, using detection channels for the first label and the second label. For example, the first nucleotide type (e.g. A) may be detected in both a first channel (e.g. configured to detect the first wavelength, such as red light) and a second channel (e.g. configured to detect the second wavelength, such as green light), the second nucleotide type (e.g. G) may not be detected in the first channel and may not be detected in the second channel, the third nucleotide type (e.g. T) may be detected in the first channel (e.g. configured to detect the first wavelength, such as red light) and may not be detected in the second channel, and the fourth nucleotide type (e.g. C) may not be detected in the first channel and may be detected in the second channel (e.g. configured to detect the second wavelength, such as green light). Although specific pairings of bases to signal types (e.g. wavelengths) and / or combinations of channels are described above, different signal types (e.g. wavelengths) and / or permutations may also be used.

[0402] In some embodiments, one channel may be used to detect four nucleobases (also known as 1-channel chemistry). For example, a first nucleotide type (e.g. A) may include a cleavable label (e.g. configured to emit a wavelength, such as green light), a second nucleotide type (e.g. G) may not include a label, a third nucleotide type (e.g. T) may include a non-cleavable label (e.g. configured to emit the wavelength, such as green light), and a fourth nucleotide type (e.g. C) may include a label-accepting site which does not include the label. A first image can then be obtained, and a subsequent treatment carried out to cleave the label attached to the first nucleotide type, and to attach the label to the label-accepting site on the fourth nucleotide type. A second image may then be obtained. For example, the first nucleotide type (e.g. A) may be detected in a channel (e.g. configured to detect the wavelength, such as green light) in the first image and not detected in the channel in the second image, the second nucleotide type (e.g. G) may not be detected in the channel in the first image and may not be detected in the channel in the second image, the third nucleotide type (e.g. T) may be detected in the channel (e.g. configured to detect the wavelength, such as green light) in the first image and may be detected in the channel (e.g. configured to detect the wavelength, such as green light) in the second image, and the fourth nucleotide type (e.g. C) may not be detected in the channel in the first image and may be detected in the channel in the second image (e.g. configured to detect the wavelength, such as green light). Although specific pairings of bases to signal types (e.g. wavelengths) and / or combinations of images are described above, different signal types (e.g. wavelengths), images and / or permutations may also be used.

[0403] Other embodiments may involve strand displacement sequencing-by-synthesis (strand displacement SBS). In such a case, a strand displacement polymerase may initiate SBS from a nick.

[0404] In one embodiment, the amplified forward strands and the amplified reverse complement strands are prepared for concurrent sequencing, or the amplified reverse strands and amplified forward complement strands are prepared for concurrent sequencing. When strands are prepared for concurrent sequencing, they are placed into a state where a sequencing process will lead to sequencing of the amplified forward strands and the amplified reverse complement strands at the same time, or sequencing of the amplified reverse strands and the amplified forward complement strands at the same time. This may include, for example, contacting sequencing primers to both the amplified forward strands and the amplified reverse complement strands, or contacting sequencing primers to both the amplified reverse strands and amplified forward complement strands. Alternatively, this may include, for example, nicking both the amplified forward complement strands and amplified reverse strands (hybridised to the amplified forward strands and reverse complement strands respectively), or nicking both the amplified reverse complement strands and amplified forward strands (hybridised to the amplified reverse strands and forward complement strands respectively). However, these non-limiting methods are examples of ways of preparing strands for (concurrent) sequencing, and the skilled person would appreciate that other methods of preparing strands for sequencing can be utilised.

[0405] In some embodiments, it is desirable to ensure that the population of amplified forward strands prepared for concurrent sequencing is substantially equal to the population of amplified reverse complement strands prepared for concurrent sequencing; and / or wherein the population of amplified reverse strands prepared for concurrent sequencing is substantially equal to the population of amplified forward complement strands prepared for concurrent sequencing.

[0406] Usually, this can be conducted by amplifying the forward strands substantially equally relative to amplifying the reverse complement strands, such that the total population of amplified forward strands is substantially equal to the total population of amplified reverse complement strands (or amplifying the reverse strands substantially equally relative to amplifying the forward complement strands, such that the total population of amplified reverse strands is substantially equal to the total population of amplified forward complement strands).

[0407] However, where the total population of amplified forward strands is not substantially equal to the total population of amplified reverse complement strands, the population of amplified forward strands prepared for concurrent sequencing can still be made to be substantially equal to the population of amplified reverse complement strands prepared for concurrent sequencing (or where the total population of amplified reverse strands is not substantially equal to the total population of amplified forward complement strands, the population of amplified reverse strands prepared for concurrent sequencing can still be made to be substantially equal to the population of amplified forward complement strands prepared for concurrent sequencing), for example by the use of blocked and unblocked sequencing primers as described in WO 2023 / 175037.

[0408] In one embodiment, the sequencing process comprises a first sequencing read and a second sequencing read. The first sequencing read and the second sequencing read may be conducted concurrently. In other words, the first sequencing read and the second sequencing read may be conducted at the same time. This leads to sequencing of the amplified forward strands and the amplified reverse complement strands simultaneously, or sequencing of the amplified reverse strands and the amplified forward complement strands simultaneously. Alternative methods of sequencing include sequencing by ligation, for example as described in US 6,306,597 or WO 06 / 084132, the contents of which are incorporated herein by reference.

[0409] Data analysis using 9 QAM

[0410] For two inserts of polynucleotide sequences (e.g. a first insert including a forward strand and a second insert including a reverse complement strand, or a first insert including a reverse strand and a second insert including a forward complement strand, as described herein), there are sixteen possible combinations of nucleobases at any given position (e.g., an A in the first insert and an A in the second insert, an A in the first insert and a T in the second insert, and so on). When the same nucleobase is present at a given position in both inserts, the light emissions associated with each target sequence during the relevant base calling cycle will be characteristic of the same nucleobase. In effect, the two inserts behave as a single insert, and the identity of the bases at that position are uniquely callable.

[0411] However, when a nucleobase of the first insert is different from a nucleobase at a corresponding position of the second insert, the signals associated with each insert in the relevant base calling cycle will be characteristic of different nucleobases. In one embodiment, a first signal coming from the first insert have substantially the same intensity as a second signal coming from the second insert. The two signals may also be co-localised, and may not be spatially and / or optically resolved. Therefore, when different nucleobases are present at corresponding positions of the two inserts, the identity of the nucleobases cannot be uniquely called from the combined signal alone. However, useful sequencing information can still be determined from these signals.

[0412] Nine distributions (or bins) of intensity values arise from the combination of two colocalised signals of substantially equal intensity.

[0413] The intensity values may be up to a scale or normalisation factor; the units of the intensity values may be arbitrary or relative (i.e., representing the ratio of the actual intensity to a reference intensity). The sum of the first signal generated from the first insert and the second signal generated from the second insert results in a combined signal. The combined signal may be captured by a first optical channel and a second optical channel. The computer system can map the combined signal generated into one of the nine bins, and thus determine sequence information relating to the added nucleobase at the first insert and the added nucleobase at the second insert.

[0414] Bins are selected based upon the combined intensity of the signals originating from each target sequence during the base calling cycle. For example, a first bin may be selected following the detection of a high-intensity (or “on / on”) signal in the first channel and a high-intensity signal in the second channel. A second bin may be selected following the detection of a high-intensity signal in the first channel and an intermediate-intensity (“on / off” or “off / on”) signal in the second channel. A third bin may be selected following the detection of a high-intensity signal in the first channel and a low-intensity or zero-intensity (“off / off”) signal in the second channel. A fourth bin may be selected following the detection of an intermediate-intensity signal in the first channel and a high-intensity signal in the second channel. A fifth bin may be selected following the detection of an intermediate-intensity signal in the first channel and an intermediate-intensity signal in the second channel. A sixth bin may be selected following the detection of an intermediate-intensity signal in the first channel and a low-intensity or zero-intensity signal in the second channel. A seventh bin may be selected following the detection of a low-intensity signal in the first channel and a high-intensity signal in the second channel. An eighth bin may be selected following the detection of a low-intensity or zero-intensity signal in the first channel and an intermediate-intensity signal in the second channel. A ninth bin may be selected following the detection of a low-intensity or zero-intensity signal in the first channel and a low-intensity signal in the second channel.

[0415] Four of the nine bins represent matches between respective nucleobases of the two inserts sensed during the cycle (the first, third, seventh and ninth bins). In response to mapping the combined signal to a bin representing a match, the computer processor may detect a match between the first insert and the second insert at the sensed position. In response to mapping the combined signal to a bin representing a match, the computer processor may base call the respective nucleobases. For example, when the combined signal is mapped to the first bin for a base calling cycle, the computer processor base calls both the added nucleobase at the first insert and the added nucleobase at the second insert as T. When the combined signal is mapped to the third bin for the base calling cycle, the processor base calls both the added nucleobase at the first insert and the added nucleobase at the second insert as A. When the combined signal is mapped to the seventh bin for the base calling cycle, the processor base calls both the added nucleobase at the first insert and the added nucleobase at the second insert as G. When the combined signal is mapped to the ninth bin for the base calling cycle, the processor base calls both the added nucleobase at the first insert and the added nucleobase at the second insert as C.

[0416] The remaining five bins are “ambiguous”. That is to say that these bins each represent more than one possible combination of first and second nucleobases. The second, fourth, sixth and eighth bins each represent two possible combinations of first and second nucleobases. The fifth bin, meanwhile, represents four possible combinations. Nevertheless, mapping the combined signal to an ambiguous bin may still allow for sequencing information to be determined. For example, the second, fourth, fifth, sixth and eighth bins represent mismatches between respective nucleobases of the two inserts sensed during the cycle. Therefore, in response to mapping the combined signal to a bin representing a mismatch, the computer processor may detect a mismatch between the first insert and the second insert at the sensed position.

[0417] In this particular example, A is configured to emit a signal in both the first channel and the second channel, C is configured to emit a signal in the first channel only, T is configured to emit a signal in the second channel only, and G does not emit a signal in either channel. However, different permutations of nucleobases can be used to achieve the same effect by performing dye swaps. For example, A may be configured to emit a signal in both the first channel and the second channel, T may be configured to emit a signal in the first channel only, C may be configured to emit a signal in the second channel only, and G may be configured to not emit a signal in either channel.

[0418] The number of classifications which may be selected based upon the combined signal intensities may be predetermined, for example based on the number of inserts expected to be present in the nucleic acid cluster. Whilst a set of nine classifications as described above is possible, the number of classifications may be greater or smaller.

[0419] In addition to identifying matches and mismatches, the mapping of the combined signal to each of the different bins (e.g. in combination with additional knowledge, such as the library preparation methods used) can provide additional information about the first insert and the second insert, or about sequences from which the first insert and the second insert were derived. For example, given the nucleic acid material input and the processing methods used to generate the nucleic acid clusters, the first insert and the second insert may be expected to be identical at a given position. In this case, the mapping of the combined signal to a bin representing a mismatch may be indicative of an error introduced during library preparation. In addition, the first insert and the second insert may be expected to be different, for example due to deliberate sequence modifications introduced during library preparation to detect modified cytosines.

[0420] As mentioned herein, the tagmentation process may involve treatment with a conversion agent. In cases where the conversion reagent is configured to convert an unmodified cytosine to uracil or a nucleobase which is read as thymine / uracil, an A-T or T-A base pair in the original molecule will result in a match (A / A or T / T) at the corresponding position of the forward and reverse complement strands. An mC-G or G-mC base pair will also result in a match (G / G or C / C) at the corresponding position of the forward and reverse complement strands. For a C-G base pair, however, the conversion of unmodified cytosine to uracil (or a nucleobase which is read as thymine / uracil) in the forward strand will result in a T at the corresponding position of the forward strand. Meanwhile, the corresponding position on the reverse complement strand will be occupied by C. Alternatively, for a G-C base pair, the conversion of unmodified cytosine to uracil (or a nucleobase which is read as thymine / uracil) in the reverse strand will result in an A at the corresponding position of the reverse complement strand. Meanwhile, the corresponding position of the forward strand will be occupied by G. Therefore, in response to mapping the combined signal to the distribution representing G / G or C / C, the presence of a modified cytosine can be determined at the corresponding position in the original polynucleotide.

[0421] In other cases where the conversion reagent is configured to convert a modified cytosine to thymine or a nucleobase which is read as thymine / uracil, an A-T or T-A base pair will result in a match (A / A or T / T) at the corresponding position of the forward and reverse complement strands. A C-G or G-C base pair will also result in a match (G / G or C / C) at the corresponding position of the forward and reverse complement strands. For a mC-G base pair, however, the conversion of 5-methylcytosine to thymine in the forward strand will result in a T at the corresponding position of the forward strand. Meanwhile, the corresponding position on the reverse complement strand will be occupied by C. Alternatively, the conversion of 5-methylcytosine to thymine in the reverse strand will result in an A at the corresponding position of the reverse complement strand. Meanwhile, the corresponding position of the forward strand will be occupied by G. Therefore, in response to mapping the combined signal to the distribution representing an A / G, G / A, T / C, or C / T mismatch, the presence of a modified cytosine can be determined at the corresponding position in the original polynucleotide.

[0422] For each base pair in the original double-stranded polynucleotide sample, it may be assumed that there are six possibilities: A-T, T-A, C-G, G-C, mC-G and G-mC. Each of these possibilities is uniquely represented by one of the plurality of classifications. According to the present methods, it is therefore possible to determine both the sequence and “methylation” status (i.e. presence of modified cytosines) of a double-stranded polynucleotide in a single sequencing run.

[0423] In addition to determining “methylation” status, it may also be possible to identify library preparation / sequencing errors.

[0424] The dye-encoding scheme may be optimised to allow for different combinations of first and second nucleobases to be resolved. This may be particularly useful where sequence modifications of a known type have been introduced into the first inserts and the second inserts. For example, where sequence modifications have been introduced that result in the conversion of unmodified cytosines to uracil or nucleobases which is read as thymine / uracil, or the conversion of modified cytosines to thymine or nucleobases which are read as thymine / uracil, the dye-encoding scheme may be selected such that the resulting combination of first and second nucleobases do not fall within the central bin (which represents four different nucleobase combinations). In the case of conversion of modified cytosines to thymine (or nucleobases which are read as thymine / uracil), a T / C or G / A mismatch between the forward and reverse complement strands is indicative of the presence of a mC-G or G-mC base pair at the corresponding position of the sample. The dye-encoding scheme may therefore be designed such that these mismatches may be resolved from other possible combinations of nucleobases. This may be achieved by detecting light emissions from A and T bases in a first illumination cycle, and from C and T bases in a second illumination cycle. In another example, light emissions may be detected from C and G bases in a first illumination cycle, and from C and T bases in a second illumination cycle. In another example, light emissions may be detected from C and A bases in a first illumination cycle, and from C and G bases in a second illumination cycle.

[0425] In the case of unmodified cytosines to uracil (or nucleobases which is read as thymine / uracil), a C / C or G / G match between the forward and reverse complement strands is indicative of the presence of a mC-G or G-mC base pair at the corresponding position of the sample. In this case, a mC-G or G-mC base pair will always be resolvable. However, the dyeencoding scheme can still be designed to optimise the resolution between unmodified bases.

[0426] The described method below allows for the determination of sequence information from two (or more) inserts (e.g. the first insert and the second insert) in a single sequencing run from a single combined signal obtained from the first insert and the second insert.

[0427] Intensity data is first obtained. The intensity data includes first intensity data and second intensity data. The first intensity data comprises a combined intensity of a first signal component obtained based upon a respective first nucleobase of the first insert and a second signal component obtained based upon a respective second nucleobase of the second insert. Similarly, the second intensity data comprises a combined intensity of a third signal component obtained based upon the respective first nucleobase of the first insert and a fourth signal component obtained based upon the respective second nucleobase of the second insert.

[0428] As such, the first insert is capable of generating a first signal comprising a first signal component and a third signal component. The second insert is capable of generating a second signal comprising a second signal component and a fourth signal component.

[0429] As described above, the first insert and the second insert may be arranged on the solid support such that signals from the first insert and the second insert are detected by a single sensing portion and / or may comprise a single cluster such that first signals and second signals from each of the respective first inserts and second inserts cannot be spatially resolved.

[0430] In one example, obtaining the intensity data comprises selecting intensity data, for example based upon a chastity score. A chastity score may be calculated as the ratio of the brightest base intensity divided by the sum of the brightest and second brightest base intensities. In one example, high-quality data corresponding to two inserts with a substantially equal intensity ratio may have a chastity score of around 0.8 to 0.9, for example 0.89-0.9.

[0431] After the intensity data has been obtained, one of a plurality of classifications is selected based on the intensity data. Each classification represents one or more possible combinations of respective first and second nucleobases, and at least one classification of the plurality of classifications represents more than one possible combination of respective first and second nucleobases. In one example, the plurality of classifications comprises nine classifications. Selecting the classification based on the first and second intensity data comprises selecting the classification based on the combined intensity of the first and second signal components and the combined intensity of the third and fourth signal components.

[0432] The method may then proceed to a step of determining sequence information of the respective first and second nucleobases based on the classification selected. The signals generated during a cycle of a sequencing are indicative of the identity of the nucleobase(s) added during sequencing (e.g. using sequencing-by-synthesis). For example, it may be determined that there is a match or a mismatch between the respective first and second nucleobases. Where it is determined that there is a match between the first and second respective nucleobases, the nucleobases may be base called. Whether there is a match or a mismatch, additional or alternative information may be obtained, as described above. It will be appreciated that there is a direct correspondence between the identity of the nucleobases that are incorporated and the identity of the complementary base at the corresponding position of the template sequence bound to the solid support. Therefore, any references herein to the base calling of respective nucleobases at the two inserts encompasses the base calling of nucleobases hybridised to the template sequences and, alternatively or additionally, the identification of the corresponding nucleobases of the template sequences.

[0433] Strand identification

[0434] As described above, a transposome complex may be attached symmetrically or asymmetrically to the surface of a substrate, such as a flow cell. Figure 1a, depicts a “symmetric” attachment, which is where all transposome complexes are attached via their transferred strand to the substrate. Fig.1b depicts an “asymmetric” attachment where one transposome, is attached by the transferred strand, and the other transposome, is attached by the (shorter) non-transferred strand to the substrate.

[0435] Methods of fragmenting and tagging a target nucleic acid and introducing a strand identifier An example of a method of fragmenting and tagging (e.g. adding on one or more adaptor sequences) is shown in Figures 2 to 7. This method may be referred to herein as “tagmentation”.

[0436] Figure 2 illustrates the process whereby (un-fragmented) target polynucleotide, such as double-stranded (ds)DNA, is added to a substrate, such as a flow cell lane containing a symmetric composition of transposomes (Fig. 2a). For the purpose of illustration only, two transposome complexes (a first and second transposome complex) are depicted. As shown in Fig.2b, the nucleic acid sample is fragmented and the 5’ ends of both strands of the dsDNA are ligated to 3’ ends of the transferred strands (longer strand in this example) of the transposome complex. As described above, fragmentation and ligation may take place at a temperature of 30°C or above. Although not shown in this example, it is possible for the process of tagmentation to occur in solution. That is, where the transposome complexes are not immobilised to a solid support. Following the step of tagmentation, the partially or fully adapted target polynucleotides may be immobilised onto a substrate, such as a flow cell for the subsequent sequencing, as outlined above.

[0437] The target polynucleotide may be referred to herein as a “template”, and the strands thereof may be referred to as “original strands” or “template strands”. Once fragmented, the target polynucleotide may be referred to herein as a “nucleic acid fragment” or “DNA fragment”.

[0438] As the 3’ ends of the dsDNA are not ligated to the non-transferred strand, there exists a gap (not shown) between the 3’ end of the nucleic acid and the 5’ end of the non-transferred strand (shorter strand in this example). As described above, this gap may be one or more bases long, for example 9 base pairs or around 9 base pairs long.

[0439] In the next step - shown in Figure 2c - the transposase enzymes (e.g. Tn5) are removed. Any method that accomplishes the removal of transposase enzymes is included within the scope of the present disclosure. Some, non-limiting, examples are described above, and include using sodium dodecyl sulfate (SDS), a proteinase or other chaotrophic detergent, or heating the substrate such as the flow cell to 60°C. Figure 2c depicts the assembly with the transposase enzymes removed. Optionally a washing step is conducted after tagmentation, wherein the washing step comprises the introduction of a washing solution to the substrate, heating the flow cell to about 60°C and then introducing an extension amplification mix to the substrate. Prior to introducing the extension amplification mix, the temperature may be reduced to at or about 38 °C. As described above, it has been found that heating the substrate (e.g. flow cell) after tagmentation may be sufficient to remove the transposase enzymes without the need for additional buffers or reagents. As described above, an example extension amplification mix includes a recombinase, a polymerase, and accessory proteins. An example washing solution may comprise an aqueous solution including a buffer agent (e.g. Tris), a salt (e.g. sodium chloride, sodium citrate, etc.), a surfactant (e.g., TWEEN polysorbates), and / or a chelating agent (e.g., EDTA). In one example, the washing solution includes water, the salt at a concentration ranging from about 25 mM to about 50 mM, the surfactant in an amount ranging from about 0.01 wt% to about 0.1 wt%, and optionally the chelating agent. The washing solution may have a relatively high pH, e.g., ranging from about 7 to about 10.

[0440] If a chaotropic detergent is used to remove the transposase enzymes, in a further embodiment, the method may comprise the addition of a chelator of the chaotropic detergent, as described above, to prevent any deleterious effect of the chaotropic detergent on subsequent enzymatic reactions.

[0441] Following this purification step to remove the transposase enzymes, the nontransferred strands are dehybridised, as shown in Figure 3a. As such, Fig. 3a illustrates solely the progenitor sequencable bridged template (attached to a first and second tranposome complex, although only a monomer from each dimer is shown for illustration). Here, as a result of tagmentation, the partially adapted target double-stranded nucleic acid sequence (also referred to herein as a nucleic acid fragment or template, such terms are interchangeable) will comprise an adaptor that comprises a first amplification domain on one strand (forward or reverse) and second amplification domain (reverse or forward respectively) - for example, P5 and P7, as described above. In the subsequent examples where we refer to P5 and P7, these are as a non-limiting example of a suitable amplification domains. The non-transferred strands are not covalently connected to the 3’ ends of the template but instead are hybridized to the transferred strand.

[0442] Fig. 3b illustrates the sequencable bridged template after a polymerase extension reaction to extend the 3’ end of the template and append the complementary sequences of the adaptors. This product is now capable of amplification into clusters. In this step, additional sequences (adapters) are added to the 3’ ends of the partially adapted fragments by an extension reaction using amplification reagents (e.g., the ExAMP reagents available from Illumina Inc.). It is to be understood that in examples where SDS or another chaotropic detergent has been chelated with a cyclodextrin, the extension reaction may take place without having to perform one or more wash cycles to remove the SDS or other chaotropic detergent from the flow cell. The extension reaction involves the addition of nucleotides in a template dependent fashion from the 3’ ends of the template nucleic acid fragment using the respective transferred strands as the template. As such, the nucleic acid fragment is extended such that it generates complementary sections. Figure 4 illustrates the process whereby (un-fragmented) nucleic acid, such as doublestranded (ds)DNA, is added to a substrate, such as a flow cell lane containing an asymmetric composition of transposomes (Fig. 4a).

[0443] As shown in Fig.4b, the nucleic acid sample is fragmented and the 5’ ends of both strands of the dsDNA are ligated to 3’ ends of the transferred strands of the transposome complex. As explained above, by an “asymmetric composition of transposomes”, is meant that the first (of which there may be more than one) transposome complex is attached to the flow cell surface through its 5’ end (e.g. through the transferred strand) and the second transposome complex (of which there may be more than one) is attached to the flow cell through its 3’ end (e.g. through the non-transferred strand).

[0444] As described above, fragmentation and ligation may take place at a temperature of 30°C or above. Again, as the 3’ ends of the dsDNA are not ligated to the non-transferred strand, there exists a gap (not shown) between the 3’ end of the nucleic acid and the 5’ end of the non-transferred strand. As described above, this gap may be one or more bases long, for example 9 base pairs or around 9 base pairs long.

[0445] In the next step - shown in Figure 4c - the transposase enzymes (e.g. Tn5) are removed. As above, any method that accomplishes the removal of transposase enzymes is included within the scope of the present disclosure. Some, non-limiting, examples are described above, and include using sodium dodecyl sulfate (SDS), a proteinase or other chaotropic detergent, or heating the substrate such as the flow cell to 60°C. Figure 4c depicts the assembly with the transposase enzymes removed. Optionally a washing step is conducted after tagmentation, wherein the washing step comprises the introduction of a washing solution to the substrate, heating the flow cell to about 60°C and then introducing an extension amplification mix to the substrate. Prior to introducing the extension amplification mix, the temperature may be reduced to about 38 °C. As described above, it has been found that heating the substrate (e.g. flow cell) after tagmentation may be sufficient to remove the transposase enzymes without the need for additional buffers or reagents. As described above, an example extension amplification mix includes a recombinase, a polymerase, and accessory proteins. An example washing solution may comprise an aqueous solution including a buffer agent (e.g. Tris), a salt (e.g. sodium chloride, sodium citrate, etc.), a surfactant (e.g., TWEEN polysorbates), and / or a chelating agent (e.g., EDTA). In one example, the washing solution includes water, the salt at a concentration ranging from about 25 mM to about 50 mM, the surfactant in an amount ranging from about 0.01 wt% to about 0.1 wt%, and optionally the chelating agent. The washing solution may have a relatively high pH, e.g., ranging from about 7 to about 10. Following removal of the transposase enzymes, the non- transferred strands are dehybridised, as shown in Figure 5a. As such, Fig. 5a illustrates solely the progenitor sequencable bridged template. Here, as a result of tagmentation, the partially adapted doublestranded nucleic acid sequence (i.e. the target polynucleotide) will comprise a first adaptor that comprises at least one amplification domain on one strand (forward or reverse) and second adaptor comprising a second amplification domain on the other strand (reverse or forward respectively) - for example, P5 and P7, as described above. In the subsequent examples where we refer to P5 and P7, these are as a non-limiting example of a suitable amplification domains. In this example, the non-transferred strands are not covalently connected to the 3’ ends of the template but instead are hybridized to the transferred strand.

[0446] Fig. 5b illustrates the ‘sequencable-bridged-template’ after a polymerase extension reaction to extend the 3’ end of the template and append the complementary sequences of the adaptors. In this step, additional sequences (referred to herein as adapters 3 and 4) are added to the 3’ ends of the partially adapted fragments by an extension reaction using amplification reagents (e.g., the ExAMP reagents available from Illumina Inc.). The resulting template strands will now contain (first or second) adaptor at their 5 and a (second or first) adaptor at their 3’ end. The adaptors may also be referred to as “adaptor structures”. It is to be understood that in examples where SDS or another chaotropic detergent has been chelated with a cyclodextrin, the extension reaction may take place without having to perform one or more wash cycles to remove the SDS or other chaotropic detergent from the flow cell. The extension reaction involves the addition of nucleotides in a template dependent fashion from the 3’ ends of the target polynucleotide using the respective transferred strands as the template. As such, the nucleic acid fragment is extended such that it generates complementary sections. This product is now capable of amplification into clusters.

[0447] Accordingly, in one aspect of the disclosure, there is provided a method of introducing a strand identifier to distinguish the descendants of a forward strand from the descendants of a reverse strand of a double-stranded target polynucleotide, the method comprising:

[0448] providing one or more first transposome complexes, wherein the first transposome complex comprises a first adaptor, one or more second transposome complexes, wherein the second transposome complex comprises a second adaptor, and a double-stranded target polynucleotide comprising a forward target strand and a reverse target strand;

[0449] ligating the first adaptor to a 5’ end of the reverse target strand using the first transposome complex and ligating the second adaptor to a 5’ end of the forward target strand using the second transposome complex, to form a partially adapted double-stranded target polynucleotide; removing a transposase enzyme from each of the first and second transposome complexes; and

[0450] conducting an extension reaction of the partially adapted double-stranded template polynucleotide to form a fully adapted double-stranded target polynucleotide, wherein the extension reaction leads to the incorporation of a third adaptor at the 3’ end of the forward target strand and a fourth adaptor at the 3’ end of the reverse target strand, wherein the third adaptor comprises a forward strand descendant identifier and the fourth adaptor comprises a reverse strand descendant identifier.

[0451] In one embodiment, the method comprises removing the non-transferred strand (for example using hear denaturation, as described further below) from the first transposase complex and removing the non-transferred strand from the second transposase complex before conducting an extension reaction.

[0452] In some embodiment, the method further comprises providing a solid support comprising first immobilised primers and second immobilised primers. In this embodiment, the one or more first transposome complexes are immobilised at a 5’ end to the solid support, and the one or more second transposome complexes immobilised at a 3’ end or a 5’ end to the solid support. More specifically, in one embodiment, the (one or more) first transposome complex is immobilised to the solid support through the 5’ end of the transferred strand and the (one or more) second transposome complex is immobilised to the solid support through the 5’ end of the transferred strand - i.e. the transposomes are symmetrically attached. Alternatively, the (one or more) first transposome complex is immobilised to the solid support through the 5’ end of the transferred strand and the (one or more) second transposome complex is optionally immobilised to the solid support through 3’ end of the nontransferred strand - i.e. the transposome complexes are asymmetrically attached.

[0453] In an alternative embodiment, the (one or more) first transposome complexes and the (one or more) second transposome complexes are in solution.

[0454] In one embodiment, the first transposome complex comprises a transferred strand and a non-transferred strand, wherein the transferred strand comprises the first adaptor, and the second transposome complex comprises a transferred strand and a non-transferred strand, wherein the transferred strand comprises the second adaptor.

[0455] The result of the extension reaction is a polynucleotide that is a fully adapted doublestranded template polynucleotide. By “fully adapted” is meant that the forward strand of the target polynucleotide comprises an adaptor at a 5’ end (which is referred to herein as the second adaptor) and an adaptor at a 3’ end (which is referred to herein as the third adaptor). Similarly, the reverse strand of the target polynucleotide comprises an adaptor at a 5’ (which is referred to herein as the first adaptor) and an adaptor at a 3’ end (which is referred to herein as the fourth adaptor). The next step is cluster generation.

[0456] In one embodiment, the first adaptor comprises (in the 5’ to 3’ direction) at least a first amplification domain, a reverse template strand identifier and a transposase recognition sequence. In another embodiment, the second adaptor comprises (in the 5’ to 3’ direction) at least a second amplification domain, a forward template strand identifier and a transposase recognition sequence.

[0457] In one embodiment, the third adaptor comprises a transposase recognition sequence, a forward strand descendant identifier and at least one amplification domain. In another embodiment, the fourth adaptor comprises a transposase recognition sequence, a reverse strand descendant identifier and at least one amplification domain.

[0458] As used herein a “transposase recognition sequence” is any sequence of nucleic acid that can be recognised by a transposase enzyme. Typically a transposase recognition sequence is or will comprise a terminal inverted repeat (TIR).

[0459] As used herein a “(forward / reverse) template strand identifier” is any sequence that can identify or otherwise distinguish (directly if sequenced, or indirectly if used to differentially bind a target-specific sequencing primer) whether a strand is the target strand or a direct descendant therefrom (rather than the reverse complement thereof). Conversely, as used herein a “(forward / reverse) descendant strand identifier” is any sequence that can identify or otherwise distinguish (directly if sequenced, or indirectly if used to differentially bind a targetspecific sequencing primer), a reverse-complement descendant strand (from a direct descendant of a target strand). In other words, the forward template strand identifier can be used to distinguish the forward strand of the target polynucleotide and all descendants thereof, whereas the forward strand descendant identifier can be used to distinguish all forward strands derived from the re-synthesis of the reverse strand (i.e. the reverse complement strand).

[0460] As explained in more detail below, the forward strand descendant identifier is non-complementary to the reverse template strand identifier. Similarly, the reverse strand descendant identifier is non-complementary to the forward template strand identifier. In particular, the forward strand descendant identifier can differ from the complement of the template forward strand identifier by one or more bases. Similarly, the reverse strand descendant identifier differs from the complement of the reverse template strand identifier by one or more bases. The skilled person will understand that there is no upper limit to the number of differential bases.

[0461] Alternatively, the reverse strand descendant identifier and / or the forward strand descendant identifier can comprise at least one modified base. Alternatively, the reverse template strand identifier or forward template strand identifier may comprise at least one modified base. This base may be a base that is often “misincorporated” during the extension reaction. For example, A is often “misincorporated” with 8-oxoguanine, and as such, the use of 8-oxoguanine as or within a template strand identifier will lead to a subsequent sequence difference between the template strand descendants and the reverse complement descendants. Alternatively, the modified base can be any base that is converted to a different base following the extension reaction. For example, a methylated cytosine that, following the extension reaction, can be converted, using a suitable conversion agent as explained above, to thymine or a nucleobase that is read as thymine or uracil.

[0462] During cluster generation, both the top strand (black dashed) and the bottom strand (grey dashed) of templates generated with the symmetric composition (Fig. 6a) or asymmetric composition (Fig. 6b) of transposomes are capable of independently generating new doublestranded bridges that are descendants of either the original top strand or bottom strand. Fig.

[0463] 6c illustrates an outcome of clustering where the bridged descendants of the top strand (black) and bottom strand (grey) are in equal proportions. Fig. 6d illustrates an outcome of clustering where the bridged descendants of the top strand (black) are in a majority. Fig. 6e illustrates an outcome of clustering where the bridged descendants of the bottom strand (grey) are in a majority. Using present methods in the art it is currently not possible to discern the top and bottom strand descendants (that is whether a sequence arose from amplification of the forward or reverse strand). However, the ability to identify the top and bottom strand is crucial when identifying errors in sequencing. Knowing that an error is present at a base and its complement in both the top and bottom strand (descendants) indicates that the error was present in the original sample. Likewise, knowing that the error is present in only the top or bottom strand (descendants) indicates an error either originating from clustering or any pre-clustering manipulation of the sample. Thus, there is a need to be able to identify strands originating from the top and bottom strands.

[0464] The present disclosure addresses this problem by introducing at least one strand identifier into at least one descendent of a forward and / or at least one descendant of a reverse strand. In one particular embodiment, a strand identifier is introduced into both the forward and reverse strand descendants. In this embodiment, the forward strand identifier and the reverse strand identifier are different, allowing descendants of the original forward strand to be distinguished, from descendants of the original reverse strand. By “original” here is meant original strand of the template polynucleotide. In one embodiment, the at least one strand identifier is introduced into at least one adaptor of the forward and / or reverse descendent.

[0465] As used herein, “top strand” may refer to the original forward strand of the target polynucleotide, while “top strand descendants” or “forward strand descendants” refer to both the forward and reverse strand amplicons of the “top strand” or original forward strand of the target polynucleotide.

[0466] Similarly, as used herein “bottom strand” may refer to the original reverse strand of the target polynucleotide, while “bottom strand descendants” or “reverse strand descendants” refer to both the forward and reverse strand amplicons of the “bottom strand” or original reverse strand of the target polynucleotide.

[0467] Fig. 7a and 7b depict an example of ‘post-tagmentation / pre-clustering’ adaptor structures according to the present disclosure that can enable top and bottom strand descendent identification and discernment: Fig. 7a shows a ‘symmetric surface transposome’ composition where the adaptors (i.e. Seq 1 or Seq 2) at both ends of the template insert are bound to the surface of a substrate via the transferred strand of the transposomes. Fig. 7b shows an ‘asymmetric surface transposome’ composition where one adaptor (Seq 1) is bound to the surface via its transferred strand and the other adaptor is bound via its non-transferred strand. In both cases the double stranded adaptors at either end of the double stranded insert include a strand identifier, either a forward strand descendent identifier or a reverse strand descendent identifier, which in this example is a non-complementary sequence of at least 1 base in length. For example, in Figure 7, Seq 4 and Seq 3’ are not complementary; this is also the case for Seq 6’ and Seq 5. Where Tn5 is the transposase used for the tagmentation reaction, Seq 7 and Seq 7’ are the transposase recognition sequence (i.e. the 19bp ME and ME’ sequences), respectfully.

[0468] As used herein, different Seq numbers - e.g. Seg 5, 7 etc. used herein denote different sequences. In one embodiment, the sequences may differ by one or more nucleotides. In another embodiment, the sequences may share no sequence identity. Alternatively, the sequences may share at least 1%, 2%, 3%, 4%, 5%, 1-%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90% or 95% overall sequence identity. The sequence identity of a variant can be determined using any number of sequence alignment programs known in the art. As an example, Emboss Stretcher from the EMBL-EBI may be used: https: / / www.ebi.ac.uk / idispatcher / psa / emboss stretcher (using default parameters: pair output format, Matrix = BLOSUM62, Gap open = 12, Gap extend = 2 for proteins; pair output format, Matrix = DNAfull, Gap open = 16, Gap extend = 4 for nucleotides).

[0469] As used herein, “ ’ “ (as in ME’, for example) denotes a complementary sequence (i.e. the complementary sequence to ME).

[0470] In one example of cluster generation, the fully adapted fragments shown in Figure 7 are denatured from one another, and loop over to hybridize to adjacent complementary immobilized surface amplification primers - i.e. Seq 1 and Seq 2 (which may also be referred to as lawn primers or first and second immobilized primers; such terms are interchangeable). A polymerase is then used to copy the templates to form double-stranded bridges. This is shown in Figure 8. When either the ‘symmetric surface transposome’ composition (Fig. 7a) or the ‘asymmetric surface transposome’ (Fig. 7b) are clustered using surface amplification primers, two types of bridges result within a cluster that differ between ‘descendants of the original (i.e. template) top (i.e. forward) strand (black; Fig. 8a) and descendants of the original (i.e. template) bottom (i.e. reverse) strand (grey; Fig.8b). These differences are in the adaptor sequences.

[0471] When the clusters are linearized by cleaving the immobilized surface primers - e.g. the Seq 2 oligonucleotide, as described above, the remaining single stranded strands are attached to the surface via the Seq 1 oligonucleotide (Fig 9a). These single strands - adapted through the Seq 1 oligonucleotide can loop over again and hybridise to adjacent complementary surface amplification primers to form new double-stranded loops. This process is called resynthesis. Following resynthesis of the bridges depicted in Fig. 8a and 8b, the cluster may be linearized again by cleaving the Seq 1 oligo and the remaining single stranded strands are now attached to the surface via the Seq 2 oligo (Fig. 9b). In both cases the top (black) (Figure 9b) and bottom (grey) strands (Figure 9a) have single strand adaptors at their 3’ ends that partially differ in their sequences. For example, in Fig. 9a, the top (black) strand has a 3’ adaptor sequence comprising Seq 2’-Seq 5’- Seq 7’ whereas the bottom (grey) strand has a 3’ adaptor sequence comprising Seq 2’-Seq 6’- Seq 7’. In other words, the presence of Seq 5 or 5’ denotes descendants from the top strand, while the presence of Seq 6 or 6’ denotes descendants from the bottom strand. As explained below, the difference in sequence between Seq 5 and Seq 6 may be exploited to selectively sequence either the top strand descendants or the bottom strand descendants by using different primers specific to Seq 5 or Seq 6 or portions thereof. Alternatively, Seq 5 and Seq 6 may be sequenced along with the template nucleic acid, and the presence of Seq 5 or Seq 6 is then subsequently used to distinguish the top and bottom strand descendants.

[0472] In one embodiment (Fig. 10a) the cluster may be sequenced using two sequencing primers P1 and P2 that bind preferentially to the top and bottom descendant strands within the cluster. For example, P1 primer may bind, partially or wholly, to only the adaptor Seq 5’-Seq 7’ sequence, and the P2 primer may bind, partially or wholly, to only the adaptor Seq 6’-Seq 7’. Thus, the P1 primer will only sequence the top (black) descendant strands and the P2 primer will only sequence the (reverse or forward sequence of the) bottom (grey) descendant strands. It will be appreciated that primers P1 and P2 maybe be added together and sequence the top and bottom descendant strands simultaneously, or they may be added sequentially and sequence the top and bottom descendant strands separately. In this latter case, the top and bottom strands will be distinguishable. Following cluster re-synthesis and linearization in preparation for the ‘paired read’, two additional primers P3 and P4 may likewise be used to sequence the top and bottom strand descendants either simultaneously or sequentially (Fig.

[0473] 10b). In this example, read 1 will sequence the reverse strand of the top and bottom descendants read 2 (following paired-end resynthesis) will sequence the forward strand of the top and bottom descendants.

[0474] In another embodiment (Fig. 11a) a sequence read may be obtained from the top and bottom strand descendants within a cluster using a primer that is common to both descendant types. For example, a primer P5 may bind partially or wholly to Seq 2’ and generate a read. The reads will differ in Seq 5’ and Seq 6’ portion of the adaptor which thus acts as strand identifiers. Although the raw signal from the sequencing instrument is mixed for Seq 5’ and Seq 6’, the strands may still be distinguishable by design. For example, if Seq 5’ reports a G which on a sequencer does not bear a fluorophore, and Seq 6’ reports a T which does bear a fluorophore, then the expectation is that only half of the strands (either the top or bottom descendants) in the cluster will generate a raw signal. If the top and bottom strands are present in equal proportion, then a 50% drop in signal intensity would be expected. The degree of change in signal intensity correlates with the ratio of top and bottom strands in the cluster.

[0475] Although it is not shown, it is appreciated that the sequencing primer (denoted as P5 in Fig. 11) could also be designed to hybridize to Seq 7 and similarly generate a read from Seq 3 and Seq 4. In this case the template insert reads can be generated simultaneously or separately to the P5 read, by a primer that binds to Seq 7’ or Seq 2’. In all the above permutations where common primers are used for sequencing the templates, the top and bottom strands cannot be quantifiably distinguished, but instead give a missed ‘signal’. Fig.

[0476] 11b depicts a similar embodiment for when the clusters are resysnthesized, linearized and generate the ‘paired read’.

[0477] Accordingly, the methods of the present disclosure comprise conducting a first sequence read using a first sequencing primer that substantially hybridises (i.e. wholly or partially) to the, or a portion thereof, of the 3’ adaptor of the amplified strand comprising the complement of the forward template strand identifier and conducting a second sequencing read using a second sequencing primer that substantially hybridises to a 3’ adaptor of the amplified strand comprising the complement of the reverse strand descendant identifier and determining a difference in sequence output.

[0478] In one embodiment, the sequencing primer may bind to either the complement of the forward template strand identifier or complement of the reverse strand descendant identifier, and optionally additionally to all or a part of the transposase recognition sequence (or its complement thereof) as shown in Figure 10a. The difference in signal between the first and second sequence reads will represent the proportion of forward strands derived from the forward strand of the template polynucleotide and the proportion of forward strands derived from the reverse strands of the target polynucleotide (the reverse complement strands). This is one example of what is meant by a “difference in sequence output” according to the present disclosure. The first and second sequencing reads may be conducted sequentially or concurrently using the methods described herein.

[0479] In an alternative embodiment, the first and second sequencing primers are identical. That is, the method may comprise conducting a first sequencing read using a universal first forward sequencing primer, wherein the first sequencing primer binds to a portion of the 3’ end of the forward strands, irrespective of whether those forward strands are derived from the forward template or are complements of the reverse strand. In one example, the universal primer may bind to the amplification domains (or complements thereof) as shown in Figure 11a. This sequencing read sequences the target polynucleotide (i.e. the insert) as well as the target and descendant strand identifiers. It is this difference in sequence - which as explained below - can be determined by a change in the raw signal and specifically a missed signal -that allows the proportion of forward strands derived from the forward strand of the template polynucleotide and the proportion of forward strands derived from the reverse strands (the reverse complement strands) to be determined. This is another example of what is meant by a “difference in sequence output” according to the present disclosure.

[0480] In addition to either of the above - i.e. either conducting a first and second sequencing read using strand-identifier specific sequencing primers, or using a universal forward sequencing primer, in a further embodiment, the method may additionally comprise conducting a 5’ sequence read using a sequencing primer that hybridises to a portion of the 5’ end of the linearised amplified strands, as described above. In one embodiment, the sequencing primer hybridises to the transposase recognition sequence (Seq 7 in Figure 11a). This sequencing read will sequence the 5’ strand identifiers (the complement of the reverse strand identifiers). This is another example of what is meant by a “difference in sequence output” according to the present disclosure.

[0481] In a further embodiment, the method may comprise conducting a second sequencing and third sequencing read. In this embodiment, the method comprises conducting paired-end re-synthesis before conducting the third and fourth sequence read.

[0482] The methods of the present disclosure may further comprise conducting a third sequence read using a third sequencing primer that substantially hybridises to a 3’ adaptor of the amplified strand comprising the complement of the reverse template strand identifier and conducting a fourth sequencing read using a fourth sequencing primer that substantially hybridises to a 3’ adaptor of the amplified strand comprising the complement of the forward strand descendant identifier and determining a difference in sequence output. In one embodiment, the sequencing primer may bind to either the complement of the reverse template strand identifier or forward strand descendant identifier, and optionally additionally to all or a part of the transposase recognition sequence (or its complement thereof) as shown in Figure 10a. The difference in signal between the first and second sequence reads will represent the proportion of reverse strands derived from the reverse strand of the template polynucleotide and the proportion of reverse strands derived from the forward strand of the target polynucleotide (the reverse complement strands). This is one example of what is meant by a “difference in sequence output” according to the present disclosure. The third and fourth sequencing reads may be conducted sequentially or concurrently using the methods described herein.

[0483] In an alternative embodiment, the first and second sequencing primers are identical. That is, the method may comprise conducting a first sequencing read using a universal first reverse sequencing primer, wherein the first sequencing primer binds to a portion of the 3’ end of the reverse strands, irrespective of whether those reverse strands are derived from the reverse template or are complements of the forward strand. In one example, the universal primer may bind to the amplification domains (or complements thereof) as shown in Figure 11b. This sequencing read sequences the target polynucleotide (i.e. the insert) as well as the target and descendant strand identifiers. It is this difference in sequence - which as explained below - can be determined by a change in the raw signal, and specifically a missed signal -that allows the proportion of reverse strands derived from the reverse strand of the target polynucleotide and the proportion of reverse strands derived from the forward strands (the reverse complement strands) to be determined. This is another example of what is meant by a “difference in sequence output” according to the present disclosure.

[0484] In addition to either of the above - i.e. either conducting a third and fourth sequencing read using strand-identifier specific sequencing primers, or using a universal reverse sequencing primer, in a further embodiment, the method may additionally comprise conducting a 5’ sequence read using a sequencing primer that hybridises to a portion of the 5’ end of the linearised amplified strands, as described above. In one embodiment, the sequencing primer hybridises to the transposase recognition sequence (Seq 7 in Figure 11b). This sequencing read will sequence the 5’ strand identifiers (the complement of the forward strand identifiers). This is another example of what is meant by a “difference in sequence output” according to the present disclosure.

[0485] Methods of introducing strand identifiers using symmetric transposome complexes.

[0486] An example of a method of preparing a nucleic acid cluster with a strand identifier using symmetric transposome complexes is shown in Figures 12 to 16. Fig. 12a depicts a symmetric composition of transposomes on a flow cell. That is, Figure 12a shows a solid support comprising a first transposome and a second transposome complex, both immobilised at their 5’ ends (i.e. by their transferred strands) to the solid support. In this example, the immobilised transferred strands comprise either a first adaptor or a second adaptor. Each adaptor comprises at least one amplification domain, as described above. As also described above, each adaptor (also referred to herein as a “adaptor structure”) further comprises either a forward strand descendent identifier or a reverse strand descendent identifier - for example, Seq 4 or Seq 5. The first and second immobilised primers are not shown. Seq 1 and Seq 2 sequence are identical or substantially identical in sequence to the immobilised primers. For example, Seq 1 may be the surface amplification primer P5 and Seq 2 may be the surface amplification primer P7, or vice versa, as described above. Seq 7 and Seq 7’ are complementary sequences and are the transposase recognition sequences (for example,, the 19bp ME and ME’ sequences).

[0487] Fig. 12b depicts a bridged product of tagmentation between the transposomes depicted in Fig. 12a following removal of the transposase enzyme. The transposase enzymes can be removed as described above. Seq 1, Seq 4, Seq 7 and Seq 2, Seq 5, Seq 7 represent the transferred strands that are covalently attached to the 5’ ends of the tagmented target polynucleotide. Seq 7’ represents the non-transferred strands that are not covalently attached to the 3’ ends of the target polynucleotide.

[0488] Fig. 12c represents the tagmented bridge product following removal of the nontransferred strands Seq 7’; for example, denaturation by heat under flow, as described above. Following removal, two strand identifiers, which in this example, are oligonucleotides containing Seq 7’ can be hybridised to the bridge product (the double-stranded target polynucleotide): a first strand identifier (e.g. a forward strand descendent identifier) which is an oligo comprising 5’- Seq 7’, Seq 3’, Seq T hybridises to adaptor 5’-Seq 1, Seq 4, Seq 7, where Seq 3’ is non complementary to Seq 4; and a second strand identifier (e.g. a reverse strand descendent identifier), which is an oligonucleotide comprising 5’ - Seq 7’, Seq 6’, Seq 2’ hybridises to adaptor 5’- Seq 2, Seq 5, Seq 7 where Seq 6’ is non complementary to Seq 5 (Fig. 12d). Fig. 12e depicts the outcome when the hybridised oligonucleotides are covalently connected to the 3’ ends of the bridge strands by means of a ‘gap-fill’ reaction employing a non-strand displacing polymerase and a ligase.

[0489] Fig. 12f depicts the nucleic acid substrates necessary for cluster amplification of the bridge template depicted in Fig. 12e. The amplification requires two immobilised surface amplification primers, Seq 1 and Seq 2. As described above the outcome of cluster amplification is a cluster that contain two types of double stranded bridges: one for the descendant strands of the original top strand (black)(Fig. 12g) and one for the descendant strands of the original bottom strand (grey)(Fig 12h). Sequencing of these clusters to identify the top and bottom strands is as described in Figs. 9a, 9b, 10a, 10b, 11a and 11b.

[0490] Fig. 13a depicts another embodiment of a symmetric composition of transposomes on a solid support, such as a flow cell. In this embodiment, the transposomes are pre-formed with full length non-transferred strands. Two transposomes (a first transposome complex and a second transposome complex) are depicted; each as a dimer wherein only one of the monomers is labelled in the figure. One monomer contains a duplex nucleic acid adaptor that bears a non-complementary double-stranded portion Seq 4 / Seq 3’; the other monomer contains a duplex DNA adaptor that also bears a non-complementary double-stranded portion Seq 5 / Seq 6’.

[0491] Fig. 13b depicts the bridged product of tagmentation between the transposomes depicted in Fig. 13a and a double-stranded target polynucleotide, following removal of the transposase enzyme and again, ‘gap-fill’ treatment to covalently connect the non-transferred strands to the 3’ ends of the bridged template.

[0492] Fig. 13c depicts the nucleic acid substrates necessary for cluster amplification of the bridge template depicted in Fig. 13b. The amplification requires two immobilised surface amplification primers, Seq 1 and Seq 2. As described above the outcome of cluster amplification is a cluster that contain two types of double stranded bridges: one for the descendant strands of the original top strand (black) and one for the descendant strands of the original bottom strand (grey) Fig. 13d. Sequencing of these clusters to identify the top and bottom strands is as described in Figs. 9a, 9b, 10a, 10b, 11a and 11b.

[0493] Fig. 14a depicts another embodiment of a surface bound symmetric transposome composition, where the transposome adaptors are partially double-stranded and form a forked adaptor. Two transposome dimers - a first and second transposome complex are depicted but only one monomer is annotated per dimer: one monomer bears an oligo comprising the sequence 5’ - Seq 1, Seq 4, Seq 7 and another oligo comprising the sequence 5’ - Seq 7’, Seq 3’, Seq 8’; these two oligos are complementary and hybridised together at the Seq 7, Seq 7’ portion; the monomer of the other dimer bears an oligo comprising the sequence 5’ - Seq 2, Seq 5, Seq 7 and another oligo comprising the sequence 5’ - Seq 7’, Seq 6’, Seq 9’; these two oligos are complementary and hybridised together at the Seq 7, Seq 7’ portion.

[0494] Accordingly, in one embodiment, the non-transferred strand of the first transposome complex comprises a fourth adaptor, wherein the fourth adaptor comprises, in the 5’ to 3’ direction, a transposase recognition sequence, a reverse strand descendant identifier and at least one amplification domain, wherein the amplification domain is a complement of the first amplification domain or a fourth amplification domain. In another embodiment, the non-transferred strand of the second transposome complex comprises a third adaptor, wherein the third adaptor comprises, in the 5’ to 3’ direction, a transposase recognition sequence, a forward strand descendant identifier and at least one amplification domain, wherein the amplification domain is a complement of the second amplification domain or a third amplification domain.

[0495] Fig. 14b depicts the bridged product of tagmentation between the transposomes depicted in Fig. 14a, following removal of the transposase enzyme and ‘gap-fill’ treatment to covalently connect the non-transferred strands to the 3’ ends of the bridged template.

[0496] Fig. 14c depicts the DNA substrates necessary for cluster amplification of the bridge template depicted in Fig. 14b. The amplification requires four surface amplification primers, Seq 1, Seq 9, Seq 2 and Seq 8. The top strand (black) is amplified using the Seq 1 and Seq 9 primers whereas the bottom strand (grey) is amplified using the Seq 2 and Seq 8 primers. As described above, the outcome of cluster amplification is a cluster that contain two types of double stranded bridges: one for the descendant strands of the original top strand (black) and one for the descendant strands of the original bottom strand (grey), Fig. 14d. Sequencing of these clusters to identify the top and bottom strands is as described in Figs. 9a, 9b, 10a, 10b, 11a and 11b.

[0497] In yet another embodiment, clusterable bridges that are products of tagmentation can be formed that are wholly complementary and double stranded in their adaptor structures. However, at least one of the adaptor bases comprises a modified base that can be converted to a different base in a reaction after tagmentation.

[0498] Fig. 15a depicts a bridged product of tagmentation between a first transferred strand bearing a sequence spanning Seq 1 and Seq 7 containing an 8-oxo-G (8-oxoguanine) base and a second transferred strand bearing a sequence spanning S2 and Seq 7 containing an 8-oxo-G base. Fig. 15b depicts the bridged product after hybridization of adaptors complementary to the transferred strand but containing a natural C base that base pairs with the 8-oxo-G bases. Fig. 15c depicts the bridged product following a gap-fill reaction to covalently attach the non-transferred strands to the 3’ ends of the top and bottom strands.

[0499] It will be understood that the product depicted in Fig. 15c can also be formed using full length transposome adaptors (i.e. transferred and non-transferred strands) that are fully complementary along their length. In such case, this is an alternative to hybridizing additional adaptors as depicted in Fig. 15b. It will also be understood that the product depicted in Fig.

[0500] 15c can also be formed by first dehybridising the Seq 7’ non-transferred strand and replacing it with a full-length complementary strand that spans from 5’ - Seq 7’ to Seq T, and 5’ - Seq 7’ to Seq 2’, followed by a gap-fill reaction, as described above. Fig. 15d depicts the nucleic acid substrates necessary for cluster amplification of the bridge template depicted in Fig. 15c. The amplification requires two immobilised surface amplification primers, Seq land Seq 2. As described above, the outcome of cluster amplification is a cluster that contain two types of double stranded bridges: one for the descendant strands of the original top strand (black) and one for the descendant strands of the original bottom strand (grey).

[0501] When the clusters are linearized by cleaving the Seq 2 oligo, the remaining single stranded strands are attached to the surface via the Seq 1 oligo (Fig 15f). Following resynthesis of the bridges depicted in Fig. 15e, the cluster may be linearized again by cleaving the Seq 1 oligo and the remaining single stranded strands are now attached to the surface via the Seq 2 oligo (Fig. 15g). In both cases, the top (black) and bottom (grey) descendant strands have single strand adaptors at their 3’ ends that partially differ in their sequences and thus can be identified by sequencing with a common sequencing primer, such as P5 or P6, as described above.

[0502] Fig. 16a to Fig. 16c depicts yet another embodiment and the generation of clusterable bridges that are products of tagmentation and that are wholly complementary and double stranded in their adaptor structures, and in which at least one of the adaptor bases comprises a modified base that can be converted to a different base in a reaction after tagmentation. In this embodiment the modified base is a methylated C in the adaptor sequence.

[0503] It will be understood that the product depicted in Fig. 16c can also be formed using transposome adaptors (i.e. transferred and non-transferred strands) that are fully complementary along their length. In such case, this is an alternative to hybridizing additional adaptors as depicted in Fig. 16b. It will also be understood that the product depicted in Fig.

[0504] 16c can also be formed by first dehybridising the Seq 7’ non-transferred strand and replacing it with a full-length complementary strand that spans from 5’ - Seq 7’ to Seq T, and 5’ - Seq 7’ to Seq 2’, followed by a gap-fill reaction, as described above.

[0505] Fig. 16d depicts the product of a reaction to convert a methylated-C base in an adaptor to a T. Such reactions are possible using a deaminase enzyme, as described above. Fig. 16e depicts the amplification products using two surface amplification primers, Seq land Seq 2. As described above, the outcome of cluster amplification is a cluster that contain two types of double stranded bridges: one for the descendant strands of the original top strand (black) and one for the descendant strands of the original bottom strand (grey).

[0506] When the clusters are linearized by cleaving the Seq 2 oligo, the remaining single stranded strands are attached to the surface via the Seq 1 oligo (Fig 16f). Following resynthesis of the bridges depicted in Fig. 16e, the cluster may be linearized again by cleaving the Seq 1 oligo and the remaining single stranded strands are now attached to the surface via the Seq 2 oligo (Fig. 16g). In both cases, the top (black) and bottom (grey) descendant strands have single strand adaptors at their 3’ ends that partially differ in their sequences and thus can be identified by sequencing with a common sequencing primer, such as P5 or P6 as described earlier.

[0507] Fig. 16h - 16j depicts an alternative workflow where a G in the transferred strands is base-paired opposite a methylated C in a connected non-transferred strand using the methods described in Fig. 16a to 16c. As described above, the outcome of cluster amplification is a cluster that contain two types of double stranded bridges: one for the descendant strands of the original top strand (black) and one for the descendant strands of the original bottom strand (grey). The two types of bridges differ in their adaptor sequences and enable strand identification as described earlier.

[0508] Fig. 16k depicts an embodiment whereby the clusterable bridge depicted in Fig. 16c can be first generated by a polymerase reaction that extends the 3’ ends of the top and bottom strands and incorporates a G base opposite the methyl-C in the transferred strand. Likewise, Fig. 161 depicts an embodiment whereby the clusterable bridge depicted in Fig. 16h can be first generated by a polymerase reaction that extends the 3’ ends of the top and bottom strands and incorporates a G base opposite the methyl-C in the transferred strand.

[0509] Methods of introducing strand identifiers using asymmetric transposome complexes.

[0510] An example of a method of preparing a nucleic acid cluster with a strand identifier using asymmetric transposome complexes is shown in Figures 17 to 21.

[0511] Fig. 17a depicts an asymmetric composition of transposomes on a flow cell. That is, Figure 17a shows a solid support comprising a first transposome and a second transposome complex, where the first transposome complex is immobilised at a 5’ end (i.e. by the transferred strands) to the solid support and the second transposome complex is immobilised at a 3’ end to the solid support (i.e. by the non-transferred strand). In this example, the transferred strands comprise either a first adaptor or a second adaptor. Each adaptor comprises at least one amplification domain, as described above. As also described above, each adaptor (also referred to herein as a “adaptor structure”) further comprises either a forward strand descendent identifier or a reverse strand descendent identifier - for example, Seq 4 or Seq 5. The first and second immobilised primers are not shown. Seq 1 and Seq 2 sequence are identical or substantially identical in sequence to the immobilised primers. For example, Seq 1 may be the surface amplification primer P5 and Seq 2 may be the surface amplification primer P7, or vice versa, as described above. Seq 7 and Seq 7’ are complementary sequences and are the transposase recognition sequences (for example, the 19bp ME and ME’ sequences). Fig. 17b depicts the tagmented bridge product following removal of the non-transferred strands Seq 7’; for example, denaturation by heat underflow, as described herein. Following removal, two adaptors (i.e. a third adaptor and a fourth adaptor) comprising Seq 7’ can be rehybridised to the bridge product: an oligo comprising 5’- Seq 7’, Seq 3’, Seq T (which may be referred to herein as adaptor 3) hybridises to adaptor 5’-Seq 1, Seq 4, Seq 7, where Seq 3’ is non complementary to Seq 4; and an oligo comprising 5’ - Seq 7’, Seq 6’, Seq 2’ (which may be referred to herein as adaptor 4) hybridises to adaptor 5’- Seq 2, Seq 5, Seq 7 where Seq 6’ is non complementary to Seq 5 (Fig. 17c).

[0512] Fig. 17d depicts the outcome when the hybridised adaptors are covalently connected to the 3’ ends of the bridge strands by means of a ‘gap-fill’ reaction employing a non-strand displacing polymerase and a ligase, as described above.

[0513] Fig. 17e depicts the nucleic acid substrates necessary for cluster amplification of the bridge template depicted in Fig. 17d. The amplification requires two immobilised surface amplification primers, Seq 1 and Seq 2. As described above, the outcome of cluster amplification is a cluster that contain two types of double stranded bridges: one for the descendant strands of the original top strand (black)(Fig. 17f) and one for the descendant strands of the original bottom strand (grey)(Fig. 17g). Sequencing of these clusters to identify the top and bottom strands is as described in Figs. 9a, 9b, 10a, 10b, 11a and 11b.

[0514] Fig. 18a depicts another embodiment of an asymmetric composition of transposomes on a solid support, such as a flow cell. In this embodiment, the transposomes are pre-formed with full length non-transferred strands. Two transposomes are depicted; each as a dimer wherein only one of the monomers is labelled in the figure. One monomer contains a duplex DNA adaptor that bears a non-complementary double-stranded portion Seq 4 / Seq 3’ (a reverse strand descendant identifier or a forward strand descendant identifier) the other monomer contains a duplex DNA adaptor that bears a non-complementary double-stranded portion Seq 5 / Seq 6 (a forward strand descendant identifier or a reverse strand descendant identifier respectively).

[0515] Fig. 18b depicts the bridged product of tagmentation between the transposomes depicted in Fig. 18a, following removal of the transposase enzyme and ‘gap-fill’ treatment to covalently connect the non-transferred strands to the 3’ ends of the bridged template.

[0516] Fig. 18c depicts the nucleic acid substrates necessary for cluster amplification of the bridge template depicted in Fig. 18b. The amplification requires two immobilised surface amplification primers, Seq 1 and Seq 2. As described above the outcome of cluster amplification is a cluster that contain two types of double stranded bridges: one for the descendant strands of the original top strand (black)(Fig 18d) and one for the descendant strands of the original bottom strand (grey)(Fig 18e). Sequencing of these clusters to identify the top and bottom strands is as described in Figs. 9a, 9b, 10a, 10b, 11a and 11b.

[0517] Fig. 19a depicts another embodiment of a surface bound asymmetric transposome composition, where the transposome adaptors are partially double-stranded and form a forked adaptor. Two transposome dimers are depicted but only monomer is annotated per dimer: one monomer bears an oligo comprising the sequence 5’ - Seq 1, Seq 4, Seq 7 and another oligo comprising the sequence 5’ - Seq 7’, Seq 3’, Seq 8’; these two oligos are complementary and hybridised together at the Seq 7, Seq 7’ portion; the monomer of the other dimer bears an oligo comprising the sequence 5’ - Seq 2, Seq 5, Seq 7 and another oligo comprising the sequence 5’ - Seq 7’, Seq 6’, Seq 9’; these two oligos are complementary and hybridised together at the Seq 7, Seq 7’ portion.

[0518] Fig. 19b depicts the bridged product of tagmentation between the transposomes depicted in Fig. 19a, following removal of the transposase enzyme and ‘gap-fill’ treatment to covalently connect the non-transferred strands to the 3’ ends of the bridged template, as described above.

[0519] Accordingly, in one embodiment, the non-transferred strand of the first transposome complex comprises a fourth adaptor, wherein the fourth adaptor comprises, in the 5’ to 3’ direction, a transposase recognition sequence, a reverse strand descendant identifier and at least one amplification domain, wherein the amplification domain is a complement of the first amplification domain or a fourth amplification domain.

[0520] In another embodiment, the non-transferred strand of the second transposome complex comprises a third adaptor, wherein the third adaptor comprises, in the 5’ to 3’ direction, a transposase recognition sequence, a forward strand descendant identifier and at least one amplification domain, wherein the amplification domain is a complement of the second amplification domain or a third amplification domain.

[0521] In an embodiment, there is provided a solid support comprising a first immobilised (surface) primer, a second immobilised (surface) primer, and optionally a third immobilised (surface) primer and a fourth immobilised (surface) primer, and a first and second transposome complex. In one embodiment, the first transposome complex (comprises a dimer, wherein each monomer) comprises a transferred and non-transferred strand. In one embodiment, the second transposome complex (comprises a dimer, wherein each monomer) comprises a transferred and non-transferred strand. The (or each) transferred strand may comprise at least one amplification domain, a target strand identifier and a transposome recognition sequence. In a further embodiment, the (or each) non-transferred strand may comprise at least one amplification domain, a descendant strand identifier and a transposome recognition sequence. Fig. 19c depicts the nucleic acid substrates necessary for cluster amplification of the bridge template depicted in Fig. 19b. In this embodiment, the amplification requires four surface amplification primers, Seq 1, Seq 9, Seq 2 and Seq 8 (i.e. first immobilised primers, second immobilised primers, third immobilised primers and fourth immobilised primers). The top strand (black) is amplified using the Seq 1 and Seq 9 primers whereas the bottom strand (grey) is amplified using the Seq 2 and Seq 8 primers. As described above, the outcome of cluster amplification is a cluster that contain two types of double stranded bridges (Fig. 19d): one for the descendant strands of the original top strand (black) and one for the descendant strands of the original bottom strand (grey). Sequencing of these clusters to identify the top and bottom strands is as described in Figs. 9a, 9b, 10a, 10b, 11a and 11b.

[0522] Fig. 20a depicts a bridged product of tagmentation using an asymmetric composition of surface-based transposomes, between a first transferred strand bearing a sequence spanning Seq 1 and Seq 7 containing an 8-oxo-G base and a second transferred strand bearing a sequence spanning S2 and Seq 7 containing an 8-oxo-G base. Fig. 20b depicts the tagmented bridge product following removal of the non-transferred strands Seq 7’; for example, denaturation by heat under flow.

[0523] Fig. 20c depicts the bridged product after hybridization of oligos complementary to the transferred strand but containing a natural C base that base pairs with the 8-oxo-G bases. Fig.

[0524] 20d depicts the bridged product following a gap-fill reaction to covalently attach the nontransferred strands to the 3’ ends of the top and bottom strands. It will be understood that the product depicted in Fig. 20d can also be formed using full length transposome adaptors that are fully complementary along their length. In such case, this is an alternative to hybridizing additional adaptors as depicted in Fig. 20c.

[0525] Fig. 20e depicts the nucleic acid substrates necessary for cluster amplification of the bridge template depicted in Fig. 20d. The amplification requires two surface amplification primers, Seq land Seq 2. As described above, the outcome of cluster amplification is a cluster that contain two types of double stranded bridges (Fig. 20f): one for the descendant strands of the original top strand (black) and one for the descendant strands of the original bottom strand (grey).

[0526] Fig. 21a to Fig. 21d depicts yet other embodiments and generation of clusterable bridges that are products of tagmentation using an asymmetric transposome composition, and that are wholly complementary and double stranded in their adaptor structures, and in which at least one of the adaptor bases comprises a modified base that can be converted to a different base in a reaction after tagmentation. In this embodiment the modified base is a methylated C in the adaptor sequence. Fig. 21b and Fig. 21c depict a method whereby the non-transferred strands Seq 7’ are dehybridised from the transferred strands (e.g. by heat denaturation as described above) and replaced by annealing full length complementary non-transferred strands (adaptors 3 and 4), followed by a gap-fill reaction to connect them to the 3’ ends of the top and bottom strands. It will be understood product depicted in Fig. 21 d can also be formed using transposome adaptors that are fully complementary along their length. In such case, this is an alternative to hybridizing additional adaptors as depicted in Fig. 21c.

[0527] Fig. 21 e depicts the product of a reaction to convert a methylated-C base in an adaptor to a T. Such reactions are possible using a deaminase enzyme, as described above. Fig. 21 f depicts the amplification products using two immobilised surface amplification primers, Seq land Seq 2. As described above the outcome of cluster amplification is a cluster that contain two types of double stranded bridges: one for the descendant strands of the original top strand (black) and one for the descendant strands of the original bottom strand (grey).

[0528] Fig. 22a and 22b depicts an alternative method where a G in the transferred strands is base-paired opposite a methylated C in a connected non-transferred strand using the methods described in Fig. 22a to 22c. As described above the outcome of cluster amplification is a cluster that contain two types of double stranded bridges (Fig. 22c: one for the descendant strands of the original top strand (black) and one for the descendant strands of the original bottom strand (grey). The two types of bridges differ in their adaptor sequences and enable strand identification as described above.

[0529] Fig. 23 depicts an embodiment whereby the clusterable bridge depicted in Fig. 21d can be first generated by a polymerase reaction that extends the 3’ ends of the top and bottom strands and incorporates a G base opposite the methyl-C in the transferred strand. Likewise, Fig. 24 depicts an embodiment whereby the clusterable bridge depicted in Fig. 22a can be first generated by a polymerase reaction that extends the 3’ ends of the top and bottom strands and incorporates a G base opposite the methyl-C in the transferred strand.

[0530] Concurrent sequencing methods for error detection and methylation analysis According to an embodiment of the present disclosure, there is provided a method of preparing polynucleotide sequences for sequencing, comprising:

[0531] providing a solid support comprising first immobilised primers and second immobilised primers, and a double-stranded target polynucleotide comprising a forward target strand and a reverse target strand, seeding both the forward target strand and the reverse target strand onto the solid support, amplifying the forward target strand and the reverse target strand to produce amplified forward strands and amplified reverse complement strands each covalently attached to a first immobilised primer, and amplified reverse strands and amplified forward complement strands each covalently attached to a second immobilised primer; and preparing the amplified forward strands and amplified reverse complement strands for concurrent sequencing, or preparing the amplified reverse strands and amplified forward complement strands for concurrent sequencing.

[0532] Previous approaches in concurrent sequencing of forward strands and reverse complement strands (or reverse strands and forward complement strands) have relied on specialised library preparations, such as tandem inserts to generate forward strands and reverse complement strands (or reverse strands and forward complement strands) on concatenated strands on a solid support, or loop-fork library preparations to generate forward strands and reverse complement strands (or reverse strands and forward complement strands) on separate strands on a solid support. By utilising direct double-stranded seeding methods, it is possible to avoid the need to utilise these specialised library preparations, allowing the detection of errors and epigenetic modifications such as methylation and hemimethylation (by comparing outputs from forward strands and reverse complement strands; or reverse strands and forward complement strands), but at higher turnaround times.

[0533] The solid support may be any solid support as described herein. In some embodiments, the solid support is a flow cell. In particular, the solid support may be a nonpatterned flow cell. A non-patterned flow cell may be especially suitable for this method as this can reduce the impact of physical occlusion and boundary effects from nanowells on the amplification process.

[0534] The steps of seeding both the forward target strand and the reverse target strand onto the solid support and amplifying the forward target strand and the reverse target strand can be conducted using any cluster generation and / or amplification step as described herein.

[0535] In one embodiment, the step of seeding both the forward target strand and the reverse target strand onto the solid support is conducted using a tagmentation reaction.

[0536] In one embodiment, the step of seeding both the forward target strand and the reverse target strand onto the solid support is conducted using a transposase (e.g. a first transposase and a second transposase). As mentioned herein, non-limiting examples of transposases include Tn5 transposase, Tn7 transposase, MuA transposase and Sleeping Beauty (SB) transposase. A transposase as presented herein can also include integrases from retrotransposons and retroviruses. The seeding of the forward target strand may be conducted using a first transposase. The seeding of the reverse target strand may be conducted using a second transposase.

[0537] In one embodiment, both the forward target strand and the reverse target strand are seeded into a same well of the solid support.

[0538] In one embodiment, a 5’-end of the forward target strand is covalently attached to a 3’-end of a first transferred strand. As described herein, the attachment (e.g. covalent attachment) may be direct (i.e. without any intermediary sequences spacing the forward target strand from the first transferred strand).

[0539] In one embodiment, a 5’-end of the first transferred strand is covalently attached to a 3’-end of a first amplification domain. As described herein, the attachment of the first transferred strand to the first amplification domain may be direct or indirect (but overall may be covalently attached). For indirect attachment (e.g. covalent attachment), the first transferred strand may be spaced from the first amplification domain by a first sequencing primer sequence and / or a first index sequence.

[0540] In one embodiment, the first transferred strand is immobilised onto the solid support. The term “immobilised” may include immobilisation via covalent bonds, or strong non-covalent bonding interactions, such as host-guest interactions (e.g. streptavidin-biotin interactions).

[0541] In one embodiment, a 5’-end of the reverse target strand is covalently attached to a 3’-end of a second transferred strand. As described herein, the attachment (e.g. covalent attachment) may be direct (i.e. without any intermediary sequences spacing the reverse target strand from the second transferred strand).

[0542] In one embodiment, a 5’-end of the second transferred strand is covalently attached to a 3’-end of a second amplification domain. As described herein, the attachment of the second transferred strand to the second amplification domain may be direct or indirect (but overall may be covalently attached). For indirect attachment (e.g. covalent attachment), the second transferred strand may be spaced from the second amplification domain by a second sequencing primer sequence and / or a second index sequence.

[0543] In some embodiments (“symmetric” double-stranded seeding), the second transferred strand is immobilised onto the solid support. With such embodiments, the step of amplifying the forward target strand and the reverse target strand may be conducted using bridge amplification and / or exclusion amplification. In alternative embodiments (“asymmetric” doublestranded seeding), the second non-transferred strand is immobilised onto the solid support. With such embodiments, the step of amplifying the forward target strand and the reverse target strand may be conducted using exclusion amplification. Again, the term “immobilised” may include immobilisation via covalent bonds, or strong non-covalent bonding interactions, such as host-guest interactions (e.g. streptavidin-biotin interactions).

[0544] In one embodiment, prior to amplification, the transposase (e.g. the first transposase and the second transposase) may be removed.

[0545] In one embodiment, prior to amplification, an extension reaction is conducted to extend the forward target strand and the reverse target strand. This step may be conducted after removal of the transposase. In embodiments where the detection of modified cytosines is desired, then the method may further comprise a step of treating the forward target strand and the reverse target strand with a conversion agent, prior to the step of amplifying the forward target strand and the reverse target strand, wherein the conversion reagent is configured to convert a modified cytosine to thymine or a nucleobase which is read as thymine / uracil, and / or wherein the conversion reagent is configured to convert an unmodified cytosine to uracil or a nucleobase which is read as thymine / uracil.

[0546] In one embodiment, the step of treating the forward target strand and the reverse target strand with a conversion agent is conducted after the step of seeding both the forward target strand and the reverse target strand onto the solid support.

[0547] In one embodiment, the conversion agent comprises an enzyme. For example, the enzyme may be a cytidine deaminase. In particular embodiments, the cytidine deaminase is a wild-type cytidine deaminase or a mutant cytidine deaminase. In still further embodiments, the cytidine deaminase is a mutant cytidine deaminase. Cytidine deaminases are especially advantageous as they are able to effect the conversion without causing significant side reactions, which may affect the accuracy for determining modified cytosines and base-calling. Mutant cytidine deaminases are particularly advantageous as they may allow greater specificity for modified cytosines (or unmodified cytosines), leading to even greater accuracy for determining modified cytosines.

[0548] Various methods of preparing the amplified forward strands and amplified reverse complement strands for concurrent sequencing, or preparing the amplified reverse strands and amplified forward complement strands for concurrent sequencing are possible, provided that the strands are put into a state where a sequencing process will lead to sequencing of the amplified forward strands and the amplified reverse complement strands at the same time, or sequencing of the amplified reverse strands and the amplified forward complement strands at the same time.

[0549] Accordingly, in some embodiments, the step of preparing the amplified forward strands and amplified reverse complement strands for concurrent sequencing comprises simultaneously contacting first sequencing primer binding sites located after a 3’-end of the amplified forward strands with first primers and second sequencing primer binding sites located after a 3’-end of the amplified reverse complement strands with second primers; or wherein the step of preparing the amplified reverse strands and amplified forward complement strands for concurrent sequencing comprises simultaneously contacting third sequencing primer binding sites located after a 3’-end of the amplified reverse strands with third primers and fourth sequencing primer binding sites located after a 3’-end of the amplified forward complement strands with fourth primers. In other embodiments, the step of preparing the amplified forward strands and amplified reverse complement strands for concurrent sequencing comprises simultaneously nicking amplified forward complement strands hybridised to the amplified forward strands and nicking amplified reverse strands hybridised to the amplified reverse complement strands; or wherein the step of preparing the amplified reverse strands and amplified forward complement strands for concurrent sequencing comprises simultaneously nicking amplified reverse complement strands hybridised to the amplified reverse strands and nicking amplified forward strands hybridised to the amplified forward complement strands.

[0550] In one embodiment, the amplified forward strands and amplified reverse complement strands, and the amplified reverse strands and amplified forward complement strands, together form a cluster on the solid support.

[0551] In some embodiments, the method comprises a step of preparing the amplified forward strands and amplified reverse complement strands for concurrent sequencing, and wherein the amplified reverse strands and the amplified forward complement strands are removed prior to concurrent sequencing. In a further embodiment, the amplified forward strands and amplified reverse complement strands form a duoclonal cluster on the solid support.

[0552] In other embodiments, the method comprises preparing the amplified reverse strands and amplified forward complement strands for concurrent sequencing, and wherein the amplified forward strands and the amplified reverse complement strands are removed prior to concurrent sequencing. In a further embodiment, the amplified reverse strands and amplified forward complement strands form a duoclonal cluster on the solid support.

[0553] In one embodiment, the population of amplified forward strands prepared for concurrent sequencing is substantially equal to the population of amplified reverse complement strands prepared for concurrent sequencing; and / or wherein the population of amplified reverse strands prepared for concurrent sequencing is substantially equal to the population of amplified forward complement strands prepared for concurrent sequencing. In a further embodiment, the total population of amplified forward strands is substantially equal to the population of amplified reverse complement strands; and / or wherein the total population of amplified reverse strands is substantially equal to the total population of amplified forward complement strands.

[0554] According to another embodiment of the present disclosure, there is provided a method of sequencing polynucleotide sequences, comprising:

[0555] preparing polynucleotides for sequencing using a method as described herein, and concurrently sequencing nucleobases in the amplified forward strands and amplified reverse complement strands, or concurrently sequencing the amplified reverse strands and amplified forward complement strands. In one embodiment, the method further comprises a step of identifying differences when comparing sequence outputs from the amplified forward strands and amplified reverse complement strands, or when comparing sequence outputs from the amplified reverse strands and amplified forward complement strands.

[0556] In one embodiment, the step of identifying differences comprises associating the differences with the presence of errors.

[0557] In some embodiments where the forward target strand and the reverse target strand have been treated with a conversion agent, the step of identifying differences comprises associating the differences with the presence of modified cytosines.

[0558] In one embodiment, the method further comprises a step of conducting paired-end reads.

[0559] Kits

[0560] Methods as described herein may be performed by a user physically. In other words, a user may themselves conduct the methods of preparing polynucleotide sequences for sequencing as described herein, and as such the methods as described herein may not need to be computer-implemented.

[0561] In another aspect of the present disclosure, there is provided a kit comprising instructions for preparing polynucleotide sequences for sequencing as described herein, and / or for sequencing polynucleotide sequences as described herein.

[0562] In an embodiment of the present disclosure, there is provided a kit for performing the methods of the invention. For example, there is provided a kit for distinguishing sequences derived from amplification of the target polynucleotide or from amplification of a reverse complement. In some embodiments the kit comprises a first and second sequencing primer, wherein the first sequencing primer is capable of hybridising to a template strand identifier or portion thereof and the second sequencing primer is capable of binding to a descendant strand identifier, or portion thereof. In one embodiment, the kit comprises a first, second, third and fourth sequencing primer, wherein the first sequencing primer is capable of hybridising to a forward target strand identifier or portion thereof, the second sequencing primer is capable of hybridising to a reverse strand descendant identifier or portion thereof, the third sequencing primer is capable of hybridising to a reverse strand target identifier or portion thereof, and the fourth sequencing primer is capable of hybridising to a forward strand descendant identifier or portion thereof. Optionally, the kit further comprises at least one amplification primer. Optionally, the kit may comprises a solid support, as described herein.

[0563] Computer programs and products

[0564] In other embodiments, methods as described herein may be performed by a computer. In other words, a computer may contain instructions to conduct the methods of preparing polynucleotide sequences for sequencing as described herein, and as such the methods as described herein may be computer-implemented.

[0565] Accordingly, in another aspect of the present disclosure, there is provided a data processing device comprising means for carrying out the methods as described herein.

[0566] The data processing device may be a polynucleotide sequencer.

[0567] The data processing device may comprise reagents used for synthesis methods as described herein.

[0568] The data processing device may comprise a solid support, such as a flow cell.

[0569] In another aspect of the present disclosure, there is provided a computer program product comprising instructions which, when the program is executed by a processor, cause the processor to carry out the methods as described herein.

[0570] In another aspect of the present disclosure, there is provided a computer-readable storage medium comprising instructions which, when executed by a processor, cause the processor to carry out the methods as described herein.

[0571] In another aspect of the present disclosure, there is provided a computer-readable data carrier having stored thereon the computer program product as described herein.

[0572] In another aspect of the present disclosure, there is provided a data carrier signal carrying the computer program product as described herein.

[0573] The various illustrative imaging or data processing techniques described in connection with the embodiments disclosed herein can be implemented as electronic hardware, computer software, or combinations of both. To illustrate this interchangeability of hardware and software, various illustrative components, blocks, modules, and steps have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends upon the particular application and design constraints imposed on the overall system. The described functionality can be implemented in varying ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of the disclosure.

[0574] The various illustrative detection systems described in connection with the embodiments disclosed herein can be implemented or performed by a machine, such as a processor configured with specific instructions, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. A processor can be a microprocessor, but in the alternative, the processor can be a controller, microcontroller, or state machine, combinations of the same, or the like. A processor can also be implemented as a combination of computing devices, e.g., a combination of a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration. For example, systems described herein may be implemented using a discrete memory chip, a portion of memory in a microprocessor, flash, EPROM, or other types of memory.

[0575] The elements of a method, process, or algorithm described in connection with the embodiments disclosed herein can be embodied directly in hardware, in a software module executed by a processor, or in a combination of the two. A software module can reside in RAM memory, flash memory, ROM memory, EPROM memory, EEPROM memory, registers, hard disk, a removable disk, a CD-ROM, or any other form of computer-readable storage medium known in the art. An exemplary storage medium can be coupled to the processor such that the processor can read information from, and write information to, the storage medium. In the alternative, the storage medium can be integral to the processor. The processor and the storage medium can reside in an ASIC. A software module can comprise computer-executable instructions which cause a hardware processor to execute the computer-executable instructions.

[0576] Computer-executable instructions may be stored in a (transitory or non-transitory) computer readable storage medium (e.g., memory, storage system, etc.) storing code, or computer readable instructions.

[0577] Additional Notes

[0578] The embodiments described herein are exemplary. Modifications, rearrangements, substitute processes, etc. may be made to these embodiments and still be encompassed within the teachings set forth herein. One or more of the steps, processes, or methods described herein may be carried out by one or more processing and / or digital devices, suitably programmed.

[0579] Conditional language used herein, such as, among others, “can,” “might,” “may,” “e.g.,” and the like, unless specifically stated otherwise, or otherwise understood within the context as used, is generally intended to convey that certain embodiments include, while other embodiments do not include, certain features, elements and / or states. Thus, such conditional language is not generally intended to imply that features, elements and / or states are in any way required for one or more embodiments or that one or more embodiments necessarily include logic for deciding, with or without author input or prompting, whether these features, elements and / or states are included or are to be performed in any particular embodiment. The terms “comprising,” “including,” “having,” “involving,” and the like are synonymous and are used inclusively, in an open-ended fashion, and do not exclude additional elements, features, acts, operations, and so forth. Also, the term “or” is used in its inclusive sense (and not in its exclusive sense) so that when used, for example, to connect a list of elements, the term “or” means one, some, or all of the elements in the list. The term “comprising” may be considered to encompass “consisting”.

[0580] Disjunctive language such as the phrase “at least one of X, Y orZ,” unless specifically stated otherwise, is otherwise understood with the context as used in general to present that an item, term, etc., may be either X, Y or Z, or any combination thereof (e.g., X, Y and / or Z). Thus, such disjunctive language is not generally intended to, and should not, imply that certain embodiments require at least one of X, at least one of Y or at least one of Z to each be present.

[0581] The terms “about” or “approximate” and the like are synonymous and are used to indicate that the value modified by the term has an understood range associated with it, where the range can be ±20%, ±15%, ±10%, ±5%, or±1%. The term “substantially” is used to indicate that a result (e.g., measurement value) is close to a targeted value, where close can mean, for example, the result is within 80% of the value, within 90% of the value, within 95% of the value, or within 99% of the value. The term “partially” is used to indicate that an effect is only in part or to a limited extent.

[0582] Unless otherwise explicitly stated, articles such as “a” or “an” should generally be interpreted to include one or more described items. Accordingly, phrases such as “a device configured to” or “a device to” are intended to include one or more recited devices. Such one or more recited devices can also be collectively configured to carry out the stated recitations. For example, “a processor to carry out recitations A, B and C” can include a first processor configured to carry out recitation A working in conjunction with a second processor configured to carry out recitations B and C.

[0583] While the above detailed description has shown, described, and pointed out novel features as applied to illustrative embodiments, it will be understood that various omissions, substitutions, and changes in the form and details of the devices or algorithms illustrated can be made without departing from the spirit of the disclosure. As will be recognized, certain embodiments described herein can be embodied within a form that does not provide all of the features and benefits set forth herein, as some features can be used or practiced separately from others. All changes which come within the meaning and range of equivalency of the claims are to be embraced within their scope.

[0584] It should be appreciated that all combinations of the foregoing concepts (provided such concepts are not mutually inconsistent) are contemplated as being part of the inventive subject matter disclosed herein. In particular, all combinations of claimed subject matter appearing at the end of this disclosure are contemplated as being part of the inventive subject matter disclosed herein.

[0585] SEQUENCE LISTING SEQ ID NO: 1: P5 sequence- AATGATACGGCGACCACCGAGATCTACAC SEQ ID NO: 2: P7 sequence CAAGCAGAAGACGGCATACGAGAT

[0586] SEQ ID NO: 3: P5 #1 Sequence AATGATACGGCGACCACCGAGAnCTACAC

[0587] where “n” is uracil in the sequence

[0588] SEQ ID NO: 4: P5 #2 sequence AATGATACGGCGACCACCGAGAnCTACAC

[0589] where “n” is alkene-thymidine (i.e., alkene-dT) in the sequence

[0590] SEQ ID NO: 5: P7 #1 sequence CAAGCAGAAGACGGCATACGAnAT

[0591] where “n” is 8-oxoguanine in the sequence

[0592] SEQ ID NO: 6: P7 #2 sequence

[0593] CAAGCAGAAGACGGC AT AC nAGAT

[0594] where “n” is 8-oxoguanine in the sequence

[0595] SEQ ID NO: 7: P15 sequence

[0596] AATGATACGGCGACCACCGAGAnCTACAC

[0597] where “n” is allyl-T

[0598] SEQ ID NO: 8: PA sequence GCTGGCACGTCCGAACGCTTCGTTAATCCGTTGAG SEQ ID NO: 9: PB sequence CGTCGTCTGCCATGGCGCTTCGGTGGATATGAACT SEQ ID NO: 10: PC sequence ACGGCCGCTAATATCAACGCGTCGAATCCGCAACT SEQ ID NO: 11: PD SequenceGCCGCGTTACGTTAGCCGGACTATTCGATGCAGC SEQ ID NO: 12: P5’ sequence (complementary to P5) GTGTAGATCTCGGTGGTCGCCGTATCATT SEQ ID NO: 13: P7’ sequence (complementary to P7) ATCTCGTATGCCGTCTTCTGCTTG SEQ ID NO: 14: Alternative P5 sequence AATGATACGGCGACCGA

[0599] SEQ ID NO: 15: Alternative P5’ sequence (complementary to alternative P5 sequence)

[0600] TCGGTCGCCGTATCATT SEQ ID NO: 16: UniProt P31941

[0601] MEASPASGPRHLMDPHI FTSNFNNGIGRHKTYLCYEVERLDNGTSVKMDQHRGFLHNQAKNL LCGFYGRHAELRFLDLVPSLQLDPAQIYRVTWFISWSPCFSWGCAGEVRAFLQENTHVRLRIFAARIY DYDPLYKEALQMLRDAGAQVSIMTYDEFKHCWDTFVDHQGCPFQPWDGLDEHSQALSGRLRAILQNQG N

Claims

CLAIMS:

1. A method of introducing a strand identifier to distinguish the descendants of a forward strand from the descendants of a reverse strand of a double-stranded target polynucleotide, the method comprising:providing one or more first transposome complexes, wherein the first transposome complex comprises a first adaptor, one or more second transposome complexes, wherein the second transposome complex comprises a second adaptor, and a double-stranded target polynucleotide comprising a forward target strand and a reverse target strand;ligating the first adaptor to a 5’ end of the reverse target strand using the first transposome complex and ligating the second adaptor to a 5’ end of the forward target strand using the second transposome complex, to form a partially adapted double-stranded target polynucleotide;removing a transposase enzyme from each of the first and second transposome complexes; andconducting an extension reaction of the partially adapted double-stranded template polynucleotide to form a fully adapted double-stranded target polynucleotide, wherein the extension reaction leads to the incorporation of a third adaptor at the 3’ end of the forward target strand and a fourth adaptor at the 3’ end of the reverse target strand, wherein the third adaptor comprises a forward strand descendant identifier and the fourth adaptor comprises a reverse strand descendant identifier.

2. The method according to claim 1, wherein the method further comprises providing a solid support comprising first immobilised primers and second immobilised primers.

3. The method according to claim 2, wherein the one or more first transposome complexes are immobilised at a 5’ end to the solid support, and the one or more second transposome complexes immobilised at a 3’ end or a 5’ end to the solid support.

4. The method according to any one of claims 1 to 3, wherein the first transposome complex comprises a transferred strand and a non-transferred strand, wherein the transferred strand comprises the first adaptor.

5. The method according to any one of claims 1 to 4, wherein the second transposome complex comprises a transferred strand and a non-transferred strand, wherein the transferred strand comprises the second adaptor.

6. The method according to claim 4 or claim 5, wherein the first transposome complex is immobilised to the solid support through the 5’ end of the transferred strand and the second transposome complex is immobilised to the solid support through the 5’ end of the transferred strand.

7. The method according to claim 4 or claim 5, wherein the first transposome complex is immobilised to the solid support through the 5’ end of the transferred strand and the second transposome complex is optionally immobilised to the solid support through 3’ end of the non-transferred strand.

8. The method according to any one of claims 1 to 7, wherein the first adaptor comprises at least a first amplification domain, a reverse template strand identifier and a transposase recognition sequence.

9. The method according to any one of claims 1 to 8, wherein the second adaptor comprises at least a second amplification domain, a forward template strand identifier and a transposase recognition sequence.

10. The method according to any one of claims 4 to 9, wherein the method comprises removing the non-transferred strand from the first transposase complex and removing the non-transferred strand from the second transposase complex before conducting an extension reaction.

11. The method according to any one of claims 1 to 10, wherein the third adaptor is covalently attached to a 3’ end of the forward strand and a fourth adaptor is covalently attached to a 3’ end of the reverse strand using a non-strand displacing polymerase and a ligase.

12. The method according to any one of claims 1 to 11, wherein the third adaptor comprises a transposase recognition sequence, a forward strand descendant identifier and at least one amplification domain.

13. The method according to any one of claims 1 to 12, wherein the fourth adaptor comprises a transposase recognition sequence, a reverse strand descendant identifier and at least one amplification domain.

14. The method according to any one of claims 1 to 13, wherein the forward strand descendant identifier is non-complementary to the reverse template strand identifier.

15. The method according to any one of claims 1 to 14, wherein the reverse strand descendant identifier is non-complementary to the forward template strand identifier.

16. The method according to any one of claims 1 to 15, wherein the forward strand descendant identifier differs from the complement of the forward template strand identifier by one or more bases.

17. The method according to any one of claims 1 to 16, wherein the reverse strand descendant identifier differs from the complement of the reverse template strand identifier by one or more bases.

18. The method according to any one of claims 1 to 12, wherein the reverse strand descendant identifier and / or the forward strand descendant identifier comprise at least onemodified base, wherein the modified base can be converted to a different base following the extension reaction.

19. The method according to any one of claims 1 to 12, wherein the reverse strand identifier and / or the forward strand template identifier comprise at least one modified base.

20. The method according to claim 19, wherein the modified base is 8-oxo-guanine.

21. The method according to claim 19, wherein the modified base is a methylated cytosine, wherein the method further comprises the step of applying a conversion agent following the extension reaction wherein the conversion agent is configured to convert the methylated cytosine to thymine or a nucleobase that is read as thymine / uracil.

22. The method according to any one of claims 4 to 21, wherein the non-transferred strand of the first transposome complex comprises the fourth adaptor, wherein the fourth adaptor comprises a transposase recognition sequence, a reverse strand descendant identifier and at least one amplification domain, wherein the amplification domain is a complement of the first amplification domain or a fourth amplification domain.

23. The method according to any one of claims 4 to 22, wherein the non-transferred strand of the second transposome complex comprises the third adaptor, wherein the third adaptor comprises a transposase recognition sequence, a forward strand descendant identifier and at least one amplification domain, wherein the amplification domain is a complement of the second amplification domain or a third amplification domain.

24. The method according to any one of claims 1 to 23, wherein the method further comprises denaturing the forward strand of the fully adapted target polynucleotide from the reverse strand of the fully adapted target polynucleotide.

25. The method according to any one of claims 1 to 24, wherein the method further comprises a step of amplifying the forward and reverse strand of the fully adapted target polynucleotide.

26. The method according to claim 25, wherein the forward strand and amplified descendants thereof, and reverse strand and amplified descendants thereof, together form a cluster on the solid support.

27. The method according to claim 25 or claim 26, wherein the method comprises a step of preparing the forward strand and amplified descendants thereof, and reverse strand and amplified descendants thereof for sequencing, wherein the method comprises removing the forward strand and amplified reverse complement descendants thereof, or removing the reverse strand and forward complement amplified descendants thereof.

28. The method according to any one of claims 2 to 27, wherein the solid support is a flow cell, optionally a non-patterned flow cell.

29. A method of distinguishing the descendants of a forward target strand from the descendants of a reverse target strand of a double-stranded target polynucleotide, the method comprising:introducing at least one strand identifier into the descendants of a forward target strand and / or introducing at least one strand identifier into the descendants of a reverse target strand of a double-stranded target polynucleotide using the method according to any one of claims 1 to 28; andsequencing nucleobases in the amplified forward strands and the amplified reverse strands, and determining a difference in sequence output within the amplified forward strand read and / or a difference in sequence output within the amplified reverse strand read; and determining a proportion of sequences derived from the forward target strand and the proportion of sequences derived from the reverse target strand based on the difference in sequence output within the amplified forward strand read and / or the difference in sequence output within the amplified reverse strand read.

30. A method of detecting an error in nucleic acid amplification of a target polynucleotide, the method comprising introducing at least one strand identifier into the descendants of a forward target strand and / or introducing at least one strand identifier into the descendants of a reverse target strand of a double-stranded target polynucleotide using the method according to any one of claims 1 to 28; andsequencing nucleobases in the amplified forward strands and the amplified reverse strands, and determining if there is a difference in sequence output within the amplified forward strand read and / or a difference in sequence output within the amplified reverse strand read; wherein the presence of a difference is indicative of an error during amplification of a target polynucleotide.

31. The method according to claim 30 or claim 31, wherein the method comprises conducting a first sequence read using a first sequencing primer that substantially hybridises to a 3’ adaptor of the amplified strand comprising the complement of the forward template strand identifier and conducting a second sequencing read using a second sequencing primer that substantially hybridises to a 3’ adaptor of the amplified strand comprising the complement of the reverse strand descendant identifier and determining a difference in sequence output.

32. The method according to claim 31, wherein the method comprises conducting a third sequence read using a first sequencing primer that substantially hybridises to a 3’ adaptor of the amplified strand comprising the complement of the reverse template strand identifier and conducting a fourth sequencing read using a fourth sequencing primer that substantially hybridises to a 3’ adaptor of the amplified strand comprising the complement of the forward strand descendant identifier and determining a difference in sequence output.

33. The method according to claim 31 or claim 32, wherein the first and second sequence reads and / or the third and fourth sequence reads are conducted sequentially or concurrently.

34. The method according to claim 33, wherein the method comprises conducting paired-end re-synthesis before conducting the third and fourth sequence read.

35. The method according to claim 29 or claim 30, wherein the method comprises: conducting a first sequence read using a first sequencing primer wherein the first sequencing primer substantially hybridises to a 3’ adaptor on a forward strand, sequencing nucleobases in the amplified forward strands, anddetermining a difference in sequence, wherein a forward strand derived from the forward target strand will have a different sequence to a reverse complement strand derived from the reverse target strand.

36. The method according to claim 29 or claim 30, wherein the method comprises: conducting a second sequence read using a second sequencing primer wherein the second sequencing primer substantially hybridises to a 3’ adaptor on a reverse strand, sequencing nucleobases in the amplified reverse strands, anddetermining a difference in sequence, wherein a reverse strand derived from the reverse target strand will have a different sequence to a forward complement strand derived from the forward target strand.

37. The method according to claim 35 or claim 36, wherein the first and second sequence reads are conducted sequentially or concurrently.

38. The method according to claim 37, wherein the method comprises conducting paired-end re-synthesis before conducting the second sequence read.

39. A method of preparing polynucleotide sequences for sequencing, comprising: providing a solid support comprising first immobilised primers and second immobilised primers, and a double-stranded target polynucleotide comprising a forward target strand and a reverse target strand,seeding both the forward target strand and the reverse target strand onto the solid support,amplifying the forward target strand and the reverse target strand to produce amplified forward strands and amplified reverse complement strands each covalently attached to a first immobilised primer, and amplified reverse strands and amplified forward complement strands each covalently attached to a second immobilised primer, andpreparing the amplified forward strands and amplified reverse complement strands for concurrent sequencing, or preparing the amplified reverse strands and amplified forward complement strands for concurrent sequencing.

40. The method according to claim 39, wherein the step of seeding both the forward target strand and the reverse target strand onto the solid support is conducted using a tagmentation reaction.

41. The method according to claim 39 or claim 40, wherein the step of seeding both the forward target strand and the reverse target strand onto the solid support is conducted using a transposase.

42. The method according to any one of claims 39 to 41, wherein both the forward target strand and the reverse target strand are seeded into a same well of the solid support.

43. The method according to any one of claims 39 to 42, wherein a 5’-end of the forward target strand is covalently attached to a 3’-end of a first transferred strand.

44. The method according to claim 43, wherein a 5’-end of the first transferred strand is covalently attached to a 3’-end of a first amplification domain.

45. The method according to any one of claims 39 to 44, wherein the first transferred strand is immobilised onto the solid support.

46. The method according to any one of claims 39 to 45, wherein a 5’-end of the reverse target strand is covalently attached to a 3’-end of a second transferred strand.

47. The method according to claim 46, wherein a 5’-end of the second transferred strand is covalently attached to a 3’-end of a second amplification domain.

48. The method according to any one of claims 39 to 47, wherein the second transferred strand is immobilised onto the solid support.

49. The method according to claim 48, wherein the step of amplifying the forward target strand and the reverse target strand is conducted using bridge amplification and / or exclusion amplification.

50. The method according to any one of claims 39 to 47, wherein the second nontransferred strand is immobilised onto the solid support.

51. The method according to claim 50, wherein the step of amplifying the forward target strand and the reverse target strand is conducted using exclusion amplification.

52. The method according to any one of claims 39 to 51, wherein the step of preparing the amplified forward strands and amplified reverse complement strands for concurrent sequencing comprises simultaneously contacting first sequencing primer binding sites located after a 3’-end of the amplified forward strands with first primers and second sequencing primer binding sites located after a 3’-end of the amplified reverse complement strands with second primers; or wherein the step of preparing the amplified reverse strands and amplified forward complement strands for concurrent sequencing comprises simultaneously contacting third sequencing primer binding sites located after a 3’-end of theamplified reverse strands with third primers and fourth sequencing primer binding sites located after a 3’-end of the amplified forward complement strands with fourth primers.

53. The method according to any one of claims 39 to 51, wherein the step of preparing the amplified forward strands and amplified reverse complement strands for concurrent sequencing comprises simultaneously nicking amplified forward complement strands hybridised to the amplified forward strands and nicking amplified reverse strands hybridised to the amplified reverse complement strands; or wherein the step of preparing the amplified reverse strands and amplified forward complement strands for concurrent sequencing comprises simultaneously nicking amplified reverse complement strands hybridised to the amplified reverse strands and nicking amplified forward strands hybridised to the amplified forward complement strands.

54. The method according to any one of claims 39 to 53, further comprising a step of treating the forward target strand and the reverse target strand with a conversion agent, prior to the step of amplifying the forward target strand and the reverse target strand, wherein the conversion reagent is configured to convert a modified cytosine to thymine or a nucleobase which is read as thymine / uracil, and / or wherein the conversion reagent is configured to convert an unmodified cytosine to uracil or a nucleobase which is read as thymine / uracil.

55. The method according to claim 54, wherein the step of treating the forward target strand and the reverse target strand with a conversion agent is conducted after the step of seeding both the forward target strand and the reverse target strand onto the solid support.

56. The method according to claim 54 or claim 55, wherein the conversion agent comprises an enzyme.

57. The method according to claim 56, wherein the enzyme is a cytidine deaminase.

58. The method according to claim 57, wherein the cytidine deaminase is a wildtype cytidine deaminase or a mutant cytidine deaminase; optionally a mutant cytidine deaminase.

59. The method according to any one of claims 39 to 58, wherein the amplified forward strands and amplified reverse complement strands, and the amplified reverse strands and amplified forward complement strands, together form a cluster on the solid support.

60. The method according to any one of claims 39 to 59, wherein the method comprises a step of preparing the amplified forward strands and amplified reverse complement strands for concurrent sequencing, and wherein the amplified reverse strands and the amplified forward complement strands are removed prior to concurrent sequencing.

61. The method according to claim 60, wherein the amplified forward strands and amplified reverse complement strands form a duoclonal cluster on the solid support.

62. The method according to any one of claims 39 to 59, wherein the method comprises preparing the amplified reverse strands and amplified forward complement strands for concurrent sequencing, and wherein the amplified forward strands and the amplified reverse complement strands are removed prior to concurrent sequencing.

63. The method according to claim 62, wherein the amplified reverse strands and amplified forward complement strands form a duoclonal cluster on the solid support.

64. The method according to any one of claims 39 to 63, wherein the population of amplified forward strands prepared for concurrent sequencing is substantially equal to the population of amplified reverse complement strands prepared for concurrent sequencing; and / or wherein the population of amplified reverse strands prepared for concurrent sequencing is substantially equal to the population of amplified forward complement strands prepared for concurrent sequencing.

65. The method according to any one of claims 39 to 64, wherein the solid support is a flow cell; optionally a non-patterned flow cell.

66. A method of sequencing polynucleotide sequences, comprising: preparing polynucleotides for sequencing using a method according to any one of claims 39 to 65, andconcurrently sequencing nucleobases in the amplified forward strands and amplified reverse complement strands, or concurrently sequencing the amplified reverse strands and amplified forward complement strands.

67. The method according to claim 66, wherein the method further comprises a step of identifying differences when comparing sequence outputs from the amplified forward strands and amplified reverse complement strands, or when comparing sequence outputs from the amplified reverse strands and amplified forward complement strands.

68. The method according to claim 67, wherein the step of identifying differences comprises associating the differences with the presence of errors.

69. The method according to any one of claims 66 to 68, wherein the step of preparing polynucleotides for sequencing is conducted using a method according to any one of claims 54 to 58, or any one of claims 59 to 65 as dependent on any one of claims 54 to 58, and the step of identifying differences comprises associating the differences with the presence of modified cytosines.

70. The method according to any one of claims 66 to 69, wherein the method further comprises a step of conducting paired-end reads.

71. A kit comprising instructions for preparing polynucleotide sequences for sequencing according to any one of claims 39 to 65, and / or for sequencing polynucleotide sequences according to any one of claims 66 to 70.

72. A data processing device comprising means for carrying out a method according to any one of claims 39 to 70.

73. The data processing device according to claim 72, wherein the data processing device is a polynucleotide sequencer.

74. A computer program product comprising instructions which, when the program is executed by a processor, cause the processor to carry out a method according to any one of claims 39 to 70.

75. A computer-readable storage medium comprising instructions which, when executed by a processor, cause the processor to carry out a method according to any one of claims 39 to 70.

76. A computer-readable data carrier having stored thereon a computer program product according to claim 74.

77. A data carrier signal carrying a computer program product according to claim 74.

Citation Information

Patent Citations

  • Methods for producing a paired tag from a nucleic acid sequence and methods of use thereof

    US20060024681A1

  • Paired end sequencing

    US20060292611A1

  • Patterned flow-cells useful for nucleic acid analysis

    US20120316086A1

  • Methods and compositions for nucleic acid sequencing

    US20130079232A1

  • DNA sequencing by parallel oligonucleotide extensions

    US6306597B1