Tagging target regions prior to nucleotide sequencing

By modifying the sequence of oligonucleotides with a bioinformatically identifiable tag on a solid surface, the method addresses SSEs in SBS, enhancing sequencing resolution and accuracy in regions prone to errors.

WO2026006314A9PCT designated stage Publication Date: 2026-02-05ILLUMINA INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/035050
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-06-28
Filing Date
2025-06-24
Publication Date
2026-02-05

AI Technical Summary

Technical Problem

Certain motifs in oligonucleotides cause sequence specific errors (SSEs) during sequencing-by-synthesis (SBS) methodologies, leading to poor sequencing resolution in regions prone to SSEs and downstream regions, which may include clinically important coding regions.

Method used

Modify the sequence of target oligonucleotides on a solid surface by introducing a bioinformatically identifiable sequence tag upstream of the region of interest, using a tagging oligonucleotide to disrupt SSE regions and facilitate monoclonal cluster formation, and verify proper sequencing through the tag.

Benefits of technology

Improves sequencing resolution downstream of the target region and corroborates the authenticity of sequences in regions of interest by disrupting SSEs and ensuring accurate determination.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025035050_05022026_PF_FP_ABST
    Figure US2025035050_05022026_PF_FP_ABST
Patent Text Reader

Abstract

Methods for preparing a solid surface for oligonucleotide sequencing including modifying a sequence of an oligonucleotide to be sequenced on a solid surface, such as on a flow of a sequencer. The modification may disrupt a sequence in a region known to produce a sequence specific error or may provide a bioinformatic tag upstream of a region of interest. The modification may be performed by hybridizing a template oligonucleotide to a surface primer, hybridizing a tagging oligonucleotide to a target region of the target oligonucleotide, where the tagging oligonucleotide includes a sequence different from the sequence of the target region, and extending the surface primer from a 3 ' end primer using a target oligonucleotide as a template, and extending the tagging oligonucleotide from a 3 ' end using the target oligonucleotide to produce the copy oligonucleotide having a sequence complementary to the target oligonucleotide except for in the tag region.
Need to check novelty before this filing date? Find Prior Art

Description

IP-2758-PCT PATENTTAGGING TARGET REGIONS PRIOR TO NUCLEOTIDE SEQUENCINGFIELD

[0001] The present disclosure relates to, among other things, oligonucleotide sequencing and / or preparing a solid surface for oligonucleotide sequencing. More particularly, the present disclosure relates to modification of a sequence of an oligonucleotide to be sequenced while the oligonucleotide is on the solid surface prior to cluster generation.INTRODUCTION

[0002] Certain motifs in oligonucleotides present challenges to sequencing of the oligonucleotides when using sequencing-by-synthesis (SBS) methodologies. Such motifs may result in “sequence specific errors” (SSEs) that result in poor sequencing resolution of the region prone to SSEs and to regions downstream of the region prone to SSEs. In many circumstances, regions downstream of the regions prone to SSEs are salient coding regions in the human genome, and reliably determining the correct sequence of such downstream regions may be clinically important.

[0003] Regardless of whether clinically important oligonucleotide regions or other regions of interest are downstream of regions prone to SSEs, correct determination of the sequences of these regions is often desired.SUMMARY

[0004] The present disclosure describes, among other things, methods for modifying the sequence of a target region of a target oligonucleotide on a solid surface prior to cluster generation in an oligonucleotide sequencing workflow. Such sequence modification may improve sequencing resolution downstream of the target region or may corroborate the authenticity of the determined sequence of a region of interest downstream of the target region.

[0005] The target region, for example, may be within a region prone to sequence specific errors (SSEs). Modification of the sequence of a region prone to SSEs may interrupt or disrupt the SSE region, which may facilitate monoclonal cluster formation and / or downstream sequencing.

[0006] The target region, for example, may be a region upstream of a clinically relevant region or other region of interest. Modification of the sequence upstream of a region of interest may include inserting a bioinformatically identifiable sequence tag upstream of the region of interest. Verification of proper sequencing of the tag may provide assurances that a sequence downstream of the tag, including the region of interest, is properly determined.

[0007] In various aspects, the present disclosure describes an oligonucleotide sequencing method or method for preparing a solid surface for oligonucleotide sequencing. The method comprises hybridizing a target oligonucleotide to a surface primer bound to a solid surface. The target oligonucleotide has a target region. The surface primer has a free 3’ end. The method further comprises hybridizing a tagging oligonucleotide to the target region of the target oligonucleotide hybridized to the surface primer. The tagging oligonucleotide comprises a sequence that differs from the complementary sequence of the target region. The tagging oligonucleotide comprises a tag region between a 3’ region and a 5’ region. The 3’ region has a sequence complementary to a sequence of a 5 ’ portion of the target region, and the 5 ’ region has a sequence complementary to a sequence of a 3 ’ portion of the target region. The method further comprises synthesizing a copy oligonucleotide by (i) extending the surface primer from the 3 ’ end and using the target oligonucleotide as a template, (ii) incorporating the tagging oligonucleotide into the copy oligonucleotide, and (iii) extending the tagging oligonucleotide from a 3’ end using the target oligonucleotide to produce the copy oligonucleotide having a sequence complementary to the target oligonucleotide except for in the tag region.

[0008] In some embodiments, the 3’ end of the surface primer is extended to a location at which the 5 ’ end of the tagging oligonucleotide hybridizes to the target region to produce a first nascent strand, and the 3’ end of the nascent first strand is ligated to the 5’ end of the tagging oligonucleotide.

[0009] In some embodiments, a blocking oligonucleotide is hybridized to the target region of the target oligonucleotide hybridized to the surface primer. The blockingoligonucleotide comprises a sequence complementary to the sequence of the target region. The blocking oligonucleotide comprises a 5 ’ region having a sequence identical to the 5 ’ region of the tagging oligonucleotide. The 3’ end of the surface primer is extended to the location at which the 5’ end of the tagging oligonucleotide hybridizes to the target region while the blocking oligonucleotide is hybridized to target region of the target oligonucleotide. The blocking oligonucleotide is removed from the target region of the target oligonucleotide prior to hybridizing the tagging oligonucleotide to the target region of the target oligonucleotide. Extending the tagging oligonucleotide from the 3’ end may occur after ligating the 3’ end of the first nascent strand to the 5’ end of the tagging oligonucleotide. The method may further comprises amplifying the copy oligonucleotide to form a cluster of oligonucleotides on the solid surface. The cluster of oligonucleotides comprises multiple copies of the copy oligonucleotide bound to the solid surface and multiple copies of a modified target oligonucleotide bound to the solid surface. The modified oligonucleotide is complementary to the copy oligonucleotide. The method may further comprise removing from the solid surface the multiple copies of the copy oligonucleotide or the multiple copies of the modified target oligonucleotide. The method may further comprises sequencing the multiple copies of the copy oligonucleotide or the multiple copies of the modified target oligonucleotide that remain bound to the solid surface.

[0010] In some embodiments, the one or more target regions comprise nucleotides within regions prone to produce sequence specific errors (SSEs). In some embodiments, the region prone to produce SSEs is a homopolymer region, such as a poly(T) or a poly(A) sequence. In some embodiments, the region prone to SSEs comprises a multi-nucleotide repeat region. The multi-nucleotide repeat region may comprise a di-nucleotide repeat region or a tri-nucleotide repeat region. The multi-nucleotide repeat region may comprise GT, GA, AC, or CT repeats. In some embodiments, the region prone to SSEs comprises a G- quadruplex region.

[0011] In some embodiments, the method further comprises amplifying the copy oligonucleotide to form a cluster of oligonucleotides on the solid surface. The cluster of oligonucleotides comprises multiple copies of the copy oligonucleotide bound to the solid surface and multiple copies of a modified target oligonucleotide bound to the solid surface. The modified oligonucleotide is complementary to the copy oligonucleotide. The method may further comprise removing from the solid surface the multiple copies of the copyoligonucleotide or the multiple copies of the modified target oligonucleotide. The method may further comprise sequencing the multiple copies of the copy oligonucleotide or the multiple copies of the modified target oligonucleotide that remain bound to the solid surface. The target region may be a region prone to sequence specific errors. Sequencing resolution downstream of the tag or modified target region may be improved.

[0012] The details of one or more embodiments are set forth in the accompanying drawings and the description below. Other features, objects, and advantages will be apparent from the description and drawings, and from the claims.

[0013] It is to be understood that both the foregoing general description and the following detailed description present embodiments of the subject matter of the present disclosure and are intended to provide an overview or framework for understanding the nature and character of the subject matter as it is claimed. The accompanying drawings are included to provide a further understanding of the subject matter and are incorporated into and constitute a part of this specification. The drawings illustrate various embodiments of the subject matter and together with the description serve to explain the principles and operations of the subject matter of the present disclosure. Additionally, the drawings and descriptions are meant to be merely illustrative and are not intended to limit the scope of the claims in any manner.

[0014] The above summary is not intended to describe each disclosed embodiment or every implementation of the present disclosure. The description that follows more particularly exemplifies illustrative embodiments. In several places throughout the disclosure, guidance is provided through lists of examples, which examples can be used in various combinations. In each instance, the recited list serves only as a representative group and should not be interpreted as an exclusive or exhaustive list.BRIEF DESCRIPTION OF THE FIGURES

[0015] The following detailed description of illustrative embodiments of the present disclosure may be best understood when read in conjunction with the following drawings.

[0016] FIG. 1 is a flow diagram illustrating a surface preparation and sequencing method consistent with some embodiments of the present disclosure.

[0017] FIG. 2 is a flow diagram illustrating a first strand synthesis method, which may be a part of a surface preparation and / or sequencing method, consistent with some embodiments of the present disclosure.

[0018] FIG. 3 is a schematic drawing illustrating a first strand synthesis method, which may be a part of a surface preparation and / or sequencing method, consistent with some embodiments of the present disclosure.

[0019] FIG. 4 is a schematic drawing illustrating a first strand synthesis method, which may be a part of a surface preparation and / or sequencing method, consistent with some embodiments of the present disclosure.

[0020] FIG. 5 is a schematic plot illustrating the relationship between homopolymer length and the percentage of reads affected during SBS.

[0021] FIG. 6 is a schematic drawing illustrating how the use of a tagging oligonucleotide during first strand synthesis consistent with some embodiments of the present disclosure can result in a modified target oligonucleotide following clustering.

[0022] FIG. 7 is a schematic drawing illustrating the structure of a G-quadruplex with four guanine nucleobases and a central cation (C+). The Watson-Crick and Hoogsteen hydrogen bonds between positions 6 and 1 and between positions 2 and 7 of adjacent nucleobases are shown by the dashed lines.

[0023] FIG. 8 is a schematic drawing illustrating a G-quadruplex region blocking an SBS polymerase and resolving of the G-quadruplex structure consistent with some embodiments of the present disclosure allows the SBS polymerase to read through the region that previously would have formed the g-quadruplex structure.

[0024] FIG. 9 is a plot showing sequencing resolution value of target oligonucleotides having homopolymers of different lengths that were (i) hybridized to primers on a flow cell and sequenced (Primer hyb), or (ii) hybridized to primers on a flow cell, clustered via amplification, and then sequenced (Clustered). Additional detail is provided in Example 1.

[0025] FIG. 10 is plot showing sequencing resolution value of target oligonucleotides having g-quad regions or polyT homopolymer regions that were uninterrupted or interrupted by substitutions or insertions. The target oligonucleotides were prepared, seeded, clustered, andsequenced using an Illumina, Inc. cBot sequencing apparatus. Resolution was determined within the g-quad region or the cycle following the homopolymer. Additional detail is provided in Example 2.

[0026] FIG. 11 is a schematic drawing illustrating a sequencing workflow for a method of using tagging oligonucleotides described herein. Additional detail is provided in Example 3.

[0027] FIG. 12 are schematic drawings illustrating of tagging oligonucleotides used in methods described herein. Additional detail is provided in Example 3.

[0028] FIG. 13 are plots of sequencing data illustrating proof of concept using a BacPack library with tagging oligonucleotides including a 1-base substitution and tagging oligonucleotides including a 5-base insertion. Additional detail is provided in Example 3.

[0029] FIG. 14 are plots of sequencing data on a BacPack library with a 29mer tagging oligonucleotide containing no flanks and a 1-base substitution. Additional detail is provided in Example 3.

[0030] FIG. 15A-B are plots showing show the effects of tagging oligonucleotides on sequencing results for homopolymers of different lengths. FIG. ISA shows percentage mismatch. FIG. 15B shows percentage Q30. Additional detail is provided in Example 3.

[0031] FIG. 16 are plots of sequencing data showing the efficiency of the tagging technology as a proportion of overlapping reads that include the tag. Additional detail is provided in Example 3.

[0032] FIG. 17 are plots of sequencing data showing the ability of IGV to resolve different tagging oligonucleotide designs. Additional detail is provided in Example 3.

[0033] FIG. 18 are plots of sequencing data showing the proof of concept of tagging oligonucleotides on sequencing data for a human library. Additional detail is provided in Example 4.

[0034] FIG. 19 are plots of sequencing data showing the effect of tagging oligonucleotide concentration on sequencing efficiency for a human library. Additional detail is provided in Example 4.

[0035] FIG. 20 is a plot showing the impact of tagging oligonucleotides on percentage soft clipping as a function of homopolymer length. Additional detail is provided in Example 4.

[0036] FIG. 21 are plot showing the effects of A homopolymers and T homopolymers on percentage soft clipping. Additional detail is provided in Example 4.

[0037] The schematic drawings are not necessarily to scale. Like numbers used in the figures refer to like components, steps and the like. However, it will be understood that the use of a number to refer to a component in a given figure is not intended to limit the component in another figure labeled with the same number. In addition, the use of different numbers to refer to components is not intended to indicate that the different numbered components cannot be the same or similar to other numbered components.Certain Definitions

[0038] Terms used herein will be understood to take on their ordinary meaning in the relevant art unless specified otherwise. Several terms used herein and their meanings are set forth below.

[0039] As used herein, "GC-rich region" refers to a series of guanosine (G) nucleotides, cytosine (C) nucleotides, or both guanosine and cytosine nucleotides on a strand of a nucleic acid. A GC-rich region can be a series of G nucleotides (a G homopolymer), a series of C nucleotides, (a C homopolymer) or a combination of both G and C nucleotides. A GC-rich region can include G nucleotides that can form one or more G-quadruplex structures. C-rich region can include a GC content of at least 40%, at least 50%, at least 60%, and 70%, at least 80%, or at least 90% on a single strand over a defined length of nucleotides. GC content can be calculated as (number of G + C nucleotides ) / (number A + T + G + C nucleotides) * 100%. The defined length can be any number, such as 25, 50, 100, or 200 nucleotides.

[0040] As used herein, the term “repeat region” refers to sequence of a single repeating nucleotide or a pattern of two or more repeating nucleotides on a strand of a nucleic acid. A repeat region includes five or more nucleotides. Examples of repeat regions include, for example, homopolymers, dinucleotide repeats, trinucleotide repeats, and tetranucleotide repeats.

[0041] As used herein, the term “amplification site” refers to a site in or on an array where one or more amplicons can be generated. An amplification site can be further configured to contain, hold or attach at least one amplicon that is generated at the site.

[0042] As used herein, the term “array” refers to a population of sites that can be differentiated from each other according to relative location. Different molecules that are at different sites of an array can be differentiated from each other according to the locations of the sites in the array. An individual site of an array can include one or more molecules of a particular type. For example, a site can include a single target nucleic acid molecule having a particular sequence or a site can include several nucleic acid molecules having the same sequence (and / or complementary sequence, thereof). The sites of an array can be different features located on the same substrate. Exemplary features include without limitation, wells in a substrate, beads (or other particles) in or on a substrate, projections from a substrate, ridges on a substrate or channels in a substrate. The sites of an array can be separate substrates each bearing a different molecule. Different molecules attached to separate substrates can be identified according to the locations of the substrates on a surface to which the substrates are associated or according to the locations of the substrates in a liquid or gel. Exemplary arrays in which separate substrates are located on a surface include, without limitation, those having beads in wells.

[0043] As used herein, the term “amplicon,” when used in reference to a nucleic acid, means the product of copying the nucleic acid, where the product has a nucleotide sequence that is the same as or complementary to at least a portion of the nucleotide sequence of the nucleic acid. An amplicon can be produced by any of a variety of amplification methods that use the nucleic acid, e.g., a target nucleic acid or an amplicon thereof, as a template including, for example, polymerase extension, polymerase chain reaction (PCR), rolling circle amplification (RCA), ligation extension, or ligation chain reaction. An amplicon can be a nucleic acid molecule having a single copy of a particular nucleotide sequence (e.g., a polymerase extension product) or multiple copies of the nucleotide sequence (e.g., a concatemeric product of RCA). A first amplicon of a target nucleic acid is typically a complementary copy. Subsequent amplicons are copies that are created, after generation of the first amplicon, from the target nucleic acid or from the first amplicon. A subsequent amplicon can have a sequence that is substantially complementary to the target nucleic acid or substantially identical to the target nucleic acid.

[0044] As used herein, the term “capture agent” refers to a material, chemical, molecule, or moiety thereof that is capable of attaching, retaining, or binding to a target molecule (e.g., a target nucleic acid). Exemplary capture agents include, without limitation, a capture nucleic acid that is complementary to at least a portion of a modified target nucleic acid (e.g., auniversal capture binding sequence), a member of a receptor-ligand binding pair (e.g., avidin, streptavidin, biotin, lectin, carbohydrate, nucleic acid binding protein, epitope, antibody, etc.) capable of binding to a modified target nucleic acid (or linking moiety attached thereto), or a chemical reagent capable of forming a covalent bond with a modified target nucleic acid (or linking moiety attached thereto). In one embodiment, a capture agent is a nucleic acid. A nucleic acid capture agent can also be used as an amplification primer.

[0045] The terms “P5” and “P7” may be used when referring to a nucleic acid capture agent. The terms “P5”’ (P5 prime) and “P7”’ (P7 prime) refer to the complements of P5 and P7, respectively. It will be understood that any suitable nucleic acid capture agent can be used in the methods presented herein, and that the use of P5 and P7 are exemplary embodiments only. Uses of nucleic acid capture agents such as P5 and P7 on flow-cells is known in the art, as exemplified by the disclosures of WO 2007 / 010251, WO 2006 / 064199, WO 2005 / 065814, WO 2015 / 106941, WO 1998 / 044151, and WO 2000 / 018957. One of skill in the art will recognize that a nucleic acid capture agent can also function as an amplification primer. For example, any suitable nucleic acid capture agent can act as a forward amplification primer, whether immobilized or in solution, and can be useful in the methods presented herein for hybridization to a sequence (e.g., a universal capture binding sequence) and amplification of a sequence. Similarly, any suitable nucleic acid capture agent can act as a reverse amplification primer, whether immobilized or in solution, and can be useful in the methods presented herein for hybridization to a sequence (e.g., a universal capture binding sequence) and amplification of a sequence. In view of the general knowledge available and the teachings of the present disclosure, one of skill in the art will understand how to design and use sequences that are suitable for capture and amplification of target nucleic acids as presented herein.

[0046] As used herein, the term “polymerase” is intended to be consistent with its use in the art and includes, for example, an enzyme that produces a complementary replicate of a nucleic acid molecule using the nucleic acid as a template strand. Typically, DNA polymerases bind to the template strand and then move down the template strand sequentially adding nucleotides to the free hydroxyl group at the 3' end of a growing strand of nucleic acid. DNA polymerases typically synthesize complementary DNA molecules from DNA templates and RNA polymerases typically synthesize RNA molecules from DNA templates (transcription). Polymerases can use a short RNA or DNA strand, called a primer, to begin strand growth. Some polymerases can displace the strand upstream of the site where they are adding bases toa chain. Such polymerases are said to be strand displacing, meaning they have an activity that removes a complementary strand from a template strand being read by the polymerase. Exemplary polymerases having strand displacing activity include, without limitation, the large fragment of Bsu (Bacillus subtilis), Bst (Bacillus stearothermophilus) polymerase, exo-Klenow polymerase or sequencing grade T7 exo-polymerase. Some polymerases degrade the strand in front of them, effectively replacing it with the growing chain behind (5' exonuclease activity). Some polymerases have an activity that degrades the strand behind them (3' exonuclease activity). Some useful polymerases have been modified, either by mutation or otherwise, to reduce or eliminate 3' and / or 5' exonuclease activity. Different polymerases can be used at different times during the sequencing process, including library production (e.g., amplification or reverse transcription), production of clonal populations of amplicons at amplification sites (e.g., a polymerase for Exclusion Amplification or Bridge Amplification), or sequencing (e.g., a polymerase that can be used with 3'-blocked nucleotides).

[0047] As used herein, the terms “nucleic acid,” “polynucleotide,” and “oligonucleotide” are used interchangeably and are intended to be consistent with its use in the art and includes naturally occurring nucleic acids and functional analogs thereof. Particularly useful functional analogs are capable of hybridizing to a nucleic acid in a sequence specific fashion or capable of being used as a template for replication of a particular nucleotide sequence. Naturally occurring nucleic acids generally have a backbone containing phosphodiester bonds. An analog structure can have an alternate backbone linkage including any of a variety of those known in the art. Naturally occurring nucleic acids generally have a deoxyribose sugar (e.g., found in deoxyribonucleic acid (DNA)) or a ribose sugar (e.g., found in ribonucleic acid (RNA)). A nucleic acid can contain any of a variety of analogs of these sugar moieties that are known in the art. A nucleic acid can include native or non-native bases. In this regard, a native deoxyribonucleic acid can have one or more bases selected from adenine, thymine, cytosine or guanine and a ribonucleic acid can have one or more bases selected from uracil, adenine, cytosine or guanine. Useful non-native bases that can be included in a nucleic acid are known in the art. In some embodiments, non-native bases that can be included in a nucleic acid include a guanine modified as described herein (e.g., a guanine present in a dGTP analog). The term “target,” when used in reference to a nucleic acid, is intended as a semantic identifier for the nucleic acid in the context of a method or composition set forth herein and does not necessarily limit the structure or function of the nucleic acid beyond what is otherwise explicitly indicated.A target nucleic acid having a universal sequence at each end, for instance a universal adapter at each end, can be referred to as a modified target nucleic acid.

[0048] As used herein, “upstream” and “downstream” are used in the context of an SBS sequence reaction in which a nucleotide that is identified earlier in a sequencing read is considered to be “upstream” of a nucleotide that is determined later in the sequencing read.

[0049] Unless otherwise specified, "a," "an," "the," and "at least one" are used interchangeably and mean one or more than one.

[0050] As used herein, “providing” in the context of a compound, complex, composition, or article means making the compound, composition, complex, or article; purchasing the compound, complex, composition or article; or otherwise obtaining the compound, composition or article.

[0051] As used in this specification and the appended claims, the term "or" is generally employed in its sense including "and / or" unless the content clearly dictates otherwise. The term "and / or" means one or all of the listed elements or a combination of any two or more of the listed elements. The use of "and / or" in some instances does not imply that the use of "or" in other instances may not mean "and / or."

[0052] The words "preferred" and "preferably" refer to embodiments of the disclosure that may afford certain benefits, under certain circumstances. However, other embodiments may also be preferred, under the same or other circumstances. Furthermore, the recitation of one or more preferred embodiments does not imply that other embodiments are not useful and is not intended to exclude other embodiments from the scope of the disclosure.

[0053] As used herein, "have," "has," "having," "include," "includes," "including," "comprise," "comprises," "comprising" or the like are used in their open-ended inclusive sense, and generally mean "include, but not limited to," "includes, but not limited to," or "including, but not limited to."

[0054] It is understood that wherever embodiments are described herein with the language "have," "has," "having," "include," "includes," "including," "comprise," "comprises," "comprising" and the like, otherwise analogous embodiments described in terms of "consisting of" and / or "consisting essentially of" are also provided. The term "consistingof" means including, and limited to, whatever follows the phrase "consisting of." That is, "consisting of" indicates that the listed elements are required or mandatory, and that no other elements may be present. The term "consisting essentially of" indicates that any elements listed after the phrase are included, and that other elements than those listed may be included provided that those elements do not interfere with or contribute to the activity or action specified in the disclosure for the listed elements.

[0055] Conditions that are "suitable" for an event to occur, or "suitable" conditions are conditions that do not prevent such events from occurring. Thus, these conditions permit, enhance, facilitate, and / or are conducive to the event.

[0056] As used herein, "providing" in the context of, for instance, an amplification or resynthesis reagent, an array, or a composition, means making the amplification or resynthesis reagent, an array, or composition, purchasing the amplification or resynthesis reagent, an array, or composition, or otherwise obtaining the amplification or resynthesis reagent, an array, or composition.

[0057] Reference throughout this specification to "one embodiment," "an embodiment," "certain embodiments," or "some embodiments," etc., means that a particular feature, configuration, composition, or characteristic described in connection with the embodiment is included in at least one embodiment of the disclosure. Thus, the appearances of such phrases in various places throughout this specification are not necessarily referring to the same embodiment of the disclosure. Furthermore, the particular features, configurations, compositions, or characteristics may be combined in any suitable manner in one or more embodiments.

[0058] Throughout this disclosure, various aspects of the disclosure can be presented in a range format. It should be understood that the description in range format is merely for convenience and brevity and should not be construed as an inflexible limitation on the scope of the disclosure. Accordingly, the description of a range should be considered to have specifically disclosed all the possible subranges as well as individual numerical values within that range. For example, description of a range such as from 1 to 6 should be considered to have specifically disclosed subranges such as from 1 to 3, from 1 to 4, from 1 to 5, from 2 to 4, from 2 to 6, from 3 to 6 etc., as well as individual numbers within that range, for example, 1, 2, 2.7, 3, 4, 5, 5.3, and 6. This applies regardless of the breadth of the range.

[0059] In the description herein particular embodiments may be described in isolation for clarity. Unless otherwise expressly specified that the features of a particular embodiment are incompatible with the features of another embodiment, certain embodiments can include a combination of compatible features described herein in connection with one or more embodiments.

[0060] For any method disclosed herein that includes discrete steps, the steps may be conducted in any feasible order. And, as appropriate, any combination of two or more steps may be conducted simultaneously.DETAILED DESCRIPTION

[0061] Reference will now be made in greater detail to various embodiments of the subject matter of the present disclosure, some embodiments of which are illustrated in the accompanying drawings.

[0062] The present disclosure describes, among other things, methods for modifying the sequence of a target region of a target oligonucleotide on a solid surface prior to cluster generation in an oligonucleotide sequencing workflow. Such sequence modification may improve sequencing resolution downstream of the target region or may corroborate the authenticity of the determined sequence of a region of interest downstream of the target region. The target region may, for example, be within a region prone to sequence specific errors (SSEs). Modification of the sequence of a region prone to SSEs may disrupt the SSE region, which may facilitate monoclonal cluster generation and / or downstream sequencing. The target region may, for example, be a region upstream of a clinically relevant region or other region of interest. Modification of the sequence upstream of a region of interest may include inserting a bioinformatically identifiable sequence tag upstream of the region of interest. Verification of proper sequencing of the tag may provide assurances that a sequence downstream of the tag, including the region of interest, is properly determined.Methods for Modifying Sequence on Solid Surface in Preparation for Sequencing

[0063] In various aspects, the present disclosure describes a method for modifying a solid surface for oligonucleotide sequencing. The surface may be modified as part of a sequencing workflow. One example of a sequencing workflow or process 100 is illustratedin FIG. 1. The depicted workflow 100 includes library preparation (110), which may include fragmenting nucleic acids to be sequenced and adding adapters to the 3’ and 5’ ends of the nucleic acid fragments. The workflow 100 includes seeding library nucleic acids on the solid surface (120). Surface primers having free 3’ ends are bound to the solid surface, and the adapters on the library nucleic acids include sequences that are complementary to at least a portion of the surface primers and hybridize to the surface primers. The library nucleic acids are preferably seeded at a concentration and under conditions configured to cause a single library nucleic acid to hybridize to a single surface primer within a region of the solid surface to allow for formation of clusters as described in more detail below. The library oligonucleotide strands that hybridize to the surface primers of the solid surface are referred to herein as “target” oligonucleotides or “target” oligonucleotide strands. “Target oligonucleotide,” “target oligonucleotide strand,” and “target strand” are used herein interchangeably. The workflow 100 further includes first strand synthesis (130), which includes synthesizing a copy strand by extending the surface primer from the free 3 ’ end using the target strand, which is hybridized to the surface primer, as a template. Synthesis of the first copy strand bound includes introducing at least one modification in the sequence relative to the target strand (the sequence of the copy strand differs from the complementary sequence of the target strand) as discussed in more detail below. “Copy oligonucleotide,” “copy oligonucleotide strand,” and “copy strand” are used herein interchangeably. A more detailed embodiment of the steps of seeding the solid surface (120) and first strand synthesis (130) as indicated by the dashed box 2 in FIG. 1 is depicted in, and described below regarding, FIG. 2.

[0064] The workflow 100 depicted in FIG. 1 further includes cluster generation (120), which includes amplification of the copy strands to form clusters of oligonucleotides on the solid surface. The clusters comprise multiple copies of the copy strand bound to the solid surface and multiple copies of modified target strand bound to the solid surface. The modified target strand is complementary to the copy strand. Because first strand synthesis of the copy strand as described herein introduces a modification relative to the target strand, amplification of the copy strand results in a modified target strand relative to the original target strand. Following amplification and cluster formation, the multiple copies of the copy strand or the multiple copies of the modified target strand may be removed from the solid surface prior to sequencing (150). The multiple copies of the copy strand or the multiplecopies of the modified target strand may be sequenced according to any suitable method, such as those well-known in the art.

[0065] Examples and embodiments of various steps illustrated in the workflow or process 100 depicted in FIG. 1 are discussed in more detail below.

[0066] Formation of clusters of oligonucleotides having the same sequence (i.e., copy strands having the same sequence or modified template strands having the same sequence) can result in better sequencing resolution than clusters with oligonucleotides having different sequences. If a cluster is formed from strands having different sequences, signals obtained during sequencing may be mixed at any given cycle and, thus, may be more difficult to resolve relative to a pure signal generated during sequencing of strands having the same sequence.

[0067] Some regions of oligonucleotides prone to sequence specific errors (SSEs), such as homopolymer regions, may result in formation of oligonucleotides having different sequences during cluster generation. The methods described herein, which include introducing a modification during first strand synthesis, may be used to counter the effects of such SSE regions.

[0068] Even if clusters are formed from identical strands, sequencing errors may occur due to phasing. Phasing occurs when different nucleotides of different identical strands are read during a cycle of sequencing. That is, sequencing of one or more strands may lag or may lead sequencing of one or more other strands, such that the sequencing of the strands in a cluster are out of phase. Sequencing resolution suffers as the proportion of strands that are out of phase increases.

[0069] Some regions of oligonucleotides prone to SSEs, such as G-quadraplexes, may increase phasing during sequencing or may, in some circumstances, prevent any sequencing from occurring. The methods described herein, which include introducing a modification during first strand synthesis, may be used to counter the effects of such SSE regions.

[0070] The methods described herein may also be used to increase confidence that a region of interest is sequenced correctly. Introducing the modification during first strand synthesis may include introducing a bioinformatically identifiable sequence tag upstream of the region of interest. Verification of proper sequencing of the tag may provide assurancesthat a sequence downstream of the tag, including the region of interest, is properly determined.

[0071] Referring now to FIG. 2, an embodiment of a seeding and first strand synthesis method 200 as illustrated in box 2 of FIG. 1 is shown. The method 200 includes hybridizing a target oligonucleotide a target region to a surface primer bound to a solid surface (220). For purposes of this disclosure, a “target region” of a target oligonucleotide is a region of the oligonucleotide for which a sequence modification is intended or results. The target oligonucleotide may be a library oligonucleotide and may comprise a sequence complementary to at least a portion of the surface primer. The surface primer has a free 3’ end.

[0072] The method 200 further includes synthesizing a copy oligonucleotide using the target oligonucleotide as a template while also introducing a modification relative to the template oligonucleotide (230). Synthesis of the copy oligonucleotide comprises extending the free 3’ end of the surface primer using the target oligonucleotide as a template. The modification relative to the template oligonucleotide is referred to herein as a “tag” region of the copy oligonucleotide. The tag region is modified relative to the target region of the target oligonucleotide. The tag region may be configured to introduce one or more nucleotide substitutions and / or insertions relative to the (complementary) sequence of the target region of the target oligonucleotide. Accordingly, when the copy oligonucleotides are amplified during cluster generation (e.g., cluster generation (140) as illustrated in FIG. 1), modified target oligonucleotides are generated where the sequence of the modified target oligonucleotides differs from the sequence of the original target oligonucleotide. More specifically, the sequence of the target region of the original target oligonucleotide differs from a corresponding modified target region of the modified target oligonucleotide.

[0073] Tag regions having sequences that differ from corresponding (complementary) sequences of the target regions may be introduced into the copy oligonucleotide during synthesis of the copy oligonucleotide in any suitable manner. Examples of processes of introducing tag regions in copy oligonucleotide are illustrated in FIGS. 3-4.

[0074] As shown in FIG. 3, a target oligonucleotide 30 is hybridized to a surface primer 20 on a solid surface 10 in Step A. The surface primer 20 has a free 3’ end. A sequence in the 3 ’ end portion of the target oligonucleotide 30 is complementary to at least a portion of thesequence of the surface primer 30. The sequence of the 3’ end portion of the target oligonucleotide 30 may comprise an adapter added during library preparation.

[0075] In Step B of FIG. 3, a tagging oligonucleotide 50 comprising a tag region 52 binds to a target region 32 of the target oligonucleotide 30. The 3’ and 5’ end regions of the tagging oligonucleotide 50 flank the tag region 52 and each comprise a nucleotide sequence complementary to a corresponding nucleotide sequence of the target region 32. That is, the 3’ end region is complementary to a 5’ portion of the target region, and the 5’ end region is complementary to a 3’ portion of the target region. The 3’ and 5’ end regions of the tagging oligonucleotide 50 may have any suitable length to achieve hybridization. For example, the 3’ and 5’ end regions of the tagging oligonucleotide 50 may each independently comprise 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20or more nucleotides. If the number of nucleotides in the 3’ and 5’ end regions is low, the temperature or salt concentration at which hybridization occurs may be relatively low, and / or the end regions may comprise nucleotides having greater binding affinity, such as locked nucleic acids (LNAs). If the number of nucleotides is larger, more stringent hybridization conditions may be tolerated. In some embodiments, the overall size of the tagging oligonucleotide 50 may be in a range from about 6 to about 100 nucleotides or more. In some embodiments, the size of the tagging nucleotide 50 is in a range from about 10 to about 50 nucleotides, such as from about 15 to about 45 nucleotides, from about 20 to about 40 nucleotides, or from about 25 to about 35 nucleotides.

[0076] In Step C of FIG. 3, the free 3’ end of the surface primer 20 is extended using the target oligonucleotide 30 as a template. Extension may be carried out by a DNA polymerase. Preferably, the polymerase does not have strand displacing activity. Preferably, the DNA polymerase does not have 5’ to 3’ exonuclease activity. Examples of DNA polymerases that do not have strand displacing activity and do not have 5’ to 3’ exonuclease activity include, but are not limited to, Q5 High Fidelity polymerase, Phusion High Fidelity polymerase, T7 DNA polymerase, Sulfolobus DNA polymerase IV, and T4 DNA polymerase. The polymerase extends the first nascent strand 3O’(a) using the target polynucleotide 30 as a template until reaching the 5’ end of the tagging oligonucleotide 50, which is hybridized to the target region 32. In the depicted embodiment, the 3’ end of the tagging oligonucleotide 50 is modified, for example with a blocking group, to prevent the polymerase from extending the tagging oligonucleotide using the target oligonucleotide as a template. Any suitable blocking group may be used, including those well-known for use in nucleic acid sequencing such as thosedescribed in more detail elsewhere herein. However, it should be understood that the 3’ end of the tagging oligonucleotide 50 may be unblocked and the polymerase may extend the tagging oligonucleotide from the 3’ end using the target oligonucleotide 50 as a template during step C.

[0077] In Step D of FIG. 3, a ligase ligates the 3’ end of the first nascent strand 3O’(a) to the 5 ’ end of the tagging oligonucleotide 50 and the 3 ’ end of the tagging oligonucleotide 50 is unblocked. Any suitable ligase may be used to ligate the 3’ end of the first nascent strand 3O’(a) to the 5’ end of the tagging oligonucleotide 50. Examples of suitable ligases include, but are not limited to, T4 DNA ligase, T3 DNA ligase, T7 DNA ligase, Taq DNA ligase, SplintR ligase, E. coli DNA ligase, ElectroLigase®, Hi-T4 DNA Ligase™, immobilized T4 DNA ligase, Salt T4 DNA Ligase®, HiFi Taq DNA ligase, and 9°N™ DNA Ligase. Examples of suitable kits and mixes that comprise suitable ligases include, but are not limited to, Blunt / TA Ligase Master Mix, Instant Sticky-end Master Mix, Quick Ligation™ Kit, NEBridge® Ligase Master Mix, and NEBNext Quick Ligation™ Module. Any suitable reagents may be used to deblock the 3’ end of the tagging oligonucleotide 50, including those well-known for use in nucleic acid sequencing such as described in more detail elsewhere herein.

[0078] In Step E of FIG. 3, a polymerase, which may be the same polymerase or a different polymerase from the polymerase used in Step B, extends the tagging oligonucleotide 50 from the unblocked, free 3 ’ end using the target oligonucleotide 50 as a template producing a second nascent portion 3O’(b). The 3’ end of the second nascent portion 3O’(b) extends to the 5’ end of the target oligonucleotide 30. In Step E, the copy oligonucleotide 30’ is produced. The copy oligonucleotide 30’ is bound to the solid surface 10 via the remnant of the surface primer 20.

[0079] In Step F of FIG. 3, the target oligonucleotide 30 is washed away, leaving the bound copy oligonucleotide 30’, which may be amplified to form clusters as well-known in surface preparation for nucleic acid sequencing and as described in more detail elsewhere herein.

[0080] In embodiments (not shown) where the 3 ’ end of the tagging oligonucleotide is not blocked in Step B and is extended in Step C, ligation of the 3’ end of the first nascent strand3O’(a) and the 5’ end of the tagging oligonucleotide 52 (Step D) produces the copy oligonucleotide 30’ bound to the solid surface 10.

[0081] In embodiments, the process of preparing the solid surface 10 in FIG. 3 may be performed by (i) flowing a composition comprising an oligonucleotide library across the solid surface 10 under conditions that allow hybridization of the target oligonucleotide 30 to hybridize to the surface primer (Step A); (ii) optionally washing away unbound library oligonucleotides; (iii) flowing a composition comprising a tagging oligonucleotide 50 across the solid surface 10 under conditions that allow hybridization of the tagging oligonucleotide 50 to the target region 32 of the target oligonucleotide (Step B); (iv) optionally washing away unbound tagging oligonucleotide 50; (v) flowing a composition comprising a polymerase and nucleotides across the solid surface 10 under conditions that allow the polymerase to extend the surface primer from the free 3’ end using the target oligonucleotide 30 as a template, where the nascent first strand 3O’(a) extends to the 5’ end of the tagging oligonucleotide 50, which is hybridized to the target region 32 of the target oligonucleotide 30 (Step C); (vi) optionally washing away the polymerase and nucleotides; (vii) flowing a composition comprising a ligase across the solid surface 10 under conditions that allow ligation of the 3’ end of the nascent first strand 3O’(a) to the 5’ end of the tagging oligonucleotide 50, optionally washing away the ligase, and flowing a composition comprising unblocking reagents across the solid surface 10 to unblock the 3’ end of the tagging oligonucleotide (Step D); (viii) flowing a composition comprising a polymerase and nucleotides across the solid surface 10 under conditions that allow the polymerase to extend the tagging oligonucleotide from the free 3’ end using the target oligonucleotide 30 as a template, until the second nascent portion 3O’(b) extends to the 5’ end of target oligonucleotide 30 (Step E); (ix) optionally washing away the polymerase and nucleotides; and (x) washing away the target oligonucleotide 30 under denaturing conditions (Step F). The process may be carried out by nucleic acid sequencing apparatus.

[0082] Another embodiment of a process for preparing a solid surface 10 is shown in FIG. 4. A target oligonucleotide 30 is hybridized to a surface primer 20 on a solid surface 10 in Step A. The surface primer 20 has a free 3’ end. A sequence in the 3’ end portion of the target oligonucleotide 30 is complementary to at least a portion of the sequence of the surface primer 30. The sequence of the 3’ end portion of the target oligonucleotide 30 may comprise an adapter added during library preparation.

[0083] In Step B of FIG. 4, a blocking oligonucleotide 40 having a sequence complementary to at least a portion of the sequence of the target region 32 of the target oligonucleotide 30 binds to the target region 32. The blocking oligonucleotide may have any suitable length. For example, the blocking oligonucleotide may comprise 5, 6, 7, 8, 9, 10, 11,12, 13, 14, 15, 16, 17, 18, 19, 20 or more nucleotides.

[0084] In Step C of FIG. 4, the free 3’ end of the surface primer 20 is extended using the target oligonucleotide 30 as a template. Extension may be carried out by a DNA polymerase. Preferably, the polymerase does not have strand displacing activity. The polymerase extends the first nascent strand 3O’(a) using the target polynucleotide 30 as a template until reaching the 5’ end of the blocking oligonucleotide 40, which is hybridized to the target region 32. The 3’ end of the blocking oligonucleotide 40 is modified, for example with a blocking group, to prevent the polymerase from extending the tagging oligonucleotide using the target oligonucleotide as a template. Any suitable blocking group may be used, including those well- known for use in nucleic acid sequencing such as those described in more detail elsewhere herein.

[0085] In Step D of FIG. 4, the blocking nucleotide is washed away under conditions that do not denature the template oligonucleotide 30 from the surface primer 20 and nascent first strand 3O’(a), and a tagging oligonucleotide 50 comprising a tag region 52 binds to a target region 32 of the target oligonucleotide 30. The 3’ and 5’ end regions of the tagging oligonucleotide 50 flank the tag region 52 and each comprise a nucleotide sequence complementary to a corresponding nucleotide sequence of the target region 32. That is, the 3’ end region is complementary to a 5’ portion of the target region, and the 5’ end region is complementary to a 3’ portion of the target region. The 3’ and 5’ end regions of the tagging oligonucleotide 50 may have any suitable length. For example, the 3’ and 5’ end regions of the tagging oligonucleotide 50 may each independently comprise 3, 4, 5, 6, 7, 8, 9, 10, 11, 12,13, 14, 15, 16, 17, 18, 19, 20 or more nucleotides. The 5’ end region of the tagging oligonucleotide 50 comprises a sequence that is identical or substantially identical to a 5’ end region of the blocking oligonucleotide 40 so that the 5’ end region of the tagging oligonucleotide 50 binds to the same portion of the target region 52 as the 5 ’ end region of the blocking oligonucleotide 40.

[0086] Step D of FIG. 4 also includes ligating the 5’ end of the tagging oligonucleotide 50 to the 3’ end of the nascent first strand 3O’(a). Any suitable ligase may be used to ligate the3’ end of the first nascent strand 3O’(a) to the 5’ end of the tagging oligonucleotide 50. Examples of suitable ligases include T4 DNA ligase, T3 DNA ligase, T7 DNA ligase, Taq DNA ligase, SplintR ligase, E. coli DNA ligase, ElectroLigase®, Hi-T4 DNA Ligase™, immobilized T4 DNA ligase, Salt T4 DNA Ligase®, HiFi Taq DNA ligase, and 9°N™ DNA Ligase. Examples of suitable kits and mixes that comprise suitable ligases include, but are not limited to, Blunt / TA Ligase Master Mix, Instant Sticky-end Master Mix, Quick Ligation™ Kit, NEBridge® Ligase Master Mix, and NEBNext Quick Ligation™ Module.

[0087] The tagging oligonucleotide 50 may or may not have a blocked 3 ’ end when first contacted with the target region 52 of the target oligonucleotide. If the tagging oligonucleotide 50 has a blocked 3’ end, Step D further includes unblocking the 3’ end of the tagging oligonucleotide 50.

[0088] In Step E of FIG. 4, a polymerase, which may be the same polymerase or a different polymerase from the polymerase used in Step C, extends the tagging oligonucleotide 50 from the 3’ end using the target oligonucleotide 50 as a template producing a second nascent portion 3O’(b). The 3’ end of the second nascent portion 3O’(b) extends to the 5’ end of the target oligonucleotide 30. In Step E, the copy oligonucleotide 30’ is produced. The copy oligonucleotide 30’ is bound to the solid surface 10 via the remnant of the surface primer 20.

[0089] In Step F of FIG. 4, the target oligonucleotide 30 is washed away, leaving the bound copy oligonucleotide 30’, which may be amplified to form clusters as well-known in surface preparation for nucleic acid sequencing and as described in more detail elsewhere herein.

[0090] In embodiments, the process of preparing the solid surface 10 in FIG. 4 may be performed by (i) flowing a composition comprising an oligonucleotide library across the solid surface 10 under conditions that allow hybridization of the target oligonucleotide 30 to hybridize to the surface primer (Step A); (ii) optionally washing away unbound library oligonucleotides; (iii) flowing a composition comprising a blocking oligonucleotide 40 across the solid surface 10 under conditions that allow hybridization of the blocking oligonucleotide 40 to the target region 32 of the target oligonucleotide (Step B); (iv) optionally washing away unbound blocking oligonucleotide 40; (v) flowing a composition comprising a polymerase and nucleotides across the solid surface 10 under conditions that allow the polymerase to extend the surface primer from the free 3’ end using the target oligonucleotide 30 as a template, wherethe nascent first strand 30 ’(a) extends to the 5’ end of the blocking oligonucleotide 40, which is hybridized to the target region 32 of the target oligonucleotide 30 (Step C); (vi) optionally washing away the polymerase and nucleotides; (vii) flowing a composition comprising a tagging oligonucleotide 50 across the solid surface 10 under conditions that allow hybridization of the tagging oligonucleotide 50 to the target region 32 of the target oligonucleotide, optionally washing away unbound tagging oligonucleotide 50, flowing a composition comprising a ligase across the solid surface 10 under conditions that allow ligation of the 3 ’ end of the nascent first strand 3O’(a) to the 5’ end of the tagging oligonucleotide 50, optionally washing away the ligase, and flowing a composition comprising unblocking reagents across the solid surface 10 to unblock the 3’ end of the tagging oligonucleotide 50 if the 3’ end of the tagging oligonucleotide 50 is blocked (Step D); (viii) flowing a composition comprising a polymerase and nucleotides across the solid surface 10 under conditions that allow the polymerase to extend the tagging oligonucleotide from the free 3’ end using the target oligonucleotide 30 as a template, until the second nascent portion 3O’(b) extends to the 5’ end of target oligonucleotide 30 (Step E); (ix) optionally washing away the polymerase and nucleotides; and (x) washing away the target oligonucleotide 30 under denaturing conditions (Step F). The process may be carried out by nucleic acid sequencing apparatus.

[0091] While the methods depicted in FIGS. 3-4 show a tagging oligonucleotide 50 with only one tag region 52, it will be understood that a tagging oligonucleotide 50 may comprise more than one tag region.

[0092] The tag region 52 may be configured to introduce an insertion and / or substitution of nucleotides relative to the target region 32 of the target oligonucleotide 30. That is, a modified target nucleotide produced during cluster generation will have a nucleotide sequence that differs from the target region 32 of the target oligonucleotide 30 due to substitution and / or insertion of one or more nucleotides.

[0093] In some embodiments, the tag region 52 is configured to introduce an insertion of 5 or more, 6, or more, 7 or more, 8 or more, 9 or more, or 10 or more nucleotides. In some embodiments, particularly with larger insertions, the tag region 52 may be configured to form a hairpin structure through self-complementarity of nucleotides in the tag region 52. By forming a hairpin structure in the tag region 52, the 3’ and 5’ end regions of the tagging oligonucleotide 50 may be positioned in proximity to one another to facilitate hybridization to the target region 32 of the target oligonucleotide 30.

[0094] In some embodiments, the tag oligonucleotide 50 may be configured to introduce a deletion relative to the target region 32 of the target oligonucleotide 30. The tagging oligonucleotide 50 may have a 3 ’ end region abutting a 5 ’ end region. The 3 ’ end region of the tagging oligonucleotide 50 may be complementary to a first sequence of the target region 32 of the target oligonucleotide 30, and the 5’ end region of the tagging oligonucleotide 50 may be complementary to a second sequence of the target region 32, wherein the first and second sequences of the target region 32 are separated. The nucleotides separating the first and second sequences of the target region 32 may be deleted when applying the methods described herein. A modified target nucleotide produced during cluster generation will have a nucleotide sequence that differs from the target region 32 of the target oligonucleotide 30 due to the deletion.

[0095] In some embodiments, the modified sequence serves as a bioinformatically identifiable tagging element that may be sequenced and recognized to identify the location of the tag region 52 or modified target region.

[0096] In some embodiments, two or more different tagging oligonucleotides 50 may be introduced to the solid surface to modify target regions 32 of two or more target oligonucleotides 30. In some embodiments, a first tagging oligonucleotides 50 is configured to modify a first target region 32 of a target strand 32 and a second tagging oligonucleotides 50 is configured to modify a second target region 32 of a second target strand 30, wherein the first target region 32 has a sequence complementary to the second target region 32. The first and second tagging oligonucleotides 50 may have complementary or partly complementary sequences. The first and second tagging oligonucleotides 50 may be introduced sequentially to the solid surface to prevent annealing between the tagging oligonucleotides 50 having at least partly complementary sequences. For example, the first tagging oligonucleotide 50 may be configured to introduce a modification to a poly(T) homopolymer region and the second tagging oligonucleotide 50 may be configured to introduce a modification to a poly(A) homopolymer region.

[0097] The target region 32 of the target oligonucleotide 30 to which the tagging oligonucleotide 50 hybridizes may be any suitable region. In embodiments, the target region 32 is upstream of a region of interest, such as a clinically relevant region. In embodiments, the target region 32 is in an SSE region. In embodiments, the modification caused by the tagging oligonucleotide 32 disrupts the SSE region. Disruption of the SSE region mayprovide increased sequencing resolution in the SSE region and downstream of the SSE region. Examples of SSE regions include repeat regions such as homopolymer tracts, G- quadraplexes (G-Quads), and the like.Repeat Regions

[0098] A repeat region may comprise a target region of a target oligonucleotide, the sequence of which may be modified in accordance with the teachings presented herein. A repeat region is a portion of an oligonucleotide, such as a target oligonucleotide, that has a sequence of a single repeating nucleotide or a pattern of two or more repeating nucleotides. Examples of repeat regions include, for example, homopolymers (poly A, poly T, poly C, poly G), dinucleotide repeats (e.g., GT repeats, GA repeats, AC repeats, CT repeats, or the like), trinucleotide repeats, and tetranucleotide repeats.

[0099] Repeat regions may be regions prone to sequence specific errors (SSEs). Sequence performance degradation around the repeat region correlates with the length (number of nucleotides) of a repeat region, with longer repeat regions resulting in greater sequence performance degradation as schematically shown in FIG. 5. The processes described herein can interrupt repeat regions, resulting in two shorter repeat regions and thereby improving sequencing resolution of the repeat regions and regions downstream of the repeat region.

[0100] For example, and as schematically shown in FIG. 6, a T homopolymer region of a target oligonucleotide 30 may be interrupted using a tagging oligonucleotide 50 configured to hybridize to the target oligonucleotide 30 and configured to cause a substitution of a C for a T in a resulting modified target oligonucleotide 30” following cluster generation when using a solid surface preparation process as described herein. The modified target oligonucleotide 30” has two shorter T homopolymer repeats on either side of the C, relative to the length of the T homopolymer repeat in the target oligonucleotide 30. While FIG. 6 shows a T C substitution, it will be understood that any suitable insertion or substitution may be introduced into a repeat region. While FIG. 6 shows modification of a T homopolymer repeat region, it will be understood that any repeat region may be interrupted when employing the processes described herein.

[0101] Repeat regions may result in SSEs during cluster formation in which amplification of oligonucleotides results in different copies having different repeat regionlengths. This can result in phasing during sequencing of the repeat region and regions downstream of the repeat region, introducing noise and increasing error rate in the sequencing analysis. Copies having different length (longer and shorter) repeat regions may result from a phenomenon known as polymerase slipping. See, for example, Viguera, et al., EMBO J. 2001 May 15; 20(10):2587-95, doi: 10.1093 / emboj / 20.10.2587; and Shinde et al., Nucleic Acids Res. 2003 Feb 1; 31(3):974-80, doi: 10.1093 / nar / gkgl78. These errors can propagate during multiple rounds of amplification, and new errors may be introduced in each round, resulting a cluster having oligonucleotides with a variety of different lengths.Accordingly, phasing and associated high error rate may result in difficulties in sequencing of such clusters.G-Quadraplex Regions

[0102] A G-quadraplex region may comprise a target region of a target oligonucleotide, the sequence of which may be modified in accordance with the teachings presented herein. Although homopolymers constitute the bulk of SSEs, G-quadruplexes account for about 5-14% of SSEs and are often upstream clinically relevant regions of the human genome. Modification of the sequence of a G-quadraplex region may disrupt the structure of G-quadraplex regions to improve sequencing fidelity downstream of a G-quadraplex region.

[0103] A G-quadruplex (also referred to as G4, G-tetrad, and G-quad) is a highly thermodynamically stable structure formed in guanine-rich regions under physiological conditions. The stability is due to both Watson-Crick and Hoogsteen hydrogen bonds and is further stabilized by monovalent cations typically present in storage and sequencing buffers (Spiegel, et al, Trends in Chemistry, 2020 Feb; 2(2): 123-136, doi: 10.1016 / j.trechm.2019.07.002. An example of a G-quadraplex structure is shown in FIG. 7. G-quadruplexes can stack to form a structure of two or more G-quadruplexes.

[0104] G-quadruplex forming sequences may have the following sequence:GmXnGmXoGmXpGm, where m is the number of G residues in each short G-tract, which are usually directly involved in G-tetrad interactions (Burge et al., Nucleic Acids Res. 2006;34(19):5402-15, doi: 10. 1093 / nar / gkl655). Xn, Xoand Xpcan be any combination of residues, including G, forming the loops Id.).

[0105] G-quadruplexes cause a loss in signal due to the SBS polymerase stalling at the presence of the G-quadruplex secondary structure as schematically shown in FIG. 8. As a result, this causes loss of signal in one orientation, leading to inability to sequence or improper sequencing downstream of a G-quadraplex region. The methods described herein may disrupt the formation of G-quadraplexes, where the method includes hybridizing a tagging oligonucleotide to a target region comprising a G-quadruplex forming sequence (or to the complementary sequence thereof) to modify the sequence such that G-quadraplexes do not form (i.e., the g-quadruplex is resolved) as schematically shown in FIG. 8.

[0106] Modification of the sequence of a G-quadraplex region may disrupt the structure of G-quadraplex regions to improve sequencing fidelity downstream of a G-quadraplex region.Regions of Interest

[0107] Not only can the methods and processes described herein be used to resolve SSEs as described above regarding repeat regions and G-quadraplexes, but also the methods and processes described herein may be used to modify the sequence of any suitable target region of a target oligonucleotide upstream of a region of interest. The modification may include inserting or creating a bioinformatically identifiable sequence tag upstream of the region of interest. Verification of proper sequencing of the tag may provide assurances that a sequence downstream of the tag, including the region of interest, is properly determined.

[0108] The region of interest may, for example, be a clinically relevant region, such as a region known to contain disease-causing or disease-contributing mutations.Solid Surface

[0109] Any suitable substrate may form a solid surface described herein. Examples of suitable substrates include glass, modified glass, functionalized glass, inorganic glasses, microspheres (e.g., inert and / or magnetic particles), plastics, polysaccharides, nylon, nitrocellulose, ceramics, resins, silica, silica-based materials, carbon, metals, an optical fiber or optical fiber bundles, polymers and multiwell (e.g., microtiter) plates. Examples of suitable plastics include acrylics, polystyrene, copolymers of styrene and other materials, polypropylene, polyethylene, polybutylene, polyurethanes and Teflon™. Examples of suitable silica-based materials include silicon and various forms of modified silicon.

[0110] In some embodiments, a substrate can be within or part of a vessel such as a well, tube, channel, cuvette, Petri plate, bottle or the like. A useful vessel is a flow-cell, for example, as described in US Pat. No. 8,241,573 or Bentley et al., Nature 456:53-59 (2008). Examples of suitable flow-cells are those that are commercially available from Illumina, Inc. (San Diego, Calif.). Another useful vessel is a well in a multiwell plate or microtiter plate.

[0111] In some embodiments, the solid surface may comprise features. The features may be present in any of a variety of suitable formats. For example, the features may comprise wells, pits, channels, ridges, raised regions, pegs, posts or the like. The features may or may not contain beads or particles. Exemplary features include wells that are present in substrates used for commercial sequencing platforms sold by 454 LifeSciences (a subsidiary of Roche, Basel Switzerland) or Ion Torrent (a subsidiary of Life Technologies, Carlsbad Calif.). Other substrates having wells include, for example, etched fiber optics and other substrates described in U.S. Pat. No. 6,266,459; U.S. Pat. No. 6,355,431; U.S. Pat. No. 6,770,441; U.S. Pat. No. 6,859,570; U.S. Pat. No. 6,210,891; U.S. Pat. No. 6,258,568; U.S. Pat. No. 6,274,320; U.S. Pat No. 8,262,900; U.S. Pat. No. 7,948,015; U.S. Pat. Pub. No. 2010 / 0137143; U.S. Pat. No. 8,349,167, or PCT Publication No. WO 00 / 63437. In several cases the substrates are exemplified in these references for applications that use beads in the wells. The well-containing substrates can be used with or without beads in the methods or compositions of the present disclosure. In some embodiments, wells of a substrate can include gel material (with or without beads) as set forth in U.S. Pat. No. 9,512,422.

[0112] The solid surface may comprise metal features on a non-metallic surface such as glass, plastic or other materials exemplified herein. A metal layer can be deposited on a surface using methods known in the art such as wet plasma etching, dry plasma etching, atomic layer deposition, ion beam etching, chemical vapor deposition, vacuum sputtering, or the like. Any of a variety of commercial instruments can be used as appropriate including, for example, the FlexAL®, OpAL®, lonfab 300Plus®, or Optofab 3000® systems (Oxford Instruments, UK). A metal layer can also be deposited by e-beam evaporation or sputtering as set forth in Thornton, Ann. Rev. Mater. Sci. 7:239-60 (1977). Metal layer deposition techniques, such as those exemplified herein, can be combined with photolithography techniques to create metal regions or patches on a surface. Exemplary methods for combining metal layer deposition techniques and photolithography techniques are provided in U.S. Pat. No. 8,778,848 and U.S. Pat. No. 8,895,249.

[0113] In some embodiments, the solid surface includes a collection of beads or other particles. The particles can be suspended in a solution or they can be located on a surface of a substrate. Examples of bead in solution are those commercialized by Luminex (Austin, Tex.). Examples of beads located on a surface include those where beads are located in wells such as a BeadChip array (Illumina Inc., San Diego Calif.) or substrates used in sequencing platforms from 454 LifeSciences (a subsidiary of Roche, Basel Switzerland) or Ion Torrent (a subsidiary of Life Technologies, Carlsbad Calif.). Other arrays having beads located on a surface are described in U.S. Pat. No. 6,266,459; U.S. Pat. No. 6,355,431; U.S. Pat. No. 6,770,441 ; U.S. Pat. No. 6,859,570; U.S. Pat. No. 6,210,891; U.S. Pat. No. 6,258,568; U.S. Pat. No. 6,274,320; US 2009 / 0026082 Al; US 2009 / 0127589 Al; US 2010 / 0137143 Al; US 2010 / 0282617 Al, or PCT Publication No. WO 00 / 63437. The beads or other solid surfaces (e.g., surface of a well or gel within a well) can be made to include surface primers.

[0114] In some embodiments, a surface primer can be attached to the solid surface or an amplification site of a solid surface. For example, the surface primer can be attached to the surface of a feature of an array. The attachment can be via an intermediate structure such as a bead, particle, or gel. An example of attachment of surface primers to an array via a gel is described in U.S. Pat. No. 8,895,249 and further exemplified by flow-cells available commercially from Illumina Inc. (San Diego, Calif.) or described in WO 2008 / 093098. Exemplary gels that can be used in the methods and apparatus set forth herein include, but are not limited to, those having a colloidal structure, such as agarose; polymer mesh structure, such as gelatin; or cross-linked polymer structure, such as polyacrylamide, SFA (see, for example, US Pat. App. Pub. No. 2011 / 0059865 Al) or PAZAM (see, for example, U.S. Prov. Pat. App. Ser. No. 61 / 753,833 and U.S. Pat. No. 9,012,022). Attachment via a bead can be achieved as exemplified in the description and cited references set forth previously herein.

[0115] Amplification sites of an array can include a plurality of surface primers capable of binding to target oligonucleotides. In typical conditions used to prepare arrays for sequencing, the nucleotide sequence of the surface primer is complementary to a sequence of one or more modified target oligonucleotides, such as a universal primer binding sequence present on a target oligonucleotide. In some embodiments, the surface primer can also function as a primer for amplification of the modified target oligonucleotide. In some embodiments, one population of surface primer includes a P5 primer or the complement thereof, and the second population of surface primer includes a P7 primer or the complement thereof.

[0116] A surface primer can be immobilized by single point covalent attachment to a solid surface at or near the 5' end of the surface primer, leaving the template-specific portion of the surface primer free to anneal to its cognate universal primer binding sequence and the 3' hydroxyl group free for extension. Any suitable covalent attachment means known in the art may be used for this purpose. The chosen attachment chemistry will depend on the nature of the solid surface, and any derivatization or functionalization applied to it. The surface primer itself may include a moiety, which may be a non-nucleotide chemical modification, to facilitate attachment. In a particular embodiment, the surface primer may include a sulphur-containing nucleophile, such as phosphorothioate or thiophosphate, at the 5' end.

[0117] In some embodiments, the features on the solid surface of an array substrate are non-contiguous, being separated by interstitial regions of the solid surface. Interstitial regions that have a substantially lower quantity or concentration of surface primers, compared to the features of the array, are advantageous. Interstitial regions that lack surface primers are particularly advantageous. For example, a relatively small amount or absence of surface primers at the interstitial regions favors localization of target oligonucleotides, and subsequently generated clusters, to desired features. In particular embodiments, the features can be concave features in a surface (e.g., wells) and the features can contain a gel material. The gel-containing features can be separated from each other by interstitial regions on the surface where the gel is substantially absent or, if present, the gel is substantially incapable of supporting localization of nucleic acids. Methods and compositions for making and using substrates having gel containing features, such as wells, are set forth in U.S. Pat. No. 9,512,422. The size of the features and / or spacing between the regions can vary such that arrays can be high density, medium density or lower density. High density arrays are characterized as having regions separated by less than about 15 pm. Medium density arrays have regions separated by about 15 to 30 pm, while low density arrays have regions separated by greater than 30 pm. An array useful in the disclosure can have regions that are separated by less than 100 pan, 50 pan, 10 pan, 5 pm, 1 pm or 0.5 pm.

[0118] In some embodiments, the solid surface comprises a patterned surface. A "patterned surface" refers to an arrangement of different regions in or on an exposed layer of a solid surface. For example, one or more of the regions can be features where one or more surface primers are present. In some embodiments, the pattern can be an x-y format of features that are in rows and columns. In some embodiments, the pattern can be a repeating arrangementof features and / or interstitial regions. In some embodiments, the pattern can be a random arrangement of features and / or interstitial regions. In some embodiments, the pattern can appear as a grid of spots or patches. The features can be located in a repeating pattern or in an irregular non-repeating pattern. Particularly useful patterns are hexagonal patterns, rectilinear patterns, grid patterns, patterns having reflective symmetry, patterns having rotational symmetry, or the like. Asymmetric patterns can also be useful. The pitch can be the same between different pairs of nearest neighbor features or the pitch can vary between different pairs of nearest neighbor features. In particular embodiments, features of an array can each have an area that is larger than about 100 nm2, 250 nm2, 500 nm2, 1 micrometer2, 2.5 micrometer2, 5 micrometer2, 10 micrometer2, 100 micrometer2, or 500 micrometer2. Alternatively, or additionally, features of an array can each have an area that is smaller than about 1 mm2, 500 micrometer2, 100 micrometer2, 25 micrometer2, 10 micrometer2, 5 micrometer2, 1 micrometer2, 500 nm2, or 100 nm2. Indeed, a region can have a size that is in a range between an upper and lower limit selected from those exemplified above. Exemplary patterned surfaces that can be used in the methods and compositions set forth herein are described in U.S. Pat. Nos. 8,778,848, 8,778,849 and 9,079,148, and U.S. Pat. Appl. Pub. No. 2014 / 0243224.

[0119] The features in a patterned surface can be wells in an array of wells (e.g., microwells or nanowells) on glass, silicon, plastic or other suitable solid surfaces with patterned, covalently-linked gel such as poly(N-(5-azidoacetamidylpentyl)acrylamide-co- acrylamide) (PAZAM, see, for example, US Pub. No. 2013 / 184796, WO 2016 / 066586, and WO 2015 / 002813). The process can create gel pads used for sequencing that can be stable over sequencing runs with a large number of cycles. The covalent linking of the polymer to the wells is helpful for maintaining the gel in the structured features throughout the lifetime of the structured substrate during a variety of uses. However, in many embodiments the gel need not be covalently linked to the wells. For example, in some conditions silane free acrylamide (SFA, see, for example, US Pat. No. 8,563,477) which is not covalently attached to any part of the structured substrate, can be used as the gel material.

[0120] In particular embodiments, a structured substrate can be made by patterning a solid surface material with wells (e.g., microwells or nanowells), coating the patterned solid surface with a gel material (e.g., PAZAM, SFA, or chemically modified variants thereof, such as the azidolyzed version of SFA (azido-SFA)) and polishing the gel coated solid surface, for examplevia chemical or mechanical polishing, thereby retaining gel in the wells but removing or inactivating substantially all of the gel from the interstitial regions on the solid surface of the structured substrate between the wells. Surface primers can be attached to gel material. A solution of modified target oligonucleotides can then be contacted with the polished substrate such that individual modified target oligonucleotides will seed individual wells via interactions with surface primers attached to the gel material; however, the target oligonucleotides will not occupy the interstitial regions due to absence or inactivity of the gel material. Amplification of the modified target oligonucleotides will be confined to the wells since absence or inactivity of gel in the interstitial regions prevents outward migration of the growing nucleic acid colony. The process can be conveniently manufactured, being scalable and utilizing conventional micro- or nanofabrication methods.Target oligonucleotides

[0121] As used herein a “target oligonucleotide” may be used interchangeably with “target nucleic acid.” A target oligonucleotide may be essentially any nucleic acid of known or unknown sequence. It may be, for example, a fragment of genomic DNA or cDNA. Sequencing may result in determination of the sequence of the whole, or a part of the target molecule. The targets can be derived from a primary nucleic acid sample that has been randomly fragmented. In one embodiment, the targets can be processed into templates suitable for amplification by the placement of universal amplification sequences, e.g., sequences present in a universal adapter.

[0122] The primary nucleic acid sample may originate in double-stranded DNA (dsDNA) form (e.g., genomic DNA fragments, amplification products and the like) from a sample or may have originated in single- stranded form from a sample, as DNA or RNA, and been converted to dsDNA form. By way of example, mRNA molecules may be copied into doublestranded cDNAs suitable for use in a method described herein using standard techniques well known in the art. The precise sequence of the polynucleotide molecules from a primary nucleic acid sample is generally not material to the disclosure, and may be known or unknown.

[0123] In one embodiment, the primary polynucleotide molecules from a primary nucleic acid sample are DNA molecules. More particularly, the primary polynucleotide molecules represent the entire genetic complement of an organism, and are genomic DNA molecules which include both intron and exon sequences, as well as non-coding regulatory sequences such as promoter and enhancer sequences. In one embodiment, particular sub-sets ofpolynucleotide sequences or genomic DNA can be used, such as, for example, particular chromosomes. Yet more particularly, the sequence of the primary polynucleotide molecules is not known. Still yet more particularly, the primary polynucleotide molecules are human genomic DNA molecules. The DNA oligonucleotides may be treated chemically or enzymatically either prior or subsequent to any random fragmentation processes, and prior or subsequent to the ligation of a universal sequence, such as universal adapter sequences.

[0124] The nucleic acid sample can include high molecular weight material such as genomic DNA (gDNA). The sample can include low molecular weight material such as nucleic acid molecules obtained from FFPE or archived DNA samples. In another embodiment, low molecular weight material includes enzymatically or mechanically fragmented DNA. The sample can include cell-free circulating DNA. A sample can include, but is not limited to, nucleic acid molecules obtained from biopsies, tumors, scrapings, swabs, blood, mucus, urine, plasma, semen, hair, laser capture micro-dissections, surgical resections, and other clinical or laboratory obtained samples. In some embodiments, the sample can be an epidemiological, agricultural, forensic or pathogenic sample.

[0125] The biological source of a sample is not intended to be limiting. In some embodiments, the sample can include nucleic acid molecules obtained from a eukaryote, such as an animal or a plant. Examples of an animal include, but are not limited to, a mammal including a human. In some embodiments, the sample can include nucleic acid molecules obtained from a prokaryote, such as a bacterium or archaeon. In some embodiments, the sample can include nucleic acid molecules obtained from a virus. In some embodiments, the source of the nucleic acid molecules may be an archived or extinct sample or species.

[0126] Random fragmentation refers to the fragmentation of a polynucleotide molecule from a primary nucleic acid sample in a non-ordered fashion by enzymatic, chemical or mechanical means. Such fragmentation methods are known in the art and use standard methods (Sambrook and Russell, Molecular Cloning, A Laboratory Manual, third edition). In one embodiment, enzymatic fragmentation can be accomplished using a process often referred to as tagmentation. Tagmentation uses a transposome complex that can include both transposon and transposase and combines into a single step fragmentation and ligation to add universal sequences that can be used as universal adapters or for the addition of other universal sequences (Gunderson et al., WO 2016 / 130704). For the sake of clarity, generating smaller fragments of a larger piece of nucleic acid via specific PCR amplification of such smaller fragments is notequivalent to fragmenting the larger piece of nucleic acid because the larger piece of nucleic acid sequence remains in intact (i.e., is not fragmented by the PCR amplification). Moreover, random fragmentation is designed to produce fragments irrespective of the sequence identity or position of nucleotides comprising and / or surrounding the break. More particularly, the random fragmentation is by mechanical means such as nebulization or sonication to produce fragments of about 50 base pairs in length to about 1500 base pairs in length, still more particularly 50-700 base pairs in length, yet more particularly 50-400 base pairs in length. Most particularly, the method is used to generate smaller fragments of from 50-150 base pairs in length.

[0127] Fragmentation of polynucleotide molecules by mechanical means (nebulization, sonication and Hydroshear, for example) results in fragments with a heterogeneous mix of blunt and 3’- and 5 '-overhanging ends. It is therefore desirable to repair the fragment ends using methods or kits (such as the Lucigen DNA terminator End Repair Kit) known in the art to generate ends that are optimal for insertion, for example, into blunt sites of cloning vectors. In a particular embodiment, the fragment ends of the population of nucleic acids are blunt ended. More particularly, the fragment ends are blunt ended and phosphorylated. The phosphate moiety can be introduced via enzymatic treatment, for example, using polynucleotide kinase.

[0128] A population of target oligonucleotides, or amplicons thereof, can have an average strand length that is desired or appropriate for a particular application of the methods or compositions set forth herein. For example, the average strand length can be less than about 100,000 nucleotides, 50,000 nucleotides, 10,000 nucleotides, 5,000 nucleotides, 1,000 nucleotides, 500 nucleotides, 100 nucleotides, or 50 nucleotides. Alternatively, or additionally, the average strand length can be greater than about 10 nucleotides, 50 nucleotides, 100 nucleotides, 500 nucleotides, 1,000 nucleotides, 5,000 nucleotides, 10,000 nucleotides, 50,000 nucleotides, or 100,000 nucleotides. The average strand length for population of target oligonucleotides, or amplicons thereof, can be in a range between a maximum and minimum value set forth above. It will be understood that amplicons generated at an amplification site (or otherwise made or used herein) can have an average strand length that is in a range between an upper and lower limit selected from those exemplified above.

[0129] In some cases, a population of target oligonucleotides can be produced under conditions or otherwise configured to have a maximum length for its members. For example, the maximum length for the members that are used in one or more steps of a method set forthherein or that are present in a particular composition can be less than 100,000 nucleotides, less than 50,000 nucleotides, less than 10,000 nucleotides, less than 5,000 nucleotides, less than 1,000 nucleotides, less than 500 nucleotides, less than 100 nucleotides, or less than 50 nucleotides. Alternatively, or additionally, a population of target oligonucleotides, or amplicons thereof, can be produced under conditions or otherwise configured to have a minimum length for its members. For example, the minimum length for the members that are used in one or more steps of a method set forth herein or that are present in a particular composition can be more than 10 nucleotides, more than 50 nucleotides, more than 100 nucleotides, more than 500 nucleotides, more than 1,000 nucleotides, more than 5,000 nucleotides, more than 10,000 nucleotides, more than 50,000 nucleotides, or more than 100,000 nucleotides. The maximum and minimum strand length for target oligonucleotides in a population can be in a range between a maximum and minimum value set forth above. It will be understood that amplicons generated at an amplification site (or otherwise made or used herein) can have maximum and / or minimum strand lengths in a range between the upper and lower limits exemplified above.

[0130] In particular embodiments, the target oligonucleotides are sized relative to the area of the amplification sites, for example, to facilitate exclusion amplification. For example, the area for each of the sites of an array can be greater than the diameter of the excluded volume of the target oligonucleotides in order to achieve exclusion amplification. Taking, for example, embodiments that use an array of features on a solid surface, the area for each of the features can be greater than the diameter of the excluded volume of the target oligonucleotides that are transported to the amplification sites. The excluded volume for a target oligonucleotide and its diameter can be determined, for example, from the length of the target oligonucleotide. Methods for determining the excluded volume of nucleic acids and the diameter of the excluded volume are described, for example, in U.S. Pat. No. 7,785,790; Rybenkov et al., Proc. Natl. Acad. Sci. U.S.A. 90: 5307-5311 (1993); Zimmerman et al., J. Mol. Biol. 222:599-620 (1991); or Sobel et al., Biopolymers 31: 1559-1564 (1991).

[0131] In a particular embodiment, the target fragment sequences are prepared with single overhanging nucleotides by, for example, activity of certain types of DNA polymerase such as Taq polymerase or Klenow exo minus polymerase which has a non-template-dependent terminal transferase activity that adds a single deoxynucleotide, for example, deoxyadenosine (A) to the 3' ends of a DNA molecule, for example, a PCR product. Such enzymes can be usedto add a single nucleotide ‘A’ to the blunt ended 3' terminus of each strand of the doublestranded target fragments. Thus, an ‘A’ could be added to the 3’ terminus of each end repaired strand of the double- stranded target fragments by reaction with Taq or Klenow exo minus polymerase, while a universal adapter polynucleotide construct could be a T-construct with a compatible ‘T’ overhang present on the 3’ terminus of each region of double stranded nucleic acid of the universal adapter. This end modification also prevents self-ligation of both vector and target such that there is a bias towards formation of the combined ligated adaptor-target- adaptor molecules.Sequencing Library Preparation

[0132] A sequencing library described herein typically includes a target oligonucleotide having a universal adapter attached one or both ends. A library of target oligonucleotides refers to the collection of target oligonucleotides containing known common sequences at their 3' and 5' ends, and may also be referred to as a 3' and 5' modified library.

[0133] Methods for attaching a universal adapter to one of both ends of a target oligonucleotide are known to the person skilled in the art. The attachment can be through standard library preparation techniques using ligation (Chesney et al. U.S. Pat. Pub. No. 2018 / 0305753 Al), through tagmentation using transposase complexes (Gunderson et al., WO 2016 / 130704), or primer extension, for instance when preparing a sample for targeted sequencing.

[0134] Target oligonucleotides are often amplified during sequencing library preparation. Amplification of modified target oligonucleotides during sequencing library preparation can be by linear amplification, exponential amplification, or both linear and exponential amplification steps. Amplification conditions useful during sequencing library preparation are routine and known to the person of ordinary skill in the art. For instance, amplification profiles (e.g., number of cycles and the temperature and time of each cycle), and concentrations of target oligonucleotides, buffers, ions, dNTPs, and polymerase are known or can be easily determined using commercially available algorithms.

[0135] In one embodiment, double-stranded target oligonucleotides from a sample, e.g., a fragmented sample, are treated by first ligating identical universal adaptor molecules to the 5' and 3' ends of the double- stranded target oligonucleotides (which may be of known, partially known or unknown sequence). In some embodiments, the identical universal adaptormolecules can be ‘mismatched adaptors’, the general features of which are defined below, and further described in Gormley et al., US 7,741,463, and Bignell et al., US 8,053,192). In some embodiments, the identical universal adaptor molecules can include fully complementary polynucleotide strands. A universal adaptor typically includes the universal primer binding sequences that aid in immobilizing the target oligonucleotides on an array for subsequent cluster generation. In one embodiment, library preparation of target oligonucleotides having universal adaptor molecules at the 5' and 3’ ends includes one or more amplification, for instance by PCR, before immobilizing the target oligonucleotides on an array for subsequent cluster generation.

[0136] In some embodiments, for instance when a universal adapter is added by tagmentation, it is desirable to modify the universal adapter present at each end of target oligonucleotides before cluster generation. The modification can occur by an amplification step, such as PCR. For instance, an initial primer extension reaction is carried out using a universal primer binding site in which extension products complementary to both strands of each target oligonucleotide are formed and add a universal primer binding sequence. The resulting primer extension products, and amplified copies thereof, collectively provide a library of modified target oligonucleotides that can be immobilized, clonally expanded to form clusters, and then sequenced. In some embodiments, a library includes target oligonucleotides originating from the same source, e.g., the same tissue, same cell, and / or same individual (for instance, a sample of cell-free DNA). The 3’ ends, and optionally the 5’ ends, of the universal adapters attached to the target oligonucleotides can include a homogeneous population or a heterogeneous population of universal primer binding sequences described herein.

[0137] Generally, amplification reactions require at least two amplification primers, often denoted 'forward' and 'reverse' primers (primer oligonucleotides) that are capable of annealing specifically to a part of the nucleic acid sequence to be amplified, e.g., a universal adapter at the ends of oligonucleotides, under conditions encountered in the primer annealing step of each cycle of an amplification reaction. It will be understood by the skilled person that if the primers contain any nucleotide sequence which does not anneal to the modified target oligonucleotides in the first amplification cycle then this sequence may be copied into the amplification products. For instance, the use of primers having universal primer binding sequences, i.e., sequences that do not anneal to the universal adapter at the ends of target oligonucleotides, the universal primer binding sequences will be incorporated into the resulting amplicon.

[0138] Amplification primers are generally single stranded polynucleotide structures. They may also contain a mixture of natural and non-natural bases and also natural and nonnatural backbone linkages, provided that any non-natural modifications does not preclude function as a primer-that being defined as the ability to anneal to a template polynucleotide strand during conditions of the amplification reaction and to act as an initiation point for synthesis of a new polynucleotide strand complementary to the template strand. Primers may additionally include non-nucleotide chemical modifications, for example phosphorothioates to increase exonuclease resistance, again provided such that modifications do not prevent primer function.

[0139] In some embodiments, the universal adapters used in the method of the disclosure are referred to as ‘mismatched’ adaptors because the adaptors include a region of sequence mismatch, i.e., they are not formed by annealing of fully complementary polynucleotide strands. Mismatched adaptors for use herein typically include at least one double- stranded region, also referred to as a region of double stranded nucleic acid, and at least one unmatched single-stranded region, also referred to as a region of single-stranded non-complementary nucleic acid strands. Mismatched adapters are routinely used in producing sequencing libraries, and the characteristics of useful mismatched adapters are known to the skilled person.

[0140] The ‘double-stranded region’ of the universal adapter is a short double-stranded region, typically including 5 or more consecutive base pairs, formed by annealing of the two partially complementary polynucleotide strands. As used herein, the term “double stranded,” when used in reference to a nucleic acid molecule, means that substantially all of the nucleotides in the nucleic acid molecule are hydrogen bonded to a complementary nucleotide. A partially double stranded nucleic acid can have at least 10%, 25%, 50%, 60%, 70%, 80%, 90% or 95% of its nucleotides hydrogen bonded to a complementary nucleotide.

[0141] The double-stranded region can form the ‘ligatable’ end of the adaptor, e.g., the end that is joined to a double-stranded target oligonucleotide in the ligation reaction. The ligatable end of the universal adaptor may be blunt or, in other embodiments, short 5' or 3' overhangs of one or more nucleotides may be present to facilitate / promote ligation. The 5’ terminal nucleotide at the ligatable end of the universal adapter is typically phosphorylated to enable phosphodiester linkage to a 3' hydroxyl group on the target polynucleotide.

[0142] The term ‘unmatched region’ refers to a region of the universal adaptor, the region of single- stranded non-complementary nucleic acid strands, wherein the sequences of the two polynucleotide strands forming the universal adaptor exhibit a degree of non-complementarity such that the two strands are not capable of fully annealing to each other under standard annealing conditions for a primer extension or PCR reaction. The unmatched region(s) may exhibit some degree of annealing under standard reaction conditions for an enzyme-catalyzed ligation reaction, provided that the two strands revert to single stranded form under annealing conditions in an amplification reaction.

[0143] A universal adapter can include at least one universal primer binding site. A universal primer binding site is a universal sequence that can be used for amplification and / or sequencing of a target oligonucleotide attached to the universal adapter. Examples of universal primer binding sites include, but are not limited to, sequences complementary to a Readl or Read2 primer.

[0144] A universal adapter can include at least one index. An index can be used as a marker characteristic of the source of particular target oligonucleotide on an array. Generally, the index is a synthetic sequence of nucleotides that is part of the universal adapter which is added to the target oligonucleotides as part of the library preparation step. Accordingly, an index is a nucleic acid sequence which is attached to each of the target oligonucleotidesof a particular sample, the presence of which is indicative of, or is used to identify, the sample or source from which the target molecules were isolated.

[0145] In some embodiments, the index may be up to 20 nucleotides in length, more preferably 1-10 nucleotides, and most preferably 4-8 nucleotides in length. For example, a four- nucleotide index gives a possibility of multiplexing 256 (44) samples on the same array, whereas a six base index enables 4,096 (46) samples to be processed on the same array.

[0146] In one embodiment, the universal primer binding sequence and / or universal primer binding site is part of the universal adapter when it is ligated to the double-stranded target fragments, and in another embodiment the universal primer binding sequence and / or universal primer binding site is added to the universal adapter after the universal adapter is ligated to the double-stranded target fragments. The addition can be accomplished using routine methods, including PCR-based methods.

[0147] The precise nucleotide sequence of the universal adapters is generally not material to the invention and may be selected by the user such that the desired sequence elements are ultimately included in the common sequences of the plurality of different modified target oligonucleotides, for example, to provide for the universal primer binding sequences and universal primer binding sites for particular sets of universal primers. Additional sequence elements may be included, for example, to provide binding sites for sequencing primers, e.g., Readl and Read2 primers, which will ultimately be used in sequencing of target oligonucleotides in the library, or products derived from amplification of the target oligonucleotides in the library, for example on a solid surface. In some embodiments, a universal adapter may include mixtures of natural and non-natural nucleotides (e.g., one or more ribonucleotides) linked by a mixture of phosphodiester and non-phosphodiester backbone linkages.

[0148] Ligation methods for adding a universal adapter to a target oligonucleotide are known in the art and use standard methods. Such methods use ligase enzymes such as DNA ligase to effect or catalyze joining of the ends of the two polynucleotide strands of, in this case, the universal adapter and the double-stranded target oligonucleotides, such that covalent linkages are formed. The universal adapter may contain a 5'-phosphate moiety to facilitate ligation to the 3'-OH present on the target fragment. The double- stranded target oligonucleotide contains a 5'-phosphate moiety, either residual from the shearing process, or added using an enzymatic treatment step, and has been end repaired, and optionally extended by an overhanging base or bases, to give a 3'-OH suitable for ligation.

[0149] As discussed herein, in one embodiment universal adaptors used in the ligation are complete and include a universal primer binding sequence and other universal sequences, e.g., a universal primer binding site and an index sequence. The resulting plurality of modified target oligonucleotides can be amplified before immobilization for sequencing. Also, as discussed herein, in one embodiment universal adaptors used in the ligation include a universal primer binding site and an index sequence, and do not include a universal primer binding sequence. The resulting plurality of modified target oligonucleotides can be further modified to include specific sequences, such as a universal primer binding sequence, and can be amplified before immobilization for sequencing.Hybridization of Target Oligonucleotides to Surface Primers and Production of ClonalClusters

[0150] In some embodiments, methods of the present disclosure include hybridizing a target oligonucleotide to a surface primer of a solid surface. In some embodiments, a solid surface comprising an array of surface primers is contacted with a single-stranded sequencing library to allow the target oligonucleotide to hybridize with the surface primer. The solid surface comprises at least one, and in some embodiments two or more populations of surface primers in an amplification site. The method includes using conditions suitable for attaching the universal adapter to one of the surface primers to result in a plurality of amplification sites that each include one member of the sequencing library. The conditions useful for the attaching are routinely used in sequencing workflows and are known to the skilled person.

[0151] In embodiments where the target oligonucleotides include at least one universal primer binding sequence and a complementary surface primer is present, sequences of the universal primer binding sequence and the complementary surface primer hybridize to result in a plurality of amplification sites that each include one member of the sequencing library. The addition of a member of a sequencing library to an amplification site is referred to as “seeding” the site. The seeding can be accomplished by use of a seeding reagent. A seeding reagent can include an array of amplification sites and a plurality of target oligonucleotides. The 3’ end of the member of the sequencing library is hybridized to a complementary surface primer that is present in the universal primer binding sequence. The skilled person will recognize that some amplification sites can include more than one member of the sequencing library at this stage and not significantly reduce the ability to obtain useful data from the subsequent sequencing reaction. The skilled person will also recognize that not all amplification sites of an array need to be occupied.

[0152] The method can further include first strand synthesis to result in immobilization of a copy oligonucleotide to a surface primer of an amplification site as described herein above. First strand synthesis and immobilization can be accomplished by extending the 3’ end of the first surface primer associated with members of the sequencing library at the amplification sites. The extending includes the incorporation of nucleotides by a DNA polymerase using the attached member of the sequencing library as a template, and results in an extended nucleic acid that is immobilized to the surface of the amplification site. The immobilization can be accomplished by use of an immobilization reagent. An immobilization reagent can include anarray of amplification sites, a plurality of target oligonucleotides, dNTPs (e.g., dATP, dTTP, dCTP, and dGTP), and a polymerase. A polymerase extends the immobilized surface primer using the nucleotide sequence of the member of the sequencing library as template, resulting in an immobilized complement of the member of the sequencing library. As described in more detail above, the methods described herein include incorporation of a tag region into the copy oligonucleotide immobilized to the solid surface. Under some conditions, such as when kinetic exclusion is used for cluster generation, seeding and first strand synthesis can occur essentially simultaneously.

[0153] The methods of the present disclosure can further include generating clonal clusters, e.g., producing a plurality of amplification sites that each include a clonal population of amplicons derived from the target oligonucleotide originally present at each amplification site (as modified by inclusion of the tag region in the immobilized copy oligonucleotide).

[0154] In one embodiment, the method can include providing an amplification reagent and an array of amplification sites that include an immobilized nucleic acid. An amplification reagent can include (i) an array of populated amplification sites (e.g., amplification sites seeded with members of a sequencing library), (ii) nucleotide triphosphates (NTPs) including dATP, dTTP, dCTP, dGTP, and a dGTP analog, and (iii) a polymerase. The amplification sites are populated with an immobilized nucleic acid that is to be clonally amplified.

[0155] In some embodiments, the nucleic acid at each amplification site is a copy oligonucleotide comprising the incorporated tag region, and the clonal amplification includes amplification subsequent to first strand synthesis of the copy oligonucleotide. The amplification reagent is reacted to produce a plurality of populated amplification sites, where the plurality of populated amplification sites each include a clonal population of amplicons, where each clonal population is derived from the copy oligonucleotide comprising the incorporated tag region produced during the first strand synthesis step.

[0156] In some embodiments an array includes two populations of surface primers (e.g., capture nucleic acids) immobilized at amplification sites. In some embodiments the amplification sites of array include one population of a first surface primer (e.g., a first capture nucleic acid) immobilized thereto, and a second surface primer (e.g., a second capture nucleic acid) can be provided in solution during the reacting. In practice, there will be a plurality of identical first surface primers and / or a plurality of identical second surfaceprimers immobilized at the amplification sites, as the amplification process may require an excess of surface primers to sustain amplification.

[0157] As will be appreciated by the person of ordinary skill in the art, any given amplification reaction requires at least one type of forward primer and at least one type of reverse primer specific for the target oligonucleotide to be amplified. However, in certain embodiments the forward and reverse primers may include target-specific portions of identical sequence and may have entirely identical nucleotide sequence and structure (including any non-nucleotide modifications). In other words, it is possible to carry out amplification at amplification sites using only one type of primer, and such single-primer methods are encompassed within the scope of the disclosure. Other embodiments may use forward and reverse primers which contain identical target- specific sequences but which differ in some other structural features. For example, one type of primer may contain a non- nucleotide modification which is not present in the other.

[0158] The production of a plurality of populated amplification sites on an array typically occurs by amplification at each amplification site. The term "solid-phase amplification" as used herein refers to any nucleic acid amplification reaction carried out on or in association with an array such that all or a portion of the amplified products are immobilized at amplification sites on the array as they are formed. In particular, the term encompasses solid-phase polymerase chain reaction (solid-phase PCR) and solid phase isothermal amplification which are reactions analogous to standard solution phase amplification, except that one or both of the forward and reverse primers include amplification primers that are immobilized on the array. Solid phase PCR covers systems such as emulsions, where one primer is anchored to, for instance a bead, and the other is in free solution, and colony formation in solid phase gel matrices wherein one primer is anchored to the array and one is in free solution.

[0159] In one embodiment, a plurality of target oligonucleotides is used to prepare clustered arrays of nucleic acid colonies, analogous to those described in U.S. Pub. No. 2005 / 0100900, U.S. Pat. No. 7,115,400, WO 00 / 18957 and WO 98 / 44151 by solid-phase amplification, such as solid-phase isothermal amplification. The terms "cluster" and "colony" are used interchangeably herein to refer to a discrete site on a solid surface including a plurality of identical immobilized nucleic acid strands and a plurality of identicalimmobilized complementary nucleic acid strands. The term "clustered array" refers to an array formed from such clusters or colonies.

[0160] Clustered arrays can be prepared using either a process of thermocycling, as described in WO 98 / 44151, or a process where the temperature is maintained as a constant, and the cycles of extension and denaturing are performed using changes of reagents. Such isothermal amplification methods include, but are not limited to, bridge amplification and exclusion amplification (ExAmp, also referred to as kinetic exclusion amplification (KEA)). Isothermal amplification methods are described in patent application numbers WO 02 / 46456, U.S. Pub. No. 2008 / 0009420, U.S. Pat. No. 8,895,249, U.S. Pub No. 2013 / 0338042, and U.S. Pat. No. 9,169,513. Isothermal amplification by exclusion amplification may be used with, for instance, the Bsu (Bacillus subtilis) DNA polymerase or large fragment of Bsu.Isothermal amplification by bridge amplification may be used with, for instance, the Bst (Bacillus stearothermophilus) DNA polymerase. Optionally, the polymerase is deficient in 5' exonuclease activity, 3' exonuclease activity, or both activities. In some embodiments, cluster generation can be accomplished using commercially available machines such as the cBot (Illumina, San Diego, CA) and certain sequencing instruments such as iSeq 100, MiniSeq, NextSeq 550 Series, NextSeq 1000 & 2000, NovaSeq 6000 Series, and NovaSeq X Series (Illumina, San Diego, CA).

[0161] It will be appreciated that any of the amplification methodologies described herein or generally known in the art may be used with universal or target-specific surface primers to amplify immobilized DNA fragments. Suitable methods for amplification include, but are not limited to, the polymerase chain reaction (PCR), strand displacement amplification (SDA), transcription mediated amplification (TMA) and nucleic acid sequencebased amplification (NASBA), as described in U.S. Pat. No. 8,003,354. The amplification methods may be employed to amplify one or more nucleic acids of interest. For example, PCR, including multiplex PCR, SDA, TMA, NASBA and the like may be utilized to amplify immobilized DNA fragments. In some embodiments, surface primers directed specifically to the polynucleotide of interest are included in the amplification reaction.

[0162] Other suitable methods for amplification of target oligonucleotides may include oligonucleotide extension and ligation, rolling circle amplification (RCA) (Lizardi et al., Nat. Genet. 19:225-232 (1998)) and oligonucleotide ligation assay (OLA) (See generally U.S. Pat. Nos. 7,582,420, 5,185,243, 5,679,524 and 5,573,907; EP 0 320 308 Bl; EP 0 336 731 B l; EP0 439 182 Bl; WO 90 / 01069; WO 89 / 12696; and WO 89 / 09835) technologies. It will be appreciated that these amplification methodologies may be designed to amplify immobilized target oligonucleotides. For example, in some embodiments, the amplification method may include ligation probe amplification or oligonucleotide ligation assay (OLA) reactions that contain surface primers directed specifically to a nucleic acid of interest. In some embodiments, the amplification method may include a primer extension- ligation reaction that contains surface primers directed specifically to the nucleic acid of interest. As a nonlimiting example of primer extension and ligation primers that may be specifically designed to amplify a nucleic acid of interest, the amplification may include surface primers used for the GoldenGate assay (Illumina, Inc., San Diego, CA) as exemplified by U.S. Pat. No. 7,582,420 and 7,611,869.

[0163] DNA nanoballs can also be used in combination with methods described herein. Methods for creating and using DNA nanoballs for genomic sequencing can be found at, for example, US patents and publications U.S. Pat. No. 7,910,354, 2009 / 0264299, 2009 / 0011943, 2009 / 0005252, 2009 / 0155781, 2009 / 0118488 and as described in, for example, Drmanac et al. (2010, Science 327(5961): 78-81). Briefly, following production of modified target oligonucleotides, the modified target oligonucleotides are circularized and amplified by rolling circle amplification (Lizardi et al., 1998. Nat. Genet. 19:225-232; US 2007 / 0099208 Al). The extended concatemeric structure of the amplicons promotes coiling creates compact DNA nanoballs. The DNA nanoballs can be captured on substrates, preferably to create an ordered or patterned array such that distance between each nanoball is maintained thereby allowing sequencing of the separate DNA nanoballs. In some embodiments such as those used by Complete Genomics (Mountain View, Calif.), consecutive rounds of adapter addition, amplification, and digestion are carried out prior to circularization to produce head to tail constructs having several target oligonucleotides separated by adapter sequences.

[0164] Exemplary isothermal amplification methods that may be used in a method of the present disclosure include, but are not limited to, Multiple Displacement Amplification (MDA) as exemplified by, for example Dean et al., Proc. Natl. Acad. Sci. USA 99:5261-66 (2002) or isothermal strand displacement nucleic acid amplification exemplified by, for example U.S. Pat. No. 6,214,587. Other non-PCR-based methods that may be used in the present disclosure include, for example, strand displacement amplification (SDA) which isdescribed in, for example Walker et al., Molecular Methods for Virus Detection, Academic Press, Inc., 1995; U.S. Pat. Nos. 5,455,166, and 5,130,238, and Walker et al., Nucl. Acids Res. 20:1691-96 (1992) or hyper-branched strand displacement amplification which is described in, for example Lage et al., Genome Res. 13:294-307 (2003). Isothermal amplification methods may be used with, for instance, the strand-displacing Phi 29 polymerase or Bst DNA polymerase large fragment, 5'->3' exo- for random primer amplification of genomic DNA. The use of these polymerases takes advantage of their high processivity and strand displacing activity. High processivity allows the polymerases to produce fragments that are 10-20 kb in length. As set forth herein, smaller fragments may be produced under isothermal conditions using polymerases having low processivity and stranddisplacing activity such as Klenow polymerase. Additional description of amplification reactions, conditions and components are set forth in detail in the disclosure of U.S. Patent No. 7,670,810.

[0165] In some embodiments, amplification sites in an array can be, but need not be, entirely clonal. Rather, for some applications, an individual amplification site can be predominantly populated with amplicons from a first modified target oligonucleotide and can also have a low level of contaminating amplicons from a second modified target oligonucleotide. An array can have one or more amplification sites that have a low level of contaminating amplicons so long as the level of contamination does not have an unacceptable impact on a subsequent use of the array. For example, when the array is to be used in a detection application, an acceptable level of contamination would be a level that does not impact signal to noise or resolution of the detection technique in an unacceptable way. Accordingly, apparent clonality will generally be relevant to a particular use or application of an array made by the methods set forth herein. Exemplary levels of contamination that can be acceptable at an individual amplification site for particular applications include, but are not limited to, at most 0.1%, 0.5%, 1%, 5%, 10% or 25% contaminating amplicons. An array can include one or more amplification sites having these exemplary levels of contaminating amplicons. For example, up to 5%, 10%, 25%, 50%, 75%, or even 100% of the amplification sites in an array can have some contaminating amplicons. It will be understood that in an array or other collection of sites, at least 50%, 75%, 80%, 85%, 90%, 95% or 99% or more of the sites can be clonal or apparently clonal.

[0166] An amplification reagent can include further components that facilitate amplicon formation, and in some cases increase the rate of amplicon formation. An example is a recombinase in isothermal reactions including exclusion amplification. A mixture of recombinase and single-stranded binding (SSB) protein is particularly useful as SSB can further facilitate amplification. Exemplary formulations for recombinase-facilitated amplification include those sold commercially as TwistAmp kits by TwistDx (Cambridge, UK). Useful components of recombinase-facilitated amplification reagent and reaction conditions are set forth in US 5,223,414 and US 7,399,590.

[0167] Another example of a component that can be included in an amplification reagent to facilitate amplicon formation and in some cases to increase the rate of amplicon formation is a helicase. Exemplary formulations for helicase-facilitated amplification include those sold commercially as IsoAmp kits from Biohelix (Beverly, MA). Further, examples of useful formulations that include a helicase protein are described in US 7,399,590 and US 7,829,284.

[0168] Yet another example of a component that can be included in an amplification reagent to facilitate amplicon formation and in some cases increase the rate of amplicon formation is an origin binding protein.

[0169] The presence of molecular crowding reagents in the solution can be used to aid exclusion amplification. Examples of useful molecular crowding reagents include, but are not limited to, polyethylene glycol (PEG), Ficoll®, dextran, or polyvinyl alcohol. Exemplary molecular crowding reagents and formulations are set forth in U.S. Pat. No. 7,399,590.

[0170] The rate at which an amplification reaction occurs can be increased by increasing the concentration or amount of one or more of the active components of an amplification reaction. For example, the amount or concentration of polymerase, nucleotide triphosphates, primers, recombinase, helicase or SSB can be increased to increase the amplification rate. In some cases, the one or more active components of an amplification reaction that are increased in amount or concentration (or otherwise manipulated in a method set forth herein) are non- nucleic acid components of the amplification reaction.

[0171] Amplification rate can also be increased in a method set forth herein by adjusting the temperature. For example, the rate of amplification at one or more amplification sites can be increased by increasing the temperature at the site(s) up to a maximum temperature where reaction rate declines due to denaturation or other adverse events. Optimal or desiredtemperatures can be determined from known properties of the amplification components in use or empirically for a given amplification reaction mixture. Such adjustments can be made based on a priori predictions of primer melting temperature (Tm) or empirically.

[0172] The rate at which an amplification reaction occurs can be increased by increasing the activity of one or more amplification reagent. For example, a cofactor that increases the extension rate of a polymerase can be added to a reaction where the polymerase is in use. In some embodiments, metal cofactors such as magnesium, zinc or manganese can be added to a polymerase reaction or betaine can be added.

[0173] In some embodiments of the methods set forth herein, it is desirable to use a population of target oligonucleotides that is double-stranded. It has been observed that amplicon formation at an array of sites under exclusion amplification conditions is efficient for double-stranded target oligonucleotides. For example, a plurality of amplification sites having clonal populations of amplicons can be more efficiently produced from double-stranded target oligonucleotides (compared to single-stranded target oligonucleotides at the same concentration) in the presence of recombinase and single-stranded binding protein. Nevertheless, it will be understood that single-stranded target oligonucleotides can be used in some embodiments of the methods set forth herein.Methods of Sequencing

[0174] An array of the present disclosure, for example, having been produced by a method set forth herein and including amplified copy and / or modified nucleic acids at amplification sites, can be used for any of a variety of applications. A particularly useful application is nucleic acid sequencing. One example is sequencing-by-synthesis (SBS). In SBS, extension of a nucleic acid surface primer along a nucleic acid template (e.g., a target oligonucleotide or amplicon thereof) is monitored to determine the sequence of nucleotides in the template. The underlying chemical process can be polymerization (e.g., as catalyzed by a polymerase enzyme). In a particular polymerase-based SBS embodiment, fluorescently labeled nucleotides are added to a surface primer (thereby extending the surface primer) in a template dependent fashion such that detection of the order and type of nucleotides added to the surface primer can be used to determine the sequence of the template. A plurality of different templates at different sites of an array set forth herein can be subjected to an SBS technique under conditions where events occurring for different templates can be distinguished due to their location in the array. Examples of DNA polymerases useful for sequencing include, but are not limited to,polymerases described in U.S. Patent No. 11,104,888, U.S. Pat. No. 11,001,816, U.S. Pat. Appl. No. 18 / 373,620; U.S. Published Patent Application No. 2023 / 0047225.

[0175] Flow cells provide a convenient format for housing an array that is produced by the methods of the present disclosure and that is subjected to an SBS or other detection technique that involves repeated delivery of reagents in cycles. For example, to initiate a first SBS cycle, one or more labeled nucleotides, DNA polymerase, etc., can be flowed into / through a flow cell that houses an array of nucleic acid templates. Those sites of an array where surface primer extension causes a labeled nucleotide to be incorporated can be detected. Optionally, the nucleotides can further include a reversible termination property that terminates further surface primer extension once a nucleotide has been added to a surface primer. For example, a nucleotide analog having a reversible terminator moiety can be added to a surface primer such that subsequent extension cannot occur until a deblocking agent is delivered to remove the moiety. Thus, for embodiments that use reversible termination, a deblocking reagent can be delivered to the flow cell (before or after detection occurs). Washes can be carried out between the various delivery steps. The cycle can then be repeated n times to extend the surface primer by n nucleotides, thereby detecting a sequence of length n. Exemplary SBS procedures, fluidic systems and detection platforms that can be readily adapted for use with an array produced by the methods of the present disclosure are described, for example, in Bentley et al., Nature 456:53-59 (2008), WO 04 / 018497; U.S. Pat. No. 7,057,026; WO 91 / 06678; WO 07 / 123,744; U.S. Pat. No. 7,329,492; U.S. Pat. No. 7,211,414; U.S. Pat. No. 7,315,019; U.S. Pat. No. 7,405,281, and U.S. Pat. No. 8,343,746. Examplary nucleotides having a reversible termination property include modifications at the 3’-OH of the nucleotide sugar moiety, such as a 3'-O-azidomethyl blocking group -CH2N3, a 3'-OH acetal blocking group, or a 3'-OH thiocarbamate blocking group (U.S. Patent No. 1 1 ,293,061 ; U.S. Published Patent Application No. 2022 / 0396832).

[0176] Other sequencing procedures that use cyclic reactions can be used, such as pyrosequencing. Pyrosequencing detects the release of inorganic pyrophosphate (PPi) as particular nucleotides are incorporated into a nascent nucleic acid strand (Ronaghi, et al., Analytical Biochemistry 242(1), 84-9 (1996); Ronaghi, Genome Res. 11(1), 3-11 (2001); Ronaghi et al. Science 281(5375), 363 (1998); U.S. Pat. No. 6,210,891; U.S. Pat. No. 6,258,568 and U.S. Pat. No. 6,274,320). In pyrosequencing, released PPi can be detected by being immediately converted to adenosine triphosphate (ATP) by ATP sulfurylase, and the level ofATP generated can be detected via luciferase-produced photons. Thus, the sequencing reaction can be monitored via a luminescence detection system. Excitation radiation sources used for fluorescence-based detection systems are not necessary for pyrosequencing procedures. Useful fluidic systems, detectors and procedures that can be used for application of pyrosequencing to arrays of the present disclosure are described, for example, in WIPO Published Pat. App. 2012 / 058096, US 2005 / 0191698 Al, U.S. Pat. No. 7,595,883, and U.S. Pat. No. 7,244,559.

[0177] Sequencing-by-ligation reactions are also useful including, for example, those described in Shendure et al. Science 309:1728-1732 (2005); U.S. Pat. No. 5,599,675; and U.S. Pat. No. 5,750,341. Some embodiments can include sequencing-by-hybridization procedures as described, for example, in Bains et al., Journal of Theoretical Biology 135(3), 303-7 (1988); Drmanac et al., Nature Biotechnology 16, 54-58 (1998); Fodor et al., Science 251(4995), 767- 773 (1995); and WO 1989 / 10977. In both sequencing-by-ligation and sequencing-by- hybridization procedures, template nucleic acids (e.g., a target oligonucleotide or amplicons thereof) that are present at sites of an array are subjected to repeated cycles of oligonucleotide delivery and detection. Fluidic systems for SBS methods as set forth herein or in references cited herein can be readily adapted for delivery of reagents for sequencing-by-ligation or sequencing-by-hybridization procedures. Typically, the oligonucleotides are fluorescently labeled and can be detected using fluorescence detectors similar to those described with regard to SBS procedures herein or in references cited herein.

[0178] Some embodiments can use methods involving the real-time monitoring of DNA polymerase activity. For example, nucleotide incorporations can be detected through fluorescence resonance energy transfer (FRET) interactions between a fluorophore-bearing polymerase and > -phosphate-labeled nucleotides, or with zeromode waveguides (ZMWs). Techniques and reagents for FRET-based sequencing are described, for example, in Levene et al. Science 299, 682-686 (2003); Lundquist et al. Opt. Lett. 33, 1026-1028 (2008); Korlach et al. Proc. Natl. Acad. Sci. USA 105, 1176-1181 (2008).

[0179] Some SBS embodiments include detection of a proton released upon incorporation of a nucleotide into an extension product. For example, sequencing based on detection of released protons can use an electrical detector and associated techniques that are commercially available from Ion Torrent (Guilford, Conn., a Life Technologies subsidiary) or sequencing methods and systems described in US 2009 / 0026082 Al; US 2009 / 0127589 Al; US 2010 / 0137143 Al; or US 2010 / 0282617 AL Methods set forth herein for amplifying targetoligonucleotides using exclusion amplification can be readily applied to substrates used for detecting protons. More specifically, methods set forth herein can be used to produce clonal populations of amplicons at the sites of the arrays that are used to detect protons.

[0180] Sequencing of templates in a cluster often includes the technique of "paired-end" or "pairwise" sequencing (U.S. Pat. No. 7,754,429 and U.S. Pat. No. 8,017,335). Paired-end sequencing is a multi-step process that allows the determination of two "reads" of sequence by sequencing both strands of a double stranded nucleic acid. The advantage of the paired- end approach is that there is significantly more information to be gained from sequencing bases from two complementary templates than from sequencing the same number of bases from each of two independent templates in a random fashion. With the use of appropriate software tools for the assembly of sequence information, it is possible to use the knowledge that the "paired-end" sequences are not completely random, but are known to occur on a single template, and are therefore linked or paired in the genome. This information greatly aids the assembly of whole genome sequences into a consensus sequence.

[0181] After production of clonal clusters, each cluster includes immobilized complementary strands. In order to provide more suitable templates for sequencing, substantially all or at least a portion of one of the immobilized strands is removed in order to generate a template which is at least partially single-stranded. The portion of the template which is single-stranded will thus be available for hybridization to a sequencing primer. The process of removing all or a portion of one immobilized strand is referred to as "linearization." There are various ways for linearization, including but not limited to enzymatic cleavage (e.g., uracil DNA glycosylase (UDG) and endonuclease VII, oxoguanine glycosylase, chemical cleavage (e.g., palladium reagents and Pd linearization, nickel reagents and Ni Pd linearization), photo-chemical cleavage. Non-limiting examples of linearization methods are disclosed in US Serial No. 18 / 473,971, filed Sep. 25, 2023; PCT Publication No. WO 2019 / 222264; US Published Patent Application No. 2019 / 0352327; WO 2007 / 010251; US Patent Application Publication No. 2009 / 0088327; and in US. Patent Publication No. 2009 / 0118128, which are incorporated by reference in their entireties.

[0182] Sequence data can be obtained from both immobilized complementary strands by performing a linearization to remove a strand attached by one surface primer, e.g., P5, obtaining a sequence read from the remaining first strand using a surface primer, copying the first strand using immobilized surface primers for strand resynthesis and repopulation of thecluster with the strand initially removed by the first linearization, releasing the first strand and sequencing the second, copied strand. In one embodiment, resynthesis and repopulation of clusters includes use of a resynthesis reagent. An resynthesis reagent can include (i) an array of amplification sites, where each amplification site includes immobilized modified target oligonucleotides, (ii) nucleotide triphosphates (dNTPs), wherein the NTPs include dATP, dTTP, dCTP, dGTP, and a dGTP analog, and (iii) a polymerase. The resynthesis reagent is reacted to produce, at each amplification site, a population of strands that are complementary to the strand sequenced during the first round. The population of complementary strands are sequenced during the second round.

[0183] The present disclosure provides integrated sequencing systems capable of making an array using one or more of the methods set forth herein. An integrated sequencing system can be capable of detecting nucleic acids on the arrays using techniques such as those described herein. Thus, an integrated sequencing system of the present disclosure can include fluidic components capable of delivering amplification reagents to an array of amplification sites such as pumps, valves, reservoirs, fluidic lines and the like. An example of useful fluidic components includes a flow cell and a cartridge. A flow cell can be configured and / or used in an integrated sequencing system to create an array of the present disclosure and to detect the array. Exemplary flow cells are described, for example, in US 2010 / 0111768 Al and U.S. Pat. No. 8,951,781. A cartridge can be configured to include the components of an amplification or resynthesis reagent in one or more chambers. As exemplified for flow cells, one or more of the fluidic components of an integrated sequencing system can be used for an amplification method and for a detection method. Taking a nucleic acid sequencing embodiment as an example, one or more of the fluidic components of an integrated sequencing system can be used for an amplification method set forth herein and for the delivery of sequencing reagents in a sequencing method, including a resynthesis method, such as those described herein. Alternatively, an integrated sequencing system can include separate fluidic systems to carry out amplification methods and to carry out detection methods and resynthesis methods. Examples of integrated sequencing systems that are capable of creating arrays of nucleic acids and also determining the sequence of the nucleic acids include, without limitation, the MiSeq™, HiSeq™, NextSeq™, MiniSeq™, NovaSeq™ and iSeq™ platforms (Illumina, Inc., San Diego, Calif.) and devices described in U.S. Pat. No. 8,951,781. Such devices can be modified to make arrays using exclusion amplification in accordance with the guidance set forth herein.

[0184] A system capable of carrying out a method set forth herein need not be integrated with a detection device. Rather, a stand-alone system or a system integrated with other devices is also possible. Fluidic components similar to those exemplified herein in the context of an integrated sequencing system can be used in such embodiments.

[0185] A system capable of carrying out a method set forth herein, whether integrated with detection capabilities or not, can include a system controller that is capable of executing a set of instructions to perform one or more steps of a method, technique or process set forth herein. For example, the instructions can direct the performance of steps for creating an array under exclusion amplification conditions. Optionally, the instructions can further direct the performance of steps for detecting nucleic acids using methods set forth previously herein. A useful system controller may include any processor-based or microprocessor-based system, including systems using microcontrollers, reduced instruction set computers (RISC), application specific integrated circuits (ASICs), field programmable gate array (FPGAs), logic circuits, and any other circuit or processor capable of executing functions described herein. A set of instructions for a system controller may be in the form of a software program. As used herein, the terms “software” and “firmware” are interchangeable, and include any computer program stored in memory for execution by a computer, including RAM memory, ROM memory, EPROM memory, EEPROM memory, and non-volatile RAM (NVRAM) memory. The software may be in various forms such as system software or application software. Further, the software may be in the form of a collection of separate programs, or a program module within a larger program or a portion of a program module. The software also may include modular programming in the form of object-oriented programming.

[0186] Several applications for arrays of the present disclosure have been exemplified herein in the context of ensemble detection, wherein multiple amplicons present at each amplification site are detected together. In alternative embodiments, a single nucleic acid, whether a target oligonucleotideor amplicon thereof, can be detected at each amplification site. For example, an amplification site can be configured to contain a single nucleic acid molecule having a target nucleotide sequence that is to be detected and a plurality of filler nucleic acids. In this example, the filler nucleic acids function to fill the capacity of the amplification site and they are not necessarily intended to be detected. The single molecule that is to be detected can be detected by a method that is capable of distinguishing the single molecule in the background of the filler nucleic acids. Any of a variety of single molecule detection techniques can be usedincluding, for example, modifications of the ensemble detection techniques set forth herein to detect the sites at increased gain or using more sensitive labels. Other examples of single molecule detection methods that can be used are set forth in U.S. 2011 / 0312529 Al; U.S. Pat. No. 9,279,154; and U.S. 2013 / 0085073 Al.

[0187] It will be understood that an array of the present disclosure, for example, having been produced by a method set forth herein, need not be used for a detection method. Rather, the array can be used to store a nucleic acid library. Accordingly, the array can be stored in a state that preserves the nucleic acids therein. For example, an array can be stored in a desiccated state, frozen state (e.g., in liquid nitrogen), or in a solution that is protective of nucleic acids. Alternatively, or additionally, the array can be used to replicate a nucleic acid library. For example, an array can be used to create replicate amplicons from one or more of the sites on the array.

[0188] Several embodiments of the disclosure have been exemplified herein with regard to transporting target oligonucleotides to amplification sites of an array and making copies of the captured target oligonucleotides at the amplification sites. Similar methods can be used for non-nucleic acid target molecules. Thus, methods set forth herein can be used with other target molecules in place of the exemplified target oligonucleotides. For example, a method of the present disclosure can be carried out to transport individual target molecules from a population of different target molecules. Each target molecule can be transported to (and in some cases captured at) an individual amplification site of an array to initiate a reaction at the site of capture. The reaction at each site can, for example, produce copies of the captured molecule or the reaction can alter the site to isolate or sequester the captured molecule. In either case, the end result can be sites of the array that are each pure with respect to the type of target molecule that is present from a population that contained different types of target molecules.

[0189] The invention is defined in the claims. However, below there is provided a non- exhaustive listing of non-limiting exemplary aspects. Any one or more of the features of these aspects may be combined with any one or more features of another example, embodiment, or aspect described herein.EXAMPLES

[0190] The present disclosure is illustrated by the following examples. It is to be understood that the particular examples, materials, amounts, and procedures are to be interpreted broadly in accordance with the scope and spirit of the disclosure as set forth herein.Example 1: Determining that homopolymer SSEs arise in the clustering stage as opposed to the SBS stage

[0191] The source of sequencing errors due to homopolymers was investigated using target oligonucleotides with 10 base pairs (bp), 20bp, and 30bp homopolymers, and a 20 bp homopolymer with a 1 base substitution in the middle of the homopolymer. The target oligonucleotides were (i) hybridized to a flow cell at high concentration and sequenced, or (ii) hybridized to a flow cell, clustered using ExAmp (representing standard clustering procedure), and then sequenced via X-Leap SBS using an Illumina, Inc. cBot sequencing apparatus. Resolution value (detected intensity / expected intensity) was calculated at the first cycle following the homopolymer sequence. As shown in FIG. 9, resolution was lower as the length of the homopolymer increased. Resolution was substantially lower for the clustered, relative to the hybridized and unclustered, target oligonucleotides with lObp, 20bp, and 30bp homopolymers, suggesting that a majority of the SSEs due to the homopolymers result from amplification during clustering. Higher resolution was observed when sequencing the target oligonucleotide having the 20 bp homopolymer with a 1 base substitution in the middle of the homopolymer in both the clustered and hybridized scenarios, illustrating that introducing breaks in SSE sequences, such as homopolymers, improves the ability to sequence through these regions with fewer errors - further reiterating the effect of the tagging technology described herein.Example 2: Proof of concept that disrupting SSE region can improve sequencing resolution

[0192] Target oligonucleotides having g-quad regions or polyT homopolymer regions that were uninterrupted or interrupted by substitutions or insertions as shown in FIG. 10 were prepared, seeded, clustered, and sequenced using an Illumina, Inc. cBot sequencing apparatus. Resolution was determined within the g-quad region or the cycle following the homopolymer. Results are show in FIG. 10, which shows improvement in resolution in all the target oligonucleotides in which the SSE region was disrupted. Larger interruptions ordisruptions (e.g., 5 bp substitution) resulted in better resolution than smaller interruptions (1 bp substitution), and insertions resulted in better resolution that substitution. Again, the results further illustrate disruption of an SSE sequence according to the methods described herein can result in substantial improvement of sequencing resolution within or downstream of the SSE.Example 3: Determining the effect of poly(A) tags and noly(T) tags

[0193] Customized tagging technology was transferred to a sequencer. Tagging oligonucleotides were introduced to the flow cell in separate stages to increase the efficiency of hybridization to the genomic library while decreasing the hybridization of tagging oligonucleotides with each other due to their complimentary nature. The experimental protocol is schematically shown in FIG. 11. Poly(A) tags were loaded first at a concentration of 0.2 pM, at a temperature ramp of 60°C to 20°C, and over the course of 3 x 6 minute pushes. Poly(T) tags were subsequently loaded at a concentration of 0.2 pM, at a temperature ramp of 50°C to 20°C, and over the course of 3 x 6 minute pushes. The first strand was then extended for 15 minutes at 30°C and 15 minutes at 50°C with 100 pL ELM4 Reagent.

[0194] An assortment of tagging oligonucleotides, including a variety of substitution and insertion lengths, were tested (FIG. 12). Tagging oligonucleotides including flanks around the SSE region were tested alongside tagging oligonucleotides with no flanks. FIG. 13 shows proof of concept sequencing data on a BacPack library with tagging oligonucleotides including a 1-base substitution and tagging oligonucleotides including a 5 -base insertion. The BacPack location was RP11-150M8: 129, 631-129, 820. FIG. 14 shows proof of concept sequencing data on a BacPack library with a 29mer tagging oligonucleotide containing no flanks and a 1-base substitution. The BacPack location was RP11-719L22: 11777-11811. All data was generated on a NextSeq2k with X-LEAP SBS Chemistry.

[0195] When compared to the control baseline, the 1-base and 5-base substitution sequencing runs showed improved quality in the region downstream of the homopolymer in the reads where the tagging oligonucleotide had been integrated. This is shown in FIG. 13 by sorting the aligned reads by modification at a given location and observing the decrease in mismatch in the downstream region. Furthermore, as shown in the SSE Scan files associated with each of these runs, the targeted homopolymer is no longer identified as an SSE in the 1-base and 5-base substitution sequencing runs, further indicating an improvement in quality at this location (FIG. 13).

[0196] Due to the flanks on the tagging oligonucleotide, the insertion or substitution sits at a defined location in the genome and, consequently, the Integrative Genomic Viewer (IGV) was able to identify this modification easily. Tagging oligonucleotides with flanks are appropriate for IGV analysis, increase hybridization efficiency, and are desired in some scenarios. However, tagging oligonucleotides for whole genome homopolymer targeting would not contain flanks to prevent the need to design an impractical number of custom- made oligonucleotides.

[0197] In the run of FIG. 14, the homopolymer is 34bp long and therefore, depending on where the 29mer tagging oligonucleotide hybridizes, the substitution can sit across a 5bp window as seen in the IGV shot (FIG. 14). As herein disclosed, reads which have successfully integrated the tag show an improvement in quality downstream of the homopolymer region.

[0198] To assess the overall effect of tagging across multiple locations in BacPack, global analysis was completed on the sequencing runs. Percentage mismatch was calculated in regions downstream of homopolymers and compared across the different runs. FIG. ISA summarizes the correlation between homopolymer size and percentage mismatch (in the 20bp downstream region) for baseline and tagging sequencing runs. FIG. 15B shows the percentage Q30 20bp downstream of the homopolymer. Analysis was completed on down- sampled BAM file, and metrics were averaged across all reads, although not all reads contained the tag.

[0199] When comparing the baseline and tagging runs across a variety of different BacPack homopolymer locations, a benefit in percentage mismatch is seen when using the tagging technology. This becomes prominent at the 29bp length point (FIG. ISA). This is the length of the tagging oligonucleotide. Therefore, only homopolymers above this length would be affected. Because the analysis pipeline is unable to isolate only the reads containing the tag and instead calculates an average percentage mismatch across all reads, the effect of tagging when using this analysis pipeline is diluted. Nevertheless, a clear benefit to percentage mismatch can still be seen when using the tagging technology.

[0200] FIG. 16 shows the efficiency of the tagging technology as a proportion of overlapping reads that include the tag, while FIG. 17 shows the ability of IGV to resolve different tagging oligonucleotide designs.Example 4: Human library sequencing

[0201] Proof of concept sequencing data was obtained on a human library with the same 29mer tagging oligonucleotide with a 1-base substitution used to generate the BacPack data from Example 3, and while using NextSeq2000 with X-LEAP SBS Chemistry under different conditions (FIG. 18). The Bed file contains 393,986 A and T homopolymers above 15bp long. FIG. 19 shows that increasing the concentration of the tagging oligonucleotide 10-fold results in a marginal improvement in hybridization efficiency. FIG. 20 shows a pipeline analysis of percentage soft clipping averaged across all reads in a region 20bp downstream of a homopolymer. The data in both these figures show that global trends in percentage soft clipping depend on hybridization efficiency.

[0202] Integrating the tagging oligonucleotide into the human genome leads to an improvement in quality in the region downstream of the homopolymer. A clear decrease in percentage soft clipping for the tagging runs when compared to the baseline is shown across multiple homopolymer locations in the human library, and is dependent on homopolymer size (FIG. 20). FIG. 21 compares percentage soft clipping between A homopolymers and T homopolymers. As herein mentioned, this analysis pipeline takes an average across all reads. Therefore, while the effect of tagging is diluted, it still shows an improvement to quality metrics. This technique is able to improve secondary metrics, such as decreasing sequence specific errors, whilst having no effect on primary metrics.

[0203] The complete disclosure of all patents, patent applications, and publications, and electronically available material (including, for instance, nucleotide sequence submissions in, e.g., GenBank and RefSeq, and amino acid sequence submissions in, e.g., SwissProt, PIR, PRF, PDB, and translations from annotated coding regions in GenBank and RefSeq) cited herein are incorporated by reference in their entirety. Supplementary materials referenced in publications (such as supplementary tables, supplementary figures, supplementary materials and methods, and / or supplementary experimental data) are likewise incorporated by reference in their entirety. In the event that any inconsistency exists between the disclosure of thepresent application and the disclosure(s) of any document incorporated herein by reference, the disclosure of the present application shall govern. The foregoing detailed description and examples have been given for clarity of understanding only. No unnecessary limitations are to be understood therefrom. The disclosure is not limited to the exact details shown and described, for variations obvious to one skilled in the art will be included within the disclosure defined by the claims.

[0204] Unless otherwise indicated, all numbers expressing quantities of components, molecular weights, and so forth used in the specification and claims are to be understood as being modified in all instances by the term "about." Accordingly, unless otherwise indicated to the contrary, the numerical parameters set forth in the specification and claims are approximations that may vary depending upon the desired properties sought to be obtained by the present disclosure. At the very least, and not as an attempt to limit the doctrine of equivalents to the scope of the claims, each numerical parameter should at least be construed in light of the number of reported significant digits and by applying ordinary rounding techniques.

[0205] Notwithstanding that the numerical ranges and parameters setting forth the broad scope of the disclosure are approximations, the numerical values set forth in the specific examples are reported as precisely as possible. All numerical values, however, inherently contain a range necessarily resulting from the standard deviation found in their respective testing measurements.

[0206] All headings are for the convenience of the reader and should not be used to limit the meaning of the text that follows the heading, unless so specified.

Claims

CLAIMSWhat is claimed is:

1. An oligonucleotide sequencing method or a method for preparing a solid surface for oligonucleotide sequencing, the method comprising: hybridizing a target oligonucleotide to a surface primer bound to a solid surface, wherein the target oligonucleotide has a target region, wherein the surface primer has a free 3’ end; hybridizing a tagging oligonucleotide to the target region of the target oligonucleotide hybridized to the surface primer, wherein the tagging oligonucleotide comprises a sequence that differs from the complementary sequence of the target region, wherein the tagging oligonucleotide comprises a tag region between a 3’ region and a 5’ region, wherein the 3’ region has a sequence complementary to a sequence of a 5’ portion of the target region, wherein the 5’ region has a sequence complementary to a sequence of a 3’ portion of the target region; and synthesizing a copy oligonucleotide by (i) extending the surface primer from the 3’ end and using the target oligonucleotide as a template, (ii) incorporating the tagging oligonucleotide into the copy oligonucleotide, and (iii) extending the tagging oligonucleotide from a 3’ end using the target oligonucleotide to produce the copy oligonucleotide having a sequence complementary to the target oligonucleotide except for in the tag region.

2. The method of claim 1, wherein extending the surface primer from the 3’ end and using the target oligonucleotide as a template comprises extending the 3 ’ end of the surface primer to a location at which the 5 ’ end of the tagging oligonucleotide hybridizes to the target region to produce a first nascent strand, and wherein incorporating the tagging oligonucleotide into the copy oligonucleotide comprises ligating the 3’ end of the nascent first strand to the 5’ end of the tagging oligonucleotide.

3. The method of claim 2, further comprising: hybridizing a blocking oligonucleotide to the target region of the target oligonucleotide hybridized to the surface primer, wherein the blocking oligonucleotide comprises a sequence complementary to the sequence of the target region, wherein the blocking oligonucleotide comprises a 5 ’ region having a sequence identical to the 5’ region of the tagging oligonucleotide, wherein the 3’ end of the surface primer is extended to the location at which the 5 ’ end of the tagging oligonucleotide hybridizes to the target region while the blocking oligonucleotide is hybridized to target region of the target oligonucleotide; and removing the blocking oligonucleotide from the target region of the target oligonucleotide prior to hybridizing the tagging oligonucleotide to the target region of the target oligonucleotide.

4. The method of claim 2 or 3, wherein extending the tagging oligonucleotide from the 3’ end occurs after ligating the 3’ end of the first nascent strand to the 5’ end of the tagging oligonucleotide.

5. The method of any one of claims 1 to 4, further comprising amplifying the copy oligonucleotide to form a cluster of oligonucleotides on the solid surface, wherein the cluster of oligonucleotides comprises multiple copies of the copy oligonucleotide bound to the solid surface and multiple copies of a modified target oligonucleotide bound to the solid surface, wherein the modified oligonucleotide is complementary to the copy oligonucleotide.

6. The method of claim 5, further comprising removing from the solid surface the multiple copies of the copy oligonucleotide or the multiple copies of the modified target oligonucleotide.

7. The method of claim 6, further comprising sequencing the multiple copies of the copy oligonucleotide or the multiple copies of the modified target oligonucleotide that remain bound to the solid surface.

8. The method of any one of claims 1 to 7, wherein the one or more target regions comprise nucleotides within regions prone to produce sequence specific errors.

9. The method of any one of claims 1 to 8, further comprising: amplifying the copy oligonucleotide to form a cluster of oligonucleotides on the solid surface, wherein the cluster of oligonucleotides comprises multiple copies of the copy oligonucleotide bound to the solid surface and multiple copies of a modified target oligonucleotide bound to the solid surface, wherein the modified oligonucleotide is complementary to the copy oligonucleotide.

10. The method of claim 9, further comprising removing from the solid surface the multiple copies of the copy oligonucleotide or the multiple copies of the modified target oligonucleotide.

11. The method of claim 10, further comprising sequencing the multiple copies of the copy oligonucleotide or the multiple copies of the modified target oligonucleotide that remain bound to the solid surface.

12. The method of claim 10, wherein the target region is a region prone to sequence specific errors.

13. The method of claim 12, wherein sequencing resolution downstream of the tag or modified target region is improved.

14. The method of claim 12 or 13, wherein the region prone to sequence specific errors comprises a homopolymer sequence.

15. The method of claim 14, wherein the homopolymer sequence comprises a poly(T) or a poly(A) sequence.

16. The method of claim 12 or 13, wherein the region prone to sequence specific errors comprises a multi-nucleotide repeat region.

17. The method of claim 16, wherein the multi-nucleotide repeat region comprises a dinucleotide repeat region or a tri-nucleotide repeat region.

18. The method of claim 16, wherein the multi-nucleotide repeat region comprises GT, GA, AC, or CT repeats.

19. The method of claim 12 or 13, wherein the region prone to sequence specific errors comprises a G-quadruplex region.

20. The method of any one of claims 1 to 19, wherein the tag region consists of a single nucleotide.

21. The method of any one of claims 1 to 19, wherein the tag region comprises more than one nucleotide.