Nucleic acid molecule sequencing systems and methods
Patent Information
- Authority / Receiving Office
- EP · EP
- Patent Type
- Applications
- Current Assignee / Owner
- ULTIMA GENOMICS INC
- Filing Date
- 2024-06-28
- Publication Date
- 2026-05-06
AI Technical Summary
Current single molecule sequencing methods face challenges with reduced light availability for signal readout, leading to incomplete incorporation of nucleotides and polymerase stalling, which limits sequencing read length and accuracy.
The method involves alternating sequences of labeled and unlabeled nucleotide flows to minimize scarring and polymerase stalling, allowing for longer sequence reads by using 100% labeled nucleotides in bright flows and unlabeled nucleotides in dark flows, and potentially resequencing to confirm base calls and fill gaps.
This approach enhances sequencing accuracy and read length by reducing scarring and polymerase stalling, enabling more reliable and complete sequence determination of nucleic acid molecules.
Smart Images

Figure US2024036204_02012025_PF_FP_ABST
Abstract
Description
[0001] NUCLEIC ACID MOLECUEE SEQUENCING SYSTEMS AND METHODS
[0002] CROSS-REFERENCE
[0003] [1] This application claims the benefit of U.S. Provisional Pat. App. Nos. 63 / 511,612 filed on June 30, 2023, 63 / 581,542 filed September 8, 2023, 63 / 563,265 filed March 8, 2024, and 63 / 632,508 filed April 10, 2024, each of which is entirely incorporated by reference herein for all purposes.
[0004] BACKGROUND
[0005] [2] Biological sample processing has various applications in the fields of molecular biology and medicine (e.g., diagnosis). For example, nucleic acid sequencing may provide information that may be used to diagnose a certain condition in a subject and in some cases tailor a treatment plan. Sequencing is widely used for molecular biology applications, including vector designs, gene therapy, vaccine design, industrial strain design and verification. Biological sample processing may involve a fluidics system and / or a detection system.
[0006] [3] Single molecule sequencing offers increasingly granular and more accurate information For example, the ability to sequence a single molecule reduces the presence of errors introduced by amplification and makes possible the sequencing of molecules that are difficult to amplify (e g., due to high GC content, secondary structures). Further, amplification can be a lengthy process, and without that requirement overall sequencing throughput can be increased. However, less light may be available from single molecule sequencing for signal readout. There is thus a need in the art for improved single molecule sequencing methods.
[0007] SUMMARY
[0008] [4] Provided herein are systems and methods for nucleic acid molecule sequencing. The present disclosure may be advantageous to improve sequencing results.
[0009] [5] An aspect of the disclosure provides a method of sequencing a template molecule, comprising: hybridizing a primer to the template molecule to form a hybridized template; determining the sequence of a first region of the template molecule by extending the primer through the first region using labeled nucleotides; extending the pnmer through a second region of the template using unlabeled nucleotides; determining the sequence of a third
[0010] -1-
[0011] SUBSTITUTE SHEET (RULE 26) region of the template molecule by extending the primer through the third region using labeled nucleotides; determining the sequence of the second region of the template molecule by comparing the sequence of the first region and the third region to a reference sequence.
[0012] [6] In some embodiments, the determining (b) the sequence of the first region comprises, for each nucleotide flow in a first plurality of nucleotide flows: providing labeled nucleotides of a respective base type; and detecting a signal indicating a presence or absence of incorporation of labeled nucleotides into the extending primer.
[0013] [7] In some embodiments, the extending (c) the primer through the second region comprises, for each nucleotide flow in a second plurality of nucleotide flows: providing unlabeled nucleotides of a respective base type.
[0014] [8] In some embodiments, the extending (c) the primer through the second region comprises, for each nucleotide flow in a second plurality of nucleotide flows: providing unlabeled nucleotides of one or more base types.
[0015] [9] In some embodiments, the determining (d) the sequence of the third region comprises, for each nucleotide flow in a third plurality of nucleotide flows: providing labeled nucleotides of a respective base type; and detecting a signal indicating a presence or absence of incorporation or labeled nucleotides into the extending primer.
[0016]
[0010] In some embodiments, the determining (e) the sequence of the second region comprises aligning the first and third regions of the template molecule to a first and third regions of the reference sequence, respectively, wherein the first and third regions are separated from each other by a second region in the reference sequence.
[0017]
[0011] In some embodiments, the sequence of the second region of the template molecule is determined as the sequence of the second region of the reference genome.
[0018]
[0012] In some embodiments, the first plurality of nucleotide flows and third plurality of nucleotide flows comprise 50 nucleotide flows, respectively. In some embodiments, the second plurality of nucleotide flows comprises 10 nucleotide flows.
[0019]
[0013] In some embodiments, the method further comprises repeating the determining (b) and extending (c) one or more times prior to determining (d) and determining (e).
[0020]
[0014] In some embodiments, the template molecule is a nucleic acid.
[0021]
[0015] In another aspect, this disclosure provides a method of sequencing a template molecule, comprising: hybridizing a primer to the template molecule to provide a hybridized template; and performing primer extension by providing: i) one or more nucleotide flows comprising labeled nucleotides of a first base type, and ii) one or more nucleotide flows.
[0022] -2-
[0023] SUBSTITUTE SHEET (RULE 26)
[0016] In some embodiments, the primer extension (b) further comprises, in each of the one or more first nucleotide flows: providing labeled nucleotides of the first base type; and detecting a signal indicating a presence or absence of incorporation of a labeled nucleotide into the extending primer.
[0024]
[0017] In some embodiments, the method further comprises (c) determining the sequence of the template molecule by aligning the detected signals to a reference sequence.
[0025]
[0018] In some embodiments, the proportion of labeled nucleotides in the one or more first nucleotide flows is 100%.
[0026]
[0019] In some embodiments, the method further comprises removing the extended primer from the template molecule; hybridizing a primer to the template molecule to provide a hybridized template; performing primer extension in a second flow cycle, wherein the second flow cycle comprises one or more second nucleotide flows comprising nucleotides of a second base type different from the first base type; removing the extended primer from the template molecule; hybridizing a primer to the template molecule to provide a hybridized template; and performing primer extension in a third flow cycle, wherein the third flow cycle comprises one or more third nucleotide flows comprising nucleotides of a third base type different from the first and second base types.
[0027]
[0020] In some embodiments, nucleotides in one or more second nucleotide flows comprise labeled nucleotides, and wherein the proportion of labeled nucleotides in the one or more second nucleotide flows is 100%. In some embodiments, nucleotides in the one or more third nucleotide flows comprise labeled nucleotides, and wherein the proportion of labeled nucleotides in the one or more third nucleotide flows is 100%.
[0028]
[0021] In some embodiments, the first flow cycle, the second flow cycle, and the third flow cycle each comprises one or more additional nucleotide flows comprising unlabeled nucleotides.
[0029]
[0022] In some embodiments, the one or more additional nucleotide flows in the first, second, and third flow cycles do not comprise the first base type, the second base type, or the third base type, respectively.
[0030]
[0023] In some embodiments, the one or more additional nucleotide flows in the first, second, and third flow cycles comprise at least one nucleotide flow comprising labeled nucleotides, wherein the labeled nucleotides are not of the first base type, the second base type, or the third base type, respectively.
[0031] -3-
[0032] SUBSTITUTE SHEET (RULE 26)
[0024] In some embodiments, the method further comprises: removing the extended primer from the template molecule; hybridizing a primer to the template molecule to provide a hybridized template; and performing primer extension in a fourth flow cycle, wherein the fourth flow cycle comprises one or more fourth nucleotide flows comprising nucleotides of a fourth base type different from the first, second, and third base types.
[0033]
[0025] In some embodiments, the template molecule is a nucleic acid.
[0034]
[0026] In another aspect, this disclosure provides a method of determining a sequence of a template molecule, comprising: hybridizing a primer to the template molecule to form a hybridized template; determining the sequence of a first region of the template molecule by extending the primer through the first region using labeled nucleotides; removing the extended primer from the template molecule; extending the primer through the first region of the template molecule using unlabeled nucleotides; and determining the sequence of a second region of the template molecule by extending the primer through the second region using labeled nucleotides.
[0035]
[0027] In some embodiments, primer extension using labeled nucleotides further comprises, in each of a plurality of nucleotide flows: providing labeled nucleotides of a respective base type; and detecting a signal indicating a presence or absence of incorporation of a labeled nucleotide into the extending primer.
[0036]
[0028] In some embodiments, primer extension using unlabeled nucleotides further comprises, in each of a plurality of nucleotide flows: providing unlabeled nucleotides of at least one base type.
[0037]
[0029] In some embodiments, the template molecule is a nucleic acid.
[0038]
[0030] In another aspect, this disclosure provides a method of sequencing a template molecule, comprising: providing a template molecule hybridized to a primer and a polymerase coupled thereto; providing a replacement polymerase; extending the primer through a first region of the template molecule until primer extension stalls; subjecting the template complex to conditions sufficient to anneal the replacement polymerase thereto; and extending the primer through a second region of the template molecule.
[0039]
[0031] In some embodiments, primer extension stalling comprises decoupling of the polymerase from the template complex
[0040]
[0032] In some embodiments, prior to the subjecting (d) the replacement polymerase is coupled to a solid support via one or more cleavable adjacent to the template molecule.
[0041] -4-
[0042] SUBSTITUTE SHEET (RULE 26)
[0033] In some embodiments, the subjecting (d) comprises cleaving the one or more cleavable moieties to decouple the replacement polymerase from the solid support.
[0043]
[0034] In some embodiments, the method further comprises providing a plurality of replacement polymerases, wherein a first subset of the plurality of replacement polymerases is coupled to a first solid support and a second subset of the plurality' of replacement polymerases is coupled to a second solid support.
[0044]
[0035] In some embodiments, the solid support is a bead. In some embodiments, the template molecule is a nucleic acid.
[0045]
[0036] In another aspect, this disclosure provides a method for increasing sequencing read quality, comprising: receiving, at one or more processors, sequencing data comprising a plurality of sequencing reads generated by extending a sequencing primer through a region of interest in a target nucleic acid molecule using a plurality of sequencing flow steps, each sequencing flow step comprising combining a hybrid with nucleotides, the hybrid comprising the sequencing primer and a nucleic acid molecule comprising the region of interest, wherein at least a portion of the nucleotides are labeled, and detecting the presence or absence of an incorporated nucleotide; filtering the sequencing data, using the one or more processors, to remove sequencing reads for which an absence of an incorporated nucleotide was detected at six or more consecutive sequencing flow steps while retaining sequencing reads for which an absence of an incorporated nucleotide was detected at six or fewer consecutive sequencing flow steps, thereby generating filtered sequencing data; determining, using the one or more processors, for each sequencing flow step of each sequencing read, a read quality metric based on one or more homopolymer probability values other than a highest homopolymer probability value; and trimming the terminus of one or more sequencing reads in the sequencing data based on the read quality' metrics for a respective sequencing read, thereby generating trimmed sequencing data.
[0046]
[0037] In some embodiments, the method further comprises generating the sequencing data, wherein each sequencing read is obtained from a respective template nucleic acid molecule.
[0047]
[0038] In some embodiments, the method further comprises calling, using the one or more processors, one or more genetic vanants using the trimmed sequencing data.
[0048]
[0039] In some embodiments, the method further comprises trimming a known adapter sequence, or a portion thereof, from one or more sequencing reads in the sequencing data.
[0049]
[0040] In some embodiments, the read quality metnc for each sequencing flow step of each sequencing read is based on a second highest homopolymer probability value.
[0050] -5-
[0051] SUBSTITUTE SHEET (RULE 26)
[0041] In some embodiments, trimming the terminus of the one or more sequencing reads in the sequencing data based on the read quality metric, thereby generating the trimmed sequencing data, comprises, for each sequencing read: determining a read quality metric moving average for the sequencing flow steps; selecting a sequencing flow step, wherein the selected sequencing flow step is the nth sequencing flow step having a moving average above a predetermined threshold, wherein n is a predefined number; and trimming at least a portion of the sequencing read comprising the selected sequencing flow step.
[0052]
[0042] In some embodiments, a predetermined number of consecutive sequencing flow steps prior to the selected sequencing flow step are trimmed In some embodiments, the predetermined number of consecutive sequencing flow steps is a multiple of four.
[0053]
[0043] In some embodiments, the method further comprises storing the trimmed sequencing data in a non-transitory computer readable medium.
[0054]
[0044] In some embodiments, the method further comprises aligning sequencing reads in the trimmed sequencing data to respective reference sequences.
[0055]
[0045] In some embodiments, the nucleotides are non-terminating nucleotides.
[0056]
[0046] Another aspect of the present disclosure provides a system comprising one or more computer processors and computer memory coupled thereto. The computer memory comprises machine executable code that, upon execution by the one or more computer processors, implements any of the methods above or elsewhere herein. Another aspect of the present disclosure provides a non-transitory computer readable medium comprising machine executable code that, upon execution by one or more computer processors, implements any of the methods above or elsewhere herein.
[0057]
[0047] Additional aspects and advantages of the present disclosure will become readily apparent to those skilled in this art from the following detailed description, wherein only illustrative instances of the present disclosure are shown and described. As will be realized, the present disclosure is capable of other and different instances, and its several details are capable of modifications in various obvious respects, all without departing from the disclosure. Accordingly, the drawings and description are to be regarded as illustrative in nature, and not as restrictive.
[0058] INCORPORATION BY REFERENCE
[0059]
[0048] All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent,
[0060] -6-
[0061] SUBSTITUTE SHEET (RULE 26) or patent application was specifically and individually indicated to be incorporated by reference. To the extent publications and patents or patent applications incorporated by reference contradict the disclosure contained in the specification, the specification is intended to supersede and / or take precedence over any such contradictory material.
[0062] BRIEF DESCRIPTION OF THE DRAWINGS
[0063]
[0049] The novel features of the invention are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present invention will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the invention are utilized, and the accompanying drawings (also "Figure" and “FIG.” herein) of which:
[0064]
[0050] FIG. 1 illustrates an example workflow for processing a sample for sequencing.
[0065]
[0051] FIG. 2 illustrates an example flow sequencing method that can be used to generate the sequencing data described herein.
[0066]
[0052] FIG. 3 illustrates an example flowgram.
[0067]
[0053] FIG. 4 illustrates examples of individually addressable locations distributed on substrates, as described herein.
[0068]
[0054] FIGs. 5A-5B each illustrate multiplexed stations in a sequencing system.
[0069]
[0055] FIG. 6 illustrates a computer system that is programmed or otherwise configured to implement methods provided herein.
[0070]
[0056] FIG. 7 shows an example image of a substrate with a hexagonal lattice of beads, as described herein.
[0071]
[0057] FIGs. 8A-8H each illustrate exemplary sequencing methods for single molecule sequencing.
[0072]
[0058] FIG. 9A illustrates an exemplary plurality of sequencing reads, in accordance with some embodiments.
[0073]
[0059] FIG. 9B illustrates a filtered set of sequencing reads, in accordance with some embodiments.
[0074]
[0060] FIG. 9C illustrates a filtered and trimmed set of sequencing reads, in accordance with some embodiments
[0075]
[0061] FIG. 9D illustrates an exemplary method for increasing sequencing read quality, in accordance with some embodiments.
[0076] -7-
[0077] SUBSTITUTE SHEET (RULE 26)
[0062] FIG. 10A illustrates three consecutive sequencing flow steps that could hypothetically occur if there is incomplete incorporation of correct nucleotides into a homopolymer.
[0078]
[0063] FIG. 10B illustrates that three consecutive sequencing flow steps cannot all yield a signal of 0 in colony sequencing.
[0079]
[0064] FIG. IOC illustrates that more than three consecutive sequencing flow steps can yield a signal of 0 in single-molecule sequencing without requiring read trimming.
[0080]
[0065] FIG. 10D illustrates the read quality metrics for an exemplary sequencing read, in accordance with some embodiments.
[0081]
[0066] FIG. 11 illustrates a mixed-reversibly terminated sequencing method.
[0082]
[0067] FIG. 11B illustrates a mixed-color sequencing method.
[0083]
[0068] FIG. 11C illustrates a non-terminated sequencing method.
[0084]
[0069] FIG. 12A-12D illustrate different workflows for loading nucleic acids using beads as spacers.
[0085]
[0070] FIG. 12E-12H illustrate different workflows for loading nucleic acids using DNA nanoballs and / or DNA origami as spacers.
[0086]
[0071] FIGs. 12I-12L illustrate different workflows for loading nucleic acids using dendrimers as spacers.
[0087]
[0072] FIG. 13 illustrates an exemplary dual strand sequencing method
[0088]
[0073] FIGs. 14A-14C illustrate a plot of mean signal vs. homopolymer length for three sequencing runs.
[0089] DETAILED DESCRIPTION
[0090]
[0074] While various embodiments of the invention have been shown and described herein, it will be obvious to those skilled in the art that such embodiments are provided by way of example only. Numerous variations, changes, and substitutions may occur to those skilled in the art without departing from the invention. It should be understood that various alternatives to the embodiments of the invention described herein may be employed.
[0091]
[0075] As used herein, the singular forms “a,” “an,” and “the” include the plural reference unless the context clearly dictates otherwise.
[0092]
[0001] When a range of values is provided, it is to be understood that each intervening value between the upper and lower limit of that range, and any other stated or intervening value in that stated range is encompassed within the scope of the present disclosure. Where the stated
[0093] -8-
[0094] SUBSTITUTE SHEET (RULE 26) range includes upper or lower limits, ranges excluding either of those included limits are also included in the present disclosure.
[0095]
[0076] The term “analyte,” as used herein, generally refers to an object that is the subject of analysis, or an object, regardless of being the subject of analysis, that is directly or indirectly analyzed during a process. An analyte may be synthetic. An analyte may be, originate from, and / or be derived from, a sample, such as a biological sample. In some examples, an analyte is or includes a molecule, macromolecule (e.g., nucleic acid, carbohydrate, protein, lipid, etc.), nucleic acid, carbohydrate, lipid, antibody, antibody fragment, antigen, peptide, polypeptide, protein, macromolecular group (e g., glycoproteins, proteoglycans, ribozymes, liposomes, etc.), cell, tissue, biological particle, or an organism, or any engineered copy or variant thereof, or any combination thereof. The term “processing an analyte,” as used herein, generally refers to one or more stages of interaction with one more samples. Processing an analyte may comprise conducting a chemical reaction, biochemical reaction, enzymatic reaction, hybridization reaction, polymerization reaction, physical reaction, any other reaction, or a combination thereof with, in the presence of, or on, the analyte. Processing an analyte may comprise physical and / or chemical manipulation of the analyte. For example, processing an analyte may comprise detection of a chemical change or physical change, addition of or subtraction of material, atoms, or molecules, molecular confirmation, detection of the presence of a fluorescent label, detection of a Forster resonance energy transfer (FRET) interaction, or inference of absence of fluorescence.
[0096]
[0077] The term “biological sample,” as used herein, generally refers to any sample derived from a subject or specimen. The biological sample can be a fluid, tissue, collection of cells (e.g., cheek swab), hair sample, or feces sample. The fluid can be blood (e.g., whole blood), saliva, urine, or sweat. The tissue can be from an organ (e.g., liver, lung, or thyroid), or a mass of cellular material, such as, for example, a tumor. The biological sample can be a cellular sample or cell-free sample. Examples of biological samples include nucleic acid molecules, amino acids, polypeptides, proteins, carbohydrates, fats, or viruses. In an example, a biological sample is a nucleic acid sample including one or more nucleic acid molecules, such as deoxyribonucleic acid (DNA) and / or ribonucleic acid (RNA). The nucleic acid sample may comprise cell-free nucleic acid molecules, such as cell-free DNA or cell-free RNA. Non-limiting examples of nucleic acids include DNA, RNA, genomic DNA or synthetic DNA / RNA or coding or non-codmg regions of a gene or gene fragment, loci (locus) defined from linkage analysis, exons, introns, messenger RNA (mRNA), transfer RNA,
[0097] -9-
[0098] SUBSTITUTE SHEET (RULE 26) ribosomal RNA (rRNA), short interfering RNA (siRNA), short-hairpin RNA (shRNA), micro-RNA (miRNA), ribozymes, cDNA, recombinant nucleic acids, branched nucleic acids, plasmids, vectors, isolated DNA of any sequence, and isolated RNA of any sequence.
[0099] Further, samples may be extracted from variety' of animal fluids containing cell free sequences, including but not limited to blood, serum, plasma, vitreous, sputum, urine, tears, perspiration, saliva, semen, mucosal excretions, mucus, spinal fluid, amniotic fluid, lymph fluid and the like. Cell free polynucleotides may be fetal in origin (via fluid taken from a pregnant subject) or may be derived from tissue of the subject itself. A biological sample may also refer to a sample engineered to mimic one or more properties (e.g., nucleic acid sequence properties, e.g., sequence identity, length, GC content, etc.) of a sample derived from a subject or specimen.
[0100]
[0078] As used herein, the term “template nucleic acid” generally refers to the nucleic acid to be sequenced. The template nucleic acid may be an analyte or be associated with an analyte. For example, the analyte can be a mRNA, and the template nucleic acid is the mRNA, or a cDNA derived from the mRNA, or other derivative thereof. In another example, the analyte can be a protein, and the template nucleic acid is an oligonucleotide that is conjugated to an antibody that binds to the protein, or derivative thereof. Examples of sequencing include single molecule sequencing or sequencing by synthesis, for example. Sequencing may comprise generating sequencing signals and / or sequencing reads. Sequencing may be performed on template nucleic acids immobilized on a support, such as a flow cell, substrate, and / or one or more beads In some cases, a template nucleic acid may be amplified to produce a colony of nucleic acid molecules attached to the support to produce amplified sequencing signals. In one example, (i) a template nucleic acid is subjected to a nucleic acid reaction, e.g., amplification, to produce a clonal population of the nucleic acid attached to a bead, the bead immobilized to a substrate, (ii) amplified sequencing signals from the immobilized bead are detected from the substrate surface during or following one or more nucleotide flows, and (iii) the sequencing signals are processed to generate sequencing reads. The substrate surface may immobilize multiple beads at distinct locations, each bead containing distinct colonies of nucleic acids, and upon detecting the substrate surface, multiple sequencing signals may be simultaneously or substantially simultaneously processed from the different immobilized beads at the distinct locations to generate multiple sequencing reads. In some sequencing methods, the nucleotide flows comprise non-termmated
[0101] -10-
[0102] SUBSTITUTE SHEET (RULE 26) nucleotides. In some sequencing methods, the nucleotide flows comprise terminated nucleotides.
[0103]
[0079] The term “nucleotide,” as used herein, generally refers to any nucleotide or nucleotide analog. The nucleotide may be naturally occurring or non-naturally occurring. The nucleotide may be a non-standard, modified, synthesized, or engineered nucleotide. The nucleotide may include a canonical base or a non-canonical base. The nucleotide may comprise an alternative base. The nucleotide may include a modified polyphosphate chain (e.g., triphosphate coupled to a fluorophore). The nucleotide may comprise a label. The nucleotide may be terminated
[0104] (e g., reversibly terminated). The nucleotide may be non-terminated (e.g., natural or modified). In some cases, nucleotides may include modifications in their phosphate moieties, including modifications to a triphosphate moiety. Nucleic acids may also be modified at the base moiety (e.g., at one or more atoms that typically are available to form a hydrogen bond with a complementary nucleotide and / or at one or more atoms that are not typically capable of forming a hydrogen bond with a complementary nucleotide), sugar moiety or phosphate backbone. Nucleic acids may also contain amine-modified groups, such as aminoallyl-dUTP (aa-dUTP) and aminohexhylacrylamide-dCTP (aha-dCTP) to allow covalent attachment of amine reactive moieties, such as N-hydroxysuccinimide esters (NHS). Alternatives to standard DNA base pairs or RNA base pairs in the oligonucleotides of the present disclosure can provide higher density in bits per cubic mm, higher safety (resistant to accidental or purposeful synthesis of natural toxins), easier discrimination in photo- programmed polymerases, or lower secondary structure. Nucleotides may be capable of reacting or bonding with detectable moieties for detection.
[0105]
[0080] The term “sequencing,” as used herein, generally refers to a process for generating or identifying a sequence of a biological molecule, such as a nucleic acid. The sequence may be a nucleic acid sequence which comprises a sequence of nucleic acid bases. Examples of sequencing include single molecule sequencing or sequencing by synthesis, for example. Sequencing may comprise generating sequencing signals and / or sequencing reads.
[0106]
[0081] The term “terminator” as used herein with respect to a nucleotide may generally refer to a moiety that is capable of terminating primer extension. A terminator may be a reversible terminator. A reversible terminator may comprise a blocking or capping group that is attached to the 3'-oxygen atom of a sugar moiety (e.g., a pentose) of a nucleotide or nucleotide analog. Such moieties are referred to as 3'-O-blocked reversible terminators. Examples of 3'-O-blocked reversible terminators include, for example, 3’-ONH2 reversible
[0107] -11-
[0108] SUBSTITUTE SHEET (RULE 26) terminators, 3'-O-allyl reversible terminators, and 3'-0-aziomethyl reversible terminators. Alternatively, a reversible terminator may comprise a blocking group in a linker (e.g., a cleavable linker) and / or dye moiety of a nucleotide analog. 3'-unblocked reversible terminators may be attached to both the base of the nucleotide analog as well as a fluorescing group (e.g., label, as described herein). Examples of 3 '-unblocked reversible terminators include, for example, the “virtual terminator” developed by Helicos BioSciences Corp, and the “lightning terminator” developed by Michael L. Metzker et al. Cleavage of a reversible terminator may be achieved by, for example, irradiating a nucleic acid molecule including the reversible terminator.
[0109]
[0082] The term “nucleotide flow” as used herein, generally refers to a temporally distinct instance of providing a nucleotide-containing reagent to a sequencing reaction space. The term “flow” as used herein, when not qualified by another reagent, generally refers to a nucleotide flow. For example, providing two flows may refer to (i) providing a nucleotide- containing reagent (e.g., an A-base-containing solution) to a sequencing reaction space at a first time point and (ii) providing a nucleotide-containing reagent (e.g., G-base-containing solution) to the sequencing reaction space at a second time point different from the first time point. A “sequencing reaction space” may be any reaction environment comprising a template nucleic acid. For example, the sequencing reaction space may be or comprise a substrate surface comprising a template nucleic acid immobilized thereto; a substrate surface comprising a bead immobilized thereto, the bead comprising a template nucleic acid immobilized thereto; or any reaction chamber or surface that comprises a template nucleic acid, which may or may not be immobilized. A nucleotide flow can have any number of base types (e.g., A, T, G, C; or U), for example 1, 2, 3, or 4 canonical base types. A “flow order,” as used herein, generally refers to the order of nucleotide flows used to sequence a template nucleic acid. A flow order may be expressed as a one-dimensional matrix or linear array of bases corresponding to the identities of, and arranged in chronological order of, the nucleotide flows provided to the sequencing reaction space:
[0110] (e.g., [A T G C A T G C A T G A T G A T G A T G C A T G C]).
[0111]
[0083] Such one-dimensional matrix or linear array of bases in the flow order may also be referred to herein as a “flow space.” A flow order may have any number of nucleotide flows. A “flow position,” as used herein, generally refers to the sequential position of a given nucleotide flow entry in the flow space (e.g., an element in the one-dimensional matrix or linear array). A “flow cycle,” as used herein, generally refers to the order of nucleotide
[0112] -12-
[0113] SUBSTITUTE SHEET (RULE 26) flow(s) of a sub-group of contiguous nucleotide flow(s) within the flow order. A flow cycle may be expressed as a one-dimensional matrix or linear array of an order of bases corresponding to the identities of, and arranged in chronological order of, the nucleotide flows provided within the sub-group of contiguous flow(s) (e.g., [A T G C], [A A T T G G C C], [A T], [A / T A / G], [A A], [A], [A T G], etc.). A flow cycle may have any number of nucleotide flows. A given flow cycle may be repeated one or more times in the flow order, consecutively or non-consecutively. Accordingly, the term “flow cycle order,” as used herein, generally refers to an ordering of flow cycles within the flow order and can be expressed in units of flow cycles. For example, where [A T G C] is identified as a 1stflow cycle, and [A T G] is identified as a 2ndflow cycle, the flow order of [A T G C A T G C A T G A T G A T G A T G C A T G C] may be described as having a flow-cycle order of [1stflow cycle; 1stflow cycle; 2ndflow cycle; 2ndflow cycle; 2ndflow cycle; 1stflow cycle; 1stflow cycle]. Alternatively or in addition, the flow cycle order may be described as [cycle 1, cycle, 2, cycle 3, cycle 4, cycle 5, cycle 6], where cycle 1 is the 1stflow cycle, cycle 2 is the 1stflow cycle, cycle 3 is the 2ndflow cycle, etc. In some cases, a flow-cycle order may be [T G C A], However, any other permutation of the nucleotides T (or U), G, C, and A may be used as a flow-cycle order.
[0114] Sample Processing Methods
[0115]
[0084] Described herein are devices, systems, methods, compositions, and kits for processing samples, such as to prepare a sample for sequencing, to sequence a sample, and / or to analyze sequencing data. FIG. 1 illustrates an example sequencing workflow 100, according to the devices, systems, methods, compositions, and kits of the present disclosure.
[0116]
[0085] Supports and / or template nucleic acids may be provided and / or prepared (101) to be compatible with downstream sequencing operations (e g , 107). A support (e g , bead) may help immobilize a template nucleic acid to a substrate, such as when the template nucleic acid is coupled to the support, and the support is in turn immobilized to the substrate. The support may further function as a binding entity' to retain derivatives molecules (e.g., amplification products) from a same template nucleic acid together for downstream processing, such as for sequencing operations. This may be useful in distinguishing a colony from other colonies
[0117] (e g., on other supports) and generating amplified sequencing signals corresponding to a template nucleic acid. A support may comprise an oligonucleotide comprising one or more functional nucleic acid sequences. The oligonucleotide may be single-stranded, doublestranded, or partially double-stranded. For example, the oligonucleotide may comprise a
[0118] -13-
[0119] SUBSTITUTE SHEET (RULE 26) capture sequence, a primer sequence, a sequencing primer sequence, a barcode sequence, a sample index sequence, a unique molecular identifier (UMI), a flow cell adapter sequence, an adapter sequence, a target sequence, a random sequence, a binding sequence (e.g., for a splint, primer, template nucleic acid, capture sequence, etc.), or any other functional sequence useful for a downstream operation, a complement thereof, or any combination thereof. The capture sequence may be configured to hybridize to a sequence of a template nucleic acid or derivative thereof. The support may comprise a plurality of oligonucleotides, for example on the order of 10, 102, 103, 104, 105, 106, 107, or more molecules The support may comprise a single species of oligonucleotide which comprise identical sequences. The support may comprise multiple species of oligonucleotides which have varying sequences. In some cases, the support comprises a single species of a primer (e.g., forw ard primer) for amplification. In some cases, the support comprises two species of primer (e.g., forward primer, reverse primer) for amplification. Devices, systems, methods, compositions, and kits for preparing and using support species are described in further detail in U.S. Patent Pub. No. 20220042072A1 and International Patent Pub. No. W02022040557A2, each of which is entirely incorporated herein by reference for all purposes.
[0120]
[0086] A support may comprise one or more capture entities, where a capture entity is configured for capture by a capturing entity. A capture entity may be coupled to or be part of an oligonucleotide coupled to the support. A capture entity may be coupled to or be part of the support. Examples of capture entity-capturing entity pairs and capturing entity-capture entity pairs include: streptavidin (SA)-biotin; complementary sequences; magnetic particle- magnetic field system; charged particle-electric field system; azide-cyclooctyne; thiol- maleimide; click chemistry pairs; cross-linking pairs; etc. The capture entity-capturing entity pair may comprise one or more chemically modified bases. A capture entity and capturing entity may bind, couple, hybridize, or otherwise associate with each other. The association may comprise formation of a covalent bond, non-covalent bond, releasable bond (e.g., cleavable bond that is cleavable upon application of a stimulus), and / or no bond. The capture entity may be capable of linking to a nucleotide. In some instances, the capturing entity may comprise a secondary capture entity, for example, for subsequent capture by a secondary captunng entity The secondary capture entity and secondary' capturing entity may comprise any one or more of the capturing mechanisms described elsewhere herein.
[0121]
[0087] A support may comprise one or more cleavable moieties, also referred to herein as excisable moieties. The cleavable moiety may be coupled to or be part of an oligonucleotide
[0122] -14-
[0123] SUBSTITUTE SHEET (RULE 26) coupled to the support. The cleavable moiety may be coupled to the support. A cleavable moiety may comprise any useful moiety that can be used to cleave an oligonucleotide (or portion thereof) from the support, or otherwise release a nucleic acid strand from the support and / or the oligonucleotide. A cleavable moiety may comprise a uracil, a ribonucleotide, methylated nucleotide, or other modified nucleotide that is excisable or cleavable using an enzyme (e.g., UDG, RNAse, APE1, MspJI, endonuclease, exonuclease, etc ). The cleavable moiety may comprise an abasic site or an analog of an abasic site (e.g., dSpacer), a dideoxyribose, a spacer, e.g., C3 spacer, hexanediol, triethylene glycol spacer (e.g., Spacer 9), hexa-ethyleneglycol spacer (e.g., Spacer 18), a photocleavable moiety, or combinations or analogs thereof. Alternatively, or in addition, the cleavable moiety may be cleavable using one or more stimuli, e.g., photo-stimulus, chemical stimulus, thermal stimulus, etc.
[0124]
[0088] The sequencing workflow 100 may not involve supports, for example when a template nucleic acid and / or its derivatives are directly attached to a substrate and amplified and / or sequenced from the substrate.
[0125]
[0089] A template nucleic acid may include an insert sequence sourced from a biological sample. The template nucleic acid may be derived from any nucleic acid of the biological sample (e.g., endogenous nucleic acid) and result from any number of processing operations, such as but not limited to fragmentation, degradation or digestion, transposition, ligation, reverse transcription, extension, replication, etc. The template nucleic acid may be singlestranded, double-stranded, or partially double-stranded. A template nucleic acid may comprise one or more functional nucleic acid sequences. For example, the template nucleic acid may comprise a capture sequence, a primer sequence, a sequencing primer sequence, a barcode sequence, a sample index sequence, a unique molecular identifier (UMI), a flow cell adapter sequence, an adapter sequence, a target sequence, a random sequence, a binding sequence (e g., for a splint, primer, template nucleic acid, capture sequence, etc.), or any other functional sequence useful for a downstream operation, a complement thereof, or any combination thereof. The template nucleic acid may comprise an adapter sequence configured to be captured by a capture sequence of an oligonucleotide coupled to a support. The one or more functional nucleic acid sequences may be disposed at one end or both ends of the insert sequence. A nucleic acid molecule comprising the insert sequence, or complement thereof, may be processed with (e g., attached to, extend from, etc.) one or more adapter molecules to generate the template nucleic acid comprising the insert sequence and one or more functional nucleic acid sequences. A template nucleic acid may comprise one or
[0126] -15-
[0127] SUBSTITUTE SHEET (RULE 26) more capture entities and / or one or more cleavable moi eties that are described elsewhere herein.
[0128]
[0090] Optionally, the supports and / or template nucleic acids may be pre-enriched (102). For example, a support comprising a distinct oligonucleotide sequence is pre-enriched to isolate from a mixture comprising support(s) that do not have the distinct oligonucleotide sequence. For example, a template nucleic acid comprising a distinct configuration (e.g., comprising a particular adapter sequence) is pre-enriched to isolate from a mixture comprising template nucleic acids that do not have the distinct configuration. In some cases, the capture entity on the supports and / or template nucleic acids are used for pre-enrichment.
[0129]
[0091] The supports and template nucleic acids may be attached (103) to generate supporttemplate complexes. A template nucleic acid may be coupled to a support via any method(s) that results in a stable association between the template nucleic acid and the support. For example, the template nucleic acid may hybridize to an oligonucleotide on the support; the template nucleic acid may be ligated to a nucleic acid coupled to the support; the template nucleic acid may hybridize to one or more intermediary molecules, such as a splint, bridge, and / or primer molecule, which hybridizes to an oligonucleotide on the support; and / or the template nucleic acid may be hybridized to an oligonucleotide on a support, which oligonucleotide comprises a primer sequence which is extended. In some cases, the respective concentrations of the supports and template nucleic acids may be adjusted such that a majority of support-template complexes are single template-attached supports (e.g., a support attached to a single template nucleic acid).
[0130]
[0092] Optionally, support-template complexes may be pre-enriched (104), wherein a support-template complex is isolated from a mixture comprising support(s) and / or template nucleic acid(s) that are not attached to each other In some cases, the capture entity on the supports and / or template nucleic acids are used for pre-enrichment.
[0131]
[0093] The template nucleic acids may be subjected to amplification reactions (105) to generate a plurality of amplification products immobilized to the support. Such amplification reactions may comprise performing polymerase chain reaction (PCR) or any other amplification methods described herein, including but not limited to emulsion PCR (ePCR or emPCR), isothermal amplification, recombinase polymerase amplification (RPA), rolling circle amplification (RCA), multiple displacement amplification (MDA), bridge amplification, template walking, etc. Amplification reactions can occur while the support is immobilized to a substrate. Amplification reactions can occur off the substrate, such as in
[0132] -16-
[0133] SUBSTITUTE SHEET (RULE 26) solution, or on a different surface or platform. Amplification reactions can occur in isolated reaction volumes, such as within multiple droplets in an emulsion during emulsion PCR (ePCR or emPCR), or in wells or tubes
[0134]
[0094] Optionally, subsequent to amplification, the supports, template nucleic acids, and / or support-template complexes may be subjected to post-amplification processing (106). Often, subsequent to amplification, a resulting mixture may comprise a mix of positive supports (e.g., those comprising a template nucleic acid molecule) and negative supports (e.g., those not attached to template nucleic acid molecules). Enrichment procedure(s) may isolate positive supports from the mixtures. Example methods of enrichment of amplified supports are described in U.S. Patent Nos. 10,900,078, U.S. Patent Pub. No. 20210079464 Al, and International Patent Pub. No. W02022040557A2, each of which is entirely incorporated by reference herein.
[0135]
[0095] The template nucleic acids may be subject to sequencing (107). The template nucleic acid(s) may be sequenced while attached to the support. Alternatively, the template nucleic acid molecules may be free of the support when sequenced and / or analyzed. The template nucleic acids may be sequenced while immobilized to a substrate, such as via a support or otherwise. Examples of substrate-based sample processing systems are described elsewhere herein. Any sequencing method may be used, for example pyrosequencing, single molecule sequencing, sequencing by synthesis (SBS), sequencing by ligation, sequencing by binding, etc.
[0136]
[0096] For example, sequencing comprises extending a sequencing primer (or growing strand) hybridized to a template nucleic acid by providing labeled nucleotide reagents, washing away unincorporated nucleotides from the reaction space, and detecting one or more signals from the labeled nucleotide reagents which are indicative of an incorporation event or lack thereof. After detection, the labels may be cleaved, and the entire process may be repeated any number of times to determine sequence information of the template nucleic acid. One or more intermediary flows may be provided intra- or inter- repeat, such as washing flows, label cleaving flows, terminator cleaving flows, reaction-completing flows (e.g., double tap flow, tnple tap flow, etc.), labeled flows (or bright flows), unlabeled flows (or dark flows), phasing flows, chemical scar capping flows, etc. A nucleotide mixture that is provided during any one flow may comprise only labeled nucleotides, only unlabeled nucleotides, or a mixture of labeled and unlabeled nucleotides. The mixture of labeled and unlabeled nucleotides may be of any fraction of labeled nucleotides, such as at least or at
[0137] -17-
[0138] SUBSTITUTE SHEET (RULE 26) most 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 99%. A nucleotide mixture that is provided during any one flow may comprise only non-terminated nucleotides, only terminated nucleotides, or a mixture of terminated and non-terminated nucleotides. When using only non-terminated nucleotides, terminator cleaving flows may be omitted from the sequencing process. When using terminated nucleotides, to proceed with the next step of extension, prior to, during, or subsequent to detection, a terminator cleaving flow may be provided to cleave blocking moieties. A nucleotide mixture that is provided during any one flow may comprise any number of canonical base types (e.g., A, T, G, C, U), such as a single canonical base type, two canonical base types, three canonical base types, four canonical base types or five canonical base types (including T and U). Different types of nucleotide bases may be flowed in any order and / or in any mixture of base types that is useful for sequencing. Various flow-based sequencing systems and methods are described in U.S. Pat. Pub. No. 2022 / 0170089 Al, which is entirely incorporated by reference herein for all purposes and described further below. Labeled nucleotides may comprise a dye, fluorophore, or quantum dot, multiples thereof, and / or combination thereof. In some cases, nucleotides of different canonical base types may be labeled and detectable at a single frequency (e.g., using the same or different dyes). In other cases, nucleotides of different canonical base types may be labeled and detectable at different frequencies (e.g., using the same or different dyes).
[0139]
[0097] Subsequent to sequencing, the sequencing signals collected and / or generated may be subjected to data analysis (108). The sequencing signals may be processed to generate base calls and / or sequencing reads. In some cases, the sequencing reads may be processed to generate diagnostics data to the biological sample, or the subject from which the biological sample was derived from. The data analysis may comprise image processing, alignment to a genome or reference genome, training and / or trained algorithms, error correction, and the like.
[0140]
[0098] While the sequencing workflow 100 with respect to FIG. 1 has been described with respect to the use of supports to bind template molecules, it will be appreciated that the different supports may be effectively replaced by using spatially distinct locations on one or more surfaces, which do not necessarily have to be the surfaces of individual supports (e.g., beads, DNA nanoballs, and / or dendrimers). For example, a first spatially distinct location on a surface may be capable of directly immobilizing a first colony of a first template nucleic acid and a second spatially distinct location on the same surface (or a different surface) may
[0141] -18-
[0142] SUBSTITUTE SHEET (RULE 26) be capable of directly immobilizing a second colony of a second template nucleic acid to distinguish from the first colony. In some cases, the surface comprising the spatially distinct locations may be a surface of the substrate on which the sample is sequenced, thus streamlining the amplification-sequencing workflow.
[0143]
[0099] It will be appreciated that in some instances, the different operations described in the sequencing workflow 100 may be performed in a different order. It will be appreciated that in some instances, one or more operations described in the sequencing workflow 100 may be omitted or replaced with other comparable operation(s). It will be appreciated that in some instances, one or more additional operations described in the sequencing workflow 100 may be performed. The different operations described with respect to sequencing workflow 100 may be performed with the help of open substrate systems described herein.
[0144] Flow-based sequencing
[0145]
[0100] Sequencing data can be generated using flow-based sequencing methods that include extending a primer bound to a template nucleic acid according to a pre-determined flow cycle and / or flow order where, in one or more flow positions, known canonical base type(s) of nucleotides (e g., A, C, G, T, U) is accessible to the extending primer. At least some of the nucleotides may include a label, which labeled nucleotides upon incorporation into the extending primer renders a detectable signal. The resulting sequence by which nucleotides are incorporated into the extended primer is expected to be the reverse complement of the sequence of the template nucleic acid. A method for sequencing can comprise using a flow sequencing method that includes (1) extending a primer using labeled nucleotides in a flow, and (2) detecting the presence or absence of a labeled nucleotide incorporated into the extending primer to generate sequencing data. Flow sequencing methods may also be referred to as “natural sequencing-by-synthesis,” “mostly natural sequencing-by -synthesis,” or “non-terminated sequencing-by-synthesis” methods. Example methods are described in U.S. Patent No. 8,772,473 and U.S. Pat. Pub. No. 2022 / 0170089A1, each of which is incorporated herein by reference in its entirety.
[0146]
[0101] In flow sequencing, iterative nucleotide flows are used to extend the primer hybridized to the template nucleic acid, with detection of incorporated nucleotides between one or more flows. The nucleotides may be, for example, non-terminating nucleotides such that more than one consecutive base can be incorporated into the extending primer strand if more than one consecutive complementary base (or homopolymer region) is present in the template strand. At least a portion of the nucleotides can be labeled so that incorporation can
[0147] -19-
[0148] SUBSTITUTE SHEET (RULE 26) be detected. Generally, only a single nucleotide type is introduced in a flow, although two or three different types of nucleotides may be simultaneously introduced in certain embodiments. This methodology can be contrasted with sequencing methods that use a reversible terminator, where primer extension is stopped after extension of every single base before the terminator is reversed (e.g., by removing a 3’ blocking group) to allow incorporation of the next succeeding base.
[0149]
[0102] FIG. 2 illustrates an example flow sequencing method that can be used to generate the sequencing data described herein. Template nucleic acids may be immobilized to a surface (e.g., the surface of a bead attached to a substrate or directly to a substrate), as described in detail herein. In this example, the template nucleic acid includes an adaptor sequence 201 followed by an insert sequence ( “ACGTTGCTA . ”). The adaptor sequence 201 can include a sequencing primer hybridization site. At operation 202, a sequencing primer 203 is hybridized to the adapter sequence 201 at the sequencing primer hybridization site. The sequencing primer 203 is then extended in a series of flows according to flow cycle 200 with flow order: [T G C A], In this example, the flow cycle 200 includes four flow steps 204, 206, 208, 210, and in a given flow step, a single base type is provided to the templateprimer hybrid. In flow step 204, nucleotides comprising labeled T nucleotides are provided; in flow step 206, nucleotides comprising labeled G nucleotides are provided; in flow step 208, nucleotides comprising labeled C nucleotides are provided; in flow step 210, nucleotides comprising labeled A nucleotides are provided Nucleotides in a single-base flow may comprise a mixture of labeled and unlabeled nucleotides of the single base. At flow step 204, a labeled T nucleotide is incorporated by the extending sequencing primer 203 opposite the A base in the template strand. Then, a signal indicative of the incorporation of the labeled T nucleotide can be detected. For example, the signal may be detected by imaging the surface the template nucleic acids are immobilized on and analyzing the resulting image(s). The sequencing platform may be washed with a wash buffer to remove unincorporated nucleotides prior to signal detection. In some cases, prior to the next flow step (e.g., 206), the label may be removed from the incorporated labeled T nucleotide (e.g., by cleaving the label from the nucleotide), before proceeding. Nucleotide flow, detection, and optionally cleavage, may be repeated according to a flow order that may or may not include repeating the flow cycle 200 for any number of times. Flow step 210 illustrates incorporation of two labeled A bases by the extending sequencing pnmer 203 opposite the two T bases in the template strand, per the non-terminated nature of the flown nucleotides. The detected signal intensity
[0150] -20-
[0151] SUBSTITUTE SHEET (RULE 26) indicating the incorporation of two A nucleotides may be greater than the signal intensity indicating the incorporation of one nucleotide. For simplicity, this Figure illustrates incorporation of two labeled A nucleotides in the same hybrid. However, flow-based sequencing may be performed on colonies of amplified molecules, e.g., each bead representing one colony, where an optically resolvable location contains multiple copies of the same template nucleic acid molecule (e.g., a location contains one amplified bead), such that the signal detected at an optically resolvable location represents an aggregate signal from the multiple copies of molecules. Thus, when using a nucleotide flow mixture containing labeled and unlabeled nucleotides of a same base type, the incorporation of the labeled nucleotides can be distributed across the multiple copies of the molecules, and the aggregate signal from the multiple copies detected. In some cases, for a majority of hybrids, at most a single labeled nucleotide may be incorporated into a single homopolymer stretch in a hybrid — the longer the homopolymer stretch, the more likely that more hybrids of the plurality of copies of hybrids in an optically resolvable location will incorporate one labeled nucleotide.
[0152]
[0103] While each flow' step in the example flow' sequencing method in FIG. 2 results in incorporation of one or more nucleotides (and thus a detected signal indicating such incorporation), it should be appreciated that not all flow steps result in incorporation of nucleotides. In some flow steps, no nucleotide base may be incorporated (for example, in the absence of a complementary base in the template).
[0153] 1104] A nucleotide mixture that is provided during any one flow may comprise only labeled nucleotides, only unlabeled nucleotides, or a mixture of labeled and unlabeled nucleotides. The mixture of labeled and unlabeled nucleotides may be of any fraction of labeled nucleotides, such as at least or at most about 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 99%. Labeled nucleotides may comprise a dye, fluorophore, or quantum dot, multiples thereof, and / or combination thereof. In some cases, nucleotides of different canonical base types may be labeled and detectable at a single frequency (e.g., using the same or different dyes). In other cases, nucleotides of different canonical base types may be labeled and detectable at different frequencies (e.g., using the same or different dyes) Labeled nucleotides may comprise an optical moiety (e g., dye, fluorophore, quantum dot, label, etc.) coupled to a nucleobase via a linker, and the label from the labeled nucleotides may be removed by cleaving the linker to remove the optical moiety. Cleaving may comprise one or
[0154] -21-
[0155] SUBSTITUTE SHEET (RULE 26) more stimuli, such as exposure to a chemical (e.g., reducing agent), an enzyme, light (e.g., UV light), or temperature change (e g., heat).
[0156]
[0105] Flow-based sequencing may comprise providing non-detected nucleotide flow(s), for example to skip sequencing of a region(s) of the template nucleic acid; to ensure completion of incorporation reactions across all template-primer hybrids in the reaction space; and / or phasing or re-phasing. A non-detected nucleotide flow may be referred to herein as a “dark flow”, “dark tap”, or “dark tap flow.” A detected nucleotide flow may be referred to herein as a “bright flow”, “bnght tap”, or “bright tap flow.” Incorporation reactions may be incomplete in the reaction space when not all available incorporation sites in the template-primer hybrids have incorporated a complementary base, such as due to reaction kinetics and / or insufficient incubation time or reagents. In some cases, single-base flows of the same canonical base type may be provided consecutively (without intervening flow of a different nucleotide base ty pe) for any number of consecutive flows, to ensure completion of incorporation reactions. A consecutive same-base flow may be referred to herein as a “double tap” or “double tap flow” if there are two consecutive flows, a “triple tap” or “triple tap flow” if there are three consecutive flows, or a “rath tap” or “rath tap flow” if there are n consecutive flows of the same base type. A double tap, triple tap, or rath tap flow may or may not be detected. Labels in a flow may or may not be removed (e g., cleaved) prior to the double tap, triple tap, or rath tap flow. Detection of labeled nucleotides from a particular flow may be performed prior to, during, or subsequent to the double tap, triple tap, or rath tap flow. Accordingly, below are non-limiting examples of flow cycles that can be used in a larger flow order of flow-based sequencing methods, which may or may not be repeated and / or mixed and matched with other flow cycles, where * after a base represents a detected flow step and / between bases represents a mixed base flow:
[0157] Single-base flow: e.g., [T* A* C * G*]
[0158] Single-base flow with double tap: e.g., [T* T A* A C* C G* G]
[0159] Mixed base flow, all labeled: e.g., [T* A* / C* / G*] Mixed base flow, some unlabeled: e.g., [T* A / C* / G] Mixed base flow, some unlabeled: e.g., [T A* / C* / G*] Skip region base flow: e.g., [T / A / C G / A / T] Three base flow cycle: e.g., [T A C]
[0160]
[0106] FIG. 3 illustrates an example flowgram of signals detected after five exemplary flow cycles of [T A C G] are performed to extend a sequencing primer, in accordance with some
[0161] -22-
[0162] SUBSTITUTE SHEET (RULE 26) cases. Each column in the flowgram corresponds to a detected flow step (e.g., 302, 306), and the values in each column collectively represent the detected signal intensity in the flow step. In each detected flow step, the flow signal can be determined from an analog signal that is detected during the sequencing process, such as a fluorescent signal of the one or more bases incorporated. In some cases, for a flow step, the detected signal intensity can be expressed in probabilistic terms. Specifically, the detected signal intensity can be expressed in a series of likelihood values corresponding to different integer homopolymer base lengths (e.g., 0 base, 1 base, 2 bases, 3 bases, etc.) for the flow position. For flow step 302, the detected signal intensity is expressed by a first likelihood value of 0.001 for 0 base, a second likelihood value of 0.9979 for 1 base, a third likelihood value of 0.001 for 3 bases, and a fourth likelihood value of 0.0001 for 4 bases. This can be interpreted to indicate that there is a high statistical likelihood that one nucleotide base has been incorporated. In this flow step, a single T was determined to be incorporated, which means there is an A in the template. Similarly, for flow step 306, the column values can collectively indicate that there is a high statistical likelihood that no base has been incorporated (with 0,9988 likelihood value for 0 bases). With similar analyses performed at each flow position, a preliminary sequence 310 (TATGGTCGTCGA) of the extending primer can be determined, and reverse complement (i.e., the template strand sequence) readily determined from the preliminary sequence. For example, the most likely sequence can be determined by selecting the base count with the highest likelihood at each flow position, as shown by the stars in the flowgram. Further, the likelihood of this sequencing data set can be determined as the product of the selected likelihood at each flow position. Accordingly, the flowgram may be formatted as a sparse matrix, with a flow signal represented by a plurality of likelihood values indicative of a plurality of base homopolymer length counts at each flow position. The homopolymer length likelihood may vary, for example, based on the noise or other artifacts present during detection of the analog signal during sequencing. In some cases, if the homopolymer length likelihood statistical parameter or likelihood is below a predetermined threshold, the parameter may be set to a predetermined non-zero value that is substantially zero (i.e., some very small value or negligible value) to aid the downstream statistical analysis further discussed herein, wherein a true zero value may give rise to a computational error or insufficiently differentiate between levels of unlikelihood, e.g., very unlikely (0.0001) and inconceivable (0). Thus, a method for sequencing may comprise generating a flowgram using analog signals (e.g., fluorescent
[0163] -23-
[0164] SUBSTITUTE SHEET (RULE 26) signals) detected from a template nucleic acid or derivative thereof and generating base calls and / or sequencing reads using the flowgram.
[0165]
[0107] As will be appreciated, in flow-based sequencing, the signal for any flow position in the sequencing data is flow order-dependent in that the same flow position for a same template nucleic acid may express different flow signals for different flow orders. Any useful predetermined flow cycles and / or flow orders may be designed to sequence a template nucleic acid and / or more accurately or precisely detect a particular type of sequence (e.g., single nucleotide polymorphisms (SNPs)) within the template nucleic acid (e.g., of a genome).
[0166]
[0108] A flowgram may be binary or non-binary. A binary flowgram detects the presence (1) or absence (0) of an incorporated nucleotide. A non-binary flowgram, such as shown in FIG. 3, can more quantitatively determine a number of incorporated nucleotides at each flow position.
[0167]
[0109] Additional sequencing schemes are described in U.S. Pat. Pub. Nos.
[0168] 2021 / 0017593 Al, 2022 / 0064728A1, and 2022 / 0154272A1, each of which is entirely incorporated herein by reference for all purposes.
[0169] Sequencing systems
[0170] [HO] The sequencing methods described herein may be performed using any sequencing platform, such as a substrate-based system. The substrate-based system may comprise a closed substrate such as a flow cell comprising one or more fluidic or microfluidic channels, wells, and / or microwells. For example, template nucleic acids on or off a bead may be immobilized to a surface in a flow cell, and reagents flowed in and out of the flow cell through channels in the flow cell to contact the template nucleic acids. The channels may be flushed with wash buffers between different reagent cy cles. The substrate-based system may comprise an open substrate. For example, template nucleic acids on or off a bead may be immobilized to a surface of an open substrate, and reagents directed to the surface, such as via nozzles (e.g., across an air gap), to contact the template nucleic acids. The open substrate may be washed with wash buffers between different reagent cycles.
[0171] [Ill] Described herein are devices, systems, and methods that use open substrates or open flow cell geometries to process a sample. The term “open substrate,” as used herein, generally refers to a substrate in which any point on an active surface of the substrate is physically accessible from a direction normal to the substrate. The devices, systems and methods may be used to facilitate any application or process involving a reaction or
[0172] -24-
[0173] SUBSTITUTE SHEET (RULE 26) interaction between two objects, such as between an analyte and a reagent or between two reagents. For example, the reaction or interaction may be chemical (e.g., polymerase reaction) or physical (e.g., displacement). The devices, systems, and methods described herein may benefit from higher efficiency, such as faster reagent delivery and lower volumes of reagents required per surface area. The devices, systems, and methods described herein may avoid contamination problems common to microfluidic channel flow cells that are fed from multiport valves which can be a source of carryover from one reagent to the next. The devices, systems, and methods may benefit from shorter completion time, use of fewer resources (e.g , various reagents), and / or reduced system costs. The open substrates or flow cell geometries may be used for any application or process, such as, but not limited to, sequencing by synthesis, sequencing by ligation, amplification, proteomics, single cell processing, barcoding, and sample preparation, as described herein.
[0174] [H2] A sample processing system may comprise a substrate, and devices and systems that perform one or more operations with or on the substrate. The sample processing system may permit highly efficient dispensing of analytes and reagents onto the substrate. The sample processing may permit highly efficient imaging of one or more analytes, or signals corresponding thereto, on the substrate. The sample processing system may comprise an imaging system comprising a detector. Substrates, detectors, and sample processing hardware that can be used in the sample processing system are described in further detail in U.S. Patent Pub. No. 20200326327A1, U S. Patent Pub. No. 20210079464 Al, International Patent Pub. No. WO2022072652A1, and U.S. Patent Pub. No. 20210354126A1, each of which is entirely incorporated herein by reference for all purposes.
[0175]
[0113] A substrate may comprise a planar or substantially planar surface. Substantially planar may refer to planarity at a micrometer level (e g., a range of unevenness on the planar surface does not exceed the micrometer scale) or nanometer level (e.g., a range of unevenness on the planar surface does not exceed the nanometer scale). Alternatively, substantially planar may refer to planarity at less than a nanometer level or greater than a micrometer level (e.g., millimeter level). A surface of the substrate may be textured or patterned. For example, the substrate may compnse grooves, troughs, hills, pillars, wells, cavities (e.g., micro-scale cavities or nano-scale cavities), channels, wedges, cuboids, cylinders, spheroids, hemispheres, etc. A substrate surface may comprise chemical groups such as amines, esters, hydroxyls, epoxides, and the like, or a combination thereof. A substrate surface may comprise any of the binders or linkers described herein, such as to help immobilize analytes
[0176] -25-
[0177] SUBSTITUTE SHEET (RULE 26) thereto. The substrate may be textured or patterned such that all features are at or above a reference level of the surface (no features below a reference level of the surface, such as a well), or such that all features are at or below a reference level of the surface (no features below a reference level of the surface, such as a pillar). In some instances, a texture of the substrate may comprise structures having a maximum dimension of at most about 500%, 400%, 300%, 200%, 100%, 90%, 80%, 70%, 60%, 50%, 40%, 30%, 20%, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, 1%, 0.1%, 0.01%, 0.001%, 0.0001%, 0 00001% of the total thickness of the substrate or a layer of the substrate. In some instances, the textures and / or patterns of the substrate may define at least part of an individually addressable location on the substrate. A textured and / or patterned substrate may be substantially planar. Alternatively, the substrate may be untextured and / or unpattemed.
[0178] |H4] The substrate may have the general form of a cylinder, a cylindrical shell or disk, a rectangular prism, or any other geometric form. The substrate may have a thickness (e.g., a minimum dimension) of at least and / or at most about 100 micrometers (pm), 200 pm, 500 pm, 1 millimeter (mm), 2 mm, 5 mm, 10 mm, 15 mm, 20 mm, 25 mm, 30 mm, 35 mm, 40 mm, 45 mm, 50 or mm. The substrate may have a first lateral dimension (such as a width for a substrate having the general form of a rectangular prism or a radius or diameter for a substrate having the general form of a cylinder) and / or a second lateral dimension (such as a length for a substrate having the general form of a rectangular prism) of at least and / or at most about 1 mm, 2 mm, 5 mm, 10 mm, 20 mm, 30 mm, 40 mm, 50 mm, 100 mm, 150 mm, 200 mm, 300 mm, 400 mm, 500 mm, 1,000 mm, 1,500 mm, 2,000 mm, 2,500 mm, 3,000 mm, 4,000 mm, 5,000 mm or more.
[0179]
[0115] The substrate may comprise a plurality of individually addressable locations. The individually addressable locations may comprise locations that are physically accessible for manipulation. The manipulation may comprise, for example, placement, extraction, reagent dispensing, seeding, heating, cooling, or agitation. The manipulation may be accomplished through, for example, localized microfluidic, pipet, optical, laser, acoustic, magnetic, and / or electromagnetic interactions with the analyte or its surroundings. The individually addressable locations may comprise locations that are digitally accessible. For example, each individually addressable location may be located, identified, and / or accessed electronically or digitally for indexing, mapping, sensing, associating with a device (e.g., detector, processor, dispenser, etc.), or otherwise processing. In some cases, the individually addressable locations may be defined by physical features of the substrate (e g., on a modified surface) to
[0180] -26-
[0181] SUBSTITUTE SHEET (RULE 26) distinguish individually addressable locations from each other and from non-individually addressable locations. In some cases, the individually addressable locations may not be defined by physical features of the substrate, and instead may be defined digitally (e.g., by indexing) and / or via the analytes and / or reagents that are loaded on the substrate (e.g., the locations in which analytes are immobilized on the substrate). The plurality of individually addressable locations may be arranged as an array, randomly, or according to any pattern, on the substrate. FIG. 4 illustrates different substrates (from a top view) comprising different arrangements of individually addressable locations 401, with panel A showing a substantially rectangular substrate with regular linear arrays, panel B showing a substantially circular substrate with regular linear arrays, and panel C showing an arbitrarily shaped substrate with irregular arrays.
[0182] |H6] The substrate may have any number of individually addressable locations, for example, on the order of 1, 101, 102, 103, 104, 105, 106, 107, 108, 109, IO10, 1011, 1012, 1013or more individually addressable locations. Each individually addressable location may have any shape or form, for example the general shape or form of a circle, oval, square, rectangle, polygonal, or non-polygonal shape when viewed from the top. A plurality of individually addressable locations can have uniform shape or form, or different shapes or forms. An individually addressable location may have any size. In some cases, an individually addressable location may have an area of at least and / or at most about 0.1, 0.2, 0.25, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1, 1.1, 1.2, 1.25, 1.3, 1.4 ,1.5, 1.6, 1.7, 1.75, 1.8, 1 9, 2, 2.25, 2.5, 2.75, 3, 3.25, 3.5, 3.75, 4, 4.25, 4.5, 4.75, 5, 5.5, 6, 7, 8, 9, 10 square micron (pm2), or more. The individually addressable locations may be distributed on a substrate with a pitch determined by the distance between the center of a first location and the center of the closest or neighboring individually addressable location. Locations may be spaced with a pitch of at least and / or at most about 0.1, 0.2, 0.25, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1, 1.1, 1.2, 1.25, 1.3, 1.4 ,1.5, 1.6, 1.7, 1.75, 1.8, 1.9, 2, 2.25, 2.5, 2.75, 3, 3.25, 3.5, 3.75, 4, 4.25, 4.5, 4.75, 5, 5.5, 6, 6.5, 7, 7.5, 8, 8.5, 9, 9.5, or 10 micron (pm). In some cases, the pitch between two individually addressable locations may be determined as a function of a size of a loading object (e.g., bead, DNA particle, dendrimer, etc.). For example, where the loading object is a bead having a maximum diameter, the pitch may be at least about the maximum diameter of the loading object.
[0183]
[0117] An individually addressable location may be capable of immobilizing thereto an analyte (e.g., a nucleic acid, a protein, a carbohydrate, etc.) or a reagent (e.g., a nucleic acid, a
[0184] -27-
[0185] SUBSTITUTE SHEET (RULE 26) probe molecule, a barcode molecule, an antibody molecule, a primer molecule, a bead, etc.). In some cases, an analyte or reagent may be immobilized to an individually addressable location via a support, such as a bead. In an example, a first bead comprising a first colony of nucleic acid molecules each comprising a first template sequence is immobilized to a first individually addressable location, and a second bead comprising a second colony of nucleic acid molecules each comprising a second template sequence is immobilized to a second individually addressable location. A substrate may comprise more than one type of individually addressable location arranged as an array, randomly, or according to any pattern, on the substrate. In some cases, different types of individually addressable locations may have different chemical, physical, and / or biological properties (e.g., hydrophobicity, charge, color, topography, size, dimensions, geometry, etc.). In some cases, an individually addressable location may comprise a distinct surface chemistry. The distinct surface chemistry may distinguish between different addressable locations and / or distinguish an individually addressable location from surrounding locations. In one example, the substrate comprises a plurality of individually addressable locations, each defined by APTMS, which are positively charged and has affinity towards an amplified bead (e.g., a bead comprising nucleic acid molecules, e.g., amplicons, immobilized thereto) which exhibits a negative charge. The locations surrounding the plurality of individually addressable locations may comprise HMDS which repels amplified beads.
[0186]
[0118] In some cases, the individually addressable locations may be indexed, e.g., spatially. Data corresponding to an indexed location, collected over multiple periods of time, may be linked to the same indexed location. In some cases, sequencing signal data collected from an indexed location, during iterations of sequencmg-by -synthesis flows, are linked to the indexed location to generate a sequencing read for an analyte immobilized at the indexed location.
[0187]
[0119] A substrate may comprise a binder or linker configured to immobilize an analyte or reagent to an individually addressable location. The binders may be integral to or added to the substrate. The binders may immobilize analytes or reagents through non-specific interactions, such as one or more of hydrophilic interactions, hydrophobic interactions, electrostatic interactions, physical interactions (for instance, adhesion to pillars or settling within wells), and the like. Alternatively or in addition, the binders may immobilize analytes or reagents through specific interactions, such as hybridization between two nucleic acid molecules (an oligonucleotide binder and a template nucleic acid). For example, the binders may comprise
[0188] -28-
[0189] SUBSTITUTE SHEET (RULE 26) one or more of antibodies, oligonucleotides, nucleic acid molecules, aptamers, affinity binding proteins, lipids, carbohydrates, and the like.
[0190]
[0120] The substrate may be rotatable about an axis, referred to herein as a rotational axis. The rotational axis may or may not be an axis through the center of the substrate. The systems, devices, and apparatus described herein may further comprise an automated or manual rotational unit configured to rotate the substrate. The rotational unit may comprise a motor and / or a rotor. For instance, the substrate may be affixed to a chuck (such as a vacuum chuck). The substrate may be rotated at a rotational speed of at least about 1 revolution per minute (rpm), at least 2 rpm, at least 5 rpm, at least 10 rpm, at least 20 rpm, at least 50 rpm, at least 100 rpm, at least 200 rpm, at least 500 rpm, at least 1,000 rpm, at least 2,000 rpm, at least 5,000 rpm, at least 10,000 rpm, or greater Alternatively or in addition, the substrate may be rotated at a rotational speed of at most about 10,000 rpm, 5,000 rpm, 2,000 rpm, 1,000 rpm, 500 rpm, 200 rpm, 100 rpm, 50 rpm, 20 rpm, 10 rpm, 5 rpm, 2 rpm, 1 rpm, or less. The substrate may be configured to rotate with different rotational velocities during different operations described herein, for example with higher velocities during reagent dispense and with lower velocities during analyte loading and imaging operations. The substrate may be configured to rotate with a rotational velocity that varies according to a time-dependent function, such as a ramp, sinusoid, pulse, or other function or combination of functions. The time-varying function may be periodic or aperiodic.
[0191]
[0121] Analytes or reagents may be immobilized to the substrate during rotation. Analytes or reagents may be dispensed onto the substrate prior to or during rotation of the substrate. When the substrate is rotated at a relatively high rotational velocity, high speed coating across the substrate may be achieved via tangential inertia directing unconstrained spinning reagents in a partially radial direction (that is, away from the axis of rotation) during rotation, a phenomenon commonly referred to as centrifugal force. In some cases, the substrate may be rotated at relatively low velocities such that reagents dispensed to a certain location do not move to another location, or moves minimally, because of the rotation, to permit controlled dispensing of reagents to desired locations. For example, bead loading may be performed with controlled dispensing. For controlled dispensing, the substrate may be rotating with a rotational frequency of no more than 60, 50, 40, 30, 25, 20, 15, 14, 13, 12, 11, 10, 9, 8, 7, 6, 5, 4, 3, 2, 1 rpm or less. In some cases the substrate may be rotating with a rotational frequency of about 5 rpm during controlled dispensing. A speed of substrate rotation may be adjusted according to the appropriate operation (e.g., high speed for spin-coating, high speed
[0192] -29-
[0193] SUBSTITUTE SHEET (RULE 26) for washing the substrate, low speed for sample loading, low speed for detection, low speed for analyte or reagent incubation, etc.).
[0194]
[0122] In some cases, the substrate may be movable in any vector or direction. For example, such motion may be non-linear (e.g., in rotation about an axis), linear (e g., on a rail track), or a hybrid of linear and non-linear motion. In some instances, the systems, devices, and apparatus described herein may further comprise a motion unit configured to move the substrate. The motion unit may comprise any mechanical component, such as a motor, rotor, actuator, linear stage, drum, roller, pulleys, etc., to move the substrate. Analytes or reagents may be immobilized to the substrate during any such motion. Analytes or reagents may be dispensed onto the substrate prior to, during, or subsequent to motion of the substrate.
[0195]
[0123] Reagents and / or analytes may be delivered to the surface of the substrate using one or more fluid nozzles. One or more nozzles may be configured to deliver fluids to the substrate as ajet, spray (or other dispersed fluid), and / or droplets. One or more nozzles may be operated to nebulize fluids prior to delivery to the substrate. For example, the fluids may be delivered as aerosol particles. In some cases, the reagents and / or analytes are delivered across a non-solid gap, such as an air gap. There may be any number of dispensing nozzles, such as 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 or more dispensing nozzles. In some cases, different reagents (e g., nucleotide solutions of different types, different probes, washing solutions, etc.) may be dispensed via different nozzles, such as to prevent contamination where each nozzle may be connected to a dedicated fluidic line or fluidic valve, which may further prevent contamination. Alternatively, some nozzles may share a fluidic line or fluidic valve, such as for pre-dispense mixing and / or to dispensing to multiple locations.
[0196]
[0124] In some cases, a solution may be dispensed on the substrate while the substrate is stationary; the substrate may then be subjected to rotation (or other motion) following the dispensing of the solution. Alternatively, the substrate may be subjected to rotation (or other motion) prior to the dispensing of the solution; the solution may then be dispensed on the substrate while the substrate is rotating (or otherwise moving). In some cases, rotation of the substrate may yield a centrifugal force (or inertial force directed away from the axis) on the solution, causing the solution to flow radially outward over the array. In this manner, rotation of the substrate may direct the solution across the array. Continued rotation of the substrate over a period of time may dispense a fluid film of a nearly constant thickness across the array.
[0197] -30-
[0198] SUBSTITUTE SHEET (RULE 26)
[0125] Reagents may be dispensed to the substrate to multiple locations, and / or multiple reagents may be dispensed to the substrate to a single location, via different mechanisms. Reagent dispensing mechanisms disclosed herein may be applicable to sample dispensing. For example, a reagent may comprise the sample. The term “loading onto a substrate,” as used herein, may refer to dispensing of the reagent or the sample to a surface of the substrate in accordance with any reagent dispensing mechanism described herein.
[0199]
[0126] In some cases, dispensing may be achieved via relative motion of the substrate and the dispenser (e.g., nozzle). For example, a reagent may be dispensed to the substrate at a first location, and thereafter travel to a second location different from the first location due to forces (e.g., centrifugal forces, centripetal forces, inertial forces, etc.) caused by motion of the substrate (e.g., rotational motion of the substrate, linear motion of the substrate, combination thereof, etc.). In another example, a reagent may be dispensed to a reference location, and the substrate may be moved relative to the reference location such that the reagent is dispensed to multiple locations of the substrate. In another example, a dispenser may be moved relative to the substrate to dispense the reagent at different locations, for example moved prior to, during, or subsequent to dispensing. In an example, a reagent is ‘painted’ onto the substrate by moving the dispenser and / or the substrate relative to each other, along a desired path on the substrate. The open substrate geometry may allow for flexible and controlled dispensing of a reagent to a desired location on the substrate. In some cases, dispensing may be achieved without relative motion between the substrate and the dispenser. For example, multiple dispensers may be used to dispense reagents to different locations, and / or multiple reagents to a single location, or a combination thereof (e.g , multiple reagents to multiple locations).
[0200] 1127] In another example, an external force (e.g., involving a pressure differential, involving physical force, involving a magnetic force, involving an electrical force, etc .), such as wind, a field-generating device, or a physical device, may be applied to one or more surfaces of the substrate to direct reagents to different locations across the substrate. In another example, the method for dispensing reagents may comprise vibration. In such an example, reagents may be distributed or dispensed onto a single region or multiple regions of the substrate. The substrate may then be subjected to vibration, which may spread the reagent to different locations across the substrate. Alternatively or in conjunction, the method may comprise using mechanical, electric, physical, or other mechanisms to dispense reagents to the substrate. For example, the solution may be dispensed onto a substrate and a physical scraper (e g., a squeegee) may be used to spread the dispensed material or spread the reagents to
[0201] -31-
[0202] SUBSTITUTE SHEET (RULE 26) different locations and / or to obtain a desired thickness or uniformity across the substrate. Beneficially, such flexible dispensing may be achieved without contamination of the reagents.
[0203]
[0128] In some instances, two or more reagents may be mixed on the surface of the substrate, such as by being dispensed at the same location and / or by directing a first reagent to travel to meet additional reagent(s). In some instances, the mixture of reagents formed on the substrate may be homogenous or substantially homogenous. The mixture of reagents may be formed at a first location on the substrate prior to dispersing the mixing of reagents to other locations on the substrate, such as at locations to meet other reagents or analytes.
[0204]
[0129] In some embodiments, one or more solutions may be delivered directly to the reaction site without substantial displacement of the one or more solution from the point of delivery. Methods of direct delivery of a solution to the reaction site may include aerosol delivery of the solution, applying the solution using an applicator, curtain-coating the solution, slot-die coating, dispensing the solution from a translating dispense probe, dispensing the solution from an array of dispense probes, dipping the substrate into the solution, or contacting the substrate to a sheet comprising the solution.
[0205]
[0130] The dispensed solution may comprise any sample or any analyte disclosed herein. The dispensed solution may comprise any reagent disclosed herein. In some cases, the solution may be a reaction mixture comprising a variety of components. In some cases, the solution may be a component of a final mixture (e.g., to be mixed after dispensing). In non-limiting examples, the solution can comprise samples, analytes, supports, beads, probes, nucleotides, oligonucleotides, labels (e.g., dyes), terminators (e g., blocking groups), other components to aid, accelerate, or decelerate a reaction (e.g., enzymes, catalysts, buffers, saline solutions, chelating agents, reducing agents, other agents, etc.), washing solution, cleavage agents, combinations thereof, deionized water, and other reagents and buffers.
[0206]
[0131] A sample may comprise beads, as described elsewhere herein, for example beads comprising single nucleic acid molecules (single- or double-stranded) or nucleic acid colonies bound thereto. In some cases, an order of magnitude of at least and / or at most about 101, 102, 103, 104, 105, 106, 107, 108, 109, IO10, 1011, 1012, 1013or more beads may be loaded on the substrate, such as to immobilize to as many individually addressable locations. In some cases, the beads may be distinguishable from one another using a property of the beads, such as color, reflectance, anisotropy, brightness, fluorescence, etc. In some cases, as described elsewhere herein, different beads may comprise different tags (e.g., nucleic acid
[0207] -32-
[0208] SUBSTITUTE SHEET (RULE 26) sequences) coupled thereto. For example, a bead may comprise an oligonucleotide molecule comprising a tag (e.g., barcode) that identifies a bead amongst a plurality of beads. FIG. 7 illustrates images of a portion of a substrate surface after loading a sample containing beads onto a substrate patterned with a substantially hexagonal lattice of individually addressable locations, where the right panel illustrates a zoomed-out image of a portion of a surface, and the left panel illustrates a zoomed-in image of a section of the portion of the surface.
[0209]
[0132] Dispense mechanisms described herein may be operated by a fluid flow unit which may be controlled by one or more controllers, individually or collectively. The fluid flow unit may comprise any of the hardware and software components described with respect to the dispense mechanisms herein.
[0210] 1133] An optical system comprising a detector may be configured to detect one or more signals from a detection area on the substrate pnor to, during, or subsequent to, the dispensing of reagents to generate an output. Signals from multiple individually addressable locations may be detected during a single detection event. Signals from the same individually addressable location may be detected in multiple instances.
[0211]
[0134] A signal may be an optical signal (e.g., fluorescent signal), electronic signal, or any detectable signal. The signal may be detected during rotation of the substrate or following termination of the rotation. The signal may be detected while the analyte is in fluid contact with a solution. The signal may be detected following washing of the solution. In some instances, after the detection, the signal may be muted, such as by cleaving a label from a probe and / or the analyte, and / or modifying the probe and / or the analyte. Such cleaving and / or modification may be performed by one or more stimuli, such as exposure to a chemical, an enzyme, light (e.g., ultraviolet light), or temperature change (e.g., heat). In some instances, the signal may otherw ise become undetectable by deactivating or changing the mode (e g , detection wavelength) of the one or more sensors, or terminating or reversing an excitation of the signal. In some instances, detection of a signal may comprise capturing an image or generating a digital output (e.g., between different images).
[0212]
[0135] The operations of (i) directing a solution to the substrate and (ii) detection of one or more signals indicative of a reaction between a probe in the solution and an analyte immobilized to the substrate, may be repeated any number of times by the system. Such operations may be repeated in an iterative manner. For example, the same analyte immobilized to a given location in the array may interact with multiple solutions in multiple cycles and for each iteration, the additional signals detected may provide incremental, or
[0213] -33-
[0214] SUBSTITUTE SHEET (RULE 26) final, data about the analyte during the processing. For example, when sequencing a nucleic acid molecule, additional signals detected for each iteration may be indicative of one or more bases in the nucleic acid sequence of the nucleic acid molecule. In some cases, multiple solutions can be provided to the substrate without intervening detection events. In some cases, multiple detection events can be performed after a single flow of solution. In some instances, a washing solution, cleaving solution (e.g., comprising cleavage agent), and / or other solutions may be directed to the substrate between each operation, between each cycle, or a certain number of times for each cycle.
[0215]
[0136] The optical system may be configured for continuous area scanning of a substrate during rotational motion of the substrate. The term “continuous area scanning (CAS),” as used herein, generally refers to a method in which an object in relative motion is imaged by repeatedly, electronically or computationally, advancing (clocking or triggering) an array sensor at a velocity that compensates for object motion in the detection plane (focal plane). CAS can produce images having a scan dimension larger than the field of the optical system. TDI scanning may be an example of CAS in which the clocking entails shifting photoelectric charge on an area sensor during signal integration. For a TDI sensor, at each clocking step, charge may be shifted by one row, with the last row being read out and digitized. Other modalities may accomplish similar function by high-speed area imaging and co-addition of digital data to synthesize a continuous or stepwise continuous scan.
[0216]
[0137] The optical system may comprise one or more sensors. The sensors may detect an image optically projected from the sample The optical system may comprise one or more optical elements. An optical element may be, for example, a lens, tube lens, prism, mirror, wave plate, filter, attenuator, grating, diaphragm, beam splitter, diffuser, polarizer, depolarizer, retroreflector, spatial light modulator, or any other optical element. The system may comprise any number of sensors. In some cases, a sensor is any detector as described herein. In some examples, the sensor may comprise image sensors, CCD cameras, CMOS cameras, TDI cameras (e.g., TDI line-scan cameras), pseudo-TDI rapid frame rate sensors, or CMOS TDI or hybrid cameras. The optical system may further comprise any one or more optical sources (e.g., lasers, LED light sources, etc ). In some cases, where there are multiple sensors, the different sensors may image the same or different regions of the rotating substrate, in some cases simultaneously. Each sensor of the plurality of sensors may be clocked at a rate appropriate for the region of the rotating substrate imaged by the sensor, which may be based on the distance of the region from the center of the rotating substrate or
[0217] -34-
[0218] SUBSTITUTE SHEET (RULE 26) the tangential velocity of the region. In some cases, multiple scan heads can be operated in parallel along different imaging paths (e.g., interleaved spiral scans, nested spiral scans, interleaved ring scans, nested ring scans). A scan head may comprise one or more of a detector element such as a camera (e.g., a TDI line-scan camera), an illumination source (e.g., as described herein), and one or more optical elements (e.g., as described herein).
[0219]
[0138] The system may further comprise one or more controllers operatively coupled to the one or more sensors, individually or collectively programmed to process optical signals from the one or more sensors, such as for each region of the rotating substrate.
[0220]
[0139] In some cases, the optical system may comprise an immersion objective lens. The immersion objective lens may be in contact with an immersion fluid that is in contact with the open substrate. The immersion fluid may comprise any suitable immersion medium for imaging (e.g., water, aqueous, organic solution). In some cases, an enclosure may partially or completely surround a sample-facing end of the optical imaging objective. The enclosure may be configured to contain the immersion fluid. The enclosure may not be in contact with the substrate; for example, a gap between the enclosure and the substrate may be filled by the fluid contained by the enclosure (e.g., the enclosure can retain the fluid via surface tension). In some cases, an electric field may be used to regulate a hydrophobicity of one or more surfaces of the container to retain at least a portion of the fluid contacting the immersion objective lens and the open substrate. In some cases, the immersion fluid may be continuously replenished or recycled via an inlet and outlet to the enclosure.
[0221]
[0140] One or more surfaces of the substrate may be exposed to and accessible from a surrounding open environment. In some cases, the surrounding open environment may be controlled and / or confined in a larger controlled environment. An open substrate may be processed within a modular local sample processing environment A barrier comprising a fluid barrier may be maintained between a sample processing environment and an exterior environment during certain processing operations, such as reagent dispensing and detecting. Systems and methods comprising a fluid barrier are described in further detail in U.S. Patent Pub. No. 20210354126A1, which is entirely incorporated herein by reference. A modular local sample processing environment may be defined by a chamber and a lid plate, where the lid plate is not in contact with the chamber, and the gap between the lid plate and the chamber may comprise the fluid barrier. The fluid barrier may comprise fluid (e.g., air) from the sample processing environment and / or the extenor environment and may have lower pressure
[0222] SUBSTITUTE SHEET (RULE 26) than the sample processing environment, the external environment, or both. The fluid in the fluid barrier may be in coherent motion or bulk motion.
[0223]
[0141] The sample processing environment may comprise therein a substrate, such as any substrate described elsewhere herein. Any operation performed on or with the substrate, as described elsewhere herein, may be performed within the sample processing environment while the fluid barrier is maintained. For example, the substrate may be rotated within the sample processing environment during various operations. In another example, fluid may be directed to the substrate while the substrate is in the sample processing environment, via a fluid handler (e.g., nozzle) that penetrates the lid plate into the sample processing environment. In another example, a detector can image the substrate while the substrate is in the sample processing environment, via a detector that penetrates the lid plate into the sample processing environment. Beneficially, the fluid barrier may help maintain temperature(s) and / or relative humidities), or ranges thereof, within the sample processing environment during various processing operations.
[0224]
[0142] The systems described herein, or any element thereof, may be environmentally controlled. For instance, the systems may be maintained at a specified temperature or humidity. For an operation, the systems (or any element thereof) may be maintained at a temperature of at least and / or at most 20 degrees Celsius (°C), 25 °C, 30 °C, 35 °C, 40 °C, 45 °C, 50 °C, 55 °C, 60 °C, 65 °C, 70 °C, 75 °C, 80 °C, 85 °C, 90 °C, 95 °C, 100 °C, or more. Different elements of the system may be maintained at different temperatures or within different temperature ranges, such as the temperatures or temperature ranges described herein. Elements of the system may be set at temperatures above the dew point to prevent condensation. Elements of the system may be set at temperatures below the dew point to collect condensation
[0225]
[0143] While examples described herein provide relative rotational motion of the substrates and / or detector systems, the substrates and / or detector systems may alternatively or additionally undergo relative non-rotational motion, such as relative linear motion, relative non-linear motion (e g., curved, arcuate, angled, etc.), and any other types of relative motion.
[0226]
[0144] An open substrate may be retained in the same or approximately the same physical location during processing of an analyte and subsequent detection of a signal associated with the processed analyte. Alternatively, different operations on or with the open substrate may be performed in different stations disposed in different physical locations. For example, a first station may be disposed above, below, adjacent to, or across from a second station. In
[0227] -36-
[0228] SUBSTITUTE SHEET (RULE 26) some cases, the different stations can be housed within an integrated housing. Alternatively, the different stations can be housed separately. In some cases, different stations may be separated by a barrier, such as a retractable barrier (e.g., sliding door). One or more different stations of a system, or portions thereof, may be subjected to different physical conditions, such as different temperatures, pressures, or atmospheric compositions. The open substrate may transition between different stations by transporting the sample processing environment comprising the chamber containing the open substrate between the different stations. One or more mechanical components or mechanisms, such as a robotic arm, elevator mechanism, actuators, rails, and the like, or other mechanisms may be used to transport the sample processing environment.
[0229]
[0145] One or more environmental units (e.g., humidifiers, heaters, heat exchangers, compressors, etc.) may be configured to, individually or collectively, regulate one or more operating conditions in one or more stations. In one example, the delivery and / or dispersal of reagents may be performed in a first station having a first operating condition, and the detection process may be performed in a second station having a second operating condition different from the first operating condition. The first station may be at a first physical location in which the open substrate is accessible to a fluid handling unit during the delivery and / or dispersal processes, and the second station may be at a second physical location in which the open substrate is accessible to the detector system.
[0230]
[0146] One or more modular sample environment systems (each having its own barrier system, e.g., fluid barrier) can be used between the different stations. In some instances, the systems described herein may be scaled up to include two or more of a same station type. For example, a sequencing system may include multiple processing and / or detection stations. FIGs. 5A-5B illustrate a system 500 that multiplexes two modular sample environment systems in a three-station system. In FIG. 5A, a first chemistry station (e.g., 520a) can operate (e.g., dispense reagents, e.g., to incorporate nucleotides to perform sequencing by synthesis) via at least a first operating unit (e.g , fluid dispenser 509a) on a first substrate
[0231] (e g., 511) in a first sample environment system (e.g., 505a) while substantially simultaneously, a detection station (e.g , 520b) can operate (e.g., scan) on a second substrate in a second sample environment system (e g., 505b) via at least a second operating unit (e g., detector 501), while substantially simultaneously, a second chemistry station (e.g., 520c) sits idle. An idle station may not operate on a substrate. An idle station (e.g., 520c) may be recharged, reloaded, replaced, cleaned, washed (e.g., to flush reagents), calibrated, reset, kept
[0232] -37-
[0233] SUBSTITUTE SHEET (RULE 26) active (e.g., power on), and / or otherwise maintained during an idle time. After an operating cycle is complete, the sample environment systems may be re-stationed, as in FIG. 5B, where the second substrate in the second sample environment system (e.g., 505b) is re-stationed from the detection station (e.g., 520b) to the second chemistry station (e.g., 520c) for operation (e.g., dispensing of reagents, e.g., to incorporate nucleotides to perform sequencing by synthesis) by the second chemistry station, and the first substrate in the first sample environment system (e.g., 505a) is re-stationed from the first chemistry station (e.g., 520a) to the detection station (e.g., 520b) for operation (e.g., scanning) by the detection station. An operating cycle may be deemed complete when operation at each active, parallel station is complete. During re-stationing, the different sample environment systems may be physically moved (e.g., along the same track or dedicated tracks, e.g., rail(s) 507) to the different stations and / or the different stations may be physically moved to the different sample environment systems. One or more components of a station, such as modular plates 503a, 503b, 503c of plate 503 (e.g., lid plate) defining a particular station(s), may be physically moved to allow a sample environment system to exit the station, enter the station, or cross through the station. During processing of a substrate at station, the environment of a sample environment region (e.g., 515) of a sample environment system (e.g., 505a) may be controlled and / or regulated according to the station’s requirements. After the next operating cycle is complete, the sample environment systems can be re-stationed again, such as back to the configuration of FIG. 5A, and this re-stationing can be repeated (e.g., between the configurations of FIGs. 5A and 5B) with each completion of an operating cycle until the required processing for a substrate is completed. In this illustrative re-stationing scheme, the detection station may be kept active (e.g., not have idle time not operating on a substrate) for all operating cycles by providing alternating different sample environment systems to the detection station for each consecutive operating cycle. Beneficially, use of the detection station is optimized. Based on different processing or equipment needs, an operator may opt to run the two chemistry stations substantially simultaneously while the detection station is kept idle.
[0234]
[0147] Beneficially, different operations within the system may be multiplexed with high flexibility and control. For example, as described herein, one or more processing stations may be operated in parallel with one or more detection stations on different substrates in different modular sample environment systems to reduce or eliminate lag between different sequences of operations (e.g., chemistry first, then detection). The modular sample environment systems
[0235] -38-
[0236] SUBSTITUTE SHEET (RULE 26) may be translated between the different stations accordingly to optimize efficient equipment use (e.g., such that the detection station is in operation almost 100% of the time). In some examples, at least 1, 2, 3, 4, 5, 6, 7, 8, 9, or more modules or stations of the sequencing system may be multiplexed. For example, 2 or more of the modules may each perform their intended function simultaneously or according to the methods described elsewhere herein. An example of this may comprise two-station multiplexing of an optics station and a chemistry station as described herein. Another example may comprise multiplexing three or more stations and process phases. For example, the method may comprise using staggered chemistry phases sharing a scanning station. The scanning station may be a high-speed scanning station. The modules or stations may be multiplexed using various sequences and configurations.
[0237] 1148] The nucleic acid sequencing systems and optical systems described herein (or any elements thereof) may be combined in a variety of architectures.
[0238] Single molecule sequencing methods
[0239]
[0149] Provided herein are devices, systems, methods, compositions, and kits that enable single molecule sequencing by synthesis. Such devices, systems, methods, compositions, and kits can be applied alternatively or in addition to the sequencing 107 operation described with respect to sequencing workflow 100 of FIG. 1 and, optionally, in the absence of preenrichment 102, 103, amplification of templates 105, or post-amplification processing 106. Such devices, systems, methods, compositions, and kits can be used in conjunction with the sample processing systems and methods, or components thereof (e g., substrates, detectors, reagent dispensing, continuous scanning, etc.) described herein.
[0240]
[0150] Single molecule sequencing provides many advantages over colony sequencing, specifically with regards to increased sequencing accuracy. In particular, for colony sequencing a template molecule must be amplified to produce a sequencing colony, and the sequencing signals (e.g., fluorescent signals) are an accumulation of signal indicative of nucleotide incorporations on the many amplified strands. Template amplification is susceptible to polymerase replication errors (e.g., inaccurate replication), and colony sequencing frequently exhibits phasing (e.g., the non-synchronous incorporation of nucleotide bases into amplified templates within a colony) and a concurrent decrease in base calling quality and hence a decrease in the confidence of sequence read accuracy. Advantageously, single-molecule sequencing bypasses these potential sources of error and hence is prone to fewer systemic sources of signal noise.
[0241] -39-
[0242] SUBSTITUTE SHEET (RULE 26)
[0151] The main source of noise in single molecule sequencing is scarring at the sites where labels are removed from incorporated nucleotides, which leads to incomplete sequencing. In typical sequencing by synthesis methods, once a labeled nucleotide has been incorporated into an extending primer and detected, the label must be removed in order to allow future cycles of nucleotide incorporation and detection. Such cleavage typically leaves scars, which tend to inhibit later incorporation events (e.g., block DNA polymerases). Thus, scarring can limit the length of sequence reads that can be obtained, a particular problem for single molecule sequencing. In colony sequencing, long sequence reads are possible because a sequence read is determined from the aggregate signals of hundreds or thousands of identical template molecules, Systems and methods described herein address these various issues.
[0243]
[0152] FIGs. 8A-8H illustrate non-limiting examples of sequencing by synthesis methods that are sufficient to enable high-accuracy single molecule sequencing.
[0244]
[0153] Compiling sequence reads across multiple sequencing runs
[0245]
[0154] In single molecule sequencing, signal cannot be aggregated from multiple copies of a template; thus, the use of partial labeling (e.g., only 10% of the incorporated nucleotides being labeled to reduce scaring) is not always practical. Instead, 100% labeled nucleotides of a particular base type may typically be used in each nucleotide flow of a sequencing run. In such cases, the sequence read may not be long (e g., due to scarring), but short sequence reads may be sufficient for some experiments (e.g., mRNA, cfDNA). Different methods may be utilized to both reduce scarring and obtain longer sequence reads. For example, in some cases, a sequence read may comprise multiple series of nucleotide flows, where a first set of nucleotide flows comprise 100% labeled nucleotides (e.g., bright flows) and a second set of nucleotide flows comprise unlabeled nucleotides (dark flows). Such sequencing will produce a sequence read including one or more gaps (e.g., sequence information will be obtained for a first region, no information will be gathered for a second intervening region (e.g., the dark flows), sequence information will be obtained for a subsequent third region, etc ). In postsequencing analysis, these non-contiguous sequencing reads may be associated with a single template molecule, and regions sequenced with unlabeled nucleotides can be inferred by alignment to reference sequences. Alternatively or in addition, a template molecule may be sequenced multiple times (e g., sequenced via a first sequencing run comprising alternating series of sets of bright and dark flows, having the extended sequencing primer removed, sequenced via a second sequencing run compnsing an offset alternating senes of sets of bright and dark flows, etc.). This resequencmg can increase the confidence in base calls for a
[0246] -40-
[0247] SUBSTITUTE SHEET (RULE 26) sequence run by obtaining repeat information for at least some regions of a template nucleic acid molecule and can enable obtaining a complete sequence for a template nucleic acid molecule without gaps. In some cases, resequencing permits flow orders without 100% labeling of nucleotides (e.g., with about 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 15%, 10%, 5%, etc. labeled nucleotides per flow). Alternatively, or in addition, multiple sequencing runs can be performed, where one, two, or three nucleotide base types are labeled at 100% (and where the remaining nucleotide base type(s) are unlabeled). In each case, after a sequencing run is completed, the extended sequencing pnmer can be removed (e.g., via thermal denaturing, chemical degradation, etc.). In cases where a sequencing read provides contiguous information for a respective region of a template nucleic acid molecule, in order to determine sequence information for all, substantially all, or a significant portion of a template nucleic acid molecule, multiple sequence reads each providing discrete information may be combined into a single consensus read.
[0248]
[0155] In FIG. SA, method 802 comprises one or more sequencing runs, each comprising a plurality of alternating series of bright and dark nucleotide flows. Bright nucleotide flows can comprise labeled nucleotides (e.g., flows with 100% labeled nucleotides). Dark nucleotide flows can comprise labeled nucleotides, where incorporation is not detected(e.g., no imaging). Alternatively or in addition, dark nucleotide flows can comprise unlabeled nucleotides. In some cases, a dark flow may comprise a single nucleotide base type. In some cases, a dark flow may comprise a plurality of nucleotide base types (e g., 2 base types, 3 base types, 4 base types, etc.). In some cases, one or more dark flows may be used (e.g., a first dark flow may comprise a first nucleotide base type comprising unlabeled nucleotides, a second dark flow may comprise three different base types comprising unlabeled nucleotides, and a third dark flow may comprise two different base types where a first base types comprises labeled nucleotides and the second base ty pe comprises unlabeled nucleotides). In some cases, it is sufficient to perform just the first sequencing run in method 802. In such cases, the sequence of regions of the template molecule interrogated with dark flows, can be determined by alignment of the resultant sequence read (e.g., comprising alternating regions with base calls and regions without - corresponding to regions interrogated with labeled and unlabeled nucleotides, respectively). In some cases, multiple sequencing runs are performed, where the series of bright and dark sequencing flows in each run are offset. For instance, the first sequencing run may comprise a first series of 50 bright flows, a first series of 10 dark
[0249] -41-
[0250] SUBSTITUTE SHEET (RULE 26) flows, a third series of 50 bright flows, etc. In some cases, the number of bright flows may be at least 4, 8, 12, 16, 20, 30, 40, 50, 60, 70, 80, 90, or 100 flows. In some cases, the number of intervening dark flows may be at least 4, 8, 12, 16, 20, or 30 flows. In some cases, the number of bright flows may be any number wherein a sequencing polymerase stalls (e.g., where incorporation of labeled nucleotides is inhibited, e.g., by scarring). In some cases, the number of dark flows may be any number that reduces polymerase stalling for subsequent incorporation of labeled nucleotides. In some cases, each series of bright flows may comprise a same number of flows. Alternatively, one or more series of bright flows in a sequencing run may comprise a different number of flows. In some cases, each series of dark flows may comprise a same number of flows. Alternatively, one or more series of dark flows in a sequencing run may comprise a different number of flows. In some cases, dark flows may each comprise a single nucleotide type (e.g., T or C, etc.). In some cases, one or more dark flows may be fast-forward flows comprising more than one nucleotide type (e.g., a A / T / C flow).
[0251]
[0156] In FIG. SB, method 804 comprises one or more sequencing runs, wherein the first sequencing run is the same as the first sequencing run in method 802. Subsequent sequencing runs may begin with a first series of dark flows (i.e., instead of a first series of bright flows). In some cases, for example in a second sequence run, a first series of dark flows may incorporate the same or approximately the same number of nucleotides as the first series of bright flows in the first sequence run. In some cases, for example in a second sequence run, a first series of dark flows may incorporate an unknown number of nucleotides (e g., where the first series of dark flows comprise multiple nucleotide types).
[0252]
[0157] In FIG. 8C, in method 806, an individual template molecule is sequenced multiple times, where each sequencing run comprises a single type of labeled nucleotide (e g , labeled T or labeled A or labeled C or labeled G). In method 806, base calls from the multiple sequence reads are compiled to form a single sequence read. In some cases, only three sequencing runs, each with a single type of labeled nucleotide, may be sufficient. In such cases, after the base calls from each run are combined into a sequence read, any uncalled loci can be assigned to the fourth nucleotide base type (e.g., if the first sequencing run has labeled A flows, the second sequencing run has labeled C flows, and the third sequencing run has labeled T flows, then after the base calls for each run are compiled into a sequence read any uncalled loci can be assigned as G bases).
[0253] -42-
[0254] SUBSTITUTE SHEET (RULE 26)
[0158] In FIG. 8D, method 808 compnses multiple sequencing runs. The first sequencing run comprises a plurality of bright sequencing flows (e.g., where each flow comprises 100% labeled nucleotides). The second sequencing run comprises a plurality of dark sequencing flows, wherein the plurality of dark flows in the second sequencing run comprises the same number of flows as in the plurality of bright flows in the first sequencing run. This pattern can be repeated, with each sequencing flow beginning with a successively larger number of dark sequencing flows and having a plurality of bright sequencing flows. Advantageously, this may be faster than the multiple sequencing flows performed in other methods 802, 804, and 806. The base calls from each sequencing run can be compiled to form a sequence read.
[0255]
[0159] Sequencing with blocked nucleotides
[0256]
[0160] In some cases, a plurality of dark flows may include one or more flows comprising blocking nucleotides. In some cases, blocking nucleotides may comprise cleavable moieties or cross-linking moieties or other linking moieties. A cross-link or other link may be cleaved to release the oligonucleotide molecules, or portion thereof, such as by applying one or more stimuli, including light stimuli, heat stimuli, chemical stimuli, magnetic stimuli, electrical stimuli, and other stimuli, or combination thereof. In some cases, a cross-linking moiety may be photo-cross-linking. A photo-cross-link may be generated by a photo cross-linking reaction. A crosslinking reagent may comprise a photolabile cross-linker, such as a nucleoside analogue 3- Cyanovinylcarbazole nucleosides (CNVK). Photolabile cross-linkers can enable ultra-fast reversible photo-cross-linking of oligonucleotides. When an oligonucleotide strand comprisesCNVK, a cross-link is formed betweenCNVK and a pyrimidine base on the complementary strand when illuminated at 365 nm. A complementary nucleotide on the oligonucleotide strand is in 5’ immediately preceding toCNVK. In some instances, an oligodeoxynucleotide (ODN) comprising 3-cyanovinylcarbazole nucleoside (CNVK) can be subjected to photoirradiation conditions to photo-cross-link a target pyrimidine and theCNVK. In some instances, irradiation is provided at 366 nm for about 1 second for photo-cross-linking to thymine, and for up to about 25 seconds for photo-cross-linking to cytosine. Complete reversal of the cross-link is achieved by illumination at 312 nm (or at 315 nm) (e.g., irradiation provided at 312 nm for about 3 minutes). Other cross-linkers include, for example,CNVD,PCX, orPCXD. Various other cross-linking reagents may be used to generate a cross-link (e.g., chemical cross-link). In some cases, heat may be applied to denature a double-stranded molecule to facilitate release the unblocked portion of the extended primer from the template molecule (e.g., to facilitate additional sequencing runs).
[0257] -43-
[0258] SUBSTITUTE SHEET (RULE 26) The heat stimulus can be combined with cross-linking reactions. For example, the unblocked portion of the extended primer may be released from the template strand by applying a heat stimulus.
[0259]
[0161] Preferentially, flows comprising denaturation-blocking nucleotides will occur later in a given plurality of dark flows (e.g., resulting in at least one incorporated blocking reagent prior to the succeeding plurality' of bright flows). FIGs. 8E and 8F illustrate schema for multiple sequencing runs using uracils as blocking nucleotides. These uracils may be used as cleavage moieties (e.g., by USER enzyme). This enables bright sequencing flows to restart from the cleaved sites instead of requiring reannealing of primer and in. Advantageously, this may increase the speed of sequencing (e.g., by decreasing the number of dark flows required in each sequencing run). For example, as illustrated in FIG. 8D, the number of flows in the second plurality of dark flows (in sequencing run 3) may be approximately twice the number of flows in the first plurality of dark flows. In contrast, in FIGs. 8E and 8F, the number of flows in each plurality of dark flows may be approximately the same.
[0260]
[0162] In these re-sequencing methods, base calls from multiple sequencing run can be compiled to form a sequence read with an overall decreased error rate. In some cases, sequencing reads corresponding to a same template molecule are identified by barcode sequences. In some cases, sequencing reads are associated by alignment to a reference sequence. In some cases, sequencing reads are identified as corresponding to a same template molecule by location (e.g., by signals from multiple sequencing runs being detected at a respective individually addressable location).
[0261]
[0163] Multiple sequencing primers
[0262]
[0164] FIG. 8G illustrates various options for performing multiple sequencing runs with a same template nucleic acid molecule. In some cases, a template nucleic acid may be particularly long and / or may contain specific regions of interest. In such cases, it may be advantageous to sequence only the regions or interest (e.g., an adapter sequence, a barcode sequence, an insert sequence, etc.).
[0263]
[0165] Multiple sequencing runs are illustrated in 814 and may be performed by annealing a first pnmer to a template, performing a first extension (e.g., first sequencing run), denatunng the first extended primer, annealing another first primer (e g., another primer with the same or substantially same sequence as the first primer) to the template, and performing a second extension (e.g., second sequencing run). In some cases, the first extension and the second extension may comprise a same flow cycle. In some cases, the flow cycle may comprise a
[0264] -44-
[0265] SUBSTITUTE SHEET (RULE 26) combination of bright sequencing flows and dark sequencing flows (e.g., as described with respect to FIGs. 8A, 8B, 8D, 8E, and 8F). In some cases, the flow cycle may comprise bright sequencing flows comprising at least one type of labeled nucleotide base (e.g., as described with respect to FIG. 8C). Alternatively, or in addition, resequencing 814 may be performed, for example, by annealing a first primer to a template, performing a first extension (e.g., first sequencing run), denaturing the first extended primer, annealing a second primer (e.g., another primer with a different sequence from the first primer and / or another primer that anneals to a different region of the template molecule), and performing a second extension (e g., second sequencing run) In some cases, a first flow cycle may be used for the first extension and a second flow cycle may be used for the second extension. In some cases, the first flow cycle and the second flow cycle are the same. In some cases, the first flow cycle and the second flow cycle are different (e.g., comprise different flow orders or different combinations of bright sequencing flows and dark sequencing flows).
[0266]
[0166] In some cases, the first primer and the second primer may hybridize to the same region of the template nucleic acid molecule. In some cases, one or both of the first and second primers may hybridize to an adapter region of the template nucleic acid molecule. In some cases, one or both of the first and second primers may hybridize to a polyT or polyA region of the template nucleic acid molecule (e.g., for RNA sequencing).
[0267]
[0167] Any number and combination of sequencing runs may be performed. For example, a first sequencing run may comprise: i) hybridizing a first pnmer to a first primer site on the template molecule, ii) extending the first primer according to a first flow cycle order (e g., to obtain first sequencing data), iii) denaturing and removing the extended first primer, iv) repeating the hybridizing, extending - where the extending may be according to the first flow cycle order or a different flow cycle order, and denaturing of another first primer one or more times, v) hybridizing a second primer to a second primer site on the template molecule, vi) extending the second primer according to a second flow cycle order (e.g., to obtain first sequencing data), vii) denaturing and removing the extended second primer, and viii) repeating the hybridizing, extending, and denaturing of another second primer one or more times. As described elsewhere, in some cases the second flow cycle order may be the same as or different from the first flow cycle order. Further, in some cases the first repeating (iv) and the second repeating (viii) may comprise using the same respective flow cycle order or a different respective flow cycle order in one or more of the repetitions.
[0268] -45-
[0269] SUBSTITUTE SHEET (RULE 26)
[0168] In some cases, the first sequence data may correspond to a first region of the template, and the second sequence data may correspond to a second region of the template. In some cases, the first region may comprise a barcode, id, adapter, or insert sequence; second region may be barcode, id, adapter, or insert sequence. In some cases, Improved seq quality for long templates.
[0270]
[0169] Single molecule paired end sequencing
[0271]
[0170] FIG. 8H illustrates a non-limiting schema for paired end sequencing comprising multiple sequencing runs. Beneficially, paired end sequencing may provide higher accuracy sequencing data for a single stranded template nucleic acid molecule (e.g., for a region of the template molecule for which overlapping sequencing information is obtained). In addition, in sequencing-by-synthesis methods, sequencing quality (see e.g., FIG. 10D) tends to degrade as the sequencing primer is extended. Beneficially, paired end sequencing may enable higher accuracy sequencing at each end of a template molecule (e.g., because the first sequencing flows or steps tend to have higher base call quality scores). Either or both sequencing runs in paired end sequencing may use non-terminated nucleotides, terminated nucleotides, or a mixture of both.
[0272]
[0171] A template molecule may be coupled to a support (e.g., a sequencing bead, dendrimer, DNA nanoball, substrate, etc.) at one end. In step 850, a first sequencing primer is annealed (e g., is hybridized) to a template molecule at a first primer binding site. The first primer binding site may be distal to the end of the template molecule coupled to the support. The first primer binding site may be adjacent to the end of the template molecule coupled to the support. For a first number of sequencing flows (e.g., a number of bright flows), or for a first region of the template molecule, labeled nucleotides are added and are incorporated into the extending first primer. In each of the first number of sequencing flows, each incorporated nucleotide may be detected. After a detection step, the labeling moiet(ies) and / or the terminating moiety may be removed (e.g., cleaved) from incorporated nucleotide.
[0273]
[0172] At least a subset or all of the nucleotides added during bright flows may be labeled. The nucleotides added in each bright flow comprise four, three, two, or one canonical base types. The nucleotides may be reversibly terminated or non-terminated. In some cases, the nucleotides of each base type may comprise a respective label moiety. In some cases, each respective label moiety may be a different fluorescent label (e.g., fluorescent moieties with different excitation / emission spectra). In some cases, nucleotides of one of the added base types are unlabeled. In some cases, for reversibly terminated sequencing, each respective
[0274] -46-
[0275] SUBSTITUTE SHEET (RULE 26) label moiety may comprise a different number of the same fluorescent label (e.g., where a first base type is labeled with one fluorescent moiety of a first type and a second base type is labeled with three fluorescent moieties of the same type).
[0276]
[0173] In step 852, for a second number of sequencing flows (e.g., a number of dark flows), or for a second region of the template molecule, nucleotides are added and are incorporated into the extending first primer. In the second number of sequencing flows, at most only a subset of incorporated nucleotides is detected, or incorporation is not detected. In some cases, nucleotides added in the second number of sequencing flows are labeled and unterminated. In some cases, at least some of the nucleotides in the second number of sequencing flows are unlabeled and unterminated. In some cases, at least some of the nucleotides in the second number of sequencing flows are terminated. The dark flows may comprise nucleotides of one, two, three, or four canonical base types. In some cases, nucleotides of one of the added base types are reversibly terminated. In some cases, nucleotides of one or more of the added base types are labeled. In some cases, nucleotides added in dark flows may be unlabeled.
[0277]
[0174] The extended first sequencing primer comprises a copied template molecule (e.g., a molecule that is a reverse complement to the template nucleic acid molecule). After the second number of sequencing flows (e.g., the dark flows), the copied template molecule and the template molecule are denatured (e.g., exposed to conditions sufficient to denature the copied template molecule from the template molecule). The copied template molecule may be immobilized to the substrate surface For example, in 850, the first sequencing primer may be conjugated to the surface (e g , also at an individually addressable location), and the template molecule (e.g., concatemer) may be annealed to the first sequencing primer such that the extended molecule, the copied template molecule is conjugated to the surface and upon denaturation the template molecule is washed away. In some cases, the template molecule may also be coupled to the surface (e.g., at the individually addressable location), and thus upon denaturation the template molecule may remain coupled to the substrate. Beneficially, retaining both the template molecule and the copied template molecule coupled to the substrate enables resequencing or multiple sequencing runs, as described elsewhere herein. In some cases, the substrate may comprise a second sequencing primer and the copied template molecule, upon denaturation, may anneal to the second sequencing primer, thus immobilizing the copied template molecule to the substrate.
[0278]
[0175] In step 854, a second sequencing primer is annealed (e.g., is hybridized) to a second sequencing pnmer binding site in the copied template molecule. The second sequencing
[0279] -47-
[0280] SUBSTITUTE SHEET (RULE 26) primer is extended along the copied template molecule via a first plurality of bright flows followed by a plurality of dark flows. For a first number of sequencing flows (e.g., a number of bright flows), or for a first region of the copied template molecule, labeled nucleotides are added and are incorporated into the extending second primer (e.g., nucleotides comprising a labeling moiety and a reversibly terminating moiety). In each of the first number of sequencing flows, each incorporated nucleotide may be detected. After detection, the labeling moiety and / or the terminating moiety is removed (e.g., cleaved) from the incorporated nucleotide. For a second number of sequencing flows (e.g., a number of dark flows), or for a second region of the template molecule, nucleotides are added and are incorporated into the extending second primer. In the second number of sequencing flows, at most only a subset of incorporated nucleotides is detected, or no incorporation is detected. That is, detection steps may be performed every n flow, where n is an integer greater than 1. N may be 2, 3, 4, 5, 6, 7, 8, 9, 10, etc. In some cases, nucleotides added in the second number of sequencing flows are unterminated. In some cases, at least some of the nucleotides in the second number of sequencing flows are unlabeled. The dark flows may comprise nucleotides of one, two, three, or four canonical base types. In some cases, nucleotides of one or more of the added base types are reversibly terminated. In some cases, nucleotides of one or more of the added base types are labeled.
[0281]
[0176] The first sequencing primer binding site is located at or adjacent to the 3’ end of the template molecule, and the second sequencing primer binding site is located at or adjacent to the 3’ end of the copied template molecule. In some cases, there will be an overlap in loci covered by bright flows in the template molecule and the copied template molecule (see ‘Overlap region’ in step 854). Such loci that have been sequenced twice with bright flows will have decreased base call error rates than loci that were only sequenced once with bright flows. Further, as sequencing progresses along a length of a template, the sequencing quality may decrease due to phasing problems (e.g., leading and lagging signals due to misincorporation and / or uncompleted extension reactions during each extension steps) — beneficially, this method permits collection of higher quality signals corresponding to both ends of the template molecule, one from each of the template and copied template molecules.
[0282]
[0177] In some cases, a method for paired end sequencing may comprise, for first strand sequencing, extending a first sequencing primer annealed to a first primer binding site in the template concatemer via bright steps (e.g., with detection) until the signal quality drops below a predetermined threshold, and then continuing to extend via dark steps (e.g., without
[0283] -48-
[0284] SUBSTITUTE SHEET (RULE 26) detection) with a strand displacing polymerase. Because the template is a concatemer, multiple first sequencing primers may be simultaneously extended from multiple first primer binding sites in the concatemer, and one extending primer molecule may eventually displace another extending primer molecule from the concatemer as the extension steps progress. The dark step extensions may be terminated, such as by incorporating a ddNTP. The extended products, which each comprises a reverse complement of the template, may comprise a reverse primer binding site. The method may further comprise, for second strand sequencing, annealing a second sequencing primer to the reverse primer binding site and extending the second sequencing primer via bright steps (e.g., with detection). The sequencing reads generated from the first and second strands may be processed as paired end reads.
[0285]
[0178] In some cases, a method for paired end sequencing may comprise, for first strand sequencing, extending a first sequencing primer annealed to a first primer binding site in the template concatemer via bright steps (e.g., with detection) until the signal quality drops. A strand displacement primer may then be annealed to a second primer binding site on the template concatemer, and extended via dark steps (e g., without detection) with a strand displacing polymerase. Because the template is a concatemer, multiple primers may be simultaneously extended from multiple primer binding sites in the concatemer, and one extending primer molecule may eventually displace another extending primer molecule from the concatemer as the extension steps progress. The dark step extensions may be terminated, such as by incorporating a ddNTP. The extended products, which each comprises a reverse complement of the template, may comprise a reverse primer binding site. The method may further comprise, for second strand sequencing, annealing a second sequencing primer to the reverse primer binding site and extending the second sequencing primer via bnght steps (e.g., with detection) The sequencing reads generated from the first and second strands may be processed as paired end reads. Addition of the strand displacement primer may accelerate the kinetics. In some cases, the reverse primer binding site may be the reverse complement of one of the first primer binding site and the second primer binding site.
[0286]
[0179] In some cases, a method for paired end sequencing may comprise: extending a first sequencing primer annealed to a first primer binding site in the template concatemer via bright steps (e g., with detection). Then, a strand displacement primer may be annealed to a second primer binding site on the template concatemer, and extended via dark steps (e.g., without detection) with a strand displacing polymerase. Because the template is a concatemer, multiple primers may be simultaneously extended from multiple primer binding
[0287] -49-
[0288] SUBSTITUTE SHEET (RULE 26) sites in the concatemer, and one extending primer molecule may eventually displace another extending primer molecule from the concatemer as the extension steps progress The dark step extensions may be terminated, such as by incorporating a ddNTP. The extended products, which each comprises a reverse complement of the template, may comprise a reverse primer binding site. The method may further comprise, for second strand sequencing, annealing a second sequencing primer to the reverse primer binding site and extending the second sequencing primer via bright steps (e.g., with detection). The sequencing reads generated from the first and second strands may be processed as paired end reads. Addition of the strand displacement primer may accelerate the kinetics. In some cases, the reverse primer binding site may be the reverse complement of one of the first primer binding site and the second primer binding site.
[0289]
[0180] In some cases, the sequencing reads generated from the first and second strands may be associated with each other based on location of sequencing (e.g., both sequencing reads may be localized to a same individually addressable location).
[0290]
[0181] A sequencing by synthesis method, which may or may not be paired end, may comprise any number of bright steps and any number of dark steps. A sequencing by synthesis method may comprise any number of bright regions (consecutive bright steps) and any number of dark regions (consecutive dark steps). In some cases, the dark steps or dark regions may be used to accelerate or fast forward through certain regions of the template during sequencing. In some cases, the dark steps or dark regions may be advantageous to correct phasing problems.
[0291]
[0182] In some cases, a method for sequencing may comprise sequencing a same template strand multiple times to generate robust sequencing data (e.g., a high-quality sequencing read) corresponding to the template strand. In some cases, a method for sequencing may comprise sequencing a same template strand multiple times and sequencing a same reverse complement strand of the template strand multiple times (e.g., both forward and reverse strands) to generate robust sequencing data (e.g., a high-quality paired end read) corresponding to the template strand. A method for re-sequencing a template strand (which may be a forward strand or reverse strand) may comprise annealing a first sequencing primer to the template strand, extending the first sequencing primer through at least a first portion of the template strand via any combination of bright steps and / or dark steps to generate first sequencing data, denaturing the extended strand from the template strand, annealing a second sequencing primer to the template strand, and extending the second sequencing primer
[0292] -50-
[0293] SUBSTITUTE SHEET (RULE 26) through at least a second portion of the template strand via any combination of bright steps and / or dark steps to generate second sequencing data, and processing (e.g., combining, comparing, matching, aligning, resolving, etc.) the first sequencing data and the second sequencing data to generate a sequencing read of the template strand. A template strand may be denatured and re-sequenced any number of times, such as about, at least about, and / or at most about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more times, such as by annealing an nth sequencing primer to the template strand, and extending the nth sequencing primer through at least an nth portion of the template strand. The different n sequencing primers may comprise the same or different sequences, which may bind to same or different primer binding sites on the template strand, respectively. The different «th portions on the template strand may refer to the same portions or different portions on the template strand. Two portions may be partially overlapping, completely overlapping (for one or both portions), or non-overlapping. The respective extensions through the template strand in the different sequencing runs may use the same or different nucleotide reagents (e.g., non-terminated nucleotides during a first sequencing run, terminated during a second sequencing run; green dye-labeled nucleotides during a first sequencing run, red dye-labeled nucleotides during a second sequencing run; labeled A-, T-, G- bases and unlabeled C-base nucleotides during a first sequencing run, labeled A-, T-, C- bases and unlabeled G-base nucleotides during a second sequencing run; 5% labeled A bases during a first sequencing run; 100% labeled A bases during a second sequencing run; etc.). The respective extensions through the template strand in the different sequencing runs may have the same flow order or flow cycle of nucleotide reagents The respective extensions through the template strand in the different sequencing runs may have different flow orders or flow cycles of nucleotide reagents (e.g., A -> T -> G -> C single base flow cycle order during a first sequencing run, T -> A -> G -> C single base flow cycle order during a second sequencing run; A / T / G / C 4-base flow cycle order during a first sequencing run; A / T / G -> A / T / C 3-base flow cycle order during a second sequencing run, etc.).
[0294] Denaturing may comprise contacting the double-stranded nucleic acid molecule with denaturing agents, such as sodium hydroxide (NaOH) or ethylene carbonate. An entire substrate may be subjected to resequencing by, after a first sequencing run, contacting the entire surface with a solution comprising a denaturing agent, contacting the entire surface with a solution comprising sequencing primers under conditions sufficient to anneal them to template nucleic acid strands immobilized to the substrate, and subjecting them to extension reactions.
[0295] -51-
[0296] SUBSTITUTE SHEET (RULE 26)
[0183] Re-sequencing
[0297]
[0184] In some cases, a method for sequencing may comprise sequencing a same template strand multiple times to generate robust sequencing data (e.g., a high-quality sequencing read) corresponding to the template strand. This may be especially important in single molecule sequencing. In some cases, a method for sequencing may comprise sequencing a same template strand multiple times and / or sequencing a same reverse complement strand of the template strand multiple times to generate robust sequencing data (e.g., a high-quality paired end read) corresponding to the template strand. A method for re-sequencing a template strand may comprise annealing a first sequencing primer to the template strand, extending the first sequencing primer through at least a first portion of the template strand via any combination of bright steps and / or dark steps to generate first sequencing data, denaturing the extended strand from the template strand, annealing a second sequencing pnmer to the template strand, and extending the second sequencing pnmer through at least a second portion of the template strand via any combination of bright steps and / or dark steps to generate second sequencing data, and processing (e.g., combining, comparing, matching, aligning, resolving, etc.) the first sequencing data and the second sequencing data to generate a sequencing read of the template strand. A template strand may be denatured and resequenced any number of times, such as about, at least about, and / or at most about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more times, such as by annealing an nth sequencing primer to the template strand and extending the nth sequencing primer through at least an nth portion of the template strand. The different n sequencing primers may comprise the same or different sequences which may bind to same or different primer binding sites on the template strand, respectively. The different nth portions on the template strand may refer to the same portions or different portions on the template strand. Two portions on the template strand (that are extended through) may be partially overlapping, completely overlapping (for one or both portions), or non-overlapping. The respective extensions through the template strand in the different sequencing runs may use the same or different sequencing methods (e.g , non-terminated sequencing during a first sequencing run, terminated during a second sequencing run; terminated sequencing during a second sequencing run; etc.). The respective extensions through the template strand in the different sequencing runs may use the same or different nucleotide reagents (e.g., non-terminated nucleotides during a first sequencing run, terminated during a second sequencing run; green dye-labeled nucleotides during a first sequencing run, red dye-labeled nucleotides during a second sequencing run; labeled A-, T-,
[0298] -52-
[0299] SUBSTITUTE SHEET (RULE 26) G- bases and unlabeled C-base nucleotides during a first sequencing run, labeled A-, T-, C- bases and unlabeled G-base nucleotides during a second sequencing run; 5% labeled A bases during a first sequencing run; 100% labeled A bases during a second sequencing run; etc.). The respective extensions through the template strand in the different sequencing runs may have the same flow order or flow cycle of nucleotide reagents. The respective extensions through the template strand in the different sequencing runs may have different flow orders or flow cycles of nucleotide reagents (e.g., A -> T -> G -> C single base flow cycle order during a first sequencing run, T -> A -> G -> C single base flow cycle order during a second sequencing run; A / T / G / C 4-base flow cycle order during a first sequencing run; A / T / G -> A / T / C 3-base flow cycle order during a second sequencing run, etc.). Denaturing may comprise contacting the double-stranded nucleic acid molecule with denaturing agents, such as sodium hydroxide (NaOH) or ethylene carbonate. An entire substrate may be subjected to resequencing by, after a first sequencing run, contacting the entire surface with a solution comprising a denaturing agent, contacting the entire surface with a solution comprising sequencing primers under conditions sufficient to anneal them to template nucleic acid strands immobilized to the substrate, and subjecting them to extension reactions. In some cases, denaturing may comprise applying heat to the double-stranded nucleic acid molecule.
[0300]
[0185] Additional sequencing schemes are described in U.S. Pat. Pub. Nos.
[0301] 2021 / 0017593 Al, 2022 / 0064728A1, and 2022 / 0154272A1, each of which is entirely incorporated herein by reference for all purposes.
[0302]
[0186] In some cases, after generating a first sequencing read with a first sequencing primer, and after denaturing a first extension product of the first sequencing primer, a second sequencing pnmer may hybridize to a portion corresponding to an adapter of the template nucleic acid molecule and be extended to generate a second sequencing read. In some cases, the second sequencing primer may also be extended through a second barcode region. Such separate read generation can relax the limitation on sheared read length in genomic DNA sequencing and allows reading mixed di-nucleosome cfDNA molecules.
[0303]
[0187] Pre-Sequencing Treatment
[0304]
[0188] In some cases, denaturing may be performed at other steps in a sequencing workflow, for example prior to sequencing or as part of loading template molecules onto a substrate.
[0305]
[0189] After the production of positive supports (e.g., sequencing beads comprising a template nucleic acid molecule coupled thereto), whether they compnse single-stranded or double-stranded nucleic acid molecules, the positive supports may be treated with conditions
[0306] -53-
[0307] SUBSTITUTE SHEET (RULE 26) for stripping (e.g., denaturing or otherwise removing sequencing primers) and / or conditions for re-hybridization of sequencing primers. In some cases, the positive supports may be enriched (e.g., isolated from negative supports) prior to loading onto a substrate. In some cases, the positive supports may not be enriched (e.g., negative supports are also present) prior to loading onto a substrate. Methods of loading such as those described with respect to FIGs. 12A-12L and elsewhere herein may be combined with any pre-sequencing treatments described herein.
[0308]
[0190] The conditions for stripping may comprise treatment with a denaturing agent, such as sodium hydroxide (NaOH) or ethylene carbonate, heating, and / or a combination thereof. The conditions for re-hybridization of sequencing primers may comprise any conditions for hybridization of nucleic acid molecules. A plurality of sequencing primers may be provided for re-hybridization to the template nucleic acid strands. The stripping and hybridization of sequencing primers may be performed simultaneously (e.g., reagents provided in the same mixture) or separately. The respective conditions for stripping and hybridization may be provided to a reaction space comprising the template nucleic acid molecules simultaneously or separately. The sequencing primers to be stripped may have a same sequence as the sequencing primers provided for re-hybridization. The sequencing primers to be stripped may have a different sequence as the sequencing primers provided for re-hybridization (e.g., where the re-hybridization sequencing primers may anneal to a same or different region of a template molecule as the sequencing primers to be stripped).
[0309]
[0191] The enriched, positive supports may be treated to stripping conditions on the substrate, i.e., after loading on the substrate. Alternatively, the enriched, positives supports may be treated to stripping conditions off the substrate, i.e., prior to loading on the substrate. The enriched, positive supports may be treated to re-hybridization conditions on the substrate, i.e., after loading on the substrate. Alternatively, the enriched, positives supports may be treated to re-hybridization conditions off the substrate, i.e., prior to loading on the substrate, such that when they are loaded, the enriched, positive supports comprise sequencing primers re-hybridized to template strands thereon.
[0310]
[0192] In some cases, the positive supports may be treated with multiple rounds of conditions for stripping, conditions for re-hybridization, and / or both in any sequence. It was unexpectedly discovered that treating enriched supports (e.g., positive supports each comprising one or more nucleic acid molecules) with conditions for stripping, conditions for
[0311] -54-
[0312] SUBSTITUTE SHEET (RULE 26) re-hybridization of sequencing primers, and / or both, prior to sequencing the nucleic acid molecules resulted in a significant improvement in sequencing quality.
[0313]
[0193] For example, provided herein is a method for sequencing data generation. The method comprises loading a plurality of beads each comprising a double-stranded template nucleic acid molecule attached thereto, onto a substrate; on the substrate, denaturing the doublestranded template nucleic acid molecules to generate a single-stranded template nucleic acid molecule attached each bead, and hybridizing a sequencing primer to one or more the of single-stranded template nucleic acid molecules; and generating the sequencing data on the single-stranded template nucleic acid molecules by extending the sequencing primers.
[0314]
[0194] Another method for sequencing data generation may comprise: loading a plurality of beads each comprising a single-stranded template nucleic acid molecule attached thereto, onto a substrate; on the substrate, subjecting the single-stranded template nucleic acid molecules to denaturing conditions, and hybridizing a sequencing primer to one or more the of single-stranded template nucleic acid molecules; and generating the sequencing data on the single-stranded template nucleic acid molecules by extending the sequencing primers.
[0315]
[0195] Another method for sequencing data generation may comprise: loading a plurality of beads each comprising a single-stranded template nucleic acid molecule attached thereto, onto a substrate; on the substrate, hybridizing a sequencing primer to one or more of the single-stranded template nucleic acid molecules, subjecting the single-stranded template nucleic acid molecules to denaturing conditions, and hybridizing a sequencing primer to one or more the of single-stranded template nucleic acid molecules; and generating the sequencing data on the double-stranded template nucleic acid molecules by extending the sequencing primers.
[0316]
[0196] Another method for sequencing data generation may comprise: loading a plurality of beads each comprising a single-stranded template nucleic acid molecule attached thereto, onto a substrate, wherein a first sequencing primer is hybridized to one or more of the singlestranded template nucleic acid molecules; on the substrate, denaturing the first sequencing primers from the single-stranded template nucleic acid molecules, and hybridizing a second sequencing primer to one or more the of single-stranded template nucleic acid molecules; and generating the sequencing data on the double-stranded template nucleic acid molecules by extending the sequencing pnmers.
[0317]
[0197] Another method for sequencing data generation may comprise: loading a plurality of beads each comprising a single-stranded template nucleic acid molecule attached thereto,
[0318] -55-
[0319] SUBSTITUTE SHEET (RULE 26) onto a substrate, wherein a first sequencing primer is hybridized to one or more of the singlestranded template nucleic acid molecules; on the substrate, generating a first set of sequencing data on the plurality of single-stranded template nucleic acid molecules by extending the first sequencing primers; on the substrate denaturing the extended first sequencing primers from the single-stranded template nucleic acid molecules, and hybridizing a second sequencing primer to one or more the of single-stranded template nucleic acid molecules; generating a second set of sequencing data on the double-stranded template nucleic acid molecules by extending the sequencing primers; and combining first and second sets of sequencing data to obtain sequencing data on the single-stranded template nucleic acid molecules.
[0320]
[0198] Another method for sequencing data generation may comprise: loading a plurality of beads each comprising a plurality of double-stranded template nucleic acid molecules attached thereto, onto a substrate; on the substrate, denaturing the double-stranded template nucleic acid molecules to generate a plurality of single-stranded template nucleic acid molecule attached each bead, and hybridizing a sequencing primer to one or more the of single-stranded template nucleic acid molecules; and generating the sequencing data on the single-stranded template nucleic acid molecules by extending the sequencing primers.
[0321]
[0199] Another method for sequencing data generation may comprise: loading a plurality of beads each comprising a plurality of single-stranded template nucleic acid molecules attached thereto, onto a substrate; on the substrate, hybndizmg a plurality of sequencing primers to the plurality of smgle-stranded template nucleic acid molecules, subjecting the plurality of single-stranded template nucleic acid molecules to denaturing conditions, and hybridizing a plurality of sequencing primers to the plurality of single-stranded template nucleic acid molecules; and generating the sequencing data on the single-stranded template nucleic acid molecules by extending the sequencing primers.
[0322]
[0200] Another method for sequencing data generation may comprise: loading a plurality of beads comprising a plurality of single-stranded template nucleic acid molecules attached thereto, onto a substrate; on the substrate, subjecting the single-stranded template nucleic acid molecules to denaturing conditions, and hybridizing a plurality of sequencing pnmers to the plurality of smgle-stranded template nucleic acid molecules; and generating the sequencing data on the plurality of single-stranded template nucleic acid molecules by extending the sequencing pnmers
[0323] -56-
[0324] SUBSTITUTE SHEET (RULE 26)
[0201] Another method for sequencing data generation may comprise: loading a plurality of beads comprising a plurality of single-stranded template nucleic acid molecules attached thereto, onto a substrate, wherein a plurality of first sequencing primers is hybridized to the plurality of single-stranded template nucleic acid molecules; on the substrate, denaturing the plurality of first sequencing primers from the plurality of single-stranded template nucleic acid molecules and re-hybridizing a plurality of second sequencing primers to the plurality of single-stranded template nucleic acid molecules; and generating the sequencing data on the plurality of single-stranded template nucleic acid molecules by extending the plurality of second sequencing primers.
[0325]
[0202] Another method for sequencing data generation may comprise: loading a plurality of beads comprising a plurality of single-stranded template nucleic acid molecules attached thereto, onto a substrate, wherein a plurality of first sequencing primers is hybridized to the plurality of single-stranded template nucleic acid molecules; generating a first set of sequencing data on the plurality' of single-stranded template nucleic acid molecules by extending the plurality of first sequencing primers; on the substrate, denaturing extension products of the plurality of first sequencing primers from the plurality of single-stranded template nucleic acid molecules and re-hybridizing a plurality of second sequencing primers to the plurality of single-stranded template nucleic acid molecules; and generating a second set of sequencing data on the plurality of single-stranded template nucleic acid molecules by extending the plurality of second sequencing primers.
[0326]
[0203] In some cases, generating sequencing data comprises performing sequencing-by- synthesis. In some cases, generating sequencing data compnses repeating a plurality of cycles of (i) extending the plurality of sequencing primers using a plurality of nucleotides comprising labeled nucleotides in a flow, and (ii) detecting the presence or absence of a labeled nucleotide incorporated into the extending plurality of sequencing primers to generate the sequencing data.
[0327]
[0204] The plurality of nucleotides may be non-terminated, reversibly terminated, or a combination thereof. The plurality of nucleotides may be nucleotides of a single base type. The plurality of nucleotides may be labeled or a combination of labeled and unlabeled.
[0328]
[0205] The substrate may be rotated pnor to, during, or subsequent to the loading the plurality of beads onto the substrate. The substrate may be rotated prior to, during, or subsequent to denaturing the plurality of double-stranded template nucleic acid molecules. The substrate may be rotated prior to, during, or subsequent to hybridizing the plurality of
[0329] -57-
[0330] SUBSTITUTE SHEET (RULE 26) sequencing pnmers to the plurality of single-stranded template nucleic acid molecules. The substrate may be rotated prior to, during, or subsequent to the extending the plurality of sequencing primers.
[0331]
[0206] In some cases, the denaturing comprises treating the plurality of double-stranded template nucleic acid molecules with sodium hydroxide (NaOH).
[0332]
[0207] In some cases, the plurality of beads are loaded onto a plurality of individually addressable locations on the substrate. A bead of the plurality of beads may comprise at least 1000 double-stranded template nucleic acid molecules of the plurality of double-stranded template nucleic acid molecules. In some case, the at least 1000 double-stranded template nucleic acid molecules are substantially identical copies.
[0333]
[0208] In some cases, combining first and second (or more) sequencing data comprises averaging analog signals (e g., signal detected from incorporated nucleotides) for each flow. In some cases, combining first and second (or more) sequencing data comprises averaging homopolymer base calls (e.g., after analysis of analog signals) for each flow. Combining sequencing data may increase the overall accuracy of a resulting sequence read.
[0334]
[0209] Single molecule data analysis methods
[0335]
[0210] Filtering single molecule sequencing reads
[0336]
[0211] FIG. 9A illustrates an exemplary plurality of sequencing reads that can be received at block 902 of FIG. 9D. In FIG. 9A, the system receives n number of sequencing reads. Each sequencing read is obtained from a flow sequencing method. In some embodiments, the sequencing reads are generated by performing one flow sequencing method on a plurality of sequencing colonies attached to the same surface, where each sequencing read corresponds to a sequencing colony. In some embodiments, the sequencing reads are generated by performing multiple flow sequencing methods. The quality of the plurality' of sequencing reads can be improved in blocks 904-908, as described below.
[0337]
[0212] At block 904, the system filters the sequencing data, by the one or more processors, to remove sequencing reads for which an absence of an incorporated nucleotide was detected at three or more consecutive sequencing flow steps, thereby generating filtered sequencing data. Specifically, the system can examine each sequencing read of the plurality of sequencing reads one by one to determine if each sequencing read needs to be filtered (i.e., excluded).
[0338] For each sequencing read, the system determines if the sequencing read indicates an absence of an incorporated nucleotide at three or more consecutive sequencing flow steps, for example, if the sequencing read indicates three consecutive sequencing flow steps yielding no
[0339] -58-
[0340] SUBSTITUTE SHEET (RULE 26) signals (“000”), four consecutive sequencing flow steps yielding no signals (“0000”), five consecutive sequencing flow steps yielding no signals (“00000”), and so on. If this is so, the sequencing read is excluded from the plurality of sequencing reads. That is, the entire length of each sequencing read is evaluated with respect to a number of consecutive sequencing flow steps. With reference to FIG. 9B, the system can examine each of the sequencing reads 1-n and exclude any sequencing read indicating an absence of an incorporated nucleotide at three or more consecutive sequencing flow steps, thus obtaining sequencing reads 1-m (where m<n).
[0341]
[0213] Typically in sequencing analysis (e g., in cases where a sequencing read is determined from colony sequencing), absence of an incorporated nucleotide at three or more consecutive sequencing flow steps is indicative of weak, incorrect, or noisy signal(s) in the flow sequencing method, and thus denotes an unreliable or damaged sequencing read. See e.g., Inti. Appl. No. W02023004421, which is hereby incorporated by reference in entirety. FIGs. 10A-10B illustrate an exemplary scenario demonstrating why an absence of an incorporated nucleotide at three consecutive sequencing flow steps cannot occur in analysis where each sequencing flow is analyzed independently. In the example flow sequencing method 1000, the flow-cycle order is T-G-C-A. Using an exemplary flow-cycle order T-G-C-A, in flow step n-1, labeled T nucleotides are combined with the hybnd; in flow step n, labeled G nucleotides are combined with the hybrid; in flow step n+1, labeled C nucleotides are combined with the hybrid; in flow step n+2, labeled A nucleotides are combined with the hybrid.
[0342]
[0214] FIG. 10A depicts a hypothetical scenario in which three consecutive sequencing flow steps, n to n+2, all yield a signal of 0 indicating an absence of an incorporated nucleotide. Specifically, in flow step n, labeled G nucleotides are not combined with the hybrid due to the A base; in flow step n+1, labeled C nucleotides are not combined with the hybrid due to the C base; in flow step n+2, labeled A nucleotides are not combined with the hybrid due to the A base.
[0343]
[0215] For the hypothetical scenario in FIG. 10A to occur, there must be a nucleotide incorporation in step n-1 as shown by 1002. This is because if there is no nucleotide incorporation in step n-1, in step n, nucleotides G would be combined with the hybrid having the base before A, rather than the hybrid having the base A.
[0344] 1216] For nucleotide incorporation to occur in step n-1 where labeled T nucleotides are applied, it follows that the base before A in the template polynucleotide must be A (as the T
[0345] -59-
[0346] SUBSTITUTE SHEET (RULE 26) base is complementary to the A base), as shown in FIG. 10B. However, if the base before A in the template polynucleotide is A, the hypothetical flow sequencing steps n to n+2 would not occur. Rather, as shown in FIG. 10B, when labeled T nucleotides are applied in step n-1, two T nucleotides are incorporated into the extending sequencing primer because the template sequence includes two consecutive A bases. Thus, the flow steps n to n+2 depicted in FIG. 10A would not occur.
[0347]
[0217] Thus, FIGs. 10A-10B demonstrate why an absence of an incorporated nucleotide at three consecutive sequencing flow steps would not (and cannot accurately) occur in normal colony sequencing. For high quality sequencing, an absence of an incorporated nucleotide can occur in at most two consecutive sequencing flow steps in cases where the sequence read is determined from colony sequencing (e g., where each flow step is analyzed independently to determine homopolymer lengths in the sequence). This is because in colony sequencing, the signal obtained from each independently addressable location is averaged across the colony of amplified, identical sequencing template molecules. Three or more consecutive flow steps resulting in 0 signal output indicate that something has gone wrong in the sequencing (e g., it is vanishingly unlikely that each and every identical template molecule in the colony would exhibit lack of correct nucleotide incorporation at the same time). In contrast, for single molecule sequencing, the signal obtained from each individually addressable location can be directly correlated to the sequence of the template molecule. Therefore, in single molecule sequencing it is possible and indeed desirable to ‘rescue’ sequence reads with 3 0-signal consecutive flow steps for downstream analysis. Table 1 illustrates the different signals that may be obtained from either a colony or single molecule sequencing method. Detected signals in the ‘normal’ (i.e., complete incorporation) situations and in the case where one or a few templates in a colony fail to incorporate a second correct nucleotide are approximately the same. However, in the case where the template in single molecule sequencing fails to incorporate a second correct ‘T’ nucleotide, there are three consecutive flow steps where 0 signal is detected. In single molecule sequencing, these three consecutive 0-signal flow steps do not indicate a damaged template / problematic sequencing read; instead, this is merely indicative of incomplete incorporation, which can be easily corrected at the next flow of the same nucleotide base type.
[0348]
[0218] Indeed, with single molecule sequencing, it is possible to use sequence reads with up to 6 consecutive 0-signal flow steps. This is because, in single molecule sequencing, if a correct number of nucleotides of a first base type are not incorporated in a respective flow,
[0349] -60-
[0350] SUBSTITUTE SHEET (RULE 26) and if successive flows for the remaining base types are 0-signal, sequencing can restart at the next flow of the first base type. Table 2 illustrates two examples with 5 and 6 consecutive flows 0-signals, and FIG. IOC illustrates an example with 5 consecutive 0-signal flows. Thus, an advantage of single molecule sequencing is that a sequence read can be ‘rescued’ even in cases of incomplete incorporation. As illustrated by the third example in Table 2, for a sequence TTAC, in a case where the second T is not incorporated in flow 1, that second T can be incorporated in flow 5, and the signals from flow 1 and flow 5 can be added.
[0351]
[0219] In Tables 1 and 2, the sequences listed are for extended sequencing primers (e.g., the nucleotides that are being incorporated at each flow). The template sequences will be the corresponding reverse complements.
[0352] -61-
[0353] SUBSTITUTE SHEET (RULE 26)
[0220] Table 1: Examples of incomplete incorporation in colony vs single molecule sequencing.
[0354] -62-
[0355] SUBSTITUTE SHEET (RULE 26)
[0356]
[0222] An absence of an incorporated nucleotide at seven or more consecutive sequencing flow steps is indicative of weak, incorrect, or noisy signal(s) in the flow sequencing method, and thus an unreliable or damaged sequencing read. For example, it may indicate that there was a base in the template sequence that had been missed (e.g., indicative of degradation of the template sequence). Thus, any sequencing read having such an absence is filtered in block 904 such that the sequencing read is not used in downstream tasks (e.g., alignment to a reference genome or portions thereof, for SNP calling, etc.).
[0357]
[0223] At block 906, the system determines, by the one or more processors, for each flow step of each sequencing read, a read quality metric. For example, with reference to FIG. 3, for each flow step (i.e., each column in the flowgram), a read quality metric (also known as regressed residual) is calculated. For example, for flow step 302, a read quality metric RQM1 is calculated; for flow step 306, RQM3 is calculated. Exemplary read quality metrics (e g., regressed residues) are illustrated in FIG. 10D.
[0358]
[0224] In some cases, the read quality metric for each flow step of each sequencing read is calculated based on a second highest homopolymer probability value (p2nd). For example, in flow step 302 in FIG. 3, the second highest probably value is 0.0010. In some cases, the read quality metric (i.e., RQMS) is calculated as:
[0359]
[0225] RQMS= log10(P2^ / e) / 10, (1)
[0360]
[0226] Where c is a scaling factor and p2nd is the second highest probability at the flow step s (e.g., representing the second most likely h-mer). In some cases, c can be set at a value between IxlO'2and IxlO'4.
[0361]
[0227] The read quality metric for a given flow step can be calculated using other techniques. In some embodiments, rather than p2nd, (1- plst) is used in the formula above. In cases in which plst + p2nd = 1, the two formula variations would yield the same read quality metric. In cases in which plst + p2nd + p3rd = 1, the two formula variations would yield different read quality metrics.
[0362]
[0228] A higher read quality metric can be indicative of a weaker signal. For example, a higher p2nd can indicate a lower plst. Because the base count associated with plst is selected a lower plst can indicate a lower confidence in the selected base count. Thus, the read quality metric is used to determine flows with low confidence, which can indicate deterioration in h-mer
[0363] -64-
[0364] SUBSTITUTE SHEET (RULE 26) determination accuracy, in a sequencing read and determine where (e.g., at which flow) to trim the sequencing read, as described below.
[0365]
[0229] It will be understood that the read quality metric could also be calculated, with appropriate modifications to the read quality metric function, using any h-mer probability value each flow step of each sequencing read (e.g., plst, p2nd, p3rd..., pnth). Calculating the read quality metric with, for example, a first highest homopolymer probability value can be performed:
[0366]
[0231] where c would be set as in equation (1).
[0367]
[0232] At block 908, the system trims the terminus of one or more sequencing reads in the sequencing data based on the read quality metrics for a respective sequencing read, thereby generating trimmed sequencing data. With reference to FIG. 9C, some of the sequencing reads 1-m are trimmed, thereby generating trimmed sequencing data.
[0368]
[0233] In some cases, if a flow sequencing step produces a read quality metric below a predetermined threshold, the system can determine that deterioration has occurred in the sequencing read (e.g., deterioration in the quality of the sequencing read). Accordingly, the system can trim the sequencing read at or before the first flow sequencing step that produces a read quality metric below the threshold
[0369]
[0234] In some cases, the system uses an average of multiple read quality values to detect determination in the sequencing read. In some embodiments, the average is a moving average. Exemplary calculation of the moving average is described with reference to FIG. 3. For example, at the third flow step, the system can calculate an average of RQM1, RQM2, and RQM3 (assuming the moving average is calculated using a sliding window of 3 flow steps); at the fourth flow step, the system can calculate an average of RQM2, RQM3, and RQM4. Thus, the moving average is a local quality measure. An example of moving average read quality is illustrated in FIG. 10D, as described below.
[0370]
[0235] In some embodiments, if the moving average exceeds a predetermined threshold, the system determines that deterioration (e g., of read quality) has occurred and trims the sequencing read accordingly. In some embodiments, if a predefined number of moving averages are above the predetermined threshold, the system determines that deterioration has occurred. For example, the flow sequencing step that triggers trimming is the nth sequencing flow step having a moving
[0371] -65-
[0372] SUBSTITUTE SHEET (RULE 26) average above a predetermined threshold, wherein n is a predefined number. That is, in some instances, the sequencing read is trimmed at the flow where the read quality moving average exceeds the predetermined threshold. In some instances, the sequencing read is trimmed at the nth-flow where the read quality moving average has exceeded the predetermined threshold. In some instances, trimming the sequencing read removes the indicated flow and all subsequent flows.
[0373]
[0236] The predetermined threshold can be a fixed value that can be tuned. For example, the predetermined threshold can be set to an average quality of the first 100 flow steps in a flow sequencing method (e.g., based on an average read quality metric for each flow across all sequencing reads). In some embodiments, the predetermined threshold is around 0.3. In some embodiments, the predetermined threshold is about 0, 0.1, 0.2, 0.3, 0.4, or 0.5. In some embodiments, the predetermined threshold is a real number between any of 0, 0. 1, 0.2, 0.3, 0.4, or 0.5. Likewise, the predetermined number n can be a tunable fixed value. For example, n can be set to 3, 5, 10, 15, or 20. In some instances, n is any whole number between 1 and 20.
[0374]
[0237] FIG. 10D illustrates the read quality metrics for an exemplary sequencing read, in accordance with some embodiments. In the depicted example, each cross indicates the read quality metric calculated at the corresponding flow step The dashed line indicates the moving average of read quality metrics. The horizontal line 1012 indicates the predetermined threshold (i.e., 0.2). If a predefined number of consecutive moving averages exceed the predetermined threshold (as shown by the bolded portion of the dashed line above the line 1012), the system determines that deterioration has occurred and therefore trims the sequencing read.
[0375]
[0238] The system then trims at least the portion of the sequencing read comprising the selected sequencing flow step. In some embodiments, a predetermined number of consecutive sequencing flow steps prior to the selected sequencing flow step are also trimmed. In some embodiments, the predetermined number of consecutive sequencing flow steps is a multiple of four (e.g., 8 previous flow steps, 12 previous flow steps, 16 previous flow steps). In other words, the system also trims multiples of 4 flow steps before the selected flow step, in addition to trimming the selected flow step.
[0376]
[0239] Thus, the trimming operation in block 908 can be dependent on at least three parameters: window length, threshold, and lag. Window length refers to the size of the sliding window in which the moving average value is calculated. Threshold refers to the predetermined threshold of
[0377] -66-
[0378] SUBSTITUTE SHEET (RULE 26) the moving average value above which the system determines that deterioration has occurred. Lag refers to the predetermined number of consecutive sequencing flow steps prior to the selected sequencing flow step that are also trimmed. In some embodiments, some or all of these parameters can be determined based on user input. In some embodiments, some or all of these parameters can be determined automatically.
[0379]
[0240] In some embodiments, the system does not calculate a read quality metric for every flow step, but rather at regular intervals (e.g., every 4 flow steps, every 8 flow steps, etc.). In some embodiments, these regular intervals will be a multiple 4. In some embodiments, the system does not calculate read quality metrics for certain flow steps in a flow sequencing method (e.g., the first 100 flow sequencing steps), for example because deterioration typically occurs during later flow steps.
[0380]
[0241] Resupply of polymerase
[0381]
[0242] A particular issue with single molecule sequencing is that polymerases have a natural on / off rate. That is, stochastically, polymerases will fall off extending double-stranded DNA, which stalls sequencing. In order for sequencing to be expeditiously restarted, additional polymerases must be provided. Additional polymerase may be provided in each sequencing flow. Alternatively, additional polymerase may be provided on a sequencing substrate. In some cases, additional polymerase may be coupled to individually addressable locations on a substrate (e.g., reversibly coupled). This coupling may be performed prior to or after loading of template for sequencing. In some cases, the additional polymerases may be provided on beads that lack templates for sequencing (e.g., blank beads). The beads may comprise double-stranded oligos, and the polymerases may be coupled thereto. In some cases, the additional polymerases may be provided on sequencing beads (e.g., beads that have templates for sequencing). In any case where polymerases are coupled to a substrate, in one or more sequencing flows, a subset of additional polymerases may be cleaved and released into solution. This increases the effective concentration of free polymerase. As the concentration of polymerase is increased, this increases the likelihood that templates lacking sequencing polymerases will anneal to an additional polymerase and sequencing will continue.
[0382]
[0243] Sequencing for Increased Homopolymer Detection Accuracy
[0383]
[0244] Accurate homopolymer detection can be an issue in sequencing-by-synthesis, especially for non-terminated nucleotide-based sequencing methods. In sequencing methods using
[0384] -67-
[0385] SUBSTITUTE SHEET (RULE 26) terminated nucleotides, in each sequencing step only a single nucleotide can be incorporated into a growing strand, which is beneficial for confidence of a base call. However, prior to incorporating a next nucleotide, the terminating moiety (and / or a labeling moiety) must be removed (typically cleaved), which can leave a chemical ‘scar’ that can, in some cases, inhibit polymerase function and hence future incorporation. In particular, the accumulation of scars is detrimental to sequencing quality and sequencing read length. In sequencing methods incorporating non-terminated nucleotides, multiple nucleotides can be incorporated into a growing strand in a single sequencing step (e.g., either bases of a single base type or bases of multiple base types), which increases sequencing speed. In cases where nucleotides of a single base type are incorporated, only the number of incorporated bases needs to be identified (e.g., the base type is known a priori). However, labeling moieties still need to be removed from nonterminated nucleotides after incorporation and detection, and this removal can also leave chemical scars that similarly negatively impact polymerase activity and future nucleotide incorporation. In colony sequencing, the scarring issue in non-terminated sequencing-by- synthesis can be addressed by, for example, fractional labeling. However, in single molecule sequencing, fractional labeling as used in colony sequencing is not typically feasible. Methods, systems, compositions, and kits for improved homopolymer accuracy are required.
[0386]
[0245] Mixed-reversibly terminated flow sequencing
[0387]
[0246] In some cases, for a bright sequencing flow, the growing strand (e g., extending primer) may be contacted with only non-terminated nucleotides — here, if the template has a homopolymer portion, the growing strand may incorporate multiple non-terminated nucleotides in a single step, and thus signals detected from incorporated labeled nucleotides may have to be further resolved to determine the length of the homopolymer. For example, relatively stronger signals may correspond to longer homopolymer length as they are indicative that more labeled nucleotides have been incorporated, and relatively weaker signals may correspond to lower homopolymer length as they are indicative that fewer labeled nucleotides have been incorporated. For example, detected signals may be algorithmically processed to distinguish a 2- mer from a 3-mer or a 4-mer from a 7-mer. However, homopolymer length determination accuracy from these signals may decrease as homopolymer lengths become longer and / or goes above a certain resolution threshold (e.g., 5-mer, 6-mer, 7-mer, 8-mer, 9-mer, 10-mer, 11-mer, 12-mer, 13-mer, 14-mer, 15-mer, 16-mer, 17-mer, 18-mer, 19-mer, 20-mer, 21-mer, etc.), such
[0388] -68-
[0389] SUBSTITUTE SHEET (RULE 26) as due to increasing quenching effects of dye moieties on incorporated labeled nucleotides, optical resolution limitations for signal collection, and / or computing limitations. Alternatively or in addition, nucleotide incorporation may be impeded by the presence of scars in the growing strand (e.g., as a result of cleaving labels from incorporated nucleotides). This can inhibit sequencing, e.g., by increasing phasing, by pausing or stopping incorporation. The present systems, methods, compositions, and kits address at least the abovementioned limitations by improving the accuracy of sequencing reads by reading a homopolymer section of a template in multiple shorter segments and by reducing the impact of scarring. The methods described herein are applicable to either sequencing single molecules or sequencing colonies of amplified template molecules.
[0390]
[0247] FIG. 11A illustrates an example of a mixed-reversibly terminated sequencing scheme. A template is hybridized to a growing strand which is ready to extend through a 6-mer poly A homopolymer portion in the template. In step (I), the first bright extension step, the growing strand is contacted with a nucleotide mixture comprising both labeled, non-terminated bases and reversibly terminated bases of T. The growing strand incorporates only two labeled, nonterminated T bases before incorporation is blocked by incorporation of a terminated T base, resulting in extending through 3 of 6 available T incorporation positions. In step (II), a first imaging is performed to collect first signals indicative of the first homopolymer segment, and then any labels and blocking moieties removed via cleaving. In step (III), the second bright extension step, step (I) is repeated where the growing strand is contacted with a nucleotide mixture comprising both labeled, non-terminated bases and reversibly terminated bases of T. This time, the growing strand incorporates only one labeled, non-terminated T base before incorporation is blocked by incorporation of a terminated T base, resulting in extending through 2 of 3 of the remaining available T incorporation positions. In step (IV), a second imaging is performed to collect second signals indicative of the second homopolymer segment, and then any labels and blocking moieties removed via cleaving. In step (V), in a dark extension step, the growing strand is contacted with unlabeled, non-terminated T bases to extend through all (in this case 1) of the remaining T incorporation positions. The data collected and / or determined from the two imaging actions (in steps (II) and (IV) respectively) may be processed (e.g., added) to determine a total homopolymer length of the homopolymer portion just sequenced. In this
[0391] -69-
[0392] SUBSTITUTE SHEET (RULE 26) illustration, a determination of at least a 5-mer homopolymer length is made from the data collected. In step (VI), steps (I)-(V) may be repeated with a next, different canonical base type.
[0393]
[0248] It will be appreciated that while this example includes only two bright extension steps ((I)-(II) and (III)-(IV)), any number of bright extension steps may be performed, which can increase the accuracy of the homopolymer length determination.
[0394]
[0249] In some cases, for single molecule sequencing and colony-based sequencing, all nonterminated bases in a bright extension step may be labeled nucleotides. The terminated bases in a bright extension step may be labeled, unlabeled, or a mixture of both. In some cases, in dark extension steps (e.g., step (V)), the growing primer strand is contacted with labeled, nonterminated bases, unlabeled non-terminated bases, or a mixture of labeled and unlabeled nonterminated bases. This may be more efficient in terms of reagent storage space (e.g., obviating the need for separate reagent storage wells for different mixtures of unterminated bases for bright and dark extension steps). Dark extension steps do not include imaging and may include either labeled or unlabeled nucleotides.
[0395]
[0250] In some cases, for colony -based sequencing, the non-terminated bases in a bright extension step may be a mixture of labeled and unlabeled nucleotides. The terminated bases in a bright extension step may be labeled, unlabeled, or a mixture of both. The mixture of labeled and unlabeled nucleotides in the non-terminated bases in the nucleotide reagent may be of any fraction of labeled nucleotides, such as at least or at most about 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 99%. The mixture of labeled and unlabeled nucleotides in the terminated bases may be of any fraction of labeled nucleotides, such as at least or at most about 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 99%. The mixture of labeled and unlabeled nucleotides in the nucleotide reagent may be of any fraction of labeled nucleotides, such as at least or at most about 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 99%. Different fractions of labeled and unlabeled nucleotides, labeled and unlabeled nucleotides in the terminated bases, and / or labeled and unlabeled nucleotides in the non-terminated bases may be different for different base types (e.g., based on expected hmer lengths and / or quenching).
[0396] -70-
[0397] SUBSTITUTE SHEET (RULE 26)
[0251] In some cases, for colony -based and single molecule sequencing, the nucleotide reagent can comprise a mixture of terminated and non-terminated nucleotides of any fraction of terminated to non-terminated nucleotides, such as or at most about 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 99%. The fraction of terminated nucleotides will influence the average number of bases incorporated in each bright extension step. For example, if the fraction of terminated nucleotides is about 10%, then the average number of incorporated bases in each extending sequencing primer may be about 10 (e.g., 9 incorporated unterminated nucleotides and 1 incorporated terminated nucleotide). Similarly, if the fraction of terminated nucleotides is about 25%, then the average number of incorporated bases may be about 4 (e g., 3 unterminated nucleotides and 1 terminated nucleotide). At most, one terminated base is expected to be incorporated in each bright extension step.
[0398]
[0252] Any number of consecutive bright extension steps of a same canonical base type may be performed, such as 2, 3, 4, 5, 6, 7, 8 or more consecutive bright extension steps of a same canonical base type. In some cases, the respective number of consecutive bright steps may differ for different nucleotide base types (e.g., 2 consecutive bright steps for A and 3 consecutive bright steps for T). In some cases, a number of consecutive bright steps may be predetermined. In some cases, a number of bright steps may be determined based on relative signal brightness in images of a same nucleotide base type (e.g., Image 1 vs Image 2 in FIG. 11 A).
[0399]
[0253] The sequencing method may comprise repeating the subjecting of a growing strand to a template to at least two consecutive bright extension steps followed by a dark extension step of the same canonical base type (e.g., A, G, C, T, U) with different bases for any number of times. For example, at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16 ,17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 150, 200, 250, 300, 350, 400, 450, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 10,000 or more times.
[0400]
[0254] A sequencing method may comprise subjecting a growing strand hybridized to a template to at least two consecutive bright extension steps followed by a dark extension step of the same canonical base type (e.g., A, G, C, T, U). For sequencing methods described herein, T and U are considered the same canonical base type. A bright extension step may comprise contacting the growing strand with a nucleotide mixture of both (1) labeled, non-terminated bases and (2) reversibly terminated bases of a same canonical base type. The reversibly terminated bases may
[0401] -71-
[0402] SUBSTITUTE SHEET (RULE 26) be labeled or unlabeled, or a mixture of both. In some cases, the last bright extension step may comprise only non-terminated bases and omit the reversibly terminated bases. Beneficially, because of the fraction of reversibly terminated bases in the mixture, when at a long homopolymer stretch in the template, the growing strand is likely to incorporate a reversibly terminated base and block incorporation of the next base before fully extending through a long homopolymer stretch. This allows generating sequencing data by collecting signals (e.g., via imaging) from shorter homopolymer segment intervals, which results in a more accurate homopolymer base call for each segment. Sequencing data generated after each of the bright extension step(s) may be processed (e.g., signals added, images added, homopolymer lengths added, etc.) to determine length information of the homopolymer stretch. For example, a total length of the homopolymer may be determined with high accuracy. In another example, a minimum length of the homopolymer may be determined with high accuracy. Any labels may be removed from the growing strand between different bright extension steps, such as via cleavage, to allow for interval imaging and more efficient incorporation of the next succeeding base. Any blocking moieties may be removed from the growing strand between different extension steps (bright or dark), such as via cleavage, to allow incorporation of the next succeeding base in the next extension step. The bright extension steps may be followed by a dark extension step of the same canonical base type to (1) extend through any remaining portions of a homopolymer stretch that was not covered by the bright extension steps to prepare for interrogation with the next base type and / or (2) catch up any strands (e.g., with a colony) that were unable to incorporate a base(s), such as due to reaction kinetics.
[0403]
[0255] Mixed-color flow sequencing and / or FRET
[0404]
[0256] FIG. 11B illustrates an example of a mixed-color non-terminated sequencing scheme. A template is hybridized to a growing strand which is ready to extend through a 6-mer poly A homopolymer portion in the template. In step (I), the bright extension step, the growing strand is contacted with a nucleotide mixture comprising a first plurality of bases labeled with a first label and a second plurality of bases labeled with a second label, where all of the bases are T. The growing strand incorporates a mixture of Ts with the first and second labels (in this case only 5 Ts are incorporated; in some cases, 6 Ts or 4 Ts may be incorporated). In step (II), a first imaging is performed to collect first signals indicative of the first label. In step (III), a second imaging is performed to collect second signals indicative of the second label, and then any labels
[0405] -72-
[0406] SUBSTITUTE SHEET (RULE 26) are removed via cleaving. In step (IV), a dark extension is performed where the growing strand is contacted with unlabeled, non-terminated T bases to extend through (in this case 1) the remaining T incorporation positions. The data collected and / or determined from the two imaging actions (in steps (II) and (III) respectively) may be processed (e.g., added) to determine a total homopolymer length of the homopolymer portion just sequenced.
[0407]
[0257] In some cases, first or second signals may further be indicative of the second or first label, respectively. For example, in some cases, the first label may be a FRET donor, and the second label may be a FRET acceptor (or the reverse). In this illustration, a determination of at least a 5-mer homopolymer length is made from the data collected. In step (V), steps (I)-(IV) may be repeated with a next, different canonical base type. It will be appreciated that while this example includes only two bright extension steps ((I)-(II) and (III)-(IV)), any number of bright extension steps may be performed, which can increase the accuracy of the homopolymer length determination.
[0408]
[0258] Beneficially, the use of at least two label types may improve homopolymer length determination. For instance, there may be less quenching between labels on incorporated nucleotides if there is a mixture of label types. In some cases, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more label types may be used.
[0409]
[0259] In some cases, for single molecule sequencing and colony-based sequencing, all nonterminated bases in a bright extension step may be labeled nucleotides. The non-terminated bases in a bright extension step may be labeled, unlabeled, or a mixture of both. In some cases, in dark extension steps (e.g., step (IV)), the growing primer strand is contacted with labeled, nonterminated bases, unlabeled non-terminated bases, or a mixture of labeled and unlabeled nonterminated bases. This may be more efficient in terms of reagent storage space (e.g., obviating the need for separate reagent storage wells for different mixtures of unterminated bases for bright and dark extension steps). Dark extension steps may not include imaging. As described with respect to FIG. 11 A, different proportions of labeled / unlabeled nucleotides may be used. In some cases, different proportions of labeled / unlabeled nucleotides for different canonical base types may be used.
[0410]
[0260] Here a method of sequencing is provided, comprising (a) contacting a growing strand hybridized to a template with a first reagent mixture comprising bases labeled with a first label type and bases labeled with a second label type, wherein the bases are of a first same canonical
[0411] -73-
[0412] SUBSTITUTE SHEET (RULE 26) base type; (b) detecting a first signal indicative of incorporation of at least a subset of the bases labeled with the first label type in the growing strand, or lack thereof, to generate first sequencing data; (c) detecting a second signal indicative of incorporation of at least a subset of the bases labeled with the second label type in the growing strand, or lack thereof, to generate second sequencing data; and (d) processing the first sequencing data and the second sequencing data to determine length information of a homopolymer sequence in the template.
[0413]
[0261] In some cases, the length information of the homopolymer sequence in the template comprises a minimum length of the homopolymer sequence. Alternatively, or in addition, the length information of the homopolymer sequence in the template comprises a total length of the homopolymer sequence.
[0414]
[0262] In some case, the method may further comprise (e) contacting the growing strand with a second reagent mixture comprising unlabeled bases of the first canonical base type. The method may further comprise repeating (a)-(e) with a second canonical base type, a third canonical base type, and / or a fourth canonical base type. These steps may be repeated any number of time suitable for determining the sequence of a nucleic acid template molecule. For example, these steps may be repeated 1, 2, 3, 4, , 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 150, 200, 300, 400, 500, 600, 700, 800, 900, 1000, or more times.
[0415]
[0263] In some cases, the first signal and the second signal may be localized to a single molecule of the template. Alternatively, the first signal and the second signal may be localized to a colony of molecules comprising the template.
[0416]
[0264] In some cases, nucleotides are unterminated. In some cases, a mixture of terminated and unterminated nucleotides may be used.
[0417]
[0265] In some cases, the template may be immobilized to a substrate surface. Alternatively or in addition, the template may be coupled to a bead that is immobilized to the substrate surface. Alternatively or in addition, the template may be coupled to a DNA nanoparticle (e.g., a DNA nanoball or DNA origami) that is immobilized to the substrate surface. Alternatively or in addition, the template may be coupled to a dendrimer that is immobilized to the substrate surface. In some cases, the substrate surface comprises at least 1,000,000 individually addressable locations and the template is immobilized to an individually addressable location in the at least 1,000,000 individually addressable locations.
[0418] -74-
[0419] SUBSTITUTE SHEET (RULE 26)
[0266] In some cases, multiple labels with similar excitation and emission spectra may be used in combination. In some cases, the use of multiple different labels may reduce quenching between labels on adjacent incorporated nucleotides. That is, in some cases, a method of sequencing is provided, comprising (a) contacting a growing strand hybridized to a template with a reagent mixture comprising bases labeled with a first label type and bases labeled with a second label type, wherein the bases are of a first same canonical base type and wherein the first label type and the second label type may be detected during a same imaging step; (b) detecting a signal indicative of incorporation of at least a subset of the bases labeled with the first label type into the growing strand, at least a subset of the bases labeled with the second label types into the growing strand, or lack thereof, to generate sequencing data; and (c) determining length information of a homopolymer sequence in the template from the sequencing data. The reagent mixture of bases labeled with the first label type and bases labeled with the second label type may be of any fraction of nucleotides with the first label type, such as at least or at most about 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 99%.
[0420]
[0267] Accumulating signal across multiple flows
[0421]
[0268] FIG. 11C illustrates an example of a single color non-terminated sequencing scheme. A template is hybridized to a growing strand which is ready to extend through a 6-mer poly A homopolymer portion in the template. In step (I), a first bright extension step is performed where the growing strand is contacted with a nucleotide mixture comprising a first plurality of bases labeled with a first label, where all of the bases are T. The growing strand incorporates a number of Ts less than the respective homopolymer portion in the template (in this case only 4 Ts are incorporated; in some cases, 0, 1, 2, 3, 4, or 5 Ts may be incorporated). In step (II), a first imaging is performed to collect first signals indicative of the first label, and then any labels are removed via cleaving. In step (III), a second bright extension step is performed where the growing strand is contacted with a nucleotide mixture comprising a second plurality of bases labeled with a first label, where all of the bases are T. The growing strand incorporates a number of Ts (in this case 2 Ts are incorporated). In step (IV), a second imaging is performed to collect second signals indicative of the first label, and then any labels are removed via cleaving. In step (V), a dark extension is performed where the growing strand is contacted with unlabeled, nonterminated T bases to extend through (in this case 0) remaining T incorporation positions. The
[0422] -75-
[0423] SUBSTITUTE SHEET (RULE 26) data collected and / or determined from the two imaging actions (in steps (II) and (IV) respectively) may be processed (e.g., added) to determine a total homopolymer length of the homopolymer portion just sequenced. In some cases, the analog signal (e.g., the fluorescence signal detected from labels on incorporated nucleotides) may be added for consecutive imaging steps for a same canonical base type. In some cases, base calls (e g., based on the respective analog signal) for each incorporation step may be added together for consecutive imaging steps for a same canonical base type. In some cases, no dark extension is performed for one or more canonical base types.
[0424]
[0269] Beneficially, the combination of information from at least two images may improve homopolymer length determination. For instance, there may be less quenching between labels on incorporated nucleotides and / or more linearity of signal strength from labels on incorporated nucleotides if signal may be accumulated across multiple flows and images. In some cases, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more flows and imaging steps may be used for one or more canonical base types. In some cases, in a sequencing run, a same number of flows and imaging steps may be used for each canonical base type (e g., for each base type, there may be two bright flows each followed by an imaging step). In some cases, in a sequencing run, a different number of flows and imaging steps may be used for one or more canonical base types (e.g., for Ts two bright flows each followed by an imaging step may be used, and for Gs three bright flows each followed by an imaging step may be used).
[0425]
[0270] In some cases, for single molecule sequencing and colony-based sequencing, all nonterminated bases in a bright extension step may be labeled nucleotides. The non-terminated bases in a bright extension step may be labeled, unlabeled, or a mixture of both. In some cases, in dark extension steps (e.g., step (IV)), the growing primer strand is contacted with labeled, nonterminated bases, unlabeled non-terminated bases, or a mixture of labeled and unlabeled nonterminated bases. This may be more efficient in terms of reagent storage space (e.g., obviating the need for separate reagent storage wells for different mixtures of unterminated bases for bright and dark extension steps). Dark extension steps do not include imaging.
[0426]
[0271] Here a method of sequencing is provided, comprising (a) contacting a growing strand hybridized to a template with a first reagent mixture comprising bases labeled with a first label type, wherein the bases are of a first same canonical base type; (b) detecting a first signal indicative of incorporation of at least a subset of the bases labeled with the first label type in the
[0427] -76-
[0428] SUBSTITUTE SHEET (RULE 26) growing strand, or lack thereof, to generate first sequencing data; (c) contacting the growing strand with a second reagent mixture comprising bases labeled with the first label type, wherein the bases are of the first canonical base type; (d) detecting a second signal indicative of incorporation of at least a subset of the bases labeled with the second label type in the growing strand, or lack thereof, to generate second sequencing data; and (e) processing the first sequencing data and the second sequencing data to determine length information of a homopolymer sequence in the template.
[0429]
[0272] In some cases, the length information of the homopolymer sequence in the template comprises a minimum length of the homopolymer sequence. Alternatively, or in addition, the length information of the homopolymer sequence in the template comprises a total length of the homopolymer sequence.
[0430]
[0273] In some case, the method may further comprise (e) contacting the growing strand with a third reagent mixture comprising unlabeled bases of the first canonical base type. The method may further comprise repeating (a)-(e) with a second canonical base type, a third canonical base type, and / or a fourth canonical base type. These steps may be repeated any number of time suitable for determining the sequence of a nucleic acid template molecule. For example, these steps may be repeated 1, 2, 3, 4, , 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 150, 200, 300, 400, 500, 600, 700, 800, 900, 1000, or more times.
[0431]
[0274] In some cases, the first signal and the second signal may be localized to a single molecule of the template. Alternatively, the first signal and the second signal may be localized to a colony of molecules comprising the template.
[0432]
[0275] In some cases, nucleotides are unterminated. In some cases, a mixture of terminated and unterminated nucleotides may be used.
[0433]
[0276] In some cases, the template may be immobilized to a substrate surface. Alternatively or in addition, the template may be coupled to a bead that is immobilized to the substrate surface.
[0434] Alternatively or in addition, the template may be coupled to a DNA nanoparticle (e.g., a DNA nanoball or DNA origami) that is immobilized to the substrate surface. Alternatively or in addition, the template may be coupled to a dendrimer that is immobilized to the substrate surface. In some cases, the substrate surface comprises at least 1,000,000 individually addressable locations and the template is immobilized to an individually addressable location in the at least 1,000,000 individually addressable locations.
[0435] -77-
[0436] SUBSTITUTE SHEET (RULE 26)
[0277] In some cases, the method may further comprise using a modified polymerase during the contacting of the growing strand with a reagent mixture In some cases the modified polymerase may be modified to incorporate at most 2, 3, 4, 5, 6 nucleotides in a homopolymer. That is, a modified polymerase with low processivity may be selected such that a low, average number of consecutive bases of a same canonical base type may be incorporated in a single sequencing flow.
[0437]
[0278] Single molecule loading methods
[0438]
[0279] Bead-based loading
[0439]
[0280] Provided herein are devices, systems, methods, compositions, and kits that use DNA nanoballs, dendrimers, and / or beads to facilitate loading of nucleic acids for sequencing on a substrate. For example, DNA nanoballs, dendrimers, and / or beads may be used as supports to immobilize nucleic acids, which supports are immobilized to a substrate. In another example, DNA nanoballs, dendrimers, and / or beads may be used as spacing and / or self-assembling objects that are used to space out and / or self-assemble nucleic acids on the substrate. Such devices, systems, methods, compositions, and kits can be applied alternatively or in addition to the sequencing workflow 100 of FIG. 1. Such devices, systems, methods, compositions, and kits can be used in conjunction with the sample processing systems and methods, or components thereof (e.g., substrates, detectors, reagent dispensing, continuous scanning, etc.) described herein.
[0440]
[0281] Nucleic acids may be loaded onto a substrate using beads, DNA nanostructures (e.g., origami), DNA nanoballs, dendrimers, or a combination thereof. FIG. 12A-12D illustrate different workflows for loading nucleic acids using beads as spacers. As shown in FIG. 12A, pre-enriched template-bead assemblies (or positive beads) may be loaded onto a substrate such that a template of a given template-bead assembly binds to the substrate. Pre-enrichment may refer to the generation of template-bead assemblies via contacting templates and beads together and then isolation of template-bead assemblies from other templates and beads that did not attach to each other. Pre-enriched template-bead assemblies may refer to the isolated template-bead assembly population. The substrate may be patterned or unpatterned. For example, the substrate may be patterned with binders that are configured to bind to templates of template-bead assemblies. In some cases, the substrate may be patterned or coated with DNA nanostructures as described elsewhere herein, where the DNA nanostructures are configured to bind to beads of template-bead assemblies. In another example, the substrate may be unpatterned such that there
[0441] -78-
[0442] SUBSTITUTE SHEET (RULE 26) is a substantially uniform coating of a surface chemistry on the substrate. The surface chemistry may comprise binders that are configured to bind to templates. For example, the surface chemistry may comprise DBCO moieties, and the template may comprise azide moieties, respectively, which can couple together (e.g., template to surface) via click chemistry. Any one or more coupling mechanisms described elsewhere herein may be used for the tempi ate- substrate binding, such as any click chemistry pair, complementary oligonucleotides that hybridize, magnetic particles that are forced together by magnetic fields, electric particles that are forced together by electric fields, specific binding, non-specific binding, electrostatic interactions, crosslinking, etc. The substrate and / or template may comprise any binder described elsewhere herein. For example, the template may comprise a first coupler of a coupling pair and the substrate may comprise a second coupler of the coupling pair. In some cases, a single template may comprise a single moiety (e.g., azide moiety, DBCO moiety, thiol moiety, oligonucleotide sequence, a crosslinking base, etc.) capable of binding to the substrate. In another example, the template may have a negative charge and a surface chemistry of the substrate may have a positive charge which is electrostatically attracted to the negative charge. In some cases, a single template may comprise a plurality of moieties (e.g., at the same location or different locations on template strand(s)) capable of binding to the substrate. A single moiety, any subset of moieties, or all of the plurality of moieties may be used to bind to the substrate. In some cases, a single binder on the substrate may bind to a single template. In some cases, multiple binders on the substrate may bind to a single template. In some cases, a template may be bound at one end to the bead and at the other end to the substrate. In other cases, a template may be bound to the bead and / or the substrate at a location that is not at the 5 ’ or 3 ’ end of a strand, for example at a base that is adj acent to a 5 ’ or 3’ end of a strand. Beneficially, the beads bound to the templates in the template-bead assemblies may function as spacers and / or self-assembling objects on the substrate which prevents one template from binding too close to another template on the substrate. For example, after loading, a template-to-template pitch (center-to-center distance) may be at least or about a bead-to-bead pitch (center-to-center distance) when the template-bead assemblies are loaded as a result of the spacing / self-assembling between the beads. In some cases, after loading, an average template-to- template pitch may be at least an average bead-to-bead pitch and / or at least an average bead diameter. In some cases, upon depositing the template-bead assemblies on the substrate, the template-bead assemblies may be permitted to self-assemble or space out on the substrate before
[0443] -79-
[0444] SUBSTITUTE SHEET (RULE 26) a binding reaction between the template and the substrate is activated via one or more stimuli (e.g., chemical reagent, catalyst, light (e.g., UV light), heat, etc.) to bind the templates to the substrate. In other cases, the template-bead assemblies may be deposited on the substrate under conditions sufficient to permit binding of the template to the substrate upon contact. After spacing out and / or self-assembling, the beads may be cleaved or otherwise removed from the templates and washed away. The cleaving may occur before or after the templates bind to the substrate. The beads may be washed away after the templates bind to the substrate.
[0445]
[0282] Alternatively, as shown in FIG. 12B, a mixture of non-pre-enriched template-bead assemblies (or positive beads) and negative beads (not bound to any templates) may be loaded onto a substrate such that a template of a given template-bead assembly binds to the substrate. In this case, the negative beads in the mixture are unable to bind to the substrate as they lack a template. Once the beads in the template-bead assemblies are cleaved and washed, the negative beads will also get washed away. In some cases, DNA nanostructures (e g., DNA origami, DNA nanoballs, etc.) and / or dendrimers may be used instead of or in addition to negative beads. In some cases, after loading, an average tempi ate-to-templ ate pitch may be at least an average bead- to-bead pitch and / or at least an average bead diameter. In some cases, upon depositing the template-bead assemblies on the substrate, the template-bead assemblies may be permitted to self-assemble or space out on the substrate before a binding reaction between the template and the substrate is activated via one or more stimuli (e.g., chemical reagent, catalyst, light (e.g., UV light), heat, etc.) to bind the templates to the substrate. In other cases, the template-bead assemblies may be deposited on the substrate under conditions sufficient to permit binding of the template to the substrate upon contact. After spacing out and / or self-assembling, the beads may be cleaved or otherwise removed from the templates and washed away. The cleaving may occur before or after the templates bind to the substrate. The beads may be washed away after the templates bind to the substrate.
[0446]
[0283] Alternatively, as shown in FIG. 12C, template-bead assemblies may be loaded onto a substrate such that a bead of a given template-bead assembly binds to the substrate. The template-bead assemblies may be pre-enriched such that only positive beads (bound to a template) are deposited on the substrate. The template-bead assemblies may be non-pre-enriched such that a mixture of positive and negative beads are deposited on the substrate. The substrate may be patterned or unpatterned. For example, the substrate may be patterned with binders that
[0447] -80-
[0448] SUBSTITUTE SHEET (RULE 26) are configured to bind to beads of template-bead assemblies. In some cases, the substrate may be patterned or coated with DNA nanostructures as described elsewhere herein, where the DNA nanostructures are configured to bind to beads of template-bead assemblies. In another example, the substrate may be unpatterned such that there is a substantially uniform coating of a surface chemistry on the substrate. The surface chemistry may comprise binders that are configured to bind to beads. Any one or more coupling mechanisms described elsewhere herein may be used for the bead-substrate binding, such as any click chemistry pair, complementary oligonucleotides that hybridize, magnetic particles that are forced by magnetic fields, electric particles that are forced by electric fields, specific binding, non-specific binding, electrostatic interactions, crosslinking, etc. The substrate and / or bead may comprise any binder described elsewhere herein. In some cases, the bead may comprise a plurality of primers that are not bound or extended into a template, which plurality of primers may be used to bind to the substrate. In some cases, the bead may comprise a first coupler of a coupling pair and the substrate may comprise a second coupler of the coupling pair. In some cases, a single bead may comprise a single moiety (e.g., azide moiety, DBCO moiety, thiol moiety, oligonucleotide sequence, a cross-linking base, etc.) capable of binding to the substrate. In another example, the bead or components attached thereto may have a negative charge and a surface chemistry of the substrate may have a positive charge which is electrostatically attracted to the negative charge. In some cases, a single bead may comprise a plurality of moieties (e.g., at the same location or different locations on template strand(s)) capable of binding to the substrate. A single moiety, any subset of moieties, or all of the plurality of moieties may be used to bind to the substrate. In some cases, a single binder on the substrate may bind to a single bead. In some cases, multiple binders on the substrate may bind to a single bead. In some cases, a template may be bound at one end to the bead. In other cases, a template may be bound to the bead at a location that is not at the 5’ or 3’ end of a strand, for example at a base that is adjacent to a 5’ or 3’ end of a strand. In this workflow, a template may not be directly bound to the substrate. Beneficially, the beads bound to the templates in the template-bead assemblies may function as spacers and / or self-assembling objects on the substrate which prevents one template from being immobilized too close to another template on the substrate. For example, after loading, a tempi ate-to-templ ate pitch (center-to-center distance) may be at least a bead-to-bead pitch (center-to-center distance) when the template-bead assemblies are loaded as a result of the spacing / self-assembling between the beads. Where non-
[0449] -81-
[0450] SUBSTITUTE SHEET (RULE 26) pre-enriched mixture of positive and negative beads are deposited, the average template-to- template pitch may be greater due to the presence of non-template-bound negative beads also loaded on the substrate. In some cases, after loading, an average template-to-template pitch may be at least an average bead-to-bead pitch and / or at least an average bead diameter. In some cases, upon depositing the template-bead assemblies on the substrate, the template-bead assemblies may be permitted to self-assemble or space out on the substrate before a binding reaction between the bead and the substrate is activated via one or more stimuli (e.g., chemical reagent, catalyst, light (e.g., UV light), heat, etc.) to bind the beads to the substrate. In other cases, the template-bead assemblies may be deposited on the substrate under conditions sufficient to permit binding of the bead to the substrate upon contact. After the beads are immobilized, in some cases, the beads may be subjected to shrinking conditions. In some cases, the bead shrinking may be permanent or substantially persistent with respect to the length of time required to generate sequence read(s) (e g., the beads may be subjected to conditions causing irreversible shrinking or fixed into a shrunk stated, e.g., by cross-linking of portions of a bead to itself). In some cases, the bead shrinking may be temporary (e.g., the beads may return to substantially their original size during some part of the sequencing). Templates attached to the beads may be further spaced apart from neighboring templates via the shrinking as on average each template is pulled closer to the center of the bead in each template-bead assembly.
[0451]
[0284] Alternatively, as shown in FIG. 12D, template-bead (e.g., here bead-template-bead) assemblies may be loaded onto a substrate such that a bead of a given template-bead assembly binds to the substrate. The template-bead assemblies may comprise a template coupled to a first bead at or near a first end and the template coupled to a second bead at or near the second end. The first bead may be a sequencing bead and may bind to the substrate. The second bead may be a spacing bead and may not bind to the substrate. In some cases, the first and second beads are the same, and it may be a matter of chance whether the first or second bead binds to the substrate. In some cases, the first and second beads may be different such that only first beads or only second bead may bind to the substrate. The template-bead assemblies may be pre-enriched such that only positive assemblies (first bead bound to a template which is in turn bound to a second bead) are deposited on the substrate. The template-bead assemblies may be non-pre- enriched such that a mixture of positive and negative assemblies (e.g., assemblies comprising negative beads and / or assemblies comprising first beads bound to templates that are not bound to
[0452] -82-
[0453] SUBSTITUTE SHEET (RULE 26) second beads and / or assemblies comprising second beads bound to templates that are not bound to first beads) are deposited on the substrate. The substrate may be patterned or unpattemed. For example, the substrate may be patterned with binders that are configured to bind to beads of template-bead assemblies. In some cases, the substrate may be patterned or coated with DNA nanostructures as described elsewhere herein, where the DNA nanostructures are configured to bind to beads of template-bead assemblies. In another example, the substrate may be unpattemed such that there is a substantially uniform coating of a surface chemistry on the substrate. The surface chemistry may comprise binders that are configured to bind to beads. Any one or more coupling mechanisms described elsewhere herein may be used for the bead-substrate binding, such as any click chemistry pair, complementary oligonucleotides that hybridize, magnetic particles that are forced by magnetic fields, electric particles that are forced by electric fields, specific binding, non-specific binding, electrostatic interactions, cross-linking, etc. The substrate and / or bead may comprise any binder described elsewhere herein. In some cases, the bead may comprise a plurality of primers that are not bound or extended into a template, which plurality of primers may be used to bind to the substrate. In some cases, the bead may comprise a first coupler of a coupling pair and the substrate may comprise a second coupler of the coupling pair. In some cases, a single bead may comprise a single moiety (e.g., azide moiety, DBCO moiety, thiol moiety, oligonucleotide sequence, a cross-linking base, etc.) capable of binding to the substrate. In another example, the bead or components attached thereto may have a negative charge and a surface chemistry of the substrate may have a positive charge which is electrostatically attracted to the negative charge. In some cases, a single bead may comprise a plurality of moieties (e.g., at the same location or different locations on template strand(s)) capable of binding to the substrate. A single moiety, any subset of moieties, or all of the plurality of moieties may be used to bind to the substrate. In some cases, a single binder on the substrate may bind to a single bead. In some cases, multiple binders on the substrate may bind to a single bead. In some cases, the second bead may be configured to not bind to the substrate (e.g., the second bead may not comprise a binder moiety). In some cases, a template may be bound at one end to the first bead and at the second end to the second bead. In other cases, a template may be bound to the first or second bead at a location that is not at the 5’ or 3’ end of a strand, for example at a base that is adjacent to a 5’ or 3’ end of a strand. In this workflow, a template may not be directly bound to the substrate. Beneficially, both the first and second beads bound to the
[0454] -83-
[0455] SUBSTITUTE SHEET (RULE 26) templates in the template-bead assemblies may function as spacers and / or the first beads may function as self-assembling objects on the substrate which prevents one template from being immobilized too close to another template on the substrate. In addition, beneficially, second beads (e.g., those beads not bound to the substrate) may function to orient templates (e.g., to maximize the distance between templates). For example, after loading, a template-to-template pitch (center-to-center distance) may be at least a bead-to-bead pitch (center-to-center distance) when the template-bead assemblies are loaded as a result of the spacing / self-assembling between the beads. In addition, after loading templates may be stretched between first and second beads, and templates may be oriented in a plane perpendicular to a plane of a surface of the substrate. After spacing out and / or self-assembling, the beads may be cleaved or otherwise removed from the templates and washed away. The cleaving may occur before or after the templates bind to the substrate. The beads may be washed away after the templates bind to the substrate. Where non- pre-enriched mixture of positive and negative beads are deposited, the average template-to- template pitch may be greater due to the presence of non-template-bound negative beads also loaded on the substrate. In some cases, after loading, an average template-to-template pitch may be at least an average bead-to-bead pitch and / or at least an average bead diameter. In some cases, upon depositing the template-bead assemblies on the substrate, the template-bead assemblies may be permitted to self-assemble or space out on the substrate before a binding reaction between the bead and the substrate is activated via one or more stimuli (e.g., chemical reagent, catalyst, light (e.g., UV light), heat, etc.) to bind the beads to the substrate. In other cases, the template-bead assemblies may be deposited on the substrate under conditions sufficient to permit binding of the bead to the substrate upon contact. After the beads are immobilized, in some cases, the beads may be subjected to shrinking conditions. In some cases, the bead shrinking may be permanent or substantially persistent with respect to the length of time required to generate sequence read(s) (e.g., the beads may be subjected to conditions causing irreversible shrinking or fixed into a shrunk stated, e.g., by cross-linking of portions of a bead to itself). In some cases, the bead shrinking may be temporary (e.g., the beads may return to substantially their original size during some part of the sequencing). Templates attached to the beads may be further spaced apart from neighboring templates via the shrinking as on average each template is pulled closer to the center of the bead in each template-bead assembly.
[0456] -84-
[0457] SUBSTITUTE SHEET (RULE 26)
[0285] In some cases, the spacer objects (e.g., the second beads) may be another type of spacer instead of beads. For example, in some cases, a nanoball-template-bead assembly may be used, where the bead binds to the surface as described with respect to FIGs. 12A-12D. Alternatively or in addition, in some cases, a dendrimer-template-bead assembly or a DNA origami-template- bead assembly may be used, where the bead binds to the surface as described with respect to FIGs. 12A-12D
[0458]
[0286] Nucleic acids may be loaded onto a substrate using DNA nanostructures (e.g., DNA nanoballs and / or DNA origami). For example, in each of the workflows described with respect to FIG. 12A-12D, the beads may be replaced with DNA nanoballs and / or DNA origami. In some cases, a combination of beads and DNA nanoballs (and / or DNA origami) may be used to load nucleic acids onto a substrate. Alternatively, or in addition, in some cases, a combination of beads and DNA nanostructures (e.g., DNA nanoballs, DNA origami, or other DNA organizations) may be used to load nucleic acids onto a substrate Alternatively or in addition, in some cases a combination of beads, DNA nanostructures, and dendrimers may be used to load nucleic acids onto a substrate. For example, in some cases, a first plurality of templates may be assembled with DNA nanoballs and a second plurality of templates may be associated with beads. In some cases, a plurality of bead-template-nanoball assemblies may be used. In some cases, a plurality of dendrimer-template-nanoball assemblies may be used. In some cases a plurality of nanoball-template-nanoball assemblies may be used. The first and second plurality of template assemblies may be loaded onto a substrate concurrently or sequentially. FIGs. 12E- 12H illustrate different workflows for loading nucleic acids using DNA nanoballs as spacers. In each case, DNA origami may be used in place of DNA nanoballs.
[0459]
[0287] As shown in FIG. 12E, template-nanoball assemblies may be loaded onto a substrate such that a template of a given template-nanoball assembly binds to the substrate. In some cases, nucleic acid nanostructures (e.g., DNA origami, or other organized nucleic acid structures), beads, dendrimers, or a combination thereof may be used instead of or in addition to nanoballs. Any one or more loading mechanisms described elsewhere herein may be used for loading template-nanoball assemblies. For example, loading of template-nanoball assemblies may be performed as described with respect to FIG. 12A. In some cases, a template may be bound at, or comprise at, one end to the nanoball and at the other end to the substrate. In other cases, a template may be bound to the nanoball and / or the substrate at a location that is not at the 5’ or 3’
[0460] -85-
[0461] SUBSTITUTE SHEET (RULE 26) end of a strand, for example at a base that is adjacent to a 5’ or 3’ end of a strand. Beneficially, the nanoballs bound to the templates in the template-nanoball assemblies may function as spacers and / or self-assembling objects on the substrate which prevents one template from binding too close to another template on the substrate. For example, after loading, a template-to- template pitch (center-to-center distance) may be at least a nanoball-to-nanoball pitch (center-to- center distance) when the template-nanoball assemblies are loaded as a result of the spacing / self- assembling between the nanoballs. In some cases, after loading, an average template-to-template pitch may be at least an average nanoball-to-nanoball pitch and / or at least an average nanoball diameter. In some cases, upon depositing the template-nanoball assemblies on the substrate, the template-nanoball assemblies may be permitted to self-assemble or space out on the substrate before a binding reaction between the template and the substrate is activated via one or more stimuli (e.g., chemical reagent, catalyst, light (e.g., UV light), heat, etc.) to bind the templates to the substrate. In other cases, the template-nanoball assemblies may be deposited on the substrate under conditions sufficient to permit binding of the template to the substrate upon contact. After spacing out and / or self-assembling, the nanoballs may be cleaved or otherwise removed from the templates and washed away. The cleaving may occur before or after the templates bind to the substrate. The nanoballs may be washed away after the templates bind to the substrate.
[0462]
[0288] In some cases, instead of cleaving nanoballs, the nanoballs may be protected (e.g., may be double-stranded) and thus unavailable for sequencing. In such cases, after loading and binding of template-nanoball assemblies, the templates may be subjected to conditions sufficient for sequencing (e g., where a sequencing primer may anneal to a template and not to nanoballs).
[0463]
[0289] Alternatively, as shown in FIG. 12F, template-nanoball assemblies may be loaded onto a substrate such that a nanoball of a given template-nanoball assembly binds to the substrate. In some cases, nucleic acid nanostructures (e.g., DNA origami, or other organized nucleic acid structures), beads, dendrimers, or a combination thereof may be used instead of or in addition to nanoballs. Any one or more loading mechanisms described elsewhere herein may be used for loading template-nanoball assemblies. For example, loading of template-nanoball assemblies may be performed as described with respect to FIG. 12B. The substrate may be patterned or unpatterned. For example, the substrate may be patterned with binders that are configured to bind to nanoballs of template-nanoball assemblies. In some cases, the substrate may be patterned or coated with DNA nanostructures, where the DNA nanostructures are configured to bind to
[0464] -86-
[0465] SUBSTITUTE SHEET (RULE 26) nanoballs of template-nanoball assemblies. In another example, the substrate may be unpattemed such that there is a substantially uniform coating of a surface chemistry on the substrate. The surface chemistry may comprise binders that are configured to bind to nanoballs. Any one or more coupling mechanisms described elsewhere herein may be used for the nanoball-substrate binding, such as any click chemistry pair, complementary oligonucleotides that hybridize, magnetic particles that are forced by magnetic fields, electric particles that are forced by electric fields, specific binding, non-specific binding, electrostatic interactions, cross-linking, etc. The substrate and / or nanoball may comprise any binder described elsewhere herein. In some cases, the nanoball may comprise a first coupler of a coupling pair and the substrate may comprise a second coupler of the coupling pair. In some cases, a single nanoball may comprise a single moiety (e.g., azide moiety, DBCO moiety, thiol moiety, oligonucleotide sequence, a crosslinking base, etc.) capable of binding to the substrate. In another example, the nanoball may have a negative charge and a surface chemistry of the substrate may have a positive charge which is electrostatically attracted to the negative charge. In some cases, a single nanoball may comprise a plurality of moieties (e.g., at the same location or different locations on template strand(s)) capable of binding to the substrate. A single moiety, any subset of moieties, or all of the plurality of moieties may be used to bind to the substrate. In some cases, a single binder on the substrate may bind to a single nanoball. In some cases, multiple binders on the substrate may bind to a single nanoball. In some cases, a template may be bound at one end to the nanoball. In other cases, a template may be bound to the nanoball at a location that is not at the 5’ or 3’ end of a strand, for example at a base that is adjacent to a 5’ or 3’ end of a strand. In this workflow, a template may not be directly bound to the substrate. Beneficially, the nanoballs bound to the templates in the template-nanoball assemblies may function as spacers and / or self-assembling objects on the substrate which prevents one template from being immobilized too close to another template on the substrate. For example, after loading, a template-to-template pitch (center-to-center distance) may be at least a nanoball-to-nanoball pitch (center-to-center distance) when the template-nanoball assemblies are loaded as a result of the spacing / self- assembling between the nanoballs. In some cases, after loading, an average template-to-template pitch may be at least an average nanoball-to-nanoball pitch and / or at least an average nanoball diameter. In some cases, upon depositing the template-nanoball assemblies on the substrate, the template-nanoball assemblies may be permitted to self-assemble or space out on the substrate
[0466] -87-
[0467] SUBSTITUTE SHEET (RULE 26) before a binding reaction between the nanoball and the substrate is activated via one or more stimuli (e.g., chemical reagent, catalyst, light (e.g., UV light), heat, etc.) to bind the nanoballs to the substrate. In other cases, the tempi ate-nanob all assemblies may be deposited on the substrate under conditions sufficient to permit binding of the nanoball to the substrate upon contact. After the nanoballs are immobilized, in some cases, the nanoballs may be subjected to shrinking conditions. In some cases, the shrinking may be permanent or persistent (e.g., persisting for at least some of a sequencing run). Templates attached to the nanoballs may be further spaced apart from neighboring templates after the shrinking as on average each template is pulled closer to the center of the nanoball in each template-nanoball assembly.
[0468]
[0290] In some cases, empty nanoballs or negative nanoballs not bound to any template may be co-deposited onto the substrate with the template-nanoball assemblies. The presence of additional negative nanoballs between template-nanoball assemblies may additionally space out the templates and increase the average template-to-template pitch (center-to-center distance).
[0469]
[0291] Alternatively, as shown in FIG. 12G, empty nanoballs not bound to templates may be loaded onto a substrate such that the nanoballs bind to the substrate, and then templates may be deposited onto the nanoball-bound substrate to bind templates to the nanoballs. In some cases, nucleic acid nanostructures (e.g., DNA origami, or other organized nucleic acid structures), beads, or a combination thereof may be used instead of or in addition to nanoballs. Unbound nanoballs may be washed away before depositing the templates. The substrate may be patterned or unpatterned. For example, the substrate may be patterned with binders that are configured to bind to nanoballs. In some cases, the substrate may be patterned or coated with DNA nanostructures as described elsewhere herein, where the DNA nanostructures are configured to bind to nanoballs. In another example, the substrate may be unpatterned such that there is a substantially uniform coating of a surface chemistry on the substrate. The surface chemistry may comprise binders that are configured to bind to nanoballs. Any one or more coupling mechanisms described elsewhere herein may be used for the nanoball-substrate binding, such as any click chemistry pair, complementary oligonucleotides that hybridize, magnetic particles that are forced by magnetic fields, electric particles that are forced by electric fields, specific binding, non-specific binding, electrostatic interactions, cross-linking, etc. The substrate and / or nanoball may comprise any binder described elsewhere herein. In some cases, the nanoball may comprise a first coupler of a coupling pair and the substrate may comprise a second coupler of the
[0470] -88-
[0471] SUBSTITUTE SHEET (RULE 26) coupling pair. In some cases, a single nanoball may comprise a single moiety (e.g., azide moiety, DBCO moiety, thiol moiety, oligonucleotide sequence, a cross-linking base, etc.) capable of binding to the substrate. In another example, the nanoball may have a negative charge and a surface chemistry of the substrate may have a positive charge which is electrostatically attracted to the negative charge. In some cases, a single nanoball may comprise a plurality of moieties (e.g., at the same location or different locations on template strand(s)) capable of binding to the substrate. A single moiety, any subset of moieties, or all of the plurality of moieties may be used to bind to the substrate. In some cases, a single binder on the substrate may bind to a single nanoball. In some cases, multiple binders on the substrate may bind to a single nanoball. Upon loading and immobilization of the nanoballs to the substrate, unbound templates may be loaded onto the nanoball-coupled substrate. In some cases, a single nanoball may comprise a single moiety (e.g., azide moiety, DBCO moiety, thiol moiety, oligonucleotide sequence, a crosslinking base, etc.) capable of binding to the template. In some cases, a single nanoball may comprise a plurality of moieties (e.g., at the same location or different locations on template strand(s)) capable of binding to the template. A single moiety, any subset of moieties, or all of the plurality of moieties may be used to bind to the template. In some cases, a single binder on the template may bind to a single nanoball. In some cases, multiple binders on the template may bind to a single nanoball. In some cases, a template may be bound at one end to the nanoball. In other cases, a template may be bound to the nanoball at a location that is not at the 5’ or 3’ end of a strand, for example at a base that is adjacent to a 5’ or 3’ end of a strand. In this workflow, a template may not be directly bound to the substrate. In some cases, each nanoball may bind to at most template. In some cases, each nanoball may bind at most 10, 9, 8, 7, 6, 5, 4, 3, 2, or 1 template. Beneficially, the nanoballs bound to the templates may function as spacers and / or selfassembling objects on the substrate which prevents one template from being immobilized too close to another template on the substrate. For example, after loading the templates, a template- to-template pitch (center-to-center distance) may be at least a nanoball-to-nanoball pitch (center- to-center distance) as a result of the spacing / self-assembling between the nanoballs. In some cases, after loading the templates, an average tempi ate-to-templ ate pitch may be at least an average nanoball-to-nanoball pitch and / or at least an average nanoball diameter. In some cases, upon depositing the nanoballs on the substrate, the nanoballs may be permitted to self-assemble or space out on the substrate before a binding reaction between the nanoball and the substrate is
[0472] -89-
[0473] SUBSTITUTE SHEET (RULE 26) activated via one or more stimuli (e.g., chemical reagent, catalyst, light (e.g., UV light), heat, etc.) to bind the nanoballs to the substrate. In other cases, the nanoballs may be deposited on the substrate under conditions sufficient to permit binding of the nanoball to the substrate upon contact. Before or after the templates are bound to the nanoballs, in some cases, the nanoballs may be subjected to shrinking conditions. Templates attached to the nanoballs may be further spaced apart from neighboring templates via the shrinking as on average each template or template binding site is pulled closer to the center of the nanoball in each template-nanoball assembly.
[0474]
[0292] Alternatively, as shown in FIG. 12H, nanoball-template-bead assemblies may be loaded onto a substrate such that a nanoball of a given nanoball-template-bead assembly binds to the substrate. The nanoball-template-bead assemblies may comprise a template coupled at or near a first end to a bead and the template coupled at or near a second end to a nanoball. The bead may function as a spacer, as described with respect to FIG. 12D. Any one or more loading mechanisms described elsewhere herein (e.g., as described with respect to FIGs. 12C, 12D, and 12F) may be used for loading nanoball-template-bead assemblies.
[0475]
[0293] Nucleic acids may be loaded onto a substrate using dendrimers. For example, in each of the workflows described with respect to FIG. 12A-12H, the beads and / or DNA nanoballs may be replaced with dendrimers. In some cases, a combination of beads and dendrimers may be used to load nucleic acids onto a substrate. Alternatively, or in addition, in some cases, a combination of beads, DNA nanoballs / DNA origami, and dendrimers may be used to load nucleic acids onto a substrate. For example, in some cases, a first plurality of templates may be assembled with dendrimers and a second plurality of templates may be associated with beads. The first and second plurality of template assemblies may be loaded onto a substrate concurrently or sequentially. FIGs. 12I-12L illustrate different workflows for loading nucleic acids using dendrimers as spacers.
[0476]
[0294] As shown in FIG. 121, template-dendrimer assemblies may be loaded onto a substrate such that a template of a given template-dendrimer assembly binds to the substrate. Any one or more loading mechanisms described elsewhere herein (e.g., as described with respect to FIGs. 12A, 12B, and 12E) may be used for loading template-dendrimer assemblies. In some cases, a template may be bound at, or comprise at, one end to the dendrimer and at the other end to the substrate. In other cases, a template may be bound to the dendrimer and / or the substrate at a
[0477] -90-
[0478] SUBSTITUTE SHEET (RULE 26) location that is not at the 5 ’ or 3 ’ end of a strand, for example at a base that is adj acent to a 5 ’ or 3’ end of a strand. Beneficially, the dendrimers bound to the templates in the template-dendrimer assemblies may function as spacers and / or self-assembling objects on the substrate which prevents one template from binding too close to another template on the substrate. For example, after loading, a template-to-template pitch (center-to-center distance) may be at least a dendrimer-to-dendrimer pitch (center-to-center distance) when the template-dendrimer assemblies are loaded as a result of the spacing / self-assembling between the dendrimers. In some cases, after loading, an average template-to-template pitch may be at least an average dendrimer- to-dendrimer pitch and / or at least an average dendrimer diameter. In some cases, upon depositing the template-dendrimer assemblies on the substrate, the template-dendrimer assemblies may be permitted to self-assemble or space out on the substrate before a binding reaction between the template and the substrate is activated via one or more stimuli (e.g., chemical reagent, catalyst, light (e.g., UV light), heat, etc.) to bind the templates to the substrate. In other cases, the template-dendrimer assemblies may be deposited on the substrate under conditions sufficient to permit binding of the template to the substrate upon contact. After spacing out and / or selfassembling, the dendrimers may be cleaved or otherwise removed from the templates and washed away. The cleaving may occur before or after the templates bind to the substrate. The dendrimers may be washed away after the templates bind to the substrate.
[0479]
[0295] Beneficially, the size of dendrimers may be tightly controlled (e.g., by regulating the polymerization and assembly of dendrimer components In addition, similar to DNA origami and DNA nanoballs, the sites where dendrimers may couple to templates may be regulated such that for each template-dendrimer assembly, the template is oriented in a same way with respect to the dendrimer. This may function to couple a template to the center of a dendrimer, thus maximizing the pitch between adjacent templates. In addition, the number and location of site(s) where a dendrimer may couple to the substrate may be regulated such that the orientation of template to substrate may be uniform or substantially uniform in a plurality of substrate-bound templatedendrimer assemblies.
[0480]
[0296] Alternatively, as shown in FIG. 12J, template-dendrimer assemblies may be loaded onto a substrate such that a dendrimer of a given template-dendrimer assembly binds to the substrate. Any one or more loading mechanisms described elsewhere herein may be used for loading template-dendrimer assemblies. For example, loading of template-dendrimer assemblies may be
[0481] -91-
[0482] SUBSTITUTE SHEET (RULE 26) performed as described with respect to FIGs. 12B and 12F. The substrate may be patterned or unpatterned. For example, the substrate may be patterned with binders that are configured to bind to dendrimers of template-dendrimer assemblies. In some cases, the substrate may be patterned or coated with DNA nanostructures, where the DNA nanostructures are configured to bind to dendrimers of template-dendrimer assemblies. In another example, the substrate may be unpatterned such that there is a substantially uniform coating of a surface chemistry on the substrate. The surface chemistry may comprise binders that are configured to bind to dendrimers. Any one or more coupling mechanisms described elsewhere herein may be used for the dendrimer-substrate binding, such as any click chemistry pair, complementary oligonucleotides that hybridize, magnetic particles that are forced by magnetic fields, electric particles that are forced by electric fields, specific binding, non-specific binding, electrostatic interactions, crosslinking, etc. The substrate and / or dendrimer may comprise any binder described elsewhere herein. In some cases, the dendrimer may comprise a first coupler of a coupling pair and the substrate may comprise a second coupler of the coupling pair. In some cases, a single dendrimer may comprise a single moiety (e.g., azide moiety, DBCO moiety, thiol moiety, oligonucleotide sequence, a cross-linking base, etc.) capable of binding to the substrate. In another example, the dendrimer may have a negative charge and a surface chemistry of the substrate may have a positive charge which is electrostatically attracted to the negative charge. In some cases, a single dendrimer may comprise a plurality of moi eties (e g., at the same location or different locations on the dendrimer) capable of binding to the substrate. A single moiety, any subset of moieties, or all of the plurality of moieties may be used to bind to the substrate. In some cases, a single binder on the substrate may bind to a single dendrimer. In some cases, multiple binders on the substrate may bind to a single dendrimer. In some cases, a template may be bound at one end to the dendrimer. In other cases, a template may be bound to the dendrimer at a location that is not at the 5’ or 3’ end of a strand, for example at a base that is adjacent to a 5’ or 3’ end of a strand. In this workflow, a template may not be directly bound to the substrate. Beneficially, the dendrimers bound to the templates in the template-dendrimer assemblies may function as spacers and / or self-assembling objects on the substrate which prevents one template from being immobilized too close to another template on the substrate. For example, after loading, a template-to-template pitch (center-to-center distance) may be at least a dendrimer-to-dendrimer pitch (center-to-center distance) when the template-dendrimer assemblies are loaded as a result
[0483] -92-
[0484] SUBSTITUTE SHEET (RULE 26) of the spacing / self-assembling between the dendrimers. In some cases, after loading, an average template-to-template pitch may be at least an average dendrimer-to-dendrimer pitch and / or at least an average dendrimer diameter. In some cases, upon depositing the template-dendrimer assemblies on the substrate, the template-dendrimer assemblies may be permitted to selfassemble or space out on the substrate before a binding reaction between the dendrimer and the substrate is activated via one or more stimuli (e.g., chemical reagent, catalyst, light (e.g., UV light), heat, etc.) to bind the dendrimers to the substrate. In other cases, the template-dendrimer assemblies may be deposited on the substrate under conditions sufficient to permit binding of the dendrimer to the substrate upon contact. After the dendrimers are immobilized, in some cases, the dendrimers may be subjected to shrinking conditions. In some cases, the shrinking may be permanent or persistent (e.g., persisting for at least some of a sequencing run). Templates attached to the dendrimers may be further spaced apart from neighboring templates after the shrinking as on average each template is pulled closer to the center of the dendrimer in each template-dendrimer assembly.
[0485]
[0297] In some cases, empty dendrimers or negative dendrimers (or negative beads) not bound to any template may be co-deposited onto the substrate with the template-dendrimer assemblies. The presence of additional negative dendrimers (or negative beads) between template-dendrimer assemblies may additionally space out the templates and increase the average template-to- template pitch (center-to-center distance).
[0486]
[0298] Alternatively, as shown in FIG. 12K, empty dendrimers not bound to templates may be loaded onto a substrate such that the dendrimers bind to the substrate, and then templates may be deposited onto the dendrimer-bound substrate to bind templates to the dendrimers. Unbound dendrimers may be washed away before depositing the templates. Any one or more loading mechanisms described elsewhere herein may be used for loading template-dendrimer assemblies. For example, loading of dendrimers may be performed as described with respect to FIGs. 12G and 12J. Upon loading and immobilization of the dendrimers to the substrate, unbound templates may be loaded onto the dendrimer-coupled substrate. In some cases, a single dendrimer may comprise a plurality of moieties (e.g., at the same location or different locations on template strand(s)) capable of binding to the template. A single moiety, any subset of moieties, or all of the plurality of moieties may be used to bind to the template. In some cases, a single binder on the template may bind to a single dendrimer. In some cases, multiple binders on the template
[0487] -93-
[0488] SUBSTITUTE SHEET (RULE 26) may bind to a single dendrimer. In some cases, a template may be bound at one end to the dendrimer.
[0489]
[0299] Alternatively, as shown in FIG. 12L, dendrimer- template-bead assemblies may be loaded onto a substrate such that a dendrimer of a given dendrimer- template-bead assembly binds to the substrate. The dendrimer- template-bead assemblies may comprise a template coupled to a bead at or near a first end and the template coupled to a dendrimer at or near another end. The bead may function as a spacer, as described with respect to FIGs. 12D and 12H. Alternatively or in addition, dendrimer-template-nanoball assemblies, dendrimer-template-DNA origami assemblies, or dendrimer-template-(additional)dendrimer assemblies may be used, where the dendrimers are bound to the substrate and the nanoballs, DNA origami, or additional dendrimers respectively, serve as spacers Loading and / or shrinking of the dendrimer- templatebead assemblies may be performed as described with respect to FIGs. 12C, 12D, and 12F-12H.
[0490]
[0300] Beneficially, the loading mechanisms described herein may immobilize templates in a spaced-apart manner which enables spatial discerning of signals collected from individual templates immobilized to the substrate during sequencing reactions, such as single molecule sequencing reactions or concatemer sequencing reactions. A template may be a concatemer molecule or a non-concatemer molecule. The spacing apart may also reduce inter-dye effects, such as quenching or FRET, between dyes coupled to different templates that may affect sequencing quality.
[0491]
[0301] It will be appreciated that another particle, support, or object may be used in place of nanoballs, dendrimers, and beads in these workflows, such as a DNA origami particle, or non- DNA objects, such as dendrimers. In some cases, any combination of particles, supports, objects, beads, DNA nanostructures, DNA origami, DNA nanoballs, and dendrimers may be used in these workflows.
[0492]
[0302] Different conditions and methods for shrinking beads, which may be applied generally to nanoballs, dendrimers, and other objects (e.g., DNA origami particle, nanoparticles, etc.), are described in further detail in U.S. Patent Pub. No. 2023 / 0340570A1 and International Pub. No. 2023 / 069648A1, each of which is incorporated by reference herein in its entirety for all purposes. For example, particles (e.g., nanoballs, beads, dendrimers, etc.) may be subjected to incubation with a buffer solution comprising a polymer, such as polyethylene glycol (PEG),
[0493] -94-
[0494] SUBSTITUTE SHEET (RULE 26) and / or a cation, such as a divalent cation, to shrink them. The substrate may be subjected to one or more washing operations, such as before, during, or after shrinking the particles.
[0495]
[0303] In some instances, the cations may be magnesium ions, calcium ions, or spermine ions (e.g., spermine14, spermine2, spermine34, spermine4+, etc.), or a combination thereof. In some cases, the cations may comprise ions of aluminum, barium, bismuth, cadmium, calcium, cesium, chromium, cobalt, copper, copper, hydrogen, iron, iron, lead, lithium, magnesium, mercury, mercury, nickel, potassium, rubidium, silver, sodium, strontium, tin, or spermine. In some cases, the cations may comprise A13+, Ba2+, Bi3+, Cd2+, Cal+, Ca2+, Csl+, CrH, Co2+, Cul+, Cu2+, H1+, Fe2+, Fe3+, Pb2+, Lil+, Mgl+, Mg2+, Hg22+, Hg2+, Ni2+, K1+, Rbl+, Agl+, Nal+, Sr2+, Sn2+, sperminel+, spermine2+, spermine3+, or spermine4+. The cation may facilitate shrinking of particle (e.g., beads, nanoballs, dendrimers, etc.) sizes. The substrate may be treated with a cation buffer solution to facilitate shrinking of particles. In some cases, a cation buffer solution may comprise about, at least about, and / or at most about 5 mM, 10 mM, 15 mM, 20 mM, 25 mM, 30 mM, 35 mM, 40 mM, about 45 mM, 50 mM, 55 mM, 60 mM, 65 mM, 70 mM, 75 mM, 80 mM, 85 mM, 90 mM, 95 mM, or 100 mM of cations. In some cases, the substrate may be treated with a PEG solution. The PEG molecule in the solution may have a molecular mass of up to about 100, 200, 300, 400, 500, 1000, 1500, 2000, 2500, 3000, 3500, 4000, 4500, 5000, 5500, 6000, 6500, 7000, 7500, 8000, 8500, 9000, 9500, 10000, 10500, 11000, 11500, 12000, 12500, 13000, 13500, 14000, 14500, 15000, 15500, 16000, 16500, 17000, 17500, 18000, 18500, 19000, 19500, 20000, or more Da. In some cases, a PEG molecule may have a molecular mass of more than about 20,000 Da. In some cases, a PEG molecule may have a molecular mass of less than about 100 Da. In some instances, a PEG molecule may have a molecular mass within a range defined by any two of the preceding values. In some cases, a PEG molecule may have a molecular weight of at least about 1 x 104, 2 x 104, 5 x 104, 1 x 105, 2 x 105, 5 x 105, 1 x 106, 2 x 106, 5 x 106, 1 x 107, 2 x 107, 5 x 107, 1 x 108or more grams per molecule (g / mol). In some cases, a PEG molecule may have a molecular weight of more than about 1 x 108g / mol. In some cases, a PEG molecule may have a molecular weight of less than about 1 x 104g / mol. In some cases, a PEG molecule may have a molecular weight within a range defined by any two of the preceding values. In some instances, the PEG concentration may be up to about 0.1%, 0.2%, 0.3%, 0.4%, 0.5%, 0.6%, 0.7%, 0.8%, 0.9%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, by weight, of a buffer solution. In some
[0496] -95-
[0497] SUBSTITUTE SHEET (RULE 26) cases, the PEG concentration may be less than about 0.1% by weight, of a buffer solution. In some cases, the PEG concentration may be more than about 50% by weight, of a buffer solution. In some instances, the PEG concentration may be a percent by weight of a buffer solution within a range defined by any two of the preceding values. In some cases, the PEG concentration may be up to about 0.1%, 0.2%, 0.3%, 0.4%, 0.5%, 0.6%, 0.7%, 0.8%, 0.9%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, by volume, of a buffer solution. In some cases, the PEG concentration may be less than about 0. 1%, by volume, of a buffer solution. In some cases, the PEG concentration may be more than about 50%, by volume, of a buffer solution. In some instances, the PEG concentration may be a percent by volume of a buffer solution within a range defined by any two of the preceding values.
[0498]
[0304] In some instances, a layer of buffer solution comprising a cation and / or polymer molecule formed on the substrate may have a thickness of from about 1 nm to about 10 nm, from about 1 nm to about 100 nm, from about 1 nm to about 1 pm, from about 1 nm to about 10 pm, from about 1 nm to about 100 pm, or from about 1 nm to about 1 mm. In some cases, a layer of buffer solution formed on the substrate may have a thickness about 1 pm to about 40 pm, from about 1 pm to about 39 pm, from about 2 pm to about 38 pm, from about 3 pm to about 37 pm, from about 4 pm to about 36 pm, from about 5 pm to about 35 pm, from about 6 pm to about 34 pm, from about 7 pm to about 33 pm, from about 8 pm to about 32 pm, from about 9 pm to about 31 pm, from about 10 pm to about 30 pm, from about 11 pm to about 29 pm, from about 12 pm to about 28 pm, from about 13 pm to about 27 pm, from about 14 pm to about 26 pm, from about 15 pm to about 25 pm, from about 16 pm to about 24 pm, from about 17 pm to about 23 pm, from about 18 pm to about 22 pm, from about 19 pm to about 21 pm, from about 1 pm to about 20 pm, from about 5 pm to about 20 pm, from about 10 pm to about 20 pm, from about 15 pm to about 20 pm, from about 10 pm to about 25 pm, from about 10 pm to about 30 pm, from about 10 pm to about 35 pm, from about 10 pm to about 40 pm, from about 10 pm to about 20 pm, from about 10 pm to about 25 pm, from about 10 pm to about 30 pm, from about 10 pm to about 35 pm, from about 10 pm to about 40 pm, from about 5 pm to about 20 pm, from about 4 pm to about 20 pm, from about 3 pm to about 20 pm, from about 2 pm to about 20 pm, or from about 1 pm to about 20 pm. In some instances, a layer of buffer solution formed on the substrate may have a thickness of at least about 0.1 nm, 0.2 nm, 0.5 nm, 1 nm, 2 nm, 5 nm, 10 nm, 20 nm,
[0499] -96-
[0500] SUBSTITUTE SHEET (RULE 26) 50 nm, 100 nm, 200 nm, 500 nm, 1 pm, 2 pm, 5 pm, 10 pm, 20 gm, 50 pm, 100 pm, 200 pm, 500 pm, 1 mm, or more than at least about 1 mm.
[0501]
[0305] In some instances, an average size of a plurality of particles is measured in full-width at half-maximum (FWHM). As used herein, the term “FWHM” refers to a size (e.g., a diameter) of a particle determined from fluorescence imaging. In some instances, FWHM is the width of an intensity profile for the imaged particle, measured at the median intensity value (e.g., amplitude) detected from the particle (e.g., from an intensity profile of the fluorescence emitted from the particle) For instance, the FWHM may be determined for one or more particles in the plurality of particles, and an average size may be determined by averaging the one or more FWHM values so determined. In some instances, an intensity line profile corresponding to a respective particle is extracted from an image of the substrate. In some such instances, the FWHM for the particle is measured directly from the intensity line profile. In some such instances, the FWHM for the particle is estimated by fitting a Gaussian to the intensity line profile. In some instances, the FWHM for the particle is determined from a gray value version of the line intensity profile of the particle. In some instances, a FWHM may be determined for a particle at multiple time points (e.g., prior to, upon, and / or subsequent to a washing operation). In some instances, an average FWHM of a plurality of particles prior to subsequent to shrinking may be about, at least about, and / or at most about 0.1 pm, 0.5 pm, 1 pm, 5 pm, 10 pm, 50 pm, 100 pm, 500 pm, 1000 pm or 1 nm, 5 nm, 10 nm, 50 nm, 100 nm, 500 nm, 1000 nm or 1 gm, 5 pm, 10 pm, 50 pm, 100 pm, 500 pm, 1000 pm or 1 mm, or more. In some cases, the average FWHM of a plurality of particles may shrink by about, at least about, and / or at most about 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more.
[0502]
[0306] In some cases, DNA particles such as DNA nanoballs and DNA origami particles may be shrunk via applying staples that bind to different strand segments, thus reducing segment-to- segment distance(s) within a molecule and providing a reduced size form of the particle. In some cases, particles such as DNA nanoballs, DNA origami, dendrimers, etc. may be subjected to shrinking via other intra-molecular or inter-molecular linking mechanisms. In one example, a dendrimer particle may comprise thiol moieties which may link with each other to form disulfide bonds or link with an intermediary molecule that comprises thiol moieties to form disulfide bonds that reduces segment-to-segment distance(s) within a molecule to provide a reduced size
[0503] -97-
[0504] SUBSTITUTE SHEET (RULE 26) form of the particle. In another example, a DNA particle may comprise cross-linkable bases (e.g., CNVK) that may link with other bases within the molecule or link with bases of an intermediary molecule to reduce segment-to-segment distance(s) within a molecule to provide a reduced size form of a particle. In another example, a DNA particle or dendrimer particle may comprise ‘click’ able moieties that may link via click chemistry with each other or with an intermediary molecule that comprises complementary ‘click’able moieties. Accordingly, methods provided herein may comprise generating particles with linking reagents (e.g., for DNA particles a modified base comprising a thiol group, a modified base comprising a click chemistry group, a modified base that is cross-linkable, an oligonucleotide sequence that binds to a staple, an oligonucleotide sequence that binds to another intramolecular oligonucleotide sequence, etc.; e g., for dendrimer particles or other non-nucleic acid particles a thiol chemistry group (e.g., a thiol linker) a click chemistry group, a moiety that is cross-linkable, etc.). For example, such linking reagents may be incorporated during synthesis of a DNA nanoball, DNA origami, or dendrimer particles. Alternatively or in addition, for DNA particles, a starting material, such as a primer that hybridizes to a circular template or an origami scaffold may comprise the linking reagent. The linking may be readily activatable, such as by providing one or more stimuli (e.g., providing staple reagents, providing light for crosslinking reaction, etc.). The linking may be reversible. The linking may be irreversible.
[0505]
[0307] In some cases, a template-nanoball assembly, as described in FIGs. 12E-12H, may be generated by coupling a template to a nanoball. In some cases, the nanoball may be generated via a primer hybridized to and extending using a circular template, the primer comprising a template-binding moiety. The template-binding moiety may comprise any coupling mechanism described elsewhere herein. Thus, a nanoball generated from the primer comprises the templatebinding moiety which can bind to the template. In some cases, the nanoball may be generated via a primer hybridized to and extending using a circular template, the primer being an extension from and / or being coupled to the template. The primer may be covalently or non-covalently coupled to the template. Thus, a nanoball generated from the primer comprises at one end the template and at the other end a concatemeric nanoball that is based off the circular template. In some cases, the circular template is distinct from the template to be sequenced.
[0506]
[0308] Single molecule dual strand sequencing
[0507] -98-
[0508] SUBSTITUTE SHEET (RULE 26)
[0309] In some cases, a single template molecule may comprise a double-stranded template with a first strand and a second strand complementary to at least a portion of the first strand. It may be beneficial for sequencing accuracy to obtain sequencing information from both strands of the double-stranded template molecule. See International Patent Appl. No. PCT / US2024 / 013236, which is hereby incorporated by reference in its entirety, for examples of simultaneously sequencing both strands from double-stranded templates. An exemplary method of retaining both strands of a double-stranded template molecule and localizing signals for both strands to a same individually addressable location is illustrated in FIG. 13.
[0509]
[0310] Provided herein is a method for sequencing (e.g., single molecule sequencing), comprising: providing a double-stranded template molecule comprising a first strand and a second strand, where the second strand is complementary to at least part of the first strand (e.g.,); contacting the double-stranded template molecule with a plurality of adapters to couple an adapter from the plurality of adapters to each end of the double-stranded template molecule to provide a double-stranded template-adapter molecule, where an adapter of the plurality of adapters comprises a hairpin comprising a double-stranded region and a single-stranded region, where the single-stranded stranded region further comprises a capture moiety; loading the double-stranded template-adapter molecule onto a substrate, where the double-stranded templateadapter molecule couples to the substrate via the capture moieties; denaturing the first strand from the second strand; hybridizing one or more primers to the double-stranded template-adapter molecule; and sequencing the first strand of the double-stranded template molecule.
[0510]
[0311] The template molecule may be from or derived from a sample. The first strand and the second strand may be hybridized to each other. The first strand and / or the second strand may comprise a 3’ or a 5’ overhang (e.g., for adapter annealing and ligation).
[0511]
[0312] Adapters may comprise a hairpin. Adapters may comprise a blocker moiety at a 3 ’ end (e.g., to prevent sequencing through the entire molecule from a single sequencing primer). Each sequencing primer may be extended at most through one strand of the double-stranded template (e.g., either the first strand or the second strand but not both).
[0512]
[0313] Coupling adapters to the ends of the double-stranded template molecule may comprise annealing and ligation of the adapters to the template. In some cases, coupling an adapter to an end of the double-stranded template molecule may comprise ligating a first end of the adapter (e.g., the 5’ end or the 3’ end) to the first or second strand of the double-stranded template
[0513] -99-
[0514] SUBSTITUTE SHEET (RULE 26) molecule and not ligating a second end of the adapter (e.g., the 3’ end or 5’ end) to the second or first strand of the double-stranded template molecule, respectively. That is, an adapter may be coupled to only one strand of the template molecule.
[0515]
[0314] In some cases, the one or more primers comprise two primers. In some cases, the one or more primers anneal to a first region (e.g., a primer attachment site) of the coupled adapters of the double-stranded template-adapter molecule. In some cases, the sequencing of the first strand and the second strand is concurrent. In some cases, the sequencing of the first strand and the second strand is sequential.
[0516]
[0315] The substrate may be patterned or unpatterned. For example, the substrate may be patterned with binders (e.g., capturing moieties) that are configured to bind to capture moieties of the template-adapter molecule. In some cases, the substrate may be patterned or coated with DNA nanostructures as described elsewhere herein, where the DNA nanostructures are configured to bind to capture entities of the template-adapter molecule. In another example, the substrate may be unpattemed such that there is a substantially uniform coating of a surface chemistry on the substrate. The surface chemistry may comprise binders that are configured to bind to template-adapter molecules. For example, the surface chemistry may comprise DBCO moieties, and the capture entities may comprise azide moieties, respectively, which can couple together (e.g., template to surface) via click chemistry. Any one or more coupling mechanisms described elsewhere herein may be used for the template-adapter-substrate binding, such as any click chemistry pair, complementary oligonucleotides that hybridize, magnetic particles that are forced together by magnetic fields, electric particles that are forced together by electric fields, specific binding, non-specific binding, electrostatic interactions, cross-linking, etc. The substrate and / or template may comprise any binder described elsewhere herein. For example, the template may comprise a first coupler of a coupling pair and the substrate may comprise a second coupler of the coupling pair. In some cases, a single template may comprise a single moiety (e.g., azide moiety, DBCO moiety, thiol moiety, oligonucleotide sequence, a cross-linking base, etc.) capable of binding to the substrate. In some case, the substrate may comprise a regular array of streptavidin and the capturing moieties of the template-adapter molecule may comprise biotin. Alternatively, the substrate may comprise a regular array of biotin and the capturing moieties of the template-adapter molecule may comprise streptavidin.
[0517] -100-
[0518] SUBSTITUTE SHEET (RULE 26)
[0316] Sequencing signals from the first stand and the second strand of the double-stranded template may be distinguishable from each other due to distance between the first strand and the second strand after denaturation In some cases, the distance is at least 50 nm, 100 nm, 200 nm, 300 nm, 400 nm, 500 nm, 600 nm, 700 nm, 800 nm, 900 nm, 1000 nm, 1100 nm, 1200 nm, 1300 nm, 1400 nm, 1500 nm, 1600 nm, 1700 nm, 1800 nm, 1900 nm, or 2000 nm. For example, for a 2000 base pair double-stranded template molecule, the first and second strands may be separated up to about 1000 nm after denaturation.
[0519]
[0317] Any sequencing method described herein may be used to sequence the double-stranded template-adapter molecule. Any sequencing analysis method described herein may be used to analyze sequencing data (e.g., sequencing signals obtained from extending a sequencing primer along the first strand or the second strand) from the double-stranded template-adapter molecule. Any loading method described herein may be used to load and couple the double-stranded template-adapter molecule to a substrate.
[0520] Computer systems
[0521]
[0318] The present disclosure provides computer control systems that are programmed to implement methods of this disclosure. FIG. 6 shows a computer system 601 that is programmed or otherwise configured to implement methods of the disclosure, such as to control the systems described herein (e.g., reagent dispensing, detecting, etc.) and collect, receive, and / or analyze sequencing information. The computer system 601 can be an electronic device of a user or a computer system that is remotely located with respect to the electronic device. The electronic device can be a mobile electronic device.
[0522]
[0319] The computer system 601 includes a central processing unit (CPU, also “processor’ and “computer processor” herein) 605, which can be a single core or multi core processor, or a plurality of processors for parallel processing. The computer system 601 also includes memory or memory location 610 (e.g., random-access memory, read-only memory, flash memory), electronic storage unit 615 (e.g., hard disk), communication interface 620 (e.g., network adapter) for communicating with one or more other systems, and peripheral devices 625, such as cache, other memory, data storage and / or electronic display adapters. The memory 610, storage unit 615, interface 620 and peripheral devices 625 are in communication with the CPU 605 through a communication bus (solid lines), such as a motherboard. The storage unit 615 can be a data
[0523] -101-
[0524] SUBSTITUTE SHEET (RULE 26) storage unit (or data repository) for storing data. The computer system 601 can be operatively coupled to a computer network (“network”) 630 with the aid of the communication interface 620. The network 630 can be the Internet, an isolated or substantially isolated internet and / or extranet, or an intranet and / or extranet that is in communication with the Internet. The network 630 in some cases is a telecommunication and / or data network. The network 630 can include one or more computer servers, which can enable distributed computing, such as cloud computing. The network 630, in some cases with the aid of the computer system 601, can implement a peer- to-peer network, which may enable devices coupled to the computer system 601 to behave as a client or a server. The CPU 605 can execute a sequence of machine-readable instructions, which can be embodied in a program or software. The instructions may be stored in a memory location, such as the memory 610. The instructions can be directed to the CPU 605, which can subsequently program or otherwise configure the CPU 605 to implement methods of the present disclosure. Examples of operations performed by the CPU 605 can include fetch, decode, execute, and writeback. The CPU 605 can be part of a circuit, such as an integrated circuit. One or more other components of the system 601 can be included in the circuit. In some cases, the circuit is an application specific integrated circuit (ASIC).
[0525]
[0320] The storage unit 615 can store files, such as drivers, libraries and saved programs. The storage unit 615 can store user data, e.g., user preferences and user programs. The computer system 601 in some cases can include one or more additional data storage units that are external to the computer system 601, such as located on a remote server that is in communication with the computer system 601 through an intranet or the Internet.
[0526]
[0321] The computer system 601 can communicate with one or more remote computer systems through the network 630. For instance, the computer system 601 can communicate with a remote computer system of a user. Examples of remote computer systems include personal computers (e.g., portable PC), slate or tablet PC’s (e g., Apple® iPad, Samsung® Galaxy Tab), telephones, Smart phones (e.g., Apple® iPhone, Android-enabled device, Blackberry®), or personal digital assistants. The user can access the computer system 601 via the network 630.
[0527]
[0322] Methods as described herein can be implemented by way of machine (e.g., computer processor) executable code stored on an electronic storage location of the computer system 601, such as, for example, on the memory 610 or electronic storage unit 615. The machine executable or machine-readable code can be provided in the form of software. During use, the code can be
[0528] -102-
[0529] SUBSTITUTE SHEET (RULE 26) executed by the processor 605. In some cases, the code can be retrieved from the storage unit 615 and stored on the memory 610 for ready access by the processor 605. In some situations, the electronic storage unit 615 can be precluded, and machine-executable instructions are stored on memory 610. The code can be pre-compiled and configured for use with a machine having a processor adapted to execute the code or can be compiled during runtime. The code can be supplied in a programming language that can be selected to enable the code to execute in a precompiled or as-compiled fashion.
[0530]
[0323] Aspects of the systems and methods provided herein, such as the computer system 601, can be embodied in programming. Various aspects of the technology may be thought of as “products” or “articles of manufacture” typically in the form of machine (or processor) executable code and / or associated data that is carried on or embodied in a type of machine readable medium. Machine-executable code can be stored on an electronic storage unit, such as memory (e.g., read-only memory, random-access memory, flash memory) or a hard disk. “Storage” type media can include any or all of the tangible memory of the computers, processors or the like, or associated modules thereof, such as various semiconductor memories, tape drives, disk drives and the like, which may provide non-transitory storage at any time for the software programming. All or portions of the software may at times be communicated through the Internet or various other telecommunication networks. Such communications, for example, may enable loading of the software from one computer or processor into another, for example, from a management server or host computer into the computer platform of an application server. Thus, another type of media that may bear the software elements includes optical, electrical, and electromagnetic waves, such as used across physical interfaces between local devices, through wired and optical landline networks and over various air-links. The physical elements that carry such waves, such as wired or wireless links, optical links or the like, also may be considered as media bearing the software. As used herein, unless restricted to non-transitory, tangible “storage” media, terms such as computer or machine “readable medium” refer to any medium that participates in providing instructions to a processor for execution.
[0531]
[0324] Hence, a machine readable medium, such as computer-executable code, may take many forms, including but not limited to, a tangible storage medium, a carrier wave medium or physical transmission medium. Non-volatile storage media include, for example, optical or magnetic disks, such as any of the storage devices in any computer(s) or the like, such as may be
[0532] -103-
[0533] SUBSTITUTE SHEET (RULE 26) used to implement the databases, etc. shown in the drawings. Volatile storage media include dynamic memory, such as main memory of such a computer platform. Tangible transmission media include coaxial cables; copper wire and fiber optics, including the wires that comprise a bus within a computer system. Carrier-wave transmission media may take the form of electric or electromagnetic signals, or acoustic or light waves such as those generated during radio frequency (RF) and infrared (IR) data communications. Common forms of computer-readable media therefore include for example: a floppy disk, a flexible disk, hard disk, magnetic tape, any other magnetic medium, a CD-ROM, DVD or DVD-ROM, any other optical medium, punch cards paper tape, any other physical storage medium with patterns of holes, a RAM, a ROM, a PROM and EPROM, a FLASH-EPROM, any other memory chip or cartridge, a carrier wave transporting data or instructions, cables or links transporting such a carrier wave, or any other medium from which a computer may read programming code and / or data. Many of these forms of computer readable media may be involved in carrying one or more sequences of one or more instructions to a processor for execution.
[0534]
[0325] The computer system 601 can include or be in communication with an electronic display 635 that comprises a user interface (UI) 640 for providing, for example, results of a nucleic acid sequence (e.g., sequence reads). Examples of UI’s include, without limitation, a graphical user interface (GUI) and web-based user interface.
[0535]
[0326] Methods and systems of the present disclosure can be implemented by way of one or more algorithms. An algorithm can be implemented by way of software upon execution by the central processing unit 605. The algorithm can, for example, perform error correction on processed sequencing signals.
[0536] EXAMPLES
[0537]
[0327] These examples are provided for illustrative purposes only and are not intended to limit the scope of the claims provided herein.
[0538] Example 1: Sequencing data for different pre-sequencing treatment protocols
[0539]
[0328] Sequencing quality metrics for sequencing runs performed with samples treated according to different pre-sequencing treatment protocols, as described elsewhere herein, were collected. The different pre-sequencing treatment protocols were the (1) Control, (2) dsDNA, (3) ssDNA, (4) ssDNA (2nd). In the Control protocol, beads comprising single-stranded template
[0540] -104-
[0541] SUBSTITUTE SHEET (RULE 26) molecules are pre-hybridized to sequencing primers, loaded to the substrate, and then sequenced. In the dsDNA protocol, beads comprising double-stranded template molecules are loaded to the substrate, second strands stripped from the template molecules and sequencing primers hybridized to the single-stranded template molecules while on the substrate, and the beads are then sequenced. In the ssDNA protocol, beads comprising single-stranded template molecules are pre-hybridized to sequencing primers, beads are loaded to the substrate, the sequencing primers are stripped from and rehybridized to the template molecules on the substrate, and the beads are then sequenced. In the ssDNA (2nd) protocol, beads comprising single-stranded template molecules are pre-hybridized to sequencing primers, beads are loaded to the substrate, the beads are sequenced for a 1sttime, the sequencing products (e.g., extended sequencing primers) are stripped and sequencing primers are rehybridized to the template molecules, and the beads are sequenced for a 2ndtime (the data shown pertains to this second run). The template molecules were sequenced using flow chemistry, as described elsewhere herein, using nonterminated, single-base flows. For each of the runs, there was comparable sequencing coverage.
[0329] Table 3 shows for each sequencing run, run type, # of beads (in millions), # of pass filter (PF) reads (in millions), PF% (% of # of PF reads in total # of beads), base error rate for 80% of the reads which are selected based on best RSQ (BER80), indel rate, lag rate, lead rate, droop rate, and misincorporation rate for the T-base.
[0542] # PF
[0543] „ „ , , . BER Indel Lag Lead Droop Misinc.
[0544] Run Type beads reads PF% ‘
[0545] 80 Rate Rate Rate Rate Rate T dsDNA 9889 7587 76.7% 0.30 0.36 0.05 0.04 0.24 026 dsDNA 9738 7605 78.1% 0.30 0.37 0.05 0.03 0.24 023 dsDNA 9195 7047 76.6% 0.28 0.35 0.06 0.04 0.24 025 dsDNA 10077 7636 75.8% 0.34 0.42 0.05 0.03 0.26 0.25 dsDNA 10282 7716 75.1% - 0.40 0.05 0.03 0.24 022 dsDNA 10316 7538 73.1% - 0.45 0.05 0.03 0.25 021 dsDNA 6572 4782 72.8% 0.26 0.36 0.05 0.04 0.24 028 dsDNA 9889 7587 76.7% 0.30 0.36 0.05 0.04 0.24 026 dsDNA 10432 7495 71.8% 0.38 - 0.06 0.04 0.24 0,22 dsDNA 10500 7640 72.8% 0.37 - 0.05 0.03 0.24 023
[0546] Averages 9689 7263 74.9% 0.31 0.38 0.05 0.03 0.24 0.24 ssDNA (2nd) 8085 5434 67.2% 0.30 0.36 0.05 0.03 0.22 021 ssDNA (2nd) 5661 3729 65.9% 0.32 0.37 0.05 0.03 0.23 021 ssDNA (2nd) 8621 5830 67.6% 0.40 0.45 0.05 0.03 0.27 027 ssDNA (2nd) 6833 4737 69.3% 0.53 0.57 0.05 0.03 0.27 0.28
[0547] -105-
[0548] SUBSTITUTE SHEET (RULE 26) ssDNA (2nd) 8169 5650 69.2% 0.49 0.55 0.06 0.04 0.26 0 30 ssDNA (2nd) 6747 4478 66.4% 0.41 0.47 0.06 0.03 0.25 028 ssDNA (2nd) 7166 4906 68.5% 0.55 0.57 0.06 0.04 0.28 0 31 ssDNA (2nd) 6380 4282 67.1% 0.43 0.47 0.05 0.04 0.27 0 30 ssDNA 9515 6535 68.7% 0.36 0.43 0.06 0.04 0.26 023 ssDNA 8263 5883 71.2% 0.40 0.45 0.06 0.05 0.27 024
[0549] Averages 7544 5146 68.1% 0.42 0.47 0.05 0.03 0.26 0.26
[0550] Control 9108 6008 66.0% 0.48 0.54 0.08 0.04 0.25 063
[0551] Control 9826 6685 68.0% 0.63 0.65 0.10 0.05 0.24 066
[0552] Control 8147 4793 58.8% 0.58 - 0.07 0.05 0.19 0.18
[0553] Control 10095 7694 76.2% 0.26 - 0.07 0.03 0.17 0 15
[0554] Control 7345 4795 65.3% 0.68 0.59 0.09 0.11 0.25 024
[0555] Control 6857 4939 72.0% 0.54 0.54 0.06 0.07 0.25 0 32
[0556] Averages 8563 5819 67.7% 0.53 0.58 0.08 0.06 0.23 0.36
[0557] Table 3 - Sequencing Quality Metrics
[0558]
[0330] As seen in Table 3, the # of beads, PF%, and droop rate (signal drooping over flows) measured for each of dsDNA, ssDNA, and ssDNA (2nd) treatment protocols were comparable with the Control treatment protocol. Unexpectedly, the sequencing runs with the dsDNA, ssDNA, and ssDNA(2nd) treatment protocols demonstrated a significant improvement in phasing quality such as shown by the lag rate and lead rate and in error quality such as shown by the BER80, indel rate, and misincorporation rate, compared to the Control treatment protocol. The average lag rate for each of the dsDNA and the two ssDNA treatment protocol runs was 0.05 and an unexpected improvement compared to the 0.08 average lag rate for the Control treatment protocol runs. The average lead rate for each of the dsDNA and the two ssDNA treatment protocol runs was 0 03 and an unexpected improvement compared to the 0.06 average lead rate for the Control treatment protocol runs. The average misincorporation rate for the dsDNA and the two ssDNA treatment protocol runs was 0.24 and 0.26, respectively, and an unexpected improvement compared to the 0 36 average misincorporation rate for the Control treatment protocol runs. The average BER80 for the dsDNA and the two ssDNA treatment protocol runs was 0.31 and 0.42, respectively, and an unexpected improvement compared to the 0.53 average BER80 for the Control treatment protocol runs. The average indel rate for the dsDNA and the two ssDNA treatment protocol runs was 0.38 and 0.47, respectively, and an unexpected improvement compared to the 0 58 average indel rate for the Control treatment protocol runs.
[0559]
[0331] Table 4 shows the total class error rate for three of the sequencing runs from Table 3, one sequencing run selected from each of the Control, dsDNA, and ssDNA treatment protocol
[0560] -106-
[0561] SUBSTITUTE SHEET (RULE 26) runs for comparison. The total class error rates are broken down into different flow bins to show the error rate progression as the number of sequencing flows progresses in each sequencing run. The total class error rates are indicative of the total error across all four bases called. It can be seen that the total class error rates are higher for the Control treatment protocol run in each flow bin compared to each of the dsDNA and ssDNA treatment protocol runs, demonstrating an unexpected improvement in error quality when sequencing runs follow the dsDNA or ssDNA treatment protocols compared to the Control treatment protocol.
[0562] Flows
[0563] Table 4 - Class Error Rates
[0564]
[0332] Table 5 shows the homopolymer base-calling error rate for the three sequencing runs compared in Table 4. The base-calling error rates are broken down into bins of different homopolymer lengths to show the error rate progression as the length of homopolymers called increases. It can be seen that for up to 8mers (homopolymers of up to 8 bases), the error rates are higher for the Control treatment protocol run compared to each of the dsDNA and ssDNA treatment protocol runs, demonstrating an unexpected improvement in homopolymer basecalling quality when sequencing runs follow the dsDNA or ssDNA treatment protocols. That is, for up to 8-mer homopolymer base-calling, following the dsDNA treatment protocol yielded best results, then the ssDNA treatment protocol, then the Control treatment protocol
[0565] H-mer length
[0566] Table 5 - Homopolymer Base-Calling Error Rates
[0567]
[0333] FIGs. 14A-14C illustrate a plot of mean signal vs. homopolymer length for the three sequencing runs compared in Tables 4 and 5, showing the plot in panels (FIG. 14A), (FIG.
[0568] 14B), and (FIG. 14C) for the Control, dsDNA, and ssDNA treatment protocols, respectively. It can be seen that the mean signals for homopolymers were higher for the dsDNA and ssDNA treatment protocol runs compared to the Control treatment protocol, which mav be the result of
[0569] -107-
[0570] SUBSTITUTE SHEET (RULE 26) improved homopolymer completion and / or an effect of lower phasing rates demonstrated in the former two treatment protocol runs, as described above.
[0571] EXEMPLARY EMBODIMENTS
[0572]
[0334] Exemplary embodiments of the methods and systems described herein include the following embodiments. The following embodiments recite non-limiting permutations of combinations of features disclosed herein. Other permutations of combinations of features are also contemplated. In particular, each of these numbered embodiments is contemplated as depending from or relating to every previous or subsequent numbered embodiment, independent of their order as listed.
[0573]
[0335] 1. A method of sequencing a template molecule, comprising: a. hybridizing a primer to the template molecule to form a hybridized template; b. determining the sequence of a first region of the template molecule by extending the primer through the first region using labeled nucleotides; c. extending the primer through a second region of the template using unlabeled nucleotides; d. determining the sequence of a third region of the template molecule by extending the primer through the third region using labeled nucleotides; e. determining the sequence of the second region of the template molecule by comparing the sequence of the first region and the third region to a reference sequence.
[0574]
[0336] 2. The method of embodiment 1, wherein the determining (b) the sequence of the first region comprises, for each nucleotide flow in a first plurality of nucleotide flows: a. providing labeled nucleotides of a respective base type; and b. detecting a signal indicating a presence or absence of incorporation of labeled nucleotides into the extending primer.
[0575]
[0337] 3. The method of embodiment 1 or embodiment 2, wherein the extending (c) the primer through the second region comprises, for each nucleotide flow in a second plurality of nucleotide flows: providing unlabeled nucleotides of a respective base type.
[0576]
[0338] 4. The method of embodiment 1 or embodiment 2, wherein the extending (c) the primer through the second region comprises, for each nucleotide flow in a second plurality of nucleotide flows: providing unlabeled nucleotides of one or more base types.
[0577] -108-
[0578] SUBSTITUTE SHEET (RULE 26)
[0339] 5. The method of any one of embodiments 1-4, wherein the determining (d) the sequence of the third region comprises, for each nucleotide flow in a third plurality of nucleotide flows: a. providing labeled nucleotides of a respective base type, and b. detecting a signal indicating a presence or absence of incorporation or labeled nucleotides into the extending primer.
[0579]
[0340] 6. The method of embodiment 1, wherein the determining (e) the sequence of the second region comprises aligning the first and third regions of the template molecule to a first and third regions of the reference sequence, respectively, wherein the first and third regions are separated from each other by a second region in the reference sequence.
[0580]
[0341] 7. The method of embodiment 6, wherein the sequence of the second region of the template molecule is determined as the sequence of the second region of the reference genome.
[0581]
[0342] 8. The method of embodiment 5, wherein the first plurality of nucleotide flows, and third plurality of nucleotide flows comprise 50 nucleotide flows, respectively.
[0582]
[0343] 9. The method of embodiment 8, wherein the second plurality of nucleotide flows comprises 10 nucleotide flows.
[0583]
[0344] 10. The method of embodiment 1, further comprising repeating the determining (b) and extending (c) one or more times prior to determining (d) and determining (e).
[0584]
[0345] 11. The method of embodiment 1, wherein the template molecule is a nucleic acid.
[0585]
[0346] 12. A method of sequencing a template molecule, comprising: a. hybridizing a primer to the template molecule to provide a hybridized template; and b. performing primer extension by providing: i) one or more nucleotide flows comprising labeled nucleotides of a first base type, and ii) one or more nucleotide flows.
[0586]
[0347] 13. The method of embodiment 12, wherein the primer extension (b) further comprises, in each of the one or more first nucleotide flows: a. providing labeled nucleotides of the first base type; and b. detecting a signal indicating a presence or absence of incorporation of a labeled nucleotide into the extending primer.
[0587]
[0348] 14. The method of embodiment 13, further comprising (c) determining the sequence of the template molecule by aligning the detected signals to a reference sequence.
[0588] -109-
[0589] SUBSTITUTE SHEET (RULE 26)
[0349] 15. The method of any one of embodiments 12-14, wherein the proportion of labeled nucleotides in the one or more first nucleotide flows is 100%.
[0590]
[0350] 16. The method of any one of embodiments 10-15, further comprising: a. removing the extended primer from the template ...
Claims
CLAIMSWHAT IS CLAIMED IS:
1. A method of sequencing a template molecule, comprising: a. hybridizing a primer to the template molecule to form a hybridized template; b. determining the sequence of a first region of the template molecule by extending the primer through the first region using labeled nucleotides; c. extending the primer through a second region of the template using unlabeled nucleotides; d. determining the sequence of a third region of the template molecule by extending the primer through the third region using labeled nucleotides; e. determining the sequence of the second region of the template molecule by comparing the sequence of the first region and the third region to a reference sequence.
2. The method of claim 1, wherein the determining (b) the sequence of the first region comprises^ for each nucleotide flow in a first plurality of nucleotide flows: a. providing labeled nucleotides of a respective base type; and b. detecting a signal indicating a presence or absence of incorporation of labeled nucleotides into the extending primer.
3. The method of claim 1 or claim 2, wherein the extending (c) the primer through the second region comprises, for each nucleotide flow in a second plurality of nucleotide flows: providing unlabeled nucleotides of a respective base type.
4. The method of claim 1 or claim 2, wherein the extending (c) the primer through the second region comprises, for each nucleotide flow in a second plurality of nucleotide flows: providing unlabeled nucleotides of one or more base types.
5. The method of any one of claims 1-4, wherein the determining (d) the sequence of the third region comprises, for each nucleotide flow in a third plurality of nucleotide flows: a. providing labeled nucleotides of a respective base type; and b. detecting a signal indicating a presence or absence of incorporation or labeled nucleotides into the extending primer.
6. The method of claim 1, wherein the determining (e) the sequence of the second region comprises aligning the first and third regions of the template molecule to a first and third-115-SUBSTITUTE SHEET (RULE 26)regions of the reference sequence, respectively, wherein the first and third regions are separated from each other by a second region in the reference sequence.
7. The method of claim 6, wherein the sequence of the second region of the template molecule is determined as the sequence of the second region of the reference genome.
8. The method of claim 5, wherein the first plurality of nucleotide flows, and third plurality of nucleotide flows comprise 50 nucleotide flows, respectively.
9. The method of claim 8, wherein the second plurality of nucleotide flows comprises 10 nucleotide flows.
10. The method of claim 1, further comprising repeating the determining (b) and extending (c) one or more times prior to determining (d) and determining (e).
11. The method of claim 1, wherein the template molecule is a nucleic acid.
12. A method of sequencing a template molecule, comprising: a. hybridizing a primer to the template molecule to provide a hybridized template; and b. performing primer extension by providing: i) one or more nucleotide flows comprising labeled nucleotides of a first base type, and ii) one or more nucleotide flows.
13. The method of claim 12, wherein the primer extension (b) further comprises, in each of the one or more first nucleotide flows: a. providing labeled nucleotides of the first base type; and b. detecting a signal indicating a presence or absence of incorporation of a labeled nucleotide into the extending primer.
14. The method of claim 13, further comprising (c) determining the sequence of the template molecule by aligning the detected signals to a reference sequence.
15. The method of any one of claims 12-14, wherein the proportion of labeled nucleotides in the one or more first nucleotide flows is 100%.
16. The method of any one of claims 10-15, further comprising: a. removing the extended primer from the template molecule; b. hybridizing a primer to the template molecule to provide a hybridized template;-116-SUBSTITUTE SHEET (RULE 26)c. performing primer extension in a second flow cycle, wherein the second flow cycle comprises one or more second nucleotide flows comprising nucleotides of a second base type different from the first base type; d. removing the extended primer from the template molecule; e. hybridizing a primer to the template molecule to provide a hybridized template; and f. performing primer extension in a third flow cycle, wherein the third flow cycle comprises one or more third nucleotide flows comprising nucleotides of a third base type different from the first and second base types.
17. The method of claim 16, wherein nucleotides in one or more second nucleotide flows comprise labeled nucleotides, and wherein the proportion of labeled nucleotides in the one or more second nucleotide flows is 100%.
18. The method of claim 16, wherein nucleotides in the one or more third nucleotide flows comprise labeled nucleotides, and wherein the proportion of labeled nucleotides in the one or more third nucleotide flows is 100%.
19. The method of any one of claims 16-18, wherein the first flow cycle, the second flow cycle, and the third flow cycle each comprises one or more additional nucleotide flows comprising unlabeled nucleotides.
20. The method of claim 19, wherein the one or more additional nucleotide flows in the first, second, and third flow cycles do not comprise the first base type, the second base type, or the third base type, respectively.21 . The method of claim 19, wherein the one or more additional nucleotide flows in the first, second, and third flow cycles comprise at least one nucleotide flow comprising labeled nucleotides, wherein the labeled nucleotides are not of the first base type, the second base type, or the third base type, respectively.
22. The method of any one of claims 16-21, further comprising: a. removing the extended primer from the template molecule; b. hybridizing a primer to the template molecule to provide a hybridized template; and-117-SUBSTITUTE SHEET (RULE 26)c. performing primer extension in a fourth flow cycle, wherein the fourth flow cycle comprises one or more fourth nucleotide flows comprising nucleotides of a fourth base type different from the first, second, and third base types.
23. The method of claim 12, wherein the template molecule is a nucleic acid.
24. A method of determining a sequence of a template molecule, comprising: a. hybridizing a primer to the template molecule to form a hybridized template; b. determining the sequence of a first region of the template molecule by extending the primer through the first region using labeled nucleotides; c. removing the extended primer from the template molecule; d. extending the primer through the first region of the template molecule using unlabeled nucleotides; and e. determining the sequence of a second region of the template molecule by extending the primer through the second region using labeled nucleotides.
25. The method of claim 24, wherein primer extension using labeled nucleotides further comprises^ in each of a plurality of nucleotide flows : providing labeled nucleotides of a respective base type; and detecting a signal indicating a presence or absence of incorporation of a labeled nucleotide into the extending primer.
26. The method of claim 24, wherein primer extension using unlabeled nucleotides further comprises, in each of a plurality of nucleotide flows: providing unlabeled nucleotides of at least one base type.
27. The method of claim 24, wherein the template molecule is a nucleic acid.
28. A method of sequencing a template molecule, comprising: a. providing a template molecule hybridized to a primer and a polymerase coupled thereto; b . providing a replacement polymerase; c. extending the primer through a first region of the template molecule until primer extension stalls; d. subjecting the template complex to conditions sufficient to anneal the replacement polymerase thereto; and e. extending the primer through a second region of the template molecule.-118-SUBSTITUTE SHEET (RULE 26)29. The method of claim 28, wherein primer extension stalling comprises decoupling of the polymerase from the template complex.
30. The method of claim 28, wherein prior to the subjecting (d) the replacement polymerase is coupled to a solid support via one or more cleavable adjacent to the template molecule.31 . The method of claim 30, wherein the subjecting (d) comprises cleaving the one or more cleavable moieties to decouple the replacement polymerase from the solid support.
32. The method of claim 28, further comprising: a. providing a plurality of replacement polymerases, wherein a first subset of the plurality of replacement polymerases is coupled to a first solid support and a second subset of the plurality of replacement polymerases is coupled to a second solid support.
33. The method of any one of claims 30-32, wherein the solid support is a bead.
34. The method of claim 28, wherein the template molecule is a nucleic acid.
35. A method for increasing sequencing read quality, comprising: a. receiving, at one or more processors, sequencing data comprising a plurality of sequencing reads generated by extending a sequencing primer through a region of interest in a target nucleic acid molecule using a plurality of sequencing flow steps, each sequencing flow step comprising combining a hybrid with nucleotides, the hybrid comprising the sequencing primer and a nucleic acid molecule comprising the region of interest, wherein at least a portion of the nucleotides are labeled, and detecting the presence or absence of an incorporated nucleotide; b. filtering the sequencing data, using the one or more processors, to remove sequencing reads for which an absence of an incorporated nucleotide was detected at six or more consecutive sequencing flow steps while retaining sequencing reads for which an absence of an incorporated nucleotide was detected at six or fewer consecutive sequencing flow steps, thereby generating filtered sequencing data; c. determining, using the one or more processors, for each sequencing flow step of each sequencing read, a read quality metric based on one or more homopolymer probability values other than a highest homopolymer probability value; and-119-SUBSTITUTE SHEET (RULE 26)d. trimming the terminus of one or more sequencing reads in the sequencing data based on the read quality metrics for a respective sequencing read, thereby generating trimmed sequencing data.
36. The method of claim 35, further comprising generating the sequencing data, wherein each sequencing read is obtained from a respective template nucleic acid molecule.
37. The method of claim 35, further comprising calling, using the one or more processors, one or more genetic variants using the trimmed sequencing data.
38. The method of claim 35, further comprising trimming a known adapter sequence, or a portion thereof, from one or more sequencing reads in the sequencing data.
39. The method of claim 35, wherein the read quality metric for each sequencing flow step of each sequencing read is based on a second highest homopolymer probability value.
40. The method of claim 35, wherein trimming the terminus of the one or more sequencing reads in the sequencing data based on the read quality metric, thereby generating the trimmed sequencing data, comprises, for each sequencing read: a. determining a read quality metric moving average for the sequencing flow steps; b. selecting a sequencing flow step, wherein the selected sequencing flow step is the nth sequencing flow step having a moving average above a predetermined threshold, wherein n is a predefined number; and c. trimming at least a portion of the sequencing read comprising the selected sequencing flow step.
41. The method of claim 35, wherein a predetermined number of consecutive sequencing flow steps prior to the selected sequencing flow step are trimmed.
42. The method of claim 35, wherein the predetermined number of consecutive sequencing flow steps is a multiple of four.
43. The method of claim 35, further comprising storing the trimmed sequencing data in a non-transitory computer readable medium.
44. The method of claim 35, further comprising aligning sequencing reads in the trimmed sequencing data to respective reference sequences.
45. The method of claim 35, wherein the nucleotides are non-terminating nucleotides.-120-SUBSTITUTE SHEET (RULE 26)