Paired-end sequencing

JP2025509660A5Pending Publication Date: 2026-03-26ILLUMINA INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-03-15
Publication Date
2026-03-26

AI Technical Summary

Technical Problem

The existing next-generation sequencing technology requires sequencing of forward and reverse strand sequences one by one, resulting in a long time to determine the full sequence of double-stranded molecules and consumes a lot of chemical reagents.

Method used

By the presence of both forward and reverse strands in the same nucleic acid cluster, forward and reverse strand sequences of double strand templates are identified and sequenced using the same sequencing run using different primers and labeled nucleotides.

Benefits of technology

The forward and reverse strand sequences of double-strand templates are achieved simultaneously in the same sequencing operation, which improves sequencing efficiency on the flow cell area, reduces the use of chemical reagents, and simplifies the sequencing workflow.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

A system and method for identifying a nucleobase in a template polynucleotide is disclosed. In an embodiment, such a method may include providing a substrate comprising a plurality of double-stranded template polynucleotides in a cluster. Each double-stranded template polynucleotide may include a first strand and a second strand. The method may further include contacting the plurality of double-stranded template polynucleotides with a first primer that binds to the first strand and a second primer that binds to the second strand. The method may further include extending the first primer and the second primer by contacting the cluster with a labeled nucleobase to form a first labeled primer and a second labeled primer. The method may further include stimulating light emission from the first and second labeled primers, where the amplitude of the signal generated by the first labeled primer is greater than the amplitude of the signal generated by the second labeled primer. The method may further include identifying the labeled nucleobases added to the first and second primers based on the amplitude of the signal generated by the labeled nucleobases.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of this disclosure relates to the field of nucleic acid sequencing. More specifically, the disclosed technology relates to using next-generation sequencing to determine the nucleotide sequence of either the forward or reverse strand of a double-stranded target polynucleotide in a single sequencing run. [Background technology]

[0002] 2. Description of Related Art In some types of next-generation sequencing (NGS) technologies, nucleic acid clusters are created on a flow cell by amplifying the original template nucleic acid strand to form a double-stranded bridge. One of the forward or reverse strands is then selectively removed, and the remaining template strand is generally single-stranded, linearized, and attached to the flow cell at only one end. Next-generation sequencing cycles can be performed as the complementary strand of the remaining template nucleic acid is synthesized, i.e., using a sequencing-by-synthesis (SBS) process.

[0003] In each sequencing cycle, a deoxyribonucleic acid analog conjugated with a fluorescent label is hybridized to a template nucleic acid, and an excitation light source is used to excite the fluorescent label on the deoxyribonucleic acid analog. A detector captures the fluorescent emission from the fluorescent label and identifies the deoxyribonucleic acid analog. As a result, the sequence of the template nucleic acid can be determined by repeatedly performing such sequencing cycles.

[0004] After a certain number of cycles (e.g., after sequencing about 500 bases), the newly synthesized complementary strand can be removed, leaving the template nucleic acid. Further amplification can be performed on the template nucleic acid to form a double-stranded bridge (i.e., reclustering), followed by selective removal of the other of the forward and reverse strands (e.g., removal of the forward strand), and sequencing of the remaining (linearized) nucleic acid using an SBS process, thus achieving paired-end sequencing. Because sequencing of the forward strand and sequencing of the reverse strand are performed sequentially, it can take a relatively long time to fully sequence both strands of a double-stranded molecule using current next-generation sequencing technologies. Summary of the Invention

[0005] In one embodiment, the disclosed technology provides systems and methods for determining the nucleobase sequence of both the forward and reverse strands of a double-stranded template polynucleotide in parallel (i.e., substantially simultaneously using the same sequencing run) while the forward and reverse strands coexist within the same nucleic acid cluster. Thus, in some embodiments, the sequencing yield per flow cell area can be doubled compared to conventional systems in which sequencing was performed sequentially.

[0006] In one aspect, the disclosed technology provides a system and method for identifying nucleobases in a template polynucleotide. In one embodiment, the disclosed method can include providing a substrate containing a plurality of double-stranded template polynucleotides in a cluster, each double-stranded template polynucleotide comprising a first strand and a second strand. The disclosed method can further include contacting the plurality of double-stranded template polynucleotides with a first primer that binds to the first strand and a second primer that binds to the second strand. The disclosed method can further include extending the first primer and the second primer by contacting the cluster with labeled nucleobases to form a first labeled primer and a second labeled primer. The disclosed method can further include stimulating light emission from the first and second labeled primers, wherein the amplitude of the signal generated by the first labeled primer is greater than the amplitude of the signal generated by the second labeled primer (or vice versa). The disclosed methods can further include identifying the labeled nucleobases added to the first primer and the second primer based on the amplitude of the signal generated by the labeled nucleobase. In some embodiments, the first primer is an index primer that hybridizes to a site adjacent to the barcode index portion of the first strand. In some embodiments, the second primer is an index primer that hybridizes to a site adjacent to the barcode index portion of the second strand. In some embodiments, the first primer is an index primer that hybridizes to a site adjacent to the barcode index portion of the first strand, and the second primer is an index primer that hybridizes to a site adjacent to the barcode index portion of the second strand.

[0007] In some embodiments, identifying the labeled nucleobases added to the first primer and identifying the labeled nucleobases added to the second primer are performed substantially simultaneously. In some embodiments, the signal generated by the first labeled primer and the signal generated by the second labeled primer are emitted from the same or substantially overlapping regions of the substrate. In some embodiments, either the first strand or the second strand of each double-stranded template polynucleotide is bound to the substrate. In some embodiments, the multiple double-stranded template polynucleotides in the cluster are generated by a bridge amplification process, an exclusion amplification process, a rolling circle amplification process, or any other suitable amplification process. In some embodiments, the substrate comprises multiple clusters of nucleic acids, and the clusters are randomly distributed on the substrate. In alternative embodiments, the clusters are arranged in a patterned array.

[0008] In some embodiments, the amplitude of the signal generated by the first labeled primer corresponds to the first amount of the first labeled primer in the cluster, and the amplitude of the signal generated by the second labeled primer corresponds to the second amount of the second labeled primer in the cluster. In some embodiments, contacting a plurality of double-stranded template polynucleotides with a first primer that binds to the first strand and a second primer that binds to the second strand comprises contacting the first strand with an unblocked first primer and contacting the second strand with a predetermined fraction of second primers having blocked 3' ends. The blocked 3' end can be formed in any manner that blocks the primer's ability to extend a nucleic acid strand, for example, by a modification in the sugar or nucleobase. In some embodiments, the blocked 3' end comprises a hairpin loop, a deoxynucleotide, a phosphate group, a propyl spacer, a modification that blocks the 3'-hydroxyl group, or an inverted nucleobase. In some embodiments, the first primer is formed from a locked nucleic acid or a peptide nucleic acid. In some embodiments, the second primer is formed from a locked nucleic acid or a peptide nucleic acid.

[0009] In some embodiments, the disclosed methods further include contacting the plurality of double-stranded template polynucleotides with a RecA-like protein or a non-nicking CRISPR-associated protein to facilitate binding of the plurality of double-stranded template polynucleotides to the first primer and the second primer. In some embodiments, extending the first primer and the second primer is catalyzed by a strand-displacing polymerase. In some embodiments, the strand-displacing polymerase comprises Klenow fragment, phi29 DNA polymerase, Bsm DNA polymerase, Bst DNA polymerase, conservative mutations thereof, or engineered forms thereof, e.g., mutations, fusions, truncations, etc. In some embodiments, the disclosed methods further include contacting the plurality of double-stranded template polynucleotides with a helicase, a single-stranded DNA-binding protein, or a mixture of oligonucleotides having random sequences to partially separate the first strand and the second strand of each double-stranded template polynucleotide.

[0010] In some embodiments, the disclosed methods further include detecting a signal generated by the first labeled primer in a first range of optical frequencies and a second range of optical frequencies, and detecting a signal generated by the second labeled primer in the first range of optical frequencies and a second range of optical frequencies, where the first range of optical frequencies and the second range of optical frequencies are not identical. For example, the first range of optical frequencies may correspond to red, e.g., 400-484 THz (or equivalently, 620-750 nm in wavelength), and the second range of optical frequencies may correspond to green, e.g., 526-606 THz (or equivalently, 495-570 nm in wavelength).

[0011] In some embodiments, the disclosed method further includes acquiring a first fluorescent image of the cluster at a first range of optical frequencies and acquiring a second fluorescent image of the cluster at a second range of optical frequencies, where the first range of optical frequencies and the second range of optical frequencies are not identical, and obtaining signals generated by the first and second labeled primers by extracting fluorescence intensities from the first and second fluorescent images of the cluster. In some examples, the first range of optical frequencies and the second range of optical frequencies may partially overlap. For example, the first range of optical frequencies may be 500-580 THz, and the second range of optical frequencies may be 540-620 THz.

[0012] In some embodiments, the disclosed method further includes extracting fluorescence intensities from the first and second fluorescence images of the same or substantially overlapping regions of the substrate. In some embodiments, identifying the labeled nucleobases attached to the first and second primers is based on a combination of the extracted fluorescence intensities from the first and second fluorescence images. In some embodiments, the combination of the identities of the labeled nucleobases attached to the first and second primers is classified as one of 16 combinations of nucleobase types based on the extracted fluorescence intensity combination and a predetermined fluorescence intensity distribution for the 16 combinations of nucleobase types. In some embodiments, the disclosed method further includes normalizing the extracted fluorescence intensities and classifying the combination of the identities of the labeled nucleobases attached to the first and second primers as one of 16 combinations of nucleobase types based on the normalized extracted fluorescence intensity combination and a predetermined normalized fluorescence intensity distribution for the 16 combinations of nucleobase types.

[0013] In some embodiments, the disclosed methods further comprise stimulating fluorescence emission from the first labeled primer and the second labeled primer in the cluster with light at a predetermined light frequency. In some embodiments, the disclosed methods further comprise stimulating fluorescence emission from the first labeled primer and the second labeled primer in the cluster with light at two predetermined light frequencies. In some embodiments, the disclosed methods further comprise identifying whether the labeled nucleobase is associated with the first strand or the second strand based on the amplitude of the signal generated by the labeled nucleobase.

[0014] In another aspect, the disclosed technology provides systems and methods for determining the sequence of a template polynucleotide. In one embodiment, the disclosed method can include hybridizing a first primer to the template polynucleotide and a second primer to the reverse complement of the template polynucleotide, where the template polynucleotide and the reverse complement of the template polynucleotide are in a substantially overlapping region of the substrate. The disclosed method can further include extending the first primer with a first labeled nucleotide analog. The disclosed method can further include extending the second primer with a second labeled nucleotide analog. The disclosed method can further include stimulating emission of light from the first and second labeled nucleotide analogs. The disclosed method can further include determining the sequence of the nucleotide in the template polynucleotide and the reverse complement of the template polynucleotide by capturing the emitted light. In some embodiments, the emission from the first and second labeled nucleotide analogs is captured substantially simultaneously. In some embodiments, the first primer is an index primer that hybridizes to a site adjacent to the barcode index portion of the template polynucleotide. In some embodiments, the second primer is an index primer that hybridizes to a site adjacent to the barcode index portion of the reverse complement of the template polynucleotide. In some embodiments, the first primer is an index primer that hybridizes to a site adjacent to the barcode index portion of the template polynucleotide, and the second primer is an index primer that hybridizes to a site adjacent to the barcode index portion of the reverse complement of the template polynucleotide.

[0015] In some embodiments, the template polynucleotide and the reverse complement of the template polynucleotide are part of a cluster of identical copies of the template polynucleotide and identical copies of the reverse complement of the template polynucleotide. In some embodiments, the cluster of identical copies of the template polynucleotide and identical copies of the reverse complement of the template polynucleotide are generated by bridge amplification, an exclusion amplification process, a rolling circle amplification process, or any other suitable amplification process. In some embodiments, the identical copies of the template polynucleotide have ends attached to the substrate by a first grafting oligonucleotide. In some embodiments, the identical copies of the reverse complement of the template polynucleotide have ends attached to the substrate by a second grafting oligonucleotide. In some embodiments, at least a portion of the reverse complement of the template polynucleotide hybridizes to a portion of the template polynucleotide. In some embodiments, the first primer is part of a first population of first primers hybridized to identical copies of the template polynucleotide, and the second primer is part of a second population of second primers hybridized to identical copies of the reverse complement of the template polynucleotide.

[0016] In some embodiments, determining the sequence of the nucleotide comprises receiving a first signal emitted at a first amplitude from a first population of first primers, receiving a second signal emitted at a second amplitude from a second population of second primers, and identifying nucleobases hybridized to the template polynucleotide and nucleobases hybridized to the reverse complement of the template polynucleotide based on a combination of the first and second signals. In some embodiments, the first and second signals are received simultaneously or substantially simultaneously, or as a combined signal.

[0017] In some embodiments, a fraction of the second population of second primers have blocked 3' ends. The blocked 3' ends can be formed in any manner that blocks the primer's ability to extend, for example, by modifications to the sugar or nucleobase. In some embodiments, the blocked 3' ends include a hairpin loop, a deoxynucleotide, a phosphate group, a propyl spacer, a modification that blocks the 3'-hydroxyl group, or an inverted nucleobase. In some embodiments, the first population of first primers have unblocked 3' ends. In some embodiments, the first primer and the second primer hybridize to the template polynucleotide and the reverse complement of the template polynucleotide, respectively, in the same reaction step. In some embodiments, extending the first primer with a first labeled nucleotide analog and extending the second primer with a second labeled nucleotide analog are performed in the same reaction step.

[0018] In some embodiments, the first labeled nucleotide analog and the second labeled nucleotide analog hybridize to the template polynucleotide and the reverse complement of the template polynucleotide, respectively, in the same reaction step. In some embodiments, the first primer and / or the second primer comprise locked nucleic acid (LNA) or peptide nucleic acid (PNA). In some embodiments, hybridization of the first primer to the template polynucleotide and hybridization of the second primer to the reverse complement of the template polynucleotide are facilitated by the presence of a RecA-like protein or a non-nicking CRISPR-associated protein. In some embodiments, extending the first primer and extending the second primer are catalyzed by a strand-displacing polymerase. In some embodiments, the strand-displacing polymerase comprises Klenow fragment, phi29 DNA polymerase, Bsm DNA polymerase, Bst DNA polymerase, conservative mutations thereof, or engineered forms thereof, e.g., mutations, fusions, truncations, etc. In some embodiments, the template polynucleotide and the reverse complement of the template polynucleotide are at least partially separated by the presence of a helicase, a single-stranded DNA binding protein, or a mixture of oligonucleotides having random sequences.

[0019] The systems, devices, kits, and methods disclosed herein each have several aspects, no single one of which is solely responsible for their desirable attributes. Numerous other embodiments are contemplated, including embodiments having fewer, additional, and / or different components, steps, features, objects, benefits, and advantages. The components, aspects, and steps may be arranged and ordered differently. After considering this discussion, and particularly after reading the section entitled "Detailed Description of the Invention," one will understand how the features of the devices and methods disclosed herein provide advantages over other known devices and methods.

[0020] It should be understood that any features of the systems disclosed herein can be combined in any desired manner and / or configuration. It should also be understood that any features of the methods disclosed herein can be combined in any desired manner. It should also be understood that any combination of method and / or system features can be used together and / or combined with any of the embodiments disclosed herein. It should also be understood that all combinations of the foregoing concepts and additional concepts discussed in more detail below are considered part of the inventive subject matter disclosed herein and can be used to achieve the benefits and advantages described herein. [Brief explanation of the drawings]

[0021] Features of examples of the present disclosure will become apparent by reference to the following detailed description and drawings, in which like reference numbers correspond to similar, if not identical, components, and for the sake of brevity, reference numbers or features having a previously mentioned function may or may not be described with reference to other drawings in which they appear. [Figure 1] FIG. 1 shows a block diagram that generally illustrates an exemplary sequencing system that can be used to practice the disclosed methods. [Figure 2] FIG. 2 shows a block diagram that schematically illustrates an exemplary imaging system that may be used with the exemplary sequencing system of FIG. 1. [Figure 3] FIG. 2 shows a functional block diagram of an exemplary computer system that may be used in the exemplary sequencing system of FIG. 1. [Figure 4] FIG. 1 shows a schematic diagram of a double-stranded DNA bridge invaded by a primer, according to one embodiment of the disclosed technology. [Figure 5A] 5 is a microscope image showing fluorescence data associated with the embodiment of FIG. 4. [Figure 5B] 5 is a microscope image showing fluorescence data associated with the embodiment of FIG. 4. [Figure 5C]5 is a microscope image showing fluorescence data associated with the embodiment of FIG. 4. [Figure 5D] 5 is a microscope image showing fluorescence data associated with the embodiment of FIG. 4. [Figure 6A] 1A-B illustrate a schematic of simultaneous sequencing of both the forward and reverse strands of a double-stranded polynucleotide template within the same cluster, according to one embodiment of the disclosed technology. [Figure 6B] 1A-B illustrate a schematic of simultaneous sequencing of both the forward and reverse strands of a double-stranded polynucleotide template within the same cluster, according to one embodiment of the disclosed technology. [Figure 7] 6C is a chart showing an exemplary dye labeling scheme that may be used with the embodiment of FIG. 6B. [Figure 8] 1 is a plot illustrating 16 distributions of signals from nucleic acid clusters, according to one embodiment of the disclosed technology. [Figure 9] FIG. 1 is a flow diagram illustrating a method for sequencing a polynucleotide according to one embodiment of the disclosed technology. DETAILED DESCRIPTION OF THE INVENTION

[0022] All patents, patent applications, and other publications, including all sequences disclosed therein and referred to herein, are expressly incorporated by reference herein to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. All cited documents are, in relevant part, incorporated herein by reference in their entirety, for the purposes indicated by the context of the citation herein. However, the citation of any document should not be construed as an admission that it is prior art to the present disclosure.

[0023] Introduction In one aspect, the disclosed technology provides systems and methods that can dramatically shorten total sequencing time and reduce the number of reagents used in next-generation sequencing workflows. Additionally, sequencing yield per flow cell area can be increased and consumable complexity can be reduced (e.g., the use of special PCR primers with chemical modifications designed for nucleic acid strand linearization can be avoided). In some embodiments, the disclosed methods enable simultaneous paired-end sequencing without the need for cluster linearization with orthogonal chemistry, cluster regeneration, and / or surface patterning, thus simplifying reaction processes and device design and increasing the efficiency of sequencing workflows.

[0024] In some embodiments, a primer for sequencing the forward strand and a primer for sequencing the reverse strand of a double-stranded DNA template are annealed / hybridized to each strand of the template in the same reaction step, reducing chemical reaction steps and thus saving time and increasing the efficiency of the sequencing-by-synthesis (SBS) workflow. For example, while the template is in the form of a non-linear dsDNA bridge attached to a flow cell, primers can simultaneously hybridize to the two template strands. Both the forward and reverse strand sequences can then be read out in the same reaction run via SBS chemical cycles.

[0025] In some embodiments, to separate the signals received from dye-labeled nucleobases hybridized to the forward and reverse strands within the same cluster, the signal from one of the strands is reduced, for example, by 50% compared to the signal generated by the other strand. This reduction in signal intensity can be achieved by blocking the addition of labeled nucleobases to some of the primers. For example, half of the primers that bind to the reverse strand can be blocked, so that no fluorescent nucleotides are added during the sequencing reaction. Thus, in this example, the overall intensity of the nucleobases added to the reverse strand is 50% lower than the intensity of the nucleobases added to the forward strand. By reviewing not only the wavelength of light emitted from the dye from each nucleic acid cluster on the flow cell, but also the intensity of this light, labeled nucleobases hybridized to the forward strand can be distinguished from labeled nucleobases hybridized to the reverse strand. This is described more fully in the following section.

[0026] In some embodiments, the disclosed technology involves obtaining sequence information using Illumina sequencing-by-synthesis and reversible terminator-based sequencing chemistry with removable fluorescent dyes (e.g., as described in Bentley et al., Nature 6:53-59

[2009] ). Short sequence reads of about tens to hundreds of base pairs can be aligned to a reference genome, and unique mappings of short sequence reads to the reference genome can be identified. Further details regarding sequencing-by-synthesis and dye-labeling methods that can be used with the disclosed technology can be found in U.S. Patent Application Publication Nos. 2007 / 0166705, 2006 / 0188901, 2006 / 0240439, 2006 / 0281109, 2005 / 0100900, 2013 / 0079232, U.S. Patent No. 7,057,026, and International Publication No. WO 2009 / 024090. Nos. 5 / 065814, 2006 / 064199, 2007 / 010251, and 2018 / 165099, as well as U.S. Patent Application No. 17 / 338590, U.S. Patent No. 7,601,499, U.S. Patent No. 9,267,173, and U.S. Patent Publication No. 2012 / 0053063, the disclosures of which are incorporated herein by reference in their entireties.

[0027] Example Sequencer Referring to FIG. 1 , a schematic diagram of an exemplary sequencing system 10 is shown including a sequencer 12 designed to sequence the genetic material of a sample 14. The sequencer can function in a variety of ways and can be based on a variety of techniques, including sequencing by primer extension using labeled nucleotides, as in the presently contemplated embodiments, as well as other sequencing techniques (e.g., sequencing by ligation or pyrosequencing). In some embodiments, the sequencer 12 progressively moves the sample through reaction and imaging cycles to progressively build oligonucleotides by binding nucleotides to templates at individual sites on the sample. In some embodiments, the sample may be prepared by a sample preparation system 16. This process may include amplification of DNA or RNA fragments on a support to generate multiple sites of DNA or RNA fragments that are sequenced by the sequencing process.Exemplary methods for generating amplified nucleic acid sites suitable for sequencing include rolling circle amplification (RCA) (Lizardi et al., Nat. Genet. 19:225-232 (1998)), bridge PCR (Adams and Kron, Method for Performing Amplification of Nucleic Acid with Two Primers Bound to a Single Solid Support, Mosaic Technologies, Inc. (Winter Hill, Mass.), Whitehead Institute for Biomedical Research, Cambridge, Mass., (1997); Adessi et al., Nucl. Acids Res. 28:E87 (2000); Pemov et al., Nucl. Acids Res. 33:e11 (2005); or U.S. Pat. No. 5,641,658), polony generation (Mitra et al., Proc. Natl. Acad. Sci. USA 100:5926-5931 (2003), Mitra et al., Anal. Biochem. 320:55-65 (2003)), or clonal amplification on beads using emulsion (Dressman et al., Proc. Natl. Acad. Sci. USA 100:8817-8822 (2003)) or ligation to bead-based adapter libraries (Brenner et al., Nat. Biotechnol. 18:630-634 (2000), Brenner et al., Proc. Natl. Acad. Sci. USA 97:1665-1670 (2000)), Reinartz et al., Brief Funct. Genomic Proteomic 1:95-104 (2002), the foregoing publications are incorporated by reference. The sample preparation system 16 may place a sample in a sample container for processing and imaging, and the sample may be in the form of an array of sites.

[0028] In some embodiments, the sequencer 12 includes a fluid control / delivery system 18 and a detection system 20. The fluid control / delivery system 18 may accept multiple process fluids, such as those indicated by reference numeral 22, for circulation of the sample through a sample container, designated by reference numeral 24, during processing. As will be understood by those skilled in the art, the process fluids may vary depending on the particular stage of sequencing. For example, in sequencing-by-synthesis (SBS) using labeled nucleotides, the process fluid introduced to the sample may include a polymerase and tagged nucleotides of four common DNA types, each nucleotide having a unique fluorescent tag and a blocking agent linked to it. The fluorescent tag allows the detection system 20 to detect which nucleotide was last added to a primer hybridized to a template nucleic acid at each site in the array, and the blocking agent prevents the addition of more than one nucleotide per cycle at each site.

[0029] At other stages of the sequencing cycle, the process fluid 22 may contain other fluids and reagents (e.g., reagents for removing extension blocks from nucleotides or cleaving nucleotide linkers to release newly extendable primer ends). For example, as reactions occur at individual sites in the sample array, the initial process fluid containing the tagged nucleotides may be washed from the sample with one or more flushing operations. The sample may then undergo detection, such as by optical imaging in the detection system 20. Reagents may then be added by the fluid control / delivery system 18 to deblock the last-added nucleotides and remove the fluorescent tag from each. The fluid control / delivery system 18 may then wash the sample again, and the sample is then prepared for a subsequent cycle of sequencing. Exemplary fluid and detection configurations that can be used in the methods and devices described herein are described in International Publication No. 07 / 123744, which is incorporated herein by reference. In some embodiments, such sequencing may continue until the quality of the data derived from the sequencing deteriorates due to a cumulative loss of yield, or until a predetermined number of cycles have been completed.

[0030] In some embodiments, the quality of the sample 24 during processing, as well as the quality of the data derived by the system, and the various parameters used to process the sample, are controlled by a quality / process control system 26. The quality / process control system 26 may include one or more programmed processors or general-purpose or application-specific computers that communicate with sensors and other processing systems in the fluid control / delivery system 18 and detection system 20. Some process parameters may be used for advanced quality and process control, for example, as part of a feedback loop that can change instrument operating parameters during the course of a sequencing run.

[0031] In some embodiments, the sequencer 12 also communicates with a system control / operator interface 28 and ultimately with a post-processing system 30. The system control / operator interface 28 can include a general-purpose or application-specific computer designed to monitor process parameters, acquired data, system settings, and the like. The operator interface can be generated by a program running locally or by a program running within the sequencer 12. In some embodiments, they can provide visual indications of the health of the sequencer's systems or subsystems, the quality of acquired data, and the like. The system control / operator interface 28 can also allow a human operator to interface with the system to adjust operations, start and pause sequencing, and perform any other desired interactions with the system hardware or software. For example, the system control / operator interface 28 can automatically perform and / or modify steps performed in a sequencing procedure without input from a human operator. Alternatively or additionally, the system control / operator interface 28 can generate recommendations regarding steps performed in a sequencing procedure and display these recommendations to a human operator. This mode can allow input from a human operator before performing and / or modifying steps in a sequencing procedure. Additionally, system control / operator interface 28 may provide the human operator with options that allow the human operator to select certain steps in the sequencing procedure to be performed automatically by sequencer 12, while requiring input from the human operator before other steps are performed and / or modified. In any event, enabling both an automated mode and an operator-interactive mode may provide increased flexibility in performing the sequencing procedure. Additionally, the combination of automation and human-controlled interaction may further enable a system capable of creating and modifying new sequencing procedures and algorithms through adaptive machine learning based on input collected from the human operator.

[0032] The post-processing system 30 may further include one or more programmed computers that receive the detected information, which may be in the form of pixelated image data, and derive sequence data from the image data. The post-processing system 30 may include image recognition algorithms that distinguish the colors (e.g., the fluorescent emission spectra) of the dyes attached to the nucleotides bound at each site as sequencing progresses (e.g., by analyzing the image data to encode specific colors and / or intensities) and record the sequence of the nucleotide at each site. The post-processing system 30 can then progressively build a sequence list for each site of the sample array, which can be further processed by various bioinformatics algorithms to establish genetic information for the extended length of material.

[0033] The sequencing system 10 can be configured to handle individual samples or can be designed for higher throughput in a manner where multiple stations are provided for delivery of reagents and other fluids and for detection to progressively build up a sequence of nucleotides. Further details can be found in U.S. Patent No. 9,797,012, which is incorporated herein by reference.

[0034] In particular, if the fluid control system 18 or the quality / process control system 26 detects that one or more operations did not perform optimally or in a desired manner, the sample may be removed from processing and reprocessed, and the schedule of such processing may be altered in real time. In embodiments where a sample is removed from a process or experiences a pause in processing of substantial duration, the sample may be placed in storage. Placing the sample in storage may include altering the sample's environment or composition to stabilize biomolecular reagents, biopolymers, or other components of the sample. Exemplary methods for altering the sample environment include, but are not limited to, lowering the temperature to stabilize sample components, adding an inert gas to reduce oxidation of sample components, and removing the sample from a light source to reduce photobleaching or photodegradation of sample components. Exemplary methods for altering sample composition include, but are not limited to, adding antioxidants, stabilizing solvents such as glycerol, changing the pH to a level that stabilizes enzymes, or removing components that degrade or alter other components. Furthermore, certain steps in the sequencing procedure may be performed before removing the sample from processing. For example, if it is determined that the sample should be removed from processing, the sample may be directed to the fluid control / delivery system 18 so that the sample can be washed before storage. These steps may be performed to ensure that information from the sample is not lost.

[0035] Furthermore, sequencing operations can be interrupted by the sequencer 12 whenever certain predetermined events occur. These events include, but are not limited to, unacceptable environmental factors such as undesirable temperature, humidity, vibration, stray light, etc.; insufficient reagent delivery or hybridization; unacceptable changes in sample temperature; unacceptable sample site number / quality / distribution; attenuated signal-to-noise ratio; insufficient image data; and the like. It should be noted that the occurrence of such events does not require an interruption of the sequencing operation. Rather, such events can be a factor weighed by the quality / process control system 26 in determining whether the sequencing operation should continue. For example, if images of a particular cycle are analyzed in real time and show a low signal for that optical channel, the image can be re-exposed using a longer exposure time, or a particular chemical treatment can be repeated. If the image shows an air bubble in the flow cell, the instrument can automatically flush more reagents to remove the bubble and then re-record the image. If the image shows low signal for a particular optical channel in one cycle due to fluidics issues, the instrument can automatically stop scanning and reagent delivery for that particular optical channel, thus saving analysis time and reagent consumption.

[0036] While the present system has been illustrated above with respect to a system in which the sample interacts with different stations through physical movement of the sample, it will be understood that the principles described herein are also applicable to systems in which the steps occurring at each station are accomplished by other means that do not require movement of the sample. For example, reagents present at the stations can be delivered to the sample by a fluidic system connected to reservoirs containing various reagents. Similarly, an optical system can be configured to detect a sample in fluid communication with one or more reagent stations. Thus, a detection step can occur before, during, or after delivery of any particular reagent described herein. Thus, a sample can be effectively removed from processing by ceasing one or more processing steps, be it fluid delivery or optical detection, without necessarily physically removing the sample from its location within the device.

[0037] The disclosed system can be used to sequentially sequence nucleic acids in multiple different samples. The disclosed system can be configured to include a sample arrangement and a station arrangement for performing sequencing steps. Samples in the sample arrangement can be arranged in a fixed order and at fixed intervals relative to each other. For example, the nucleic acid array arrangement can be arranged along the outer edge of a circular table. Similarly, the stations can be arranged in a fixed order and at fixed intervals relative to each other. For example, the stations can be arranged in a circular arrangement having a perimeter corresponding to the layout for the sample array arrangement. Each of the stations can be configured to perform a different operation in a sequencing protocol. The sample array and station arrangement can be moved relative to each other so that the stations perform the desired steps of the reaction scheme at each reaction site. The relative positions and relative movement schedule of the stations can be correlated with the order and duration of the reaction steps in the sequencing reaction scheme, such that a single sequencing reaction cycle is completed when the sample array completes a cycle of interaction with the full set of stations. For example, primers hybridized to nucleic acid targets on an array can be extended, detected, and deblocked by the addition of a single nucleotide, respectively, when the order of stations, the spacing between stations, and the speed of passage through the array correspond to the order of reagent delivery and reaction time of a complete sequencing reaction cycle.

[0038] According to the above configuration, each lap (or complete rotation in embodiments where a circular table is used) completed by an individual sample array can correspond to the determination of a single nucleotide for each target nucleic acid on the array (e.g., including the uptake, imaging, cleavage, and deblocking steps performed in each cycle of a sequencing run). Furthermore, several sample arrays present in the system (e.g., on the circular table) simultaneously move along similar repeating laps through the system, thereby resulting in continuous sequencing by the system. Using the disclosed systems or methods, reagents can be actively delivered or removed from a first sample array according to the first reaction step of a sequencing cycle, while incubation, or some other reaction step within the cycle, occurs for a second sample array. Thus, a set of stations can be configured in a spatial and temporal relationship with the arrangement of sample arrays to allow reactions to occur simultaneously on multiple sample arrays, thereby enabling continuous and simultaneous sequencing, even when the sample arrays are subjected to different steps of a sequencing cycle at any given time. Such a circular system may be used when chemistry and imaging times are unbalanced. For small flow cells that require only a short scan time, a system can have several flow cells running in parallel to optimize the time the instrument spends acquiring data. If imaging time and chemistry time are equal, a system sequencing a sample on a single flow cell will spend half the time performing the chemistry cycle rather than the imaging cycle; thus, a system capable of processing two flow cells can have one on the chemistry cycle and one on the imaging cycle. If imaging time is one-tenth the chemistry processing time, a system can have ten flow cells in various stages of chemistry processing while continuously acquiring data.

[0039] In some embodiments, the disclosed system is configured to allow a first sample array to be replaced with a second sample array while the system is continuously sequencing nucleic acids on a third sample array. Thus, a first sample array can be individually added to or removed from the system without interrupting the sequencing reaction occurring on another sample array, thereby enabling continuous sequencing of a set of sample arrays. Furthermore, sequencing runs of different lengths can be performed consecutively and simultaneously in the system, as individual sample arrays can complete different numbers of laps through the system, and sample arrays can be added or removed from the system in an independent manner so that reactions occurring at other sites are not disrupted.

[0040] FIG. 2 illustrates an exemplary detection station 38 capable of detecting nucleotides added to sites on an array and that can be used in conjunction with the exemplary sequencing system of FIG. 1 . As discussed above, a sample can be moved to two or more stations on the device located at different physical locations, or one or more steps can be performed on a sample in communication with one or more stations without necessarily being moved to different locations. Therefore, a description herein of a particular station should be understood to refer to stations in various configurations, regardless of whether the sample moves between stations, whether the station moves to the sample, or whether the station and sample are stationary relative to each other. In the embodiment shown in FIG. 2 , one or more light sources 46 provide a light beam that is directed to conditioning optics 48. The light source 46 may include one or more lasers, with multiple lasers being used to detect dyes that fluoresce at different corresponding wavelengths. The light source can direct a beam to conditioning optics 48 for beam filtering and shaping in the conditioning optics. For example, in embodiments contemplated herein, conditioning optics 48 combines beams from multiple lasers to generate a substantially linear radiation beam that is transmitted to focusing optics 50. The laser modules may further include measurement components that record the power output of each laser. The power measurement can be used as a feedback mechanism to control the length of time an image is recorded to obtain a uniform exposure energy, and therefore signal, for each image. If the measurement components detect a laser module failure, the instrument may flush the sample into a "holding buffer" to store the sample until errors in the lasers can be corrected.

[0041] The sample 24 is positioned on a sample positioning system 52, which can appropriately position the sample in three dimensions and move the sample for progressive imaging of sites on the sample array. In embodiments contemplated herein, focusing optics 50 confocally directs radiation to one or more surfaces of the array where individual sites to be sequenced are located. Depending on the wavelength of light in the focused beam, a retrobeam of radiation is returned from the sample due to fluorescence of dyes attached to nucleotides at each site.

[0042] The retrobeam is then returned through retrobeam optics 54, which can filter the beam, such as by separating different wavelengths within the beam and directing these separated beams to one or more cameras 56. The cameras 56 can be based on any suitable technology, such as a charge-coupled device that generates pixelated image data based on photons impinging on a location within the device. In some embodiments, the cameras 56 can include a CMOS sensor. In some embodiments, the cameras 56 can include one or more autofocus cameras. In some embodiments, the cameras 56 can include one or more time-delay-integration (TDI) cameras. The cameras generate image data, which is then transferred to image processing circuitry 58. In some embodiments, the processing circuitry 58 can perform various operations, such as analog-to-digital conversion, scaling, filtering, and correlating the data in multiple frames, to properly and accurately image multiple sites at specific locations on the sample. The image processing circuitry 58 can store the image data and ultimately transfer the image data to a post-processing system 30, where array data can be derived from the image data. Exemplary detection devices that can be used in the detection station include those described, for example, in U.S. Patent Application Publication No. 2007 / 0114362 (U.S. Patent Application No. 11 / 286,309) and WO 07 / 123744, each of which is incorporated herein by reference.

[0043] A computer system 106, such as that shown in Figure 3, can be used to implement the system control / operator interface 28 and post-processing system 30 of the exemplary sequencing system 10 of Figure 1. As shown in Figure 3, the computer system 106 can include functionality for controlling the optical / fluidics system and for determining the nucleic acid base sequence of a polynucleotide.

[0044] In one embodiment, the computer system 106 includes a processor 202 in electrical communication with a memory 204, a storage device 206, and a communication interface 208. The processor 202 can be configured to execute instructions to cause the fluidics system 104 to deliver reagents to the flow cell 114 during a sequencing reaction. The processor 202 can execute instructions to control the light source 120 of the optical system 102 to generate light near a predetermined wavelength. The processor 202 can execute instructions to control the detector 126 of the optical system 102 and receive data from the detector 126. The processor 202 can execute instructions to process data received from the detector 126, such as a fluorescent image, and to determine the nucleotide sequence of a polynucleotide based on the data received from the detector 126. The memory 204 can be configured to store instructions for configuring the processor 202 to perform the functions of the computer system 106 when the sequencing system 100 is powered on. The storage device 206 can store instructions for configuring the processor 202 to perform the functions of the computer system 106 when the sequencing system 100 is powered off. The communication interface 208 may be configured to facilitate communication between the computer system 106 , the optical system 102 , and the fluid system 104 .

[0045] Computer system 106 may include a user interface 210 configured to communicate with a display device (not shown) for displaying sequencing results of sequencing system 100. User interface 210 may be configured to receive input from a user of sequencing system 100. Optical system interface 212 and fluidic system interface 214 of computer system 106 may be configured to control optical system 102 and fluidic system 104 through communication links 108a and 108b illustrated in FIG. 1A. For example, optical system interface 212 may communicate with computer interface 110 of optical system 102 via communication link 108a.

[0046] The computer system 106 may include a nucleic acid base determiner 216 configured to determine the nucleotide sequence of the polynucleotide using data received from the detector 126. The nucleic acid base determiner 216 may include one or more of a template generator 218, a position register 220, an intensity extractor 222, an intensity corrector 224, a base caller 226, and a quality score determiner 228. The template generator 218 may be configured to generate a template of the position of the polynucleotide cluster within the flow cell 114 using the fluorescent image captured by the detector 126. The position register 220 may be configured to register the position of the polynucleotide cluster within the flow cell 114 within the fluorescent image captured by the detector 126 based on the position template generated by the template generator 218. The intensity extractor 222 may be configured to extract the intensity of the fluorescent emission from the fluorescent image to generate extracted intensities. For example, peak intensity values ​​found at diffraction-limited spots of DNA clusters may be extracted from the image and used to represent the signal of the DNA cluster. In another example, the total intensity contained within the diffraction-limited spot of the DNA cluster can be extracted from the image and used to represent the signal of the DNA cluster. Alternatively, intensity estimation can be performed through the use of equalization and channel estimation.

[0047] The intensity corrector 224 can be configured to reduce or eliminate noise or aberrations inherent in the sequencing reaction or optical system. For example, the intensity can be affected by laser intensity fluctuations, DNA cluster shape / size variations, uneven illumination, optical distortions or aberrations, and / or phasing / pre-phasing occurring within the DNA cluster. In some embodiments, the intensity corrector 224 can phase-correct or pre-phase-correct the extracted intensities. In some embodiments, the intensity corrector 224 can normalize the extracted fluorescence intensities to reduce or eliminate the effects of DNA cluster size variations. For example, each DNA template can contain the same calibration oligonucleotide. Thus, the extracted fluorescence intensity of a cluster obtained from sequencing known nucleotides in the calibration oligonucleotide can be used as a normalization factor for that cluster. The intensity corrector 224 can divide the extracted fluorescence intensity of that cluster obtained from sequencing nucleotides in other regions of the DNA template by the normalization factor to obtain a normalized extracted fluorescence intensity. The base corrector 226 can be configured to determine the nucleobase of the polynucleotide from the corrected intensities. The bases of the polynucleotide determined by the base caller 226 can be associated with a quality score determined by the quality score determiner 228. Quality scoring refers to the process of assigning a quality score to each base call. To assess the quality of a base call from a sequencing read, an exemplary process may include calculating a set of predictor values ​​for the base call and using the predictor values ​​to look up a quality score in a quality table. The quality scores may be presented in any suitable format that allows a user to determine the probability of error for any given base call. In some embodiments, the quality score is presented as a numeric value. For example, a quality score may be shown as QXX, where XX is the score and indicates the probability that this particular call is 10 -XX / 10This means that the base sequence has a probability of error of 0.01%. Thus, by way of example, Q30 is equivalent to an error rate of 1 in 1000, or 0.1%, and Q40 is equivalent to an error rate of 1 in 10,000, or 0.01%. The error rate can be calculated using a control nucleic acid. Furthermore, some metric displays can include the error rate per cycle. In some embodiments, the quality table is generated using a calibration data set, where the calibration set represents run and sequence variability. Further details of the calculations that can be performed by the nucleobase determiner, calculation of error rate, and quality score can be found in U.S. Patent No. 8,392,126, U.S. Patent Application Publication No. 2020 / 0080142, and U.S. Patent Application Publication No. 2012 / 0020537, each of which is incorporated herein by reference in its entirety.

[0048] Sequencing without cluster linearization or reclustering Figure 4 shows how primers can invade a double-stranded molecule bound to a substrate at both ends, and nucleotides can still be added during an NGS run. This example was the first step in demonstrating that two primers can invade a double-stranded molecule and have effective performance in an NGS sequencing run. As shown, Figure 4 is a schematic diagram of a double-stranded DNA bridge on a solid support. The double-stranded DNA includes a first strand 401 and a second strand 402 that are complementary to each other. Strands 401 and 402 of the double-stranded DNA have both ends attached to a solid support 430 via graft sequence oligonucleotides 437 and 438, thus forming a double-stranded DNA bridge on the solid support. As shown in Figure 4, a "Lead 1" sequencing primer 410, which is complementary to a portion of the first strand 401, can invade the double-stranded bridge and bind to one strand. Nucleotide analogs can be added to "lead 1" primer 410 via a polymerase reaction to form an extended primer containing fluorescent label 455 for fluorescent imaging. Thus, although strand 401 is in double-stranded bridge form and has not been linearized, the nucleobase sequence of strand 401 can be determined by the SBS process. Similarly, strand 402 can be sequenced in double-stranded bridge form without linearization. As described below in connection with FIG. 6B, in some embodiments, strand 401 and strand 402 can be sequenced simultaneously.

[0049] In some embodiments of the disclosed sequencing methods, strand 401 and / or strand 402 do not need to be linearized, i.e., chemically cleaved; therefore, special PCR primers with chemical modifications (e.g., a P5 primer with deoxyuridine (dU) as the cleavage site or a P7 primer with 8-oxo-guanine (8-oxoG) nucleotide as the cleavage site; see U.S. Patent Application Publication No. 2019 / 0352327, which is incorporated herein by reference) are not required in the bridge PCR process. Furthermore, in some embodiments, because strand 401 and / or strand 402 do not need to be linearized, the use of a linearization reagent mixture can be avoided. Also, in some embodiments, if strand 401 and strand 402 can be sequenced simultaneously, post-linearization cluster regeneration (reclustering) and the associated use of a relinearization reagent mixture to regenerate one ssDNA template after reading the other template can be avoided. Thus, the disclosed technology can provide a more cost-effective sequencing modality by saving on various chemical reagents.

[0050] To enable DNA strands to be sequenced in a bridged form without linearization, certain biochemical processes (e.g., those occurring in cells during DNA recombination or DNA synthesis) can be used to enable hybridization of sequencing primers to the strand and subsequent primer extension reactions on the non-linear strand. In some embodiments, helicases can be used to catalyze the processive unwinding of double-stranded DNA. In some embodiments, non-nicking CRISPR-associated proteins can be used to facilitate primer binding to the strand. In some embodiments, recA-like proteins (e.g., rec 233, rec T.th, etc.) can be used to coat sequencing primers, allowing them to efficiently invade double-stranded DNA bridges. In some embodiments, polymerases with strand displacement activity (e.g., Phi29, Bsm, Bst, Bsu, Klenow large fragment, etc., or conservative mutations thereof) can be used to extend primers during the SBS process. In some embodiments, the strands may be stabilized in a local single-stranded form by using single-stranded DNA binding proteins or libraries of short random DNA oligos to reduce dsDNA reannealing. In some embodiments, the sequencing primers may be formed from locked nucleic acids (LNA) or peptide nucleic acids (PNA) that can bind strongly to DNA strands 401 and / or 402.

[0051] 5A-5D are fluorescence microscopy images showing proof-of-principle data related to the embodiment of FIG. 4. The images were taken on a flow cell with multiple randomly seeded clusters containing unlinearized double-stranded DNA bridges after the first SBS cycle to incorporate dye-labeled nucleotide analogs into the clusters. In particular, an Illumina MiSeq™ system was used. In the system, four different fluorescent color channels correspond to four different bases. FIG. 5A is an image taken in "Channel T," FIG. 5B is an image taken in "Channel C," FIG. 5C is an image taken in "Channel A," and FIG. 5D is an image taken in "Channel G." A RecA-like protein was used to facilitate primer invasion of the dsDNA bridge. An Illumina polymerase from the MiSeq™ system was used to incorporate the dye-labeled nucleotide analogs into the sequencing primer. Successful strand-displacing SBS reactions are evident from brighter spots (compared to background), which also represent fluorophore blinking events observed in the fluorescence detector, e.g., spots 501, 502, 503, 504, 505, 506, 507, 508, 509, 510, 511, and 512. The brighter spots are diffraction-limited images of clusters that have successfully incorporated fluorescent dye-labeled nucleotide analogs. Illumina polymerases can be further engineered to fully enable and optimize strand-displacing activity.

[0052] 6A and 6B illustrate sequencing of both the forward and reverse strands of a double-stranded polynucleotide template within the same cluster and in the same sequencing run. As shown in FIG. 6A, a cluster of clonal copies of a double-stranded polynucleotide bridge can be formed on a solid support 630, for example, by bridge PCR amplification from a double-stranded polynucleotide template. The double-stranded polynucleotide template can include a first strand 601 (e.g., the forward strand) and a second strand 602 (e.g., the reverse strand) that are complementary to each other. Multiple copies of strands 601 and 602 can have both ends attached to the solid support 630, thus forming a double-stranded polynucleotide bridge on the solid support, similar to the schematic diagram shown in FIG. 4.

[0053] 6B shows that both the forward and reverse strands within a cluster of double-stranded polynucleotide bridges can be simultaneously sequenced using a primer specific to the first strand 601 and a primer specific to the second strand 602 in the same reaction run. Because the fluorescent signal associated with the extended first strand sequencing (lead 1) primer and the fluorescent signal associated with the extended second strand sequencing (lead 2) primer are emitted from fluorescent labels that coexist within the same cluster, the signals may not be optically resolved. Therefore, a method for determining whether a fluorescent signal is associated with the extended first strand sequencing primer or the extended second strand sequencing primer is needed to accurately determine the nucleic acid sequence of both the first and second strands, at least when the dye-labeled nucleotide analog in the extended first strand sequencing primer is not the same as the dye-labeled nucleotide analog in the extended second strand sequencing primer (e.g., when an "A" is added in the first strand 601 and a "C" is added in the second strand 602).

[0054] In some embodiments, whether a fluorescent signal is associated with the first strand or the second strand can be determined by using distinguishable levels of signal intensity. In one example, a mixture of non-extendable (e.g., terminated or blocked) and extendable versions of a primer can be provided and used to sequence one of the strands (e.g., second strand 602), while the primer used to sequence the other of the strands (e.g., first strand 601) contains only the extendable version. As shown in FIG. 6B, both an extendable Read 2 primer 6200 and a non-extendable Read 2 primer 6206 are used to bind to the second strand 602, and all molecules of the Read 1 primer 610 used to bind to the first strand 601 are extendable. For example, a predetermined fraction of the Read 2 primer molecules (e.g., one-quarter, one-third, half, two-thirds, etc., or any value therebetween) can be chemically blocked and therefore non-extendable by a polymerase.

[0055] In some embodiments, the extendable Read 2 primer 6200, the non-extendable Read 2 primer 6206, and the Read 1 primer 610 may be provided to the flow cell as a mixture. The primers may hybridize to strands within a cluster, and excess primers in the fluid may be washed away. Under some conditions, for example, when both the Read 1 and Read 2 primers are provided to a cluster at saturating concentrations, and when the number of copies of the first strand 601 is approximately the same as the number of copies of the second strand 602 within the cluster, the number of extendable Read 2 primers 6200 still bound to the cluster may be smaller than the number of (extendable) Read 1 primers still bound to the cluster. As a result, after the SBS process, the number of extended Read 2 primers within a cluster may be smaller than the number of extended Read 1 primers. Therefore, the number of dye-labeled nucleotides associated with the second strand 602 within a cluster may be smaller than the number of dye-labeled nucleotides associated with the first strand 601. In a regime where signal intensity correlates with the number of dyes in a cluster, after receiving fluorescent excitation, the signal emitted from a labeled nucleotide associated with the second strand 602 within the cluster may be weaker compared to the signal emitted from a labeled nucleotide associated with the first strand 601. Thus, the weaker signal and the nucleotide identity it represents may be determined as associated with the second strand 602.

[0056] In some embodiments, the non-extendable lead 2 primer 6206 can have a blocked 3' end. In some instances, the blocked 3' end can include a hairpin loop, a modification that blocks the 3'-hydroxyl group, or a phosphate group.

[0057] In another example, the blocked 3' end may contain a dideoxynucleotide, such as dideoxycytidine (ddC), exemplified below, which functions as a 3' chain terminator that prevents 3' extension by a DNA polymerase.

[0058] [ka]

[0059] In yet another example, the blocked 3' end can comprise an inverted nucleobase, which when incorporated at the 3' end of an oligo can result in a 3'-3' linkage that inhibits both degradation by 3' exonucleases and extension by DNA polymerases. As an example, a 3' inverted dT is illustrated below.

[0060] [ka]

[0061] In yet another example, the blocked 3' end can include a C3 propyl spacer, as exemplified below: The spacer C3 incorporated at the 3' end of the oligo can function as an effective blocking agent for polymerase extension reactions.

[0062] [ka]

[0063] While the above example may show that some of the Lead 2 primers are blocked, reducing the intensity of the fluorescent signal received from labeled nucleobases hybridized to the second strand, it should be understood that any mechanism for distinguishing the intensity of fluorescent labels attached to either the first strand or the second strand is contemplated to be within the scope of the present invention. For example, in an alternative embodiment, the Lead 2 primers are unblocked, but a proportion or percentage of the Lead 1 primers contain blocking groups such that labeled nucleobases are not attached to these Lead 1 primers.

[0064] FIG. 7 illustrates an exemplary dye-labeling scheme that can be used with the embodiment of the disclosed technology shown in FIG. 6B. As shown in FIG. 7, different types of nucleotide analogs can be labeled with different fluorescent labels / dyes having different absorption and / or emission spectra. For example, dGTP is unlabeled, dATP is labeled with a first label / dye, dCTP is labeled with a second label / dye, and dTTP is labeled with a third label / dye. Thus, different types of nucleotide analogs can generate different fluorescent emission characteristics after excitation by a light source. Alternative dye-labeling schemes can be contemplated. For example, dTTP can be labeled with two different dyes, or a mixture of two single-labeled dTTPs can be used. In some embodiments, the absorption spectrum of the dye allows it to be excited by a single light source of a predetermined wavelength, such as a "blue" laser at approximately 450 nm. However, embodiments are not limited to light sources generating light of this particular wavelength; other wavelengths corresponding to red, green, violet, or other available wavelengths of light are contemplated. In other embodiments, two or more light sources can be used to excite the dyes if their absorption spectra are sufficiently different.

[0065] A first fluorescent label / dye may have an emission spectrum that can be captured in a first image taken in a first light channel (e.g., "Image 1" in FIG. 7). A second fluorescent label / dye may have an emission spectrum that can be captured in a second image taken in a second light channel different from the first light channel (e.g., "Image 2" in FIG. 7). A third fluorescent label / dye may have an emission spectrum that can be captured in both the first and second light channels (e.g., both "Image 1" and "Image 2" in FIG. 7). As a result, in the example shown in FIG. 7, dTTP may be identified as appearing at a sufficiently high intensity in both the first and second images. dGTP may appear as having zero or very low (e.g., below a cutoff value) intensity in both images. dATP may be identified as appearing at a sufficiently high intensity in the first image but at a very low (e.g., below a cutoff value) intensity in the second image. dCTP may be identified as appearing at a sufficiently high intensity in the second image but at a very low intensity (e.g., below a cutoff value) in the first image. The fluorescent dyes conjugated to the four types of nucleotide analogs are exemplary only and are not intended to be limiting. In other embodiments, the nucleotide analog not conjugated to any fluorescent dye may be dTTP, dCTP, or dATP. In other embodiments, the nucleotide analog conjugated to the first fluorescent dye may be dGTP, dCTP, or dTTP. In other embodiments, the nucleotide analog conjugated to the second fluorescent dye may be dGTP, dTTP, or dATP. In other embodiments, the nucleotide analog conjugated to the third fluorescent dye may be dGTP, dATP, or dCTP. In other embodiments, the fluorescent dyes may be bound during the secondary color generation step by a set of fluorescently labeled proteins (e.g., antibodies) that specifically bind to different nucleotide bases directly or to ligands / adapters linked to nucleotides.For example, fluorescently labeled streptavidin may be used to recognize and bind to biotin adaptors linked to nucleotides, or fluorescently labeled anti-digoxigenin may be used to recognize and bind to digoxigenin adaptors linked to nucleotides.

[0066] In some embodiments, the nucleotide analogs used in the disclosed sequencing systems can be fully functionalized nucleotides. The linker located between the nucleotide base and the fluorescent molecule can contain one or more cleavable groups. Prior to a subsequent sequencing cycle, the fluorescent label can be removed from the nucleotide analog by cleavage of the linker. For example, the linker attaching the fluorescent label to the nucleotide analog can contain an azide and / or alkoxy group, e.g., on the same carbon, so that the linker can be cleaved after each incorporation cycle with a phosphine reagent, thereby releasing the fluorescent label. The nucleotide triphosphate can be reversibly blocked at the 3' position to allow sequencing control, and only a single nucleotide analog can be added to each extendible primer-polynucleotide in each cycle. For example, the 3' ribose position of the nucleotide analog can contain both an alkoxy and an azide functional group that can be removed by cleavage with a phosphine reagent, thereby generating a nucleotide that can be further extended. Prior to a subsequent sequencing cycle, the reversible 3' block can be removed and another nucleotide analog can be added to each extendible primer-polynucleotide.

[0067] In some embodiments, the fluorescent label is selected from the group consisting of polymethine derivatives, coumarin derivatives, benzopyran derivatives, chromenoquinoline derivatives, and compounds containing bis-boron heterocycles, such as BOPPY and BOPYPY. In some embodiments, the fluorescent label is attached to the nucleotide via a cleavable linker. In some further embodiments, the labeled nucleotide may have a fluorescent label attached to the C5 position of a pyrimidine base or to the C7 position of a 7-deazapurine base, optionally via a linker moiety. For example, the nucleobase may be 7-deazaadenine, and a dye attached to the 7-deazaadenine at the C7 position, optionally via a cleavable linker. The nucleobase may be 7-deazaguanine, and a dye attached to the 7-deazaguanine at the C7 position, optionally via a cleavable linker. The nucleobase may be cytosine, and a dye attached to the cytosine at the C5 position, optionally via a cleavable linker. As another example, the nucleobase can be thymine or uracil, and the dye is attached to the thymine or uracil at the C5 position, optionally through a cleavable linker. In some further embodiments, the cleavable linker can contain a chemical moiety similar to or the same as the reversible terminator 3' hydroxy blocking group, such that the 3' hydroxy blocking group and the cleavable linker can be removed under the same reaction conditions or in a single chemical reaction. Non-limiting examples of cleavable linkers include the LN3 linker, the sPA linker, and the AOL linker, each of which is exemplified below.

[0068] [ka]

[0069] [ka]

[0070] In some embodiments, the nucleotides are selected from the group consisting of a dGTP analog, a dTTP analog, a dUTP analog, a dCTP analog, and a dATP analog. In some embodiments, the first nucleotide is a first reversibly blocked nucleotide triphosphate (rbNTP), the second nucleotide is a second rbNTP, the third nucleotide is a third rbNTP, and the fourth nucleotide is a fourth rbNTP, and each of the first nucleotide, the second nucleotide, the third nucleotide, and the fourth nucleotide is a different type of nucleotide from each other. In some embodiments, the four rbNTPs are selected from the group consisting of rbATP, rbTTP, rbUTP, rbCTP, and rbGTP. In some embodiments, each of the four rbNTPs comprises a modified base and a reversible terminator 3'-blocking group. Non-limiting examples of 3'-blocking groups include azidomethyl ( * -CH2N3), substituted azidomethyl (e.g., * -CH(CHF2)N3 or * -CH(CH2F)N3) and * -CH2-O-CH2-CH=CH2, where an asterisk * indicates a point attachment to the 3' oxygen of the ribose or deoxyribose ring of the nucleotide.

[0071] Further details regarding dyes and fully functionalized nucleotides can be found in U.S. Patent Application Publication Nos. 2018 / 0094140 and 2020 / 0277670, International Patent Application Publication No. 2017 / 051201, and U.S. Provisional Patent Application Nos. 63 / 057758 and 63 / 127061, the disclosures of which are incorporated herein by reference in their entireties.

[0072] FIG. 8 is a scatter plot showing 16 example distributions of signals from a nucleic acid cluster according to an embodiment of the disclosed technology shown in FIG. 6B, which in one example may be implemented with the dye-labeling scheme shown in FIG. 7. As described in connection with FIG. 6B, in one embodiment, within the same cluster, the fluorescent signal coming from a set of extended Read 1 primers 610 is brighter than the fluorescent signal coming from a set of extended unblocked Read 2 primers 6200. The scatter plot in FIG. 8 shows 16 distributions (or bins) of intensity values ​​from the combination of the brighter and dimmer signals of the cluster; the two signals may be co-localized and may not be optically resolved. The intensity values ​​shown in FIG. 8 may be scaled or normalized, and the units of the intensity values ​​may be arbitrary or relative (i.e., represent the ratio of actual intensity to reference intensity). The sum of the brighter signal from the extended Read 1 primers 610 and the dimmer signal from the extended unblocked Read 2 primers 6200 results in a combined signal. The combined signal can be captured by a first optical channel and a second optical channel (e.g., the "Image 1" channel and the "Image 2" channel in FIG. 7). Because the brighter signal can be A, T, C, or G and the dimmer signal can be A, T, C, or G, there are 16 possibilities for the combined signal, which correspond to 16 distinguishable patterns when optically captured according to the embodiment described in connection with FIG. 7. That is, each of the 16 possibilities corresponds to a bin shown in FIG. 8. The computer system can map the combined signal from the cluster to one of the 16 bins and thus determine the added nucleobase in extended Read 1 primer 610 and the added nucleobase in extended unblocked Read 2 primer 6200, respectively.

[0073] For example, when the combined signal is mapped to bin 812 for a base calling cycle, the computer processor base calls both the added nucleobase in the extended Read 1 primer 610 and the added nucleobase in the extended unblocked Read 2 primer 6200 as a C. When the combined signal is mapped to bin 814 for a base calling cycle, the processor base calls the added nucleobase in the extended Read 1 primer 610 as a C and the added nucleobase in the extended unblocked Read 2 primer 6200 as a T. When the combined signal is mapped to bin 816 for a base calling cycle, the processor base calls the added nucleobase in the extended Read 1 primer 610 as a C and the added nucleobase in the extended unblocked Read 2 primer 6200 as a G. When the combined signal is mapped to bin 818 for the base calling cycle, the processor base calls the added nucleobase in the extended Read 1 primer 610 as a C and the added nucleobase in the extended unblocked Read 2 primer 6200 as an A.

[0074] When the combined signal is mapped to bin 822 for a base calling cycle, the processor base calls the added nucleobase in the extended Read 1 primer 610 as T and the added nucleobase in the extended unblocked Read 2 primer 6200 as C. When the combined signal is mapped to bin 824 for a base calling cycle, the processor base calls both the added nucleobase in the extended Read 1 primer 610 and the added nucleobase in the extended unblocked Read 2 primer 6200 as T. When the combined signal is mapped to bin 826 for a base calling cycle, the processor base calls the added nucleobase in the extended Read 1 primer 610 as T and the added nucleobase in the extended unblocked Read 2 primer 6200 as G. When the combined signal is mapped to bin 828 for the base calling cycle, the processor base calls the added nucleobase in the extended Read 1 primer 610 as T and the added nucleobase in the extended unblocked Read 2 primer 6200 as A.

[0075] When the combined signal is mapped to bin 832 for a base calling cycle, the processor base calls the added nucleobase in the extended Read 1 primer 610 as G and the added nucleobase in the extended unblocked Read 2 primer 6200 as C. When the combined signal is mapped to bin 834 for a base calling cycle, the processor base calls the added nucleobase in the extended Read 1 primer 610 as G and the added nucleobase in the extended unblocked Read 2 primer 6200 as T. When the combined signal is mapped to bin 836 for a base calling cycle, the processor base calls both the added nucleobase in the extended Read 1 primer 610 and the added nucleobase in the extended unblocked Read 2 primer 6200 as G. When the combined signal is mapped to bin 838 for the base calling cycle, the processor base calls the added nucleobase in the extended Read 1 primer 610 as G and the added nucleobase in the extended unblocked Read 2 primer 6200 as A.

[0076] When the combined signal is mapped to bin 842 for a base calling cycle, the processor base calls the added nucleobase in the extended Read 1 primer 610 as A and the added nucleobase in the extended unblocked Read 2 primer 6200 as C. When the combined signal is mapped to bin 844 for a base calling cycle, the processor base calls the added nucleobase in the extended Read 1 primer 610 as A and the added nucleobase in the extended unblocked Read 2 primer 6200 as T. When the combined signal is mapped to bin 846 for a base calling cycle, the processor base calls the added nucleobase in the extended Read 1 primer 610 as A and the added nucleobase in the extended unblocked Read 2 primer 6200 as G. When the combined signal is mapped to bin 848 for the base calling cycle, the processor base calls both the added nucleobase in the extended Read 1 primer 610 and the added nucleobase in the extended unblocked Read 2 primer 6200 as A. Further details regarding performing base calling based on a scatter plot with 16 bins can be found in U.S. Patent Application Publication No. 2019 / 0212294, the disclosure of which is incorporated herein by reference.

[0077] Simplified sequencing workflow 9 is a flow diagram illustrating a method 900 for sequencing polynucleotides that can employ an embodiment of the disclosed technology according to FIG. 6B. The described method allows for simultaneous sequencing of both the forward and reverse strands of a template dsDNA without the need for cluster linearization or regeneration, thus requiring less sequencing reagent consumption, allowing for faster generation of data from both strands. Furthermore, the simplified method can reduce the number of workflow steps while providing the same yield compared to existing next-generation sequencing methods. Thus, the simplified method can result in reduced sequencing run times.

[0078] As shown in FIG. 9 , the disclosed method 900 may begin at block 901. The method may then proceed to block 910, where default oligografting is performed, which may include attaching oligonucleotide anchor / graft sequences to the planar, optically transparent surface of a flow cell. The method may then proceed to block 920, where a DNA library is generated from the sample, where template polynucleotides in the sample may be end-repaired to generate 5′-phosphorylated blunt ends, and the polymerase activity of the Klenow fragment may be used to add a single A base to the 3′ ends of the blunt-phosphorylated nucleic acid fragments. This addition prepares the nucleic acid fragments for ligation to oligonucleotide adaptors, which have a single T-base overhang at their 3′ ends to increase ligation efficiency. The adaptor oligonucleotides are complementary to the flow cell anchor oligos.

[0079] After DNA library generation, the method can then proceed to block 930, where the double-stranded DNA library is denatured to generate single-stranded template polynucleotides for seeding onto the flow cell. The method can then proceed to block 940, where clustering is performed from the single-stranded template polynucleotides. Under limiting dilution conditions, adapter-modified single-stranded template polynucleotides are added to the flow cell and immobilized by hybridization to the anchor oligo. The attached nucleic acid fragments are extended and the bridges are amplified to create an ultra-high-density sequencing flow cell with hundreds of millions of clusters, each containing approximately 1,000 copies of the same template. Details regarding enrichment of nucleic acids using cluster amplification can be found in Kozarewa et al., Nature Methods 6:291-295 (2009), which is incorporated herein by reference.

[0080] After cluster generation, without performing cluster linearization, the method can proceed directly to block 950, where the Read 1 primer and the Read 2 primer are simultaneously hybridized / annealed to both the forward and reverse strands of the dsDNA bridge of the nucleic acid cluster on the flow cell. The method can then proceed to block 960, where both the forward and reverse strands of the dsDNA bridge are simultaneously sequenced. Sequencing proceeds by extending the Read 1 primer and the unblocked Read 2 primer to generate nucleic acid base reads. In each cycle, fluorescently labeled nucleotides compete to add to the growing strand of the extended primer. Only one is incorporated at the primer position based on the sequence of the template strand. After nucleotide addition, the cluster is excited by a light source, and a characteristic fluorescent signal is emitted. The emission spectrum and signal intensity uniquely determine the base call. Hundreds of millions of nucleic acid clusters, or even thousands to tens of thousands of clusters, may be sequenced in a massively parallel manner. After sequencing the dsDNA bridge on the flow cell, the method can end at block 970.

[0081] sample Thus, in some embodiments, the sample comprises or consists of purified or isolated polynucleotides derived from a tissue sample, a biological fluid sample, a cell sample, or the like. Suitable biological fluid samples include, but are not limited to, blood, plasma, serum, sweat, tears, phlegm, urine, sputum, ear fluid, lymph, saliva, cerebrospinal fluid, ravage, bone marrow suspension, vaginal fluid, cervical lavage, cerebral fluid, peritoneal fluid, milk, respiratory, intestinal, and genitourinary tract secretions, amniotic fluid, milk, and leukapheresis samples. In some embodiments, the sample is a sample that can be readily obtained by non-invasive procedures, such as blood, plasma, serum, sweat, tears, phlegm, urine, sputum, ear fluid, saliva, or feces. In certain embodiments, the sample is a peripheral blood sample or the plasma and / or serum fraction of a peripheral blood sample. In other embodiments, the biological sample is a swab or smear, a biopsy specimen, or a cell culture. In another embodiment, the sample is a mixture of two or more biological samples, e.g., the biological sample can include two or more of a biological fluid sample, a tissue sample, and a cell culture sample. As used herein, the terms "blood," "plasma," and "serum" expressly encompass fractions thereof or processed portions thereof. Similarly, if a sample is obtained from a biopsy, swab, smear, etc., the term "sample" expressly encompasses processed fractions or portions obtained from the biopsy, swab, smear, etc.

[0082] In certain embodiments, samples include, but are not limited to, samples from different individuals, samples from different developmental stages of the same individual or different individuals, samples from different diseased individuals (e.g., individuals with cancer or suspected of having a genetic disease), normal individuals, samples obtained at different stages of a disease in an individual, samples obtained from individuals receiving different treatments for a disease, samples from individuals exposed to different environmental factors, samples from individuals predisposed to a disease condition, samples from individuals exposed to an infectious agent, etc.

[0083] In one exemplary, but non-limiting, embodiment, the sample is a maternal sample obtained from a pregnant woman, e.g., a pregnant woman. The maternal sample can be a tissue sample, a biological fluid sample, or a cell sample. In another exemplary, but non-limiting, embodiment, the maternal sample is a mixture of two or more biological samples, e.g., the biological sample can include two or more of a biological fluid sample, a tissue sample, and a cell culture sample.

[0084] In certain embodiments, samples can also be obtained from in vitro cultured tissues, cells, or other polynucleotide-containing sources. Cultured samples can be taken from sources including, but not limited to, cultures (e.g., tissues or cells) maintained in different media and conditions (e.g., pH, pressure, or temperature), cultures (e.g., tissues or cells) maintained for different time periods, cultures (e.g., tissues or cells) treated with different factors or reagents (e.g., drug candidates or modulators), or cultures of different types of tissues and / or cells.

[0085] In some embodiments, use of the disclosed sequencing techniques does not involve the preparation of a sequencing library. In other embodiments, the sequencing techniques contemplated herein involve the preparation of a sequencing library. In one exemplary approach, the preparation of a sequencing library involves the production of a random collection of adapter-modified DNA fragments (e.g., polynucleotides) that are ready to be sequenced.

[0086] Polynucleotide sequencing libraries can be prepared from DNA or RNA, including equivalents or analogs of either DNA or cDNA, such as DNA or cDNA, which is a complementary or copy DNA produced from an RNA template by the action of reverse transcriptase. Polynucleotides can be derived from double-stranded forms (e.g., dsDNA, such as genomic DNA fragments, cDNA, PCR amplification products, etc.), or in certain embodiments, polynucleotides can be derived from single-stranded forms (e.g., ssDNA, RNA, etc.) and converted to dsDNA form. By way of example, in certain embodiments, single-stranded mRNA molecules can be copied into double-stranded cDNA suitable for use in preparing a sequencing library. The exact sequence of the primary polynucleotide molecules is generally not critical to the library preparation method and can be known or unknown. In one embodiment, the polynucleotide molecules are DNA molecules. More specifically, in certain embodiments, the polynucleotide molecule represents the entire genetic complement of an organism or substantially the entire genetic complement of an organism and is a genomic DNA molecule (e.g., cellular DNA, cell-free DNA (cfDNA), etc.), which typically includes intron and exon sequences (coding sequences), as well as non-coding regulatory sequences such as promoter and enhancer sequences. In certain embodiments, the primary polynucleotide molecule comprises a human genomic DNA molecule, e.g., a cfDNA molecule present in the peripheral blood of a pregnant subject.

[0087] Methods for isolating nucleic acids from biological sources can vary depending on the nature of the source. Those skilled in the art can easily isolate nucleic acids from sources as required for the methods described herein. In some cases, it may be advantageous to fragment large nucleic acid molecules (e.g., cellular genomic DNA) in a nucleic acid sample to obtain polynucleotides of a desired size range. Fragmentation can be random or specific, for example, as achieved using restriction endonuclease digestion. Methods for random fragmentation can include, for example, limited DNase digestion, alkaline treatment, and physical shearing. Fragmentation can also be achieved by any of a number of methods known to those skilled in the art. For example, fragmentation can be achieved by mechanical means, including, but not limited to, nebulization, sonication, and hydroshearing.

[0088] In some embodiments, sample nucleic acid is obtained from unfragmented cfDNA.For example, cfDNA typically exists as fragments of less than about 300 base pairs, so fragmentation is typically not required to use cfDNA sample to generate sequencing library.

[0089] Typically, polynucleotides, whether forcibly fragmented (e.g., in vitro fragmented) or naturally occurring as fragments, are converted to blunt-ended DNA with a 5'-phosphate and a 3'-hydroxyl. Standard protocols, such as those for sequencing using Illumina platforms, instruct users to purify the end-repaired product before dA-tailing the end-repaired sample DNA and then purify the dA-tailed product before the adapter-ligation step of library preparation.

[0090] In various embodiments, verification of sample integrity and sample tracking can be achieved by sequencing a mixture of sample genomic nucleic acid, e.g., cfDNA, and accompanying marker nucleic acid that has been introduced into the sample, e.g., prior to processing.

[0091] Computer Systems In some embodiments, the disclosed systems and methods may involve approaches for shifting or distributing certain sequence data analysis functions and sequence data storage to a cloud computing environment or cloud-based network. User interaction with sequencing data, genomic data, or other types of biological data may be mediated through a central hub that stores the data and controls access to various interactions with the data. In some embodiments, the cloud computing environment may also provide shared protocols, analytical methods, libraries, sequence data, and distributed processing for sequencing, analysis, and reporting. In some embodiments, the cloud computing environment facilitates user correction or annotation of sequence data. In some embodiments, the systems and methods may be implemented in a computer browser, on-demand, or online.

[0092] In some embodiments, software written to perform the methods described herein is stored on some form of computer readable medium, such as memory, CD-ROM, DVD-ROM, memory stick, flash drive, hard drive, SSD hard drive, server, mainframe storage system, etc.

[0093] In some embodiments, the method may be written in any of a variety of suitable programming languages, for example, compiled languages ​​such as C, C#, C++, Fortran, and Java. Other programming languages ​​may include scripting languages ​​such as Perl, MatLab, SAS, SPSS, Python, Ruby, Pascal, Delphi, R, and PHP. In some embodiments, the method is written in C, C#, C++, Fortran, Java, Perl, R, Java, or Python. In some embodiments, the method may be a stand-alone application having data entry and data display modules. Alternatively, the method may be a computer software product, and distributed objects may include classes that include applications that include the computational methods described herein.

[0094] In some embodiments, the methods can be incorporated into existing data analysis software, such as that found on sequencing instruments. Software comprising the computer-implemented methods described herein can be installed directly on a computer system or indirectly stored on a computer-readable medium and loaded onto the computer system as needed. Additionally, the methods can be located on a computer that is remote from where the data is generated, such as software found on a server maintained at a separate location from where the data is generated, such as that provided by a third-party service provider.

[0095] The assay instrument, desktop computer, laptop computer, or server may include a processor in operative communication with accessible memory containing instructions for implementing the systems and methods. In some embodiments, the desktop computer or laptop computer is in operative communication with one or more computer-readable storage media or devices and / or output devices. The assay instrument, desktop computer, and laptop computer may operate under many different computer-based operating languages, such as those utilized by Apple-based or PC-based computer systems. The assay instrument, desktop computer, and / or laptop computer and / or server system may further provide a computer interface for creating or modifying experimental definitions and / or conditions, viewing data results, and monitoring experimental progress. In some embodiments, the output device may be a graphic user interface, such as a computer monitor or computer screen, a printer, a portable device such as a handheld digital assistant (i.e., PDA, Blackberry, iPhone), a tablet computer (e.g., iPAD), a hard drive, a server, a memory stick, a flash drive, or the like.

[0096] The computer-readable storage device or medium may be any device, such as a server, mainframe, supercomputer, magnetic tape system, etc. In some embodiments, the storage device may be located in close proximity to the assay device, e.g., adjacent to or in close proximity to the assay device. For example, the storage device may be located relative to the assay device in the same room, in the same building, in an adjacent building, on the same floor within a building, on a different floor within a building, etc. In some embodiments, the storage device may be located external to or distal to the assay device. For example, the storage device may be located in a different part of a city, in a different city, in a different state, or in a different country relative to the assay device. In embodiments in which the storage device is located distal to the assay device, communication between the assay device and one or more of a desktop, laptop, or server is typically via an internet connection, either wirelessly or via a network cable via an access point. In some embodiments, the storage device may be maintained and managed by an individual or entity directly associated with the assay device, while in other embodiments, the storage device may be maintained and managed by a third party, typically in a location distal to the individual or entity associated with the assay device. In the embodiments described herein, the output device can be any device for visualizing data.

[0097] The assay instrument, desktop, laptop, and / or server system may be used to store and / or retrieve computer-implemented software programs incorporating computer code for executing and implementing the computational methods described herein, data for use in implementing the computational methods, etc. One or more of the assay instrument, desktop, laptop, and / or server may include one or more computer-readable storage media for storing and / or retrieving software programs incorporating computer code for executing and implementing the computational methods described herein, data for use in implementing the computational methods, etc. Computer-readable storage media may include, but are not limited to, one or more of a hard drive, SSD hard drive, CD-ROM drive, DVD-ROM drive, floppy disk, tape, flash memory stick, or card. Furthermore, a network, including the Internet, may be a computer-readable storage medium. In some embodiments, a computer-readable storage medium refers to computational resource storage accessible by a computer network via the Internet or a corporate network provided by a service provider, rather than, for example, from a local desktop or laptop computer at a remote location to the assay instrument.

[0098] In some embodiments, computer-implemented software programs incorporating computer code for executing and implementing the computational methods described herein, computer-readable storage media for storing and / or retrieving data used in implementing the computational methods, etc., are operated and maintained by a service provider in operative communication with the assay instrument, desktop, laptop, and / or server system via an internet or network connection.

[0099] In some embodiments, a hardware platform for providing a computing environment includes a processor (i.e., CPU) where processor time and memory layout, such as random access memory (i.e., RAM), are system considerations. For example, smaller computer systems offer cheaper, faster processors and greater memory and storage capabilities. In some embodiments, a graphics processing unit (GPU) can be used. In some embodiments, a hardware platform for performing the computational methods described herein includes one or more computer systems with one or more processors. In some embodiments, smaller computers are clustered together to create a supercomputer network.

[0100] In some embodiments, the computational methods described herein are implemented on a collection of inter- or intra-connected computer systems (i.e., grid technology) that may cooperatively run various operating systems. For example, the CONDOR framework (University of Wisconsin-Madison) and systems available from United Devices are illustrative of the cooperation of multiple independent computer systems for the purpose of handling large amounts of data. These systems may provide a Perl interface for submitting, monitoring, and managing large sequence analysis jobs on a cluster in a serial or parallel configuration.

[0101] definition Unless otherwise defined, technical and scientific terms used in this disclosure have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. See, for example, Singleton et al., Dictionary of Microbiology and Molecular Biology 2nd ed., J. Wiley & Sons (New York, NY 1994); Sambrook et al., Molecular Cloning, A Laboratory Manual, Cold Spring Harbor Press (Cold Spring Harbor, NY 1989). For purposes of this disclosure, the following terms are defined below.

[0102] As used herein, the term "cluster" or "clamp" refers to a group of molecules (e.g., a group of DNA or a group of signals). In some embodiments, the signals of a cluster are derived from different features. In some embodiments, a signal clump represents a physical area covered by a single amplification oligonucleotide. In various examples, the physical area may be a tile, subtile, lane, or sublane on a flow cell. Ideally, each signal clump can be observed as several signals. Thus, overlapping signals can be detected from the same signal clump. In some embodiments, a signal cluster or clump may include one or more signals or spots corresponding to a particular feature. When used in connection with a microarray device or other molecular analysis device, a cluster may include one or more signals that together occupy a physical area occupied by an amplified oligonucleotide (or other polynucleotide or polypeptide having the same or similar sequence). For example, if the feature is an amplified oligonucleotide, the cluster may be the physical area covered by a single amplified oligonucleotide. In other embodiments, a signal cluster or clump need not strictly correspond to a feature. For example, a spurious noise signal may be included in a signal cluster, but not necessarily within a feature area. For example, a cluster of signals from four cycles of a sequencing reaction may include at least four signals.

[0103] As used herein, a "flow cell" can include a device having a lid extending over a reaction structure and forming a flow channel therebetween that is in communication with multiple reaction sites of the reaction structure, or can include a detection device configured to detect a specified reaction occurring at or near the reaction sites. The flow cell can include a solid-state optical detection or "imaging" device, such as a charge-coupled device (CCD) or a complementary metal-oxide semiconductor (CMOS) (photo) detection device. As one specific example, the flow cell can be fluidically configured and electrically coupled to a cartridge (with an integrated pump) that can be configured to fluidly and / or electrically couple to a bioassay system. The cartridge and / or bioassay system can deliver reaction solutions to the reaction sites of the flow cell according to a predetermined protocol (e.g., sequencing by synthesis) and perform multiple imaging events. For example, the cartridge and / or bioassay system can direct one or more reaction solutions through the flow channel of the flow cell and thereby along the reaction sites. At least one of the reaction solutions may contain four types of nucleotides with the same or different fluorescent labels. The nucleotides may be bound to reaction sites of the flow cell, such as corresponding oligonucleotides of the reaction sites. The reaction sites may then be illuminated using an excitation light source (e.g., a solid-state light source such as a light-emitting diode (LED)) to illuminate the cartridge and / or bioassay system. The excitation light may have a predetermined wavelength or multiple wavelengths, including a range of wavelengths. Fluorescent labels excited by the incident excitation light may provide an emission signal (e.g., light of one or more wavelengths different from the excitation light, and potentially different from each other) that can be detected by a light sensor in the flow cell.

[0104] The flow cells described herein can be configured to carry out a variety of biological or chemical processes. More specifically, the flow cells described herein can be used in a variety of processes and systems in which it is desirable to detect events, properties, qualities, or characteristics indicative of a specified reaction. For example, the flow cells described herein can include or be integrated with optical detection devices, biosensors, and their components, as well as bioassay systems operating in conjunction with the biosensors. The flow cells can be configured to facilitate multiple specified reactions that can be detected individually or collectively. The flow cells can be configured to perform multiple cycles in which multiple specified reactions occur in parallel. For example, the flow cells can be used to sequence high-density arrays of DNA features through repeated cycles of enzymatic operation and light or image detection / capture. Thus, the flow cells can be in fluid communication with one or more microfluidic channels that deliver reagents or other reaction components in a reaction solution to the reaction sites of the flow cell. The reaction sites can be provided or spaced in a predetermined manner, such as a uniform or repeating pattern. Alternatively, the reaction sites can be randomly distributed. Each reaction site can be associated with one or more light guides and one or more light sensors that detect light from the associated reaction site. In one example, the light guide includes one or more filters for filtering certain wavelengths of light. The light guide can be, for example, an absorptive filter (e.g., an organic absorptive filter) such that the filter material absorbs certain wavelengths (or ranges of wavelengths) and allows at least one predetermined wavelength (or range of wavelengths) to pass therethrough. In some flow cells, the reaction site can be located within a reaction recess or chamber within which a designated reaction can be at least partially compartmentalized.

[0105] As used herein, the terms "spot radius" or "cluster radius" refer to a defined radius that encompasses a diffraction-limited spot or cluster of signals. Thus, by defining a cluster radius larger or smaller, a larger number of signals can fall within the radius for subsequent ordering and selection. The cluster radius can be defined in terms of any distance measure, such as pixels, meters, millimeters, or any other useful measure of distance.

[0106] As used herein, a "signal" refers to a detectable event, such as, for example, luminescence, light emission, etc., in an image. Thus, in some embodiments, a signal can represent any detectable luminescence (i.e., a "spot") captured in an image. Thus, as used herein, a "signal" can refer to actual emissions from analyte features, or to spurious emissions that do not correlate with actual features. Thus, a signal may result from noise and may be subsequently discarded as not representative of actual analyte features on the test strip.

[0107] As used herein, the "intensity" of emitted light refers to the intensity of light transmitted per unit area, where the area is measured on a plane perpendicular to the direction of propagation of the light beam, and the intensity is the amount of energy transmitted per unit time. In some embodiments, signal "intensity," "amplitude," "magnitude," or "level" may be used synonymously with signal intensity. In some embodiments, the image captured by the detector approximates or is proportional to an intensity map integrated over a certain amount of time. In some embodiments, the signal of a diffraction-limited spot of DNA clusters is extracted from the image as the total intensity contained in the spot up to a factor of the integration time. For example, the signal of a DNA cluster may be defined as the intensity contained within the spot radius of the DNA cluster up to a factor of the integration time. In other embodiments, the peak intensity value found within the spot radius may be used to represent the signal of the DNA cluster up to a factor of the integration time.

[0108] As used herein, the process of registering a template of signal locations onto a given image is referred to as “registration,” and the process of determining the intensity or amplitude value of each signal in the template for a given image is referred to as “intensity extraction.” For registration, the methods and systems provided herein can take advantage of the random nature of signal clamp locations by using image correlation to register the template to the image.

[0109] As used herein, a "nucleotide" comprises a nitrogen-containing heterocyclic base, a sugar, and one or more phosphate groups. Nucleotides are the monomeric units of nucleic acid sequences. Examples of nucleotides include ribonucleotides and deoxyribonucleotides. In ribonucleotides (RNA), the sugar is ribose, while in deoxyribonucleotides (DNA), the sugar is deoxyribose, i.e., a sugar lacking the hydroxyl group at the 2' position of the ribose. The nitrogen-containing heterocyclic base can be a purine base or a pyrimidine base. Purine bases include adenine (A) and guanine (G), as well as modified derivatives or analogs thereof. Pyrimidine bases include cytosine (C), thymine (T), and uracil (U), as well as modified derivatives or analogs thereof. The C-1 atom of deoxyribose is linked to the N-1 atom of a pyrimidine or the N-9 atom of a purine. The phosphate group can be mono-, di-, or triphosphate. These nucleotides may be naturally occurring nucleotides, although it should be further understood that non-naturally occurring nucleotides, modified nucleotides or analogs of the aforementioned nucleotides may also be used.

[0110] As used herein, a "nucleobase" is a heterocyclic base, such as adenine, guanine, cytosine, thymine, uracil, inosine, xanthine, hypoxanthine, or a heterocyclic derivative, analog, or tautomer thereof. Nucleobases can be naturally occurring or synthetic. Non-limiting examples of nucleobases include adenine, guanine, thymine, cytosine, uracil, xanthine, hypoxanthine, 8-azapurine, purine substituted at the 8-position with methyl or bromine, 9-oxo-N6-methyladenine, 2-aminoadenine, 7-deazaxanthine, 7-deazaguanine, 7-deaza-adenine, N4-ethanocytosine, 2,6-diaminopurine, N6-ethano-2,6-diaminopurine, 5-methylcytosine, 5-(C3-C6)-alkynylcytosine, 5-fluorouracil, and 5-bromouracil. , thiouracil, pseudoisocytosine, 2-hydroxy-5-methyl-4-triazolopyridine, isocytosine, isoguanine, inosine, 7,8-dimethylalloxazine, 6-dihydrothymine, 5,6-dihydrouracil, 4-methyl-indole, ethenoadenine, and the non-naturally occurring nucleobases described in U.S. Pat. Nos. 5,432,272 and 6,150,510, and WO 92 / 002258, WO 93 / 10820, WO 94 / 22892, and WO 94 / 24144, and Fasman (Practical Handbook of Biochemistry and Molecular Biology, pp. 385-394, 1989, CRC Press, Boca Raton, LO), all of which are incorporated herein by reference in their entireties.

[0111] The terms "nucleic acid" or "polynucleotide" refer to deoxyribonucleotide or ribonucleotide polymers in single- or double-stranded form and, unless otherwise limited, encompass known analogs of natural nucleotides that hybridize to nucleic acids in a manner similar to naturally occurring nucleotides, such as peptide nucleic acids (PNAs) and phosphorothioate DNA. Unless otherwise specified, a particular nucleic acid sequence includes its complementary sequence. Nucleotides include, but are not limited to, ATP, dATP, CTP, dCTP, GTP, dGTP, UTP, TTP, dUTP, 5-methyl-CTP, 5-methyl-dCTP, ITP, dITP, 2-amino-adenosine-TP, 2-amino-deoxyadenosine-TP, 2-thiothymidine triphosphate, pyrrolo-pyrimidine triphosphate, and 2-thiocytidine, as well as alpha-thiotriphosphate for all of the above, and 2'-O-methyl-ribonucleotide triphosphates for all of the above bases. Modified bases include, but are not limited to, 5-Br-UTP, 5-Br-dUTP, 5-F-UTP, 5-F-dUTP, 5-propynyl-dCTP, and 5-propynyl-dUTP.

[0112] The polymerases used are generally enzymes for joining 3'-OH 5'-triphosphate nucleotides, oligomers, and their analogs. Polymerases include DNA-dependent DNA polymerase, DNA-dependent RNA polymerase, RNA-dependent DNA polymerase, RNA-dependent RNA polymerase, T7 DNA polymerase, T3 DNA polymerase, T4 DNA polymerase, T7 RNA polymerase, T3 RNA polymerase, SP6 RNA polymerase, DNA polymerase I, Klenow fragment, Thermophilus aquaticus DNA polymerase, Tth DNA polymerase, VentR® DNA polymerase (New England Biolabs), Deep VentR® DNA polymerase (New England Biolabs), Bst DNA polymerase large fragment, Stoeffel fragment, 90N DNA polymerase, 90N DNA polymerase, Pfu DNA polymerase, TfI DNA polymerase, Tth DNA polymerase, RepliPHI Phi29 polymerase, and TIi Examples of suitable polymerases include, but are not limited to, DNA polymerases, eukaryotic DNA polymerase beta, telomerase, Therminator™ polymerase (New England Biolabs), KOD HiFi™ DNA polymerase (Novagen), KOD1 DNA polymerase, Q-beta replicase, terminal transferase, AMV reverse transcriptase, M-MLV reverse transcriptase, Phi6 reverse transcriptase, HIV-1 reverse transcriptase, novel polymerases discovered through bioprospecting, and polymerases cited in U.S. Patent Application Publication No. 2007 / 0048748, U.S. Patent Nos. 6,329,178, 6,602,695, and 6,395,524 (incorporated by reference). These polymerases include wild-type, mutant isoforms, and engineered variants. "Encode" or "parse" are verbs that refer to transferring from one format to another, and refer to transferring the genetic information of a target template sequence into a reporter configuration.

[0113] Nucleosides and nucleotides may be labeled at sites on the sugar or nucleobase. Dyes may be attached to any position on the nucleotide base, for example, via a linker. In certain embodiments, Watson-Crick base pairing can still be performed on the resulting analogs. Specific nucleobase labeling sites include the C5 position of pyrimidine bases or the C7 position of 7-deazapurine bases. A linker group may be used to covalently attach a dye to a nucleoside or nucleotide. As used herein, the term "covalently attached" or "covalently bonded" refers to the formation of a chemical bond characterized by the sharing of electron pairs between atoms. For example, a covalently bonded polymer coating refers to a polymer coating that forms a chemical bond with the functionalized surface of a substrate, as compared to attachment to the surface by other means, such as adhesion or electrostatic interactions. It will be understood that a polymer covalently attached to a surface may be attached by means in addition to covalent bonding.

[0114] A variety of different types of linkers can be used, having different lengths and chemical properties. The term "linker" encompasses any moiety useful for connecting one or more molecules or compounds to each other, to other components of a reaction mixture, and / or to a reaction site. For example, a linker can attach a reporter molecule or "label" (e.g., a fluorescent dye) to a reaction component. In certain embodiments, the linker is a member selected from substituted or unsubstituted alkyl (e.g., 2-5 carbon chains), substituted or unsubstituted heteroalkyl, substituted or unsubstituted aryl, substituted or unsubstituted heteroaryl, substituted or unsubstituted cycloalkyl, and substituted or unsubstituted heterocycloalkyl. In one example, the linker moiety is selected from straight and branched carbon chains, optionally containing at least one heteroatom (e.g., at least one functional group such as ether, thioether, amide, sulfonamide, carbonate, carbamate, urea, and thiourea), and optionally containing at least one aromatic, heteroaromatic, or non-aromatic ring structure (e.g., cycloalkyl, phenyl). In certain embodiments, molecules with trifunctional linking capabilities are used, including, but not limited to, cynuric chloride, mealamine, diaminopropanoic acid, aspartic acid, cysteine, glutamic acid, pyroglutamic acid, S-acetylmercaptosuccinic anhydride, carbobenzoxylidine, histidine, lysine, serine, homoserine, tyrosine, piperidinyl-1,1-aminocarboxylic acid, diaminobenzoic acid, etc. In certain specific embodiments, hydrophilic PEG (polyethylene glycol) linkers are used.

[0115] In certain embodiments, the linker is derived from a molecule containing at least two reactive functional groups (e.g., one at each end), which can react with complementary reactive functional groups on various reactive components or can be used to immobilize one or more reactive components at a reactive site. "Reactive functional group," as used herein, includes olefins, acetylenes, alcohols, phenols, ethers, oxides, halides, aldehydes, ketones, carboxylic acids, esters, amides, cyanates, isocyanates, thiocyanates, isothiocyanates, amines, hydrazines, hydrazones, hydrazides, diazos, diazonium, nitros, nitriles, mercaptans, sulfides, disulfides, sulfoxides, sulfones, sulfonic acids, sulfinic acids, acetals, ketones, and the like. Reactive functional groups include, but are not limited to, ols, anhydrides, sulfates, sulfenic acid isonitriles, amidines, imides, imidates, nitrones, hydroxylamines, oximes, hydroxamic acids, thiohydroxamic acids, allenes, orthoesters, sulfites, enamines, ynamines, ureas, isoureas, semicarbazides, carbodiimides, carbamates, imines, azides, azo compounds, azoxy compounds, and nitroso compounds. Reactive functional groups also include those used to prepare bioconjugates, such as N-hydroxysuccinimide esters, maleimides, and the like.

[0116] The cleavable linker can be, by way of non-limiting example, an electrophilically cleavable linker, a nucleophilically cleavable linker, a photocleavable linker, a linker cleavable under reducing conditions (e.g., a disulfide- or azide-containing linker), a linker cleavable under oxidative conditions, a linker cleavable by the use of a safety lock linker, or a linker cleavable by an elimination mechanism. The use of a cleavable linker to attach the dye compound to the substrate moiety allows for the label to be removed after detection, if desired, to avoid any interfering signals in downstream steps.

[0117] In some embodiments, one or more dye molecules or labeling molecules may be attached to a nucleotide base through non-covalent interactions or a combination of covalent and non-covalent interactions via multiple intervening molecules. In one example, nucleotides or nucleotide analogs newly incorporated by a polymerase synthesizing a target polynucleotide are initially unlabeled. One or more fluorescent labels can be introduced into the nucleotide or nucleotide analog by binding to a labeled affinity reagent containing one or more fluorescent dyes. The use of unlabeled nucleotides and affinity reagents in sequencing by synthesis is disclosed in U.S. Publication No. 2013 / 0079232, which is incorporated herein by reference. For example, one, two, three, or each of four different types of nucleotides (e.g., dATP, dCTP, dGTP, and dTTP or dUTP) in a reaction mixture may be initially unlabeled. Each of the four types of nucleotides (e.g., dNTPs) may have a 3' hydroxy blocking group to ensure that only a single base can be added by the polymerase to the 3' end of a copy polynucleotide being synthesized from the target polynucleotide. After incorporation of an unlabeled nucleotide, an affinity reagent that specifically binds to the incorporated dNTP can then be introduced to provide a labeled extension product containing the incorporated dNTP. The affinity reagent can be designed to specifically bind to the incorporated dNTP through, for example, an antibody-antigen interaction or a ligand-receptor interaction. The dNTP can be modified to contain a specific antigen that pairs with a specific antibody contained in the corresponding affinity reagent. Thus, one, two, three, or each of the four different types of nucleotides can be specifically labeled via their corresponding affinity reagents.In some embodiments, the affinity reagent may comprise a small molecule or protein tag capable of binding to a hapten moiety of a nucleotide (e.g., streptavidin-biotin, anti-DIG and DIG, anti-DNP and DNP), an antibody (including, but not limited to, antibody binding fragments, single-chain antibodies, bispecific antibodies, etc.), an aptamer, a knottin, an affimer, or any other known agent that binds to the incorporated nucleotide with appropriate specificity and affinity. In some embodiments, the hapten moiety of an unlabeled nucleotide may be attached to the nucleobase via a cleavable linker, which may be cleaved under the same reaction conditions as those for removing the 3' blocking group. In some embodiments, a single affinity reagent may be labeled with multiple copies of the same fluorescent dye, e.g., 1, 2, 3, 4, 5, 6, 8, 10, 12, or 15 copies of the same dye. In further embodiments, each affinity reagent may be labeled with a different number of copies of the same fluorescent dye. In some embodiments, the first affinity reagent may be labeled with a first number of first fluorescent dyes, the second affinity reagent may be labeled with a second number of second fluorescent dyes, the third affinity reagent may be labeled with a third number of third fluorescent dyes, and the fourth affinity reagent may be labeled with a fourth number of fourth fluorescent dyes. In some embodiments, each affinity reagent may be labeled with a different combination of one or more types of dyes, where each type of dye has a certain copy number. In some embodiments, different affinity reagents may be labeled with different dyes that can be excited by the same light source, but each dye has a distinct fluorescence intensity or a distinct emission spectrum. In some embodiments, different affinity reagents may be labeled with the same dyes at different molar ratios to produce measurable differences in their fluorescence intensities.

[0118] The nucleotide analogs may be conjugated or associated with one or more optically detectable labels to provide a detectable signal. In some embodiments, the optically detectable labels may be fluorescent compounds, such as small molecule fluorescent labels. Suitable fluorescent molecules (fluorophores) as fluorescent labels include 1,5 IAEDANS; 1,8-ANS; 4-methylumbelliferone; 5-carboxy-2,7-dichlorofluorescein; 5-carboxyfluorescein (5-FAM); fluorescein amidite (FAM); 5-carboxyfluorescein; tetrachloro-6-carboxyfluorescein (TET); hexachloro-6-carboxyfluorescein (HEX); 2,7-dimethoxy-4,5-dichloro-6-carboxyfluorescein (JOE); VIC®; NED®; tetramethylrhodamine (TMR); 5-carboxytetramethylrhodamine (5-TAMRA); 5-HAT (hydroxytryptamine); 5-hydroxytryptamine (HAT); 5-ROX (carboxy-X-rhodamine); 6-carboxyrhodamine 6G; 6-JOE; Light Cycler® Red 610; Light Cycler® Red 640; Light Cycler® Red 670; Light Cycler® Red 705; 7-amino-4-methylcoumarin; 7-aminoactinomycin D (7-AAD); 7-hydroxy-4-methylcoumarin; 9-amino-6-chloro-2-methoxyacridine; 6-methoxy-N-(4-aminoalkyl)quinolinium bromide hydrochloride (ABQ); Acid Fuchsin; ACMA (9-amino-6-chloro-2-methoxyacridine); Acridine Orange; Acridine Red; Acridine Yellow; Acriflavine; Acriflavine Feulgen SITSA; AFP-Autofluorescent Protein- (Quantum Biotechnologies); Texas Red; Texas Red-X Conjugate; Thiadicarbocyanine (DiSC3); Thiazine Red R; Thiazole Orange; Thioflavin 5; Thioflavin S; Thioflavin TCN; Thiolite; Thiozole Orange; Tinopol CBS (Calcofluor White);TMR;TO-PRO-1;TO-PRO-3;TO-PRO-5; TOTO-1; TOTO-3; TriColor (PE-Cy5); TRITC (Tetramethylrhodamine-lsoThioCyanate, tetramethylrhodamine-isothiocyanate); True Blue; TruRed; Ultralight; Uranine B; Uvitex SFC; WW 781; X-rhodamine; X-rhodamine-5-(and-6)-isothiocyanate (5(6)-XRITC); Xylene Orange; Y66F; Y66H; Y66W; YO-PRO-1; YO-PRO-3; YOYO-1; YOYO-3; Sybr Green; Thiazole Orange; and other interchelating dyes covering a wide spectrum, including Alexa Fluor 350 and Alexa Fluor Members of the Alexa Fluor® dye series (Molecular Probes / Invitrogen) that match the primary output wavelengths of common excitation sources, such as 405, 430, 488, 500, 514, 532, 546, 555, 568, 594, 610, 633, 635, 647, 660, 680, 700, and 750; members of the Cy Dye fluorophore series (GE Healthcare) that also cover a broad spectrum, such as Cy3, Cy3B, Cy3.5, Cy5, Cy5.5, and Cy7; and Oyster® dye fluorophores (Denovo) such as Oyster-500, -550, -556, 645, 650, and 656. members of the DY-Labels series (Dyomics) having an absorption maximum in the range of 418 nm (DY-415) to 844 nm (DY-831), e.g., DY-415, -495, -505, -547, -548, -549, -550, -554, -555, -556, -560, -590, -610, -615, -630, -631, -632, -633, - 634, -635, -636, -647, -648, -649, -650, -651, -652, -675, -676, -677, -680, -681, -682, -700, -701, -730, -731 , -732, -734, -750, -751, -752, -776, -780, -781, -782, -831, -480XL, -481XL, -485XL, -510XL, -520XL, -521XL;Examples of fluorescent labels include, but are not limited to, members of the ATTO series (ATTO-TEC GmbH) of fluorescent labels, such as ATTO 390, 425, 465, 488, 495, 520, 532, 550, 565, 590, 594, 610, 611X, 620, 633, 635, 637, 647, 647N, 655, 680, 700, 725, and 740; and dyes from the CAL Fluor® series or Quasar® series (Biosearch Technologies), such as CAL Fluor® Gold 540, CAL Fluor® Orange 560, Quasar® 570, CAL Fluor® Red 590, CAL Fluor® Red 610, CAL Fluor® Red 635, Quasar® 570, and Quasar® 670. In some embodiments, the first optically detectable label interacts with a second optically detectable moiety to alter the detectable signal, for example, via fluorescence resonance energy transfer ("FRET" (Förster resonance energy transfer); also known as Förster resonance energy transfer);

[0119] Fluorescent labels utilized by the systems and methods disclosed herein can have different peak absorption wavelengths, for example, ranging from 400 nm to 800 nm. In some embodiments, the peak absorption wavelength of a fluorescent label can be approximately at or between 400, 410, 420, 430, 440, 450, 460, 470, 480, 490, 500, 510, 520, 530, 540, 550, 560, 570, 580, 590, 600, 610, 620, 630, 640, 650, 660, 670, 680, 690, 700, 710, 720, 730, 740, 750, 760, 770, 780, 790, 800 nm, or a number or range between any two of these values. In some embodiments, the peak absorption wavelength of the fluorescent label can be at least or at most 400, 410, 420, 430, 440, 450, 460, 470, 480, 490, 500, 510, 520, 530, 540, 550, 560, 570, 580, 590, 600, 610, 620, 630, 640, 650, 660, 670, 680, 690, 700, 710, 720, 730, 740, 750, 760, 770, 780, 790, or 800 nm.

[0120] Fluorescent labels can have different peak emission wavelengths, for example, ranging from 400 nm to 800 nm. In some embodiments, the peak emission wavelength of a fluorescent label can be approximately at or between 400, 410, 420, 430, 440, 450, 460, 470, 480, 490, 500, 510, 520, 530, 540, 550, 560, 570, 580, 590, 600, 610, 620, 630, 640, 650, 660, 670, 680, 690, 700, 710, 720, 730, 740, 750, 760, 770, 780, 790, 800 nm, or a number or range between any two of these values. In some embodiments, the peak emission wavelength of the fluorescent label can be at least or at most 400, 410, 420, 430, 440, 450, 460, 470, 480, 490, 500, 510, 520, 530, 540, 550, 560, 570, 580, 590, 600, 610, 620, 630, 640, 650, 660, 670, 680, 690, 700, 710, 720, 730, 740, 750, 760, 770, 780, 790, or 800 nm.

[0121] Fluorescent labels can have different Stokes shifts, for example, ranging from 10 nm to 200 nm. In some embodiments, the Stokes shift can be, or can be approximately, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200 nm, or a number or range between any two of these values. In some embodiments, the Stokes shift can be at least or at most 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, or 200 nm.

[0122] In some embodiments, the distance between the peak emission wavelengths of any two fluorescent labels can vary, for example, in the range of 10 nm to 200 nm. In some embodiments, the distance between the peak emission wavelengths of any two fluorescent labels can be, or can be approximately, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, or 200 nm, or a number or range between any two of these values. In some embodiments, the distance between the peak emission wavelengths of any two fluorescent labels can be at least or at most 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, or 200 nm.

[0123] A "light source" may be any device capable of emitting energy along the electromagnetic spectrum. The light source may be a source of visible light (VIS), ultraviolet light (UV), and / or infrared light (IR). "Visible light" (VIS) generally refers to the band of electromagnetic radiation having wavelengths from about 400 nm to about 750 nm. "Ultraviolet (UV) light" generally refers to electromagnetic radiation having wavelengths shorter than those of visible light, or in the range of about 10 nm to about 400 nm. "Infrared light" or infrared radiation (IR) generally refers to electromagnetic radiation having wavelengths greater than the VIS range, or in the range of about 750 nm to about 50,000 nm. A light source may also provide full-spectrum light. A light source can output light from a selected wavelength or range of wavelengths. In some embodiments of the present invention, a light source may be configured to provide light above or below a predetermined wavelength, or may provide light within a predetermined range. A light source may be used in combination with a filter to selectively transmit or block light of selected wavelengths from the light source. The light source can be connected to a power source by one or more electrical connectors. An array of light sources can be connected in series or parallel to the power source. The power source can be a battery, a vehicle electrical system, or a building electrical system. The light source can be connected to the power source through control electronics (control circuitry). The control electronics can include one or more switches. The one or more switches can be automated, controlled by a sensor, timer, or other input, or controlled by a user, or a combination thereof. For example, a user can operate a switch to turn on a UV light source. The light source can be applied on a constant basis until it is turned off, or it can be pulsed (repeated on / off cycles) until it is turned off. In some embodiments, the light source can be switched from a continuously on state to a pulsed state, or vice versa. In some embodiments, the light source can be configured to get brighter or dimmer over time.

[0124] For operation, the light source may be connected to a power source capable of providing sufficient intensity to illuminate the sample. The control electronics may be used to switch the intensity on or off based on input from a user or some other input, and may also be used to modulate the intensity to a suitable level (e.g., to control the brightness of the output light). The control electronics may be configured to turn the light source on and off as needed. The control electronics may include switches for manual, automatic, or semi-automatic operation of the light source. The one or more switches may be, for example, a transistor, a relay, or an electromechanical switch. In some embodiments, the control circuitry may further include an AC-DC and / or DC-DC converter for converting a voltage from a voltage source to an appropriate voltage for the light source. The control circuitry may also include a DC-DC regulator for regulating the voltage. The control circuitry may further include a timer and / or other circuit elements for applying a voltage to the optical filter for a fixed period following receipt of the input. The switch may be activated manually, automatically in response to a predetermined condition, or using a timer. For example, the control electronics may process information such as user input, stored instructions, etc.

[0125] One or more of a plurality of light sources may be provided. In some embodiments, each of the plurality of light sources may be the same. Alternatively, one or more of the light sources may vary. The light characteristics of the light emitted by the light sources may be the same or different. The plurality of light sources may or may not be independently controllable. One or more characteristics of the light sources may or may not be controlled, including, but not limited to, whether the light source is on or off, the brightness of the light source, the wavelength of the light, the intensity of the light, the angle of illumination, the position of the light source, or any combination thereof.

[0126] In some embodiments, the light output from the light source may be about 350 to about 750 nm, or any amount or range therebetween, such as about 350 nm to about 360, 370, 380, 390, 400, 410, 420, 430, or about 450 nm, or any amount or range therebetween. In other embodiments, the light from the light source may be about 550 to about 700 nm, or any amount or range therebetween, such as about 550 to about 560, 570, 580, 590, 600, 610, 620, 630, 640, 650, 660, 670, 680, 690, or about 700 nm, or any amount or range therebetween. In some embodiments, the wavelength of the light produced by the light source may vary, for example, from 400 nm to 800 nm. In some embodiments, the wavelength of the light produced by the light source can be, or can be approximately, 400, 410, 420, 430, 440, 450, 460, 470, 480, 490, 500, 510, 520, 530, 540, 550, 560, 570, 580, 590, 600, 610, 620, 630, 640, 650, 660, 670, 680, 690, 700, 710, 720, 730, 740, 750, 760, 770, 780, 790, 800 nm, or a number or range between any two of these values. In some embodiments, the wavelength of light generated by the light source can be at least or at most 400, 410, 420, 430, 440, 450, 460, 470, 480, 490, 500, 510, 520, 530, 540, 550, 560, 570, 580, 590, 600, 610, 620, 630, 640, 650, 660, 670, 680, 690, 700, 710, 720, 730, 740, 750, 760, 770, 780, 790, or 800 nm. The light source may be capable of emitting electromagnetic radiation in any spectrum. In some embodiments, the light source may have a wavelength between 10 nm and 100 μm. In some embodiments, the wavelength of the light can be between 100 nm and 5000 nm, between 300 nm and 1000 nm, or between 400 nm and 800 nm.In some embodiments, the wavelength of the light may be less than and / or equal to 10 nm, 100 nm, 200 nm, 300 nm, 400 nm, 500 nm, 600 nm, 700 nm, 800 nm, 900 nm, 1000 nm, 1100 nm, 1200 nm, 1300 nm, 1500 nm, 1750 nm, 2000 nm, 2500 nm, 3000 nm, 4000 nm, or 5000 nm.

[0127] In one example, the light source may be a light emitting diode (LED) (e.g., a gallium arsenide (GaAs) LED, an aluminum gallium arsenide (AlGaAs) LED, a gallium arsenide phosphide (GaAsP) LED, an aluminum gallium indium phosphide (AlGaInP) LED, a gallium (III) phosphide (GaP) LED, an indium gallium nitride (InGaN) / gallium (III) nitride (GaN) LED, or an aluminum gallium phosphide (AlGaP) LED). In another example, the light source can be a laser, e.g., a vertical cavity surface emitting laser (VCSEL), or other suitable light emitter, such as an indium-gallium-aluminum-phosphide (InGaAlP) laser, a gallium-arsenide phosphide / gallium phosphide (GaAsP / GaP) laser, or a gallium-aluminum-arsenide / gallium-aluminum-arsenide (GaAlAs / GaAs) laser.Other examples of light sources may include, but are not limited to, electronically excited light sources (e.g., cathodoluminescence, electronically stimulated luminescence (ESL) bulbs, cathode ray tubes (CRT monitors), Nixie tubes), incandescent light sources (e.g., carbon button lamps, conventional incandescent light bulbs, halogen lamps, glow bars, Nernst lamps), electroluminescent (EL) light sources (e.g., light emitting diodes, organic light emitting diodes, polymer light emitting diodes, solid state lighting, LED lamps, electroluminescent sheets, electroluminescent wire), gas discharge light sources (e.g., fluorescent lamps, induction lighting, hollow cathode lamps, neon and argon lamps, plasma lamps, xenon flash lamps), or high intensity discharge light sources (e.g., carbon arc lamps, ceramic discharge metal halide lamps, hydrogen iodide arc lamps, mercury vapor lamps, metal halide lamps, sodium vapor lamps, xenon arc lamps). Alternatively, the light source may be a bioluminescent, chemiluminescent, phosphorescent, or fluorescent light source.

[0128] As used herein, an "optical channel" is a predetermined profile of optical frequencies (or equivalently, wavelengths). For example, a first optical channel may have wavelengths between 500 nm and 600 nm. To capture an image in the first optical channel, a detector responsive only to light between 500 nm and 600 nm may be used, or a bandpass filter with a transmission window of 500 nm to 600 nm may be used to filter the incident light to a detector responsive to light between 300 nm and 800 nm. A second optical channel may have wavelengths between 300 nm and 450 nm and between 850 nm and 900 nm. To capture an image in the second optical channel, a detector responsive to light between 300 nm and 450 nm and another detector responsive to light between 850 nm and 900 nm may be used, and the detection signals of the two detectors may then be combined. Alternatively, to capture an image in the second optical channel, a bandstop filter that rejects light between 300 nm and 900 nm may be used in front of a detector responsive to light between 451 nm and 849 nm.

[0129] Additional Notes The embodiments described herein are exemplary. Modifications, rearrangements, alternative processes, etc. may be made to these embodiments and still fall within the teachings described herein. One or more of the steps, processes, or methods described herein may be performed by one or more suitably programmed processing and / or digital devices.

[0130] The various illustrative imaging or data processing techniques described in connection with the embodiments disclosed herein can be implemented as electronic hardware, computer software, or a combination of both. To illustrate this interchangeability of hardware and software, various illustrative components, blocks, modules, and steps have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends on the particular application and design constraints imposed on the overall system. The described functionality may be implemented in various ways for each particular application, and such implementation decisions should not be interpreted as causing a departure from the scope of the present disclosure.

[0131] The various exemplary detection systems described in connection with the embodiments disclosed herein may be implemented or performed by a mechanical apparatus such as a processor configured with specific instructions, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. The processor may be a microprocessor, but alternatively, the processor may be a controller, microcontroller, or state machine, combinations thereof, etc. A processor may also be implemented as a combination of computing devices, such as a DSP and a microprocessor, multiple microprocessors, one or more microprocessors in association with a DSP core, or any other such configuration. For example, the systems described herein may be implemented using discrete memory chips, a portion of memory within a microprocessor, flash, EPROM, or other types of memory.

[0132] Elements of a method, process, or algorithm described in connection with the embodiments disclosed herein may be embodied directly in hardware, in a software module executed by a processor, or in a combination of the two. The software module may reside in RAM memory, flash memory, ROM memory, EPROM memory, EEPROM memory, registers, a hard disk, a removable disk, a CD-ROM, or any other form of computer-readable storage medium known in the art. An exemplary storage medium may be coupled to the processor such that the processor can read information from, and write information to, the storage medium. Alternatively, the storage medium may be integral to the processor. The processor and the storage medium may reside in an ASIC. The software module may include computer-executable instructions that cause a hardware processor to execute the computer-executable instructions.

[0133] Unless otherwise indicated, conditional language used herein, such as "can," "might," "may," "eg," and the like, unless otherwise indicated and as otherwise understood within the context of use, is intended to generally convey that certain embodiments are included, and that other embodiments do not include certain features, elements, and / or steps. Thus, such conditional language does not generally imply that features, elements, and / or conditions are required in any manner for one or more embodiments, or that one or more embodiments necessarily include logic for determining or prompting, with or without author input, whether those features, elements, and / or conditions are included or performed in any particular embodiment. Terms such as "comprising," "including," "having," and "involving" are synonymous and used in an inclusive, open-ended manner and do not exclude additional elements, features, acts, operations, etc. Also, the term "or," when used to connect, for example, a list of elements, is used in its inclusive sense (rather than its exclusive sense), so as to mean one, some, or all of the elements in the list.

[0134] Disjunctive language, such as the phrase "at least one of X, Y, or Z," is understood differently in context as it is generally used to indicate that an item, term, etc. can be either X, Y, or Z, or any combination thereof (e.g., X, Y, and / or Z), unless specifically stated otherwise. Thus, such disjunctive language is generally not intended to, and should not, imply that a particular embodiment requires that at least one of X, at least one of Y, or at least one of Z, respectively, be present.

[0135] Terms such as "about" or "approximately" are synonymous and are used to indicate that the value modified by the term has an understood range associated with it, which may be ±20%, ±15%, ±10%, ±5%, or ±1%. The term "substantially" is used to indicate that a result (e.g., a measurement) is close to a target value, where close may mean, for example, that the result is within 80% of the value, within 90% of the value, within 95% of the value, or within 99% of the value. The term "partially" is used to indicate that an effect is only partial or to a limited extent.

[0136] Unless otherwise noted, articles such as "a" or "an" should generally be construed to include one or more listed items. Thus, phrases such as "a device configured to" or "a device for" are intended to include one or more listed devices. Such one or more listed devices may also be collectively configured to perform the described detailed descriptions. For example, "a processor configured to perform detailed descriptions A, B, and C" may include a first processor to perform operations in conjunction with a second processor configured to perform detailed description A and to perform operations in conjunction with a second processor configured to perform detailed descriptions B and C.

[0137] While the foregoing detailed description has illustrated, described, and pointed out novel features applied to the exemplary embodiments, it will be understood that various omissions, substitutions, and changes in the form and details of the devices or algorithms shown may be made without departing from the spirit of the present disclosure. It will be recognized that certain embodiments described herein may be embodied in forms that do not provide all of the features and advantages described herein, since some features may be used or practiced separately from others. All changes that come within the meaning and range of equivalency of the claims are intended to be embraced within their scope.

[0138] It is to be understood that all combinations of the foregoing concepts (provided such concepts are not mutually inconsistent) are intended to be part of the inventive subject matter disclosed herein. In particular, all combinations of claimed subject matter appearing at the end of this disclosure are contemplated to be part of the inventive subject matter disclosed herein.

Claims

1. A method for identifying nucleic acid bases in a template polynucleotide, To provide a substrate containing multiple double-stranded template polynucleotides within a cluster, wherein each double-stranded template polynucleotide comprises a first strand and a second strand. The plurality of double-stranded template polynucleotides are brought into contact with a first primer that binds to the first strand and a second primer that binds to the second strand, The first primer and the second primer are extended by contacting the cluster with the labeled nucleic acid base to form the first labeled primer and the second labeled primer, Stimulating luminescence from the first and second labeled primers such that the amplitude of the signal generated by the first labeled primer is greater than the amplitude of the signal generated by the second labeled primer. A method comprising identifying the labeled nucleic acid bases attached to the first primer and the second primer based on the amplitude of the signal generated by the labeled nucleic acid bases.

2. The method according to claim 1, wherein the identification of the labeled nucleic acid base attached to the first primer and the identification of the labeled nucleic acid base attached to the second primer are performed substantially simultaneously.

3. The method according to claim 1 or 2, wherein the signal generated by the first labeled primer and the signal generated by the second labeled primer are emitted from the same region or substantially overlapping regions of the substrate.

4. The method according to claim 1, wherein the amplitude of the signal generated by the first labeled primer corresponds to a first amount of the first labeled primer in the cluster, and the amplitude of the signal generated by the second labeled primer corresponds to a second amount of the second labeled primer in the cluster.

5. The method according to claim 1, wherein contacting the plurality of double-stranded template polynucleotides with a first primer bound to the first strand and a second primer bound to the second strand includes contacting the first strand with an unblocked first primer and contacting the second strand with a predetermined fraction of the second primer having a blocked 3' end.

6. The method according to claim 1, comprising contacting the plurality of double-stranded template polynucleotides with a RecA-like protein or a non-nicking CRISPR-related protein to facilitate the binding of the plurality of double-stranded template polynucleotides to the first primer and the second primer.

7. The method according to claim 1, comprising contacting the plurality of double-stranded template polynucleotides with a helicase, a single-stranded DNA-binding protein, or a mixture of oligonucleotides having a random sequence to partially separate the first and second strands of each double-stranded template polynucleotide.

8. The signal generated by the first labeled primer is detected in a first range of optical frequencies and a second range of optical frequencies. The method includes detecting the signal generated by the second labeled primer in a first range of optical frequencies and a second range of optical frequencies, The method according to claim 1, wherein the first range of optical frequencies and the second range of optical frequencies are not the same.

9. To acquire a first fluorescence image of the cluster in a first range of optical frequencies, Acquiring a second fluorescence image of the cluster in a second range of optical frequencies, wherein the first range of optical frequencies and the second range of optical frequencies are not the same. The method according to claim 1, comprising obtaining the signals generated by the first and second labeled primers by extracting the fluorescence intensity from the first and second fluorescence images of the cluster.

10. The method according to claim 9, wherein the identification of the labeled nucleic acid bases attached to the first primer and the second primer is based on a combination of the extracted fluorescence intensities from the first and second fluorescence images.

11. The method according to claim 10, wherein the combination of identity of the labeled nucleic acid bases attached to the first primer and the second primer is classified as one of 16 combinations of nucleic acid base types based on the combination of extracted fluorescence intensities and a predetermined fluorescence intensity distribution for the 16 combinations of nucleic acid base types.

12. The extracted fluorescence intensity is normalized, The method according to claim 9, comprising classifying the combination of identity of the labeled nucleic acid bases attached to the first primer and the second primer as one of 16 combinations of nucleic acid base types based on the normalized extracted fluorescence intensity combination and a predetermined normalized fluorescence intensity distribution for the 16 combinations of nucleic acid base types.

13. The method according to claim 1, comprising stimulating fluorescence emission from the first labeled primer and the second labeled primer within the cluster with light at two predetermined optical frequencies.

14. The method according to claim 1, further comprising identifying whether the labeled nucleic acid base is associated with the first or second strand based on the amplitude of the signal generated by the labeled nucleic acid base.

15. A method for determining the sequence of a template polynucleotide, The hybridization involves hybridizing a first primer to the template polynucleotide and a second primer to the reverse complement of the template polynucleotide, wherein the template polynucleotide and its reverse complement are located in substantially overlapping regions of the substrate, and at least a portion of the reverse complement of the template polynucleotide hybridizes with a portion of the template polynucleotide. The first primer is extended with the first labeled nucleotide analog, The second primer is extended with a second labeled nucleotide analog, To stimulate luminescence from the first and second labeled nucleotide analogs, A method comprising determining the sequences of nucleotides in the template polynucleotide and the reverse complement of the template polynucleotide by capturing the luminescence.

16. The method according to claim 15, wherein the template polynucleotide and the reverse complement of the template polynucleotide are part of a cluster of identical copies of the template polynucleotide and identical copies of the reverse complement of the template polynucleotide.

17. The method according to claim 1 or 16, wherein the cluster is generated by bridge amplification.

18. The method according to claim 16, wherein the same copy of the template polynucleotide has a terminal bonded to the substrate by a first graft oligonucleotide, and the same copy of the reverse complement of the template polynucleotide has a terminal bonded to the substrate by a second graft oligonucleotide.

19. The method according to claim 15, wherein the first primer is part of a first group of first primers hybridized to the same copy of the template polynucleotide, and the second primer is part of a second group of second primers hybridized to the same copy of the reverse complement of the template polynucleotide.

20. Determining the sequence of the nucleotides is Receiving a first signal emitted with a first amplitude from a first group of the first primers, Receiving a second signal emitted with a second amplitude from a second group of the second primers, The method according to claim 19, comprising identifying the nucleic acid base hybridized to the template polynucleotide and the nucleic acid base hybridized to the reverse complement of the template polynucleotide based on the combination of the first and second signals.

21. The method according to claim 19, wherein the first group of primers has an unblocked 3' end, and the fraction of the second group of second primers has a blocked 3' end.

22. The method according to claim 5 or 21, wherein the blocked 3' end comprises a hairpin loop, a deoxynucleotide, a phosphate group, a propyl spacer, a modification that blocks a 3'-hydroxyl group, or an inverted nucleic acid base.

23. The method according to claim 15, wherein the first primer and the second primer each hybridize to the template polynucleotide and the reverse complement of the template polynucleotide in the same reaction step.

24. The method according to claim 15, wherein extending the first primer with the first labeled nucleotide analog and extending the second primer with the second labeled nucleotide analog are performed in the same reaction step.

25. The method according to claim 24, wherein the first labeled nucleotide analog and the second labeled nucleotide analog each hybridize to the template polynucleotide and the reverse complement of the template polynucleotide in the same reaction step.

26. The method according to claim 1 or 19, wherein the first primer and / or the second primer comprises locked nucleic acid (LNA) or peptide nucleic acid (PNA).

27. The method according to claim 15, wherein the hybridization of the first primer to the template polynucleotide and the hybridization of the second primer to the reverse complement of the template polynucleotide are facilitated by the presence of a RecA-like protein or a non-nicking CRISPR-related protein.

28. The method according to claim 1 or 19, wherein the extension of the first primer and the extension of the second primer are catalyzed by a chain-substituted polymerase.

29. The method according to claim 28, wherein the strand substitution polymerase comprises a Klenow fragment, phi29 DNA polymerase, Bsm DNA polymerase, Bst DNA polymerase, or a conserved variant thereof.

30. The method according to claim 15, wherein the template polynucleotide and the reverse complement of the template polynucleotide are at least partially separated by the presence of a helicase, a single-stranded DNA-binding protein, or a mixture of oligonucleotides having a random sequence.