Nucleic acid sequencing by emergence

The method of transient probe binding and optical imaging for nucleic acid sequencing addresses the inefficiencies of current technologies by enabling long read lengths with high accuracy and reduced costs and time.

JP7856276B2Active Publication Date: 2026-05-11XGENOMES CORP
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
XGENOMES CORP
Filing Date
2018-11-29
Publication Date
2026-05-11

AI Technical Summary

Technical Problem

Current nucleic acid sequencing methods face challenges in achieving long read lengths efficiently, cost-effectively, and accurately, with existing technologies either being too expensive, time-consuming, or suffering from low processing capacity and accuracy issues.

Method used

A method involving transient binding of molecular probes to nucleic acids on a test substrate, using oligonucleotide probes to form heteroduplexes, which are then imaged and measured to determine the sequence, allowing for high-resolution sequencing without the need for extensive sample preparation or reagent changes.

Benefits of technology

Enables long read lengths with high accuracy and efficiency, reducing sequencing costs and time by utilizing transient probe binding and optical imaging techniques.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007856276000001
    Figure 0007856276000001
  • Figure 0007856276000002
    Figure 0007856276000002
  • Figure 0007856276000003
    Figure 0007856276000003
Patent Text Reader

Abstract

A system and method for nucleic acid sequencing is provided. The nucleic acid is immobilized on a test substrate in a double-stranded, linearized, extended form and then denatured to single strands on the substrate to obtain adjacent immobilized first and second strands of the nucleic acid. The strands are exposed to each pool of oligonucleotide probes in a set of probes under conditions that allow the probes to form heteroduplexes with the corresponding complementary portions of the immobilized first or second strands, thereby generating each instance of optical activity. An imager measures the location and duration of this optical activity on the substrate. The exposure and measurement are repeated for the probes in the set of probes, thereby obtaining multiple sets of positions. The nucleic acid sequence is determined from the multiple sets of positions through compilation of the positions in the set. [Selection diagram] None
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] Cross-reference of related applications This application claims priority to U.S. Patent Application No. 62 / 591,850, titled “Sequencing by Emergence,” filed on 29 November 2017, which is incorporated herein by reference.

[0002] This disclosure generally relates to systems and methods for sequencing nucleic acids via the transient binding of probes to one or more polynucleotides. [Background technology]

[0003] background DNA sequencing was initially made possible by gel electrophoresis-based methods, namely dideoxy chain arrest (e.g., Sanger et al., Proc. Natl. Acad. Sci. 74:5463-5467, 1977) and chemical decomposition (e.g., Maxam et al., Proc. Natl. Acad. Sci. 74:560-564, 1977). Both of these methods for sequencing nucleotides were time-consuming and expensive. Nevertheless, the former, after spending hundreds of millions of dollars over a period of more than a decade, led to the sequencing of the first human genome.

[0004] As the dream of personalized medicine draws closer to reality, there is a growing demand for inexpensive, large-scale methods for sequencing individual human genomes (Mir, Sequencing Genomes: From Individuals to Populations, Briefings in Functional Genomics and Proteomics, 8: 367-378, 2009). Several sequencing methods that avoid gel electrophoresis (and are, secondly, less expensive) have been developed as “next-generation sequencing.” One such sequencing method using reversible terminators (implemented by Illumina Inc.) is the most promising. The most advanced form of Sanger sequencing, and currently the most promising detection method used in Illumina’s technology, involves fluorescence. Other possible means of detecting single nucleotide insertions include detection utilizing proton emission (e.g., via field-effect transistors, ion currents through nanopores, and electron microscopes). Illumina chemistry involves cyclic addition of nucleotides using reversible terminators (Canard et al., Metzker Nucleic Acids Research 22:4259-4267, 1994), and it supports fluorescent labeling (Bentley et al., Nature 456:53-59, 2008). Illumina sequencing starts with clonal amplification of a single genome molecule, substantial sample pretreatment is required to convert the target genome into a library, which is then clonally amplified as clusters.

[0005] However, two methods have since emerged that circumvent the requirement of pre-sequencing amplification. Both new methods perform fluorescence sequencing (SbS) by synthesis of single-molecule DNA. The first method, by HelicosBio (now SeqLL), performs stepwise SbS including reversible termination (Harris et al., Science, 320:106-9, 2008). The second method, SMRT sequencing by Pacific Biosciences, utilizes labeling on terminal phosphates, which are the innate leaving groups of the nucleotide incorporation reaction, allowing sequencing to be performed continuously without the need to change reagents. One drawback of this approach is its low processing capacity, as the detector must remain fixed in a single field of view (e.g., Levene et al., Science 299:682-686, 2003 and Eid Et al., Science, 323:133-8, 2009). A method somewhat similar to the sequencing approach used by PCI Bioscience, currently under development by Genia (now at Roche), involves detecting SbS via nanopores rather than optical methods.

[0006] The most commonly used sequencing methods have limited read lengths, increasing both the cost of sequencing and the difficulty of assembling the resulting reads. Sanger sequencing yields read lengths in the range of 1000 bases (e.g., Kchouk et al., Biol. Med. 9:395, 2017). Roche 454 sequencing and Ion Torrent both have read lengths in the range of several hundred bases. Illumina sequencing, which initially started with reads of about 25 bases, now typically produces 150–300 base pair reads. However, because each base in the read length requires the supply of fresh reagent, sequencing 250 bases rather than 25 bases requires 10 times longer processing time and 10 times more expensive reagents. In recent years, the standard read length for Illumina instruments has been reduced to approximately 150 bases, likely because longer reads are susceptible to phasing (where molecules within a cluster become out of synchronization), which can lead to errors in the technique.

[0007] The longest read lengths achievable with commercially available systems are obtained through nanopore chain sequencing from Oxford Nanopores Technology (ONT) and sequencing from Pacific Bioscience (PacBio) (e.g., Kchouk et al., Biol. Med. 9:395, 2017). The latter routinely yields reads of approximately 10,000 base pairs on average, while the former, though very rarely, can yield reads of several hundred kilobase pairs (e.g., Laver et al., Biomol. Det. Quant. 3:1-8, 2015). While these long read lengths are desirable for alignment, they come at the cost of accuracy. Due to the often very low accuracy, these methods cannot be used as standalone sequencing technologies for most human sequencing applications, although they can be used as aids to Illumina sequencing. Furthermore, the processing power of existing long-read technologies is too low for routine human genome-scale sequencing.

[0008] Besides ONT and PacBio sequencing, there are several other approaches that are not sequencing techniques in themselves, but are sample preparation approaches that provide a scaffold for constructing long reads to assist Illumina short-read sequencing techniques. One of these is a droplet-based technique developed by 10X Genomics, which isolates 100-200kb fragments (e.g., the average fragment length range after extraction) within a droplet, processes them into a library of shorter fragments, and each fragment contains a sequence identification tag specific to the 100-200kb from which it originates, which can be deconvolved into approximately 50-200kb buckets when sequencing a genome from multiple droplets (Goodwin et al., Nat. Rev. Genetics 17:333-351, 2016). Another approach, developed by Biono Genomics, involves extending DNA by exposure to nicking endonucleases to introduce nicks into the DNA. The method fluorescently detects nicking points to provide a molecular map or scaffold. Currently, this method has not been developed to have a density high enough to assist in genome assembly, but it still provides direct visualization of the genome, can detect large structural variations, and can determine long-range haplotypes.

[0009] Despite the development of different sequencing methods and the general trend toward reducing sequencing costs, the size of the human genome continues to result in high sequencing costs for patients. Each human genome is organized into 46 chromosomes, ranging from approximately 50 megabases to 250 megabases. NGS sequencing methods still face numerous challenges that impact performance, including reliance on reference genomes, which can substantially increase the time required for analysis (discussed, e.g., Kulkarni et al., Comput Struct Biotechnol J. 15:471-477, 2017).

[0010] Given the above background, what is needed in this field of technology are devices, systems, and methods for providing standalone sequencing techniques that are efficient in terms of reagent use and time, and that provide long haplotype degradation reads without compromising accuracy.

[0011] The information disclosed in the background section is solely for the purpose of advancing the understanding of the general background, and should not be considered as an acknowledgment or suggestion that this information constitutes prior art already known to those skilled in the art. [Overview of the project] [Problems that the invention aims to solve]

[0012] overview This disclosure addresses the demand in the art for devices, systems, and methods to provide improved nucleic acid sequencing technologies. In one broad embodiment, this disclosure includes a method for identifying at least one unit of a multi-unit molecule by conjugating a molecular probe to one or more units of the molecule. This disclosure is based on the detection of single-molecule interactions between one or more species of molecular probes and molecules. In some embodiments, the probe transiently conjugates to at least one unit of the molecule. In some embodiments, the probe repeatedly conjugates to at least one unit of the molecule. In some embodiments, the molecular entity is located with nanometer accuracy on a polymer, surface, or matrix. [Means for solving the problem]

[0013] In one embodiment, a method for sequencing nucleic acids is disclosed herein. The method comprises (a) immobilizing a nucleic acid in a double-stranded linearized extension form onto a test substrate to form an immobilized extension double-stranded nucleic acid. The method further comprises (b) denaturing the immobilized extension double-stranded nucleic acid into a single-stranded form on the test substrate to obtain an immobilized first strand and an immobilized second strand of nucleic acid, where each base of the immobilized second strand is positioned adjacent to the corresponding complementary base of the immobilized first strand. The method then comprises (c) exposing the immobilized first strand and the immobilized second strand to each pool of each oligonucleotide probe in a set of oligonucleotide probes, where each oligonucleotide probe in the set of oligonucleotide probes is of a predetermined sequence and length. Exposure (c) occurs under conditions that cause each individual probe in each pool of each oligonucleotide probe to bind to a portion of the immobilized first strand or immobilized second strand complementary to each oligonucleotide probe and to form each heteroduplex, thereby producing each instance of optical activity. The method then proceeds to (d) measure the location and duration on the test substrate of each instance of optical activity occurring during exposure (c) using a two-dimensional imager. The method then proceeds to (e) repeat the exposure (c) and measurement (d) for each oligonucleotide probe in the set of oligonucleotide probes, thereby obtaining multiple sets of locations on the test substrate. Each set of locations on the test substrate corresponds to one oligonucleotide probe in the set of oligonucleotide probes. The method further includes (f) determining the sequence of at least a portion of the nucleic acids from the multiple sets of locations on the test substrate by compiling the locations on the test substrate represented by the multiple sets of locations.

[0014] In some embodiments, exposure (c) occurs under conditions that transiently and reversibly bind each individual probe from each pool of each oligonucleotide probe to a portion of a first or second immobilized chain complementary to the individual probe, thereby forming each heteroduplex and producing an instance of optical activity. In some embodiments, exposure (c) occurs repeatedly under conditions that transiently and reversibly bind each individual probe from each pool of each oligonucleotide probe to a portion of a first or second immobilized chain complementary to the individual probe, thereby forming each heteroduplex and producing each instance of optical activity. In some such embodiments, each oligonucleotide probe in the set of oligonucleotide probes is bound to a label (e.g., a dye, fluorescent nanoparticles, or light-scattering particles).

[0015] In some embodiments, the exposure method of claim 1 is in the presence of a first label in the form of an intercalating dye, where each oligonucleotide probe in a set of oligonucleotide probes is conjugated with a second label, and the first and second labels have overlapping donor emission spectra and acceptor excitation spectra, which cause one of the first and second labels to fluoresce when the first and second labels are close to each other, and each instance of optical activity originates near the intercalating dye, intercalating each heteroduplex between the oligonucleotide and the immobilized first or immobilized second chain to the second label.

[0016] In some embodiments, exposure is in the presence of a first label in the form of an intercalating dye, where each oligonucleotide probe in a set of oligonucleotide probes is bound to a second label, the first label causing the second label to fluoresce when the first and second labels are in close proximity to each other, and each optically active instance originates near the intercalating dye, intercalating each heteroduplex between the oligonucleotide and the immobilized first or immobilized second chain to the second label.

[0017] In some embodiments, exposure is in the presence of a first label in the form of an intercalating dye, where each oligonucleotide probe in a set of oligonucleotide probes is bound to a second label, the second label causing the first label to fluoresce when the first and second labels are close to each other, and each optically active instance originates near the intercalating dye, intercalating each heteroduplex between the oligonucleotide and the immobilized first or immobilized second chain to the second label.

[0018] In some embodiments, exposure is in the presence of an intercalating dye, and each instance of optical activity is derived from the fluorescence of the intercalating dye intercalating each heteroduplex between the oligonucleotide and the immobilized first or immobilized second chain. In such embodiments, each instance of optical activity is greater than the fluorescence of the intercalating dye before intercalating each heteroduplex.

[0019] In some embodiments, more than one oligonucleotide probe in a set of oligonucleotide probes is exposed to a first strand and a second strand immobilized during a single instance of exposure (c), and each different oligonucleotide probe of the set of oligonucleotide probes exposed to the first strand and the second strand immobilized during a single instance of exposure (c) is associated with a different label. In some such embodiments, in a set of oligonucleotide probes where a first oligonucleotide probe is associated with a first label, a first pool is exposed to a first strand and a second strand immobilized during a single instance of exposure (c), and in a set of oligonucleotide probes where a second oligonucleotide probe is associated with a second label, a second pool is exposed to a first strand and a second strand immobilized during a single instance of exposure (c), and the first label and the second label are different. Alternatively, in a set of oligonucleotide probes where a first oligonucleotide probe is associated with a first label, a first pool is exposed to a first strand and a second strand immobilized during a single instance of exposure (c), and in a set of oligonucleotide probes where a second oligonucleotide probe is associated with a second label, a second pool is exposed to a first strand and a second strand immobilized during a single instance of exposure (c), and in a set of oligonucleotide probes where a third oligonucleotide probe is associated with a third label, a third pool is exposed to a first strand and a second strand immobilized during a single instance of exposure (c), and the first label, the second label, and the third label are each different.

[0020] In some embodiments, iteration (e), exposure (c), and measurement (d) are each performed for each single oligonucleotide probe in a set of oligonucleotide probes.

[0021] In some embodiments, exposure (c) is performed at a first temperature for a first oligonucleotide probe in a set of oligonucleotide probes, and iterations (e), exposure (c), and measurement (d) include performing exposure (c) and measurement (d) for the first oligonucleotide at a second temperature.

[0022] In some embodiments, exposure (c) is performed at a first temperature for a first oligonucleotide probe in a set of oligonucleotide probes, and instances of iterations (e), exposure (c), and measurement (d) include performing exposure (c) and measurement (d) for the first oligonucleotide at each of a plurality of different temperatures. The method further includes constructing a melting curve for the first oligonucleotide probe using the measured locations and durations of optical activity recorded by measurement (d) for the first temperature and each of the plurality of different temperatures.

[0023] In some embodiments, the set of oligonucleotide probes comprises multiple subsets of oligonucleotide probes, and replication (e), exposure (c), and measurement (d) are performed for each respective subset of oligonucleotide probes in the multiple subsets of oligonucleotide probes. In some such embodiments, each respective subset of oligonucleotide probes comprises two or more probes different from the set of oligonucleotide probes. Alternatively, each respective subset of oligonucleotide probes comprises four or more probes different from the set of oligonucleotide probes. In some such embodiments, the set of oligonucleotide probes consists of four subsets of oligonucleotide probes. In some embodiments, the method further comprises dividing the set of oligonucleotide probes into multiple subsets of oligonucleotide probes based on the melting temperature of each oligonucleotide probe calculated or experimentally derived, wherein oligonucleotide probes having similar melting temperatures are placed in the same subset of oligonucleotide probes by the division, and the temperature or duration of the instance of exposure (c) is determined by the average melting temperature of the oligonucleotide probes in the corresponding subset of oligonucleotide probes. Furthermore, in some embodiments, the method further includes dividing a set of oligonucleotide probes into multiple subsets of oligonucleotide probes based on the sequence of each oligonucleotide probe, thereby placing oligonucleotide probes having overlapping sequences into different subsets.

[0024] In some embodiments, measuring a location on a test substrate involves identifying and fitting each instance of optical activity with a fitting function to identify and fit the center of each instance of optical activity in a frame of data obtained by a two-dimensional imager, where the center of each instance of optical activity is considered to be the location of each instance of optical activity on the test substrate. In some such embodiments, the fitting function is a Gaussian function, a first moment function, a gradient-based approach, or a Fourier transform.

[0025] In some embodiments, each instance of optical activity persists across multiple frames measured by a two-dimensional imager, and the measurement of the location on the test substrate involves identifying and fitting each instance of optical activity across multiple frames using a fitting function to identify the center of each instance of optical activity across multiple frames, where the center of each instance of optical activity is considered to be the location of each instance of optical activity on the test substrate across multiple frames. In some such embodiments, the fitting function is a Gaussian function, a first moment function, a gradient-based approach, or a Fourier transform.

[0026] In some embodiments, measuring a location on a test substrate involves inputting a frame of data measured by a two-dimensional imager into a trained convolutional neural network, where the frame of data includes each instance of optical activity among multiple instances of optical activity, and each instance of optical activity among multiple instances of optical activity corresponds to an individual probe coupled to a fixed first chain or a fixed second chain, and in response to the input, the trained convolutional neural network identifies the location on the test substrate of one or more instances of optical activity among multiple instances of optical activity.

[0027] In some embodiments, the measurement resolves the center of each instance of optical activity to a position on the test substrate with a positioning accuracy of at least 20 nm, at least 2 nm, at least 60 nm, or at least 6 nm.

[0028] In some embodiments, the measurement resolves the center of each instance of optical activity to a position on the test substrate, which is a subdiffraction-limited position.

[0029] In some embodiments, the location and duration measurements (d) of each instance of optical activity on the test substrate are measured to exceed 5,000 photons at that location, exceed 50,000 photons at that location, or exceed 200,000 photons in that case.

[0030] In some embodiments, each instance of optical activity is greater than a predetermined number of standard deviations (e.g., greater than 3, 4, 5, 6, 7, 8, 9, or 10) above the background observed on the test substrate.

[0031] In some embodiments, each oligonucleotide probe in a plurality of oligonucleotide probes contains a unique N-mer sequence, where N is an integer in the set {1, 2, 3, 4, 5, 6, 7, 8, and 9}, and all unique N-mer sequences of length N are represented by the plurality of oligonucleotide probes. In some such embodiments, the unique N-mer sequence contains one or more nucleotide positions occupied by one or more degenerate nucleotides. In some such embodiments, each degenerate nucleotide position in one or more nucleotide positions is occupied by a universal base (e.g., 2'-deoxyinosine). In some such embodiments, the unique N-mer sequence has a single degenerate nucleotide position adjacent to the 5' side and a single degenerate nucleotide position adjacent to the 3' side. Alternatively, the single degenerate nucleotide on the 5' side and the single degenerate nucleotide on the 3' side are each 2'-deoxyinosine.

[0032] In some embodiments, the nucleic acid is at least 140 base pairs long, and determination (f) determines the sequence coverage of more than 70% of the nucleic acid sequence. In some embodiments, the nucleic acid is at least 140 base pairs long, and determination (f) determines the sequence coverage of more than 90% of the nucleic acid sequence. In some embodiments, the nucleic acid is at least 140 base pairs long, and determination (f) determines the sequence coverage of more than 99% of the nucleic acid sequence. In some embodiments, determination (f) determines the sequence coverage of more than 99% of the nucleic acid sequence.

[0033] In some embodiments, the nucleic acid is at least 10,000 base pairs long or at least 1,000,000 base pairs long.

[0034] In some embodiments, the test substrate is cleaned before repeating exposure (c) and measurement (d), thereby removing each oligonucleotide probe from the test substrate before exposing the test substrate to another oligonucleotide probe in the set of oligonucleotide probes.

[0035] In some embodiments, immobilization (a) includes coating a nucleic acid onto a test substrate by molecular combing (receding meniscus), flow-stretching nanoconfinement, or electrostretching.

[0036] In some embodiments, each instance of optical activity has an observation metric that satisfies a predetermined threshold. In some such embodiments, the observation metric includes duration, signal-to-noise ratio, photon count, or intensity. In some embodiments, the predetermined threshold distinguishes between (i) a first form of binding, where each residue of the unique N-mer sequence binds to a complementary base in the immobilized first or second strand of nucleic acid, and (ii) a second form of binding, where there is at least one mismatch between the unique N-mer sequence and the sequence in the immobilized first or second strand of nucleic acid to which each oligonucleotide probe binds to form each instance of optical activity.

[0037] In some embodiments, each oligonucleotide probe in a set of oligonucleotide probes has its own corresponding predetermined threshold. In some such embodiments, the predetermined threshold for each oligonucleotide probe in a set of oligonucleotide probes is derived from a training dataset. For example, in some embodiments, the predetermined threshold for each oligonucleotide probe in a set of oligonucleotide probes is derived from a training dataset, and the training set includes, for each oligonucleotide probe in a set of oligonucleotide probes, a measure of the observational metric when bound to a reference sequence, such that each residue of the unique N-mer sequence of each oligonucleotide probe binds to a complementary base in the reference sequence. In some such embodiments, the reference sequence is immobilized on a reference substrate. Alternatively, the reference sequence is included with the nucleic acid and immobilized on a test substrate. In some embodiments, the reference sequence includes all or part of the genome of PhiX174, M13, lambda phage, T7 phage, or E. coli, budding yeast, or fission yeast. In some embodiments, the reference sequence is a synthetic construct of a known sequence. In some embodiments, the reference sequence includes all or part of rabbit globin RNA.

[0038] In some embodiments, each oligonucleotide probe in a set of oligonucleotide probes produces a first instance of optical activity by binding to a complementary portion of a fixed first chain, and a second instance of optical activity by binding to a complementary portion of a fixed second chain.

[0039] In some embodiments, each oligonucleotide probe in a set of oligonucleotide probes produces two or more first instances of optical activity by binding to two or more complementary moieties of a fixed first chain, and two or more second instances of optical activity by binding to two or more complementary moieties of a fixed second chain.

[0040] In some embodiments, each oligonucleotide probe binds three or more times to a portion of an immobilized first or second chain complementary to each oligonucleotide probe during exposure (c), thereby obtaining three or more instances of optical activity, each instance of optical activity representing one binding event among multiple binding events.

[0041] In some embodiments, each oligonucleotide probe binds five or more times to a portion of a complementary immobilized first or second chain during exposure (c), thereby yielding five or more instances of optical activity, each instance of optical activity representing one binding event among multiple binding events.

[0042] In some embodiments, each oligonucleotide probe binds to a portion of a complementary immobilized first or second chain to each oligonucleotide probe more than 10 times during exposure (c), thereby obtaining more than 10 instances of optical activity, each instance of optical activity representing one binding event among multiple binding events.

[0043] In some embodiments, exposure (c) occurs for a period of 5 minutes or less, 2 minutes or less, or 1 minute or less.

[0044] In some embodiments, exposure (c) occurs over one or more frames of a two-dimensional imager, over two or more frames of a two-dimensional imager, over 500 or more frames of a two-dimensional imager, or over 5000 or more frames of a two-dimensional imager.

[0045] In some embodiments, exposure (c) is performed for a first oligonucleotide probe in a set of oligonucleotide probes for a first period, and repetition (e), exposure (c), and measurement (d) comprises performing exposure (c) for a second oligonucleotide for a second period, where the first period is longer than the second period.

[0046] In some embodiments, exposure (c) is performed for a first oligonucleotide probe in a set of oligonucleotide probes for a first number of frames of a two-dimensional imager, and repetition (e), exposure (c), and measurement (d) comprises performing exposure (c) for a second oligonucleotide for a second number of frames of the two-dimensional imager, where the first number of frames is greater than the second number of frames.

[0047] In some embodiments, each oligonucleotide probe in a set of oligonucleotide probes is of the same length.

[0048] In some embodiments, each oligonucleotide probe in a set of oligonucleotide probes is of the same length M, where M is a positive integer greater than or equal to 2 (for example, M is 2, 3, 4, 5, 6, 7, 8, 9, 10, or greater than 10), and the determination (f) of the sequence of at least a portion of the nucleic acid from a set of multiple positions on a test substrate further utilizes the overlapping sequences of the oligonucleotide probes represented by the set of multiple positions. In some such embodiments, each oligonucleotide probe in a set of oligonucleotide probes shares M-1 sequence homology with another oligonucleotide probe in a set of oligonucleotide probes. In some such embodiments, determining the sequence of at least a portion of the nucleic acid from a set of multiple positions on a test substrate includes determining a first tiling path corresponding to a fixed first strand and a second tiling path corresponding to a fixed second strand. In some such embodiments, a break in the first tiling path is resolved using the corresponding portion of the second tiling path, or a break in the first or second tiling path is resolved using a reference sequence. Alternatively, a discontinuity in the first or second tiling path is resolved using a corresponding portion of a third or fourth tiling path obtained from another instance of the nucleic acid. In some such embodiments, the confidence in sequence assignment is increased using the corresponding portions of the first and second tiling paths. Alternatively, the confidence in sequence assignment is increased using a corresponding portion of a third or fourth tiling path obtained from another instance of the nucleic acid.

[0049] In some embodiments, the duration of the exposure (c) instance is determined by the estimated melting temperature of each alkyl group probe in the set of alkyl group probes used in the exposure (c) instance.

[0050] In some embodiments, the method further comprises (f) exposing an immobilized double strand or an immobilized first strand and an immobilized second strand to an antibody, affimer, nanomodifier, aptamer, or methyl-binding protein to determine modifications to the nucleic acid from multiple sets of positions on a test substrate, or to correlate them with the sequence of a portion of the nucleic acid.

[0051] In some embodiments, the test substrate is a two-dimensional surface. In some such embodiments, the two-dimensional surface is coated with a gel or matrix.

[0052] In some embodiments, the test substrate is a cell, a three-dimensional matrix, or a gel.

[0053] In some embodiments, the test substrate is bound to a sequence-specific oligonucleotide probe before immobilization (a), and immobilization (a) includes capturing nucleic acids on the test substrate using the sequence-specific oligonucleotide probe bound to the test substrate.

[0054] In some embodiments, the nucleic acid is in a solution containing additional cellular components, and fixation (a) or denaturation (b) further includes washing the test substrate after the nucleic acid has been fixed onto the test substrate and before exposure (c), thereby separating and purifying the additional cellular components from the nucleic acid.

[0055] In some embodiments, the test substrate is passivated prior to exposure (c) with polyethylene glycol, bovine serum albumin-biotin-streptavidin, casein, bovine serum albumin (BSA), one or more different tRNAs, one or more different deoxyribonucleotides, one or more different ribonucleotides, salmon sperm DNA, Pluronic F-127, Tween-20, hydrogen silsesquioxane (HSQ), or any combination thereof.

[0056] In some embodiments, the test substrate is coated with a vinylsilane coating containing 7-octenyltrichlorosilane before fixation (a).

[0057] Another aspect of the present disclosure is a method for sequencing a nucleic acid, comprising: (a) immobilizing the nucleic acid on a test substrate in a linearized extension form, thereby forming an immobilized extension nucleic acid; and (b) exposing the immobilized extension nucleic acid to each pool of each oligonucleotide probe in a set of oligonucleotide probes, where each oligonucleotide probe in the set of oligonucleotide probes has a predetermined sequence and length, and the exposure (b) occurs transiently and reversibly to each portion of the immobilized nucleic acid complementary to each oligonucleotide probe, under conditions that enable the individual probes in each pool of each oligonucleotide probe, thereby producing each instance of optical activity. The method provides a method comprising: (c) measuring the location and duration on a test substrate of each instance of optical activity occurring during exposure (b) using a two-dimensional imager; (d) repeating exposure (b) and measurement (c) for each oligonucleotide probe in a set of oligonucleotide probes to obtain a plurality of sets of positions on the test substrate, where each set of positions on the test substrate corresponds to one oligonucleotide probe in the set of oligonucleotide probes; and (e) determining the sequence of at least a portion of the nucleic acid from the plurality of sets of positions on the test substrate by compiling the positions on the test substrate represented by the plurality of sets of positions. In some such embodiments, the nucleic acid is a double-stranded nucleic acid, and the method further comprises denaturing the immobilized double-stranded nucleic acid into a single-stranded form on a test substrate to obtain an immobilized first strand and an immobilized second strand of nucleic acid, where the immobilized second strand is complementary to the immobilized first strand. In some embodiments, the nucleic acid is single-stranded RNA.

[0058] Another aspect of the present disclosure provides a method for analyzing nucleic acids, comprising: (a) immobilizing nucleic acids in a double-stranded form on a test substrate to form immobilized double-stranded nucleic acids; (b) denaturing the immobilized double-stranded nucleic acids in a single-stranded form on a test substrate to obtain an immobilized first strand and an immobilized second strand of nucleic acid, wherein the immobilized second strand is complementary to the immobilized first strand; and (c) exposing the immobilized first strand and the immobilized second strand to one or more oligonucleotide probes to determine whether one or more oligonucleotide probes bind to the immobilized first strand or to the immobilized second strand. [Brief explanation of the drawing]

[0059] [Figure 1A] Various embodiments of this disclosure collectively represent exemplary system topologies, each including a polymer containing multiple probes that participate in binding events, and a computer storage medium that retrieves and stores information related to the localization and sequence identification of binding events, and subsequently performs further analysis to determine the polymer sequence. [Figure 1B] Various embodiments of this disclosure collectively represent exemplary system topologies, each including a polymer containing multiple probes that participate in binding events, and a computer storage medium that retrieves and stores information related to the localization and sequence identification of binding events, and subsequently performs further analysis to determine the polymer sequence. [Figure 2A] Various embodiments of this disclosure collectively provide flowcharts of steps and features of methods for determining the arrangement and / or structural characteristics of a target polymer. [Figure 2B] Various embodiments of this disclosure collectively provide flowcharts of steps and features of methods for determining the arrangement and / or structural characteristics of a target polymer. [Figure 3] Various embodiments of this disclosure provide flowcharts of steps and features of additional methods for determining the arrangement and / or structural characteristics of a target polymer. [Figure 4]Various embodiments of this disclosure provide flowcharts of steps and features of additional methods for determining the arrangement and / or structural characteristics of a target polymer. [Figure 5A] Various embodiments of this disclosure collectively illustrate examples of transient probe binding to polynucleotides. [Figure 5B] Various embodiments of this disclosure collectively illustrate examples of transient probe binding to polynucleotides. [Figure 5C] Various embodiments of this disclosure collectively illustrate examples of transient probe binding to polynucleotides. [Figure 6A] Various embodiments of this disclosure collectively illustrate examples of k-mer probes of different lengths that bind to target polynucleotides. [Figure 6B] Various embodiments of this disclosure collectively illustrate examples of k-mer probes of different lengths that bind to target polynucleotides. [Figure 7A] Various embodiments of this disclosure collectively illustrate examples of using reference oligonucleotides having a continuous periodicity of oligonucleotide sets. [Figure 7B] Various embodiments of this disclosure collectively illustrate examples of using reference oligonucleotides having a continuous periodicity of oligonucleotide sets. [Figure 7C] Various embodiments of this disclosure collectively illustrate examples of using reference oligonucleotides having a continuous periodicity of oligonucleotide sets. [Figure 8A] Various embodiments of this disclosure collectively illustrate examples of applying different sets of probes to a single reference molecule. [Figure 8B] Various embodiments of this disclosure collectively illustrate examples of applying different sets of probes to a single reference molecule. [Figure 8C] Various embodiments of this disclosure collectively illustrate examples of applying different sets of probes to a single reference molecule. [Figure 9A] Various embodiments of this disclosure collectively illustrate examples of transient coupling when multiple types of probes are used. [Figure 9B] Various embodiments of this disclosure collectively illustrate examples of transient coupling when multiple types of probes are used. [Figure 9C] Various embodiments of this disclosure collectively illustrate examples of transient coupling when multiple types of probes are used. [Figure 10A] Various embodiments of this disclosure collectively illustrate examples in which the number of transient coupling events collected correlates with the degree of probe localization that can be achieved. [Figure 10B] Various embodiments of this disclosure collectively illustrate examples in which the number of transient coupling events collected correlates with the degree of probe localization that can be achieved. [Figure 11A] The various embodiments of this disclosure collectively illustrate examples of tiling probes. [Figure 11B] The various embodiments of this disclosure collectively illustrate examples of tiling probes. [Figure 12A] Various embodiments of this disclosure collectively illustrate examples of transient binding of directly labeled probes. [Figure 12B] Various embodiments of this disclosure collectively illustrate examples of transient binding of directly labeled probes. [Figure 12C] Various embodiments of this disclosure collectively illustrate examples of transient binding of directly labeled probes. [Figure 13A] Various embodiments of this disclosure collectively illustrate examples of transient probe binding in the presence of intercalating dyes. [Figure 13B] Various embodiments of this disclosure collectively illustrate examples of transient probe binding in the presence of intercalating dyes. [Figure 13C] Various embodiments of this disclosure collectively illustrate examples of transient probe binding in the presence of intercalating dyes. [Figure 14A] Various embodiments of this disclosure collectively illustrate examples of different probe labeling techniques. [Figure 14B]Various embodiments of this disclosure collectively illustrate examples of different probe labeling techniques. [Figure 14C] Various embodiments of this disclosure collectively illustrate examples of different probe labeling techniques. [Figure 14D] Various embodiments of this disclosure collectively illustrate examples of different probe labeling techniques. [Figure 14E] Various embodiments of this disclosure collectively illustrate examples of different probe labeling techniques. [Figure 15] Various embodiments of this disclosure collectively illustrate examples of transient binding of probes to denatured and examined double-stranded DNA. [Figure 16A] The various embodiments of this disclosure collectively illustrate examples of cell lysis and nucleic acid fixation and extension. [Figure 16B] The various embodiments of this disclosure collectively illustrate examples of cell lysis and nucleic acid fixation and extension. [Figure 17] Various embodiments of this disclosure illustrate exemplary microfluidic structures that capture single cells and optionally provide nucleic acid extraction, extension, and sequencing from the cells. [Figure 18] Various embodiments of this disclosure illustrate exemplary microfluidic structures that provide different ID tags to individual cells. [Figure 19] Various embodiments of this disclosure illustrate examples of sequencing polynucleotides from individual cells. [Figure 20A] Various embodiments of this disclosure collectively illustrate exemplary device layouts for performing transient probe-coupled imaging. [Figure 20B] Various embodiments of this disclosure collectively illustrate exemplary device layouts for performing transient probe-coupled imaging. [Figure 21] Various embodiments of this disclosure illustrate exemplary capillary tubing containing reagents separated by an air gap. [Figure 22A] Various embodiments of this disclosure collectively illustrate examples of fluorescence. [Figure 22B] Various embodiments of this disclosure collectively illustrate examples of fluorescence. [Figure 22C] Various embodiments of this disclosure collectively illustrate examples of fluorescence. [Figure 22D] Various embodiments of this disclosure collectively illustrate examples of fluorescence. [Figure 22E] Various embodiments of this disclosure collectively illustrate examples of fluorescence. [Figure 23A] Various embodiments of this disclosure collectively illustrate examples of fluorescence. [Figure 23B] Various embodiments of this disclosure collectively illustrate examples of fluorescence. [Figure 23C] Various embodiments of this disclosure collectively illustrate examples of fluorescence. [Figure 24] Various embodiments of this disclosure demonstrate transient binding to synthesized denatured double-stranded DNA. [Modes for carrying out the invention]

[0060] Detailed description Embodiments are described in detail here, and these embodiments are shown in the accompanying drawings. Numerous specific details are provided in the following detailed description to provide a thorough understanding of the disclosure. However, it will be apparent to those skilled in the art that the disclosure can be practiced without these specific details. In other examples, well-known methods, procedures, components, circuits, and networks are not described in detail so as not to unnecessarily obscure the aspects of the embodiments.

[0061] definition The terminology used in this disclosure is intended solely to describe specific embodiments and is not intended to limit the invention. The singular forms “a,” “an,” and “the” used in the specification and appended claims also include the plural forms unless otherwise explicitly indicated in the context. It will also be understood that the terms “and / or” as used herein refer to and encompass one or any possible combination of the terms listed in relation. It will further be understood that the terms “encompassing” and / or “containing,” as used herein, indicate the presence of the described features, integers, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0062] As used herein, the term "if" may be interpreted, depending on the context, as meaning "if," "when," "in response to a decision," or "in response to detection." Similarly, the phrase "if determined" or "[the stated condition or event] is detected" may be interpreted, depending on the context, as meaning "at the time of a decision," "in response to a decision," "[the state of] detection," or "[the state of] detection."

[0063] The term "or" shall mean inclusive "or" rather than exclusive "or." That is, unless otherwise specified or as is evident from the context, the phrase "X uses A or B" shall mean any of the natural inclusive arrangements. That is, the phrase "X uses A or B" fits any of the following examples: X uses A; X uses B; or X uses both A and B. In addition, the articles "a" and "an" used in this application and the attached claims shall generally be interpreted as meaning "one or more" unless otherwise specified or as is evident from the context directed toward the singular.

[0064] While terms such as "first," "second," etc., may be used herein to describe various elements, it should also be understood that these elements should not be limited by these terms. These terms are used merely to distinguish one element from another. For example, without departing the scope of this disclosure, the first filter may be referred to as the second filter, and similarly, the second filter may be referred to as the first filter. The first filter and the second filter are both filters, but they are not the same filter.

[0065] As used herein, the terms “about” or “approximately” may mean an acceptable range of error for a particular value as determined by those skilled in the art, which may in part depend on how the value is measured or determined, for example, the limits of the measuring system. For example, “about” may mean a standard deviation of 1 or more than 1 per practice in the art. “About” may mean a range of ±20%, ±10%, ±5%, or ±1% of a given value. The terms “about” or “approximately” may mean within 10 times, 5 times, or 2 times a value. If a particular value is described in this application and claims, unless otherwise specified, the term “about” should be inferred to mean an acceptable range of error for the particular value. The term “about” may have a meaning that is generally understood by those skilled in the art. The term “about” may mean ±10%. The term “about” may mean ±5%.

[0066] As used herein, the terms “nucleic acid,” “nucleic acid molecule,” and “polynucleotide” are interchangeable. These terms refer to nucleic acids in any compositional form, including deoxyribonucleic acid (DNA, e.g., complementary DNA (cDNA), genomic DNA (gDNA), etc.), ribonucleic acid (RNA, e.g., messenger RNA (mRNA), small inhibitory RNA (siRNA), ribosomal RNA (rRNA), transfer RNA (tRNA), microRNA (mRNA), RNA highly expressed by the fetus or placenta, etc.), and / or DNA or RNA analogues (e.g., including base analogues, sugar analogues, and / or non-native backbone, etc.), RNA / DNA hybrids, and polyamide nucleic acids (PNA), all of which may be single-stranded or double-stranded. Unless otherwise specified, nucleic acids may contain known analogues of natural nucleotides, some of which may function in a manner similar to naturally occurring nucleotides. Nucleic acids can be in any form useful for carrying out the steps herein (e.g., linear, circular, supercoiled, single-stranded, double-stranded, etc.). In some examples, nucleic acids are plasmids, phages, autonomously replicating sequences (ARS), centromeres, artificial chromosomes, chromosomes, or other nucleic acids that can or may be replicated in vitro or in host cells, cells, cell nuclei, or the cytoplasm of cells, in certain embodiments. In some embodiments, nucleic acids may be derived from a single chromosome or a fragment thereof (e.g., a nucleic acid sample derived from one chromosome of a sample obtained from a diploid organism). Nucleic acid molecules may include the full length of a native polynucleotide (e.g., a long non-coding (lnc)RNA, mRNA, chromosome, mitochondrial DNA, or polynucleotide fragment). The polynucleotide fragment must be at least 200 nucleotides long, preferably at least several thousand nucleotides long. More preferably, in the case of genomic DNA, the polynucleotide fragment will be several hundred kilobases to several megabases long.

[0067] In certain embodiments, nucleic acids include nucleosomes, fragments or portions of nucleosomes, or nucleosome-like structures. Sometimes, nucleic acids include proteins (e.g., histones, DNA-binding proteins). Nucleic acids analyzed by the processes described herein are sometimes substantially isolated and not substantially associated with proteins or other molecules. Nucleic acids also include derivatives, variants, and analogues of RNA or DNA synthesized, replicated, or amplified from single-stranded ("sense" or "antisense," "plus" or "-" strands, "forward" reading frame or "reverse" reading frame) and double-stranded polynucleotides. Deoxyribonucleotides include deoxyadenosine, deoxycytidine, deoxyguanosine, and deoxythymidine. In the case of RNA, the base cytosine is replaced with uracil, and the 2' position of the sugar contains a hydroxyl moiety. In some embodiments, nucleic acids are prepared using nucleic acids obtained from a subject as a template.

[0068] As used herein, the terms “endpoint” or “terminus” (or simply “end”) may refer to the genomic coordinates or genomic identity or nucleotide identity of the outermost base at, for example, the tip of a cell-free DNA molecule, such as a plasma DNA molecule. The terminus may correspond to any end of the DNA molecule. In this form, when referring to the origin and end of a DNA molecule, both may correspond to the endpoint. In some embodiments, one terminus is the genomic coordinates or nucleotide identity of the outermost base at one end of a cell-free DNA molecule detected or determined by an analytical method, such as massively parallel sequencing or next-generation sequencing, single-molecule sequencing, double-stranded or single-stranded DNA sequencing library preparation protocol, polymerase chain reaction (PCR), or microarray. In some embodiments, such in vitro techniques may alter the true in vivo physical end(s) of the cell-free DNA molecule. Therefore, each detectable end may represent a biologically true end, or the end may be one or more nucleotides inward from the original end of the molecule, or one or more nucleotides extended from the end, e.g., the 5' blunting and 3' filling of the overhang of a non-blunt end double-stranded DNA molecule by a Klenow fragment. Genomic identity or genomic coordinates of end locations may be derived from the results of alignment of sequence reads to a human reference genome, e.g., hg19. It may be derived from a catalog of indices or codes representing the original coordinates of the human genome. It may, non-limitingly, refer to the location or nucleotide identity on a cell-free DNA molecule read by target-specific probes, minisequencing, or DNA amplification. The term “genomic location” may refer to a nucleotide location in a polynucleotide (e.g., a gene, plasmid, nucleic acid fragment, viral DNA fragment). The term “genomic location” is not limited to a nucleotide location within a genome (e.g., a haploid set of chromosomes in a gamete or microorganism, or in each cell of a multicellular organism).

[0069] As used herein, “mutation,” “single nucleotide variant,” “single nucleotide polymorphism,” and “variant” refer to a detectable change in the genetic material of one or more cells. In specific examples, one or more mutations may be found in cancer cells and can identify cancer cells (e.g., driver and passenger mutations). Mutations can be inherited from apparent cells to daughter cells. Those skilled in the art will notice that a gene mutation in a parent cell (e.g., a driver mutation) may induce additional, different mutations (e.g., passenger mutations) in daughter cells. Mutations or variants generally occur in nucleic acids. In specific examples, a mutation may be a detectable change in one or more deoxyribonucleic acids or fragments thereof. Mutations generally refer to nucleotides that are added, deleted, substituted, inverted, or transposed at a new position in a nucleic acid. Mutations may be spontaneous or experimentally induced. Mutations in the sequence of a specific tissue are an example of “tissue-specific alleles.” For example, tumors may have mutations that result in alleles at a gene locus that do not occur in normal cells. Another example of a "tissue-specific allele" is the fetal-specific allele, which occurs in fetal tissue but not in maternal tissue. The term "allele" can be used interchangeably with "mutation" in some cases.

[0070] The term "transient binding" means that the reagent or probe reversibly binds to a binding site on a polynucleotide, and the probe does not typically remain attached to that binding site. This provides useful information about the location of the binding site during the course of analysis. Typically, one reagent or probe binds to an immobilized polymer and then detaches from the polymer after a short residence time. The same or another reagent or probe then binds to the polymer at a different site. In some embodiments, multiple binding sites along the polymer are also simultaneously bound to multiple reagents or probes. In some examples, different probes bind to overlapping binding sites. This step of reversibly binding a reagent or probe to a polymer is repeated many times during the course of analysis. The location, frequency, residence time, and photon emission of such binding events ultimately provide a map of the polymer's chemical structure. In fact, the transient nature of these binding events allows for the detection of numerous such binding events. If probes remained bound for a long period, each probe would inhibit the binding of other probes.

[0071] The term "repeated binding" refers to the fact that the same binding site in a polymer is bound multiple times during the analysis process by the same binding reagent or probe, or by the same chemical species of the binding reagent or probe. Typically, one reagent binds to the site, then dissociates, another reagent binds, then dissociates, and so on, until a map of the polymer is developed. This repeated binding improves the sensitivity and accuracy of the information obtained from the probe. More photons are accumulated, and multiple independent binding events increase the probability of detecting an actual signal. Sensitivity increases when a signal is too weak to be called over background noise if it is detected only once. In such cases, the signal becomes callable if it is persistently detected (for example, confidence that the signal is real increases when the same signal is detected multiple times). The accuracy of calling binding sites increases because multiple readings of the information confirm one reading with another.

[0072] As used herein, the term "probe" may include oligonucleotides to which any optically fluorescent labeling is attached. In some embodiments, the probe is a peptide or polypeptide optionally labeled with a fluorescent dye or a fluorescent or light-scattering molecule. These probes are used to determine the location of a binding site to either nucleic acids or proteins.

[0073] As used herein, the terms “oligonucleotide” and “oligo” refer to short nucleic acid sequences. In some examples, an oligo is of a defined size, for example, each oligo being the length of k nucleotides (also referred to herein as a “k-mer”). Typically, oligo sizes include trimers, tetramers, pentamers, hexamers, and so on. Oligos are also referred to herein as N-mers.

[0074] As used herein, the term “label” encompasses a single detectable entity (e.g., an entity emitting a wavelength) or multiple detectable entities. In some embodiments, the label transiently binds to a nucleic acid or is bound to a probe. Different types of labels will blink fluorescence, vary photon emission, or switch a photoelectric switch on and off. Different labels are used in different imaging techniques. In detail, some labels are uniquely adapted to different types of fluorescence microscopy. In some embodiments, fluorescent labels fluoresce at different wavelengths and also have different durations. In some embodiments, background fluorescence is present in the imaging field. In some such embodiments, such background is removed from the analysis by rejecting the early time window of fluorescence through scattering. If the label is present at one end of the probe (e.g., the 3' end of an oligo probe), the accuracy of localization corresponds to that end of the probe (e.g., the 3' end of the probe sequence and the 5' end of the target sequence). The clear, transient, fluctuating, or flashing behavior of the label allows for identification of whether the attached probe is on or off binding to the binding site.

[0075] As used herein, the term “flap” refers to an entity that acts as a receptor for the binding of a second entity. These two entities may include a molecular binding pair. Such a binding pair may include a nucleic acid binding pair. In some embodiments, the flap includes an extension of an oligo- or polynucleotide sequence that binds to a labeled oligonucleotide. Such a binding between the flap and the oligonucleotide must be substantially stable during the process of imaging the transient binding of a portion of the probe that binds to the target.

[0076] The terms “extended,” “expanded,” “stretched,” “linearized,” and “straightened” can be used interchangeably. More specifically, the term “extended polynucleotide” (or “expanded polynucleotide,” etc.) refers to a nucleic acid molecule that has been attached to a surface or matrix by some means and subsequently stretched into a linear form. Generally, these terms mean that the binding sites along the polynucleotide are separated by a physical distance that correlates somewhat with the number of nucleotides between them (e.g., polynucleotides are straight). Some inaccuracy in the extent that the physical distance coincides with the number of bases can be tolerated.

[0077] As used herein, the term “imaging” encompasses both two-dimensional arrays and two-dimensional scanning detectors. In most examples, the imaging techniques used herein will invariably include a fluorescence-activated material (e.g., a laser of a suitable wavelength) and a fluorescence detector.

[0078] As used herein, the term “sequence bit” refers to one or several bases (e.g., 1 to 9 bases in length) of a sequence. In particular, in some embodiments, the sequence corresponds to the length of the oligo(or peptide) used for transient binding. Thus, in such embodiments, the sequence refers to a region of the target polynucleotide.

[0079] As used herein, the term "haplotype" typically refers to a set of mutations that are present together from birth. This occurs because the set of mutations are located close together on a polynucleotide or chromosome. In some cases, a haplotype includes one or more single nucleotide polymorphisms (SNPs). In some cases, a haplotype includes one or more alleles.

[0080] As used herein, the term "methyl-binding protein" refers to a protein containing a methyl-CpG binding domain, which comprises approximately 70 nucleotide residues. Such domains have low affinity for unmethylated regions of DNA and can therefore be used to identify the location of methylated nucleic acids. Some common methyl-binding proteins include MeCP2, MBD1, and MBD2. However, there are various different proteins that contain methyl-CpG binding domains (e.g., Rolloff et al., BMC Genomics 4:1, 2003).

[0081] As used herein, the term "nanobody" refers to a proprietary set of proteins containing only heavy-chain antibody fragments. These are highly stable proteins that can be designed to have sequence homology similar to various human antibodies, thus enabling specific targeting of cell types or regions within the body. A review of nanobody can be found in Bannas et al., Frontiers in Immu. 8:1603, 2017.

[0082] As used herein, the term "affimer" refers to a non-antibody-binding protein. These are highly customizable proteins having two peptide loops and an N-terminal sequence, randomized in some embodiments to provide affinity and specificity to a desired protein target. Thus, in some embodiments, affimers are used to identify a desired sequence or structural region within a protein. In some such embodiments, affimers are used to identify the expression, location, and interaction of many different types of proteins (see, for example, Tiede et al., ELife 6:E24903, 2017).

[0083] As used herein, the term "aptamer" refers to another category of binding molecules that are highly versatile and customizable. Aptamers contain nucleotide and / or peptide regions. Typically, a random set of possible aptamer sequences is generated, and then the desired sequence that binds to the specific target molecule of interest is selected. Aptamers possess additional features beyond their stability and flexibility that make them more desirable than other categories of binding proteins (see, e.g., Song et al., Sensors 12:612-631, 2012 and Dunn et al., Nat. Rev. Chem. 1:0076, 2017).

[0084] Multiple embodiments are described below with reference to illustrative applications for illustrative purposes. It should be understood that numerous specific details, relationships, and methods are provided to provide a complete understanding of the features described herein. However, those skilled in the art will notice that the features described herein can be practiced without using one or more of the specific details or other methods. Since some actions may be performed in different orders and / or simultaneously with other actions or events, the features described herein are not limited to the illustrated order of actions or events. Furthermore, not all illustrated actions or events are required to perform the methodology of the features described herein.

[0085] Exemplary System Embodiments Details of an exemplary system are described here in conjunction with Figure 1. Figure 1 is a block diagram of system 100 in several executions. Device 100 in several executions includes one or more processing units (CPUs) 102 (also referred to as processors or processing cores), one or more network interfaces 104, a user interface 106, non-persistent memory 111, persistent memory 112, and one or more communication booths 114 for interconnecting these components. One or more communication booths 114 may include circuits (sometimes called chipsets) that interconnect and control communication between system components. Typical non-persistent memory 111 includes high-speed random-access memory such as DRAM, SRAM, DDR RAM, ROM, EEPROM, and flash memory, while typical persistent memory 112 includes CD-ROMs, digital versatile disks (DVDs) or other optical storage devices, magnetic cassettes, magnetic tapes, magnetic disk storage devices or other magnetic storage devices, magnetic disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid-state storage devices. The persistent memory 112 may include one or more storage devices located remotely from the CPU(s) 102. The persistent memory 112 and the non-volatile memory devices(s) within the non-persistent memory 112 include non-temporary computer-readable storage media. In some executions, the non-persistent memory 111, or non-temporary computer-readable storage media, sometimes in cooperation with the persistent memory 112, stores the following programs, modules, and data structures, or subsets thereof: • An optional operating system 116, including procedures for handling various basic system services and for performing hardware-dependent tasks; Optional network communication module (or other instruction) 118, or communication network, for connecting system 100 to other devices; • An optical activity detection module 120 for retrieving information about each target molecule 130; • Information on each of the multiple binding sites 140 for each target molecule 130; • Information about each coupling event 142 in multiple coupling events for each coupling site 140, including at least (i) a period 144 and (ii) the number of photons emitted 146; • Sequencing module 150 for determining the sequence of each of the 130 target molecules; • Information about each binding site 140 in multiple binding sites for each target molecule 130, including at least (i) base call 152 and (ii) probability 154; • Selective information regarding the reference genome 160 for each target molecule 130; and • Selective information regarding complementary chains 170 for each of the 130 target molecules.

[0086] In various executions, one or more of the previously identified elements are stored in one or more of the previously mentioned memory devices, corresponding to a set of instructions for performing the functions described above. The previously identified modules, data, or programs (e.g., a set of instructions) do not need to be executed as another software program, procedure, dataset, or module, and therefore various subsets of these modules and data may be combined or otherwise reorganized in various executions. In some executions, non-persistent memory 111 may store a subset of the modules and data structures identified above. In some further embodiments, the memory stores additional modules and data structures not described above. In some embodiments, one or more of the previously identified elements are stored in a computer system other than the visualization system 100, addressable by the visualization system 100, so that all or part of such data can be retrieved when the visualization system 100 needs to.

[0087] Examples of network communication modules 118 include, but are not limited to, the World Wide Web (WWW), intranets and / or wireless networks such as cellular networks, wireless local area networks (LANs) and / or metropolitan area networks (MANs), and other devices using wireless communication. Wireless communication may use one of several communication standards, protocols and technologies, including Global System for Mobile Communication (GSM), Enhanced Data GSM Environment (EDGE), High Speed ​​Downlink Packet Access (HSDPA), High Speed ​​Uplink Packet Access (HSUPA), Evolution Data-Only (EV-DO), HSPA, HSPA+, and Dual-Cell. HSPA (DC-HSPADA), Long-Term Evolution (LTE), Near Field Communication (NFC), Wideband Code Division Multiple Access (W-CDMA), Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Bluetooth, Wireless Fidelity (Wi-Fi) (e.g., IEEE 802.11a, IEEE 802.11ac, IEEE 802.11ax, IEEE 802.11b, IEEE 802.11g and / or IEEE 802.11n), Voice over Internet Protocol (VoIP), Wi-MAX, Protocol for Email (e.g., Internet Message Access) This includes, but is not limited to, any other suitable communication protocols, including the IMAP protocol and / or Post Office Protocol (POP), instant messaging (e.g., Extensible Messaging and Presence Protocol (XMPP), Session Initiation Protocol and Presence Leveraging Extension for Instant Messaging (SIMPLE), Instant Message and Presence Service (IMPS)), and / or Short Message Service (SMS), or any other communication protocols not yet developed as of the filing date of this disclosure.

[0088] Figure 1 represents "System 100," which is intended to be a functional description of various features that may be present in a computer system, rather than merely a structural schematic of the execution described herein. In fact, and as will be apparent to those skilled in the art, the separately illustrated items can be combined, and some items can be separated. Furthermore, while Figure 1 represents specific data and modules in non-persistent memory 111, some or all of this data and modules may be in persistent memory 112. In some embodiments, memory 111 and / or 112 store additional modules and data structures not described above.

[0089] The system described herein has been disclosed with reference to Figure 1, but the method described herein will now be described in detail with reference to Figures 2A, 2B, 3, and 4.

[0090] Block 202 provides a method for determining the chemical structure of a molecule. The goal of this disclosure is to enable single-nucleotide degradation sequencing of nucleic acids. In some embodiments, a method is provided for characterizing the interaction between one or more probes and a molecule. The method comprises adding one or more probe species to a molecule under conditions that transiently bind one or more probe species to the molecule. The method proceeds by continuously monitoring individual binding events on the molecule with a detector and recording each binding event over a period of time. Data from each binding event are analyzed to determine one or more features of the interaction.

[0091] In some embodiments, a method for determining the identity of polymers is provided. In some embodiments, a method for determining the identity of cells or tissues is provided. In some embodiments, a method for determining the identity of living organisms is provided. In some embodiments, a method for determining the identity of solids is provided. In some embodiments, the method is applied to single-cell sequencing.

[0092] Target polymers In some embodiments, the molecule is a nucleic acid, preferably a native polynucleotide. In various embodiments, the method further comprises extracting a single target polynucleotide molecule from a cell, organelle, chromosome, virus, exosome, or bodily fluid as an intact target polynucleotide.

[0093] In some embodiments, the polymer is a short polynucleotide (e.g., <1 kilobase or <300 bases). In some embodiments, the short polynucleotide is 100–200 bases, 150–250 bases, 200–350 bases, or 100–500 bases long, as is found in cell-free DNA in bodily fluids such as urine and blood.

[0094] In some embodiments, the nucleic acid is at least 10,000 base pairs long. In some embodiments, the nucleic acid is at least 1,000,000 base pairs long.

[0095] In various embodiments, the single target polynucleotide is a chromosome. In various embodiments, the single target polynucleotide is about 10 2 , 10 3 , 10 4 , 10 5 , 10 6 , 10 7 , 10 8 , or 10 9 It is the base length.

[0096] In some embodiments, the method enables the analysis of the amino acid sequence on a target protein. In some embodiments, a method for analyzing the amino acid sequence on a target polypeptide is provided. In some embodiments, a method for analyzing peptide modifications in addition to the amino acid sequence on a target polynucleotide is provided. In some embodiments, the molecular entity is a polymer containing at least 5 units. In such embodiments, the binding probe is a molecular probe including oligonucleotides, antibodies, affimers, nanobodies, aptamer-binding proteins, or small molecules.

[0097] In such embodiments, each of the 20 amino acids is bound by a corresponding specific probe, such as an N-recogniin, nanobody, antibody, or aptamer. The binding of each probe is specific to the corresponding amino acid in the polypeptide chain. In some embodiments, the order of subunits in the polypeptide is determined. In some embodiments, binding is to a substitute for the binding site. In some embodiments, the substitute is a tag attached to a specific amino acid or peptide sequence, and transient binding would be to the substitute tag.

[0098] In some embodiments, the molecules are heterogeneous molecules. In some embodiments, the heterogeneous molecules include supramolecular structures. In some embodiments, the method enables the identification and ordering of chemical structural units for heterogeneous polymers. Such embodiments include extending a polymer and attaching multiple probes to identify chemical structures at multiple sites along the extended polymer. Extending the heteropolymer enables subdiffraction-level (e.g., nanometer) localization of probe attachment sites.

[0099] In several embodiments, methods are provided for sequencing polymers by binding probes that recognize polymer subunits. Typically, binding of a single probe is insufficient to sequence a polymer. For example, Figure 1A shows an embodiment in which sequencing of polymer 130 is based on measuring transient interactions with a repertoire of probes 182 (e.g., interactions between a repertoire of denatured polynucleotides and oligonucleotides, or interactions between a panel of denatured polypeptides and nanobodies or affimers).

[0100] Extraction and / or preparation of target polymers In some embodiments, it is necessary to isolate the target cells from other cells that are not present before nucleic acid extraction is performed. In one such example, circulating tumor cells or circulating fetal cells are isolated from blood (e.g., by using cell surface markers for affinity capture). In some embodiments, if the interest is to detect and analyze polynucleotides from microbial cells, it is necessary to isolate microbial cells from human cells. In some embodiments, a wide range of microorganisms are affinity-captured using opsonins and isolated from mammalian cells. In addition, in some embodiments, differential lysis is performed. Mammalian cells are first lysed under relatively mild conditions. Microbial cells are typically harder than mammalian cells, and therefore they remain intact during the lysis of mammalian cells. The lysed mammalian cell fragments are washed away. Then, microbial cells are lysed using harsher conditions. Subsequently, the target microbial polynucleotides are selectively sequenced.

[0101] In some embodiments, the target nucleic acid is extracted from the cell before sequencing. In alternative embodiments, sequencing (e.g., of chromosomal DNA) is performed inside the cell, where the chromosomal DNA follows a complex trajectory in its quiescent phase. Stable binding of oligos in situ has been demonstrated by Beliveau et al., Nature Communications 6:7147 (2015). Such in-situ binding of oligos in three-dimensional space and their nanometer-level positioning enable the determination of the arrangement and structural arrangement of chromosomal molecules within the cell.

[0102] Target polynucleotides often exist in their native folded state. In one such example, genomic DNA is highly condensed within the chromosome, while RNA forms secondary structures. In some embodiments, long polynucleotides are obtained during extraction from a biological sample (e.g., by preserving the substantially native length of the polynucleotide). In some embodiments, the polynucleotides are straightened so that their location along the length is traceable with little or no obscuration. Ideally, the target polynucleotides are straightened, elongated, or stretched before or after straightening.

[0103] The method is particularly suitable for sequencing very long polymer lengths, where the native length or a substantial proportion thereof is preserved (e.g., an entire chromosome or approximately 1 megabase fraction in DNA). However, common molecular biology methods result in unintended fragmentation of DNA. Pipetting and vortexing, for example, induce shear forces that disrupt DNA molecules. Nuclease contamination can lead to nucleic acid degradation. In some embodiments, the native length, or a substantial high molecular weight (HMW) fragment of the native length, is preserved before immobilization, extension, and sequencing begin.

[0104] In some embodiments, polynucleotides are intentionally fragmented into relatively homogeneous, long lengths (e.g., about 1 Mb) before sequencing proceeds. In some embodiments, polynucleotides are fragmented into relatively homogeneous, long lengths after or during fixation or extension. In some embodiments, fragmentation is performed enzymatically. In some embodiments, fragmentation is performed physically. In some embodiments, physical fragmentation is performed by sonication. In some embodiments, physical fragmentation is performed by ion bombardment or radiation. In some embodiments, physical fragmentation is performed by electromagnetic radiation. In some embodiments, physical fragmentation is performed by UV irradiation. In some embodiments, the dose of UV irradiation is controlled to perform fragmentation to a given length. In some embodiments, physical fragmentation is performed by a combination of UV irradiation and staining with a dye (e.g., YOYO-1). In some embodiments, the fragmentation step is stopped by a physical action or reagent addition. In some embodiments, the reagent used to stop the fragmentation step is a reducing agent such as beta-mercaptoethanol (BME).

[0105] Fragmentation and sequencing based on radiation dose If the field of view of a two-dimensional sensor allows for observation of the complete megabase length of DNA in one dimension of the sensor, it is efficient to generate 1 Mb length genomic DNA. It should also be noted that reducing the size of chromosome fragments minimizes strand tangling, allowing for the acquisition of the maximum length DNA in an extended and well-isolated form.

[0106] A method for sequencing long chromosomal subfragments, including the following steps: i) A step of staining chromosomal double-stranded DNA with a dye that intercalates between double-stranded base pairs; ii) Exposing stained chromosomal DNA to a predetermined dose of electromagnetic radiation to produce sub-fragments of chromosomal DNA within a desired size range; iii) A step of extending and fixing the stained chromosomal subfragment DNA on the surface; iv) A step in which the stained chromosome subfragment is denatured to disrupt base pairs, thereby releasing intercalating dyes; v) Exposing the obtained bled, extended, and fixed single strands to a repertoire of oligonucleotides of a given length and sequence; vi) The step of determining the binding site along the faded extended single strand of each oligonucleotide in the repertoire. vii) A step to compile the binding sites of all oligos in the repertoire and obtain the complete sequencing of the chromosomal subfragment.

[0107] In some of the earlier embodiments, staining occurs while the chromosomes are inside the cell. In some of the earlier embodiments, the labeled oligonucleotides are only partially labeled because more staining agent is added and intercalates within the duplex when the duplex is formed. In some of the earlier embodiments, a dose of electromagnetic radiation is applied that can decolorize the stain, in addition to denaturation, if applicable. In some of the earlier embodiments, the predetermined dose is achieved by manipulating the intensity and duration of exposure and by stopping fragmentation with chemical exposure, the chemical exposure being a reducing agent such as beta-mercaptoethanol. In some of the earlier embodiments, the dose is predetermined to create a Poisson distribution of fragments around 1 Mb in length.

[0108] Methods of fixing and immobilizing Block 204: Nucleic acids are immobilized on a test substrate in a double-stranded linearized extension form, thereby forming immobilized extended double-stranded nucleic acids. The molecules may be immobilized on the surface or matrix. In some embodiments, fragmented or native polymers are immobilized. In some embodiments, the immobilized double-stranded linearized nucleic acids follow curved or winding trajectories rather than straight ones.

[0109] In some embodiments, immobilization involves coating a test substrate with nucleic acids by molecular combing (retraction meniscus), flow extension, nanoconfinement, or electrostretching. In some embodiments, coating of nucleic acids with a substrate further includes a UV crosslinking step in which the nucleic acids are covalently bonded to the substrate. In some embodiments, coating does not require UV crosslinking of the nucleic acids, and the nucleic acids are bonded to the substrate by other means (e.g., hydrophobic interactions, hydrogen bonding, etc.).

[0110] Immobilization (e.g., fixation) of a polynucleotide at only one end causes the polynucleotide to elongate and contract in a non-cooperative manner. Therefore, regardless of the extension method used, the degree of elongation along the length of the polymer cannot be guaranteed to be at any specific location in the target. In some embodiments, it is necessary that the relative positions of multiple locations along the polymer remain constant. In such embodiments, the extended molecule must be immobilized or fixed to a surface by multiple contact points along its length so that surface elongation may be used (e.g., ACS Nano. 2015 Jan 27;9(1):809-16) (e.g., as performed in the molecular combing technique in Michaelet et al, Science 277:1518-1523, 1997; see also Molecular Combing of DNA: Methods and Applications, Journal of Self-Assembly and Molecular Electronics (SAME) 1:125-148).

[0111] In some embodiments, the array of polynucleotides is immobilized on a surface, and in some embodiments, the polynucleotides in the array are far enough apart to be individually resolved by diffraction-limited imaging. In some embodiments, the polynucleotides are given on the surface in an ordered manner so that the molecules are packed to the maximum extent within a given surface area and do not overlap. In some embodiments, this is done by creating a patterned surface (e.g., an aligned arrangement of hydrophobic patches or strips at such locations where the ends of the polynucleotides bind). In some embodiments, the polynucleotides in the array are not far enough apart to be individually resolved by diffraction-limited imaging and are resolved individually by super-resolution imaging.

[0112] In some embodiments, polynucleotides are organized in a DNA curtain (Greene et al., Methods Enzymol. 472:293-315, 2010). This is particularly useful for long polynucleotides. In such embodiments, transient binding is recorded while a DNA strand attached to one end is extended by fluid or electrophoretic forces, or after both ends of the strand are captured. In some embodiments, if many copies of the same sequence form multiple polynucleotides in a DNA curtain, the sequence is assembled from multiple polynucleotides rather than from a single polynucleotide, based on the binding pattern in an aggregate. In some embodiments, both ends of a polynucleotide bind to pads (e.g., areas of the surface that adhere to the polynucleotide more than other parts of the surface location), with each end binding to a different pad. In such embodiments, a single linear polynucleotide binds to two pads to hold the elongated arrangement of the polynucleotide in place and allow for the formation of an aligned array of equally spaced, non-overlapping or non-interacting polynucleotides. In some embodiments, only one polynucleotide occupies an individual pad. In some embodiments, when the pads are fixed using the Poisson process, some pads are not occupied by polynucleotides, some are occupied by one polynucleotide, and some are occupied by more than one polynucleotide.

[0113] In some embodiments, target molecules are captured on an aligned supramolecular scaffold (e.g., a DNA origami structure). In some embodiments, the scaffold structure starts freely in solution and utilizes solution-phase dynamics to capture target molecules. Once occupied, the scaffold anchors to the surface or self-assembles and becomes fixed to the surface. The aligned array allows for efficient subdiffraction packing of molecules, resulting in a high density of molecules per field of view (high-density array). Single-molecule localization methods enable super-resolution of polynucleotides in high-density arrays (e.g., point-to-point distances of 40 nm or less).

[0114] In some embodiments, hairpins are ligated to the ends of the duplex template (sometimes after smoothing the ends of the nucleic acid). In some embodiments, the hairpins contain biotin to immobilize the nucleic acid on the surface. In alternative embodiments, the hairpins serve to covalently link the two strands of the duplex. In some such embodiments, an oligo(olio)d(T) is added to the other end of the nucleic acid for surface capture. After denaturation, both strands of the nucleic acid are available for interaction with the oligo.

[0115] In some embodiments, the aligned arrays take the form of individual scaffolds that connect to one another to form a larger DNA lattice (see, for example, Woo and Rothemund, Nature Communications, 5: 4889). In some such embodiments, the individual small scaffolds are fixed to one another by base pairing. They then provide a higher-order nanostructure array for the sequencing step of the present disclosure. In some embodiments, capture sites are arranged at a 10 nm pitch in the aligned two-dimensional lattice. With complete occupancy, such a lattice has the ability to capture approximately 1 trillion molecules per square centimeter.

[0116] In some embodiments, the capture sites in the lattice are arranged in an aligned two-dimensional lattice at pitches of 5 nm, 10 nm, 15 nm, 30 nm, or 50 nm.

[0117] In some embodiments, aligned arrays are fabricated using nanofluidics. In one such example, an array of nanotrenches or nanogrooves (e.g., 100 nm wide and 150 nm deep) has a non-smooth surface that serves to order long polynucleotides. In such embodiments, the appearance of one polynucleotide in the nanotrench or nanogroove prevents the entry of another polynucleotide. In another embodiment, a nanopit array is used, where segments of long polynucleotides are located within pits, and intervening long segments are spread between the pits.

[0118] In some embodiments, super-resolution imaging and precise sequencing are possible even with high-density polynucleotides. For example, in some embodiments, only a subset of polynucleotides is of interest (e.g., targeted sequencing). In such embodiments, when targeted sequencing is performed and polynucleotides are attached to a surface or matrix at a higher density than usual, only a subset of polynucleotides from a complex sample (e.g., whole genome or transcriptome) needs to be analyzed. In such embodiments, even if multiple polynucleotides are present within the diffraction-limited space, if a signal is detected, it is highly likely that it originates from only one of the target loci, and that this locus is not within the diffraction-limited distance of another such locus to which the probe is simultaneously bound. The required distance between each polynucleotide undergoing targeted sequencing correlates with the percentage of polynucleotides to be targeted. For example, if less than 5% of the polynucleotides are targeted, the polynucleotide density is 20 times greater than if the entire polynucleotide sequence were desired. In some embodiments of targeted sequencing, imaging time is shorter than when the whole genome is analyzed (for example, in the previous example, targeted sequencing imaging can be 10 times faster than whole genome sequencing).

[0119] In some embodiments, the test substrate is bound to a sequence-specific oligonucleotide probe before immobilization, and immobilization involves capturing nucleic acids on the test substrate using the sequence-specific oligonucleotide probe bound to the test substrate. In some embodiments, the nucleic acid is bound at its 5' end. In some embodiments, the nucleic acid is bound at its 3' end. In another embodiment, if two other probes are present on the substrate, one probe will bind to the first end of the nucleic acid and the other probe will bind to the second end of the nucleic acid. In examples where two probes are used, it is also necessary to have prior information regarding the length of the nucleic acid. In some embodiments, the nucleic acid is first cleaved with a predetermined endonuclease.

[0120] In various embodiments, the target polynucleotide is extracted or embedded in a gel or matrix before fixation (e.g., as described in Shag et al., Nature Protocols 7:467-478, 2012). In one such non-limiting example, the polynucleotide is deposited in a flow channel containing a medium undergoing a liquid-gel transition. The polynucleotide is first extended and distributed in the liquid phase and then fixed by a phase transition to the solid / gel phase (e.g., by heating, or, in the case of a polyacrylamide gel, by cofactor addition or time). In some embodiments, the polynucleotide is extended in the solid / gel phase.

[0121] In some alternative embodiments, the probe itself is immobilized on a surface or matrix. In such embodiments, one or more target molecules (e.g., polynucleotides) are suspended in solution and transiently bound to the immobilized probe. In some embodiments, polynucleotides are captured using a spatially addressable array of oligonucleotides. In some embodiments, short polynucleotides (e.g., less than 300 nucleotides), such as cell-free DNA or microRNA, or relatively short polynucleotides (e.g., less than 10,000 nucleotides), such as mRNA, are randomly immobilized on a surface by capturing modified or unmodified ends with a suitable capture molecule. In some embodiments, the short or relatively short polynucleotides interact with the surface multiple times, and sequencing is performed in a direction parallel to the surface. This resolves the organization of splicing isoforms. For example, in some isoforms, the locations of repeating or shuffled exons are depicted.

[0122] In some embodiments, the immobilized probe contains a common sequence for annealing to the polynucleotide. Such embodiments are particularly useful when the target polynucleotide has a common sequence, preferably at one or both ends. In some embodiments, the polynucleotide is single-stranded and has a common sequence such as a poly-A tail. In one such example, native mRNA supporting the poly-A tail is captured by a lone oligo-d(T) probe on the surface. In some embodiments, particularly those in which short DNA is analyzed, the ends of the polynucleotide are adapted to interact with a capture molecule on the surface / matrix.

[0123] In some embodiments, polynucleotides form double helixes with sticky ends created by restriction enzymes. In some non-limiting examples, restriction enzymes at rare sites (e.g., PMme1 or NOT1) are used to create long fragments of polynucleotides, each fragment containing a common terminal sequence. In some embodiments, fitting is carried out using terminal transferases. In other embodiments, ligation or tagmentation is used to introduce adapters for Illumina sequencing. This allows users to prepare samples using established Illumina protocols, which are then captured and sequenced by the methods described herein. In such embodiments, polynucleotides are preferably captured before amplification, which tends to introduce errors and biases.

[0124] Extension method In most embodiments, the polynucleotide or other target molecule must be attached to a surface or matrix for the elongation to occur. In some embodiments, the elongation of the nucleic acid is made equal to, longer than, or shorter than its crystallographic length (e.g., a 0.34 nm separation from one base to the next is known). In some embodiments, the polynucleotide is elongated beyond its crystallographic length.

[0125] In some embodiments, polynucleotides are elongated by molecular combing (e.g., described in Michaelet et al., Science 277:1518-1523, 1997, and Deen et al., ACS Nano 9:809-816, 2015). This allows for the parallel elongation and unidirectional alignment of millions and billions of molecules. In some embodiments, molecular combing is carried out by washing a solution containing the desired nucleic acid onto a substrate and then retracting the meniscus of the solution. Before retracting the meniscus, the nucleic acid forms covalent or other interactions with the substrate. As the solution retracts, the nucleic acid is pulled in the same direction as the meniscus (e.g., by surface retention), but if the strength of the interaction between the nucleic acid and the substrate is sufficient to overcome the surface retention force, the nucleic acid is elongated in a uniform manner in the direction that retracts the meniscus. In some embodiments, molecular combing is carried out as described in Kaykov et al., Sci Reports. 6:19636 (2016), which is incorporated herein by reference in whole. In other embodiments, molecular combing is carried out in multiple channels (e.g., of a microfluidic device) using the method described in Petit et al. Nano Letters 3:1141-1146 (2003) or various modifications thereof.

[0126] The shape of the air / water interface determines the orientation of the extended polynucleotides that are extended by molecular combing. In some embodiments, the polynucleotides are extended perpendicular to the air / water interface. In some embodiments, the target polynucleotides are attached to the surface without modification of one or both ends. In some embodiments, if the ends of the double-stranded nucleic acid are captured by hydrophobic interactions, extension in the receding meniscus denatures a portion of the duplex, forming further hydrophobic interactions with the surface.

[0127] In some embodiments, polynucleotides are extended by molecular threading (e.g., Payne et al., PLoS ONE 8(7):E69058, 2013). In some embodiments, molecular threading is performed after the target has been denatured into a single strand (e.g., by chemical denaturation, temperature, and enzymes). In some embodiments, polynucleotides are ligated at one end and then extended in a fluid flow (e.g., Greene et al., Methods in Enzymology, 327: 293-315).

[0128] In various embodiments, target polynucleotide molecules reside within microfluidic channels. In one such example, polynucleotides flow into the microfluidic channel or are extracted into the flow channel from one or more chromosomes, exosomes, nuclei, or cells. In some embodiments, rather than inserting polynucleotides into nanochannels via micro- or nanofluidic flow cells, polynucleotides are inserted into open-top channels by constructing the channel in such a way that the surface on which the channel wall is formed is electrically deflected (see, e.g., Asanov et al., Anal Chem. 1998 Mar. 15; 70(6):1156-6). In one such example, a positive bias is applied to the surface, thereby attracting negatively charged polynucleotides to the nanochannel. At the same time, the channel wall ridges do not contain deflection, thereby making it less likely that the polynucleotides themselves will accumulate in the ridges.

[0129] In some embodiments, stretching is due to hydrodynamic drag. In one such example, polynucleotides are stretched by cross-flow in a nanoslit (Marie et al., Proc Natl Acad Sci USA 110:4893-8, 2013). In some embodiments, nucleic acid stretching is due to nanoconfinement in a flow channel. Flow-stretching nanoconfinement involves stretching nucleic acids into a linear conformation by a flow gradient, which is commonly carried out within microfluidic devices. The nanoconfinement portion of this stretching method typically refers to a narrow region of the microfluidic device. The use of a narrow region or channel helps overcome the problem of molecular uniqueness (e.g., the tendency of individual nucleic acids or other polymers to adopt multiple conformations during stretching). One problem with flow-stretching methods is that the flow is not always applied uniformly along the nucleic acid molecule. This can result in nucleic acids exhibiting a wide range of stretching lengths. In some embodiments, flow-stretching methods involve stretching flow and / or hydrodynamic drag. In some embodiments, polynucleotides are attracted into nanochannels, and one or more polynucleotides are nanoconfined within the channel, thereby extending it. In some embodiments, after nanoconfinement, the polynucleotides are deposited on a deflected surface or on a topcoat or matrix on the surface.

[0130] Several methods exist for applying positive or negative bias to a surface. In one such example, the surface is fabricated or coated with a non-contaminating material, or passivated with lipids (e.g., lipid bilayers), bovine serum albumin (BSA), casein, or various PEG derivatives. Passivation serves to prevent polynucleotide blockade in any portion of the channel, thereby enabling extension. In some embodiments, the surface also contains indium tin oxide (ITO).

[0131] In some embodiments, zwitterionic POPC (1-palmitoyl-2-oleoyl-sn-glycero-3-phosphocholine) lipids are coated onto the surface of nanofluidic channels together with 1% Lissamine™ Rhodamine B 1,2-dihexadecanoyl-sn-glycero-3-phosphoethanolamine to create lipid bilayers (LBLs) on the surface. The addition of triethylammonium salt (Rhodamine-DHPE) lipids enables observation of LBL formation by fluorescence microscopy. The lipid bilayer passivation methods used in some embodiments of this disclosure are described in Persson et al., Nano Lett. 12:2260-2265, 2012.

[0132] In some embodiments, the stretching of one or more polynucleotides is carried out by electrophoresis. In some embodiments, the polynucleotides are ligated at one end and then stretched by an electric field (e.g., Giese et al., Nature Biotechnology 26: 317-325, 2008). Electrostretching of nucleic acids is predictable given the fact that nucleic acids are highly negatively charged molecules. Methods of electrostretching, such as those described by Randall et al. 2006, Lab Chip. 6, 516-522, involve drawing the nucleic acid through microchannels by electric current (and inducing the orientation of the nucleic acid molecule). In some embodiments, electrostretching is performed in or without a gel. One benefit of using a gel is that it limits the three-dimensional space available to the nucleic acid, thereby helping to overcome molecular uniqueness. A general advantage of electrostretching over pressure-driven stretching methods such as nanoconfining is the absence of shear forces that would break down the nucleic acid molecule.

[0133] In some embodiments, when multiple polynucleotides are present on a single surface, the polynucleotides may not be aligned in the same orientation or may not be straight (e.g., the polynucleotides may adhere to the surface or pass through the gel in curved trajectories). In such embodiments, two or more of the multiple polynucleotides may overlap, increasing the likelihood of confusion regarding probe positioning along the length of each polynucleotide. While the same sequencing information can be obtained from curved sequences as from straight, well-aligned molecules, the image processing task required to handle sequencing information from curved sequences requires greater computational power than the information obtained from straight, properly-aligned molecules.

[0134] In embodiments where one or more polynucleotides are extended in a direction parallel to a flat surface, their lengths are imaged by a series of adjacent pixels in a two-dimensional array detector such as a CMOS or CCD camera. In some embodiments, one or more polynucleotides are extended in a direction perpendicular to the surface. In some embodiments, the polynucleotides are imaged by optical sheet microscopy, sipinning disk confocal microscopy, three-dimensional super-resolution microscopy, three-dimensional single-molecule localization, or laser scanning disk confocal microscopy or a variation thereof. In some embodiments, the polynucleotides are extended at an oblique angle to the surface. In some embodiments, the polynucleotides may be imaged by a two-dimensional detector, and the image may be processed by Single Molecule Localization algorithm software (e.g., Fiji / ImageJ plug-in ThunderSTORM as described in Ovesny et al., BioInform. 30:2389-2390, 2014).

[0135] DNA extraction and isolation from single cells before fixation and extension. In some embodiments, traps for single cells are designed within a microfluidic structure to hold individual cells in one location while allowing nucleic acid contents to be released (e.g., by using the device designs of WO / 2012 / 056192 or WO / 2012 / 055415). In some embodiments, instead of extracting and extending polynucleotides in nanochannels, the coverslip or foil used to seal the micro / nanofluidic structure is coated with polyvinylsilane to enable molecular combing (e.g., by fluid movement as described by Petit et al., Nano Letters 3:1141-1146. 2003). Mild conditions inside the fluid chip allow for long-term retention of the extracted polynucleotides.

[0136] Numerous different approaches are available for extracting biopolymers from single cells or nuclei (e.g., several suitable methods are summarized in Kim et al., Integr Biol 1(10), 574-86, 2009). In some non-limiting examples, cells are treated with KCl to remove the cell membrane. Cells are lysed by adding a hypotonic solution. In some embodiments, each cell is isolated separately, the DNA of each cell is extracted separately, and then each set of DNA is sequenced separately in a microfluidic vessel or device. In some embodiments, extraction occurs by treating one or more cells with a surfactant and / or protease. In some embodiments, a chelating agent (e.g., EDTA) is provided in the lysis solution to capture the divalent cations required by the nuclease (thus reducing nuclease activity).

[0137] In some embodiments, the nucleus and extranuclear components of a single cell are extracted separately by the following method: One or more cells are provided to a supply channel of a microfluidic device. One or more cells are then captured, where each cell is captured by a trapping structure. A first lysis buffer is added to the solution, which lyses the cell membrane but helps maintain the integrity of the cell nucleus. Upon addition of the first lysis buffer, the extranuclear components of one or more cells are released into the flow cell, where the released RNA is immobilized. One or more nuclei are then lysed by supplying a second lysis buffer. The addition of the second lysis buffer causes the components of one or more nuclei (e.g., genomic DNA) to be released into the flow cell, where the DNA is then immobilized. The extracellular and intracellular components of one or more cells are immobilized at different locations in the same flow cell or in different flow cells within the same device.

[0138] The diagrams in Figures 16A and 16B illustrate a microfluidic structure for capturing and isolating multiple single cells. Cell 1602 is captured by cell trap 1606 in flow cell 2004. In some embodiments, after the cell is captured, a lysis reagent is flowed. After lysis, the polynucleotides are distributed near the capture region, remaining isolated from the polynucleotides extracted from other cells. In some embodiments, induction by electrophoresis is performed to manipulate the nucleic acids (e.g., utilizing charge 1610), as shown in Figure 16B. Lysis will release nucleic acids 1608 from cells 1602 and nucleus 1604. The nucleic acids 1608 remain where they were when cells 1602 were trapped (e.g., relative to cell trap 1606). The trap is the size of a single cell (e.g., 2–10 μM). In some embodiments, the channel binding the microdroplet and cell is larger than 2 μM or 10 μM. In some embodiments, the distance between the bifurcated channel and the trap is 1 to 1000 microns.

[0139] Extraction and extension of high molecular weight DNA on the surface Various methods for extending HMW polynucleotides are used in different embodiments (e.g., ACS Nano. 9(1):809-16, 2015). In one such example, extension on a surface is carried out in a flow cell (e.g., using the approach described by Petit and Carbeck in Nano. Lett. 3: 1141-1146, 2003). In addition to the fluid approach, in some embodiments, polynucleotides are extended using an electric field as disclosed in Giess et al., Nature Biotechnology 26, 317-325 (2008). Multiple approaches are available for extending polynucleotides when they are not attached to a surface (e.g., Frietag et al., Biomicrofluidics, 9(4):044114 (2015); Marie et al., Proc Natl Acad Sci USA 110:4893-8, 2013).

[0140] Instead of using DNA in gel plugs, chromosomes suitable for loading onto the chip are prepared by the polyamine method described by Cram et al., Methods Cell Sci., 2002, 24, 27-35, and pipetteed directly into the device. In some such embodiments, proteins that bind to DNA in the chromosomes are digested using proteases to release substantially naked DNA, which is then fixed and extended as described above.

[0141] Sample processing for localized preservation of leads In embodiments where very long regions or polymers are sequenced, any degradation of the polymer can significantly reduce the accuracy of the overall sequencing. Methods to promote the retention of the entire extended polymer are described below.

[0142] Polynucleotides can become damaged during extraction, storage, or preparation. Nicks and adducts can form in native double-stranded genomic DNA molecules. This is especially true when the polynucleotides of the sample originate from FFPE material. Therefore, in some embodiments, a DNA repair solution is introduced before or after DNA immobilization. In some embodiments, this is carried out after DNA extraction into a gel plug. In some embodiments, the repair solution contains DNA endonucleases, kinases, and other DNA modifying enzymes. In some embodiments, the repair solution contains polymerases and ligases. In some embodiments, the repair solution is a pre-PCR kit from New England Biolabs. In some embodiments, such methods are largely carried out as described in Karimi-Busheri et al., Nucleic Acids Res. Oct 1;26(19):4395-400, 1998 and Kunkel et al., Proc. Natl Acad Sci. USA,78, 6734-6738, 1981.

[0143] In some embodiments, a gel overlay is applied after the polynucleotide has been extended. In some such embodiments, after extension and denaturation on the surface, the polynucleotide (double-stranded or denatured) is covered with a gel layer. Alternatively, the polynucleotide is extended while already in a gel environment (e.g., as described above). In some embodiments, after the polynucleotide has been extended, it is flowed into the gel. For example, in some embodiments, the polynucleotide is attached to a surface at one end and extended in a flow stream or by an electric current in electrophoresis, causing the surrounding medium to flow into the gel. In some embodiments, this occurs by including acrylamide, ammonium persulfate, and TEMED in the flow stream. Such compounds solidify to form polyacrylamide. In alternative embodiments, a heat-responsive gel is applied. In some embodiments, the ends of the polyacrylamide are modified with Acrydite, which polymerizes with acrylamide. In some such embodiments, considering the negative backbone of the native polynucleotide, an electric field is applied to extend the polynucleotide toward the positive electrode.

[0144] In some embodiments, nucleic acids are extracted from cells in a gel plug or gel layer to preserve DNA integrity, and then an AC electric field is applied to extend the DNA in the gel. If this is carried out on the top gel layer of a coverslip, the method of the present invention can be applied to the extended DNA to detect transient oligobonds.

[0145] In some embodiments, the sample is crosslinked into the matrix of its environment. In one example, this is the intracellular environment. For example, when sequencing is performed in situ within a cell, polynucleotides are crosslinked into the intracellular matrix using heterobifunctional crosslinkers. This occurs when sequencing is applied directly into the cell using techniques such as FISSEQ (Lee et al., Science 343:1360-1363, 2014).

[0146] Much of the damage occurs during the process of extracting biomolecules from cells and tissues and then handling them before analysis. In the case of DNA, handling methods that lead to loss of integrity include pipetting, vortexing, freeze-thaw, and heating. In some embodiments, mechanical stress is minimized by methods such as those disclosed in ChemBioChem, 11:340-343 (2010). In addition, high concentrations of divalent cations, EDTA, EGTA, or gallic acid (and their analogs and derivatives) inhibit nuclease degradation. In some embodiments, a 2:1 ratio of sample to divalent cation weight is sufficient to inhibit nucleases even in samples such as feces where extreme nucleases are present.

[0147] To maintain the integrity of nucleic acids (e.g., to prevent damage to DNA or breakdown into smaller fragments), in some embodiments, it is desirable to hold biomacromolecules such as DNA in their natural protective environments, such as chromosomes, mitochondria, cells, nuclei, or exosomes. In embodiments where the nucleic acid is already outside the protective environment, it is desirable to place the nucleic acid within a protective environment, such as a gel or microdroplet. In some embodiments, the nucleic acid is released from the protective environment by physically bringing it into close proximity to the sequencing site (e.g., part of a fluid system or flow cell from which sequencing data is acquired). Thus, in some embodiments, the biomacromolecule (e.g., nucleic acid, protein) is provided within a protective entity that holds the biomacromolecule in a near-native state (e.g., native length), and the protective entity containing the biomacromolecule is brought near the sequencing site, after which the biomacromolecule is released into or near the sequencing area. In some embodiments, the present invention provides an agarose gel containing genomic DNA, wherein the agarose gel holds substantial fragments of genomic DNA larger than 200 kb in length, and comprises placing the agarose containing genomic DNA near an environment in which the DNA is sequenced (e.g., a surface, gel, or matrix), releasing the genomic DNA from the agarose into the environment (or near an environment in which further transport and handling are minimized), and performing sequencing. Release into the sequencing environment may be by application of an electric field or by digestion of the gel in agarose.

[0148] Polymer modification Block 206 The immobilized extended double-stranded nucleic acid is then denatured into a single-stranded form on a test substrate, thereby obtaining an immobilized first strand and an immobilized second strand of the nucleic acid. Each base of the immobilized second strand is located adjacent to the corresponding complementary base of the immobilized first strand. In some embodiments, denaturation is carried out by first extending or elongating the polynucleotide, and then separating the two strands by adding a denaturation solution.

[0149] In some embodiments, denaturation is chemical denaturation involving one or more reagents (e.g., 0.5 M NaOH, DMSO, formamide, urea, etc.). In some embodiments, denaturation is thermal denaturation (e.g., by heating the sample to 85°C or higher). In some embodiments, denaturation is enzymatic denaturation, such as by the use of a helicase or other enzyme having helicase activity. In some embodiments, polynucleotides are denatured by physical processes such as interaction with a surface or elongation beyond a critical length. In some embodiments, denaturation is all or part of the process.

[0150] In some embodiments, the attachment of the probe to modifications on repeating units of the polymer (e.g., phosphorylation on nucleotides or polypeptides in polynucleotides) is performed before an optional denaturation step.

[0151] In some embodiments, no selective denaturation of double-stranded polynucleotides is performed at all. In some such embodiments, the probe must be able to anneal to the duplex structure. For example, in some embodiments, the probe binds to individual strands of the duplex (e.g., using a PNA probe) through strand intrusion by inducing excessive breathing of the duplex, by recognizing the sequence in the duplex through a zinc-finger protein, or by using a Cas9 or similar protein that lyses the duplex and binds the guide RNA. In some embodiments, the guide RNA is provided as a gRNA containing an interlogging probe sequence and a repertoire of nuclear sequences.

[0152] In some embodiments, the double-stranded target contains a nick (e.g., a natural nick or one produced by DNase1 treatment). In such embodiments, under reaction conditions, one strand may fray or detach from the other (e.g., transiently denature), or spontaneous base pair breathing may occur. This transiently binds the probe before it is replaced by the native strand.

[0153] In some embodiments, a single double-stranded target polynucleotide is denatured so that each of the duplex strands is available for oligonucleotide binding. In some embodiments, the single polynucleotide is damaged and repaired by a denaturation step in sequencing or by another step (e.g., by the addition of a suitable DNA polymerase).

[0154] In some preferred embodiments, immobilization and linearization of double-stranded genomic DNA (during preparation for transient binding on the surface) include molecular combing, UV crosslinking of DNA to the surface, optional wetting, denaturation of double-stranded DNA through exposure to a chemical denaturant (e.g., alkaline solution, DMSO, etc.), optional exposure to an acidic solution after washing, and optional exposure to a preconditioning buffer.

[0155] Probe annealing Block 208 After an optional denaturation step, the method is followed by exposing the immobilized first and immobilized second strands to each pool of each oligonucleotide probe in a set of oligonucleotide probes, where each oligonucleotide probe in the set of oligonucleotide probes has a predetermined sequence and length. The exposure occurs under conditions that allow the individual probes in each pool of each oligonucleotide probe to bind to and form each heteroduplex with each portion (or more portions) of the immobilized first or immobilized second strand that is complementary to each oligonucleotide probe, thereby producing each instance of optical activity.

[0156] Figures 5A, 5B, and 5C illustrate examples of transiently conjugating different probes to a single polymer 502. Each probe (e.g., 504, 506, and 508) contains a specific interlogging sequence (e.g., a nucleotide or peptide sequence). After applying probe 504 to the polynucleotide 502, probe 504 is washed away from the polymer 502 in one or more washing steps. Probes 506 and 508 are then removed using similar washing steps.

[0157] Probe design and target In some embodiments, a probe is provided to a target polynucleotide in a solution. If the volume of the solution is sufficient to settle the polynucleotide on the surface or matrix, the probe can come into contact with the polynucleotide through diffusion and molecular collision. In some embodiments, the solution is agitated to bring the probe into contact with one or more polynucleotides. In some embodiments, the probe-containing solution is replaced to bring a fresh probe to the surface. In some embodiments, an electric field is used to attract the probe to the surface, for example, a positively deflected surface attracting negatively charged oligonucleotides.

[0158] In some embodiments, the target comprises a polynucleotide sequence, and the probe binding portion comprises, for example, an interlogging portion of a trimer, tetramer, pentamer, or hexamer oligonucleotide sequence, optionally one or more degenerate or universal positions, and optionally nucleotide spacers (e.g., one or more T nucleotides) or basic or non-nucleotide portions. As shown in Figures 6A and 6B, similar binding occurs along polynucleotide 602, regardless of the size of the oligo probes used (e.g., 604 and 610). The main inherent difference between oligos of different k-mer lengths is that the k-mer length determines the length of the binding site to which each probe binds (e.g., the trimer probe 604 primarily binds to 3-nucleotide length sites such as 606, and the pentamer probe 610 primarily binds to 5-nucleotide length sites such as 610).

[0159] In Figure 6A, trimer oligo probes are typically short. Such short sequences are usually not used as probes because they cannot bind stably unless very low temperatures and long incubation times are used. However, such probes can form transient binding to target polynucleotides, as required by the detection methods described herein. Furthermore, the shorter the oligonucleotide probe sequence, the fewer oligonucleotides present in the repertoire. For example, only 64 oligonucleotide sequences are needed for a complete repertoire of trimer oligos, while 256 oligonucleotide sequences are needed for a complete tetramer repertoire. Additionally, extremely short probes are modified in some embodiments to increase the melting temperature and, in some embodiments, to include degenerate (e.g., N) nucleotides. For example, four N nucleotides increase the stability of the trimer oligo to that of a heptamer.

[0160] In Figure 6B, this diagram shows the binding of the pentamer to the exact match position (612-3), the single-nucleotide mismatch position (612-2), and the double-nucleotide mismatch position (612-1).

[0161] The binding of any single probe is not sufficient to sequence a polynucleotide. In some embodiments, a complete repertoire of probes is required to reconstruct the polynucleotide sequence. Information about the location of oligo-binding sites, the temporally separated binding of probes to overlapping binding sites, the partial binding of mismatches between oligos and target nucleotides, the frequency of binding, and the duration of binding all contribute to sequence inference. For extended or elongated polynucleotides, the location of probe binding along the length of the polynucleotide contributes to the construction of a robust sequence. For double-stranded polynucleotides, a more reliable sequence arises from the simultaneous sequencing of both strands of the duplex (e.g., both complementary strands).

[0162] In some embodiments, a common reference probe sequence is added along with each of the oligonucleotide probes in the repertoire. For example, in Figures 7A, 7B, and 7C, the common reference probe 704 binds to the same binding site 708 on the target polynucleotide 702, independently of any additional probes included in the probe set (e.g., 706, 712, and 716). The presence of the reference probe 704 does not inhibit the binding of other probes to each binding site (e.g., 710, 714, 718, 720, and 722).

[0163] As shown in Figure 7C, binding sites 718, 720, and 722 demonstrate how individual probes (716-1, 716-2, and 716-3) bind to all possible sites, even when these sites overlap. In Figures 7A, 7B, and 7C, the probe sequences are represented by trimers. However, similar methods can be performed equally well with probes such as tetramers, pentamers, and hexamers.

[0164] In some embodiments, a set of oligonucleotide probes is a complete repertoire of oligos (e.g., each oligo of a given length). For example, the entire set of 1024 individual pentamers is coded and included in a specific repertoire according to one embodiment of the present disclosure. In some embodiments, a repertoire of multiple lengths is provided. In some embodiments, a set of oligonucleotide probes is a tiling series of oligo probes. In some embodiments, a set of oligonucleotide probes is a panel of oligo probes. In certain applications in synthetic biology (e.g., DNA data storage), sequencing involves finding an order of specific blocks of a sequence, where the blocks are designed to code desired data.

[0165] As shown in Figures 8A, 8B, and 8C, multiple probe sets (e.g., 804, 806, and 808) are applied to an arbitrary target polymer 802 in some embodiments. Each probe type will preferentially bind to a complementary binding site. In many embodiments, washing with a buffer between each cycle helps remove probes from previous sets.

[0166] In some embodiments, the probe for nucleic acid sequencing is an oligonucleotide, and the probe for epimodifications is a modification-binding protein or peptide (e.g., a methyl-binding protein such as MBD1) or an anti-modification antibody (e.g., an anti-methyl C antibody). In some embodiments, the oligo probe targets a specific site in the genome (e.g., a site with a known mutation). As shown in Figures 9A, 9B, and 9C, both the oligonucleotides (e.g., 804, 806, and 808) and the alternate probes (e.g., 902) are applied simultaneously (and by multiple cycles) to a polynucleotide or polymer 802 in some embodiments. A method for determining the target site of interest is provided by Liu et al., BMC Genomics 9: 509 (2008), which is incorporated herein by reference.

[0167] In some embodiments, each of the repertoire probes, or a subset of the repertoire probes, is applied one after the other (for example, one or a subset of bindings is detected first, then removed, then the next is added, detected and removed, and then the process proceeds, etc.). In some embodiments, all or a subset of the repertoire binding probes are added simultaneously, each binding probe is linked to a marker that codes entirely or partially for its identity, and the code for each binding probe is decoded by detection.

[0168] As shown in Figures 11A and 11B, a tiling series of probes is used in some embodiments to obtain information about the binding sites of multiple probes. In Figure 11A, the first tiling set 1104 is applied to the target polynucleotide 1102. Each probe in the subset of probes in the first tiling set 1108 contains one base 1108, thereby providing a 5x coverage of that one nucleotide in the target polynucleotide 1102. The coverage will be proportional to the kmer length of the probes in the tiling series (for example, a set of trimer oligos will provide a 3x coverage of each base in the target polynucleotide).

[0169] In some embodiments, when a set of oligonucleotide probes is tiled along a target nucleotide, problems can arise if there are discontinuities in the tiling path. For example, when using a pentameric oligonucleotide set, there are no oligos that can bind to one or more extensions of a sequence in the target molecule longer than five bases. In this case, in some embodiments, one or more approaches are taken. First, if the target polynucleotide contains a double-stranded nucleic acid, one or more base assignments follow a sequence(s) obtained from the complementary strand of the duplex. Second, if multiple copies of the target molecule are available, one or more base assignments depend on other copies of the same sequence on other copies of the target molecule. Third, in some embodiments, if a reference sequence is available, one or more base assignments follow the reference sequence, and the bases are annotated to indicate that they are artificially implanted from the reference sequence.

[0170] In some embodiments, certain probes are omitted from the repertoire for various reasons. For example, some probe sequences exhibit problematic interactions with themselves, with other probes in the repertoire, or with polynucleotides, such as self-complementarity or palindromic sequences (e.g., known probabilistic indiscriminate binding). In some embodiments, a minimum number of informative probes is determined for each type of polynucleotide. Within a complete repertoire of oligos, half of the oligos are perfectly complementary to the other half. In some embodiments, it is ensured that these complementary pairs (and others with problematic substantial complementarity) are not added to the polynucleotide simultaneously, but rather assigned to different probe subsets. In some embodiments, when both sense and antisense single-stranded DNA are present, sequencing is performed on only one member of each oligo complementary pair. Sequencing information obtained from both the sense and antisense strands is combined to construct the complete sequence. However, this method is undesirable because it loses the advantages given by sequencing both strands of a double-stranded polynucleotide simultaneously.

[0171] In some embodiments, the oligos include libraries prepared using conventional microarray synthesis. In some embodiments, the microarray library includes oligos that systematically bind to specific target regions of the genome. In some embodiments, the microarray library includes oligos that systematically bind to locations a specific distance away from the entire polynucleotide. For example, a library containing 1 million oligos may include oligos designed to bind approximately every 3,000 bases. Similarly, a library containing 10 million oligos may be designed to bind approximately every 300 bases, and a library containing 30 million oligos may be designed to bind every 100 bases. In some embodiments, the oligo sequences are computationally designed based on a reference genome sequence.

[0172] In some embodiments, the targeted portion of the genome is a specific locus. In other embodiments, the targeted portion of the genome is a panel or gene of interchromosomal loci (e.g., cancer-related genes) identified by genome-wide association studies. In some embodiments, the targeted loci are also dark matter of the genome, heterochromosomal regions of the genome that are typically repetitive, and complex loci surrounding repetitive regions. Such regions include telomeres, centromeres, the short arms of terminal centromere chromosomes, and other less complex regions of the genome. Traditional sequencing methods cannot address the repetitive portions of the genome, but if nanometer precision is high enough, the method can address these regions comprehensively.

[0173] In some embodiments, each oligonucleotide in a plurality of oligonucleotide probes contains a unique N-mer sequence, where N is an integer {1, 2, 3, 4, 5, 6, 7, 8, and 9} in the set, and all unique N-mer sequences of length N are represented by the plurality of oligonucleotide probes.

[0174] The longer the oligo used to construct the probe, the more likely it is that a palindromic or foldback sequence affecting that oligo will function as an efficient probe. In some embodiments, binding efficiency is substantially improved by reducing the length of such oligos by removing one or more degenerate bases. For this reason, the use of shorter interlogging sequences (e.g., tetramers) is advantageous. However, shorter probe sequences also exhibit less stable binding (e.g., lower binding temperature). In some embodiments, the binding stability of the oligo is enhanced using specific stable base modifications or oligo conjugates (e.g., stilbene caps). In some embodiments, fully modified trimers or tetramers (e.g., locked nucleic acids (LNAs)) are used.

[0175] In some embodiments, the unique N-mer sequence includes one or more nucleotide positions occupied by one or more degenerate nucleotides. In some embodiments, the degenerate position includes one of four nucleotides, and versions having each of the four nucleotides are provided in the reaction mix. In some embodiments, each degenerate nucleotide position at one or more nucleotide positions is occupied by a universal base. In some embodiments, the universal base is 2'-deoxyinosine. In some embodiments, the unique N-mer sequence is flanked at a single degenerate nucleotide position at the 5' end and flanked at a single degenerate nucleotide position at the 3' end. In some embodiments, the 5' single degenerate nucleotide and the 3' single degenerate nucleotide are each 2'-deoxyinosine.

[0176] In some embodiments, each oligonucleotide probe in a set of oligonucleotide probes has the same length M. In some embodiments, M is a positive integer greater than or equal to 2. Determining the sequence of at least a portion of the nucleic acid from multiple sets of positions on a test substrate (f) further utilizes the overlapping sequences of the oligonucleotide probes represented by the multiple sets of positions. In some embodiments, each oligonucleotide probe in a set of oligonucleotide probes shares M-1 sequence homology with another oligonucleotide probe in the set of oligonucleotide probes.

[0177] Probe labeling In some embodiments, each oligonucleotide probe in a set of oligonucleotide probes is bound to a label. Figures 14A–E illustrate different methods of labeling the probes. In some embodiments, the label is a dye, fluorescent nanoparticles, or light-scattering particles. In some embodiments, probe 1402 is directly bound to label 1406. In some embodiments, probe 1402 is indirectly labeled via a flap sequence 1410 containing a sequence 1408-B complementary to the sequence on oligo 1408-A.

[0178] Many types of organic dyes with desirable properties are available for labeling, some having high photostability and / or high quantum efficiency and / or minimal dark state and / or high solubility and / or low nonspecific binding. Atto542 is a suitable dye possessing several desirable qualities. Cy3B is a very bright dye, and Cy3 is also effective. Several dyes, such as the red dyes Atto655 and Atto647N, allow for the avoidance of wavelengths at which autofluorescence from cells or intracellular materials is commonly observed. Many types of nanoparticles are available for labeling. Beyond fluorescently labeled latex particles, this disclosure uses gold or silver particles, semiconductor nanocrystals, and nanodiamonds as nanoparticle labels. Nanodiamonds are particularly suitable as labels in some embodiments. Nanodiamonds emit light with high quantum efficiency (QE), have high photostability and a long fluorescence lifetime (e.g., approximately 20 ns, which can be used to reduce background observed from light scattering and / or autofluorescence), and are small (e.g., approximately 40 nm in diameter). DNA nanostructures and nanoballs can be exceptionally bright labels by incorporating organic dyes into their structures or by wiping away labels such as intercalating dyes.

[0179] In some embodiments, each indirect label explicitly indicates that the base identity is encoded in the sequence interlogging portion of the probe. In some embodiments, the label comprises one or more molecules of nucleic acid intercalating dyes. In some embodiments, the label comprises one or more types of dye molecules, fluorescent nanoparticles, or light-scattering particles. In some embodiments, it is preferable that the label does not rapidly photobleach in order to allow for longer imaging times.

[0180] Figures 12A, 12B, and 12C show the transient on-off binding of oligonucleotide 1204 with fluorescent label 1202 attached to target polynucleotide 1206. Label 1202 fluoresces regardless of whether probe 1204 binds to the binding site of target polynucleotide 1206. Similarly, Figures 13A, 13B, and 13C show the transient on-off binding of unlabeled oligonucleotide probe 1306. The binding event is detected by the intercalation of a dye (e.g., YOYO-1) into dubrex 1304 transiently formed from solution 1302. The intercalating dye exhibits a significant increase in fluorescence when bound to the double-stranded nucleic acid compared to free suspension in solution.

[0181] In some embodiments, the target-binding probe is not directly labeled. In some such embodiments, the probe contains a flap. In some embodiments, constructing (e.g., encoding) an oligonucleotide involves ligating a specific sequence unit to one end of each kmer in the set of oligonucleotides (e.g., a flap sequence). Each unit of the coding sequence on the flap acts as a docking site for different fluorescently labeled probes. To encode a 5-base probe sequence, the flap on the probe contains five different binding sites, each of which is, for example, a different DNA base sequence tangent to the next site. For example, the first position on the flap is adjacent to the probe sequence (the portion that binds to the polynucleotide target), the second position is adjacent to the first position, and so on. An advantage of using probe-flaps for sequencing is that each of the various probe-flaps is ligated to a set of fluorescently labeled oligos to generate a unique identifier tag for the probe sequence. In some embodiments, this is done by using four different labeled oligo sequences complementary to each position on the flap (e.g., 16 different labels in total).

[0182] In some embodiments, probes with A, C, T, and G defined are coded in such a way that the label reports only one defined nucleotide at a specific position in the oligonucleotide (and the remaining positions are degenerate). This requires just four color codings, one color per nucleotide.

[0183] In some embodiments, only one fluorophora is used throughout the process. In such embodiments, each cycle is divided into four subcycles, in each of which one of four bases is individually added at a designated position (e.g., position 1), followed by the addition of the next. In each cycle, the probe carries the same label. In this execution, the entire repertoire is used up in 20 cycles, resulting in significant time savings.

[0184] In some embodiments, the first base in the sequence is encoded by the first unit in the flap, the second base by the second unit, and so on. The order of the units in the flap corresponds to the order of the base sequence in the oligo. Different fluorescent labels are then docked onto each of the units (by complementary base pairing). In one example, the first position emits light at wavelengths of 500 nm to 530 nm, the second at wavelengths of 550 nm to 580 nm, the third at 600 nm to 630 nm, the fourth at 650 nm to 680 nm, and the fifth at 700 nm to 730 nm. The identity of the base at each location is then encoded, for example, by a fluorescent lifetime label. In one such example, the label corresponding to A has a longer lifetime C, C has a longer lifetime than G, and G has a longer lifetime than T. In the previous example, base A at position 1 would emit light at 500nm to 530nm and have the longest lifetime, while base G at position 3 would emit light at 600nm to 630nm and have the third longest lifetime, and so on.

[0185] As shown in Figure 14E, probe 1402 will contain sequence 1408A, which corresponds to sequence 1408-B. Sequence 1408-B is attached to flap region 1410. As an example of possible sequences that may result in the complete construct in Figure 14E, each of the four positions of 1410 is defined by the sequences AAAA (e.g., the position complementary to 1412), CCCC (e.g., the position complementary to 1414), GGGG (e.g., the position complementary to 1416), and TTTT (e.g., the position complementary to 1418), respectively. Thus the complete flap sequence would be 5'-AAAACCCCGGGGTTTT-3'. Each position is then encoded by a specific emission wavelength, and the four different bases that can be at that position are encoded by four different fluorescent lifetime-labeled oligonucleotides, where the lifetime / luminance ratio corresponds to the specific position and base coding within probe 1402 itself.

[0186] An example of appropriate code is as follows: • Position 1-A, Base code - TTTT - Emission peak 510, Lifetime / Brightness #1 • Position 1-C Base code - TTTT - Emission peak 510, lifetime / brightness #2 • Position 1-G, Base code - TTTT - Emission peak 510, Lifetime / Brightness #3 • Position 1-T base code - TTTT - Emission peak 510, lifetime / brightness #4 • Position 2-A, Base code - GGGG - Emission peak 560, Lifetime / Brightness #1 • Position 2-C base code - GGGG - Emission peak 560, lifetime / brightness #2 • Position 2-G, Base code -GGGG-, Emission peak 560, Lifetime / Brightness #3 • Position 2-T base code - GGGG - Emission peak 560, lifetime / brightness #4 • Position 3-A Base code - CCCC - Emission peak 610, lifetime / brightness #1 • Position 3-C base code - CCCC - Emission peak 610, lifetime / brightness #2 • Position 3-G base code - CCCC - Emission peak 610, lifetime / brightness #3 • Position 3-T base code - CCCC - Emission peak 610, lifetime / brightness #4 • Position 4-A, Base code - AAAA - Emission peak 660, Lifetime / Brightness #1 • Position 4-C Base code - GGGG - Emission peak 660, lifetime / brightness #2 • Position 4-G, Base code -GGGG-, Emission peak 660, Lifetime / Brightness #3 • Position 4-T base code - GGGG - Emission peak 660, lifetime / brightness #4

[0187] Alternatively, the four positions are coded by fluorescence lifetime, and the base is coded by fluorescence emission wavelength. In some embodiments, other measurable physical attributes may be used instead of coding, or in combination with wavelength and lifetime, if suitable. For example, the polarization or luminance of the emission can be measured to increase the repertoire of codes available for inclusion in the flap.

[0188] In some embodiments, toehold probes (e.g., described in Levesque et al., Nature Methods 10:865-867, 2013) are used. These probes are partially double-stranded and competitively destabilized when bound to mismatched targets (e.g., detailed in Chen et al., Nature Chemistry 5, 782-789, 2013). In some embodiments, toehold probes are used alone. In some embodiments, toehold probes are used to ensure correct hybridization. In some embodiments, toehold probes are used to facilitate the off-reaction of other probes bound to the target polynucleotide.

[0189] An example of a label excited by a common excitation line is a quantum dot. In some such embodiments of this example, Qdot 525, Qdot 565, Qdot 605, and Qdot 655 are selected so that each of the four nucleotides is different. Alternatively, four different laser lines are used to excite organic fluorophores, and their detected emission is split by an image splitter. In some other embodiments, the emission wavelengths are the same for two or more organic dyes, but the fluorescence lifetimes are different. Those skilled in the art will be able to conceive of numerous different coding and detection schemes without excessive effort and experimentation.

[0190] In some embodiments, different oligos in the repertoire are not added individually, but rather coded and pooled together. The simplest step up from one color and one oligo at a time is two colors and two oligos at a time. With one dye for each of the five oligos, it is reasonable to expect to pool up to approximately five oligos at a time, taking advantage of the direct detection of five identifiable single-dye flavors.

[0191] In more complex examples, the number of flavors and codes increases. For example, 64 different codes are required to individually code for each base in the complete repertoire of a trimer. Also, by example, 1024 different codes are required to individually code for each base in the complete repertoire of a pentamer. Such a large number of codes are achieved by having one code per oligo composed of multiple dye flavors. In some embodiments, a smaller set of codes is used to code a subset of the repertoire (subrepertoire), for example, in some examples, 64 codes are used to code 16 subsets of the complete 1024 sequence repertoire of a pentamer.

[0192] In some embodiments, a large repertoire of oligocodes can be obtained in numerous ways. For example, in some embodiments, beads are loaded with code-specific dyes, or the code based on DNA nanostructures contains dyes emitting different fluorescence wavelengths at optimal intervals (e.g., Lin et al., Nature Chemistry 4: 832-839, 2012). For example, Figures 14C and 14D show the use of beads 1412 supporting fluorescent labels 1414. In Figure 14C, the labels 1414 are coated onto the beads 1412. In Figure 14D, the labels 1414 are encapsulated within the beads 1412. In some embodiments, each label 1414 is a different type of fluorescent molecule. In some embodiments, all labels 1414 are the same type of fluorescent molecule (e.g., Cy3).

[0193] In some embodiments, a coding scheme is used, where a modular code is used to describe the position and identity of the bases in the oligonucleotide. In some embodiments, this is done by adding a coating arm to the probe, which includes a combination of labels that identify the probe. For example, if a library of each possible pentameric oligonucleotide probe is to be coded, the arm has five sites, each site corresponding to each of the five nucleobases in the pentamer, and each of the five sites is bound to one of the five distinguishable species. In one such example, a fluorophore with a specific peak emission wavelength corresponds to each of the positions (e.g., 500 nm for position 1, 550 nm for position 2, 600 nm for position 3, 650 nm for position 4, and 700 nm for position 5), and four fluorophores with the same wavelength but different fluorescence lifetimes code for each of the four bases at their respective positions.

[0194] In some embodiments, different labels or other binding reagents on the oligo are encoded by emission wavelength. In some embodiments, different labels are encoded by fluorescence lifetime. In some embodiments, different labels are encoded by fluorescence polarization. In some embodiments, different labels are encoded by a combination of wavelength and fluorescence lifetime.

[0195] In some embodiments, different labels are encoded by repeated on-off hybridization rates. Different binding probes with different binding-dissociation constants are used. In some embodiments, probes are encoded by fluorescence intensity. In some embodiments, probes are fluorescence intensities encoded by having a different number of attached non-self-quenching fluorophores. Individual fluorophores typically need to be sufficiently separated so as not to quench. In some embodiments, this is accomplished by utilizing rigid linkers or DNA nanostructures to hold labels in place at appropriate distances from each other.

[0196] An alternative embodiment for encoding by fluorescence intensity is to use dye variants that have similar emission spectra but different quantum yields or other measurable optical properties. For example, Cy3B with excitation / emission 558 / 572 is substantially brighter (e.g., quantum yield 0.67) than Cy3 with excitation / emission 550 / 570 and a quantum yield of 0.15, but has a similar absorption / emission spectrum. In some such embodiments, a 532 nm laser is used to excite both dyes. Another suitable dye is Cy3.5 (excitation / emission 591 / 604 nm), which has an upward-shifted excitation and emission spectrum but is nevertheless excited by a 532 nm laser. However, excitation at that wavelength is not optimal for Cy3.5, and the emission of the dye is not expected to be as bright with a bandpass filter for Cy3. Atto532 with excitation / emission 532 / 553 has a quantum yield of 0.9, and is expected to be bright because the 532 nm laser hits this dye in the sweet spot.

[0197] Another approach to obtaining multiple codes using a single excitation wavelength is to measure the emission lifetime of the dyes. In one example according to such embodiments, a set including Alexa Fluor 546, Cy3B, Alexa Fluor 555, and Alexa Fluor 555 is used. In some examples, other dye sets are more useful. In some embodiments, the repertoire of codes is also expanded by utilizing FRET pairs and by measuring the polarization of the emission. Another way to increase the number of labels is by coding with multiple colors.

[0198] Figure 15 shows an example of fluorescence from transient binding of oligonucleotide probes to polynucleotides. Selected frames from the time series (e.g., frame numbers 1, 20, 40, 60, 80, 100) indicate the presence (e.g., black spots) and absence (e.g., white areas) of signals at specific sites, which are indicators of on-off binding. Each frame shows fluorescence from multiple bound probes along the polynucleotide. The aggregate image shows the aggregate fluorescence from all past frames, indicating all sites where the oligonucleotide probes bound.

[0199] Transient binding of probes to target polynucleotides Probe binding is a mechanical process, and each bound probe always has some probability of becoming unbound (determined by various factors, including temperature and salt concentration). Thus, there is always an opportunity to replace one probe with another. For example, in one embodiment, a complementary strand of the probe is used, which causes continuous competition between annealing to the target DNA extended on the surface and to the complementary strand in solution. In another embodiment, the probe has three parts: the first part is complementary to the target, the second part is partially complementary to the target and partially complementary to the oligo in solution, and the third part is complementary to the oligo in solution. In some embodiments, gathering information about the precise spatial location of units of chemical structure helps determine the structure and / or arrangement of the polymer. In some embodiments, the location of the probe binding site is determined with nanometer-scale or even sub-nanometer-scale precision (e.g., by using a single-molecule localization algorithm). In some embodiments, multiple physically closer observed binding sites are resolvable by diffraction-limited optical imaging, and the binding events are resolved because they are temporally separated. The nucleic acid sequence is determined based on the identity of the probe that binds to each site.

[0200] Exposure occurs under conditions that allow each individual probe in each pool of oligonucleotide probes to transiently and reversibly bind to a portion of the immobilized first or second chain complementary to the individual probe, forming each heteroduplex, thereby producing an instance of optical activity. In some embodiments, residence time (e.g., duration and / or maintenance of binding by a particular probe) is used to determine whether the binding event is a perfect match, mismatch, or spurious.

[0201] In some embodiments, exposure occurs under conditions that allow each individual probe in each pool of oligonucleotide probes to transiently and reversibly bind to a portion of a complementary immobilized first or second chain, thereby forming each heteroduplex, and thereby repeatedly producing each instance of optical activity.

[0202] In some embodiments, sequencing involves subjecting the extended polynucleotide to transient interactions from each of a complete sequence repertoire of probes provided sequentially (the solution containing one probe sequence is removed, and the solution containing the next probe is added). In some embodiments, the binding of each probe is performed under conditions that cause transient binding of the probe. For example, binding is performed at 25°C for one probe and 30°C for the next. The probes may also be bound in a set, and all of the transiently binding probes may be collected in a set and used together in substantially the same manner. In some such embodiments, each probe sequence in the set is differentially labeled or differentially encoded.

[0203] In some embodiments, transient binding is performed in a buffer containing small amounts of divalent cations but no monovalent cations. In some embodiments, the buffer comprises 5 mM Tris-HCl, 10 mM magnesium chloride, mm EDTA, 0.05% Tween-20, and pH 8. In some embodiments, the buffer comprises magnesium chloride in concentrations of less than 1 nM, less than 5 nM, less than 10 nM, or less than 15 nM.

[0204] In some embodiments, multiple conditions are used to facilitate transient binding. In some embodiments, for each pentamer species from the entire repertoire of probe species, e.g., a repertoire of 1024 possible pentamers, certain conditions are used for one probe species depending on its Tm, and other conditions are used for another probe depending on its Tm, and so on. In some embodiments, only 512 non-complementary pentamers are provided (e.g., because both target polynucleotide chains are present in the sample). In some embodiments, each probe addition comprises a mixture of probes containing five specific bases and two degenerate bases (so a heptomer is 16), all labeled with the same label that functions as a single pentamer in terms of its ability to interlog the sequence. The degenerate bases add stability without increasing the complexity of the probe set.

[0205] In some embodiments, the same conditions are provided to multiple probes sharing the same or similar Tm. In some such embodiments, each probe in the repertoire includes a different coded label (or label depending on what is being identified). In such examples, the temperature is maintained through multiple probe exchanges and then increased for the next series of probes sharing the same or similar Tm.

[0206] In some embodiments, during the probe binding period, the temperature is modified so that the probe binding behavior is determined at one or more temperatures. In some embodiments, a similar melt curve is performed, where the binding behavior or binding pattern to the target polymer is correlated with a stepwise series of temperatures over a selected range (e.g., 10°C to 65°C or 1°C to 35°C).

[0207] In some embodiments, Tm is calculated, for example, by a neighbor method parameter. In other embodiments, Tm is derived empirically. For example, the optimal melting temperature range is derived by running a melting curve (e.g., measuring the degree of melting by absorption over various temperatures). In some embodiments, the composition of the probe set is designed according to their theoretically matching Tm, verified by empirical testing. In some embodiments, bonding is performed at a temperature substantially below Tm (e.g., up to 33°C lower than the calculated Tm). In some embodiments, an empirically defined optimal temperature for each oligo is used for bonding each oligo in sequencing.

[0208] In some embodiments, instead of modifying the temperature for oligonucleotide probes having different Tm values, or in addition to that, the concentrations of the probe and / or salt are modified, and / or the pH is modified. In some embodiments, the electrical bias of the surface is repeatedly switched between positive and negative to actively promote transient binding between the probe and one or more target molecules.

[0209] In some embodiments, the concentration of the oligo used is adjusted according to the AT vs. GC content of the oligo sequence. In some embodiments, a higher oligo concentration is provided for oligos with a higher GC content. In some embodiments, a buffer (e.g., CTAB, betaine, or a chaotropic reagent, such as a buffer containing tetramethylammonium chloride (TMACl)) is used at a concentration of 2.5 M to 4 M to equalize the effect of the base composition.

[0210] In some embodiments, probes are non-uniformly distributed across the sample (e.g., the flow chamber, slide, polynucleotide(s) length, and / or aligned array of polynucleotides) due to probabilistic effects or the design of the sequencing chamber (e.g., vortices in the flow cell trapping probes at the corners of the nanochannels or against the walls of the nanochannels). Localized probe deficiencies are addressed by ensuring efficient mixing or agitation of the probe solution. In some embodiments, this is done using acoustic waves by including disturbance-generating particles in the solution and / or by assembling a flow cell (e.g., a herringbone pattern on one or more surfaces) that generates turbulence. In addition, due to the laminar flow present in the flow cell, there is typically little mixing, and the solution near the surface mixes only slightly with the bulk solution. This poses problems when removing reagent / bound probes present near the surface and bringing fresh reagent / probes to the surface. This can be countered by implementing the disturbance-generating approaches described above and / or by performing large-scale fluid flow / exchange across the entire surface. In some embodiments, after the target molecules are arranged, non-fluorescent beads or spheres are attached to the surface to give the surface landscape a rough texture. This creates vortices and flows necessary for more effective mixing and / or exchange of fluids near the surface.

[0211] In some embodiments, the entire repertoire or a subset is added together. In some such embodiments, a buffer (e.g., TMACl or guanidine thiocyanate, as described in U.S. Patent Application Publication 2004 / 0058349) is used to equalize the effect of the base composition. In some embodiments, probe species having the same or similar Tm are added together. In some embodiments, the probe species added together are not differentially labeled. In some embodiments, the probe species added together are differentially labeled. In some embodiments, the differential labeling is a luminescent label having different luminance, lifetime or wavelength, and / or combinations of such physical properties.

[0212] In some embodiments, two or more oligonucleotides are used together, and their binding sites are determined without using conditions to distinguish between the signals of different oligonucleotides (e.g., the oligonucleotides are labeled with the same color). When both duplex strands are available, obtaining binding site data from both strands allows for identification between two or more oligonucleotides as part of the assembly algorithm. In some embodiments, one or more reference probes are added along with each probe in the repertoire, and the assembly algorithm can then use the binding sites of such reference probes to scaffold or set up the sequence assembly.

[0213] In one alternative embodiment, the probes bind stably, but an external trigger that switches the environment to off mode controls their transient nature. In a less restrictive embodiment, the trigger is heat, pH, electric field, or reagent exchange that releases the probes. The probes can then re-bind once the environment is switched back to on mode. In some embodiments, if binding does not saturate all sites in the first round of binding, oligos in the second cycle of binding bind to a different set of sites different from those in the first cycle. In some embodiments, these cycles are performed multiple times at a controllable rate.

[0214] In some embodiments, the transient coupling lasts for 1 millisecond or less, 50 milliseconds or less, 500 milliseconds or less, 1 microsecond or less, 10 microseconds or less, 50 microseconds or less, 500 microseconds or less, 1 second or less, and for 2 seconds or less, 5 seconds or less, or 10 seconds or less.

[0215] In transient coupling approaches that ensure a continuous supply of fresh probes, photobleaching of fluorophores does not pose a major problem, and sophisticated field diaphragms or Powell lenses are not required to limit emission. Therefore, the selection of fluorophores (or the provision of anti-bleaching agents, redox systems) is not critical, and in some such embodiments, relatively simple optical systems can be constructed (for example, f-stops to prevent illumination of molecules not in the camera's field of view are not highly necessary).

[0216] In some embodiments, another advantage of transient binding is that multiple measurements are performed at each binding site along the polynucleotide, which can increase the reliability in detection accuracy. For example, in some cases, due to the typical probabilistic nature of molecular processes, the probe may bind to the wrong location. With transiently bound probes, such detached and isolated binding events can be discarded, and only those binding events proven by multiple detected interactions are accepted as validated detection events for sequencing purposes.

[0217] Detection of transient binding and location of binding sites Transient binding is an essential component that enables localized subdiffraction levels. Each probe in a set of transiently binding probes has a probability at any given time that it will bind to the target molecule or be present in solution. Therefore, not all binding sites will be bound to a probe at any given time. This allows for the detection of binding events at sites closer than the diffraction limit of light (e.g., two sites on the target molecule separated by only 10 nm). For example, if the sequence AAGCTT is repeated after 60 bases, it means the repeating sequences are approximately 20 nm apart (if the target is stretched and straightened to a Watson-Crick base length of approximately 0.34 nm). 20 nanometers is normally indistinguishable by optical imaging. However, if the probe binds to two sites at different time intervals between imaging, they can be detected individually. This enables super-resolution imaging of binding events. Nanometer-scale precision is particularly important for resolving repeats and determining their number.

[0218] In some embodiments, multiple binding events to a site in the target are determined by analyzing data from a repertoire rather than from a single probe sequence, and by considering events arising from partially overlapping sequences. In one example, the same (actually, subnanometer-level) site is bound by hexamer probes ATTAAG and TTAAGC, which share a common 5-base sequence, each validating the other, and extend the sequence of one base on either side of the 5-base sequence. In some examples, the bases on each side of the 5-base sequence are mismatches (terminal mismatches are typically expected to be more tolerable than internal mismatches), and only the 5-base sequence present in both binding events is validated.

[0219] In some alternative embodiments, transient single-molecule bonding is detected by non-optical methods. In some embodiments, the non-optical method is an electrical method. In some embodiments, transient single-molecule bonding is detected by non-fluorescence methods, where direct excitation is absent, and rather bioluminescence or chemiluminescence mechanisms are used.

[0220] In some embodiments, each base in the target nucleic acid is interrogated by multiple oligonucleotides whose sequences overlap. This repeated sampling of each base allows for the detection of rare single-nucleotide variants or mutations in the target polynucleotide.

[0221] Some embodiments of this disclosure take into account the repertoire of binding interactions (beyond the threshold binding period) that each oligonucleotide had with polynucleotides under analysis. In some embodiments, sequencing involves not only stitching or reconstructing sequences from perfect matches, but also obtaining sequences by first analyzing the binding tendencies of each oligo. In some embodiments, transient bindings are recorded as a means of detection but are not used to improve localization.

[0222] Imaging technology that detects optical activity and determines the location of the binding site. Block 214: A two-dimensional imager is used to measure the location and duration of each instance of optical activity occurring during exposure on the test substrate.

[0223] Measuring locations on the test substrate involves inputting frames of data measured by a two-dimensional imager into a trained convolutional neural network. Each data frame contains each instance of optical activity within a plurality of instances of optical activity. Each instance of optical activity within the plurality of instances of optical activity corresponds to an individual probe coupled to a fixed first chain or a portion of a fixed second chain. In response to the input, the trained convolutional neural network identifies the location on the test substrate of one or more instances of optical activity within the plurality of instances of optical activity.

[0224] In some embodiments, the detector is a two-dimensional detector, and the binding event is localized to nanometer-scale accuracy (e.g., by using a single-molecule localization algorithm). In some embodiments, the interaction feature includes the duration of each binding event, which corresponds to the affinity between the probe(s) and the molecule. In some embodiments, the feature is a location on a surface or matrix, which corresponds to the location within an array of specific molecules (e.g., polynucleotides corresponding to specific gene sequences).

[0225] In some embodiments, each instance of optical activity has an observation metric that satisfies a predetermined threshold. In some embodiments, the observation metric includes duration, signal-to-noise ratio, photon count, or intensity. In some embodiments, the predetermined threshold is satisfied if each instance of optical activity is observed in one frame. In some embodiments, the predetermined threshold is satisfied if each instance of optical activity is relatively low and each instance of optical activity is observed in one-tenth of a frame.

[0226] In some embodiments, a predetermined threshold distinguishes between (i) a first form of binding, where each residue of the unique N-mer sequence binds to a complementary base in the immobilized first or second strand of nucleic acid, and (ii) a second form of binding, where at least one mismatch exists between the unique N-mer sequence and the sequence in the immobilized first or second strand of nucleic acid to which each oligonucleotide probe binds to form each optically active instance.

[0227] In some embodiments, each oligonucleotide probe in a set of oligonucleotide probes has its own corresponding predetermined threshold.

[0228] In some embodiments, a predetermined threshold is determined based on observing one or more, two or more, three or more, four or more, five or more, or six or more binding events at specific locations along the polynucleotide.

[0229] In some embodiments, a predetermined threshold for each oligonucleotide probe in a set of oligonucleotide probes is derived from a training dataset (e.g., a dataset derived from information obtained by applying a method to sequencing lambda phages).

[0230] In some embodiments, a predetermined threshold for each oligonucleotide probe in a set of oligonucleotide probes is derived from a training dataset. The training set includes, for each oligonucleotide probe in the set of oligonucleotide probes, a measure of observational metrics for each oligonucleotide probe in binding to a reference sequence, such that each residue of the unique N-mer sequence of each oligonucleotide probe binds to a complementary base in the reference sequence.

[0231] In some embodiments, the reference sequence is immobilized on a reference substrate. In some embodiments, the reference sequence is included with the nucleic acid and immobilized on a test substrate. In some embodiments, the reference sequence includes all or part of the genome of PhiX174, M13, lambda phage, T7 phage, E. coli, budding yeast, or fission yeast. In some embodiments, the reference sequence is a synthetic construct of a known sequence. In some embodiments, the reference sequence includes all or part of rabbit globin RNA (for example, if the nucleic acid contains RNA, or if only one strand of a polynucleotide is sequenced).

[0232] In some embodiments, exposure is in the presence of a first label in the form of an intercalating dye. Each oligonucleotide probe in a set of oligonucleotide probes is conjugated with a second label. The first and second labels have overlapping donor emission spectra and acceptor excitation spectra, causing one of the first and second labels to fluoresce when the first and second labels are close to each other. Each instance of optical activity originates near the intercalating dye, intercalating each heteroduplex between the oligonucleotide and the immobilized first or immobilized second chain to the second label. In some embodiments, exposure and fluorescence involve a Förster resonance energy transfer (FRET) method. In such embodiments, the intercalating dye comprises a FRET donor, and the second label comprises a FRET acceptor.

[0233] In some embodiments, the signal is detected by FRET from an intercalating dye to a label on the probe or target sequence. In some embodiments, after the target is immobilized, all ends of the target molecule are labeled by a terminal transferase incorporating a fluorescently labeled nucleotide, for example, that acts as a FRET partner. In some such embodiments, the probe is labeled at one end with Cy3B or Atto 542 labeling.

[0234] In some embodiments, FRET is replaced by photoactivation. In such embodiments, the donor (e.g., a label on a template) contains a photoactivator, and the acceptor (e.g., a label on an oligonucleotide) becomes a fluorophore in an inactive or dark state (e.g., Cy5 can be darkened by caging with 1 mg / mL NaBH4 in 20 mM Tris, 2 mM EDTA, and 50 mM NaCl at pH 7.5 prior to fluorescence imaging experiments). In such embodiments, the fluorescence of the darkened fluorophore is switched on when in the vicinity of the activator.

[0235] In some embodiments, the exposure is in the presence of a first label in the form of an intercalating dye (e.g., a photoactivator). Each oligonucleotide probe in a set of oligonucleotide probes is conjugated to a second label (e.g., a darkened fluorophore). The first label causes the second label to fluoresce when the first and second labels are in proximity to each other. Each instance of optical activity is derived from the vicinity of an intercalating dye that intercalates each heteroduplex between a first strand immobilized to the oligonucleotide and a second strand immobilized to the oligonucleotide, to the second label.

[0236] In some embodiments, the exposure is in the presence of a first label in the form of an intercalating dye (e.g., a darkened fluorophore). Each oligonucleotide probe in a set of oligonucleotide probes is conjugated to a second label (e.g., a photoactivator). The second label causes the first label to fluoresce when the first and second labels are in proximity to each other. Each instance of optical activity is derived from the vicinity of an intercalating dye that intercalates each heteroduplex between a first strand immobilized to the oligonucleotide and a second strand immobilized to the oligonucleotide, to the second label.

[0237] In some embodiments, the exposure is in the presence of an intercalating dye. Each instance of optical activity intercalates each heteroduplex between the oligonucleotide and the immobilized first strand or the immobilized second strand from the fluorescence of the intercalating dye, where each instance of optical activity is greater than its fluorescence before the intercalating dye intercalates each heteroduplex. The enhanced fluorescence (more than 100-fold) of one or more dyes intercalating into the duplex provides a point-source-like signal for the single-molecule localization algorithm, enabling precise determination of the location of the binding site. The intercalating dye intercalates into the duplex, resulting in a significant number of heteroduplex binding events for each binding site that are reliably detected and precisely localized.

[0238] In some embodiments, each oligonucleotide probe in a set of oligonucleotide probes generates a first instance of optical activity by binding to a complementary portion of the immobilized first strand and a second instance of optical activity by binding to a complementary portion of the immobilized second strand. In some embodiments, a portion of the immobilized first strand generates an instance of optical activity by binding of its complementary oligonucleotide probe, and a portion of the immobilized second strand complementary to a portion of the immobilized first strand generates another instance of optical activity by binding of its complementary oligonucleotide probe.

[0239] In some embodiments, each oligonucleotide probe in a set of oligonucleotide probes generates two or more first instances of optical activity by binding to two or more complementary portions of the immobilized first strand and two or more second instances of optical activity by binding to two or more complementary portions of the immobilized second strand.

[0240] In some embodiments, each oligonucleotide probe binds three or more times to a portion of a complementary immobilized first or second chain during exposure, thereby resulting in three or more instances of optical activity, each instance of optical activity representing one binding event among multiple binding events.

[0241] In some embodiments, each oligonucleotide probe binds five or more times to a portion of a complementary immobilized first or second chain during exposure, thereby resulting in five or more instances of optical activity. Each instance of optical activity represents one binding event among multiple binding events.

[0242] In some embodiments, each oligonucleotide probe binds to a portion of a complementary immobilized first or second chain more than 10 times during exposure, thereby resulting in more than 10 instances of optical activity, each instance of optical activity representing one binding event among multiple binding events.

[0243] In some embodiments, exposure occurs for a period of 5 minutes or less, 4 minutes or less, 3 minutes or less, 2 minutes or less, or 1 minute or less.

[0244] In some embodiments, exposure occurs over one or more frames of the two-dimensional imager. In some embodiments, exposure occurs over two or more frames of the two-dimensional imager. In some embodiments, exposure occurs over 500 or more frames of the two-dimensional imager. In some embodiments, exposure occurs over 5,000 or more frames of the two-dimensional imager. In some embodiments, when the optical activity is low (e.g., there are almost no instances of probe coupling), one frame of transient coupling is sufficient to localize the signal.

[0245] In some embodiments, the duration of the exposure instance is determined by the estimated melting temperature of each oligonucleotide probe in the set of oligonucleotide probes used for the exposure instance.

[0246] In some embodiments, optical activity includes fluorescence emission from a label. Each label is excited, and the corresponding emission wavelength is detected separately using different filters in a filter wheel. In some embodiments, emission lifetime is measured using a fluorescence lifetime imaging (FLIM) system. Alternatively, the wavelength is split and emitted to different quanta of a single sensor, or to four different sensors. A method for splitting the spectrum at a CCD pixel using a prism is described in Lundquit et al., Opt Lett., 33:1026-8, 2008. In some embodiments, a spectrophotographer is also used. Alternatively, in some embodiments, the emission wavelength is combined with the luminance level to provide information regarding the residence time of the probe at the binding site.

[0247] Several detection methods, such as scanning probe microscopy (including fast atomic force microscopy) and electron microscopy, can resolve nanometer distances when polynucleotide molecules are extended in the detection plane. However, these methods do not provide information about the optical activity of fluorophores. Several optical imaging techniques exist for detecting fluorescent molecules with super-resolution accuracy. These include stimulated emission suppression microscopy (STED), probabilistic reconstruction optical microscopy (STORM), super-resolution fluctuation imaging (SOFI), single-molecule localization microscopy (SMLM), and total internal reflection fluorescence (TIRF) microscopy. In some embodiments, the SMLM approach, which is most similar to point accumulation (PAINT) in nanoscale topography, is preferred. These methods typically require one or more lasers for exciting the fluorophores, a focus detection / fixation mechanism, a CCD camera, a suitable objective lens, a relay lens, and a mirror. In some embodiments, the detection step includes capturing multiple image frames (e.g., a movie or video) to record the probe binding on and off.

[0248] The SMLM method relies on high photon counts. High photon counts improve the accuracy of determining the centroid of the fluorophores generated in a Gaussian pattern, but the requirement for high photon counts is also related to long-term image acquisition and reliance on bright, photostable fluorophores. High solution concentrations of probes can be achieved without inducing harmful background by using quenched probe molecular beacons or by having two or more labels of the same type, for example, one on each side of the oligo. In such embodiments, these labels are quenched in solution via inter-dye interactions. However, once the labels are bound to their targets, they become separable and fluoresce brightly, making the labels easier to detect.

[0249] In some embodiments, the on-rate of a probe is manipulated (e.g., increased) by increasing the probe concentration, increasing the temperature, or increasing molecular crowding (e.g., by including PEG400, PEG800, etc., in the solution). The off-rate can be increased by reducing the thermal stability of the probe by modifying its chemical composition, adding destabilizing adducts, or, in the case of oligonucleotides, reducing their length. In some embodiments, the off-rate is also facilitated by increasing the temperature, decreasing the salt concentration (e.g., increasing stringency), or modifying the pH.

[0250] In some embodiments, the probe concentration used is increased by making the probe essentially nonfluorescent until they bind. One way to do this is by having the binding induce a photoactivation event. Another way is to quench the label until binding occurs (e.g., molecular beacon). Another way is for the signal to be detected as a result of an energy transfer event (e.g., FRET, CRET, BRET). In one embodiment, a biopolymer on the surface carries a donor and the probe carries an acceptor, or vice versa. In another embodiment, an intercalating dye is provided in solution, and a FRET interaction exists between the intercalating dye and the probe upon binding of the labeled probe. An example of an intercalating dye is YOYO-1, and an example of a label on the probe is ATTO655. In another embodiment, the intercalating dye is used without a FRET mechanism, and both the single-stranded target sequence and the probe sequence on the surface are unlabeled, and the signal is detected only if binding creates a double-stranded structure through which the intercalating dye intercalates. Intercalating dyes, depending on their identity, are 100 or 1000 times less bright if they do not intercalate into DNA and are absent from the solution. In some embodiments, TIRF or thin-layer oblique illumination (HILO) microscopy (e.g., Mertz et al., J. of Biomedical Optics, 15(1): 016027, 2010) is used to eliminate any background signal from the intercalating dye in the solution.

[0251] However, in some embodiments, highly labeled probes induce high background fluorescence, which makes detection of the surface signal difficult. In some embodiments, this is addressed by a DNA dye or intercalating dye that labels the duplex formed on the surface. The dye will not intercalate if the target is single-stranded or has a single-stranded probe, but will intercalate if a duplex is formed between the probe and the target. In some embodiments, the probe is unlabeled, and the detected signal is solely due to the intercalating dye. In some embodiments, the probe is labeled with a label that acts as a FRET partner to the intercalating dye or DNA dye. In some embodiments, the intercalating dye is a donor that is coupled with acceptors of different wavelengths, thereby encoding the probe with multiple fluorophores.

[0252] In some embodiments, the detection step includes detecting multiple binding events to each complementary site. In some embodiments, the multiple events originate from the same probe molecule that binds on or off, or are replaced by another molecule with the same specificity (e.g., it is specific to the same sequence or molecular structure), and this occurs multiple times. In some embodiments, binding on or off is not affected by changes in conditions. For example, both binding on and binding off occur under the same conditions (e.g., salt concentration, temperature, etc.) due to a weak probe-target interaction.

[0253] In some embodiments, sequencing is performed by imaging multiple on-off binding events at multiple locations on a single target polynucleotide with probe lengths that are shorter, the same length, or no more than 10 times shorter. In such embodiments, the longer target polynucleotide is fragmented, or a panel of fragments is pre-selected and arranged on the surface so that each polynucleotide molecule can be resolved individually. In these examples, the frequency or duration of probe binding to specific locations is used to determine whether the probe corresponds to the target sequence. The frequency or duration of probe-orb binding determines whether the probe corresponds to all or part of the target sequence (including the remaining mismatched bases).

[0254] The appearance of side-by-side overlap between target polynucleotides is detected in some embodiments by increased fluorescence from a DNA dye. In some embodiments where no dye is used, the overlap is detected by an increased frequency of apparent binding sites along the segmentation. For example, in some cases where diffraction-limited molecules appear to overlap optically but do not actually overlap physically, they are super-resolved using single-molecule localization as described elsewhere in this disclosure. Where end-on-end overlap occurs, in some embodiments, parallel polynucleotides are distinguished from true continuous lengths using labels that mark the ends of the polynucleotides. In some embodiments, such optical chimeras are discarded as artifacts if many copies of the genome are predicted and only one appearance of the apparent chimera is found. Similarly, in embodiments where molecular ends (diffraction-limited) appear to overlap optically but do not physically overlap, they are resolved by the methods of this disclosure. In some embodiments, signals emanating from very close labels are degraded because the localization is so precise.

[0255] In some embodiments, sequencing is performed by imaging multiple on-off binding events at multiple locations on a single target polynucleotide longer than the probe. In some embodiments, the locations of probe-binding events on the single polynucleotide are determined. In some embodiments, the locations of probe-binding events on the single polynucleotide are determined by extending the target polynucleotide, thereby detecting and resolving different locations along its length.

[0256] In some embodiments, distinguishing the optical activity of an unbound probe from that of a probe bound to a target molecule requires rejection or removal of the signal from the unbound probe. In some such embodiments, this is performed, for example, by using an evanescent field or a waveguide for illumination, or by utilizing FRET pair labeling, or by utilizing photoactivation, to detect the probe at a specific location (e.g., Hylkje et al., Biophys J. 2015; 108(4): 949-956).

[0257] In some embodiments, the probe is unlabeled, but interaction with the target is detected by a DNA dye such as intercalating dye 1302, which intercalates duplex and initiates fluorescence emission 1304 when or if binding occurs (e.g., shown in Figures 13A–13C). In some embodiments, one or more intercalating dyes intercalate duplex at any given time. In some embodiments, when an intercalating dye is intercalated, the fluorescence emitted therefrom is orders of magnitude greater than the fluorescence emitted by the intercalating dye free-floating in solution. For example, the signal from an intercalated YOYO-1 dye is about 100 times greater than the signal from a free YOYO-1 dye in solution. In some embodiments, when brightly stained (or partially photobleached) double-stranded polynucleotides are imaged, individual signals along the observed polynucleotides may correspond to a single intercalating dye molecule. To facilitate the exchange of the YOYO-1 dye in the duplex and to obtain a bright signal, a redox-oxidation system (ROX) containing methyl viologen and ascorbic acid is provided in a binding buffer in some embodiments.

[0258] In some embodiments, single molecule sequencing by detecting incorporation of nucleotides labeled with a single dye molecule (e.g., as performed in Helicos and PacBio sequencing) introduces errors when the dye is not detected. In some instances, this is because the dye has photobleached, the cumulative signal detected is weak due to dye blinking, the emission of the dye is too weak, or the dye has entered a long-lived dark photophysical state. In some embodiments, this is overcome by a plurality of alternative methods. The first is to label the dye with a robust individual dye having suitable photophysical properties (e.g., Cy3B). Another is to provide buffer conditions and additives that reduce photobleaching and dark photophysical states (e.g., beta-mercaptoethanol, Trolox, vitamin C and its derivatives, redox systems). Another is to minimize exposure to light (e.g., having a more sensitive detector that requires shorter exposure, or providing stroboscopic illumination). The second is to label with nanoparticles such as quantum dots (e.g., Qdot655), fluorospheres, nanodiamonds, plasmon resonance particles, light scattering particles instead of a single dye. Another is to have more dyes per nucleotide rather than a single dye (e.g., as shown in FIGS. 14C and 14D). In this case, the plurality of dyes 1414 are arranged in a way that minimizes self-quenching (e.g., using a rigid nanostructure 1412 such as a DNA origami with fairly well-spaced dyes) or are arrayed linearly with a rigid linker with spacing.

[0259] In some embodiments, the detection error rate is further reduced (and the useful lifetime of the signal is increased) in the presence of a solution of one or more compounds selected from urea, ascorbic acid or its salts, and isoascorbic acid or its salts, beta-mercaptoethanol (BME), DTT, redox systems, or Trolox.

[0260] In some embodiments, transient binding of the probe to the target molecule is sufficient to reduce errors due to the photophysical properties of the dye. The information obtained during the imaging step is a collection of many on / off interactions of different labeled probes. Therefore, even if one label is photobleached or in a dark state, the label on another bound probe that arrives on the molecule will not be photobleached or in a dark state, and thus in some embodiments, will provide information about the location of their binding sites.

[0261] In some embodiments, the signal from the label at each transient coupling event is projected through an optical path (typically providing magnification) to extend to more than one pixel of a 2D detector. The point spreading function (PSF) of the signal is plotted, and the centroid of the PSF is imaged as the precise location of the signal. In some embodiments, this localization is performed to sub-diffraction (e.g., super-resolution) and even to sub-nanometer accuracy. The accuracy of localization is inversely proportional to the number of photons collected. Therefore, the more photons emitted per second by the fluorescent label, or the longer the photons are collected, the higher the accuracy.

[0262] In one example, as shown in Figures 10A and 10B, both the number of binding events and the number of photons collected at each binding site correlate with the degree of localization performed. For the target polymer 1002, the minimum number of binding events 1004-1 and the minimum number of photons 1008-1 recorded at the binding site correlate with the least accurate localization 1006-1 and 1010-1, respectively. As the number of binding events 1004-2, 1004-3, or the number of photons recorded 1008-2, 1008-3 increases at the binding site, the degree of localization increases, respectively, to 1006-2, 1006-3, and 1010-2, 1010-3. In Figure 10A, different numbers of probabilistic binding events (e.g., 1004-1, 1004-2, 1004-4) of the labeled probe on polynucleotide 1002 result in different degrees of probe localization (1006-1, 1006-2, 1006-3), where more binding events (e.g., 1004-2) correlate with a higher degree of localization (e.g., 1006-2), and fewer binding events (e.g., 1004-1) correlate with a lower degree of localization (e.g., 1006-1). In Figure 10B, similarly, different numbers of detected photons (e.g., 1008-1, 1008-2, and 1008-3) result in different degrees of localization (1010-1, 1010-2, and 1010-3, respectively).

[0263] In alternative embodiments, the signal from the label at each transient binding event is not projected through the optical magnification path. Instead, the substrate (typically an optically transparent surface on which the target molecule resides) is directly coupled to a two-dimensional detection array. When the pixels of the detection array are small (e.g., less than 1 square micron), one-to-one signal projection on the surface allows the binding signal to be localized with an accuracy of at least 1 micron. In some embodiments where the nucleic acid is sufficiently elongated (e.g., 2 kilobases of a polynucleotide are elongated to a length of 1 micron), signals as far as 2 kilobases apart are degraded. For example, in the case of a hexamer probe where the signal is expected to occur every 4096 bases or every 2 microns, this degradation would be sufficient to clearly localize individual binding sites. Signals that partially enter between two pixels provide intermediate locations (e.g., if the signal enters between two pixels, the degradation could be 500 nm at 1 square micron). In some embodiments, the substrate is physically translated (e.g., in increments of 100 nm) with respect to the two-dimensional array detector to provide higher resolution. In such embodiments, the device is smaller (or thinner) because it does not require lenses or interlenticular space. In some embodiments, the substrate translation also provides a direct translation of molecular memory readouts to electronic readouts that are more compatible with existing computers and databases.

[0264] In some embodiments, the capture frame rate and data transfer rate are increased compared to standard microscopy techniques to capture rapid transient couplings. In some embodiments, the process speed is increased by coupling high frame detection with a high-density probe. However, individual exposures remain at the minimum threshold exposure to reduce the electronic noise associated with each exposure. The accumulated electronic noise from a 200-millisecond exposure would be less than that from two 10-millisecond exposures.

[0265] Faster CMOS cameras that enable even faster imaging are now available. For example, the Andor Zyla Plus can achieve up to 398 frames per second at 512 x 1024 square pixels using only a USB 3.0 connection, which is even faster than a limited region of interest (ROI) or CameraLink connection.

[0266] An alternative approach to obtaining rapid imaging is to use a Garbo mirror or digital micromirror to transmit temporally incremented images to different sensors. The correct order of the movie frames is then reconstructed by interleaving the frames from the different sensors according to the acquisition time.

[0267] The transient binding process can be accelerated by adjusting various biochemical parameters, such as salt concentration. Several high-frame-rate cameras exist that can be used to adapt the binding rate, often with a limited field of view to obtain a faster readout from a subset of pixels. One alternative approach is to use a galvanometer mirror to temporally distribute the continuous signal to different regions or different centers of a single sensor. The latter allows the use of the sensor's full field of view but increases the overall temporal resolution when the distributed signals are compiled.

[0268] Building a dataset of multiple combined events Block 218: Exposure and measurement are repeated for each oligonucleotide probe in the set of oligonucleotide probes, thereby obtaining multiple sets of positions on the test substrate, each set of positions on the test substrate corresponding to one oligonucleotide probe in the set of oligonucleotide probes.

[0269] In some embodiments, the set of oligonucleotide probes comprises multiple subsets of oligonucleotide probes, and repeated exposure and measurement are performed for each subset of oligonucleotide probes in the multiple subsets of oligonucleotide probes.

[0270] In some embodiments, each subset of oligonucleotide probes contains two or more different probes from a set of oligonucleotide probes. In some embodiments, each subset of oligonucleotide probes contains four or more different probes from a set of oligonucleotide probes. In some embodiments, the set of oligonucleotide probes consists of four subsets of oligonucleotide probes.

[0271] In some embodiments, the method further includes dividing a set of oligonucleotide probes into multiple subsets of oligonucleotide probes based on the calculated or experimentally obtained melting temperature of each oligonucleotide probe. Oligonucleotide probes having similar melting temperatures are placed in the same subset of oligonucleotide probes by the division. Furthermore, the temperature or duration of an instance of exposure is determined by the average melting temperature of the oligonucleotide probes in the corresponding subset of oligonucleotide probes.

[0272] In some embodiments, the method further comprises dividing a set of oligonucleotide probes into multiple subsets of oligonucleotide probes based on the sequence of each oligonucleotide probe, wherein oligonucleotide probes having overlapping sequences are placed in different subsets.

[0273] In some embodiments, repeated exposure and measurement are performed for each single oligonucleotide probe in the oligonucleotide probe set.

[0274] In some embodiments, exposure is performed on a first oligonucleotide probe in a set of oligonucleotide probes at a first temperature, and the repetition of exposure and measurement includes performing exposure and measurement on the first oligonucleotide at a second temperature.

[0275] In some embodiments, exposure is performed on a first oligonucleotide probe in a set of oligonucleotide probes at a first temperature. Instances of repeated exposure and measurement include performing exposure and measurement on the first oligonucleotide at each of several different temperatures. The method further includes constructing a melt curve for the first oligonucleotide probe using the measured locations and periods of optical activity recorded by the first temperature and measurements for each of the several different temperatures.

[0276] In some embodiments, the test substrate is washed before repeated exposure and measurement, thereby removing one or more oligonucleotide probes from the test substrate before exposing the test substrate to another set of oligonucleotide probes. Optionally, the probes are first replaced with one or more washing solutions, and then the next set of probes is added.

[0277] In some embodiments, measuring locations on a test substrate involves identifying and fitting each instance of optical activity using a fitting function to identify and fit the center of each instance of optical activity in a frame of data obtained by a two-dimensional imager. The center of each instance of optical activity is considered to be the location of each instance of optical activity on the test substrate.

[0278] In some embodiments, the fitting function is a Gaussian function, a first moment function, a gradient-based approach, or a Fourier transform. While a Gaussian fit is merely an approximation of the PSF of a microscope, in some embodiments, the addition of a spline (e.g., a cubic spline) or a Fourier transform approach improves the accuracy in determining the centroid of the PSF (see, for example, Babcock et al., Sci Rep. 7:552, 2017 and Zhang et al., 46:1819-1829, 2007).

[0279] After data processing, single-molecule locations identify which probes from sets 1-5 have the same positional footprint on the polynucleotide (e.g., which bind to the same nanometer location) (e.g., by detected color). In one example, a nanometer-sized location is defined with 1 nm center accuracy (±0.5 nm), and all probes whose PSF centroids fall within the same 1 nm are thus binned together. Each single defined oligo species must be bound multiple times (e.g., depending on the number of photons emitted and collected) to enable precise positioning relative to the nanometer (or sub-nanometer) centroid.

[0280] In some embodiments, nanometer or sub-nanometer size localization determines, for example, for the oligo sequence 5'-CGACT-3', that the first base is A, the second is G, the third is T, the fourth is C, and the fifth is T. Such a pattern suggests the target sequence of 5'-CGACT-3', and therefore 1024 pentameric oligo probes defined for all single bases are applied or tested in just 5 periods, where each period includes both oligo addition and washing steps. In such execution, the concentration of each specific oligo in the set is lower than when used alone. In this example, data acquisition is taken longer to reach a threshold number of binding events. Also, degenerate oligos at higher concentrations than specific oligos are used in some embodiments. In some embodiments, this coding scheme is performed by direct labeling of the probe, or by synthesizing or conjugating the label at, for example, the 3' or 5' of the oligo. However, in some alternative embodiments, this is accomplished by indirect labeling (for example, by attaching flap sequences to each labeled oligo).

[0281] In some embodiments, the location of each oligo is precisely defined by determining the PSF of multiple events for that location, and then confirmed by partial sequence overlap from offset events (and, if available, data from the complementary strand of the duplex). This embodiment relies heavily on single-molecule localization of probe bindings down to 1 or several nanometers.

[0282] In some embodiments, each instance of optical activity persists across multiple frames measured by a two-dimensional imager. Measuring the location on the test substrate involves identifying and fitting each instance of optical activity across multiple frames using a fitting function to identify the center of each instance of optical activity across multiple frames. The center of each instance of optical activity is considered to be the location of each instance of optical activity on the test substrate across multiple frames. In some embodiments, the fitting function finds the center of each frame individually among the multiple frames. In other embodiments, the fitting function finds the center in each frame, or collectively, across multiple frames.

[0283] In some embodiments, the fit includes tracking steps where a location is immediately adjacent (e.g., within 0.5 pixels) in the next frame, averaging them together, weighting them by their brightness, and assuming that these are the same combined event. However, if there are events separated by multiple frames (e.g., with at least 5-frame gaps, at least 10-frame gaps, at least 25-frame gaps, at least 50-frame gaps, or at least 100-frame gaps between combined events), the fit function assumes they are different combined events. Tracking different combined events helps improve reliability in array assignment.

[0284] In some embodiments, the measurement resolves the center of each instance of optical activity to a position on the test substrate with a positioning accuracy of at least 20 nm. In some embodiments, the measurement resolves the center of each instance of optical activity to a position on the test substrate with a positioning accuracy of at least 2 nm, at least 60 nm, and at least 6 nm. In some embodiments, the measurement resolves the center of each instance of optical activity to a position on the test substrate with a positioning accuracy between 2 nm and 100 nm. In some embodiments, the measurement resolves the center of each instance of optical activity to a position on the test substrate, where the position is a sub-diffraction-limited position. In some embodiments, the resolution is more limited than the accuracy.

[0285] In some embodiments, measurements of location and duration for each instance of optical activity on the test substrate are performed at more than 5,000 photons at that location. In some embodiments, measurements of location and duration for each instance of optical activity on the test substrate are performed at more than 50,000 photons at that location or more than 200,000 photons at that location.

[0286] Each dye has a maximum photon generation rate (e.g., 1 kHz to 1 MHz). For some dyes, it is only possible to measure 200,000 photons per second. The typical lifetime of a dye is 10 nanoseconds. In some embodiments, measurements of location and duration of each instance of optical activity on a test substrate measure more than 1,000,000 photons at that location.

[0287] In some cases, certain outlier sequences bind in a non-Watson-Crick manner, or short motifs lead to excessively high on-rates or low off-rates. For example, some purine-polypyrimidine interactions between RNA and DNA are very strong (e.g., RNA motifs such as AGG). These not only have lower off-rates but also higher on-rates due to more stable nucleating sequences. In some cases, binding occurs from outliers that do not necessarily follow certain known rules. In some embodiments, algorithms are used to identify or consider predictions of such outliers.

[0288] In some embodiments, each instance of optical activity is greater than a predetermined number of standard deviations above the background observed on the test substrate (e.g., a standard deviation greater than 3, 4, 5, 6, 7, 8, 9, or 10).

[0289] In some embodiments, exposure is carried out for a first oligonucleotide probe in a set of oligonucleotide probes during a first period. In some such embodiments, the repetition of exposure and measurement includes carrying out exposure for a second oligonucleotide during a second period. The first period is longer than the second period.

[0290] In some embodiments, exposure is performed on a first oligonucleotide probe in a set of oligonucleotide probes for a first number of frames of a two-dimensional imager. In some such embodiments, the repetition of exposure and measurement includes performing exposure on a second oligonucleotide for a second number of frames of the two-dimensional imager. The first number of frames is greater than the second number of frames.

[0291] In some embodiments, complementary probes in one or more tiling sets are used to bind to each of the denatured duplex strands. As shown in Figure 11B, it is possible to determine the sequence of at least a portion of the nucleic acid from multiple sets of positions on a test substrate, including determining a first tiling path 1114 corresponding to the fixed first strand 1110 and a second tiling path 1116 corresponding to the fixed second strand 1112.

[0292] In some embodiments, a discontinuation in the first tiling path is resolved using a corresponding portion of the second tiling path. In some embodiments, a discontinuation in the first or second tiling path is resolved using a reference sequence. In some embodiments, a discontinuation in the first or second tiling path is resolved using a corresponding portion of a third or fourth tiling path obtained from another instance of the nucleic acid.

[0293] In some embodiments, the confidence in sequence assignment of each binding site sequence is increased using corresponding portions of the first and second tiling passes. In some embodiments, the confidence in sequence assignment of that sequence is increased using corresponding portions of a third or fourth tiling pass obtained from another instance of the nucleic acid.

[0294] Alignment or assembly of arrays Block 222 The sequence of at least a portion of the nucleic acid is determined from multiple sets of positions on the test substrate by compiling the positions on the test substrate, which are represented by multiple sets of positions.

[0295] Preferably, adjacent sequences are obtained via de novo assembly. However, in some embodiments, reference sequences are also used to facilitate assembly. This allows the de novo assembly to be constructed. When complete genome sequencing requires the synthesis of information from multiple molecules (ideally molecules obtained from the same parental chromosome) that connect the same segment of the genome, an algorithm is needed to process the information from multiple molecules. One type of algorithm aligns molecules based on sequences common to multiple molecules, filling in gaps in each molecule by attribution from the co-aligned molecules whose regions are covered (for example, a gap in one molecule is covered by reads in another co-aligned molecule).

[0296] In some embodiments, a shotgun assembly method (e.g., described in Schuler et al., Science 274:540-546, 1996) is adapted to perform assembly using the sequence assignment obtained as described herein. The advantage of the present invention's method over shotgun sequencing is that multiple reads are pre-assembled because they are collected from a full-length intact target molecule (e.g., the locations of the reads relative to each other are known and the lengths of the gaps between reads are known). In various embodiments, a reference genome is used to facilitate the assembly of either or both long-range genomic structures or short-range polynucleotide sequences. In some embodiments, reads are partially de novo assembled, then aligned to a reference, and then the reference-assisted assembly is further de novo assembled. In some embodiments, various reference assemblies are used to provide some guidance for genome assembly. However, in typical embodiments, information obtained from actual molecules (especially when it is confirmed by two or more molecules) is given greater weight than any information from the reference sequence.

[0297] In some embodiments, targets from which sequence bits are obtained are aligned based on segmentation of sequence overlaps between targets, creating longer computer contigs and ultimately the sequence of the entire chromosome.

[0298] In some embodiments, the identity of a polynucleotide is determined by the pattern of probe binding along its length. In some embodiments, the identity is that of an RNA species or RNA isoform. In some embodiments, the identity is that of the polynucleotide at its corresponding reference.

[0299] In some embodiments, the accuracy or precision of localization is insufficient to stitch the sequence bits together. In some embodiments, a subset of probes is found to join within a specific spatial relationship, but in some embodiments, it is difficult to reliably determine their exact order from the localization data. In some embodiments, the resolution is diffraction-limited. In some embodiments, a short range of sequences or diffraction-limited spots within a spatial relationship are assembled by the sequence overlap of probes within that spatial relationship or spot. Thus, the short range sequence is assembled by utilizing information about how individual sequences of a subset of oligos overlap, for example. In some embodiments, the short range sequence constructed in this manner is then stitched together with a longer range sequence based on the order on the polynucleotide. Thus, the long range sequence is obtained by joining short range sequences obtained from adjacent or overlapping spots.

[0300] In some embodiments (for example, in the case of a target polynucleotide that is naturally double-stranded), the reference sequence and sequence information obtained for the complementary strand are used to facilitate sequence assignment.

[0301] In some embodiments, the nucleic acid is at least 140 base pairs long, and the determination determines the sequence coverage of more than 70% of nucleic acid sequences. In some embodiments, the nucleic acid is at least 140 base pairs long, and the determination determines the sequence coverage of more than 90% of nucleic acid sequences. In some embodiments, the nucleic acid is at least 140 base pairs long, and the determination determines the sequence coverage of more than 99% of nucleic acid sequences. In some embodiments, the determination determines the sequence coverage of more than 99% of nucleic acid sequences.

[0302] Nonspecific or mismatched binding events Generally, sequencing assumes that the target porin nucleotide contains a nucleotide complementary to the nucleotide it binds to. However, this is not always the case. Binding mismatch errors are an example of when this assumption does not apply. Nevertheless, mismatches are useful in determining the target sequence if they occur according to known rules or behaviors. The use of short oligonucleotides (e.g., pentamers) means that a single mismatch has a significant impact on stability, since one base accounts for 20% of the pentamer length. For this reason, under appropriate conditions, impeccable specificity can be obtained with short oligoprobes. Still, mismatches can occur, and due to the probabilistic nature of molecular interactions, their binding durations may, in some cases, be indistinguishable from bindings where all five bases are specific. However, algorithms used to perform base (or sequence) calling and assembly often account for the occurrence of mismatches. Many types of mismatches are predictable and follow specific rules. Some of these rules are derived from theoretical evidence, while others are derived experimentally (see, for example, Maskos and Southern, Nucleic Acids Res 21(20): 4663-4669, 2013; Williams et al., Nucleic Acids Res 22:1365-1367, 1994).

[0303] The effect of nonspecific binding to the surface is mitigated by the non-persistence of the probe, and binding to nonspecific sites is not persistent. When an imager occupies a nonspecific (e.g., not on a complementary target sequence) binding site, it may decolorize, but in some cases, it remains in place and blocks further binding to that site (e.g., interaction by G quadruple formation). Typically, most nonspecific binding sites that interfere with the resolution of the imager binding to the target polynucleotide are occupied and decolorized early in imaging, allowing the on / off binding of the imager to the polynucleotide site to be easily observed thereafter. Thus, in one embodiment, high laser power is used to decolorize the probe that initially occupies a nonspecific binding site, and in some cases no image is taken at this time, and then the laser power is optionally reduced to initiate imaging and capture the on-off binding to the polynucleotide. After initial nonspecific binding, further nonspecific bindings are less frequent (as faded probes often remain adhered to nonspecific binding sites) and, in some embodiments, are computationally filtered out by applying a threshold, which is considered specific binding to, for example, docking sites. Binding to the same site must be persistent, occurring at least five, preferably ten, times at the same site. Typically, approximately 20 specific binding events to docking sites are detected.

[0304] Another way to filter out nonspecific bindings is that the fluorophore signal must correlate with the position of the linear chain of the target molecule being extended on the surface. In some embodiments, the position of the linear chain can be determined by directly staining the linear chain or by inserting the line into a persistent binding site. In general, signals that do not fall along the line are discarded in some embodiments, whether they are persistent or not. Similarly, when supramolecular lattices are used, binding events that do not correlate with known lattice structures are discarded in some embodiments.

[0305] Multiple binding events also increase specificity. For example, consensus is obtained from multiple calls rather than establishing identity of a region or sequence detected by a single "call." Furthermore, multiple binding events to a target region or sequence allow for the distinction between binding to the actual location and non-specific binding events, where binding (within the threshold period) is less likely to occur multiple times at the same location. It has also been observed that measuring multiple binding events over time allows for the accumulation of non-specific binding events on a faded surface, after which non-specific binding is rarely detected again. This is thought to be because even if the signal from non-specific binding fades, the non-specific binding site remains occupied or blocked.

[0306] In some embodiments, sequencing is complicated by mismatches and nonspecific binding on polynucleotides. To avoid the effects of nonspecific binding or outlier events, in some embodiments, the method prioritizes signals based on location and persistence. Location-based priority is predicted depending on whether the probe coexists on an extended polymer or on a supramolecular lattice (e.g., a DNA origami grid), including locations within a lattice structure. Binding persistence-based priority relates to the duration and frequency of binding, and a priority list is used to determine the likelihood of a perfect match, partial match, or nonspecific binding. The suitability of the signal is determined using the priority established for each binding probe in the panel or repertoire.

[0307] In some embodiments, priority is used to facilitate signal validation and base calling by determining whether the signal duration is greater than a predetermined threshold, whether the signal repetition or frequency is greater than a predetermined threshold, whether the signal correlates with the location of the target molecule, and / or whether the number of collected photons is greater than a predetermined threshold. In some embodiments, if the answer to any of these decisions is true, the signal is accepted as real (e.g., not a mismatch or nonspecific binding event).

[0308] In some embodiments, mismatches are distinguished by their temporal binding pattern and are therefore considered a secondary layer of sequence information. In such embodiments, if a binding signal is determined to be a mismatch by its temporal binding characteristics, the sequence bit is bioinformatically cleaved to remove the presumed mismatched base, and the remaining sequence bit is added to the sequence reconstruction. Since mismatches are most likely to occur at the ends of hybridizing oligos, the temporal binding characteristics cause one or more bases to be cleaved from the ends in some embodiments. The decision of which bases to cleave is, in some embodiments, informed by information from other oligo tilings across the same sequence space.

[0309] In some embodiments, signals that do not appear to be reversible are unfavorably weighted because they have the opportunity or likelihood of corresponding to nonspecific signals (e.g., due to the adhesion of fluorescent contaminants to the surface).

[0310] Blocks 302-304 provide another method for sequencing nucleic acids, which includes immobilizing nucleic acids on a test substrate in a linearized extension form, thereby forming immobilized extension nucleic acids. The nucleic acids are immobilized on the substrate by any one of the above methods.

[0311] Isolating single cells on the surface and extracting both DNA and RNA. RNA and / or DNA can be isolated and sequenced from a single cell. In some embodiments, if the goal is to sequence DNA, sequencing is started after adding RNAse to the sample. In some embodiments, if the objective is to sequence RNA, sequencing is started after adding DNAse to the sample. In some embodiments, if both cytoplasmic and nuclear nucleic acids are to be analyzed, they are extracted differentially or sequentially. In some embodiments, the cell membrane (not the cornea) is first disrupted to release and collect cytoplasmic nucleic acids. Then, the cornea is disrupted to release nuclear nucleic acids. In some embodiments, proteins and polypeptides are collected as part of the cytoplasmic fraction. In some embodiments, RNA is collected as part of the cytoplasmic fraction. In some embodiments, DNA is collected as part of the nuclear fraction. In some embodiments, the cytoplasmic and nuclear fractions are extracted together. In some embodiments, mRNA and genomic DNA are differentially captured after extraction. For example, mRNA is captured by an oligo-dT probe attached to its surface. This can occur in the first part of the flow cell, where the DNA is captured in the second part of the flow cell, which has a hydrophobic vinylsilane coating on which the ends of the DNA can be trapped (e.g., presumably by hydrophobic interactions).

[0312] Positively charged surfaces such as poly(L)lysine (PLL) (e.g., obtained from Microsurfaces Inc. or coated in-house) are known to be able to bind to cell membranes. In some embodiments, low-height flow channels (e.g., less than 30 microns) are used to increase the chances of cells colliding with the surface. The number of collisions is increased in some embodiments by introducing turbulence using a herringbone pattern in the flow cell sealing. In some embodiments, cell adhesion does not need to be efficient, as it is desirable in such embodiments that cells are dispersed at a low density on the surface (e.g., to ensure sufficient space between cells so that RNA and DNA extracted from each individual cell remain spatially separated). In some embodiments, cells are ruptured using protease treatment so that both the cell and nuclear membrane are destroyed (e.g., so that the cell contents are released into the medium and captured on the surface near the isolated cells). Once immobilized, the DNA and RNA are extended in some embodiments. In some embodiments, the extension buffer flows in one direction across the coverslip surface (e.g., stretching and aligning DNA and RNA polynucleotides in the direction of the fluid flow). In some embodiments, the conditions are adjusted (e.g., temperature, extension buffer composition, and flow physical force) to denature most of the RNA's secondary / tertiary structure so that the RNA is available for binding to an antibody. Once the RNA has been extended in a denatured form, it is possible to switch from the denaturation buffer to the binding buffer.

[0313] Alternatively, RNA is extracted and immobilized by first disrupting the cell membrane and inducing a unidirectional flow. The nuclear membrane is then disrupted using a protease, inducing a flow in the opposite direction. In some embodiments, DNA is fragmented before or after release using, for example, rare-cutting restriction enzymes (e.g., NOT1, PMME1). This fragmentation assists in the release of DNA and allows for the isolation and skimming of individual strands. It is ensured that the immobilized cells are sufficiently separated and that the system is set up so that the RNA and DNA extracted from each cell do not mix. In some embodiments, this is assisted by inducing a transition from liquid to gel before, after, or between cell rupture.

[0314] In some embodiments, the nucleic acid is a double-stranded nucleic acid. In such embodiments, the method further includes denaturing a double-stranded nucleic acid immobilized on a test substrate into a single-stranded form. The nucleic acid must be in a single-stranded form for sequencing to proceed. Once the immobilized double-stranded nucleic acid has been denatured, both an immobilized first strand and an immobilized second strand of the nucleic acid are obtained. The immobilized second strand is complementary to the immobilized first strand.

[0315] In some embodiments, nucleic acids are single-stranded (e.g., mRNA, lncRNA, microRNA). In some embodiments where the nucleic acid is single-stranded RNA, denaturation is not required before the sequencing method can proceed.

[0316] In some embodiments, the sample contains a single-stranded polynucleotide that does not have a native complementary strand nearby. The binding sites for each oligo in the repertoire along the polynucleotide are compiled. In some embodiments, the sequence is reconstructed by assembling all the sequencing bits according to their locations and stitching them together.

[0317] RNA elongation The elongation of nucleic acids on a charged surface is affected by the cation concentration of the solution. At low salt concentrations, single-stranded RNA, negatively charged along its backbone, will bind to the surface randomly along its length.

[0318] Several methods exist for denaturing RNA and extending it into a linear form. In some embodiments, the RNA is initially encouraged to enter a spherical form (e.g., by utilizing a high salt concentration). In some such embodiments, the ends of each RNA molecule (e.g., particularly the poly-A tail) become more interacting. Once the RNA is bound in a spherical form, a different buffer (e.g., a denaturing buffer) is applied to the flow cell in some embodiments.

[0319] In an alternative embodiment, the surface is pre-coated with oligo-d(T) to capture the poly-A tail of mRNA (e.g., described in Ozsolak et al., Cell 143:1018-1029, 2010). The poly-A tail is typically a region that should relatively lack secondary structure (e.g., because they are homopolymers). Since the poly-A tail is relatively long in higher eukaryotes (250-3000 nucleotides), in some embodiments, long oligo-d(T) capture probes are designed so that hybridization is carried out under relatively high stringency (e.g., high temperature and / or high salt conditions) sufficient to dissolve a significant fraction of intramolecular base pairings in the RNA. After binding, in some embodiments, the remaining change of the RNA structure from spherical to linear is carried out by utilizing denaturing conditions that disrupt intramolecular base pairings in the RNA but are not sufficient to invalidate the capture, and by fluid flow or electrophoretic force.

[0320] Block 310 In some embodiments, an immobilized extended nucleic acid is exposed to each pool of each oligonucleotide probe in a set of oligonucleotide probes. Each oligonucleotide probe in the set of oligonucleotide probes has a predetermined sequence and length, and the exposure occurs under conditions that allow for transient and reversible exposure to each portion of the immobilized nucleic acid in which the individual probes of each pool of each oligonucleotide probe are complementary to each oligonucleotide probe, thereby producing each instance of optical activity.

[0321] Block 312 In some embodiments, a two-dimensional imager is used to measure the location and duration of each instance of optical activity occurring on the test substrate during exposure.

[0322] Block 314 In some embodiments, exposure and measurement are repeated for each oligonucleotide probe in a set of oligonucleotide probes, thereby obtaining multiple sets of positions on a test substrate, each set of positions on the test substrate corresponding to one oligonucleotide probe in the set of oligonucleotide probes.

[0323] Block 316 In some embodiments, the sequence of at least a portion of the nucleic acid is determined from a plurality of sets of positions on a test substrate by compiling positions on the test substrate represented by a plurality of sets of positions.

[0324] RNA sequencing Although RNA is typically shorter than genomic DNA, sequencing RNA from one end to the other using current technology is challenging. Nevertheless, determining the organization of the entire mRNA sequence is crucial for alternative splicing. In some embodiments, mRNA is captured by binding of the poly(A) tail with immobilized oligo(d(T)), and its secondary structure is removed by applied elongation force (e.g., >400 pN) and denaturing conditions (e.g., including formamide or 7M or 8M urea), allowing it to be extended on the surface. This then allows for transient binding of a binding reagent (e.g., exon-specific). Due to the short length of RNA, resolving and identifying exons using the single-molecule localization methods described herein is beneficial. In some embodiments, a few scattered binding events in the RNA are sufficient to determine the order and identity of exons in the mRNA for specific mRNA isoforms.

[0325] Double-stranded consensus The method for obtaining sequence information from a sample molecule is as follows: i) Provide a first color marker to the first oligo. Provide a second color marker to the second oligo, where the second oligo is a complementary sequence to the first oligo. ii) Extend, fix, and denature double-stranded nucleic acid molecules on a substrate. iii) Expose both the first and second oligonucleotides to the denatured nucleic acid of ii. iv) Determine the binding sites of the first and second oligos. v) If multiple bond locations exist, that location is considered correct. vi) Multiple sites along the extended nucleic acid are bound.

[0326] In some embodiments, the oligos bind transiently and reversibly. In some embodiments, the first and second oligos are part of a competing repertoire of first and second oligos of a given length, and steps ii-iii are repeated for each first and second oligo pair in the repertoire to sequence the entire nucleic acid.

[0327] In some embodiments, multiple modifications are necessary to ensure that the two colors coexist optically when they should. This includes correcting for chromatic aberration. In some such embodiments, two oligonucleotides are added together, but the chemical properties of the modified oligonucleotides are used with non-self-pairing analog bases to prevent them from annealing to each other and thus neutralize their effects, where modified G cannot pair with modified C in the complementary oligonucleotide but can pair with unmodified C on the target nucleic acid, modified A cannot pair with modified T in the complementary oligonucleotide but can pair with unmodified T, and so on. Thus, in such embodiments, the first and second oligonucleotides are modified such that the first oligonucleotide cannot form a base pair with the second oligonucleotide.

[0328] In some embodiments, the first and second oligos are not added together, but one is added after the other.

[0329] In such embodiments, one oligonucleotide is added after the other, and a washing step is performed between them. In this case, the two oligonucleotides of the complementary pair are labeled with the same color, and no correction is needed for chromatic aberration. Furthermore, there is no possibility of the two oligonucleotides binding to each other.

[0330] In some embodiments, nucleic acids are further exposed to first and second oligos until the entire repertoire of oligos is depleted.

[0331] In some embodiments, the second oligo is added as the next oligo after the first oligo, before any other oligo pairs in the repertoire are added. In some embodiments, the second oligo is not added as the next oligo before any other oligo pairs in the repertoire are added.

[0332] Examples of such embodiments include the following method for obtaining sequence information from molecules in a sample: i) Extend, fix, and denature double-stranded nucleic acid molecules on a substrate. ii) Expose the first labeled oligo to the denatured nucleic acid of i), and detect and record the site of its binding. iii) Remove the first labeled oligo by washing. iv) Expose the second labeled oligo to the denatured nucleic acid from i) and detect and record the binding site. v) If necessary, correct for drift between recordings in ii) and iv). vi) If the recorded positions of the joins obtained in ii~iv coexist, the sequence information obtained in this way regarding the sequence of locations is considered correct.

[0333] In some embodiments, the first and second oligos are part of a competing repertoire of first and second oligos of a given length, and steps ii-iii are repeated for each first and second oligo pair in the repertoire to sequence the entire nucleic acid.

[0334] Coexistence tells us that we are looking at the same sequence locus. Furthermore, a probe targeting the sense strand can be expected to identify the central base using four differentially labeled oligos, and a probe targeting the antisense strand can be expected to identify the central base using four differentially labeled oligos having sequences complementary to the sense strand probe. In order to obtain a verified base call of the central position, the data for the sense strand must confirm the data for the second strand. Therefore, if the oligo with the central A base binds to the sense strand, the oligo with the central T base must bind to the antisense strand.

[0335] Obtaining such confirmation or consensus on the sense and antisense strands also helps overcome ambiguity for binding via G:T or G:U fluctuation base pairing. If this occurs on the sense strand, it is unlikely to produce a signal on the antisense strand because C:A is less likely to form a base pair.

[0336] In some embodiments, modified G bases or T / U bases can be used in the probe to prevent the formation of fluctuating base pairs. In some other embodiments, the reconstruction algorithm takes into account the possibility of fluctuating base pair formation, particularly when confirmation at the C:G base pair is not present on the complementary strand and the location correlates with an oligo bond to the complementary strand that forms an A:T base pair. In some embodiments, 7-deazaguanidine, which has the ability to form only two hydrogen bonds instead of three, is used as the G modification to reduce the stability of the base pairs it forms, as well as the occurrence rate of G quadruples and their (and their promiscuous bonds).

[0337] Simultaneous duplex consensus assembly In some embodiments, both strands of the double helix are present and exposed to the oligonucleotides as described above while in proximity. In some embodiments, it is not possible to distinguish from the transient optical signal detected which of the two complementary strands each oligo in each set of oligonucleotides is bound to. For example, if the binding sites along each polynucleotide for each oligo in each set of oligonucleotides along a polynucleotide are compiled, two probes with different sequences may appear to be bound to the same site. These oligos must have complementary sequences, and the difficulty lies in determining which strand each of the two bound oligos binds to, which is required beforehand to accurately compile the sequences for the polynucleotide.

[0338] To determine whether a single binding event occurs on one strand or the other, the complete set of optical activities obtained must be considered. For example, if two tiling series of oligos cover the positional relationship in question, one of the two tiling series to which the signal belongs is assigned based on which series the oligo sequence producing the signal overlaps with. In some embodiments, the sequence is then reconstructed by first constructing each of the two tiling series using the binding site and sequence overlap. The two tiling series are then aligned as reverse complementary strands, and base assignments at each site are only permitted if the two strands are complete reverse complementary strands at each of those sites (e.g., thereby providing a duplex consensus sequence).

[0339] In some embodiments, sequencing mismatches are stopped as being ambiguous base calls, where one of two possibilities, such as from independent mismatch binding events, needs to be corroborated by an additional layer of information. In some embodiments, once a duplex consensus is obtained, a conventional (multimolecule) consensus is determined by comparing data from other polynucleotides that cover the same region of the genome (e.g., when information on binding sites from multiple cells is available). One problem with such an approach is the possibility of polynucleotides containing haplotype sequences.

[0340] Alternatively, in some embodiments, the consensus of individual strands is obtained before the duplex consensus of the individual strands is obtained. In such embodiments, the sequences of each of the duplex strands are obtained simultaneously. This is performed in some embodiments without the need for additional sample preparation steps and, unlike current NGS methods, differentially tags the two strands of the duplex with molecular barcodes (e.g., as described in Salk et al., Proc. Natl. Acad. Sci. 109(36), 2012).

[0341] Simultaneous sequence acquisition of both the sense and antisense strands is preferably compared to nanopore 2D or 1D 2 consensus sequencing. These alternative methods require the sequence obtained for one strand of the duplex before the sequence of the second strand is obtained. In some embodiments, duplex consensus sequencing provides an accuracy in the range of, for example, one error per million bases (compared to the less than sufficient accuracy of 10 6 ~10 2 ~10 3 of other NGS approaches). This makes the method highly suitable for the requirements for resolving rare variants that indicate a cancerous state (e.g., those present in cell-free DNA) or are present at low frequency in a tumor cell population.

[0342] Single-cell resolution sequencing In various embodiments, the method further includes sequencing the genome of a single cell. In some embodiments, the single cell does not involve adhesion from other cells. In some embodiments, the single cell is attached to other cells in a cluster or tissue. In some embodiments, such cells are deaggregated into individual non-adherent cells.

[0343] In some embodiments, cells are deaggregated before they move fluidly into the entrance of a structure (e.g., a flow cell or microwell) into which the polynucleotides are extended (e.g., by using a pipette). In some embodiments, deaggregation is performed by pipetting the cells, by adding proteases, sonic treatment, or physical agitation. In some embodiments, cells are deaggregated after they have moved fluidly into the structure into which they are extended.

[0344] In some embodiments, a single cell is isolated, and polynucleotides are released from the single cell, so that all polynucleotides originating from the same cell remain located in close proximity to each other and in a different location from where the contents of other cells are located. In some embodiments, a trap structure described in Di Carlo et al., Lab Chip 6:1445-1449, 2006 is used.

[0345] In some embodiments, either a microfluidic structure can be used to capture and isolate multiple single cells (e.g., when the traps are separated, as shown in Figures 16A and 16B), or a structure can be used to capture multiple non-isolated cells (e.g., when the traps are continuous). In some embodiments, the traps are the size of a single cell (e.g., 2 μM to 10 μM). In some embodiments, the flow cell is several hundred microns to several millimeters in length and about 30 microns deep.

[0346] In some embodiments, as shown in Figure 17 as an example, a single cell flows into a delivery channel 1702, is trapped 1704, releases nucleotides, and is then extended. In some embodiments, a cell 1602 is lysed 1706, and then the cell nucleus is lysed through a second lysis step 1708, thus sequentially releasing extracellular and intracellular polynucleotides 1608. In some cases, both extracellular and intracellular polynucleotides are released using a single lysis step. After release, the polynucleotides 1608 are immobilized and extended along the length of the flow cell 2004. In some embodiments, the trap is the size of a single cell (e.g., width 2 μM to 10 μM). In one embodiment, the dimensions of the trap are 4.3 μM wide at the bottom, 6 μM high at the middle, 8 μM at the top, and 33 μm deep, and the device is fabricated from cyclic olefin (COC) using injection molding.

[0347] In some embodiments, single cells are lysed into individual channels, and each individual cell is reacted with a unique tag sequence by transposase-mediated integration, after which the polynucleotides are aggregated and sequenced in the same mixture. In some embodiments, the transposase complex is transfected into the cells or is contained within droplets that are dissolved in droplets containing cells.

[0348] In some embodiments, the aggregates are small clusters of cells, and in some embodiments, the entire cluster is tagged with the same sequencing tag. In some embodiments, the cells are not aggregated, and there are no suspended cells such as circulating tumor cells (CTCs) or circulating fetal cells.

[0349] Single-cell sequencing presents a problem with cytosine-thymidine single-nucleotide variants induced by spontaneous cytosine deamination after cell lysis. This can be overcome by pre-treating the sample with uracil N-glycosylase (UNG) before sequencing (e.g., Chen et al., Mol Diagn Ther. 18(5): 587-593, 2014).

[0350] Haplotype Identification In various embodiments, the above method is used to sequence haplotypes. Haplotype sequencing involves sequencing a first target polynucleotide spanning a haplotype in a diploid genome using the method described herein. A second target polynucleotide spanning a second haplotype region in the diploid genome must also be sequenced. The first and second target polynucleotides may originate from different copies of homologous chromosomes. The sequences of the first and second target polynucleotides are compared to determine the haplotypes on the first and second target polynucleotides.

[0351] Thus, the single-molecule reads and assemblies obtained from the embodiments are classified as haplotype-specific. The only instance where haplotype-specific information is not readily available over long ranges is when the assemblies are intermittent. In such embodiments, the read locations are still provided. Even in such situations, if multiple polynucleotides covering the same segment of the genome are analyzed, the haplotype can be computationally determined.

[0352] In some embodiments, homologous molecules are separated according to haplotype or parental chromosome specificity. The visual nature of the information actually obtained physically or visually by the methods of this disclosure may indicate a specific haplotype. In some embodiments, haplotype resolution enables the performance of improved genetic or ancestral studies. In other embodiments, haplotype resolution enables the performance of better histological typing. In some embodiments, haplotype resolution or detection of a specific haplotype enables the performance of diagnostics.

[0353] Simultaneous sequencing of polynucleotides from multiple cells In various embodiments, polynucleotides from multiple cells (or nuclei) are arranged using the previously described method, such that each polynucleotide holds information about the cell of origin.

[0354] In certain embodiments, transposon-mediated sequence insertions are interposed within cells, and each insertion includes a unique ID sequence tag as a marker for the cell of origin. In other embodiments, transposon-mediated insertions occur within a container in which single cells are isolated, such containers including agarose beads, oil-water droplets, etc. The unique tag indicates that all polynucleotides carrying the tag must originate from the same cell. All genomic DNA and / or RNA are then extracted, mixed, and extended. Subsequently, when sequencing according to embodiments of the present invention (or any other sequencing method) is performed on the polynucleotides, reading the ID sequence tags indicates which cell the polynucleotides originated from. It is preferable to keep the tags identifying the cells short. For 10,000 cells (e.g., from tumor microbiology), approximately 65,000 unique sequences are provided by 8-nucleotide-length identification sequences, and around 1 million unique sequences are provided by 10-nucleotide-length identification sequences.

[0355] In some embodiments, individual cells are tagged with an identity (ID) tag. As shown in Figure 19, in some embodiments, the identity tag is integrated into a polynucleotide by tagmentation, while the reagent is provided directly into a single cell, dissolved in a cell 1802, or in a microdroplet that engulfs the cell 1802. Each cell receives a different ID tag (from a large repertoire, e.g., over one million possible tags). After the microdroplet and cell 1804 are fused, the ID tag is integrated into the polynucleotide within the individual cell. The contents of the individual cells are mixed in a flow cell 2004. Sequencing (e.g., by the method disclosed herein) then reveals which cell the specific polynucleotide originated from. In an alternative embodiment, the microdroplet engulfs the cell and delivers the tagging reagent to the cell (e.g., by diffusion into the cell or by rupturing the cell contents into the microdroplet).

[0356] This same indexing principle can be applied to non-cellular samples (e.g., from different organisms) when the objective is to mix samples and sequence them together, but to recover sequence information belonging to each individual sample.

[0357] Furthermore, when multiple cells are sequenced, the diversity and frequency of haplotypes in a cell population can be determined. In some embodiments, genomic heterogeneity within a population can be analyzed without the need to aggregate the contents of single cells, because if the molecules are long enough, different chromosomes, long chromosomal segments, or haplotypes present in the cell population can be determined. This does not indicate that two haplotypes coexist in a cell, but it reports the diversity (or haplotypes) of genomic structural types and their frequencies, as well as which abnormal structural variants are present.

[0358] In some embodiments, when the polynucleotide is RNA and the cDNA copy is sequenced, tagging involves cDNA synthesis with primers containing the tag sequence. When the RNA is sequenced directly, the tag is added by ligation of the tag to the 3' RNA terminal using T4 RNA ligase. An alternative tagging method is to extend the RNA or DNA with a terminal transferase along with more than one of the four A, C, G, and T bases so that each individual polynucleotide probabilistically obtains a unique sequence of nucleotide tails.

[0359] In some embodiments, tag sequences are distributed across multiple sites to maintain shorter sequence lengths, allowing more sequence reads to be used for sequencing the polynucleotide sequence itself. Herein, for example, three short identity sequences are introduced into each cell or container. The origin of the polynucleotide is then determined from bits of tags distributed along the polynucleotide. In this case, a bit of tag read from one location is insufficient to determine the originating cell, but multiple tag bits are sufficient to make the determination.

[0360] Detection of structural variants In some embodiments, the differences between the detected sequence and the reference genome include substitutions, insertions, deletions, and structural variations. More specifically, if the reference sequence is not assembled by the method of this disclosure, repeats will be compressed and rearrangements will be expanded.

[0361] In some embodiments, the orientation of a series of sequence reads along a polynucleotide is reported to indicate whether or not a reversal event has occurred. One or more reads that are oriented in the opposite direction to other reads compared to a reference indicate reversal.

[0362] In some embodiments, the presence of one or more reads that are unexpected in the context of other nearby reads indicates a transposition or translocation compared to the reference. The location of the read in the reference indicates which part of the genome has shifted to another. In some examples, the read at the new location is a duplication rather than a translocation.

[0363] In some embodiments, repeat regions or copy number variations can also be detected. Repeat occurrences of paralogous mutations bearing reads or related reads are observed as multiple or very similar reads occurring at multiple locations in the genome. These multiple locations are, in some examples, closely packed together (e.g., as in satellite DNA), or they are, in other cases, dispersed across the genome (e.g., as in pseudogenes). The methods of this disclosure are applicable to short tandem repeats (STRS), variable number tandem repeats (VNTR), trinucleotide repeats, and the like. The absence or repetition of specific reads indicates that a deletion or amplification has occurred, respectively. In some embodiments, the methods are applicable, particularly in the presence of multiple and / or complex transpositions within a polynucleotide. Because this method is based on the analysis of a single polynucleotide, in some embodiments, the structural variants described above can be resolved to the point of rare occurrence in a small number of cells, e.g., as few as 1% of cells from a population.

[0364] Similarly, in some embodiments, segmental duplications or duplicons are correctly located within the genome. Segmental duplicons are typically long regions in DNA sequences of nearly identical sequences (e.g., longer than 1 kilobase). These segmental duplications induce numerous structural mutations in individual genomes, including somatic mutations. Segmental duplicons may be located in the distal regions of the genome. With current next-generation sequencing, it is difficult to determine which segmental duplicon is generating the reads (i.e., complicating assembly). In some embodiments of this disclosure, sequence reads are obtained as long molecules (e.g., in the 0.1–10 megabase range), and it is possible to determine the genomic context of a duplicon by using the reads to determine which segment of the genome is adjacent to the specific segment of the genome corresponding to the duplicon.

[0365] Breakpoints in structural variants are precisely localized in some embodiments of this disclosure. In some embodiments, it is possible to detect the fusion of two parts of the genome, and the precise individual reads where the breakpoint occurs are determined. The sequence reads collected as described herein contain chimeras of the two fused regions, where all sequences on one side of the breakpoint correspond to one of the fused segments, and the sequences on the other side correspond to the other of the fused segments. This provides high reliability in determining the breakpoint, even in cases where the structure is complex around the breakpoint. In some embodiments, precise chromosomal breakpoint information is used to understand disease mechanisms, to detect the occurrence of specific translocations, or to diagnose diseases.

[0366] Location of epigenomic modifications In some embodiments, the method further includes exposing an immobilized double-stranded nucleic acid or an immobilized first strand and an immobilized second strand to an antibody, affimer, nanobody, aptamer, or methyl-binding protein to determine modifications to the nucleic acid or to correlate a portion of the nucleic acid sequence from a set of positions on a test substrate. Some antibodies bind to double-stranded or single-stranded nucleic acids. Methyl-binding proteins are expected to bind to double-stranded polynucleotides.

[0367] In some embodiments, native polynucleotides do not require processing before they are presented for sequencing. This allows the method to integrate epigenomic information with sequence information because the chemical modifications of the DNA remain in place. Preferably, the polynucleotides are well aligned in a directional manner, thereby making the imaging process, base calling, and assembly imaging relatively easy, resulting in a low sequence error rate and high coverage. Although several embodiments for carrying out this disclosure are described, each is implemented such that the burden of sample preparation is completely or almost completely eliminated.

[0368] These methods, in some embodiments, are performed on genomic DNA without amplification, thus avoiding amplification bias and errors, and preserving and detecting epigenomic marks (e.g., orthogonally to sequence acquisition). In some cases, when nucleic acids are methylated, it is useful to determine this using sequence-specific methods. For example, one method for identifying a fetus from maternal DNA is to methylate the former at the target locus. This is useful for non-invasive prenatal testing (NIPT).

[0369] Multiple types of methylation are possible, including alkylation of carbon-5 (C5), which in mammals yield several cytosine variants: C5-methylcytosine (5-mC), C5-hydroxymethylcytosine (5-hmC), C5-formylcytosine, and C5-carboxylcytosine. Eukaryotes and prokaryotes also methylate adenine to N6-methyladenine (6-mA). N4-methylcytosine is also abundant in prokaryotes.

[0370] Antibodies are available or can be constructed for each of these modifications, and any other modifications deemed interesting. Affimers, nanobodies, or aptamers targeting modifications are particularly relevant due to their potential for smaller footprints. Any reference to antibodies in this invention should be interpreted as including affimers, nanobodies, aptamers, and any similar reagents. In addition, other naturally occurring DNA-binding proteins, such as methyl proteins (MBD1, MBD2, etc.), are used in some embodiments.

[0371] Methylation analysis is performed orthogonally to sequencing in some embodiments. In some embodiments, this is performed before sequencing. As an example, an anti-methyl C antibody or methyl-binding protein (methyl-binding domain (MBD) protein family includes MeCP2, MBD1, MBD2, and MBD4) or peptide (based on MBD1) is bound to polynucleotides in some embodiments, and their sites are detected via labeling before they are removed (e.g., by adding a high-salt buffer, chaotropic reagent, SDS, protease, urea, and / or heparin). Preferably, the reagent binds transiently by using a transient binding buffer to promote on-off binding, or the reagent is designed to bind transiently. Similar approaches are used for other polynucleotide modifications such as hydroxymethylation or DNA damage sites, in which case antibodies are available or can be produced. Sequencing is initiated after the modification sites are detected and the modification-binding reagents are removed. In some embodiments, the target polynucleotide is denatured into a single strand, and then anti-methyl and anti-hydroxymethyl antibodies are added. The method is highly sensitive and capable of detecting a single modification on a long polynucleotide.

[0372] Figure 19 shows the extraction and extension of DNA and RNA from single cells, as well as differential labeling of DNA and RNA (e.g., by antibodies against mC and m6A, respectively). Cells 1602 are immobilized on the surface and then lysed 1902. The nucleic acids 1608 released from the nucleus 1604 by lysis are immobilized and extended 1904. The nucleic acids are then exposed to and conjugated with antibodies containing attached DNA tags 1910 and 1912. In some embodiments, the tag is a fluorescent dye or oligonucleotide docking sequence for single-molecule localization based on DNA PAINT. In some embodiments, instead of using a tag and DNA PAINT, the antibody or other conjugating protein is directly fluorescently labeled with a single or multiple fluorescent labels. In examples where the antibody is encoded, examples of labeling are shown in Figures 14A, 14C, and 14D. Epimodification analysis of both DNA and RNA is linked to their sequences using the sequencing methods described herein in some embodiments.

[0373] In some embodiments, in addition to detecting methylation by binding to a protein, the presence of methylation at the binding site is detected by the differential oligonucleotide binding behavior when modification is present at the target nucleic acid site, compared to when no modification is present.

[0374] In some embodiments, methylation is detected using bisulfite treatment. Herein, after passing through a repertoire, bisulfite treatment is used to convert unmethylated cytosine to uracil, and then the repertoire is applied again. If a nucleotide position read as C before bisulfite treatment is read as U after bisulfite treatment, it can be considered unmethylated.

[0375] There is no reference epigenome for DNA modifications such as methylation. For it to be useful, it is necessary to link known polynucleotide methylation maps to sequence-based maps. Therefore, in some embodiments, epimapping methods correlate sequence bits obtained by oligobinding to provide relevance to epimaps. In addition to sequence reads, other types of methylation information are also linked in some embodiments. This includes, as non-limiting examples, maps based on nickel endonucleases, maps based on oligobinding, and denaturation and denaturation-regeneration maps. In some embodiments, polynucleotides are mapped using transient binding of one or more oligos. In addition to functional modifications to the genome, the same approach is applied in some embodiments to other features to map to the genome, such as DNA damage sites and protein or ligand binding sites.

[0376] In preferred disclosures, either base sequencing or epigenomic sequencing is performed first. In some embodiments, both are performed simultaneously. For example, an antibody against a specific epimodification is differentially encoded from an oligo in some embodiments. In such embodiments, conditions that facilitate the transient binding of both types of probes are utilized (e.g., low salt concentration).

[0377] In some embodiments, when polynucleotides include chromosomes or chromatin, antibodies are used on the chromosomes or chromatin to detect modifications to DNA and also to histones (e.g., histone acetylation and methylation). The locations of these modifications are determined by transient binding of the antibody to the locations on the chromosome or chromatin. In some embodiments, the antibodies are labeled with oligotags and do not bind transiently, but rather are permanently or semi-permanently immobilized at their binding sites. In such embodiments, the antibodies include oligotags, and the locations of these antibody binding sites are detected by utilizing transient binding of complementary oligos to oligos on the antibody tags.

[0378] Isolation and analysis of cell-free nucleic acids Some of the DNA or RNA most readily available for diagnosis are found outside cells in bodily fluids or feces. Such nucleic acids are often shed from cells within the body. Cell-free DNA circulating in the blood is used in prenatal testing for trisomy 21 and other chromosomal and genomic disorders. It is also a means of detecting tumor-derived DNA and other DNA or RNA that are markers for certain pathological conditions. However, the molecules are typically present in small segments (e.g., in the range of about 200 base pairs in blood, and shorter in urine). The copy number of a genomic region is determined by comparing it to the number of reads aligned to a specific region of reference, compared to other parts of the genome.

[0379] In some embodiments, the methods of the present disclosure are applied to the enumeration or analysis of cell-free DNA sequences by two approaches. The first involves immobilizing short nucleic acids before or after denaturation. A transient binding reagent is used to interrogate the nucleic acids to determine the identity of the nucleic acid, its copy number, whether mutations or specific SNP alleles are present, and whether the detected sequences are methylated or carry other modifications (biomarkers).

[0380] A second approach involves ligating small nucleic acid fragments (for example, cell-free nucleic acids after isolation from a biological sample). Ligation allows for the extension of the combined nucleic acids. Ligation is performed by polishing the ends of the DNA and performing blunt end ligation. Alternatively, blood or cell-free DNA may be split into two alicots, one alicot tailed with poly(A) (using terminal transferase) and the other alicot tailed with poly(T).

[0381] The resulting concatenation is then subjected to sequencing. The resulting "super" sequence reads are then compared against a reference to extract individual reads. Each individual read is extracted computationally and then processed in the same manner as other short reads.

[0382] In some embodiments, the biological sample includes feces, which is a medium containing numerous exonucleases that degrade nucleic acids. In such embodiments, a high concentration of a divalent cation chelating agent (e.g., EDTA), required by the exonucleases to function, is used to preserve the DNA sufficiently intact and enable sequencing. In some embodiments, cell-free nucleic acids are detached from cells via encapsulation in exosomes. The exosomes are isolated by ultracentrifugation or by using a spin column (Quiagen), and the DNA or RNA contained therein is collected and sequenced.

[0383] In some embodiments, methylation information is obtained from cell-free nucleic acids according to the methods described above.

[0384] Combination of sequencing technologies In some embodiments, the methods described herein are combined with other sequencing techniques. In some embodiments, sequencing by a second method is initiated on the same molecule following sequencing by transient binding. For example, a longer and more stable oligonucleotide is bound to initiate sequencing by synthesis. In some embodiments, the method is stopped if complete genome sequencing is not achieved and is used to provide a scaffold for short-read sequencing, such as that of Illumina. In this case, it is advantageous to perform Illumina library preparation by excluding the PCR amplification step to obtain higher genome coverage. One advantage of some of these embodiments is that the required sequencing coverage is halved, for example, from about 40 times to 20 times. In some embodiments, this is due to the additional sequencing performed by this method and the location information provided by this method. In some embodiments, longer-lasting and more stable oligos, which may be optically labeled, may be bound to a target to mark a specific region of interest in the genome (e.g., the BRCA1 locus) before or simultaneously with (preferably with a different label) short sequencing of the oligos in part or all of the sequencing process.

[0385] Machine learning methods In some embodiments, artificial intelligence or machine learning is used to learn the behavior of members of a repertoire when testing against polymers with known sequences (e.g., polynucleotides) and / or when cross-validating polynucleotide sequences with data obtained from other methods. In some embodiments, the learning algorithm considers the complete behavior of this probe against one or more polynucleotide targets, including probe binding sites specific to one or more conditions or contexts. The more sequencing is performed on the same or different samples, the more reliable the knowledge from machine learning becomes. What is learned from machine learning can be applied to a variety of other assays, particularly emerging sequencing based on transient binding, as well as those involving oligo-oligo / polynucleotide interactions (e.g., sequencing by hybridization).

[0386] In some embodiments, artificial intelligence or machine learning is trained by providing experimentally obtained binding pattern data to bind a complete repertoire of short oligos (e.g., trimers, tetramers, pentamers, or hexamers) to one or more polynucleotides of a known sequence. The training data for each oligo includes the binding site, binding duration, and the number of binding events over a given period. After this training, the machine learning algorithm can be applied to the polynucleotides of the sequence to be determined and, based on its learning, assemble the sequence of polynucleotides. In some embodiments, the machine learning algorithm is also provided with a reference sequence.

[0387] In some embodiments, the array assembly algorithm includes both machine learning elements and non-machine learning elements.

[0388] In some embodiments, instead of computer algorithms learning from experimentally obtained binding patterns, binding patterns are obtained through simulation. For example, in some embodiments, simulations are performed on transient binding of oligos from a repertoire to polynucleotides of known sequences. The simulations are based on models of the behavior of each oligo obtained from experimental or published data. For example, predictions of binding stability are obtained according to the nearest neighbor method (e.g., SantaLucia et al., Biochemistry 35, 3555-3562 (1996) and Breslauer et al., Proc. Natl. Acad. Sci. 83: 3746-3750, 1986). In some embodiments, the behavior of mismatches is known (e.g., a mismatch of G binding to A may be similar to or stronger than the interaction of T binding to A) or experimentally induced. Furthermore, in some embodiments, unusually high binding strengths of some short subsequences of the oligo (e.g., GGA or ACC) are known. In some embodiments, a machine learning algorithm is trained on simulated data and then used to determine the sequence of an unknown sequence when it is interrogated by a complete repertoire of short oligos.

[0389] In some embodiments, data on oligos from a repertoire or panel (location, binding duration, signal intensity, etc.) are inserted into a machine learning algorithm trained on one or more known sequences, preferably (tens, hundreds, or thousands). The machine learning algorithm is then applied to create a dataset from the sequence in question, and the machine learning algorithm creates the sequence of the unknown sequence in question. Training of algorithms for sequencing organisms with relatively small or less complex genomes (e.g., for bacteria, bacteriophages, etc.) should be performed on organisms of that type. For organisms with larger or more complex genomes (e.g., S. pombe or humans), especially those with repetitive DNA regions, training should be performed on organisms of that type. For long-range assembly of megabase fragments to the full chromosome length, training is performed on similar organisms in some embodiments so that the specific aspects of the genome are represented during training. For example, the human genome is diploid and exhibits large sequence regions with segmental duplication. Other genomes of interest, particularly many agriculturally important plant species, have highly complex genomes. For example, wheat and other grains have highly polyploid genomes.

[0390] In some embodiments, a machine learning-based sequence reconstruction approach includes (a) providing information on the binding behavior of each oligo in a repertoire collected from one or more training datasets; (b) providing information on the physical binding of each oligo in the repertoire to a polynucleotide whose sequence is determined; and (c) providing information for each oligo on the binding site and / or binding duration and / or the number of times binding occurs (e.g., duration of binding repeats) at each location.

[0391] In some embodiments, the sequence of a specific experiment is first processed by a non-machine learning algorithm. The output sequence of the first algorithm is then used to train a machine learning algorithm, thereby training on the actual experimentally derived sequence of the same exact molecule. In some embodiments, the sequence assembly algorithm includes a Bayesian approach. In some embodiments, the data derived from the methods of this disclosure are fed into an algorithm of the type described in WO2010075570 and, optionally, combined with other types of genomic or sequencing data.

[0392] In some embodiments, sequences are extracted from the data in numerous ways. At one end of the spectrum of sequence reconstruction methods, sequences are obtained simply by aligning monomers or strings, as monomer localization or monomer strings are very precise (nanometer or subnanometer size). At the other end of the spectrum, the data are used to eliminate various hypotheses about the sequences. For example, one hypothesis is that the sequences correspond to known individual genome sequences. The algorithm determines where the data diverges from individual genomes. In another example, the hypothesis is that the sequences correspond to known genome sequences of "normal" somatic cells. The algorithm determines where the data from putative tumor cells diverges from sequences of "normal" somatic cells.

[0393] In one embodiment of this disclosure, a training set comprising one or more known target polynucleotides (e.g., lambda phage DNA, or a synthetic construct comprising a supersequence containing complementary strands to each oligo in the repertoire) is used for tested iterative binding of each oligonucleotide from the repertoire. In some embodiments, a machine learning algorithm is used to determine the binding and mismatch characteristics of the oligo probe. Thus, counterintuitively, mismatch binding is seen as a way to provide further data used to assemble sequences and / or to add confidence to sequences.

[0394] Sequencing equipment and devices Sequencing methods have general instrumentation requirements. Essentially, the instrumentation must be capable of imaging and reagent exchange. Imaging requirements include one or more of the following: objective lenses, relay lenses, beam splitters, mirrors, filters, and cameras or point detectors. Cameras include CCD or array CMOS detectors. Point detectors include photomultiplier tubes (PMTs) or avalanche photodiodes (APDs). In some examples, high-speed cameras are used. Any other configurations are adapted depending on the form of the method. For example, the light source (e.g., lamp, LED, or laser), the connection of illumination to the substrate (e.g., prism, grating, sol-gel, lens, movable stage, or movable objective lens), the mechanism for moving the sample relative to the imager, sample mixing / stirring, temperature control, and electrical control are each independently adapted to the different embodiments disclosed herein.

[0395] For single-molecule applications, illumination is preferably via evanescent wave generation, for example, by bringing laser light to the edge of the substrate at an appropriate angle, through prism-based total internal reflection, objective lens-based total internal reflection, grating-based waveguide, hydrogel-based waveguide, or evanescent mode waveguide. In some examples, the waveguide includes a core layer and a first coating layer. Alternatively, illumination includes HILO illumination or an optical sheet. In some single-molecule instruments, the effects of light scattering are mitigated by utilizing the synchronization of pulsed illumination and time-gating detection, and in this specification, light scattering is gated out. In some embodiments, dark-field illumination is used. Some instruments are configured for fluorescence lifetime measurement.

[0396] In some embodiments, the apparatus also includes means for extracting polynucleotides from cells, nuclei, organelles, chromosomes, and the like.

[0397] The instrument suitable for most embodiments is the Illumina Genome Analyzer IIx. This instrument includes a prism-based TIR, a 20x dry objective lens, a light scrambler, 532nm and 660nm lasers, an infrared laser-based focusing system, an emission filter wheel, a Photometrix CoolSnap CCD camera, temperature control, and a syringe pump-based system for reagent exchange. Improvements to this instrument with alternative camera combinations enable better single-molecule sequencing in some embodiments. For example, the sensor preferably has low electron noise of less than 2e. The sensor also has a large number of pixels. The syringe pump-based reagent exchange system is replaced in some embodiments with a pressure-driven flow-based system. The system is used in some embodiments with a compatible Illumina flow cell or a conventional flow cell adapted to accommodate the instrument's active or modified pumping.

[0398] Alternatively, a motorized Nikon Ti-E microscope coupled with a laser bed (laser depending on the label selection) or a laser system from a genome analyzer and a light scrambler, an EM CCD camera (e.g., Hamamatsu ImageEM) or a scientific CMOS (e.g., Hamamatsu Orca FLASH), and optionally temperature control may be used. In some embodiments, consumer sensors rather than scientific sensors are used. This has the ability to dramatically reduce sequencing costs. This is coupled with a pressure-driven or syringe-driven pump system and a specially designed flow cell. In some embodiments, the flow cell is made of glass or plastic, each with its own advantages and disadvantages. In some embodiments, the flow cell is made using microfabrication methods with cyclic olefin copolymers (COCs), e.g., TOPAS, other plastics or PDMS, or silicon or glass. In some embodiments, injection molding of thermoplastic resins provides low-cost routers for industrial-scale manufacturing. In some optical configurations, the thermoplastic resin needs to have good optical properties with minimal internal fluorescence. Polymers that do not contain aromatic or conjugated systems should ideally be excluded, as they are expected to have considerable internal fluorescence. Zeonor 1060R, Topas 5013, and PMMA-VSUVT (e.g., described in U.S. Patent No. 8,057,852) have reasonable optical properties in the green and red wavelength ranges (e.g., for Cy3 and Cy5), with Zeonor 1060R reported to have the most suitable properties. In some embodiments, it is possible to adhere thermoplastic resins over large areas in microfluidic devices (e.g., reported by Sun et al., Microfluidics and Nanofluidics, 19(4), 913-922, 2015). In some embodiments, a glass cover glass to which the biopolymer is attached is adhered to the thermoplastic fluid structure.

[0399] Alternatively, a manually operated flow cell is used at the top of the microscope. In some embodiments, this is constructed by fabricating the flow cell using a double-sided adhesive sheet, laser-cut to have a channel of appropriate dimensions, and sandwiched between a coverslip and a glass slide. From one reagent exchange cycle to another, the flow cell remains on the instrument / microscope, allowing for positioning between frames. In some embodiments, a motorized stage with a linear coder is used to ensure that the stage moves when imaging large areas, returning to the correct position. Fiduciary markers are used to maintain correct positioning. In this example, it is preferable to have alignment markings, such as etching in the flow cell that is optically detected or beads immobilized on the surface within the flow cell. If a polynucleotide backbone is stained (e.g., by YOYO-1), their fixed, known positions are used to align images from one frame to the next.

[0400] In one embodiment, a lighting mechanism utilizing laser or LED illumination (e.g., as described in U.S. Patent No. 7,175,811 and Ramachandran et al., Scientific Reports 3:2133, 2013) is coupled with an optional heating mechanism and reagent exchange system to perform the method described herein. In some embodiments, a smartphone-based imaging setup (ACS Nano 7:9147) is coupled with an optional temperature control module and reagent exchange system. In such embodiments, it is primarily the camera of the phone used, but other embodiments such as the illumination and vibration capabilities of an iPhone or other smartphone device may also be used.

[0401] Figures 20A and 20B show possible devices for performing transient probe-coupled imaging as described herein, using a flow cell 2004 and an integrated optical layout. Reagents are delivered as packets of reagent / buffer 2008 separated by an air gap 2022. Figure 20A shows an exemplary layout (e.g., TIRF setup) in which an evanescent wave 2010 is generated by combining laser light 2014 transmitted through a prism 2016. In some embodiments, the reaction temperature is controlled by an integrated thermal control 2012 (e.g., in one example, a transparent substrate 2024 contains electrically coupled indium tin oxide, thereby changing the temperature of the entire substrate 2024). Reagents are delivered as a continuous flow of reagent / buffer 2008. Laser light 2014 is coupled using a lattice waveguide 2020 or a photon structure to generate the evanescent field 2010. In some embodiments, thermal control is provided by a space-covering block 2026.

[0402] The layout configuration shown in Figure 20A is compatible with the layout configuration shown in Figure 20B. Alternatively, for example, objective lens-style TIRFs, light-guided TIRFs, and condenser-style TIRFs may be used. Continuous or air-gap reagent delivery is controlled in some embodiments by a syringe pump or pressure-driven flow. In the air-gap method, all of the reagent 2008 is pre-loaded into the capillary / piping 2102 (e.g., shown in Figure 21) or channel and delivered by pushing or pulling from a syringe pump or pressure-controlled system. The air gap 2022 contains air or a gas such as nitrogen or a liquid that is immiscible with aqueous solutions. Molecular combing and reagent delivery can also be performed using the air gap 2022. A fluid device (e.g., a fluid vessel, cartridge, or tip) includes a flow cell area where polynucleotide immobilization and optionally extension are performed, a reagent storage area, an inlet, an outlet, and an optional structure that forms a polynucleotide extraction and evanescent field. In some embodiments, the device is made of glass, plastic, or a glass-plastic hybrid. In some embodiments, a thermal and electrical conductive element (e.g., metal) is integrated into the glass and / or plastic component. In some embodiments, the fluid vessel is a well. In some embodiments, the fluid vessel is a flow cell. In some embodiments, the surface is coated with one or more chemical layers, biochemical layers (e.g., BSA-biotin, streptavidin), liquid layers, hydrogels, or gel layers. Subsequently, a 22×22 mm cover glass is coated with vinylsilane (available from BioTechniques 45:649-658, 2008 or Genomic Vision), or the cover glass is spin-coated with 1.5% Zeonex in a chlorobenzene solution.The substrate may also be coated with 2% 3-aminopropyltriethoxysilane (APTES) or polylysine, and extension occurs by electrostatic interaction in HEPES buffer at pH 7.5–8. Alternatively, the silanated coverslip is spin-coated or dip-coated in a 1–8% polyacrylamide solution containing bisacrylamide and TEMED. For this purpose, and using vinyl silane-coated coverslips, the coverslip can be coated with 10% (v / v) 3-methacrylateoxypropyltrimethoxysilane (Bind Silane; Pharmacia Biotech) in acetone for 1 hour. Polyacrylamide coatings can also be obtained as described (Liu Q et al. Biomacromolecules, 2012, 13(4), pp 1086–1092). Several hydrogel coatings that can be used are described and referenced in Mateescu et al. Membranes 2012, 2, 40–69.

[0403] Nucleic acids can also be elongated in an agarose gel by applying an alternating current (AC) electric field. DNA molecules can be electrophoresed onto the gel, or DNA can be mixed with molten agarose and then incorporated into the agarose. An AC field with a frequency of approximately 10 Hz is then applied, and an electric field strength of 200–400 V / cm is used. Elongation can be performed in an agarose gel concentration range of 0.5–3%. In some examples, the surface is coated with BSA-biotin in a flow channel or well, and then streptavidin or NeutrAvidin is added. Using this coated coverslip, double-stranded genomic DNA can be elongated by first binding the DNA in a pH 7.5 buffer and then elongating the DNA in a pH 8.5 buffer. In some examples, a streptavidin-coated coverslip is used to capture and immobilize the nucleic acid strands, but elongation is not performed. Thus, nucleic acids are attached to one end, while the other end dangles in the solution.

[0404] In some embodiments, a more integrated monolithic device is constructed for sequencing, rather than using various microscope-like components of optical sequencing systems such as GAIIx. In such embodiments, polynucleotides are attached to a sensor array and / or adjacent substrates and optionally directly extended. Direct detection on the sensor array has been demonstrated for DNA hybridization to the array (e.g., Lamture et al., Nucleic Acid Research 22:2121-2125, 1994). In some embodiments, the sensor is time-gated to reduce background fluorescence by Rayleigh scattering, which has a shorter lifetime compared to emission from a fluorescent dye.

[0405] In one embodiment, the sensor is a CMOS detector. In several embodiments, multiple colors are detected (for example, as described in U.S. Patent Application Publication 2009 / 0194799). In several embodiments, the detector is a Foveon detector (for example, as described in U.S. Patent No. 6,727,521). In several embodiments, the sensor array is an array of triple junction diodes (for example, as described in U.S. Patent No. 9,105,537).

[0406] In some embodiments, the reagent / buffer is delivered to the flow cell in a single dose (e.g., via a blister pack). Each blister in the pack contains a different oligonucleotide from a repertoire. Without any mixing or contamination between the oligonucleotides, the first blister is punctured and the nucleic acid is exposed to its contents. In some embodiments, after a washing step is applied, the process moves to the next blister. This serves to physically separate different sets of oligonucleotides, thereby reducing background noise where oligonucleotides from previous sets remain in the imaging field.

[0407] In some embodiments, sequencing occurs in the same device or monolithic structure from which cells are expelled and / or polynucleotides are extracted. In some embodiments, all reagents required to perform the method are pre-loaded into the fluid device before the analysis begins. In some embodiments, reagents (e.g., probes) are present in a dry state in the device and are wetted and dissolved before the reaction proceeds.

[0408] Examples Example 1: Sample preparation for sequencing Step 1: Extraction of long genomic DNA NA12878 or NA18507 cells (Coriell Biorepository) are cultured and harvested. The cells are mixed with low-melting-temperature agarose heated to 60°C. The mixture is poured into a gel mold (e.g., purchased from Bio-Rad) and set in a gel plug, resulting in approximately 4 × 10⁶ cells. 7 Prepare cells (this number can be larger or smaller depending on the desired density of polynucleotides). Lyse the cells in the gel plug by immersing it in a solution containing proteinase K. Gently wash the gel plug with TE buffer (e.g., in a 15 mL Falcon tube filled with washing buffer, leaving small bubbles to aid mixing, on a test tube rotating apparatus). Place the plug in a depression of approximately 1.6 mL volume and extract the DNA by digesting it with agarase enzyme. Apply a pH 5.5 solution of 0.5 M MES to the digested DNA. Perform this step using the FiberPrep kit (Genomic Vision, France) and associated protocols to obtain DNA molecules with an average length of 300 kb. Alternatively, obtain the genomic DNA itself extracted from these cell lines from Corriel and pipette it directly into a pH 5.5 solution of 0.5 M MES using a wide-bore pipette (giving an average spacing of less than 1 μM at approximately 10 μL in 1.2 mL).

[0409] Step 2: Molecular extension on the surface The final part of Step 1 involves adding the extracted polynucleotides to a depression containing a pH 5.5 solution of 0.5 M MES. A vinylsilane-coated substrate coverslip (e.g., Genomic Vision CombiSlips) is immersed in the depression and incubated for 1–10 minutes (depending on the desired polynucleotide density). The coverslip is then gently peeled off using a mechanical puller, such as a syringe pump with a clip attached to grasp it (or using Genomic Vision FiberComb). The DNA on the coverslip is crosslinked to the surface using a crosslinker (Stratagene, USA) with an energy of 10,000 microjoules. If the process is performed carefully, it yields high molecular weight (HMW) polynucleotides, where the length present in the polynucleotide population exceeds 1 Mb, or molecules approximately 10 Mb long, with an average surface-extended length of 200–300 kb. With more careful optimization, the average length can be moved to the megabase range (see the section on megabase range combing above).

[0410] As an alternative, as mentioned earlier, pre-extracted DNA (e.g., Novagen cat. No. 70572-3 or Promega human male genomic DNA) is used, which contains a good proportion of genomic molecules larger than 50 kb. When resolving high fractions individually using diffraction-limited imaging, a concentration of approximately 0.2–0.5 ng / μL and an immersion time of approximately 5 minutes is sufficient to provide molecular density, as specified herein.

[0411] Step 3: Flow cell preparation The coverslip is pressed onto a flow cell gasket made from a double-sided adhesive 3M sheet already attached to the glass slide. The gasket (with protective layers on both sides on the double-sided adhesive sheet) is fabricated using a laser cutter to create one or more flow channels. The length of the flow channels is longer than the length of the coverslip, so that when the coverslip is placed in the center of the flow channels, one portion of the channel at each end not covered by the coverslip is used as an inlet and outlet for distributing fluid into and out of the flow channels, respectively. The fluid passes over an extended polynucleotide adhered to the vinylsilane surface. The fluid flows through the channels using a safety swab stick (Johnsons, USA) at one end, and aspiration occurs when the fluid is drawn in by pipetting at the other end. The channels are pre-moistened with phosphate-buffered saline-Tween and phosphate-buffered saline (PBS washing solution).

[0412] Step 4: Denaturation of double-stranded DNA Before the next oligo can be added, the previous oligo must be efficiently washed away; this can be done by changing the buffer up to four times and, if necessary, removing persistent binding with a denaturing agent such as DMSO or an alkaline solution. Double-stranded DNA is denatured by flushing it through a flow cell with alkali (0.5 M NaOH) and incubating it at room temperature for approximately 20–60 minutes. This is followed by a PBS / PBST wash. Alternatively, incubation can also be performed by 1 hour in 1 M HCl followed by a PBS / PBST wash.

[0413] Step 5: Passivation Depending on the situation, a blocking buffer such as BlockAid (Invitrogen, USA) may be added and incubated for approximately 5-15 minutes. This is followed by a PBS / PBST wash.

[0414] Example 2: Sequencing by transient binding of oligonucleotides to denatured polynucleotides Step 1: Addition of oligonucleotides under transient binding conditions. The flow cell is pre-conditioned with PBST and, optionally, with buffer A (10 mM Tris-HCl, 100 mM NaCl, 0.05% Tween-20, pH 7.5). Approximately 1–10 nM of each oligonucleotide is applied to the extended denatured polynucleotide in buffer B (5 mM Tris-HCl, 10 mM MgCl2, 1 mM EDTA, 0.05% Tween-20, pH 8) or buffer B+ (5 mM Tris-HCl, 10 mM MgCl2, 1 mM EDTA, 0.05% Tween-20, pH 8, 1 mM PCA, 1 mM PCD, 1 mM Trolox). The oligo lengths are typically in the range of 5–7 nucleotides, and the reaction temperature depends on the Tm of the oligo. One probe type used by the inventors was of the general formula 5'-Cy3-NXXXXXN-3' (where X is a specified base and N is a degenerate position), with LNA nucleotides at positions 1, 2, 4, 6, and 7, and DNA nucleotides at positions 3 and 5, purchased from Sigma Proligo, and as previously used by Pihlak et al. The binding temperature was linked to the Tm of each oligo sequence.

[0415] After washing with A+ and B+ solutions, transient binding of oligonucleotides is performed at room temperature with LNA DNA chimeric oligo 3004 NTgGcGN (uppercase is LNA, lowercase is DNA nucleotide) using oligos in B+ solution at concentrations between 0.5 and 100 nM (typically between 3 nm and 10 nm). Different temperatures and / or salt conditions (and concentrations) are used for different oligo sequences according to their Tm and binding behavior. When the FRET mechanism is used for detection, fairly high concentrations of oligos up to 1 μM may be used. In some embodiments, FRET is between an intercalating dye molecule (appropriately, YOYO-1, Sytox Green, Sytox Orange, Sybr Gold, etc.; in dilutions of 1 / 1000 to 1 / 10,000 depending on which intercalating dye from Life Technologies is used) that intercalates transiently into the duplex, and a label on the oligo. In some embodiments, the intercalating dye is used directly as a label without FRET. In this case, the oligonucleotides are not labeled. Not only are they inexpensive, but unlabeled oligonucleotides can be used at higher concentrations than labeled oligonucleotides because the background from the intercalating dye during heteroduplex formation is 100 to 1000 brighter than that of unintercalating dyes (for example, depending on which intercalant is used).

[0416] Step 2: Imaging - Capture multiple frames The flow channel is installed in an inverted microscope (e.g., Nikon Ti-E) equipped with Perfect Focus, a TIRF attachment and TIRF objective lens laser, and a Hamamatsu 512x512 back-illuminated EMCCD camera. The probe is added to buffer B+ and, if necessary, assisted in imaging.

[0417] A probe bound to a polynucleotide disposed on the surface is illuminated by evanescent waves generated by the total internal reflection of a 75-400mW laser beam (e.g., 532nm green light) tuned through a fiber optic scrambler (Point Source) with a TIRF angle of approximately 1500 degrees, via a 1.49 NA 100x Nikon oil immersion objective lens on a Nikon Ti-E with a TIRF attachment. The image is collected through the same lens with a further magnification of 1.5x and projected onto a Hamamatsu ImageEM camera via a dichroic mirror and emission filter. 5000-30,000 frames of 50-200 milliseconds are acquired with an EM gain of 100-140 using Perfect Focus. Preferably, a high laser power (e.g., 400mW) is used early for a few seconds to decolorize the initial nonspecific binding, reducing the blanket of most of the signal from the surface to a lower density and resolving individual binding events. The laser power is then optionally reduced.

[0418] Figures 22A–22E show examples of illumination of probes transiently binding to target polynucleotides. In these figures, the target polynucleotides are derived from human DNA. Dark spots indicate the fluorescent regions of the probe, with darker spots indicating more regions that were more frequently bound by the probe (e.g., more photons were collected). Figures 22A–22E are images (e.g., video) from a time series captured during sequencing of a single target polynucleotide. Points 2202, 2204, 2206, and 2208 are shown throughout the time series as examples of regions in the polynucleotide that were bound with some intensity over time (e.g., because different sets of oligonucleotides were exposed to the target polynucleotide).

[0419] An imaging buffer is added. In some embodiments, the imaging buffer is supplemented or replaced with a buffer containing beta-mercaptoethanol, an enzymatic redox system, and / or ascorbate and gallic acid. Fluorophores are detected along the line, indicating that binding has occurred. If the flow cell is prepared with more than one channel, one of the channels is stained with a YOYO-1 intercalating dye to check the polynucleotide density and the quality of the polynucleotide extension (e.g., using Intensilight or 488 nm laser illumination).

[0420] Step 3: Image acquisition - Move to another location (optional step) By moving the cover glass (mounted on the slide holder of the Nikon Ti-e as part of the flow cell) relative to the objective lens (i.e., the CCD), another location is imaged. Imaging is performed at multiple other locations, thereby capturing probes that bind to polynucleotides or portions of polynucleotides given at different locations (outside the field of view of the CCD at the first location). Image data from each location is stored in computer memory.

[0421] Step 4: Add the next set of oligosaccharides Add the next set of oligonucleotides and repeat steps 1-3 until the entire polynucleotide is sequenced.

[0422] Step 5: Determining the location and identity of the joint The location of each fluorescence point signal is detected, and the location of the pixel from which fluorescence from the bound label is projected is recorded. The identity of the bound oligonucleotide is determined by determining which labeled oligonucleotide is bound, for example, by wavelength selection using optical filters. The fluorophore is detected by multiple filters, and in this case, the emission signature of each fluorophore in the filter set is used to determine the identity of the fluorophore and, by extension, the oligonucleotide. If the flow cell is prepared with more than one channel, one of the channels is stained with a YOYO-1 intercalating dye to check the polynucleotide density and the quality of the polynucleotide extension (for example, using Intensilight or 488 nm laser illumination). One or more images or movies are captured for each fluorescence wavelength used to label the oligonucleotide.

[0423] Step 6: Data Processing If both strands of the duplex remain attached to the surface, oligo binding occurs simultaneously at complementary locations on both strands of the double helix. The entire dataset is then analyzed to identify sets of oligos that provide signals closely localized to specific locations on the nucleic acid, their locations confirmed by overlapping oligo sequences corresponding to selected points in the polynucleotide; this then reveals two overlapping tiling series for each oligo. The tiling series to which the next signal in the positional relationship fits indicates the strand to which it is bound.

[0424] Since the strands remain fixed to the surface, the binding sites recorded for each oligo can be overlaid using a software script that runs an algorithm. This yields a signal indicating that the oligo binding sites fall within a framework of tiling paths of the two oligonucleotide sequences, i.e., separate (but complementary) paths for each strand of the denatured duplex. Each tiling path, if complete, spans the entire length of its strand. The tiled sequences for each strand are then compared to provide a double-stranded (also known as 2D) consensus sequence. If a gap exists in one of the tiling paths, the sequence of the complementary tiling path is taken. In some embodiments, the sequence is compared to the same sequence or reference in multiple copies to assist in base assignment and close the gap.

[0425] Example 3: Detection of epimarker locations on polynucleotides Depending on the circumstances, transient binding of an epigenomic conjugation reagent is performed before (or sometimes after or during) the oligobinding step. Depending on the reagent used, binding is performed before or after denaturation. In the case of anti-methyl C antibodies, binding is performed on denatured DNA, while in the case of methyl-binding proteins, binding is performed on double-stranded DNA before any denaturation step.

[0426] Step 1 - Transient binding of methyl bonding reagent After denaturation, the flow cell is flushed with PBS washing solution, and the Cy3B-labeled anti-methyl antibody 3D3 clone (Diagenode) is added to the PBS.

[0427] Alternatively, before denaturation, flush the flow cell with PBS and add Cy3B-labeled MBD1.

[0428] Imaging will be performed as previously described for transient oligobinding.

[0429] Step 2: Detachment of methyl bond reagent Typically, epianalysis is performed before sequencing. Therefore, methyl binding reagents are sometimes flushed off before the polynucleotides, before sequencing begins. This is done by running PBS / PBST and / or high-salt buffer and SDS through multiple cycles, and then checking by imaging that removal has occurred. If it is clear that a negligible amount of binding reagent remains, the remaining reagent is removed through harsher treatment, such as with a chaotropic salt like GuCl.

[0430] Step 3: Data Correlation Epigenomics data is obtained after sequencing, and correlations are performed between sequencing binding sites to correlate epi-binding sites and provide the chronological relationship of methylation sequences.

[0431] Example 4: Fluorescence collected from transient binding to lambda phage DNA Figures 23A, 23B, and 23C show examples of transient binding events. They collectively demonstrate transient binding of oligo-IDLin2621, Cy3-labeled 5'NAgCgGN3', at a concentration of 1.5 nM in buffer B+ at room temperature. The target polynucleotide is a lambda phage genome manually combed onto a vinylsilane surface (Genomic Vision) in MES pH 5.5 buffer + 0.1 M NaCl. A 400 mW laser at 523 nM was used through a point source fiber optic scrambler. Fluorescence was collected with a 532 nm excitation band, a 100x TIRF objective lens, and a TIRF attachment with additional magnifications of 1.49 NA and 1.5x, and in multi-chroic mode. Vibration isolation was not performed. Images were captured with Perfect Focus on a Hamamatsu ImageEM 512x512 with a 100 EM gain setting. 10,000 frames were collected at 100 ms. The concentration of Cy3 in the oligonucleotide probe set was approximately 250 nM–300 nM. Figure 23A shows fluorescence collected before cross-correlation drift correction in ThunderSTORM. Figure 23B shows fluorescence collected after cross-correlation drift correction with a scale bar. Figure 23C shows fluorescence in an enlarged area of ​​Figure 23B. Figure 23C shows a long polynucleotide chain detected by persistent binding of Lin2621 to multiple locations. From the images, it is clear that the target polynucleotide chain was immobilized and extended to the imaging surface at a distance closer than the diffraction limit of Cy3 emission.

[0432] Example 5: Fluorescence collected from transient binding in synthetic DNA Figure 24 shows examples of fluorescence data collected from three different polynucleotide strands. Multiple probing and washing steps are shown with synthetic 3 kilobase denatured double-stranded DNA. The synthetic DNA was denatured by combing it in MES pH 5.5 on a vinylsilane surface. A series of binding and washing steps were performed, video was recorded, and the results were processed in ImageJ using THunderSTORM. The strands (1, 2, and 3) from the three examples were cropped from super-resolution images of the following experimental series performed at ambient temperature with 10 nM oligo in buffer B+: oligo 3004 binding, washing, oligo 2879 binding, washing, oligo 3006 binding, washing, and oligo 3004 binding (again). This demonstrates that binding maps can be derived from transient binding, binding patterns can be erased by washing, and then different binding patterns can be obtained with different oligos on the same first and second strands of synthetic DNA. The regression to oligo 3004 at the end of this series, and its similarity to the pattern used at the beginning of the series, demonstrates the robustness of the process without any arbitrary optimization trials.

[0433] The experimentally determined binding sites correspond to three of the four possible perfect-match binding sites shown by duplex chains 1 and 3, as well as all four binding sites and one significant mismatch site shown by duplex chain 2. A second probing at oligo 3004 is observed to show a clearer signal, likely due to a lower mismatch. This is consistent with the possibility of a slight temperature increase due to heating from prolonged exposure to laser light.

[0434] The oligo sequences used in this experiment are as follows (uppercase bases are locked nucleic acids (LNAs)): Oligo(Olio)3004: 5'cy3 NTgGcGN Oligo2879: 5' cy3 NGgCgAN Oligo3006: 5' cy3 NTgGgCN:

[0435] The sequence listing for the 3kbp synthetic template (below the document) is as follows:

[0436] Example 6: Integrated isolation, nucleic acid extraction, and sequencing of single cells Step 1: Design and fabricate the microfluidic structure. The microchannels are designed to accommodate cells derived from human cancer cell lines with a typical diameter of 15 μm, thereby giving the microfluidic network a minimum depth and width of 33 μm. The device includes an inlet for the cells and an inlet for a buffer that merges into a single channel and supplies to a single-cell trap (shown in Figure 17). At the intersection between the cell inlet and the buffer inlet, the cells are aligned along the sidewall of the supply channel where one or more traps are located. Each trap is a simple constriction sized to capture cells derived from human cancer cell lines. The constrictions for the cell traps have a trapezoidal cross-section: it is 4.3 μm wide at the bottom, 6 μm wide in the middle of the depth, and 8 μm wide at the top, with a depth of 33 μm. Each cell trap connects the supply channel to a branching point, with one side being a discard channel (not shown in Figure 17) and the other side being a channel containing a flow stretch section (for nucleic acid extension and sequencing), one of which is present for each cell. The flow stretch section consists of a channel with a width of 20 μm (or up to 2 mm), a length of 450 μm, and a depth of 100 nm (or up to 2 μm). In some embodiments, the flow stretch channel is narrower at the start and expands to the dimensions described.

[0437] Step 2: Device Fabrication The device is fabricated by replicating nickel shims using injection molding of TOPAS 5013 (TOPAS). Briefly, a silicon master is fabricated by UV lithography and reactive ion etching. A 100 nm NiV seeding layer is deposited, and nickel is electroplated to a final thickness of 330 μm. The Si master is removed by chemical etching in KOH. Injection molding is performed using a melting temperature of 250°C, a mold temperature of 120°C, a maximum holding pressure of 1,500 bar for 2 seconds, and an injection speed varying between 20 cm³ / s and 45 cm³ / s. Finally, a coverslip (1.5) is bonded to the device, or the device is sealed using a 150 μm TOPAS foil by a combination of UV and heat treatment under a maximum pressure of 0.51 MPa. The surface roughness of the foil is reduced by compressing the foil between two flat nickel plates electroplated from a silicon wafer for 20 minutes at 140°C and 5.1 MPa before sealing the device. This ensures that the device lid is optically flat, enabling high-NA optical microscopy. The device is mounted on an inverted fluorescence microscope (Nikon Ti-E) equipped with an oil immersion TIRF objective lens (100x / NA 1.49) and an EMCCD camera (Hamamatsu ImageEM 512). Fluid is driven through the device at a pressure in the range of 0-10 mbar using a pressure control device (MFCS, Fluigent). Ethanol is poured into the device, then degassed, and FACSFlow Sheath Fluid (BD Biosciences) is added to all microchannels except the microchannel connected to the flow stretch device. Selective addition is performed by maintaining positive pressure at the inlet of the supply channel into which the solution is introduced, applying negative pressure or suction at the outlet of the waste channel, and applying positive pressure at the outlet of the flow stretch channel. A buffer suitable for single-molecule imaging and electrophoresis (0.5×TBE + 0.5% v / v Triton-X100 + 1% v / v beta-mercaptoethanol: BME) is added to the channel of the flow stretch device.This buffer prevents DNA adhesion in the flow-stretch section and suppresses electroosmotic flow that could resist the introduction of extracted DNA when the height of the ...

Claims

1. A method for sequencing nucleic acids, (a) Immobilizing the nucleic acid on a test substrate in a double-stranded linearized extension form, thereby forming an immobilized extended double-stranded nucleic acid; (b) Denature the immobilized extended double-stranded nucleic acid into a single-stranded form on the test substrate, thereby obtaining an immobilized first strand and an immobilized second strand of the nucleic acid, wherein each base of the immobilized second strand is located adjacent to the corresponding complementary base of the immobilized first strand; (c) Exposure the immobilized first chain and the immobilized second chain to each oligonucleotide probe in a set of oligonucleotide probes, wherein each oligonucleotide probe in the set of oligonucleotide probes is of a predetermined sequence and length and includes a label, the label being selected from the group consisting of dyes, fluorescent nanoparticles, light scattering particles and FRET partners, and the exposure (c) occurs under conditions that cause each individual probe of each oligonucleotide probe to bind to a portion of the immobilized first chain or the immobilized second chain complementary to each oligonucleotide probe and to form each heteroduplex, thereby causing each generation of optical activity; (d) Using a two-dimensional imager, measure the location and duration of each occurrence of optical activity on the test substrate during exposure (c); (e) Repeating exposure (c) and measurement (d) for each oligonucleotide probe in the set of oligonucleotide probes, thereby obtaining a plurality of sets of positions on the test substrate, wherein each set of positions on the test substrate corresponds to one oligonucleotide probe in the set of oligonucleotide probes; and (f) Determining the sequence of at least a portion of the nucleic acid from the set of positions on the test substrate by compiling the oligonucleotide probe sequence on the test substrate, which is represented by the set of positions on the test substrate. Methods that include...

2. The method according to claim 1, wherein the exposure (c) occurs under conditions that transiently and reversibly bind each individual probe of each pool of each oligonucleotide probe to a portion of the immobilized first or immobilized second chain complementary to the individual probe and form each heteroduplex, thereby generating optical activity.

3. The method according to claim 1, wherein the exposure (c) occurs under conditions that repeatedly, transiently, and reversibly bind each individual probe of each pool of each oligonucleotide probe to a portion of the immobilized first or immobilized second chain that is complementary to the individual probe, and that each heteroduplex is formed, thereby repeatedly causing each generation of optical activity.

4. The method according to claim 1, wherein in step (c) of exposure, each oligonucleotide probe in the set of oligonucleotide probes is an oligonucleotide probe that is bound to a label.

5. The aforementioned exposure occurs in the presence of a first label in the form of a dye. Each oligonucleotide probe in the set of oligonucleotide probes is an oligonucleotide probe conjugated with a second label, and The first sign causes the second sign to fluoresce when the first sign and the second sign are close to each other. The method according to claim 1.

6. The method according to claim 1, wherein one or more oligonucleotide probes in the set of oligonucleotide probes are exposed to the immobilized first strand and the immobilized second strand during the exposure (c).

7. The method according to claim 1, wherein different oligonucleotide probes in the set of oligonucleotide probes exposed to the immobilized first strand and the immobilized second strand during the exposure (c) are associated with different labels.

8. The method according to claim 1, wherein the exposure (c) is performed on the first oligonucleotide probe in the set of oligonucleotide probes at a first temperature, and the repetition (e) of the exposure (c) and the measurement (d) is performed on the first oligonucleotide probe at a second temperature.

9. The exposure (c) is performed on the first oligonucleotide probe in the set of oligonucleotide probes at a first temperature. The repetition (e) of the exposure (c) and measurement (d) includes performing the exposure (c) and measurement (d) for the first oligonucleotide probe at each of a plurality of different temperatures, and The further includes constructing a melting curve for the first oligonucleotide probe using the locations and periods of optical activity recorded by the measurement (d) for the first temperature and each of the plurality of different temperatures, The method according to claim 1.

10. The method according to claim 1, wherein the set of oligonucleotide probes comprises a plurality of subsets of the oligonucleotide probes, and the repetition of exposure (c) and measurement (d) is performed for each respective subset of oligonucleotide probes in the plurality of subsets of oligonucleotide probes.

11. Measuring the location on the test substrate includes identifying and fitting each occurrence of optical activity using a Gaussian function or Fourier transform in order to identify the center of each occurrence of optical activity in the image frame of the data obtained by the two-dimensional imager and to fit it to the location on the test substrate, and The center of each occurrence of optical activity is considered to be the position of each occurrence of optical activity on the test substrate. The method according to claim 1.

12. Each occurrence of optical activity persists across multiple image frames measured by the two-dimensional imager. Measuring the location on the test substrate includes identifying each occurrence of optical activity across the plurality of image frames using a Gaussian function or Fourier transform and fitting it to the location on the test substrate in order to identify the center of each occurrence of optical activity across the plurality of image frames. The center of each occurrence of optical activity is considered to be the position of each occurrence of optical activity on the test substrate across the plurality of image frames. The method according to claim 1.

13. Measuring the location on the test board includes inputting image frames of the data measured by the two-dimensional imager into a trained convolutional neural network. The image frame of the data includes each of the occurrences of optical activity among a plurality of occurrences of optical activity, Each occurrence of optical activity in the plurality of occurrences of optical activity corresponds to an individual probe that binds to a portion of the fixed first chain or the fixed second chain. In response to the input, the trained convolutional neural network identifies the location on the test substrate of one or more occurrences of optical activity in the plurality of occurrences of optical activity. The method according to claim 1.

14. The method according to any one of claims 11 to 12, wherein the measurement resolves the center of each occurrence of optical activity at a position on the test substrate with a positional accuracy of at least 20 nm.

15. The method according to any one of claims 1 to 14, wherein the measurement of the location and duration on the test substrate for each occurrence of optical activity (d) is performed to measure more than 5,000 photons at the location.

16. The method according to any one of claims 1 to 15, wherein the standard deviation of each occurrence of optical activity is greater than a predetermined number of standard deviations with respect to the background observed on the test substrate.

17. The method according to claim 16, wherein the predetermined number of standard deviations is greater than 3.

18. The method according to any one of claims 1 to 17, wherein each oligonucleotide probe in the plurality of oligonucleotide probes contains a unique N-mer sequence, where N is an integer in the set {1, 2, 3, 4, 5, 6, 7, 8, 9, and 10} and all unique N-mer sequences of length N are represented by the plurality of oligonucleotide probes.

19. The method according to claim 18, wherein the unique N-mer sequence includes one or more nucleotide positions occupied by one or more degenerate nucleotide positions.

20. The method according to any one of claims 1 to 19, wherein the test substrate is washed before repeating the exposure (c) and measurement (d), thereby removing each oligonucleotide probe from the test substrate before exposing the test substrate to another oligonucleotide probe in the set of oligonucleotide probes.

21. The method according to any one of claims 1 to 20, wherein each occurrence of optical activity has an observation metric that satisfies a predetermined threshold, the observation metric includes the duration of binding / optical activity, signal-to-noise ratio, photon count, or fluorescence intensity.

22. Each oligonucleotide probe in the set of oligonucleotide probes has the same length M. M is a positive integer of 2 or more (bases), Determining the sequence of at least a portion of the nucleic acids from the plurality of sets of positions on the test substrate (f) further utilizes the overlap sequence of the oligonucleotide probe represented by the plurality of sets of positions. The method according to any one of claims 1 to 21.

23. The method according to claim 22, wherein determining the sequences of at least some of the nucleic acids from the plurality of sets of positions on the test substrate includes determining a first tiling path corresponding to the fixed first strand and a second tiling path corresponding to the fixed second strand.

24. A method for sequencing nucleic acids, (a) Immobilizing the linearized extended nucleic acid onto a test substrate to form an immobilized extended nucleic acid; (b) Exposing the immobilized extended nucleic acid to each pool of each oligonucleotide probe in a set of oligonucleotide probes, wherein each oligonucleotide probe in the set of oligonucleotide probes is of a predetermined sequence and length and includes a label, the label being selected from the group consisting of dyes, fluorescent nanoparticles, light scattering particles and FRET partners, and the exposure (b) occurs under conditions that enable each probe in each pool of each oligonucleotide probe to transiently and reversibly bind to a portion of the immobilized nucleic acid complementary to each oligonucleotide probe, thereby causing each generation of optical activity; (c) Using a two-dimensional imager, measure the location and duration of each occurrence of optical activity that occurs during the exposure (b); (d) Repeat the exposure (b) and measurement (c) for each oligonucleotide probe in the set of oligonucleotide probes to obtain a plurality of sets of positions on the test substrate, each set of positions on the test substrate corresponding to one oligonucleotide probe in the set of oligonucleotide probes; and (e) Determining the sequence of at least a portion of the nucleic acid from the set of positions on the test substrate by compiling the oligonucleotide probe sequence on the test substrate, which is represented by the set of positions on the test substrate. Methods that include...

25. A method for sequencing nucleic acids, (a) Immobilizing the linearized extended nucleic acid onto a test substrate to form an immobilized extended nucleic acid; (b) Exposing the immobilized extended nucleic acid to each pool of each oligonucleotide probe in a set of oligonucleotide probes, wherein each oligonucleotide probe in the set of oligonucleotide probes is of a predetermined sequence and length and is label-free, and the exposure (b) occurs under conditions that enable each probe in each pool of each oligonucleotide probe to transiently and reversibly bind to a portion of the immobilized nucleic acid that is complementary to each oligonucleotide probe, thereby causing each generation of optical activity; (c) Using a two-dimensional imager, measure the location and duration of each occurrence of optical activity that occurs during the exposure (b); (d) Repeat the exposure (b) and measurement (c) for each oligonucleotide probe in the set of oligonucleotide probes to obtain a plurality of sets of positions on the test substrate, each set of positions on the test substrate corresponding to one oligonucleotide probe in the set of oligonucleotide probes; and (e) Determining the sequence of at least a portion of the nucleic acid from the set of positions on the test substrate by compiling the oligonucleotide probe sequence on the test substrate, which is represented by the set of positions on the test substrate. Methods that include...

26. The method according to any one of claims 1 to 25, wherein the oligonucleotide probe is a probe comprising one or more modified nucleotides including LNA and PNA.

27. ​​The method according to any one of claims 1 to 26, wherein the fixation in (a) comprises coating the nucleic acid onto the test substrate by molecular combing (receding meniscus), flow stretching, nanoconfinement, or electrostretching.