High-throughput nucleic acid sequencing with single-molecule sensor arrays

JP2024156659A5Active Publication Date: 2025-06-27F HOFFMANN LA ROCHE & CO AG +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024103840
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2020-04-21
Filing Date
2024-06-27
Publication Date
2025-06-27
Estimated Expiration
2041-04-21

AI Technical Summary

Technical Problem

Current DNA sequencing technologies face challenges with low error rates and limited read lengths, particularly in single molecule sequencers, which often exhibit static and dynamic heterogeneity, making them unsuitable for high-precision diagnostics.

Method used

The development of single molecule array sequencing (SMAS) devices that utilize a plurality of sensors to detect labels attached to nucleotides in a nucleic acid strand, combined with error correction methods to reduce errors and improve sequencing accuracy.

Benefits of technology

SMAS devices achieve higher throughput, lower error rates, and longer read lengths compared to cluster-based approaches, enhancing the precision of nucleic acid sequencing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

To provide single-molecule array sequencing (SMAS) devices and systems.SOLUTION: Each sensor of an array of sensors of an SMAS device is capable of detecting labels attached to nucleotides incorporated into a single nucleic acid strand bound to a respective binding site. Each sensor can detect a single label (e.g., fluorescent, magnetic, organometallic, charged molecule, etc.) attached to the incorporated nucleotide. Also disclosed are methods of using SMAS devices and systems for highly-scalable nucleic acid (e.g., DNA) sequencing based on sequencing by synthesis (SBS) of multiple instances of clonally amplified DNA immobilized on such SMAS devices. Also disclosed are error correction methods that mitigate errors (e.g., errant label detections or non-detections) made in sequencing individual nucleic acid strands.SELECTED DRAWING: Figure 48B
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Application No. 63 / 013,236, entitled “High Throughput DNA Sequencing Using Single Molecule Sensor Arrays,” filed April 21, 2020 (Attorney Docket No. ROA-1002P-US / P36083-US), the entire contents of which are incorporated herein by reference. This application also incorporates by reference in its entirety for all purposes PCT Application No. PCT / US20 / 27290, filed April 8, 2020 (Attorney Docket No. ROA-1000-WO / P35097-WO), entitled “Nucleic Acid Sequencing by Synthesis Using Magnetic Sensor Arrays,” which was published on October 15, 2020 as WO 2020 / 210370, and PCT Application No. PCT / US2021 / 021274, filed March 7, 2021 (Attorney Docket No. ROA-1001-WO / P35967-WO), entitled “Magnetic Sensor Arrays for Nucleic Acid Sequencing and Methods of Making and Using Same.” [Background technology]

[0002] background Commercially successful approaches to DNA sequencing involve either the synthesis and analysis of clonal deoxyribonucleic acid (DNA) clusters or the detection of individual DNA molecules. Cluster sequencers exhibit error rates low enough for diagnostic applications, but are very limited in read length due to the nature of error propagation in molecular ensembles. Single molecule sequencers can generate significantly longer reads, but often exhibit static and dynamic heterogeneity that results in errors too large for high-precision diagnostics.

[0003] Thus, there is a need to improve DNA sequencing, and nucleic acid sequencing in general, to enable longer reads with lower error rates. Summary of the Invention

[0004] overview This summary represents a non-limiting embodiment of the present disclosure.

[0005] Disclosed herein are embodiments of single molecule array sequencing (SMAS) devices and systems. Each sensor of a plurality of sensors in a sensor array of the SMAS device detects a label attached to a nucleotide incorporated into a single nucleic acid strand bound to a respective binding site. Each sensor can detect a single label (e.g., fluorescent, magnetic, organometallic, charged molecule, etc.) attached to the incorporated nucleotide. Also disclosed are methods of using the SMAS device and systems for highly scalable nucleic acid (e.g., DNA) sequencing based on sequencing-by-synthesis (SBS) of multiple instances of clonally amplified DNA immobilized on such SMAS devices. Also disclosed are error correction methods that mitigate errors (e.g., erroneous label detection or non-detection) that occur when sequencing individual nucleic acid strands.

[0006] In some embodiments, an apparatus for sequencing nucleic acids comprises a fluid chamber, a plurality of S magnetic sensors configured to detect labels present in the fluid chamber, and at least one processor. The fluid chamber comprises a plurality of S binding sites, each of the S binding sites configured to bind to no more than one strand of nucleic acid. Each of the S magnetic sensors senses a respective strand of nucleic acid bound to a respective binding site of the S binding sites. The at least one processor is configured to execute one or more machine-executable instructions, the instructions, when executed, for each of a plurality of M interrogation steps of a sequencing procedure. In the interrogation step, the at least one processor is caused to, for each of the S magnetic sensors, (a) acquire a respective characteristic of the respective magnetic sensor, the respective characteristic being indicative of the presence or absence of the at least one landmark, and (b) determine, at least in part based on the acquired respective characteristic, whether the respective magnetic sensor detected the presence or absence of the at least one landmark during the interrogation step.

[0007] In some embodiments, the system includes a plurality of S binding sites, each of the S binding sites configured to bind to no more than one strand of nucleic acid, a plurality of S sensors (e.g., magnetic, optical, etc.) configured to detect a label, and at least one processor. Each of the S sensors is configured to sense a respective strand of nucleic acid bound to a respective binding site of the S binding sites. The at least one processor is configured to execute one or more machine-executable instructions that, when executed, cause the at least one processor to, in each interrogation step of the plurality of M interrogation steps of the sequencing procedure, for each of the S sensors: (a) acquire a respective characteristic of the respective sensor, the respective characteristic being indicative of the presence or absence of the at least one label; and (b) determine, at least in part, based on the acquired respective characteristic, whether the respective sensor detected the presence or absence of the at least one label during the interrogation step. Additionally, the one or more machine-executable instructions, when executed, further include causing the at least one processor to perform an error correction procedure on the at least one record, the at least one record including results of the sequencing procedure for at least a subset of the S sensors in each of the M interrogation steps.

[0008] In some embodiments, a method for sequencing a plurality of S nucleic acid strands using a SMAS device includes: (a) binding the S nucleic acid strands to S binding sites; (b) performing a sequencing procedure including M interrogation steps to generate S records, each of the S records capturing M detection results of a respective one of the S sensors, each of the M detection results indicating whether a respective one of the S sensors detected at least one label in the fluid chamber during a respective one of the M interrogation steps; and (c) applying an error correction procedure to at least a subset of the S records to infer a nucleic acid sequence of at least one of the S nucleic acid strands.

[0009] Some embodiments are methods for reducing errors in sequencing data generated as a result of a nucleic acid sequencing procedure using a single molecule sensor array, the single molecule sensor array having a plurality of sensors, each of the plurality of sensors associated with a respective binding site of a plurality of binding sites, each of the plurality of binding sites configured to bind to no more than one strand of nucleic acid being sequenced. In some such embodiments, the method includes: (a) identifying a plurality of records in the sequencing data, each of the plurality of records capturing a respective sequencing result for a respective instance of a first strand of nucleic acid, each of the plurality of records having a plurality of entries, each of the plurality of entries indicating, for a respective one of a plurality of interrogation steps of a nucleic acid sequencing procedure, either (i) a label was detected by a respective sensor associated with a respective instance of a nucleic acid of the first strand, or (ii) a label was not detected by a respective sensor associated with a respective instance of a nucleic acid of the first strand; (b) determining a plurality of candidate sequences for the first nucleic acid strand based on the plurality of records, each of the plurality of candidate sequences predicting at least a portion of a nucleic acid sequence of the first nucleic acid strand; and (c) determining from among the plurality of candidate sequences as at least a portion of the nucleic acid sequence of the first strand of nucleic acid. and identifying a particular candidate sequence of the plurality of candidate sequences that is most likely to be correct.

[0010] The disclosed sequencing and error correction apparatus, systems, and methods promise potentially higher throughput, lower error rates, and longer read lengths compared to cluster-based approaches. [Brief description of the drawings]

[0011] The objects, features and advantages of the present disclosure will become readily apparent from the following description of specific embodiments taken in conjunction with the accompanying drawings. [Figure 1] FIG. 1 illustrates a portion of a magnetic sensor according to some embodiments. [Figure 2A] 4 illustrates the resistance of a magnetoresistive (MR) sensor that may be used in accordance with some embodiments. [Figure 2B] 4 illustrates the resistance of a magnetoresistive (MR) sensor that may be used in accordance with some embodiments. [Figure 3A] FIG. 3A illustrates a spin torque oscillator (STO) sensor that may be used in accordance with some embodiments. [Figure 3B] FIG. 3B shows the experimental response of STO under exemplary conditions. [Figure 3C] 1 shows a short nanosecond field pulse of STO that may be used according to some embodiments. [Figure 3D] 1 shows a short nanosecond field pulse of STO that may be used according to some embodiments. [Figure 4A] FIG. 4A shows a single sensor of a cluster sequencer used to sense a number N of clonally amplified DNA strands in its vicinity. [Figure 4B] FIG. 4B illustrates an exemplary plurality of S single-molecule sensors each used by a SMAS device to monitor a respective single-stranded DNA (ssDNA), according to some embodiments. [Figure 5A] FIG. 5A is a block diagram illustrating components of an exemplary SMAS device for nucleic acid sequencing, according to some embodiments. [Figure 5B] 1 illustrates a portion of an exemplary SMAS device for nucleic acid sequencing, according to some embodiments. [Figure 5C] 1 illustrates a portion of an exemplary SMAS device for nucleic acid sequencing, according to some embodiments. [Figure 5D] 1 illustrates a portion of an exemplary SMAS device for nucleic acid sequencing, according to some embodiments. [Figure 5E] FIG. 5E illustrates a square lattice (or grid) pattern of a sensor according to some embodiments. [Figure 6A] FIG. 6A shows a sensor, a coiled DNA strand, and a label according to some embodiments. [Figure 6B] FIG. 6B shows exemplary dimensions of a sensor, an elongated DNA strand, and a label, according to some embodiments. [Figure 7A] FIG. 7A illustrates an example geometry for estimating the fill limit of a sensor array of a SMAS device, according to some embodiments. [Figure 7B] FIG. 7B shows sensors of a SMAS device arranged in a square lattice, according to some embodiments. [Figure 8A] 1 illustrates sensors of a SMAS device arranged in a hexagonal pattern, according to some embodiments. [Figure 8B] 1 illustrates sensors of a SMAS device arranged in a hexagonal pattern, according to some embodiments. [Figure 9A] FIG. 9A shows an example geometry for estimating the fill limit of a sensor array of a SMAS device, according to some embodiments. [Figure 9B] FIG. 9B shows sensors of a SMAS device arranged in a hexagonal lattice, according to some embodiments. [Figure 10] FIG. 10 compares the density of an exemplary SMAS implementation with state-of-the-art cluster sequencers. [Figure 11] FIG. 11 illustrates an exemplary method for sequencing multiple nucleic acid strands using a SMAS device, according to some embodiments. [Figure 12] FIG. 12 is a flow diagram of a sequencing procedure using an additive approach, according to some embodiments. [Figure 13] FIG. 13 is a diagram illustrating an additive sequencing protocol according to some embodiments. [Figure 14] FIG. 14 is a flow diagram of a sequencing procedure using a subtractive approach, according to some embodiments. [Figure 15] FIG. 15 is a diagram illustrating a subtractive sequencing protocol according to some embodiments. [Figure 16]FIG. 16 is a flow diagram of a sequencing procedure using a modified additive approach, according to some embodiments. [Figure 17] FIG. 17 illustrates a modified additive sequencing protocol according to some embodiments. [Figure 18A] FIG. 18A shows failed nucleotide incorporation (FNI) for a cluster sequencer. [Figure 18B] FIG. 18B shows FNI for the SMAS device. [Figure 18C] FIG. 18C shows failure to remove labels (FLR) for a cluster sequencer. [Figure 18D] FIG. 18D shows the FLR for the SMAS device. [Figure 18E] FIG. 18E shows failure to remove nucleotides (FNR) for the cluster sequencer. [Figure 18F] FIG. 18F shows the FNR for the SMAS device. [Figure 18G] FIG. 18G shows a failure to detect (FLD) nucleotide for a cluster sequencer. [Figure 18H] FIG. 18H shows the FLD for the SMAS device. [Figure 19] FIG. 19 is a flow diagram of an exemplary sequencing procedure using a modified additive approach with FLR and FNI error detection, according to some embodiments. [Figure 20] FIG. 20 shows an example of a record with FNI and FLR errors. [Figure 21] FIG. 21 shows the expected signal levels detected by a cluster sequencer sensor that captures the behavior of molecular ensembles during the sequencing procedure. [Figure 22] FIG. 22 illustrates how a SMAS device provides better accuracy when using error correction techniques, according to some embodiments. [Diagram 23]FIG. 13 illustrates the correction of FNI errors by deleting a run of four "no label detected" entries in a record of detection results from a sequencing procedure, according to some embodiments. [Figure 24] FIG. 24 shows the results of an exemplary SBS reaction according to some embodiments. [Diagram 25] FIG. 25 shows the effect of large cluster size on the base calling accuracy of a cluster sequencer. [Figure 26-1] FIG. 26 illustrates deterministic error correction of FLR and FNI errors according to some embodiments. [Figure 26-2] FIG. 26 illustrates deterministic error correction of FLR and FNI errors according to some embodiments. [Figure 27] FIG. 27 shows the FNI, FLR, and FNR errors of the detection data. [Figure 28] FIG. 28 shows FLR error correction and base calling from data generated by a SMAS device, according to some embodiments. [Figure 29] FIG. 29 shows FNI error correction and base calling from data generated by a SMAS device, according to some embodiments. [Figure 30-1] FIG. 30 is a diagram illustrating error correction and base calling from data generated by a SMAS device, according to some embodiments. [Figure 30-2] FIG. 30 is a diagram illustrating error correction and base calling from data generated by a SMAS device, according to some embodiments. [Diagram 31] FIG. 31 shows the FNI, FLR, FNR, and FLD errors in an exemplary detection result from a SMAS device. [Figure 32-1] FIG. 32 is a diagram illustrating the application of error correction procedures to data captured during SBS by a SMAS device, according to some embodiments. [Figure 32-2] FIG. 32 is a diagram illustrating the application of error correction procedures to data captured during SBS by a SMAS device, according to some embodiments. [Diagram 33] FIG. 33 is a flow diagram illustrating an error correction procedure according to some embodiments. [Figure 34A] Figure 34A shows the average signal intensity at the interrogation step, where a matching nucleotide is introduced and, from successful incorporation, a label is detected. [Figure 34B] FIG. 34B shows the function fit to the measured intensities from the cluster model. [Diagram 35] FIG. 35 is a plot of the probability function of the cluster sequencer. [Diagram 36] FIG. 36 shows the discrete probability functions of the cluster sequencer. [Figure 37A] FIG. 37A shows the intensity plot of the cluster sequencer. [Figure 37B] FIG. 37B shows the probability distribution function of the cluster sequencer. [Figure 38A] Plot the probability function of the cluster sequencer. [Figure 38B] Plot the probability function of the cluster sequencer. [Figure 39] FIG. 39 shows the Nr parameter space of the cluster sequencer under various conditions. [Figure 40A] FIG. 40A shows the calculated probability of the cluster sequencer along the Q30 contour for various combinations Nr. [Figure 40B] FIG. 40B plots the cumulative error probability calculated for the cluster sequencer. [Diagram 41] Figure 41 shows that the cumulative probability of an incorrect base call at position 150 is less than 1 in 100.

number

number

number

number

[0012] For ease of understanding, the same reference numbers have been used, whenever possible, to designate identical elements common to the figures. It is contemplated that elements disclosed in one embodiment can be beneficially utilized in other embodiments without specific recitation. Moreover, a description of an element in the context of one drawing is also applicable to other drawings showing that element. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0013] Detailed Description Although some of the descriptions and examples herein relate to DNA sequencing, it should be understood that the present disclosure applies generally to nucleic acid sequencing.

[0014] Terminology and Notation As used herein, the term "strand" refers to a single nucleic acid strand (e.g., ssDNA). The terms "strand" and "fragment" are used interchangeably when referring to nucleic acids.

[0015] As used herein, the term "plurality" means two or more, but not necessarily all. Thus, a plurality of sensors means only at least two sensors, but not necessarily all sensors in a sensor array or sequencing device / system. Similarly, a plurality of binding sites means only at least two binding sites, but not necessarily all binding sites in a sequencing device / system.

[0016] As used herein, the term "instance" when referring to a nucleic acid strand refers to a template nucleic acid strand or a copy thereof (e.g., produced by an amplification or replication process). Ideally, a copy of a template nucleic acid strand is identical to the template strand, but as known in the art, the copies are not necessarily identical due to replication / amplification errors. It is understood that a replica produced by amplification is still considered a copy of the original nucleic acid strand, even if the amplification procedure introduces errors. Thus, all instances of a strand are ideally identical to each other, but may not be.

[0017] As used herein, the term "interrogation cycle" refers to one cycle of a nucleic acid sequencing procedure during which all possible nucleotides are introduced to determine which, if any, nucleotides are incorporated into the strand being sequenced. For example, in a DNA sequencing procedure, adenine (A), thymine (T), cytosine (C), and guanine (G) are all tested in some (arbitrary) order (these do not have to be the same for each interrogation cycle). As described in more detail below, depending on the sequencing procedure selected, more than one label per strand may be detected during one sequencing cycle.

[0018] As used herein, the term "interrogation step" refers to a step or set of steps in a sequencing procedure during which it is determined whether one or more sensors of a sequencing device detect a label. For DNA sequencing cycling through A, T, C, and G, there are four interrogation steps per interrogation cycle (one for each nucleotide). For a sensor in use, each interrogation step results in a single determination of whether the sensor detects a label or not.

[0019] As used herein, the term "detection result" refers to a value that indicates either (a) that a label was detected during the interrogation step, or (b) that a label was not detected during the interrogation step. In some embodiments, the detection result is a binary value (e.g., 0 or 1). The detection result may be derived from other data (e.g., a signal representing resistance, frequency, intensity, etc.; a measurement of resistance, frequency, intensity, etc.).

[0020] As used herein, the term "record" refers to a stored representation of a single sensor's detection results. If the selected sequencing procedure has M interrogation steps, then upon completion of the sequencing procedure, each record will have M detection results. The S sensor records may be stored in a single file (e.g., as a table with S rows and M columns, or S columns and M rows), or a separate file may be created for each sensor record.

[0021] As used herein with respect to detection results contained within a record, the term "run" means a sequence of consecutive identical values.

[0022] The terms "sensor" and "sensing element" are used interchangeably herein.

[0023] The variable S is used herein to refer to a number of sensors in a plurality of sensors. The S sensors may be sensing instances of the same chain or may be sensing instances of different chains.

[0024] The variable K is used herein to refer to a number of sensors in a plurality of sensors that all sense an instance of the same strand.

[0025] sign The nucleic acid sequencing methods described herein use labeled nucleotide precursors that contain cleavable labels. These cleavable labels can be, for example, magnetic, fluorescent, organometallic, or charged molecules. could be.

[0026] Each label may comprise, for example, a magnetic nanoparticle, such as a molecule, a superparamagnetic nanoparticle, or a ferromagnetic particle. The magnetic labels may be nanoparticles with high magnetic anisotropy. Examples of nanoparticles with high magnetic anisotropy include, but are not limited to, Fe3O4, FePt, FePd, and CoPt. To facilitate chemical binding to nucleotides, the particles may be synthesized and coated with SiO2. See, for example, M. Aslam, L. Fu, S. Li, and VPDravid, "Silica encapsulation and magnetic properties of FePt nanoparticles," Journal of Colloid and Interface Science, Volume 290, Issue 2, 15 October 2005, pp. 444-449. Because magnetic labels of this size have a permanent magnetic moment whose orientation fluctuates randomly on a very short time scale, some embodiments described further below rely on a highly sensitive sensing scheme to detect magnetic field fluctuations caused by the presence of the magnetic label.

[0027] Each label may, for example, comprise a fluorophore. Fluorescent labels are well known in the art and are suitable for use in the present disclosure.

[0028] The label may include, for example, an organometallic compound. As will be understood, an organometallic compound is any member of a class of substances that contains at least one metal-carbon bond in which the carbon is part of an organic group. Organometallic compounds include, for example, Gillman's reagent (including lithium, copper), Grinar's reagent (including magnesium), tetracarbonylnickel, ferrocene (including transition metals), organolithium compounds (e.g., n-butyllithium (n-BuLi)), organozinc compounds (e.g., diethylzinc (Et2Zn)), organotin compounds (e.g., tributyltin hydride (Bu3SnH)), organoborane compounds (e.g., triethylborane (Et3B)), organoaluminum compounds (e.g., trimethylaluminum (Me3Al)), and the like.

[0029] The label may include, for example, a charged molecule.

[0030] There are several ways to attach labels to nucleotide precursors and cleave the labels after incorporation of the nucleotide precursor. For example, labels may be attached to bases, in which case they may be chemically cleaved. As another example, labels may be attached to phosphates, in which case they may be cleaved by a polymerase, or, if attached via a linker, may be cleaved by cleaving the linker.

[0031] In some embodiments, the label is linked to a nitrogenous base (e.g., A, C, T, G, or derivatives) of the nucleotide precursor. After incorporation of the nucleotide precursor and detection by a sequencing instrument (e.g., as described in more detail below), the label is cleaved from the incorporated nucleotide.

[0032] In some embodiments, the label is attached via a cleavable linker. Cleavable linkers are known in the art and are described, for example, in U.S. Pat. No. 7,057,026, U.S. Pat. No. 7,414,116, and continuations and improvements thereof. In some embodiments, the label is attached to the 5-position of the pyrimidine or the 7-position of the purine via a linker that includes an allyl or azide group. In other embodiments, the linker includes a disulfide, an indole, or a Sieber group. The linker includes an alkyl (C 1~6 ) or alkoxy(C 1~6 ), nitro, cyano, fluoro groups or groups having similar properties. Linkers can be cleaved by water-soluble phosphine or phosphine-based transition metal-containing catalysts. Other linkers and linker cleavage mechanisms are known in the art. For example, linkers containing trityl, p-alkoxybenzyl esters and p-alkoxybenzyl amides, as well as tert-butyloxycarbonyl (Boc) groups and acetal systems can be cleaved under acidic conditions by proton releasing cleavage agents. Thioacetal or other sulfur-containing linkers can be cleaved using thiophilic metals such as nickel, silver, or mercury. Cleavage protecting groups can also be considered to prepare suitable linker molecules. Ester and disulfide-containing linkers can be cleaved under reductive conditions. Linkers containing triisopropylsilane (TIPS) or t-butyldimethylsilane (TBDMS) can be cleaved in the presence of F ions. Photocleavable linkers that are cleaved by wavelengths that do not affect other components of the reaction mixture include linkers containing O-nitrobenzyl groups. Linkers containing benzyloxycarbonyl groups can be cleaved by Pd-based catalysis.

[0033] In some embodiments, the nucleotide precursor comprises a label attached to a polyphosphate moiety, for example, as described in U.S. Pat. Nos. 7,405,281 and 8,058,031. Briefly, the nucleotide precursor comprises a nucleoside moiety and a chain of three or more phosphate groups, in which one or more of the oxygen atoms are optionally replaced, for example, by S. The label can be attached to the α, β, γ or more phosphate groups (if present) directly or via a linker. In some embodiments, the label is attached to the phosphate group via a non-covalent linker, for example, as described in U.S. Pat. No. 8,252,910. In some embodiments, the linker is a hydrocarbon selected from substituted or unsubstituted alkyl, substituted or unsubstituted heteroalkyl, substituted or unsubstituted aryl, substituted or unsubstituted heteroaryl, substituted or unsubstituted cycloalkyl, and substituted or unsubstituted heterocycloalkyl, see, for example, U.S. Pat. No. 8,367,813. The linker may also comprise a nucleic acid chain. See, for example, U.S. Pat. No. 9,464,107.

[0034] In embodiments where the label is linked to a phosphate group, the nucleotide precursor is incorporated into the nascent strand by a nucleic acid polymerase that also cleaves and releases the detectable label. In some embodiments, the label is removed by cleaving the linker, e.g., as described in U.S. Pat. No. 9,587,275.

[0035] In some embodiments, the nucleotide precursor is a non-extendable "terminator" nucleotide, i.e., a nucleotide having a 3'-terminus blocked from addition of the next nucleotide by a blocking "terminator" group. The blocking group is a reversible terminator that can be removed to continue the chain synthesis process described herein. The attachment of removable blocking groups to nucleotide precursors is known in the art. See, for example, U.S. Pat. Nos. 7,541,444, 8,071,739, and continuations and improvements thereof. Briefly, the blocking group can include an allyl group that can be cleaved by reaction with a metal-allyl complex in the presence of a phosphine or nitrogen-phosphine ligand in aqueous solution. Other examples of reversible terminator nucleotides for use in sequencing by synthesis include modified nucleotides described in International Application No. PCT / US2019 / 066670, entitled "3' Protected Nucleotides," filed on December 16, 2019, and published as WO 2020 / 131759.

[0036] Sensors The characteristics and capabilities of the sensors used in the nucleic acid sequencing devices, systems and methods described herein depend on the choice of label used. The sensors can be, for example, magnetic sensors (e.g. The sensor may be a magnetic sensor (e.g., for detecting magnetic nanoparticles, organometallic compounds, etc.) or an optical sensor (e.g., for detecting fluorophores). It is understood that other types of sensors may be suitable for detecting various types of labels, and the examples described herein are not intended to be limiting. Generally speaking, the disclosed devices, systems, and methods may use any type of label that can be detected by a selected type of sensor, and conversely, the disclosed devices, systems, and methods may use any type of sensor that can detect the presence (and absence) of a selected type of label.

[0037] Reference number 105 is used herein generally for single molecule sensors, regardless of their type (and regardless of the type of label they detect), and reference number 15 is used for sensors that sense clusters of nucleic acid strands.

[0038] Magnetic Sensor Some embodiments disclosed herein use a magnetic sensor to detect the presence of a magnetic label (e.g., a magnetic nanoparticle, an organometallic complex, a charged molecule, etc.) attached to a nucleotide precursor. FIG. 1 shows a portion of a magnetic sensor 105 according to some embodiments. The exemplary magnetic sensor 105 of FIG. 1 has a bottom surface 108 and a top surface 109 and comprises three layers, e.g., two ferromagnetic layers 106A, 106B separated by a non-magnetic spacer layer 107. The non-magnetic spacer layer 107 may be a metallic material, e.g., copper or silver, in which case the structure is called a spin valve (SV), or an insulator, e.g., alumina or magnesium oxide, in which case the structure is called a magnetic tunnel junction (MTJ). Suitable materials for use in the ferromagnetic layers 106A, 106B include, e.g., alloys of Co, Ni, and Fe (sometimes mixed with other elements). In some embodiments, the ferromagnetic layers 106A, 106B are designed such that their magnetic moments are oriented in the plane of the film or perpendicular to the plane of the film. 1 for purposes such as interface smoothing, texturing, and protection from processes used to pattern the device in which the sensor 105 is incorporated; however, the active area of ​​the magnetic sensor 105 is in this three-layer structure. Thus, components in contact with the magnetic sensor 105 may be in contact with one of the three layers 106A, 106B, or 107, or may be in contact with another portion of the magnetic sensor 105.

[0039] As shown in Figures 2A and 2B, the resistance of the MR sensor is proportional to 1-cos(θ), where θ is the angle between the moments of the two ferromagnetic layers 106A, 106B shown in Figure 1. To maximize the signal generated by the magnetic field and provide a linear response of the magnetic sensor 105 to the applied magnetic field, the magnetic sensor 105 may be designed such that in the absence of a magnetic field, the moments of the two ferromagnetic layers 106A, 106B are oriented at π / 2 radians or 90 degrees relative to each other. This orientation can be achieved by any number of methods known in the art. For example, one solution is to use an antiferromagnet to "pin" the magnetization direction of one of the ferromagnetic layers (either 106A or 106B, referred to as "FM1") by an effect called exchange bias, and then coat the sensor with a bilayer with an insulating layer and a permanent magnet. The insulating layer avoids electrical shorting of the magnetic sensor 105, and the permanent magnet provides a "hard bias" magnetic field perpendicular to the pinning direction of FM1, which then rotates the second ferromagnet (either 106B or 106A, referred to as "FM2") to produce the desired configuration. A magnetic field parallel to FM1 then rotates FM2 around this 90 degree configuration, and the change in resistance results in a voltage signal that can be calibrated to measure the magnetic field acting on the magnetic sensor 105. In this way, the magnetic sensor 105 functions as a magnetic field-voltage transducer.

[0040] The example discussed immediately above assumes that the moments are oriented in the plane of the film at 90 degrees to each other. It should be noted that although the use of oriented ferromagnetic layers has been described, alternatively, a perpendicular configuration can be achieved by orienting the moment of one of the ferromagnetic layers 106A, 106B out of the plane of the film, which can be achieved using what is called perpendicular magnetic anisotropy (PMA).

[0041] In some embodiments, the magnetic sensor 105 uses a quantum mechanical effect known as spin transfer torque. In such devices, a current passing through one ferromagnetic layer 106A (or 106B) in a SV or MTJ preferentially transmits electrons with spins parallel to the moment of the layer, while electrons with spins antiparallel are more likely to be reflected. In this way, the current is spin polarized such that there are more electrons of one spin type than the other. This spin polarized current then interacts with the second ferromagnetic layer 106B (or 106A) and exerts a torque on the moment of the layer. This torque can, in different circumstances, cause the moment of the second ferromagnetic layer 106B (or 106A) to precess around the effective magnetic field acting on the ferromagnetic material, or cause the moment to reversibly switch between two orientations defined by the uniaxial anisotropy induced in the system. The resulting spin torque oscillators (STOs) are frequency tunable by varying the magnetic field acting on them. Therefore, they have the ability to act as a magnetic field vs. frequency (or phase) transducer (thereby generating an AC signal with frequency), as shown in FIG. 3A, which illustrates the concept of using an STO sensor. FIG. 3B shows the experimental response of an STO through a delay detection circuit when an AC magnetic field with a frequency of 1 GHz and a peak-to-peak amplitude of 5 mT is applied across the STO. This result, as well as the results shown in FIGS. 3C and 3D for short nanosecond magnetic field pulses, show how these oscillators can be used as nanoscale magnetic field detectors. Further details can be found in T. Nagasawa, H. Suto, K. Kudo, T. Yang, K. Mizushima, and R. Sato, ``Delay detection of frequency modulation signal from a spin-torque oscillator under a nanosecond-pulsed magnetic field. field,''Journal of Applied Physics,Vol.111,07C908(2012).

[0042] Optical Sensors Some nucleic acid sequencing approaches use fluorescent labels. In such approaches, the nucleic acid molecules being sequenced are immobilized on a solid support and the binding of fluorescently labeled target molecules (e.g., nucleotides) to the molecules is monitored. An optical instrument, such as an excitation and reading device for fluorescence, provides light of a specific wavelength to excite the fluorescent labels and detects the fluorescence from the labels emitted at a slightly different wavelength. Since the beam path of the excitation light must be at least partially different from the beam path of the fluorescence, spectral separation may be achieved using excitation and emission filters (whose spectra do not significantly overlap) and / or vertical or side illumination may be used.

[0043] Optical sensor and sequencing devices and methods that use fluorescent labels (eg, fluorophores) are well known in the art.

[0044] Amplification / Replication Nucleic acid sequencers generally rely on an amplification (or replication) process to generate multiple nucleic acid instances from a single nucleic acid strand (e.g., instances of single-stranded DNA strands (ssDNA) from one single DNA molecule). Polymerase chain reaction (PCR) is a well-known method for amplifying double-stranded DNA that allows for the replication of significant amounts of DNA from a small initial amount.

[0045] Cluster Sequencing Instrument Some sequencing devices, referred to herein as cluster (CLUS) devices, use amplification techniques to form localized clusters of many DNA strands. For example, one single-stranded DNA is used as a template, and PCR amplification generates thousands or millions of instances of the DNA sequence in a localized area. By immobilizing at least a portion of the PCR primers to a solid support, the generated DNA molecules can be immobilized in localized clusters to form identifiable "clones." The generated DNA clusters can include ssDNA. Clonal amplification techniques include, for example, bridge PCR, emulsion PCR, including bead-based emulsion PCR. For bridge amplification, single DNA molecules are amplified to form DNA clusters by in situ PCR using primers attached to a solid surface, such as a glass slide. Each DNA cluster is a physically separated "clone" consisting of an instance of a DNA strand. In emulsion PCR-based clonal amplification, single DNA molecules are clonally amplified in emulsion droplets. In some methods, the DNA strands are attached to microbeads within the droplets. Clonal amplification of single molecules can also be performed in separate microwells.

[0046] As used herein, the term "cluster" refers to a localized cluster of nucleic acid strands, ideally with identical sequences, generated from clonal amplification. If the nucleic acid is DNA, the cluster comprises (ideally) identical DNA strands (or fragments) bound to a solid support. For example, the cluster can be generated on a spot on a glass slide, or can be bound to a microbead, microwell, or other microparticle.

[0047] The use of CLUS instruments for fluorescence-based DNA sequencing is well known.

[0048] A sequencing device using an array of magnetic sensors for cluster-based nucleic acid sequencing is described, for example, in PCT Application No. PCT / US2021 / 021274, entitled "Magnetic Sensor Arrays for Nucleic Acid Sequencing, and Methods of Making and Using the Same," filed on March 7, 2021 (Attorney Docket No. ROA-1001-WO / P35967-WO).

[0049] FIG. 4A shows a single sensor 15 of a CLUS device used to sense a number N of clonally amplified DNA strands 101 in its vicinity. Sensor 15 can be, for example, a magnetic sensor for sensing magnetic labels attached to incorporated nucleotides. For convenience, FIG. 4A shows strand 101 in contact with sensor 15, but it should be understood that there can be a barrier (e.g., an insulating layer) between sensor 15 and strand 100. Sensor 15 can be, for example, a magnetic sensor as described in PCT Application No. PCT / US2021 / 021274, supra.

[0050] State-of-the-art commercially available CLUS devices, such as those that sense fluorescent labels, can use hundreds of millions of sensors 15, each sensing many instances of each amplified DNA strand 101. One drawback of some CLUS devices is that achieving optimal cluster density can be critical for high-quality sequencing. Specifically, the use of large clusters tends to provide higher data quality but lower data output, while the use of small clusters can result in run failures, poor run performance, lower Q30 scores, introduction of sequencing artifacts, and lower total data output. To mitigate these issues, newer CLUS devices use patterned flow cells with separate nanocells for cluster generation. These nanocells are organized in a hexagonal arrangement to more efficiently utilize the flow cell surface area.

[0051] Single Molecule Array Sequencing Instrument Single molecule array sequencing devices (referred to herein as "SMAS devices") are an alternative to CLUS devices. In contrast to CLUS devices, which sense and sequence localized clusters of multiple instances of a single nucleic acid strand, SMAS devices use sensors that sense and sequence individual strands of nucleic acid individually. Generally speaking, in SMAS devices, the sensors do not sense more than one physical nucleic acid strand, but different sensors sense instances of the same strand. In other words, there are multiple instances of a nucleic acid strand, but each sensed strand is sensed by a different respective sensor. Depending on the amplification technique used, the individual strands may be randomly distributed throughout the fluid chamber of the SMAS device, or may be located in a more localized area. As will be further described below, the location of an instance of a particular strand can be identified, and an error correction procedure can be applied to the detection results corresponding to the instance before calling the bases to improve sequencing accuracy compared to CLUS devices. Furthermore, compared to CLUS devices, for a reasonable chemical failure rate, SMAS devices require fewer instances of each nucleic acid strand to be sequenced to achieve accurate sequencing results.

[0052] FIG. 4B illustrates an exemplary plurality of S single molecule sensors 105 each used by the SMAS device to monitor a respective single stranded DNA (ssDNA) 101. Each of the plurality of S sensors 105 may be, for example, a magnetic sensor, an optical sensor, or the like. FIG. 4B illustrates five single molecule sensors 105A, 105B, 105C, 105D, and 105E, each sensing a respective DNA strand 101 (which may be instances of the same DNA strand or instances of different DNA strands). Each sensor 105 may be, for example, a nanoscale sensor so small that only a single DNA strand 101 can bind to a binding site associated with the sensor 105. (For convenience, FIG. 4B illustrates strand 101 in contact with sensor 105, but as described further below, in some embodiments, strand 100 is bound to an individual binding site, each of which is associated with a respective sensor 105.)

[0053] Consider clonally amplified DNA bound to a solid surface containing an array of densely packed sensors 105, as shown in FIG. 4B. The DNA can be replicated by solid-phase amplification (SPA) to create clusters of monoclonal DNA, with each strand sensed by a different sensor 105, or the DNA can be amplified in bulk and then immobilized on the surface of the SMAS device. If the DNA is amplified (e.g., by SPA) on the surface of the fluid chamber of the SMAS device, sensors 105A, 105B, 105C, 105D, 105E can sense instances of the clonal DNA. Alternatively, if the DNA is amplified in the bulk off the device and added to the fluid chamber of the SMAS device, the amplified DNA strands 101 can be more randomly distributed among the sensors 105.

[0054] 5A is a block diagram illustrating components of an exemplary SMAS device 100 for nucleic acid sequencing, according to some embodiments. As shown, device 100 comprises a sensor array 110 coupled to circuitry 120 coupled to at least one processor 130. Sensor array 110 comprises a plurality of sensors 105 (e.g., magnetic, optical, etc.), which may be arranged in any suitable manner, as described further below. The characteristics and properties of the sensors 105 in sensor array 110 depend on the type of label used for sequencing.

[0055] The circuitry 120 may, for example, comprise one or more lines that allow the sensors 105 in the sensor array 110 to be interrogated by at least one processor 130 (e.g., with the aid of other components known in the art, such as current sources). For example, during operation, the processor(s) 130 may cause the circuitry 120 to apply current to such lines to detect a characteristic of at least one of the multiple sensors 105 in the sensor array 110, the characteristic being the presence or absence of a label within range of the sensor 105. In other words, the characteristic (e.g., resistance, frequency, voltage, signal level, etc.) is indicative of whether the sensor 105 has detected at least one indicator or has not detected an indicator. For example, the at least one processor 130 may evaluate a value of the characteristic (e.g., frequency, wavelength, magnetic field, resistance, noise level, intensity, color of light, etc.) and determine that an indicator has been detected (or not) based on a comparison of the value of the characteristic to a threshold or baseline value (e.g., by determining whether the value of the characteristic of the sensor 105 meets or exceeds the threshold). As another example, the at least one processor 130 may compare the acquired characteristic of the sensor 105 to a previously detected value of the characteristic (e.g., a baseline value of the sensor 105) and base the determination of whether an indicator has been detected on a change in the value of the characteristic (e.g., a change in the magnetic field, resistance, noise level, frequency, wavelength, intensity, color of light, etc.). For example, as described further below in the discussion of Figure 19, at least one processor 130 can evaluate a characteristic obtained from a sensor 105 to detect whether a sensor 105 that detected a label during a first interrogation step of a sequencing procedure still detects the label after a cleavage step that should have removed the label. Similarly, at least one processor 130 can evaluate a change in the characteristic between interrogation steps to determine whether a sensor 105 (a) did not detect the label during either interrogation step, (b) detected the label during both interrogation steps, (c) did not detect the label during the first interrogation step but detected the label during a subsequent interrogation step, and / or (d) detected the label during the first interrogation step but did not detect the label during a subsequent interrogation step.

[0056] The detected characteristic depends on the type of label used in the sequencing procedure. The label may be, for example, fluorescent, in which case the sensor 105 may be, for example, an optical sensor capable of detecting the wavelength, frequency, modulation frequency, color, or intensity of light emitted by the fluorescent label. Optical sensors suitable for detecting fluorescent labels are well known in the art. When the label used in the nucleic acid sequencing procedure is fluorescent, in some embodiments, the circuit 120 enables the at least one processor 130 to detect deviations or variations in the light (or electromagnetic energy) detected by some or all of the sensors 105 in the sensor array 110.

[0057] The label may be, for example, magnetic (e.g., magnetic nanoparticles, organometallic compounds, charged molecules, etc.), in which case the sensor 105 may be a magnetic sensor capable of detecting magnetic properties. Magnetic sensors are described in the applicant's previously filed patent applications, including, for example, PCT Application No. PCT / US20 / 27290, filed April 8, 2020, and published October 15, 2020 as International Publication No. WO 2020 / 210370, entitled "Nucleic Acid Sequencing by Synthesis Using Magnetic Sensor Arrays" (Attorney Docket No. ROA-1000-WO / P35097-WO). In some embodiments in which the label is magnetic, the sensor 105 is a magnetoresistive (MR) sensor capable of detecting, for example, a magnetic field or resistance, a change in a magnetic field or a change in resistance, or a noise level. In some embodiments, each of the sensors 105 of the sensor array 110 is a thin-film device that uses the MR effect to detect a magnetic label attached to a nucleotide incorporated into a single strand of a nucleic acid bound to a respective binding site. The sensor 105 can operate as a potentiometer having a resistance that changes as the strength and / or direction of the sensed magnetic field changes. In some embodiments that use magnetic labels, the sensor 105 comprises a magnetic oscillator (e.g., a spin torque oscillator (STO)) and the characteristic indicative of whether at least one label has been detected is the frequency of a signal, or a change in the frequency of the signal, associated with or generated by the magnetic oscillator.

[0058] If the labels used in the nucleic acid sequencing procedure are magnetic, in some embodiments, at least one processor 130, with the aid of circuitry 120, detects deviations or variations in the magnetic environment of some or all of the sensors 105 in the sensor array 110. For example, An MR-type sensor 105 in the absence should have relatively small noise above a certain frequency compared to a sensor 105 in the presence of a magnetic label, since magnetic field fluctuations from the magnetic label cause fluctuations in the moment of the sensing ferromagnet. These fluctuations can be measured using heterodyne detection (e.g., by measuring the noise power density) or by directly measuring the voltage of the sensor 105 and can be evaluated using a comparator circuit to compare it with another sensor element that does not sense the binding site. If the sensor 105 includes an STO element, the fluctuating magnetic field from the magnetic label causes a jump in the phase of the sensor 105 due to an instantaneous change in frequency, which can be detected using a phase detection circuit. Another option is to design it so that the STO only oscillates within a small magnetic field range, so that the presence of the magnetic label turns off the oscillation.

[0059] It should be understood that the examples of labels and sensors 105 provided above are merely illustrative. In general, any type of label capable of labeling a nucleotide precursor can be used with an array 110 of any type of sensor 105 capable of detecting that type of label.

[0060] Figures 5B, 5C, and 5D show a portion of an exemplary SMAS device 100 for nucleic acid sequencing, according to some embodiments. The exemplary SMAS device 100 uses magnetic labels and a magnetic sensor 105. Figure 5B is a top view of device 100. Figure 5C is a cross-sectional view at the dashed line marked "5C" in Figure 5B, and Figure 5D is a cross-sectional view at the dashed line marked "5D" in Figure 5B.

[0061] The exemplary device 100 shown in Figures 5B, 5C, and 5D includes a sensor array 110 for sensing magnetic signatures in a fluid chamber 115. The sensor array 110 includes a plurality of magnetic sensors 105, with 16 sensors 105 shown in the array 110 in Figure 5B. It should be understood that implementations of the SMAS device 100 can include any number of sensors 105 (e.g., hundreds, thousands, or millions of sensors 105). To avoid obscuring the drawing, only seven of the sensors 105 are labeled in Figure 5B, namely sensors 105A, 105B, 105C, 105D, 105E, 105F, and 105G. As explained above, the magnetic sensors 105 detect the presence or absence of magnetic signatures. That is, each magnetic sensor 105 detects whether at least one magnetic signature is present in its vicinity.

[0062] 5C and 5D in conjunction with FIG. 5B, each sensor 105 is shown in the exemplary embodiment of the device 100 as having a cylindrical shape. However, it should be understood that in general, the sensors 105 can have any suitable shape. For example, the sensors 105 can be three-dimensional rectangular prisms. Furthermore, different sensors 105 can have different shapes (e.g., some can be rectangular prisms and other can be cylindrical, etc.). It should be understood that the drawings are merely exemplary.

[0063] As shown in Figures 5C and 5D, the device 100 comprises a fluid chamber 115. The fluid chamber 115 comprises a plurality of binding sites 116 (e.g., S binding sites 116). In some embodiments, the fluid chamber 115 holds fluids (e.g., nucleotide precursors and other fluids) used during a nucleic acid sequencing procedure. However, it should be understood that embodiments in which the fluid chamber 115 does not hold fluids are contemplated and are within the scope of the present disclosure. For example, the binding sites 116 may be disposed on a removable (or movable) part (e.g., a panel, plate, slide, etc.) that may be immersed in reagents and other fluids after the nucleic acid strands bind to the binding sites 116, and may then be positioned such that the sensor 105 can detect the label. Thus, although the name of the fluid chamber 115 suggests that it holds a fluid, it is not required that the fluid chamber 115 hold a fluid.

[0064] As shown in Figures 5B, 5C, and 5D, each of the sensors 105 is associated with a respective binding site 116. (For simplicity, in this document, the binding sites are generally labeled with the reference number 116. Individual binding sites are labeled with the reference number 116 followed by a letter.) In other words, there is a one-to-one relationship between the sensors 105 and the binding sites 116. As shown in Figure 5B, sensor 105A is associated with binding site 116A, sensor 105B is associated with binding site 116B, sensor 105C is associated with binding site 116C, sensor 105D is associated with binding site 116D, sensor 105E is associated with binding site 116E, sensor 105F is associated with binding site 116F, and sensor 105G is associated with binding site 116G. Each of the other unlabeled sensors 105 shown in Figure 5B is also associated with a respective binding site 116. 5B, 5C, and 5D, each sensor 105 is shown disposed beneath its respective binding site 116, however, it should be understood that the binding sites 116 may be in other locations relative to their respective sensors 105. For example, the binding sites 116 may be on the sides of the respective sensors 105.

[0065] Each of the binding sites 116 is configured to bind one or less strands of nucleic acid (e.g., ssDNA) to the SMAS device 100 within the fluid chamber 115. In other words, each binding site 116 has properties and / or characteristics that allow for single-stranded and only single-stranded nucleic acid to be bound for sensing (and sequencing) by the respective sensor 105. Each sensor 105 can then detect a label attached to a nucleotide incorporated into the nucleic acid strand bound to the binding site 116 during a nucleic acid sequencing procedure, as discussed further below. In some embodiments, the binding site 116 has a structure (or structures) configured to immobilize the nucleic acid to the binding site 116. For example, the structure (or structures) can comprise a cavity or a ridge. Although FIGS. 5C and 5D show the binding site 116 as extending from a surface of the fluid chamber 115, it should be understood that the binding site 116 may be flush with or etched into the surface of the fluid chamber 115.

[0066] The binding sites 116 can have any suitable size and shape that facilitates the attachment of single-stranded and only single-stranded nucleic acids to each binding site 116. For example, the shape of the binding sites can be similar to or identical to the shape of the sensor 105 (e.g., if the sensor 105 is cylindrical in three dimensions, the binding sites 116 can be cylindrical and protrude from the surface of the fluid chamber 115 or form a fluid reservoir within the surface of the fluid chamber 115, with a radius that can be larger, smaller, or the same size as the radius of the respective sensor 105; if the sensor 105 is a rectangular prism in three dimensions, the binding sites 116 can be rectangular prisms with a surface 116 that is larger, smaller, or the same size as the nearest portion of the sensor 105, etc.). In general, the binding sites 116 and the surface of the fluid chamber 115 can have any shape and characteristics that facilitate the attachment of a single nucleic acid strand to each binding site 116 and allow the sensor 105 to detect a label attached to the nucleotide incorporated at each binding site 116.

[0067] 5C and 5D show a sealed fluid chamber 115 with a top extending in the xy plane, the fluid chamber 115 need not be sealed. In some embodiments, the surface of the fluid chamber 115 has properties and characteristics that protect the sensor 105 from any fluid present in the fluid chamber 115, but still allow the nucleic acid strand to bind to the binding site 116, and the sensor 105 detects the label attached to the nucleotide incorporated in the nucleic acid strand bound to the binding site 116. The material of the fluid chamber 115 (and possibly the material of the binding site 116) may be or include an insulator. In some embodiments, the surface of the fluid chamber 115 comprises an organic polymer, a metal, or a silicate. The fluid chamber 115 may comprise, for example, a metal oxide, silicon dioxide, polypropylene, gold, glass, or silicon. The thickness of the surface of the fluid chamber 115 may be such that the sensor 105 can detect the label attached to the nucleotide incorporated in the nucleic acid strand bound to the binding site 116 in the fluid chamber 115. The thickness of the surface can be selected to allow detection of magnetic labels bound to the sensor. In some embodiments, the surface is about 3-20 nm thick, such that each sensor 105 is about 5 nm to about 50 nm from any labels bound to nucleotides incorporated into the nucleic acid strand bound to the respective binding site 116 of the sensor 105. It should be understood that these values ​​are merely exemplary. It will be understood that an implementation can have a fluid chamber 115 with a thicker or thinner surface.

[0068] The circuit 120 of the device 100 may include one or more lines 125. In some embodiments, each of the multiple sensors 105 is coupled to at least one line 125. In the example shown in Figures 5B, 5C, and 5D, the device 100 includes eight lines 125A, 125B, 125C, 125D, 125E, 125F, 125G, and 125H. (For simplicity, this document generally refers to the row with reference number 125. Individual rows are given the reference number 125 followed by a letter.) A pair of lines 125 can be used to access (e.g., interrogate) each sensor 105. In the exemplary embodiment shown in Figures 5B, 5C, and 5D, each sensor 105 of the sensor array 110 is coupled to two lines 125. For example, sensor 105A is coupled to lines 125A and 125H; sensor 105B is coupled to lines 125B and 125H; sensor 105C is coupled to lines 125C and 125H; sensor 105D is coupled to lines 125D and 125H. Sensor 105E is coupled to lines 125D and 125E; sensor 105F is coupled to lines 125D and 125F; sensor 105G is coupled to lines 125D and 125G. In the exemplary embodiment of Figures 5B, 5C, and 5D, lines 125A, 125B, 125C, and 125D are shown to be below magnetic sensor 105, and lines 125E, 125F, 125G, and 125H are shown to be above magnetic sensor 105. Figure 5C shows sensor 105E associated with lines 125D and 125E, sensor 105F associated with lines 125D and 125F, sensor 105G associated with lines 125D and 125G, and sensor 105D associated with lines 125D and 125H. Figure 5D shows sensor 105D associated with lines 125D and 125H, sensor 105C associated with lines 125C and 125H, sensor 105B associated with lines 125B and 125H, and sensor 105A associated with lines 125A and 125H.

[0069] 5B, 5C, and 5D are arranged in a rectangular pattern in sensor array 110. (It should be understood that a square pattern is a special case of a rectangular pattern.) Each of the lines 125 identifies a row or column of sensor array 110. For example, each of lines 125A, 125B, 125C, and 125D identifies a different row of sensor array 110, and each of lines 125E, 125F, 125G, and 125H identifies a different column of sensor array 110. As shown in FIG. 5C, each of lines 125E, 125F, 125G, and 125H contacts one of the sensors 105 along the cross section (i.e., line 125E contacts the top of sensor 105E, line 125F contacts the top of sensor 105F, line 125G contacts the top of sensor 105G, and line 125H contacts the top of sensor 105D), and line 125D contacts the bottom of each of sensors 105E, 105F, 105G, and 105D. Similarly, as shown in FIG. 5D, lines 125A, 125B, 125C, and 125D each contact the bottom of one of the sensors 105 along the cross section (i.e., line 125A contacts the bottom of sensor 105A, line 125B contacts the bottom of sensor 105B, line 125C contacts the bottom of sensor 105C, and line 125D contacts the bottom of sensor 105D), and line 125H contacts the top of each of sensors 105D, 105C, 105B, and 105A.

[0070] The portions of the lines 125 that connect the sensors 105 and the sensor array 110 are 5B using dashed lines to indicate that the sensor 105 may be embedded within the fluid chamber 115 in which it may be enclosed. As explained above, the sensor 105 may be protected (e.g., by an insulator) from the contents of the fluid chamber 115 in which it may be enclosed. Thus, it should be understood that the various illustrated components (e.g., at the lines 125, the sensor 105, the binding site 116, etc.) are not necessarily visible in a physical instantiation of the device 100 (e.g., they may be embedded in or covered by a protective material such as an insulator).

[0071] In some embodiments, some or all of the binding sites 116 are present in nanowells or trenches of the line 125 that pass through the sensor 105. For example, as shown in the example of FIG. 5D, the line 125H may be thinner above the sensor 105 than between the sensors 105. For example, the line 125H has a first thickness above the sensor 105D, a second, larger thickness between the sensors 105D and 105C, and a first thickness above the sensor 105C. Such configurations can be advantageously manufactured using conventional thin film manufacturing methods (e.g., by depositing a material, applying a mask to the deposited material, and removing (e.g., by etching) a portion of the deposited material according to the mask). Both the binding sites 116 and the nanowells, if present, can be manufactured using conventional techniques.

[0072] For ease of explanation, Figures 5B, 5C, and 5D show an exemplary device 100 with only 16 sensors 105 in the sensor array 110, only 16 corresponding binding sites 116, and 8 lines 125. It should be understood that the device 100 can have fewer or more sensors 105 in the sensor array 110, and therefore more or fewer binding sites 116. Similarly, an embodiment with lines 125 can have more or fewer lines 125. In general, any configuration of sensors 105 and binding sites 116 that allows the sensor 105 to detect labels attached to nucleotides incorporated into a single nucleic acid strand bound to the binding sites 116 can be used. Similarly, any configuration of one or more lines 125, or some other mechanism, that allows the sensor 105 to determine whether it has sensed one or more labels can be used. The examples presented herein are not intended to be limiting.

[0073] As explained above, the sensor 105 shown in Figures 5B, 5C, and 5D may be a magnetic sensor 105. The sensor 105 is therefore in close proximity to the binding site 116 and therefore to the nucleic acid strand bound to the binding site 116. It should be understood that the appropriate location of the sensor array 110 relative to the binding site 116 depends in part on the type of label used and therefore the type of sensor 105 used. For example, if the label is a fluorophore and the sensor 105 is an optical sensor, it may be appropriate for the sensor array 110 to be spaced apart from the binding site 116 (e.g., located above the binding site 116).

[0074] 5B, 5C, and 5D (and other figures herein) show sensors 105 and binding sites 116 in a one-to-one relationship, it should be understood that each binding site 116 may be sensed by multiple sensors 105. A feature that distinguishes SMAS device 100 from CLUS devices is that the sensors 105 of SMAS device 100 do not sense multiple nucleic acid strand instances. If SMAS device 100 has more sensors 105 than binding sites 116, it may be possible for at least some nucleic acid strands to be sensed by multiple sensors 105 (e.g., to improve accuracy of label detection).

[0075] 5B, 5C, and 5D is a rectangular array, with the sensors 105 arranged in rows and columns. In other words, the sensors 105 of the sensor array 110 are arranged in a rectangular grid. In some embodiments, adjacent rows and columns of the rectangular grid pattern are equidistant from one another, such that , as shown in FIG. 5E, the sensors 105 are arranged in a square grid (or lattice) pattern. In an embodiment in which the sensors 105 are arranged in a square grid pattern, each sensor 105 has up to four nearest neighbors. For example, as shown in FIG. 5E, sensor 105A has four nearest neighbors labeled 105B, 105C, 105D, and 105E. The nearest sensors 105 are nearest neighbor distances 112 away, as shown in FIG. 5E. Thus, each of sensors 105B, 105C, 105D, and 105E is distance 112 away from sensor 105A.

[0076] A commercially viable SMAS device 100 can use high precision nanoscale fabrication of densely packed nanoscale sensors 105 capable of recognizing individual labels. The size of the functionalized binding site 116 can be similar to the size of DNA with attached labels such that multiple strands cannot bind to the same binding site 116 or be sensed by the same sensor 105. A well-established metric for assessing the commercial competitiveness of a sequencer is how tightly DNA strands can be packed within the fluid chamber 115.

[0077] An appropriate value for the nearest neighbor distance 112, which can be used to determine the size of the SMAS device 100 and / or the maximum number of sensors 105 that can fit within a selected size SMAS device 100, can then be determined based on the characteristics of the sensors 105, the length of the nucleic acid strand that the device 100 is to sequence, and the characteristics of the labels used. For example, the total length of the nucleic acid strand and the size of the labels used can provide physical limitations on how close two sensors 105 can be placed within the SMAS device 100. In some embodiments, the size of the sensors 105 can be limited by the nanoscale patterning capabilities of the process used to manufacture the SMAS device 100. For example, with the technology available at the time of writing, the size of each magnetic sensor 105 (e.g., the diameter of the sensor 105 in the xy plane, assuming a cylindrical sensor 105) can be on the order of 20 nm. Assuming that the type of nucleic acid to be sequenced is DNA and that it is desired to sequence fragments up to 150 base pairs (bp) in length, the maximum length of the DNA strand 101 to be sequenced is about 50 nm in the extended state, but the ssDNA conformation can change between the extended and coiled states, as shown in FIG. 6A, depending on the ionic strength of the buffer. Since the labels 102 are involved in single molecule reactions, the labels 102 should have molecular dimensions. In the case of a SMAS device 100 using a magnetic sensor 105, the labels 102 can be, for example, superparamagnetic nanoparticles, organometallic compounds, or any other functional molecular group that can be detected by a nanoscale magnetic sensor 105. Thus, each label 102 is assumed to have a size of about 10 nm or less. With these assumptions, FIG. 6B shows the relative dimensions of the magnetic sensor 105, the DNA strand 101 in its extended state, and the magnetic labels 102.

[0078] A practical SMAS device 100 that uses magnetic sensors 105 to detect magnetic nanoparticles used as labels 102 can be implemented using existing technology. For the sake of discussion, assume that only labels 102 within 20 nm of the edge of the sensor 105 are detected. The detection range of each sensor 105 is small because the magnetic labels 102 that may be selected for nucleic acid sequencing applications (e.g., superparamagnetic nanoparticles, organometallic compounds, etc.) do not generate significant perturbations to the detected magnetic field. Labels 102 attached to nucleotides incorporated into ssDNA bound to the binding site 116 of a particular sensor 105 can temporarily reside outside the range of the respective sensor 105, but because ssDNA adopts various conformational states during the detection process, it is desirable that the labels are not permitted to reach the sensitivity space (detection region) of adjacent sensors 105 when the ssDNA adopts its fully extended state.

[0079] A practical sensor loading limit for SMAS device 100 is, for example, when the labels are superparamagnetic nanoparticles (e.g., iron oxide, iron platinum, etc.) and the sensor array 110 of SMAS device 100 is non-volatile. The SMAS device 100 can be derived assuming a rectangular (e.g., square) array of magnetic tunnel junctions (MTJs) similar to those used in photocatalytic data storage applications. In this case, a region of or immediately adjacent to each nanoscale sensor 105 can be functionalized to act as a respective binding site 116. A simple geometry for estimating the packing limit of a sensor array of a SMAS device 100 is shown in FIG. 7A, which shows two sensors 105A, 105B. Each sensor 105A, 105B is assumed to have a cylindrical shape for convenience only, is assumed to have a diameter of about 20 nm (as explained above), and is assumed to be able to detect any label within 20 nm of its edge. The sensing area boundary 111 is indicated by the inner dashed line shown in FIG. 7A. Sensor 105A senses DNA strand 101A bound to its binding site, and sensor 105B senses DNA strand 101B bound to its binding site. The maximum reach of the labels 102A, 102B when bound to nucleotides incorporated into strands 101A, 101B (e.g., when the 150 base DNA strand is in a fully uncoiled state) is indicated by the outer dash-dotted circle 103. For accurate sequencing results, it is desirable for each sensor 105 to detect only the labels 102 bound to nucleotides incorporated into the DNA strand 101 bound to the respective binding site 116 of the sensor 105. Thus, using the assumptions described above, the minimum nearest neighbor distance 112 between sensors 105 to avoid crosstalk (e.g., detecting a label 102 bound to a nucleotide incorporated into a nucleic acid strand 101 bound to the binding site 116 of another sensor 105) is about 100 nm.

[0080] In some embodiments of the SMAS device 100, the sensors 105 (e.g., MTJs) are arranged in a square lattice that is compatible with existing cross-point MRAM sensor geometries, as shown in FIG. 7B. The area of ​​the unit cell 114 is 10 4 nm 2 , which allows each DNA strand 101 to have approximately 10 4 nm 2 This allows the device to extend over an area of ​​approximately 10 10book / cm 2 This results in a DNA surface density of 10 for the SMAS device 100. Assuming that we use at least 10 instances of each strand 101 in the sensor array 110, 9 Unique chain / cm 2 can be sequenced simultaneously, generating 150 Gbase (1 billion × 150 bp DNA strand length) of information per square centimeter of the sensor array 110. In the ideal case (e.g., with a low chemical failure rate, only three DNA instances are needed, as discussed further below), approximately 3.3 × 10 9 Different strands / cm 2 can be sequenced simultaneously, generating approximately 500 Gbase of data per square centimeter of the sensor array 110.

[0081] As a specific example, a SMAS device 100 with a configuration similar to a single Toshiba 4 Gbit density STT-MRAM chip first introduced at the International Electron Device Meeting (IEDM) in 2016 could generate approximately 600 Gbase of high quality data. The minimum distance 112 between sensors 105 on the Toshiba platform is 90 nm, only slightly below the estimated minimum distance 112 of 100 nm derived above. Thus, crosstalk using a configuration similar to the Toshiba platform is likely to be low even for 150 base long ssDNA, although shorter fragments could be sequenced to further reduce crosstalk.

[0082] It should be understood that the arrangement of sensors 105 in a grid pattern (e.g., a square grid as shown in FIG. 7B) is one of many possible arrangements. Those skilled in the art will appreciate that other arrangements of sensors 105 are possible and are within the scope of the disclosure herein. For example, sensors 105 may be arranged in a hexagonal pattern as shown in FIG. 8A, which illustrates a top view of SMAS device 100. The exemplary SMAS device 100 shown in FIG. 8A includes a sensor array 110 for sensing indicators 102 within fluid chamber 115. Sensor array 110 includes a plurality of sensors 105, with 16 sensors 105 shown. Implementations of device 100 may include any number of sensors 105 (e.g., hundreds, thousands, millions, etc.). It should be understood that sensors 105 may have any suitable size and shape. To avoid obscuring the drawing, only two of the sensors 105, namely sensors 105A and 105B, are labeled in FIG. 8A. As explained above, the sensor 105 may be, for example, a magnetic sensor (e.g., to detect the effect of magnetism or magnetic nanoparticles). In general, the sensor 105 may have any suitable size and shape, as explained above in the description of at least FIGS. 5B, 5C, and 5D.

[0083] As shown in FIG. 8A, each of the sensors 105 is associated with a respective binding site 116. In other words, there is a one-to-one relationship between the sensors 105 and the binding sites 116. As shown in FIG. 8A, the sensor 105A is associated with the binding site 116A, the sensor 105B is associated with the binding site 116B, and each of the other unlabeled sensors 105 is also associated with a respective binding site 116. In the exemplary embodiment of FIG. 8A, each sensor 105 is shown disposed under its respective binding site 116, but it should be understood that the binding sites 116 may be in other locations relative to their respective sensors 105. For example, the binding sites 116 may be on the sides of the respective sensors 105. The description of the binding sites 116 in at least the description of FIG. 5B, 5C, and 5D applies to FIG. 8A and other figures showing the binding sites 116 and will not be repeated here.

[0084] The exemplary SMAS device 100 of Figure 8A also includes fluid chamber 115 described above in the description of Figures 5B, 5C, and 5D. Those descriptions also apply to Figure 8A and will not be repeated here.

[0085] The circuit 120 of the device 100 of FIG. 8A may include one or more lines 125. Each of the lines 125 in the exemplary embodiment of FIG. 8A identifies a row or a diagonal column of the sensor array 110. For example, each of the lines 125A, 125B, 125C, and 125D identifies a different row of the sensor array 110, and each of the lines 125E, 125F, 125G, and 125H identifies a different diagonal column of the sensor array 110. In the example shown in FIG. 8A, the device 100 has eight lines 125A, 125B, 125C, 125D, 125E, 125F, 125G, and 125H, and pairs of the lines 125 can be used to access individual sensors 105. For example, the lines 125A and 125H can be used to access the sensor 105A, and the lines 125B and 125H can be used to access the sensor 105B. The lines 125 may be directed below and / or above the sensor 105, as described in the description of Figures 5B, 5C, and 5D, among others.

[0086] 8A shows an exemplary device 100 having only 16 sensors 105, 16 corresponding binding sites 116, and 8 lines 125 in the sensor array 110, it should be understood that the SMAS device 100 can have fewer or more sensors 105 in the sensor array 110, and therefore more or fewer binding sites 116. Additionally, the SMAS device 100 can have more or fewer lines 125. In general, any configuration of sensors 105 and binding sites 116 that allows the sensor 105 to detect labels attached to nucleotides incorporated into a single nucleic acid strand bound to the binding sites 116 can be used. Similarly, any configuration of one or more lines 125, or some other mechanism, that allows a determination of whether the sensor 105 has sensed one or more labels can be used.

[0087] As shown in Figure 8B, when the sensors 105 are arranged in a hexagonal pattern, each sensor 105 has up to six nearest neighbors that are all at a nearest neighbor distance 112. In other words, each sensor 105 is nearest distance 112 from each of the other six sensors 105 that are closest to it. For example, as shown in Figure 8B, the unlabeled sensor 105 in the center of the figure has six nearest neighbors labeled 105A, 105B, 105C, 105D, 105E, and 105F. 1, has sensors 105, all spaced a nearest neighbor distance 112 apart.

[0088] The packing limit of the binding sites 116 of a SMAS device 100 using an optical sensor and fluorescent labels 102 (e.g., fluorophores) with a hexagonal pattern of binding sites 116 can be derived. Assuming that the labels 102 are fluorophores, the binding sites 116 are in a hexagonal pattern, and the sensor array 110 is distant from the binding sites 116, single molecule fluorescence from the labels 102 can be projected into the far field where it can be detected by the sensor array 110 with the light-sensitive sensor 105. Using single molecule super-resolution imaging techniques such as those described in CGGalbraith and JAGalbraith, "Super-resolution microscopy at a glance," Journal of Cell Science, Vol. 124(10), 1607-11 (2011), the location of individual fluorophore labels 102 within the SMAS device 100 can be resolved. Because the DNA packing dimensions are well below the diffraction limit, the location of the fluorophore labels 102 can be resolved. Although this type of detection can be somewhat complex and / or expensive, this technique has recently been introduced into commercial sequencing systems to improve the throughput of cluster-based sequencers. Furthermore, this technique may be implemented in the imaging of large single molecule arrays in the near future.

[0089] A simple geometry for estimating the packing limit of binding sites 116 arranged in a hexagonal pattern in a SMAS device 100 using fluorophore labels 102 is shown in Figure 9A. DNA strand 101A is bound to binding site 116A, and DNA strand 101B is bound to binding site 116B. (Sensor 105 is not shown in Figure 9A because it is assumed that sensor array 110 is remote from the binding sites.) The maximum reach of labels 102A, 102B when bound to an incorporated nucleotide (e.g., when a DNA strand with 150 bases is in its fully uncoiled state) is indicated by dash-dotted circle 103. To avoid crosstalk, fluorophore labels 102 attached to adjacent binding sites 116 are not allowed to occupy overlapping space during the imaging process, for example, fluorophore label 102A attached to a particular binding site 116A should not be allowed to reach the space accessible to fluorophore label 102B attached to the adjacent binding site 116B, as the ssDNA 101A explores its allowed conformational state. This restriction also helps to avoid fluorescence quenching. Assuming that fluorophore labels 102 are used, the binding sites 116 can be tightly packed in a hexagonal lattice, as shown in Figure 9B. Assuming that the maximum length of the 150 bp DNA strand 101 is 50 nm, the size of the fluorophore labels 102 is 10 nm, the minimum distance from the center of each binding site 116 to its edge is 20 nm, and each DNA strand 101 binds to the center of its respective binding site 116, the minimum distance 112 is 140 nm. Therefore, as shown in FIG. 9B, all DNA strands 101 have a total of 1.7×10 4 nm 2 This allows the occupancy of a unit cell 114 having an area of ​​5.9×10 9 chain / cm 2 , or 5.9×10 if approximately 10 instances of each DNA strand are present in the SMAS device 100. 8 Unique chain / cm 2The SMAS device 100 produces approximately 90 Gbase of data from every square centimeter of the sensor array 110. In the best case scenario, where only three DNA copies are needed, the sensor array 110 produces approximately 2×10 9 unique DNA strands / cm 2 With this, the SMAS device 100 can generate approximately 300 Gb of data from every square centimeter of the sensor array 110.

[0090] The above discussion of hexagonal arrays was in the context of fluorophore labels 102 and optical sensors 105. It is also possible to use hexagonal arrays of magnetic sensors 105. The sensor packing limit for a SMAS device 100 having a hexagonal arrangement of binding sites 116 and magnetic sensors 105 can be derived as described above in the description of Figures 7A and 7B. For the magnetic sensors 105, the nearest neighbor distance 112 is approximately 100 nm, which is approximately 100 nm over a unit cell area of ​​100 nm (see Figure 9B). 14 is about 8.7 x 10 3 nm 2 This means that.

[0091] FIG. 10 compares the density of the SMAS implementations described in the context of FIGS. 7A and 7B (magnetic labels 102 and magnetic sensors 105) and FIGS. 9A and 9B (fluorescent labels 102 and optical sensors 105) with that of the current state-of-the-art CLUS sequencer. For discussion, assume that the pitch of the nanocell array of the patterned flow cell is about 500 nm. As shown in the left panel of FIG. 10, the nanocells of the CLUS sequencer are arranged in a hexagonal lattice with a lattice constant of 500 nm. Each nanocell holds about 50 to about 200 identical DNA strands (e.g., generated by solid-phase bridge amplification). The top right of FIG. 10 shows a hexagonal SMAS lattice using fluorophore labeling and super-resolution imaging (e.g., as described in connection with FIGS. 9A and 9B), and the bottom right of FIG. 10 shows a square SMAS lattice using superparamagnetic nanoparticle labeling and a sensor array 110 of MTJs (e.g., as described in connection with FIGS. 7A and 7B). The three representations in FIG. 10 are scaled proportionally to show how the SMAS lattice configuration compares to the CLUS configuration. The black hexagons (left and top right) and squares (bottom right) indicate unit cells that hold the minimum number of individual molecules required to call the sequence of a nucleic acid strand. The ideal case, described in more detail below, where only three DNA strands are required for successful base calling is shown for the SMAS lattice. Note that in the SMAS instance (right side of FIG. 10), the DNA instances are randomly distributed throughout the sensor array 110 and their locations can be identified during the first sequencing cycle, as discussed further below.

[0092] As shown in Fig. 10, the area of ​​the unit cell of the CLUS device is 2.2 × 10 5 nm 2 4.6 × 10 8 Clusters / cm 2 This corresponds to a DNA cluster density of about 100 Gb / cm. With the above assumptions, the CLUS sequencer will generate about 70 Gbase of data per square centimeter of sensing area. In contrast, in the ideal case where only three instances of a strand are used, the SMAS device 100 will generate about 500 Gb / cm 2(magnetic sensor 105 (e.g., MTJ) and magnetic label 102 (e.g., superparamagnetic nanoparticles)) and about 300 Gb / cm 2 (Optical sensor 105 (super-resolution imaging) and fluorescent label 102) data are generated. The results of the exemplary embodiments of the CLUS sequencer and SMAS device 100 are summarized in the table below. The table estimates the sequencing throughput assuming only three instances of each DNA strand and ten instances of each DNA strand for the SMAS embodiment. [Table 1]

[0093] The above table shows that the SMAS device 100 outperforms the state-of-the-art CLUS device when the number of DNA instances used for the algorithmic error correction described further below is small (e.g., <10). As the error correction procedure relies on more instances of each ssDNA, the SMAS device 100 begins to behave like a CLUS device, and there is little or no benefit to sensing individual molecules rather than clusters. Fluorescent SMAS essentially represents the limit of reducing clusters to single molecules. One approach to reduce the cost of sequencing is to reduce the cluster size and move DNA clusters closer together in order to obtain more information from a fixed sensing area. This approach While sequencing reduces the amount of reagents needed to perform the sequencing chemistry, it also significantly increases the complexity and cost of imaging hardware by constantly pushing the limits of what is currently possible with commercially available optical instruments. This strategy is an uphill battle because inscaling is only possible through parallel improvements in chemistry. This is because the problems become larger for each reaction as clusters get smaller, and chemical defects that occur stochastically at the single molecule level become more pronounced and intolerable.

[0094] The cost of implementing super-resolution imaging in a CLUS device makes the SMAS device 100, and particularly the SMAS device 100 using magnetic sensors 105 and magnetic labels, an alternative to potentially destructive sequencing. The SMAS device 100 disclosed herein, and particularly the SMAS device using magnetic sensors 105, promises superior throughput at significantly lower equipment costs by leveraging technology and mass production developed by the large scale semiconductor and data storage industries.

[0095] SMAS sequencing protocol As explained above, when the SMAS device 100 is used for nucleic acid sequencing, the nucleic acid strands can be amplified either before or after the nucleic acid is added to the SMAS device 100 (e.g., using bridge amplification). Regardless of how the nucleic acid is amplified, the strands can be sequenced by SBS (e.g., by synthesizing dsDNA from ssDNA) one base at a time. The SMAS sequencing protocol is described assuming that the nucleic acid being sequenced is DNA. It should be understood that the disclosed protocol can be modified for sequencing other nucleic acids. Such modifications are within the capabilities of one of ordinary skill in the art, given the disclosure herein.

[0096] To simplify the analysis and illustrate the advantages of using the disclosed SMAS device 100 rather than a CLUS sequencer, consider a DNA sequencing protocol in which a single type of label (e.g., molecular, fluorescent, magnetic, etc.) is attached to all four nucleotides (A, T, C, and G). In other words, some type of identical label is attached to each of the four nucleotides (e.g., if the selected labels 102 are particles of FePt, then each of A, T, C, and G is labeled with an FePt particle). These labeled nucleotides are then incorporated into the DNA strand one base at a time using termination chemistry, e.g., once a nucleotide is incorporated, the label 102 is cleaved off before the polymerase moves to the next base. The sensor 105 detects the labels 102 attached to the nucleotides.

[0097] An exemplary method 200 for sequencing a plurality of nucleic acid strands (e.g., ssDNA) using SMAS device 100 is shown in FIG. 11. At 202, the method begins. At 204, optionally, one or more nucleic acid strands may be amplified before being added to SMAS device 100. At 206, a plurality of S nucleic acid strands are bound to a plurality of S binding sites 116 of SMAS device 100 (wherein plurality includes at least two, but not necessarily all, of binding sites 116 of SMAS device 100). Optionally, at 208, the nucleic acid strands are amplified (e.g., via bridge amplification, which may be performed in addition to or instead of the amplification at 204). At 210, a sequencing procedure is performed. The sequencing procedure may be, for example, a subtractive approach, a subtractive approach, or a modified additive approach, as described further below. The sequencing procedure performed at 210 produces S records, each of the S records capturing a number M of detection results for one of the plurality of S sensors (where again, plurality includes at least two but not necessarily all sensors 105 of SMAS device 100, and the M detection results may include as little as one detection result, some subset of the total number of detection results obtained during the sequencing procedure, or all of the detection results obtained during the sequencing procedure). Each of the M detection results indicates whether the sensor 105 to which the record corresponds detected at least one label during each of the M interrogation steps. The detection results may be stored in a record, which may be stored in a memory. At 212, an error correction procedure is performed, as described further below. The error correction procedure may include deterministic and / or probabilistic error correction techniques. The error correction procedure may be performed, for example, by at least one processor 130 of SMAS device 100. Alternatively, it may be performed by a processor external to SMAS device 100 (e.g., an off-device processor, such as in an external computer). The error correction procedure may be performed as the sequencing procedure is progressing (e.g., in real-time or near real-time), or may be performed at a later time. At 214, the method 200 ends.

[0098] As described above, at 210, various protocols can be implemented to read nucleic acid sequences (e.g., DNA sequences) using the SMAS device 100. For simplicity of analysis, it is assumed that the multiple S-sensors 105 of the SMAS device 100 detect only the presence or absence of the labels 102 and do not distinguish between nucleotides based on the detected signal level. As a result, in some embodiments, the recording of the detection results of each sensor 105 includes only a "yes" or "no" (or 1 / 0 or any other binary indicator) indication of whether the sensor 105 detected a label or did not detect a label during a particular interrogation step. It should be understood that other approaches are possible and are within the scope of the disclosure herein. For example, different labels 102 can be attached to different nucleotides. As another example, rather than a binary "yes" or "no" determination, a value of a property can be detected (e.g., resistance, frequency, intensity, etc.) and / or recorded, and a decision can be made based on that criterion as to whether the label was detected. For example, instead of simply having 0 and 1 (or "no" and "yes") as possible outputs of a sequencing procedure, using different labels for different nucleotides allows one of five labels to be obtained: 0 (no label detected), level 1 (label 1 detected), level 2 (label 2 detected), level 3 (label 3 detected), and level 4 (label 4 detected). In such a case, a range of detected properties can be defined to distinguish whether a label was detected at all, and if so, which label was detected (e.g., if the value of the property is between 0 and a first value, it is determined that no label was detected; if the value of the property is between a first value and a second value, it is determined that a first label was detected; if the value of the property is between a second value and a third value, it is determined that a second label was detected, etc.).

[0099] Below is a description of three example DNA sequencing protocols, each including repeated interrogation cycles, each having four interrogation steps. During each interrogation cycle, four binary "yes" or "no" questions are answered for each ssDNA being sequenced. One interrogation step asks the question "Is the detected base an adenine?" ("A?") is answered. Another interrogation step asks the question "Is the detected base a thymine?" ("T?") is answered. Another interrogation step asks the question "Is the detected base a cytosine?" ("C?") is answered. And another interrogation step asks the question "Is the detected base a guanine?" ("G?") is answered. A record of the detection results obtained during the sequencing procedure can be made as the interrogation cycles, including A? → T? → C? → G?, are repeated. The described order in which nucleotides are introduced and bases are detected is arbitrary (meaning that the order of the interrogation steps is arbitrary) and the order in which bases are tested in the examples herein (A?⇒T?⇒C?⇒G?) is merely exemplary.

[0100] Additive approach In the additive approach, the sensor 105 detects nanoscale labels 102 attached to nucleotides with cleavable linkers. All four nucleotides have the same kind of label 102 (e.g., molecular, fluorescent, magnetic, etc.) and use the same type of cleavable linker. Four detection results are obtained, one of which indicates the absence of an error and the other indicates the presence of multiple S nuclei. The interrogation cycle resulting in label detection for each of the acid chains 101 comprises the following steps according to one embodiment:

[0101] 1. Obtain baseline characteristics of each of the multiple S sensors 105 of the SMAS device 100 (which may be all or less than all of the sensors 105 in the sensor array 110) (e.g., by measuring a baseline signal at each of the multiple S sensors 105).

[0102] 2. Labeled A nucleotides are introduced and incorporated. Unbound labeled molecules are washed away.

[0103] 3. Interrogation step 1: Obtain a characteristic of each of the plurality of S-sensors 105 (e.g., by detecting a signal at each of the plurality of S-sensors 105) and determine whether each sensor 105 detected at least one indicator. Store the detection results for each sensor 105 in a location in the record corresponding to interrogation step 1 of the current interrogation cycle.

[0104] 4. Introduce and incorporate labeled T nucleotides. Wash away unbound labeled molecules.

[0105] 5. Interrogation step 2: Obtain a characteristic of each of the plurality of S-sensors 105 (e.g., by detecting a signal at each of the plurality of S-sensors 105) and determine whether each sensor 105 detected at least one indicator. Store the detection results for each sensor 105 in a location in the record corresponding to interrogation step 2 of the current interrogation cycle.

[0106] 6. Introduce and incorporate labeled C nucleotides. Wash away unbound labeled molecules.

[0107] 7. Interrogation step 3: Obtain a characteristic of each of the plurality of S-sensors 105 (e.g., by detecting a signal at each of the plurality of S-sensors 105) and determine whether each sensor 105 detected at least one indicator. Store the detection results for each sensor 105 in a location in the record corresponding to interrogation step 3 of the current interrogation cycle.

[0108] 8. Introduce and incorporate labeled G nucleotides. Wash away unbound labeled molecules.

[0109] 9. Interrogation step 4: Obtain a characteristic of each of the plurality of S-sensors 105 (e.g., by detecting a signal at each of the plurality of S-sensors 105) and determine whether each sensor 105 detected at least one indicator. Store the detection results for each sensor 105 in a location in the record corresponding to interrogation step 4 of the current interrogation cycle.

[0110] 10. Cleave the labels from the A, T, C and G nucleotides and wash away.

[0111] Steps 1 through 10 can then be repeated for the next interrogation cycle. It should be understood that the particular order of steps 1 through 10 is exemplary, and furthermore, the number and numbering of steps 1 through 10 is for convenience and can be changed. As an example, as explained above, the order in which nucleotides are introduced is arbitrary. As another example, although steps 2, 4, 6, and 8 include the introduction and incorporation of nucleotides and washing away unbound nucleotides as single steps, it should be understood that each of steps 2, 4, 6, and 8 can be divided into a series of smaller steps. Similarly, steps 3, 5, 7, and 9 can be further divided into a series of smaller steps (e.g., acquiring a property, determining whether a label is detected, and storing the detection results). Conversely, steps can also be combined (e.g., steps 2 and 3 can be combined, steps 4 and 5 can be combined, etc.).

[0112] If there is a high probability that no errors will occur during any query cycle of the additive approach, It should be understood that in this case, it is possible to call (determine) each base of each individual strand as soon as the label is detected. For example, referring to the above steps, for a particular sensor 105, in interrogation step 1 with a labeled A nucleotide, if the obtained characteristic indicates that the sensor 105 has detected the label, storing the detection result may result in calling for that sensor 105 (and binding site 116) the base complementary to A (T). Similarly, for a particular sensor 105, in interrogation step 2 with a labeled T nucleotide, if the obtained characteristic indicates that the sensor 105 has detected the label, storing the detection result may result in calling for that sensor 105 (and binding site 116) the base complementary to T (A). Similarly, for a particular sensor 105, in interrogation step 3 with a labeled C nucleotide, if the obtained characteristic indicates that the sensor 105 has detected the label, storing the detection result may result in calling for that sensor 105 (and binding site 116) the base complementary to C (G). Finally, for a particular sensor 105, if the resulting characteristic in interrogation step 4, which includes a labeled G nucleotide, indicates that the sensor 105 detects the label, then storing the detection result may result in calling the base complementary to G (C) for that sensor 105 (and binding site 116). However, as described in more detail below, there are several types of errors that can occur during the sequencing procedure (e.g., during an additive approach), and therefore in some embodiments, records are made during the sequencing procedure to record label detection / non-detection during each interrogation step of each interrogation cycle. An error correction procedure may then be applied to some or all of the records before calling the base.

[0113] FIG. 12 is a flow diagram of a sequencing procedure 220 using an additive approach, according to some embodiments. The sequencing procedure 220 can be, for example, a sequencing procedure performed in step 210 of an exemplary method 200 of sequencing a plurality of nucleic acid strands (e.g., ssDNA) using the SMAS device 100 shown and described in the description of FIG. 11. At 222, the sequencing procedure 220 begins. At 224, a baseline characteristic of each of the S-sensors 105 is obtained (e.g., by at least one processor 130 of the SMAS device 100 with the aid of the circuitry 120). Once the interrogation cycle begins, at 226, a first labeled nucleotide is selected (e.g., with reference to steps 1-10 above, the first labeled nucleotide can be A). At 228, the selected labeled nucleotide is introduced into the fluid chamber 115, where the nucleotide potentially becomes incorporated into the nucleic acid strand bound to the binding site 116. At 230, unbound nucleotides are washed away. At 232, a characteristic is obtained from each of the plurality of S sensors, and a detection result (e.g., a label detected or not detected) is determined for each of the plurality of S sensors 105. At 234, the S detection results are recorded in S records (e.g., as a 1 indicating a label was detected or as a 0 indicating a label was not detected). At 236, it is determined whether the last tested nucleotide was the last nucleotide of the interrogation cycle. For the example order of nucleotide testing assumed in steps 1-10 above, it is determined (e.g., by at least one processor 130) at 236 whether G was the last tested nucleotide. If not, at 238, the next labeled nucleotide to be tested in the interrogation cycle is selected, and steps 228 through 236 are repeated until it is determined at 236 that the last tested nucleotide was the last nucleotide of the interrogation cycle. At 240, the label is cleaved and washed away. At 242, it is determined (e.g., by at least one processor 130) whether the last completed interrogation cycle was the last interrogation cycle of the sequencing procedure 220.For example, the at least one processor 130 can determine whether enough detection results have been recorded to allow the at least one processor 130 (or some other processing entity, such as an external processor) to call the target number of bases (e.g., 150 bases). If not, the sequencing procedure 220 returns to step 224. If so, the sequencing procedure 220 returns to step 244, again as described above. As noted, the order in which the nucleotides are tested is arbitrary.

[0114] In the exemplary case of DNA sequencing, an additive sequencing protocol involving four nucleotide incorporations and one label cleavage reaction is summarized in FIG. 13. The leftmost panel of FIG. 13 shows a sensor array 110 with a total of 100 individual sensors 105, which are shown as squares. For illustrative purposes, each of the 100 binding sites 116 in the sensor array 110 is assumed to hold a respective DNA strand, and each DNA strand is sensed by a respective sensor 105 (in other words, there is a one-to-one relationship between the binding sites 116 and the sensors 105). Some of the DNA strands may be copies of others. Labeled nucleotides are added to the fluid chamber 115 one at a time, and the labels are cleaved simultaneously after the nucleotide is incorporated. If there are no errors, base calling can be achieved after five reactions, i.e., four nucleotide incorporations and one base cleavage reaction. If errors occur, an error correction procedure can be applied, as described below.

[0115] Subtractive Approach In the subtractive approach, a sensor 105 detects nanoscale labels 102 attached to nucleotides with cleavable linkers. All four types of nucleotides have the same type of label (e.g., molecular, fluorescent, magnetic, etc.), but each has a different type of cleavable linker. An interrogation cycle that results in four detection results in the absence of errors, one of which is a label detection for each of the multiple S nucleic acid strands 101 in the absence of errors, in one embodiment includes the following steps:

[0116] 1. Simultaneously introduce labeled A, T, C and G nucleotides and incorporate and rinse out unbound labeled molecules. Obtain a baseline characteristic of each of the plurality of S sensors 105 (e.g., by detecting a signal at each of the plurality of S sensors 105). In the absence of errors, all sensors 105 will detect the label.

[0117] 2. Interrogation step 1: Introduce a reagent (e.g., an enzyme) that cleaves the label only from the first nucleotide, e.g., A, rinse, and obtain a characteristic (e.g., measure a signal) at each of the plurality, S, of sensors 105. Determine which sensors 105 are no longer detecting the label (e.g., based on a change in baseline characteristic). Store the detection results for each sensor 105 in a location in the record that corresponds to interrogation step 1 of the current interrogation cycle.

[0118] 3. Interrogation step 2: Introduce a reagent that cleaves the label only from the second nucleotide, e.g., T, rinse, and obtain a characteristic (e.g., measure a signal) at each of the plurality of S sensors 105. Determine which sensors 105 are no longer detecting the label (e.g., based on a change in baseline characteristic). Store the detection results for each sensor 105 in a location in the record that corresponds to interrogation step 2 of the current interrogation cycle.

[0119] 4. Interrogation step 3: Introduce a reagent that cleaves the label only from the third nucleotide, e.g., C, rinse, and obtain a characteristic (e.g., measure a signal) at each of the plurality, S, of sensors 105. Determine which sensors 105 are no longer detecting the label (e.g., based on a change in baseline characteristic). Store the detection results for each sensor 105 in a location in the record that corresponds to interrogation step 3 of the current interrogation cycle.

[0120] 5. Interrogation step 4: Introduce a reagent that cleaves the label only from the fourth nucleotide, e.g., G, rinse, and obtain a characteristic (e.g., measure a signal) at each of the plurality of S sensors 105. Determine which sensors 105 are no longer detecting the label (e.g., based on a change in baseline characteristic). Store the detection results for each sensor 105 in a location in the record that corresponds to interrogation step 4 of the current interrogation cycle.

[0121] Steps 1 through 5 can be repeated for the next interrogation cycle. It should be understood that the particular order of steps 1 through 5 is exemplary, and furthermore, the number and numbering of steps 1 through 5 is for convenience and can be changed. As an example, as explained above, the order in which the nucleotides are cleaved is arbitrary. Similarly, in step 1, the nucleotides can be introduced sequentially (not necessarily simultaneously). As another example, although interrogation steps 1, 2, 3, and 4 include introducing reagents, rinsing, acquiring properties, determining which sensors are no longer detecting (or still detecting) the label, and storing the results as a single step, it should be understood that each interrogation step can be divided into a series of smaller steps.

[0122] It should be understood that if there is a high probability of no errors during any interrogation cycle of the subtractive approach, it is possible to call (determine) each base of each individual strand as soon as label removal (absence of label) is first detected. For example, referring to the above steps, for a particular sensor 105, in interrogation step 1 with a labeled A nucleotide, if the resulting characteristic indicates that the sensor 105 is no longer detecting the label, storing the detection result may result in calling the base complementary to A (T) for that sensor 105 (and binding site 116). Similarly, for a particular sensor 105, in interrogation step 2 with a labeled T nucleotide, if the resulting characteristic indicates that the sensor 105 is no longer detecting the label, storing the detection result may result in calling the base complementary to T (A) for that sensor 105 (and binding site 116). Similarly, for a particular sensor 105, if the resulting characteristic in interrogation step 3 with a labeled C nucleotide indicates that the sensor 105 is no longer detecting the label, storing the detection result may result in calling the base complementary to C (G) for that sensor 105 (and binding site 116). Finally, for a particular sensor 105, if the resulting characteristic in interrogation step 4 with a labeled G nucleotide indicates that the sensor 105 is no longer detecting the label, storing the detection result may result in calling the base complementary to G (C) for that sensor 105 (and binding site 116). However, as explained in more detail below, there are several types of errors that can occur during the sequencing procedure (e.g., during subtractive approaches), and therefore, in some embodiments, records are made during the sequencing procedure to record label detection / non-detection during each interrogation step of each interrogation cycle. An error correction procedure may then be applied to some or all of the records before calling the base.

[0123] FIG. 14 is a flow diagram of a sequencing procedure 250 using a subtractive approach, according to some embodiments. The sequencing procedure 250 may be, for example, a sequencing procedure performed in step 210 of an exemplary method 200 of sequencing a plurality of nucleic acid strands (e.g., ssDNA) using the SMAS device 100 shown and described in the description of FIG. 11. At 250, the sequencing procedure starts at 252. At 254, all of the labeled nucleotides are introduced into the fluid chamber 115, and the nucleotides are incorporated into the nucleic acid strands bound to the S binding sites 116. At 256, the unbound nucleotides are washed away. At 258, a baseline characteristic of each of the S sensors 105 is obtained (e.g., by at least one processor 130 of the SMAS device 100 with the aid of the circuitry 120). Assuming that a nucleotide has been incorporated into the nucleic acid strand bound to each of the S binding sites, the obtained characteristic represents the characteristic of the sensor 105 when detecting at least one label. At 260, one of the cleavable linkers is selected for cleavage (or, equivalently, one of the nucleotides is selected). At 262, the label attached to the selected nucleotide is cleaved and rinsed. Assuming there are no errors, following step 262, the sensor 105 sensing the nucleic acid strand incorporating the tested nucleotide (e.g., one with a label attached by the selected cleavable linker) exhibits a change in a property (e.g., a change in a signal associated with or generated by the sensor 105). At 264, the plurality of S sensors A characteristic is obtained from each of the plurality of S sensors 105 and a detection result (e.g., a label detected or not detected) is determined for each of the plurality of S sensors 105. The S detection results are recorded in S records (e.g., as a 1 indicating a label was detected or as a 0 indicating a label was not detected) at 266. It is determined at 268 whether the last tested nucleotide was the last nucleotide of the interrogation cycle. For the example order of nucleotide testing assumed in steps 1-5 above, it is determined (e.g., by at least one processor 130) at 268 whether G was the last tested nucleotide. If not, then at 270 the next cleavable linker (or equivalently, the next tested nucleotide) to be cleaved in the interrogation cycle is selected and steps 262 through 268 are repeated until it is determined at 268 that the last cleaved linker (or equivalently, the last tested nucleotide) was the last linker (or nucleotide) of the interrogation cycle. At 272, it is determined (e.g., by the at least one processor 130) whether the last completed interrogation cycle is the last interrogation cycle of the sequencing procedure 250. For example, the at least one processor 130 can determine whether enough detection results have been recorded to allow the at least one processor 130 (or some other processing entity, such as an external processor) to call the target number of bases (e.g., 150 bases). If not, the sequencing procedure 250 returns to step 254. If so, the sequencing procedure 250 returns to step 274. Again, as explained above, the order in which the nucleotides are tested is arbitrary.

[0124] In the exemplary case of DNA sequencing, a subtractive sequencing protocol involving one nucleotide incorporation and four base cleavage reactions is summarized in FIG. 15. The leftmost panel of FIG. 15 shows a sensor array 110 with a total of 100 individual sensors 105, which are shown as squares. For illustrative purposes, each of the 100 binding sites 116 in the sensor array 110 is assumed to hold a respective DNA strand, and each DNA strand is sensed by a respective sensor 105 (in other words, there is a one-to-one relationship between the binding sites 116 and the sensors 105). Some of the DNA strands may be copies of others. All four labeled nucleotides are added simultaneously to the fluid chamber 115, and the labels are removed after incorporation, one nucleotide (e.g., the cleavable linker) at a time. If there are no errors, base calling can be achieved after five reactions, i.e., one nucleotide incorporation and four base cleavage reactions. If errors occur, an error correction procedure described below can be applied.

[0125] Modified Additive Approach In a modified additive approach, the sensor 105 detects nanoscale labels 102 attached to nucleotides with cleavable linkers. All four types of nucleotides have the same type of label 102 (e.g., molecular, fluorescent, magnetic, etc.) and use the same type of cleavable linker. The labeled nucleotides are added separately and the presence of the label 102 is detected after addition of each nucleotide. In the absence of errors, an interrogation cycle that results in four detection results, at least one of which is a label detection, for each of the plurality S of nucleic acid strands 101, includes, in one embodiment, the following steps:

[0126] 1. Obtain baseline characteristics for each of the multiple S sensors 105 of the SMAS device 100 (which may be all or less than all of the sensors 105 in the sensor array 110) (e.g., by measuring a baseline signal at each of the multiple S sensors 105).

[0127] 2. Introduce and incorporate the first labeled nucleotide, e.g., labeled A nucleotide. Wash away unbound labeled molecules.

[0128] 3. Interrogation step 1: Obtain a characteristic of each of the plurality of S-sensors 105 (e.g., by detecting a signal at each of the plurality of S-sensors 105) and determine whether each sensor 105 detected at least one indicator. Store the detection results for each sensor 105 in a location in the record corresponding to interrogation step 1 of the current interrogation cycle.

[0129] 4. Cut off the sign and wash it off.

[0130] 5. A second labeled nucleotide, e.g. a labeled T nucleotide, is introduced and incorporated. Unbound labeled molecules are washed away.

[0131] 6. Interrogation step 2: Obtain a characteristic of each of the plurality of S-sensors 105 (e.g., by detecting a signal at each of the plurality of S-sensors 105) and determine whether each sensor 105 detected at least one indicator. Store the detection results for each sensor 105 in a location in the record corresponding to interrogation step 2 of the current interrogation cycle.

[0132] 7. Cut off the sign and wash it off.

[0133] 8. Introduce and incorporate a third labeled nucleotide, e.g., a labeled C nucleotide. Wash away unbound labeled molecules.

[0134] 9. Interrogation step 3: Obtain a characteristic of each of the plurality of S-sensors 105 (e.g., by detecting a signal at each of the plurality of S-sensors 105) and determine whether each sensor 105 detected at least one indicator. Store the detection results for each sensor 105 in a location in the record corresponding to interrogation step 3 of the current interrogation cycle.

[0135] 10. Cut off the sign and wash it off.

[0136] 11. Introduce and incorporate a fourth labeled nucleotide, e.g., a labeled G nucleotide. Wash away unbound labeled molecules.

[0137] 12. Interrogation step 4: Obtain a characteristic of each of the plurality of S-sensors 105 (e.g., by detecting a signal at each of the plurality of S-sensors 105) and determine whether each sensor 105 detected at least one indicator. Store the detection results for each sensor 105 in a location in the record corresponding to interrogation step 4 of the current interrogation cycle.

[0138] 13. Cut off the sign and wash it off.

[0139] Steps 1 through 13 can then be repeated for the next interrogation cycle. It should be understood that the particular order of steps 1 through 13 is exemplary, and furthermore, the number and numbering of steps 1 through 13 is for convenience and can be changed. As an example, as explained above, the order in which nucleotides are introduced is arbitrary. As another example, while steps 2, 5, 8, and 11 include the introduction and incorporation of nucleotides and washing away unbound nucleotides as single steps, it should be understood that each of steps 2, 5, 8, and 11 can be divided into a series of smaller steps. Similarly, steps 3, 6, 9, and 12 (interrogation steps 1, 2, 3, and 4) can be further divided into a series of smaller steps (e.g., acquiring a property, determining whether a label is detected, and storing the detection results). Conversely, steps can also be combined (e.g., steps 2 and 3 can be combined, steps 3 and 4 can be combined, steps 2-4 can be combined, steps 5 and 6 can be combined, steps 6 and 7 can be combined, steps 5-10 can be combined, steps 6-12 can be combined, steps 6-14 can be combined, steps 6-16 can be combined, steps 6-18 can be combined, steps 6-19 can be combined, steps 7-20 can be combined, steps 7-21 can be combined, steps 7-22 can be combined, steps 7-23 can be combined, steps 7-24 can be combined, steps 7-26 can be combined, steps 7-28 can be combined, steps 7-29 can be combined, steps 8-30 can be combined, steps 8-31 can be combined, steps 8-32 can be combined, steps 8-33 can be combined, steps 8-34 can be combined, steps 8-35 can be combined, steps 8-36 can be combined, steps 8-37 can be combined, steps 8- 7 can be combined, etc.

[0140] It should be understood that if there is a high probability of no error during any interrogation cycle of the modified additive approach, it is possible to call (determine) each base of each individual strand as soon as the label is detected. For example, referring to the above steps, for a particular sensor 105, in interrogation step 1 with a labeled A nucleotide, if the obtained characteristic indicates that the sensor 105 has detected the label, storing the detection result may result in calling the base complementary to A (T) for that sensor 105 (and binding site 116). Similarly, for a particular sensor 105, in interrogation step 2 with a labeled T nucleotide, if the obtained characteristic indicates that the sensor 105 has detected the label, storing the detection result may result in calling the base complementary to T (A) for that sensor 105 (and binding site 116). Similarly, for a particular sensor 105, in interrogation step 3 with a labeled C nucleotide, if the obtained characteristic indicates that the sensor 105 has detected the label, storing the detection result may result in calling the base complementary to C (G) for that sensor 105 (and binding site 116). Finally, for a particular sensor 105, if the resulting characteristic in interrogation step 4, which includes a labeled G nucleotide, indicates that the sensor 105 detects the label, then storing the detection result may result in calling the base complementary to G (C) for that sensor 105 (and binding site 116). However, as described in more detail below, there are several types of errors that can occur during the sequencing procedure (e.g., during an additive approach), and therefore in some embodiments, records are made during the sequencing procedure to record label detection / non-detection during each interrogation step of each interrogation cycle. An error correction procedure may then be applied to some or all of the records before calling the base.

[0141] FIG. 16 is a flow diagram of a sequencing procedure 350 using a modified additive approach, according to some embodiments. The sequencing procedure 350 can be, for example, a sequencing procedure performed in step 210 of an exemplary method 200 of sequencing a plurality of nucleic acid strands (e.g., ssDNA) using the SMAS device 100 shown and described in the description of FIG. 11. At 352, the sequencing procedure 350 begins. At 354, a baseline characteristic of each of the S-sensors 105 is obtained (e.g., by at least one processor 130 of the SMAS device 100 with the aid of the circuitry 120). Once the interrogation cycle begins, at 356, a first labeled nucleotide is selected (e.g., with reference to steps 1-13 above, the first labeled nucleotide can be A). At 358, the selected labeled nucleotide is introduced into the fluid chamber 115, where the nucleotide potentially becomes incorporated into the nucleic acid strand bound to the binding site 116. At 360, unbound nucleotides are washed away. At 362, a characteristic is obtained from each of the plurality of S sensors, and a detection result (e.g., a label detected or not detected) is determined for each of the plurality of S sensors 105. At 364, the S detection results are recorded in S records (e.g., as a 1 indicating a label was detected or as a 0 indicating a label was not detected). At 366, the label is cleaved and washed away. At 368, it is determined whether the last tested nucleotide was the last nucleotide of the interrogation cycle. For the example order of nucleotide testing assumed in steps 1-13 above, it is determined (e.g., by at least one processor 130) whether G was the last tested nucleotide. If not, at 370, the next labeled nucleotide to be tested in the interrogation cycle is selected, and steps 358 through 368 are repeated until it is determined at 368 that the last tested nucleotide was the last nucleotide of the interrogation cycle. At 372, it is determined (e.g., by at least one processor 130) whether the last completed interrogation cycle was the last interrogation cycle of the sequencing procedure 350.For example, at least one processor 130 may be a target for at least one processor 130 (or some other processing entity, such as an external processor). It can be determined whether enough detection results have been recorded to allow the number of bases (e.g., 150 bases) to be called. If not, the sequencing procedure 350 returns to step 354. If so, the sequencing procedure 350 returns to step 374. Again, as explained above, the order in which the nucleotides are tested is arbitrary.

[0142] In the exemplary case of DNA sequencing, a modified additive sequence determination protocol including four nucleotide incorporation and four base cleavage reactions is shown in FIG. 17. The leftmost panel of FIG. 17 shows a sensor array 110 with a total of 100 individual sensors 105, which are shown as squares. For illustrative purposes, each of the 100 binding sites 116 in the sensor array 110 is assumed to hold a respective DNA strand, and each DNA strand is sensed by a respective sensor 105 (in other words, there is a one-to-one relationship between the binding sites 116 and the sensors 105). Some of the DNA strands may be copies of others. As shown and described, labeled nucleotides are added to the fluid chamber 115 one at a time, and the labels are cleaved after incorporation and label detection. In the absence of errors, on average, a base call can be achieved after five reactions, i.e., 2.5 nucleotide incorporation and 2.5 base cleavage reactions.

[0143] Thus, in the absence of errors, for DNA sequencing, the modified additive approach results in at least one base call per ssDNA after 8 reactions (4 nucleotide incorporations and 4 base cleavage), testing all bases. However, on average, a base call is made after only 5 reactions (2.5 nucleotide incorporations and 2.5 base cleavage). Because the label is removed after the incorporation of every nucleotide, multiple nucleotides can be incorporated and called during a single A? → T? → C? → G? interrogation cycle. Specifically, for an unknown ssDNA sequence, there is a 1 / 4 chance that the unknown base is T. If the base happens to be T, it is detected in the third step after one incorporation and one base cleavage reaction when an A nucleotide is introduced. There is a 1 / 4 chance that the unknown base is A. If the base happens to be A, it is detected in the fifth step of the interrogation cycle A? → T?, when a T nucleotide is introduced and two incorporations and two cleavages are performed. There is a 1 / 4 chance that the unknown base is G. If the base happens to be G, it will be detected in step 7 of the query cycle A? → T? → C? when a C nucleotide is introduced and 3 incorporations and 3 cleavages have been performed. Finally, there is a 1 in 4 chance that the unknown base is C. If the base happens to be C, it will be detected in step 11 of the query cycle A? → T? → C? → G? when a C nucleotide is introduced and 4 incorporations and 4 cleavages have been performed. Thus, it takes an average of 2.5 queries (5 reactions) to call a single unknown base.

number

[0144] Sources of sequencing errors Ideally, a sequencing procedure, whether CLUS or SMAS device 100, is error-free. In other words, for example, nucleotides are always properly labeled, nucleotides are always correctly incorporated into DNA, all labels are successfully cleaved during the cleavage step, all cleaved labels are successfully washed away, etc. In practice, however, errors can occur during any sequencing procedure. This section investigates the causes of sequencing errors in both CLUS and SMAS devices 100 and describes error mitigation strategies for SMAS device 100. As described further below, error correction methods can be used to improve the sequencing accuracy of SMAS device 100.

[0145] The modified additive approach described above is a good model for explaining how errors propagate in both the CLUS and SMAS devices 100 because it is a conceptually simple (symmetric in that each nucleotide is treated the same way) sequencing procedure. Assuming that the nanoscale labels are attached to the nucleotides via cleavable linkers, there are four possible sources of error. Each error occurs at a rate, denoted as r, with a value between 0 and 1. The four sources of error are:

[0146] Failed Nucleotide Incorporation (FNI): Failed Nucleotide Incorporation (FNI) occurs when a properly labeled nucleotide molecule does not reach the ssDNA binding site or when the polymerase is unable to incorporate it. Figure 18 A shows an FNI of the CLUS instrument sequencing five instances of ssDNA. Following the flow of complementary nucleotides, only three of the five ssDNAs incorporated the labeled nucleotide (shown as having a magnetic label). Thus, two of the five nucleotides

number

[0147] Failure to Remove Label (FLR): A failure to remove label (FLR) occurs when a labeled nucleotide molecule is incorporated but the label is not removed after label detection because the cleavage reagent does not reach or is unable to cleave the linker. Figure 18C shows an FLR for the CLUS device described above in the description of Figure 18A. After incorporation of the complementary nucleotide and rinsing to remove unbound nucleotides, detection of the label, and cleavage and rinsing of the label, one label is attached to one of the ssDNA instances.

number

number

[0148] Failure of Nucleotide Removal (FNR): A failure of nucleotide removal (FNR) occurs when a labeled nucleotide, whether complementary or non-complementary, binds non-specifically to the binding site 116 and / or the surface of the sensor 105. FIG. 18E shows the FNR for the CLUS device described above in the description of FIG. 18A. After nucleotide flow and rinsing to remove unbound nucleotides, two bad nucleotides and their labels remain on the surface of the binding site. Similarly, in FIG. 18F, which shows the FNR for the SMAS device 100 described above in the description of FIG. 18B, after nucleotide flow and rinsing to remove unbound nucleotides, one bad nucleotide remains on the surface of binding site 116A and another bad nucleotide remains on the surface of binding site 116D. In this example, for both the CLUS device and the SMAS device 100,

number

[0149] Failure to detect label (FLD): A failure to detect label (FLD) occurs when the correct complementary nucleotide is incorporated but the label is not detected because it is missing or the sensor is unable to recognize it. Figure 18G shows an FLD for the CLUS device described above in the description of Figure 18A. After incorporation of the complementary nucleotide and rinsing to remove unbound nucleotides, two of the instances of ssDNA incorporate the complementary nucleotide but lack the label.

number

number

[0150] Although Figures 18A-18H show the labels as magnets, thereby suggesting magnetic labels and magnetic sensors, as described above, it should be understood that the labels may be any type of detectable label (e.g., fluorescent, magnetic, etc.) and the sensors may be any type of sensor capable of detecting a selected type of label (e.g., optical, magnetic, organometallic, charged molecule, etc.).

[0151] The four error types (FNI, FLR, FNR and FLD) have the same ratio r

number

number

number

[0152] Cluster sequencers versus single molecule array sequencers: Qualitative comparison and error correction Two types of error correction are disclosed herein, referred to as deterministic error correction and probabilistic error correction. SMAS device 100 may use either or both types of error correction, as described further below.

[0153] As explained above, the modified additive approach is a good model for explaining how errors propagate and how the disclosed error correction algorithms can be implemented. It should be understood that the disclosed error mitigation algorithms can also be applied when other sequencing approaches are used, such as additive or subtractive approaches.

[0154] for example,

number

number

[0155] When using the SMAS device 100, FLR errors can be detected and removed, whether in real time during the sequencing procedure or at some point thereafter. FLR errors can be detected by obtaining the characteristics of each of the S-sensors 105 after cleaving and rinsing the labels. FNI errors can be detected by examining the records of each sensor 105 and identifying interrogation cycles in which that sensor 105 failed to detect the label. Thus, the modified additive approach can be adjusted to add these detection steps, according to one embodiment, as follows:

[0156] 1. Obtain baseline characteristics for each of the multiple S sensors 105 of the SMAS device 100 (which may be all or less than all of the sensors 105 in the sensor array 110) (e.g., by measuring a baseline signal at each of the multiple S sensors 105).

[0157] 2. Introduce and incorporate the first labeled nucleotide, e.g., labeled A nucleotide. Wash away unbound labeled molecules.

[0158] 3. Interrogation step 1: Obtain a characteristic of each of the plurality of S-sensors 105 (e.g., by detecting a signal at each of the plurality of S-sensors 105) and determine whether each sensor 105 detected at least one indicator. Store the detection results for each sensor 105 in a location in the record corresponding to interrogation step 1 of the current interrogation cycle.

[0159] 4. Cut off the sign and wash it off.

[0160] 5. Determine a signature for each of the plurality S of sensors 105 that detected the label in step 3. If the signature obtained for any of those sensors 105 indicates that the sensor 105 is still detecting the label, then the chemistry was not able to cleave the label (e.g., there is an FLR error for that sensor).

[0161] 6. A second labeled nucleotide, e.g. a labeled T nucleotide, is introduced and incorporated. Unbound labeled molecules are washed away.

[0162] 7. Interrogation step 2: Obtain a characteristic of each of the plurality of S-sensors 105 (e.g., by detecting a signal at each of the plurality of S-sensors 105) and determine whether each sensor 105 detected at least one indicator. Store the detection results for each sensor 105 in a location in the record corresponding to interrogation step 2 of the current interrogation cycle.

[0163] 8. Cut off the sign and wash it off.

[0164] 9. Determine a signature for each of the plurality S of sensors 105 that detected the label in step 7. If the signature obtained for any of those sensors 105 indicates that the sensor 105 is still detecting the label, then the chemistry was not able to cleave the label (e.g., there is an FLR error for that sensor).

[0165] 10. Introduce and incorporate a third labeled nucleotide, e.g., a labeled C nucleotide. Wash away unbound labeled molecules.

[0166] 11. Interrogation step 3: Obtain a characteristic of each of the plurality of S-sensors 105 (e.g., by detecting a signal at each of the plurality of S-sensors 105) and determine whether each sensor 105 detected at least one indicator. Store the detection results for each sensor 105 in a location in the record corresponding to interrogation step 3 of the current interrogation cycle.

[0167] 12. Cut off the sign and wash it off.

[0168] 13. Determine a signature for each of the plurality S of sensors 105 that detected the label in step 11. If the signature obtained for any of those sensors 105 indicates that the sensor 105 is still detecting the label, then the chemistry was not able to cleave the label (e.g., there is an FLR error for that sensor).

[0169] 14. Introduce and incorporate a fourth labeled nucleotide, e.g., a labeled G nucleotide. Wash away unbound labeled molecules.

[0170] 15. Interrogation step 4: Obtain a characteristic of each of the plurality of S sensors 105 (e.g., by detecting a signal at each of the plurality of S sensors 105) and determine whether each sensor 105 detected at least one label. Store the detection result for each sensor 105 in a location in the record that corresponds to interrogation step 4 of the current interrogation cycle. If there are sensors 105 that do not have an assigned base for the interrogation cycle (e.g., sensors 105 that failed to detect A, T, C, or G during the interrogation cycle), the chemistry was unable to incorporate a nucleotide (e.g., FNI exists for these sensors 105).

[0171] 16. Cut off the sign and wash it off.

[0172] 17. Calculate the characteristics of each of the plurality of S sensors 105 that detected the marker in step 15. If the resulting profile for any of those sensors 105 indicates that the sensor 105 is still detecting the label, then the chemistry was not able to cleave the label (e.g., for that sensor, an FLR error is present).

[0173] Steps 1 through 17 can then be repeated for the next interrogation cycle (e.g., to estimate the next base or to re-read the current base if the previous interrogation cycle failed to read the current base). It should be understood that the particular order of steps 1 through 17 is exemplary, and furthermore, the number and numbering of steps 1 through 17 is for convenience and can be changed. As an example, as explained above, the order in which nucleotides are introduced is arbitrary. As another example, although steps 2, 6, 10, and 14 include the introduction and incorporation of nucleotides and washing away unbound nucleotides as single steps, it should be understood that each of steps 2, 6, 10, and 14 can be divided into a series of smaller steps. Similarly, steps 3, 7, 11, and 15 (interrogation steps 1, 2, 3, and 4) can be further divided into a series of smaller steps (e.g., acquiring a characteristic, determining whether a label was detected, and storing the detection result). Similarly, label 15 includes identifying FNI errors, but that task can be a separate step. Conversely, steps can also be combined (for example, some or all of steps 2 to 5, some or all of steps 6 to 9, some or all of steps 10 to 13, some or all of steps 14 to 17, etc.).

[0174] FIG. 19 is a flow diagram of an exemplary sequencing procedure 400 using a modified additive approach with FLR and FNI error detection, according to some embodiments. The sequencing procedure 400 may be, for example, a sequencing procedure performed in step 210 of an exemplary method 200 of sequencing a plurality of nucleic acid strands (e.g., ssDNA) using the SMAS device 100 shown and described in the description of FIG. 11. At 402, the sequencing procedure 400 begins. At 404, a baseline characteristic of each of the S sensors 105 is obtained (e.g., by at least one processor 130 of the SMAS device 100 with the aid of the circuitry 120). Once an interrogation cycle begins, at 406, a first labeled nucleotide is selected (e.g., with reference to steps 1-17 above, the first labeled nucleotide may be A). At 408, the selected labeled nucleotide is introduced into the fluid chamber 115, where the nucleotide potentially becomes incorporated into the nucleic acid strand bound to the binding site 116. At 410, unbound nucleotides are washed away. At 412, a characteristic is obtained from each of the plurality of S sensors, and a detection result (e.g., a detected or not detected label) is determined for each of the plurality of S sensors 105. At 414, the S detection results are recorded in S records (e.g., as a 1 indicating a label was detected or as a 0 indicating a label was not detected). At 416, the label is cut and washed away. At 418, a characteristic is obtained for the sensors 105 that detected the label during steps 412 / 414. At 420, it is determined whether any of the sensors 105 that detected the label during steps 412 / 414 are still detecting the label. If so, at 422, it is determined that an FLR error was detected for the sensor 105 that still detected at least one label, even if the label was cut and rinsed away at 416. The sequencing procedure 400 then continues to 424. If, in 420, it is determined (eg, by at least one processor 130) that none of the sensors 105 that detected a label during steps 412 / 414 have yet detected a label, the sequencing procedure also continues to 424.At 424, it is determined whether the last tested nucleotide was the last nucleotide of the query cycle. For the example order of nucleotide testing assumed in steps 1-17 above, it is determined (e.g., by at least one processor 130) at 368 whether G was the last tested nucleotide. If not, at 426, the next labeled nucleotide to be tested in the query cycle is selected, and at 424, the last tested nucleotide is determined to be the last nucleotide of the query cycle. Steps 408-420 (and 422, if applicable) are repeated until it is determined that the S sensors 105 failed to detect the label during the last completed interrogation cycle. At 428, FNI errors are detected for the S sensors 105 that failed to detect the label during the last completed interrogation cycle. At 430, it is determined (e.g., by the at least one processor 130) whether the last completed interrogation cycle is the last interrogation cycle of the sequencing procedure 400. For example, the at least one processor 130 may determine whether enough detection results have been recorded to enable the at least one processor 130 (or some other processing entity, such as an external processor) to call the target number of bases (e.g., 150 bases). If not, the sequencing procedure 400 returns to step 404. If so, the sequencing procedure 400 ends at 432. Again, as explained above, the order in which the nucleotides are tested is arbitrary.

[0175] Mitigating FNI and FLR errors To illustrate the impact of FNI and FLR errors on the CLUS and SMAS devices 100, each type of sequencer is used to call an exemplary DNA sequence in which FNI and FLR errors occur randomly when the sequence is read using the modified additive approach of SBS described above. The error rates are

number

number

[0176] FIG. 21 shows the expected signal levels detected by the CLUS device sensor that captures the behavior of molecular ensembles during the sequencing procedure. In each interrogation step, the CLUS device sensor can detect four signal intensity levels of the molecular ensemble (composed of three ssDNAs): 0 labels, 1 label, 2 labels, or 3 labels. The CLUS device sequencing procedure considers the combined signal of the ensemble and cannot distinguish when the reaction on an individual strand fails. A base is called in a particular interrogation step whenever the CLUS device sensor senses at least two labels. This threshold can be represented by a criterion, where a base is called when the CLUS sensor signal level is greater than 1.5. As FIG. 21 shows, a high rate of chemical failures results in significant base calling errors and very low base calling accuracy. The CLUS device approach results in only 6 out of 21 bases (about 29%) being called as following the true sequence. This level of accuracy is only slightly better than random guessing, which has an accuracy of 25% (because with 4 bases, there is a 1 in 4 chance of guessing the base correctly). Furthermore, the CLUS instrument cannot tell the difference between a successful and a failed chemistry, nor can it tell the location of the FNI (dashed circle) or FLR (backslashed circle) errors shown in FIG. 20. In the case of the CLUS instrument, the exact location of the FLR error is obscured by ensemble averaging. To slightly improve the quality of the base calls on the CLUS instrument, In addition, probabilistic error correction algorithms can only be implemented by essentially making educated guesses about the locations of base insertions, deletions, and substitution sites. Exemplary algorithms are described, for example, in A. Cacho et al., "A Comparison of of Base-calling Algorithms for Illumina Sequencing Technology,'Briefs in Bioinformatics, Vol.17(5), 786-795, 2016; W. C. Kao et al.,'BayesCall: A model-based base-calling algorithm for high-throughput short-read sequencing,'Genome Res., Vol.19(10), 1884-1895, 2009; and C. Ledergerber and C. Dessimoz,'Base-calling for next-generation sequencing platforms,'Brief Bioinform., Vol.12, 489-97, 2011.

[0177] FIG. 22 illustrates how the SMAS device 100 can provide better accuracy when using the error correction techniques described herein. As explained above, FLR errors that occur during a sequencing procedure can be detected during the sequencing procedure. Specifically, the SMAS device 100 knows (or can find) the location of the FLR because a characteristic of each sensor 105 (e.g., signal level) is obtained and recorded after the label is cleaved and washed away, but before the next nucleotide is introduced. The FLR error can be corrected by treating it as "label not detected" when making a base call. In other words, if the record of the sequencing procedure contains a binary (e.g., 0 / 1) entry for each interrogation step, the FLR can be corrected by changing the values ​​in those interrogation steps from a "detected" value to a "not detected" value. As a specific example, if 0 represents a label not detected and 1 represents a label detected, then prior to error correction, the FLR in the mth interrogation step is represented by a 1 at the mth position in the record. That error can be corrected by changing the value of 1 at the mth position in the record to a value of 0. The top part of Figure 22 shows the detection results of each of the three sensors 105 of the SMAS device 100 before error correction to remove the FLR errors. The bottom part of Figure 22 shows the result of correcting the FLR errors before calling the bases.

[0178] The modified additive sequencing procedure using the SMAS device 100 ensures that more than half (

number

[0179] When using SMAS device 100, FNI errors can also be corrected because the integration failure creates a distinctive signature in the detection results of SMAS sensor 105 (e.g., , in a record consisting of label detection / non-detection by the sensor 105 during the sequencing procedure). In particular, FNI errors in the modified additive approach result in runs of zeros (or other "no label detected" detection results) for four or more consecutive interrogation steps. As explained in the description of FIG. 19, some FNI errors can be detected by identifying that a particular sensor 105 did not detect a label during an interrogation cycle. It should be understood that FNI errors can also "span" multiple interrogation cycles. For example, during a first interrogation cycle that includes A?⇒T?⇒C?⇒G? interrogation steps, if a particular sensor 105 detects a label during the A? interrogation step, it will not detect any label until the C? interrogation step of the next interrogation cycle. The C? interrogation step follows the A? interrogation step in the exemplary interrogation cycle, and since the modified additive approach is used in that sequencing cycle, the C? interrogation step of the first interrogation cycle should result in a label being detected. Note that step 428 of FIG. 19 does not result in an FNI error going undetected during either the first or second interrogation cycle, since neither interrogation cycle resulted in a label being undetected by the particular sensor 105. However, inspection of the detection result record reveals the presence of an FNI error. An FNI error can be deterministically corrected by deleting a run of zeros (four in the case of DNA sequencing) to align the bad strand with a strand that is not affected by the FNI error. FIG. 23 illustrates the correction of an FNI error by deleting a run of four "label not detected" entries in the detection result record from the sequencing procedure. As shown in FIG. 23, the FNI error correction results in a perfect alignment between the called sequence and the true sequence.

[0180] Qualitative analysis of a simplified model system with a limited set of errors suggests that the use of the SMAS device 100 for nucleic acid sequencing is significantly superior to the use of the CLUS device, at least when the number of instances K of the sequenced DNA strand is low and the chemical failure rate is high. To set up a framework for a quantitative comparison of the two platforms, we consider below how cluster size (for the CLUS device) and the number of instances sequenced (for the SMAS device 100) affect base calling accuracy. For both FNI and FLR errors,

number

number

[0181] FIG. 25 shows the effect of a large cluster size N on the base calling accuracy of the CLUS device. FIG. 25 shows the expected signal levels detected by the CLUS device sensor, which captures the behavior of the molecular ensemble during the sequencing procedure. In each interrogation step, the CLUS device sensor can detect any of the 12 signal intensity levels of the molecular ensemble (11 ssDNAs), i.e., 0 to 11 labels detected. When the signal level detected by the CLUS sensor is greater than 5.5, a base is called in a particular interrogation step. As FIG. 25 shows, the failed chemistry results in base calling errors, with only 11 of the 18 called bases (about 61%) following the true sequence.

[0182] Comparing Figure 25 with Figure 21,

number

number

number

number

[0183] FIG. 26 illustrates, in accordance with some embodiments,

number

number

[0184] Thus, when the SMAS device 100 is used with deterministic error correction, it can provide a perfect match between the true sequence and the called sequence when only FNI and FLR errors occur. Moreover, when only FNI and FLR errors occur, it is in fact possible to call an error-free sequence using only a single sensor 105, reading a single ssDNA with the deterministic error correction techniques described above (e.g., changing the FLR to "label not detected" and / or deleting a "label not detected" run of a specified length (e.g., 4) from the detection result record).

[0185] However, when FNR and / or FDL errors are introduced, using only deterministic error correction is generally unlikely to eliminate all errors in the detection result record. In order to address FNR and / or FDL errors, probabilistic error correction can be included in addition to or instead of deterministic error correction.

[0186] Mitigating FNI, FLR and FNR errors This section also includes the FNR error in the analysis. The impact of such errors on sequencing accuracy is comparable to the impact of FNI and FLR due to the averaging inherent in the CLUS device's detection of labels in clusters of instances of nucleic acids. FNR errors are highly detrimental to the performance of sequencing methods using SMAS device 100 because FNR errors cannot be deterministically corrected. (Note that FNR errors cannot be corrected by themselves at all in CLUS devices; instead, CLUS devices rely on ensemble behavior to mitigate the impact of FLR and other types of errors.)

[0187] FIG. 27 illustrates the problems introduced by FNR into an example sequence (TAG CAA GGT CCG CTA CTG GCA GAC TGG), assuming that the errors of FNI, FLR, and now also FNR, occur randomly over the course of 18 interrogation cycles of the A? → T? → C? → G? interrogation process. For purposes of example,

number

number

[0188] By applying probabilistic error correction, error correction can be improved to mitigate FNR errors in addition to FLR and FNI errors. For example, note the thymine interrogation step at position 2 (interrogation step 2 of interrogation cycle 1). Sensors S1 and S3 detect the label, but S2 does not. S2 does not detect the label because of an FNR error in both sensors S1 and S3 simultaneously, or because of an FNI error in sensor S2. If the probability of each error is r, then the probability that an FNR error occurred simultaneously in both sensors S1 and S3 is r. 2and the probability of an FNI error for sensor S2 is r. An error correction algorithm (e.g., executed by at least one processor 130 or another processor) assumes that the more likely event occurred (that there was an FNI error in sensor S2) and removes all entries in positions 2-5 from the data record capturing the detection results from sensor S2, shifting the S2 detection results in the S2 record. As a result, the detection results in the S2 record are realigned with the detection results generated by sensors S1 and S3, as shown in the upper portion of FIG. 30 labeled "A." Position 4 ("A" in FIG. 30) The previous (before deletion) G label detection in the portion labeled ( ) can be attributed to FNR since sensors S1 and S3 do not detect the label at position 4 (interrogation step 4 of interrogation cycle 1).

[0189] As shown in the portion labeled "E" of FIG. 30, the same error correction procedure can be performed from left to right at positions 13 (labeled "B" as shown in the portion of FIG. 30), 32 (labeled "C"), and 46 (labeled "D") to show the progressive improvement of the alignment between the S1, S2, and S3 records of the detection results. The portion labeled "E" of FIG. 30 shows that the implementation of multiple probabilistic error correction steps aligns the outputs of all sensors S1, S2, and S3, but does not appear to improve the alignment between the called sequence and the true sequence. Even after error correction, only 9 out of 20 (45%) bases are correctly called. In other words, base calling errors still occur. Specifically, following the error correction procedure, all three sensors S1, S2, and S3 report the detected labels in the interrogation step where the labels are detected, but some sensors also detect the labels erroneously incorporated by FNR at positions 10, 22, 40, and 50 (shown in the continuation of FIG. 30).

[0190] When base calling occurs when more than half of the sensors 105 agree in their detection results (after error correction), a thymine insertion error occurs at sequence position 8 (interrogation step 22), and sensors S1 and S3 both detect a label attached to a non-complementary nucleotide during the same interrogation step. (It should be understood that the presence of a thymine insertion error at position 8 is known because the erroneous data was generated for illustration and is known. In one embodiment, the sensors 105 only indicate whether a label was detected during the interrogation step, and do not indicate whether the detection (or lack of detection) is correct or incorrect. Thus, in one embodiment, an error in interrogation step 22 is essentially indistinguishable from a correct detection result.) The properly aligned true and called sequences clearly display the location of the single erroneous base insertion and can be represented as follows: Error: |insert True sequence: TAG CAA G*G TCC GCT ACT GGC Sequence called: TAG CAA GTG TCC GCT ACT GGC *Insertion position

[0191] This insertion error can be corrected if the base calling rules are modified to require all three sensors S1, S2, and S3 to match. With such rules, all three sensors S1, S2, and S3 would need to simultaneously suffer an FNR error to cause an incorrect base call. The probability of such an event is r 3 Only.

number

[0192] Mitigation of FNI errors, FLR errors, FNR errors and FLD errors The general error correction strategy used in some embodiments considers and mitigates all four types of chemical errors that cause FNI, FLR, FNR, and FLD errors. Figure 31 shows an example sequence (TAG CAA GGT CCG CTA CTG GCA GAC TGG) and shows the FNI, FLR, FNR, and FLD errors. Now assume that FLD errors also occur randomly during the 18 interrogation cycles of the A? ⇒ T? ⇒ C? ⇒ G? interrogation process. For the purpose of creating many errors in the sequencing data to provide a vehicle for illustrating an exemplary error correction procedure, we assume a very high average error rate of 1 in 5 failed reactions (

number

[0193] Under the exemplary conditions and assumptions made here, simply considering the data records made by SBS using the SMAS device 100, it is not possible to distinguish between correct nucleotide incorporation and FNR, nor between correct nucleotide non-incorporation and FNI. FLR errors can be detected and corrected deterministically as previously described (by checking the sensor 105 after cleaving and washing away the label, and treating FLR as "label not detected"), but FNR errors cannot be identified because they cannot be distinguished from correct detection events, and FNI and FLD errors cannot be identified because they cannot be distinguished from correct nucleotide non-incorporation. Nevertheless, error mitigation can still be achieved using probabilistic error correction techniques. For example, as explained above, if fewer than all sensors S1, S2, and S3 detect or do not detect a label during a particular interrogation step, the probability of two (or more) events can be calculated, the event with the highest probability can be assumed to be correct, and appropriate error correction steps can be taken.

[0194] FIG. 32 illustrates the application of the error correction procedure to data acquired during SBS under the above conditions and assumptions. The portion of FIG. 32 labeled "A" is the raw data before FLR error removal. Assuming that the signal level of sensor 105 is checked after the label is cleaved and rinsed off as described above, the location of the FLR error is known. The FLR error can be completely eliminated using deterministic error correction, i.e., by changing the "label detected" value (e.g., 1 or "yes") to a "label not detected" (e.g., 0 or "no") value in the data record at the location corresponding to the interrogation step where the FLR error was detected. Note that during interrogation cycle 15 shown in FIG. 31, an FLD error is followed by an FLR error in the data for sensor S2. In other words, sensor S2 was unable to detect the label of the incorporated nucleotide during the first interrogation step of cycle 15. If the label is cleaved after the first interrogation step of cycle 15 and before the second interrogation step of cycle 15, the signal level of sensor S2 is confirmed. This check reveals the presence of a label on sensor S2, which is known to be an FLR error because all labels should have been disconnected and rinsed off after the last interrogation step, so even if an FLR error follows another error, it is detectable and can be removed.

[0195] The portion of FIG. 32 labeled "B" shows a record of the detection results after removal of the FLR error by deterministic error correction, applied as described above. The data record shown in "B" is now shown

number

[0196] To illustrate how probabilistic error correction can be applied, the following table shows the data records of FIG. 32 for the first five interrogation cycles (interrogation steps 1-20) of three sensors S1, S2, and S3 after the FLR error has been removed (e.g., from the record labeled "B" in FIG. 32). In other words, the table below shows the first 20 detection results after deterministic error correction to remove the FLR error. For interrogation steps in which the sensor detected a landmark, the table contains a value of 1, and for interrogation cycles in which the sensor did not detect a landmark, the table contains a value of 0. [Table 2]

[0197] As explained above, simple majority voting after removal of FLR errors results in only 8 of the 17 bases being called correctly, as shown in the portion of Figure 32 labeled "B." Probabilistic error correction can provide significant improvement, as will be explained below.

[0198] Take interrogation step 2 as an example: sensors S1 and S3 both detected the label (entry 1s in the table above), but sensor S2 did not (table entry 0). Thus, either sensors S1 and S3 are both incorrect, or sensor S2 is incorrect. By taking into account the probability of various events that can result in each of these outcomes, error correction algorithms can mitigate errors in the sequencing data. Specifically, because the FLR has been removed from the data record, the only way that sensors S1 and S3 can both incorrectly detect a label during interrogation step 2 is if they both suffered an FNR error during that interrogation step. If the probability of an FNR error is r, then the probability that sensors S1 and S3 both suffer an FNR error during a single interrogation step is r. 2 For the purposes of this example,

number

[0199] If sensor S2 is wrong, it is because it failed to detect the label due to either an FLD error or an FNI error. Recall that an FLD error occurs when the correct complementary nucleotide is incorporated but it lacks a label or the sensor is unable to detect it, and an FNI error occurs when the correct complementary nucleotide is never incorporated during a sequencing cycle. FLD and FNI errors are mutually exclusive (i.e., a sensor can only suffer one of them at a time, never both). Thus, if the probability of each type of error is r, then the probability that sensor S2 suffered either an FLD or an FNI error is 2r. In our example,

number

number

[0200] As explained above, sensor S2 may be in error due to either an FLD error or an FNI error. Following an FLD error, the DNA strand sensed by sensor S2 remains "in sync" or "aligned" with the DNA strands sensed by sensors S1 and S3. In other words, if interrogation step m sequences the 40th base of the DNA strands sensed by each of sensors S1, S2, and S3, then interrogation step m

number

[0201] In some embodiments, the action taken by the error correction algorithm depends in part on the examination of the candidate error correction data assuming that each of the two types of errors has occurred separately. In other words, the detection result record can be modified to correct the error assuming that it was caused by the FLD error to generate a first candidate correction data record, and the detection result record can be modified to correct the error assuming that it was caused by the FNI error to generate a second candidate correction data record. The two candidate correction data records can then be examined and / or analyzed and / or compared to determine which one is more likely to be correct. To correct the FLD error, the "sign not detected" indication is flipped to a "sign detected" indication. To correct the FNI error, the data entry is shifted by four places (e.g., to the left as the data record is presented in the example herein).

[0202] To illustrate the specific example of interrogation step 2 in an exemplary data record, the first candidate correction data record, Option A, affects the output of sensor S2, assuming the (presumed) error was an FLD error. The presumed error is corrected by flipping the bit for interrogation step 2 in the record of sensor S2 from 0 to 1 with the bold underlined value "1", as shown in the table for Option A below. [Table 3]

[0203] The second candidate correction data record, Option B, assumes that the error affecting the output of sensor S2 was an FNI error. That presumed error is corrected by deleting the data recorded during interrogation steps 2, 3, 4, and 5 from the sensor S2 data entry and "resynchronizing" or "realigning" the data record corresponding to sensor S2 with the data records of sensors S1 and S3, resulting in the following table (values ​​previously in locations 21-24 shift to locations 17-20): The entries in the table for Option B that have been corrected by the error correction algorithm are shown in bold, underlined type. [Table 4]

[0204] Options A and B can then be compared and / or analyzed to determine which is more likely to be correct, and one of the options may be discarded. For example, a processor (e.g., at least one processor 130 or another processor) can determine a value of a metric for each candidate corrected data record and determine whether option A or option B is more likely to be correct based at least in part on the comparison of the metrics. One example of a metric is the number of interrogation steps starting from the one after the current interrogation step that has now been corrected, where interrogation step J is the further away location in the data record where the indicia detection results of all three (or, more generally, K) sensors match. For example, using this metric and setting the value of J to 8, option A has a metric value of 3 and option B has a metric value of 6. In some embodiments, based solely on this result, it is assumed that option B is more likely to be correct and option A is discarded because option B has a metric value significantly greater than option A's metric value. In some embodiments, one of the two options is discarded only if the value of its metric exceeds the value of the other option's metric by some threshold (e.g., a percentage, an amount (e.g., at least twice as much, at least 1.5 times as much, etc.), etc.). In some embodiments, option A is kept and no options are discarded until later.

[0205] In some embodiments, the contribution to the metric value is weighted based on the distance of the data being considered from the currently corrected current interrogation step. For example, because the likelihood of additional errors introduced into the data record increases as more bases are sequenced (e.g., the likelihood that some error occurs for one of the K sensors between interrogation steps 3 and 40 is greater than the likelihood that some error occurs for one of the K sensors between interrogation steps 3 and 6), the metric may assume that closer data entries are more likely to be correct than more distant data entries, and thus give more weight to data entries closer to the currently corrected data entry than to more distant data entries. The weighting may be, for example, For example, the weighting may be linear or non-linear. By way of example only, for a metric having contributions from data up to 12 query steps away, the contribution from a query step within 4 query steps of the currently corrected data may be given a weight of 1, the contribution from a query step between 5 and 8 query steps of the currently corrected data may be given a weight of 0.5, and the contribution from a query step between 9 and 12 query steps of the currently corrected data may be given a weight of 0.2. It should be understood that many possible metrics, with or without weighting, can be used and those provided above are merely exemplary and are not intended to be limiting.

[0206] It should also be understood that while the above metrics use the number of interrogation steps starting from the one after the current interrogation step that is currently corrected and the further interrogation step J position in the matching data record where the indicator detection results of all three (or more generally, K) sensors match, they could equally well use the number of interrogation steps starting from the one after the current interrogation step that is currently corrected and the further interrogation step J position in the data record where the indicator detection results of all three (or more generally, K) sensors do not match. In this case, a higher value of the metric indicates more inconsistency between the sensor data entries, and thus the candidate corrected data record is more likely to be correct the lower the value of the metric. As will be apparent to one skilled in the art, adjustments can be made to any weightings that are applied.

[0207] It should also be understood that it is not necessary to discard one of the possible options after the correction of an estimated error in the data record. For example, following the (estimated) correction of the (estimated) error in interrogation step 2 in the record of sensor S2, both options A and B can be retained, and further error detection and correction is performed on both in parallel. Similarly, multiple options of candidate sequences can be determined and / or evaluated / compared each time an estimated error is corrected. Running metric values ​​can be maintained for each possible option / candidate sequence at each step of the error correction procedure, and the most likely candidate sequence can be determined at some point (e.g. after all candidate options have been determined and evaluated (e.g. against each other), or after some additional number of interrogation steps, etc.).

[0208] Furthermore, in the above example, the possibility that both sensors S1 and S3 incorrectly detected the sign was immediately discarded because the probability of that event is significantly lower (given the assumptions herein) than the probability that sensor S2 is wrong, but the same procedure can instead be followed as for sensor S2. In other words, option C in interrogation step 2 can be determined assuming that both sensors S1 and S3 suffer from FNR error and sensor S2 is correct. In this case, the metric can be adjusted to take into account the likelihood of various possible outcomes (e.g., by "penalizing" the metric for option C based on the probability that both sensors S1 and S3 suffer from FNR error (e.g., multiplying the metric by the ratio of the probability that both sensors S1 and S3 are wrong to the probability that sensor S2 is wrong)).

[0209] It should be appreciated that the error correction methodology described herein can be leveraged in several ways to improve the accuracy of nucleic acid sequencing using the SMAS device 100. Assuming sufficient computational power, an implementation (e.g., using at least one processor 130, or another processor or processors) can determine and evaluate an exhaustive set of candidate sequences to which error correction has been applied, and then select from among them the candidate sequence that is most likely to be correct. To reduce computational complexity, an implementation can also make a decision during the error correction process to eliminate candidate error correction sequences (or potential error sources) that are deemed unlikely to be sufficiently correct (e.g., option C in the example above), and retain only the candidate error correction sequences that are more likely to be correct. The disclosed It will be appreciated that the flexibility of the principles described makes them suitable for error mitigation in systems with a wide variety of computational capabilities. Returning to the example above, assuming option B was the only option that was retained after error correction was applied to the data from query step 2, the corrected data would appear as follows: [Table 5]

[0210] The next interrogation step where the three sensors S1, S2, and S3 do not match is at interrogation step 5. Again, sensor S2 does not match sensors S1 and S3 in the same way as at interrogation step 2. In some embodiments, the error correction algorithm determines that (a) the probability that sensor S2 is wrong is greater than the probability that both sensors S1 and S3 are wrong, and (b) at interrogation step 5 sensor S2 has suffered either an FNI error or an FLD error. Again, two options can be made, one assuming the error is an FLD error (corrected by flipping a bit) and the other assuming the error is an FNI error (corrected by shifting the data by four places). The corrected data record is shown below: Option A (Probable FLD Error Corrected): [Table 6] Option B (Probable FNI error corrected): [Table 7]

[0211] Again, metrics for options A and B may be calculated, one of the options may be discarded, or both may be retained. As an example, assume option A is retained, resulting in the following error corrected data: [Table 8]

[0212] The next interrogation step where the sensor data do not match is interrogation step 10, where sensor S1 detected a landmark but neither sensor S2 nor sensor S3 detected one. Because FLR errors have been removed from the data record, the only way sensor S1 could erroneously detect a landmark during interrogation step 10 is if it suffered an FNR error during that interrogation step. The probability of an FNR error is r. If sensors S2 and S3 are both in error, it is because (a) they both suffered an FNI error, (b) they both suffered an FLD error, or (c) one of them suffered an FNI error and the other suffered an FLD error. The probability of any of the mutually exclusive events (a), (b), or (c) is 4r 2 Therefore, in some embodiments, it is assumed that the more likely event has occurred, i.e., sensor S1 has suffered an FNR error (assumed value of r

number

[0213] The error correction procedure can continue as described throughout the remaining data record. The portion of Figure 32 labeled "C" shows the results of an example. As shown, following application of probabilistic error correction as described above, 16 of the 20 bases (80%) are correctly called.

[0214] FIG. 33 is a flow diagram illustrating an error correction procedure 450, according to some embodiments. The error correction procedure 450 may be, for example, the error correction procedure 212 shown in FIG. 11, or may be executed by a processor (e.g., at least one processor 130 shown in FIG. 5A or FIG. 50 described below). At 452, the error correction procedure 450 begins. At 454, a plurality of records are identified in the sequencing data generated as a result of a nucleic acid sequencing procedure using the SMAS device 100. Each of the identified plurality of records includes a plurality of entries, each of which captures a detection result for one instance of a particular strand of nucleic acid. Thus, if the number of identified records is K, then each of the K records includes one entry for each detection result per interrogation step of the sequencing procedure. Each detection result indicates that during the interrogation step, either (a) a label was detected by the corresponding sensor 105, or (b) a label was not detected by the corresponding sensor 105. The plurality of records can be identified in several ways. For example, as described further below, different unique barcodes can be linked to the primer ends of nucleic acid strands to read known sequences during cycles of the sequencing procedure. Thus, multiple records can be identified by searching the sequencing data for barcodes associated with a particular strand of nucleic acid. As another example, a common sequence of entries can be identified in the sequence data (e.g., within entries that document the detection results of the first approximately 35 interrogation steps of the sequencing procedure).

[0215] At 456, a plurality of candidate sequences are determined for the particular strand of the nucleic acid based on the plurality of records. Each of the plurality of candidate sequences comprises at least a portion of the nucleic acid sequence of the particular strand of the nucleic acid, e.g., In some embodiments, determining the plurality of candidate sequences includes identifying a particular interrogation step in the plurality of records where a first sensor detected each of the labels and a second sensor did not detect either of the labels, and establishing two candidate sequences, one of which assumes that the first sensor correctly detected each of the labels and the other of which assumes that the first sensor incorrectly detected each of the labels. In some embodiments, determining the plurality of candidate sequences includes identifying a particular interrogation step in the plurality of records where a first sensor detected each of the labels and a second sensor did not detect either of the labels, and establishing two candidate sequences, one of which assumes that the second sensor did not incorrectly detect either of the labels, and the other of which assumes that the second sensor did not correctly detect either of the labels. In some embodiments, determining the plurality of candidate sequences includes identifying a series of consecutive entries (e.g., four entries) in at least one of the plurality of records indicating that the labels were not detected, and deleting the series of consecutive entries indicating that the labels were not detected from at least one of the plurality of records. In some embodiments, each of the plurality of entries is a first binary value (indicating that the label was detected) or a second binary value (indicating that the label was not detected), and determining the plurality of candidate sequences includes identifying runs of the second binary value (e.g., 4) in at least one of the plurality of records and removing runs of the second binary value from at least one of the plurality of records.

[0216] At 458, a particular candidate sequence of the plurality of candidate nucleic acid sequences is identified from among the plurality of candidate sequences as the sequence most likely to be correct. In some embodiments, identifying the particular candidate sequence of the plurality of candidate sequences most likely to be correct comprises determining or estimating which of the plurality of candidate sequences is most likely to be correct. In some embodiments, identifying the particular candidate sequence of the plurality of candidate sequences most likely to be correct comprises determining a respective metric for each of the candidate sequences and selecting the particular candidate sequence as most likely to be correct based at least in part on the respective metric and a criterion (e.g., minimum likelihood of occurrence, threshold likelihood of occurrence). In some embodiments, identifying the particular candidate sequence of the plurality of candidate sequences most likely to be correct comprises identifying a number of results for a particular interrogation step represented by a number of records (e.g., either more than half of the sensors 105 detected the label or more than half of the sensors 105 did not detect the label). In some embodiments, identifying a particular candidate sequence among the plurality of candidate sequences that is most likely to be correct includes determining a respective likelihood of occurrence for each of the plurality of candidate sequences, and selecting the particular candidate sequence based on its respective likelihood of occurrence satisfying a constraint (e.g., a minimum probability). In some embodiments, the particular candidate sequence that is most likely to occur among the candidate sequences is identified as being most likely to be correct. In some embodiments, one or more candidate sequences are eliminated based on a known constraint, such as knowledge that a particular sequence of bases is impossible. For example, it may be known from the origin or source of the nucleic acid (e.g., human) that a particular sequence of bases is impossible, and thus candidate sequences having such impossible sequences may be eliminated from further consideration.

[0217] At 460, the error correction procedure 450 ends.

[0218] It should be understood that probabilistic error correction will only be successful if the most likely scenario identified (e.g., identified at 458 in Figure 33) is in fact the correct scenario. When chemical failure rates are high, as in the examples described herein, there may be multiple scenarios that are equally likely to occur (or close to each other in probability of occurrence), in which case more sophisticated bioinformatics tools can be used. For example, candidate sequences may be identified based on knowledge of the source of the nucleic acid being sequenced (e.g., considering the source / origin of the nucleic acid). Considering the above, the error correction process may be removed (based on the knowledge that a particular sequence of bases is not possible). Nevertheless, if correctly implemented as described herein, the error correction process results in the correct alignment of the sensor 105 output. In the example shown in FIG. 32, after removal of FNI and FLR, all three sensors S1, S2, and S3 report a label in the correct detection interrogation step where the label is detected, but the sensors do not match at many interrogation positions (5, 10, 13, 20 22, 27, 32 40, 41, 48, and 50) where they either detect a label that was erroneously incorporated by FNR or cannot detect a label due to FLD. Calling a base when more than half of the sensors 105 in the aligned sequences match results in a thymine insertion at sequence position 8 (interrogation step 22) and a guanine deletion at position 13 (interrogation step 32). Properly aligned true and called sequences, clearly showing the positions of the base insertions and deletions, can be presented as follows: Error: Insertion Deletion True sequence: TAG CAA G * G TCC G CT ACT GGC Sequence called: TAG CAA G T G TCC * CT ACT GGC

[0219] As will be understood in light of the disclosure herein, accidental FNR and FLD cause insertion and deletion errors that cannot be algorithmically corrected and remain undiscovered if the true sequence is not known. In other words, if more than half of the single molecule sensors 105 in the aligned sequence give an incorrect answer, a base is miscalled. The probability of such an event depends on the rate at which chemical defects occur (the value of r). As explained above, the examples presented herein use high error rates to illustrate the application of error correction techniques. The error rate in a practical implementation must be significantly lower, thereby reducing the likelihood that the error correction procedure will not be able to correct the error. The disclosed error correction techniques can be used to properly align multiple sensor 105 outputs in the interrogation step. This can be achieved using a deep understanding of the physical origin of possible error types (e.g., knowledge that a particular sequence is not possible for the source nucleic acid), their average occurrence rates, and their signatures in the sensor sequence output. If the chemical error rate is high and the error signatures are unclear, error correction algorithms can be computationally intensive and difficult to implement. The following discussion explains how the probability of an incorrect base call depends on the read length, the cluster size N (for CLUS devices), the number of sensors K that sense instances of the same nucleic acid strand (for SMAS devices 100), and the failed chemical error rate.

[0220] Typical quantitative results from a cluster sequencer A simple quantitative model is developed here to estimate the probability of an incorrect base call in a cluster sequencer using the modified additive sequencing protocol introduced above. The different types of errors (FNI, FLR, FNR, and FLD) are expressed as a function of the ratio r (here:

number

number

number

[0221] This background signal is generated by out-of-phase nucleic acid strands that incorporate non-complementary nucleotides at the in-phase positions of the ensemble average.

number

number

number

[0222] As shown by Figures 34A and 34B, during the initial sequencing query (C small), the 〈1〉 and 〈0〉 states are well separated, but they rapidly approach a mean value N / 2 according to the functional form expressed by Equations 1(a) and (b). Also, because error occurrences are random independent events, the measured signals of the two states are discretely distributed around their ensemble mean values ​​〈1〉 and 〈0〉. Specifically, the probability that the measured ON state intensity of cluster size N is k when the ensemble mean is 〈1〉 is given by the Poisson distribution:

number

number

number

number

number

number

number

number

number

number

number

[0223] FIG. 37A shows

number

number

number

number

number

number

number

[0224] In general, P C、N、r The probability of an incorrect base call in a sequencing query number C for a cluster size N and chemical failure rate r, denoted as

number

number

number

number

[0225] Figures 38A and 38B plot Equations 4(a) and 4(b) as a function of C for various combinations of N and r.

number

number

number

number

[0226] Figure 39 shows the position 150

number

[0227] Currently, the benchmark in the sequencing industry is the ability to read 150 consecutive bases with a 1 in 1,000 chance of making an incorrect base call at position 150. This is commonly referred to as Q30, but for detecting rare mutations in high-precision diagnostics, a much larger sequencing quality factor of Q40 and even Q50 with longer read lengths is desirable. P in Equation 3(a) and (b) C、N、r A general expression for CNr can be used to fully explore the CNr parameter space and estimate the error tolerance and cluster size requirements of any sequencing metric.

number

number

number

[0228] Figure 40A shows all

number

number

number

number

number

number

[0229] Finally, the cumulative probability of an incorrect base call at position 150 (in some embodiments the target read length) is less than 1 in 100.

number

number

number

number

number

number

number

number

number

number

[0230] Typical quantitative results from single molecule array sequencers To compare the CLUS and SMAS platforms, a simple quantitative model is developed to estimate the probability of an incorrect base call in the SMAS device 100. Unlike the ensemble case applicable to the CLUS device (discussed above), where little or no error correction can be performed, the SMAS device 100 performs detections corresponding to individual nucleic acid molecules. The ability to individually sequence and record the output allows for the development and implementation of powerful techniques to identify and eliminate at least some of the errors in the resulting data record(s). As disclosed herein, one or more error correction techniques can be applied to the data generated from the sequencing procedure (e.g., SBS) before base calling is performed to identify and correct errors in the detection results and improve the accuracy of the called sequence. In particular, alignment of the detection results from multiple sensors 105 during some or all of the interrogation steps of the sequencing procedure can be improved. Even if an error correction algorithm is successful in correctly aligning the multiple sensor detection results, incorrect base calls can still be made. As explained above, occasional FNR and FLD errors can cause insertion and deletion errors that may not be corrected. Depending on the number of errors in the data record (determined in part by the chemical failure rate), the error correction process can be complex and computationally intensive, but it will be understood that modern processors have sufficient computing power to perform even the most computationally intensive of the disclosed techniques.

[0231] In what follows, we consider the general case of K single molecule sensors 105 of a SMAS device 100, each capable of monitoring a single instance of clonal DNA. Similar to the analysis of the CLUS device above, we assume that four types of errors (FNI, FLR, FNR and FLD) occur randomly during the sequencing procedure and are distributed throughout the interrogation process.

[0232] As described above, in some embodiments, a probabilistic error correction algorithm is implemented (e.g., by at least one processor 130, which may be included in the SMAS device 100 or may be external to the SMAS device 100). In some embodiments, the probabilistic error correction algorithm improves the alignment of the detection results of at least some of the sensors 105 in the data record. In some embodiments, some or all of the error correction algorithm is implemented after some or all of the interrogation steps are completed and some or all of the data are acquired. As previously described, the error correction procedure essentially eliminates FNI and FLR, as well as some FLD. The algorithmic realignment of the detection results of the sensors 105 also makes the probability of making an incorrect base call independent of the number C of interrogation steps. Also, since the error correction algorithm realigns the detection results of at least some of the sensors 105 in the data record(s), thereby correcting at least some of the errors, the effective error rate is smaller than that of rCLUS. After application of the exemplary error correction algorithm, in some embodiments, a base is miscalled only if more than half of the K sensors 105 in the algorithmically aligned sequence give an incorrect result.

[0233] The probability of making an incorrect base call (P K、r ) is simply a function of (a) the number K of sensors 105 that sequence instances of the same nucleic acid molecule (which may be less than all sensors 105 in the sensor array 110), and (b) the chemical failure rate r. Similar to the approach taken in the analysis of the CLUS device described above, the value of K is restricted to odd values ​​to avoid cases where exactly half of the sensors 105 do not match the other half. The probability of making an incorrect base call is given by:

number

number

number

number

number

number

number

number

[0234] for example,

number

number

number

number

[0235] As was done above for the CLUS instrument, for the SMAS instrument 100, the Kr parameter space is searched below to identify regions where the probability of an incorrect base call at any query position is less than 1 in 100 (Q20), 1 in 1,000 (Q30), 1 in 10,000 (Q40), and 1 in 100,000 (Q50). K、r 42 shows the results of calculations of the Kr parameter space where the probability of an incorrect base call in Qr is less than 1 in 100 (Q20), 1 in 1,000 (Q30), 1 in 10,000 (Q40), and 1 in 100,000 (Q50). As shown in FIG. 42, if the number K of single molecule sensors 105 sensing instances of the same nucleic acid molecule is 11 and the required sequencing accuracy is Q30, the acceptable chemical failure rate is

number

[0236] As a comparison with FIG. 39 shows, the tolerable error rate for the SMAS device 100 is significantly greater than the tolerable rate for the CLUS device, but the results alone do not justify the CLUS device (P C、N、r ) is very low during the initial interrogation step, and the threshold interrogation step C th The two platforms do not compare equally, as the number of incorrect base calls (P K、r ) remains constant throughout the interrogation process, thus resulting in a large cumulative error.

[0237] A fairer way to compare the performance of the CLUS and SMAS instruments 100 is to compare the cumulative error probability of these two instrument types. Equation 5(b) above represents the cumulative error probability of the CLUS instrument. The cumulative error probability of the SMAS instrument 100 can also be derived. The probability of making an incorrect base call per interrogation step C is P K、r Since (Equation 6) the probability of making a correct call is

number

number

number

number

[0238] Figures 43A and 43B show the cumulative probability of an incorrect base call at position 150 for the CLUS and SMAS devices 100. Equation Figure 5(b) can be used, for example, to calculate the probability that a CLUS device will make an incorrect base call at any base position up to 150. Figure 43A shows the Kr parameter space for the CLUS device, showing that the cumulative probability of an incorrect base call at position 150 is less than 1 in 100 for the CLUS device.

number

number

number

number

number

number

number

number

[0239] A comparison of Figures 43A and 43B reveals that the SMAS device 100 is a potentially superior sequencing platform than the CLUS device. The SMAS device 100 can have a smaller footprint (e.g., as described in the discussion of Figures 7A, 7B, 9A, 9B, and 10) and can have a much higher error tolerance than the CLUS device. The use of the SMAS device 100 promises higher throughput, lower error rates, and longer read lengths compared to the CLUS devices, which are larger and rely on large molecular ensembles. The development of a commercially viable SMAS device 100 and / or system can use some or all of the following: (a) high-precision nanoscale fabrication of densely packed sensors 105 capable of recognizing individual labels, (b) optimization of chemical processes to reduce error rates to acceptable levels, and / or (c) availability of effective bioinformatics tools to adjust the alignment in the data records of the sequencing data from at least some of the nanoscale sensors 105 by stochastically eliminating errors.

[0240] Exemplary SMAS Sequencing Procedure

[0241] As explained above, improving the sequencing throughput of a CLUS device can be achieved by decreasing the cluster size N (thereby packing more clusters onto the device), which can be difficult if the failure rate of the sequencing chemistry is also reduced. In contrast, a possible implementation of an error-tolerant, ultra-high throughput SMAS device 100 using a large array of single molecule binding sites 116 according to some embodiments is shown below. For purposes of example, it is assumed that the SMAS device 100 sequences DNA, although it should be understood that in general any type of nucleic acid can be sequenced.

[0242] Figures 44 and 45 show an exemplary sample preparation and loading process 500 according to some embodiments. Figure 44 is a flow diagram showing the process 500, and Figure 45 shows the results of various steps of the process 500. In some embodiments, the sample preparation and loading process 500 begins at 502. At 504, DNA extraction and purification is performed, resulting in several extracted DNA fragments 505, as shown in Figure 45. At 506, an adapter complementary to a primer is ligated to one end (e.g., 3') of the extracted DNA to generate strand 507, as shown in Figure 45. At 508, PCR (or some other replicating technique) is performed to generate multiple (ideally identical) instances of the extracted strands, as shown in Figure 45 as 509. At 510, a molecular linker capable of generating strong bonds (e.g., by click chemistry) to a chemically functionalized surface (binding site 116) of the fluid chamber 115 of the SMAS device 100 is attached to the other end (e.g., 5') of the ssDNA fragment, thereby generating strand 511 shown in FIG. 45. At 512, the functionalized strand 511 is loaded into the fluid chamber 115, randomly interspersed among the binding sites 116, and binds to the binding sites. As shown in the rightmost portion of FIG. 45, each binding site 116 supports only one DNA strand. (It should be understood that while each binding site 116 can support one or less strands, there is no requirement that all binding sites 116 support a DNA strand. Fewer than all of the binding sites 116 of the SMAS device 100 may be used, whether intentionally or accidentally.) Assuming that the extracted DNA fragments 503 are different from one another, as a result of the sample preparation and loading process 500, multiple instances of each of the extracted DNA fragments 505 are present in the fluid chamber 115, but their locations are unknown. At 514, the exemplary sample preparation and loading process 500 ends.

[0243] An advantage of the exemplary sample preparation and loading process 500 is that it simplifies DNA amplification, which can be performed in bulk, off-device, using (for example) conventional PCR, before the DNA strands are added to the SMAS device 100. In contrast, when a CLUS device is used, amplification (e.g., bridge amplification) is performed only after the DNA fragments are added to the CLUS device to create an array of contiguous clusters of amplified DNA.

[0244] After the sample preparation and loading process 500 has been performed, base calling can be performed, for example, using the additive approach, the subtractive approach, or the modified additive approach introduced above. Figures 46A, 46B, and 46C show simulated detection results (sensors 105 detect labels) using the modified additive approach during three exemplary interrogation cycles (A?⇒T?⇒C?⇒G?, respectively, for a total of 12 interrogation steps) performed by an exemplary SMAS device 100 having a sensor array 110 with 20 sensors 105 (and 20 binding sites 116) arranged in 4 rows and 5 columns. Multiple instances of the four different DNA strands are randomly distributed throughout the sensor array 110, but their specific locations within the sensor array 110 and their sequences are not initially known.

[0245] FIG. 47 shows how the detection data shown in FIGS. 46A, 46B, and 46C can be sorted to call bases and reveal the positions of different DNA strands. FIG. 47 provides a table showing the output of all sensors 105 in an exemplary array in each interrogation step, and the resulting base calls resulting in the called sequences. The right portion of FIG. 47 sorts the sensors 105 to group the detection results of sensors 105 that sense instances of the same DNA strand. As shown in FIG. 47, the following four sequences are called: GCT (strand number 1), TAG (strand number 2), ACG (strand number 3), and TTA (strand number 4).

[0246] If an error (FNI, FLR, FNR or FLD) occurs during the interrogation step, some of the detection results (label detected or no label detected) will be erroneous, and as long as the identity of the sensors 105 sensing instances of the same DNA strand is determined, the above-mentioned deterministic and / or probabilistic error detection and / or correction techniques can be implemented to detect and eliminate at least some of the errors. Recall that instances of a particular DNA strand may be bound to binding sites 116 scattered throughout the fluid chamber 115, and their locations are generally not known when the sequencing process begins. Once the process begins, during each interrogation step, each of the plurality of S sensors 105 detects a label at its respective binding site 116. To perform error correction, a subgroup of S sensors 105 that are sequencing instances of the same nucleic acid strand is identified.

[0247] Consider a very large sensor array 110 (e.g., 4 billion binding sites 116 and 4 billion respective sensors 105) with 400 million different DNA strands, each about 150 bases long. This means that there are about 10 instances of each unique DNA strand randomly distributed throughout the fluid chambers 115 (and binding sites 116 and sensor array 110). Also, for the example, assume that the sequences are random. Assuming a reasonably low error rate r, after the first interrogation cycle, almost all of the binding sites 116 (and sensors 105) that hold (sensing) DNA instances starting with A have been identified, as well as binding sites that hold (sensing) T, binding sites that hold (sensing) C, and binding sites that hold (sensing) G. Approximately 10 9 The sensors 105 detect the label indicating that the first base is A, and 9 The sensors 105 detect the label indicating that the first base is T, and 9 The sensors detect the label indicating that the first base is C, and 9 This sensor detects a label that indicates that the first base is G. After the second interrogation cycle, nearly all of the binding sites 116 (and sensors 105) that hold (sense) DNA instances starting from all 16 possible combinations (AA, AT, AC, AG, TA, TT, TC, TG, CA, CT, CC, CG, GA, GT, GC, and GG) have been identified. 8 sensors detect the label indicating that the first and second bases are AA, and approximately 2.5×10 8 The sensors detect the label indicating that the first and second bases are AT, and approximately 2.5 x 10 8 sensors detect the labels that indicate that the first and second bases are AC, etc. In general, there are some number D of label detections (or, assuming a modified additive approach is used for sequencing,

number

number

number

number

[0248] While the confidence that the correct set of binding sites 116 has been identified increases with the number of interrogation steps, so does the probability of making a detection error (e.g., falsely detecting a label or falsely failing to detect a label). Multiple errors can occur during the initial interrogation cycles while binding sites 116 carrying instances of the same strand are identified. Results derived for the CLUS apparatus suggest that this is not a problem. For example, Figure 38A shows that the probability of the CLUS apparatus making an incorrect base call during the initial interrogation steps is very small, and it is only above a threshold C that the probability of error increases sharply. th Also, recall that, since the SMAS device 100 simply reports an ensemble result by summing the results of the individual sensors 105, when no error correction is applied, the base calling accuracy of the SMAS device 100 is the same as that of the CLUS device.

[0249] For example, consider the 4 billion sensor array example above, and imagine a set of 11 sensors 105 (

number

number

[0250] If the chemical error rate is expected or known to be too high, such that errors may become problematic in the first approximately 35 interrogation steps, alternative approaches can be used to help identify binding sites 116 that have instances of the same DNA strand. For example, different unique barcodes can be ligated to the primer ends of a subset of the extracted DNA so that a known sequence is read during the initial sequencing cycles. FIG. 49 illustrates the use of barcodes in sample preparation and DNA loading, according to some embodiments. As shown in FIG. 49, unique barcodes are ligated to the extracted DNA to facilitate recognition of sites that hold instances of the same DNA in the presence of sequencing errors. For example, FIG. 49 illustrates four unique DNA strands, each assigned a unique barcode (e.g., strand 1 is assigned barcode 119A, strand 2 is assigned barcode 119B, strand 3 is assigned barcode 119C, and strand 4 is assigned barcode 119D). If the barcodes are significantly different from each other, they should be easily identifiable even if the chemical failure rate is very high. As will be appreciated, the appropriate number of unique barcodes can be high for high throughput diagnostic applications.

[0251] The exemplary 4 billion sensor SMAS device 100 described herein would be considered a fairly high throughput sequencer by current standards. Such a SMAS device 100 would provide approximately 150 gigabases (Gb) of reads during a single run, comparable to the output of state-of-the-art high-end sequencing systems deployed in 2020.

[0252] It should be understood that there are many ways to implement the devices, systems, and methods disclosed herein. For example, a system for nucleic acid sequencing may consist of a single device (e.g., SMAS device 100 that includes all of the hardware and software to perform the disclosed operations) or may include SMAS device 100 and other components that together perform the disclosed operations. For example, a system may include SMAS device 100 that performs a nucleic acid sequencing procedure and stores detection results from the sequencing procedure, and at least one processor external to SMAS device 100 that performs error detection and correction and base calling on the stored detection results. and a second processor (eg, in an external computer).

[0253] FIG. 50 illustrates an exemplary system 160 according to some embodiments. The system 160 comprises (i.e., includes, but is not limited to) a fluid chamber 115, a plurality of S sensors 105, and at least one processor 130. Optionally, the system 160 comprises a memory 170 for storing records (e.g., one or more files having binary entries documenting whether each of the plurality of S sensors 105 detected or did not detect at least one label during each of a plurality of interrogation cycles) including detection results obtained during the sequencing procedure. As shown by the dashed lines in FIG. 50, when the system 160 includes the memory 170, the at least one processor 130 can be communicatively coupled to the memory 170 such that the at least one processor 130 can store data in and / or retrieve data from the memory 170.

[0254] The fluid chamber 115 comprises a plurality of S binding sites, each of which is configured to bind to one or less strands of nucleic acid to be sequenced. FIG. 50 shows four binding sites 116, but it is understood that the system 160 can comprise more or fewer binding sites 116. Each of the S sensors 105 is configured to detect a label present in the fluid chamber 115. FIG. 50 shows four sensors 105, but it is understood that the system 160 can comprise more or fewer sensors 105. When the system 160 is operating, each of the S sensors 105 detects a label attached to a nucleotide incorporated in a respective strand of nucleic acid bound to each of the binding sites 116 of the S binding sites 116. As mentioned above, the sensor 105 can be a magnetic sensor, an optical sensor, or any other type of sensor capable of detecting the label used to label the nucleotide. The fluid chamber 115, the sensor 105, and the binding sites 116 have been described in detail above. These descriptions apply to FIG. 50 and will not be repeated here.

[0255] The at least one processor 130 is configured to execute one or more machine-executable instructions, which, when executed, cause the at least one processor 130 to perform a sequencing procedure that includes multiple interrogation steps (e.g., as described in the context of any of Figs. 11, 12, 14, 16, 44). Specifically, in operation, during an interrogation step of the sequencing procedure, the at least one processor 130 acquires a respective characteristic of each of the S sensors 105 (represented by the dashed lines between the at least one processor 130 and sensors 105A, 105B, 105C, and 105D). Each characteristic indicates whether the sensor 105 has detected a label (e.g., indicative of the presence or absence of at least one label). The at least one processor 130 can interpret the acquired characteristics to determine whether the sensor 105 detects the presence of a label. Based at least in part on the acquired respective characteristics, the at least one processor 130 records whether the respective sensor has detected the presence or absence of at least one label during the interrogation step. The at least one processor 130 is also configured to perform an error correction procedure on the at least one record containing the results of the sequencing procedure. The error correction procedure may operate on some or all of the records generated by the sequencing procedure, and may operate on detection results from some or all of the interrogation steps of the sequencing procedure. For example, as described above, to apply the error correction procedure, the at least one processor may identify and apply a deterministic or probabilistic error correction to a subset of K records, each of the K records in the subset corresponding to detection results from a sensor 105 that senses an instance of the same nucleic acid strand. The sequencing procedure and the error correction procedure are described in detail above. These descriptions apply to the system and at least one processor 130 of FIG. 50 and will not be repeated here.

[0256] At least one processor 130 may be a general purpose or special purpose processor (or a set of processing cores). The sensor 105 may be implemented by a programmable logic controller (PLC) and may thus execute an array of programmed instructions to accomplish various operations relating to obtaining characteristics of the sensor 105, performing error correction procedures, and / or interacting with a user, system operator, or other system components.

[0257] At least one processor 130 of system 160 may be a single processor (e.g., in SMAS device 100) or may comprise multiple processors, which may be co-located (e.g., in SMAS device 100) or physically separated from one another. For example, a first portion of at least one processor 130 may be comprised in SMAS device 100, and a second portion of at least one processor 130 may be external to SMAS device 100. In an embodiment in which at least one processor 130 comprises a first and a second portion, the first portion may be responsible for obtaining a characteristic of sensor 105, determining based on the characteristic whether sensor 105 detected an indicator during an interrogation cycle, and recording (e.g., in memory 170) whether each of S sensors 105 detected the presence or absence of at least one indicator during an interrogation cycle, and the second portion may be responsible for obtaining a record of the detection results and performing error correction procedures. Alternatively, the first portion may be responsible for obtaining characteristics of the sensors 105, determining based on the characteristics whether each of the sensors 105 detected at least one label during an interrogation cycle, and providing an indication of whether the sensor 105 detected the label to another entity via a communication interface (e.g., a wireless or wired interface such as Ethernet, Wi-Fi, etc.). In such an implementation, the second portion of the at least one processor 130 may be responsible for obtaining a record of the detection results provided by the first portion of the at least one processor 130 (e.g., a file having binary entries documenting whether each of the plurality S sensors 105 detected or did not detect at least one label during each interrogation cycle), performing error correction procedures, and calling bases.

[0258] In the foregoing description and in the accompanying drawings, specific terms are set forth to provide a thorough understanding of the disclosed embodiments. In some cases, the terms or drawings may imply specific details that are not required to practice the invention.

[0259] Well-known components have been shown in block diagram form and / or have not been described in detail, or in some cases not been described at all, to avoid unnecessarily obscuring the present disclosure.

[0260] The section headings provided in the detailed description are merely for convenience or reference and are not intended to be limiting. The section headings do not define, limit, construe, or describe the scope or extent of such sections. In addition, while various specific embodiments have been disclosed, it will be apparent that various changes and modifications can be made without departing from the broader spirit and scope of the present disclosure. For example, any feature or aspect of an embodiment can be applied in combination with any other of the embodiments or in place of the corresponding feature or aspect.

[0261] Some of the techniques and methods disclosed herein (e.g., obtaining detection results from the sensor 105, performing error correction procedures, etc.) and / or user interfaces for configuring and managing them may be implemented by machine execution of one or more sequencing instructions (including associated data necessary for proper execution of the instructions). Such instructions may be stored on one or more computer-readable media for subsequent retrieval and execution within one or more processors of a special purpose or general purpose computer system or consumer electronic device or appliance. Computer-readable media on which such instructions and data may be embodied include various These include, but are not limited to, forms of non-volatile storage media (e.g., optical, magnetic, or semiconductor storage media), as well as carrier waves that can be used to transfer such instructions and data via wireless, optical, or wired signal media, or any combination thereof. Examples of transfer of such instructions and data via a carrier wave include, but are not limited to, transfer (upload, download, email, etc.) over the Internet and / or other computer networks via one or more data transfer protocols (e.g., HTTP, FTP, SMTP, etc.).

[0262] Unless otherwise expressly defined herein, all terms should be given the broadest possible interpretation, including the meaning implied by the specification and drawings, as well as the meaning understood by a person skilled in the art and / or as defined in dictionaries, treatises, etc. As expressly set forth herein, some terms may not be consistent with their ordinary or customary meaning.

[0263] As used in this specification and the appended claims, the singular forms "a," "an," and "the" do not exclude plural referents unless specifically stated otherwise. The word "or" should be interpreted as inclusive unless specifically stated otherwise. Thus, the phrase "A or B" should be interpreted as meaning all of the following: "both A and B," "A but not B," and "B but not A." Any use of "and / or" herein does not imply that the word "or" by itself implies exclusivity.

[0264] As used in this specification and the appended claims, phrases of the form "at least one of A, B, and C," "at least one of A, B, or C," "one or more of A, B, or C," and "one or more of A, B, and C" are interchangeable and each encompass all of the following meanings: "A only," "B only," "C only," "A and B but not C," "A and C but not B," "B and C but not A," and "all of A, B, and C."

[0265] To the extent the terms "include," "having," "has," "with," and variations thereof are used in the detailed description or claims, such terms are intended to be inclusive in a similar manner to the term "comprising," i.e., to mean "including but not limited to."

[0266] The terms "exemplary" and "embodiment" are used to describe an example, not a preference or requirement.

[0267] The term "coupled" is used herein to denote both a direct connection / attachment, and a connection / attachment via one or more intervening elements or structures.

[0268] The terms "over," "under," "between," and "on" as used herein refer to the relative location of one feature with respect to another feature. For example, a feature that is disposed "on" or "under" another feature may be in direct contact with the other feature or may have intervening materials. Additionally, a feature that is disposed "between" two features may be in direct contact with the two features or may have one or more intervening features or materials. In contrast, a first feature "on" a second feature is in contact with that second feature.

[0269] The term "substantially" is used to describe a structure, configuration, dimension, etc. that is largely or approximately as stated, although manufacturing tolerances and the like may result in situations where in practice the structure, configuration, dimension, etc. is not always or necessarily exactly as stated. For example, describing two lengths as "substantially equal" means that the two lengths are the same for all practical purposes, although they may not (and need not) be exactly equal on a sufficiently small scale. As another example, a structure that is "substantially vertical" is considered to be vertical for all practical purposes, even if it is not exactly 90 degrees to the horizontal.

[0270] The drawings are not necessarily to scale, and dimensions, shapes and sizes of features may vary substantially from how they are shown in the drawings.

[0271] While specific embodiments have been disclosed, it will be apparent that various changes and modifications can be made without departing from the broader spirit and scope of the present disclosure. For example, any feature or aspect of an embodiment can be applied in combination with any other of the embodiments or in place of the corresponding feature or aspect, at least where practicable. The specification and drawings are therefore to be regarded in an illustrative and not a restrictive sense.

Claims

1. 1. A method for detecting and correcting errors made during a sequencing procedure performed using a single molecule array sequencing device comprising a plurality of sensors, each of the plurality of sensors configured to detect no more than one molecule at a time, the method comprising: detecting a label removal failure error made by a first sensor of the plurality of sensors, the label removal failure error being made during the sequencing procedure; correcting the label removal failure error in a record of a detection result from the sequencing procedure, the detection result being from the first sensor; A method comprising:

2. 2. The method of claim 1, wherein the step of detecting the failed to remove indicator error performed by the first sensor of the plurality of sensors comprises: determining that the first sensor of the plurality of sensors has detected a label after a cleavage and washing step in the sequencing procedure and prior to introducing a labeled nucleotide in a next interrogation step, the next interrogation step being the first interrogation step after the cleavage and washing steps; A method comprising:

3. 3. The method of claim 2, wherein the step of correcting the failed tag removal error comprises: The method includes, in recording the detection result, changing a value corresponding to the next inquiry step from a first value indicating detection of the marker to a second value indicating not detection of the marker.

4. 2. The method of claim 1, wherein the step of detecting the failed to remove indicator error performed by the first sensor of the plurality of sensors comprises: determining that the first sensor of the plurality of sensors detects a label after a cleavage and washing step in the sequencing procedure.

5. 2. The method of claim 1, wherein the step of detecting the failed to remove indicator error performed by the first sensor of the plurality of sensors comprises: (a) detecting a feature of the first sensor at a first time following introduction of a labeled nucleotide into a fluidic channel of the single molecule array sequencing device, the labeled nucleotide comprising a plurality of labels, the fluidic channel allowing the labeled nucleotide to be incorporated by a molecule within a detection region of the first sensor; (b) after step (a), determining that the first sensor has detected at least one indicium based at least in part on the characteristics; and (c) performing a cutting and flushing process after step (b); (d) after step (c), detecting the characteristic of the first sensor at a second time; (e) after step (d), determining that the characteristic of the first sensor detected at the second time and the characteristic of the first sensor detected at the first time are substantially identical; A method comprising:

6. 6. The method of claim 5, wherein the step of detecting the failed to remove indicator error comprises: In recording the detection result, the method includes a step of changing a value corresponding to a next interrogation step of the sequencing procedure from a first value indicating detection of a label to a second value indicating that the label was not detected, the next interrogation step being the first interrogation step after the cutting and washing steps.

7. 7. The method of claim 1, further comprising: After correcting said label removal failure errors, the method comprises estimating a sequence using probabilistic error correction.

8. 8. The method of claim 7, wherein the step of estimating the sequence using probabilistic error correction comprises: identifying a step of the sequencing procedure where a first detection result at the first sensor, a second detection result at the second sensor, and a third detection result at the third sensor are inconsistent; identifying a first error scenario in which the first detection result, the second detection result, and the third detection result coincide, where the first detection result is assumed to be correct; determining a probability that the first error scenario will occur; identifying a second error scenario consistent with the first detection result, the second detection result, and the third detection result, where the first detection result is assumed to be incorrect; determining a probability that the second error scenario will occur; modifying the detection record based on a comparison of the probability that the first error scenario occurs and the probability that the second error scenario occurs; A method comprising:

9. 9. The method of claim 8, wherein modifying the detection result record based on the comparison of the probability that the first error scenario occurs and the probability that the second error scenario occurs comprises: in response to the probability that the second error scenario occurs being greater than the probability that the first error scenario occurs; changing an entry in the record of detection results of the first sensor from a first value indicating that a sign was detected to a second value indicating that no sign was detected; changing the entry in the record of the detection results of the first sensor from the second value, indicating that no sign was detected, to the first value, indicating that a sign was detected; A method comprising:

10. 10. The method of claim 9, further comprising the step of modifying the detection result record based on the comparison of the probability that the first error scenario occurs and the probability that the second error scenario occurs, further comprising: modifying an entry in the record of detection results for at least one of the second sensor or the third sensor in response to the probability that the first error scenario occurs being higher than the probability that the second error scenario occurs.

11. 9. The method of claim 8, wherein modifying the record of detection results based on a comparison of the probability that the first error scenario occurs and the probability that the second error scenario occurs comprises removing consecutive identical entries from the record.

12. 1. A system comprising: a plurality of S binding sites, each of the S binding sites configured to bind to no more than one strand of nucleic acid to be sequenced; a plurality of S sensors configured to detect labels attached to nucleotides incorporated into strands of nucleic acid bound to the plurality of S binding sites, each of the S sensors for sensing a respective strand of nucleic acid bound to a respective binding site of the S binding sites; at least one processor configured to execute one or more machine-executable instructions that, when executed, cause the at least one processor to: (a) acquiring a characteristic of the sensor during an interrogation step of a sequencing procedure, the characteristic being indicative of the presence or absence of at least one label attached to a nucleotide incorporated into each strand of the nucleic acid bound to the respective binding site; (b) determining whether the sensor detects the at least one label attached to the nucleotide incorporated in each strand of the nucleic acid bound to the respective binding site based at least in part on the characteristics obtained in step (a); (c) detecting the characteristic of the sensor after a cutting step performed after step (a) and before a next interrogation step of the sequencing procedure; (d) indicating said label removal failure error by determining whether said sensor still detects at least one label attached to a nucleotide incorporated into each strand of said nucleic acid bound to said respective binding site based at least in part on said characteristic obtained in step (c); (e) in response to a determination in step (b) that the sensor detects at least one label attached to a nucleotide incorporated into each strand of the nucleic acid bound to the respective binding site, and in step (d) that the sensor still detects at least one label attached to a nucleotide incorporated into each strand of the nucleic acid bound to the respective binding site after the cleavage step, correcting the label removal failure error by changing the sensor's indication of "label detected" to an indication of "no label detected" in the next interrogation step after step (d); (f) repeating steps (a) through (e) for subsequent interrogation steps of a plurality of interrogation steps in the sequencing procedure; A system that executes the above.

13. 13. The system of claim 12, wherein the one or more machine-executable instructions, when executed, further cause the at least one processor to: After step (f), identifying a plurality of candidate sequences associated with the particular nucleic acid strand instance; determining a respective metric for each of the plurality of candidate sequences; selecting, based at least in part on said respective metrics and criteria, the particular candidate sequence that is most likely to be correct; A system that executes the above.

14. 14. The system of claim 13, wherein the one or more machine-executable instructions, when executed, further cause the at least one processor to: eliminating at least one of said plurality of candidate sequences based on known constraints on nucleic acid sequences in said particular nucleic acid strand.

15. 15. The system of claim 14, wherein the known constraint is knowledge indicating that the at least one of the plurality of candidate sequences cannot occur in nature.

16. 16. A system according to any one of claims 12 to 15, wherein the one or more machine-executable instructions, when executed, further cause the at least one processor to: generating a record, the record including results of the sequencing procedure for at least a subset of the S sensors, the record including a set of binary values, a first binary value indicating that at least one indicator was detected and a second binary value indicating that no indicator was detected; and when the one or more machine-executable instructions are executed, further causing the at least one processor to: identifying a consecutive number for said second binary values ​​in said record, said number being equal to a number of interrogation steps per sequencing cycle; deleting said consecutive numbers for said second binary value from said record; A system that executes the above.

17. 16. A system as claimed in any one of claims 12 to 15, wherein each sequencing cycle in the sequencing procedure has P interrogation steps, and wherein the one or more machine-executable instructions, when executed, further cause the at least one processor to: identifying a set of P consecutive indications that no indicator was detected by a first sensor of the plurality of S sensors; deleting the set of P consecutive indications that no indicator was detected by the first sensor of the plurality of S sensors; A system that executes the above.

18. 18. A system according to any one of claims 12 to 17, wherein the one or more machine-executable instructions, when executed, further cause the at least one processor to: The system performs a step of modifying at least one entry of a record based on a number of results of a particular interrogation step, said at least one entry of said record corresponding to said particular interrogation step.

19. 1. An apparatus for sequencing a nucleic acid, comprising: a fluid chamber comprising a plurality of S binding sites, each of the S binding sites configured to bind to no more than one strand of nucleic acid to be sequenced; a plurality of S magnetic sensors configured to detect labels present in the fluid chamber, each of the S magnetic sensors for sensing a respective strand of nucleic acid bound to a respective binding site of the plurality of S binding sites; at least one processor configured to execute one or more machine-executable instructions; the instructions, when executed, cause the at least one processor to, for each magnetic sensor of the plurality of S magnetic sensors: (a) acquiring a characteristic of the magnetic sensor during an interrogation step of a sequencing procedure, the characteristic being indicative of the presence or absence of at least one label attached to a nucleotide incorporated into each strand of the nucleic acid bound to the respective binding site; (b) determining whether the magnetic sensor detects at least one label attached to a nucleotide incorporated into each strand of the nucleic acid bound to each binding site based at least in part on the characteristics obtained in step (a); (c) detecting a characteristic of the magnetic sensor after a cutting step performed after step (a) and before a next interrogation step of the sequencing procedure; (d) determining, based at least in part on the characteristics obtained in step (c), whether the magnetic sensor still detects at least one label attached to a nucleotide incorporated into each strand of the nucleic acid bound to the respective binding site, thereby indicating a label removal failure error; (e) in response to a determination in step (b) that the magnetic sensor detects at least one label attached to a nucleotide incorporated into each strand of the nucleic acid bound to the respective binding site, and in response to a determination in step (d) that the magnetic sensor still detects at least one label attached to a nucleotide incorporated into each strand of the nucleic acid bound to the respective binding site after the cleavage step, correcting the label removal failure error by changing the magnetic sensor's indication of "label detected" to an indication of "no label detected" in the next interrogation step after step (d); A device that performs the above.

20. 20. The apparatus of claim 19, wherein the step of determining whether the magnetic sensor detects at least one landmark comprises: determining whether the obtained characteristic of the magnetic sensor meets or exceeds a threshold; comparing the obtained characteristic of the magnetic sensor with a previously detected value of the characteristic; 13. An apparatus comprising:

21. 21. The apparatus of claim 19 or 20, wherein the one or more machine-executable instructions, when executed by the at least one processor, further cause the at least one processor to:

2. An apparatus for causing a signal processor to perform a step of performing an error correction procedure on at least one record, the at least one record including results of the sequencing procedure for at least a subset of the plurality of S magnetic sensors at each of a plurality of M interrogation steps.

22. 22. The apparatus of claim 21, wherein the step of performing the error correction procedure on the at least one record comprises: identifying a plurality of candidate sequences associated with a particular instance of a nucleic acid strand based at least in part on said at least one record; determining or predicting which of said plurality of candidate sequences is most likely correct; 13. An apparatus comprising:

23. 23. The apparatus of claim 22, wherein the step of determining or estimating which of the plurality of candidate sequences is most likely correct comprises: determining a respective metric for each of the plurality of candidate sequences; selecting, based at least in part on said respective metrics and criteria, the particular candidate sequence that is most likely to be correct; 13. An apparatus comprising:

24. 24. The apparatus of claim 22 or 23, wherein the step of determining or estimating which of the plurality of candidate sequences is most likely to be correct comprises: eliminating at least one of said plurality of candidate sequences based on known constraints on nucleic acid sequences in said particular nucleic acid strand.

25. 25. The apparatus of claim 24, wherein the known constraint is that a particular base sequence in the at least one of the plurality of candidate sequences cannot occur in nature.

26. 26. Apparatus according to any one of claims 21 to 25, wherein each sequencing cycle in the sequencing procedure comprises P interrogation steps, and the step of performing the error detection procedure on the at least one record comprises: identifying a set of P consecutive indications in said at least one record that no label was detected; deleting said set of P consecutive indications that no label was detected from said at least one record; 13. An apparatus comprising:

27. 27. Apparatus according to any one of claims 21 to 26, wherein the step of performing the error detection procedure on the at least one record comprises: modifying at least one entry of said at least one record based on a number of results of a particular interrogation step, said at least one entry of said at least one record corresponding to said particular interrogation step.

28. 20. A method of sequencing a plurality of S nucleic acid strands using the apparatus of claim 19, comprising the steps of: binding the plurality of S nucleic acid strands to the S binding sites; performing a sequencing procedure comprising M interrogation steps to capture M detection results at a respective one of the plurality of S magnetic sensors, each of the M detection results indicating whether the respective one of the plurality of S magnetic sensors detected at least one label in the fluid chamber during a respective one of the M interrogation steps; and wherein performing the sequencing procedure includes, by the at least one processor, executing the one or more machine-executable instructions that cause the at least one processor to perform steps (a) to (e) for each magnetic sensor of the plurality of S magnetic sensors.

29. 29. The method of claim 28, wherein each sequencing cycle in the sequencing procedure comprises P interrogation steps, the method further comprising: generating S records, each of the S records capturing results of the sequencing procedure for a respective one of the S sensors at each of the M interrogation steps; identifying a set of P consecutive indications in at least one record of at least a subset of the S records that no label was detected; deleting said set of P consecutive indications that no label was detected from said at least one record; A method comprising:

30. 30. The method of claim 29, further comprising correcting one or more of the at least a subset of the S records based at least in part on a probability of a nucleotide incorporation failure (FNI) error, a nucleotide removal failure (FNR) error, and / or a label detection failure (FLD) error.

31. 31. The method of claim 29 or 30, wherein the at least one subset of the S records comprises an odd number of at least three records representing sequencing results of instances of a first nucleic acid strand, the method further comprising: identifying a number of detection results for a particular interrogation step in said at least one subset of said S records; calling or not calling bases of the first nucleic acid strand based at least in part on the multiple detection results for a particular interrogation step; A method comprising:

32. 32. The method of claim 29, further comprising the steps of: for a selected one of the M detection results: calling the at least one base of the plurality of S nucleic acid strands in response to the selected detection result in more than half of the at least one subset of the S records indicating detection of the at least one label in the fluid chamber.

33. 13. A method of reducing errors in sequencing data generated as a result of a nucleic acid sequencing procedure using a single molecule sensor array, the single molecule sensor array having a plurality of sensors, each of the plurality of sensors associated with a respective binding site of a plurality of binding sites, each of the plurality of binding sites configured to bind to no more than one strand of nucleic acid to be sequenced, the sequencing data indicating for each sensor of the plurality of sensors (i) whether the sensor detected at least one label in each interrogation step, and (ii) whether the sensor continued to detect at least one label after a cleavage step performed during the interrogation step, the method comprising: Identifying a plurality of records in the sequencing data, each of the plurality of records capturing a respective sequencing result for a respective instance of a first strand of nucleic acid, each of the plurality of records having a plurality of entries, each of the plurality of entries indicating, for a respective one of a plurality of interrogation steps of the nucleic acid sequencing procedure, either (a) that a label was detected by a respective sensor associated with the respective instance of the first strand of nucleic acid, or (b) that no label was detected by the respective sensor associated with the respective instance of the first strand of nucleic acid; identifying at least one failed to remove marker (FLR) error in the plurality of records by identifying at least one interrogation step immediately after the sensor continues to detect at least one marker after the disconnection step and after the sensor detects at least one marker during a immediately preceding interrogation step; correcting the at least one FLR error by changing at least one entry in the plurality of records from an "indicator detected" entry to an "no indicator detected" entry, the at least one entry corresponding to the at least one interrogation step immediately after the disconnection step and after the sensor detected the at least one indicator during the immediately preceding interrogation step, thereby generating a corrected plurality of records; determining a plurality of candidate sequences for the first strand of the nucleic acid based on the plurality of corrected records, each of the plurality of candidate sequences predicting at least a portion of a nucleic acid sequence of the first strand of the nucleic acid; identifying, from among said plurality of candidate sequences, a particular candidate sequence of said plurality of candidate sequences that is most likely to be correct as said at least a portion of said nucleic acid sequence of said first strand of nucleic acid; A method comprising:

34. 34. The method of claim 33, wherein the step of identifying a plurality of records comprises: searching the sequencing data for a barcode associated with the first strand of nucleic acid; identifying a consensus sequence of entries in each of said plurality of records; A method comprising:

35. 35. The method of claim 33 or 34, wherein determining the plurality of candidate sequences for the first strand of nucleic acid comprises: identifying particular interrogation steps within said modified plurality of records where a first sensor detected a respective indicator and a second sensor did not detect any indicator; establishing a first candidate sequence in which the first sensor should detect the landmarks and assuming that the first sensor correctly detects the respective landmark; establishing a second candidate sequence in which the first sensor assumes that it should not detect any of the labels and has erroneously detected the respective label; A method comprising:

36. 35. The method of claim 33 or 34, wherein determining the plurality of candidate sequences for the first strand of nucleic acid comprises: identifying particular interrogation steps within said modified plurality of records where a first sensor detected a respective indicator and a second sensor did not detect any indicator; establishing a first candidate sequence assuming that the second sensor should detect a label and failed to erroneously detect any labels; Establishing a second candidate sequence assuming that the second sensor should not detect the label and fails to correctly detect any label; A method comprising:

37. 37. The method of claim 33, wherein each sequencing cycle in the sequencing procedure has P interrogation steps, and determining the plurality of candidate sequences for the first strand of nucleic acid comprises: identifying a set of P consecutive entries in at least one of the modified plurality of records that indicate that no indicators were detected; deleting the set of P consecutive entries indicating that no indicator was detected from the at least one of the modified plurality of records; A method comprising:

38. 38. The method of any of claims 33 to 37, wherein the at least a portion of the nucleic acid sequence of the first strand of nucleic acid is a single base, and identifying the particular candidate sequence of the plurality of candidate sequences that is most likely to be correct comprises identifying a number of results for a particular query step represented by the corrected plurality of records.

39. 39. The method of any of claims 33 to 38, wherein the step of identifying the particular candidate sequence from the plurality of candidate sequences that is most likely to be correct comprises: determining a respective metric for each of the plurality of candidate sequences; selecting, based at least in part on said respective metrics and criteria, the particular candidate sequence that is most likely to be correct; A method comprising: