High throughput nucleic acid sequencing with single molecule sensor array
Through a single-molecular array sequencing (SMAS) device and system, combined with sensor arrays and error correction methods, the problem of high error rates in existing DNA sequencing is solved, and nucleic acid sequencing with higher throughput and longer reads is achieved, improving the accuracy of sequencing.
Patent Information
- Application Number
- CN202510641465.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2020-04-21
- Filing Date
- 2021-04-21
- Publication Date
- 2025-08-19
AI Technical Summary
Existing DNA sequencing methods have high error rates, limiting the read length, especially in diagnostic applications, which are difficult to achieve high accuracy.
Single-molecular array sequencing (SMAS) devices and systems are used to detect the labeling of a single nucleotide using sensor arrays, and combined with error correction methods, synthesized sequencing (SBS) is performed through SMAS devices and systems to alleviate errors in the sequencing process.
Nucleic acid sequencing with higher throughput, lower error rate and longer read length is achieved, improving the accuracy and reliability of sequencing.
Smart Images

Figure CN120502365A_ABST
Abstract
Description
[0001] This application is a divisional application of the invention patent application with the application date of April 21, 2021, application number "202180034742.4", and name "High-throughput nucleic acid sequencing with single-molecule sensor array".
[0002] Cross-reference to related applications
[0003] This application claims priority to U.S. Provisional Application No. 63 / 013,236, filed on April 21, 2020, and entitled “HIGH-THROUGHPUT DNA SEQUENCING WITH SINGLE-MOLECULE SENSOR-ARRAYS” (Agent Docket No. ROA-1002P-US / P36083-US), and this application incorporates the contents of U.S. Provisional Application No. 63 / 013,236 by reference in its entirety. This application also incorporates by reference for all purposes PCT Application No. PCT / US20 / 27290, filed on April 8, 2020, entitled “NUCLEIC ACID SEQUNCING BY SYNTHESIS USING MAGNETIC SENSOR ARRAYS” (Attorney Docket No. ROA-1000-WO / P35097-WO), published on October 15, 2020 as WO 2020 / 210370, and PCT Application No. PCT / US20 / 27290, filed on April 8, 2020, entitled “NUCLEIC ACID SEQUNCING BY SYNTHESIS USING MAGNETIC SENSOR ARRAYS” (Attorney Docket No. ROA-1000-WO / P35097-WO), published on October 15, 2020 as WO 2020 / 210370, and PCT Application No. PCT / US20 / 27290, filed on March 7, 2021, entitled “MAGNETIC SENSOR ARRAYS FOR NUCLEIC ACID SEQUNCING AND METHODS OF MAKING AND USING The full text of PCT application No. PCT / US2021 / 021274 (hereinafter referred to as "THEM") (attorney docket No. ROA-1001-WO / P35967-WO) is available for downloading. Technical Field
[0004] The present application generally relates to high-throughput nucleic acid sequencing with single-molecule sensor arrays. Background Art
[0005] Commercially successful DNA sequencing methods involve the synthesis and analysis of clonal deoxyribonucleic acid (DNA) clusters or the detection of individual DNA molecules. Although cluster sequencers exhibit sufficiently low error rates for diagnostic applications, their read lengths are significantly limited due to the propagation of errors in molecular ensembles. Single-molecule sequencers can produce significantly longer reads, but typically exhibit static and dynamic heterogeneity that results in errors that are too large for high-precision diagnostics.
[0006] Therefore, there is a need to improve DNA sequencing, and nucleic acid sequencing in general, to achieve longer reads with lower error rates. Summary of the Invention
[0007] This summary represents non-limiting embodiments of the present disclosure.
[0008] Disclosed herein are embodiments of single molecule array sequencing (SMAS) devices and systems. Each sensor in a plurality of sensors within a sensor array of a SMAS device detects a label attached to a nucleotide incorporated into a single nucleic acid strand bound to a respective binding site. Each sensor can detect a single label (e.g., fluorescent, magnetic, organometallic, charged molecule, etc.) attached to the incorporated nucleotide. Also disclosed are methods for performing highly tunable nucleic acid (e.g., DNA) sequencing based on synthesis sequencing (SBS) of multiple instances of clonally amplified DNA immobilized on such a SMAS device using the SMAS device and system. Also disclosed are error correction methods that mitigate errors (e.g., detection or non-detection of erroneous labels) generated in sequencing individual nucleic acid strands.
[0009] In some embodiments, a device for sequencing nucleic acids includes a fluid chamber, a plurality of S magnetic sensors configured to detect a tag present in the fluid chamber, and at least one processor. The fluid chamber includes a plurality of S binding sites, each of the S binding sites being configured to bind to no more than one nucleic acid chain. Each of the S magnetic sensors senses a respective chain of nucleic acid bound to a respective binding site of the S binding sites. The at least one processor is configured to execute one or more machine-executable instructions that, when executed, cause the at least one processor to (a) obtain a respective characteristic of a respective magnetic sensor for each of the S magnetic sensors in each of the M query steps of the sequencer, wherein the respective characteristic indicates the presence or absence of at least one tag, and (b) determine, at least in part based on the respective characteristic obtained, whether the respective magnetic sensor detected the presence or absence of at least one tag during the query step.
[0010] In some embodiments, a system comprises a plurality of S binding sites (each of the S binding sites being configured to bind no more than one nucleic acid strand), a plurality of S sensors (e.g., magnetic, optical, etc.) configured to detect a label, and at least one processor. Each of the S sensors is configured to sense a respective nucleic acid strand bound to a respective binding site of the S binding sites. The at least one processor is configured to execute one or more machine-executable instructions that, when executed, cause the at least one processor to, for each of the S sensors in each of a plurality of M query steps of a sequencing process, (a) obtain a respective characteristic of the respective sensor, wherein the respective characteristic indicates the presence or absence of at least one label, and (b) determine, based at least in part on the obtained respective characteristic, whether the respective sensor detected the presence or absence of the at least one label during the query step. Additionally, when executed, the one or more machine-executable instructions further cause the at least one processor to perform an error correction procedure on at least one record containing the results of the sequencing process for at least a subset of the S sensors in each of the M query steps.
[0011] In some embodiments, a method for sequencing a plurality of S nucleic acid chains using a SMAS device includes (a) binding the S nucleic acid chains to S binding sites, (b) performing a sequencing procedure including M query steps to generate S records, each of the S records capturing M detection results of a respective one of the S sensors, each of the M detection results indicating whether a respective one of the S sensors detected at least one marker in a fluid chamber during a respective one of the M query steps, and (c) applying an error correction procedure to at least one subset of the S records to estimate the nucleic acid sequence of at least one of the S nucleic acid chains.
[0012] Some embodiments are a method of mitigating errors in sequencing data generated by a nucleic acid sequencing procedure using a single molecule sensor array having a plurality of sensors, each of the plurality of sensors being associated with a respective binding site of a plurality of binding sites, each of the plurality of binding sites being configured to bind no more than one nucleic acid strand to be sequenced. In some such embodiments, the method includes (a) identifying a plurality of records in sequencing data, each of the plurality of records capturing a respective sequencing result for a respective instance of a first strand of nucleic acid, each of the plurality of records having a plurality of entries, each of the plurality of entries indicating a respective step of a plurality of query steps for a nucleic acid sequencing procedure, (i) detection of a label by a respective sensor associated with a respective instance of the first strand of nucleic acid, or (ii) absence of detection of the label by a respective sensor associated with a respective instance of the first strand of nucleic acid; (b) determining a plurality of candidate sequences for the first strand of nucleic acid based on the plurality of records, each of the plurality of candidate sequences estimating at least a portion of the nucleic acid sequence of the first strand of nucleic acid; and (c) identifying a particular candidate sequence from the plurality of candidate sequences as the at least a portion of the nucleic acid sequence of the first strand of nucleic acid, the particular candidate sequence being the most likely correct from the plurality of candidate sequences.
[0013] Compared to cluster-based approaches, the disclosed sequencing and error correction apparatus, systems, and methods are expected to achieve higher throughput, lower error rates, and longer read lengths. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] The objects, features and advantages of the present invention will become readily apparent from the following description of certain embodiments taken in conjunction with the accompanying drawings, in which:
[0015] Figure 1 A portion of a magnetic sensor according to some embodiments is described.
[0016] Figure 2A and 2B The resistance of a magnetoresistive (MR) sensor is described, which may be used according to some embodiments.
[0017] Figure 3A A spin torque oscillator (STO) sensor is described, which may be used in accordance with some embodiments.
[0018] Figure 3B Shows the experimental response of STO under example conditions.
[0019] Figure 3C and 3D Short nanosecond field pulses of STO are described, which may be used according to some embodiments.
[0020] Figure 4AA single sensor of a cluster sequencing device is illustrated for sensing some N clonally amplified DNA strands in its vicinity.
[0021] Figure 4B An exemplary plurality of S single-molecule sensors is illustrated, each sensor being used to monitor a respective single-stranded DNA (ssDNA) via a SMAS device according to some embodiments.
[0022] Figure 5A is a block diagram showing components of an exemplary SMAS device for nucleic acid sequencing according to some embodiments.
[0023] Figure 5B 、 5C 5D and 5D illustrate portions of an exemplary SMAS device for nucleic acid sequencing according to some embodiments.
[0024] Figure 5E Illustrates a square grid (or lattice) pattern of a sensor according to some embodiments.
[0025] Figure 6A A sensor, a DNA strand in a helical state, and a label are illustrated according to some embodiments.
[0026] Figure 6B Exemplary dimensions of sensors, long DNA strands, and labels according to some embodiments are illustrated.
[0027] Figure 7A An exemplary geometric arrangement for estimating sensor array packaging limits of a SMAS device is described, according to some embodiments.
[0028] Figure 7B The sensors of a SMAS device are illustrated arranged in a square grid, according to some embodiments.
[0029] Figure 8A and 8B Sensors of a SMAS device are illustrated arranged in a hexagonal pattern, according to some embodiments.
[0030] Figure 9A An exemplary geometric arrangement for estimating sensor array packaging limits of a SMAS device is described, according to some embodiments.
[0031] Figure 9B Sensors of a SMAS device are illustrated arranged in a hexagonal lattice, according to some embodiments.
[0032] Figure 10 Compare the density of the exemplary SMAS implementation to the current state-of-the-art cluster sequencing devices.
[0033] Figure 11An exemplary method for sequencing multiple nucleic acid strands using a SMAS device according to some embodiments is described.
[0034] Figure 12 is a flow chart of a sequencing procedure using an additive approach according to some embodiments.
[0035] Figure 13 An additive sequencing scheme according to some embodiments is described.
[0036] Figure 14 is a flow chart of a sequencing procedure using a subtractive approach according to some embodiments.
[0037] Figure 15 A subtractive sequencing scheme according to some embodiments is described.
[0038] Figure 16 is a flow chart of a sequencing procedure using a modified additive method according to some embodiments.
[0039] Figure 17 An improved additive sequencing scheme according to some embodiments is described.
[0040] Figure 18A Illustration of failed nucleotide incorporation (FNI) of a cluster sequencer.
[0041] Figure 18B Describe the FNI of the SMAS device.
[0042] Figure 18C Illustrates failed mark removal (FLR) of the cluster sequencer.
[0043] Figure 18D Demonstrating FLR of the SMAS device.
[0044] Figure 18E Illustration of failed nucleotide removal (FNR) of a cluster sequencing device.
[0045] Figure 18F The FNR of the SMAS device is described.
[0046] Figure 18G Illustration of failed nucleotide deletion (FLD) in cluster sequencing apparatus.
[0047] Figure 18H Demonstrating the FLD of the SMAS device.
[0048] Figure 19 is a flow chart of an exemplary sequencing procedure using the modified additive method with FLR and FNI error detection according to some embodiments.
[0049] Figure 20 Displays instance records with FNI and FLR errors.
[0050] Figure 21 Illustrate the expected signal levels detected by cluster sequencer sensors that capture the behavior of molecular collectives during the sequencing process.
[0051] Figure 22 Describes how SMAS devices provide better accuracy when using error correction techniques according to some embodiments.
[0052] Figure 23 FNI error correction by deleting strings of four "not detected marker" entries in the record of detection results from a sequencer is described in accordance with some embodiments.
[0053] Figure 24 The results of an exemplary SBS reaction according to some embodiments are illustrated.
[0054] Figure 25 The effect of larger cluster size on the base-calling accuracy of cluster sequencers is illustrated.
[0055] Figure 26 Deterministic error correction of FLR and FNI errors according to some embodiments is described.
[0056] Figure 27 Explain FNI, FLR, and FNR errors in the detection data.
[0057] Figure 28 FLR error correction and base calling of data generated by a SMAS device according to some embodiments are described.
[0058] Figure 29 FNI error correction and base calling of data generated by a SMAS device are described according to some embodiments.
[0059] Figure 30 Error correction and base calling of data generated by a SMAS device according to some embodiments are described.
[0060] Figure 31 Illustrate FNI, FLR, FNR, and FLD errors in exemplary detection results from a SMAS device.
[0061] Figure 32 The application of error correction procedures to data captured by a SMAS device during SBS according to some embodiments is described.
[0062] Figure 33 A flowchart illustrating an error correction process according to some embodiments is shown.
[0063] Figure 34AIllustrates the average signal intensity at the query step, where the label should be detected because a matching nucleotide was introduced and successfully incorporated.
[0064] Figure 34B Illustrate the function fit of the measured intensities from the cluster model.
[0065] Figure 35 Plot the probability function of the cluster sequencer.
[0066] Figure 36 Describe the discrete probability function of the cluster sequencer.
[0067] Figure 37A Illustrate the intensity curve of the cluster sequencer.
[0068] Figure 37B Describe the probability distribution function of the cluster sequencer.
[0069] Figure 38A and 38B Plot the probability function of the cluster sequencer.
[0070] Figure 39 The Nr parameter space of the cluster sequencer under various conditions is illustrated.
[0071] Figure 40A The calculated probabilities of cluster sequencers along the Q30 contour are shown for various Nr combinations.
[0072] Figure 40B The calculated cumulative error probability of a cluster sequencer is shown.
[0073] Figure 41 Describing the Nr parameter space for a cluster sequencer where the cumulative probability of an incorrect base call at position 150 is less than or equal to 1 in 100 1 in 1,000 1 in 10,000 and 1 in 100,000
[0074] Figure 42 Calculation results for the Kr parameter space for a SMAS device are illustrated, where the probability of incorrect base calling at each query step is less than 1 in 100 (Q20), 1 in 1,000 (Q30), 1 in 10,000 (Q40), and 1 in 100,000 (Q50), according to some embodiments.
[0075] Figure 43A and 43B The cumulative probability of an incorrect base call at position 150 is shown for a cluster sequencing device and a SMAS device, according to some embodiments.
[0076] Figure 44 and 45 An exemplary sample preparation and loading process according to some embodiments is described.
[0077] Figure 46A 、 46B 46C illustrate simulation test results of an exemplary SMAS device according to some embodiments.
[0078] Figure 47 Description of some embodiments Figure 46A 、 46B and how the detection data in 46C can be rearranged to identify bases and show the positions of different DNA strands.
[0079] Figure 48A and 48B Plotted are the calculated probabilities of incorrect base calls as a function of the number of query steps C and the chemical failure rate r.
[0080] Figure 49 Depicted is the use of barcodes in sample preparation and DNA loading according to some embodiments.
[0081] Figure 50 An exemplary system 160 is described in accordance with some embodiments.
[0082] To facilitate understanding, identical reference numerals have been used, where possible, to designate identical components that are common to the figures. It is contemplated that components disclosed in one embodiment may be beneficially employed in other embodiments without specific recitation. Furthermore, descriptions of components within the context of one figure may apply to other figures illustrating the components. DETAILED DESCRIPTION
[0083] Some of the description and examples herein are in the context of DNA sequencing, but it should be understood that the disclosure is generally applicable to nucleic acid sequencing.
[0084] Terminology and Notes
[0085] As used herein, the term "strand" refers to a single nucleic acid chain (eg, ssDNA). When referring to nucleic acids, the terms "strand" and "fragment" are used interchangeably.
[0086] As used herein, the term "plurality" means two or more, but not necessarily all. Thus, a plurality of sensors simply means at least two sensors, not necessarily all sensors in a sensor array or sequencing device / system. Similarly, a plurality of binding sites simply means at least two binding sites, not necessarily all binding sites in a sequencing device / system.
[0087] As used herein, the term "instance" when referring to a nucleic acid strand means a template nucleic acid strand or a copy thereof (e.g., produced by an amplification or replication process). Ideally, a copy of a template nucleic acid strand is identical to the template strand, but as is known in the art, copies are not necessarily identical due to replication / amplification errors. It should be understood that even if errors are introduced by the amplification procedure, the duplicates produced by amplification are still considered copies of the original nucleic acid strand. Thus, all instances of a strand are ideally identical to each other but may not be identical.
[0088] As used herein, the term "query cycle" refers to a single cycle of a nucleic acid sequencing process during which all possible nucleotides are introduced to determine which, if any, are incorporated into the sequenced strand. For example, for a DNA sequencing process, all adenines (A), thymines (T), cytosines (C), and guanines (G) are tested in a certain (arbitrary) order (which need not be the same for each query cycle). As described in detail below, depending on the selected sequencing process, more than one label may be detected per strand during a single sequencing cycle.
[0089] As used herein, the term "query step" refers to a step or set of steps in a sequencing process during which a determination is made as to whether one or more sensors of a sequencer are detecting a marker. For DNA sequencing cycles spanning all A's, T's, C's, and G's, there are four query steps per query cycle (one for each nucleotide). For a sensor in use, each query step results in a single determination as to whether the sensor is detecting a marker.
[0090] As used herein, the term "detection result" refers to a value indicating that: (a) a marker was detected during a query step or (b) a marker was not detected during a query step. In some embodiments, the detection result is a binary value (e.g., 0 or 1). The detection result can be derived from other data (e.g., a signal representing resistance, frequency, intensity, etc.; a measurement of resistance, frequency, intensity, etc.).
[0091] As used herein, the term "record" refers to a stored representation of the detection results of a single sensor. If the selected sequencer has M query steps, then after the sequencer is completed, each record will have M detection results. The records for S sensors can be stored in a single file (e.g., as a table with S rows and M columns, or S columns and M rows), or separate files can be created for each sensor's record.
[0092] As used herein, with respect to test results contained in a record, the term "string" means a consecutive sequence of identical values.
[0093] The terms "sensor" and "sensing component" are used interchangeably herein.
[0094] The variable S is used herein to refer to the number of sensors in a plurality of sensors. The S sensors may be instances sensing the same chain, or they may be instances sensing different chains.
[0095] The variable K is used herein to refer to the number of sensors in a plurality of sensors that all sense instances of the same chain.
[0096] mark
[0097] The methods for nucleic acid sequencing described herein use labeled nucleotide precursors comprising a cleavable label. These cleavable labels can be, for example, magnetic, fluorescent, organometallic, or charged molecules.
[0098] Each label can comprise, for example, a magnetic nanoparticle, such as a molecule, a superparamagnetic nanoparticle, or a ferromagnetic particle. The magnetic label can be a nanoparticle with high magnetic anisotropy. Examples of nanoparticles with high magnetic anisotropy include, but are not limited to, Fe3O4, FePt, FePd, and CoPt. To facilitate chemical binding to nucleotides, the particles can be synthesized and coated with SiO2. See, for example, M. Aslam, L. Fu, S. Li, and VP Dravid, "Silica encapsulation and magnetic properties of FePt nanoparticles," Journal of Colloid and Interface Science, Vol. 290, No. 2, October 15, 2005, pp. 444-449. Because magnetic markers of this size have a permanent magnetic moment whose direction fluctuates randomly on very short time scales, some embodiments described further below rely on sensitive sensing schemes that detect fluctuations in the magnetic field due to the presence of the magnetic marker.
[0099] Each label can comprise, for example, a fluorophore. Fluorescent labels are well known in the art and are suitable for use with the present disclosure.
[0100] The label may comprise, for example, an organometallic compound. As is understood, an organometallic compound is any member of a class of substances comprising at least one metal-carbon bond, wherein the carbon is part of an organic group. Examples of organometallic compounds include Gilman reagent (which comprises lithium and copper), Grinard reagent (which comprises magnesium), nickel tetracarbonyl, and ferrocene (which comprise a transition metal), organolithium compounds such as n-butyllithium (n-BuLi), organozinc compounds such as diethylzinc (EtZn), organotin compounds such as tributyltin hydride (BuSnH), organoborane compounds such as triethylborane (EtB), and organoaluminum compounds such as trimethylaluminum (MeAl).
[0101] Labels can comprise, for example, charged molecules.
[0102] There are a variety of methods by which a label can be attached to a nucleotide precursor and cleaved after incorporation into the nucleotide precursor. For example, a label can be attached to a base, in which case it can be chemically cleaved. As another example, a label can be attached to a phosphate, in which case it can be cleaved by a polymerase, or if attached via a linker, by cleaving the linker.
[0103] In some embodiments, the label is attached to a nitrogenous base (e.g., A, C, T, G, or a derivative) of the nucleotide precursor. Following incorporation of the nucleotide precursor and detection by a sequencing device (e.g., as described in further detail below), the label is cleaved from the incorporated nucleotide.
[0104] In some embodiments, the label is attached via a cleavable linker. Cleavable linkers are known in the art and have been described, for example, in U.S. Pat. Nos. 7,057,026 and 7,414,116, and their continuations and improvements. In some embodiments, the label is attached to the 5 position in a pyrimidine or the 7 position in a purine via a linker comprising an allyl or azide group. In other embodiments, the linker comprises a disulfide bond, an indole, or a Sieber group. The linker may further comprise one or more groups selected from alkyl (C 1-6 ) or alkoxy (C 1-6), nitro, cyano, fluoro groups or substituents with groups of similar properties. In simple terms, the linker can be cleaved by a water-soluble phosphine or a catalyst containing a phosphine-based transition metal. Other linkers and linker cleavage mechanisms are known in the art. For example, linkers comprising trityl, p-alkoxybenzyl esters and p-alkoxybenzyl amides and tert-butyloxycarbonyl (Boc) groups and acetal systems can be cleaved by proton-releasing cleavage agents under acidic conditions. Thioacetals or other sulfur-containing linkers can be cleaved using thiophilic metals (such as nickel, silver or mercury). Cleavable protecting groups can also be considered for the preparation of suitable linker molecules. Ester-containing linkers and disulfide-containing linkers can be cleaved under reducing conditions. Linkers containing triisopropylsilane (TIPS) or tert-butyldimethylsilane (TBDMS) can be cleaved in the presence of F ions. Photocleavable linkers that can be cleaved by wavelengths that do not affect other components of the reaction mixture include linkers containing O-nitrobenzyl groups. Linkers containing benzyloxycarbonyl groups can be cleaved by catalysts based on Pd.
[0105] In some embodiments, the nucleotide precursor comprises a label attached to a phosphate moiety, as described, for example, in U.S. Patent Nos. 7,405,281 and 8,058,031. Briefly, the nucleotide precursor comprises a nucleoside moiety and a chain of 3 or more phosphate groups, wherein one or more of the oxygen atoms are optionally substituted, for example, with S. The label may be attached directly or via a linker to an α, β, γ or higher phosphate group (if present). In some embodiments, the label is attached to a phosphate group via a non-covalent linker, as described, for example, in U.S. Patent No. 8,252,910. In some embodiments, the linker is a hydrocarbon selected from the group consisting of substituted or unsubstituted alkyl, substituted or unsubstituted heteroalkyl, substituted or unsubstituted aryl, substituted or unsubstituted heteroaryl, substituted or unsubstituted cycloalkyl, and substituted or unsubstituted heterocycloalkyl; see, for example, U.S. Patent No. 8,367,813. The linker may also comprise a nucleic acid chain; see, for example, U.S. Patent No. 9,464,107.
[0106] In embodiments where the label is attached to a phosphate group, the nucleotide precursor is incorporated into the nascent chain by a nucleic acid polymerase that also cleaves and releases the detectable label. In some embodiments, the label is removed by cleavage of the linker, e.g., as described in U.S. Pat. No. 9,587,275.
[0107] In some embodiments, the nucleotide precursor is a non-extendable "terminator" nucleotide, that is, a nucleotide whose 3' end is blocked by a blocking "terminator" group and cannot add the next nucleotide. The blocking group is a reversible terminator that can be removed to continue the chain synthesis process as described herein. Attaching removable blocking groups to nucleotide precursors is known in the art. See, for example, U.S. Patents Nos. 7,541,444 and 8,071,739 and their successors and improvements. In short, the blocking group may comprise an allyl group that can be cleaved by reacting with a metal-allyl complex in the presence of a phosphine or nitrogen-phosphine ligand in aqueous solution. Other examples of reversible terminator nucleotides for use in synthetic sequencing include modified nucleotides described in International Application No. PCT / US2019 / 066670, filed on December 16, 2019, and entitled “3’-protected Nucleotides,” which was published as WO / 2020 / 131759.
[0108] sensor
[0109] The characteristics and capabilities of the sensors used in the nucleic acid sequencing devices, systems, and methods described herein depend on the choice of label used. The sensor can be, for example, a magnetic sensor (to detect, for example, magnetic nanoparticles, organometallic compounds, etc.) or an optical sensor (to detect, for example, fluorophores). It should be understood that other types of sensors can be suitable for detecting various types of labels, and the examples described herein are not intended to be limiting. In general, the disclosed devices, systems, and methods can use any type of label that can be detected by a selected type of sensor, and conversely, the disclosed devices, systems, and methods can use any type of sensor that can detect the presence (and absence) of a selected type of label.
[0110] Reference numeral 105 is used herein generically for single molecule sensors, regardless of the type of those single molecule sensors (and regardless of the type of label they detect). Reference numeral 15 is used for sensors that sense nucleic acid clusters.
[0111] Magnetic sensor
[0112] Some embodiments disclosed herein use magnetic sensors to detect the presence of magnetic labels (eg, magnetic nanoparticles, organometallic complexes, charged molecules, etc.) coupled to nucleotide precursors. Figure 1 A portion of the magnetic sensor 105 is illustrated according to some embodiments. Figure 1The exemplary magnetic sensor 105 has a bottom surface 108 and a top surface 109 and includes three layers, such as two ferromagnetic layers 106A, 106B separated by a non-magnetic spacer layer 107. The non-magnetic spacer layer 107 can be, for example, a metallic material such as, for example, copper or silver, in which case the structure is called a spin valve (SV), or it can be an insulator such as, for example, aluminum oxide or magnesium oxide, in which case the structure is called a magnetic tunnel junction (MTJ). Suitable materials for the ferromagnetic layers 106A, 106B include alloys such as Co, Ni, and Fe (sometimes mixed with other elements). In some embodiments, the ferromagnetic layers 106A, 106B are engineered so that their magnetic moments are oriented in the plane of the film or perpendicular to the plane of the film. Additionally, materials can be used as shown in FIG. Figure 1 Although the three layers 106A, 106B, and 107 are deposited below and above the three layers in the structure for purposes such as interface smoothing, texturing, and protection from processes used to pattern the device incorporating the sensor 105, the active area of the magnetic sensor 105 is located within this three-layer structure. Therefore, a component in contact with the magnetic sensor 105 may be in contact with one of the three layers 106A, 106B, or 107, or it may be in contact with another portion of the magnetic sensor 105.
[0113] like Figure 2A and 2B As shown in Figure 2, the resistance of the MR sensor is proportional to 1-cos(θ), where θ is the value shown in Figure 2. Figure 1 The angle between the magnetic moments of the two ferromagnetic layers 106A and 106B in the sensor 105 is π / 2 radians, or 90 degrees relative to each other in the absence of a magnetic field. This orientation can be achieved by a number of methods known in the art. For example, one solution is to use an antiferromagnet to "fix" the magnetization direction of one of the ferromagnetic layers (106A or 106B, designated "FM1") through a process called exchange bias, and then coat the sensor with a double layer comprising an insulating layer and a permanent magnet. The insulating layer prevents electrical shorting of the magnetic sensor 105, and the permanent magnet provides a "hard bias" magnetic field perpendicular to the fixed direction of FM1, which then causes the second ferromagnetic layer (106B or 106A, designated "FM2") to rotate and produce the desired configuration. A magnetic field parallel to FM1 then causes FM2 to rotate about this 90 degree configuration, and the change in resistance results in a voltage signal that can be calibrated to measure the magnetic field acting on the magnetic sensor 105. In this way, the magnetic sensor 105 acts as a magnetic field to voltage converter.
[0114] It should be noted that although the example discussed immediately above describes the use of ferromagnets whose magnetic moments are oriented at 90 degrees relative to each other in the plane of the film, a perpendicular configuration can alternatively be achieved by orienting the magnetic moment of one of the ferromagnetic layers 106A, 106B out of the plane of the film, which orientation can be achieved using what is known as perpendicular magnetic anisotropy (PMA).
[0115] In some embodiments, the magnetic sensor 105 uses a quantum mechanical effect called spin transfer torque. In such a device, a current passing through one ferromagnetic layer 106A (or 106B) in an SV or MTJ preferentially allows electrons with a spin parallel to the magnetic moment of the layer to be transmitted, while electrons with antiparallel spins are more likely to be reflected. In this way, the current becomes spin-polarized, with more electrons of one spin type than the other spin type. This spin-polarized current then interacts with the second ferromagnetic layer 106B (or 106A), exerting a torque on the magnetic moment of the layer. This torque can, in different circumstances, cause the magnetic moment of the second ferromagnetic layer 106B (or 106A) to precess around the effective magnetic field acting on the ferromagnet, or it can cause the magnetic moment to reversibly switch between two orientations defined by the uniaxial anisotropy induced in the system. The resulting spin torque oscillator (STO) can be tuned in frequency by changing the magnetic field acting on it. Therefore, it has the ability to act as a magnetic field to frequency (or phase) converter (thereby generating an AC signal with a frequency), such as Figure 3A , which illustrates the concept of using an STO sensor. Figure 3B The experimental response of an STO through a delay detection circuit is shown when an AC magnetic field with a frequency of 1 GHz and a peak-to-peak amplitude of 5 mT is applied across the STO. This result is similar to those shown for short nanosecond field pulses. Figure 3C and 3D The results in
[15] demonstrate how these oscillators can be used as nanoscale magnetic field detectors. Further details can be found in T. Nagasawa, H. Suto, K. Kudo, T. Yang, K. Mizushima, and R. Sato, "Delay detection of frequency modulation signal from a spin-torque oscillator under a nanosecond-pulsed magnetic field," Journal of Applied Physics, vol. 111, 07C908 (2012).
[0116] Optical sensors
[0117] Some nucleic acid sequencing methods use fluorescent labels. In such methods, the sequenced nucleic acid molecules are fixed on a solid support and the binding of fluorescently labeled target molecules (e.g., nucleotides) to the molecules is monitored. An optical instrument (e.g., an excitation and reading device for fluorescence) provides light of a certain wavelength to excite the fluorescent label and detects fluorescence emitted from the label at a slightly different wavelength. Because the beam path (light path) of the excitation light must be at least partially different from the beam path (light path) of the fluorescence, excitation and emission filters (whose spectra do not significantly overlap) can be used to achieve spectral separation, and / or vertical or side illumination can be used.
[0118] Optical sensors and sequencing devices and methods using fluorescent labels, such as fluorophores, are well known in the art.
[0119] Amplification / replication
[0120] Nucleic acid sequencing devices generally rely on an amplification (or replication) process to produce a large number of nucleic acid instances from a single nucleic acid strand (e.g., instances of single-stranded DNA (ssDNA) from a DNA molecule). The polymerase chain reaction (PCR) is a well-known method for amplifying double-stranded DNA, which enables the replication of large quantities of DNA from small initial amounts.
[0121] Cluster Sequencer
[0122] Some sequencing devices (referred to herein as cluster (CLUS) devices) use amplification technology to form local clusters of many DNA chains. For example, a DNA chain is used as a template, and PCR amplification produces thousands or millions of DNA sequence instances in a local area. At least a portion of the PCR primers is fixed to a solid support, which allows the generated DNA molecules to be fixed to local clusters to form distinguishable "pure lines". The DNA clusters generated may include ssDNA. Examples of clonal amplification techniques include bridge PCR and emulsion PCR, including emulsion PCR based on microbeads. For bridge amplification, primers attached to a solid surface (such as a glass slide) are used to amplify single DNA molecules by in situ PCR to form DNA clusters. Each DNA cluster is a physically separated "pure line" composed of instances of a DNA chain. For clonal amplification based on emulsion PCR, single DNA molecules are cloned and amplified in emulsion droplets. In some methods, DNA chains are attached to microbeads inside the droplets. Clonal amplification of single molecules can also be carried out in separate micropores.
[0123] As used herein, the term "cluster" refers to a localized grouping of nucleic acid chains, ideally having identical sequences, that is produced by clonal amplification. When the nucleic acid is DNA, the cluster comprises identical DNA chains (or fragments) attached (ideally) to a solid support. For example, clusters can be generated on spots on a glass slide or attached to microbeads, microwells, or other microparticles.
[0124] The use of CLUS devices for fluorescence-based DNA sequencing is well known.
[0125] A sequencing device for nucleic acid sequencing using clusters using a magnetic sensor array is described, for example, in PCT application No. PCT / US2021 / 021274, filed on March 7, 2021, and entitled “MAGNETIC SENSORARRAYS FOR NUCLEIC ACID SEQUENCING AND METHODS OF MAKING AND USING THEM” (Agent File No. ROA-1001-WO / P35967-WO).
[0126] Figure 4A A single sensor 15 of a CLUS device is illustrated for sensing a number of N clonally amplified DNA chains 101 in its vicinity. The sensor 15 may be, for example, a magnetic sensor to sense magnetic labels attached to incorporated nucleotides. For convenience, Figure 4A Chain 101 is shown in contact with sensor 15, but it should be understood that a barrier (such as an insulating layer) may be present between sensor 15 and chain 100. Sensor 15 may be a magnetic sensor, for example, as described in PCT Application No. PCT / US2021 / 021274 referenced above.
[0127] Current state-of-the-art commercial CLUS devices (such as those that sense fluorescent markers) can use hundreds of millions of sensors 15, each sensing many instances of individually amplified DNA chains 101. One drawback of some CLUS devices is that achieving optimal cluster density can be crucial for high-quality sequencing. Specifically, using large clusters tends to provide higher data quality but reduces data output, while using small clusters can lead to run failures, poor run performance, lower Q30 scores, introduction of sequencing artifacts, and reduced overall data output. To alleviate these problems, newer CLUS devices use patterned flow cells with different nanopores for cluster generation. These nanopores are organized into a hexagonal arrangement to more efficiently use the flow cell surface area.
[0128] Single-molecule array sequencing device
[0129] Single molecule array sequencers (referred to herein as "SMAS devices") are an alternative to CLUS devices. In contrast to CLUS devices, which sense and sequence local clusters of multiple instances of a single nucleic acid chain, SMAS devices use sensors that sense and sequence individual chains of nucleic acids individually. Generally speaking, in a SMAS device, no sensor senses more than one physical nucleic acid chain, but different sensors sense instances of the same chain. In other words, there are multiple instances of a nucleic acid chain, but each sensed chain is sensed by a different individual sensor. Depending on the amplification technology used, the individual chains may be randomly distributed in the fluid chamber of the SMAS device, or they may be located in more localized areas. As discussed further below, the positions of instances of a particular chain can be identified, and error correction procedures can be applied to the detection results corresponding to the instances before base identification to improve the accuracy of sequencing relative to the CLUS device. In addition, relative to the CLUS device, for a reasonable chemical failure rate, the SMAS device requires fewer instances of each nucleic acid chain to be sequenced to achieve accurate sequencing results.
[0130] Figure 4B An exemplary plurality of S single-molecule sensors 105 are illustrated, each of which is used to monitor a respective single-stranded DNA (ssDNA) 101 via a SMAS device. Each of the plurality of S sensors 105 can be, for example, a magnetic sensor, an optical sensor, or the like. Figure 4B Five single-molecule sensors 105A, 105B, 105C, 105D, and 105E are illustrated, each of which senses a respective DNA strand 101 (which may be instances of the same DNA strand, or instances of different DNA strands). Each sensor 105 may be, for example, a nanoscale sensor that is so small that only a single DNA strand 101 can bind to a binding site associated with the sensor 105. (For convenience, Figure 4B The strand 101 is shown in contact with a sensor 105, but as further described below, in some embodiments, the strand 100 is attached to individual binding sites, each of which is associated with a respective sensor 105.
[0131] Consider clonally amplified DNA bound to a solid surface comprising an array of densely packed sensors 105, such as Figure 4B As shown in FIG. DNA can be replicated by solid phase amplification (SPA) to create monoclonal DNA clusters, with each strand intended to be sensed by a different sensor 105, or the DNA can be amplified in large quantities and then immobilized on the surface of the SMAS device. If the DNA is amplified on the surface of the fluid chamber of the SMAS device (e.g., by SPA), sensors 105A, 105B, 105C, 105D, and 105E can sense instances of clonal DNA. Alternatively, if the DNA is amplified in large quantities outside the device and added to the fluid chamber of the SMAS device, the amplified DNA strands 101 can be more randomly distributed among the sensors 105.
[0132] Figure 5A A block diagram illustrating components of an exemplary SMAS device 100 for nucleic acid sequencing, according to some embodiments, is shown. As shown, device 100 includes a sensor array 110 coupled to circuitry 120, which is coupled to at least one processor 130. Sensor array 110 includes a plurality of sensors 105 (e.g., magnetic sensors, optical sensors, etc.), which can be arranged in any suitable manner, as further described below. The characteristics and properties of sensors 105 in sensor array 110 depend on the type of label used for sequencing.
[0133] The circuitry 120 may include, for example, one or more wires that allow the sensors 105 in the sensor array 110 to be interrogated by the at least one processor 130 (e.g., with the aid of other components known in the art, such as current sources, etc.). For example, in operation, the processor 130 may cause the circuitry 120 to apply a current to such wires to detect a characteristic of at least one of the plurality of sensors 105 in the sensor array 110, wherein the characteristic indicates the presence or absence of a marker within range of the sensor 105. In other words, the characteristic (e.g., resistance, frequency, voltage, signal level, etc.) indicates that the sensor 105 has detected at least one marker or has not detected any marker. For example, the at least one processor 130 may assess the value of the characteristic (e.g., frequency, wavelength, magnetic field, resistance, noise level, intensity, color of light, etc.) and determine that a marker has been detected (or not detected) based on a comparison of the characteristic value to a threshold value (e.g., by determining whether the characteristic value of the sensor 105 meets or exceeds the threshold value) or a baseline value. As another example, the at least one processor 130 may compare the obtained characteristic of the sensor 105 with a previously detected characteristic value (e.g., a baseline value of the sensor 105) and base the determination of whether a marker is detected or not detected on a change in the characteristic value (e.g., a change in magnetic field, resistance, noise level, frequency, wavelength, intensity, color of light, etc.). For example, as shown below in Figure 19 As further described in the discussion of , the at least one processor 130 can evaluate characteristics obtained from the sensor 105 to detect whether the sensor 105 that detected a marker during a first interrogation step of the sequencing procedure still detects the marker after a cutting step in which the marker should have been removed. Similarly, the at least one processor 130 can evaluate changes in characteristics from one interrogation step to the next interrogation step to determine whether the sensor 105 (a) did not detect the marker during any interrogation step, (b) detected the marker during both interrogation steps, (c) did not detect the marker during the first interrogation step but detected the marker during a subsequent interrogation step, and / or (d) detected the marker during the first interrogation step but did not detect the marker during a subsequent interrogation step.
[0134] The characteristic detected depends on the type of label used in the sequencing procedure. The label may be, for example, fluorescent, in which case sensor 105 may be an optical detector capable of detecting, for example, the wavelength, frequency, modulation frequency, color, or intensity of light emitted by the fluorescent label. Optical sensors suitable for detecting fluorescent labels are well known in the art. Where the label used in the nucleic acid sequencing procedure is fluorescent, in some embodiments, circuitry 120 allows at least one processor 130 to detect deviations or fluctuations in the light (or electromagnetic energy) detected by some or all of the sensors 105 in sensor array 110.
[0135] The label can be, for example, magnetic (e.g., magnetic nanoparticles, organometallic compounds, charged molecules, etc.), in which case the sensor 105 can be a magnetic sensor that can detect magnetic properties. Magnetic sensors have been described in applicant's previously filed patent applications, including, for example, PCT application No. PCT / US20 / 27290, entitled "NUCLEIC ACID SEQUENCING BY SYNTHESIS USING MAGNETICSENSOR ARRAYS," filed on April 8, 2020 (Attorney Docket No. ROA-1000-WO / P35097-WO) and published on October 15, 2020 as WO2020 / 210370. In some embodiments where the label is magnetic, the sensor 105 is a magnetoresistive (MR) sensor that can detect, for example, a magnetic field or resistance, a change in a magnetic field or a change in resistance, or a noise level. In some embodiments, each of the sensors 105 of the sensor array 110 is a thin-film device that uses the MR effect to detect magnetic labels attached to nucleotides incorporated into a single strand of nucleic acid bound to a respective binding site. Sensor 105 can function as a potentiometer whose resistance changes with the strength and / or direction of a sensed magnetic field. In some embodiments using magnetic labels, sensor 105 includes a magnetic oscillator (e.g., a spin torque oscillator (STO)), and the characteristic indicating whether at least one label is detected is the frequency of a signal associated with or generated by the magnetic oscillator, or a change in the frequency of the signal.
[0136] In the case where the label used in the nucleic acid sequencing procedure is magnetic, in some embodiments, the at least one processor 130, with the help of the circuit 120, detects deviations or fluctuations in the magnetic environment of some or all sensors 105 in the sensor array 110. For example, compared to sensors 105 with magnetic labels, MR-type sensors 105 without magnetic labels should have relatively low noise above a certain frequency because field fluctuations from the magnetic labels will cause fluctuations in the magnetic moment of the sensing ferromagnet. These fluctuations can be assessed using heterodyne detection (e.g., by measuring noise power density) or by directly measuring the voltage of the sensor 105 and using a comparator circuit to compare it with another sensor component that does not sense the binding site. In the case where the sensor 105 includes an STO component, the fluctuating magnetic field from the magnetic label will cause a phase jump in the sensor 105 due to the instantaneous change in frequency, which can be detected using a phase detection circuit. Another option is to design the STO so that it oscillates only within a small magnetic field range, so that the presence of the magnetic label will shut down the oscillation.
[0137] It should be understood that the examples of labels and sensors 105 provided above are exemplary only. In general, any type of label that can label a nucleotide precursor can be used with an array 110 of any type of sensor 105 that can detect that type of label.
[0138] Figure 5B 、 5C 5D and 5D illustrate portions of an exemplary SMAS device 100 for nucleic acid sequencing according to some embodiments. The exemplary SMAS device 100 utilizes magnetic labels and magnetic sensors 105. Figure 5B is a top view of the device 100 .
[0139] Figure 5C It is by Figure 5B a cross-sectional view at the position indicated by the long dashed line marked as "5C" in FIG. , and Figure 5D It is by Figure 5B A cross-sectional view at the position indicated by the long dashed line marked “5D”.
[0140] Shown in Figure 5B 、 5C The exemplary device 100 in FIG5D includes a sensor array 110 for sensing magnetic markers within a fluid chamber 115. The sensor array 110 includes a plurality of magnetic sensors 105, wherein Figure 5B Sixteen sensors 105 are shown in the array 110. It should be understood that embodiments of the SMAS device 100 may include many sensors 105 (e.g., hundreds, thousands, or millions of sensors 105). To avoid cluttering the drawings, Figure 5BOnly seven of the marker sensors 105 are shown, namely sensors 105A, 105B, 105C, 105D, 105E, 105F, and 105G. As described above, the magnetic sensors 105 detect the presence or absence of magnetic markers. In other words, each of the magnetic sensors 105 detects whether there is at least one magnetic marker in its vicinity.
[0141] Reference Figure 5C and 5D Combine Figure 5B In the exemplary embodiment of device 100, each sensor 105 is shown as having a cylindrical shape. However, it should be understood that, in general, sensors 105 may have any suitable shape. For example, sensors 105 may be rectangular parallelepipeds in three dimensions. Furthermore, different sensors 105 may have different shapes (e.g., some may be rectangular parallelepipeds and others may be cylindrical, etc.). It should be understood that the figures are exemplary only.
[0142] like Figure 5C and 5D As shown in , device 100 includes a fluid chamber 115. Fluid chamber 115 includes a plurality of binding sites 116 (e.g., three binding sites 116). In some embodiments, fluid chamber 115 contains fluids (e.g., nucleotide precursors and other fluids) used during the nucleic acid sequencing procedure. However, it should be understood that embodiments in which fluid chamber 115 does not contain fluids are contemplated and within the scope of this disclosure. For example, binding sites 116 may be arranged on a removable (or movable) portion (e.g., a panel, a plate, a slide, etc.), which may be immersed in reagents and other fluids after the nucleic acid chain has been attached to binding sites 116 and then placed so that sensor 105 can detect the label. Therefore, although the name of fluid chamber 115 indicates that it contains fluids, it is not necessary for fluid chamber 115 to contain fluids.
[0143] like Figure 5B 、 5C As shown in 5D and 5D, each of the sensors 105 is associated with a respective binding site 116. (For simplicity, this document generally refers to a binding site by the reference numeral 116. Individual binding sites are given the reference numeral 116 followed by a letter.) In other words, the sensors 105 and binding sites 116 are in a one-to-one relationship. Figure 5BAs shown in FIG, sensor 105A is associated with binding site 116A, sensor 105B is associated with binding site 116B, sensor 105C is associated with binding site 116C, sensor 105D is associated with binding site 116D, sensor 105E is associated with binding site 116E, sensor 105F is associated with binding site 116F, and sensor 105G is associated with binding site 116G. Figure 5B Each of the other unlabeled sensors 105 in is also associated with a respective binding site 116. Figure 5B 、 5C 5D , each sensor 105 is shown disposed below its respective binding site 116, but it should be understood that the binding site 116 may be located elsewhere relative to its respective sensor 105. For example, the binding site 116 may be located to the side of its respective sensor 105.
[0144] Each of the binding sites 116 is configured to allow no more than one nucleic acid strand (e.g., ssDNA) to bind to the fluid chamber 115 of the SMAS device 100. In other words, each binding site 116 has properties and / or characteristics that allow one and only one strand of nucleic acid to bind thereto for sensing (and sequencing) by the respective sensor 105. Thereafter, the respective sensor 105 can detect the label attached to the nucleotide incorporated into the nucleic acid strand bound to the binding site 116 during a nucleic acid sequencing procedure, as discussed further below. In some embodiments, the binding site 116 has a structure (or structures) configured to anchor the nucleic acid to the binding site 116. For example, the structure (or structures) may include a cavity or a ridge. Figure 5C and 5D The binding sites 116 are illustrated as extending from the surface of the fluid chamber 115 , but it should be understood that the binding sites 116 may be flush with or etched into the surface of the fluid chamber 115 .
[0145] Binding sites 116 can have any suitable size and shape that facilitates attachment of one and only one strand of nucleic acid to each binding site 116. For example, the shape of the binding site can be similar to or identical to the shape of sensor 105 (e.g., if sensor 105 is cylindrical in three dimensions, binding site 116 can also be a cylinder protruding from or forming a fluid reservoir within the surface of fluid chamber 115, with a radius that is larger, smaller, or the same size as the radius of the respective sensor 105; if sensor 105 is a rectangular parallelepiped in three dimensions, binding site 116 can also be a rectangular parallelepiped with a surface that is larger, smaller, or the same size as the proximal portion of sensor 105, etc.). In general, binding sites 116 and the surface of fluid chamber 115 can have any shape and characteristics that facilitate attachment of a single nucleic acid strand to each binding site 116 and allow sensor 105 to detect a label attached to an incorporated nucleotide at its respective binding site 116.
[0146] Figure 5C and 5D The enclosed fluid chamber 115 is illustrated with a top portion extending in the xy plane, but the fluid chamber 115 need not be enclosed. In some embodiments, the surface of the fluid chamber 115 has properties and characteristics that protect the sensor 105 from any fluid in the fluid chamber 115, while still allowing the nucleic acid chain to bind to the binding site 116 and allowing the sensor 105 to detect the label attached to the nucleotide incorporated into the nucleic acid chain attached to the binding site 116. The material of the fluid chamber 115 (and possibly the material of the binding site 116) can be an insulator or include an insulator. In some embodiments, the surface of the fluid chamber 115 includes an organic polymer, a metal, or a silicate. The fluid chamber 115 may include, for example, a metal oxide, silicon dioxide, polypropylene, gold, glass, or silicon. The thickness of the surface of the fluid chamber 115 can be selected so that the sensor 105 can detect the magnetic label attached to the nucleotide incorporated into the nucleic acid chain bound to the binding site 116 within the fluid chamber 115. In some embodiments, the surface is about 3 to 20 nm thick such that each sensor 105 is between about 5 nm and about 50 nm from any label attached to a nucleotide incorporated into a nucleic acid strand that binds to a corresponding binding site 116 of the sensor 105. It should be understood that these values are exemplary only. It should be understood that embodiments can have fluid chambers 115 with thicker or thinner surfaces.
[0147] The circuit 120 of the device 100 may include one or more wires 125. In some embodiments, each of the plurality of sensors 105 is coupled to at least one wire 125. Figure 5B 、 5C5D, the device 100 includes eight lines 125A, 125B, 125C, 125D, 125E, 125F, 125G, and 125H. (For simplicity, this document generally refers to lines by the reference numeral 125. Individual lines are given the reference numeral 125 followed by a letter.) Lines 125 can be used to access (e.g., interrogate) individual sensors 105. In the example shown in FIG. Figure 5B 、 5C In the exemplary embodiment of Figures 5 and 5D, each sensor 105 of the sensor array 110 is coupled to two lines 125. For example, sensor 105A is coupled to lines 125A and 125H; sensor 105B is coupled to lines 125B and 125H; sensor 105C is coupled to lines 125C and 125H; sensor 105D is coupled to lines 125D and 125H; sensor 105E is coupled to lines 125D and 125E; sensor 105F is coupled to lines 125D and 125F; and sensor 105G is coupled to lines 125D and 125G. Figure 5B 、 5C In the exemplary embodiments of FIGS. 5A and 5D , display lines 125A, 125B, 125C, and 125D are located below the magnetic sensor 105 , and display lines 125E, 125F, 125G, and 125H are located above the magnetic sensor 105 . Figure 5C Sensor 105E is shown with respect to lines 125D and 125E, sensor 105F is shown with respect to lines 125D and 125F, sensor 105G is shown with respect to lines 125D and 125G, and sensor 105D is shown with respect to lines 125D and 125H. Figure 5D Sensor 105D is shown with respect to lines 125D and 125H, sensor 105C is shown with respect to lines 125C and 125H, sensor 105B is shown with respect to lines 125B and 125H, and sensor 105A is shown with respect to lines 125A and 125H.
[0148] Figure 5B 、 5C The sensors 105 of the exemplary SMAS device 100 of FIG. 5 and 5D are arranged in a rectangular pattern sensor array 110. (It should be understood that a square pattern is a special case of a rectangular pattern.) Each of the lines 125 identifies a row or column of the sensor array 110. For example, each of the lines 125A, 125B, 125C, and 125D identifies a different row of the sensor array 110, and each of the lines 125E, 125F, 125G, and 125H identifies a different column of the sensor array 110. Figure 5C, each of lines 125E, 125F, 125G, and 125H is in contact with one of the sensors 105 along the cross-section (that is, line 125E is in contact with the top of sensor 105E, line 125F is in contact with the top of sensor 105F, line 125G is in contact with the top of sensor 105G, and line 125H is in contact with the top of sensor 105D), and line 125D is in contact with the bottom of each of sensors 105E, 105F, 105G, and 105D. Similarly, and as Figure 5D As shown in FIG, each of lines 125A, 125B, 125C, and 125D contacts the bottom of one of the sensors 105 along the cross-section (that is, line 125A contacts the bottom of sensor 105A, line 125B contacts the bottom of sensor 105B, line 125C contacts the bottom of sensor 105C, and line 125D contacts the bottom of sensor 105D), and line 125H contacts the top of each of sensors 105D, 105C, 105B, and 105A.
[0149] Figure 5B The sensor 105 and the portion of the wire 125 connected to the sensor array 110 are depicted using dashed lines to indicate that they can be embedded within the device 100. As described above, the sensor 105 can be protected (e.g., by an insulator) from the contents of the fluid chamber 115, which itself can be sealed. Thus, it should be understood that the various illustrated components (e.g., wire 125, sensor 105, binding site 116, etc.) are not necessarily visible in a physical instantiation of the device 100 (e.g., they can be embedded in or covered by a protective material such as an insulator).
[0150] In some embodiments, some or all of the binding sites 116 reside in nanopores or grooves in the wire 125 that passes through the sensor 105. For example, Figure 5D As shown in the example of FIG, line 125H can be thinner above sensor 105 than between sensors 105. For example, line 125H has a first thickness above sensor 105D, a second, larger thickness between sensors 105D and 105C, and a first thickness above sensor 105C. This configuration can be advantageously fabricated using conventional thin film fabrication methods (e.g., by depositing material, applying a mask to the deposited material, and removing (e.g., by etching) some of the deposited material in accordance with the mask). Binding site 116 and, if present, the nanopore can both be fabricated using conventional techniques.
[0151] To simplify the explanation, Figure 5B 、 5C5D and 5D illustrate an exemplary device 100 having only sixteen sensors 105, only sixteen individual binding sites 116, and eight wires 125 in the sensor array 110. It should be appreciated that the device 100 may have fewer or more sensors 105 in the sensor array 110, and therefore, may have more or fewer binding sites 116. Similarly, embodiments including wires 125 may have more or fewer wires 125. In general, any configuration of sensors 105 and binding sites 116 that allows the sensor 105 to detect a label attached to a nucleotide incorporated into a single nucleic acid strand attached to a binding site 116 may be used. Similarly, any configuration of one or more wires 125 or some other mechanism that allows for determining whether a sensor 105 has sensed one or more labels may be used. The examples presented herein are not intended to be limiting.
[0152] As explained above, the Figure 5B 、 5C The sensors 105 in 5D and 5D can be magnetic sensors 105. Thus, the sensors 105 are in close proximity to the binding sites 116, and therefore, are also in close proximity to the nucleic acid strands bound to the binding sites 116. It will be appreciated that the appropriate location of the sensor array 110 relative to the binding sites 116 depends in part on the type of label used, and therefore on the type of sensor 105 used. For example, if the label is a fluorophore and the sensor 105 is an optical sensor, then it may be appropriate for the sensor array 110 to be located away from the binding sites 116 (e.g., above the binding sites 116).
[0153] although Figure 5B 、 5C 5D and 5D (and other figures herein) illustrate sensors 105 and binding sites 116 in a one-to-one relationship, but it should be understood that each binding site 116 can be sensed by more than one sensor 105. A characteristic that distinguishes the SMAS device 100 from a CLUS device is that the sensors 105 of the SMAS device 100 do not sense more than one instance of a nucleic acid strand. If the SMAS device 100 has more sensors 105 than binding sites 116, it may be feasible to sense at least some nucleic acid strands by multiple sensors 105 (e.g., to improve the accuracy of label detection).
[0154] Shown and described in Figure 5B 、 5C The exemplary sensor array 110 in the context of 5D is a rectangular array in which the sensors 105 are arranged in rows and columns. In other words, the plurality of sensors 105 of the sensor array 110 are arranged in a rectangular grid pattern. In some embodiments, adjacent rows and columns of the rectangular grid pattern are equidistant from each other, resulting in the sensors 105 being arranged in a square grid (or lattice) pattern, such as Figure 5EIn an embodiment where the sensors 105 are arranged in a square grid pattern, each sensor 105 has up to four nearest neighbors. For example, Figure 5E As shown in FIG, sensor 105A has four nearest neighbors, labeled 105B, 105C, 105D, and 105E. Figure 5E , the closest sensor 105 is a nearest neighbor distance 112. Thus, each of sensors 105B, 105C, 105D, and 105E is a distance 112 from sensor 105A.
[0155] A commercially viable SMAS device 100 can be fabricated using high-precision nanoscale fabrication of densely packed nanoscale sensors 105 capable of identifying individual tags. The size of the functionalized binding sites 116 can be similar to the size of, for example, the DNA to which the tags are attached, so that multiple strands cannot bind to the same binding site 116 or be sensed by the same sensor 105. A recognized metric for evaluating the commercial competitiveness of sequencers is the density with which DNA strands are packed together in the fluid chamber 115.
[0156] A suitable value for the nearest neighbor distance 112 can be determined based on the properties of the sensor 105, the length of the nucleic acid strands to be sequenced by the device 100, and the properties of the labels used. This suitable value can then be used to determine the size of the SMAS device 100 and / or the maximum number of sensors 105 that can be assembled within a SMAS device 100 of a selected size. For example, the combined length of the nucleic acid strands and the size of the labels to be used can provide a physical limit on how closely two sensors 105 can be positioned within the SMAS device 100. In some embodiments, the size of the sensor 105 can be limited by the nanoscale patterning capabilities of the process used to fabricate the SMAS device 100. For example, using the technology available at the time of writing, the size of each magnetic sensor 105 (e.g., the diameter of the sensor 105 in the xy plane, assuming a cylindrical sensor 105) can be approximately 20 nm. Assuming that the type of nucleic acid to be sequenced is DNA, and it is desired to sequence fragments up to 150 base pairs (bp) in length, the maximum length of the DNA chain 101 to be sequenced in the elongated state is about 50 nm, although the ssDNA conformation can vary between elongated and helical, as shown in FIG. Figure 6A As shown in , it depends on the ionic strength of the buffer. Because the marker 102 participates in a unimolecular reaction, the marker 102 should have a molecular size. For the SMAS device 100 using the magnetic sensor 105, the marker 102 can be, for example, a superparamagnetic nanoparticle, an organometallic compound, or any other functional molecular group that can be detected by the nanoscale magnetic sensor 105. Therefore, it is assumed that each marker 102 has a size of no greater than about 10 nm. Under these assumptions, Figure 6BThe relative sizes of the magnetic sensor 105 , the DNA strand 101 in its elongated state, and the magnetic label 102 are shown.
[0157] The actual SMAS device 100 using magnetic sensors 105 to detect magnetic nanoparticles used as labels 102 can be implemented using existing technology. For the sake of demonstration, it is assumed that only labels 102 within 20 nm of the edge of the sensor 105 are detected. The detection range of each sensor 105 is small because the magnetic labels 102 (e.g., superparamagnetic nanoparticles, organometallic compounds, etc.) that can be selected for nucleic acid sequencing applications do not produce significant perturbations to the detected magnetic field. Although labels 102 attached to nucleotides incorporated into ssDNA that bind to the binding site 116 of a particular sensor 105 can temporarily reside outside the range of a respective sensor 105 because the ssDNA assumes various conformational states during the detection process, it is desirable that the labels not be allowed to reach the sensitive space (detection area) of an adjacent sensor 105 when the ssDNA assumes its fully elongated state.
[0158] The sensor packing limit of a practical SMAS device 100 can be derived, for example, assuming that the labeled nanoparticles are superparamagnetic (e.g., iron oxide, iron platinum, etc.), and that the sensor array 110 of the SMAS device 100 is a rectangular (e.g., square) array of magnetic tunnel junctions (MTJs) similar to those used in non-volatile data storage applications. In this case, an area of each nanoscale sensor 105 or its immediate vicinity can be functionalized to serve as a respective binding site 116. A simple geometric arrangement for estimating the sensor array packing limit of the SMAS device 100 is shown in FIG. Figure 7A , which shows two sensors 105A, 105B. Assume that each sensor 105A, 105B (assumed to have a cylindrical shape for convenience only) has a diameter of about 20 nm (as described above) and is capable of detecting any mark within 20 nm from its edge. The sensing area boundary 111 is shown in Figure 7A101B). The inner dashed line in FIG. Sensor 105A senses DNA strand 101A bound to its binding site, and sensor 105B senses DNA strand 101B bound to its binding site. The maximum reach of labels 102A and 102B when attached to nucleotides incorporated into strands 101A and 101B (e.g., when a 150-base DNA strand is in its fully unhelical state) is shown by the outer dashed circle 103. For accurate sequencing results, it is desirable for each sensor 105 to detect only labels 102 attached to nucleotides incorporated into DNA strand 101 bound to its respective binding site 116. Therefore, under the assumptions described above, the minimum nearest neighbor distance 112 between sensors 105 to avoid crosstalk (e.g., detecting labels 102 attached to nucleotides incorporated into nucleic acid strand 101 bound to the binding site 116 of another sensor 105) is approximately 100 nm.
[0159] In some embodiments of the SMAS device 100, the sensors 105 (eg, MTJs) are arranged in a square grid that is compatible with existing cross-point MRAM sensor geometries, such as Figure 7B The area of the unit grid 114 is 10 4 nm 2 , which allows each DNA strand 101 to extend through approximately 10 4 nm 2 area, which yields a DNA surface density of approximately 10 10 Chains / cm 2 Assuming that at least ten instances of each chain 101 are used in the sensor array 110, approximately 10 9 Unique chains / cm 2 , generating 150Gbase (1 billion x 150bp DNA chain length) of information per square centimeter of sensor array 110. Under ideal conditions (e.g., when the chemical failure rate is low, only three DNA samples are needed, as discussed further below), approximately 3.3 x 10 9 Different chains / cm 2 , and each square centimeter of the sensor array 110 can generate approximately 500Gbase data.
[0160] As a specific example, a SMAS device 100 having a configuration similar to that of a single Toshiba 4Gbit density STT-MRAM chip first introduced at the International Electron Devices Meeting (IEDM) in 2016 can potentially generate approximately 600Gbase of high-quality data. The minimum distance 112 between sensors 105 of the Toshiba platform is 90nm, which is only slightly lower than the estimated minimum distance 112 of 100nm derived above. Therefore, crosstalk using a configuration similar to the Toshiba platform can still be low even for ssDNA of 150 bases in length, but shorter fragments can be sequenced to reduce crosstalk even further.
[0161] It should be understood that the sensor 105 is arranged in a grid pattern (e.g., as shown in FIG. Figure 7B The arrangement of the sensors 105 in a square grid is one of many possible arrangements. One of ordinary skill in the art will appreciate that other arrangements of the sensors 105 are possible and within the scope of the disclosure herein. For example, the sensors 105 may be arranged in a hexagonal pattern, such as Figure 8A , which shows a top view of the SMAS device 100. Figure 8A The exemplary SMAS device 100 in FIG. 1 includes a sensor array 110 for sensing markers 102 within a fluid chamber 115. The sensor array 110 includes a plurality of sensors 105, of which sixteen sensors 105 are shown. It should be understood that implementations of the device 100 may include any number of sensors 105 (e.g., hundreds, thousands, millions, etc.). To avoid cluttering the drawings, the following diagrams are provided. Figure 8A Only two of the sensors 105 are marked, namely sensors 105A and 105B. As explained above, the sensors 105 may be, for example, magnetic sensors (e.g. to detect the effects of magnetism or magnetic nanoparticles). Figure 5B 、 5C As described in the discussion of 5D and 5D, in general, sensor 105 can have any suitable size and shape.
[0162] like Figure 8A As shown in FIG, each of the sensors 105 is associated with a respective binding site 116. In other words, the sensors 105 and the binding sites 116 are in a one-to-one relationship. Figure 8A As shown in FIG, sensor 105A is associated with binding site 116A, sensor 105B is associated with binding site 116B, and each of the other unlabeled sensors 105 is also associated with a respective binding site 116. Figure 8AIn the example embodiment of FIG. 1 , each sensor 105 is shown disposed below its respective binding site 116, but it should be understood that the binding site 116 may be located in other positions relative to its respective sensor 105. For example, the binding site 116 may be located to the side of its respective sensor 105. Figure 5B 、 5C The discussion of binding site 116 in the description of 5D applies to Figure 8A and other diagrams showing the binding site 116 and are not repeated here.
[0163] Figure 8A The exemplary SMAS device 100 also includes the Figure 5B 、 5C and the fluid chamber 115 discussed in 5D. Those descriptions also apply to Figure 8A And will not be repeated here.
[0164] Figure 8A The circuitry 120 of the device 100 may include one or more lines 125 . Figure 8A Each of the lines 125 in the exemplary embodiment of FIG. 1 identifies a row or diagonal column of the sensor array 110. For example, each of the lines 125A, 125B, 125C, and 125D identifies a different row of the sensor array 110, and each of the lines 125E, 125F, 125G, and 125H identifies a different diagonal column of the sensor array 110. Figure 8A In the example of FIG, device 100 has eight wires 125A, 125B, 125C, 125D, 125E, 125F, 125G, and 125H, and pairs of wires 125 can be used to access individual sensors 105. For example, wires 125A and 125H can be used to access sensor 105A, and wires 125B and 125H can be used to access sensor 105B. Wires 125 can be oriented below and / or above sensors 105, such as Figure 5B 、 5C and described in the discussion of 5D et al.
[0165] although Figure 8AWhile an exemplary device 100 is illustrated having only sixteen sensors 105, only sixteen corresponding binding sites 116, and eight wires 125 in the sensor array 110, it should be understood that the SMAS device 100 may have fewer or more sensors 105 in the sensor array 110, and therefore, may have more or fewer binding sites 116. Furthermore, the SMAS device 100 may have more or fewer wires 125. Generally speaking, any configuration of sensors 105 and binding sites 116 that allows the sensor 105 to detect a label attached to a nucleotide incorporated into a single nucleic acid strand attached to a binding site 116 may be used. Similarly, any configuration of one or more wires 125, or some other mechanism that allows for determining whether a sensor 105 has sensed one or more labels, may be used.
[0166] like Figure 8B As shown in FIG, when the sensors 105 are arranged in a hexagonal pattern, each sensor 105 has at most six nearest neighbors, all at a nearest neighbor distance 112. In other words, each sensor 105 is at a nearest neighbor distance 112 from each of the six other sensors 105 closest to it. For example, Figure 8B As shown in FIG. 1 , the unlabeled sensor 105 in the middle of the drawing has six nearest neighbor sensors 105 , labeled 105A, 105B, 105C, 105D, 105E, and 105F, all of which are a nearest neighbor distance 112 away.
[0167] The packing limit of binding sites 116 can be derived for a SMAS device 100 using an optical sensor and fluorescent labels 102 (e.g., fluorophores) and having a hexagonal pattern of binding sites 116. Assuming that the labels 102 are fluorophores, the binding sites 116 are in a hexagonal pattern, and the sensor array 110 is far from the binding sites 116, single-molecule fluorescence from the labels 102 can be projected into the far field where it can be detected by the sensor array 110 including the photosensors 105. Single-molecule super-resolution imaging techniques (such as those described in CG Galbraith and JA Galbraith, "Super-resolution microscopy at a glance," Journal of Cell Science, Vol. 124(10), 1607-11 (2011)) can be used to resolve the location of individual fluorophore labels 102 in the SMAS device 100. Because the DNA packaging size is well below the diffraction limit, the position of the fluorophore label 102 can be resolved. Although this type of detection can be somewhat complex and / or expensive, the technology has recently been introduced into commercial sequencing systems to improve the throughput of cluster-based sequencers. In addition, the technology may be implemented in the imaging of large single-molecule arrays in the near future.
[0168] A simple geometric arrangement for estimating the packing limit of binding sites 116 located in a hexagonal pattern in a SMAS device 100 using fluorophore labels 102 is shown in Figure 9A DNA strand 101A is bound to binding site 116A, and DNA strand 101B is bound to binding site 116B. (Sensor 105 is not shown.) Figure 9A , because the sensor array 110 is assumed to be far away from the binding sites. ) The maximum reach of the labels 102A, 102B (e.g., when a DNA chain having 150 bases is in its fully non-helical state) (when attached to an incorporated nucleotide) is represented by the dot-dashed circle 103. To avoid crosstalk, fluorophore labels 102 attached to adjacent binding sites 116 are not allowed to occupy overlapping space during the imaging process, e.g., a fluorophore label 102A attached to a particular binding site 116A should not be allowed to reach a space accessible to a fluorophore label 102B attached to an adjacent binding site 116B when the ssDNA 101A is exploring its allowed conformational state. Such restriction also helps to avoid fluorescence quenching. Assuming that fluorophore labels 102 are used, the binding sites 116 can be densely packed in a hexagonal lattice, as shown in FIG. Figure 9BAssume that the maximum length of the 150 bp DNA strand 101 is 50 nm, the size of the fluorophore label 102 is 10 nm, the minimum distance from the center of each binding site 116 to its edge is 20 nm, and each DNA strand 101 binds to the center of its respective binding site 116, and the minimum distance 112 is 140 nm. Therefore, Figure 9B As shown in the figure, each DNA chain 101 is allowed to occupy 1.7×10 4 nm 2 The area of the unit grid is 114, which yields 5.9×10 9 Chains / cm 2 , or 5.9×10 if there are approximately 10 instances of each DNA strand in the SMAS device 100 8 Unique chains / cm 2 The SMAS device 100 will generate approximately 90 Gbase of data per square centimeter of the sensor array 110. In the best case scenario, when only 3 DNA replicas are required, the sensor array 110 maintains approximately 2×10 9 Unique DNA chains / cm 2 , and the SMAS device 100 is capable of generating approximately 300 Gb of data per square centimeter from the sensor array 110.
[0169] The above discussion of hexagonal arrays is in the context of fluorophore labels 102 and optical sensors 105. Hexagonal arrangements of magnetic sensors 105 can also be used. Figure 7A and 7B The sensor packaging limit of the SMAS device 100 with a hexagonal arrangement of binding sites 116 and magnetic sensors 105 is obtained as described in the discussion of . For the magnetic sensors 105, the nearest neighbor distance 112 is about 100 nm, which means that the (hexagonal) unit grid area 114 (see Figure 9B ) is about 8.7×10 3 nm 2 .
[0170] Figure 10 Comparison described in Figure 7A and 7B (magnetic marker 102 and magnetic sensor 105) and Figure 9A and 9B The density of the SMAS implementation in the context of (fluorescent markers 102 and optical sensors 105) is comparable to the density of the current state-of-the-art CLUS sequencer. For the sake of argument, it is assumed that the pitch of the nanopore array of the patterned flow cell is about 500 nm. Figure 10As shown in the left-hand panel of FIG, the nanopores of the CLUS sequencer are arranged in a hexagonal lattice with a lattice constant of 500 nm. Each nanopore holds about 50 to about 200 identical DNA strands (generated, for example, by solid-phase bridge amplification). Figure 10 The upper right-hand side of the figure shows the hexagonal SMAS lattice using fluorophore labeling and super-resolution imaging (e.g., Figure 9A and 9B ), and Figure 10 The lower right hand side of FIG shows a square SMAS grid of a sensor array 110 using superparamagnetic nanoparticle labels and MTJs (e.g., Figure 7A and 7B (as described in the text of ). Figure 10 The three in the figure are scaled to show how the SMAS lattice structure compares to the CLUS structure. The black hexagons (left and upper right) and squares (lower right) mark the unit lattice that maintains the minimum number of individual molecules required to recognize the sequence of nucleic acid chains. For the SMAS lattice, the ideal case in which only three DNA chains are required for successful base recognition is illustrated, which is discussed in further detail below. It should be noted that in the SMAS case ( Figure 10 ), DNA instances are randomly distributed throughout the sensor array 110, and their positions can be identified during the first sequencing cycle, as discussed further below.
[0171] like Figure 10 As shown in the figure, the area of the unit grid of the CLUS device is 2.2×10 5 nm 2 , which corresponds to 4.6×10 8 Clusters / cm 2 Using the assumptions made above, the CLUS sequencer generates approximately 70 Gb / cm² of data per square centimeter of sensing area. Conversely, under ideal conditions, when only three instances of a chain are used, the SMAS device 100 generates approximately 500 Gb / cm². 2 (magnetic sensor 105 (e.g., MTJ) and magnetic label 102 (e.g., superparamagnetic nanoparticles)) and about 300 Gb / cm 2 (Optical sensor 105 (super-resolution imaging) and fluorescent labeling 102) data. Results for exemplary implementations of the CLUS sequencer and SMAS device 100 are summarized in the table below, which estimates sequencing throughput assuming only three instances per DNA strand and ten instances per DNA strand for the SMAS implementation.
[0172]
[0173] The table above shows that the SMAS device 100 outperforms the current state-of-the-art CLUS device when the number of DNA instances used for algorithmic error correction, described further below, is small (e.g., <10). Because the error correction process relies on more instances of each ssDNA, the SMAS device 100 begins to behave like a CLUS device and, unlike sensing clusters, offers little benefit in sensing individual molecules. Fluorescent SMAS essentially represents a limitation in reducing clusters to single molecules. One approach to reducing sequencing costs is to reduce cluster size and pack DNA clusters closer together to extract more information from the fixed sensing area. While this approach reduces the amount of reagents required to run the sequencing chemistry, it also significantly increases the complexity and cost of the imaging hardware by continuously pushing the limits of what is currently possible with commercial optical instrumentation. This strategy is a challenging task because scaling is impossible without concurrent improvements in the chemistry. This is because as clusters become smaller, each reaction becomes increasingly important, and chemical failures that occur randomly at the single-molecule level become more pronounced and difficult to tolerate.
[0174] The cost of implementing super-resolution imaging in a CLUS device makes the SMAS device 100, and particularly those using magnetic sensors 105 and magnetic labels, a possible destructive sequencing alternative. The SMAS devices 100 disclosed herein, and particularly those using magnetic sensors 105, ensure excellent throughput at significantly lower instrument costs by utilizing technology developed by the large-scale semiconductor and data storage industries and high-volume manufacturing.
[0175] SMAS sequencing scheme
[0176] As described above, when the SMAS device 100 is used for nucleic acid sequencing, the nucleic acid strands can be amplified before or after the nucleic acids are added to the SMAS device 100 (e.g., using bridge amplification). Regardless of how the nucleic acids are amplified, the strands can be sequenced one base at a time by SBS (e.g., by synthesizing dsDNA from ssDNA). The SMAS sequencing protocol is described assuming that the nucleic acid being sequenced is DNA. It will be appreciated that the disclosed protocol can be modified for sequencing other nucleic acids. Such modifications will be within the capabilities of one of ordinary skill in the art, given an understanding of the disclosure herein.
[0177] To simplify analysis and illustrate the benefits of using the disclosed SMAS device 100 rather than a CLUS sequencer, consider a DNA sequencing scenario in which a single type of label (e.g., molecular, fluorescent, magnetic, etc.) is attached to all four nucleotides (A, T, C, and G). In other words, the same label of a certain type is attached to each of the four nucleotides (e.g., if the selected label 102 is an FePt particle, then each of A, T, C, and G is labeled with an FePt particle). These labeled nucleotides are then incorporated into the DNA chain one base at a time using termination chemistry, for example, once the nucleotide is incorporated, the label 102 is cleaved before the polymerase moves on to the next base. A sensor 105 detects the label 102 attached to the nucleotide.
[0178] An exemplary method 200 for sequencing multiple nucleotide chains (e.g., ssDNA) using the SMAS device 100 is shown in FIG. Figure 11 At 202, the method begins. At 204, one or more nucleic acid strands are optionally amplified before being added to the SMAS device 100. At 206, a plurality of S nucleic acid strands are bound to a plurality of S binding sites 116 of the SMAS device 100 (wherein the plurality includes at least two, but not necessarily all, binding sites 116 of the SMAS device 100). Optionally, at 208, the nucleic acid strands are amplified (e.g., via bridge amplification, which may be in addition to or in addition to the amplification at 204). At 210, a sequencing procedure is performed. The sequencing procedure may be, for example, an additive method, a subtractive method, or a modified additive method, as described further below. The sequencing procedure performed at 210 generates S records, each of the S records capturing M detection results from one of the plurality of S sensors (wherein, again, the plurality includes at least two, but not necessarily all, sensors 105 of the SMAS device 100, and the M detection results may include as few as one detection result, a subset of the total number of detection results obtained during the sequencing procedure, or all detection results obtained during the sequencing procedure). Each of the M detection results indicates whether the sensor 105 corresponding to the record detected at least one marker during a respective step of the M query steps. The M detection results may be stored in a record, which may be stored in a memory. At 212, an error correction procedure is performed, as further described below. The error correction procedure may include deterministic and / or probabilistic error correction techniques. The error correction procedure may be performed, for example, by at least one processor 130 of the SMAS device 100. Alternatively, it may be performed by a processor external to the SMAS device 100 (e.g., an off-device processor, such as in an external computer). The error correction procedure may be performed while the sequencing procedure is being performed (e.g., in real time or near real time), or it may be performed at a later time. At 214, the method 200 ends.
[0179] As described above, at 210, the SMAS device 100 can be used to implement various schemes for reading nucleic acid sequences (e.g., DNA sequences). To simplify the analysis, it is assumed that the plurality of S sensors 105 of the SMAS device 100 only detect the presence or absence of the marker 102 and do not distinguish nucleotides based on the detected signal level. Therefore, in some embodiments, the record of the detection result of each sensor 105 only includes a "yes" or "no" (or 1 / 0 or any other binary indicator) indication that the sensor 105 detected or did not detect the marker during a particular query step. It should be understood that other methods are possible and within the scope of the present disclosure. For example, different markers 102 can be attached to different nucleotides. As another example, rather than a binary "yes" or "no" decision, a characteristic value (e.g., resistance, frequency, intensity, etc.) can be detected and / or recorded, and a decision on whether the marker was detected can be made based on this basis. For example, instead of using only 0 and 1 (or "no" and "yes") as possible outputs of a sequencing program, using different labels for different nucleotides can result in one of the following five levels: 0 (label not detected), level 1 (label 1 detected), level 2 (label 2 detected), level 3 (label 3 detected), and level 4 (label 4 detected). In this case, the range of the detected property can be limited to distinguish whether the label is detected at all and, if detected, which label is detected (e.g., if the property value is between 0 and a first value, the label is determined to be not detected; if the property value is between the first and second values, the first label is determined to be detected; if the property value is between the second and third values, the second label is determined to be detected; etc.).
[0180] Below is a description of three examples of DNA sequencing protocols, each of which includes repeated query cycles, each query cycle having four query steps. During each query cycle, four binary "yes" or "no" questions are answered for each ssDNA sequenced. In one query step, the question "Is the detected base adenine?" ("A?") is answered. In another query step, the question "Is the detected base thymine?" ("T?") is answered. In another query step, the question "Is the detected base cytosine?" ("C?") is answered. And in another query step, the question "Is the detected base guanine?" ("G?") is answered. The record of the detection results obtained during the sequencing procedure can be established as a query cycle, including repeated query cycles. Query Step. It should be understood that the order in which nucleotides are introduced and bases are detected is arbitrary (meaning that the order of the query step is arbitrary), and that the order in which the bases are tested in the examples herein are This is for illustrative purposes only.
[0181] Additive method
[0182] In the additive method, the sensor 105 detects nanoscale labels 102 bound to nucleotides with cleavable linkers. All four types of nucleotides carry the same type of label 102 (e.g., molecular, fluorescent, magnetic, etc.) and use the same type of cleavable linker. According to one embodiment, a query cycle that will produce four detection results (one of which (in the absence of errors) will be a label detection for each of the plurality of three nucleic acid chains 101) involves the following steps:
[0183] 1. Obtain baseline characteristics for each of the plurality of S sensors 105 (which may be all or less than all of the sensors 105 in the sensor array 110 ) of the SMAS device 100 (e.g., by determining a baseline signal at each of the plurality of S sensors 105 ).
[0184] 2. Labeled A nucleotides are introduced and incorporated. Unbound labeled molecules are washed away.
[0185] 3. Query step 1: Get the characteristics of each of the plurality of S sensors 105 (e.g., by detecting a signal at each of the plurality of S sensors 105) and determine whether each sensor 105 detects at least one tag. Save the detection result of each sensor 105 at a location in the record corresponding to query step 1 of the current query cycle.
[0186] 4. Labeled T nucleotides are introduced and incorporated. Unbound labeled molecules are washed away.
[0187] 5. Query step 2: Obtain characteristics of each of the plurality of S sensors 105 (e.g., by detecting a signal at each of the plurality of S sensors 105) and determine whether each sensor 105 detects at least one tag. Save the detection result of each sensor 105 at a location in the record corresponding to query step 2 of the current query cycle.
[0188] 6. Introduce and incorporate labeled C nucleotides. Wash away unbound labeled molecules.
[0189] 7. Query step 3: Obtain characteristics of each of the plurality of S sensors 105 (e.g., by detecting a signal at each of the plurality of S sensors 105) and determine whether each sensor 105 detects at least one tag. Save the detection result of each sensor 105 at a location in the record corresponding to query step 3 of the current query cycle.
[0190] 8. Introduce and incorporate labeled G nucleotides. Wash away unbound labeled molecules.
[0191] 9. Query step 4: Obtain characteristics of each of the plurality of S sensors 105 (e.g., by detecting a signal at each of the plurality of S sensors 105) and determine whether each sensor 105 detects at least one tag. Save the detection result of each sensor 105 at a location in the record corresponding to query step 4 of the current query cycle.
[0192] 10. Cut and wash away the labels of A, T, C and G nucleotides.
[0193] Then steps 1 to 10 can be repeated for the next query cycle. It should be understood that some of the orderings in steps 1 to 10 are exemplary, and further, the quantity and numbering of steps 1 to 10 are for convenience and can be modified. As an example, and as previously described, the order in which nucleotides are introduced is arbitrary. As another example, steps 2, 4, 6, and 8 include introducing and incorporating nucleotides, and washing out unbound nucleotides in a single step, but it should be understood that each of steps 2, 4, 6, and 8 can be divided into a series of smaller steps. Similarly, steps 3, 5, 7, and 9 can be further divided into a series of smaller steps (e.g., obtaining characteristics, determining whether a tag is detected, and preserving the detection results). On the contrary, steps can be combined (e.g., steps 2 and 3 can be combined, steps 4 and 5 can be combined, etc.).
[0194] It should be understood that if no errors occur during any query cycle of the additive method, the individual bases of the individual strands can be identified (determined) once the label is detected. For example, referring to the above steps, if at query step 1 involving a labeled A nucleotide, the resulting characteristic for a particular sensor 105 indicates that the sensor 105 detected the label, then saving the detection result can be equivalent to identifying the base (T) complementary to A for the detector 105 (and binding site 116). Similarly, if at query step 2 involving a labeled T nucleotide, the resulting characteristic for the particular sensor 105 indicates that the sensor 105 detected the label, then saving the detection result can be equivalent to identifying the base (A) complementary to T for the detector 105 (and binding site 116). Similarly, if at query step 3 involving a labeled C nucleotide, the resulting characteristic for the particular sensor 105 indicates that the sensor 105 detected the label, then saving the detection result can be equivalent to identifying the base (G) complementary to C for the detector 105 (and binding site 116). Finally, if, at query step 4 involving a labeled G nucleotide, for a particular sensor 105, the resulting characteristic indicates that the sensor 105 detected the label, then saving the detection result can be equivalent to identifying the base (C) complementary to the G for that detector 105 (and binding site 116). However, as described in further detail below, several types of errors can occur during the sequencing process (e.g., during the additive method), and therefore, in some embodiments, a record is created during the sequencing process to record the detection / non-detection of the label during each query step of each query cycle. An error correction process can then be applied to some or all of the records before the base is identified.
[0195] Figure 12 is a flow chart of a sequencer 220 using an additive method according to some embodiments. The sequencer 220 may be, for example, as shown and described in Figure 11The sequencing process is performed at step 210 of the exemplary method 200 for sequencing multiple nucleic acid strands (e.g., ssDNA) using the SMAS device 100 discussed above. At 222, the sequencing process 220 begins. At 224, baseline characteristics are obtained for each of the three sensors 105 (e.g., by at least one processor 130 of the SMAS device 100, with the aid of circuitry 120). When the interrogation cycle begins, at 226, a first labeled nucleotide is selected (e.g., referring to steps 1 through 10 above, the first labeled nucleotide would be A). At 228, the selected labeled nucleotide is introduced into the fluid chamber 115 and potentially incorporated into the nucleic acid strand bound to the binding site 116. At 230, unbound nucleotides are flushed away. At 232, characteristics are obtained from each of the plurality of three sensors, and a determination is made as to whether the detection result of each of the plurality of three sensors 105 is determined (e.g., whether a label is detected or not). At 234, S detection results are recorded in S records (e.g., with 1 indicating that a mark has been detected or with 0 indicating that a mark has not been detected). At 236, it is determined whether the nucleotide tested last is the last nucleotide of the query cycle. For the example ordering of the nucleotide test assumed in steps 1 to 10 above, it is determined at 236 (e.g., by at least one processor 130) whether G is the nucleotide tested last. If not, then at 238, select the next labeled nucleotide to be tested in the query cycle, and repeat steps 228 to 236 until the nucleotide tested last is determined to be the last nucleotide of the query cycle at 236. At 240, the mark is cut and rinsed away. At 242, it is determined whether the query cycle completed last (e.g., by at least one processor 130) is the last query cycle of the sequencer 220. For example, the at least one processor 130 can determine whether enough detection results have been recorded so that at least one processor 130 (or some other processing entities, such as an external processor) can determine the bases (e.g., 150 bases) of the target number. If not, the sequencing process 220 returns to step 224. If so, the sequencing process 220 ends at 244. Again, as explained above, the order of the test nucleotides is arbitrary.
[0196] The additive sequencing protocol, which in the exemplary case of DNA sequencing comprises four nucleotide incorporations and one label cleavage reaction, is summarized in Figure 13 middle. Figure 13The leftmost panel of illustrates a sensor array 110 having a total of 100 individual sensors 105, which are shown in squares. For illustrative purposes, it is assumed that each of the 100 binding sites 116 in the sensor array 110 holds a separate DNA strand, and each DNA strand is sensed by a separate sensor 105 (in other words, the binding sites 116 and the sensors 105 are in a one-to-one relationship). Some DNA strands may be copies of other DNAs. Labeled nucleotides are added to the fluid chamber 115 one type at a time, and the labels are cut simultaneously after the nucleotides are incorporated. In the absence of errors, base recognition can be completed after five reactions (that is, four nucleotide incorporations and one base cleavage reaction). If an error occurs, an error correction procedure as described below can be applied.
[0197] Subtractive methods
[0198] In the subtractive method, the sensor 105 detects nanoscale labels 102 bound to nucleotides with cleavable linkers. All four types of nucleotides carry the same type of label (e.g., molecular, fluorescent, magnetic, etc.), but each has a different type of cleavable linker. In one embodiment, a query cycle that, in the absence of errors, will produce four detection results (one of which will (in the absence of errors) be a label detection for each of the plurality of three nucleic acid chains 101) involves the following steps:
[0199] 1. Simultaneously introduce labeled A, T, C, and G nucleotides, incorporate, and wash away unbound labeled molecules. Obtain a baseline characteristic for each of the plurality of S sensors 105 (e.g., by detecting a signal at each of the plurality of S sensors 105). In the absence of errors, all sensors 105 will detect the label.
[0200] 2. Query Step 1: Introduce a reagent (e.g., an enzyme) that cleaves the tag only from the first nucleotide (e.g., A), wash, and obtain a characteristic (e.g., a measurement signal) at each of the plurality of S sensors 105. Determine (e.g., based on a change in baseline characteristic) which sensors 105 are no longer detecting the tag. Save the detection result for each sensor 105 at the location in the record corresponding to Query Step 1 for the current query cycle.
[0201] 3. Query Step 2: Introduce a reagent that cleaves the label only from the second nucleotide (e.g., T), wash, and obtain a characteristic (e.g., assay signal) at each of the plurality of S sensors 105. Determine (e.g., based on a change in baseline characteristic) which sensors 105 are no longer detecting the label. Save the detection result for each sensor 105 at the location in the record corresponding to Query Step 2 for the current query cycle.
[0202] 4. Query Step 3: Introduce a reagent that cleaves the label only from the third nucleotide (e.g., C), wash, and obtain a characteristic (e.g., assay signal) at each of the plurality of S sensors 105. Determine (e.g., based on a change in baseline characteristic) which sensors 105 are no longer detecting the label. Save the detection result for each sensor 105 at the location in the record corresponding to Query Step 3 for the current query cycle.
[0203] 5. Query Step 4: Introduce a reagent that cleaves the label only from the fourth nucleotide (e.g., G), wash, and obtain a characteristic (e.g., assay signal) at each of the plurality of S sensors 105. Determine (e.g., based on a change in baseline characteristic) which sensors 105 are no longer detecting the label. Save the detection result for each sensor 105 at the location in the record corresponding to Query Step 4 for the current query cycle.
[0204] For the next query cycle, steps 1 through 5 can be repeated. It should be understood that the ordering of certain steps 1 through 5 is exemplary, and further, the number and numbering of steps 1 through 5 are for convenience and can be modified. As an example, and as previously explained, the order in which nucleotides are cleaved is arbitrary. Similarly, in step 1, nucleotides can be introduced subsequently (not necessarily simultaneously). As another example, query steps 1, 2, 3, and 4 include introducing reagents, washing, obtaining characteristics, determining which sensors are no longer (or still) detecting the marker, and saving the results in a single step, but it should be understood that each query step can be divided into a series of smaller steps.
[0205] It should be understood that if no errors occur during any query cycle of the subtractive method, the individual bases of the individual strands can be identified (determined) once the removal of the label (absence of the label) is first detected. For example, referring to the above steps, if at query step 1 involving a labeled A nucleotide, for a particular sensor 105, the resulting characteristic indicates that the sensor 105 is no longer detecting the label, then saving the detection result can be equivalent to identifying the base (T) complementary to A for that detector 105 (and binding site 116). Similarly, if at query step 2 involving a labeled T nucleotide, for a particular sensor 105, the resulting characteristic indicates that the sensor 105 is no longer detecting the label, then saving the detection result can be equivalent to identifying the base (A) complementary to T for that detector 105 (and binding site 116). Similarly, if, at query step 3 involving a labeled C nucleotide, the resulting characteristic for a particular sensor 105 indicates that the sensor 105 is no longer detecting the label, then saving the detection result can be equivalent to identifying the base (G) complementary to C for that detector 105 (and binding site 116). Finally, if, at query step 4 involving a labeled G nucleotide, the resulting characteristic for that particular sensor 105 indicates that the sensor 105 is no longer detecting the label, then saving the detection result can be equivalent to identifying the base (C) complementary to G for that detector 105 (and binding site 116). However, as described in further detail below, several types of errors can occur during the sequencing process (e.g., during a subtractive method), and therefore, in some embodiments, a log is created during the sequencing process to record whether a label was detected / not detected during each query step of each query cycle. Error correction procedures can then be applied to some or all of the logs before the base is identified.
[0206] Figure 14 is a flow chart of a sequencing process 250 using a subtractive approach according to some embodiments. The sequencing process 250 may be, for example, as shown and described in Figure 11The sequencing procedure is performed at step 210 of the exemplary method 200 for sequencing multiple nucleic acid strands (e.g., ssDNA) using the SMAS device 100 discussed above. Sequencing procedure 250 begins at 252. At 254, all labeled nucleotides are introduced into fluid chamber 115 and incorporated into the nucleic acid strands bound to the three binding sites 116. At 256, unbound nucleotides are flushed away. At 258, baseline characteristics are obtained for each of the three sensors 105 (e.g., by at least one processor 130 of the SMAS device 100, with the aid of circuitry 120). Assuming that nucleotides have been incorporated into the nucleic acid strands bound to each of the three binding sites, the obtained characteristics represent the characteristics of the sensor 105 when it is detecting at least one label. At 260, one of the cleavable linkers is selected for cleavage (or, equivalently, one of the nucleotides is selected). At 262, the label attached to the selected nucleotide is cleaved and flushed away. Assuming there are no errors, after step 262, the sensors 105 that sense nucleic acid chains that incorporate the tested nucleotides (e.g., nucleic acid chains to which the label is attached via the selected cleavable linker) will exhibit a change in a characteristic (e.g., a change in a signal associated with or generated by the sensor 105). At 264, a characteristic is obtained from each of the plurality of S sensors, and the detection result of each of the plurality of S sensors 105 is determined (e.g., whether the label is detected or not detected). At 266, the S detection results are recorded in S records (e.g., with a 1 indicating that the label is detected or a 0 indicating that the label is not detected). At 268, it is determined whether the last tested nucleotide is the last nucleotide of the query cycle. For the example ordering of nucleotide tests assumed in steps 1 to 5 above, it will be determined at 268 (e.g., by at least one processor 130) whether G is the last tested nucleotide. If not, then at 270 the next cleavable linker to be cleaved in the query cycle (or equivalently, the next nucleotide to be tested) is selected, and steps 262 through 268 are repeated until it is determined at 268 that the last linker cleaved (or equivalently, the last nucleotide tested) is the last linker (or nucleotide) of the query cycle. At 272, a determination is made (e.g., by at least one processor 130) whether the last completed query cycle was the last query cycle of the sequencer 250. For example, the at least one processor 130 may determine whether enough detection results have been recorded to enable the at least one processor 130 (or some other processing entity, such as an external processor) to identify the target number of bases (e.g., 150 bases). If not, the sequencer 250 returns to step 254. If so, the sequencer 250 ends at 274. Again, as described above, the order in which the nucleotides are tested is arbitrary.
[0207] Subtractive sequencing protocols, which in the exemplary case of DNA sequencing comprise one nucleotide incorporation and four base cleavage reactions, are summarized in Figure 15 middle. Figure 15 The leftmost panel of illustrates a sensor array 110 having a total of 100 individual sensors 105, which are shown in squares. For illustrative purposes, it is assumed that each of the 100 binding sites 116 in the sensor array 110 holds a separate DNA strand, and each DNA strand is sensed by a separate sensor 105 (in other words, the binding sites 116 and the sensors 105 are in a one-to-one relationship). Some DNA strands may be copies of other DNAs. All four types of labeled nucleotides are added to the fluid chamber 115 simultaneously, and after incorporation, the labels are removed one type of nucleotide at a time (e.g., a cleavable linker). In the absence of errors, base recognition can be completed after five reactions (that is, one nucleotide incorporation and four base cleavage reactions). If an error occurs, an error correction procedure as described below can be applied.
[0208] Improved additive method
[0209] In the modified additive method, a sensor 105 detects nanoscale labels 102 bound to nucleotides with cleavable linkers. All four types of nucleotides carry the same type of label 102 (e.g., molecular, fluorescent, magnetic, etc.) and use the same type of cleavable linker. The labeled nucleotides are added separately, and after each nucleotide is added, the presence of the label 102 is detected. In one embodiment, a query cycle that, in the absence of errors, will produce four detection results (at least one of which will be a label detection for each of the plurality of three nucleic acid chains 101) involves the following steps:
[0210] 1. Obtain baseline characteristics for each of the plurality of S sensors 105 (which may be all or less than all of the sensors 105 in the sensor array 110 ) of the SMAS device 100 (e.g., by determining a baseline signal at each of the plurality of S sensors 105 ).
[0211] 2. Introduce and incorporate the first labeled nucleotide, such as labeled A nucleotide. Wash away unbound labeled molecules.
[0212] 3. Query step 1: Get the characteristics of each of the plurality of S sensors 105 (e.g., by detecting a signal at each of the plurality of S sensors 105) and determine whether each sensor 105 detects at least one tag. Save the detection result of each sensor 105 at a location in the record corresponding to query step 1 of the current query cycle.
[0213] 4. Cut and rinse off the markings.
[0214] 5. Introduce and incorporate a second labeled nucleotide, such as a labeled T nucleotide. Wash away unbound labeled molecules.
[0215] 6. Query step 2: Obtain characteristics of each of the plurality of S sensors 105 (e.g., by detecting a signal at each of the plurality of S sensors 105) and determine whether each sensor 105 detects at least one tag. Save the detection result of each sensor 105 at a location in the record corresponding to query step 2 of the current query cycle.
[0216] 7. Cut and rinse off the markings.
[0217] 8. Introduce and incorporate a third labeled nucleotide, such as a labeled C nucleotide. Wash away unbound labeled molecules.
[0218] 9. Query step 3: Obtain characteristics of each of the plurality of S sensors 105 (e.g., by detecting a signal at each of the plurality of S sensors 105) and determine whether each sensor 105 detects at least one tag. Save the detection result of each sensor 105 at a location in the record corresponding to query step 3 of the current query cycle.
[0219] 10. Cut and rinse off the markings.
[0220] 11. Introduce and incorporate a fourth labeled nucleotide, such as a labeled G nucleotide. Wash away unbound labeled molecules.
[0221] 12. Query step 4: Obtain characteristics of each of the plurality of S sensors 105 (e.g., by detecting a signal at each of the plurality of S sensors 105) and determine whether each sensor 105 detects at least one tag. Save the detection result of each sensor 105 at a location in the record corresponding to query step 4 of the current query cycle.
[0222] 13. Cut and rinse off the markings.
[0223] Then, for the next query cycle, steps 1 to 13 can be repeated. It should be understood that some of the ordering in steps 1 to 13 is exemplary, and further, the quantity and numbering of steps 1 to 13 are for convenience and can be modified. As an example, and as described above, the order in which nucleotides are introduced is arbitrary. As another example, steps 2, 5, 8 and 11 include introducing and incorporating nucleotides, and washing out unbound nucleotides in a single step, but it should be understood that each of steps 2, 5, 8 and 11 can be divided into a series of smaller steps. Similarly, steps 3, 6, 9 and 12 (respectively query steps 1, 2, 3 and 4) can be further divided into a series of smaller steps (such as obtaining characteristics, determining whether a mark is detected, and preserving detection results). On the contrary, steps can be combined (such as steps 2 and 3 can be combined, steps 3 and 4 can be combined, steps 2 to 4 can be combined, steps 5 and 6 can be combined, steps 6 and 7 can be combined, steps 5 to 7 can be combined, etc.).
[0224] It should be understood that if no errors occur during any query cycle of the improved additive method, the individual bases of each strand can be identified (determined) once the label is detected. For example, referring to the above steps, if at query step 1 involving a labeled A nucleotide, the resulting characteristic for a particular sensor 105 indicates that the sensor 105 detected the label, then saving the detection result can be equivalent to identifying the base (T) complementary to A for the detector 105 (and binding site 116). Similarly, if at query step 2 involving a labeled T nucleotide, the resulting characteristic for the particular sensor 105 indicates that the sensor 105 detected the label, then saving the detection result can be equivalent to identifying the base (A) complementary to T for the detector 105 (and binding site 116). Similarly, if at query step 3 involving a labeled C nucleotide, the resulting characteristic for the particular sensor 105 indicates that the sensor 105 detected the label, then saving the detection result can be equivalent to identifying the base (G) complementary to C for the detector 105 (and binding site 116). Finally, if, at query step 4 involving a labeled G nucleotide, for a particular sensor 105, the resulting characteristic indicates that the sensor 105 detected the label, then saving the detection result can be equivalent to identifying the base (C) complementary to the G for that detector 105 (and binding site 116). However, as described in further detail below, several types of errors can occur during the sequencing process (e.g., during the additive method), and therefore, in some embodiments, a record is created during the sequencing process to record the detection / non-detection of the label during each query step of each query cycle. An error correction process can then be applied to some or all of the records before the base is identified.
[0225] Figure 16is a flow chart of a sequencing process 350 using a modified additive method according to some embodiments. The sequencing process 350 may be, for example, as shown and described in Figure 11 The sequencing process is performed at step 210 of the exemplary method 200 for sequencing multiple nucleic acid strands (e.g., ssDNA) using the SMAS device 100 discussed above. At 352, the sequencing process 350 begins. At 354, baseline characteristics are obtained for each of the three sensors 105 (e.g., by at least one processor 130 of the SMAS device 100, with the aid of circuitry 120). When the query cycle begins, at 356, a first labeled nucleotide is selected (e.g., referring to steps 1 through 13 above, the first labeled nucleotide would be A). At 358, the selected labeled nucleotide is introduced into the fluid chamber 115 and potentially incorporated into the nucleic acid strand bound to the binding site 116. At 360, unbound nucleotides are flushed away. At 362, characteristics are obtained from each of the plurality of three sensors, and a determination is made as to whether the detection result of each of the plurality of three sensors 105 is determined (e.g., detection of a label or non-detection of a label). At 364, S detection results are recorded in S records (e.g., with 1 indicating that a mark is detected or with 0 indicating that a mark is not detected). At 366, the mark is cut and rinsed away. At 368, it is determined whether the nucleotide tested last is the last nucleotide of the query cycle. For the example sorting of the nucleotide test assumed in steps 1 to 13 above, it is determined at 368 (e.g., by at least one processor 130) whether G is the nucleotide tested last. If not, then at 370, the next labeled nucleotide to be tested in the query cycle is selected, and steps 358 to 368 are repeated until the nucleotide tested last is determined to be the last nucleotide of the query cycle at 368. At 372, it is determined whether the query cycle completed last (e.g., by at least one processor 130) is the last query cycle of the sequencer 350. For example, the at least one processor 130 can determine whether enough detection results have been recorded to enable at least one processor 130 (or some other processing entities, such as an external processor) to identify a target number of bases (e.g., 150 bases). If not, the sequencing process 350 returns to step 354. If so, the sequencing process 350 ends at 374. Again, as explained above, the order of the test nucleotides is arbitrary.
[0226] An improved additive sequencing protocol (which, in the exemplary case of DNA sequencing, comprises four nucleotide incorporation and four base cleavage reactions) is described in Figure 17 middle. Figure 17The leftmost panel of FIG illustrates a sensor array 110 having a total of 100 individual sensors 105, which are shown in squares. For illustrative purposes, it is assumed that each of the 100 binding sites 116 in the sensor array 110 holds a separate DNA strand, and each DNA strand is sensed by a separate sensor 105 (in other words, the binding sites 116 and sensors 105 are in a one-to-one relationship). Some DNA strands may be copies of other DNA strands. As shown and described, labeled nucleotides are added to the fluid chamber 115 one type at a time, and the label is cleaved after incorporation and label detection. In the absence of errors, base recognition can be completed after an average of 5 reactions (that is, 2.5 nucleotide incorporations and 2.5 base cleavage reactions).
[0227] Thus, in the absence of errors, for DNA sequencing, the modified additive method produces at least one base call per ssDNA after 8 reactions (4 nucleotide incorporations and 4 base cleavages) to test all bases. However, on average, base calls are made after only 5 reactions (2.5 nucleotide incorporations and 2.5 base cleavages). Because the label is removed after each nucleotide is introduced, it is possible to perform base calls in a single Multiple nucleotides are incorporated and identified during the query cycle. Specifically, in an unknown ssDNA sequence, there is a one-in-four chance that the unknown base is a T. If the base is exactly a T, it will be detected after one incorporation and one base cleavage reaction when an A nucleotide is introduced at the third step. There is a one-in-four chance that the unknown base is an A. If the base is exactly an A, it will be detected during the query cycle. The fifth step is detected when a T nucleotide has been introduced and two introductions and two cuts have been performed. The probability that the unknown base is a G is one in four. If the base is exactly G, then the query loop will be The seventh step is detected when a C nucleotide has been introduced and three introductions and three cleavages have been performed. Finally, there is a one in four chance that the unknown base is a C. If the base is exactly C, then the query loop will be The eleventh step is detected when a C nucleotide has been introduced and four introductions and four cleavages have been performed. Therefore, an average of 2.5 queries (5 reactions) are required. Alternatively, if the unknown 4-base sequence of a particular ssDNA happens to be the best-case ATCG (for the chosen order of introduced nucleotides assumed for this example), only one query cycle is required. A total of 8 reactions (4 nucleotide incorporations and 4 base cleavages), or 2 reactions per base call. However, if the unknown sequence happens to be, for example, GCTA, GGCT, GCTT, GGGG, etc., four query cycles are required, each including all This results in a total of 32 reactions (16 nucleotide incorporations and 16 base cleavages), or 8 reactions per base call. However, on average, for a random DNA sequence, 2.5 queries or 5 reactions (2.5 nucleotide incorporations and 2.5 base cleavages) are required to make one base call.
[0228] Sources of Sequencing Errors
[0229] Ideally, the sequencing process, whether in a CLUS device or a SMAS device 100, would be error-free. In other words, for example, nucleotides would always be correctly labeled, nucleotides would always be correctly incorporated into DNA, all labels would be successfully cleaved during the cleavage step, and all cleaved labels would be successfully washed away. However, in reality, errors can occur during any sequencing process. This section explores the sources of sequencing errors in both the CLUS device and the SMAS device 100 and describes error mitigation strategies for the SMAS device 100. As further described below, error correction methods can be used to improve the sequencing accuracy of the SMAS device 100.
[0230] Because the modified additive method described above is a conceptually simple (and symmetrical, since each nucleotide is processed in the same way) sequencing procedure, it is a good model for illustrating how errors propagate in both the CLUS device and the SMAS device 100. Considering four sources of error, it is assumed that the nanoscale tags are attached to the nucleotides via cleavable linkers. Each error occurs at a ratio denoted as r, which has a value between 0 and 1. The four sources of error are:
[0231] Failed nucleotide incorporation (FNI) Failed nucleotide incorporation (FNI) occurs when a correctly labeled nucleotide molecule has not yet reached the ssDNA binding site or the polymerase fails to incorporate it. Figure 18A The FNI of the CLUS device is shown for five examples of sequencing ssDNA. After the flow of complementary nucleotides, only three of the five ssDNAs have incorporated labeled nucleotides (illustrated with magnetic labels). Thus, two out of five nucleotides (r=0.4) cannot be incorporated. Figure 18B The FNI of the SMAS device 100 is illustrated. Each of the five binding sites 116 holds an instance of ssDNA. After the flow of complementary nucleotides, only three of the five ssDNAs (those bound to binding sites 116A, 116B, and 116C) have incorporated labeled nucleotides (illustrated as having magnetic labels for example purposes only). Furthermore, two of the five ssDNA instances (r = 0.4) were unable to incorporate nucleotides.
[0232] Failed Mark Removal (FLR)Failed label removal (FLR) occurs when a labeled nucleotide molecule is incorporated but the label is not removed after label detection because the cleavage reagent has not reached the linker or fails to cleave it. Figure 18C Explanation: Figure 18A FLR of the CLUS device described in the discussion of . After incorporation of complementary nucleotides and washing to remove unincorporated nucleotides, detection of the label, and cleavage and washing of the label, one label remains attached to one of the ssDNA instances (r = 0.2). Similarly, in Figure 18D The description above is in Figure 18B In the FLR of the SMAS device 100 described in the discussion of , after incorporation of complementary nucleotides and washing to remove unincorporated nucleotides, detection of the label, and cleavage and washing of the label (e.g., steps 2 to 4, 5 to 7, 8 to 10, and / or 11 to 13 described above), the label remains attached to the ssDNA at binding site 116A (r = 0.2).
[0233] Failed nucleotide removal (FNR) Failed nucleotide removal (FNR) occurs when labeled nucleotides (whether complementary or non-complementary) bind non-specifically to the binding site 116 and / or the surface of the sensor 105. Figure 18E Explanation: Figure 18A An example of a FNR of a CLUS device described in . After the flow of nucleotides and flushing to remove unbound nucleotides, two bad nucleotides and their labels remain on the surface of the binding site. Similarly, in Figure 18F In the figure, the above Figure 18B In the FNR of the SMAS device 100 described in the discussion of FIG, after nucleotide flow and flushing to remove unbound nucleotides, one poor nucleotide remains on the surface of binding site 116A and another poor nucleotide remains on the surface of binding site 116D. In this example, r = 0.4 for both the CLUS device and the SMAS device 100.
[0234] Failed Label Detection (FLD) : Failed label detection (FLD) occurs when the correct complementary nucleotide is incorporated, but the label is not detected because it is missing or the sensor fails to recognize it. Figure 18G Explanation: Figure 18A After incorporation of complementary nucleotides and washing to remove unincorporated nucleotides, two of the ssDNA instances had incorporated complementary nucleotides, but the label was missing (r = 0.4). Similarly, in Figure 18H In the figure, the above Figure 18BThe FLD of the SMAS device 100 described in the discussion should be attached to the labeled deletion (r = 0.4) of the nucleotides incorporated into the ssDNA at the binding sites 116C and 116D after incorporating the complementary nucleotides and rinsing to remove unbound nucleotides (e.g., steps 2, 5, 8, or 11 described above).
[0235] Figures 18A to 18H The label is described as a magnet, thus indicating magnetic labels and magnetic sensors, but it should be understood that, as described above, the label can be any type of detectable label (e.g., fluorescent, magnetic, etc.) and the sensor can be any type of sensor capable of detecting the selected type of label (e.g., optical, magnetic, organometallic, charged molecules, etc.).
[0236] Assume that four error types (FNI, FLR, FNR, and FLD) occur at the same rate r, where 0 < r < 1; for example, if r = 0.01, then on average 1 in 100 cases fails. Also assume that the sensor 105 of the SMAS device 100 (e.g., a nanoscale sensor 105) can detect a single label almost every time, and the response of the large cluster sensors used in the CLUS device is linear; for example, the sensors of the CLUS device can distinguish between N and N + 1 labeled strands for all values of N.
[0237] Cluster sequencers and single-molecule array sequencers: qualitative comparison and error correction
[0238] Two types of error correction are disclosed herein, referred to as deterministic error correction and probabilistic error correction. The SMAS device 100 can use one or both types of error correction, as further described below.
[0239] As described above, the modified additive method is a good model for illustrating how errors propagate and how the disclosed error correction algorithms can be implemented. It should be understood that the disclosed error mitigation algorithms can also be applied when using other sequencing methods, such as additive methods or subtractive methods.
[0240] Consider a CLUS device and a SMAS device 100 using the modified additive method sequencing procedure, which has a large error rate of r = 0.1 (e.g., 1 failure in 10 reactions) and a small number of instances of the (ideally identical) strands; for example, N = K = 3, where the variable N represents the cluster size used in the CLUS device, and the variable K represents the number of sensors 105 of the SMAS device 100 sensing the same DNA strand instance. (As described previously, the K sensors can be close to each other, or they can be spread within the sensor array 110). To describe an embodiment of deterministic error correction, initially only consider FNI and FLR errors. Then consider FNI, FLR, and FLD errors, and describe the error mitigation strategy. Finally, consider all four types of errors, and describe the error correction procedure for resolving all four types of errors.
[0241] When using the SMAS device 100, FLR errors can be detected and removed, either during the sequencing process or at a later time in real time. FLR errors can be detected by characterizing each of the three sensors 105 after cutting and flushing the marker. FNI errors can be detected by examining the records of each sensor 105 and identifying query cycles in which the sensor 105 failed to detect any marker. Therefore, the modified additive method can be adjusted to add these detection steps as follows according to one embodiment:
[0242] 1. Obtain baseline characteristics for each of the plurality of S sensors 105 (which may be all or less than all of the sensors 105 in the sensor array 110 ) of the SMAS device 100 (e.g., by determining a baseline signal at each of the plurality of S sensors 105 ).
[0243] 2. Introduce and incorporate the first labeled nucleotide, such as labeled A nucleotide. Wash away unbound labeled molecules.
[0244] 3. Query step 1: Get the characteristics of each of the plurality of S sensors 105 (e.g., by detecting a signal at each of the plurality of S sensors 105) and determine whether each sensor 105 detects at least one tag. Save the detection result of each sensor 105 at a location in the record corresponding to query step 1 of the current query cycle.
[0245] 4. Cut and rinse off the markings.
[0246] 5. Get a characteristic for each of the plurality of S sensors 105 that detected the tag in step 3. If the obtained characteristic for any of those sensors 105 indicates that the sensor 105 is still detecting the tag, then the chemistry failed to cleave the tag (e.g., there is an FLR error for that sensor).
[0247] 6. Introduce and incorporate a second labeled nucleotide, such as a labeled T nucleotide. Wash away unbound labeled molecules.
[0248] 7. Query step 2: Get the characteristics of each of the plurality of S sensors 105 (e.g., by detecting a signal at each of the plurality of S sensors 105) and determine whether each sensor 105 detects at least one tag. Save the detection result of each sensor 105 at a location in the record corresponding to query step 2 of the current query cycle.
[0249] 8. Cut and rinse off the markings.
[0250] 9. Get a characteristic for each of the plurality of S sensors 105 that detected the tag in step 7. If the obtained characteristic for any of those sensors 105 indicates that the sensor 105 is still detecting the tag, then the chemistry failed to cleave the tag (e.g., there is an FLR error for that sensor).
[0251] 10. Introduce and incorporate a third labeled nucleotide, such as a labeled C nucleotide. Wash away unbound labeled molecules.
[0252] 11. Query step 3: Obtain characteristics of each of the plurality of S sensors 105 (e.g., by detecting a signal at each of the plurality of S sensors 105) and determine whether each sensor 105 detects at least one tag. Save the detection result of each sensor 105 at a location in the record corresponding to query step 3 of the current query cycle.
[0253] 12. Cut and rinse off the markings.
[0254] 13. Get a characteristic for each of the plurality of S sensors 105 that detected the tag in step 11. If the obtained characteristic for any of those sensors 105 indicates that the sensor 105 is still detecting the tag, then the chemistry failed to cleave the tag (e.g., there is an FLR error for that sensor).
[0255] 14. Introduce and incorporate a fourth labeled nucleotide, such as a labeled G nucleotide. Wash away unbound labeled molecules.
[0256] 15. Query Step 4: Determine the characteristics of each of the plurality of S sensors 105 (e.g., by detecting a signal at each of the plurality of S sensors 105) and determine whether each sensor 105 detects at least one marker. The detection result of each sensor 105 is stored in a location in the record corresponding to Query Step 4 for the current query cycle. If there are sensors 105 that were not assigned a base for the query cycle (e.g., sensors 105 that were unable to detect an A, T, C, or G during the query cycle), then the chemistry was unable to incorporate the nucleotide (e.g., FNI was present for these sensors 105).
[0257] 16. Cut and rinse off the markings.
[0258] 17. Get a characteristic for each of the plurality of S sensors 105 that detected the tag in step 15. If the obtained characteristic for any of those sensors 105 indicates that the sensor 105 is still detecting the tag, then the chemistry failed to cleave the tag (e.g., there is an FLR error for that sensor).
[0259] Can then repeat step 1 to 17 for the next query cycle (for example, to estimate the next base or to read the current base again if the previous query cycle cannot read the current base). It should be understood that some of the ordering in steps 1 to 17 is exemplary, and further, the quantity and numbering of steps 1 to 17 are for convenience and can be modified. As an example, and as described above, the order in which nucleotides are introduced is arbitrary. As another example, steps 2, 6, 10 and 14 include introducing and incorporating nucleotides, and washing away unbound nucleotides in a single step, but it should be understood that each of steps 2, 6, 10 and 14 can be divided into a series of smaller steps. Similarly, steps 3, 7, 11 and 15 (respectively query steps 1, 2, 3 and 4) can be further divided into a series of smaller steps (for example, obtaining characteristics, determining whether a tag is detected, and preserving test results). Similarly, although step 15 includes identifying FNI errors, the task can be performed in a separate step. Rather, steps may be combined (eg, some or all of steps 2 to 5, some or all of steps 6 to 9, some or all of steps 10 to 13, some or all of steps 14 to 17, etc.).
[0260] Figure 19 is a flow chart of an exemplary sequencing process 400 using the improved additive method with FLR and FNI error detection according to some embodiments. The sequencing process 400 may be, for example, as shown and described in Figure 11The sequencing process is performed at step 210 of the exemplary method 200 for sequencing multiple nucleic acid strands (e.g., ssDNA) using the SMAS device 100 discussed above. At 402, the sequencing process 400 begins. At 404, baseline characteristics are obtained for each of the three sensors 105 (e.g., by at least one processor 130 of the SMAS device 100, with the aid of circuitry 120). When the query cycle begins, at 406, a first labeled nucleotide is selected (e.g., referring to steps 1 through 17 above, the first labeled nucleotide would be A). At 408, the selected labeled nucleotide is introduced into the fluid chamber 115 and potentially incorporated into the nucleic acid strand bound to the binding site 116. At 410, unbound nucleotides are flushed away. At 412, characteristics are obtained from each of the plurality of three sensors, and a determination is made as to whether the detection result of each of the plurality of three sensors 105 is determined (e.g., whether a label is detected or not). At 414, the S detection results are recorded in S records (e.g., as a 1 indicating that a tag was detected or a 0 indicating that a tag was not detected). At 416, the tag is cleaved and flushed away. At 418, the characteristics of those sensors 105 that detected the tag during steps 412 / 414 are obtained. At 420, a determination is made as to whether any of the sensors 105 that detected the tag during steps 412 / 414 are still detecting the tag. If so, at 422, a determination is made as to whether an FLR error has been detected for the sensor 105 that is still detecting at least one tag, even though the tag was cleaved and flushed away at 416. The sequencing process 400 then continues to 424. If it is determined at 420 (e.g., by at least one processor 130) that none of the sensors 105 that detected the tag during steps 412 / 414 are still detecting the tag, the sequencing process also continues to 424. At 424, a determination is made as to whether the last nucleotide tested is the last nucleotide of the query cycle. For the example sequence of nucleotide tests hypothesized in steps 1 through 17 above, a determination is made at 368 (e.g., by at least one processor 130) as to whether G is the last nucleotide tested. If not, the next labeled nucleotide to be tested in the query cycle is selected at 426, and steps 408 through 420 (and, if applicable, 422) are repeated until, at 424, it is determined that the last nucleotide tested is the last nucleotide of the query cycle. At 428, FNI errors are detected for those of the S sensors 105 that were unable to detect any markers during the last completed query cycle. At 430, a determination is made (e.g., by at least one processor 130) as to whether the last completed query cycle was the last query cycle of the sequencer 400. For example, the at least one processor 130 may determine whether sufficient detection results have been recorded to enable the at least one processor 130 (or some other processing entity, such as an external processor) to identify the target number of bases (e.g., 150 bases). If not, the sequencer 400 returns to step 404.If so, the sequencing process 400 ends at 432. Again, as explained above, the order of the test nucleotides is arbitrary.
[0261] Mitigating FNI and FLR errors
[0262] To illustrate the effects of FNI and FLR errors on the CLUS device and the SMAS device 100, each type of sequencer was used to identify an exemplary DNA sequence in which FNI and FLR errors occurred randomly when reading the sequence using the modified additive method of SBS described above. Assume that the error rates of FNI and FLR errors are both An exemplary sequence is: TAG CAA GGT CCGCTA CTG GCA GAC TGG. Figure 20 Shown throughout 18 In the query loop of the query step, There are two types of errors generated. Figure 20 As shown in Figure 1, approximately 1 in 10 reactions failed, and for the three ssDNA samples sequenced, errors were evenly distributed between FNI and FLR errors. The model scenario represents one of many possible scenarios for collective behavior. The consequences of FNI and FLR errors on base call accuracy were analyzed for the cases when the three DNA strands were placed on a single sensor of a CLUS device and when they were placed on three discrete nanoscale sensors 105 of a SMAS device 100.
[0263] Figure 21 Illustrate the expected signal levels detected by the CLUS device sensor, which captures the behavior of the molecular collective during the sequencing procedure. At each query step, the CLUS device sensor can detect four signal intensity levels of the molecular collective (composed of three ssDNAs): that is, 0 tags, 1 tag, 2 tags or 3 tags are detected. The sequencing process of the CLUS device takes into account the combined signal of the collective and cannot distinguish when the reaction to an individual chain fails. Whenever the CLUS device sensor senses at least two tags, a base is identified at a specific query step. The threshold can be represented by a decision criterion: when the CLUS sensor signal level is greater than 1.5, a base is identified. As Figure 21 This indicates that a high rate of chemical failures can lead to significant base calling errors and very low base calling accuracy. The CLUS device method resulted in only 6 of 21 (about 29%) bases being called that matched the true sequence. This level of accuracy is only slightly better than random guessing with 25% accuracy (since there are 4 bases, there is a one in four chance of correctly guessing a base). In addition, the CLUS device cannot distinguish between successful and failed chemical reactions, nor is it aware of the Figure 20The positions of FNI (dashed circles) or FLR (circles filled with backslashes) errors in the CLUS apparatus are shown. For the CLUS apparatus, collective averaging obscures the exact positions of FLR errors. Probabilistic error correction algorithms can only be implemented to slightly improve the quality of base calls by essentially making educated guesses about the positions of base insertions, deletions, and substitutions. Exemplary algorithms are described, for example, in A. Cacho et al., "A Comparison of Base-calling Algorithms for Illumina Sequencing Technology," Briefings in Bioinformatics, Vol. 17(5), 786-795, 2016; W. C. Cao et al., "BayesCall: A model-based base-calling algorithm for high-throughput short-read sequencing," Genome Res., Vol. 19(10), 1884-1895, 2009; and C. Ledergerber and C. Dessimoz, "Base-calling for next-generation sequencing platforms," Briefings in Bioinformatics, Vol. 17(5), 786-795, 2016. Bioinform.), Vol. 12, 489–97, 2011.
[0264] Figure 22This article describes how the SMAS device 100 can provide improved accuracy when using the error correction techniques described herein. As described above, FLR errors occurring during the sequencing process can be detected during the sequencing process. Specifically, the SMAS device 100 knows (or can find) the position of the FLR because it obtains the characteristics (e.g., signal level) of each sensor 105 and records them after the label is cleaved and washed away and before the next nucleotide is introduced. FLR errors can be corrected by treating them as "label not detected" during base calling. In other words, if the record of the sequencing process contains binary (e.g., 0 / 1) entries for each query step, FLR can be corrected by changing the values at those query steps from "detected" to "not detected." As a specific example, if 0 represents label not detected and 1 represents label detected, then, before error correction, the FLR at the mth query step will be represented by a 1 in the mth position of the record. This error can be corrected by changing the value of 1 at the mth position in the record to a value of 0. Figure 22 The top portion of illustrates the detection results of each of the three sensors 105 of the SMAS device 100 before error correction to remove the FLR error. Figure 22 The lower part shows the results of correcting FLR errors before calling bases.
[0265] The improved additive sequencing procedure using the SMAS device 100 allows a base to be identified during a particular query step when more than half of the K sensors 105 (in the example where K = 3 (two or three sensors 105)) detect a marker during the query step. However, unlike the CLUS device, the SMAS device 100 collects considerably more information because it detects the presence or absence of the marker at each of the multiple (assumed to be 3 in the example) binding sites 116 and at each query step of the sequencing procedure. Therefore, using the SMAS device 100 may result in fewer base calls being made, but those calls result in a significantly more accurate estimated sequence than those called by the CLUS device. Specifically, for the exemplary sequence, once the FLR errors (e.g., Figure 22 ), using the SMAS device 100 resulted in 11 of the 16 (approximately 69%) identified bases being consistent with the true sequence. Figure 21 and 22 This demonstrates that the consequences of chemical failure on base calling accuracy are significantly different for the two types of sequencing devices, with the SMAS device 100 providing better accuracy.
[0266] When using the SMAS device 100, FNI errors can also be corrected because failed incorporation creates a characteristic signature in the SMAS sensor 105 detection results (e.g., in a record consisting of detected / not detected markers by the sensor 105 during a sequencing procedure). In particular, FNI errors in the modified additive method result in a string (continuous sequence) of zeros (or other "not detected marker" detection results) for four or more consecutive interrogation steps. Figure 19 As explained in the discussion of , some FNI errors can be detected by identifying that a particular sensor 105 does not detect any tags during a query cycle. It should be understood that FNI errors can also "span" multiple query cycles. For example, assuming that During the first query cycle of the query steps, a particular sensor 105 detects a tag during the A? query step, and then does not detect any tags until the C? query step of the next query cycle. Because the C? query step follows the A? query step in the exemplary query cycle, and the modified additive method is used as the sequencing cycle, the C? query step of the first query cycle should have resulted in the detection of a tag. It should be noted that Figure 19 Step 428 will not result in any FNI errors being detected during the first or second query cycles, since neither query cycle will result in a particular sensor 105 failing to detect a tag. However, inspection of the test result record will reveal the presence of an FNI error. FNI errors can be deterministically corrected by deleting strings (four in the case of DNA sequencing) of zeros to align the bad strand with a strand not affected by the FNI error. Figure 23 The FNI error is corrected by deleting several strings of four "not detected flag" entries in the record of the test results of the sequencer. Figure 23 As shown in , FNI error correction results in a perfect alignment between the called and true sequences.
[0267] Qualitative analysis of a simplified model system with a finite set of errors indicates that nucleic acid sequencing using the SMAS device 100 greatly outperforms that using the CLUS device, at least when the number of instances K of sequenced DNA strands is small and the chemical failure rate is high. To set the framework for a quantitative comparison of the two platforms, we explore below how cluster size (for the CLUS device) and the number of instances sequenced (for the SMAS device 100) affect base calling accuracy. For both FNI and FLR errors, consider the case where N=K=11 and r=0.1. Assume that the sensor is reading the same instance sequence considered above (TAG CAA GGT CCG CTACTG GCA GAC TGG) and 18 The query loop of the query step randomly caused chemical errors in FNI and FLR. Figure 24 Description of the high chemical failure rate ( or 10%) of the DNA chain based on 11 exemplary SBS reaction results. Figure 24 As shown in Figure 2, approximately 1 in 10 reactions failed.
[0268] Figure 25 This figure illustrates the effect of larger cluster size N on the base calling accuracy of the CLUS device. Figure 25 The expected signal level detected by the CLUS device sensor, which captures the behavior of the molecular collective during the sequencing process, is shown. At each query step, the CLUS device sensor can detect any of twelve signal intensity levels for the molecular collective (eleven ssDNAs), that is, 0 to 11 markers are detected. When the signal level detected by the CLUS sensor is greater than 5.5, the base is identified at the specific query step. Figure 25 It was shown that the failed chemistry resulted in base calling errors: only 11 of the 18 (about 61%) called bases matched the true sequence.
[0269] Figure 25 and Figure 21 Comparisons indicate that the CLUS system achieves better accuracy with N = 11 than when N = 3. Specifically, increasing the cluster size N leads to a significant reduction in base call errors. While with N = 3, only approximately 29% of the identified bases were consistent with the true sequence, increasing the cluster size to N = 11 increases the consistency rate to approximately 61% because the CLUS system benefits from the collective behavior of a larger group. Current state-of-the-art commercial CLUS-based sequencers operate with cluster arrays that hold approximately 100 DNA strand instances.
[0270] Figure 26 The results are shown for the case where the SMAS device 100 is used with K=11 (in other words, 11 instances of ssDNA, each detected by a different sensor 105) and deterministic error correction of FLR and FNI errors according to some embodiments. A base is identified at a particular query step when more than half (e.g., at least 6 for K=11) of the sensors 105 detect a marker. Figure 26 As shown, implementing deterministic FLR error correction (middle) and FNI error correction (bottom) as described above results in a perfect alignment between the identified sequence and the true sequence. It should be noted that without error detection / correction, the identified sequence based on data from the SMAS device 100 would be identical to the sequence identified using data from the CLUS device, since the SMAS device 100 without error correction simply recreates the collective result by summing all individual sensor results. The ability to detect and correct errors in sequencing data provides the SMAS device 100 with an advantage over the CLUS device.
[0271] Therefore, if only FNI and FLR errors occur, using the SMAS device 100 in conjunction with deterministic error correction can result in perfect agreement between the true sequence and the identified sequence. Furthermore, if only FNI and FLR errors occur, an error-free sequence can actually be identified using only a single sensor 105 reading a single ssDNA and the deterministic error correction techniques discussed above (e.g., changing FLR to "not detected flag" and / or deleting a number of strings of a specified length (e.g., 4) of "not detected flags" from the recording of the test results).
[0272] However, when FNR and / or FDL errors are introduced, it is generally unlikely that all errors in the test result record can be eliminated using only deterministic error correction. To address FNR and / or FDL errors, probabilistic error correction may be included in addition to or in place of deterministic error correction.
[0273] Mitigating FNI, FLR, and FNR errors
[0274] This section further includes FNR errors in the analysis. The impact of this type of error on the base calling accuracy of the CLUS device is equivalent to the impact of FNR and FLR due to the inherent averaging when the CLUS device detects labels in a cluster of nucleic acid instances. FNR errors are significantly more detrimental to the performance of sequencing methods using the SMAS device 100 because FNR errors cannot be corrected deterministically. (It should be noted that the CLUS device itself cannot correct FNR errors at all. Instead, the CLUS device relies on collective behavior to mitigate the impact of FLR and other types of errors.)
[0275] Figure 27 To illustrate the problem introduced by the FNR error in the exemplary sequence (TAG CAA GGT CCG CTA CTG GCA GAC TGG), assume that FNI, FLR and now FNR errors are present in 18 The query step occurs randomly during the query cycle. For example purposes, assuming K = 3 (that is, each of the three binding sites 116 holds an instance of a specific ssDNA, and each of the three respective sensors 105 senses a respective one of the three ssDNA instances), there are 15 failures in an average of 100 reactions (r = 0.15, which is a large chemical failure rate), and the errors are evenly distributed between FNI errors, FLR errors, and FNR errors. Under the example conditions and assumptions made here, given only the data record created by SBS using the SMAS device 100, it is impossible to distinguish correctly detected events ( Figure 27 ) and FNR (circles filled with forward slashes). Figure 28Illustrates the results when a base is identified under a marker detected in more than half (at least 2 out of 3) of the sensors S1, S2, and S3. While FLR errors can be deterministically corrected (as described above, by treating them as "mark not detected"), FNR errors cannot be identified because they are indistinguishable from correct marker detection events. Therefore, in this example, only 8 of the 17 (about 47%) identified bases match the true sequence. Therefore, the introduction of FNR errors makes deterministic FNI error correction more challenging because FNR errors destroy the string of four or more "mark not detected" detection results that might otherwise have been removed. If FNI error correction is implemented unprocessed by deleting strings of four zeros in an attempt to align the bad chain with the chain not affected by the error, then sequencing accuracy will not be improved. In fact, if Figure 29 As shown in , for this example, base calling accuracy appears to have deteriorated because after removing the strings of four "undetected marker" calls, only 9 out of 20 (45%) base calls were consistent with the true sequence.
[0276] Error correction can be improved by applying probabilistic error correction to mitigate FNR errors in addition to FLR and FNI errors. For example, note the thymine query step at position 2 (query step 2 of query cycle 1). Sensors S1 and S3 detect the marker, but S2 does not. S2 cannot detect the marker because of either simultaneous FNR errors at sensors S1 and S3 or due to a FNI error at sensor S2. Assuming the probability of each error is r, the probability of simultaneous FNR errors at sensors S1 and S3 is r 2 , and the probability of an FNI error at sensor S2 is r. The error correction algorithm (e.g., performed by at least one processor 130 or another processor) assumes that the more likely event (the presence of an FNI error at sensor S2) has occurred and deletes all entries in the S2 record that shift the S2 detection results from positions 2 to 5 from the data record that captures the detection results from sensor S2. Therefore, the detection results in the S2 record are compared with the detection results generated by sensors S1 and S3, as shown in FIG. Figure 30 Previously (before deletion) at position 4 (in Figure 30 The detection of the G marker at position 4 (in the portion marked “A” of ) can now be attributed to FNR since sensors S1 and S3 did not detect the marker in position 4 (query step 4 of query loop 1).
[0277] Can be in position 13 (such as Figure 30 The same error correction process is performed from left to right at 32 (labeled "C") and 46 (labeled "D") to show the gradual improvement of the alignment between the S1, S2 and S3 records of the test results, as shown in the portion labeled "B" of the figure. Figure 30 As described in the section marked "E". Figure 30 The portion marked "E" indicates that despite performing multiple probabilistic error correction steps to align the outputs of all sensors S1, S2, and S3, it does not seem to improve the alignment between the identified sequence and the true sequence. Even after error correction, only 9 bases out of 20 (45%) were correctly identified. In other words, base calling errors still occurred. Specifically, after the error correction procedure, all three sensors S1, S2, and S3 reported that they had detected the marker at the query step where the marker should have been detected, but some of the sensors also detected markers at positions 10, 22, 40, and 50 (shown in Figure 2). Figure 30 Marks incorrectly incorporated by FNR at ).
[0278] When the detection results of more than half of the sensors 105 are consistent (after error correction), the base is identified as resulting in a thymine insertion error at sequence position 8 (query step 22), where both sensors S1 and S3 detected a label bound to a non-complementary nucleotide during the same query step. (It should be understood that the reason why a thymine insertion error is known at position 8 is because the error data is established for illustrative purposes and is known. In one embodiment, the sensors 105 only indicate whether the label was detected during the query step, and do not indicate whether the detection (or lack of detection) is correct or incorrect. Therefore, in one embodiment, the error at query step 22 will be essentially indistinguishable from a correct detection result.) The correctly aligned true sequence and the identified sequence clearly showing the position of the single incorrect base insertion can be presented as:
[0279] Error: |Insert Actual sequence: TAG CAA G*G TCC GCT ACT GGC
[0280] Recognized sequence: TAG CAA GTG TCC GCT ACT GGC
[0281] *Insert position
[0282] If the base calling rule is modified to require that all three sensors S1, S2, and S3 agree, this insertion error can be corrected. For this rule, all three sensors S1, S2, and S3 must encounter an FNR error at the same time to result in an incorrect base call. The probability of this event is only r 3Assuming r = 0.05, all three sensors S1, S2, and S3 encounter FNR events during the same query step only 125 times out of 100,000 queries on average (or a probability of 0.000125), which is very low even for the extremely high error rate used in the current example. However, if FLD errors also occur, implementing this rule can lead to incorrect identification, as discussed further below.
[0283] Mitigating FNI, FLR, FNR, and FLD errors
[0284] The general error correction strategy used in some embodiments addresses and mitigates all four types of chemical failures that lead to FNI, FLR, FNR, and FLD errors. Figure 31 Explain the example sequence (TAG CAAGGT CCG CTACTG GCA GACTGG), assuming FNI, FLR, FNR and now also FLD errors in 18 For the purpose of creating many errors in the sequencing data to provide a medium for illustrating an exemplary error correction procedure, a very high average error rate of 1 in 5 failed reactions ( (or a 20% error rate), and also assumed that errors are evenly distributed between FNI errors, FLR errors, FNR errors, and FLD errors. Thus, approximately 20 out of 100 reactions fail, and these failures are equally distributed among the four error types. It will be appreciated that this high error rate is unlikely to occur in practice, and therefore the difficulty of the examples considered here is likely much higher than would be encountered in a real-world implementation.
[0285] Under the example conditions and assumptions made here, given only the data record established by SBS using the SMAS device 100, it is impossible to distinguish between the incorporation of a correct nucleotide and FNR, nor can it distinguish between the non-incorporation of a correct nucleotide and FNI. Although FLR errors can be deterministically detected and corrected as described above (by checking the sensor 105 after cutting and washing away the mark, and treating FLR as "no mark detected"), FNR errors cannot be identified because they cannot be distinguished from correctly detected events, and FNI and FLD errors cannot be identified because they cannot be distinguished from not incorporating the correct nucleotide. However, error mitigation can still be accomplished using probabilistic error correction techniques. For example, as described above, when fewer than all sensors S1, S2, and S3 detect or do not detect a mark during a particular query step, the probabilities of two (or more) events can be calculated, and the event with the highest probability can be assumed to be the correct event, and an appropriate error correction step can be adopted.
[0286] Figure 32The application of the error correction procedure to data captured during SBS under the conditions and assumptions described above is illustrated. Figure 32 The portion marked "A" of FIG is the original data before the FLR error is removed. Assuming that the location of the FLR error is known by examining the sensor 105 signal level after cutting and washing out the mark as described above. The FLR error can be completely removed using deterministic error correction, that is, by changing the "mark detected" value (e.g., 1 or "yes") in the data record corresponding to the location of the query step where the FLR error was detected to a "mark not detected" value (e.g., 0 or "no"). It should be noted that in the data record shown in FIG, the FLR error is not detected. Figure 31 During query cycle 15 in
[15] , an FLD error in the data from sensor S2 is followed by an FLR error. In other words, sensor S2 is unable to detect the label of the incorporated nucleotide during the first query step of query cycle 15. The signal level of sensor S2 is checked when the label is cleaved after the first query step of cycle 15 and before the second query step of cycle 15. This check reveals the presence of label at sensor S2, which would be considered an FLR error because all label should have been cleaved and flushed away after the last query step. Therefore, even when an FLR error follows another error, it is detectable and can be removed.
[0287] Figure 32 The portion of FIGURE labeled "B" shows a record of the detection results after FLR errors have been removed via deterministic error correction, applied as previously described. The data record shown in "B" now only includes an indication of whether a marker was detected or not detected by each of sensors S1, S2, S3 at each of the (4×18) query steps shown. (It will be appreciated that the record is comparable to Figure 32 As explained above, it is not known from these records which "mark detected" indications are correct and which are FNR errors, and it is not known which "mark not detected" indications are correct and which are FNI or FLD errors. Therefore, probabilistic error correction can be used to estimate the sequence.
[0288] To illustrate how probabilistic error correction can be applied, the following table shows Figure 32 The FLR error has been removed (e.g. from Figure 32 The following table shows the data records for the first five query cycles (query steps 1 to 20) of the three sensors S1, S2, and S3 after the last detection (the record marked "B" in the table). In other words, the following table shows the first 20 detection results after deterministic error correction removes FLR errors. For query steps where the sensor detects a mark, the table contains a value of 1, and for query cycles where the sensor does not detect a mark, the table contains a value of 0:
[0289]
[0290] As explained above, a simple majority vote after removing the FLR errors would result in only 8 of the 17 bases being correctly identified, as Figure 32 Probabilistic error correction can provide significant improvements as described below.
[0291] Considering query step 2 as an example, sensors S1 and S3 both detected the tag (entry 1 in the table above), but sensor S2 did not detect the tag (table entry 0). Therefore, either sensors S1 and S3 are erroneous, or sensor S2 is erroneous. By considering the probabilities of various events that could lead to each of these outcomes, the error correction algorithm can mitigate errors in the sequenced data. Specifically, because the FLR has been removed from the data record, the only way for sensors S1 and S3 to both incorrectly detect the tag during query step 2 is if both experienced an FNR error during that query step. If the probability of an FNR error is r, then the probability that sensors S1 and S3 both experienced an FNR error during a single query step is r. 2 For the purposes of this example, a high error rate of r = 0.2 is assumed, and therefore the probability that both sensors S1 and S3 incorrectly detected a marker during query step 2 is 0.04.
[0292] If sensor S2 is erroneous, it is because sensor S2 failed to detect the label due to an FLD error or an FNI error. Recall that an FLD error occurs when the correct complementary nucleotide is incorporated, but the label is missing or the sensor fails to detect its label, and an FNI error occurs when the correct complementary nucleotide is not incorporated at all during the sequencing cycle. FLD and FNI errors are mutually exclusive (that is, a sensor can only encounter one of them at a time, never both). Therefore, assuming the probability of each type of error is r, the probability of sensor S2 encountering a FLD error or an FNI error is 2r. For the example herein, a high error rate of r = 0.2 has been assumed, so the probability of sensor S2 being erroneous during query step 2 is 0.4. Comparing the probability of sensor S2 being erroneous during query step 2 with the probability of both sensors S1 and S3 being erroneous, since 0.4 >> 0.04, the probability of sensor S2 being erroneous is greater. In some embodiments, the error correction algorithm assumes the more likely event, meaning that sensor S2 is erroneous and the probability of both sensors S1 and S3 being erroneous is discarded from further consideration.
[0293] As explained above, sensor S2 may be erroneous due to either an FLD error or an FNI error. Following an FLD error, the DNA strand sensed by sensor S2 will remain "synchronized" or "aligned" with the DNA strands sensed by sensors S1 and S3. In other words, if query step m sequences the 40th base of a DNA strand sensed by each of sensors S1, S2, and S3, query step m+1 will sequence the 41st base of each strand, even if one of the sensors (e.g., sensor S2) experienced an FLD error during query step m. On the other hand, the consequence of an FNI error is that the DNA strand sensed by the sensor experiencing the FNI error becomes "out of sync" or "misaligned" with the DNA strands sensed by the sensor that did not experience the FNI error. In the current example, if the error at query step 2 is due to FNI (e.g., it will be "behind" the DNA chain sensed by sensors S1 and S3 in four query steps, which will be the next incorporation of the complementary nucleotide), then the DNA chain sensed by sensor S2 will become out of sync with the DNA chain sensed by sensors S1 and S3.
[0294] In some embodiments, the action taken by the error correction algorithm depends in part on an examination of the candidate error-corrected data, each assuming that two types of errors have occurred. In other words, the record of the test results can be modified to correct the error, assuming that the error is due to an FLD error, to produce a first candidate corrected data record, and the record of the test results can be modified to correct the error, assuming that the error is due to an FNI error, to produce a second candidate corrected data record. The two candidate corrected data records can then be examined and / or analyzed and / or compared to determine which is more likely to be correct. To correct an FLD error, the "not detected flag" indication is flipped to a "detected flag" indication. To correct an FNI error, the data entry is shifted four positions (e.g., to the left as presented as a data record in the example herein).
[0295] To illustrate the specific example of query step 2 in the example data record, the first candidate corrected data record option A assumes that the (assumed) error affecting the output of sensor S2 is an FLD error. The assumed error is corrected by flipping the query step 2 bit in the record for sensor S2 from 0 to 1, as shown in the bold, underlined value "1" in the table option A below:
[0296]
[0297] The second candidate corrected data record, Option B, assumes that the error affecting the output of sensor S2 is an FNI error. The assumed error is corrected by deleting the data recorded during query steps 2, 3, 4, and 5 from the sensor S2 data entry so that the data record corresponding to sensor S2 is "resynchronized" or "re-compared" with the data records of sensors S1 and S3, resulting in the following table (the values originally at positions 21 to 24 are shifted to positions 17 to 20). The Option B table entries modified by the error correction algorithm are shown in bold, underlined font:
[0298]
[0299]
[0300] Options A and B can then be compared and / or analyzed to determine which is more likely to be correct, and one of the options can be discarded. For example, a processor (e.g., at least one processor 130 or another processor) can determine a metric value for each candidate corrected data record and, based at least in part on a comparison of the metrics, determine which of Options A and B is more likely to be correct. An example of a metric is the number of query steps following the query step after the current, now-corrected query step and the position of the further query step J in the data record where the marker detection results from all three (or more generally, K) sensors are consistent. For example, using this metric and setting the value J to 8, Option A has a metric value of 3 and Option B has a metric value of 6. In some embodiments, based solely on this result, it is assumed that because Option B's metric value is significantly greater than Option A's, Option B is more likely to be correct, and Option A is discarded. In some embodiments, one of the two options is discarded only if its metric value exceeds the other option's metric value by a certain threshold (e.g., a percentage, an amount (e.g., at least twice, at least 1.5 times greater, etc.)). In some embodiments, Option A is retained and not discarded until a later time.
[0301] In some embodiments, contributions to the metric are weighted based on the distance from the data considered in the current, now-corrected query step. For example, because the likelihood of additional errors introduced into the data record increases as more bases are sequenced (e.g., the likelihood of a certain type of error occurring in one of the K sensors between query step 3 and query step 40 is greater than the likelihood of a certain type of error occurring in one of the K sensors between query step 3 and query step 6), the metric may assume that closer data entries are more likely to be correct than more distant data entries, and therefore give more weight to data entries closer to the now-corrected data than to those more distant data entries. The weighting may be, for example, linear or nonlinear. As just one example, for a metric with data contributions up to 12 query steps away, query step contributions within four query steps of the now-corrected data may be assigned a weight of 1, query step contributions from five to eight query steps of the now-corrected data may be assigned a weight of 0.5, and query step contributions from nine to twelve query steps of the now-corrected data may be assigned a weight of 0.2. It should be appreciated that many possible metrics may be used, with or without weighting, and that those provided above are merely exemplary and are not intended to be limiting.
[0302] It should also be appreciated that while the metric described above uses the query step number from the query step after the now-corrected current query step and the position of the further query step J in the data record where the marker detection results of all three (or more generally, K) sensors agree, it could equivalently use the query step number from the query step after the now-corrected current query step and the position of the further query step J in the data record where the marker detection results of all three (or more generally, K) sensors disagree. In this case, a larger metric value would indicate more mismatches between the sensor data entries, and thus the candidate corrected data record would be more likely to be correct for a lower metric value. As will be apparent to one of ordinary skill, any weighting to be applied can be adjusted.
[0303] It should also be understood that after correcting a hypothetical error in a data record, one of the possible options need not be discarded. For example, after correcting a hypothetical error at query step 2 in the record of sensor S2, both options A and B can be retained, and further error detection and correction performed on both in parallel. Similarly, each time a hypothetical error is corrected, multiple options for candidate sequences can be determined and / or evaluated / compared. A running metric for each possible option / candidate sequence can be maintained at each step of the error correction process, and the most likely candidate sequence can be determined at some point (e.g., after all candidate options have been determined and evaluated (e.g., relative to each other), or after some additional number of query steps, etc.).
[0304] Furthermore, although in the above example, the probability of both sensors S1 and S3 falsely detecting a flag is immediately discarded because the probability of such an event (given the assumptions herein) is significantly lower than the probability that sensor S2 is false, the same procedure can alternatively be followed for sensor S2. In other words, option C at query step 2 can be determined, assuming that both sensors S1 and S3 experience FNR errors, and sensor S2 is correct. In this case, the metric can be adjusted to account for the likelihood of various possible outcomes (e.g., by "penalizing" the metric for option C based on the probability that both sensors S1 and S3 experience FNR errors (e.g., multiplying the metric by the ratio of the probability that both sensors S1 and S3 are false to the probability that sensor S2 is false)).
[0305] It will be appreciated that the error correction methods described herein can be used in a variety of ways to improve the accuracy of nucleic acid sequencing using the SMAS device 100. Assuming sufficient computing power, an embodiment (e.g., using at least one processor 130 or another processor or processors) can determine and evaluate an exhaustive set of candidate sequences to which error correction is applied, and then select the candidate sequence that is most likely to be correct from among them. To reduce computational complexity, an embodiment can also make decisions during the error correction process to eliminate candidate error-corrected sequences (or potential sources of error) that are considered unlikely to be correct (e.g., option C in the example above) and retain only those candidate error-corrected sequences that are more likely to be correct. It will be appreciated that the flexibility of the disclosed principles makes them suitable for error mitigation in systems with a variety of computing power.
[0306] Returning to the above example, assume that option B is the only option that remains after applying error correction to the data from query step 2, and the corrected data appears as follows:
[0307]
[0308] The next query step for the three sensors S1, S2, and S3 that are inconsistent is at query step 5. Once again, sensor S2 is inconsistent with sensors S1 and S3 in the same manner as in query step 2. In some embodiments, the error correction algorithm determines that (a) the probability that sensor S2 is erroneous is greater than the probability that both sensors S1 and S3 are erroneous, and (b) sensor S2 encountered an FNI error or an FLD error at query step 5. Once again, two options can be established, one assuming the error is an FLD error (corrected by flipping a bit) and the other assuming the error is FNI (corrected by shifting the data four positions). The corrected data record is shown below:
[0309] Option A (assuming FLD errors are corrected):
[0310]
[0311] Option B (assuming FNI errors are corrected):
[0312]
[0313] Once again, metrics for options A and B can be calculated, and one of the options can be discarded, or both can be retained. For example, assume that option A is retained, resulting in the following error-corrected data:
[0314]
[0315] The next query step for inconsistent sensor data is query step 10. Here, sensor S1 detects a marker, but neither sensor S2 nor sensor S3 detects the marker. Because FLR errors have been removed from the data record, the only way sensor S1 could have falsely detected a marker during query step 10 is if it encountered an FNR error during that query step. The probability of an FNR error is r. If sensors S2 and S3 are both wrong, it is because (a) both encountered an FNI error, (b) both encountered an FLD error, or (c) one encountered an FNI error and the other encountered an FLD error. The probability of any of the mutually exclusive events (a), (b), or (c) is 4r 2 Therefore, in some embodiments, it is assumed that the more likely event occurs, that is, the sensor S1 suffers from an FNR error (since r>>4r for the assumed r value). 2 ). As explained above, FNR errors can be corrected by flipping the data entries from "marker detected" values to "marker not detected" values, which results in the following table:
[0316]
[0317] The error correction procedure may continue as described throughout the remainder of the data record. Figure 32 The portion labeled "C" of FIG. 5 shows the results for this example. As shown, after applying the probabilistic error correction as described above, 16 of the 20 (80%) bases were correctly identified.
[0318] Figure 33 FIG. 4 is a flow chart illustrating an error correction process 450 according to some embodiments. The error correction process 450 may be, for example, the process described in Figure 11 The error correction program 212 in the Figure 5A or Figure 50At least one processor 130 in 452 is performed. At 452, the error correction program 450 starts. At 454, multiple records are identified in the sequencing data generated by the nucleic acid sequencing program using the SMAS device 100. Each of the multiple records identified contains multiple entries, each of which captures the detection results of an instance of a specific chain of nucleic acid. Therefore, if the number of records identified is K, each of the K records contains one entry / detection result / sequencing program query step. Each detection result indicates that during the query step, (a) the label was detected by the corresponding sensor 105, or (b) the label was not detected by the corresponding sensor 105. The multiple records can be identified in a variety of ways. For example, as further described below, different unique barcodes can be spliced to the primer end of the nucleic acid chain so that a known sequence is read during the cycle of the sequencing program. Therefore, the multiple records can be identified by searching the sequencing data of the barcode associated with the specific chain of nucleic acid. As another example, a common sequence of entries can be identified in the sequencing data (eg, within entries recording the detection results of the first approximately 35 query steps of a sequencing process).
[0319] At 456, a plurality of candidate sequences for a particular strand of nucleic acid are determined based on the plurality of records. Each of the plurality of candidate sequences estimates at least a portion (e.g., down to a base) of a nucleic acid sequence for a particular strand of nucleic acid. In some embodiments, determining the plurality of candidate sequences comprises identifying a particular query step within the plurality of records at which the first sensor detected a respective marker and the second sensor did not detect any marker, and establishing two candidate sequences, one of the two candidate sequences assuming that the first sensor correctly detected the respective marker and the second of the two candidate sequences assuming that the first sensor incorrectly detected the respective marker. In some embodiments, determining the plurality of candidate sequences comprises identifying a particular query step within the plurality of records at which the first sensor detected a respective marker and the second sensor did not detect any marker, and establishing two candidate sequences, one of the two candidate sequences assuming that the second sensor incorrectly did not detect any marker and the second of the two candidate sequences assuming that the second sensor correctly did not detect any marker. In some embodiments, determining the plurality of candidate sequences comprises identifying a set of consecutive entries (e.g., four entries) in at least one of the plurality of records indicating that a marker was not detected, and deleting the set of consecutive entries indicating that a marker was not detected from at least one of the plurality of records. In some embodiments, each of the plurality of entries is a first binary value (indicating that a marker was detected) or a second binary value (indicating that a marker was not detected), and determining the plurality of candidate sequences comprises identifying a string of (e.g., four) second binary values in at least one of the plurality of records, and deleting the string of second binary values from at least one of the plurality of records.
[0320] At 458, a particular candidate sequence from the plurality of candidate nucleic acid sequences is identified as the sequence most likely to be correct from among the plurality of candidate sequences. In some embodiments, identifying the particular candidate sequence that is most likely to be correct from among the plurality of candidate sequences includes determining or estimating which of the plurality of candidate sequences has the highest probability of being correct. In some embodiments, identifying the particular candidate sequence that is most likely to be correct from among the plurality of candidate sequences includes determining a respective metric for each of the candidate sequences and selecting the particular candidate sequence as the most likely to be correct based, at least in part, on the respective metric and a criterion (e.g., minimum likelihood of occurrence, threshold likelihood of occurrence). In some embodiments, identifying the particular candidate sequence that is most likely to be correct from among the plurality of candidate sequences includes identifying a majority of results for a particular query step represented by a plurality of records (e.g., more than half of the sensors 105 detected a tag or more than half of the sensors 105 did not detect a tag). In some embodiments, identifying the particular candidate sequence that is most likely to be correct from among the plurality of candidate sequences includes determining a respective likelihood of occurrence for each of the plurality of candidate sequences and selecting the particular candidate sequence based on its respective likelihood of occurrence satisfying a constraint (e.g., minimum probability). In some embodiments, the particular candidate sequence from the candidate sequences with the highest likelihood of occurrence is identified as the most likely to be correct. In some embodiments, one or more of the candidate sequences are eliminated based on known constraints, such as the knowledge that a particular sequence of bases is improbable. For example, it may be known from the origin or source of the nucleic acid (e.g., human) that a particular sequence of bases is improbable, and thus candidate sequences with such improbable sequences may be eliminated from further consideration.
[0321] At 460, the error correction procedure 450 ends.
[0322] It should be understood that only when the most likely scenario is identified (e.g. Figure 33 Probabilistic error correction is successful only when the chemical failure rate is high, as in the example described herein, there may be multiple scenarios where the same possibility exists (or where their probability of occurrence is close to each other), in which case more sophisticated bioinformatics tools may be employed. For example, candidate sequences may be eliminated based on knowledge of the source of the nucleic acid being sequenced (e.g., knowledge that a particular sequence of bases is unlikely based on the source / origin of a given nucleic acid). However, if implemented correctly as described herein, the error correction process results in a correct alignment of the sensor 105 output. In the example shown at Figure 32In the example shown in FIG, after removing FNR and FLR, all three sensors S1, S2, and S3 report a marker at the correct detection query step where the marker should be detected, but the sensors disagree at many query positions (5, 10, 13, 20, 22, 27, 32, 40, 41, 48, and 50), where the sensors detect markers that were incorrectly incorporated by FNR or fail to detect markers due to FLR. When more than half of the aligned sequences are identical among sensors 105, the recognition of bases results in a thymine insertion at sequence position 8 (query step 22) and a guanine deletion at position 13 (query step 32). The true sequence and the recognized sequence, which clearly show the correct alignment of the base insertion and deletion positions, can be presented as: Error: Insertion deletion True sequence TAG CAA G * G TCC G CT ACT GGC
[0323] Recognized sequence: TAG CAA G T G TCC * CT ACT GGC
[0324] As will be appreciated in light of the disclosure herein, coincidental FNR and FLD result in insertion and deletion errors that cannot be algorithmically corrected and will remain undetected if the true sequence is unknown. In other words, when more than half of the single-molecule sensors 105 in the aligned sequence give incorrect answers, bases are incorrectly identified. The probability of such events depends on the ratio (r-value) of chemical failures. As described above, the examples presented herein use high error rates to illustrate the application of error correction techniques. The error rates in actual implementations should be significantly reduced, thereby reducing the possibility that the error correction program cannot correct errors. The disclosed error correction techniques can be used to correctly compare multiple sensor 105 outputs at the query step. This can be achieved using a deep understanding of the physical origins of possible error types (e.g., knowledge that certain sequences are impossible for source nucleic acids), their average occurrence rates, and their signatures in the sensor sequence outputs. If the chemical error rate is very high and the signatures of the errors are obscured, the error correction algorithm can be computationally intensive and difficult to implement. The following discussion describes how the probability of an incorrect base call depends on the read length, cluster size N (for CLUS devices), the number K of sensors sensing instances of the same nucleic acid strand (for SMAS devices 100), and the failed chemical error rate.
[0325] General quantitative results of cluster sequencers
[0326] This paper develops a simple quantitative model for estimating the probability of incorrect base identification in a cluster sequencer employing the improved additive ordering scheme introduced above. It is assumed that various types of errors (FNI, FLR, FNR, and FLD) occur randomly at a rate r throughout the clusters, where 0 < r < 1. Initially, the cluster strands are in-phase with each other (e.g., synchronized, aligned, out-of-sync), and the detected signal is proportional to the cluster size (N). A signal is detected when a complementary-labeled nucleotide is introduced and successfully incorporated. When a non-complementary nucleotide is introduced during a query cycle with a query step, no signal should be detected. Errors occur at a rate r, which results in an increasing number of strands out-of-phase (out-of-sync) with the collective average. This reduces the intensity (or amplitude) of the collective signal when a complementary nucleotide is incorporated and increases the intensity or amplitude of the background signal when a non-complementary nucleotide is introduced. The average signal intensity at a query step where a labeled nucleotide should be detected because a matching nucleotide has been introduced and successfully incorporated (ON-State) is given by:
[0327]
[0328] where C is the detection query step (or number). Similarly, the intensity at a query step where a labeled nucleotide should not be detected because a non-complementary nucleotide has been introduced (OFF-State) is given by:
[0329]
[0330] This background signal is generated by out-of-phase nucleic acid strands that incorporate nucleotides non-complementary to the in-phase position of the collective average. The functions of Equations 1(a) and (b) are plotted in Figure 34A for N = 11 and r = 0.1. Figure 34B Illustrating how the functions fit the intensity measurements of the previously described cluster model instance. As illustrated, bases are correctly identified up to but frequent errors occur at larger C values.
[0331] As illustrated by Figure 34A and 34B during early sequencing queries (C small), the <1> and <0> states are completely separated, but they quickly approach the average value of N / 2 following the functional form represented by Equations 1(a) and (b). Furthermore, since error occurrences are random and independent events, the signal measurements of the two states are discretely distributed around their collective averages <1> and <0>. Specifically, the probability that the ON-State intensity measurement of the cluster size N is k when the collective average is <1> is given by the Poisson distribution:
[0332]
[0333] Similarly, when the collective average is <0> The probability that the OFF-State intensity record value of the same cluster is k is:
[0334]
[0335] Probability function P <1> (k) and P <0> (k), N = 11, r = 0.1 and C = 0, 5, 10, 15 and 20, plotted on Figure 35 The figure shows two Poisson distributions, and the tails overlap more and more as C increases. Under the two discrete distributions P <0或1> The sum of all possible values of (k) equals 1:
[0336]
[0337] Base calling errors occur when an ON-State is mistaken for an OFF-State or vice versa. Figure 36 Explain the ON-State P at different sequence query steps C = 0, 5, 10, 15 and 20 <1> (k) and OFF-State P <0> (k) (N = 11 and r = 0.1) is a discrete probability function. The source of incorrect base calls is P <0> The tail of (k) is shown as a patterned dot when it extends above the middle value of N / 2 or as a dot at P <1> (k) The extension below the median value of N / 2 is shown as a dotted circle. The tail of the ON-State distribution extends significantly below (exist Figure 36 Incorrect <1> ) or the tail of the OFF-State distribution (incorrect <0> ) extends above When , the probability of making an incorrect base call becomes very large.
[0338] Figure 37A Shown are plots of average ON-State and OFF-State intensities as a function of C (r = 0.1) and cluster size for N = 11 (top) and N = 101 (bottom). Figure 37B The OFF-State probability distribution function P is shown for cluster sizes C = 1, 10, 20, 30, and 40 (r = 0.1) and N = 11 (top) and N = 101 (bottom). <0> (k). Increasing cluster size by reducing P <0> (k) The relative width of the distribution (this increases the distance from P <1> (k)) and delay the occurrence of base call errors.
[0339] In general, the probability of incorrect base calling at the number of sequencing queries C (for cluster size N and chemical failure rate r) (denoted as P C,N,r ) is the sum of the probabilities of incorrectly identifying the OFF-State, that is, for values of k exceeding k=(N+1) / 2, it is P <0> (k) values. These are Figure 36 and 37B Increasing the cluster size N increases the initial interval between two discrete distribution peaks and delays the occurrence of base call errors. To simplify further discussion, only the case where the cluster size N is odd is considered to avoid the ambiguity introduced when the detected signal is N / 2 (which is neither ON-State nor OFF-State). For odd values of N, P C,N,r is given by:
[0340]
[0341] Or, P C,N,r is the sum of the probabilities of incorrectly identifying the ON-State, that is, for values of k below k = (N-1) / 2, it is P <1> The sum of (k) values ( Figure 36 ), which is given by:
[0342]
[0343] Figure 38A and 38B Plot Equations 4(a) and 4(b) as a function of C for various combinations of N and r. Figure 38A Plot calculation P C,N,r (C) function, r = 0.1 and N = 11, 51, 101 and 151, and Figure 38B Plot calculation P C,N,r (C) function, N = 101 and r = 0.1, 0.05 and 0.01. The figure shows the C th The probability of incorrect base calls increases significantly. Figure 38A and 38B Indicates that as C approaches infinity, P C,N,r Close to 0.5. Figure 38A and 38B The graph in Figure 1 shows the behavior of a sequencer that analyzes a collection of molecules (e.g., a CLUS device). When C is small, the probability of incorrect base recognition (P C,N,r ) is still low, but it is below a certain threshold (C th ) increases significantly, and the threshold is determined by the values of N and r parameters. As C tends to infinity, P C,N,rClose to 0.5, at which point the intensity of the ON-State is equal to the intensity of the OFF-State, and there is a 1 in 20 chance of an incorrect base call. C,N,r It depends largely on these three parameters C, N, and r. The dependence on C is particularly important because C th Limits the number of consecutive bases that can be identified before the probability of an error becomes too great.
[0344] Figure 39 Describe the Nr parameter space, where the probability of an incorrect base call at position 150 (P C=375,N,r ) is lower than 1 in 100 (Q20), 1 in 1,000 (Q30), 1 in 10,000 (Q40) and 1 in 100,000 (Q50). Increasing the cluster size N, or reducing the chemical failure rate r, and setting the threshold C th Push to a higher C value, but if Figure 39 Medium quantification shows that cluster sizes are quite large and the allowed chemical error rate must be small to make DNA sequencers suitable for diagnostic applications.
[0345] Currently, the benchmark in the sequencing industry is the ability to read 150 consecutive bases with a 1 in 1,000 probability of an incorrect base call at position 150. This is generally referred to as Q30, but significantly larger sequencing quality factors of Q40 and even Q50 and longer read lengths are needed to detect rare mutations in high-accuracy diagnostics. C,N,r The general representation of fully explores the CNr parameter space and can be used to estimate the error tolerance and cluster size requirements of any ordering metric. Figure 39 The region of the Nr parameter space is shown, where at position 150 The probability of incorrect base calling is less than 1 in 100 (Q20), 1 in 1,000 (Q30), 1 in 10,000 (Q40), and 1 in 100,000 (Q50). For example, if the average cluster size N in the sequencing array is 100 molecules and the required sequencing accuracy is Q30, with a 150 bp long read segment The permissible chemical failure rate is r≤0.002641, meaning that only 26 or fewer failures are permitted per 10,000 individual single-molecule reactions on the sequencer array at any sequencing query step. If the required accuracy is Q50, only 19 or fewer errors per 10,000 reactions are permitted. If the average cluster size N is reduced to 10 molecules, this number drops to approximately 6 (Q30) and approximately 1 (Q50) per 10,000 reactions.
[0346] Figure 40A Shows calculated P for various Nr combinations along the Q30 contour C,N,r(C), marked by crosses (“+” signs) in the inset, all intersections are at P C,N,r (C=375)=0.001. The figure shows that increasing the cluster size N not only improves the chemical failure tolerance, but also increases the chemical failure tolerance by increasing the threshold C th Pushing to higher C values delays the occurrence of base call errors, which results in a reduction in cumulative errors. If the probability of making an incorrect base call at query cycle C is P C,N,r , then the probability of correct recognition is (1-P C,N,r ). The probability of correct identification of C continuously is:
[0347]
[0348] The probability of not making C correct base calls in row form (which is the same as or less than the probability of making at least one error at any query cycle C) (or the cumulative error probability is given by:
[0349]
[0350] Among them, P j,N,r It is given by Equation 4(a) or (b). Figure 40B Calculate the cumulative error probability by drawing along the same contour line It also shows that larger clusters produce lower cumulative errors.
[0351] Finally, the indicator is calculated and a marker is plotted where the cumulative probability of an incorrect base call at position 150 (in some embodiments, the target read length) is less than or equal to 1 in 100. 1 in 1,000 1 in 10,000 and 1 in 100,000 The Nr region of the parameter space. Figure 41 Illustrating the cumulative probability of an incorrect base call at position 150 Less than or equal to 1 / 100 1 in 1,000 1 in 10,000 and 1 in 100,000 Nr parameter space. Figure 41 The figure in quantifies that a CLUS sequencer can include large DNA cluster sizes N to benefit from collective behavior, and that it may require extremely reliable chemistry (only a few dozen failures per 10,000 reactions) for high-precision diagnostic applications. More specifically, if the average cluster in the sequencing array remains, for example, 100 molecules, and a particular sequencing application tolerates 1 in 1,000 With a cumulative base call error probability of , only about 22 or fewer failures in 10,000 individual single-molecule reactions on the sequencer array are allowed at any sequencing query step. Figure 41 The graph in illustrates that increasing sequencing throughput by reducing the cluster size N and packing more clusters into the sensing area can only be achieved with parallel improvements in sequencing chemistry. The required rate of improvement accelerates as the cluster size N becomes smaller, and CLUS devices can no longer benefit from large collective behavior.
[0352] General quantitative results of single-molecule array sequencers
[0353] To compare the CLUS and SMAS platforms, a simple quantitative model was developed to estimate the probability of incorrect base calls in the SMAS device 100. Unlike the collective case applicable to the CLUS device (described above), in which error correction is rarely or impossible to implement, the ability of the SMAS device 100 to sequence and record the detection results corresponding to individual nucleic acid molecules individually allows the development and implementation of powerful techniques for identifying and eliminating at least some errors in the resulting data records. One or more error correction techniques disclosed herein can be applied to data generated from a sequencing program (e.g., SBS) before base calling to identify and correct errors in the detection results to improve the accuracy of the identified sequence. Specifically, the alignment of detection results from multiple sensors 105 at some or all query steps of the sequencing program can be improved. Even if the error correction algorithm successfully correctly aligns multiple sensor detection results, incorrect base calls may still be made. As explained above, coincidental FNR errors and FLD errors can lead to insertion and deletion errors that may be uncorrectable. Depending on the number of errors in the data record (which is determined in part by the chemical failure rate), the error correction process can be complex and computationally intensive, but it will be appreciated that modern processors have sufficient computational power to perform even the most computationally intensive disclosed techniques.
[0354] Below, we consider the general case of K single-molecule sensors 105 of a SMAS device 100, each capable of monitoring a single instance of clonal DNA. As in the analysis of the CLUS device above, we assume that four types of errors (FNI, FLR, FNR, and FLD) occur randomly during the sequencing process and are distributed throughout the query step.
[0355] As described above, in some embodiments, a probabilistic error correction algorithm is implemented (e.g., by at least one processor 130, which may be included in or external to the SMAS device 100). In some embodiments, the probabilistic error correction algorithm improves the alignment of at least some sensor 105 detection results in the data record. In some embodiments, some or all of the error correction algorithm is implemented after some or all query steps have been completed and some or all data has been captured. As previously described, the error correction process substantially eliminates FNI and FLR, as well as some FLD. Algorithmic re-alignment of sensor 105 detection results also makes the probability of making an incorrect base call independent of the number of query steps C. Furthermore, because the error correction algorithm re-aligns at least some sensor 105 detection results in the data record, thereby correcting at least some errors, the effective error rate r is less than in the CLUS case. After applying the exemplary error correction algorithm, in some embodiments, a base is incorrectly called only when more than half of the K sensors 105 in the algorithmically aligned sequence give incorrect results.
[0356] The probability of making an incorrect base call (P K,r ) is a function only of (a) the number K of sensors 105 that sequence instances of the same nucleic acid molecule (which may be less than all sensors 105 in the sensor array 110) and (b) the chemical failure rate r. Similar to the approach taken for the analysis of the CLUS device above, K values are constrained to odd values to avoid situations where exactly half of the sensors 105 disagree with the other half. The probability of making an incorrect base call is given by:
[0357]
[0358] in For example, if K=3,
[0359]
[0360] In the example of K=3, multiplication The term "f" describes a situation in which two of the three sensors 105 simultaneously encounter an error at a particular query step (e.g., they incorrectly detect a label (FLR, FNR) or incorrectly do not detect a label (FNI, FLD)), thereby forcing an incorrect base call. Denoting the three sensors 105 as S1, S2, and S3, this scenario occurs when: (1) S1 and S2 encounter an error simultaneously, (2) S1 and S3 encounter an error simultaneously, or (3) S2 and S3 encounter an error simultaneously. The term explains the impossible case where all three sensors S1, S2, and S3 encounter an error simultaneously, which also results in an incorrect base call. Since the largest term in the polynomial expansion is r K-1Since 0 < r < 1, the probability of incorrect base identification is significantly reduced by increasing the number of single - molecule sensors 105 (i.e., increasing the K value).
[0361] For example, if r = 0.1, then P K=3,r=0.1 = 0.029, which means the probability of incorrect base identification is about three in one hundred. In other words, on average about 4.35 out of 150 base identifications will be incorrect, which is too high for some diagnostic applications. To sequence using three nanoscale sensors 105 with Q30 (P K,r = 0.001), the chemical failure rate would need to be reduced to r = 0.01837, meaning that only about 19 out of 1,000 queries would be allowed to be incorrect. However, if the number of sensors 105 (K value) is increased to 11, more than 12 failures out of one hundred reactions can be tolerated.
[0362] As done above for the CLUS device, the K - r parameter space is explored below for the SMAS device 100 to identify regions where the probability of incorrect base identification at any query position is less than one in one hundred (Q20), one in one thousand (Q30), one in ten thousand (Q40), and one in one hundred thousand (Q50). Figure 42 Illustrate the computational results of the K - r parameter space where the probability of incorrect base identification (P K,r ) at each query step is less than one in one hundred (Q20), one in one thousand (Q30), one in ten thousand (Q40), and one in one hundred thousand (Q50). As Figure 42 shown, if the number K of single - molecule sensors 105 sensing instances of the same nucleic acid molecule is 11 and the required sequencing accuracy is Q30, the allowed chemical failure rate is which means that up to about 13 failures out of 100 individual single - molecule reactions among those 11 sensors 105 are allowed. If the required accuracy is Q50, about 6 or fewer errors out of every 100 reactions among the 11 sensors 105 are allowed.
[0363] As indicated by the comparison with Figure 39 , the allowed error rate of the SMAS device 100 is significantly greater than the allowed rate for the CLUS device. However, the results alone do not fairly compare the two platforms because the probability of incorrect base identification (P C,N,r ) in the CLUS device is extremely low during early query steps and suddenly increases at the threshold query step C th . This phenomenon is discussed in conjunction with Figure 39 . On the other hand, for the SMAS device 100, the probability of incorrect base identification (P K,r) remains constant throughout the query steps and thus leads to a large cumulative error.
[0364] A more fair way to compare the performance of the CLUS device and the SMAS device 100 is to compare the cumulative error probability of the two device types. Equation 5(b) above represents the cumulative error probability of the CLUS device. The cumulative error probability of the SMAS device 100 can also be derived. The probability of making an incorrect base call at each query step C is P K,r (Equation 6), and therefore the probability of correct recognition is (1-P K,r The probability of correctly identifying C in row form is (1-P K,r ) C , and the cumulative error probability for
[0365]
[0366] Figure 43A and 43B The cumulative probabilities of incorrect base calls at position 150 are shown for the CLUS device and the SMAS device 100. Equation 5(b) can be used, for example, to calculate the probability of an incorrect base call at any base position less than or equal to 150 by the CLUS device. Figure 43A The Lr parameter space of the CLUS device is shown and the cumulative probability of an incorrect base call at position 150 for the CLUS device is marked as less than or equal to 1 in 100. 1 in 1,000 1 in 10,000 and 1 in 100,000 area. Figure 43B Equation (8) is evaluated and the Kr parameter space is displayed, which marks where the cumulative probability of an incorrect base call at position 150 for the SMAS device 100 is less than or equal to 1 in 100 1 in 1,000 1 in 10,000 and 1 in 100,000 area.
[0367] Figure 43A and 43B The comparison shows that the SMAS device 100 is a potentially superior sequencing platform to the CLUS device. The SMAS device 100 may have a smaller footprint (e.g., Figure 7A 、 7B, 9A, 9B, and 10) and can be more error-tolerant than CLUS devices. Using a SMAS device 100 allows for higher throughput, lower error rates, and longer read lengths compared to CLUS devices, which are larger and rely on macromolecular ensembles. Development of a commercially viable SMAS device 100 and / or system can utilize some or all of the following: (a) high-precision nanoscale fabrication of densely packed sensors 105 capable of recognizing individual markers, (b) optimization of chemical steps to reduce error rates to acceptable levels, and / or (c) use of effective bioinformatics tools to adjust the alignment of sequencing data from at least some nanoscale sensors 105 in a data record by probabilistically eliminating errors.
[0368] Exemplary SMAS Sequencing Procedure
[0369] As explained above, improvements in sequencing throughput for CLUS devices can be achieved by reducing the cluster size N (thereby packing more clusters into the device) if the sequencing chemistry failure rate is also reduced, which can be challenging. In contrast, the following presents a feasible implementation of an error-tolerant, ultra-high-throughput SMAS device 100 using a large array of single-molecule binding sites 116, according to some embodiments. For example, it is assumed that the SMAS device 100 sequences DNA, but it should be understood that, in general, any type of nucleic acid can be sequenced.
[0370] Figure 44 and 45 An exemplary sample preparation and loading process 500 is illustrated according to some embodiments. Figure 44 is a flow chart illustrating process 500, and Figure 45 The results of each step of process 500 are described. In some embodiments, the sample preparation and loading process 500 begins at 502. At 504, DNA extraction and purification are performed, which results in several extracted DNA fragments 505, such as Figure 45 At 506, an adapter complementary to the primer is spliced to one end (e.g., 3') of the extracted DNA to produce the Figure 45 At 508, PCR (or some other replication technique) is performed to produce multiple (ideally, identical) instances of the extracted strand, such as Figure 45 At 510, a molecular linker capable of establishing a strong bond (e.g., by click chemistry) on the chemically functionalized surface of the fluid chamber 115 (binding site 116) of the SMAS device 100 is attached to the other end (e.g., 5′) of the ssDNA fragment, thereby generating the ssDNA fragment shown in FIG. Figure 45 At 512, the functionalized chains are loaded into the fluid chamber 115 and randomly dispersed among the binding sites 116 and bound to the binding sites 116. Figure 45 As shown in the rightmost portion of , each of the binding sites 116 supports no more than a single DNA strand. (Although each binding site 116 can support no more than one strand, it should be understood that not every binding site 116 must support a DNA strand. Whether intentional or accidental, fewer than all of the binding sites 116 of the SMAS device 100 may be used.) Assuming that the extracted DNA fragments 503 are different from each other, due to the sample preparation and loading process 500, multiple instances of each of the extracted DNA fragments 505 will be present within the fluid chamber 115, but their locations are unknown. At 514, the exemplary sample preparation and loading process 500 ends.
[0371] A benefit of the exemplary sample preparation and loading process 500 is that it simplifies DNA amplification, which can be performed in bulk off-device using, for example, conventional PCR, before adding the DNA strands to the SMAS device 100. In contrast, when using a CLUS device, amplification (e.g., bridge amplification) is performed only after the DNA fragments have been added to the CLUS device in order to create a continuous cluster array of amplified DNA.
[0372] After the sample preparation and loading process 500 has been performed, base calling can be performed using, for example, the additive method, the subtractive method, or the modified additive method described above. Figure 46A 、 46B 46C illustrate three exemplary interrogation cycles (for a total of 12 interrogation steps, each 100 is 100) performed using an example SMAS device 100 having a sensor array 110 of 20 sensors 105 (and 20 binding sites 116) arranged in four rows and five columns. ) during a simulated detection period using the modified additive method (sensor 105 detects a label). Multiple instances of four different DNA strands are randomly distributed throughout sensor array 110, but their specific locations within sensor array 110 and their sequences are initially unknown.
[0373] Figure 47 Instructions on how to rearrange instructions Figure 46A 、 46B and 46C detection data to identify bases and show the positions of different DNA strands. Figure 47 Tables are provided showing the output of each sensor 105 in the exemplary array at individual query steps and the resulting base calls that led to the identified sequences. Figure 47 The right-hand portion of is used to reorder the sensors 105 so as to group the detection results of the sensors 105 that sense instances of the same DNA strand. Figure 47 As shown in FIG, four sequences were identified: GCT (chain #1), TAG (chain #2), ACG (chain #3), and TTA (chain #4).
[0374] If an error (FNI, FLR, FNR, or FLD) occurs during the interrogation step, some detection results (label detected or not detected) will be incorrect, and the deterministic and / or probabilistic error detection and / or correction techniques described above can be implemented to detect and eliminate at least some errors, as long as the identity of those sensors 105 sensing instances of the same DNA strand is determined. Recall that instances of a particular DNA strand can be attached to binding sites 116 dispersed throughout the fluid chamber 115, and their locations are generally unknown at the beginning of the sequencing process. Once the process is initiated, each of the plurality of S sensors 105 detects the label at its respective binding site 116 during each interrogation step. For error correction, a subset of the S sensors 105 that sequenced instances of the same nucleic acid strand is identified.
[0375] Consider an extremely large sensor array 110 with 400 million different DNA strands (e.g., 4 billion binding sites 116 and 4 billion individual sensors 105), each DNA strand being approximately 150 bases long. This means that there are approximately 10 instances of each unique DNA strand randomly distributed throughout the fluid chamber 115 (and binding sites 116, and sensor array 110). For the sake of example, it is also assumed that the sequence is random. Assuming a reasonably low error rate r, after the first query cycle, almost all binding sites 116 (and sensors 105) that hold (sensed) DNA instances beginning with A will be identified, those that hold (sensed) T will be identified, and those that hold (sensed) C, and those that hold (sensed) G. Approximately 10 9 The sensor 105 will detect a marker indicating that the first base is A, about 10 9 The sensor 105 will detect a marker indicating that the first base is T, about 10 9 The sensor will detect the label indicating that the first base is C, and about 10 9 The sensor will detect a label indicating that the first base is a G. After the second interrogation cycle, nearly all binding sites 116 (and sensors 105) that hold (sens) DNA instances starting with all 16 possible combinations (AA, AT, AC, AG, TA, TT, TC, TG, CA, CT, CC, CG, GA, GT, GC, and GG) will be identified. 8 The sensor will detect the label indicating that the first and second bases are AA, about 2.5×10 8 The sensor will detect the label indicating that the first and second bases are AT, about 2.5×10 8 The sensors will detect labels indicating that the first and second bases are AC. In general, at some number D of label detections (or assuming that a modified additive method is used for sequencing), the After a query step), all 4 DNA chains that keep starting with a sequence of some D-base length will be identified. D =4 2C / 5 This means that the average size of a group of sensors 105 sensing instances of the same DNA strand in a SMAS device 100 having an array of 4 billion sensors 110 is 4×10 9 / (4 2C / 5 Since our example has an average of about 10 instances per unique chain, we will perform In the present invention, the present invention provides the method for the identification of the binding site 116 of the example of the specific chain.The method of the present invention is to identify the position of the binding site 116 of the example of the specific chain.Assuming that the improved additive method is used, about 14 bases will be identified during the process.Because the human genome is not random, and not all mathematically possible sequences are displayed, in fact, significantly fewer query steps may be needed for diagnostic applications.If during DNA extraction, targeting a specific group of genes, then the identity (position) of the binding site 116 of the example of the same DNA chain can be determined with even fewer steps, which further reduces the quantity of the possible sequence of bases and is conducive to the identification of binding site 116.
[0376] The confidence that the correct set of binding sites 116 has been identified increases with the number of query steps, but the probability of detection errors (e.g., incorrectly detecting a tag or incorrectly not detecting a tag) also increases. During the initial query cycle, multiple errors may occur while identifying binding sites 116 that remain instances of the same chain. The results obtained with the CLUS device indicate that this may not be a problem. For example, Figure 38A It is shown that the probability of incorrect base calls made by the CLUS device during the early query steps is extremely small and only occurs when the threshold C is reached. th Also recall that if no error correction is applied, the base calling accuracy of the SMAS device 100 is the same as that of the CLUS device, since the SMAS device 100 will simply report a collective result by summing the individual sensor 105 results.
[0377] Consider, for example, the 4 billion sensor array example above and consider a set of 11 sensors 105 (K=11) monitoring an instance of a particular DNA strand, randomly distributed throughout the binding sites 116. Now, consider them collectively (K=N=11), as if the binding sites 116 are forming a cluster and only measuring the combined properties (e.g., signal) of their individual sensors 105. Figure 48A and 48B Plot the calculated probability P of making an incorrect base call C,N,r , which is given by Equations 4(a) and (b) as a function of the number of query steps C and the chemical failure rate r. Figure 48AThe curve in the Cr space is marked with P C,N,r Approximate location of the threshold for the sudden increase. Figure 48B Is displayed in Figure 48A The top view of the contour plot in Figure 1 clearly indicates the chemical failure tolerance of a SMAS device 100 containing 4 billion sensors, which averages approximately 10 instances per DNA strand. The positions (identities) of the approximately 10 binding sites 116 (and sensors 105) holding (sensing) instances of each unique DNA strand can be reliably determined as long as the error probability remains low over approximately 35 interrogation steps. This limits the maximum allowable chemical failure rate to 0.013, meaning that 13 detection events in 1,000 will be tolerated. Figure 48A and 48B The calculations in indicate that if the chemical failure rate remains below approximately 13 incorrectly detected events per 1,000, the 4 billion sensor SMAS device 100 should be able to establish the positions of all instances of all billion different DNA strands within the fluid chamber 115 (and in the binding sites 116 and in the sensors 105). Once those positions are established, the error correction techniques described herein can be immediately implemented to eliminate errors that occur during the remaining approximately 340 interrogation steps (assuming the use of the modified additive method).
[0378] If the expected or known chemical error rate is too high, such that errors are likely to plague the first approximately 35 interrogation steps, alternative methods can be used to help identify binding sites for instances carrying the same DNA strand 116. For example, different unique barcodes can be spliced onto the ends of primers in a subset of extracted DNA so that known sequences are read during early sequencing cycles. Figure 49 The use of barcodes in sample preparation and DNA loading according to some embodiments is described. Figure 49 As shown in , unique barcodes are spliced into the extracted DNA to facilitate identification of sites that retain instances of identical DNA in the presence of sequencing errors. For example, Figure 49 Four unique DNA strands are shown, each of which is assigned a unique barcode (e.g., barcode 119A is assigned to strand 1, barcode 119B is assigned to strand 2, barcode 119C is assigned to strand 3, and barcode 119D is assigned to strand 4). If the barcodes are significantly different from each other, they should be easily identifiable even if the chemical failure rate is very high. As will be appreciated, for high-throughput diagnostic applications, the suitable number of unique barcodes can be very high.
[0379] The exemplary 4 billion sensor SMAS device 100 described herein is considered a fairly high-throughput sequencer by current standards. Such a SMAS device 100 provides approximately 150 gigabase (Gb) reads during a single run, which is comparable to the output of current state-of-the-art high-end sequencing systems introduced in 2020.
[0380] It should be understood that there are many ways to implement the devices, systems, and methods disclosed herein. For example, a system for nucleic acid sequencing can consist of a single device (e.g., SMAS device 100, which includes all hardware and software to perform the disclosed operations), or it can include SMAS device 100 and other components that together perform the disclosed operations. For example, the system can include SMAS device 100 and at least one processor external to SMAS device 100 (e.g., in an external computer), wherein SMAS device 100 performs a nucleic acid sequencing process and stores the detection results from the sequencing process, and wherein the at least one processor performs error detection and correction on the stored detection results and identifies bases.
[0381] Figure 50 An exemplary system 160 according to some embodiments is described. The system 160 includes (that is, includes (but is not limited to)) a fluid chamber 115, a plurality of S sensors 105, and at least one processor 130. Optionally, the system 160 includes a memory 170 for storing records containing detection results obtained during the sequencing process (e.g., one or more files having binary entries recording whether each of the plurality of S sensors 105 detected or did not detect at least one marker during each of a plurality of query cycles). Figure 50 As shown by the dashed line in , if the system 160 includes a memory 170 , the at least one processor 130 may be communicatively coupled to the memory 170 such that the at least one processor 130 may store data in the memory 170 and / or retrieve data from the memory 170 .
[0382] The fluid chamber 115 includes a plurality of S binding sites, each of which is configured to bind no more than one nucleic acid strand to be sequenced. Figure 50 Four binding sites 116 are shown, but it should be appreciated that the system 160 may include more or fewer binding sites 116. Each of the three sensors 105 is configured to detect a label present in the fluid chamber 115. Figure 50 Four sensors 105 are shown, but it should be understood that the system 160 may include more or fewer sensors 105. When the system 160 is in operation, each of the three sensors 105 detects a label attached to a nucleotide in a respective strand of a nucleic acid that binds to a respective binding site 116 of the three binding sites 116. As previously described, the sensors 105 may be magnetic sensors, optical sensors, or any other type of sensor that can detect a label for a labeled nucleotide. The fluid chamber 115, sensors 105, and binding sites 116 are described in detail above. Those descriptions apply to Figure 50 And will not be repeated here.
[0383] The at least one processor 130 is configured to execute one or more machine-executable instructions. When the instructions are executed, the at least one processor 130 performs a sequenced procedure including a plurality of query steps (e.g., Figure 11 、 12 , 14, 16, 44, and the like. Specifically, in operation, during a query step of a sequencing procedure, the at least one processor 130 obtains a respective characteristic of each of the three sensors 105 (represented by the dashed lines between the at least one processor 130 and the sensors 105A, 105B, 105C, and 105D). The respective characteristic indicates whether the sensor 105 detected or did not detect a marker (e.g., it indicates the presence or absence of at least one marker). The at least one processor 130 may interpret the obtained characteristic to determine whether the sensor 105 detected or did not detect the presence of the marker. Based at least in part on the obtained respective characteristic, the at least one processor 130 records whether the respective sensor detected the presence or absence of the at least one marker during the query step. The at least one processor 130 is also configured to perform an error correction procedure on at least one record containing the results of the sequencing procedure. The error correction procedure may operate on some or all records generated by the sequencing procedure, and it may operate on detection results from some or all query steps of the sequencing procedure. For example, as described above, to apply the error correction procedure, the at least one processor may identify a subset of K records and apply deterministic or probabilistic error correction thereto, wherein each of the K records in the subset corresponds to a detection result from the sensor 105 of an instance sensing the same nucleic acid strand. The sequencing procedure and error correction procedure are described in detail above. Those descriptions apply to Figure 50 The system and at least one processor 130 are described in detail and will not be repeated here.
[0384] The at least one processor 130 may be implemented by a general-purpose or special-purpose processor (or group of processing cores) and may therefore execute a series of programmed instructions to perform various operations related to obtaining sensor 105 characteristics, performing error correction procedures, and / or interacting with a user, system operator, or other system components.
[0385] The at least one processor 130 of the system 160 may be a single processor (e.g., within the SMAS device 100), or it may include multiple processors, which may be co-located (e.g., within the SMAS device 100) or physically separate from one another. For example, a first portion of the at least one processor 130 may be included within the SMAS device 100, and a second portion of the at least one processor 130 may be external to the SMAS device 100. In embodiments where the at least one processor 130 includes a first portion and a second portion, the first portion may be responsible for obtaining characteristics of the sensors 105, determining whether the sensors 105 detected a marker during a polling cycle based on the characteristics, and recording (e.g., in the memory 170) whether each of the three sensors 105 detected the presence or absence of at least one marker during the polling cycle, while the second portion is responsible for obtaining a record of the detection results and performing error correction procedures. Alternatively, the first portion may be responsible for obtaining characteristics of the sensors 105, determining whether each of the sensors 105 detected at least one marker during a query cycle based on the characteristics, and providing an indication of whether the sensor 105 detected the marker to another entity via a communication interface (e.g., a wireless or wired interface, such as Ethernet, Wi-Fi, etc.). In such an embodiment, the second portion of the at least one processor 130 may be responsible for obtaining a record of the detection results provided by the first portion of the at least one processor 130 (e.g., a file having a binary entry recording whether each of the plurality of S sensors 105 detected or did not detect at least one marker during each query cycle), performing an error correction procedure, and identifying the base.
[0386] In the foregoing description and in the accompanying drawings, specific terms have been set forth to provide a thorough understanding of the disclosed embodiments. In some instances, the terms or drawings may imply specific details that are not necessary to practice the invention.
[0387] To avoid unnecessarily obscuring the present invention, well-known components are shown in block diagram form and / or in some cases not discussed in detail at all.
[0388] The section headings provided in the embodiments are for convenience or reference only and are not intended to be limiting. Section headings in no way define, limit, interpret, or describe the scope or extent of these sections. Furthermore, while various specific embodiments have been disclosed, it will be apparent that various modifications and variations may be made to the invention without departing from the broader spirit and scope of the invention. For example, features or aspects of any of the described embodiments may be used in combination with or in place of corresponding features or aspects of any other of the described embodiments.
[0389] Certain techniques and methods disclosed herein (e.g., obtaining detection results from sensor 105, performing error correction procedures, etc.) and / or user interfaces for constructing and managing the same can be implemented by a machine executing one or more sequences of instructions (including associated data required for proper instruction execution). Such instructions can be recorded on one or more computer-readable media for later retrieval and execution within one or more processors of a special-purpose or general-purpose computer system or a consumer electronic device or appliance. Computer-readable media in which such instructions and data may be embodied include, but are not limited to, various forms of non-volatile storage media (e.g., optical, magnetic, or semiconductor storage media) and carrier waves that can be used to transmit such instructions and data via wireless, optical, or wired signaling media, or any combination thereof. Examples of transmission of such instructions and data via carrier waves include, but are not limited to, transmission (upload, download, electronic mail, etc.) via the Internet and / or other computer networks using one or more data transmission protocols (e.g., HTTP, FTP, SMTP, etc.).
[0390] Unless expressly defined otherwise herein, all terms are intended to be given their broadest possible interpretation, including the meanings encompassed by this specification and the drawings and as understood by those skilled in the art and / or as defined in dictionaries, treatises, etc. As expressly stated herein, some terms may not be consistent with their ordinary or customary meanings.
[0391] As used in this specification and the appended claims, the singular forms "a," "an," and "the" do not exclude plural referents unless otherwise indicated. Unless otherwise indicated, the word "or" is to be construed as inclusive. Thus, the phrase "A or B" is to be interpreted as meaning all of the following: "A and B," "A but not B," and "B but not A." Any use of "and / or" herein does not imply that the word "or" alone is exclusive.
[0392] As used in this specification and the appended claims, phrases of the form "at least one of A, B, and C," "at least one of A, B, or C," "one or more of A, B, or C," and "one or more of A, B, and C" are interchangeable and each encompasses all of the following meanings: "only A," "only B," "only C," "A and B but not C," "A and C but not B," "B and C but not A," and "all of A, B, and C."
[0393] To the extent that the terms "include(s)," "having," "has," "with," and variations thereof are used in the detailed description or claims, such terms are intended to be inclusive in a manner similar to the term "comprising," that is, to mean "including, but not limited to,"
[0394] The terms "exemplary" and "embodiment" are used to indicate examples, not preferences or requirements.
[0395] The term "coupled" is used herein to mean both directly connected / attached as well as connected / attached through one or more intermediate components or structures.
[0396] The terms "above," "below," "between," and "over" are used herein to refer to the relative position of one feature with respect to other features. For example, a feature positioned "above" or "below" another feature may be in direct contact with the other feature or may have intervening materials. Additionally, a feature positioned "between" two features may be in direct contact with both features or may have one or more intervening features or materials. In contrast, a first feature positioned "above" a second feature is in contact with the second feature.
[0397] The term "substantially" is used to describe a structure, construction, dimension, etc. that is largely or nearly as described, but due to manufacturing tolerances and the like, practical situations may arise where the structure, construction, dimension, etc. is not always or necessarily exactly as described. For example, describing two lengths as "substantially equal" means that the two lengths are the same for all practical purposes, but on a sufficiently small scale they may not (and need not) be exactly equal. As another example, a structure that is "substantially vertical" will be considered vertical for all practical purposes even if it is not exactly 90 degrees relative to the horizontal.
[0398] The drawings are not necessarily drawn to scale, and the dimensions, shapes, and sizes of features may vary substantially from how they are depicted in the drawings.
[0399] Although specific embodiments have been disclosed, it will be apparent that various modifications and changes may be made thereto without departing from the broader spirit and scope of the invention. For example, features or aspects of any of the described embodiments may be combined with or used in place of corresponding features or aspects of any other of the described embodiments, at least where practicable. Accordingly, the specification and drawings should be regarded in an illustrative rather than a restrictive sense.
Claims
1. A method for detecting and correcting errors generated during a sequencing procedure, the sequencing procedure being performed using a single molecule array sequencer, the single molecule array sequencer comprising a plurality of sensors, each sensor of the plurality of sensors being configured to detect no more than one molecule at a time, the method comprising: detecting a failed marker removal error generated by a first sensor of the plurality of sensors, the failed marker removal error generated during a sequencing procedure; and The failed marker removal error is corrected in a record of a test result from the sequencer, the test result from the first sensor.
2. The method of claim 1 , wherein detecting a failed marker removal error generated by a first sensor of the plurality of sensors comprises: The first of the plurality of sensors is determined to have detected a label after a cleavage and wash step of the sequencing procedure and before introduction of a labeled nucleotide in a next interrogation step, the next interrogation step being the first interrogation step after the cleavage and wash step.
3. The method of claim 2, wherein correcting the failed tag removal error comprises: A value corresponding to the next query step is changed in the record of the detection result from a first value indicating that the marker is detected to a second value indicating that the marker is not detected.
4. The method of claim 1 , wherein detecting a failed marker removal error generated by the first sensor of the plurality of sensors comprises: After the cutting and flushing steps of the sequencer, it is determined that the first sensor of the plurality of sensors detects a marker.
5. The method of claim 1 , wherein detecting a failed marker removal error generated by a first sensor of the plurality of sensors comprises: (a) after introducing a labeled nucleotide into a fluid channel of the single molecule array sequencing device, the labeled nucleotide comprising a plurality of labels, the fluid channel allowing the labeled nucleotide to be incorporated by a molecule within a sensing region of the first sensor, thereby first detecting a characteristic of the first sensor; (b) after (a), determining that the first sensor detects at least one marker based at least in part on the characteristic; (c) after (b), performing a cutting and washing step; (d) after step (c), detecting the characteristics of the first sensor for a second time; (e) After (d), determining that the characteristic of the first sensor is detected for the second time is substantially the same as the characteristic of the first sensor is detected for the first time.
6. The method of claim 5, wherein correcting the failed tag removal error comprises: Changing a value corresponding to a next query step of the sequencing procedure in a record of the detection result from a first value indicating detection of a change to a marker to a second value indicating no detection of a marker, the next query step being the first query step after the cutting and washing step.
7. The method according to any one of claims 1 to 6, further comprising: After correcting the failed marker removal errors, probabilistic error correction is performed to estimate the sequence.
8. The method of claim 7, wherein performing probabilistic error correction to estimate the sequence comprises: a step of identifying, in the sequence program, inconsistencies among a first detection result of the first sensor, a second detection result of the second sensor, and a third detection result of the third sensor; identifying a first error scenario consistent with the first detection result, a second detection result, and a third detection result, wherein it is assumed that the first detection result is correct; determining the likelihood of the first error scenario occurring; identifying a second error scenario consistent with the first test result, the second test result, and a third test result, wherein the first test result is assumed to be incorrect; determining a probability of occurrence of the second error scenario; and The recording of the detection results is adjusted based on a comparison of the likelihood of the first error scenario occurring and the likelihood of the second error scenario occurring.
9. The method of claim 8, wherein adjusting the recording of the detection result based on the comparison of the likelihood of the first error scenario occurring and the likelihood of the second error scenario occurring comprises: In response to the likelihood of the second error scenario occurring being higher than the likelihood of the first error scenario occurring: changing an entry in a record of detection results of the first sensor from a first value indicating that a marker was detected to a second value indicating that a marker was not detected, or An entry in a record of detection results of the first sensor is changed from a second value indicating that a marker was not detected to a first value indicating that a marker was detected.
10. The method of claim 9, wherein adjusting the recording of the detection result based on the comparison of the likelihood of the first error scenario occurring and the likelihood of the second error scenario occurring further comprises: In response to the likelihood of the first error scenario occurring being higher than the likelihood of the second error scenario occurring, an entry in a record of detection results of at least one of the second sensor or the third sensor is altered.
11. The method of claim 8, wherein adjusting the recording of the detection result based on the comparison of the likelihood of the first error scenario occurring and the likelihood of the second error scenario occurring comprises: Delete consecutive identical entries from the record.
12. A system comprising: a plurality of three binding sites, each of the plurality of three binding sites being configured to bind no more than one nucleic acid strand to be sequenced; a plurality of three sensors configured to detect a label attached to a nucleotide incorporated into a nucleic acid strand bound to the plurality of three binding sites, each of the plurality of three sensors for sensing a respective strand of nucleic acid bound to a respective binding site of the plurality of three binding sites; and At least one processor configured to execute one or more machine-executable instructions that, when executed, cause the at least one processor to perform, for each sensor in the plurality of S sensors: (a) In the query step of the sequencer, obtaining a characteristic of the sensor, wherein the characteristic is indicative of the presence or absence of at least one label attached to a nucleotide incorporated into a respective strand of the nucleic acid bound to a respective binding site, and (b) determining, based at least in part on the characteristic obtained in step (a), whether the sensor detects at least one label attached to a nucleotide incorporated into a respective strand of the nucleic acid bound to the respective binding site, (c) detecting the characteristic of the sensor after the cutting process following step (a) and before the next query step of the sequencing process; (d) determining, based at least in part on the characteristic obtained in step (c), whether the sensor is still detecting at least one label attached to a nucleotide incorporated into a respective strand of nucleic acid bound to a respective binding site, thereby indicating an error in failed label removal, (e) in response to determining in step (b) that the sensor is detecting at least one label attached to a nucleotide incorporated into a respective strand of the nucleic acid bound to the respective binding site, and in response to determining in step (d) that the sensor is still detecting at least one label attached to a nucleotide incorporated into a respective strand of the nucleic acid bound to the respective binding site after the cleavage procedure, correcting the failed label removal error by changing a "label detected" indication of the sensor to a "label not detected" indication in a next query step after step (d), and (f) Repeating steps (a) through (e) for a subsequent query step in a plurality of query steps of the sequencing process.
13. The system according to claim 12, wherein: When executed, the one or more machine-executable instructions further cause the at least one processor to: After step (f), identifying a plurality of candidate sequences associated with instances of the particular nucleic acid strand; determining a respective metric for each of the plurality of candidate sequences; and Based at least in part on the respective metrics and criteria, a particular candidate sequence is selected as most likely to be correct.
14. The system according to claim 13, wherein: When executed, the one or more machine-executable instructions further cause the at least one processor to: At least one of the plurality of candidate sequences is eliminated based on a predetermined constraint on the nucleic acid sequence of the specific nucleic acid strand.
15. The system of claim 14, wherein the predetermined constraint is the knowledge that at least one of the plurality of candidate sequences is unlikely to occur naturally.
16. The system according to any one of claims 12 to 15, wherein: When executed, the one or more machine-executable instructions further cause the at least one processor to: creating a record comprising results of a sequencing procedure for at least a subset of the three sensors, wherein the record comprises a set of binary values, wherein a first binary value indicates that at least one marker was detected and a second binary value indicates that the marker was not detected, and wherein, when executed, the one or more machine-executable instructions further cause the at least one processor to: identifying a consecutive number of second binary values in said record, wherein said number is equal to the number of query steps per sequencing cycle, and The consecutive number of second binary values is deleted from the record.
17. The system of any one of claims 12 to 15, wherein each sequencing loop of the sequencing program has a P query step, and wherein, When executed, the one or more machine-executable instructions further cause the at least one processor to: A set of P consecutive indications that a first sensor of the plurality of S sensors did not detect a marker is identified, and the set of P consecutive indications that the first sensor of the plurality of S sensors did not detect a marker is deleted.
18. The system according to any one of claims 12 to 15, wherein: When executed, the one or more machine-executable instructions further cause the at least one processor to: At least one entry of a record is modified based on a majority result of a particular query step, the at least one entry of the record corresponding to the particular query step.
19. An apparatus for sequencing nucleic acids, the apparatus comprising: a fluid chamber comprising a plurality of S binding sites, each of the plurality of S binding sites being configured to bind no more than one nucleic acid strand to be sequenced; a plurality of S magnetic sensors configured to detect a label present in the fluid chamber, each of the plurality of S magnetic sensors for sensing a respective strand of nucleic acid bound to a respective binding site of the plurality of S binding sites; and At least one processor configured to execute one or more machine-executable instructions that, when executed, cause the at least one processor to perform, for each magnetic sensor in the plurality of S magnetic sensors: (a) obtaining a property of the magnetic sensor during a query step of a sequencing process, wherein the property indicates the presence or absence of at least one label attached to a nucleotide incorporated into a respective strand of the nucleic acid bound to a respective binding site; (b) determining, based at least in part on the characteristics obtained in step (a), whether the magnetic sensor detects at least one label attached to a nucleotide incorporated into a respective strand of the nucleic acid bound to the respective binding site; (c) detecting the characteristic of the magnetic sensor after the cutting process following step (a) and before the next query step of the sequencing process; (d) determining, based at least in part on the characteristic obtained in step (c), whether the magnetic sensor is still detecting at least one label attached to a nucleotide incorporated into a respective strand of nucleic acid bound to a respective binding site, thereby indicating a failed label removal error, and (e) in response to determining in step (b) that the magnetic sensor is detecting at least one label attached to a nucleotide incorporated into a respective strand of the nucleic acid bound to the respective binding site, and in response to determining in step (d) that the magnetic sensor is still detecting at least one label attached to a nucleotide incorporated into a respective strand of the nucleic acid bound to the respective binding site after the cleavage procedure, correcting the failed label removal error by changing a “label detected” indication of the magnetic sensor to a “label not detected” indication in a next query step after step (d).
20. The apparatus of claim 19, wherein determining whether the magnetic sensor detects the at least one marker comprises: determining whether the obtained characteristic of the magnetic sensor meets or exceeds a threshold, or The resulting characteristic of the magnetic sensor is compared to a previously detected value of the characteristic.
21. The apparatus according to claim 19, wherein When executed by the at least one processor, the one or more machine-executable instructions further cause the at least one processor to: An error correction procedure is performed on at least one record comprising a result of a sequencing procedure for at least a subset of the plurality of S magnetic sensors at each of a plurality of M query steps.
22. The system of claim 21 , wherein performing the error correction procedure on the at least one record comprises: identifying a plurality of candidate sequences associated with instances of a particular nucleic acid strand based on at least a portion of the at least one record, and A determination is made or estimated as to which of the plurality of candidate sequences is most likely to be correct.
23. The apparatus of claim 22, wherein determining or estimating which of the plurality of candidate sequences is most likely to be correct comprises: determining a respective metric for each of the plurality of candidate sequences; and Based at least in part on the respective metrics and criteria, a particular candidate sequence is selected as most likely to be correct.
24. The apparatus of claim 22, wherein determining or estimating which of the plurality of candidate sequences is most likely to be correct comprises eliminating at least one of the plurality of candidate sequences based on a predetermined constraint on the nucleic acid sequence of the particular nucleic acid strand.
25. The apparatus of claim 24, wherein the predetermined constraint is that a specific sequence of bases in at least one of the plurality of candidate sequences is unlikely to occur naturally.
26. The apparatus of any one of claims 21 to 25, wherein each sequencing cycle of the sequencing procedure has a P query step, and wherein performing the error correction procedure on the at least one record comprises: identifying in the at least one record a set of P consecutive indications that no markers were detected, and The set P of consecutive indications for which no mark was detected is deleted from the at least one record.
27. The apparatus of any one of claims 21 to 25, wherein performing the error correction procedure on the at least one record comprises: At least one entry of the at least one record is modified based on a majority result of a specific query step, the at least one entry of the at least one record corresponding to the specific query step.
28. A method for sequencing a plurality of three nucleic acid chains using the apparatus of claim 19, the method comprising: Binding the plurality of S nucleic acid chains to the plurality of S binding sites; and performing a sequencing procedure including M query steps to capture M detection results of respective sensors of the plurality of S magnetic sensors, each of the M detection results indicating whether the respective sensor of the plurality of S magnetic sensors detected at least one mark in the fluid chamber during a respective query step of the M query steps, wherein performing the sequencing procedure includes: the at least one processor executing the one or more machine-executable instructions to perform steps (a) to (e) for each magnetic sensor of the plurality of S magnetic sensors.
29. The method of claim 28, wherein each sequencing loop of the sequencing procedure has a P query step, and wherein the method further comprises: creating S records, each of the S records capturing a result of the sequencing procedure for each of the S sensors in each of the M query steps; identifying, in at least one record of said at least one subset of said S records, a set of P consecutive indications that no marker was detected, and The set P of consecutive indications for which no mark was detected is deleted from the at least one record.
30. The method of claim 29, further comprising adjusting one or more of the at least one subset of the S records based at least in part on a likelihood of a failed nucleotide incorporation (FNI) error, a failed nucleotide removal (FNR) error, and / or a failed label detection (FLD) error.
31. The method of claim 29, wherein the at least one subset of the S records comprises an odd number of at least three records representing sequencing results of instances of a first nucleic acid strand, and wherein the method further comprises: identifying a majority of detection results for a particular query step in said at least one subset of said S records; and Based at least in part on the majority of detection results, a base of the first nucleic acid strand is identified or not identified for the particular query step.
32. The method according to any one of claims 29 to 31, further comprising, for a selected test result of the M test results: In response to selected detection results in more than half of the at least one subset of the S records indicating deletion of at least one tag in the fluid chamber, a base of at least one of the plurality of S nucleic acid strands is identified.
33. A method for mitigating errors in sequencing data generated by a nucleic acid sequencing procedure using a single molecule sensor array, the single molecule sensor array having a plurality of sensors, each of the plurality of sensors being associated with a respective binding site of a plurality of binding sites, each of the plurality of binding sites being configured to bind no more than one nucleic acid strand to be sequenced, the sequencing data indicating, for each sensor of the plurality of sensors, whether (i) the sensor detected at least one marker at each interrogation step, and (ii) the sensor continued to detect at least one marker after a cleavage procedure performed between interrogation steps, the method comprising: identifying a plurality of records in the sequencing data, each of the plurality of records capturing a respective sequencing result for a respective instance of a first strand of nucleic acid, each of the plurality of records having a plurality of entries, each of the plurality of entries indicating, for a respective query step in a plurality of query steps of the nucleic acid sequencing procedure, either (a) detection of a marker by a respective sensor associated with the respective instance of the first strand of nucleic acid, or (b) non-detection of a marker by the respective sensor associated with the respective instance of the first strand of nucleic acid; identifying at least one failed marker removal FLR error in the plurality of records by identifying at least one interrogation step immediately after the sensor continues to detect the at least one marker after the cutting procedure and immediately after the sensor detects the at least one marker during an immediately preceding interrogation step; correcting at least one FLR error by changing at least one entry in the plurality of records from a "mark detected" entry to a "mark not detected" entry, the at least one entry corresponding to the at least one query step after the sensor continues to detect the at least one mark following the cutting procedure and after the sensor detected the at least one mark during an immediately preceding query step, thereby establishing an adjusted plurality of records; determining a plurality of candidate sequences for the first strand of nucleic acid based on the adjusted plurality of records, each of the plurality of candidate sequences estimating at least a portion of the nucleic acid sequence of the first strand of nucleic acid; and A particular candidate sequence of the plurality of candidate sequences is identified as at least a portion of the nucleic acid sequence of the first strand of nucleic acid, the particular candidate sequence being most likely to be correct among the plurality of candidate sequences.
34. The method of claim 33, wherein identifying the plurality of records comprises: Searching the sequencing data for the barcode associated with the first strand of nucleic acid, or. A common sequence of entries in each of the plurality of records is identified.
35. The method of claim 33, wherein determining the plurality of candidate sequences of the first strand of nucleic acid comprises: identifying a particular query step in the adjusted plurality of records at which the first sensor detected a respective marker and the second sensor did not detect any marker; establishing a first candidate sequence, the first candidate sequence assuming that the first sensor should detect a marker and correctly detects the respective marker; and A second candidate sequence is established assuming that the first sensor should not detect any markers but incorrectly detected the respective markers.
36. The method of claim 33, wherein determining the plurality of candidate sequences of the first strand of nucleic acid comprises: identifying a particular query step in the adjusted plurality of records at which the first sensor detected a respective marker and the second sensor did not detect any marker; establishing a first candidate sequence that assumes the second sensor should detect a marker but incorrectly does not detect any marker; and A second candidate sequence is established that assumes that the second sensor should not detect a marker and correctly does not detect any marker.
37. The method of any one of claims 33 to 36, wherein each sequencing cycle in the nucleic acid sequencing process has a P query step, and wherein determining the plurality of candidate sequences of the first strand of nucleic acid comprises: identifying a set of P consecutive entries in at least one of the adjusted plurality of records indicating that a marker was not detected, and The set of P consecutive entries indicating no detected markers are deleted from the at least one of the adjusted plurality of records.
38. The method of any one of claims 33 to 36, wherein the at least a portion of the nucleic acid sequence of the first strand of nucleic acid is a single base, and wherein identifying the particular candidate sequence that is most likely to be correct among the plurality of candidate sequences comprises identifying a majority of the results of the particular query step represented by the adjusted plurality of records.
39. The method of any one of claims 33 to 36, wherein identifying the particular candidate sequence among the plurality of candidate sequences that is most likely to be correct comprises: determining a respective metric for each of the plurality of candidate sequences; and Based at least in part on the respective metrics and criteria, a particular candidate sequence is selected as most likely to be correct.
Citation Information
Patent Citations
Labelled nucleotides
US7057026B2
Fluorescent nucleotide analogs and uses therefor
US7405281B2
Labelled nucleotides
US7414116B2
Modified nucleotides
US7541444B2
Labeled nucleotide analogs and uses therefor
US8058031B2