High-throughput nucleic acid sequencing using single-molecule sensor arrays
Single-molecule array sequencing with magnetic sensors and error correction methods addresses the limitations of existing technologies by enhancing throughput, reducing errors, and extending read lengths in DNA sequencing.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- F HOFFMANN LA ROCHE & CO AG
- Filing Date
- 2024-06-27
- Publication Date
- 2026-04-17
AI Technical Summary
Existing DNA sequencing technologies face challenges in achieving longer read lengths with lower error rates, as cluster sequencers have limited read lengths due to error propagation, while single-molecule sequencers introduce errors that are too large for high-precision diagnostics.
The use of single-molecule array sequencing (SMAS) apparatus with magnetic sensors that detect labels bound to nucleotides in individual nucleic acid strands, combined with error correction methods to mitigate sequencing errors.
This approach enables higher throughput, lower error rates, and longer read lengths compared to cluster-based methods, improving sequencing accuracy and precision.
Smart Images

Figure 0007847614000130 
Figure 0007847614000131 
Figure 0007847614000132
Abstract
Description
Technical Field
[0001] Cross - reference to related applications This application claims priority to U.S. Provisional Application No. 63 / 013,236, entitled "High - Throughput DNA Sequencing Using a Single - Molecule Sensor Array" (Attorney Docket No. ROA - 1002P - US / P36083 - US), filed on April 21, 2020, and incorporates the entire content thereof by reference herein. This application also incorporates by reference in its entirety for all purposes PCT Application No. PCT / US20 / 27290, entitled "Nucleic Acid Sequencing by Synthesis Using a Magnetic Sensor Array" (Attorney Docket No. ROA - 1000 - WO / P35097 - WO), filed on April 8, 2020, and published as International Publication No. 2020 / 210370 on October 15, 2020, and PCT Application No. PCT / US2021 / 021274, entitled "Magnetic Sensor Arrays for Nucleic Acid Sequencing, and Methods of Making and Using the Same" (Attorney Docket No. ROA - 1001 - WO / P35967 - WO), filed on March 7, 2021.
Background Art
[0002] Background Commercially successful approaches to DNA sequencing involve either the synthesis and analysis of clonal deoxyribonucleic acid (DNA) clusters or the detection of individual DNA molecules. Cluster sequencers exhibit error rates low enough for diagnostic applications, but have very limited read lengths due to the nature of error propagation in molecular ensembles. Single - molecule sequencers can generate considerably longer reads, but often exhibit static and dynamic heterogeneities that introduce errors too large for high - precision diagnostics.
[0003] Therefore, there is generally a need to improve DNA sequencing and nucleic acid sequencing in order to enable longer reads with lower error rates.
Summary of the Invention
[0004] Summary This summary represents non-limiting embodiments of the present disclosure.
[0005] Embodiments of single-molecule array sequencing (SMAS) apparatus and systems are disclosed herein. Each of the multiple sensors in the sensor array of the SMAS apparatus detects a label bound to a nucleotide incorporated into a single nucleic acid strand bound to its respective binding site. Each sensor can detect a single label (e.g., fluorescent, magnetic, organometallic, charged molecule, etc.) bound to the incorporated nucleotide. Methods for using the SMAS apparatus and systems for highly scalable nucleic acid (e.g., DNA) sequencing based on sequencing by synthesis (SBS) of multiple instances of clonal amplified DNA immobilized on such SMAS apparatus are also disclosed. Error correction methods for mitigating errors (e.g., false or false label detection) that occur when sequencing individual nucleic acid strands are also disclosed.
[0006] In some embodiments, the apparatus for sequencing nucleic acids comprises a fluid chamber, a plurality of S magnetic sensors configured to detect labels present in the fluid chamber, and at least one processor. The fluid chamber comprises a plurality of S binding sites, each of which is configured to bind to one or fewer nucleic acid strands. Each of the S magnetic sensors senses each strand of nucleic acid bound to each of the S binding sites. At least one processor is configured to execute one or more machine-executable instructions, which, once executed, are used to query each of a plurality of M query steps in the sequencing procedure. In the query process, for each of the S magnetic sensors, at least one processor is made to (a) acquire the respective characteristics of each magnetic sensor, wherein each characteristic indicates the presence or absence of at least one marker, and (b) determine, at least partially, whether each magnetic sensor detected the presence or absence of at least one marker during the query process.
[0007] In some embodiments, the system comprises a plurality of S sensors (e.g., magnetic, optical, etc.) configured to detect labels, each of which S binding sites is configured to bind to one or fewer strands of nucleic acid, and at least one processor. Each of the S sensors is configured to sense each strand of nucleic acid bound to each of the S binding sites. At least one processor is configured to execute one or more machine-executable instructions, which, when executed, cause at least one processor to perform, for each of the S sensors, in each of a plurality of M query steps of a sequencing procedure, (a) obtain each characteristic of each sensor, wherein each characteristic indicates the presence or absence of at least one label, and (b) determine, at least in part, based on each obtained characteristic, whether each sensor detected the presence or absence of at least one label during the query step. In addition, one or more machine-executable instructions, when executed, further include causing at least one processor to perform an error correction procedure on at least one record, the at least one record including the results of an array determination procedure for at least a subset of S sensors in each of M query steps.
[0008] In some embodiments, a method for sequencing a plurality of S nucleic acid strands using a SMAS device includes (a) binding S nucleic acid strands to S binding sites; (b) generating S records by performing a sequencing procedure comprising M query steps, each of which S records incorporates M detection results from one of the S sensors, and each of the M detection results indicates whether, during each of the M query steps, one of the S sensors detected at least one label in the fluid chamber; and (c) applying an error correction procedure to at least a subset of the S records in order to estimate the nucleic acid sequence of at least one of the S nucleic acid strands.
[0009] Some embodiments are methods for mitigating errors in sequencing data generated as a result of a nucleic acid sequencing procedure using a single-molecule sensor array, wherein the single-molecule sensor array has multiple sensors, each of which is associated with each of multiple binding sites, and each of the multiple binding sites is configured to bind to one or fewer strands of nucleic acid being sequenced. In some such embodiments, the method is to (a) identify a plurality of records in sequencing data, each of which records incorporates the respective sequencing result for each instance of the first strand of nucleic acid, each of which records has a plurality of entries, each of which entries indicates, for each of the multiple query steps of the nucleic acid sequencing procedure, either (i) a label was detected by the respective sensor associated with each instance of the first strand of nucleic acid, or (ii) a label was not detected by the respective sensor associated with each instance of the first strand of nucleic acid; (b) determine a plurality of candidate sequences for a first nucleic acid strand based on the plurality of records, each of which candidate sequences estimates at least a portion of the nucleic acid sequence of the first nucleic acid strand; and (c) select from the plurality of candidate sequences as at least a portion of the nucleic acid sequence of the first strand of nucleic acid. This includes identifying a specific candidate sequence from multiple candidate sequences that are most likely to be correct.
[0010] The disclosed sequencing and error correction apparatus, systems, and methods promise potentially higher throughput, lower error rates, and longer read lengths compared to cluster-based approaches. [Brief explanation of the drawing]
[0011] The purposes, features, and advantages of this disclosure will be readily apparent from the following description of specific embodiments, in conjunction with the accompanying drawings. [Figure 1] Figure 1 shows some examples of magnetic sensors according to several embodiments. [Figure 2A] The resistance of a magnetoresistive (MR) sensor, which can be used according to several embodiments, is shown. [Figure 2B] The resistance of a magnetoresistive (MR) sensor, which can be used according to several embodiments, is shown. [Figure 3A] Figure 3A shows a spin torque oscillator (STO) sensor that can be used according to several embodiments. [Figure 3B] Figure 3B shows the experimental response of STO under exemplary conditions. [Figure 3C] This shows a short nanosecond field pulse of STO that can be used according to several embodiments. [Figure 3D] This shows a short nanosecond field pulse of STO that can be used according to several embodiments. [Figure 4A] Figure 4A shows a single sensor in a cluster sequencing device used to sense several N cloned amplified DNA strands in its vicinity. [Figure 4B] Figure 4B shows an example of several S single-molecule sensors, each used by a SMAS instrument to monitor individual single-stranded DNA (ssDNA), according to several embodiments. [Figure 5A] Figure 5A is a block diagram showing the components of an exemplary SMAS apparatus for nucleic acid sequencing according to several embodiments. [Figure 5B] Some exemplary SMAS apparatuses for nucleic acid sequencing, according to several embodiments, are shown. [Figure 5C] Some exemplary SMAS apparatuses for nucleic acid sequencing, according to several embodiments, are shown. [Figure 5D] Some exemplary SMAS apparatuses for nucleic acid sequencing, according to several embodiments, are shown. [Figure 5E] Figure 5E shows the square grid (or grid) patterns of the sensor according to several embodiments. [Figure 6A] Figure 6A shows sensors, coiled DNA strands, and labels according to several embodiments. [Figure 6B] Figure 6B shows exemplary dimensions of the sensor, elongated DNA strand, and label according to several embodiments. [Figure 7A] Figure 7A shows exemplary geometric arrangements for estimating the packing limit of the sensor array of a SMAS device according to several embodiments. [Figure 7B] Figure 7B shows sensors in a SMAS device arranged in a square grid according to several embodiments. [Figure 8A] This shows sensors in a SMAS device arranged in a hexagonal pattern according to several embodiments. [Figure 8B] This shows sensors in a SMAS device arranged in a hexagonal pattern according to several embodiments. [Figure 9A] Figure 9A shows exemplary geometric arrangements for estimating the packing limit of the sensor array of a SMAS device according to several embodiments. [Figure 9B] Figure 9B shows sensors in a SMAS device arranged in a hexagonal grid according to several embodiments. [Figure 10] Figure 10 compares the density of an exemplary SMAS implementation with that of a state-of-the-art cluster sequencing system. [Figure 11] Figure 11 shows exemplary methods for sequencing multiple nucleic acid strands using a SMAS instrument, according to several embodiments. [Figure 12] Figure 12 is a flowchart of a sequencing procedure using an additive approach, according to several embodiments. [Figure 13] Figure 13 shows additive sequencing protocols according to several embodiments. [Figure 14] Figure 14 is a flowchart of a sequencing procedure using a subtractive approach, according to several embodiments. [Figure 15] Figure 15 shows subtractive sequencing protocols according to several embodiments. [Figure 16]Figure 16 is a flowchart of a sequencing procedure using a modified additive approach, according to several embodiments. [Figure 17] Figure 17 shows a modified additive sequencing protocol according to several embodiments. [Figure 18A] Figure 18A shows nucleotide incorporation failure (FNI) in the cluster sequencing machine. [Figure 18B] Figure 18B shows the FNI for the SMAS system. [Figure 18C] Figure 18C shows the label removal failure (FLR) for the cluster sequencing instrument. [Figure 18D] Figure 18D shows the FLR for the SMAS system. [Figure 18E] Figure 18E shows nucleotide removal failures (FNRs) for cluster sequencing. [Figure 18F] Figure 18F shows the FNR for the SMAS system. [Figure 18G] Figure 18G shows nucleotide detection failure (FLD) for the cluster sequencing instrument. [Figure 18H] Figure 18H shows the FLD for the SMAS system. [Figure 19] Figure 19 is a flowchart illustrating an exemplary sequencing procedure using a modified additive approach with FLR and FNI error detection, according to several embodiments. [Figure 20] Figure 20 shows an example of a recording with FNI and FLR errors. [Figure 21] Figure 21 shows the predicted signal levels detected by the cluster sequencing instrument sensor, which captures the behavior of the molecular ensemble during the sequencing procedure. [Figure 22] Figure 22 illustrates how SMAS devices provide better accuracy when using error correction techniques, according to several embodiments. [Figure 23]This figure shows how, in several embodiments, FNI errors can be corrected by removing the four "no label detected" entries from the run in the recording of detection results from the sequencing procedure. [Figure 24] Figure 24 shows exemplary SBS reaction results from several embodiments. [Figure 25] Figure 25 shows the effect of large cluster sizes on the nucleotide recall accuracy of the cluster sequencing instrument. [Figure 26-1] Figure 26 shows deterministic error correction of FLR and FNI errors in several embodiments. [Figure 26-2] Figure 26 shows deterministic error correction of FLR and FNI errors in several embodiments. [Figure 27] Figure 27 shows the FNI, FLR, and FNR errors of the detection data. [Figure 28] Figure 28 shows FLR error correction and base calling from data generated by a SMAS instrument in several embodiments. [Figure 29] Figure 29 shows FNI error correction and base calling from data generated by a SMAS instrument in several embodiments. [Figure 30-1] Figure 30 shows error correction and base calling from data generated by a SMAS instrument, according to several embodiments. [Figure 30-2] Figure 30 shows error correction and base calling from data generated by a SMAS instrument, according to several embodiments. [Figure 31] Figure 31 shows exemplary detection results from a SMAS device, including FNI, FLR, FNR, and FLD errors. [Figure 32-1] Figure 32 shows the application of error correction procedures to data captured during SBS by a SMAS device, according to several embodiments. [Figure 32-2] Figure 32 shows the application of error correction procedures to data captured during SBS by a SMAS device, according to several embodiments. [Figure 33] Figure 33 is a flowchart illustrating error correction procedures according to several embodiments. [Figure 34A] Figure 34A shows the average signal intensity during the query step where the label is detected, following the successful introduction and integration of the matching nucleotide. [Figure 34B] Figure 34B shows the function fit from the cluster model to the measured intensity. [Figure 35] Figure 35 is a plot of the probability function of the cluster sequencing machine. [Figure 36] Figure 36 shows the discrete probability function of the cluster sequencing device. [Figure 37A] Figure 37A shows the intensity plot of the cluster sequencing device. [Figure 37B] Figure 37B shows the probability distribution function of the cluster sequencing device. [Figure 38A] Plot the probability function of the cluster sequencing device. [Figure 38B] Plot the probability function of the cluster sequencing device. [Figure 39] Figure 39 shows the Nr parameter space of the cluster sequencing system under various conditions. [Figure 40A] Figure 40A shows the calculated probabilities of the cluster sequencing device along the Q30 contour for various combinations of Nr. [Figure 40B] Figure 40B plots the cumulative error probabilities calculated for the cluster sequencing machine. [Figure 41] Figure 41 shows that the cumulative probability of an incorrect base call at position 150 is less than 1 in 100.
number
number
number
number
[0012] For ease of understanding, the same reference numerals are used to indicate identical elements common to the figures, where possible. Elements disclosed in one embodiment are intended to be usefully utilized in other embodiments without specific description. Furthermore, the description of an element in the context of one drawing is applicable to other drawings showing that element. [Modes for carrying out the invention]
[0013] Detailed explanation While some of the descriptions and examples in this specification relate to DNA sequencing, it should be understood that this disclosure applies generally to nucleic acid sequencing.
[0014] Terminology and Notation As used herein, the term “chain” refers to a single nucleic acid chain (e.g., ssDNA). The terms “chain” and “fragment” are used interchangeably when referring to nucleic acids.
[0015] As used herein, the term “multiple” means two or more, but not necessarily all. Therefore, “multiple sensors” means at least two sensors, but not necessarily all sensors in a sensor array or sequencing device / system. Similarly, “multiple binding sites” means at least two binding sites, but not necessarily all binding sites in a sequencing device / system.
[0016] As used herein, the term “instance” when referring to a nucleic acid strand means the template nucleic acid strand or a copy thereof (e.g., produced by an amplification or replication process). Ideally, a copy of a template nucleic acid strand is identical to the template strand, but as is known in the art, copies are not necessarily identical due to replication / amplification errors. It is understood that a replica produced by amplification is still considered a copy of the original nucleic acid strand, even if the amplification procedure introduced errors. Thus, all instances of a strand are, ideally, identical to one another, but may not be.
[0017] As used herein, the term “query cycle” refers to a single cycle of a nucleic acid sequencing procedure during which all possible nucleotides are introduced and, if present, which nucleotides are incorporated into the sequenced strand. For example, in a DNA sequencing procedure, adenine (A), thymine (T), cytosine (C), and guanine (G) are all tested in some (any) order (these do not need to be the same for each query cycle). Depending on the selected sequencing procedure, more than one marker may be detected per strand during a single sequencing cycle, as will be described in detail below.
[0018] As used herein, the term “query step” refers to a step or set of steps in a sequencing procedure during which it is determined whether one or more sensors in the sequencing instrument have detected the label. In the case of a DNA sequencing cycling through all A, T, C, and G, there are four query steps per query cycle (one per nucleotide). For a sensor in use, each query step results in a single determination of whether that sensor has detected the label or not.
[0019] As used herein, the term “detection result” refers to a value indicating either (a) that a label was detected during the query process, or (b) that a label was not detected during the query process. In some embodiments, the detection result is a binary value (e.g., 0 or 1). The detection result may be derived from other data (e.g., signals representing resistance, frequency, intensity, etc.; measured values of resistance, frequency, intensity, etc.).
[0020] As used herein, the term “record” refers to a stored representation of the detection results of a single sensor. If a selected sequencing procedure has M query steps, then upon completion of the sequencing procedure, each record will have M detection results. Records from S sensors may be stored in a single file (for example, as a table with S rows and M columns, or as a table with S columns and M rows), or separate files may be created for each sensor's record.
[0021] As used herein with respect to detection results contained within a record, the term "run" means a sequence of identical consecutive values.
[0022] The terms "sensor" and "sensing element" are used interchangeably in this specification.
[0023] In this specification, the variable S is used to refer to several sensors within a group of sensors. Sensor S may be sensing instances on the same chain or sensing instances on different chains.
[0024] In this specification, the variable K is used to refer to several sensors within a group of sensors that all sense instances of the same strand.
[0025] sign The nucleic acid sequencing method described herein uses a labeled nucleotide precursor containing a cleavable label. These cleavable labels are, for example, magnetic, fluorescent, organometallic, or charged molecules. could be.
[0026] Each label may include, for example, magnetic nanoparticles, such as molecules, superparamagnetic nanoparticles, or ferromagnetic particles. The magnetic labels can be nanoparticles with high magnetic anisotropy. Examples of nanoparticles with high magnetic anisotropy include, but are not limited to, Fe3O4, FePt, FePd, and CoPt. To facilitate chemical bonding to nucleotides, particles can be synthesized and coated with SiO2. See, for example, M. Aslam, L. Fu, S. Li, and VPDravid, "Silica encapsulation and magnetic properties of FePt nanoparticles," Journal of Colloid and Interface Science, Volume 290, Issue 2, 15 October 2005, pp. 444-449. Because magnetic labels of this size have a permanent magnetic moment whose direction fluctuates randomly on very short timescales, some embodiments further described below rely on highly sensitive sensing schemes to detect the magnetic field fluctuations caused by the presence of the magnetic label.
[0027] Each label may include, for example, a fluorophore. Fluorescent labels are well known in the art and are suitable for use in the disclosure herein.
[0028] The labels may include, for example, organometallic compounds. As is understood, organometallic compounds are any member of the class of substances that contain at least one metal-carbon bond in which carbon is part of an organic group. Examples of organometallic compounds include Gilman reagents (containing lithium and copper), Grinal reagents (containing magnesium), tetracarbonyl nickel, ferrocene (containing transition metals), organolithium compounds (e.g., n-butyllithium (n-BuLi)), organozinc compounds (e.g., diethylzinc (Et2Zn)), organotin compounds (e.g., tributyltin hydride (Bu3SnH)), organoborane compounds (e.g., triethylborane (Et3B)), organoaluminum compounds (e.g., trimethylaluminum (Me3Al)), and the like.
[0029] The label may, for example, contain charged molecules.
[0030] There are several methods for attaching a label to a nucleotide precursor and then cleaving the label after the nucleotide precursor has been incorporated. For example, the label may be attached to a base, in which case it may be chemically cleaved. Alternatively, the label may be attached to a phosphate group, in which case they may be cleaved by polymerase, or, if attached via a linker, by cleaving the linker.
[0031] In some embodiments, the label is linked to a nitrogenous base (e.g., A, C, T, G, or derivative) of the nucleotide precursor. After the nucleotide precursor is incorporated and detected by a sequencing device (e.g., as described in more detail below), the label is cleaved from the incorporated nucleotide.
[0032] In some embodiments, the label is attached via a cleavable linker. Cleavable linkers are known in the art and are described, for example, in U.S. Patent No. 7,057,026, U.S. Patent No. 7,414,116, and their continuations and improvements. In some embodiments, the label is attached to the 5-position of the pyrimidine or the 7-position of the purine via a linker containing an allyl or azide group. In other embodiments, the linker contains a disulfide, indole, or siever group. The linker is alkyl(C) 1~6 ) or alkoxy(C 1~6 ), may further contain one or more substituents selected from nitro, cyano, fluoro groups or groups having similar properties. In short, Linkers can be cleaved by water-soluble phosphine or phosphine-based transition metal-containing catalysts. Other linkers and linker cleavage mechanisms are known in the art. For example, linkers containing trityl, p-alkoxybenzyl esters and p-alkoxybenzylamides, as well as tert-butyloxycarbonyl (Boc) groups and acetal systems, can be cleaved under acidic conditions with proton-releasing cleavage agents. Thioacetals or other sulfur-containing linkers can be cleaved using thiophilic metals such as nickel, silver, or mercury. Cleavage-protecting groups can also be considered to prepare suitable linker molecules. Ester and disulfide-containing linkers can be cleaved under reductive conditions. Linkers containing triisopropylsilane (TIPS) or t-butyldimethylsilane (TBDMS) can be cleaved in the presence of F ions. Photocleavable linkers that are cleaved at wavelengths that do not affect other components of the reaction mixture include linkers containing O-nitrobenzyl groups. Linkers containing benzyloxycarbonyl groups can be cleaved by Pd-based catalysts.
[0033] In some embodiments, the nucleotide precursor includes a label attached to the polyphosphate moiety, as described, for example, in U.S. Patent No. 7,405,281 and U.S. Patent No. 8,058,031. In short, the nucleotide precursor comprises a nucleoside moiety and a chain of three or more phosphate groups, where one or more oxygen atoms are optionally substituted, for example, with S. The label may be attached directly to the α, β, γ, or higher phosphate groups (if present) or via a linker. In some embodiments, the label is attached to the phosphate groups via a non-covalent linker, as described, for example, in U.S. Patent No. 8,252,910. In some embodiments, the linker is a hydrocarbon selected from substituted or unsubstituted alkyls, substituted or unsubstituted heteroalkyls, substituted or unsubstituted allyls, substituted or unsubstituted heteroallyls, substituted or unsubstituted cycloalkyls, and substituted or unsubstituted heterocycloalkyls, see, for example, U.S. Patent No. 8,367,813. The linker may also include a nucleic acid chain; see, for example, U.S. Patent No. 9,464,107.
[0034] In embodiments where the label is linked to a phosphate group, the nucleotide precursor is incorporated into the nascent chain by a nucleic acid polymerase that also cleaves and releases the detectable label. In some embodiments, the label is removed by cleaving the linker, for example, as described in U.S. Patent No. 9,587,275.
[0035] In some embodiments, the nucleotide precursor is an unextendable “terminator” nucleotide, i.e., a nucleotide having a 3' end blocked from the addition of the next nucleotide by a blocking “terminator” group. The blocking group is a reversible terminator that can be removed in order to continue the chain synthesis process described herein. Attaching a removable blocking group to a nucleotide precursor is known in the art. See, for example, U.S. Patent No. 7,541,444, U.S. Patent No. 8,071,739, and their continuations and improvements. Briefly, the blocking group may include an allyl group that can be cleaved by reacting with a metal-allyl complex in aqueous solution in the presence of a phosphine or nitrogen-phosphine ligand. Other examples of reversible terminator nucleotides used for synthetic sequencing include the modified nucleotides described in international application PCT / US2019 / 066670, titled “3' protected nucleotides,” which was filed on 16 December 2019 and published as international publication 2020 / 131759.
[0036] sensor The characteristics and capabilities of the sensors used in the nucleic acid sequencing apparatus, systems, and methods described herein depend on the selection of the label used. The sensors may include, for example, magnetic sensors (e.g., For example, the sensors may be magnetic nanoparticles, organometallic compounds, etc., or optical sensors (for example, to detect fluorophores). It should be understood that other types of sensors may be suitable for detecting various types of labels, and the examples described herein are not intended to be limiting. Generally speaking, the disclosed apparatus, systems, and methods may use any type of label that can be detected by a selected type of sensor, and conversely, the disclosed apparatus, systems, and methods may use any type of sensor that can detect the presence (and absence) of a selected type of label.
[0037] Reference numeral 105 is used herein for single-molecule sensors in general, regardless of the type of single-molecule sensor they are (and the type of label they detect). Reference numeral 15 is used for sensors that sense clusters of nucleic acid strands.
[0038] Magnetic sensor Several embodiments disclosed herein utilize a magnetic sensor to detect the presence of a magnetic label (e.g., magnetic nanoparticles, organometallic complexes, charged molecules, etc.) bound to a nucleotide precursor. Figure 1 shows a portion of a magnetic sensor 105 according to several embodiments. The exemplary magnetic sensor 105 in Figure 1 has a bottom surface 108 and a top surface 109 and comprises three layers, for example, two ferromagnetic layers 106A, 106B separated by a non-magnetic spacer layer 107. The non-magnetic spacer layer 107 may be a metallic material such as copper or silver, in which case the structure is called a spin valve (SV), or it may be an insulator such as alumina or magnesium oxide, in which case the structure is called a magnetic tunnel junction (MTJ). Suitable materials for use in the ferromagnetic layers 106A, 106B include, for example, alloys of Co, Ni, and Fe (which may also be mixed with other elements). In some embodiments, the ferromagnetic layers 106A, 106B are designed so that their magnetic moments are oriented in the plane of the film or perpendicular to the plane of the film. Additional materials may be deposited both below and above the three layers 106A, 106B, and 107 shown in Figure 1 to serve purposes such as interface smoothing, texturing, and protection from processes used to pattern the device into which the sensor 105 is incorporated, but the active area of the magnetic sensor 105 is located in this three-layer structure. Therefore, components in contact with the magnetic sensor 105 may be in contact with one of the three layers 106A, 106B, or 107, or they may be in contact with other parts of the magnetic sensor 105.
[0039] As shown in Figures 2A and 2B, the resistance of the MR sensor is proportional to 1-cos(θ), where θ is the angle between the moments of the two ferromagnetic layers 106A and 106B shown in Figure 1. To maximize the signal generated by the magnetic field and provide a linear response of the magnetic sensor 105 to the applied magnetic field, the magnetic sensor 105 may be designed such that the moments of the two ferromagnetic layers 106A and 106B are oriented to each other at π / 2 radians or 90 degrees in the absence of a magnetic field. This orientation can be achieved by any number of methods known in the art. For example, one solution is to use an antiferromagnet to "pin" the magnetization direction of one of the ferromagnetic layers (either 106A or 106B, called "FM1") by an effect called exchange bias, and then coat the sensor with a double layer having an insulating layer and a permanent magnet. The insulating layer prevents electrical short circuits in the magnetic sensor 105, and the permanent magnet supplies a "hard bias" magnetic field perpendicular to the pinning direction of FM1, which then rotates the second ferromagnetic material (either 106B or 106A, called "FM2") to produce the desired configuration. A magnetic field parallel to FM1 then rotates FM2 around this 90-degree configuration, and the change in resistance yields a voltage signal that can be calibrated to measure the magnetic field acting on the magnetic sensor 105. In this way, the magnetic sensor 105 functions as a magnetic field voltage transducer.
[0040] The example discussed above involves those moments being arranged in the plane of the film at 90 degrees to each other. While the use of directed ferromagnetic materials has been described, it should be noted that, alternatively, a vertical configuration can be achieved by orienting the moment of one of the ferromagnetic layers 106A and 106B out of the plane of the film, which can be achieved using something called perpendicular magnetic anisotropy (PMA).
[0041] In some embodiments, the magnetic sensor 105 utilizes a quantum mechanical effect known as spin transfer torque. In such a device, a current passing through one ferromagnetic layer 106A (or 106B) in the SV or MTJ preferentially transmits electrons having spins parallel to the layer's moment, while electrons with antiparallel spins are likely to be reflected. In this way, the current is spin-polarized such that there are more electrons of one spin type than the other. This spin-polarized current then interacts with a second ferromagnetic layer 106B (or 106A), exerting a torque on the layer's moment. This torque can, under different circumstances, cause the moment of the second ferromagnetic layer 106B (or 106A) to precess around an effective magnetic field acting on the ferromagnet, or it can cause the moment to reversibly switch between two orientations defined by uniaxial anisotropy induced in the system. The resulting spin torque oscillators (STOs) are frequency-tunable by changing the magnetic field acting on them. Therefore, they have the ability to act as magnetic field versus frequency (or phase) transducers (thereby generating an AC signal with frequency), as shown in Figure 3A illustrating the concept of using an STO sensor. Figure 3B shows the experimental response of an STO through a delay detection circuit when an AC magnetic field with a frequency of 1 GHz and a peak-to-peak amplitude of 5 mT is applied across the STO. This result, as well as the results shown in Figures 3C and 3D for short nanosecond magnetic field pulses, demonstrates how these oscillators can be used as nanoscale magnetic field detectors. Further details can be found in "Delay detection of frequency modulation signal from a spin-torque oscillator under a nanosecond-pulsed magnetic field" by T. Nagasawa, H. Suto, K. Kudo, T. Yang, K. Mizushima, and R. Sato. This can be found in field,''Journal of Applied Physics, Vol.111, 07C908 (2012).
[0042] Optical sensor Some nucleic acid sequencing approaches use fluorescent labeling. In such approaches, the nucleic acid molecule to be sequenced is immobilized on a solid support, and the binding of a fluorescently labeled target molecule (e.g., nucleotide) to the molecule is monitored. Optical instruments, such as excitation and reading devices for fluorescence, provide light of a specific wavelength to excite the fluorescent label and detect fluorescence from the label emitted at a slightly different wavelength. Since the beam path of the excitation light must be at least partially different from the beam path of the fluorescence, spectral separation may be achieved using excitation and emission filters (whose spectra do not significantly overlap), and / or vertical or side illumination may be used.
[0043] Optical sensors and array determination devices and methods using fluorescent labels (e.g., fluorophores) are well known in the art.
[0044] Amplification / Replication Nucleic acid sequencing machines generally generate numerous nucleic acid instances from a single nucleic acid strand through an amplification (or replication) process (e.g., instances of single-stranded DNA (ssDNA) from a single DNA molecule). Polymerase chain reaction (PCR) is a well-known method for amplifying double-stranded DNA, enabling the replication of a significant amount of DNA from a small initial amount.
[0045] Cluster sequencing device Several sequencing devices, referred to herein as cluster (CLUS) devices, use amplification techniques to form localized clusters of many DNA strands. For example, a single strand of DNA is used as a template, and PCR amplification generates thousands or millions of instances of the DNA sequence in a local region. By immobilizing at least a portion of the PCR primers onto a solid support, the generated DNA molecules can be immobilized onto the local clusters to form identifiable "clones." The generated DNA clusters may contain ssDNA. Examples of clonal amplification techniques include, for example, bridge PCR and emulsion PCR, including bead-based emulsion PCR. For bridge amplification, a single DNA molecule is amplified to form a DNA cluster by in situ PCR using primers bound to a solid surface such as a glass slide. Each DNA cluster is a physically isolated "clone" consisting of instances of a DNA strand. In emulsion PCR-based clonal amplification, a single DNA molecule is cloned in an emulsion droplet. In some methods, the DNA strand is bound to microbeads in the droplet. Clonal amplification of a single molecule can also be performed in separate microwells.
[0046] As used herein, the term “cluster” refers to a localized cluster of nucleic acid strands that are ideally identical in sequence, generated from clonal amplification. When the nucleic acid is DNA, the cluster contains (ideally) identical DNA strands (or fragments) bound to a solid support. For example, clusters may be generated on a spot on a glass slide or bound to microbeads, microwells, or other microparticles.
[0047] The use of CLUS instruments for fluorescence-based DNA sequencing is well known.
[0048] A sequencing apparatus that uses an array of magnetic sensors for nucleic acid sequencing using clusters is described, for example, in PCT application number PCT / US2021 / 021274, titled "Magnetic sensor array for nucleic acid sequencing, and method for fabricating and using the same" (agent reference number ROA-1001-WO / P35967-WO), filed on March 7, 2021.
[0049] Figure 4A shows a single sensor 15 of the CLUS device used to sense several N clone-amplified DNA strands 101 in its vicinity. Sensor 15 may be, for example, a magnetic sensor for sensing magnetic labels attached to incorporated nucleotides. For convenience, Figure 4A shows strand 101 in contact with sensor 15, but it should be understood that a barrier (e.g., an insulating layer) may be present between sensor 15 and strand 100. Sensor 15 may be a magnetic sensor, for example, as described in PCT application number PCT / US2021 / 021274.
[0050] State-of-the-art commercially available CLUS instruments, such as those that sense fluorescent labels, can utilize hundreds of millions of sensors, each sensing many instances of its respective amplified DNA strand 101. One drawback of some CLUS instruments is that achieving optimal cluster density can be crucial for high-quality sequencing. Specifically, the use of large clusters tends to provide higher data quality but lower data output, while the use of small clusters can result in run failures, insufficient performance, lower Q30 scores, introduction of sequencing artifacts, and lower total data output. To mitigate these issues, newer CLUS instruments use patterned flow cells with separate nanocells for cluster generation. These nanocells are organized in a hexagonal arrangement to more efficiently utilize the flow cell surface area.
[0051] Single-molecule array sequencing system Single-molecule array sequencing (SMAS) systems (referred to herein as “SMAS systems”) are an alternative to CLUS systems. In contrast to CLUS systems, which sense and sequence local clusters of multiple instances of a single nucleic acid strand, SMAS systems use sensors that sense and sequence individual strands of nucleic acid individually. Generally speaking, in an SMAS system, a sensor does not sense more than one physical nucleic acid strand, but different sensors sense instances of the same strand. In other words, multiple instances of a nucleic acid strand exist, but each sensed strand is sensed by a different sensor. Depending on the amplification technique used, individual strands may be randomly distributed throughout the fluid chamber of the SMAS system or located in more localized areas. As will be further described below, the location of specific strand instances can be identified, and error correction procedures can be applied to the detection results corresponding to the instances before calling bases to improve sequencing accuracy compared to CLUS systems. Furthermore, compared to CLUS systems, due to a reasonable chemical failure rate, SMAS systems require fewer instances of each nucleic acid strand being sequenced to achieve accurate sequencing results.
[0052] Figure 4B shows an exemplary set of S single-molecule sensors 105, each used by the SMAS instrument to monitor each single-stranded DNA (ssDNA) 101. Each of the S sensors 105 may be, for example, a magnetic sensor, an optical sensor, etc. Figure 4B shows five single-molecule sensors 105A, 105B, 105C, 105D, and 105E, each sensing its own DNA strand 101 (which may be an instance of the same DNA strand or an instance of a different DNA strand). Each sensor 105 may be a nanoscale sensor so small that, for example, only a single DNA strand 101 can bind to the binding site associated with the sensor 105. (For convenience, Figure 4B shows the strands 101 in contact with the sensor 105, but as will be further described below, in some embodiments, the strands 100 are bound to individual binding sites, each of which is associated with its respective sensor 105.)
[0053] As shown in Figure 4B, consider clone-amplified DNA bound to a solid surface containing a densely packed array of sensors 105. The DNA can be replicated by solid-phase amplification (SPA) to create clusters of monoclonal DNA, with each strand being sensed by a different sensor 105, or the DNA can be amplified in bulk and then immobilized on the surface of the SMAS apparatus. If the DNA is amplified on the surface of the fluid chamber of the SMAS apparatus (e.g., by SPA), sensors 105A, 105B, 105C, 105D, and 105E can sense instances of clone-amplified DNA. Alternatively, if the DNA is amplified in a bulk-off apparatus and added to the fluid chamber of the SMAS apparatus, the amplified DNA strands 101 may be more randomly distributed among the sensors 105.
[0054] Figure 5A is a block diagram showing components of an exemplary SMAS apparatus 100 for nucleic acid sequencing according to several embodiments. As shown, the apparatus 100 comprises a sensor array 110 coupled to a circuit 120 coupled to at least one processor 130. The sensor array 110 comprises a plurality of sensors 105 (e.g., magnetic, optical, etc.) which can be arranged in any suitable manner, as will be further described below. The characteristics and properties of the sensors 105 in the sensor array 110 depend on the type of label used for sequencing.
[0055] The circuit 120 may comprise one or more lines that enable, for example, sensors 105 in the sensor array 110 to be investigated by at least one processor 130 (with the help of other components known in the art, such as a current source). For example, during operation, the processor(s) 130 may cause the circuit 120 to apply current to such lines to detect a characteristic of at least one of the multiple sensors 105 in the sensor array 110, the characteristic being whether a mark is present or absent within the range of the sensor 105. This indicates that, in other words, the characteristics (e.g., resistance, frequency, voltage, signal level, etc.) indicate whether the sensor 105 detected at least one indicator or not. For example, at least one processor 130 can evaluate the value of the characteristic (e.g., frequency, wavelength, magnetic field, resistance, noise level, intensity, light color, etc.) and determine whether an indicator was detected (or not) based on a comparison of the characteristic value with a threshold or baseline value (e.g., by determining whether the characteristic value of sensor 105 meets or exceeds a threshold). As another example, at least one processor 130 may compare the acquired characteristic of sensor 105 with a previously detected value of the characteristic (e.g., a baseline value of sensor 105) and determine whether an indicator was detected based on a change in the characteristic value (e.g., a change in magnetic field, resistance, noise level, frequency, wavelength, intensity, light color, etc.). For example, as further described below in the discussion of Figure 19, at least one processor 130 can evaluate the characteristics obtained from sensor 105 to determine whether sensor 105, which detected a label during the first query step of the sequencing procedure, still detects the label after the cutting step, which should have removed the label. Similarly, at least one processor 130 can evaluate the changes in characteristics between query steps to determine whether sensor 105 (a) did not detect a label during either query step, (b) detected a label during both query steps, (c) did not detect a label during the first query step but detected a label during a subsequent query step, and / or (d) detected a label during the first query step but did not detect a label during a subsequent query step.
[0056] The features detected depend on the type of label used in the sequencing procedure. The label may be fluorescent, for example, in which case the sensor 105 may be an optical sensor capable of detecting, for example, the wavelength, frequency, modulation frequency, color, or intensity of light emitted by the fluorescent label. Optical sensors suitable for detecting fluorescent labels are well known in the art. When the label used in the nucleic acid sequencing procedure is fluorescent, in some embodiments the circuit 120 enables at least one processor 130 to detect deviations or fluctuations in the light (or electromagnetic energy) detected by some or all of the sensors 105 in the sensor array 110.
[0057] The label may be, for example, magnetic (e.g., magnetic nanoparticles, organometallic compounds, charged molecules, etc.), in which case the sensor 105 may be a magnetic sensor capable of detecting magnetic properties. Magnetic sensors are described, for example, in the present applicant's previously filed patent applications, including PCT application number PCT / US20 / 27290, filed on April 8, 2020, and published on October 15, 2020, as international publication no. 2020 / 210370, entitled “Synthesis of nucleic acid sequencing using a magnetic sensor array” (agency reference number ROA-1000-WO / P35097-WO). In some embodiments where the label is magnetic, the sensor 105 is a magnetoresistive (MR) sensor capable of detecting, for example, a magnetic field or resistance, a change in the magnetic field or a change in resistance, or a noise level. In some embodiments, each of the sensors 105 of the sensor array 110 is a thin-film device that uses the MR effect to detect a magnetic label bound to a nucleotide incorporated into a single strand of nucleic acid bound to its respective binding site. Sensor 105 can operate as a potentiometer having a resistance that changes as the intensity and / or direction of the sensed magnetic field changes. In some embodiments using magnetic labels, sensor 105 comprises a magnetic oscillator (e.g., a spin-torque oscillator (STO)) and the characteristic indicating whether at least one label has been detected is the frequency of a signal associated with or generated by the magnetic oscillator, or a change in the frequency of the signal.
[0058] If the label used in the nucleic acid sequencing procedure is magnetic, in some embodiments, at least one processor 130, with the help of circuit 120, detects deviations or fluctuations in the magnetic environment of some or all of the sensors 105 in the sensor array 110. For example, the magnetic label The MR-type sensor 105 in the absence of the magnetic label should have relatively low noise above a certain frequency compared to the sensor 105 in the presence of the magnetic label, because fluctuations in the magnetic field from the magnetic label cause fluctuations in the moment of the sensing ferromagnet. These fluctuations can be measured using heterodyne detection (e.g., by measuring the noise power density) or by directly measuring the voltage of the sensor 105 and can be evaluated using a comparator circuit to compare it with another sensor element that does not sense the coupling site. If the sensor 105 includes an STO element, the fluctuating magnetic field from the magnetic label causes a phase jump in the sensor 105 due to instantaneous changes in frequency, which can be detected using a phase detection circuit. Another option is to design the STO to oscillate only within a small magnetic field range so that the presence of the magnetic label turns off the oscillation.
[0059] It should be understood that the examples of labels and sensors 105 provided above are for illustrative purposes only. In general, any type of label capable of labeling nucleotide precursors can be used with any type of sensor 105 array 110 capable of detecting that type of label.
[0060] Figures 5B, 5C, and 5D show exemplary SMAS apparatus 100 for nucleic acid sequencing according to several embodiments. The exemplary SMAS apparatus 100 uses magnetic labels and magnetic sensors 105. Figure 5B is a top view of the apparatus 100. Figure 5C is a cross-sectional view at the location of the dashed line labeled "5C" in Figure 5B, and Figure 5D is a cross-sectional view at the location of the dashed line labeled "5D" in Figure 5B.
[0061] The exemplary apparatus 100 shown in Figures 5B, 5C, and 5D includes a sensor array 110 for sensing magnetic labels in a fluid chamber 115. The sensor array 110 includes a plurality of magnetic sensors 105, with 16 sensors 105 shown in the array 110 in Figure 5B. It should be understood that the implementation of the SMAS apparatus 100 can include any number of sensors 105 (e.g., hundreds, thousands, or millions of sensors 105). To avoid obscuring the drawing, only seven of the sensors 105, namely sensors 105A, 105B, 105C, 105D, 105E, 105F, and 105G, are labeled in Figure 5B. As described above, the magnetic sensors 105 detect the presence or absence of magnetic labels; that is, each magnetic sensor 105 detects whether at least one magnetic label is present in its vicinity.
[0062] Referring here to Figures 5C and 5D in relation to Figure 5B, each sensor 105 is shown in exemplary embodiments of the device 100 as having a cylindrical shape. However, it should be understood that, in general, the sensor 105 can have any suitable shape. For example, the sensor 105 may be a three-dimensional rectangular parallelepiped. Furthermore, different sensors 105 may have different shapes (for example, some may be rectangular parallelepipeds and others cylindrical, etc.). It should be understood that the drawings are for illustrative purposes only.
[0063] As shown in Figures 5C and 5D, the apparatus 100 includes a fluid chamber 115. The fluid chamber 115 comprises a plurality of binding sites 116 (e.g., S binding sites 116). In some embodiments, the fluid chamber 115 holds fluids used during the nucleic acid sequencing procedure (e.g., nucleotide precursors and other fluids). However, it should be understood that embodiments in which the fluid chamber 115 does not hold fluid are contemplated and are within the scope of the disclosure herein. For example, the binding sites 116 may be placed on a removable (or movable) portion (e.g., a panel, plate, slide, etc.) which may be immersed in reagents and other fluids after the nucleic acid chain has bound to the binding sites 116, and then positioned so that the sensor 105 can detect the label. Thus, although the name fluid chamber 115 suggests that it holds fluid, it is not essential that the fluid chamber 115 holds fluid.
[0064] As shown in Figures 5B, 5C, and 5D, each of the sensors 105 is associated with its respective coupling site 116. (For simplicity, in this document, coupling sites are generally given the reference number 116. Individual coupling sites are given a letter after the reference number 116.) In other words, there is a one-to-one relationship between sensors 105 and coupling sites 116. As shown in Figure 5B, sensor 105A is associated with coupling site 116A, sensor 105B is associated with coupling site 116B, sensor 105C is associated with coupling site 116C, sensor 105D is associated with coupling site 116D, sensor 105E is associated with coupling site 116E, sensor 105F is associated with coupling site 116F, and sensor 105G is associated with coupling site 116G. Each of the other unlabeled sensors 105 shown in Figure 5B is also associated with its respective coupling site 116. In the exemplary embodiments shown in Figures 5B, 5C, and 5D, each sensor 105 is shown positioned below its respective coupling portion 116, but it should be understood that the coupling portion 116 may be in a different position relative to each of those sensors 105. For example, the coupling portion 116 may be on the side of each sensor 105.
[0065] Each of the binding sites 116 is configured to bind one or fewer nucleic acid strands (e.g., ssDNA) to the SMAS apparatus 100 in the fluid chamber 115. In other words, each binding site 116 has properties and / or features that allow single-stranded and single-stranded nucleic acids to bind for sensing (and sequencing) by the respective sensors 105. The respective sensors 105 can then detect labels attached to nucleotides incorporated into the nucleic acid strand bound to the binding site 116 during the nucleic acid sequencing procedure, as will be further discussed below. In some embodiments, the binding site 116 has a structure (or more structures) configured to fix the nucleic acid to the binding site 116. For example, the structure (or more structures) may include cavities or protrusions. Figures 5C and 5D show the binding sites 116 extending from the surface of the fluid chamber 115, but it should be understood that the binding sites 116 may be coplanar with the surface of the fluid chamber 115 or may be etched.
[0066] The binding sites 116 can have any suitable size and shape to facilitate the attachment of single-stranded and single-stranded nucleic acids to each binding site 116. For example, the shape of the binding site may be similar to or identical to the shape of the sensor 105 (for example, if the sensor 105 is three-dimensional and cylindrical, the binding site 116 may be cylindrical, protruding from the surface of the fluid chamber 115 or forming a fluid container within the surface of the fluid chamber 115, with a radius that may be larger than, smaller than, or the same size as the radius of each sensor 105. If the sensor 105 is three-dimensional and rectangular, the binding site 116 may be a rectangular parallelepiped having a surface 116 that is larger than, smaller than, or the same size as the nearest part of the sensor 105, etc.). In general, the surfaces of the binding sites 116 and the fluid chamber 115 can have any shape and properties to facilitate the binding of a single nucleic acid strand to each binding site 116, enabling the sensor 105 to detect a label attached to a nucleotide incorporated at each binding site 116.
[0067] Figures 5C and 5D show a sealed fluid chamber 115 having a apex extending in the xy plane, although the fluid chamber 115 does not need to be sealed. In some embodiments, the surface of the fluid chamber 115 has properties and characteristics that protect the sensor 105 from any fluid inside the fluid chamber 115, but still allow the nucleic acid chain to bind to the binding site 116, and the sensor 105 detects a label attached to a nucleotide incorporated into the nucleic acid chain bound to the binding site 116. The material of the fluid chamber 115 (and optionally the material of the binding site 116) may be an insulator or may contain an insulator. In some embodiments, the surface of the fluid chamber 115 contains an organic polymer, a metal, or a silicate. The fluid chamber 115 may contain, for example, a metal oxide, silicon dioxide, polypropylene, gold, glass, or silicon. The thickness of the surface of the fluid chamber 115 allows the sensor 105 to detect the nucleotide incorporated into the nucleic acid chain bound to the binding site 116 in the fluid chamber 115. The magnetic label bound to the sensor can be selected to detect it. In some embodiments, the surface thickness of each sensor 105 is approximately 3 to 20 nm, so that the distance from any label bound to a nucleotide incorporated into a nucleic acid chain bound to each binding site 116 of the sensor 105 is approximately 5 nm to 50 nm. It should be understood that these values are for illustrative purposes only. It will be understood that one implementation configuration may have a fluid chamber 115 with a thicker or thinner surface.
[0068] The circuit 120 of the device 100 may have one or more lines 125. In some embodiments, each of a plurality of sensors 105 is coupled to at least one line 125. In the example shown in Figures 5B, 5C, and 5D, the device 100 has eight lines 125A, 125B, 125C, 125D, 125E, 125F, 125G, and 125H. (For simplicity, this document generally refers to the line with reference number 125. Each line is given reference number 125 followed by a letter.) A pair of lines 125 can be used to access (e.g., query) individual sensors 105. In the exemplary embodiment shown in Figures 5B, 5C, and 5D, each sensor 105 of the sensor array 110 is coupled to two lines 125. For example, sensor 105A is coupled to lines 125A and 125H; sensor 105B is coupled to lines 125B and 125H; sensor 105C is coupled to lines 125C and 125H; sensor 105D is coupled to lines 125D and 125H; sensor 105E is coupled to lines 125D and 125E; sensor 105F is coupled to lines 125D and 125F; sensor 105G is coupled to lines 125D and 125G. In the exemplary embodiments of Figures 5B, 5C, and 5D, lines 125A, 125B, 125C, and 125D are shown to be located below the magnetic sensor 105, and lines 125E, 125F, 125G, and 125H are shown to be located above the magnetic sensor 105. Figure 5C shows sensors 105E associated with lines 125D and 125E, sensors 105F associated with lines 125D and 125F, sensors 105G associated with lines 125D and 125G, and sensors 105D associated with lines 125D and 125H. Figure 5D shows sensors 105D associated with lines 125D and 125H, sensors 105C associated with lines 125C and 125H, sensors 105B associated with lines 125B and 125H, and sensors 105A associated with lines 125A and 125H.
[0069] The sensors 105 of the exemplary SMAS device 100 in Figures 5B, 5C, and 5D are arranged in a rectangular pattern sensor array 110. (It should be understood that a square pattern is a special case of a rectangular pattern.) Each of the lines 125 identifies a row or column of the sensor array 110. For example, each of lines 125A, 125B, 125C, and 125D identifies a different row of the sensor array 110, and each of lines 125E, 125F, 125G, and 125H identifies a different column of the sensor array 110. As shown in Figure 5C, each of lines 125E, 125F, 125G, and 125H is in contact with one of the sensors 105 along its cross-section (i.e., line 125E is in contact with the top of sensor 105E, line 125F is in contact with the top of sensor 105F, line 125G is in contact with the top of sensor 105G, and line 125H is in contact with the top of sensor 105D), and line 125D is in contact with the bottom of each of sensors 105E, 105F, 105G, and 105D. Similarly, as shown in Figure 5D, each of lines 125A, 125B, 125C, and 125D is in contact with the bottom of one of the sensors 105 along its cross-section (i.e., line 125A is in contact with the bottom of sensor 105A, line 125B is in contact with the bottom of sensor 105B, line 125C is in contact with the bottom of sensor 105C, and line 125D is in contact with the bottom of sensor 105D), and line 125H is in contact with the top of each of sensors 105D, 105C, 105B, and 105A.
[0070] The portion of line 125 that connects to sensor 105 and sensor array 110 is such that they are part of the device A dashed line is used in Figure 5B to show that it can be embedded within 100. As described above, the sensor 105 can be protected (e.g., by an insulator) from the contents of the fluid chamber 115 in which it may be enclosed. It should be understood that various illustrated components (e.g., line 125, sensor 105, coupling portion 116, etc.) are not necessarily visible in the physical instance of the device 100 (e.g., they may be embedded in or covered by protective material such as an insulator).
[0071] In some embodiments, part or all of the bonding site 116 is located within a nanowell or trench of line 125 passing through sensor 105. For example, as shown in the example in Figure 5D, line 125H may be thinner on sensor 105 than between sensors 105. For example, line 125H has a first thickness above sensor 105D, a second greater thickness between sensors 105D and 105C, and a first thickness above sensor 105C. Such configurations can be advantageously manufactured using conventional thin-film manufacturing methods (e.g., by depositing a material, applying a mask to the deposited material, and removing a portion of the deposited material according to the mask (e.g., by etching)). Both the bonding site 116 and, if present, the nanowells can be manufactured using conventional techniques.
[0072] For simplicity of explanation, Figures 5B, 5C, and 5D show an exemplary apparatus 100 having only 16 sensors 105, only 16 corresponding binding sites 116, and 8 lines 125 in a sensor array 110. It should be understood that apparatus 100 may have fewer or more sensors 105 in the sensor array 110, and therefore more or fewer binding sites 116. Similarly, embodiments having lines 125 may have more or fewer lines 125. In general, any configuration of sensors 105 and binding sites 116 may be used that enables sensors 105 to detect labels bound to nucleotides incorporated into a single nucleic acid chain bound to the binding sites 116. Similarly, any configuration of one or more lines 125, or any other mechanism, may be used that enables determination of whether sensors 105 have sensed one or more labels. The examples presented herein are not intended to be limiting.
[0073] As explained above, the sensor 105 shown in Figures 5B, 5C, and 5D may be a magnetic sensor 105. Therefore, the sensor 105 is close to the binding site 116 and, consequently, close to the nucleic acid chain bound to the binding site 116. It should be understood that the appropriate position of the sensor array 110 relative to the binding site 116 depends in part on the type of label used and, consequently, the type of sensor 105 used. For example, if the label is a fluorophore and the sensor 105 is an optical sensor, it may be appropriate for the sensor array 110 to be located away from the binding site 116 (e.g., above the binding site 116).
[0074] Figures 5B, 5C, and 5D (and other drawings herein) show a one-to-one relationship between the sensors 105 and the binding sites 116, but it should be understood that each binding site 116 can be sensed by multiple sensors 105. A distinguishing feature of the SMAS device 100 from the CLUS device is that the sensors 105 of the SMAS device 100 do not sense multiple nucleic acid strand instances. If the SMAS device 100 has more sensors 105 than binding sites 116, at least some nucleic acid strands may be sensed by multiple sensors 105 (for example, to improve the accuracy of label detection).
[0075] The exemplary sensor array 110 illustrated and described in the situations of Figures 5B, 5C, and 5D is a rectangular array in which the sensors 105 are arranged in rows and columns. In other words, the multiple sensors 105 of the sensor array 110 are arranged in a rectangular grid. In some embodiments, adjacent rows and columns of the rectangular grid pattern are equidistant from each other, and as a result As shown in Figure 5E, the sensors 105 are arranged in a square grid (or lattice) pattern. In embodiments where the sensors 105 are arranged in a square lattice pattern, each sensor 105 has up to four nearest neighbors. For example, as shown in Figure 5E, sensor 105A has four nearest neighbors labeled 105B, 105C, 105D, and 105E. The closest sensor 105 is located at a nearest neighbor distance of 112, as shown in Figure 5E. Thus, each of sensors 105B, 105C, 105D, and 105E is located at a distance of 112 from sensor 105A.
[0076] A commercially viable SMAS apparatus 100 can utilize high-precision nanoscale fabrication of high-density packed nanoscale sensors 105 capable of recognizing individual labels. The size of the functionalized binding sites 116 may be similar to the size of DNA with labels bound in such a way that multiple strands cannot bind to the same binding site 116 or be sensed by the same sensor 105. A well-established metric for evaluating the commercial competitiveness of a sequencer is how densely DNA strands can be packed within the fluid chamber 115.
[0077] Next, a suitable value for the nearest neighbor distance 112, which can be used to determine the size of the SMAS device 100 and / or the maximum number of sensors 105 that can fit within a selected-size SMAS device 100, can be determined based on the characteristics of the sensors 105, the length of the nucleic acid strands that the device 100 intends to sequence, and the characteristics of the labels used. For example, the total length of the nucleic acid strands and the size of the labels used can provide a physical limit on how close two sensors 105 can be placed within the SMAS device 100. In some embodiments, the size of the sensors 105 may be limited by the nanoscale patterning capability of the process used to manufacture the SMAS device 100. For example, using techniques available at the time of writing, the size of each magnetic sensor 105 (e.g., assuming a cylindrical sensor 105, the diameter of the sensor 105 in the xy plane) may be around 20 nm. Assuming that the nucleic acid to be sequenced is DNA and that it is desirable to sequence fragments up to 150 base pairs (bp) in length, the maximum length of the DNA strand 101 to be sequenced is approximately 50 nm in the extended state, but the ssDNA conformation can change between the extended and coiled states depending on the ionic strength of the buffer, as shown in Figure 6A. Since the label 102 is involved in a single-molecule reaction, the label 102 should have molecular dimensions. In the case of a SMAS apparatus 100 using a magnetic sensor 105, the label 102 can be, for example, superparamagnetic nanoparticles, organometallic compounds, or any other functional molecular group that can be detected by the nanoscale magnetic sensor 105. Therefore, it is assumed that each label 102 has a size of approximately 10 nm or less. Using these assumptions, Figure 6B shows the relative dimensions of the magnetic sensor 105, its extended DNA strand 101, and the magnetic label 102.
[0078] A practical SMAS apparatus 100 that uses magnetic sensors 105 to detect magnetic nanoparticles used as labels 102 can be implemented using existing technologies. For the sake of discussion, let us assume that only labels 102 within 20 nm of the edge of the sensor 105 are detected. Magnetic labels 102 that may be selected for nucleic acid sequencing applications (e.g., superparamagnetic nanoparticles, organometallic compounds, etc.) do not generate significant perturbations to the detected magnetic field, so the detection range of each sensor 105 is small. Labels 102 bound to nucleotides incorporated into ssDNA bound to the binding site 116 of a particular sensor 105 can temporarily exist outside the range of each sensor 105, but since ssDNA takes on various conformations during the detection process, it is desirable that the labels not be allowed to reach the sensitive space (detection region) of adjacent sensors 105 when the ssDNA is in its fully extended state.
[0079] The sensor packing limit of a practical SMAS device 100 is, for example, when the label is a superparamagnetic nanoparticle (e.g., iron oxide, platinum iron, etc.) and the sensor array 110 of the SMAS device 100 is non-volatile. It can be derived by assuming a rectangular (e.g., square) array of magnetic tunnel junctions (MTJs) similar to those used in transient data storage applications. In this case, the region of each nanoscale sensor 105 or its immediate vicinity can be functionalized to function as the respective binding site 116. A simple geometric arrangement for estimating the packing limit of the sensor array of the SMAS device 100 is shown in Figure 7A, which shows two sensors 105A and 105B. Each sensor 105A and 105B is assumed to have a cylindrical shape for convenience only, to have a diameter of approximately 20 nm (as described above), and to be able to detect any label within 20 nm of its edge. The sensing region boundary 111 is indicated by the inner dashed line shown in Figure 7A. Sensor 105A senses DNA strand 101A bound to its binding site, and sensor 105B senses DNA strand 101B bound to its binding site. The maximum reach of labels 102A and 102B when bound to nucleotides incorporated into strands 101A and 101B (for example, when a 150-base DNA strand is not fully coiled) is indicated by the outer dotted circle 103. To ensure accuracy in sequencing results, it is desirable that each sensor 105 detects only the label 102 bound to the nucleotide incorporated into DNA strand 101 bound to its respective binding site 116. Therefore, using the above assumption, the minimum nearest neighbor distance 112 between sensors 105 to avoid crosstalk (for example, detecting the label 102 bound to the nucleotide incorporated into nucleic acid strand 101 bound to the binding site 116 of another sensor 105) is approximately 100 nm.
[0080] In some embodiments of the SMAS device 100, the sensor 105 (e.g., MTJ) is arranged in a square grid compatible with existing crosspoint MRAM sensor shapes, as shown in Figure 7B. The area of the unit cell 114 is 10 4 nm 2 Therefore, each DNA strand 101 is approximately 10 4 nm 2 This makes it possible to extend across the entire area, which means approximately 10 10per cm 2 the DNA surface density of the SMAS device 100 is obtained. Assuming that at least 10 instances of each strand 101 within the sensor array 110 are used, then 10 9 unique strands per cm 2 can be sequenced simultaneously, generating 150 Gbase (1.5 billion x 100 bp DNA strand length) of information per square centimeter of the sensor array 110. In an ideal scenario (e.g., when the chemical failure rate is low, only 3 DNA instances are required, as discussed further below), approximately 3.3 x 10 9 different strands per cm 2 can be sequenced simultaneously, and approximately 500 Gbase of data per square centimeter of the sensor array 110 can be generated.
[0081] As a specific example, the SMAS device 100 having a configuration similar to a single Toshiba 4 Gbit density STT-MRAM chip first introduced at the International Electron Devices Meeting (IEDM) in 2016 has the potential to generate approximately 600 Gbase of high-quality data. The minimum distance 112 between the sensors 105 of the Toshiba platform is 90 nm, just slightly below the estimated minimum distance 112 of 100 nm derived above. Thus, crosstalk using a configuration similar to the Toshiba platform may be low even for 150-base-long ssDNA, but shorter fragments can be sequenced to further reduce crosstalk.
[0082] It should be understood that the arrangement of the sensors 105 in a grid pattern (e.g., a square grid as shown in Figure 7B) is one of many possible arrangements. Other arrangements of the sensors 105 are also possible and will be understood by those skilled in the art to be within the scope of the disclosure herein. For example, the sensors 105 may be arranged in a hexagonal pattern, as shown in Figure 8A, which shows a top view of the SMAS apparatus 100. The exemplary SMAS apparatus 100 shown in Figure 8A comprises a sensor array 110 for sensing a marker 102 in a fluid chamber 115. The sensor array 110 includes a plurality of sensors 105, with 16 sensors 105 shown. Implementations of the apparatus 100 can include any number of sensors 105 (e.g., hundreds, thousands, millions, etc.). It is important to understand what is possible. To avoid obscuring the drawings, in Figure 8A only two of the sensors 105, namely sensors 105A and 105B, are labeled. As described above, sensor 105 may be, for example, a magnetic sensor (e.g., to detect the effect of magnetism or magnetic nanoparticles). Generally, sensor 105 can have any suitable size and shape, as described above in the descriptions of at least Figures 5B, 5C, and 5D.
[0083] As shown in Figure 8A, each of the sensors 105 is associated with its respective coupling site 116. In other words, there is a one-to-one relationship between the sensors 105 and the coupling sites 116. As shown in Figure 8A, sensor 105A is associated with coupling site 116A, sensor 105B is associated with coupling site 116B, and each of the other unlabeled sensors 105 is also associated with its respective coupling site 116. In the exemplary embodiment of Figure 8A, each sensor 105 is shown positioned below its respective coupling site 116, but it should be understood that the coupling sites 116 may be in other positions relative to each of those sensors 105. For example, the coupling sites 116 may be on the side of each sensor 105. At least the descriptions of the coupling sites 116 in the descriptions of Figures 5B, 5C, and 5D apply to Figure 8A and the other figures showing the coupling sites 116 and will not be repeated here.
[0084] The exemplary SMAS apparatus 100 in Figure 8A also includes the fluid chamber 115 described in the descriptions of Figures 5B, 5C, and 5D. These descriptions also apply to Figure 8A and will not be repeated here.
[0085] The circuit 120 of the device 100 in Figure 8A may include one or more lines 125. Each of the lines 125 in the exemplary embodiment of Figure 8A identifies a row or diagonal column of the sensor array 110. For example, each of lines 125A, 125B, 125C, and 125D identifies a different row of the sensor array 110, and each of lines 125E, 125F, 125G, and 125H identifies a different diagonal column of the sensor array 110. In the example shown in Figure 8A, the device 100 has eight lines 125A, 125B, 125C, 125D, 125E, 125F, 125G, and 125H, and pairs of lines 125 can be used to access individual sensors 105. For example, lines 125A and 125H can be used to access sensor 105A, and lines 125B and 125H can be used to access sensor 105B. Line 125 may be directed below and / or above sensor 105, in particular, as described in the descriptions of Figures 5B, 5C, and 5D.
[0086] Figure 8A shows an exemplary apparatus 100 having only 16 sensors 105, only 16 corresponding binding sites 116, and 8 lines 125 in a sensor array 110. However, it should be understood that the SMAS apparatus 100 may have fewer or more sensors 105 in the sensor array 110, and therefore more or fewer binding sites 116. Furthermore, the SMAS apparatus 100 may have more or fewer lines 125. In general, any configuration of sensors 105 and binding sites 116 can be used to enable the sensors 105 to detect labels bound to nucleotides incorporated into a single nucleic acid chain bound to the binding sites 116. Similarly, any configuration of one or more lines 125, or any other mechanism, can be used to enable the determination of whether the sensor 105 has sensed one or more labels.
[0087] As shown in Figure 8B, when the sensors 105 are arranged in a hexagonal pattern, each sensor 105 has up to six nearest neighbors, each at a nearest neighbor distance of 112. In other words, each sensor 105 is at a nearest neighbor distance of 112 from each of the other six sensors 105 that are closest to it. For example, as shown in Figure 8B, the unlabeled sensor 105 in the center of the figure has six nearest neighbors labeled 105A, 105B, 105C, 105D, 105E, and 105F. It has sensors 105, all of which are located at a nearest distance of 112.
[0088] The packing limit of the binding site 116 can be derived for a SMAS apparatus 100 using an optical sensor with a hexagonal pattern of binding sites 116 and a fluorescent label 102 (e.g., a fluorophore). Assuming that the label 102 is a fluorophore, the binding site 116 has a hexagonal pattern, and the sensor array 110 is far from the binding site 116, single-molecule fluorescence from the label 102 can be projected into a far field that can be detected by the sensor array 110 equipped with a photosensitive sensor 105. The position of individual fluorophore labels 102 within the SMAS apparatus 100 can be resolved using single-molecule super-resolution imaging techniques, such as those described by CGGalbraith and JAGalbraith, "Super-resolution microscopy at a glance," Journal of Cell Science, Vol. 124(10), 1607-11 (2011). The position of the fluorophore labels 102 can be resolved because the DNA packing dimensions are far below the diffraction limit. While this type of detection can be somewhat complex and / or expensive, the technique has recently been incorporated into commercially available sequencing systems to improve the throughput of cluster-based sequencers. Furthermore, this technique may be implemented in imaging large single-molecule arrays in the near future.
[0089] Figure 9A shows a simple geometric arrangement for estimating the packing limit of the binding sites 116 located in a hexagonal pattern in a SMAS instrument 100 using fluorophore label 102. DNA strand 101A is bound to binding site 116A, and DNA strand 101B is bound to binding site 116B. (Sensor 105 is not shown in Figure 9A because it is assumed that the sensor array 110 is far from the binding sites.) The maximum reach of labels 102A and 102B when bound to the incorporated nucleotides (e.g., when a DNA strand with 150 bases is in its completely uncoiled state) is indicated by the dashed-dotted circle 103. To avoid crosstalk, fluorophore labels 102 bound to adjacent binding sites 116 are not allowed to occupy overlapping space during the imaging process. For example, fluorophore label 102A bound to a particular binding site 116A should not be allowed to reach the space accessible to fluorophore label 102B bound to the adjacent binding site 116B, as ssDNA 101A will explore its acceptable conformation. This restriction also helps to avoid fluorescence quenching. Assuming the use of fluorophore labels 102, the binding sites 116 can be densely packed into a hexagonal lattice, as shown in Figure 9B. Assuming the maximum length of a 150 bp DNA strand 101 is 50 nm, the size of the fluorophore label 102 is 10 nm, the minimum distance from the center to the edge of each binding site 116 is 20 nm, and each DNA strand 101 binds to the center of its respective binding site 116, the minimum distance 112 is 140 nm. Therefore, as shown in Figure 9B, all DNA strands 101 are 1.7 × 10⁻⁶. 4 nm 2 It can occupy a unit cell 114 having an area of 5.9 × 10 9 chain / cm 2 Or, if approximately 10 instances of each DNA strand are present in the SMAS instrument 100, then 5.9 × 10 8 Unique chain / cm 2The DNA surface density is obtained. The SMAS instrument 100 generates approximately 90 Gbases of data from every square centimeter of the sensor array 110. In the best-case scenario where only three DNA replications are required, the sensor array 110 is approximately 2 × 10⁶ 9 intrinsic DNA strands / cm 2 By holding this, the SMAS device 100 can generate approximately 300 Gb of data from every square centimeter of the sensor array 110.
[0090] The above discussion of the hexagonal array was related to the fluorophore label 102 and the optical sensor 105. It is also possible to use a hexagonal arrangement of magnetic sensors 105. The sensor packing limit of the SMAS apparatus 100 having a hexagonal arrangement of the bonding site 116 and magnetic sensors 105 can be derived as described above in the explanation of Figures 7A and 7B. In the case of magnetic sensors 105, the nearest neighbor distance 112 is approximately 100 nm, which corresponds to (see Figure 9B) a unit cell area of 1 14 is approximately 8.7 × 10 3 nm 2 It means that.
[0091] Figure 10 compares the density of SMAS implementations described in the context of Figures 7A and 7B (magnetic labeling 102 and magnetic sensor 105) and Figures 9A and 9B (fluorescent labeling 102 and optical sensor 105) with the density of current state-of-the-art CLUS sequencers. For the purposes of discussion, we assume that the pitch of the nanocell array of patterned flow cells is approximately 500 nm. As shown in the left panel of Figure 10, the nanocells of the CLUS sequencer are arranged in a hexagonal lattice with a lattice constant of 500 nm. Each nanocell holds approximately 50 to 200 identical DNA strands (e.g., generated by solid-phase bridge amplification). The upper right of Figure 10 shows a hexagonal SMAS lattice using fluorophore labeling and super-resolution imaging (e.g., described in relation to Figures 9A and 9B), and the lower right of Figure 10 shows a square SMAS lattice using superparamagnetic nanoparticle labeling and MTJ sensor array 110 (e.g., described in relation to Figures 7A and 7B). The three representations in Figure 10 are scaled proportionally to show how the SMAS grid configuration compares to the CLUS configuration. The black hexagons (left and upper right) and squares (lower right) represent unit cells that hold the minimum number of individual molecules required to call the sequence of a nucleic acid strand. The ideal case, where only three DNA strands are required for successful base calling, is shown for the SMAS grid, which will be discussed in more detail below. Note that in the SMAS instance (right side of Figure 10), the DNA instances are randomly distributed across the entire sensor array 110, and their locations can be identified during the first sequencing cycle, as will be discussed further below.
[0092] As shown in Figure 10, the area of the CLUS unit cell is 2.2 × 10 5 nm 2 Therefore, 4.6 × 10 8 cluster / cm 2 This corresponds to the DNA cluster density. Based on the above assumptions, the CLUS sequencer generates approximately 70 Gbases of data per square centimeter of the sensing area. In contrast, in the ideal case where only three instances of the strand are used, the SMAS instrument 100 generates approximately 500 Gb / cm². 2(Magnetic sensor 105 (e.g., MTJ) and magnetic label 102 (e.g., superparamagnetic nanoparticles)) and approximately 300 Gb / cm 2 Data is generated from the optical sensor 105 (super-resolution imaging) and the fluorescent label 102. The results of exemplary embodiments of the CLUS sequencer and SMAS instrument 100 are summarized in the table below. The table estimates the sequencing throughput assuming only three instances of each DNA strand, and assuming 10 instances of each DNA strand for the SMAS embodiment. [Table 1]
[0093] The table above shows that the SMAS instrument 100 performs better than the standard CLUS instrument when the number of DNA instances used for algorithmic error correction, which will be further explained below, is small (e.g., <10). As the error correction procedure relies on more instances of each ssDNA, the SMAS instrument 100 begins to behave like a CLUS instrument, and there is little or no benefit in sensing individual molecules rather than clusters. Fluorescent SMAS essentially represents the limit of reducing clusters to single molecules. One approach to reduce the cost of sequencing is to reduce the cluster size and bring DNA clusters closer together in order to obtain more information from fixed sensing regions. This approach While this approach reduces the amount of reagents needed to perform sequencing chemistry, it also significantly increases the complexity and cost of imaging hardware by constantly pushing the limits of what is currently possible with commercially available optical instruments. This strategy is a difficult struggle because inscaling cannot be achieved without improving the chemistry in parallel. This is because as clusters become smaller, the problems per reaction become larger, and chemical defects that occur probabilistically at the single-molecule level become more apparent and unacceptable.
[0094] The cost of implementing super-resolution imaging in a CLUS system makes the SMAS system 100, particularly the SMAS system 100 using the magnetic sensor 105 and magnetic label, potentially an alternative to destructive sequencing. The SMAS system 100 disclosed herein, particularly the SMAS system using the magnetic sensor 105, promises superior throughput at significantly lower equipment costs by leveraging technologies and mass production developed by the large-scale semiconductor and data storage industries.
[0095] SMAS sequencing protocol As described above, when the SMAS instrument 100 is used for nucleic acid sequencing, the nucleic acid strand may be amplified either before or after the nucleic acid is added to the SMAS instrument 100 (e.g., by using bridge amplification). Regardless of how the nucleic acid is amplified, the strand can be sequenced one base at a time by SBS (e.g., by synthesizing dsDNA from ssDNA). The SMAS sequencing protocol is described assuming that the nucleic acid being sequenced is DNA. It should be understood that the disclosed protocol may be modified for sequencing other nucleic acids. Such modifications are within the scope of the skills of those skilled in the art, provided that the disclosure herein is understood.
[0096] To simplify the analysis and illustrate the advantages of using the disclosed SMAS instrument 100 instead of a CLUS sequencer, we consider a DNA sequencing protocol in which a single type of label (e.g., molecular, fluorescent, magnetic, etc.) is bound to all four nucleotides (A, T, C, and G). In other words, several types of identical labels are attached to each of the four nucleotides (e.g., if the selected label 102 is an FePt particle, then each of A, T, C, and G is labeled with an FePt particle). These labeled nucleotides are then incorporated into the DNA strand one base at a time using termination chemistry, for example, so that once a nucleotide is incorporated, label 102 is cleaved before the polymerase can move to the next base. Sensor 105 detects the label 102 bound to the nucleotide.
[0097] Figure 11 shows an exemplary method 200 for sequencing multiple nucleic acid strands (e.g., ssDNA) using a SMAS instrument 100. In 202, the method is initiated. In 204, one or more nucleic acid strands may be optionally amplified before being added to the SMAS instrument 100. In 206, the multiple S nucleic acid strands are bound to multiple S binding sites 116 of the SMAS instrument 100 (multiple including at least two, but not necessarily all, of the binding sites 116 of the SMAS instrument 100). Optionally, in 208, the nucleic acid strands are amplified (e.g., via bridge amplification which can be performed in addition to, or instead of, the amplification in 204). In 210, a sequencing procedure is performed. The sequencing procedure may be, for example, a subtractive approach, a subtractive approach, or a modified additive approach, as further described below. The sequencing procedure performed in 210 generates S records, each of which captures several M detection results for one of several S sensors (wherein multiple includes at least two but not all of the sensors 105 of the SMAS device 100, and the M detection results may include just one detection result, a subset of the total number of detection results obtained during the sequencing procedure, or all of the detection results obtained during the sequencing procedure). Each of the M detection results indicates whether the sensor 105 corresponding to the record detected at least one marker during each of the M query steps. The detection results can be stored in a record that can be stored in memory. In step 212, an error correction procedure is performed, as will be described later. The error correction procedure may include deterministic and / or probabilistic error correction techniques. The error correction procedure may be performed, for example, by at least one processor 130 of the SMAS device 100. Alternatively, it may be performed by a processor outside the SMAS device 100 (e.g., an off-device processor in an external computer). The error correction procedure may be performed while the sequence determination procedure is in progress (e.g., in real time or near real time) or afterward. In step 214, the method 200 terminates.
[0098] As described above, in 210, various protocols can be implemented to read nucleic acid sequences (e.g., DNA sequences) using the SMAS instrument 100. For the sake of simplicity of analysis, it is assumed that the multiple S sensors 105 of the SMAS instrument 100 detect only the presence or absence of the label 102 and do not distinguish nucleotides based on the detected signal level. As a result, in some embodiments, the recording of the detection result of each sensor 105 includes only a "yes" or "no" (or 1 / 0 or any other binary indicator) indication of whether the sensor 105 detected the label or not during a particular query step. It should be understood that other approaches are also possible and are within the scope of the disclosure herein. For example, different labels 102 can be bound to different nucleotides. Another example is that instead of a binary "yes" or "no" decision, the value of a characteristic (e.g., resistance, frequency, intensity, etc.) can be detected and / or recorded and a decision can be made based on that criterion regarding whether the label was detected. For example, instead of having simply 0 and 1 (or "no" and "yes") as possible outputs of a sequencing procedure, using different labels for different nucleotides can yield one of five labels: 0 (no label detected), level 1 (label 1 detected), level 2 (label 2 detected), level 3 (label 3 detected), and level 4 (label 4 detected). In such cases, a range of detected characteristics can be defined to distinguish whether any labels were detected at all, and if so, which labels were detected (for example, if the characteristic value is between 0 and a first value, it is determined that no labels were detected; if the characteristic value is between a first and a second value, it is determined that the first label was detected; if the characteristic value is between a second and a third value, it is determined that the second label was detected, etc.).
[0099] The following describes three example DNA sequencing protocols, each containing a repeating query cycle, with each query cycle having four query steps. During each query cycle, four binary "yes" or "no" questions are answered for each ssDNA being sequenced. One query step asks, "Is the detected base adenine?" ("A?") is answered. Another query step asks, "Is the detected base thymine?" ("T?") is answered. Another query step asks, "Is the detected base cytosine?" ("C?") is answered. And yet another query step asks, "Is the detected base guanine?" ("G?") is answered. A record of the detection results obtained during the sequencing procedure can be created as the query cycle containing A?⇒T?⇒C?⇒G? is repeated. The order in which nucleotides are introduced and bases are detected is arbitrary (meaning the order of the query steps is arbitrary), and the order in which bases are tested in the examples herein (A? ⇒ T? ⇒ C? ⇒ G?) is merely illustrative.
[0100] Additive approach In the additive approach, sensor 105 detects nanoscale labels 102 bound to nucleotides having cleavable linkers. All four types of nucleotides have the same type of label 102 (e.g., molecular, fluorescent, magnetic, etc.) and use the same type of cleavable linker. If four detection results are obtained and one of them does not contain an error, multiple S nuclei are detected. A query cycle for label detection of each of the acid chains 101 includes, according to one embodiment, the following steps:
[0101] 1. Obtain the baseline characteristics of each of the multiple S sensors 105 of the SMAS device 100 (for example, by measuring the baseline signal at each of the multiple S sensors 105) (this may include all or fewer of the sensors 105 in the sensor array 110).
[0102] 2. Introduce and incorporate the labeled A nucleotide. Wash away any unbound labeled molecules.
[0103] 3. Inquiry process 1: The characteristics of each of the multiple S sensors 105 (for example, by detecting a signal in each of the multiple S sensors 105) are acquired, and it is determined whether each sensor 105 detected at least one marker. The detection result for each sensor 105 is stored in the record at the location corresponding to inquiry process 1 of the current inquiry cycle.
[0104] 4. Introduce and incorporate the labeled T nucleotide. Wash away any unbound labeled molecules.
[0105] 5. Inquiry process 2: The characteristics of each of the multiple S sensors 105 (for example, by detecting a signal in each of the multiple S sensors 105) are acquired, and it is determined whether each sensor 105 detected at least one marker. The detection result for each sensor 105 is stored in the record at the location corresponding to inquiry process 2 of the current inquiry cycle.
[0106] 6. Introduce and incorporate the labeled C nucleotide. Wash away any unbound labeled molecules.
[0107] 7. Inquiry step 3: The characteristics of each of the multiple S sensors 105 (for example, by detecting a signal in each of the multiple S sensors 105) are acquired, and it is determined whether each sensor 105 detected at least one marker. The detection result for each sensor 105 is stored in the record at the location corresponding to inquiry step 3 of the current inquiry cycle.
[0108] 8. Introduce and incorporate the labeled G nucleotide. Wash away any unbound labeled molecules.
[0109] 9. Inquiry step 4: The characteristics of each of the multiple S sensors 105 (for example, by detecting a signal in each of the multiple S sensors 105) are acquired, and it is determined whether each sensor 105 detected at least one marker. The detection result for each sensor 105 is stored in the location in the record corresponding to inquiry step 4 of the current inquiry cycle.
[0110] 10. Cleave the labels from the A, T, C, and G nucleotides and wash them away.
[0111] Subsequently, steps 1 through 10 can be repeated for the next query cycle. It should be understood that the specific order of steps 1 through 10 is illustrative, and furthermore, the numbers and numbering of steps 1 through 10 are for convenience only and can be changed. For example, as explained earlier, the order in which nucleotides are introduced is arbitrary. As another example, steps 2, 4, 6, and 8 include the introduction and incorporation of nucleotides, as well as the rinsing of unbound nucleotides, as single steps, but it should be understood that each of steps 2, 4, 6, and 8 can be divided into a series of smaller steps. Similarly, steps 3, 5, 7, and 9 can be further divided into a series of smaller steps (e.g., characterization, determining whether the label was detected, and storing the detection results). Conversely, steps can also be combined (e.g., steps 2 and 3 can be combined, steps 4 and 5 can be combined, etc.).
[0112] In a case where an error is unlikely to occur during any query cycle of an additive approach It should be understood that as soon as the label is detected, it is possible to call (determine) each base of each individual strand. For example, referring to the above steps, for a particular sensor 105, if the obtained characteristic in query step 1, which includes the labeled A nucleotide, indicates that the sensor 105 has detected the label, then saving the detection result may result in calling a base complementary to A(T) for that sensor 105 (and binding site 116). Similarly, for a particular sensor 105, if the obtained characteristic in query step 2, which includes the labeled T nucleotide, indicates that the sensor 105 has detected the label, then saving the detection result may result in calling a base complementary to T(A) for that sensor 105 (and binding site 116). Likewise, for a particular sensor 105, if the obtained characteristic in query step 3, which includes the labeled C nucleotide, indicates that the sensor 105 has detected the label, then saving the detection result may result in calling a base complementary to C(G) for that sensor 105 (and binding site 116). Finally, for a particular sensor 105, if the characteristics obtained in query step 4, which includes the labeled G nucleotide, indicate that the sensor 105 has detected the label, then saving the detection result may result in calling a base complementary to G(C) for that sensor 105 (and binding site 116). However, as will be explained in more detail below, there are several types of errors that can occur during the sequencing procedure (e.g., during an additive approach), and therefore, in some embodiments, a record is made during the sequencing procedure to record the detection / non-detection of the label during each query step in each query cycle. An error correction procedure can then be applied to some or all of the record before calling a base.
[0113] Figure 12 is a flowchart of a sequencing procedure 220 using an additive approach according to several embodiments. The sequencing procedure 220 may be a sequencing procedure performed in step 210 of an exemplary method 200 for sequencing multiple nucleic acid strands (e.g., ssDNA) using the SMAS instrument 100 shown and described in Figure 11. At 222, the sequencing procedure 220 is initiated. At 224, baseline characteristics of each S sensor 105 are obtained (e.g., by at least one processor 130 of the SMAS instrument 100 with the help of circuit 120). As the query cycle begins, at 226, a first labeled nucleotide is selected (e.g., referring to steps 1-10 above, the first labeled nucleotide may be A). At 228, the selected labeled nucleotide is introduced into a fluid chamber 115 and potentially incorporated into the nucleic acid strand to which the nucleotide is bound at the binding site 116. At 230, unbound nucleotides are washed away. In step 232, characteristics are obtained from each of the S sensors, and a detection result (e.g., a detected or undetected label) is determined for each of the S sensors 105. In step 234, the S detection results are recorded in S records (e.g., as 1 indicating that a label was detected, or as 0 indicating that a label was not detected). In step 236, it is determined whether the last tested nucleotide was the last nucleotide of the query cycle. For example of the sequence of nucleotide tests assumed in steps 1-10 above, in step 236 (e.g., by at least one processor 130) it is determined whether G is the last nucleotide tested. Otherwise, in step 238, the next labeled nucleotide to be tested in the query cycle is selected, and steps 228-236 are repeated until it is determined in step 236 that the last tested nucleotide is the last nucleotide of the query cycle. In step 240, the label is cleaved and washed away. In step 242, it is determined (e.g., by at least one processor 130) whether the last completed query cycle is the last query cycle of the sequencing procedure 220.For example, at least one processor 130 can determine whether enough detection results have been recorded to allow at least one processor 130 (or any other processing entity such as an external processor) to call the target number of bases (e.g., 150 bases). If not, the sequencing procedure 220 returns to step 224. If it does, the sequencing procedure 220 returns to step 244. Here again, as explained above. As explained, the order in which the nucleotides are tested is arbitrary.
[0114] In an exemplary case of DNA sequencing, an additive sequencing protocol involving four nucleotide incorporations and one label cleavage reaction is summarized in Figure 13. The leftmost panel of Figure 13 shows a sensor array 110 having a total of 100 individual sensors 105, which are shown as squares. For illustrative purposes, it is assumed that each of the 100 binding sites 116 in the sensor array 110 holds its respective DNA strand, and each DNA strand is sensed by its respective sensor 105 (in other words, there is a one-to-one relationship between binding sites 116 and sensors 105). Some DNA strands may be copies of others. Labeled nucleotides are added to the fluid chamber 115 one type at a time, and the labels are cleaved simultaneously after the nucleotides have been incorporated. If there are no errors, base calling can be achieved after five reactions, i.e., four nucleotide incorporations and one base cleavage reaction. If errors occur, the error correction procedure described below can be applied.
[0115] Subtractive approach In a subtractive approach, the sensor 105 detects nanoscale labels 102 bound to nucleotides having cleavable linkers. All four types of nucleotides have the same type of label (e.g., molecular, fluorescent, magnetic, etc.), but each has a different type of cleavable linker. If no errors are present, four detection results are obtained, one of which, if no errors are present, is a label detection for each of the multiple S nucleic acid strands 101. The query cycle, in one embodiment, includes the following steps:
[0116] 1. The labeled A, T, C, and G nucleotides are introduced simultaneously, the unbound labeled molecules are incorporated, and the mixture is rinsed. The baseline characteristics of each of the S sensors 105 are acquired (for example, by detecting a signal in each of the S sensors 105). If there are no errors, all sensors 105 will detect the label.
[0117] 2. Query step 1: A reagent (e.g., enzyme) is introduced to cleave the label from only the first nucleotide, e.g., A, and rinse, and characteristics (e.g., measuring the signal) are obtained for each of the S sensors 105. It is determined which sensor 105 no longer detects the label (e.g., based on changes in baseline characteristics). The detection results for each sensor 105 are stored in the record at the location corresponding to query step 1 of the current query cycle.
[0118] 3. Query step 2: A reagent is introduced to cleave the label from only the second nucleotide, e.g., T, and rinse, and each of the S sensors 105 is characterized (e.g., the signal is measured). It is determined which sensor 105 no longer detects the label (e.g., based on changes in baseline characteristics). The detection result for each sensor 105 is stored in the record at the location corresponding to query step 2 of the current query cycle.
[0119] 4. Query step 3: A reagent is introduced to cleave the label from a third nucleotide, e.g., only C, and rinse, and each of the S sensors 105 is characterized (e.g., the signal is measured). It is determined which sensor 105 no longer detects the label (e.g., based on changes in baseline characteristics). The detection result for each sensor 105 is stored in the record at the location corresponding to query step 3 of the current query cycle.
[0120] 5. Query step 4: A reagent is introduced to cleave the label from only the fourth nucleotide, e.g., G, and rinse, and characteristics (e.g., measure the signal) are obtained for each of the S sensors 105. It is determined which sensor 105 no longer detects the label (e.g., based on changes in baseline characteristics). The detection results for each sensor 105 are stored in the record at the location corresponding to query step 4 of the current query cycle.
[0121] Steps 1 through 5 can be repeated for the next query cycle. The specific order of steps 1-5 is illustrative, and the numbering and numbering of steps 1-5 are for convenience only and can be changed. For example, as explained earlier, the order in which the nucleotides are cleaved is arbitrary. Similarly, in step 1, the nucleotides can be introduced sequentially (but not necessarily simultaneously). As another example, query steps 1, 2, 3, and 4 include reagent introduction, rinsing, characterization, determination of which sensors no longer detect (or still detect) the label, and result storage as a single step, but each query step can be divided into a series of smaller steps.
[0122] It should be understood that if it is unlikely that an error will occur during any query cycle of the subtractive approach, it is possible to call (determine) each base of each individual strand as soon as label removal (absence of label) is first detected. For example, referring to the above steps, for a particular sensor 105, if the obtained property in query step 1, which includes the labeled A nucleotide, indicates that sensor 105 no longer detects the label, then saving the detection result could mean calling a base complementary to A(T) for that sensor 105 (and binding site 116). Similarly, for a particular sensor 105, if the obtained property in query step 2, which includes the labeled T nucleotide, indicates that sensor 105 no longer detects the label, then saving the detection result could mean calling a base complementary to T(A) for that sensor 105 (and binding site 116). Similarly, for a particular sensor 105, if the properties obtained in query step 3, which includes the labeled C nucleotide, indicate that the sensor 105 no longer detects the label, saving the detection result may result in calling a base complementary to C(G) for that sensor 105 (and binding site 116). Finally, for a particular sensor 105, if the properties obtained in query step 4, which includes the labeled G nucleotide, indicate that the sensor 105 no longer detects the label, saving the detection result may result in calling a base complementary to G(C) for that sensor 105 (and binding site 116). However, as will be explained in more detail below, there are several types of errors that can occur during the sequencing procedure (e.g., during a subtractive approach), and therefore, in some embodiments, a record is made during the sequencing procedure to record the detection / non-detection of the label during each query step in each query cycle. An error correction procedure can then be applied to some or all of the record before calling a base.
[0123] Figure 14 is a flowchart of sequencing procedure 250 using a subtractive approach according to several embodiments. Sequence procedure 250 may be a sequencing procedure performed in step 210 of exemplary method 200, which sequences multiple nucleic acid strands (e.g., ssDNA) using the SMAS instrument 100 shown and described in Figure 11. In 250, the sequence begins in sequencing procedure 252. In 254, all labeled nucleotides are introduced into the fluid chamber 115 and incorporated into nucleic acid strands to which the nucleotides are bound at S binding sites 116. In 256, unbound nucleotides are washed away. In 258, baseline characteristics of each S sensor 105 are obtained (e.g., by at least one processor 130 of the SMAS instrument 100 with the help of circuit 120). Assuming that nucleotides are incorporated into nucleic acid strands to which the nucleotides are bound at each of the S binding sites, the obtained characteristics represent the characteristics of the sensor 105 when detecting at least one label. In step 260, one of the cleavable linkers is selected for cleavage (or, equivalently, one of the nucleotides is selected). In step 262, the label attached to the selected nucleotide is cleaved and rinsed. Assuming no errors, following step 262, a sensor 105 that senses a nucleic acid strand incorporating the tested nucleotide (e.g., the one to which the label was attached by the selected cleavable linker) exhibits a change in properties (e.g., a change in the signal associated with or generated by the sensor 105). In step 264, multiple S sensors are selected. Characteristics are obtained from each of the S sensors 105, and a detection result (e.g., a detected label or a label that was not detected) is determined for each of the S sensors 105. In 266, the S detection results are recorded in S records (e.g., as 1 indicating that a label was detected, or as 0 indicating that a label was not detected). In 268, it is determined whether the last tested nucleotide was the last nucleotide in the query cycle. For example of the sequence of nucleotide tests assumed in steps 1-5 above, in 268 (e.g., by at least one processor 130) it is determined whether G is the last tested nucleotide. Otherwise, in 270, the next cleavable linker (or equivalently, the next nucleotide to be tested) to be cleaved in the query cycle is selected, and steps 262 to 268 are repeated until it is determined in 268 that the last cleaved linker (or equivalently, the last tested nucleotide) is the last linker (or nucleotide) in the query cycle. In step 272, it is determined whether the last completed query cycle (for example, by at least one processor 130) is the last query cycle of the sequencing procedure 250. For example, at least one processor 130 can determine whether at least one processor 130 (or any other processing entity such as an external processor) has recorded enough detection results to allow it to call the target number of bases (e.g., 150 bases). If not, the sequencing procedure 250 returns to step 254. If it does, the sequencing procedure 250 returns to step 274. Again, as described above, the order in which the nucleotides are tested is arbitrary.
[0124] In an exemplary case of DNA sequencing, a subtractive sequencing protocol involving one nucleotide incorporation and four base cleavage reactions is summarized in Figure 15. The leftmost panel of Figure 15 shows a sensor array 110 having a total of 100 individual sensors 105, which are shown as squares. For illustrative purposes, it is assumed that each of the 100 binding sites 116 in the sensor array 110 holds its respective DNA strand, and each DNA strand is sensed by its respective sensor 105 (in other words, there is a one-to-one relationship between the binding sites 116 and the sensors 105). Some of the DNA strands may be copies of others. All four types of labeled nucleotides are added simultaneously to the fluid chamber 115, and the labels are removed after incorporation, removing one type of nucleotide (e.g., a cleavable linker) at a time. If there are no errors, base calling can be achieved after five reactions, i.e., one nucleotide incorporation and four base cleavage reactions. If errors occur, the error correction procedure described below can be applied.
[0125] Modified additive approach In the modified additive approach, sensor 105 detects nanoscale labels 102 bound to nucleotides having cleavable linkers. All four types of nucleotides have the same type of label 102 (e.g., molecular, fluorescent, magnetic, etc.) and use the same type of cleavable linker. The labeled nucleotides are added separately, and the presence of label 102 is detected after the addition of each nucleotide. In the absence of errors, a query cycle that yields four detection results, in which at least one of each of several S nucleic acid strands 101 is a labeled detection, includes the following steps in one embodiment:
[0126] 1. Obtain baseline characteristics for each of the multiple S sensors 105 of the SMAS device 100 (for example, by measuring the baseline signal at each of the multiple S sensors 105) (this may include all or fewer of the sensors 105 in the sensor array 110).
[0127] 2. Introduce and incorporate the first labeled nucleotide, for example, labeled A nucleotide. Wash away any unbound labeled molecules.
[0128] 3. Inquiry process 1: The characteristics of each of the multiple S sensors 105 (for example, by detecting a signal in each of the multiple S sensors 105) are acquired, and it is determined whether each sensor 105 detected at least one marker. The detection result for each sensor 105 is stored in the record at the location corresponding to inquiry process 1 of the current inquiry cycle.
[0129] 4. Cut off the label and wash it away.
[0130] 5. Introduce and incorporate a second labeled nucleotide, such as a labeled T nucleotide. Wash away any unbound labeled molecules.
[0131] 6. Inquiry process 2: The characteristics of each of the multiple S sensors 105 (for example, by detecting a signal in each of the multiple S sensors 105) are acquired, and it is determined whether each sensor 105 detected at least one marker. The detection result for each sensor 105 is stored in the record at the location corresponding to inquiry process 2 of the current inquiry cycle.
[0132] 7. Cut off the label and rinse it away.
[0133] 8. Introduce and incorporate a third labeled nucleotide, such as a labeled C nucleotide. Wash away any unbound labeled molecules.
[0134] 9. Inquiry step 3: The characteristics of each of the multiple S sensors 105 (for example, by detecting a signal in each of the multiple S sensors 105) are acquired, and it is determined whether each sensor 105 detected at least one marker. The detection result for each sensor 105 is stored in the record at the location corresponding to inquiry step 3 of the current inquiry cycle.
[0135] 10. Cut off the label and rinse it away.
[0136] 11. Introduce and incorporate a fourth labeled nucleotide, such as a labeled G nucleotide. Wash away any unbound labeled molecules.
[0137] 12. Inquiry step 4: The characteristics of each of the multiple S sensors 105 (for example, by detecting a signal in each of the multiple S sensors 105) are acquired, and it is determined whether each sensor 105 detected at least one marker. The detection result for each sensor 105 is stored in the record at the location corresponding to inquiry step 4 of the current inquiry cycle.
[0138] 13. Cut off the label and wash it away.
[0139] Subsequently, steps 1 through 13 can be repeated for the next query cycle. It should be understood that the specific order of steps 1 through 13 is illustrative, and furthermore, the numbers and numbering of steps 1 through 13 are for convenience only and can be changed. For example, as explained earlier, the order in which nucleotides are introduced is arbitrary. As another example, steps 2, 5, 8, and 11 include the introduction and incorporation of nucleotides, as well as the rinsing of unbound nucleotides, as single steps, but it should be understood that each of steps 2, 5, 8, and 11 can be divided into a series of smaller steps. Similarly, steps 3, 6, 9, and 12 (query steps 1, 2, 3, and 4) can be further divided into a series of smaller steps (e.g., characterization, determination of whether the label was detected, and storage of the detection results). Conversely, steps can also be combined (e.g., steps 2 and 3 can be combined, steps 3 and 4 can be combined, steps 2-4 can be combined, steps 5 and 6 can be combined, steps 6 and 7 can be combined, steps 5- (For example, 7 can be combined.)
[0140] It should be understood that, given the high probability that no errors will occur during any query cycle of the modified additive approach, it is possible to call (determine) each base of the individual strands as soon as the label is detected. For example, referring to the above steps, for a particular sensor 105, if the obtained property in query step 1, which includes the labeled A nucleotide, indicates that the sensor 105 has detected the label, then saving the detection result could mean calling a base complementary to A(T) for that sensor 105 (and binding site 116). Similarly, for a particular sensor 105, if the obtained property in query step 2, which includes the labeled T nucleotide, indicates that the sensor 105 has detected the label, then saving the detection result could mean calling a base complementary to T(A) for that sensor 105 (and binding site 116). Likewise, for a particular sensor 105, if the obtained property in query step 3, which includes the labeled C nucleotide, indicates that the sensor 105 has detected the label, then saving the detection result could mean calling a base complementary to C(G) for that sensor 105 (and binding site 116). Finally, for a particular sensor 105, if the characteristics obtained in query step 4, which includes the labeled G nucleotide, indicate that the sensor 105 has detected the label, then saving the detection result may result in calling a base complementary to G(C) for that sensor 105 (and binding site 116). However, as will be explained in more detail below, there are several types of errors that can occur during the sequencing procedure (e.g., during an additive approach), and therefore, in some embodiments, a record is made during the sequencing procedure to record the detection / non-detection of the label during each query step in each query cycle. An error correction procedure can then be applied to some or all of the record before calling a base.
[0141] Figure 16 is a flowchart of sequencing procedure 350 using a modified additive approach according to several embodiments. Sequence procedure 350 may be a sequencing procedure performed in step 210 of exemplary method 200 for sequencing multiple nucleic acid strands (e.g., ssDNA) using the SMAS instrument 100 shown and described in Figure 11. In 352, sequencing procedure 350 is initiated. In 354, baseline characteristics of each S sensor 105 are obtained (e.g., by at least one processor 130 of the SMAS instrument 100 with the help of circuit 120). As the query cycle begins, in 356, a first labeled nucleotide is selected (e.g., referring to steps 1-13 above, the first labeled nucleotide may be A). In 358, the selected labeled nucleotide is introduced into the fluid chamber 115 and potentially incorporated into the nucleic acid strand to which the nucleotide is bound at the binding site 116. In 360, unbound nucleotides are washed away. In step 362, characteristics are obtained from each of the S sensors, and a detection result (e.g., a detected or undetected label) is determined for each of the S sensors 105. In step 364, the S detection results are recorded in S records (e.g., as 1 indicating that a label was detected, or as 0 indicating that a label was not detected). In step 366, the labels are cut and washed away. In step 368, it is determined whether the last tested nucleotide was the last nucleotide of the query cycle. For example of the sequence of nucleotide tests assumed in steps 1-13 above, in step 368 (e.g., by at least one processor 130) it is determined whether G is the last nucleotide tested. Otherwise, in step 370, the next labeled nucleotide to be tested in the query cycle is selected, and steps 358-368 are repeated until it is determined in step 368 that the last tested nucleotide is the last nucleotide of the query cycle. In step 372, it is determined (e.g., by at least one processor 130) whether the last completed query cycle is the last query cycle of the sequencing procedure 350.For example, the goal is at least one processor 130 (or some other processing entity such as an external processor). It can be determined whether enough detection results have been recorded to allow the recall of a base number (e.g., 150 bases). If not, sequencing procedure 350 returns to step 354. If so, sequencing procedure 350 returns to step 374. Again, as described above, the order in which the nucleotides are tested is arbitrary.
[0142] In an exemplary case of DNA sequencing, a modified additive sequencing protocol involving four nucleotide incorporations and four base cleavage reactions is shown in Figure 17. The leftmost panel of Figure 17 shows a sensor array 110 having a total of 100 individual sensors 105, which are shown as squares. For illustrative purposes, it is assumed that each of the 100 binding sites 116 in the sensor array 110 holds its respective DNA strand, and each DNA strand is sensed by its respective sensor 105 (in other words, there is a one-to-one relationship between binding sites 116 and sensors 105). Some DNA strands may be copies of others. As illustrated and described, the labeled nucleotides are added to the fluid chamber 115 one type at a time, and the labels are cleaved after incorporation and label detection. In error-free conditions, on average, base calling can be achieved after 5 reactions, i.e., 2.5 nucleotide incorporations and 2.5 base cleavage reactions.
[0143] Therefore, assuming no errors, for DNA sequencing, the modified additive approach produces at least one base call per ssDNA after 8 reactions (4 nucleotide incorporations and 4 base cleavages), testing for all bases. However, on average, base calls occur after only 5 reactions (2.5 nucleotide incorporations and 2.5 base cleavages). Since the label is removed after the introduction of all nucleotides, multiple nucleotides can be incorporated and called during a single A?⇒T?⇒C?⇒G? query cycle. Specifically, in an unknown ssDNA sequence, there is a 1 / 4 probability that the unknown base is T. If the base happens to be T, it is detected in the third step after one incorporation and one base cleavage reaction when the A nucleotide is introduced. There is a 1 / 4 probability that the unknown base is A. If the base happens to be A, it is detected in the fifth step of the A?⇒T? query cycle when the T nucleotide is introduced and two incorporations and two cleavages are performed. There is a 1 / 4 probability that the unknown base is G. If the base happens to be G, it is detected at step 7 of the query cycle A?⇒T?⇒C? after the C nucleotide is introduced and 3 incorporations and 3 cleavages are performed. Finally, there is a 1 / 4 probability that the unknown base is C. If the base happens to be C, it is detected at step 11 of the query cycle A?⇒T?⇒C?⇒G? after the C nucleotide is introduced and 4 incorporations and 4 cleavages are performed. Therefore, on average, 2.5 queries (5 reactions) are required to call a single unknown base.
number
[0144] Cause of sequence determination error Ideally, sequencing procedures, whether using a CLUS instrument or a SMAS instrument 100, should be error-free. In other words, for example, nucleotides should always be properly labeled, nucleotides should always be correctly incorporated into DNA, all labels should be successfully cleaved during the cleavage process, and all cleaved labels should be successfully washed away. However, in reality, errors can occur during any sequencing procedure. This section investigates the causes of sequencing errors in both CLUS and SMAS instruments 100 and describes error mitigation strategies for the SMAS instrument 100. Error correction methods can be used to improve the sequencing accuracy of the SMAS instrument 100, as will be further described below.
[0145] The modified additive approach described above is a conceptually simple sequencing procedure (symmetric in that each nucleotide is processed in the same way), and therefore a good model for explaining how errors propagate in both CLUS and SMAS instruments. Assuming that nanoscale labels are bound to nucleotides via cleavable linkers, four causes of errors are possible. Each error occurs in a ratio denoted as r, having a value between 0 and 1. The four causes of errors are as follows:
[0146] Failed Nucleotide Incorporation (FNI): Failed nucleotide incorporation (FNI) occurs when the labeled nucleotide molecule does not properly reach the ssDNA binding site or when the polymerase is unable to incorporate it. Figure 18A shows FNI in a CLUS instrument sequencing five instances of ssDNA. Following the flow of complementary nucleotides, only three of the five ssDNAs incorporated the labeled nucleotide (shown as having magnetic labeling). Therefore, two of the five nucleotides
number
[0147] Label Removal Failure (FLR): Label removal failure (FLR) occurs when a labeled nucleotide molecule is incorporated, but the cleavage reagent does not reach the linker or is unable to cleave it, so the label is not removed after label detection. Figure 18C shows the FLR of the CLUS instrument described in Figure 18A. After the incorporation of complementary nucleotides and rinsing to remove unbound nucleotides, label detection, and label cleavage and rinsing, one label is one of the ssDNA instances.
number
number
[0148] Nucleotide Removal Failure (FNR): Nucleotide removal failure (FNR) occurs when labeled nucleotides, whether complementary or non-complementary, bind nonspecifically to the surface of binding site 116 and / or sensor 105. Figure 18E shows the FNR of the CLUS apparatus described in the description of Figure 18A. After the nucleotide flow and rinsing to remove unbound nucleotides, two defective nucleotides and their labels remain on the surface of the binding site. Similarly, in Figure 18F, which shows the FNR of the SMAS apparatus 100 described in the description of Figure 18B, after the nucleotide flow and rinsing to remove unbound nucleotides, one defective nucleotide remains on the surface of binding site 116A and another defective nucleotide remains on the surface of binding site 116D. In this example, for both the CLUS apparatus and the SMAS apparatus 100,
number
[0149] Label Detection Failure (FLD): Label detection failure (FLD) occurs when the correct complementary nucleotide is incorporated, but the label is not detected because it is missing or the sensor cannot recognize it. Figure 18G shows an FLD of the CLUS instrument described in Figure 18A. After incorporating the complementary nucleotide and rinsing to remove unbound nucleotides, two of the ssDNA instances have incorporated the complementary nucleotide, but the label is missing.
number
number
[0150] Figures 18A to 18H show the label as a magnet, thereby suggesting a magnetic label and magnetic sensor. However, as explained above, the label may be any type of detectable label (e.g., fluorescent, magnetic, etc.), and the sensor may be any type of sensor capable of detecting selected types of labels (e.g., optical, magnetic, organometallic, charged molecules, etc.).
[0151] The four error types (FNI, FLR, FNR, and FLD) have the same ratio r
number
number
number
[0152] Cluster sequencers vs. single-molecule array sequencers: Qualitative comparison and error correction This specification discloses two types of error correction, called deterministic error correction and stochastic error correction. The SMAS apparatus 100 may use one or both types of error correction, as will be further described below.
[0153] As explained above, the modified additive approach is a good model for illustrating how errors propagate and how the disclosed error correction algorithm can be implemented. It should be understood that the disclosed error mitigation algorithm can also be applied when other sequencing approaches, such as additive or subtractive approaches, are used.
[0154] for example,
number
number
[0155] When using the SMAS device 100, FLR errors can be detected and removed in real time during the sequencing procedure or at some point thereafter. FLR errors can be detected by obtaining the characteristics of each S sensor 105 after cutting and rinsing the labels. FNI errors can be detected by examining the records of each sensor 105 and identifying the query cycle in which that sensor 105 failed to detect the labels. Thus, the modified additive approach can be adjusted to add these detection steps according to one embodiment as follows:
[0156] 1. Obtain baseline characteristics for each of the multiple S sensors 105 of the SMAS device 100 (for example, by measuring the baseline signal at each of the multiple S sensors 105) (this may include all or fewer of the sensors 105 in the sensor array 110).
[0157] 2. Introduce and incorporate the first labeled nucleotide, for example, labeled A nucleotide. Wash away any unbound labeled molecules.
[0158] 3. Inquiry process 1: The characteristics of each of the multiple S sensors 105 (for example, by detecting a signal in each of the multiple S sensors 105) are acquired, and it is determined whether each sensor 105 detected at least one marker. The detection result for each sensor 105 is stored in the record at the location corresponding to inquiry process 1 of the current inquiry cycle.
[0159] 4. Cut off the label and wash it away.
[0160] 5. Determine the characteristics of each of the S sensors 105 that detected the label in step 3. If the characteristics obtained for any of those sensors 105 indicate that the sensor 105 is still detecting the label, then chemistry failed to cut the label (for example, an FLR error exists for that sensor).
[0161] 6. Introduce and incorporate a second labeled nucleotide, such as a labeled T nucleotide. Wash away any unbound labeled molecules.
[0162] 7. Inquiry process 2: The characteristics of each of the multiple S sensors 105 (for example, by detecting a signal in each of the multiple S sensors 105) are acquired, and it is determined whether each sensor 105 detected at least one marker. The detection result for each sensor 105 is stored in the record at the location corresponding to inquiry process 2 of the current inquiry cycle.
[0163] 8. Cut off the label and rinse it away.
[0164] 9. Determine the characteristics of each of the S sensors 105 that detected the label in step 7. If the characteristics obtained for any of those sensors 105 indicate that the sensor 105 is still detecting the label, then chemistry failed to decrypt the label (for example, an FLR error exists for that sensor).
[0165] 10. Introduce and incorporate a third labeled nucleotide, such as a labeled C nucleotide. Wash away any unbound labeled molecules.
[0166] 11. Inquiry step 3: The characteristics of each of the multiple S sensors 105 (for example, by detecting a signal in each of the multiple S sensors 105) are acquired, and it is determined whether each sensor 105 detected at least one marker. The detection result for each sensor 105 is stored in the record at the location corresponding to inquiry step 3 of the current inquiry cycle.
[0167] 12. Cut off the label and rinse it away.
[0168] 13. Determine the characteristics of each of the S sensors 105 that detected the label in step 11. If the characteristics obtained for any of those sensors 105 indicate that the sensor 105 is still detecting the label, then chemistry failed to cut the label (for example, an FLR error exists for that sensor).
[0169] 14. Introduce and incorporate a fourth labeled nucleotide, such as a labeled G nucleotide. Wash away any unbound labeled molecules.
[0170] 15. Query step 4: The characteristics of each of the multiple S sensors 105 (e.g., by detecting a signal in each of the multiple S sensors 105) are obtained, and it is determined whether each sensor 105 detected at least one label. The detection result for each sensor 105 is stored at the location in the record corresponding to query step 4 of the current query cycle. If there are sensors 105 that do not have a base assigned for the query cycle (e.g., sensors 105 that could not detect A, T, C, or G during the query cycle), then chemistry was unable to incorporate the nucleotide (e.g., FNI exists for these sensors 105).
[0171] 16. Cut off the label and wash it away.
[0172] 17. For each of the S sensors 105 that detected the label in step 15, the characteristics are as follows: Desired. If the characteristics obtained for any of those sensors 105 indicate that the sensor 105 is still detecting the label, the chemistry was unable to cleave the label (e.g., for that sensor, there is a FLR error).
[0173] Then, for the next interrogation cycle (e.g., if the previous interrogation cycle failed to read the current base, to estimate the next base, or to reread the current base), steps 1 to 17 can be repeated. It should be understood that the specific order of steps 1 to 17 is exemplary, and further, the number and numbering of steps 1 to 17 are for convenience and can be changed. As an example, as previously described, the order in which nucleotides are introduced is arbitrary. As another example, steps 2, 6, 10, and 14 include nucleotide introduction and incorporation, and flushing of unbound nucleotides as a single step, but it should be understood that each of steps 2, 6, 10, and 14 can be divided into a series of smaller steps. Similarly, steps 3, 7, 11, and 15 (interrogation steps 1, 2, 3, and 4) can be further divided into a series of smaller steps (e.g., obtaining characteristics, determining whether a label has been detected, and saving the detection result). Similarly, label 15 includes identifying an FNI error, but that task can be made a separate step. Conversely, steps can also be combined (e.g., some or all of steps 2 - 5, some or all of steps 6 - 9, some or all of steps 10 - 13, some or all of steps 14 - 17, etc.).
[0174] Figure 19 is a flowchart of an exemplary sequencing procedure 400 using a modified additive approach with FLR and FNI error detection, according to several embodiments. Sequencing procedure 400 may be a sequencing procedure performed in step 210 of an exemplary method 200 for sequencing multiple nucleic acid strands (e.g., ssDNA) using a SMAS instrument 100 as shown and described in Figure 11. Sequencing procedure 400 starts in 402. In 404, baseline characteristics of each S sensor 105 are obtained (e.g., by at least one processor 130 of the SMAS instrument 100 with the help of circuit 120). As the query cycle begins, in 406, a first labeled nucleotide is selected (e.g., referring to steps 1-17 above, the first labeled nucleotide may be A). In 408, the selected labeled nucleotide is introduced into a fluid chamber 115 and potentially incorporated into the nucleic acid strand to which the nucleotide is bound at the binding site 116. In 410, unbound nucleotides are washed away. In step 412, characteristics are obtained from each of the S sensors, and a detection result (e.g., a detected label or a label that was not detected) is determined for each of the S sensors 105. In step 414, the S detection results are recorded in S records (e.g., as 1 indicating that a label was detected, or as 0 indicating that a label was not detected). In step 416, the labels are cut and washed away. In step 418, characteristics are obtained for the sensors 105 that detected labels during steps 412 / 414. In step 420, it is determined whether any of the sensors 105 that detected labels during steps 412 / 414 are still detecting labels. If so, in step 422, it is determined that an FLR error has been detected for the sensors 105 that are still detecting at least one label, even though the labels were cut and washed away in step 416. The sequence determination procedure 400 then continues to step 424. If, in step 420, it is determined (for example, by at least one processor 130) that none of the sensors 105 that detected a label during steps 412 / 414 have yet detected a label, the sequence determination procedure also proceeds to step 424.At 424, it is determined whether the last nucleotide tested was the last nucleotide of the interrogation cycle. For an example of the order of nucleotide testing assumed in the above steps 1-17, at 368 (e.g., by at least one processor 130), it is determined whether G was the last nucleotide tested. Otherwise, at 426, the next labeled nucleotide to be tested in the interrogation cycle is selected, and at 424 it is determined whether the last nucleotide tested was the last nucleotide of the interrogation cycle. Steps 408-420 (and 422 if applicable) are repeated until it is determined at 424. At 428, FNI errors of S sensors 105 in which no label could be detected during the last completed interrogation cycle are detected. At 430, it is determined (e.g., by at least one processor 130) whether the last completed interrogation cycle was the last interrogation cycle of the sequencing procedure 400. For example, at least one processor 130 can determine whether sufficient detection results have been recorded to enable at least one processor 130 (or some other processing entity such as an external processor) to call a target number of bases (e.g., 150 bases). If not, the sequencing procedure 400 returns to step 404. If so, the sequencing procedure 400 ends at 432. Here too, as explained above, the order in which nucleotides are tested is arbitrary.
[0175] Reduction of FNI and FLR Errors To illustrate the impact of FNI and FLR errors on the CLUS device and the SMAS device 100, for each type of sequencer, an exemplary DNA sequence is called in which FNI and FLR errors occur randomly when the sequence is read using the modified additive approach of the above SBS. The error rate is for both FNI and FLR errors
Number
number
[0176] Figure 21 shows the predicted signal levels detected by the CLUS instrument sensor, which captures the behavior of the molecular ensemble during the sequencing procedure. In each query step, the CLUS instrument sensor can detect four signal intensity levels for the molecular ensemble (consisting of three ssDNAs): 0 labels, 1 label, 2 labels, or 3 labels. The CLUS instrument sequencing procedure considers the binding signal of the ensemble and cannot distinguish when the reaction for individual strands fails. A base is called in a particular query step whenever the CLUS instrument sensor senses at least two labels. This threshold can be expressed by a criterion, where a base is called when the CLUS sensor signal level is greater than 1.5. As Figure 21 shows, a high rate of chemical failures leads to significant base calling errors and very low base calling accuracy. The CLUS instrument approach yields only 6 out of 21 bases (approximately 29%) that are called according to the true sequence. This level of accuracy is only slightly better than random guessing with 25% accuracy (since, for 4 bases, the probability of correctly guessing a base is 1 in 4). Furthermore, the CLUS instrument cannot distinguish between successful and failed chemical reactions, nor can it indicate the error location of FNI (dashed circle) or FLR (inverted diagonal circle) as shown in Figure 20. In the case of the CLUS instrument, the exact location of the FLR error becomes obscured by ensemble averaging. To slightly improve the quality of base calling of the CLUS instrument... Therefore, only probabilistic error correction algorithms can be implemented by essentially making empirically based inferences about the locations of base insertions, deletions, and substitution sites. An exemplary algorithm is, for example, A. Cacho et al., "A Comparison." of Base-calling Algorithms for Illumina This is described in "Sequencing Technology," Briefings in Bioinformatics, Vol.17(5), 786-795, 2016; WCKao et al., "BayesCall: A model-based base-calling algorithm for high-throughput short-read sequencing," Genome Res., Vol.19(10), 1884-1895, 2009; and C. Ledergerber and C. Dessimoz, "Base-calling for next-generation sequencing platforms," Brief Bioinform., Vol.12, 489-97, 2011.
[0177] Figure 22 illustrates how the SMAS instrument 100 can provide better accuracy when using the error correction techniques described herein. As explained above, FLR errors occurring during the sequencing procedure can be detected during the sequencing procedure. Specifically, the SMAS instrument 100 knows (or can find) the location of the FLR because the characteristics of each sensor 105 (e.g., signal level) are obtained and recorded after the label has been cut and washed away, and before the next nucleotide is introduced. The FLR error can be corrected by treating it as "label not detected" when making a base call. In other words, if the recording of the sequencing procedure contains binary (e.g., 0 / 1) entries for each query step, the FLR can be corrected by changing the values in those query steps from "detected" to "not detected". As a specific example, if 0 represents that the label was not detected and 1 represents that the label was detected, then before error correction, the FLR in the m-th query step is represented as 1 at the m-th position in the recording. That error can be corrected by changing the value of 1 at the m-th position in the recording to a value of 0. The top of Figure 22 shows the detection results of each of the three sensors 105 of the SMAS instrument 100 before error correction to remove the FLR error. The bottom of Figure 22 shows the results after correcting the FLR error before calling the base.
[0178] The modified additive sequencing procedure using the SMAS device 100 involves more than half of the K sensors 105 (
number
[0179] When using the SMAS device 100, the built-in failure creates a characteristic signature in the detection result of the SMAS sensor 105, so FNI errors can also be corrected (for example). (In the record consisting of the detection / non-detection of the label by sensor 105 during the sequencing procedure). In particular, FNI errors in the modified additive approach result in runs (sequential sequences) of zero (or other "no label detection" detection results) for four or more consecutive query steps. As explained in the description of Figure 19, some FNI errors can be detected by identifying that a particular sensor 105 did not detect a label during a query cycle. It should also be understood that FNI errors can "span" multiple query cycles. For example, in a first query cycle including the A?⇒T?⇒C?⇒G? query steps, if a particular sensor 105 detects a label during the A? query step, it will not detect any labels until the C? query step of the next query cycle. The C? query step follows the A? query step in the exemplary query cycle, and since the modified additive approach is used in the sequencing cycle, a label should be detected as a result of the C? query step in the first query cycle. It should be noted that step 428 in Figure 19 does not result in no label being detected by the specific sensor 105 during either the first or second query cycle, thus ensuring that no FNI error is detected during either the first or second query cycle. However, examination of the detection results record reveals the presence of an FNI error. The FNI error can be definitively corrected by removing zero (4 in the case of DNA sequencing) runs to align the bad strand with the strand unaffected by the FNI error. Figure 23 shows the correction of the FNI error by removing four "no label detected" entries from the detection results record from the sequencing procedure. As shown in Figure 23, the FNI error correction results in a perfect alignment between the calling sequence and the true sequence.
[0180] A qualitative analysis of a simplified model system with a limited set of errors suggests that the use of the SMAS instrument 100 for nucleic acid sequencing is far superior to that of the CLUS instrument, at least when the number of instance Ks of the sequenced DNA strand is small and the chemical failure rate is high. To establish a framework for a quantitative comparison of the two platforms, we examine below how cluster size (for the CLUS instrument) and the number of instances sequenced (for the SMAS instrument 100) affect base calling accuracy. For both FNI and FLR errors,
number
number
[0181] Figure 25 shows the effect of a large cluster size N on the base calling accuracy of the CLUS device. Figure 25 shows the predicted signal levels detected by the CLUS device sensor that captures the behavior of the molecular ensemble during the sequencing procedure. In each interrogation step, the CLUS device sensor can detect 12 signal intensity levels (11 ssDNAs) of the molecular ensemble, i.e., any of the detected 0 to 11 labels. When the signal level detected by the CLUS sensor is greater than 5.5, a base is called in a particular interrogation step. As shown in Figure 25, failed chemistry results in base calling errors, and only 11 out of 18 called bases (about 61%) follow the true sequence.
[0182] Comparing Figure 25 with Figure 21,
Number
Number
Number
Number
[0183] Figure 26 shows, according to some embodiments,
Number
number
[0184] Therefore, when the SMAS instrument 100 is used with deterministic error correction, a perfect match between the true sequence and the called sequence can be obtained if only FNI and FLR errors occur. Furthermore, if only FNI and FLR errors occur, it is actually possible to call an error-free sequence using only a single sensor 105, reading a single ssDNA along with the deterministic error correction techniques described above (e.g., changing FLR to "no label detected" and / or removing a specified length (e.g., 4) of "no label detected" runs from the detection result record).
[0185] However, when FNR and / or FDL errors are introduced, using only deterministic error correction is generally unlikely to eliminate all errors in the detection results record. To address FNR and / or FDL errors, probabilistic error correction can be included in addition to, or instead of, deterministic error correction.
[0186] Mitigation of FNI, FLR, and FNR errors This section further includes FNR errors in the analysis. CLUS instrument base calling. The impact of such errors on accuracy is equivalent to that of FNI and FLR due to the averaging inherent in the detection of labels in clusters of nucleic acid instances in the CLUS instrument. FNR errors are considerably detrimental to the performance of sequencing methods using the SMAS instrument 100 because they cannot be definitively corrected. (Note that FNR errors cannot be corrected at all by themselves in the CLUS instrument; instead, the CLUS instrument relies on ensemble behavior to mitigate the effects of FLR and other types of errors.)
[0187] Figure 27 illustrates the problem introduced by FNR into the exemplary sequence (TAG CAA GGT CCG CTA CTG GCA GAC TGG), assuming that FNI, FLR errors, and here FNR errors also occur randomly during the 18 query cycles of the A?⇒T?⇒C?⇒G? query process. For the sake of the example,
number
number
[0188] By applying probabilistic error correction, error correction can be improved to mitigate FNR errors in addition to FLR and FNI errors. For example, consider the thymine query process at position 2 (query process 2 of query cycle 1). Sensors S1 and S3 detect the marker, but S2 does not. S2 does not detect the marker because an FNR error occurred simultaneously in both sensors S1 and S3, or because an FNI error occurred in sensor S2. If the probability of each error is r, the probability that an FNR error occurred simultaneously in both sensors S1 and S3 is r. 2The probability of an FNI error in sensor S2 is r. An error correction algorithm (e.g., executed by at least one processor 130 or another processor) assumes that a more likely event occurred (there was an FNI error in sensor S2) and shifts the S2 detection results in the S2 recording by deleting all entries for positions 2-5 from the data recording that captures the detection results from sensor S2. As a result, the detection results in the S2 recording are realigned with the detection results generated by sensors S1 and S3, as shown at the top of Figure 30, labeled "A". Position 4 ("A" in Figure 30) The previous (before deletion) G-label detection in the section indicated as shown may be due to FNR, as sensors S1 and S3 did not detect the label at position 4 (inquiry step 4 of inquiry cycle 1).
[0189] As shown in the section labeled "E" in Figure 30, the same error correction procedure can be performed from left to right at positions 13 (labeled "B" as shown in the section of Figure 30), 32 (labeled "C"), and 46 (labeled "D") to demonstrate a gradual improvement in the alignment between the S1, S2, and S3 recordings of the detection results. The section labeled "E" in Figure 30 shows that while performing multiple probabilistic error correction steps aligns the outputs of all sensors S1, S2, and S3, it does not appear to improve the alignment between the called sequences and the true sequences. Even after error correction, only 9 out of 20 (45%) of the bases are correctly called. In other words, base calling errors still occur. Specifically, following the error correction procedure, all three sensors S1, S2, and S3 report the detected labels in the query step where the labels are detected, but some sensors also detect labels incorrectly incorporated by the FNR at positions 10, 22, 40, and 50 (as shown in the continuation of Figure 30).
[0190] When more than half of the sensors 105 call the bases when their detection results match (after error correction), a thymine insertion error occurs at sequence position 8 (query step 22), and both sensors S1 and S3 detect the label bound to a non-complementary nucleotide during the same query step. (It should be understood that we know there is a thymine insertion error at position 8 because the erroneous data is created for illustrative purposes and is known. In one embodiment, sensor 105 only indicates whether the label was detected during the query step, not whether the detection (or lack thereof) is correct or incorrect. Therefore, in one embodiment, the error in query step 22 is essentially indistinguishable from a correct detection result.) The properly aligned true sequence and the called sequence clearly show the location of a single erroneous base insertion and can be represented as follows: Error: | Insert True sequence: TAG CAA G*G TCC GCT ACT GGC Called array: TAG CAA GTG TCC GCT ACT GGC * Insertion position
[0191] This insertion error can be corrected if the base calling rule is modified to require that all three sensors S1, S2, and S3 match. Under such a rule, all three sensors S1, S2, and S3 must simultaneously suffer an FNR error to cause an incorrect base calling. The probability of such an event is r 3 That is all.
number
[0192] Mitigation of FNI, FLR, FNR, and FLD errors Common error correction strategies used in some embodiments consider and mitigate all four types of chemical errors that cause FNI errors, FLR errors, FNR errors, and FLD errors. Figure 31 shows an exemplary sequence (TAG CAA GGT CCG CTA CTG GCA GAC TGG) and illustrates the FNI, FLR, FNR errors, and here. Let's assume that FLD errors also occur randomly during the 18 query cycles of the A?⇒T?⇒C?⇒G? query process. To provide a vehicle for illustrating an exemplary error correction procedure, we will create many errors in the sequencing data, with a very high average error rate of 1 in 5 failed reactions.
number
[0193] Under the exemplary conditions and assumptions made herein, considering only the data records created by SBS using the SMAS instrument 100, it is not possible to distinguish between correct nucleotide incorporation and FNR, nor between correct nucleotide non-incorporation and FNI. While FLR errors can be deterministically detected and corrected as previously described (by checking sensor 105 after cutting and washing away the label and treating FLR as "no label detected"), FNR errors cannot be identified because they are indistinguishable from correct detection events, and FNI and FLD errors cannot be identified because they are indistinguishable from correct nucleotide non-incorporation. Nevertheless, error mitigation can still be achieved using probabilistic error correction techniques. For example, as described above, if fewer sensors S1, S2, and S3 than all sensors S1, S2, and S3 detect or do not detect the label during a particular query step, the probabilities of two (or more) events can be calculated, the event with the highest probability can be assumed to be correct, and an appropriate error correction step can be performed.
[0194] Figure 32 illustrates the application of the error correction procedure to the data captured during SBS under the above conditions and assumptions. The portion of Figure 32 labeled "A" is the raw data before FLR error removal. Assuming that the signal level of sensor 105 is checked after the label is cleaved and rinsed off, as described above, the location of the FLR error is known. The FLR error can be completely eliminated using deterministic error correction, i.e., by changing the "label detected" value (e.g., 1 or "yes") in the data recording at the location corresponding to the query step where the FLR error was detected to a "label not detected" value (e.g., 0 or "no"). Note that during the query cycle 15 shown in Figure 31, an FLR error occurs in the data for sensor S2 following an FLD error. In other words, sensor S2 was unable to detect the label of the incorporated nucleotide during the first query step of the 15th cycle. After the first query step of the 15th cycle and before the second query step of the 15th cycle, if the label is cleaved, the signal level of sensor S2 is confirmed. This check reveals the presence of a label in sensor S2, which is known to be an FLR error because all labels should have been cut off and rinsed away after the last query step. Therefore, even if an FLR error follows another error, it can be detected and eliminated.
[0195] The section labeled "B" in Figure 32 shows the detection results after the removal of FLR errors by deterministic error correction, as applied as described above. The data record shown in "B" is shown here.
number
[0196] To illustrate how probabilistic error correction can be applied, the following table shows the data recordings from Figure 32 for the first five query cycles (query steps 1-20) of the three sensors S1, S2, and S3 after the FLR error has been removed (e.g., from the recording labeled "B" in Figure 32). In other words, the following table shows the first 20 detection results after deterministic error correction to remove the FLR error. The table contains values of 1 for query steps in which the sensor detected the label, and values of 0 for query cycles in which the sensor did not detect the label. [Table 2]
[0197] As explained above, simple majority voting after removing FLR errors results in only 8 out of 17 bases being correctly called, as shown in the portion labeled "B" in Figure 32. Probabilistic error correction can provide a significant improvement, as will be discussed later.
[0198] Taking query step 2 as an example, both sensors S1 and S3 detected the label (entry 1s in the table above), but sensor S2 did not (entry 0 in the table). Therefore, either both sensors S1 and S3 are wrong, or sensor S2 is wrong. By taking into account the probabilities of various events that could lead to each of these results, the error correction algorithm can mitigate errors in the sequence determination data. Specifically, since FLR is removed from the data recording, the only way for both sensors S1 and S3 to misdetect the label during query step 2 is if both suffer an FNR error during that query step. If the probability of an FNR error is r, the probability that both sensors S1 and S3 suffer an FNR error during a single query step is r 2 Therefore, for the purpose of this example,
number
[0199] If sensor S2 is incorrect, it is because sensor S2 failed to detect the label due to either an FLD error or an FNI error. Recall that an FLD error occurs when the correct complementary nucleotide is incorporated but lacks the label or the sensor cannot detect the label, and an FNI error occurs when the correct complementary nucleotide is not incorporated at all during the sequencing cycle. FLD and FNI errors are mutually exclusive (i.e., a sensor can only suffer one of them at a time, and never both). Therefore, if the probability of each type of error is r, the probability that sensor S2 suffered either an FLD error or an FNI error is 2r. In this example,
number
number
[0200] As explained above, sensor S2 may be wrong due to either an FLD error or an FNI error. Following an FLD error, the DNA strand sensed by sensor S2 remains "synchronized" or "aligned" with the DNA strands sensed by sensors S1 and S3. In other words, if query step m sequences the 40th base of the DNA strands sensed by each of sensors S1, S2, and S3, then query step m
number
[0201] In some embodiments, the actions taken by the error correction algorithm rely in part on examining candidate error correction data, each separately assuming that one of two types of errors has occurred. In other words, the detection result record can be modified to correct an error, assuming it was caused by an FLD error to generate a first candidate correction data record, and the detection result record can be modified separately to correct an error, assuming it was caused by an FNI error to generate a second candidate correction data record. The two candidate correction data records can then be examined and / or analyzed and / or compared to determine which is more likely to be correct. To correct an FLD error, the “sign not detected” indication is reversed to the “sign detected” indication. To correct an FNI error, the data entry is shifted by only four locations (for example, to the left, as the data record is presented in the example herein).
[0202] To illustrate a specific example of query step 2 in an exemplary data recording, option A, the first candidate corrected data recording, affects the output of sensor S2, assuming the (estimated) error was an FLD error. The estimated error is corrected by inverting the bit for query step 2 in the recording of sensor S2 from 0 to 1 with the bold, underlined value "1", as shown in the table for option A below. [Table 3]
[0203] Option B, the second candidate corrected data record, assumes that the error affecting the output of sensor S2 was an FNI error. This estimated error is corrected by deleting the data recorded between query steps 2, 3, 4, and 5 from the sensor S2 data entry, thereby "resynchronizing" or "realigning" the data record corresponding to sensor S2 with the data records of sensors S1 and S3, resulting in the following table (the values in previous locations 21-24 are shifted to locations 17-20). Table entries for Option B corrected by the error correction algorithm are indicated by a type in bold and underlined. [Table 4]
[0204] Next, options A and B can be compared and / or analyzed to determine which is more likely to be correct, and in some cases one of the options can be discarded. For example, a processor (e.g., at least one processor 130 or another) can determine the value of a metric for each candidate correction data record and, at least in part, determine which of options A and B is more likely to be correct based on a comparison of the metrics. An example of a metric is the number of query steps starting from the current query step that has now been corrected, where query step J is located further away in the data record than the three (or more generally, K) sensor signature detection results match. For example, using this metric and setting the value of J to 8, the metric value for option A is 3 and the metric value for option B is 6. In some embodiments, based solely on this result, it is assumed that option B is more likely to be correct and option A is discarded because the metric value for option B is significantly larger than the metric value for option A. In some embodiments, one of the two options is discarded only if the value of its metric exceeds a certain threshold (e.g., a percentage, a quantity (e.g., at least twice, at least 1.5 times, etc.)) of the value of the metric of the other option. In some embodiments, option A is retained and the option is not discarded until later.
[0205] In some embodiments, contributions to the metric value are weighted based on the distance of the data being considered from the currently corrected query step. For example, since the likelihood of an additional error introduced into the data recording increases as more bases are sequenced (e.g., the likelihood of some error occurring for one of the K sensors between query step 3 and query step 40 is greater than the likelihood of some error occurring for one of the K sensors between query step 3 and query step 6), the metric can assume that closer data entries are more likely to be correct than further away data entries, and therefore more weight is given to data entries closer to the currently corrected data entry than to further away data entries. The weighting is, for example, The metrics may be linear or nonlinear. As just one example, for a metric with contributions from data far from up to 12 query steps, the contribution from query steps within the four currently corrected data may be weighted 1, the contribution from query steps between the fifth and eighth currently corrected data may be weighted 0.5, and the contribution from query steps between the ninth and twelfth currently corrected data may be weighted 0.2. Many possible metrics can be used, with or without weighting, and it should be understood that those provided above are merely examples and not intended to be limiting.
[0206] The above metrics use the number of query steps starting after the currently corrected query step and the furthest query step J position in the matching data record where the label detection results of all three (or more generally, K) sensors match. However, it should be understood that they can also be used equivalently to the number of query steps starting after the currently corrected query step and the furthest query step J position in the data record where the label detection results of all three (or more generally, K) sensors do not match. In this case, a larger metric value indicates more mismatches between sensor data entries, and therefore, a lower metric value indicates a higher probability of a candidate corrected data record being correct. As will be apparent to those skilled in the art, adjustments can be made to any weighting applied.
[0207] It should also be understood that it is not necessary to discard one of the possible options after correcting the estimated error in the data recording. For example, following the (estimated) correction of the (estimated) error in query step 2 in the recording of sensor S2, both options A and B can be retained, and furthermore, error detection and correction can be performed in parallel for both. Similarly, each time an estimated error is corrected, multiple options of candidate sequences can be determined and / or evaluated / compared. Running metric values can be maintained for each possible option / candidate sequence at each step of the error correction procedure, and the most likely candidate sequence can be determined at some point (e.g., after all candidate options have been determined and evaluated (e.g., against each other), or after a number of additional query steps, etc.).
[0208] Furthermore, in the example above, the possibility that both sensors S1 and S3 misdetected the label was immediately discarded because (given the assumptions herein) the probability of that event is significantly lower than the probability that sensor S2 is wrong, but instead the same procedure can be followed for sensor S2. In other words, option C in query step 2 can be determined by assuming that both sensors S1 and S3 suffer an FNR error and sensor S2 is correct. In this case, the metrics can be adjusted to account for the likelihood of various possible outcomes (for example, by “penalizing” the metrics of option C based on the probability that both sensors S1 and S3 suffer an FNR error (e.g., multiplying the metrics by the ratio of the probability that both sensors S1 and S3 are wrong to the probability that sensor S2 is wrong)).
[0209] It should be understood that the error correction methodologies described herein can be utilized in several ways to improve the accuracy of nucleic acid sequencing using the SMAS instrument 100. Assuming sufficient computational power, an implementation (e.g., using at least one processor 130, or another or more processors) can determine and evaluate an exhaustive set of candidate sequences to which error correction has been applied, and then select from among them the candidate sequence that is most likely to be correct. To reduce computational complexity, an implementation can also make a decision during the error correction process to eliminate candidate error-corrected sequences (or potential error sources) that are considered to have a low probability of being correct (e.g., option C in the above example), and retain only the candidate error-corrected sequences that are more likely to be correct. It should be understood that the flexibility of these principles makes them suitable for error mitigation in systems with a wide variety of computing powers. Returning to the example above, assuming that option B was the only option retained after error correction was applied to the data from query step 2, the corrected data would appear as follows: [Table 5]
[0210] The next query step where the three sensors S1, S2, and S3 do not match is query step 5. Here again, sensor S2 does not match sensors S1 and S3 in the same way as in query step 2. In some embodiments, the error correction algorithm determines that (a) the probability that sensor S2 is wrong is greater than the probability that both sensors S1 and S3 are wrong, and (b) in query step 5, sensor S2 has suffered either an FNI error or an FLD error. Here again, two options can be made: one assumes the error is an FLD error (corrected by inverting bits), and the other assumes the error is an FNI (corrected by shifting the data by four positions). The corrected data record is shown below: Option A (Estimated FLD error corrected): [Table 6] Option B (Estimated FNI error corrected): [Table 7]
[0211] Here, the metrics for options A and B may be calculated, one of the options may be discarded, or both may be retained. As an example, assume that option A is retained and the following error-corrected data is obtained. [Table 8]
[0212] The next query step where the sensor data does not match is query step 10. Here, sensor S1 detected the label, but neither sensor S2 nor sensor S3 detected it. Since FLR errors have been removed from the data recording, the only way sensor S1 could have misdetected the label during query step 10 is if it suffered an FNR error during that query step. The probability of an FNR error is r. If both sensors S2 and S3 are wrong, it is because (a) both suffered an FNI error, (b) both suffered an FLD error, or (c) one suffered an FNI error and the other suffered an FLD error. The probability of any of the mutually exclusive events (a), (b), or (c) is 4r. 2 Therefore, in some embodiments, it is assumed that a more likely event occurred, namely that sensor S1 suffered an FNR error (assumed value of r).
number
[0213] The error correction procedure can be continued as described throughout the rest of the data recording. The portion labeled "C" in Figure 32 shows the result of the example. As shown, following the application of probabilistic error correction as described above, 16 of the 20 bases (80%) are correctly identified.
[0214] Figure 33 is a flowchart showing an error correction procedure 450 in several embodiments. The error correction procedure 450 may be, for example, the error correction procedure 212 shown in Figure 11, or it may be performed by a processor (for example, at least one processor 130 shown in Figure 5A or Figure 50, described below). In 452, the error correction procedure 450 begins. In 454, multiple records are identified in the sequencing data generated as a result of a nucleic acid sequencing procedure using the SMAS instrument 100. Each of the identified multiple records contains multiple entries, each capturing a detection result for one instance of a particular strand of nucleic acid. Thus, if the number of identified records is K, each of the K records contains one entry for each detection result per query step of the sequencing procedure. Each detection result indicates that during the query step, either (a) the label was detected by the corresponding sensor 105 or (b) the label was not detected by the corresponding sensor 105. Multiple records can be identified in several ways. For example, as will be further explained below, different unique barcodes can be ligated to the primer ends of nucleic acid strands to read known sequences during the sequencing procedure cycle. Therefore, multiple records can be identified by searching the sequencing data for barcodes associated with a particular strand of nucleic acid. As another example, common sequences of entries can be identified in the sequence data (for example, within entries that document the detection results of the first approximately 35 query steps of the sequencing procedure).
[0215] In step 456, based on multiple records, multiple candidate sequences are determined for a specific strand of nucleic acid. Each of the multiple candidate sequences is at least a portion of the nucleic acid sequence of the specific strand of nucleic acid (for example) In some embodiments, determining multiple candidate sequences involves identifying specific query steps in multiple records where a first sensor detects each label and a second sensor detects none of the labels, and establishing two candidate sequences, one of which assumes that the first sensor correctly detected each label, and the other assumes that the first sensor incorrectly detected each label. In some embodiments, determining multiple candidate sequences involves identifying specific query steps in multiple records where a first sensor detects each label and a second sensor detects none of the labels, and establishing two candidate sequences, one of which assumes that the second sensor did not incorrectly detect any of the labels, and the other assumes that the second sensor did not correctly detect any of the labels. In some embodiments, determining multiple candidate sequences involves identifying a series of consecutive entries (e.g., four entries) in at least one of the multiple records indicating that no labels were detected, and removing a series of consecutive entries from at least one of the multiple records indicating that no labels were detected. In some embodiments, each of a plurality of entries is either a first binary value (indicating that a marker was detected) or a second binary value (indicating that no marker was detected), and determining a plurality of candidate sequences includes identifying a run of the second binary value in at least one of the plurality of records (e.g., 4) and removing a sequence of the second binary value from at least one of the plurality of records.
[0216] In 458, a specific candidate sequence among multiple candidate nucleic acid sequences is identified as the sequence most likely to be correct from among the multiple candidate sequences. In some embodiments, identifying a specific candidate sequence most likely to be correct from among multiple candidate sequences includes determining or estimating which of the multiple candidate sequences is most likely to be correct. In some embodiments, identifying a specific candidate sequence most likely to be correct from among multiple candidate sequences includes determining a metric for each of the candidate sequences and selecting a specific candidate sequence as most likely to be correct, at least in part, based on the respective metrics and criteria (e.g., minimum likelihood of occurrence, threshold likelihood of occurrence). In some embodiments, identifying a specific candidate sequence most likely to be correct from among multiple candidate sequences includes identifying a number of results for a particular query step represented by multiple records (e.g., more than half of the sensors 105 detected the label, or more than half of the sensors 105 did not detect the label). In some embodiments, identifying a particular candidate sequence from among several candidate sequences that are most likely to be correct involves determining the likelihood of occurrence for each of the candidate sequences and selecting a particular candidate sequence based on that likelihood of occurrence satisfying constraints (e.g., minimum probability). In some embodiments, the particular candidate sequence that is most likely to occur among the candidate sequences is identified as the most likely to be correct. In some embodiments, one or more candidate sequences are eliminated based on known constraints, such as the knowledge that a particular sequence of bases is impossible. For example, it may be known from the origin or source of the nucleic acid (e.g., human) that a particular sequence of bases is impossible, and therefore, candidate sequences having such impossible sequences can be eliminated from further consideration.
[0217] At step 460, the error correction procedure 450 is completed.
[0218] It should be understood that probabilistic error correction is only successful if the most likely identified scenario (e.g., the identification in 458 in Figure 33) is indeed the correct scenario. In cases of high chemical failure rates, as in the examples described herein, there may be multiple scenarios with comparable probability of occurrence (or nearly equal probability of occurrence), in which case more sophisticated bioinformatics tools can be used. For example, candidate sequences can be based on knowledge of the source of the nucleic acid being sequenced (e.g., considering the source / origin of the nucleic acid). Considering this, certain sequences of bases may be removed (based on the knowledge that such sequences are impossible). Nevertheless, if performed correctly as described herein, the error correction process results in the correct alignment of the sensor 105 output. In the example shown in Figure 32, after the removal of FNI and FLR, all three sensors S1, S2, and S3 report the label in the correct detection query steps where the label is detected, but the sensors do not match at many query positions (5, 10, 13, 20, 22, 27, 32, 40, 41, 48, and 50) where the label is either incorrectly incorporated by FNR or cannot be detected due to FLD. When calling the bases when more than half of sensor 105 matches in the aligned sequences, a thymine insertion at sequence position 8 (query step 22) and a guanine deletion at position 13 (query step 32) occur. The properly aligned true sequences and calling sequences clearly showing the insertion and deletion positions of the bases can be presented as follows: Error: Insertion / Deletion True sequence: TAG CAA G * G TCC G CT ACT GGC Called array: TAG CAA G T G TCC * CT ACT GGC
[0219] As can be understood in light of the disclosure herein, accidental FNRs and FLDs result in insertion and deletion errors that cannot be algorithmically corrected and remain undiscovered if the true sequence is unknown. In other words, a base is misrepresented if more than half of the single-molecule sensors 105 in the aligned sequence give incorrect answers. The probability of such an event depends on the ratio (value of r) at which chemical defects occur. As described above, the examples presented herein use high error rates to illustrate the application of error correction techniques. Error rates in actual embodiments must be significantly lower so as to reduce the likelihood that the error correction procedure will fail to correct the error. The disclosed error correction techniques can be used to properly align multiple sensor 105 outputs in a querying step. This can be achieved using a deep understanding of the physical origin of possible error types (e.g., knowledge that a particular sequence is impossible for a source nucleic acid), their average occurrence rates, and their signatures in the sensor sequence outputs. When the chemical error rate is high and the error signatures are unclear, the error correction algorithms can be computationally intensive and difficult to implement. The following discussion explains how the probability of incorrect base calling depends on the read length, cluster size N (for CLUS instruments), the number of sensors K that detect instances of the same nucleic acid strand (for SMAS instruments 100), and the rate of chemical errors that fail.
[0220] Typical quantitative results from a cluster sequencer A simple quantitative model is developed here to estimate the probability of false base calls in a cluster sequencer using the modified additive sequencing protocol introduced above. Various types of errors (FNI, FLR, FNR, and FLD) are given by the ratio of r (here,
number
number
number
[0221] This background signal is generated by heterophase nucleic acid strands that incorporate non-complementary nucleotides at the in-phase positions of the ensemble average. As shown in Figure 34A,
number
number
number
[0222] As shown in Figures 34A and 34B, during the initial sequencing query (C small), the states <1> and <0> are well separated, but they rapidly approach the mean value N / 2 according to the functional morphology represented by equations 1(a) and (b). Furthermore, since error occurrences are random and independent events, the measured signals of the two states are discretely distributed around their ensemble mean values <1> and <0>. Specifically, the probability that the intensity of the measured ON state for cluster size N is k, given the ensemble mean of <1>, is given by a Poisson distribution:
number
number
number
number
number
number
number
number
number
number
number
[0223] Figure 37A shows,
number
number
number
number
number
number
number
[0224] Generally, P C、N、r The probability of an incorrect base call in sequencing query number C for cluster size N and chemical failure rate r, expressed as above, is the sum of the probabilities of the OFF state being incorrectly called, i.e., the above
number
number
number
number
[0225] Figures 38A and 38B plot equations 4(a) and 4(b) as functions of C for various combinations of N and r. Figure 38A shows,
number
number
number
number
[0226] Figure 39 shows position 150
number
[0227] Currently, the benchmark in the sequencing industry is the ability to read 150 consecutive bases with a probability of 1 in 1,000 incorrect base calls at position 150. This is commonly referred to as Q30, but for detecting rare mutations in high-precision diagnostics, a considerably larger sequencing quality factor of Q40 and even Q50 with longer read lengths is desirable. P in Equations 3(a) and (b) C、N、r A general expression for this can be used to fully explore the CNr parameter space and estimate the error tolerance and cluster size requirements for any sequencing metric. Figure 39 shows position 150
number
number
number
[0228] Figure 40A is all
number
number
number
number
number
number
[0229] Finally, the cumulative probability of an incorrect base call at position 150 (in some embodiments, the target read length) is less than 1 in 100.
number
number
number
number
number
number
number
number
number
number
[0230] Typical quantitative results from single-molecule array sequencers To compare the CLUS platform and the SMAS platform, a simple quantitative model is developed to estimate the probability of incorrect base calls in the SMAS instrument 100. Unlike the ensemble applicable to the CLUS instrument (described above), where error correction is little to no, the SMAS instrument 100 is applicable to individual nucleic acid molecules. The ability to sequence and record results individually enables the development and implementation of robust techniques for identifying and eliminating at least some of the errors in the resulting data recording(s). As disclosed herein, one or more error correction techniques may be applied to data generated from a sequencing procedure (e.g., SBS) before base calling is performed to identify and correct errors in the detection results and improve the accuracy of the called sequences. Specifically, the alignment of detection results from multiple sensors 105 can be improved in some or all of the querying steps of the sequencing procedure. Even if the error correction algorithm succeeds in correctly aligning the detection results from multiple sensors, incorrect base calling may still occur. As described above, accidental FNR and FLD errors can lead to insertion and deletion errors that may not be corrected. Depending on the number of errors in the data recording (partially determined by the chemical failure rate), the error correction process may be complex and computationally intensive, but it will be understood that modern processors have sufficient computing power to perform even the most computationally intensive techniques disclosed.
[0231] Below, we consider a typical example of the K single-molecule sensor 105 of the SMAS instrument 100, each capable of monitoring a single instance of cloned DNA. Similar to the analysis of the CLUS instrument described above, we assume that four types of errors (FNI, FLR, FNR, and FLD) occur randomly during the sequencing procedure and are distributed throughout the querying process.
[0232] As described above, in some embodiments, a probabilistic error correction algorithm is implemented (for example, by at least one processor 130 which may be included in or outside the SMAS device 100). In some embodiments, the probabilistic error correction algorithm improves the alignment of the detection results of at least some sensors 105 in the data recording. In some embodiments, some or all of the error correction algorithm is performed after some or all of the query steps are completed and some or all of the data has been acquired. As previously mentioned, the error correction procedure essentially eliminates FNI and FLR, as well as some FLD. The algorithmic realignment of the detection results of sensors 105 also makes the probability of making an incorrect base call independent of the query step number C. Also, since the error correction algorithm realigns the detection results of at least some sensors 105 in the data recording, thereby correcting at least some of the errors, the effective error rate is smaller than in the case of rCLUS. After the application of the exemplary error correction algorithm, in some embodiments, a base is incorrectly called only if more than half of the K sensors 105 in the algorithmically aligned sequence give incorrect results.
[0233] The probability of making an incorrect base call (P K、r ) is simply a function of (a) the number K of sensors 105 that sequence instances of the same nucleic acid molecule (which may be less than all sensors 105 in the sensor array 110), and (b) the chemical failure rate r. Similar to the approach adopted for the analysis of the CLUS instrument described above, the value of K is restricted to an odd number to avoid the case where exactly half of the sensors 105 do not match the other half. The probability of making an incorrect base call is given by:
number
number
number
number
number
number
number
number
[0234] for example,
number
number
number
number
[0235] As described above for the CLUS instrument, for the SMAS instrument 100, the Kr parameter space is explored below to identify regions where the probability of an incorrect base call at any query position is lower than 1 in 100 (Q20), 1 in 1,000 (Q30), 1 in 10,000 (Q40), and 1 in 100,000 (Q50). Figure 42 shows all query steps (P K、r The calculation results for the Kr parameter space where the probability of incorrect base calling is less than 1 in 100 (Q20), 1 in 1,000 (Q30), 1 in 10,000 (Q40), and 1 in 100,000 (Q50) are shown. As shown in Figure 42, when the number K of single-molecule sensors 105 that sense instances of the same nucleic acid molecule is 11 and the required sequencing accuracy is Q30, the acceptable chemical failure rate is
number
[0236] As shown in the comparison with Figure 39, the acceptable error rate for the SMAS device 100 is considerably larger than the acceptable rate for the CLUS device, but this result alone does not mean that the CLUS device (P C、N、r The probability of making an incorrect base call in the threshold query step C is very low during the initial query step, th The sudden increase in the two platforms means they cannot be compared equally. This phenomenon is explained in relation to Figure 39. On the other hand, in the case of the SMAS instrument 100, incorrect base calling (P K、r The probability of ) remains constant throughout the entire query process, and therefore the cumulative error becomes large.
[0237] A fairer way to compare the performance of the CLUS instrument and the SMAS instrument 100 is to compare the cumulative error probabilities of these two instrument types. Equation 5(b) above represents the cumulative error probability of the CLUS instrument. The cumulative error probability of the SMAS instrument 100 can also be derived. The probability of making an incorrect base call for each query step C is P K、r (Equation 6) Therefore, the probability of making a correct call is
number
number
number
number
[0238] Figures 43A and 43B show the cumulative probabilities of incorrect base calling at position 150 for the CLUS and SMAS instruments 100. Equation Figure 5(b) can be used, for example, to calculate the probability that a CLUS instrument makes an incorrect base calling at any base position less than or equal to 150. Figure 43A shows the Kr parameter space for the CLUS instrument, where the cumulative probability of incorrect base calling at position 150 is less than 1 in 100 CLUS instruments.
number
number
number
number
number
number
number
number
[0239] A comparison of Figures 43A and 43B reveals that the SMAS instrument 100 is a potentially superior sequencing platform compared to the CLUS instrument. The SMAS instrument 100 can have a smaller footprint (as discussed in, for example, Figures 7A, 7B, 9A, 9B, and 10) and can have significantly higher error tolerance than the CLUS instrument. The use of the SMAS instrument 100 guarantees higher throughput, lower error rates, and longer readout lengths compared to the CLUS instrument, which relies on larger molecular ensembles. The development of a commercially viable SMAS instrument 100 and / or system could utilize some or all of the availability of effective bioinformatics tools to adjust the alignment in data recording of sequencing data from at least several nanoscale sensors 105 by (a) high-precision nanoscale fabrication of densely packed sensors 105 capable of recognizing individual labels, (b) optimization of chemical processes to reduce error rates to an acceptable level, and / or (c) probabilistically eliminating errors.
[0240] Exemplary SMAS sequencing procedure
[0241] As explained above, improving the sequencing throughput of a CLUS instrument can be achieved by reducing the cluster size N (therefore packing more clusters into the instrument), which can be difficult, if it also reduces the failure rate of sequencing chemistry. In contrast, the following are feasible implementations of an error-tolerant, ultra-high-throughput SMAS instrument 100 using a large array of single-molecule binding sites 116, according to several embodiments. For the purposes of this example, the SMAS instrument 100 is assumed to sequence DNA, but it should be understood that in general, any type of nucleic acid can be sequenced.
[0242] Figures 44 and 45 show exemplary sample preparation and loading processes 500 according to several embodiments. Figure 44 is a flow chart of process 500, and Figure 45 shows the results of various steps of process 500. In some embodiments, the sample preparation and loading process 500 begins at 502. At 504, DNA extraction and purification are performed to obtain several extracted DNA fragments 505, as shown in Figure 45. At 506, an adapter complementary to the primer is ligated to one end (e.g., 3') of the extracted DNA to produce a strand 507 shown in Figure 45. At 508, PCR (or some other replication technique) is performed to produce multiple (ideally identical) instances of the extracted strand, as shown as 509 in Figure 45. In 510, a molecular linker capable of creating a strong bond (e.g., by click chemistry) to the chemically functionalized surface (binding site 116) of the fluid chamber 115 of the SMAS apparatus 100 binds to the other end (e.g., 5') of the ssDNA fragment, thereby generating the strand 511 shown in Figure 45. In 512, the functionalized strand 511 is loaded into the fluid chamber 115, randomly scattered among the binding sites 116, and binds to the binding sites. As shown in the far right of Figure 45, each binding site 116 supports only one DNA strand. (Each binding site 116 can support one or fewer strands, but it should be understood that there is no requirement that all binding sites 116 must support DNA strands. Fewer than all of the binding sites 116 of the SMAS apparatus 100 may be used, whether intentionally or accidentally.) Assuming that the extracted DNA fragments 503 are distinct from one another, as a result of the sample preparation and loading process 500, multiple instances of each extracted DNA fragment 505 are present in the fluid chamber 115, but their positions are unknown. In 514, the exemplary sample preparation and loading process 500 is completed.
[0243] The advantage of the exemplary sample preparation and loading process 500 is that it simplifies DNA amplification, which can be performed on a bulk off-instrument using (e.g.) conventional PCR, before the DNA strands are added to the SMAS instrument 100. In contrast, when a CLUS instrument is used, amplification (e.g., bridge amplification) is performed only after the DNA fragments have been added to the CLUS instrument to create an array of sequential clusters of amplified DNA.
[0244] After the sample preparation and loading process 500 has been performed, base calling can be performed using, for example, the additive approach, subtractive approach, or modified additive approach introduced above. Figures 46A, 46B, and 46C show simulated detection results (sensors 105 detect labels) using a modified additive approach during three exemplary query cycles (A?⇒T?⇒C?⇒G?, each with a total of 12 query steps) performed by an exemplary SMAS instrument 100 having a sensor array 110 with 20 sensors 105 (and 20 binding sites 116) arranged in 4 rows and 5 columns. Multiple instances of four different DNA strands are randomly distributed throughout the sensor array 110, but their specific locations within the sensor array 110 and their sequences are initially unknown.
[0245] Figure 47 shows how the detection data shown in Figures 46A, 46B, and 46C can be sorted to call bases and to reveal the locations of different DNA strands. Figure 47 provides a table showing the output of all sensors 105 in an exemplary array in each query step, and the resulting base calls that yield the calling sequences. The right-hand portion of Figure 47 sorts the sensors 105 to group the detection results of sensors 105 that sense instances of the same DNA strand. As shown in Figure 47, the following four sequences are called: GCT (strand number 1), TAG (strand number 2), ACG (strand number 3), and TTA (strand number 4).
[0246] If an error (FNI, FLR, FNR, or FLD) occurs during the querying process, some of the detection results (whether a label was detected or not) are incorrect, and at least some errors can be detected and eliminated by performing the deterministic and / or probabilistic error detection and / or correction techniques described above, as long as the identification information of the sensors 105 sensing instances of the same DNA strand is determined. It should be noted that instances of a particular DNA strand may be bound to binding sites 116 scattered throughout the fluid chamber 115, and their locations are generally unknown when the sequencing process is initiated. Once the process is started, between each querying step, each of the multiple S sensors 105 detects a label at its respective binding site 116. To perform error correction, a subgroup of S sensors 105 that are sequencing instances of the same nucleic acid strand is identified.
[0247] Consider a very large sensor array 110 (e.g., 4 billion binding sites 116 and 4 billion individual sensors 105) with 400 million distinct DNA strands, each approximately 150 base pairs long. This means there are approximately 10 instances of each unique DNA strand randomly distributed throughout the fluid chamber 115 (and the binding sites 116 and sensor array 110). Also, for the example, assume the sequences are random. Assuming a reasonably low error rate r, after the first query cycle, almost all of the binding sites 116 (and sensors 105) that hold (sens) DNA instances beginning with A are identified as well as the binding sites that hold (sens) T, C, and G. 9 The sensor 105 detects a label indicating that the first base is A, and approximately 10 9 The sensor 105 detects a label indicating that the first base is T, and approximately 10 9 The sensors detect a label indicating that the first base is C, and approximately 10 9 Each sensor detects a label indicating that the first base is G. After the second query cycle, almost all of the binding sites 116 (and sensors 105) that hold (sens) DNA instances, starting with all 16 possible combinations (AA, AT, AC, AG, TA, TT, TC, TG, CA, CT, CC, CG, GA, GT, GC, and GG), have been identified. Approximately 2.5 × 10⁻⁶ 8 Each sensor detects a label indicating that the first and second bases are AA, resulting in approximately 2.5 × 10⁻¹⁶ units. 8 Each sensor detects a label indicating that the first and second bases are AT, resulting in approximately 2.5 × 10⁻⁶ 8 For example, a single sensor detects a label indicating that the first and second bases are AC. Generally, a certain number of D labels are used for sequencing (or a modified additive approach is assumed to be used).
number
number
number
number
[0248] The confidence that the correct set of binding sites 116 has been identified increases with the number of query steps, but the probability of detection errors also increases (e.g., misdetecting or failing to detect a label). Multiple errors can occur during the first query cycle while binding sites 116 holding instances of the same strand are being identified. The results derived for the CLUS instrument suggest that this is not a problem. For example, Figure 38A shows that the probability of the CLUS instrument making an incorrect base call during the initial query steps is very small, and the probability of errors increases sharply only at threshold C. th This indicates that this only occurs when it reaches a certain point. Also, since the SMAS instrument 100 simply reports the ensemble result by summing the results of the individual sensors 105, it should be noted that the base calling accuracy of the SMAS instrument 100 is the same as that of the CLUS instrument when error correction is not applied.
[0249] For example, consider the above example of a 4 billion sensor array, with a set of 11 sensors 105 that monitor instances of a specific DNA strand randomly distributed throughout the binding site 116.
number
number
[0250] If the chemical error rate is expected or known to be too high, and as a result errors may be problematic in the first approximately 35 query steps, an alternative approach can be used to assist in identifying binding sites 116 that have instances of the same DNA strand. For example, different unique barcodes can be ligated to the primer ends of a subset of extracted DNA so that known sequences can be read during the initial sequencing cycle. Figure 49 shows the use of barcodes in sample preparation and DNA loading according to several embodiments. As shown in Figure 49, unique barcodes are ligated to the extracted DNA to facilitate the recognition of sites that hold instances of the same DNA in the presence of sequencing errors. For example, Figure 49 shows four unique DNA strands, each assigned a unique barcode (e.g., strand 1 is assigned barcode 119A, strand 2 is assigned barcode 119B, strand 3 is assigned barcode 119C, and strand 4 is assigned barcode 119D). If the barcodes are significantly different from each other, the barcodes should be easily identifiable even in the case of a very high chemical failure rate. As is understood, a sufficient number of unique barcodes can be expensive for high-throughput diagnostic applications.
[0251] The 4 billion sensor SMAS system 100 described herein is considered a fairly high-throughput sequencer by current standards. Such a SMAS system 100 delivers approximately 150 gigabase (Gb) reads per run, comparable to the output of state-of-the-art high-end sequencing systems introduced in 2020.
[0252] It should be understood that there are many ways to implement the apparatus, systems, and methods disclosed herein. For example, a system for nucleic acid sequencing may consist of a single apparatus (e.g., SMAS apparatus 100 including all the hardware and software for performing the disclosed operations), or it may consist of the SMAS apparatus 100 and other components that perform the disclosed operations together. For example, the system may consist of the SMAS apparatus 100 which performs a nucleic acid sequencing procedure and stores the detection results from the sequencing procedure, and an external component which performs error detection and correction on the stored detection results and retrieves the bases. It may also have another processor (for example, in an external computer).
[0253] Figure 50 shows an exemplary system 160 according to several embodiments. The system 160 comprises a fluid chamber 115, a plurality of S sensors 105, and at least one processor 130 (i.e., including, but not limited to, these). Optionally, the system 160 includes a memory 170 for storing records containing detection results obtained during the sequencing procedure (e.g., one or more files having binary entries documenting whether each of the plurality of S sensors 105 detected or did not detect at least one indicator between each of a plurality of query cycles). If the system 160 includes the memory 170, as shown by the dashed line in Figure 50, at least one processor 130 can be communicatively connected to the memory 170 so that the at least one processor 130 can store data in the memory 170 and / or retrieve data from the memory 170.
[0254] The fluid chamber 115 comprises S binding sites, each of which is configured to bind to one or fewer strands of nucleic acid to be sequenced. Figure 50 shows four binding sites 116, but it should be understood that the system 160 may have more or fewer binding sites 116. Each of the S sensors 105 is configured to detect a label present in the fluid chamber 115. Figure 50 shows four sensors 105, but it should be understood that the system 160 may have more or fewer sensors 105. When the system 160 is operating, each of the S sensors 105 detects a label bound to a nucleotide incorporated into each strand of nucleic acid bound to each of the S binding sites 116. As previously stated, the sensor 105 may be a magnetic sensor, an optical sensor, or any other type of sensor capable of detecting a label used to label nucleotides. The fluid chamber 115, the sensors 105, and the binding sites 116 have been described in detail above. These descriptions apply to Figure 50 and are not repeated here.
[0255] At least one processor 130 is configured to execute one or more machine-executable instructions. When executed, the instructions cause at least one processor 130 to perform a sequence determination procedure that includes several query steps (as described, for example, in the context of Figures 11, 12, 14, 16, and 44). Specifically, during operation, during the query steps of the sequence determination procedure, at least one processor 130 acquires each of the characteristics of each of the S sensors 105 (represented by dashed lines between at least one processor 130 and sensors 105A, 105B, 105C, and 105D). Each characteristic indicates whether the sensor 105 detected a label (e.g., the presence or absence of at least one label). At least one processor 130 can interpret the acquired characteristics to determine whether the sensor 105 detected the presence or absence of a label. Based at least in part on each acquired characteristic, at least one processor 130 records whether each sensor detected the presence or absence of at least one label during the query steps. At least one processor 130 is also configured to perform an error correction procedure on at least one record containing the results of a sequencing procedure. The error correction procedure may operate on some or all of the records generated by the sequencing procedure, or on detection results from some or all of the query steps of the sequencing procedure. For example, as described above, in order to apply an error correction procedure, at least one processor may identify deterministic or stochastic error correction and apply it to a subset of K records, each of the K records in the subset corresponding to a detection result from a sensor 105 that senses instances of the same nucleic acid strand. The sequencing procedure and the error correction procedure have been described in detail above. These descriptions apply to the system and at least one processor 130 in Figure 50 and will not be repeated here.
[0256] At least one processor 130 is a general-purpose or dedicated processor (or a set of processing cores) It may be implemented by (t), and thus, by executing a programmed sequence of instructions, various operations related to acquiring the characteristics of the sensor 105, executing error correction procedures, and / or interacting with a user, system operator, or other system components can be achieved.
[0257] At least one processor 130 of system 160 may be a single processor (for example, in the SMAS device 100) or may comprise multiple processors, which may be located in the same place (for example, in the SMAS device 100) or may be physically separated from each other. For example, a first part of at least one processor 130 may be provided in the SMAS device 100, and a second part of at least one processor 130 may be outside the SMAS device 100. In embodiments in which at least one processor 130 comprises first and second parts, the first part may be responsible for acquiring characteristics of sensor 105, determining based on the characteristics whether sensor 105 detected a sign during the query cycle, and recording (for example, in memory 170) whether each of S sensors 105 detected the presence or absence of at least one sign during the query cycle, while the second part may be responsible for acquiring a record of the detection results and performing error correction procedures. Alternatively, the first part may be responsible for acquiring the characteristics of the sensor 105, determining based on the characteristics whether each of the sensors 105 detected at least one marker during the query cycle, and providing an indication of whether the sensor 105 detected a marker to another entity via a communication interface (e.g., a wireless or wired interface such as Ethernet or Wi-Fi). In such an implementation, the second part of at least one processor 130 may be responsible for acquiring a record of the detection results provided by the first part of at least one processor 130 (e.g., a file having binary entries documenting whether each of the S sensors 105 detected or did not detect at least one marker during each query cycle), performing error correction procedures, and calling bases.
[0258] The preceding description and accompanying drawings use certain terms to provide a complete understanding of the disclosed embodiments. In some cases, terms or drawings may refer to certain details that are not necessary to carry out the invention.
[0259] To avoid unnecessarily obscuring this disclosure, well-known components are shown in block diagram form and / or not described in detail, or in some cases not described at all.
[0260] Section headings provided in the detailed description are for convenience or reference only and are not intended to limit. Section headings do not define, limit, interpret, or describe the scope or scope of such sections. Furthermore, while various specific embodiments are disclosed, it will be apparent that various modifications and variations can be made without departing from the broader spirit and scope of this disclosure. For example, any feature or aspect of an embodiment can be applied in combination with or instead of any other feature or aspect of the embodiment.
[0261] Some of the techniques and methods disclosed herein (e.g., obtaining detection results from sensor 105, performing error correction procedures, etc.) and / or user interfaces for configuring and managing them can be implemented by the machine execution of one or more sequence instructions (including relevant data necessary for the execution of the appropriate instructions). Such instructions can be recorded on one or more computer-readable media for later retrieval and execution within one or more processors of a dedicated or general-purpose computer system or consumer electronic device or equipment. Various computer-readable media can embody such instructions and data. Examples of non-volatile storage media in various forms (e.g., optical, magnetic, or semiconductor storage media), and carriers that can be used to transfer such instructions and data via wireless, optical, or wired signaling media or any combination thereof, include, but are not limited to, carriers. Examples of transfer of such instructions and data via carriers include, but are not limited to, transfers via the Internet and / or other computer networks via one or more data transfer protocols (e.g., HTTP, FTP, SMTP, etc.) (e.g., uploads, downloads, emails, etc.).
[0262] Unless otherwise explicitly defined herein, all terms should be given the broadest possible interpretation, including the meaning implied by this specification and the drawings, the meaning understood by those skilled in the art, and / or the meaning defined in dictionaries, papers, etc. As expressly stated herein, some terms may not correspond to their usual or conventional meanings.
[0263] Where used herein and in the appended claims, the singular forms “a,” “an,” and “the” do not exclude multiple references unless otherwise specified. The word “or” should be interpreted as inclusive unless otherwise specified. Thus, the phrase “A or B” should be interpreted as meaning all of the following: “both A and B,” “A but not B,” and “B but not A.” No use of “and / or” herein implies that the word “or” alone implies exclusivity.
[0264] As used herein and in the appended claims, the phrases of the form “at least one of A, B, and C,” “at least one of A, B, or C,” “one or more of A, B, or C,” and “one or more of A, B, and C” are interchangeable and each encompasses all of the following meanings: “A only,” “B only,” “C only,” “A and B but not C,” “A and C but not B,” “B and C but not A,” and “all of A, B, and C.”
[0265] To the extent that the terms “include,” “having,” “has,” and “with,” and their variations thereto are used in a detailed description or in the claims, such terms are intended to be inclusive in the same way as the term “comprising,” that is, to mean “including but not limited to.”
[0266] The terms “exemplary” and “embodiment” are used to represent examples, not preferences or requirements.
[0267] The term “joined” is used herein to refer to direct connection / attachment and connection / attachment via one or more intervening elements or structures.
[0268] The terms “over,” “under,” “between,” and “on” in this specification refer to the relative position of one feature to another feature. For example, a feature positioned “over” or “below” another feature may be in direct contact with the other feature or may have intervening material. Furthermore, a feature positioned “between” two features may be in direct contact with both features or may have one or more intervening features or materials. In contrast, a first feature “on” a second feature is in contact with that second feature.
[0269] The term "substantially" is used to describe structures, configurations, dimensions, etc., that are largely or nearly as described, but due to manufacturing tolerances, etc., the structure, configuration, dimensions, etc., may not always or necessarily be as precise as described. For example, describing two lengths as "substantially equal" means that the two lengths are the same for all practical purposes, but they do not have to be (and do not need to be) exactly equal on a sufficiently small scale. As another example, a structure that is "substantially perpendicular" is considered perpendicular for all practical purposes, even if it is not exactly 90 degrees to the horizontal.
[0270] The drawings are not necessarily to scale, and the dimensions, shapes, and sizes of features may differ substantially from those shown in the drawings.
[0271] While specific embodiments are disclosed, it will be apparent that various modifications and variations can be made without departing from the broader spirit and scope of this disclosure. For example, any feature or aspect of an embodiment can be applied in combination with or instead of any other embodiment, at least where feasible. Accordingly, this specification and the drawings should be considered illustrative rather than restrictive.
Claims
1. A method for detecting and correcting errors occurring during a sequencing procedure performed using a single-molecule array sequencing apparatus equipped with multiple sensors, wherein each of the multiple sensors is configured to detect one or fewer molecules at a time, and the method A step of detecting a label removal failure error performed by a first sensor among the plurality of sensors, wherein the label removal failure error occurred during the sequence determination procedure. A step of correcting the label removal failure error in the recording of the detection result from the sequence determination procedure, wherein the detection result is from the first sensor, Includes, A method comprising the step of detecting a label removal failure error performed by a first sensor among the plurality of sensors, the step of determining that the first sensor among the plurality of sensors has detected a label after the step of cleaving and washing away the label in the sequencing procedure and before introducing the labeled nucleotide in the next query step, wherein the next query step is the first query step after the step of cleaving and washing away the label.
2. In the method according to claim 1, the step of correcting the label removal failure error is: A method comprising the step of changing the value corresponding to the next query step in the recording of the detection result from a first value indicating the detection of the sign to a second value indicating that the sign was not detected.
3. In the method according to claim 1, the step of detecting the label removal failure error performed by the first sensor among the plurality of sensors is: (a) A step of detecting the characteristics of the first sensor at a first time after introducing a labeled nucleotide into the fluid channel of the single-molecule array sequencing apparatus, The labeled nucleotide comprises a plurality of labels, and the fluid channel allows the labeled nucleotide to be incorporated by a molecule within the detection region of the first sensor. The step of indicating whether the characteristic is such that at least one marker is present within the range of the first sensor or no marker is present within the range of the first sensor, (b) After step (a), a step of determining that the first sensor has detected at least one sign, based at least in part on the characteristics, (c) After step (b), a step of cutting and washing away the label, (d) After step (c), a step of detecting the characteristics of the first sensor at a second time, (e) After step (d), a step of determining that the characteristics of the first sensor detected at the second time and the characteristics of the first sensor detected at the first time are the same, Methods that include...
4. In the method according to claim 3, the step of detecting the label removal failure error is: A method comprising the step of changing, in the recording of the detection result, a value corresponding to the next query step of the sequencing procedure from a first value indicating the detection of a label to a second value indicating that the label was not detected, wherein the next query step is the first query step after the step of cutting and washing away the label.
5. A method according to any one of claims 1 to 4, further, A step of estimating the sequence using probabilistic error correction, following the step of correcting the aforementioned label removal failure error, A step of identifying a first error scenario that is consistent with a first detection result for a first sensor, a second detection result for a second sensor, and a third detection result for a third sensor, wherein the first detection result is assumed to be correct. The steps include determining the probability of the first error scenario occurring, A step of identifying a second error scenario that matches the first detection result of the first sensor, the second detection result of the second sensor, and the third detection result of the third sensor, wherein the first detection result is assumed to be incorrect. The steps include determining the probability of the second error scenario occurring, The steps include correcting the record of the detection result based on a comparison of the probability of the first error scenario occurring and the probability of the second error scenario occurring, A method that includes steps.
6. The method according to claim 5, wherein the step of correcting the record of the detection result based on the comparison of the probability of the first error scenario occurring and the probability of the second error scenario occurring is: Depending on the fact that the probability of the second error scenario occurring is higher than the probability of the first error scenario occurring, The steps include changing the entry in the recording of the detection result of the first sensor from a first value indicating that a sign was detected to a second value indicating that no sign was detected, The steps include changing the entry in the recording of the detection result of the first sensor from the second value indicating that no sign was detected to the first value indicating that a sign was detected, Methods that include...
7. The method according to claim 6, further comprising the step of correcting the record of the detection result based on the comparison of the probability of the first error scenario occurring and the probability of the second error scenario occurring, A method comprising the step of modifying an entry in the recording of the detection result for at least one of the second or third sensors, depending on whether the probability of the first error scenario occurring is higher than the probability of the second error scenario occurring.
8. A method according to claim 5, wherein the step of correcting the record of the detection results includes the step of deleting a plurality of consecutive identical entries from the record, each of which indicates that no indicator was detected.
9. It is a system, A plurality of S binding sites, each of which is configured to bind to one or fewer nucleic acid strands to be sequenced, A plurality of S sensors configured to detect labels attached to nucleotides incorporated into nucleic acid chains bound to the plurality of S binding sites, wherein each of the S sensors is for sensing each nucleic acid chain bound to each of the S binding sites, The system comprises at least one processor configured to execute one or more machine-executable instructions, wherein when the instruction is executed, the at least one processor is instructed to perform the following actions for each of the plurality of S sensors: (a) In the query step of the sequencing procedure, a step of acquiring the characteristics of the sensor, wherein the characteristics indicate the presence or absence of at least one label attached to each nucleotide incorporated into each strand of the nucleic acid bound to each of the binding sites, (b) A step of determining whether the sensor has detected the at least one label attached to the nucleotide incorporated into each strand of the nucleic acid bound to each binding site, based at least in part on the characteristics obtained in step (a); (c) After the mark cutting step performed after step (a) and before the next query step of the sequence determination procedure, a step of detecting the characteristics of the sensor, (d) A step in which a label removal failure error is indicated by determining, at least partially, based on the characteristics obtained in step (c), whether the sensor is still detecting at least one label attached to a nucleotide incorporated into each strand of the nucleic acid bound to each binding site, (e) In step (b), the sensor determines that it has detected at least one label attached to a nucleotide incorporated into each strand of the nucleic acid bound to each of the binding sites, and in step (d), the sensor determines that it has still detected at least one label attached to a nucleotide incorporated into each strand of the nucleic acid bound to each of the binding sites after the label cleavage step, and in the next query step after step (d), the sensor corrects the label removal failure error by changing the instruction "label detected" to "no labels detected". (f) For subsequent query steps in the sequence determination procedure, a step which repeats steps (a) to (e), A system that executes an action.
10. In the system according to claim 9, when one or more machine-executable instructions are executed, the at least one processor is further given the following: A step following step (f) above, to identify a plurality of candidate sequences associated with an instance of a particular nucleic acid strand, wherein each instance is a template nucleic acid strand or a copy thereof, A step of determining a metric for each of the aforementioned plurality of candidate sequences, wherein each metric provides an indicator of the likelihood of occurrence. The steps include selecting a specific candidate sequence that is most likely to be correct by applying a judgment criterion to each of the aforementioned metrics, A system that executes an action.
11. In the system according to claim 10, when one or more machine-executable instructions are executed, the at least one processor is further given the following: Based on known constraints on the nucleic acid sequence in the aforementioned specific nucleic acid chain, the step of eliminating at least one of the plurality of candidate sequences is performed. A system in which the known constraint is the knowledge that at least one of the multiple candidate sequences cannot occur naturally.
12. In the system according to any one of claims 9 to 11, when one or more machine-executable instructions are executed, the system further provides to at least one processor: A step of generating a record, wherein the record includes the results of the array determination procedure for at least one subset of the S sensors, the record includes a set of binary values, where a first binary value indicates that at least one indicator was detected, and a second binary value indicates that no indicator was detected, and when one or more machine-executable instructions are executed, the at least one processor is further instructed to do the following: A step of identifying a consecutive number of the second binary value in the record, wherein the number is equal to the number of query steps per sequence determination cycle, The steps include deleting the consecutive numbers for the second binary value from the record, A system that executes an action.
13. In the system according to any one of claims 9 to 12, each sequence determination cycle in the sequence determination procedure has P query steps, and when one or more machine-executable instructions are executed, the system further provides to at least one processor: The steps include identifying a set of P consecutive indications that no indicator was detected by the first sensor among the plurality of S sensors, The steps include deleting the set of P consecutive indications that none of the indicators were detected by the first sensor among the plurality of S sensors, A system that executes an action.
14. In the system according to any one of claims 9 to 13, when one or more machine-executable instructions are executed, the system further provides to at least one processor: A system that performs the step of modifying at least one entry in a record to reflect the result of a majority vote for a particular query step, wherein the result of the majority vote indicates whether more than half of the plurality of S sensors detected a sign, and the at least one entry in the record corresponds to the particular query step.
15. A device for sequencing nucleic acids, A fluid chamber comprising a plurality of S binding sites, wherein each of the S binding sites is configured to bind to one or fewer nucleic acid chains to be sequenced, A plurality of S magnetic sensors configured to detect a label present in the fluid chamber, wherein each of the S magnetic sensors is for sensing each strand of nucleic acid bound to each of the plurality of S binding sites, A processor configured to execute one or more machine-executable instructions, The system is equipped with such that when the instruction is executed, it instructs at least one processor to provide information for each of the plurality of S magnetic sensors. (a) A step in the query step of the sequencing procedure, wherein the characteristics of the magnetic sensor are to indicate the presence or absence of at least one label attached to each nucleotide incorporated into each strand of the nucleic acid bound to each of the binding sites, (b) A step of determining whether the magnetic sensor has detected at least one label attached to a nucleotide incorporated into each strand of the nucleic acid bound to each binding site, based at least in part on the characteristics obtained in step (a); (c) After the mark cutting step performed after step (a) and before the next query step of the arrangement determination procedure, a step of detecting the characteristics of the magnetic sensor, (d) A step of determining, at least partially, based on the characteristics obtained in step (c), whether the magnetic sensor still detects at least one label attached to a nucleotide incorporated into each strand of the nucleic acid bound to each binding site, wherein the result indicates a label removal failure error. (e) In step (b), the magnetic sensor detects at least one label attached to a nucleotide incorporated into each strand of the nucleic acid bound to each binding site, and in step (d), in response to the determination that the magnetic sensor still detects at least one label attached to a nucleotide incorporated into each strand of the nucleic acid bound to each binding site after the cleavage step, the step of correcting the label removal failure error by changing the instruction of the magnetic sensor in the next query step after step (d) from "label detected" to "no labels detected", A device that performs an action.
16. In the apparatus according to claim 15, the step of determining whether the magnetic sensor has detected at least one sign is: A step of determining whether the characteristics obtained for the magnetic sensor meet or exceed a threshold, or A step of comparing the characteristics obtained for the magnetic sensor with the baseline values of the characteristics of the magnetic sensor. A device including a device.
17. In the apparatus according to claim 15 or 16, when one or more machine-executable instructions are executed by the at least one processor, the at least one processor is further given the following: An apparatus for performing an error correction procedure on at least one record, wherein the at least one record includes the results of the array determination procedure for at least one subset of the plurality of S magnetic sensors in each of a plurality of M query steps.
18. In the apparatus according to claim 17, the step of performing the error correction procedure for at least one record is: A step of identifying a plurality of candidate sequences associated with an instance of a particular nucleic acid strand, based on at least a portion of the at least one record, wherein each instance is a template nucleic acid strand or a copy thereof. The steps include determining or estimating which of the aforementioned multiple candidate sequences is most likely to be correct, A device including a device.
19. In the apparatus according to claim 18, the step of determining or estimating which of the plurality of candidate sequences is most likely to be correct is: A step of determining a metric for each of the aforementioned multiple candidate sequences, wherein each metric provides an indicator of the likelihood of occurrence. The steps include selecting a specific candidate sequence that is most likely to be correct by applying a judgment criterion to each of the aforementioned metrics, A device including a device.
20. In the apparatus according to claim 18 or 19, the step of determining or estimating which of the plurality of candidate sequences is most likely to be correct is: The step includes eliminating at least one of the plurality of candidate sequences based on known constraints on the nucleic acid sequence in the particular nucleic acid chain, The known constraint is that a specific base sequence in at least one of the multiple candidate sequences cannot occur spontaneously.
21. In the apparatus according to any one of claims 17 to 20, each sequence determination cycle in the sequence determination procedure has P query steps, and the step of performing the error correction procedure for at least one record is The steps include identifying a set of P consecutive indications in at least one of the records that no indicators were detected, The steps include deleting the set of P consecutive indications from the at least one record that no indicators were detected, A device including a device.
22. In the apparatus according to any one of claims 17 to 21, the step of performing the error correction procedure for at least one record is: Apparatus comprising the step of modifying at least one entry of the at least one record to reflect the result of a majority vote of a particular query step, wherein the at least one entry of the at least one record corresponds to the particular query step and the result of the majority vote indicates whether more than half of the plurality of S sensors detected a sign.
23. A method for sequencing a plurality of S nucleic acid strands using the apparatus described in claim 15, The steps of binding the plurality of S nucleic acid chains to the S binding sites, A step of performing an array determination procedure including M query steps in order to capture M detection results from each of the plurality of S magnetic sensors, wherein each of the M detection results indicates whether each of the plurality of S magnetic sensors detected at least one sign in the fluid chamber during each of the M query steps. A method comprising the step of performing the arrangement determination procedure, wherein the step of the at least one processor executes one or more machine-executable instructions causing each of the plurality of S magnetic sensors to perform steps (a) through (e).
24. The method according to claim 23, wherein each sequencing cycle in the sequencing procedure has P query steps, and the method further A step of generating S records, wherein each of the S records captures the result of the arrangement determination procedure for each of the M query steps, The steps include identifying a set of P consecutive indications indicating that no indicators were detected in at least one record of at least one subset of the S records, The steps include deleting the set of P consecutive indications from the at least one record that no indicators were detected, Methods that include...
25. A method according to claim 24, further comprising the step of correcting one or more of at least one subsets of the S records based at least in part on the probabilities of nucleotide incorporation failure (FNI) errors, nucleotide removal failure (FNR) errors, and / or label detection failure (FLD) errors.
26. The method according to claim 24 or 25, wherein the at least one subset of the S records comprises an odd number of at least three records representing the sequencing results of instances of a first nucleic acid strand, each instance being a template nucleic acid strand or a copy thereof, and the method further, A step of identifying the majority vote detection result for a particular query process in at least one subset of the S records, A step of calling or not calling a base of the first nucleic acid strand, based at least in part on the numerous detection results for a particular query step, A method comprising the above, wherein the majority vote detection result indicates whether more than half of the plurality of S sensors detected the sign.
27. A method according to any one of claims 24 to 26, further comprising: A method comprising the step of calling a base in at least one of the plurality of S nucleic acid strands in response to the selected detection results in more than half of the at least one subset of the S records indicating the detection of the at least one label in the fluid chamber.
28. A method for mitigating errors in sequencing data generated as a result of a nucleic acid sequencing procedure using a single-molecule sensor array, wherein the single-molecule sensor array has a plurality of sensors, each of the plurality of sensors is associated with each of a plurality of binding sites, each of the plurality of binding sites is configured to bind to one or fewer nucleic acid strands being sequenced, and the sequencing data indicates, for each of the plurality of sensors, (i) in each query step, whether the sensor detected at least one label, and (ii) after a label cleavage step performed between query steps, the sensor continued to detect at least one label, and the method is: In the sequence determination data, the step of identifying multiple records is as follows: A step in which each of the plurality of records captures the respective sequencing result for each instance of the first strand of nucleic acid, each instance being a template nucleic acid strand or a copy thereof, each of the plurality of records having a plurality of entries, each of the plurality of entries indicating, for each of the plurality of query steps of the nucleic acid sequencing procedure, either (a) a label was detected by the respective sensor associated with each instance of the first strand of nucleic acid, or (b) no label was detected by the respective sensor associated with each instance of the first strand of nucleic acid. Identifying at least one flag removal failure (FLR) error in the plurality of records by identifying at least one query process immediately after the sensor continues to detect at least one flag after the flag cutting process, and after the sensor detects at least one flag during the preceding query process, The step of correcting at least one FLR error by changing at least one entry in the plurality of records from an entry stating "signature detected" to an entry stating "no signs detected," in response to identifying at least one signature removal failure (FLR) error in the plurality of records, corresponding to the at least one query step immediately after the at least one signature has been detected after the signature removal step and after the sensor has detected the at least one signature during the preceding query step, thereby generating the corrected plurality of records. A step of determining a plurality of candidate sequences for the first strand of the nucleic acid based on the plurality of corrected records, wherein each of the plurality of candidate sequences estimates at least a portion of the nucleic acid sequence of the first strand of the nucleic acid. The steps include identifying a specific candidate sequence from among the plurality of candidate sequences that is most likely to be correct, as at least a part of the nucleic acid sequence of the first strand of the nucleic acid, Methods that include...
29. In the method of claim 28, the step of identifying the plurality of records is, A step of searching for the sequence determination data for a barcode associated with the first strand of the nucleic acid, A step of identifying a common sequence of entries in each of the aforementioned multiple records, Methods that include...
30. In the method according to claim 28 or 29, the step of determining the plurality of candidate sequences for the first strand of the nucleic acid is: The steps include identifying within the modified plurality of records a specific query process in which the first sensor detects each marker and the second sensor does not detect any marker, The first step is to establish a first candidate sequence in which it is assumed that the first sensor should detect the signs and that each of the signs has been correctly detected, The steps include establishing a second candidate sequence in which the first sensor should not detect any of the markers, and which is assumed to have misdetected each of the markers, Methods that include...
31. In the method according to claim 28 or 29, the step of determining the plurality of candidate sequences for the first strand of the nucleic acid is: The steps include identifying within the modified plurality of records a specific query process in which the first sensor detects each marker and the second sensor does not detect any marker, The steps include establishing a first candidate sequence in which it is assumed that the second sensor should detect a sign and was unable to incorrectly detect any sign, The steps include establishing a second candidate sequence assuming that the second sensor should not detect any signs and was unable to correctly detect any arbitrary signs, Methods that include...
32. In the method according to any one of claims 28 to 31, each sequencing cycle in the sequencing procedure has P query steps, and the step of determining the plurality of candidate sequences for the first strand of the nucleic acid is The steps include identifying a set of P consecutive entries in at least one of the modified records that indicate no indicator was detected, The steps include deleting the set of P consecutive entries indicating that no indicator was detected in at least one of the modified records, Methods that include...
33. A method according to any one of claims 28 to 32, wherein at least a portion of the nucleic acid sequence of the first strand of the nucleic acid is a single nucleotide, and the step of identifying the particular candidate sequence that is most likely to be correct among a plurality of candidate sequences includes the step of identifying the result of a majority vote for a particular query step represented by the plurality of modified records, A method for indicating whether the result of the majority vote indicates that more than half of the multiple sensors detected the sign.
34. The method according to any one of claims 28 to 33, wherein the step of identifying the particular candidate sequence that is most likely to be correct among the plurality of candidate sequences is: A step of determining a metric for each of the aforementioned multiple candidate sequences, wherein each metric provides an indicator of the likelihood of occurrence. The steps include selecting a specific candidate sequence that is most likely to be correct by applying a judgment criterion to each of the aforementioned metrics, Methods that include...
Citation Information
Patent Citations
Sequencing method
JP2004501668A
Chase ligation sequencing method
JP2010539982A
Systems and methods for automated, reusable, and parallel biological reactions
JP2013539978A
Systems and methods using magnetically responsive sensors to determine genetic information
JP2018525980A
Detection of biomaterial
WO2000034522A2