Compositions, methods and systems for detecting nucleotides
By designing engineered nucleotide molecules and nanopore sensor technologies, the problem of signal interference and noise in nucleic acid sequencing is solved, the accuracy and read length of sequencing are improved, and efficient nucleotide detection is achieved.
Patent Information
- Application Number
- CN202380083724.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-10-05
- Filing Date
- 2023-10-04
- Publication Date
- 2025-07-11
AI Technical Summary
When detecting nucleotide sequences, existing nucleic acid sequencing methods have problems with detection signal interference and signal noise. Especially in reversible termination sequencing technology, the cleavage of the tag is not completely possible to leave residues, affecting the accuracy and read length of the sequencing.
An engineered nucleotide molecule is designed, including pentose, bases coupled to pentose, polyphosphate chains, protective groups and identifier parts. The identifier parts coupled to pentose are ensured that the identifier parts can be effectively removed after nucleotide incorporation, reducing interference during sequencing, and using nanopore sensors to detect signal changes.
It realizes reducing sequencing signal interference during nucleic acid sequencing, improves sequencing accuracy and read length, can detect nucleotide incorporation in real time, and enhances the reliability and resolution of sequencing.
Smart Images

Figure CN120303413A_ABST
Abstract
Description
[0001] Cross-reference
[0002] This application claims the benefit of U.S. Provisional Patent Application No. 63 / 413,305, filed on October 5, 2022, which is hereby incorporated by reference in its entirety. Background of the Invention
[0003] Nucleic acid sequencing is the process of determining the nucleotide sequence in a nucleic acid sample. Specific nucleic acid sequence information can be used to discover or identify genetic diseases, diagnose infectious diseases, and develop and monitor treatment methods.
[0004] A variety of nucleic acid sequencing methods have been studied, such as electrophoresis, sequencing by hybridization, mass spectrometry-based methods, ligation sequencing, and sequencing by synthesis (SBS). Summary of the Invention
[0005] The present disclosure provides methods and systems for analyzing a sample (e.g., a nucleic acid sample derived from a biological sample).
[0006] In one aspect, the present disclosure provides an engineered nucleotide molecule comprising: a pentose; a base coupled to the pentose, wherein the base is selected from adenine, guanine, cytosine, thymine, uracil, and analogs thereof; a polyphosphate chain coupled to the pentose, wherein the polyphosphate chain comprises two or more phosphate groups; a protecting group coupled to the pentose, wherein the protecting group is configured to inhibit the coupling of additional nucleotides to the engineered nucleotide molecule; and an identifier moiety coupled to the pentose, wherein the identifier moiety is specific to the engineered nucleotide molecule, and wherein the identifier moiety is directly coupled to the polyphosphate chain.
[0007] In some embodiments of any of the engineered nucleotide molecules disclosed herein, the pentose is deoxyribose.
[0008] In some embodiments of any of the engineered nucleotide molecules disclosed herein, the polyphosphate chain comprises three or more phosphate groups. In some embodiments of any of the engineered nucleotide molecules disclosed herein, the polyphosphate chain comprises four or more phosphate groups. In some embodiments of any of the engineered nucleotide molecules disclosed herein, the polyphosphate chain comprises six phosphate groups.
[0009] In some embodiments of any of the engineered nucleotide molecules disclosed herein, a hydroxyl group is located at the 3'-position of the pentose.
[0010] In some embodiments of any of the engineered nucleotide molecules disclosed herein, a protecting group is coupled to a hydroxyl group of a pentose. In some embodiments of any of the engineered nucleotide molecules disclosed herein, the protecting group comprises an allyl or an azide. In some embodiments of any of the engineered nucleotide molecules disclosed herein, the protecting group is removable from the engineered nucleotide molecule.
[0011] In some embodiments of any of the engineered nucleotide molecules disclosed herein, the identifier moiety is removable from the engineered nucleotide molecule. In some embodiments of any of the engineered nucleotide molecules disclosed herein, the identifier moiety comprises a polynucleotide. In some embodiments of any of the engineered nucleotide molecules disclosed herein, the identifier moiety comprises a non - polynucleotide / non - polypeptide polymer.
[0012] In some embodiments of any of the engineered nucleotide molecules disclosed herein, the polynucleotide has a length of at least about 5 bases. In some embodiments of any of the engineered nucleotide molecules disclosed herein, the polynucleotide has a length of at least about 10 bases. In some embodiments of any of the engineered nucleotide molecules disclosed herein, the polynucleotide has a length of at least about 20 bases. In some embodiments of any of the engineered nucleotide molecules disclosed herein, the polynucleotide has a length of at least about 30 bases.
[0013] In some embodiments of any of the engineered nucleotide molecules disclosed herein, the polynucleotide comprises a polyN selected from polyA, polyT, polyC, polyG, polyU, and variants thereof.
[0014] In another aspect, the present disclosure provides a method for analyzing a target nucleic acid molecule, comprising: (a) providing a complex comprising (i) a target nucleic acid molecule and (ii) a primer nucleic acid molecule that exhibits complementarity to a portion of the target nucleic acid molecule; (b) contacting the complex with an engineered nucleotide molecule to produce a growing strand that is coupled to the primer nucleic acid molecule, wherein the growing strand exhibits sequence complementarity to another portion of the target nucleic acid molecule, and wherein the engineered nucleotide molecule comprises: a pentose; a base coupled to the pentose, wherein the base is selected from adenine, guanine, cytosine, thymine, uracil, and analogs thereof; a polyphosphate chain coupled to the pentose, wherein the polyphosphate chain comprises two or more phosphate groups; a protecting group coupled to the pentose, wherein the protecting group is configured to inhibit the coupling of additional nucleotides to the engineered nucleotide; and an identifier moiety coupled to the pentose, wherein the identifier moiety is specific for the engineered nucleotide, and wherein the identifier moiety is directly coupled to the polyphosphate chain.
[0015] In some embodiments of any of the methods disclosed herein, the method further comprises using a sensor portion to detect (i) contact or (ii) the production of a growing strand.
[0016] In some embodiments of any of the methods disclosed herein, detecting comprises measuring one or more signals indicative of impedance or a change in impedance in the sensor portion upon (i) contact or (ii) the production of a growing strand.
[0017] In some embodiments of any of the methods disclosed herein, the method further comprises contacting a complex with the sensor portion to incorporate at least a portion of an engineered nucleotide molecule as part of the growing strand.
[0018] In some embodiments of any of the methods disclosed herein, the sensor portion comprises a pore or an enzyme. In some embodiments of any of the methods disclosed herein, the sensor portion comprises a pore and an enzyme coupled to the pore. In some embodiments of any of the methods disclosed herein, the pore is part of a nanopore protein. In some embodiments of any of the methods disclosed herein, the pore is part of a solid-state nanopore.
[0019] In some embodiments of any of the methods disclosed herein, the enzyme comprises a polymerase.
[0020] In some embodiments of any of the methods disclosed herein, the method further comprises removing a protecting group from the pentose after (b).
[0021] In some embodiments of any of the methods disclosed herein, the method further comprises coupling an additional nucleotide to the engineered nucleotide after the removal.
[0022] In some embodiments of any of the methods disclosed herein, the removal of the protecting group comprises an enzymatic reaction.
[0023] In some embodiments of any of the methods disclosed herein, the removal of the protecting group comprises a non-enzymatic chemical reaction.
[0024] In some embodiments of any of the methods disclosed herein, the method further comprises removing an identifier portion from the pentose after (b).
[0025] In some embodiments of any of the methods disclosed herein, the pentose is deoxyribose.
[0026] In some embodiments of any of the methods disclosed herein, the polyphosphate chain comprises three or more phosphate groups.
[0027] In some embodiments of any of the methods disclosed herein, the polyphosphate chain comprises four or more phosphate groups.
[0028] In some embodiments of any of the methods disclosed herein, a hydroxyl group is located at the 3'-position of a pentose. In some embodiments of any of the methods disclosed herein, a protecting group is coupled to a hydroxyl group of a pentose. In some embodiments of any of the methods disclosed herein, the protecting group is removable from the pentose. In some embodiments of any of the methods disclosed herein, the protecting group comprises an allyl or an azide.
[0029] In some embodiments of any of the methods disclosed herein, the identifier portion is removable from the engineered nucleotide molecule.
[0030] In some embodiments of any of the methods disclosed herein, the identifier portion comprises a polynucleotide sequence that does not exhibit complementarity to at least a portion of a target nucleic acid molecule.
[0031] In some embodiments of any of the methods disclosed herein, the identifier portion comprises a polynucleotide. In some embodiments of any of the methods disclosed herein, the identifier portion comprises a non-polynucleotide / non-polypeptide polymer. In some embodiments of any of the methods disclosed herein, the polynucleotide has a length of at least about 5 bases. In some embodiments of any of the methods disclosed herein, the polynucleotide has a length of at least about 10 bases. In some embodiments of any of the methods disclosed herein, the polynucleotide has a length of at least about 20 bases. In some embodiments of any of the methods disclosed herein, the polynucleotide has a length of at least about 30 bases. In some embodiments of any of the methods disclosed herein, the polynucleotide comprises a polyN selected from polyA, polyT, polyC, polyG, polyU, and variants thereof.
[0032] Another aspect of the present disclosure provides a non-transitory computer-readable medium comprising machine-executable code that, when executed by one or more computer processors, implements any of the methods described above or elsewhere herein.
[0033] Another aspect of the present disclosure provides a system that includes one or more computer processors and a computer memory coupled thereto. The computer memory contains machine-executable code that, when executed by one or more computer processors, implements any of the methods described above or elsewhere in this document.
[0034] In the following detailed description, which illustrates and describes only illustrative embodiments of the present disclosure, other aspects and advantages of the present disclosure will become apparent to those skilled in the art. As will be appreciated, the present disclosure is capable of other different embodiments, and several details thereof can be modified in various obvious aspects without departing from the present disclosure. Accordingly, the drawings and the detailed description are to be regarded as illustrative in nature and not restrictive.
[0035] Incorporated by reference
[0036] All publications, patents, and patent applications mentioned in this specification are hereby incorporated by reference as if each publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. In the event of conflict between the incorporated publications and patents or patent applications and the disclosure contained herein, this specification is intended to supersede and / or take precedence over any such conflicting material. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] The novel features of the present disclosure are specifically set forth in the appended claims. A better understanding of the features and advantages of the present disclosure will be obtained by reference to the following detailed description of illustrative embodiments in which the principles of the present disclosure are utilized, along with the accompanying drawings (also referred to herein as "figures"). In these
[0038] In the figures:
[0039] Figure 1 Examples of engineered nucleotide molecules according to some embodiments are schematically shown.
[0040] Figure 2 Another example of an engineered nucleotide molecule according to some embodiments is schematically shown.
[0041] Figure 3 A computer system programmed or otherwise configured to implement the methods provided herein is shown.
[0042] Figure 4 An exemplary method of analyzing a target nucleic acid molecule according to some embodiments is shown. DETAILED DESCRIPTION
[0043] Although the present disclosure has shown and described various embodiments of the present invention, it will be readily apparent to those skilled in the art that such embodiments are provided by way of example only. Many changes, alterations, and substitutions may be contemplated by those skilled in the art without departing from the present disclosure. It should be understood that various alternatives to the embodiments of the present invention described herein may be employed.
[0044] Whenever the terms “at least,” “greater than,” or “greater than or equal to” precede the first value in a series of two or more numerical values, the terms “at least,” “greater than,” or “greater than or equal to” apply to each value in the series. For example, greater than or equal to 1, 2, or 3 is equivalent to greater than or equal to 1, greater than or equal to 2, or greater than or equal to 3.
[0045] Whenever the terms “not exceeding,” “less than,” or “less than or equal to” precede the first value in a series of two or more numerical values, the terms “not exceeding,” “less than,” or “less than or equal to” apply to each value in the series. For example, less than or equal to 3, 2, or 1 is equivalent to less than or equal to 3, less than or equal to 2, or less than or equal to 1.
[0046] As used in the specification and claims, unless the context clearly indicates otherwise, the singular forms “a,” “an,” and “the” may include plural referents. For example, the term “sequencing sensor” may include multiple sequencing sensors.
[0047] As used interchangeably herein, the terms “about” and “approximately” may refer to within an acceptable error range of a particular value as determined by one of ordinary skill in the art, which will depend in part on how the value is measured or determined, i.e., the limitations of the measurement system. For example, in accordance with the practice in the art, “about” may mean within 1 standard deviation or greater than 1 standard deviation. Alternatively, “about” may mean a range of up to 20%, up to 10%, up to 5%, or up to 1% of a given value. Or, particularly for biological systems or processes, the term may mean within an order of magnitude of a value, such as within 5-fold or within 2-fold. Where a particular value is described, unless otherwise stated, it should be assumed that the term “about” may mean within the acceptable error range of that particular value.
[0048] As used interchangeably herein, the terms “protecting group,” “blocking group,” and “reversible terminator” generally refer to any atom or group of atoms added to a molecule to prevent unwanted chemical reactions of existing groups in the molecule. For example, to ensure the single incorporation of a complementary nucleotide opposite a base of a target nucleic acid molecule being sequenced by synthetic sequencing (SBS), a protecting group may be added to an engineered nucleotide molecule incorporated into the growing strand (e.g., at the 3′-hydroxyl of the deoxyribose of the engineered nucleotide molecule). The protecting group of the engineered nucleotide molecule may be removed (e.g., by an enzymatic reaction, a chemical reaction, electromagnetic radiation, etc.) under reaction conditions that do not interfere with the integrity of the target nucleic acid molecule being sequenced, either simultaneously with or after the incorporation of the engineered nucleotide molecule into the growing strand (e.g., by an enzyme, such as a polymerase). With the incorporation of the next engineered nucleotide molecule with a protecting group, the SBS sequencing cycle can continue accordingly.
[0049] As used interchangeably herein, the terms “identifier moiety,” “tag,” and “label” generally refer to a molecule that is directly or indirectly detectable and that is directly or indirectly conjugated to a target compound or composition to be detected (e.g., a nucleotide molecule). The identifier moiety can be detectable per se (e.g., radioisotope labeled or fluorescently labeled), or, if an enzyme label, can catalyze a chemical change in a detectable substrate compound or composition. In some cases, the presence or absence of the identifier moiety can be detected by measuring an electrochemical property (e.g., capacitance, resistance, impedance, conductivity, voltage, etc.) of an electrochemical cell (e.g., a nanopore sensor) upon addition or removal of the identifier moiety. The identifier moiety can be suitable for small-scale detection, or more suitable for high-throughput screening. Thus, non-limiting examples of the identifier moiety can include radioisotopes, fluorescent dyes, chemiluminescent compounds, bioluminescent compounds, dyes, polynucleotides, polypeptides (e.g., enzymes, fluorescent proteins, etc.), and non-polynucleotide / non-polypeptide polymers. The identifier moiety can be simply detected. Alternatively or additionally, the identifier moiety can be quantified.
[0050] As used interchangeably herein, the terms "polynucleotide", "oligonucleotide", "oligomer", and "nucleic acid" generally refer to polymeric forms of nucleotides of any length, which can be deoxyribonucleotides, ribonucleotides, or analogs thereof, and can be in single-stranded, double-stranded, or multi-stranded form. Polynucleotides can be exogenous or endogenous to a cell. Polynucleotides can exist in a cell-free environment. Polynucleotides can be genes or fragments thereof. Polynucleotides can be deoxyribonucleic acid (DNA). Polynucleotides can be ribonucleic acid (RNA). Polynucleotides can have any three-dimensional structure and can perform any function. Polynucleotides can contain one or more analogs (e.g., altered backbone, sugar, or nucleobase). If present, the nucleotide structure can be modified either before or after polymer assembly. Some non-limiting examples of analogs include: 5-bromouracil, peptide nucleic acid, xeno nucleic acid, morpholino, locked nucleic acid, glycol nucleic acid, threose nucleic acid, dideoxynucleotide, cordycepin, 7-deaza-GTP, fluorophores (e.g., rhodamine or fluorescein linked to sugar), thiol-containing nucleotides, biotin-linked nucleotides, fluorescent base analogs, CpG islands, methyl-7-guanosine, methylated nucleotides, inosine, thiouridine, pseudouridine, dihydrouridine, brachidial, and wybutosine. Non-limiting examples of polynucleotides include coding or non-coding regions of genes or gene fragments, loci defined by linkage analysis, exons, introns, messenger RNA (mRNA), transfer RNA (tRNA), ribosomal RNA (rRNA), short interfering RNA (siRNA), short hairpin RNA (shRNA), microRNA (miRNA), ribozymes, complementary DNA (cDNA, such as double-stranded cDNA (dd-cDNA) or single-stranded cDNA (ss-cDNA)), circulating tumor DNA (ctDNA), damaged DNA, recombinant polynucleotides, branched polynucleotides, plasmids, vectors, isolated DNA of any sequence, isolated RNA of any sequence, cell-free polynucleotides (including cell-free DNA (cfDNA) and cell-free RNA (cfRNA)), nucleic acid probes (e.g., fluorescence in situ hybridization (FISH) probes), and primers. The sequence of nucleotides can be interrupted by non-nucleotide components. Polynucleotides can contain one or more modified nucleotides, such as methylated nucleotides and nucleotide analogs. The sequence of nucleotides can be interrupted by non-nucleotide components. Polynucleotides can be further modified after polymerization, such as by conjugation with a labeling component.
[0051] As used interchangeably herein, the terms "complementary," "complementary sequence," "complementary to," and "complementarity" generally refer to a sequence that is fully complementary to and hybridizable with a given sequence. A sequence that hybridizes to a given nucleic acid is referred to as the "complementary sequence" or "reverse complementary sequence" of the given molecule, provided that its base sequence over a given region is capable of binding complementarily to the base sequence of its binding partner such that, for example, adenine (A)-thymine (T), A-uracil (U), guanine (G)-cytosine (C), and G-U base pairs are formed. Generally, a first sequence that is hybridizable to a second sequence can hybridize specifically or selectively to the second sequence such that, during a hybridization reaction, it preferably hybridizes to the second sequence or group of second sequences as compared to hybridization to non-target sequences (e.g., is more thermodynamically stable under given conditions such as stringent conditions commonly used in the art). Generally, hybridizable sequences share a degree of sequence complementarity over all or part of their respective lengths, such as complementarity between about 25% and about 100%, including at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, and 100% sequence complementarity. The respective lengths can include regions having at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 11, at least about 12, at least about 13, at least about 14, at least about 15, at least about 16, at least about 17, at least about 18, at least about 19, at least about 20, at least about 21, at least about 22, at least about 23, at least about 24, at least about 25, at least about 30, at least about 35, at least about 40, at least about 45, at least about 50, or more nucleotides.For purposes such as assessing the percentage of complementarity, sequence identity can be measured by any suitable alignment algorithm, including but not limited to the Needleman-Wunsch algorithm (see, for example, the EMBOSS Needle aligner available at www.ebi.ac.uk / Tools / psa / emboss_needle / nucleotide.html, optionally using default settings), the BLAST algorithm (see, for example, the BLAST alignment tool available at blast.ncbi.nlm.nih.gov / Blast.cgi, optionally using default settings), or the Smith-Waterman algorithm (see, for example, the EMBOSS Water aligner available at www.ebi.ac.uk / Tools / psa / emboss_water / nucleotide.html, optionally using default settings). Any suitable parameters of the selected algorithm, including default parameters, can be used to evaluate the optimal alignment.
[0052] Complementarity can be complete or substantial / sufficient. Complete complementarity between two nucleic acids can mean that the two nucleic acids can form a duplex, where each base in the duplex binds to a complementary base through Watson-Crick pairing. Substantial or sufficient complementarity can mean that the sequence in one strand is not completely and / or not perfectly complementary to the sequence in the opposite strand, but under a set of hybridization conditions (e.g., salt concentration and temperature), sufficient binding occurs between the bases on the two strands to form a stable hybrid complex. Such conditions can be predicted by using the sequence and standard mathematical calculations to predict the Tm of the hybridizing strands or by empirically determining the Tm using conventional methods.
[0053] As used herein, the term "hybridization" generally refers to a reaction in which one or more polynucleotides react to form a complex that is stabilized by hydrogen bonds between the bases of the nucleotide residues. The hydrogen bonds can occur through Watson Crick base pairing, Hoogstein binding, or in any other sequence-specific manner according to base complementarity. The complex can comprise two strands forming a duplex structure, three or more strands forming a multistranded complex, a self-hybridizing strand, or any combination thereof. The hybridization reaction can constitute a step in a broader process, such as initiating PCR or the enzymatic cleavage of polynucleotides by endonucleases. A second sequence complementary to a first sequence can be referred to as the "complementary sequence" of the first sequence. The term "hybridizable" applied to a polynucleotide generally refers to the ability of the polynucleotide to form a complex that is stabilized by hydrogen bonds between the bases of the nucleotide residues in a hybridization reaction.
[0054] As used herein, the term "polymerase" generally refers to an enzyme (e.g., natural or synthetic) capable of catalyzing a polymerization reaction. Examples of polymerases can include nucleic acid polymerases (e.g., deoxyribonucleic acid (DNA) polymerases or ribonucleic acid (RNA) polymerases) and transcriptases (e.g., reverse transcriptase). A polymerase can be a polymerization enzyme. The term "DNA polymerase" generally refers to an enzyme capable of catalyzing the polymerization of DNA.
[0055] As used herein, the term "sequencing" generally refers to a procedure for determining the order in which nucleotides occur in a target nucleotide sequence. Sequencing methods can include high-throughput sequencing, such as next-generation sequencing (NGS). Sequencing can be whole-genome sequencing or targeted sequencing. Sequencing can be single-molecule sequencing or massively parallel sequencing. Next-generation sequencing methods can obtain millions of sequences in a single run. In one example, one or more nanopore sequencing methods can be used for sequencing, such as, for example, synthesis sequencing, ligation sequencing, or cleavage sequencing.
[0056] As used herein, the term "nanopore" generally refers to a pore, channel, or pathway formed or otherwise provided in a membrane. The membrane can be an organic membrane, such as a lipid bilayer, or a synthetic membrane, such as a membrane formed from a polymeric material such as a protein nanopore. The membrane can be a solid-state membrane (e.g., a silicon substrate). The nanopore can be disposed adjacent to or in proximity to a sensing circuit or an electrode of a sensing circuit (e.g., a complementary metal-oxide semiconductor (CMOS) or a field-effect transistor (FET) circuit) and coupled thereto. The nanopore can be part of a sensing circuit. The nanopore can have a characteristic width or diameter, e.g., from about 0.1 nanometers (nm) to 1000 nm. The nanopore can be a biological nanopore, a solid-state nanopore, a hybrid bio-solid-state nanopore, a variant thereof, or a combination thereof. Examples of biological nanopores include, but are not limited to, OmpG from the genera Escherichia coli (E. coli), Salmonella, Shigella, and Pseudomonas, as well as α-hemolysin from Staphylococcus aureus (S. aureus), MspA from Mycobacterium smegmatis (M. smegmatis), functional variants thereof, or combinations thereof. Sequencing can include forward sequencing and / or reverse sequencing. Examples of solid-state nanopores include, but are not limited to, silicon nitride, silicon oxide, graphene, molybdenum sulfide, functional variants thereof, or combinations thereof. Solid-state nanopores can be fabricated by high-energy beams, imprinting (e.g., nanoimprinting), laser ablation, chemical etching, plasma etching (e.g., oxygen plasma etching), etc.
[0057] As used interchangeably herein, the terms “nanopore sequencing” and “nanopore-based sequencing” generally refer to methods for determining the sequence of a polynucleotide by means of a nanopore. In some cases, the sequence of the polynucleotide can be determined in a template-dependent manner.
[0058] As used herein, the term “real-time” generally refers to an event (e.g., an operation, a process, a measurement, a detection, etc.) that occurs almost immediately or within a short time after another event (e.g., addition of a nucleobase, growing strand generation, etc.), where the short time is, for example, at least about 0.0001 milliseconds (ms), at least about 0.0005 ms, at least about 0.001 ms, at least about 0.005 ms, at least about 0.01 ms, at least about 0.05 ms, at least about 0.1 ms, at least about 0.5 ms, at least about 1 ms, at least about 5 ms, at least about 0.01 seconds, at least about 0.05 seconds, at least about 0.1 seconds, at least about 0.5 seconds, at least about 1 second or longer. In some cases, a real-time event can occur almost immediately after another event or within a short time after another event, where the short time is, for example, at most about 1 second, at most about 0.5 seconds, at most about 0.1 seconds, at most about 0.05 seconds, at most about 0.01 seconds, at most about 5 ms, at most about 1 ms, at most about 0.5 ms, at most about 0.1 ms, at most about 0.05 ms, at most about 0.01 ms, at most about 0.005 ms, at most about 0.001 ms, at most about 0.0005 ms, at most about 0.0001 ms or shorter.
[0059] As used herein, the term “sample” generally refers to any sample that can contain one or more components (e.g., nucleic acid molecules) for processing or analysis. The sample can be a biological sample. The sample can be a cell or tissue sample. The sample can be a cell-free sample, such as blood (e.g., whole blood), plasma, serum, sweat, saliva, or urine. The sample can be obtained in vivo or cultured in vitro.
[0060] The term “substituted” refers to a functional group as described herein, such as an alkyl or a hydrocarbyl group, in which at least one bond to a hydrogen atom is replaced by a bond to a non-hydrogen atom or a non-carbon atom, provided that normal valences are maintained and the substitution results in a stable compound. Substituted groups also include groups in which one or more bonds to a carbon atom or a hydrogen atom are replaced by one or more bonds to a heteroatom (including double or triple bonds). Non-limiting examples of substituents include functional groups as described herein, such as N, to form -CN.
[0061] Reversible termination sequencing technology is an SBS method that detects the sequence of a nucleic acid template by stepwise extension of a growing nucleic acid strand. Reversible termination sequencing can include modifying nucleotide molecules with: (i) attaching a removable label (e.g., a fluorescent label) to the base of the nucleotide molecule via a linker, and / or (ii) attaching a reversible terminator (e.g., a protecting group) to the 3’O position of the sugar. However, such modification on the base may leave a trace after at least partial cleavage of the linker carrying the label, which may interfere with the current polymerization step or any subsequent polymerization step, and / or result in a shorter read length. In some examples, the cleavage of the label may be incomplete, leaving a residual label on the growing nucleic acid strand, which may generate background noise when detecting signals from subsequently added nucleotide molecules, especially in consensus sequencing. Accordingly, there is a recognized unmet need for engineered nucleotide molecules comprising a detectable label for reversible termination sequencing, wherein after incorporation of a portion of the engineered nucleotide molecule (e.g., the sugar coupled to the base) into a polynucleotide sequence (e.g., the growing strand produced by a polymerase during SBS), the engineered nucleotide molecule can become substantially (e.g., completely) free of (i) the detectable label and (ii) any linker used to conjugate the detectable label to the engineered nucleotide molecule.
[0062] Aspects of the present disclosure provide an engineered nucleotide molecule, a composition thereof, and a method of using the same, wherein the engineered nucleotide molecule has a protecting group and an identifier moiety. In some embodiments, the engineered nucleotide molecule can comprise a protecting group (e.g., at the 3’O position of the sugar) coupled to the sugar (e.g., a pentose) of the engineered nucleotide molecule and an identifier moiety linked to the polyphosphate chain of the engineered nucleotide molecule, thereby enabling incorporation of only one engineered nucleotide molecule and detection of signals from one engineered nucleotide molecule in each cycle (e.g., during SBS). In some embodiments, to incorporate the engineered nucleotide molecule into the growing strand (or primer nucleic acid molecule), the 3’OH of the growing strand can attack the α-phosphate of the polyphosphate chain of the engineered nucleotide molecule to be incorporated, thereby forming a phosphodiester bond and releasing other polyphosphate groups comprising the identifier moiety. Accordingly, linking the identifier moiety to the phosphate about to be released during the polymerization step (e.g., released naturally by the same mechanism as the polymerization step) can enhance sequencing or can not interfere with sequencing, by, for example, (i) having little to no residual trace on the remaining portion of the engineered nucleotide molecule, and / or (ii) having little to no residual identifier moiety on the growing strand. Accordingly, any signal detected during incorporation of the engineered nucleotide molecule can be attributed solely to the newly added engineered nucleotide molecule and not to any previously added nucleobase.
[0063] I. Engineered Nucleotide Molecule
[0064] On the one hand, the present disclosure provides an engineered nucleotide molecule, a composition thereof, a method of using the same (e.g., for sequencing a target nucleic acid molecule), and a system for analyzing a target nucleic acid molecule. The engineered nucleotide molecule can include a sugar (e.g., pentose), a base coupled to the sugar, a polyphosphate chain coupled to the sugar, a protecting group coupled to the sugar, and an identifier moiety coupled to the sugar. The identifier moiety can be coupled to the sugar through the polyphosphate chain. Alternatively, the identifier moiety can be coupled to a different portion of the engineered nucleotide molecule (e.g., coupled to the base).
[0065] In some embodiments, the sugar can be, for example, pentose, hexose, glucose, fructose, or galactose. In some embodiments, the sugar can be pentose, such as ribose, deoxyribose, arabinofuranose, lyxofuranose, or xylofuranose. In some embodiments, the pentose can be ribose (e.g., for growing an RNA strand). In some embodiments, the pentose can be deoxyribose (e.g., for growing a DNA strand).
[0066] In some embodiments, the base can be selected from adenine (A), guanine (G), cytosine (C), thymine (T), uracil (U), and analogs thereof. Non-limiting examples of such base analogs can include 5-aza-uracil, 2-thio-5-aza-uracil, 2-thio-uracil, 5-hydroxy-uracil, 3-methyl-uracil, 5-carboxymethyl-uracil, 5-propynyl-uracil, 5-tauromethyl-uracil, 5-tauromethyl-2-thio-uracil, 1-tauromethyl-4-thio-uracil, 5-methyl-uracil, dihydrouracil, 2-thio-dihydro-uracil, 5-bromouracil, 2-methoxy-uracil, 2-methoxy-4-thio-uracil, 5-aza-cytosine, 3-methyl-cytosine, N4-acetyl-cytosine, 5-formyl-cytosine, N4-methyl-cytosine, 5-hydroxymethyl-cytosine, pyrrolo-cytosine, 2-thio-cytosine, 2-thio-5-methyl-cytosine, 2-methoxy-cytosine, 2-methoxy-5-methyl-cytosine, hypoxanthine, 1-methyl-hypoxanthine, 7-methyl-hypoxanthine, deazahypoxanthine, 7-deaza-guanine, 7-deaza-8-aza-guanine, 6-thio-guanine, 6-thio-7-deaza-guanine, 6-thio-7-deaza-8-aza-guanine, 7-methyl-guanine, 6-thio-7-methyl-guanine, 6-methoxy-guanine, 1-methyl-guanine, N2-methyl-guanine, N2,N2-dimethyl-guanine, 8-oxo-guanine, 7-methyl-8-oxo-guanine, 1-methyl-6-thio-guanine, N2-methyl-6-thio-guanine, N2,N2-dimethyl-6-thio-guanine, 2-aminopurine, 2,6-diaminopurine, 7-deaza-adenine, 7-deaza-8-aza-adenine, 2-aminoadenine, 7-deaza-2-aminopurine, 7-deaza-8-aza-2-aminopurine, 7-deaza-2,6-diaminopurine, 6-mercaptopurine, 7-deaza-8-aza-2,6-diaminopurine, 1-methyl-adenine, N6-methyl-adenine, N6-isopentenyl-adenine, N6-(cis-hydroxyisopentenyl)-adenine, 2-methylthio-N6-(cis-hydroxyisopentenyl)adenine, N6-glycylcarbamoyl-adenine, N6-threonylcarbamoyl-adenine, 2-methylthio-N6-threonylcarbamoyl-adenine, N6,N6-dimethyladenine, 7-methyl-adenine, 2-methylthio-adenine, 2-methoxy-adenine, 4-O-ethylthymine, pyrazolopyrimidine, and any substituted analogs thereof.
[0067] In some embodiments, the protecting group can be coupled through the hydroxyl group of the pentose. In some embodiments, the hydroxyl group can be located at the 3'-position of the pentose. Alternatively or additionally, the hydroxyl group can also be located at the 2'-position of the pentose (e.g., for ribose). The protecting group can be any suitable group capable of coupling to the pentose and can be cleaved by any suitable reaction to regenerate the hydroxyl group. Non-limiting examples of the protecting group can include allyl, azide, azo, amine, cyanoethyl, dimethylethyl, dimethylacetamidine, azidomethyl, phenoxyacetyl, alkyldithiomethyl, methoxyacetyl, acetyl, tosylate, phosphate, nitrate, 4-methoxytetrahydropyranyl, tetrahydropyranyl, 4-methoxytetrahydropyranyl, tetrahydropyranyl, 5-methyltetrahydrofuranyl, 5-methyltetrahydropyranyl, tetrahydropyranyl, tetrahydrofuranyl, methoxytetrahydropyranyl, 2-nitrobenzyl or any substituted analog thereof. When the engineered nucleotide molecule is coupled to the growing nucleic acid chain, the protecting group can inhibit the coupling of additional nucleotides to the growing nucleic acid chain.
[0068] In some embodiments, the protecting group of the terminal engineered nucleotide molecule on the nucleic acid chain can be cleaved to regenerate the hydroxyl group on the engineered nucleotide molecule (e.g., on the sugar of the engineered nucleotide molecule) to allow subsequent addition of another nucleotide molecule (e.g., another engineered nucleotide molecule as disclosed herein) to the growing nucleic acid chain. The protecting group can be cleaved by any suitable reaction, such as an enzymatic reaction (e.g., by Bacillus stearothermophilus DNA polymerase I), a non-enzymatic chemical reaction (e.g., with phosphite, sodium dithionite, palladium-catalyzed reaction), a thermal reaction (e.g., in a polymerase chain reaction (PCR) buffer containing 50 mM KCl, 1.5 mM MgCl2, 20 mM Tris (pH 8.4, 25 °C)), or a photocleavage reaction (e.g., exposure to electromagnetic radiation, such as ultraviolet (UV) light).
[0069] In some embodiments, the protecting group of the pentose hydroxyl group (e.g., the hydroxyl group of the 3'-OH of deoxyribose) of the engineered nucleotide molecule disclosed herein can be cleaved by an enzyme different from the polymerase that extends (e.g., polymerizes) the growing nucleic acid chain. Alternatively, the protecting group can be cleaved by the same polymerase that effects the extension.
[0070] In some embodiments, when an engineered nucleotide molecule disclosed herein is sufficiently close to the sensor portion of an electrochemical cell (e.g., nanopore sensor, nanoporeless sensor, etc.) disclosed herein, the size of its identifier portion can be large enough to induce a change in the electrochemical properties (e.g., capacitance, resistance, impedance, conductivity, voltage, etc.) of the electrochemical cell. In some cases, the change in the electrochemical properties can occur and be detectable before, during, or after the identifier portion is released from the engineered polynucleotide molecule. For example, the engineered nucleotide molecule can be brought to the nanopore sensor by a polymerase extending a growing nucleic acid strand, and this complex of the engineered nucleotide molecule with the polymerase, the growing nucleic acid strand, and / or the target nucleic acid molecule to be analyzed can be sufficient to induce a change in the electrochemical properties (e.g., a change in the capacitance of the nanopore sensor). Thus, in some cases, since the size of the identifier portion can determine the detection or analysis of the target nucleic acid molecule, the identifier portion may not need to be a fluorescent molecule.
[0071] In some embodiments, the identifier portion can comprise a polynucleotide sequence that does not exhibit complementarity to at least a portion of the target nucleic acid molecule. In some embodiments, the polynucleotide sequence can exhibit sequence identity to the polynucleotide sequence of the target nucleic acid molecule that is less than or equal to about 90%, less than or equal to about 80%, less than or equal to about 70%, less than or equal to about 60%, less than or equal to about 50%, less than or equal to about 40%, less than or equal to about 30%, less than or equal to about 20%, less than or equal to about 10%, less than or equal to about 9%, less than or equal to about 8%, less than or equal to about 7%, less than or equal to about 6%, less than or equal to about 5%, less than or equal to about 4%, less than or equal to about 3%, less than or equal to about 2%, less than or equal to about 1%, less than or equal to about 0.5%, or less than or equal to about 0.1%.
[0072] The polynucleotide sequence of the identifier portion may have a length of at least about 5 bases, at least about 10 bases, at least about 15 bases, at least about 20 bases, at least about 25 bases, at least about 30 bases, at least about 35 bases, at least about 40 bases, at least about 45 bases, at least about 50 bases, at least about 55 bases, at least about 60 bases, at least about 65 bases, at least about 70 bases, at least about 75 bases, at least about 80 bases, at least about 85 bases, at least about 90 bases, at least about 95 bases, at least about 100 bases, at least about 110 bases, at least about 120 bases, at least about 130 bases, at least about 140 bases, at least about 150 bases, at least about 160 bases, at least about 170 bases, at least about 180 bases, at least about 190 bases, at least about 200 bases or more bases. The length of the polynucleotide sequence of the identifier portion can be at most about 200 bases, at most about 190 bases, at most about 180 bases, at most about 170 bases, at most about 160 bases, at most about 150 bases, at most about 140 bases, at most about 130 bases, at most about 120 bases, at most about 110 bases, at most about 100 bases, at most about 95 bases, at most about 90 bases, at most about 85 bases, at most about 80 bases, at most about 75 bases, at most about 70 bases, at most about 65 bases, at most about 60 bases, at most about 55 bases, at most about 50 bases, at most about 45 bases, at most about 40 bases, at most about 35 bases, at most about 30 bases, at most about 25 bases, at most about 20 bases, at most about 15 bases, at most about 10 bases, at most about 5 bases or less.
[0073] In some embodiments, the polynucleotide sequence of the identifier portion may comprise polyN (e.g., T40, A40, A10, or T10). The polyN may be characterized by: (i) two or more identical bases (e.g., TTTT) or (ii) two or more consecutive groups of identical bases (e.g., a polynucleotide, such as ATATATAT). The group of identical bases may comprise at least two different bases, at least three different bases, at least four different bases, at least five different bases, or more. The group of identical bases may comprise at most five different bases, at most four different bases, at most three different bases, or at most two different bases. The length of the group of identical bases may be at least about 2 bases, at least about 3 bases, at least about 4 bases, at least about 5 bases, at least about 6 bases, at least about 7 bases, at least about 8 bases, at least about 9 bases, at least about 10 bases, or more. The length of the group of identical bases may be at most about 10 bases, at most about 9 bases, at most about 8 bases, at most about 7 bases, at most about 6 bases, at most about 5 bases, at most about 4 bases, at most about 3 bases, or at most about 2 bases. Non-limiting examples of polyN may include polyA, polyT, polyC, polyG, polyU, or a polynucleotide (e.g., polyAT, polyCG, polyAG, polyCT, polyAC, polyTG, polyAU). The length of polyN may be at least about 5 bases, at least about 10 bases, at least about 15 bases, at least about 20 bases, at least about 25 bases, at least about 30 bases, at least about 35 bases, at least about 40 bases, at least about 45 bases, at least about 50 bases, at least about 55 bases, at least about 60 bases, at least about 65 bases, at least about 70 bases, at least about 75 bases, at least about 80 bases, at least about 85 bases, at least about 90 bases, at least about 95 bases, at least about 100 bases, at least about 110 bases, at least about 120 bases, at least about 130 bases, at least about 140 bases, at least about 150 bases, at least about 160 bases. Bases, at least about 170 bases, at least about 180 bases, at least about 190 bases, at least about 200 bases, or more.The length of polyN can be at most about 200 bases, at most about 190 bases, at most about 180 bases, at most about 170 bases, at most about 160 bases, at most about 150 bases, at most about 140 bases, at most about 130 bases, at most about 120 bases, at most about 110 bases, at most about 100 bases, at most about 95 bases, at most about 90 bases, at most about 85 bases, at most about 80 bases, at most about 75 bases, at most about 70 bases, at most about 65 bases, at most about 60 bases, at most about 55 bases, at most about 50 bases, at most about 45 bases, at most about 40 bases, at most about 35 bases, at most about 30 bases, about 25 bases, at most about 20 bases, at most about 15 bases, at most about 10 bases, at most about 5 bases or fewer.
[0074] In some embodiments, the identifier portion can include a radioisotope, a fluorescent label, a chemiluminescent label, a bioluminescent label, and an enzyme label.Non-limiting examples of the identifier portion (e.g., a fluorescent label) can include fluorescein, 5-carboxyfluorescein (FAM), 2′7′-dimethoxy-4′5-dichloro-6-carboxyfluorescein (JOE), rhodamine, 6-carboxyrhodamine (R6G), N,N,N′,N′-tetramethyl-6-carboxyrhodamine (TAMRA), 6-carboxy-X-rhodamine (ROX), 4-(4′-dimethylaminophenylazo)benzoic acid (DABCYL), cascade blue, Oregon green, Texas red, cyanine, and 5-(2′-aminoethyl)aminonaphthalene-1-sulfonic acid (EDANS), [R6G]dUTP, [TAMRA]dUTP, [R110]dCTP, [R6G]dCTP, [TAMRA]dCTP, [JOE]ddATP, [R6G]ddATP, [FAM]ddCTP, [R110]ddCTP, [TAMRA]ddGTP, [ROX]ddTTP, [dR6G]ddATP, [dR110]ddCTP, [dTAMRA]ddGTP, and [dROX]ddTTP available from Perkin Elmer, Foster City, Calif.; FluoroLink deoxynucleotides, FluoroLink Cy3-dCTP, FluoroLink Cy5-dCTP, FluoroLink Fluor X-dCTP, FluoroLink Cy3-dUTP, and FluoroLink Cy5-dUTP available from Amersham, Arlington Heights, Ill.; fluorescein-15-dATP, fluorescein-12-dUTP, tetramethyl-rhodamine-6-dUTP, IR770-9-dATP, fluorescein-12-ddUTP, fluorescein-12-UTP, and fluorescein-15-2′-dATP available from Boehringer Mannheim, Indianapolis, Ind.; and chromosome-labeled nucleotides, BODIPY-FL-14-UTP, BODIPY-FL-4-UTP, BODIPY-TMR-14-UTP, BODIPY-TMR-14-dUTP, BODIPY-TR-14-UTP, BODIPY-TR-14-dUTP, cascade blue-7-UTP, cascade blue-7-dUTP, fluorescein-12-UTP, fluorescein-12-dUTP, Oregon green 488-5-dUTP, rhodamine green-5-UTP, rhodamine green-5-dUTP, tetramethylrhodamine-6-UTP, tetramethylrhodamine-6-dUTP, Texas red-5-UTP, Texas red-5-dUTP, and Texas red-12-dUTP available from Molecular Probes, Eugene, Oreg.
[0075] In some embodiments, the identifier portion may comprise a polymer that is not a polypeptide or polynucleotide. In some embodiments, the polymer is substantially soluble under aqueous conditions. Non-limiting examples of polymers (e.g., polymer chains or portions thereof that do not comprise polynucleotide sequences or polypeptide sequences) include polyethylene glycol, polyvinylimine, polyacrylamide, polyacrylic acid, polyvinyl alcohol, or ionomers. In some embodiments, the polymer may be a homopolymer. In some embodiments, the polymer may be a copolymer.
[0076] In some embodiments, the molecular weight of the identifier portion can be from about 50 Daltons (Da) to about 500 Da, from about 50 Da to about 1 kilodalton (kDa), from about 50 Da to about 2 kDa, from about 50 Da to about 5 kDa, from about 50 Da to about 10 kDa, from about 50 Da to about 15 kDa, from about 50 Da to about 20 kDa, from about 50 Da to about 25 kDa, from about 50 Da to about 30 kDa, from about 50 Da to about 35 kDa, from about 50 Da to about 40 kDa, from about 50 Da to about 50 kDa, from about 50 Da to about 60 kDa, from about 50 Da to about 70 kDa, from about 50 Da to about 80 kDa, from about 50 Da to about 90 kDa, from about 50 Da to about 100 kDa, from about 100 Da to about 10 kDa, from about 100 Da to about 15 kDa, from about 100 Da to about 20 kDa, from about 100 Da to about 25 kDa, from about 100 Da to about 30 kDa, from about 100 Da to about 35 kDa, from about 100 Da to about 40 kDa, from about 100 Da to about 50 kDa, from about 100 Da to about 60 kDa, from about 100 Da to about 70 kDa, from about 100 Da to about 80 kDa, from about 100 Da to about 90 kDa, from about 100 Da to about 100 kDa, from about 200 Da to about 10 kDa, from about 200 Da to about 15 kDa, from about 200 Da to about 20 kDa, from about 200 Da to about 25 kDa, from about 200 Da to about 30 kDa, from about 200 Da to about 35 kDa, from about 200 Da to about 40 kDa, from about 200 Da to about 50 kDa, from about 200 Da to about 60 kDa, from about 200 Da to about 70 kDa, from about 200 Da to about 80 kDa, from about 200 Da to about 90 kDa, from about 200 Da to about 100 kDa, from about 500 Da to about 10 kDa, from about 500 Da to about 15 kDa, from about 500 Da to about 20 kDa, from about 500 Da to about 25 kDa, from about 500 Da to about 30 kDa, from about 500 Da to about 35 kDa, from about 500 Da to about 40 kDa, from about 500 Da to about 50 kDa, from about 500 Da to about 60 kDa, from about 500 Da to about 70 kDa, from about 500 Da to about 80 kDa, from about 500 Da to about 90 kDa, from about 500 Da to about 100 kDa, from about 1 kDa to about 10 kDa, from about 1 kDa to about 15 kDa, from about 1 kDa to about 20 kDa, from about 1 kDa to about 25 kDa, from about 1 kDa to about 30 kDa, from about 1 kDa to about 35 kDa, from about 1 kDa to about 40 kDa, from about 1 kDa to about 50 kDa, from about 1 kDa to about 60 kDa, from about 1 kDa to about 70 kDa, from about 1 kDa to about 80 kDa, from about 1 kDa to about 90 kDa, from about 1 kDa to about 100 kDa, from about 2 kDa to about 10 kDa, from about 2 kDa to about 15 kDa,from about 2 kDa to about 20 kDa, from about 2 kDa to about 25 kDa, from about 2 kDa to about 30 kDa, from about 2 kDa to about 35 kDa, from about 2 kDa to about 40 kDa, from about 2 kDa to about 50 kDa, from about 2 kDa to about 60 kDa, from about 2 kDa to about 70 kDa, from about 2 kDa to about 80 kDa, from about 2 kDa to about 90 kDa, from about 2 kDa to about 100 kDa, from about 5 kDa to about 10 kDa, from about 5 kDa to about 15 kDa, from about 5 kDa to about 20 kDa, from about 5 kDa to about 25 kDa, from about 5 kDa to about 30 kDa, from about 5 kDa to about 35 kDa, from about 5 kDa to about 40 kDa, from about 5 kDa to about 50 kDa, from about 5 kDa to about 60 kDa, from about 5 kDa to about 70 kDa, from about 5 kDa to about 80 kDa, from about 5 kDa to about 90 kDa, from about 5 kDa to about 100 kDa, from about 10 kDa to about 15 kDa, from about 10 kDa to about 20 kDa, from about 10 kDa to about 25 kDa, from about 10 kDa to about 30 kDa, from about 10 kDa to about 35 kDa, from about 10 kDa to about 40 kDa, from about 10 kDa to about 50 kDa, from about 10 kDa to about 60 kDa, from about 10 kDa to about 70 kDa, from about 10 kDa to about 80 kDa, from about 10 kDa to about 90 kDa, from about 10 kDa to about 100 kDa, from about 1 kDa to about 25 kDa, from about 20 kDa to about 30 kDa, from about 20 kDa to about 35 kDa, from about 20 kDa to about 40 kDa, from about 20 kDa to about 50 kDa, from about 20 kDa to about 60 kDa, from about 20 kDa to about 70 kDa, from about 20 kDa to about 80 kDa, from about 20 kDa to about 90 kDa, from about 20 kDa to about 100 kDa, from about 30 kDa to about 40 kDa, from about 30 kDa to about 50 kDa, from about 30 kDa to about 60 kDa, from about 30 kDa to about 70 kDa, from about 30 kDa to about 80 kDa, from about 30 kDa to about 90 kDa, from about 30 kDa to about 100 kDa, from about 50 kDa to about 60 kDa, from about 50 kDa to about 70 kDa, from about 50 kDa to about 80 kDa, from about 50 kDa to about 90 kDa, from about 50 kDa to about 100 kDa, from about 60 kDa to about 70 kDa, from about 60 kDa to about 80 kDa, from about 60 kDa to about 90 kDa, from about 60 kDa to about 100 kDa, from about 70 kDa to about 80 kDa, from about 70 kDa to about 90 kDa, from about 70 kDa to about 100 kDa, from about 80 kDa to about 90 kDa, from about 80 kDa to about 100 kDa, and from about 90 kDa to about 100 kDa.
[0077] In some embodiments, an engineered nucleotide molecule can initially include an identifier portion, e.g., directly coupled to at least one additional portion of the engineered nucleotide molecule, such as a phosphate group on a polyphosphate chain. In some cases, the identifier portion can be coupled to the additional portion of the engineered nucleotide molecule via a linker. In some cases, when the identifier portion is excised from the engineered nucleotide molecule (e.g., during polymerization), at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% or substantially about 100% of the identifier portion or a combination of the identifier portion and the linker (e.g., as measured by molecular weight) can be excised or removed from the engineered nucleotide molecule, leaving at most about 20%, at most about 15%, at most about 10%, at most about 9%, at most about 8%, at most about 7%, at most about 6%, at most about 5%, at most about 4%, at most about 3%, at most about 2%, at most about 1%, or substantially about 0% of the identifier portion or a combination of the identifier portion and the linker in the engineered nucleotide molecule.
[0078] In some embodiments, an identifier portion as disclosed herein can be coupled to a phosphate of a polyphosphate chain via a linker portion. The linker portion can be hydrophobic or hydrophilic. Non-limiting examples of the linker portion include esters, ethers, thioethers, ethylene glycols, alkylene groups, alkenylene groups, alkynylene groups, heteroalkylene groups, cycloalkylene groups, heterocycloalkylene groups, arylene groups, heteroarylene groups, and heterocycloalkylidene groups, where any of the groups can be substituted or unsubstituted. Alternatively, the identifier portion can be directly coupled to the phosphate of the polyphosphate chain without a separate linker.
[0079] In some embodiments, the length of the polyphosphate chain can be at least about 2 phosphates, at least about 3 phosphates, at least about 4 phosphates, at least about 5 phosphates, at least about 6 phosphates, at least about 7 phosphates, at least about 8 phosphates, at least about 9 phosphates, at least about 10 phosphates, at least about 15 phosphates, at least about 20 phosphates or more. The length of the polyphosphate chain can be at most about 20 phosphates, at most about 15 phosphates, at most about 10 phosphates, at most about 9 phosphates, at most about 8 phosphates, at most about 7 phosphates, at most about 6 phosphates, at most about 5 phosphates, at most about 4 phosphates, or at most about 3 phosphates.
[0080] In some embodiments, the engineered nucleotide molecule can comprise a single phosphate group or moiety (e.g., a non-polyphosphate chain) coupled to a pentose, and the identifier moiety can be directly coupled to the single phosphate moiety. In some cases, such engineered nucleotide molecules can be sufficient to facilitate a polymerization reaction in which the identifier moiety is excised and at least the pentose coupled to a base is added to a growing nucleic acid site, e.g., a growing nucleic acid chain.
[0081] In some embodiments, the polyphosphate chain can comprise at least (i) a first phosphate closest to the sugar of the engineered nucleotide molecule (e.g., an alpha-phosphate or α-phosphate), and (ii) a second phosphate second closest to the sugar and directly coupled to the α-phosphate (e.g., a beta-phosphate or β-phosphate). The identifier moiety can be coupled (e.g., directly conjugated) to the β-phosphate or any subsequent phosphate coupled thereto. For example, a subsequent phosphate group can be a third phosphate (e.g., a gamma-phosphate or γ-phosphate). In some embodiments, the identifier moiety can be coupled to the terminal phosphate group of the polyphosphate chain (e.g., coupled to the γ-phosphate of a triphosphate). Alternatively or additionally, the identifier moiety can be coupled to a non-terminal phosphate group of the polyphosphate chain (e.g., coupled to the β-phosphate of a triphosphate). At least one identifier moiety can be coupled to one of the phosphate groups of the polyphosphate chains disclosed herein, and one of the phosphate groups can include a β-phosphate, γ-phosphate, δ phosphate (or phosphate at position 4), ε phosphate (or phosphate at position 5), ζ phosphate (or phosphate at position 6), η phosphate (or phosphate at position 7), θ phosphate (or phosphate at position 8), ι phosphate (or phosphate at position 9), κ phosphate (or phosphate at position 10), phosphate at position 10, phosphate at position 11, phosphate at position 12, phosphate at position 13, phosphate at position 14, phosphate at position 15, phosphate at position 20, or any subsequent available phosphate.
[0082] In some embodiments, an identifier moiety that can be released (e.g., cleaved) from the engineered nucleotide molecule can be detected, e.g., by a sensor moiety (e.g., a nanopore sensor) disclosed herein. For example, after release, detecting the released identifier moiety (e.g., when it is near a nanopore or upon entry into a nanopore sensor) can be used to determine whether the engineered nucleotide molecule has been fully incorporated into the growing chain. Alternatively, accurate detection (e.g., sequence calling) of such incorporation may not require separate detection of the released identifier moiety (e.g., in addition to any measurements made during incorporation of the engineered nucleotide molecule into the growing chain).
[0083] In one aspect, the present disclosure provides a method of analyzing a target nucleic acid molecule using an engineered nucleotide molecule and a sensor moiety (e.g., a sequencing sensor).
[0084] In some embodiments, a method of analyzing a target nucleic acid molecule includes: a) providing a complex comprising the target nucleic acid molecule and a primer nucleic acid molecule that exhibits complementarity to a portion of the target nucleic acid molecule; and b) contacting the complex with an engineered nucleotide molecule to produce a growing strand coupled to the primer nucleic acid molecule, wherein the growing strand exhibits sequence complementarity to an additional portion of the target nucleic acid molecule, and wherein the engineered nucleotide molecule comprises a pentose sugar, a base coupled to the pentose sugar, a polyphosphate chain coupled to the pentose sugar, a protecting group coupled to the pentose sugar, and an identifier moiety coupled to the pentose sugar.
[0085] In some embodiments, the method further comprises (c) using a sensor moiety to obtain sequence information of at least a portion of the growing strand to analyze the additional portion of the target nucleic acid molecule.
[0086] Figure 4 An exemplary method of analyzing a target nucleic acid molecule is shown. At operation 401, method 400 includes providing a complex comprising the target nucleic acid molecule and a primer nucleic acid molecule. At operation 402, method 400 includes contacting the complex with an engineered nucleotide molecule to produce a growing strand. At operation 403, method 400 includes using a sensor moiety to obtain sequence information of at least a portion of the growing strand.
[0087] In some embodiments, the engineered nucleotide molecule can include: (i) a first type of engineered nucleotide molecule that includes a first type of identifier moiety coupled to the pentose sugar via a first type of linker; (ii) a second type of engineered nucleotide molecule that includes a second type of identifier moiety coupled to the pentose sugar via a second type of linker; (iii) a third type of engineered nucleotide molecule that includes a third type of identifier moiety coupled to the pentose sugar via a third type of linker; and (iv) a fourth type of engineered nucleotide molecule that includes a fourth type of identifier moiety coupled to the pentose sugar via a fourth type of linker.
[0088] In some embodiments, the first type of identifier moiety, the second type of identifier moiety, the third type of identifier moiety, and the fourth type of identifier moiety can be the same type of identifier moiety.
[0089] In some embodiments, the first type of linker, the second type of linker, the third type of linker, and the fourth type of linker can be the same type of linker.
[0090] In some embodiments, (c) comprises detecting the identifier moiety when the identifier moiety associates with the polymerase. In some embodiments, (c) comprises detecting the identifier moiety when the identifier moiety is cleaved from the polyphosphate chain and a growing nucleic acid strand is produced. In some embodiments, (c) comprises detecting the identifier moiety when the identifier moiety is in the vicinity of the sensor moiety. In some embodiments, (c) comprises detecting the identifier moiety when the identifier moiety translocates to and through the sensor moiety.
[0091] In some embodiments, the time between detecting the identifier moiety and (i) the association of the identifier moiety with the polymerase, (ii) the cleavage of the identifier moiety, (iii) the production of the growing nucleic acid strand, (iv) bringing the identifier moiety to the vicinity of the sensor moiety, or (v) the translocation of the identifier moiety to and through the sensor moiety is at most about 5 minutes (min), at most about 4 min, at most about 3 min, at most about 2 min, at most about 1 min, at most about 50 seconds (s), at most about 40 s, at most about 30 s, at most about 20 s, at most about 10 s, at most about 1 s, at most about 900 milliseconds (ms), at most about 800 ms, at most about 700 ms, at most about 600 ms, at most about 500 ms, at most about 400 ms, at most about 300 ms, at most about 200 ms, at most about 100 ms, at most about 50 ms, at most about 10 ms, at most about 1 ms, at most about 900 microseconds (μs), at most about 800 μs, at most about 700 μs, at most about 600 μs, at most about 500 μs, at most about 400 μs, at most about 300 μs, at most about 200 μs, at most about 100 μs, at most about 50 μs, at most about 10 μs, at most about 1 μs, at most about 900 nanoseconds (ns), at most about 800 ns, at most about 700 ns, at most about 600 ns, at most about 500 ns, at most about 400 ns, at most about 300 ns, at most about 200 ns, at most about 100 ns, at most about 90 ns, at most about 80 ns, at most about 70 ns, at most about 60 ns, at most about 50 ns, at most about 40 ns, at most about 30 ns, at most about 20 ns, at most about 10 ns, at most about 9 ns, at most about 8 ns, at most about 7 ns, at most about 6 ns, at most about 5 ns, at most about 4 ns, at most about 3 ns, at most about 2 ns, at most about 1 ns or less.
[0092] In some embodiments, detection of the identifier portion occurs substantially in real time relative to: (i) association of the identifier portion with a polymerase; (ii) cleavage of the identifier portion; (iii) generation of a growing nucleic acid strand; (iv) bringing the identifier portion near a sensor portion; or (v) translocation of the identifier portion into and through a sensor portion. In some embodiments, detection of the identifier portion occurs immediately after or within a short time after: (i) association of the identifier portion with a polymerase; (ii) cleavage of the identifier portion; (iii) generation of a growing nucleic acid strand; (iv) bringing the identifier portion near a sensor portion; or (v) translocation of the identifier portion into and through a sensor portion. In some embodiments, the short time period is at most about 1 ms, at most about 900 μs, at most about 800 μs, at most about 700 μs, at most about 600 μs, at most about 500 μs, at most about 400 μs, at most about 300 μs, at most about 200 μs, at most about 100 μs, at most about 50 μs, at most about 10 μs, at most about 1 μs, at most about 900 ns, at most about 800 ns, at most about 700 ns, at most about 600 ns, at most about 500 ns, at most about 400 ns, at most about 300 ns, at most about 200 ns, at most about 100 ns, at most about 90 ns, at most about 80 ns, at most about 70 ns, at most about 60 ns, at most about 50 ns, at most about 40 ns, at most about 30 ns, at most about 20 ns, at most about 10 ns, at most about 9 ns, at most about 8 ns, at most about 7 ns, at most about 6 ns, at most about 5 ns, at most about 4 ns, at most about 3 ns, at most about 2 ns, at most about 1 ns or less.
[0093] In one aspect, the present disclosure provides a system for analyzing a target nucleic acid molecule. The system can include a sensor portion configured to detect one or more signals indicative of an electrical property (e.g., capacitance, resistance, impedance, conductivity, voltage, or a change thereof) in the sensor portion when at least a portion of the target molecule binds to or approaches at least a portion of the sensor portion. In some cases, the electrical property can be impedance or a change in impedance. The one or more signals can be used to analyze or identify the target molecule.
[0094] The system can include at least one sensor portion disclosed herein. The system can include at least 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30, at least about 40, at least about 50, at least about 60, at least about 70, at least about 80, at least about 90, at least about 100, at least about 200, at least about 300, at least about 400, at least about 500, at least about 600, at least about 700, at least about 800, at least about 900, at least about 1000, or more sensor portions. The system can include at most about 1,000, at most about 900, at most about 800, at most about 700, at most about 600, at most about 500, at most about 400, at most about 300, at most about 200, at most about 100, at most about 90, at most about 80, at most about 70, at most about 60, at most about 50, at most about 40, at most about 30, at most about 20, at most about 10, at most about 9, at most about 8, at most about 7, at most about 6, at most about 5, at most about 4, at most about 3, at most about 2, or fewer sensor portions.
[0095] The detected signal indicating impedance or impedance change in the sensor caused by the target molecule can be a single measurement. Alternatively, the detected signal can be the median or average of multiple measurements.
[0096] When one or more signals indicative of impedance or impedance changes in the sensor portion are detected, at least a portion of the target molecule can bind to the binding portion of the sensor portion. The binding portion can be configured to bind to at least a portion of the target molecule (e.g., nucleotides, amino acids, small molecules, ions, etc.). The sensor portions disclosed herein can include at least one binding portion. The sensor portion can include at least 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30, at least about 40, at least about 50, at least about 60, at least about 70, at least about 80, at least about 90, at least about 100, at least about 200, at least about 300, at least about 400, at least about 500, at least about 600, at least about 700, at least about 800, at least about 900, at least about 1,000 or more binding portions. The sensor portion can include at most about 1,000, at most about 900, at most about 800, at most about 700, at most about 600, at most about 500, at most about 400, at most about 300, at most about 200, at most about 100, at most about 90, at most about 80, at most about 70, at most about 60, at most about 50, at most about 40, at most about 30, at most about 20, at most about 10, at most about 9, at most about 8, at most about 7, at most about 6, at most about 5, at most about 4, at most about 3, at most about 2 or fewer binding portions.
[0097] In some embodiments, the sensor portion includes a pore and / or an enzyme. In some embodiments, the pore is part of a nanopore protein. In some embodiments, the pore is part of a solid-state nanopore. In some embodiments, the enzyme is coupled to the pore. In some embodiments, the enzyme is a polymerase.
[0098] In some embodiments, the sensor portion can be configured to measure the fluorescence signal of a fluorescent label.
[0099] When the complex contacts an engineered nucleotide molecule, the identifier portion on the engineered nucleotide molecule can interact with the sensor portion and generate a signal. When the identifier portion associates with a polymerase, the identifier portion can detect the identifier portion. Alternatively, when the identifier portion is cleaved from a polyphosphate chain and a growing nucleic acid chain is produced, the sensor portion can measure the signal. In some embodiments, the identifier portion can be large enough such that when the identifier portion is near the sensor portion, the sensor portion can detect its presence. In some embodiments, when the identifier portion translocates into and through the sensor portion, the identifier portion can be detected.
[0100] When an engineered nucleotide molecule is added to a growing nucleic acid strand, a protecting group can inhibit the coupling of additional nucleotides to the growing nucleic acid strand. Thus, the sensor portion can determine the nucleotide type on the target nucleic acid molecule within one cycle. In some embodiments, the protecting group on the terminal nucleotide molecule on the growing strand can be cleaved to regenerate the hydroxyl group on the nucleotide for subsequent addition of another nucleotide molecule. In some embodiments, the protecting group can be cleaved by any suitable reaction, such as an enzymatic reaction (e.g., by Bacillus stearothermophilus DNA polymerase I), a non-enzymatic chemical reaction (e.g., with phosphine, sodium dithionite, palladium-catalyzed reaction), a thermal reaction (e.g., in a PCR buffer containing 50 mM KCl, 1.5 mM MgCl2, 20 mM Tris (pH 8.4, 25 °C)), or a photocleavage reaction (e.g., upon exposure to electromagnetic radiation such as ultraviolet (UV) light).
[0101] In some embodiments, the protecting group is excised from the engineered nucleotide molecule at least about 1 ns, at least about 5 ns, at least about 10 ns, at least about 50 ns, at least about 100 ns, at least about 500 ns, at least about 1 μs, at least about 10 μs, at least about 50 μs, at least about 100 μs, at least about 500 μs, at least about 1 ms, at least about 10 ms, at least about 50 ms, at least about 100 ms, at least about 500 ms, at least about 1 s, at least about 10 s, at least about 50 s, at least about 100 s, or longer after the engineered nucleotide molecule to which the protecting group is attached is added to the growing nucleic acid strand.
[0102] In some embodiments, the protecting group is cleaved from the engineered nucleotide molecule at least about 1 ns, at least about 5 ns, at least about 10 ns, at least about 50 ns, at least about 100 ns, at least about 500 ns, at least about 1 μs, at least about 10 μs, at least about 50 μs, at least about 100 μs, at least about 500 μs, at least about 1 ms, at least about 10 ms, at least about 50 ms, at least about 100 ms, at least about 500 ms, at least about 1 s, at least about 10 s, at least about 50 s, at least about 100 s, or longer after the sensor portion detects the identifier portion. In some embodiments, the protecting group is excised from the engineered nucleotide molecule when the sensor portion detects the identifier portion.
[0103] After the hydroxyl group is regenerated, step a) can be restarted, thereby allowing determination of the next nucleotide on the target nucleic acid molecule. This process can be repeated until the full length or a desired length sequence of the target nucleic acid molecule is determined.
[0104] II. Sample
[0105] In some embodiments, the target nucleic acid molecule can be derived from a sample of interest (e.g., a biological sample from a subject).
[0106] As disclosed herein, the sample for analysis can comprise multiple polynucleotides. The polynucleotides can be single-stranded DNA, double-stranded DNA, or a combination thereof. The polynucleotides can comprise genomic DNA, genomic cDNA, cell-free DNA, cell-free cDNA, or any combination of the foregoing.
[0107] The polynucleotides can include cell-free DNA, circulating tumor DNA, genomic DNA, and DNA from formalin-fixed and paraffin-embedded (FFPE) samples. In some instances, DNA extracted from FFPE samples may be damaged, and such damaged DNA can be repaired using available FFPE DNA repair kits. The sample can comprise any suitable DNA and / or cDNA sample, such as urine, feces, blood, saliva, tissue, biopsy, body fluid, or tumor cells.
[0108] The polynucleotide sample can be from any suitable source. For example, the sample can be obtained from a patient, animal, plant, or the environment, such as a natural or artificial atmosphere, water system, soil, atmospheric pathogen collection system, subsurface sediment, groundwater, or sewage treatment plant.
[0109] The polynucleotides from the sample can include one or more different polynucleotides, e.g., DNA, RNA, ribosomal RNA (rRNA), transfer RNA (tRNA), microRNA (miRNA), messenger RNA (mRNA), fragments of any of the foregoing, or any combination of any of the foregoing. The sample can comprise DNA. The sample can comprise genomic DNA. The sample can comprise mitochondrial DNA, chloroplast DNA, plasmid DNA, bacterial artificial chromosome, yeast artificial chromosome, oligonucleotide tag, or any combination of any of the foregoing.
[0110] The polynucleotides can be single-stranded, double-stranded, or a combination thereof. The polynucleotides can be single-stranded polynucleotides, in which double-stranded polynucleotides may or may not be present.
[0111] The starting amount of polynucleotide in the sample can be, for example, less than about 50 ng, such as less than about 45 ng, less than about 40 ng, less than about 35 ng, less than about 30 ng, less than about 25 ng, less than about 20 ng, less than about 15 ng, less than about 10 ng, less than about 5 ng, less than about 4 ng, less than about 3 ng, less than about 2 ng, less than about 1 ng, less than about 0.5 ng, less than about 0.1 ng or less. The starting amount of polynucleotide in the sample can be, for example, more than about 0.1 ng, such as more than about 0.5 ng, more than about 1 ng, more than about 2 ng, more than about 3 ng, more than about 4 ng, more than about 5 ng, more than about 10 ng, more than about 15 ng, more than about 20 ng, more than about 25 ng, more than about 30 ng, more than about 35 ng, more than about 40 ng, more than about 45 ng, more than about 50 ng or more. The amount of starting polynucleotide can be, for example, from about 0.1 ng to about 100 ng, from about 1 ng to about 75 ng, from about 5 ng to about 50 ng or from about 10 ng to about 20 ng.
[0112] The polynucleotide in the sample can be single-stranded when obtained, or can be made single-stranded by treatment (e.g., denaturation). The polynucleotide can be subjected to subsequent steps (e.g., circularization and amplification) without an extraction step and / or without a purification step. For example, a fluid sample can be processed to remove cells without an extraction step to produce a purified fluid sample and a cell sample, and then the polynucleotide can be isolated from the purified fluid sample. A variety of methods for isolating polynucleotides can be used, such as by precipitation or non-specific binding to a substrate, followed by washing the substrate to release the bound polynucleotide. If the polynucleotide is isolated from the sample without a cell extraction step, the polynucleotide will be mainly extracellular or "cell-free" polynucleotides, which may correspond to dead or damaged cells. The identity of such cells can be used to characterize, for example, the cells or cell populations from which they are derived in a microbial community.
[0113] The sample can be from a subject. The subject can be any suitable organism, including, for example, plants, animals, fungi, protists, prokaryotes without nuclei, viruses, mitochondria, and chloroplasts. The sample polynucleotide can be isolated from the subject, such as a cell sample, a tissue sample, a body fluid sample, or an organ sample or a cell culture derived from any of these, including, for example, a cultured cell line, a biopsy, a blood sample, a buccal swab, or a fluid sample containing cells such as saliva. The subject can be an animal, such as a cow, a pig, a mouse, a rat, a chicken, a cat, a dog, or a mammal, such as a human. The sample can contain tumor cells, such as in a sample from a tumor tissue of a subject.
[0114] Other examples of sample sources can include blood, urine, feces, nasal cavity, lungs, intestine, other body fluids or excreta, derivatives thereof, or combinations thereof.
[0115] Samples from a single individual can be divided into multiple separate samples, such as 2, 3, 4, 5, 6, 7, 8, 9, 10 or more separate samples, which are independently subjected to the methods of the present disclosure, for example, analyzed in duplicate, triplicate, quadruplicate or more. When the sample is from a subject, the reference sequence can also be derived from the subject, such as a consensus sequence from the analyzed sample or a polynucleotide sequence from another sample or tissue of the same subject. For example, ctDNA mutations in a blood sample can be analyzed, and cellular DNA from another sample of the subject (such as a buccal or skin sample) can be analyzed to determine the reference sequence.
[0116] The polynucleotide can be extracted from the cells in the sample or not extracted from the cells in the sample according to any suitable method.
[0117] The plurality of polynucleotides can include cell-free polynucleotides, such as cell-free DNA (cfDNA) or circulating tumor DNA (ctDNA). Cell-free DNA circulates in healthy and diseased individuals. CfDNA (ctDNA) from tumors is not limited to any particular cancer type but appears to be a common finding in different malignancies. The concentration of cell-free circulating DNA in the plasma of control subjects may be lower compared to patients with or suspected of having a disease. In one example, the concentration of cell-free circulating DNA in the plasma can be, for example, 14 ng / mL to 18 ng / mL in control subjects and 18 ng / mL to 318 ng / mL in neoplastic patients.
[0118] Apoptosis and necrotic cell death may contribute to the production of cell-free circulating DNA in body fluids. For example, significantly elevated levels of circulating DNA can be observed in the plasma of patients with prostate cancer and patients with other prostate diseases such as benign prostatic hyperplasia and prostatitis. In addition, circulating tumor DNA may be present in fluids from the primary organ of the tumor. In one example, breast cancer detection can be achieved in ductal lavage fluid; colorectal cancer detection can be achieved in feces; lung cancer detection can be achieved in sputum, and prostate cancer detection can be achieved in urine or ejaculate. Cell-free DNA can be obtained from a variety of sources. Exemplary sources can be blood samples of a subject. However, cfDNA or other fragmented DNA can come from a variety of other sources, including, for example, urine and fecal samples can be sources of cfDNA including ctDNA.
[0119] III. Nanopore
[0120] A system for analyzing a target nucleic acid molecule can include a reaction chamber that contains one or more nanopore devices. The nanopore devices can be individually addressable nanopore devices. An individually addressable nanopore can be individually readable. An individually addressable nanopore can be individually writable. An individually addressable nanopore can be individually readable and individually writable. The system can include one or more computer processors for facilitating sample preparation and the various operations of the present disclosure, such as polynucleotide sequencing. The processor can be coupled to the nanopore devices.
[0121] The nanopore devices can include a plurality of individually addressable sensing electrodes. Each sensing electrode can include a membrane adjacent to the electrode and one or more nanopores in the membrane. The nanopores can be in a membrane such as a lipid bilayer that is adjacent to or in sensing proximity to an electrode that is part of or coupled to an integrated circuit. The nanopores can be associated with a single electrode and a sensing integrated circuit or with multiple electrodes and a sensing integrated circuit. The nanopores can include solid-state nanopores. The nanopore devices can include reference electrodes.
[0122] In some cases, the sensor portion can be configured to detect one or more signals indicative of impedance or impedance changes (e.g., between a sensing electrode and a reference electrode) when at least a portion of an engineered nucleotide molecule binds to at least a portion of the sensor portion (e.g., a sensing electrode). Alternatively, the sensor portion can be configured to detect one or more signals indicative of impedance or impedance changes (e.g., between a sensing electrode and a reference electrode) when at least a portion of an engineered nucleotide molecule is unbound but in proximity to at least a portion of the sensor portion (e.g., a sensing electrode).
[0123] The sensor portion of the present disclosure can be configured to detect more signals indicative of impedance or impedance change (e.g., between a sensing electrode and a reference electrode) when the distance between (i) at least a portion of an engineered nucleotide molecule and (ii) the sensing electrode is at least about 0.1 nm, at least about 0.5 nm, at least about 1 nm, at least about 2 nm, at least about 3 nm, at least about 4 nm, at least about 5 nm, at least about 6 nm, at least about 7 nm, at least about 8 nm, at least about 9 nm, at least about 10 nm, at least about 20 nm, at least about 30 nm, at least about 40 nm, at least about 50 nm, at least about 60 nm, at least about 70 nm, at least about 80 nm, at least about 90 nm, at least about 100 nm, at least about 200 nm, at least about 300 nm, at least about 400 nm, at least about 500 nm, at least about 600 nm, at least about 700 nm, at least about 800 nm, at least about 900 nm, at least about 1 mm, at least about 2 mm, at least about 3 mm, at least about 4 mm, at least about 5 mm, at least about 6 mm, at least about 7 mm, at least about 8 mm, at least about 9 mm, at least about 10 mm, at least about 20 mm, at least about 30 mm, at least about 40 mm, at least about 50 mm, at least about 60 mm, at least about 70 mm, at least about 80 mm, at least about 90 mm, at least about 100 mm, at least about 200 mm, at least about 300 mm, at least about 400 mm, at least about 500 mm, at least about 600 mm, at least about 700 mm, at least about 800 mm, at least about 900 mm, at least about 1,000 mm or greater.The sensor portion of the present disclosure can be configured to detect more signals indicative of impedance or impedance change (e.g., between a sensing electrode and a reference electrode) when the distance between (i) at least a portion of an engineered nucleotide molecule and (ii) the sensing electrode is at most about 1,000 μm, at most about 900 μm, at most about 800 μm, at most about 700 μm, at most about 600 μm, at most about 500 μm, at most about 400 μm, at most about 300 μm, at most about 200 μm, at most about 100 μm, at most about 90 μm, at most about 80 μm, at most about 70 μm, at most about 60 μm, at most about 50 μm, at most about 40 μm, at most about 30 μm, at most about 20 μm, at most about 10 μm, at most about 9 μm, at most about 8 μm, at most about 7 μm, at most about 6 μm, at most about 5 μm, at most about 4 μm, at most about 3 μm, at most about 2 μm, at most about 1 μm, at most about 900 nm, at most about 800 nm, at most about 700 nm, at most about 600 nm, at most about 500 nm, at most about 400 nm, at most about 300 nm, at most about 200 nm, at most about 100 nm, at most about 90 nm, at most about 80 nm, at most about 70 nm, at most about 60 nm, at most about 50 nm, at most about 40 nm, at most about 30 nm, at most about 20 nm, at most about 10 nm, at most about 9 nm, at most about 8 nm, at most about 7 nm, at most about 6 nm, at most about 5 nm, at most about 4 nm, at most about 3 nm, at most about 2 nm, at most about 1 nm, at most about 0.5 nm, at most about 0.1 nm, or less.
[0124] The sensor portion of the present disclosure can be configured to detect more signals indicative of impedance or impedance change (e.g., between a sensing electrode and a reference electrode) when an engineered nucleotide molecule is within a predetermined space that is close to or adjacent to the sensing electrode. The predetermined space can be characterized as having at least about 0.1 nm 2 , at least about 0.5 nm 2 , at least about 1 nm, at least about 2 nm 2 , at least about 3 nm 2 , at least about 4 nm 2 , at least about 5 nm 2 , at least about 6 nm 2 , at least about 7 nm 2 , at least about 8 nm 2 , at least about 9 nm 2 , at least about 10 nm 2 , at least about 20 nm 2 , at least about 30 nm 2 , at least about 40 nm 2 , at least about 50 nm 2 , at least about 60 nm 2, at least about 70 nm 2 , at least about 80 nm 2 , at least about 90 nm 2 , at least about 100 nm 2 , at least about 200 nm 2 , at least about 300 nm 2 , at least about 400 nm 2 , at least about 500 nm 2 , at least about 600 nm 2 , at least about 700 nm 2 , at least about 800 nm 2 , at least about 900 nm 2 , at least about 1 μm 2 , at least about 2 μm 2 , at least about 3 μm 2 , at least about 4 μm 2 , at least about 5 μm 2 , at least about 6 μm 2 , at least about 7 μm 2 , at least about 8 μm 2 , at least about 9 μm 2 , at least about 10 μm 2 , at least about 20 μm 2 , at least about 30 μm 2 , at least about 40 μm 2 , at least about 50 μm 2 , at least about 60 μm 2 , at least about 70 μm 2 , at least about 80 μm 2 , at least about 90 μm 2 , at least about 100 μm 2 , at least about 200 μm 2 , at least about 300 μm 2 , at least about 400 μm 2 , at least about 500 μm 2 , at least about 600 μm 2 , at least about 700 μm 2 , at least about 800 μm 2 , at least about 900 μm 2 , at least about 1,000 μm 2 or a larger volume. The predetermined space may be characterized by having at most about 1,000 μm 2 , at most about 900 μm 2 , at most about 800 μm 2 , at most about 700 μm 2 , at most about 600 μm 2 , at most about 500 μm2 , up to about 400 μm 2 , up to about 300 μm 2 , up to about 200 μm 2 , up to about 100 μm 2 , up to about 90 μm 2 , up to about 80 μm 2 , up to about 70 μm 2 , up to about 60 μm 2 , up to about 50 μm 2 , up to about 40 μm 2 , up to about 30 μm 2 , up to about 20 μm 2 , up to about 10 μm 2 , up to about 9 μm 2 , up to about 8 μm 2 , up to about 7 μm 2 , up to about 6 μm 2 , up to about 5 μm 2 , up to about 4 μm 2 , up to about 3 μm 2 , up to about 2 μm 2 , up to about 1 μm 2 , up to about 900 nm 2 , up to about 800 nm 2 , up to about 700 nm 2 , up to about 600 nm 2 , up to about 500 nm 2 , up to about 400 nm 2 , up to about 300 nm 2 , up to about 200 nm 2 , up to about 100 nm 2 , up to about 90 nm 2 , up to about 80 nm 2 , up to about 70 nm 2 , up to about 60 nm 2 , up to about 50 nm 2 , up to about 40 nm 2 , up to about 30 nm 2 , up to about 20 nm 2 , up to about 10 nm 2 , up to about 9 nm 2 , up to about 8 nm 2 , up to about 7 nm 2 , about 6 nm 2 , up to about 5 nm 2 , up to about 4 nm 2 , up to about 3 nm 2 , up to about 2 nm2 、 up to about 1 nm 2 、 up to about 0.5 nm 2 、 up to about 0.1 nm 2 or smaller volume.
[0125] The devices and systems for use in the methods provided by the present disclosure can accurately detect individual nucleotide incorporation events, such as when a nucleotide is incorporated into a growing strand complementary to a template. Enzymes, such as DNA polymerase, RNA polymerase, and / or ligase, can be involved in incorporating nucleotides into a growing polynucleotide chain. Enzymes such as polymerase can produce a polynucleotide chain.
[0126] The added nucleotide can be complementary to the corresponding template polynucleotide chain that hybridizes to the growing strand. The nucleotide can include a tag or tag substance coupled to any position of the nucleotide, including but not limited to the phosphate of the nucleotide such as the γ-phosphate, the sugar, or the nitrogenous base moiety. In some cases, during nucleotide tag incorporation, the tag is detected when the tag associates with the polymerase. The tag can be detected until it translocates through the nanopore after nucleotide incorporation and subsequent cleavage and / or release of the tag. The nucleotide incorporation event can release the tag from the nucleotide, and the tag passes through the nanopore and is detected. The tag can be released by the polymerase or cleaved / released in any suitable manner, including but not limited to cleavage by an enzyme located near the polymerase. In this way, since a unique tag is released from each type of nucleotide (i.e., adenine, cytosine, guanine, thymine, or uracil), the incorporated base (i.e., A, C, G, T, or U) can be identified. In non-release nucleotide incorporation events, the tag coupled to the incorporated nucleotide is detected by means of the nanopore. In some instances, the tag can move through or near the nanopore and is detected by means of the nanopore.
[0127] The methods and systems of the present disclosure can achieve the detection of polynucleotide incorporation events, for example, at a resolution of at least 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30, at least about 40, at least about 50, at least about 100, at least about 500, at least about 1000, at least about 5000, at least about 10000, at least about 50000, at least about 100000 or more nucleotide bases within a given time period. For example, a nanopore device can be used to detect individual polynucleotide incorporation events, each event associated with an individual nucleic acid base. In other instances, a nanopore device can be used to detect events associated with multiple bases. For example, the signal sensed by the nanopore device can be a combined signal from at least about 2, at least about 3, at least about 4, or at least about 5 bases.
[0128] In some sequencing methods, the tag does not pass through the nanopore. The tag can be detected by the nanopore and exit the nanopore without passing through it, such as exiting from the opposite direction of the tag entering the nanopore. The sequencing device can be configured to actively eject the tag from the nanopore.
[0129] In some sequencing methods, the tag is not released after a nucleotide incorporation event. The nucleotide incorporation event can present the tag to the nanopore without releasing the tag. The tag can be detected by the nanopore without being released. The tag can be attached to the nucleotide by a linker that is long enough to present the tag to the nanopore for detection.
[0130] In some embodiments, nucleic acids can be sequenced by sequential addition and / or removal of engineered nucleotide molecules.
[0131] When a nucleotide incorporation event occurs, the nanopore can be detected in real time. Enzymes (such as DNA polymerase) attached to or near the nanopore can facilitate the passage of the polynucleotide through or near the nanopore. The nucleotide incorporation event, or the incorporation of multiple nucleotides, can release or present one or more tags that can be detected by the nanopore. Detection can occur when the tag passes through or near the nanopore, when the tag resides in the nanopore and / or when presenting the tag to the nanopore. In some cases, enzymes attached to or near the nanopore can assist in detecting the tag when incorporating one or more nucleotides.
[0132] The nanopore can be formed or otherwise embedded in a membrane disposed adjacent to a sensing electrode of a sensing circuit (such as an integrated circuit). The integrated circuit can be an application-specific integrated circuit (ASIC). The integrated circuit can be a field-effect transistor or a complementary metal-oxide semiconductor (CMOS). The sensing circuit can be located in a chip or other device having the nanopore, or outside the chip or device, such as in an off-chip configuration.
[0133] When the nucleic acid or tag passes through or near the nanopore, the sensing circuit detects an electrical signal associated with the nucleic acid or tag. The nucleic acid can be a subunit of a larger chain. The tag can be a byproduct of a nucleotide incorporation event or other interaction between the tagged nucleic acid and a substance near or at the nanopore (such as an enzyme that cleaves the tag from the nucleic acid). The tag can remain attached to the nucleotide. The detected signal can be collected and stored in a storage location and then used to construct a nucleic acid sequence. The collected signal can be processed to interpret any anomalies in the detected signal, such as errors.
[0134] Nanopores can be used for indirect sequencing of polynucleotides, in some cases by electrical detection. Indirect sequencing can be any method in which the nucleotides incorporated into the growing strand do not pass through the nanopore. The polynucleotide can pass at any suitable distance from and / or proximal to the nanopore, which in some cases is a distance such that a tag released from a nucleotide incorporation event is detected in the nanopore.
[0135] By-products of nucleotide incorporation events can be detected by the nanopore. A nucleotide incorporation event refers to the incorporation of a nucleotide into a growing polynucleotide chain. The by-products may be associated with the incorporation of a given type of nucleotide. Nucleotide incorporation events can be catalyzed by an enzyme (such as DNA polymerase) and use base pair interactions with a template molecule to select available nucleotides for incorporation at each position.
[0136] Tagged nucleotides or nucleotide analogs can be used to sequence a nucleic acid sample. In some examples, methods for sequencing a nucleic acid molecule include: (a) incorporating (e.g., polymerizing) tagged nucleotides, where the tag associated with an individual nucleotide is released upon incorporation, and (b) detecting the released tag with a nanopore. In some cases, the method further includes guiding a tag attached to or released from an individual nucleotide through the nanopore. The released or attached tag can be guided by any suitable technique, which in some cases is by means of an enzyme (or molecular motor) and / or a voltage difference across the pore. Alternatively, the released or attached tag can be guided through the nanopore without using an enzyme. For example, as described herein, the tag can be guided by a voltage difference across the nanopore.
[0137] Labels can be detected by means of a nanopore device having at least one nanopore in a membrane. The label can associate with an individual labeled nucleotide during the incorporation of the individual labeled nucleotide. The nanopore device can detect the label associated with the individual labeled nucleotide during the incorporation process. Labeled nucleotides, whether incorporated into a growing nucleic acid chain or not, can be detected, determined, or distinguished by the nanopore device within a given time period, in some cases by means of the electrodes of the nanopore device and / or the nanopore. The time for the nanopore device to detect the label can be shorter than (significantly shorter in some cases) the time the label and / or the nucleotide coupled to the label is held by an enzyme (such as an enzyme that facilitates the incorporation of nucleotides into a nucleic acid chain (e.g., polymerase)). During the time period when the incorporated labeled nucleotide associates with the enzyme, the label can be detected multiple times by the electrodes. For example, during the time period when the incorporated labeled nucleotide associates with the enzyme, the label can be detected by the electrodes at least 1 time, at least about 2 times, at least about 3 times, at least about 4 times, at least about 5 times, at least about 6 times, at least about 7 times, at least about 8 times, at least about 9 times, at least about 10 times, at least about 20 times, at least about 30 times, at least about 40 times, at least about 50 times, at least about 60 times, at least about 70 times, at least about 80 times, at least about 90 times, at least about 100 times, at least about 200 times, at least about 300 times, at least about 400 times, at least about 500 times, at least about 1000 times, at least about 10,000 times, at least about 100,000 times, or at least about 1,000,000 times.
[0138] Labels associated with individual nucleotides can be detected by the nanopore without being released from the nucleotide upon incorporation. The label can be detected without being released from the incorporated nucleotide during the synthesis of a nucleic acid chain complementary to the target strand. The label can be attached to the nucleotide via a linker such that the label is presented to the nanopore (e.g., the label hangs in at least a portion of the nanopore or otherwise extends through at least a portion of the nanopore). The length of the linker can be long enough to allow the label to extend to or through at least a portion of the nanopore. In some cases, the label is presented to (i.e., moved into) the nanopore by a voltage difference. Other ways of presenting the label into the pore can also be suitable (e.g., using an enzyme, a magnet, an electric field, a pressure difference). In some cases, no active force is applied to the label (i.e., the label diffuses into the nanopore).
[0139] A chip for sequencing a nucleic acid sample can include a plurality of individually addressable nanopores. The individually addressable nanopores among the plurality of individually addressable nanopores can contain at least one nanopore formed in a membrane disposed adjacent to an integrated circuit. Each individually addressable nanopore is capable of detecting a label associated with an individual nucleotide. Nucleotides can be incorporated (e.g., polymerized), and the label can not be released from the nucleotide upon incorporation.
[0140] The tag can be presented to the nanopore and released from the nucleotide after a nucleotide incorporation event. The released tag can pass through the nanopore. In some cases, the tag does not pass through the nanopore. The tag released in a nucleotide incorporation event is distinguished from a tag that can pass through the nanopore but is not fully released during its residence time in the nanopore after the nucleotide incorporation event. In some cases, a tag that resides in the nanopore for at least 100 milliseconds (ms) is released in a nucleotide incorporation event, and a tag that resides in the nanopore for less than 100 ms is not released in a nucleotide incorporation event. The tag can be captured and / or guided through the nanopore by a second enzyme or protein (e.g., a nucleic acid binding protein). The second enzyme can cleave the tag at the time of nucleotide incorporation (e.g., during or after this period). The linker between the tag and the nucleotide can be cleaved.
[0141] An incorporated nucleic acid can be detected by the nanopore and / or for a shorter time compared to an unincorporated nucleotide. Alternatively, an incorporated nucleic acid can be detected by the nanopore for a longer period of time compared to an unincorporated nucleotide. As described herein, the difference and / or ratio between these times can be used to determine whether a nucleotide detected by the nanopore has been incorporated.
[0142] The detection time can be based on the free flow of the nucleotide through the nanopore; an unincorporated nucleotide can reside in or near the nanopore for a period of time between 1 nanosecond (ns) and 100 ms or between 1 ns and 50 ms, while an incorporated nucleotide can reside in or near the nanopore for a period of time between 50 ms and 500 ms or between 100 ms and 200 ms. The period of time can vary depending on the processing conditions; however, the residence time of the incorporated nucleotide can be greater than the residence time of the unincorporated nucleotide.
[0143] IV. Polymerase
[0144] DNA polymerase can bind to the 3' end of the nick of the nucleic acid (NA) molecule disclosed herein (e.g., the 3' end of the heterologous nick of the circularized NA molecule). DNA sequencing can be accomplished by amplifying and transcribing a polynucleotide near the nanopore and the tagged nucleotide using an enzyme such as DNA polymerase. The sequencing method can involve incorporating or polymerizing a tagged nucleotide using a polymerase such as DNA polymerase or a transcriptase. The polymerase can be mutated to enable it to receive the tagged nucleotide. The polymerase can also be mutated to increase the time for the nanopore to detect the tag.
[0145] For example, a sequencing enzyme can be any suitable enzyme that produces a polynucleotide chain through a phosphodiester bond of a nucleotide. For example, a DNA polymerase can be a 9°N m™ polymerase or a variant thereof, Escherichia coli DNA polymerase I, bacteriophage T4 DNA polymerase, Sequenase, Taq DNA polymerase, 9°N m™ polymerase (exo-) A485L / Y409V, Φ29 DNA polymerase, Bst DNA polymerase, or a variant, mutant, or homolog of any of the foregoing. A homolog can have any suitable percentage homology, e.g., at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or at least 95% sequence identity.
[0146] In some instances, for nanopore sequencing, a polymerase can be attached to or positioned near a nanopore. Suitable methods for attaching a polymerase to a nanopore include crosslinking the enzyme to the nanopore or near the nanopore, such as by forming an intramolecular disulfide bond. The nanopore and the enzyme can also be a fusion, e.g., a fusion encoded by a single polypeptide chain. Methods for generating a fusion protein can include fusing the coding sequence of the enzyme in-frame and adjacent to the coding sequence of the nanopore and expressing the fusion sequence from a single promoter. A polymerase can be attached or coupled to a nanopore using a molecular staple or a protein finger. The polymerase can be attached to the nanopore through an intermediate molecule, e.g., biotin conjugated to the enzyme and the nanopore, where a streptavidin tetramer links the two biotins. The intermediate molecule can be referred to as a linker.
[0147] A sequencing enzyme can also be attached to a nanopore with an antibody. Proteins that form a covalent bond with each other can be used to attach a polymerase to a nanopore. A phosphatase or an enzyme that cleaves a tag from a nucleotide can also be attached to a nanopore.
[0148] A polymerase can be mutated relative to a non-mutated polymerase to facilitate and / or increase the efficiency of incorporation of a tagged nucleotide into a growing polynucleotide by the mutated polymerase. The polymerase can be mutated to allow a nucleotide analog (such as a tagged nucleotide) to better enter the active site region of the polymerase and / or mutated to match the nucleotide analog in the active region.
[0149] Other mutations, such as amino acid substitutions, insertions, deletions, and / or exogenous features of polymerization, can result in enhanced metal ion coordination, reduced exonuclease activity, reduced reaction rates for one or more steps of the polymerase kinetic cycle, reduced branching ratio, altered cofactor selectivity, increased yield, increased thermal stability, increased accuracy, increased speed, increased read length, increased salt tolerance, relative to a non-mutated polymerase.
[0150] Suitable polymerases can have kinetic rate characteristics suitable for detecting tags via nanopores. Rate characteristics generally refer to the overall rate of nucleotide incorporation and / or the rate of any step of nucleotide incorporation, such as nucleotide addition, enzyme isomerization (e.g., becoming a closed state or isomerizing from a closed state), cofactor binding or release, product release, incorporation of polynucleotides into the growing polynucleotide, or the rate of translocation.
[0151] The polymerase can be adapted to permit detection of sequencing events. The rate characteristics of the polymerase can be such that a tag is loaded into (and / or detected by) the nanopore for an average of 0.1 milliseconds (ms), 1 ms, 5 ms, 10 ms, 20 ms, 30 ms, 40 ms, 50 ms, 60 ms, 80 ms, 100 ms, 120 ms, 140 ms, 160 ms, 180 ms, 200 ms, 220 ms, 240 ms, 260 ms, 280 ms, 300 ms, 400 ms, 500 ms, 600 ms, 800 ms, or 1000 ms. For example, the rate characteristics of the polymerase can be such that a tag is loaded into the nanopore and / or detected by the nanopore for at least 5 ms, at least 10 ms, at least 20 ms, at least 30 ms, at least 40 ms, at least 50 ms, at least 60 ms, at least 80 ms, at least 100 ms, at least 120 ms, at least 140 ms, at least 160 ms, at least 180 ms, at least 200 ms, at least 220 ms, at least 240 ms, at least 260 ms, at least 280 ms, at least 300 ms, at least 400 ms, at least 500 ms, at least 600 ms, at least 800 ms, or at least 1000 ms. The nanopore can detect the tag between an average of 80 ms and 260 ms, between 100 ms and 200 ms, or between 100 ms and 150 ms.
[0152] The nanopore / polymerase complex can be configured to permit detection of one or more events associated with amplification and transcription of circular polynucleotides. The one or more events can be kinetically observable and / or non-kinetically observable, such as nucleotides migrating through the nanopore without contacting the polymerase.
[0153] In some cases, the polymerase reaction exhibits two kinetic steps starting from an intermediate where a nucleotide or polyphosphate moiety binds to the polymerase, and two kinetic steps starting from an intermediate where the nucleotide and polyphosphate moiety do not bind to the polymerase. The two kinetic steps can include enzyme isomerization, nucleotide incorporation, and product release. In some cases, the two kinetic steps are template translocation and nucleotide binding.
[0154] Suitable polymerases can exhibit strong or enhanced strand displacement.
[0155] V. Identification of Sequence Variants
[0156] The methods provided by the present disclosure can be used to identify sequence variants in a polynucleotide sample. If the sequence differences occur in at least two different polynucleotides, e.g., two different circular polynucleotides, then the sequence differences between the sequencing reads and the reference sequence are referred to as true sequence variants, which can be distinguished by having different junctions. Since the positions and types of sequence variants due to amplification or sequencing errors are unlikely to be exactly replicated on two different polynucleotides containing the same target sequence, including this validation parameter can reduce the background of error sequence variants while increasing the sensitivity and accuracy of detecting actual sequence variations in the sample. The frequency of the sequence variant can be less than 5%, less than 4%, less than 3%, less than 2%, less than 1.5%, less than 1%, less than 0.75%, less than 0.5%, less than 0.25%, less than 0.1%, less than 0.075%, less than 0.05%, less than 0.04%, less than 0.03%, less than 0.02%, less than 0.01%, less than 0.005%, less than 0.001%, or a lower frequency that is sufficiently above the background value to allow accurate identification. The occurrence frequency of the sequence variant may be less than 0.1%. When the frequency of the sequence variant is statistically significantly higher than the background error rate, e.g., the p-value is less than 0.05, less than 0.01, less than 0.001, or less than 0.0001, then the frequency of the sequence variant can be sufficiently above the background. When the frequency is at least 2-fold, at least 3-fold, at least 4-fold, at least 5-fold, at least 6-fold, at least 7-fold, at least 8-fold, at least 9-fold, at least 10-fold, at least 25-fold, at least 50-fold, at least 100-fold, or more higher than the background error rate, the frequency of the sequence variant can be sufficiently above the background. The background error rate for accurately determining the sequence at a given position can be less than 1%, less than 0.5%, less than 0.1%, less than 0.05%, less than 0.01%, less than 0.005%, less than 0.001%, or less than 0.0005%.
[0157] Identifying sequence variants can include optimally aligning one or more sequencing reads with a reference sequence to identify the differences between the two, and identifying the junctions. Alignment can include placing one sequence along another sequence, iteratively introducing gaps along each sequence, scoring the degree of match between the two sequences, and repeating at different positions along the reference sequence. The match with the highest score is considered the alignment and represents an inference of the relatedness between the sequences.
[0158] The reference sequence for comparison with the sequencing reads is a reference genome, such as the genome of a member of the same species as the subject. The reference genome can be complete or incomplete. The reference genome can consist only of the region containing the target polynucleotide, for example, the region from the reference genome or the consensus sequence generated from the sequencing reads in the analysis. The reference sequence can contain or be composed of the polynucleotide sequences of one or more organisms, such as sequences from one or more bacteria, archaea, viruses, protists, fungi, or other organisms. The reference sequence can consist only of a part of the reference genome, such as the region corresponding to one or more target sequences in the analysis. For example, for detecting a pathogen, the reference genome can be the entire genome of the pathogen, or a part of it for identification, such as a specific strain or serotype. The sequencing reads can be aligned with multiple different reference sequences for screening multiple different organisms or strains.
[0159] VI. Therapeutic Applications
[0160] The methods, systems, and compositions provided herein can be directed to one or more therapeutic applications, such as for characterizing a patient sample and optionally diagnosing a condition of a subject. Therapeutic applications can include informing a patient of treatment options that may be most responsive to them and / or informing the treatment of a subject in need of therapeutic intervention based on the results of the methods provided in this disclosure.
[0161] For example, the methods provided in this disclosure can be used to diagnose the presence, progression, and / or metastasis of a tumor, such as when the polynucleotide being analyzed contains, consists of, or is composed of cfDNA, ctDNA, or fragmented tumor DNA. For example, the efficacy of tumor treatment in a subject can be monitored by monitoring ctDNA over time, a decrease in ctDNA can be used as an indication of treatment efficacy, and an increase in ctDNA can inform the selection of a different treatment and / or a different dose. Other uses include assessing organ rejection in a transplant recipient, such as using an increase in the amount of circulating DNA corresponding to the transplant donor genome as an early indicator of transplant rejection, and genotyping / haplotyping of pathogen infections (such as viral or bacterial infections). Detection of sequence variants in circulating fetal DNA can be used to diagnose a condition of the fetus.
[0162] The methods provided in this disclosure can include diagnosing a subject based on sequencing results, such as diagnosing a subject with a disease associated with a detected causal genetic variant, or reporting the likelihood that a patient has or will develop such a disease.
[0163] Causal genetic variants can include sequence variants associated with a particular type or stage of cancer or cancer having a particular characteristic such as metastatic potential, drug resistance, and / or drug responsiveness. The methods provided by the present disclosure can be used to inform treatment decisions, guide, and monitor cancer treatment. For example, the treatment efficacy can be monitored by comparing ctDNA samples of a patient before, during, and after treatment, the treatment including a particular molecular targeted therapy such as a monoclonal drug, a chemotherapeutic drug, a radiation regimen, and any combination of the foregoing methods. For example, ctDNA can be monitored to see if certain mutations increase or decrease after treatment, or if new mutations appear, which can enable a doctor to modify the treatment regimen in a shorter time than monitoring methods that track a patient's symptoms. The methods can include diagnosing a subject based on the results of polynucleotide sequencing, such as diagnosing a subject with a particular stage or type of cancer associated with a detected sequence variant, or reporting the likelihood that a patient has or will develop such cancer.
[0164] For example, for a therapy specifically targeted to a patient based on a molecular marker, the patient can be tested to find out if certain mutations are present in their tumor, and these mutations can be used to predict the response or drug resistance to the therapy and guide the decision whether to use the therapy. Detecting and monitoring ctDNA during treatment helps guide treatment selection.
[0165] Sequence variants associated with one or more cancers can be used for diagnosis, prognosis, or treatment decisions. For example, suitable target sequences of oncological significance include alterations of the TP53 gene, the ALK gene, the KRAS gene, the PIK3CA gene, the BRAF gene, the EGFR gene, and the KIT gene. The target sequences can be specifically amplified, and / or the sequence variants of the target sequences can be specifically analyzed to see if they may be all or part of a cancer-related gene.
[0166] The methods provided by the present disclosure can be used to discover new rare mutations associated with one or more cancer types, stages, or cancer characteristics. For example, in a population of individuals sharing an analyzed characteristic such as a particular disease, cancer type, and / or cancer stage, the methods provided by the present disclosure can be used to identify sequence variants reflecting mutations of a particular gene or gene part. The identified sequence variants that occur at a statistically significantly higher frequency in the group of individuals sharing the characteristic compared to individuals without the characteristic can have an association with the characteristic. Then, the identified sequence variants or types of sequence variants can be used to diagnose or treat individuals found to carry them.
[0167] Additional therapeutic applications can include use in non-invasive fetal diagnosis. Fetal DNA can be found in the blood of pregnant women. The methods provided by this disclosure can be used to identify sequence variants in circulating fetal DNA and can thus be used to diagnose one or more genetic diseases in the fetus, such as diseases associated with one or more causal genetic variants. Examples of causal genetic variants include trisomies, cystic fibrosis, sickle cell anemia, and Tay-Sachs disease. The mother can provide a control sample and a blood sample for comparison. The control sample can be any suitable tissue and can then be sequenced to provide a reference sequence. The cfDNA sequences corresponding to the fetal genomic DNA can then be identified as sequence variants relative to the maternal reference. The father can also provide a reference sample to aid in the identification of fetal sequences and sequence variants.
[0168] Diverse therapeutic applications can include detecting exogenous polynucleotides, including detection from pathogens such as bacteria, viruses, fungi, and microorganisms, and this information can suggest treatment.
[0169] VII. Computer Systems
[0170] This disclosure provides computer systems for implementing the methods of this disclosure. Figure 3 Shown is computer system 1101, which is programmed or otherwise configured to communicate with and regulate various aspects of the sequencing of this disclosure. Computer system 1101 can regulate various operations of the sensor portion, for example, detecting one or more signals indicative of impedance or impedance changes in the sensor portion when at least a portion of the target nucleic acid molecules bind to the binding portion of the sensor portion. Computer system 1101 can be an electronic device of the user or a computer system remotely located relative to the electronic device. The electronic device can be a mobile electronic device.
[0171] The computer system 1101 includes a central processing unit (CPU, also referred to herein as "processor" and "computer processor") 1105, which can be a single-core or multi-core processor, or multiple processors for parallel processing. The computer system 1101 also includes a memory or memory location 1110 (e.g., random access memory, read-only memory, flash memory), an electronic storage unit 1115 (e.g., hard disk), a communication interface 1120 (e.g., network adapter) for communicating with one or more other systems, and peripheral devices 1125 such as caches, other memories, data storage, and / or electronic display adapters. The memory 1110, storage unit 1115, interface 1120, and peripheral devices 1125 communicate with the CPU 1105 via a communication bus (solid lines), such as a motherboard. The storage unit 1115 can be a data storage unit (or data repository) for storing data. The computer system 1101 can be operably coupled to a computer network ("network") 1130 via the communication interface 1120. The network 1130 can be the Internet, an intranet, and / or an extranet, or an intranet and / or extranet that communicates with the Internet. In some cases, the network 1130 is a telecommunications and / or data network. The network 1130 can include one or more computer servers, which can implement distributed computing, such as cloud computing. In some cases, the network 1130 can implement a peer-to-peer network with the help of the computer system 1101, which can enable devices coupled to the computer system 1101 to act as clients or servers.
[0172] The CPU 1105 can execute a series of machine-readable instructions, which can be embodied in a program or software. The instructions can be stored in a memory location, such as the memory 1110. The instructions can be directed to the CPU 1105, and subsequently, the CPU 1105 can be programmed or otherwise configured to implement the methods of the present disclosure. Examples of operations performed by the CPU 1105 can include fetching, decoding, executing, and writing back.
[0173] The CPU 1105 can be part of a circuit, such as an integrated circuit. One or more other components of the system 1101 can be included in the circuit. In some cases, the circuit is an application-specific integrated circuit (ASIC).
[0174] The storage unit 1115 can store files, such as drivers, libraries, and saved programs. The storage unit 1115 can store user data, e.g., user preferences and user programs. In some cases, the computer system 1101 can include one or more additional data storage units external to the computer system 1101, such as on a remote server that communicates with the computer system 1101 via an intranet or the Internet.
[0175] Computer system 1101 can communicate with one or more remote computer systems via network 1130. For example, computer system 1101 can communicate with a user's remote computer system. Examples of remote computer systems include personal computers (e.g., portable PCs), tablets or tablet computers (e.g., Apple iPad, Samsung Galaxy Tab), telephones, smartphones (e.g., Apple iPhone, Android-enabled devices, Blackberry ), or personal digital assistants. A user can access computer system 1101 via network 1130.
[0176] The methods described herein can be implemented by way of machine (e.g., computer processor) executable code stored on an electronic storage location of computer system 1101 (e.g., memory 1110 or electronic storage unit 1115). The machine executable or machine readable code can be provided in software form. During use, the code can be executed by processor 1105. In some cases, the code can be retrieved from storage unit 1115 and stored in memory 1110 for ready access by processor 1105. In some cases, the electronic storage unit 1115 may not be included and the machine executable instructions are stored in memory 1110.
[0177] The code can be pre-compiled and configured to be used with a machine having a processor suitable for executing the code, or can be compiled at runtime. The code can be provided in a programming language, and the programming language can be selected to enable the code to be executed in a pre-compiled or just-in-time compiled manner.
[0178] Aspects of the systems and methods provided herein, such as computer system 1101, may be embodied as programming. Various aspects of technology may be considered “products” or “articles of manufacture,” typically in the form of machine (or processor) executable code and / or associated data carried or embodied on a machine-readable medium. The machine executable code may be stored on an electronic storage unit, such as a memory (e.g., read-only memory, random access memory, flash memory) or a hard disk. A “storage” type of medium may include any or all tangible memories of a computer, processor, etc., or associated modules, such as various semiconductor memories, tape drives, disk drives, etc., which may provide non-transitory storage for software programming at any time. All or part of the software can sometimes be communicated via the Internet or other various telecommunication networks. For example, such communication can enable software to be loaded from one computer or processor to another, e.g., from a management server or host to a computer platform of an application server. Thus, another type of medium that can carry software elements includes light waves, electric waves, and electromagnetic waves, such as those used between physical interfaces of local devices, via wired and optical landline networks, and over various air links. Physical elements carrying such waves, such as wired or wireless links, optical links, etc., can also be considered media carrying software. As used herein, unless restricted to non-transitory, tangible “storage” media, terms such as computer or machine “readable media” refer to any medium that participates in providing instructions to a processor for execution.
[0179] Thus, machine-readable media, such as computer executable code, can take many forms, including but not limited to tangible storage media, carrier media, or physical transmission media. Non-volatile storage media include, for example, optical or magnetic disks, such as any storage device in any computer, etc., which may be used, for example, to implement databases shown in the drawings. Volatile storage media include dynamic memory, such as the main memory of such computer platforms. Tangible transmission media include coaxial cables; copper wire and fiber optics, including the wires that make up the bus within a computer system. Carrier transmission media may take the form of electrical or electromagnetic signals, or may take the form of acoustic or light waves, such as those generated during radio frequency (RF) and infrared (IR) data communications. Thus, common forms of computer-readable media include, for example: floppy disks, flexible disks, hard disks, magnetic tape, any other magnetic medium, CD-ROM, DVD or DVD-ROM, any other optical medium, punched cards, paper tape, any other physical storage medium with punched patterns, RAM, ROM, PROM, and EPROM, FLASH-EPROM, any other memory chip or cartridge, carriers that carry data or instructions, cables or links that carry such carriers, or any other medium from which a computer can read programming code and / or data. Many of these forms of computer-readable media may involve carrying one or more sequences of one or more instructions to a processor for execution.
[0180] The computer system 1101 may include or communicate with an electronic display 1135 that includes a user interface (UI) 1140 to provide, for example, (i) the progress of sequencing, and (ii) sequencing information obtained from the sequencing. Examples of the UI include, but are not limited to, a graphical user interface (GUI) and a web-based user interface.
[0181] The methods and systems of the present disclosure may be implemented by one or more algorithms. The algorithms may be implemented by software when executed by a central processing unit 1105. For example, the algorithms may determine sequence reads of a target nucleic acid.
[0182] Although the preferred embodiments of the present disclosure have been shown and described herein, it will be apparent to those skilled in the art that these embodiments are provided by way of example only. The present disclosure is not intended to be limited by the specific examples provided in the specification. Although the present disclosure has been described with reference to the foregoing specification, the description and illustration of the embodiments herein are not meant to be construed in a limiting sense. Various variations, changes, and alternatives will now occur to those skilled in the art without departing from the present disclosure. In addition, it should be understood that all aspects of the present disclosure are not limited to the specific descriptions, configurations, or relative proportions set forth herein that depend on various conditions and variables. It should be understood that various alternatives of the embodiments of the present disclosure described herein may be used to implement the present disclosure. Accordingly, it is contemplated that the present disclosure will also cover any such alternatives, modifications, variations, or equivalents. The following claims are intended to define the scope of the present disclosure and thus cover the methods and structures within the scope of these claims and their equivalents.
[0183] Examples
[0184] Example 1: 3’-O-azidomethyl-deoxyadenosine-5’-hexaphosphate-T40 (3’-O-N3-dA6P-T40)
[0185] Figure 1 An example of an engineered nucleotide molecule disclosed herein is schematically shown. The engineered nucleotide molecule 3’-O-N3-dA6P-T40 has a deoxyribose, an A base coupled to the sugar, a hexaphosphate coupled to the sugar, a T40 coupled to the sugar through the hexaphosphate, and an azidomethyl coupled to the 3’-O of the sugar. Figure 1 The molecule shown may be used in any sequencing method disclosed herein.
[0186] Example 2: 3’-O-allyl-deoxycytidine-5’-triphosphate-T40 (3’-O-allyl-dCTP-T40)
[0187] Figure 2An example of an engineered nucleotide molecule disclosed herein is schematically shown. The engineered nucleotide molecule 3'-O-allyl-dCTP-T40 has a deoxyribose, a C base coupled to the sugar, a triphosphate coupled to the sugar, a T40 coupled to the sugar through the triphosphate, and an allyl coupled to the 3'-O of the sugar. Figure 2 The molecules shown can be used in any sequencing method disclosed herein.
[0188] Although the preferred embodiments of the present disclosure have been shown and described herein, it will be apparent to those skilled in the art that these embodiments are provided by way of example only. The present disclosure is not intended to be limited by the specific examples provided in the specification. Although the present disclosure has been described with reference to the foregoing specification, the description and illustration of the embodiments herein are not meant to be construed in a limiting sense. Various variations, changes, and alternatives will now occur to those skilled in the art without departing from the present disclosure. In addition, it should be understood that all aspects of the present disclosure are not limited to the specific descriptions, configurations, or relative proportions set forth herein that depend on various conditions and variables. It should be understood that various alternatives to the embodiments of the present disclosure described herein can be used to practice the present disclosure. It is, therefore, contemplated that the present disclosure will also cover any such alternatives, modifications, variations, or equivalents. The following claims are intended to define the scope of the present disclosure and thereby cover the methods and structures within the scope of these claims and their equivalents.
Claims
1. An engineered nucleotide molecule, comprising: A pentose; A base coupled to the pentose, wherein the base is selected from adenine, guanine, cytosine, thymine, uracil, and analogs thereof; A polyphosphate chain coupled to the pentose, wherein the polyphosphate chain comprises two or more phosphate groups; A protecting group coupled to the pentose, wherein the protecting group is configured to inhibit the coupling of additional nucleotides to the engineered nucleotide molecule; and An identifier moiety coupled to the pentose, wherein the identifier moiety is specific to the engineered nucleotide molecule, and wherein the identifier moiety is directly coupled to the polyphosphate chain.
2. The engineered nucleotide molecule according to claim 1, wherein the pentose is deoxyribose.
3. The engineered nucleotide molecule according to claim 1, wherein the polyphosphate chain comprises three or more phosphate groups.
4. The engineered nucleotide molecule according to claim 1, wherein the polyphosphate chain comprises four or more phosphate groups.
5. The engineered nucleotide molecule according to claim 1, wherein the protecting group is coupled to a hydroxyl group of the pentose.
6. The engineered nucleotide molecule according to claim 5, wherein the hydroxyl group is located at the 3' position of the pentose.
7. The engineered nucleotide molecule according to claim 1, wherein the identifier moiety is removable from the engineered nucleotide molecule.
8. The engineered nucleotide molecule according to claim 1, wherein the identifier moiety comprises a polynucleotide.
9. The engineered nucleotide molecule according to claim 1, wherein the identifier moiety comprises a non - polynucleotide / non - polypeptide polymer.
10. The engineered nucleotide molecule according to claim 8, wherein the polynucleotide has a length of at least about 5 bases.
11. The engineered nucleotide molecule according to claim 8, wherein the polynucleotide has a length of at least about 10 bases.
12. The engineered nucleotide molecule according to claim 8, wherein the polynucleotide has a length of at least about 20 bases.
13. The engineered nucleotide molecule according to claim 8, wherein the polynucleotide has a length of at least about 30 bases.
14. The engineered nucleotide molecule according to claim 8, wherein the polynucleotide comprises a polyN selected from polyA, polyT, polyC, polyG, polyU, and variants thereof.
15. The engineered nucleotide molecule according to claim 1, wherein the protecting group comprises allyl or azide.
16. The engineered nucleotide molecule according to claim 1, wherein the protecting group is removable from the engineered nucleotide molecule.
17. A method of analyzing a target nucleic acid molecule, comprising: (a) Providing a complex, the complex comprising (i) the target nucleic acid molecule and (ii) a primer nucleic acid molecule that exhibits complementarity to a portion of the target nucleic acid molecule; And (b) contacting the complex with an engineered nucleotide molecule to produce a growing strand coupled to the primer nucleic acid molecule, wherein the growing strand exhibits sequence complementarity to an additional portion of the target nucleic acid molecule, and wherein the engineered nucleotide molecule comprises: a pentose; a base coupled to the pentose, wherein the base is selected from adenine, guanine, cytosine, thymine, uracil, and analogs thereof; a polyphosphate chain coupled to the pentose, wherein the polyphosphate chain comprises two or more phosphate groups; a protecting group coupled to the pentose, wherein the protecting group is configured to inhibit the coupling of additional nucleotides to the engineered nucleotide; and an identifier moiety coupled to the pentose, wherein the identifier moiety is specific to the engineered nucleotide, and wherein the identifier moiety is directly coupled to the polyphosphate chain.
18. The method of claim 17, further comprising detecting (i) the contacting or (ii) the production of the growing strand using a sensor moiety.
19. The method of claim 18, wherein the detecting comprises measuring one or more signals indicative of impedance or impedance changes in the sensor moiety at (i) the contacting or (ii) the production of the growing strand.
20. The method of claim 18, further comprising contacting the complex with the sensor moiety to incorporate at least a portion of the engineered nucleotide molecule as part of the growing strand.
21. The method of claim 18, wherein the sensor moiety comprises a pore or an enzyme.
22. The method of claim 20, wherein the sensor moiety comprises the pore and the enzyme coupled to the pore.
23. The method of claim 20, wherein the pore is part of a nanopore protein.
24. The method of claim 20, wherein the pore is part of a solid-state nanopore.
25. The method of claim 20, wherein the enzyme comprises a polymerase.
26. The method of claim 17, further comprising removing the protecting group from the pentose after (b).
27. The method of claim 26, further comprising coupling an additional nucleotide to the engineered nucleotide after the removing.
28. The method of claim 26, wherein the removing comprises an enzymatic reaction.
29. The method of claim 26, wherein the removing comprises a chemical reaction without an enzyme.
30. The method of claim 17, further comprising removing the identifier moiety from the pentose after (b).
31. The method of claim 17, wherein the pentose is deoxyribose.
32. The method of claim 17, wherein the polyphosphate chain comprises three or more phosphate groups.
33. The method of claim 17, wherein the polyphosphate chain comprises four or more phosphate groups.
34. The method of claim 17, wherein the protecting group is coupled to a hydroxyl group of the pentose.
35. The method according to claim 34, wherein the hydroxyl group is located at the 3'-position of the pentose.
36. The method according to claim 17, wherein the protecting group is removable from the pentose.
37. The method according to claim 17, wherein the identifier moiety is removable from the engineered nucleotide molecule.
38. The method according to claim 17, wherein the identifier moiety comprises a polynucleotide sequence that does not exhibit complementarity to at least a portion of the target nucleic acid molecule.
39. The method according to claim 17, wherein the identifier moiety comprises a polynucleotide.
40. The method according to claim 39, wherein the polynucleotide has a length of at least about 5 bases.
41. The method according to claim 39, wherein the polynucleotide has a length of at least about 10 bases.
42. The method according to claim 39, wherein the polynucleotide has a length of at least about 20 bases.
43. The method according to claim 39, wherein the polynucleotide has a length of at least about 30 bases.
44. The method according to claim 39, wherein the polynucleotide comprises a polyN selected from polyA, polyT, polyC, polyG, polyU, and variants thereof.
45. The method according to claim 17, wherein the identifier moiety comprises a non-polynucleotide / non-polypeptide polymer.
46. The method according to claim 17, wherein the protecting group comprises an allyl group or an azide.