Systems and methods for analyzing target molecules

By guiding the flow of additional regions of the primer nucleic acid molecules in the sensor hole, combining the action of the enzyme, a growth chain complementary to the target nucleic acid molecules is generated, and the efficiency and accuracy of the sequence analysis of target nucleic acid molecules in the prior art is solved, and efficient detection of rare sequence variants and early pathological mutations is achieved.

CN120418447APending Publication Date: 2025-08-01AXBIO INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380086884.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-10-21
Filing Date
2023-10-19
Publication Date
2025-08-01

AI Technical Summary

Technical Problem

The prior art is difficult to efficiently and accurately analyze the sequence information of target nucleic acid molecules, especially when detecting rare sequence variants and early pathological mutations, and lacks effective methods and equipment.

Method used

By providing a complex containing a target nucleic acid molecule and a primer nucleic acid molecule, the analysis of the target nucleic acid molecule is achieved using the pore structure in the sensor and the cooperation of the enzyme. The method includes guiding the flow of additional regions of the primer nucleic acid molecule through the sensor hole, and identifying the sequence information of the target nucleic acid molecule through electrical signal changes, generating growth chains complementary to the target nucleic acid molecule, and obtaining its sequence information.

Benefits of technology

It realizes efficient and accurate analysis of target nucleic acid molecules, can detect rare sequence variants and early pathological mutations, and improves the efficiency and accuracy of nucleic acid sequencing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120418447A_ABST
    Figure CN120418447A_ABST
Patent Text Reader

Abstract

The present disclosure provides systems and methods for analyzing or identifying target nucleic acid molecules. For example, a system of the present disclosure may comprise (a) a complex comprising a target nucleic acid molecule and a primer nucleic acid molecule, where the primer nucleic acid molecule comprises (i) a region complementary to a portion of the target nucleic acid molecule and (ii) an additional region that is not complementary to the target nucleic acid molecule, and (b) after the additional region has flowed through a well of a sensor, the primer nucleic acid molecule is selectively coupled to the target nucleic acid molecule. The additional region is identified using the sensor, thereby analyzing the target nucleic acid molecule. The sensor may be configured to detect one or more signals indicative of an impedance or a change in impedance in the sensor when at least a portion of the target molecule is bound to the binding unit. The one or more signals may be used to analyze or identify a target molecule.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - Reference to Related Applications

[0002] This application claims the benefit of U.S. Provisional Application No. 63 / 418,204, filed Oct. 21, 2022, which is hereby incorporated by reference in its entirety. Background of the Invention

[0003] Nucleic acid sequencing can be used to provide sequence information of a nucleic acid sample. Such sequence information can be helpful for diagnosing or treating a condition (e.g., a disease) of a subject (e.g., an individual, a patient, etc.). For example, the nucleic acid sequence information of a subject can be used to identify, diagnose, or develop treatments for one or more genetic diseases. In another example, the nucleic acid sequence information of one or more pathogens can guide the treatment of one or more infectious diseases.

[0004] The detection of one or more rare sequence variants (e.g., mutations) can be valuable for healthcare. The detection of rare sequence variants can be important for the early detection of one or more pathological mutations. Detecting one or more cancer-related mutations (e.g., point mutations) in a clinical sample can improve chemotherapy for recurrent patients or the identification of one or more minimal residual diseases in tumor cell detection. In addition, such mutation detection can be important for assessing exposure to environmental mutagens, monitoring endogenous DNA repair, or studying the accumulation of one or more somatic mutations in aging individuals. Detecting or rare sequence variants can enhance prenatal diagnosis and enable the characterization of fetal cells present in maternal blood. Summary of the Invention

[0005] One aspect of the present disclosure provides a method for analyzing a target nucleic acid molecule, comprising: (a) providing a complex comprising the target nucleic acid molecule and a primer nucleic acid molecule, wherein the primer nucleic acid molecule comprises (i) a region complementary to a portion of the target nucleic acid molecule and (ii) an additional region that is not complementary to the target nucleic acid molecule; and (b) identifying the additional region using a sensor after the additional region has flowed through a pore of the sensor, thereby analyzing the target nucleic acid molecule.

[0006] In some embodiments of any of the methods provided herein, the method further comprises contacting the complex with an enzyme operably coupled to the pore. In some embodiments of any of the methods provided herein, the contacting enables the flow of the additional region through the pore.

[0007] In some embodiments of any of the methods provided herein, the method further comprises, in (b), alternately varying the electrical signal in the sensor to direct the additional region to flow through the pore two or more times. In some embodiments of any of the methods provided herein, the method further comprises, in (b): (i) performing an extension reaction of the target nucleic acid molecule from the primer nucleic acid molecule to generate a growing strand coupled to the primer nucleic acid molecule, wherein the growing strand exhibits sequence complementarity to an additional portion of the target nucleic acid molecule; and (ii) obtaining sequence information of at least a portion of the growing strand by directing at least a portion of the growing strand through the pore of the sensor.

[0008] In some embodiments of any of the methods provided herein, in (a), the region of the primer nucleic acid molecule flanks (1) the additional region of the primer nucleic acid molecule and (2) the growing strand. In some embodiments of any of the methods provided herein, in (b), the obtaining of the sequence information comprises identifying a sequence read that contains characteristic sequence information indicative of at least a portion of the additional region of the primer nucleic acid molecule. In some embodiments of any of the methods provided herein, in (b), directing the additional region of the primer nucleic acid molecule through the pore before at least a portion of the growing strand. In some embodiments of any of the methods provided herein, in (b), the directing comprises alternately varying the electrical signal in the sensor to direct at least a portion of the growing strand through the pore two or more times.

[0009] In some embodiments of any of the methods provided herein, the additional region comprises at least 5 bases. In some embodiments of any of the methods provided herein, the additional region comprises at least 10 bases. In some embodiments of any of the methods provided herein, the additional region comprises at least 20 bases. In some embodiments of any of the methods provided herein, the additional region comprises a net positive charge. In some embodiments of any of the methods provided herein, the additional region comprises a net negative charge.

[0010] In some embodiments of any of the methods provided herein, the pore is part of a nanopore protein. In some embodiments of any of the methods provided herein, the pore is part of a solid-state nanopore. In some embodiments of any of the methods provided herein, the target nucleic acid molecule is a circular nucleic acid molecule. In some embodiments of any of the methods provided herein, in (c), the obtaining comprises detecting one or more signals indicative of impedance or a change in impedance in the sensor when guiding at least the additional portion through the pore. In some embodiments of any of the methods provided herein, the pore is embedded in a membrane. In some embodiments of any of the methods provided herein, the primer nucleic acid molecule comprises a barcode.

[0011] Another aspect of the present disclosure provides a system for analyzing a target nucleic acid molecule, comprising: a primer nucleic acid molecule comprising: (i) a region complementary to a portion of the target nucleic acid molecule and (ii) an additional region not complementary to the target nucleic acid molecule, wherein the primer nucleic acid molecule is configured to form a complex with the target nucleic acid molecule; a sensor comprising a pore, wherein the sensor is configured to guide the additional region to flow through the pore; and a controller operably coupled to the sensor, wherein the controller is configured to identify the additional region flowing through the pore, thereby analyzing the target nucleic acid molecule.

[0012] In some embodiments of any of the systems provided herein, the system further comprises an enzyme operably coupled to the pore, wherein the enzyme is configured to contact the complex. In some embodiments of any of the systems provided herein, the contact between the enzyme and the complex is configured to enable the additional region to flow through the pore. In some embodiments of any of the systems provided herein, the controller is further configured to alternately vary an electrical signal in the sensor to guide the additional region to flow through the pore two or more times.

[0013] In some embodiments of any of the systems provided herein, the sensor is further configured to perform an extension reaction of the target nucleic acid molecule from the primer nucleic acid molecule to generate a growing strand coupled to the primer nucleic acid molecule, wherein the growing strand exhibits sequence complementarity to an additional portion of the target nucleic acid molecule, and wherein the controller is further configured to obtain sequence information of at least the portion of the growing strand by guiding at least a portion of the growing strand through the pore of the sensor.

[0014] In some embodiments of any of the systems provided herein, the region of the primer nucleic acid molecule flanks (1) the additional region of the primer nucleic acid molecule and (2) the growing strand. In some embodiments of any of the systems provided herein, the controller is configured to identify sequence reads that contain characteristic sequence information indicative of at least a portion of the additional region of the primer nucleic acid molecule. In some embodiments of any of the systems provided herein, the additional region of the primer nucleic acid molecule is guided through the pore before at least the portion of the growing strand. In some embodiments of any of the systems provided herein, the controller is configured to alternately vary an electrical signal in the sensor to guide at least the portion of the growing strand through the pore two or more times.

[0015] In some embodiments of any of the systems provided herein, the additional region contains at least 5 bases. In some embodiments of any of the systems provided herein, the additional region contains at least 10 bases. In some embodiments of any of the systems provided herein, the additional region contains at least 20 bases. In some embodiments of any of the systems provided herein, the additional region contains a net positive charge. In some embodiments of any of the systems provided herein, the additional region contains a net negative charge.

[0016] In some embodiments of any of the systems provided herein, the pore is part of a nanopore protein. In some embodiments of any of the systems provided herein, the pore is part of a solid-state nanopore. In some embodiments of any of the systems provided herein, the target nucleic acid molecule is a circular nucleic acid molecule. In some embodiments of any of the systems provided herein, the controller is configured to detect one or more signals indicative of impedance or impedance changes in the sensor when guiding at least the additional portion through the pore. In some embodiments of any of the systems provided herein, the pore is embedded in a membrane. In some embodiments of any of the systems provided herein, the primer nucleic acid molecule contains a barcode.

[0017] Other aspects and advantages of the present disclosure will become apparent to those of ordinary skill in the art from the following detailed description, which illustrates and describes only illustrative embodiments of the present disclosure. As will be recognized, the present disclosure is capable of other different embodiments and its several details can be modified in various obvious aspects without departing from the present disclosure. Accordingly, the drawings and the detailed description are to be regarded as illustrative in nature and not restrictive.

[0018] Incorporation by Reference

[0019] All publications, patents, and patent applications mentioned in the specification are hereby incorporated by reference as if each publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. In the event of conflict between the incorporated publications and patents or patent applications and the disclosure contained herein, the specification is intended to supersede and / or take precedence over any such conflicting material. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] The novel features of the present disclosure are set forth specifically in the appended claims. A better understanding of the features and advantages of the present disclosure will be obtained by reference to the following detailed description of illustrative embodiments that utilize the principles of the present disclosure and the accompanying drawings (also referred to herein as "Figure" and "FIG."), in which:

[0021] Figure 1 An example of a sensor for analyzing or identifying a target molecule is schematically shown.

[0022] Figure 2 A computer system programmed or otherwise configured to implement the methods provided herein is shown.

[0023] Figure 3 An example process for analyzing or identifying a target molecule is shown.

[0024] Figure 4 Another example process for analyzing or identifying a target molecule is shown. DETAILED DESCRIPTION

[0025] Although various embodiments of the present invention have been shown and described herein, it will be readily apparent to those skilled in the art that such embodiments are provided by way of example only. Many variations, changes, and substitutions may be contemplated by those skilled in the art without departing from the present invention. It should be understood that various alternatives to the embodiments of the present invention described herein may be employed.

[0026] Whenever the terms "at least", "greater than", or "greater than or equal to" precede the first of a series of two or more numerical values, the terms "at least", "greater than", or "greater than or equal to" apply to each numerical value in the series. For example, greater than or equal to 1, 2, or 3 is equivalent to greater than or equal to 1, greater than or equal to 2, or greater than or equal to 3.

[0027] Whenever the terms "not greater than", "less than", or "less than or equal to" precede the first of a series of two or more numerical values, the terms "not greater than", "less than", or "less than or equal to" apply to each numerical value in the series. For example, less than or equal to 3, 2, or 1 is equivalent to less than or equal to 3, less than or equal to 2, or less than or equal to 1.

[0028] Unless the context clearly dictates otherwise, the singular forms "a", "an", and "the" as used in this specification and the claims may include plural referents. For example, the term "transmembrane receptor" may include multiple transmembrane receptors.

[0029] As used interchangeably herein, the terms "about" and "approximately" generally refer to an acceptable error range around a particular value as determined by one of ordinary skill in the art, which will depend in part on how the value is measured or determined, i.e., the limitations of the measuring system. For example, "about" can mean within 1 or more standard deviations, per the practice in the art. Alternatively, "about" can mean a range of up to 20%, up to 10%, up to 5%, or up to 1% of a given value. Or, particularly with respect to biological systems or processes, the term can mean within an order of magnitude of a value, such as within 5-fold or 2-fold of a value. Where a particular value is described in this application and the claims, unless otherwise stated, the term "about" can mean within the acceptable error range of that particular value.

[0030] As used herein, the term "dielectric material" generally refers to an electrically insulating material that can be polarized by the action of an applied electric field. When a dielectric material is placed in an electric field, charge may not flow through the dielectric material. In some cases, in such an electric field, the dielectric material may exhibit dielectric polarization, where positive charges may move in the direction of the electric field and negative charges may move in the opposite direction, thereby creating an internal electric field within the dielectric material that can partially compensate the external electric field. Examples of dielectric materials may include, but are not limited to, polyester, polyethylene, polypropylene, cloth (such as nylon), paper, laminates, glass, self-assembled monolayers (SAMs), etc.

[0031] As used herein, the term "biomolecule" generally refers to any molecule, its derivatives, or its functional variants found in biological systems. Biomolecules can be naturally occurring or the result of external perturbations to the system (e.g., disease, poisoning, genetic manipulation, etc.), as well as their synthetic analogs and derivatives. Non-limiting examples of biomolecules may include amino acids (naturally occurring or synthetic), peptides, polypeptides, glycosylated and non-glycosylated proteins (e.g., polyclonal and monoclonal antibodies, receptors, interferons, enzymes, etc.), nucleosides, nucleotides, oligonucleotides (e.g., DNA, RNA, PNA oligomers), polynucleotides (e.g., DNA, cDNA, RNA, etc.), carbohydrates, hormones, haptens, steroids, toxins, etc. Biomolecules can be isolated from natural sources or they can be synthetic.

[0032] As used herein, the term "cell" generally refers to a biological cell or cell derivative. A cell can be the basic structural, functional, and / or biological unit of a living organism. A cell can be derived from any organism having one or more cells. Some non-limiting examples include: prokaryotic cells, eukaryotic cells, bacterial cells, archaeal cells, cells of unicellular eukaryotic organisms, protozoan cells, cells from plants (e.g., cells from plant crops, fruits, vegetables, grains, soybeans, corn, maize, wheat, seeds, tomatoes, rice, cassava, sugarcane, pumpkins, hay, potatoes, cotton, hemp, tobacco, flowering plants, conifers, gymnosperms, ferns, lycophytes, hornworts, liverworts, mosses), algal cells (e.g., Botryococcus braunii, Chlamydomonas reinhardtii, Nannochloropsis gaditana, Chlorella pyrenoidosa, Sargassum patens C. Agardh, etc.), seaweeds (e.g., kelp), fungal cells (e.g., yeast cells, cells from mushrooms), animal cells, cells from invertebrates (e.g., fruit flies, cnidarians, echinoderms, nematodes, etc.), cells from vertebrates (e.g., fish, amphibians, reptiles, birds, mammals), cells from mammals (e.g., pigs, cows, goats, sheep, rodents, rats, mice, non-human primates, humans, etc.), and the like. Sometimes cells do not originate from natural organisms (e.g., cells can be synthetic, sometimes referred to as artificial cells).

[0033] As used interchangeably herein, the terms "nucleotide", "nucleobase", and "base" generally refer to a base-sugar-phosphate combination. Nucleotides can include synthetic nucleotides. Nucleotides can include synthetic nucleotide analogs. Nucleotides can be monomeric units of nucleic acid sequences (e.g., deoxyribonucleic acid (DNA) and ribonucleic acid (RNA)). The term nucleotide can include ribonucleoside triphosphates (adenosine triphosphate (ATP), uridine triphosphate (UTP), cytidine triphosphate (CTP), guanosine triphosphate (GTP), uridine triphosphate (UTP)) and deoxyribonucleoside triphosphates (such as dATP, dCTP, dITP, dUTP, dGTP, dTTP) or derivatives thereof. Such derivatives can include, for example, [αS]dATP, 7-deaza-dGTP and 7-deaza-dATP, as well as nucleotide derivatives that confer nuclease resistance on nucleic acid molecules containing them. As used herein, the term nucleotide generally refers to dideoxynucleoside triphosphates (ddNTPs) and derivatives thereof. Illustrative examples of dideoxynucleoside triphosphates can include, but are not limited to, ddATP, ddCTP, ddGTP, ddITP, and ddTTP. Nucleotides can be unlabeled or detectably labeled (e.g., tagged with a label). Labeling can also be performed with quantum dots. Detectable labels can include, for example, radioisotopes, fluorescent labels, chemiluminescent labels, bioluminescent labels, and enzyme labels. Fluorescent labels of nucleotides can include, but are not limited to, fluorescein, 5-carboxyfluorescein (FAM), 2′7′-dimethoxy-4′5-dichloro-6-carboxyfluorescein (JOE), rhodamine, 6-carboxyrhodamine (R6G), N,N,N′,N′-tetramethyl-6-carboxyrhodamine (TAMRA), 6-carboxy-X-rhodamine (ROX), 4-(4′-dimethylaminophenylazo)benzoic acid (DABCYL), cascade blue, Oregon green, Texas red, cyanine, and 5-(2′-aminoethyl)aminonaphthalene-1-sulfonic acid (EDANS).Specific examples of fluorescently labeled nucleotides may include [R6G]dUTP, [TAMRA]dUTP, [R110]dCTP, [R6G]dCTP, [TAMRA]dCTP, [JOE]ddATP, [R6G]ddATP, [FAM]ddCTP, [R110]ddCTP, [TAMRA]ddGTP, [ROX]ddTTP, [dR6G]ddATP, [dR110]ddCTP, [dTAMRA]ddGTP, and [dROX]ddTTP, which are available from Perkin Elmer, Foster City, Calif.; FluoroLink deoxynucleotides, FluoroLink Cy3-dCTP, FluoroLink Cy5-dCTP, FluoroLink FluorX-dCTP, FluoroLink Cy3-dUTP, and FluoroLink Cy5-dUTP, which are available from Amersham, Arlington Heights, Ill.; fluorescein-15-dATP, fluorescein-12-dUTP, tetramethyl-rhodamine-6-dUTP, IR770-9-dATP, fluorescein-12-ddUTP, fluorescein-12-UTP, and fluorescein-15-2'-dATP, which are available from Boehringer Mannheim, Indianapolis, Ind.; and chromosome-labeled nucleotides, BODIPY-FL-14-UTP, BODIPY-FL-4-UTP, BODIPY-TMR-14-UTP, BODIPY-TMR-14-dUTP, BODIPY-TR-14-UTP, BODIPY-TR-14-dUTP, cascade blue-7-UTP, cascade blue-7-dUTP, fluorescein-12-UTP, fluorescein-12-dUTP, Oregon Green 488-5-dUTP, rhodamine green-5-UTP, rhodamine green-5-dUTP, tetramethylrhodamine-6-UTP, tetramethylrhodamine-6-dUTP, Texas Red-5-UTP, Texas Red-5-dUTP, and Texas Red-12-dUTP, which are available from Molecular Probes, Eugene, Oreg. Nucleotides can also be labeled or tagged by chemical modification. Chemically modified single nucleotides can be biotin-dNTPs. Some non-limiting examples of biotinylated dNTPs may include biotin-dATP (e.g., bio-N6-ddATP, biotin-14-dATP), biotin-dCTP (e.g., biotin-11-dCTP, biotin-14-dCTP), and biotin-dUTP (e.g., biotin-11-dUTP, biotin-16-dUTP, biotin-20-dUTP).

[0034] The naturally occurring nucleotides guanine, cytosine, adenine, thymine, and uracil may be abbreviated as G, C, A, T, and U, respectively. A nucleotide may include any subunit that can be incorporated into a growing nucleic acid chain. Such a subunit may be A, C, G, T, or U, or any other subunit that is specific for one or more complementary A, C, G, T, or U or that is complementary to a purine (i.e., A or G or a variant thereof) or a pyrimidine (i.e., C, T, or U or a variant thereof). The subunit may be an individual nucleic acid base or a group of bases (e.g., AA, TA, AT, GC, CG, CT, TC, GT, GT, TG, AC, CA, or the uracil counterparts thereof).

[0035] As used herein, the terms “circularize” and “circularization” generally refer to promoting or generating a structure in which the two ends of a polynucleotide molecule are coupled to each other (e.g., by covalent and / or hydrogen bonds). The ends of the polynucleotide molecule may be directly coupled to each other. Alternatively, the ends of the polynucleotide molecule may be indirectly coupled to each other through a coupling moiety (or linker) (e.g., at least one separate molecule capable of binding to each of the two ends of the polynucleotide molecule).

[0036] As used interchangeably herein, the terms "polynucleotide", "oligonucleotide", "oligomer", and "nucleic acid" generally refer to polymeric forms of nucleotides of any length, whether deoxyribonucleotides or ribonucleotides or analogs thereof, whether single-stranded, double-stranded, or multi-stranded. Polynucleotides can be exogenous or endogenous to a cell. Polynucleotides can exist in a cell-free environment. Polynucleotides can be genes or fragments thereof. Polynucleotides can be DNA. Polynucleotides can be RNA. Polynucleotides can have any three-dimensional structure and can perform any function. Polynucleotides can include one or more analogs (e.g., altered backbone, sugar, or nucleobase). In the presence of modifications, nucleotide structures can be modified either before or after polymer assembly. Some non-limiting examples of analogs include: 5-bromouracil, peptide nucleic acid, xeno nucleic acid, morpholino, locked nucleic acid, glycol nucleic acid, threose nucleic acid, dideoxynucleotide, cordycepin, 7-deaza-GTP, fluorophores (e.g., rhodamine or fluorescein linked to a sugar), thiol-containing nucleotides, biotin-linked nucleotides, fluorescent base analogs, CpG islands, methyl-7-guanosine, methylated nucleotides, inosine, thiouridine, pseudouridine, dihydrouridine, wybutosine, and queuosine. Non-limiting examples of polynucleotides include coding or non-coding regions of genes or gene fragments, loci defined by linkage analysis, exons, introns, messenger RNA (mRNA), transfer RNA (tRNA), ribosomal RNA (rRNA), short interfering RNA (siRNA), short hairpin RNA (shRNA), microRNA (miRNA), ribozymes, complementary DNA (cDNA, such as double-stranded cDNA (dd-cDNA) or single-stranded cDNA (ss-cDNA)), circulating tumor DNA (ctDNA), damaged DNA, recombinant polynucleotides, branched polynucleotides, plasmids, vectors, isolated DNA of any sequence, isolated RNA of any sequence, cell-free polynucleotides (including cell-free DNA (cfDNA) and cell-free RNA (cfRNA)), nucleic acid probes (e.g., fluorescence in situ hybridization (FISH) probes), and primers. The sequence of nucleotides can be interrupted by non-nucleotide components. Polynucleotides can contain one or more modified nucleotides, such as methylated nucleotides and nucleotide analogs. In the presence of modifications, nucleotide structures can be modified either before or after polymer assembly. The sequence of nucleotides can be interrupted by non-nucleotide components. Polynucleotides can be further modified after polymerization, such as by conjugation with a labeling component.

[0037] As used herein, the term "gene" generally refers to nucleic acids (e.g., DNA such as genomic DNA and cDNA) involved in encoding an RNA transcript and its corresponding nucleotide sequence. The term "gene" as used herein with respect to genomic DNA may include intervening non-coding regions as well as regulatory regions and may include 5' and 3' termini. In some uses, the term includes transcribed sequences including 5' and 3' untranslated regions (5'-UTR and 3'-UTR), exons, and introns. In some genes, the transcribed region will contain an "open reading frame" encoding a polypeptide. In some uses of the term, "gene" includes only the coding sequence necessary to encode a polypeptide (e.g., the "open reading frame" or "coding region"). A gene may not encode a polypeptide, such as ribosomal RNA (rRNA) genes and transfer RNA (tRNA) genes. The term "gene" may include not only transcribed sequences but also non-transcribed regions including upstream and downstream regulatory regions, enhancers, and promoters. A gene may be an "endogenous gene" or native gene in its natural location in an organism's genome. A gene may be an "exogenous gene" or non-native gene. A non-native gene may be a gene that is not normally found in the host organism but is introduced into the host organism by gene transfer (e.g., a transgene). A non-native gene may be a naturally occurring nucleic acid or polypeptide sequence containing a mutation, insertion, and / or deletion (e.g., a non-native sequence).

[0038] As used herein, the term "mutation" generally refers to a change in the nucleotide sequence of a normally conserved nucleic acid sequence resulting in the formation of a mutant different from the normal (unchanged) or wild-type sequence. Prior to sequencing, the location (e.g., relative to a gene or sample polynucleotide) and sequence of a mutation may be undetermined. Alternatively, the location (e.g., relative to a gene or sample polynucleotide) and sequence of a mutation may have been determined prior to sequencing, in which case sequencing may be performed to detect the presence of the mutation in the sample polynucleotide. Mutations may include base pair substitutions (e.g., single nucleotide substitutions) and frameshift mutations. Frameshift mutations may require the insertion or deletion of one to several nucleotide pairs.

[0039] As used herein, the term "probe" generally refers to a nucleotide or polynucleotide labeled with a label (e.g., a fluorescent label), which can be used to detect or identify its corresponding target nucleotide or polynucleotide by hybridizing with the corresponding target sequence in a hybridization reaction. As used interchangeably herein, the terms "nucleotide probe", "nucleotide label", and "labeled nucleotide" generally refer to a probe having a single nucleotide. As used interchangeably herein, the terms "polynucleotide probe", "polynucleotide label", and "labeled polynucleotide" generally refer to a probe having a polynucleotide. The polynucleotide probe can be labeled with at least one label (e.g., one label for each nucleotide of the polynucleotide probe). The probe can hybridize with one or more target nucleotides or polynucleotides. The polynucleotide probe can be completely complementary to one or more target polynucleotides in a sample, or can contain one or more nucleotides that are not complementary (i.e., mismatched) to one or more nucleotides of one or more target polynucleotides in the sample.

[0040] In some embodiments, the label can be a redox species. As used herein, the term "redox species" generally refers to a molecule or compound or a portion thereof (e.g., a molecular or functional portion of the molecule or compound) that can be oxidized and / or reduced (i.e., "redoxed") during or after an electrical stimulus (e.g., during or after application of a potential), or can undergo a Faraday reaction. In one example, the redox species can include one or more molecular moieties that accept and / or provide one or more electrons depending on their redox state. In some cases, the redox species can form part of a small molecule, compound, polymeric molecule (e.g., a molecular moiety), or can exist as a single molecule or compound. Examples of redox species can include imidazolium, pyrrolidinium, tetraalkylammonium, [OTf]-, [FAP]-, [PF6]-, [BF4]-, [DCA]-, [NTf2]-, [FSI]-, [B(CN)4]-, ferrocene (Fc), its derivatives, its functional variants, and combinations thereof. Examples of Fc derivatives can include methylferrocene, dimethylferrocene, ethylferrocene, propylferrocene, n-butylferrocene, tert-butylferrocene, and 1,1-dicarboxyferrocene.

[0041] The same nucleotide can be labeled with the same label. Optionally, the same nucleotide can be labeled with different labels. For example, the first nucleotide A can be labeled with a first label, and the second nucleotide A can be labeled with a second label, where the first label and the second label are different. Using multiple labels for the same nucleotide can help resolve the sequence of identical nucleotides presented in a sequential manner in cases where the sensor can detect and distinguish the first label and the second label from each other.

[0042] As used interchangeably herein, the terms "complementary", "complementary sequence", "complementary to", and "complementarity" generally refer to a sequence that is fully complementary to and hybridizable with a given sequence. A sequence that hybridizes to a given nucleic acid is called a "complementary sequence" or "reverse complementary sequence" of the given molecule, provided that its base sequence in a given region is capable of binding complementarily to the base sequence of its binding partner such that, for example, A-T, A-U, G-C, and G-U base pairs are formed. Generally, a first sequence that is hybridizable with a second sequence can hybridize specifically or selectively with the second sequence such that, during a hybridization reaction, it preferably hybridizes with the second sequence or a group of second sequences relative to hybridization with non-target sequences (e.g., is thermodynamically more stable under given conditions (such as stringent conditions commonly used in the art)). Generally, hybridizable sequences share a degree of sequence complementarity over all or a portion of their respective lengths, such as complementarity between 25% and 100%, including at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, and at least 100% sequence complementarity. The respective lengths can include regions having at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, or more nucleotides.

[0043] For purposes such as assessing the percentage of complementarity, sequence identity can be measured by any suitable alignment algorithm, including but not limited to the Needleman-Wunsch algorithm (see, for example, the EMBOSS Needle aligner available at www.ebi.ac.uk / Tools / psa / emboss_needle / nucleotide.html, optionally using default settings), the BLAST algorithm (see, for example, the BLAST alignment tool available at blast.ncbi.nlm.nih.gov / Blast.cgi, optionally using default settings), or the Smith-Waterman algorithm (see, for example, the EMBOSS Water aligner available at www.ebi.ac.uk / Tools / psa / emboss_water / nucleotide.html, optionally using default settings). Any suitable parameters of the selected algorithm, including default parameters, can be used to evaluate the optimal alignment.

[0044] Complementarity can be complete or substantial / enough. Complete complementarity between two nucleic acids can mean that the two nucleic acids can form a duplex, where each base in the duplex binds to a complementary base through Watson-Crick pairing. Substantial or enough complementarity can mean that the sequence in one strand is not completely and / or not perfectly complementary to the sequence in the opposite strand, but under a set of hybridization conditions (e.g., salt concentration and temperature), sufficient binding occurs between the bases on the two strands to form a stable hybrid complex. Such conditions can be predicted by using the sequence and standard mathematical calculations to predict the Tm of the hybridizing strands or by empirically determining the Tm using conventional methods.

[0045] As used herein, the term "hybridization" generally refers to a reaction in which one or more polynucleotides react to form a complex that is stabilized by hydrogen bonds between the bases of the nucleotide residues. The hydrogen bonds can occur through Watson Crick base pairing, Hoogstein binding, or in any other sequence-specific manner according to base complementarity. The complex can comprise two strands forming a duplex structure, three or more strands forming a multistranded complex, a self-hybridizing strand, or any combination thereof. The hybridization reaction can constitute a step in a broader process, such as initiating PCR or the enzymatic cleavage of polynucleotides by endonucleases. A second sequence complementary to a first sequence can be referred to as the "complementary sequence" of the first sequence. The term "hybridizable" as applied to a polynucleotide generally refers to the ability of the polynucleotide to form a complex that is stabilized by hydrogen bonds between the bases of the nucleotide residues in a hybridization reaction.

[0046] As used herein, the term "target polynucleotide" generally refers to a polynucleotide in a nucleic acid molecule or population of nucleic acid molecules having a target sequence, where the presence, number, and / or nucleotide sequence, or variations in one or more of them, is to be determined. The term "target sequence" generally refers to a nucleic acid sequence on a single strand of a nucleic acid. The target sequence can be part of a gene, a regulatory sequence, genomic DNA, cDNA, ctDNA, RNA (including mRNA, miRNA, rRNA), or others. The target sequence can be a target sequence from a sample or a secondary target, such as a product of an amplification reaction. The target polynucleotide can be part of a gene (or a fragment thereof) containing one or more mutations.

[0047] As used herein, the term "stringent conditions" generally refers to one or more hybridization conditions under which a nucleic acid having complementarity to a target sequence hybridizes primarily to the target sequence and substantially not to non-target sequences. Stringent conditions can be sequence-dependent and can vary depending on many factors. In some cases, the longer the sequence, the higher the temperature at which the sequence can hybridize specifically to its target sequence.

[0048] As used herein, the term "polymerase" generally refers to an enzyme (e.g., natural or synthetic) capable of catalyzing a polymerization reaction. Examples of polymerases can include nucleic acid polymerases (e.g., DNA polymerase or RNA polymerase) and transcriptases (e.g., reverse transcriptase). A polymerase can be a polymerization enzyme. The term "DNA polymerase" generally refers to an enzyme capable of catalyzing the polymerization of DNA.

[0049] As used herein, the term "linked polymerase" generally refers to a polymerase (such as a DNA polymerase) coupled to (e.g., fused to) a linker, where the linker can be capable of coupling (e.g., binding or conjugating to) another entity (e.g., a nanopore, such as a protein nanopore or a solid-state nanopore).

[0050] As used interchangeably herein, the terms “sequence variant” and “sequencing variant” generally refer to any sequence change relative to one or more reference sequences. Typically, for a given population of individuals for which a reference sequence is provided, the sequence variant occurs at a lower frequency than the reference sequence. For example, a particular bacterial genus may have a consensus reference sequence for the 16S rRNA gene, but individual species within that genus may have one or more sequence variants within the gene or a portion of the gene, which can be used to identify that species within the bacterial population. As a further example, when optimally aligned, the sequences of multiple individuals of the same species or multiple sequencing reads of the same individual may yield a consensus sequence, and sequence variants with respect to that consensus sequence can be used to identify mutants indicative of a hazardous contamination within the population. Generally, a “consensus sequence” refers to a nucleotide sequence that reflects the most common base selection at each position in the sequence, in which the relevant nucleic acid sequences have been subjected to in-depth mathematical and / or sequence analysis, such as optimal sequence alignment according to any of a variety of sequence alignment algorithms. The reference sequence can be a single reference sequence, such as a determined genomic sequence of a single individual. The reference sequence can be a consensus sequence formed by aligning multiple sequences, such known sequences being genomic sequences of multiple individuals serving as a reference population or multiple sequencing reads of polynucleotides from the same individual. The reference sequence can be a consensus sequence formed by optimally aligning sequences from the sample being analyzed, such that the sequence variant represents a change relative to the corresponding sequence in the same sample. Sequence variants can occur at low frequencies in the population (also referred to as “rare” sequence variants). For example, sequence variants can occur at a frequency less than or equal to 5%, 4%, 3%, 2%, 1.5%, 1%, 0.75%, 0.5%, 0.25%, 0.1%, 0.075%, 0.05%, 0.04%, 0.03%, 0.02%, 0.01%, 0.005%, 0.001% or lower. Sequence variants can occur at a frequency less than or equal to 0.1%.

[0051] A sequence variant can be any variation relative to a reference sequence. Sequence variations can include alterations, insertions or deletions of a single nucleotide or multiple nucleotides (such as 2, 3, 4, 5, 6, 7, 8, 9, 10 or more nucleotides). When a sequence variant contains two or more nucleotide differences, the different nucleotides can be contiguous or non-contiguous with each other. Examples of sequence variant types include single nucleotide polymorphisms (SNPs), deletion / insertion polymorphisms (DIPs), copy number variants (CNVs), short tandem repeats (STRs), simple sequence repeats (SSRs), variable number tandem repeats (VNTRs), amplified fragment length polymorphisms (AFLPs), retrotransposon-based insertion polymorphisms, sequence-specific amplified polymorphisms, and epigenetic marker differences (e.g., methylation differences) that can be detected as sequence variants.

[0052] As used herein, the term "sequencing" generally refers to procedures for determining the order in which nucleotides occur in a target nucleotide sequence. Sequencing methods can include high-throughput sequencing, such as next-generation sequencing (NGS). Sequencing can be whole-genome sequencing or targeted sequencing. Sequencing can be single-molecule sequencing or massively parallel sequencing. Next-generation sequencing methods can obtain millions of sequences in a single run. In one example, one or more nanopore sequencing methods can be used for sequencing, such as, for example, synthesis sequencing, ligation sequencing, or cleavage sequencing.

[0053] As used herein, the term "nanopore" generally refers to a pore, channel, or pathway formed or otherwise provided in a membrane. The membrane can be an organic membrane, such as a lipid bilayer, or a synthetic membrane, such as a membrane formed from a polymeric material such as a protein nanopore. The membrane can be a solid-state membrane (e.g., a silicon substrate). The nanopore can be disposed adjacent to or proximate to a sensing circuit or an electrode of a sensing circuit (e.g., a complementary metal-oxide semiconductor (CMOS) or field-effect transistor (FET) circuit). The nanopore can be part of the sensing circuit. The nanopore can have a characteristic width or diameter, e.g., from about 0.1 nanometers (nm) to 1000 nm. The nanopore can be a biological nanopore, a solid-state nanopore, a hybrid bio-solid-state nanopore, variants thereof, or combinations thereof. Examples of biological nanopores include, but are not limited to, OmpG from the genera Escherichia coli (E. coli), Salmonella, Shigella, and Pseudomonas, and alpha-hemolysin (α-hemolysin) from the genus Staphylococcus aureus (S. aureus), MspA from the genus Mycobacterium smegmatis (M. smegmatis), functional variants thereof, or combinations thereof. Sequencing can include forward sequencing and / or reverse sequencing. Examples of solid-state nanopores include, but are not limited to, silicon nitride, silicon oxide, graphene, molybdenum sulfide, functional variants thereof, or combinations thereof. Solid-state nanopores can be fabricated by high-energy beams, imprinting (e.g., nanoimprinting), laser ablation, chemical etching, plasma etching (e.g., oxygen plasma etching), etc.

[0054] As used herein, the term "nanopore sequencing complex" generally refers to a nanopore that is linked or coupled to an enzyme, such as a polymerase, which in turn is associated with a polymer, such as a polynucleotide template. The nanopore sequencing complex can be located in a membrane, such as a lipid bilayer, where its function is to identify polymer components, such as nucleotides or amino acids.

[0055] As used interchangeably herein, the terms “nanopore sequencing” and “nanopore-based sequencing” generally refer to methods for determining the sequence of a polynucleotide by means of a nanopore. In some cases, the sequence of a polynucleotide can be determined in a template-dependent manner. In some cases, the methods, systems or compositions disclosed herein may not be limited to any particular nanopore sequencing method, system or device.

[0056] As used herein, the term "barcode" generally refers to a defined nucleic acid sequence that permits identification of certain characteristics of a polynucleotide associated with the barcode (e.g., a polynucleotide that contains at least a portion of the barcode or a polynucleotide that is complementary to at least a portion of the barcode). In some instances, the characteristic of the polynucleotide to be identified can be the sample from which the polynucleotide is derived. The length of the barcode can be at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 11, at least about 12, at least about 13, at least about 14, at least about 15, at least about 16, at least about 17, at least about 18, at least about 19, at least about 20, or more nucleotides. The length of the barcode can be at most about 20, at most about 19, at most about 18, at most about 17, at most about 16, at most about 15, at most about 14, at most about 13, at most about 12, at most about 11, at most about 10, at most about 9, at most about 8, at most about 7, at most about 6, at most about 5, at most about 4, at most about 3, or at most about 2 nucleotides. A barcode associated with a polynucleotide from a first sample can be different from (e.g., a different sequence and / or a different length) a barcode associated with a polynucleotide from a second sample that is different from the first sample. In such cases, identification of the barcode in the respective polynucleotide can aid in identification of the sample source of one or more polynucleotides. Thus, different samples with different barcodes can be analyzed (e.g., sequenced) together (e.g., in batches), and separated at least in part based on the barcodes during the analysis. In some instances, the barcode can be accurately identified even after one or more nucleotide mutations, insertions, or deletions (e.g., at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, or more nucleotide mutations, insertions, or deletions) in the barcode sequence. Multiple polynucleotides from the same sample can have the same barcode. Alternatively, multiple polynucleotides from the same sample can have different barcodes. A first barcode can differ from a second barcode by at least three nucleotide positions, such as at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, or more nucleotide positions. Multiple barcodes can be represented in a sample library, each sample containing a polynucleotide that contains one or more barcodes that are different from barcodes contained in polynucleotides from other samples in the library.A polynucleotide sample containing one or more barcodes can be pooled according to the barcode sequences to which they are ligated, such that all four nucleotide bases A, G, C, and T appear approximately evenly in the library along one or more positions of each barcode (such as at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, or more positions or all positions of the barcode). In some instances, the methods of the present disclosure can include identifying the sample from which a target polynucleotide is derived based on the barcode sequence ligated to the target polynucleotide. The barcode can contain a nucleic acid sequence that can serve as an identifier of the sample from which the target polynucleotide is derived when ligated to the target polynucleotide. In one instance, an oligonucleotide primer (e.g., an amplification primer) can contain one or more barcodes. In another instance, a nucleic acid molecule can be conjugated (e.g., ligated) to an adapter nucleic acid (e.g., for circularization), and the adapter nucleic acid can contain one or more barcodes. The barcode can be a non-natural polynucleotide sequence, e.g., a synthetic polynucleotide sequence.

[0057] As used interchangeably herein, the terms "real-time" and "real time" generally refer to an event (e.g., an operation, a process, a measurement, a detection, etc.) that occurs almost immediately or within a short time period after another event (e.g., addition of a nucleobase, generation of a growing strand, etc.), such as within at least about 0.0001 milliseconds (ms), at least about 0.0005 ms, at least about 0.001 ms, at least about 0.005 ms, at least about 0.01 ms, at least about 0.05 ms, at least about 0.1 ms, at least about 0.5 ms, at least about 1 ms, at least about 5 ms, at least about 0.01 seconds, at least about 0.05 seconds, at least about 0.1 seconds, at least about 0.5 seconds, at least about 1 second, or longer. In some cases, a real-time event can occur almost immediately or within a short time period after another event, such as within at most about 1 second, at most about 0.5 seconds, at most about 0.1 seconds, at most about 0.05 seconds, at most about 0.01 seconds, at most about 5 ms, at most about 1 ms, at most about 0.5 ms, at most about 0.1 ms, at most about 0.05 ms, at most about 0.01 ms, at most about 0.005 ms, at most about 0.001 ms, at most about 0.0005 ms, at most about 0.0001 ms, or shorter.

[0058] As used herein, the term "sample" generally refers to any sample that can contain one or more components (e.g., nucleic acid molecules) for processing or analysis. The sample can be a biological sample. The sample can be a cell or tissue sample. The sample can be a cell-free sample, such as blood (e.g., whole blood), plasma, serum, sweat, saliva, or urine. The sample can be obtained in vivo or cultured in vitro.

[0059] As used herein, the term "subject" generally refers to an individual or entity from which a sample is derived, e.g., a vertebrate (e.g., a mammal such as a human) or an invertebrate. The mammal can be a mouse, a monkey, a human, a farm animal (e.g., a cow, a sheep, a pig, or a chicken) or a pet (e.g., a cat or a dog). The subject can be a plant. The subject can be a patient. The subject can be free of symptoms of a disease (e.g., cancer). Alternatively, the subject can have symptoms of a disease.

[0060] As used herein, the term "deconvolution" generally refers to identifying an individual (e.g., a single nucleotide or a specific nucleotide sequence) in a library or a collection by detecting the presence of known associated indicators (e.g., characteristics, markers, identifiers, etc.). In some cases, deconvolution of data measured by a sensor provided herein can be used to identify (i) at least a portion of a non-complementary region of a primer nucleic acid molecule and / or (ii) one or more nucleotides in a growing strand that are at least partially complementary to a target nucleic acid molecule. The deconvolution function for such deconvolution can be a fixed deconvolution function. Alternatively or in addition, the deconvolution function can be derived as part of an optimal fitting algorithm.

[0061] For example, a first nucleotide sequence (e.g., ACCT) can generate a first sensor signal (e.g., an impedance signal) having a first sensing signature, and another nucleotide sequence (e.g., GACC) can generate a second sensor signal having a second sensing signature. When the sensor detects the first sensor signal, the deconvolution method can identify (e.g., with a certain degree of certainty) that it is the first sensor signal having the first sensing signature, and thus can deconvolve the first sensor signal and assign the first nucleotide sequence (e.g., ACCT) as the detected base sequence. Similarly, when the sensor detects the second sensor signal, the deconvolution method can identify (e.g., with a certain degree of certainty) that it is the second sensor signal having the second sensing signature, and thus can deconvolve the second sensor signal and assign the second nucleotide sequence (e.g., GACC) as the detected base sequence.

[0062] I. Systems and Methods for Detecting and Analyzing Target Molecules

[0063] In one aspect, the present disclosure provides a method of processing (e.g., analyzing) a target nucleic acid (NA) molecule. The method can include providing a complex (e.g., a nucleic acid molecule complex) comprising the target nucleic acid molecule and a primer nucleic acid (NA) molecule. The primer nucleic acid molecule can comprise a plurality of regions. The plurality of regions can comprise (i) a region complementary to a portion of the target nucleic acid molecule and (ii) an additional region that is not complementary to the target nucleic acid molecule. The method can include analyzing the target nucleic acid molecule by using a sensor to identify the additional region after the additional region has flowed through a pore of the sensor.

[0064] The target nucleic acid molecule can be from a nucleic acid sample (NA sample). The NA sample can be a biological sample (e.g., at least a part of a bodily sample from a subject), and the target nucleic acid molecule can be at least a part of the biological sample (e.g., without further in vitro processing such as amplification by polymerase chain reaction). For example, the target nucleic acid molecule can be at least a part of a native nucleic acid molecule from a biological sample (e.g., a chromosomal fragment, cell-free DNA, etc.). Alternatively, the target nucleic acid molecule can be a synthetic molecule derived from a biological sample (e.g., a replicated, amplified, or modified variant of a native nucleic acid molecule). The target nucleic acid molecule can be a linear NA molecule or a circular NA molecule.

[0065] In the complex, the target NA molecule and the primer NA molecule can bind together. The target NA molecule and the primer NA molecule can be heterologous to each other (e.g., not from the same source or not from the same biological sample of a subject). Alternatively, the target NA molecule and the primer NA molecule can be homologous. The target NA molecule and the primer NA molecule can bind by a covalent bond. Alternatively or in addition, the target NA molecule and the primer NA molecule can bind by an ionic bond. Alternatively, or in addition, the target NA molecule and the primer NA molecule can bind by a hydrogen bond.

[0066] The target NA molecule and the primer NA molecule can be fully complementary. Alternatively, the target NA molecule and the primer NA molecule can be partially complementary (e.g., at least one region in the primer NA molecule is complementary to the target NA molecule and at least one region in the primer NA molecule is not complementary to the target NA molecule, and / or at least one region in the target NA molecule is complementary to the primer NA molecule and at least one region in the NA molecule is not complementary to the primer NA molecule), but not fully complementary to each other.

[0067] The length of a non-complementary region (e.g., a non-complementary region of a target NA molecule or a primer NA molecule when compared to each other) can comprise at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 11, at least about 12, at least about 13, at least about 14, at least about 15, at least about 16, at least about 17, at least about 18, at least about 19, at least about 20, at least about 25, at least about 30, at least about 35, at least about 40, at least about 45, at least about 50, at least about 55, at least about 60, at least about 65, at least about 70, at least about 75, at least about 80, at least about 85, at least about 90, at least about 95, or at least about 100 nucleobases (or bases, which can be used interchangeably herein). The length of the non-complementary region can comprise at most about 100, at most about 95, at most about 90, at most about 85, at most about 80, at most about 75, at most about 70, at most about 65, at most about 60, at most about 55, at most about 50, at most about 45, at most about 40, at most about 35, at most about 30, at most about 25, at most about 20, at most about 19, at most about 18, at most about 17, at most about 16, at most about 15, at most about 14, at most about 13, at most about 12, at most about 11, at most about 10, at most about 9, at most about 8, at most about 7, at most about 6, at most about 5, at most about 4, at most about 3, at most about 2, or at most about 1 base. In some cases, the non-complementary region can comprise multiple bases provided herein, and the multiple bases can be a continuous polynucleotide sequence or a non-continuous polynucleotide sequence.

[0068] The length of the complementary region (e.g., the complementary region of a target NA molecule or a primer NA molecule when compared to each other) can comprise at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 11, at least about 12, at least about 13, at least about 14, at least about 15, at least about 16, at least about 17, at least about 18, at least about 19, at least about 20, at least about 25, at least about 30, at least about 35, at least about 40, at least about 45, at least about 50, at least about 55, at least about 60, at least about 65, at least about 70, at least about 80, at least about 85, at least about 90, or at least about 100 nucleobases (or bases, which can be used interchangeably herein). The length of the complementary region can comprise at most about 100, at most about 95, at most about 90, at most about 85, at most about 80, at most about 75, at most about 70, at most about 65, at most about 60, at most about 55, at most about 50, at most about 45, at most about 40, at most about 35, at most about 30, at most about 25, at most about 20, at most about 19, at most about 18, at most about 17, at most about 16, at most about 15, at most about 14, at most about 13, at most about 12, at most about 11, at most about 10, at most about 9, at most about 8, at most about 7, at most about 6, at most about 5, at most about 4, at most about 3, at most about 2, or at most about 1 base. In some cases, the complementary region can comprise multiple bases provided herein, and the multiple bases can be a continuous polynucleotide sequence or a discontinuous polynucleotide sequence.

[0069] The complementary region of the primer NA molecule can be located downstream of the non-complementary region of the primer NA molecule (e.g., after the 3' end of the non-complementary region of the primer NA molecule). Alternatively, the complementary region of the primer NA molecule can be located upstream of the non-complementary region of the primer NA molecule (e.g., before the 5' end of the non-complementary region of the primer NA molecule).

[0070] The complementary region of the primer NA molecule can be substantially (e.g., completely) complementary to a portion of the target nucleic acid molecule, such that, for example, a complementary double-stranded portion is formed in the complex. Alternatively, the complementary region of the primer NA molecule can comprise a continuous polynucleotide sequence that is not completely complementary to a portion of the target nucleic acid molecule (e.g., forming A to T or U pairs, and G to C pairs throughout the continuous polynucleotide sequence). For example, in the primer NA molecule, the continuous polynucleotide sequence that is not complementary to the target NA molecule can be located between two NA regions (e.g., two continuous polynucleotide sequences), each of which is complementary to a different portion of the target NA molecule. The length of such a continuous polynucleotide sequence of the primer NA molecule can be at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, or at least about 10, at least about 12, at least about 13, at least about 14, at least about 15, at least about 20, at least about 25, at least about 30, at least about 35, at least about 40, at least about 45, at least about 50, at least about 60, at least about 80, at least about 100, or more nucleotides. The length of the continuous polynucleotide sequence of the primer NA molecule can be at most about 100, at most about 80, at most about 60, at most about 50, at most about 45, at most about 40, at most about 35, at most about 30, at most about 25, at most about 20, at most about 15, at most about 14, at most about 13, at most about 12, at most about 11, at most about 10, at most about 9, at most about 8, at most about 7, at most about 6, at most about 5, at most about 4, at most about 3, or at most about 2 nucleotides. Alternatively or in addition, the continuous polynucleotide sequence of the primer NA molecule can be partially complementary, such that, for example, less than a threshold portion (e.g., less than about 50%, less than about 40%, less than about 30%, less than about 20%, less than about 10%, less than about 5%, or less) of the continuous polynucleotide sequence forms complementary pairs with a portion of the target nucleic acid molecule.

[0071] The primer NA molecule can have one or more regions that are not complementary to the target NA molecule (e.g., one or more polynucleotide regions that are not directly adjacent to each other). The primer can have at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, or at least about 10 regions that are not complementary to the target NA molecule. The primer NA molecule can have at most about 10, at most about 9, at most about 8, at most about 7, at most about 6, at most about 5, at most about 4, at most about 3, at most about 2, or at most about 1 region that is not complementary to the target NA molecule.

[0072] The primer NA molecules disclosed herein may comprise (e.g., as an additional region of the primer NA molecule, or as part of an additional region) n repeats of a single nucleotide (e.g., A n , T n , C n , G n , etc.). For example, the additional region of the primer NA molecule may comprise T n , such as T 10 , T 20 , T 30 , T 40 , etc. Alternatively, or in addition, the primer NA molecule may comprise n repeats of a polynucleotide sequence, where the polynucleotide sequence may be a dinucleotide (e.g., two different and consecutive nucleotides, such as (AT) n , (GC) n , etc.), trinucleotide, tetranucleotide, pentanucleotide, etc. In some instances, the primer NA molecule may comprise n repeats, where n may be at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, or more. In some instances, the primer NA molecule may comprise n repeats, where n may be at most 100, at most 90, at most 80, at most 70, at most 60, at most 50, at most 45, at most 40, at most 35, at most 30, at most 25, at most 20, at most 19, at most 18, at most 17, at most 16, at most 15, at most 14, at most 13, at most 12, at most 11, at most 10, at most 9, at most 8, at most 7, at most 6, at most 5, at most 4, at most 3, or at most 2.

[0073] The primer NA molecules disclosed herein may comprise (e.g., as an additional region of the primer NA molecule, or as part of an additional region) at least a portion of a barcode. The complementary region of the primer NA molecule may comprise a barcode or a portion of a barcode. Alternatively, or in addition, the non-complementary region of the primer NA molecule may comprise a barcode or a portion of a barcode. Primer NA molecules for sequencing different target NA molecules may have the same barcode. Alternatively, or in addition, primer NA molecules for sequencing different target NA molecules may have different barcodes. For example, a plurality of different barcodes may be a set of unique polynucleotide sequences such that (i) one primer NA molecule or its complementary product can be distinguished from (ii) another primer NA molecule or its complementary product by its unique barcode (e.g., after analyzing the sequencing results). The primer NA molecule may comprise at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, or more barcodes. The primer NA molecule may comprise at most 10, at most 9, at most 8, at most 7, at most 6, at most 5, at most 4, at most 3, at most 2, or at most 1 barcode. Alternatively, the primer NA molecule may not comprise a barcode.

[0074] The primer NA molecule may facilitate an extension reaction of the target NA molecule (e.g., in the presence of an enzyme such as a polymerase), thereby generating a growing strand coupled to the primer NA molecule. The growing strand may be generated when the complex is adjacent to the sensor (e.g., coupled to the sensor, such as by binding moiety-nanopore unit binding). Initiation of the generation of the growing strand may occur before, during, or after at least a portion of the additional region flows through the pore of the sensor. The growing strand coupled to the primer NA molecule may be complementary to at least a portion of the target NA molecule. The growing strand coupled to the primer NA molecule may be partially complementary to the target NA molecule (e.g., having a non-complementary portion). In some instances, the complementary region of the primer NA molecule may flank (e.g., be located therebetween) the non-complementary region of the primer NA molecule and the growing strand.

[0075] In a primer NA molecule (e.g., before forming a complex with a target nucleic acid molecule, before generating a growing strand, etc.), the number of nucleobases in a region complementary to a portion of the target nucleic acid molecule can be at least or at most about 1, at least or at most about 2, at least or at most about 3, at least or at most about 4, at least or at most about 5, at least or at most about 6, at least or at most about 7, at least or at most about 8, at least or at most about 9, at least or at most about 10, at least or at most about 15, at least or at most about 20, at least or at most about 25, at least or at most about 30, at least or at most about 40, at least or at most about 50, at least or at most about 60, at least or at most about 70, at least or at most about 80, at least or at most about 90, or at least or at most about 100 nucleobases more than the number of nucleobases in an additional region that is not complementary to the target nucleic acid molecule. In such a primer NA molecule, the size of the region complementary to a portion of the target nucleic acid molecule can be at least or at most about 1%, at least or at most about 2%, at least or at most about 5%, at least or at most about 10%, at least or at most about 15%, at least or at most about 20%, at least or at most about 30%, at least or at most about 40%, at least or at most about 50%, at least or at most about 60%, at least or at most about 70%, at least or at most about 80%, at least or at most about 90%, at least or at most about 100%, at least or at most about 120%, at least or at most about 150%, at least or at most about 200%, at least or at most about 300%, at least or at most about 400%, at least or at most about 500%, or at least or at most about 1,000% larger than the size of the additional region that is not complementary to the target nucleic acid molecule.

[0076] Alternatively, in a primer NA molecule (e.g., before forming a complex with a target nucleic acid molecule, before generating a growing strand, etc.), the number of nucleobases in an additional region that is not complementary to the target nucleic acid molecule can be at least or at most about 1, at least or at most about 2, at least or at most about 3, at least or at most about 4, at least or at most about 5, at least or at most about 6, at least or at most about 7, at least or at most about 8, at least or at most about 9, at least or at most about 10, at least or at most about 15, at least or at most about 20, at least or at most about 25, at least or at most about 30, at least or at most about 40, at least or at most about 50, at least or at most about 60, at least or at most about 70, at least or at most about 80, at least or at most about 90, or at least or at most about 100 nucleobases more than the number of nucleobases in a region that is complementary to a portion of the target nucleic acid molecule. In such primer NA molecules, the size of the additional region that is not complementary to the target nucleic acid molecule can be at least or at most about 1%, at least or at most about 2%, at least or at most about 5%, at least or at most about 10%, at least or at most about 15%, at least or at most about 20%, at least or at most about 30%, at least or at most about 40%, at least or at most about 50%, at least or at most about 60%, at least or at most about 70%, at least or at most about 80%, at least or at most about 90%, at least or at most about 100%, at least or at most about 120%, at least or at most about 150%, at least or at most about 200%, at least or at most about 300%, at least or at most about 400%, at least or at most about 500%, or at least or at most about 1,000% larger than the size of the region that is complementary to a portion of the target nucleic acid molecule.

[0077] A nucleic acid molecule (e.g., a target NA molecule or a primer NA molecule) can be circular. Alternatively, the nucleic acid molecule can be non-circular. For example, the nucleic acid molecule can be linear. In some instances, one or both of the target NA molecule and the primer NA molecule in the complex can be linear. Alternatively or in addition, one or both of the target NA molecule and the primer NA molecule in the complex can be circular. The nucleic acid molecule can be single-stranded. Alternatively, the nucleic acid molecule can be double-stranded.

[0078] A nucleic acid molecule (e.g., a target NA molecule or a primer NA molecule) can have a net positive charge. In some cases, the net charge of the nucleic acid molecule can be +1, +2, +3, +4, +5, +6, +7, +8, +9, or +10. Alternatively, the nucleic acid molecule can have a net negative charge. In some cases, the net charge of the nucleic acid molecule can be -1, -2, -3, -4, -5, -6, -7, -8, -9, or -10. Alternatively, the nucleic acid molecule can have a neutral charge (e.g., a charge of 0 or no charge). For example, the primer NA molecule of the NA molecule complex can have a net positive charge. Alternatively, the primer NA molecule of the NA molecule complex can have a net negative charge. Alternatively, the primer NA molecule of the NA molecule complex can have a neutral charge. In another example, the complementary portion of the NA molecule complex can have a net positive charge. Alternatively, the complementary portion of the NA molecule complex can have a net negative charge. Alternatively, the complementary portion of the NA molecule complex can have a neutral charge. In a different example, the non-complementary portion of the NA molecule complex can have a net positive charge. Alternatively, the non-complementary portion of the NA molecule complex can have a net negative charge. Alternatively, the non-complementary portion of the NA molecule complex can have a neutral charge.

[0079] The target nucleic acid molecule can be a circular NA molecule. The circular NA molecule can be derived from a linear NA molecule (e.g., a native linear NA molecule, or a replicated or modified variant thereof), which is circularized by an adaptor (e.g., a synthetic polynucleotide sequence). Such circularization can be a connection between the linear NA molecule and the adaptor, e.g., non-enzymatic, chemical conjugation, or enzymatic conjugation. At least a portion of the primer NA molecule can exhibit complementarity to at least a portion of the adaptor. Alternatively, the primer NA molecule can not exhibit complementarity to the adaptor. The adaptor can comprise at least about 2, at least about 5, at least about 10, at least about 15, at least about 20, at least about 25, at least about 30, at least about 40, at least about 50, at least about 100, or more nucleobases. The adaptor can comprise at most about 100, at most about 50, at most about 40, at most about 30, at most about 25, at most about 20, at most about 15, at most about 10, at most about 5, or at most about 2 nucleobases.

[0080] The methods provided herein can include processing (e.g., analyzing or identifying) a target molecule, such as a target nucleic acid molecule. One or more sensors can be utilized for processing. The sensor can be configured to detect one or more signals (e.g., a current, voltage, impedance, or a change thereof in the sensor) when at least a portion of the target molecule binds to or approaches at least a portion of the sensor. The one or more signals can be used to analyze or identify the target molecule. For example, the one or more signals can indicate an impedance or a change in impedance in the sensor.

[0081] One or more sensors may comprise at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30, at least about 40, at least about 50, at least about 60, at least about 70, at least about 80, at least about 90, at least about 100, at least about 200, at least about 300, at least about 400, at least about 500, at least about 600, at least about 700, at least about 800, at least about 900, at least about 1,000 or more sensors (e.g., within a single sensor chip, such as within a single sequencing chip). One or more sensors may comprise at most about 1,000, at most about 900, at most about 800, at most about 700, at most about 600, at most about 500, at most about 400, at most about 300, at most about 200, at most about 100, at most about 90, at most about 80, at most about 70, at most about 60, at most about 50, at most about 40, at most about 30, at most about 20, at most about 10, at most about 9, at most about 8, at most about 7, at most about 6, at most about 5, at most about 4, at most about 3, at most about 2 or fewer sensors.

[0082] A detection signal from one or more sensors caused by a target molecule (e.g., indicative of an impedance or impedance change in the sensor) can be a single measurement. Alternatively, the detection signal can be a median or average of multiple measurements, e.g., at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 15, at least about 20 or more measurements, or at most about 20, at most about 15, at most about 10, at most about 9, at most about 8, at most about 7, at most about 6, at most about 5, at most about 4, at most about 3, or at most about 2 measurements.

[0083] In some cases, the sensor may comprise a nanopore protein. The nanopore protein may comprise a pore. Non-limiting examples of nanopore proteins may include porins (e.g., MspA porin), hemolysin, aerolysin, transporter proteins, channel proteins, and proteins comprising a β-barrel domain structure. In some cases, the sensor may comprise a solid-state nanopore (e.g., a non-protein nanopore).

[0084] In some cases, at least a portion of the nanopore may be embedded in a membrane. The membrane can be a lipid bilayer. Alternatively, the membrane can be a single lipid layer. In some instances, the membrane is a biological membrane (e.g., a cell membrane).

[0085] To detect one or more signals (e.g., indicative of impedance or impedance change in a sensor), at least a portion of the complex can bind (e.g., directly or indirectly) to a binding portion of the sensor. The binding portion can be configured to bind to at least a portion of a target molecule (e.g., nucleotide, amino acid, small molecule, ion, etc.). Alternatively, or in addition, the binding portion can be configured to bind to at least a portion of a primer nucleic acid molecule. In one instance, at least a portion of the complementary regions of the target nucleic acid molecule and the primer nucleic acid molecule can bind through the binding portion.

[0086] A sensor and a binding unit can be operably coupled to each other as a single unit (e.g., nanopore sensor - polymerase binding unit fusion polypeptide). In such a case, the number of sensors and the number of binding units can be the same (e.g., within a single sensor chip, such as within a single sequencing chip). Alternatively, a single sensor (e.g., nanopore) can be operably coupled to multiple binding units (e.g., multiple polymerases).

[0087] The binding portion can comprise an enzyme such that at least a portion of the nucleic acid complex can contact the enzyme, e.g., before an additional region of the primer nucleic acid molecule flows through the pore of the sensor. Non-limiting examples of such enzymes can include DNA polymerase, RNA polymerase, DNA primase, DNA helicase, DNA ligase, or topoisomerase. The enzyme can be operably coupled to the sensor (e.g., the nanopore of the sensor). The enzyme can be covalently coupled to the sensor (e.g., covalently linked to the sensor as a fusion polypeptide construct). Alternatively, the enzyme can be non-covalently coupled to the sensor.

[0088] The complex provided herein can be formed prior to contact with the sensor (e.g., the binding portion of the sensor). Alternatively, a portion of the components that make up the complex (e.g., one but not both of the target nucleic acid molecule and the primer nucleic acid molecule) can contact the sensor (e.g., the binding portion of the sensor) prior to forming the complex.

[0089] Contact of the enzyme with at least a portion of the NA complex can enable a portion of the complex to flow into the pore of the sensor. A portion of the complex that enters the pore of the sensor can be at least a portion of the target NA molecule. Alternatively, a portion of the complex that enters the pore of the sensor can be at least a portion of the primer NA molecule. In some cases, a portion of the primer NA molecule that enters the pore of the sensor can be the complementary region. Alternatively, a portion of the primer NA molecule that enters the pore of the sensor can be the non-complementary region. For example, after at least a portion of the complex binds to the enzyme coupled to the nanopore, the non-complementary region of the primer nucleic acid molecule can be directed into or flow through the pore of the nanopore.

[0090] NA molecular complexes can form before a portion of the complex flows into the pores of the sensor. Alternatively, NA molecular complexes can form after a portion of the complementary regions of the nucleic acids flow into the pores.

[0091] In some cases, the contact of the binding moiety (e.g., an enzyme) with the NA complex may be sufficient to effect the flow of a portion of the complex (e.g., an additional region of a primer NA molecule) into the pores of the sensor. For example, after contact, a portion of the complex (e.g., a portion of a primer NA molecule, such as a non-complementary additional region) can flow naturally into the pores of the sensor, e.g., without any additional external force. Without wishing to be bound by theory, the additional non-complementary region of the primer NA molecule may be attracted towards the pores of the sensor, which is at least in part due to the close proximity between the additional non-complementary region of the primer NA molecule and the pores of the sensor. Alternatively, the contact may not be sufficient to effect the flow, and an additional external force (e.g., an electric field between two electrodes located at opposite ends of the pore provided herein) may be required to drive the flow.

[0092] In some cases, a portion of the complex entering the pores (e.g., the non-complementary region of a primer nucleic acid molecule) can be recruited towards the pores or can be directed into the pores of the sensor by controlling the electrical properties of the sensor (e.g., controlling the electric field on the sensor). The electric field can be enhanced to recruit / attract a portion of the complex towards and / or into the pores. Alternatively, the electric field can be weakened to recruit / attract a portion of the complex towards and / or into the pores. In some cases, when changing the electric field to recruit / attract the non-complementary region of the primer nucleic acid molecule towards the pores of the sensor, the electric field on the sensor can be continuously controlled (e.g., continuously enhanced or weakened) to effect the flow of at least a portion of (i) the remaining portion of the primer NA molecule and / or (ii) the growing strand coupled to the primer NA molecule through the pores. Alternatively, the electric field can be kept constant while at least a portion of (i) the remaining portion of the primer NA molecule and / or (ii) the growing strand flows through the pores. However, in another alternative, it is also possible to allow at least a portion of (i) the remaining portion of the primer NA molecule and / or (ii) the growing strand to flow through the pores without controlling the electric field (e.g., due to the inherent action of the pores such as nanopore proteins).

[0093] An electrical signal can be applied to the sensors provided herein. The electrical signal of the sensor can be constant. Alternatively, the electrical signal on the sensor can be an alternating current signal. In some instances, the electrical signal or a change thereof can direct a portion of the NA complex into the pores of the sensor. Alternatively, or in addition, the electrical signal or a change thereof can direct a portion of the NA complex to flow through the pores of the sensor. Alternatively, or in addition, the electrical signal or a change thereof can direct a portion of the NA complex to move reversely through the sensor (e.g., move in one direction and then move in the opposite direction). Alternatively, or in addition, the electrical signal or a change thereof can direct a portion of the NA complex to stop moving relative to the pores of the sensor (e.g., get stuck or immobile within the pores of the sensor).

[0094] The electrical signal or a change thereof can direct at least a portion (e.g., an appended non-complementary region) of the primer NA molecule to pass through the pores of the sensor at least about 1 time, at least about 2 times, at least about 3 times, at least about 4 times, at least about 5 times, at least about 6 times, at least about 7 times, at least about 8 times, at least about 9 times, at least about 10 times or more. The electrical signal can direct at least a portion of the primer NA molecule to pass through the pores of the sensor at most about 10 times, at most about 9 times, at most about 8 times, at most about 7 times, at most about 6 times, at most about 5 times, at most about 4 times, at most about 3 times, at most about 2 times or less.

[0095] In one aspect, the present disclosure provides a system for implementing any of the methods provided herein. The system (e.g., apparatus) can include a sensor provided herein for analyzing a target nucleic acid molecule. The system can further include a controller for controlling one or more (or any) operations of analyzing the target nucleic acid molecule. The controller can include a computer processor. The controller can be operably coupled to the sensor. The controller can be configured to control the electrical signal or a change thereof across the sensor (e.g., across the nanopore). The controller can obtain data indicative of one or more signal readings (e.g., current, voltage, impedance, or a change thereof in the sensor), and identify (e.g., sequence) the nucleic acid molecule flowing through the pore at least in part based on the analysis of the data. In some instances, the controller can identify the primer NA molecule (e.g., identify the barcode of the primer NA molecule). The controller can identify a portion of the primer that is not complementary to the target NA molecule. Alternatively or in addition, the controller can identify a portion of the primer NA molecule that is complementary to the target NA molecule. Alternatively or in addition, the controller can identify the growing strand coupled to the primer molecule.

[0096] Alternatively or in addition, the controller can be non-operably coupled to the sensor; rather, the controller can be operably coupled to the membrane embedded in the sensor.

[0097] The controller can be configured to control the electrical signal of the sensor. The controller can be configured to turn on and off the electrical signal. The controller can be configured to vary (e.g., enhance or attenuate) the electrical signal. The controller can be configured to alternately vary the electrical signal. The controller can be configured to maintain the electrical signal. The controller can be configured to identify the sequence reads of nucleic acid molecules flowing through the pore at least in part based on data indicating one or more signal readings from the sensor. The controller can identify the sequence reads based on characteristic sequence information (e.g., at least in part based on data indicating one or more signal readings from the sensor).

[0098] The growing strand coupled to the primer NA molecule can be directed through the pore of the sensor. The non-complementary region of the primer NA molecule can be directed through the pore of the sensor before at least a portion of the growing strand enters and / or flows through the pore of the sensor. Alternatively or in addition, the non-complementary region of the primer NA molecule can be directed through the pore of the sensor after at least a portion of the growing strand enters and / or flows through the pore of the sensor. The non-complementary region of the primer NA molecule can be directed through the pore of the sensor before the complementary region of the primer NA molecule enters and / or flows through the pore of the sensor. Alternatively or in addition, the non-complementary region of the primer NA molecule can be directed through the pore of the sensor after the complementary region of the primer NA molecule enters or flows through the pore of the sensor. The complementary region of the primer NA molecule can be directed through the pore of the sensor before at least a portion of the growing strand enters and / or flows through the pore of the sensor. Alternatively or in addition, the complementary region of the primer NA molecule can be directed through the pore of the sensor after at least a portion of the growing strand enters or flows through the pore of the sensor.

[0099] For example, the non-complementary region of the primer NA molecule, the complementary region of the primer NA molecule, and at least a portion of the growing strand coupled to the complementary region of the primer NA molecule (e.g., by the polymerase provided herein) can enter and flow through the nanopore in a sequential manner (e.g., in the order described herein), and thus, the controller can obtain signal data from the nanopore sensor and identify (e.g., sequence) the non-complementary region of the primer NA molecule, the complementary region of the primer NA molecule, and at least a portion of the growing strand in a sequential manner, respectively.

[0100] Sequence information related to the target nucleic acid molecule can be obtained from the NA molecule complex. In some instances, the sequence information can be obtained by sequencing at least a portion of the primer NA molecule (e.g., at least a portion of the non-complementary region) and at least a portion of the growing strand coupled to the primer NA molecule. In some instances, the sequence information of the entire primer NA molecule and at least a portion of the growing strand can be obtained.

[0101] Sequence information can include identifying sequence reads from data indicative of one or more signal readings from a sensor. The sequence reads can have characteristic sequence information indicative of at least a portion of a non-complementary region of a primer NA molecule. Alternatively or in addition, the sequence reads can have characteristic sequence information indicative of at least a portion of a complementary region of a primer NA molecule and / or a growing strand coupled to the primer NA molecule.

[0102] Figure 1 An example of a system (e.g., for direct read sequencing) and method thereof that includes a nanopore embedded in a membrane bilayer is schematically shown. A polymerase is operably coupled to the nanopore. A target nucleic acid molecule (e.g., a circular template nucleic acid molecule) 105 is coupled (e.g., by annealing) to a linear primer nucleic acid molecule that includes (i) a complementary region 110 corresponding to at least a portion of the target nucleic acid molecule and (ii) a non-complementary single-stranded overhang 120, thereby forming a nucleic acid complex. See Figure 1 Left portion. The nucleic acid complex contacts the polymerase coupled to the nanopore sensor. Before, during, or after the nucleic acid complex contacts the polymerase, the non-complementary region 120 enters through the nanopore, and the complementary region 110 enters through the nanopore after the non-complementary region 120 enters. See Figure 1 Middle portion. Based on the template, the polymerase synthesizes a growing strand (130a and 130b) that is coupled to the complementary region 110 of the primer nucleic acid molecule. The non-complementary overhang guides the linear primer NA molecule into and / or through the nanopore. An initial portion of the growing strand 130a flows through the nanopore, and at least a portion of a subsequent portion of the growing strand 130b can flow through the nanopore, and a signal or a change thereof from the nanopore sensor (e.g., impedance or impedance change of the nanopore sensor) can be detected by the nanopore sensor, and the resulting data can be used to identify the target nucleic acid molecule.

[0103] In some cases, impedance or impedance changes can be detected by applying a constant voltage (e.g., a sinusoidal voltage perturbation) while measuring a current (e.g., a current change). The impedance value (Z) can be measured by dividing the value of the applied voltage (V) by the value of the measured current (I). The number of impedance measurements made can be at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, or more. The number of impedance measurements made can be at most 20, at most 19, at most 18, at most 17, at most 16, at most 15, at most 14, at most 13, at most 12, at most 11, at most 10, at most 9, at most 8, at most 7, at most 6, at most 5, at most 4, at most 3, at most 2, or at most 1.

[0104] There can be a time lag between a first impedance measurement and a second impedance measurement. The first impedance measurement and the second impedance measurement can measure the impedance or its changes of two different nucleobases incorporated into a growing strand, respectively. Alternatively, the first impedance measurement and the second impedance measurement can measure the impedance or its changes of the same nucleobase incorporated into a growing strand. For example, multiple impedance measurements (e.g., at least 2, 3, 4, 5, 6, 7, 8, 19, 10, 15, 20, 25, 30, 40, 50 times or more) can be made on the same nucleobase incorporated into a growing strand. In some cases, the time lag between the first impedance measurement and the second impedance measurement can be at least about 0.1 millisecond, at least about 0.2 millisecond, at least about 0.3 millisecond, at least about 0.4 millisecond, at least about 0.5 millisecond, at least about 0.6 millisecond, at least about 0.7 millisecond, at least about 0.8 millisecond, at least about 0.9 millisecond, at least about 1 millisecond, at least about 1.5 milliseconds, at least about 2 milliseconds, at least about 3 milliseconds, at least about 4 milliseconds, at least about 5 milliseconds, at least about 6 milliseconds, at least about 7 milliseconds, at least about 8 milliseconds, at least about 9 milliseconds, at least about 10 milliseconds, or longer. In some cases, the time lag between the first impedance measurement and the second impedance measurement can be at most about 10 milliseconds, at most about 9 milliseconds, at most about 8 milliseconds, at most about 7 milliseconds, at most about 6 milliseconds, at most about 5 milliseconds, at most about 4 milliseconds, at most about 3 milliseconds, at most about 2 milliseconds, at most about 1.5 milliseconds, at most about 1 millisecond, at most about 0.9 millisecond, at most about 0.8 millisecond, at most about 0.7 millisecond, at most about 0.6 millisecond, at most about 0.5 millisecond, at most about 0.4 millisecond, at most about 0.3 millisecond, at most about 0.2 millisecond, at most about 0.1 millisecond, or shorter.

[0105] The impedance or change in impedance detected, for example, between a sensing electrode and a reference electrode can be at least about 1 microohm, 2 microohms, 3 microohms, 4 microohms, 5 microohms, 6 microohms, 7 microohms, 8 microohms, 9 microohms, 10 microohms, 20 microohms, 30 microohms, 40 microohms, 50 microohms, 60 microohms, 70 microohms, 80 microohms, 90 microohms, 100 microohms, 200 microohms, 300 microohms, 400 microohms, 500 microohms, 600 microohms, 700 microohms, 800 microohms, 900 microohms, 1 milliohm, 2 milliohms, 3 milliohms, 4 milliohms, 5 milliohms, 6 milliohms, 7 milliohms, 8 milliohms, 9 milliohms, 10 milliohms, 20 milliohms, 30 milliohms, 40 milliohms, 50 milliohms, 60 milliohms, 70 milliohms, 80 milliohms, 90 milliohms, 100 milliohms, 200 milliohms, 300 milliohms, 400 milliohms, 500 milliohms, 600 milliohms, 700 milliohms, 800 milliohms, 900 milliohms, 1 ohm, 2 ohms, 3 ohms, 4 ohms, 5 ohms, 6 ohms, 7 ohms, 8 ohms, 9 ohms, 10 ohms, 20 ohms, 30 ohms, 40 ohms, 50 ohms, 60 ohms, 70 ohms, 80 ohms, 90 ohms, 100 ohms, 200 ohms, 300 ohms, 400 ohms, 500 ohms, 600 ohms, 700 ohms, 800 ohms, 900 ohms, 1 kiloohm, 2 kiloohms, 3 kiloohms, 4 kiloohms, 5 kiloohms, 6 kiloohms, 7 kiloohms, 8 kiloohms, 9 kiloohms, 10 kiloohms, 20 kiloohms, 30 kiloohms, 40 kiloohms, 50 kiloohms, 60 kiloohms, 70 kiloohms, 80 kiloohms, 90 kiloohms, 100 kiloohms, 200 kiloohms, 300 kiloohms, 400 kiloohms, 500 kiloohms, 600 kiloohms, 700 kiloohms, 800 kiloohms, 900 kiloohms, 1,000 kiloohms, 10,000 kiloohms, 100,000 kiloohms, 1 gigaohm, 10 gigaohms, 50 gigaohms, 100 gigaohms or greater.The impedance or change in impedance detected, for example, between a sensing electrode and a reference electrode can be at most about 100 gigaohms, 50 gigaohms, 10 gigaohms, 1 gigaohm, 100,000 kiloohms, 10,000 kiloohms, 1,000 kiloohms, 900 kiloohms, 800 kiloohms, 700 kiloohms, 600 kiloohms, 500 kiloohms, 400 kiloohms, 300 kiloohms, 200 kiloohms, 100 kiloohms, 90 kiloohms, 80 kiloohms, 70 kiloohms, 60 kiloohms, 50 kiloohms, 40 kiloohms, 30 kiloohms, 20 kiloohms, 10 kiloohms, 9 kiloohms, 8 kiloohms, 7 kiloohms, 6 kiloohms, 5 kiloohms, 4 kiloohms, 3 kiloohms, 2 kiloohms, 1 kiloohm, 900 ohms, 800 ohms, 700 ohms, 600 ohms, 500 ohms, 400 ohms, 300 ohms, 200 ohms, 100 ohms, 90 ohms, 80 ohms, 70 ohms, 60 ohms, 50 ohms, 40 ohms, 30 ohms, 20 ohms, 10 ohms, 9 ohms, 8 ohms, 7 ohms, 6 ohms, 5 ohms, 4 ohms, 3 ohms, 2 ohms, 1 ohm, 900 milliohms, 800 milliohms, 700 milliohms, 600 milliohms, 500 milliohms, 400 milliohms, 300 milliohms, 200 milliohms, 100 milliohms, 90 milliohms, 80 milliohms, 70 milliohms, 60 milliohms, 50 milliohms, 40 milliohms, 30 milliohms, 20 milliohms, 10 milliohms, 9 milliohms, 8 milliohms, 7 milliohms, 6 milliohms, 5 milliohms, 4 milliohms, 3 milliohms, 2 milliohms, 1 milliohm, 900 microohms, 800 microohms, 700 microohms, 600 microohms, 500 microohms, 400 microohms, 300 microohms, 200 microohms, 100 microohms, 90 microohms, 80 microohms, 70 microohms, 60 microohms, 50 microohms, 40 microohms, 30 microohms, 20 microohms, 10 microohms, 9 microohms, 8 microohms, 7 microohms, 6 microohms, 5 microohms, 4 microohms, 3 microohms, 2 microohms, 1 microohm or less.

[0106] The impedance or change in impedance detected, for example, between a sensing electrode and a reference electrode can be a measurement (e.g., a single measurement, multiple measurements to yield an average of multiple measurements) that is made over a period of at least about 1 nanosecond, 2 nanoseconds, 3 nanoseconds, 4 nanoseconds, 5 nanoseconds, 6 nanoseconds, 7 nanoseconds, 8 nanoseconds, 9 nanoseconds, 10 nanoseconds, 20 nanoseconds, 30 nanoseconds, 40 nanoseconds, 50 nanoseconds, 60 nanoseconds, 70 nanoseconds, 80 nanoseconds, 90 nanoseconds, 100 nanoseconds, 200 nanoseconds, 300 nanoseconds, 400 nanoseconds, 500 nanoseconds, 600 nanoseconds, 700 nanoseconds, 800 nanoseconds, 900 nanoseconds, 1 microsecond, 2 microseconds, 3 microseconds, 4 microseconds, 5 microseconds, 6 microseconds, 7 microseconds, 8 microseconds, 9 microseconds, 10 microseconds, 20 microseconds, 30 microseconds, 40 microseconds, 50 microseconds, 60 microseconds, 70 microseconds, 80 microseconds, 90 microseconds, 100 microseconds, 200 microseconds, 300 microseconds, 400 microseconds, 500 microseconds, 600 microseconds, 700 microseconds, 800 microseconds, 900 microseconds, 1 millisecond, 2 milliseconds, 3 milliseconds, 4 milliseconds, 5 milliseconds, 6 milliseconds, 7 milliseconds, 8 milliseconds, 9 milliseconds, 10 milliseconds, 20 milliseconds, 30 milliseconds, 40 milliseconds, 50 milliseconds, 60 milliseconds, 70 milliseconds, 80 milliseconds, 90 milliseconds, 100 milliseconds, 200 milliseconds, 300 milliseconds, 400 milliseconds, 500 milliseconds, 600 milliseconds, 700 milliseconds, 800 milliseconds, 900 milliseconds, 1 second, 2 seconds, 3 seconds, 4 seconds, 5 seconds, 6 seconds, 7 seconds, 8 seconds, 9 seconds, 10 seconds or greater.The impedance or impedance change detected, for example, between a sensing electrode and a reference electrode can be a measurement (e.g., a single measurement, multiple measurements to yield an average of multiple measurements) that is made over a period of time of at most about 10 seconds, 9 seconds, 8 seconds, 7 seconds, 6 seconds, 5 seconds, 4 seconds, 3 seconds, 2 seconds, 1 second, 900 milliseconds, 800 milliseconds, 700 milliseconds, 600 milliseconds, 500 milliseconds, 400 milliseconds, 300 milliseconds, 200 milliseconds, 100 milliseconds, 90 milliseconds, 80 milliseconds, 70 milliseconds, 60 milliseconds, 50 milliseconds, 40 milliseconds, 30 milliseconds, 20 milliseconds, 10 milliseconds, 9 milliseconds, 8 milliseconds, 7 milliseconds, 6 milliseconds, 5 milliseconds, 4 milliseconds, 3 milliseconds, 2 milliseconds, 1 millisecond, 900 microseconds, 800 microseconds, 700 microseconds, 600 microseconds, 500 microseconds, 400 microseconds, 300 microseconds, 200 microseconds, 100 microseconds, 90 microseconds, 80 microseconds, 70 microseconds, 60 microseconds, 50 microseconds, 40 microseconds, 30 microseconds, 20 microseconds, 10 microseconds, 9 microseconds, 8 microseconds, 7 microseconds, 6 microseconds, 5 microseconds, 4 microseconds, 3 microseconds, 2 microseconds, 1 microsecond, 900 nanoseconds, 800 nanoseconds, 700 nanoseconds, 600 nanoseconds, 500 nanoseconds, 400 nanoseconds, 300 nanoseconds, 200 nanoseconds, 100 nanoseconds, 90 nanoseconds, 80 nanoseconds, 70 nanoseconds, 60 nanoseconds, 50 nanoseconds, 40 nanoseconds, 30 nanoseconds, 20 nanoseconds, 10 nanoseconds, 9 nanoseconds, 8 nanoseconds, 7 nanoseconds, 6 nanoseconds, 5 nanoseconds, 4 nanoseconds, 3 nanoseconds, 2 nanoseconds, 1 nanosecond or less.

[0107] As provided herein, the present disclosure provides a system for analyzing or identifying a target molecule. The system can have the necessary components to be configured for practicing or implementing any of the methods disclosed herein. The system can include a sensor that includes a sensing electrode and a reference electrode that are in electrical communication with each other. The sensor can include a dielectric material coupled to the sensing electrode and covering a first portion of the surface of the sensing electrode. The sensor can include a conductive material coupled to the sensing electrode and covering a second portion of the surface of the sensing electrode. The sensor can include a binding unit coupled to the conductive material, wherein the binding unit is configured to bind to the target molecule. The sensor can be configured to detect one or more signals indicative of an impedance or impedance change in the sensor when at least a portion of the target molecule binds to the binding unit. The one or more signals can be used to analyze or identify the target molecule. The conductive material can be a bond (e.g., a chemical bond), or can include a linking unit (e.g., a nanorod, a peptide, a small molecule, etc.) of any desired size (e.g., length, cross-sectional diameter or area, volume, etc.). In an alternative aspect, the binding unit can be directly coupled to the sensing electrode. However, in different aspects, the binding unit can be coupled to at least a portion of the dielectric material that is coupled to the sensing electrode.

[0108] One or more signals may indicate (i) a resistance in the sensor or a change thereof, (ii) a capacitance in the sensor or a change thereof, or (iii) an inductance in the sensor or a change thereof. One or more signals may indicate at least two of the following: (i) a resistance in the sensor or a change thereof, (ii) a capacitance in the sensor or a change thereof, and (iii) an inductance in the sensor or a change thereof. One or more signals may indicate (i) a resistance in the sensor or a change thereof, (ii) a capacitance in the sensor or a change thereof, or (iii) an inductance in the sensor and a change thereof.

[0109] One or more signals may be a current or a voltage. One or more signals may be a current and a voltage. One or more signals may not be a tunneling current.

[0110] The first portion of the sensing electrode covered by the dielectric material may be at least 50 percent (%) of the surface of the sensing electrode. In some cases, the first portion of the sensing electrode may be at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 99% or more of the surface of the sensing electrode. In some cases, the first portion of the sensing electrode may be at most 100%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50% or less of the surface of the sensing electrode.

[0111] The second portion of the sensing electrode may be at most 50% of the surface of the sensing electrode. In some cases, the second portion of the sensing electrode may be at most 50%, 45%, 40%, 35%, 30%, 25%, 20%, 15%, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, 1% or less of the surface of the sensing electrode. In some cases, the second portion of the sensing electrode may be at least 1%, 2%, 3%, 4%, 5%, 7%, 8%, 9%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50% or more of the surface of the sensing electrode.

[0112] The average cross-sectional dimension of the sensing electrode (or the surface area coupled to the dielectric material and the conductive material) may be greater than the average size of the target molecule by no more than 100 times. In some cases, the average cross-sectional dimension of the sensing electrode may be greater than the average size of the target molecule by up to 100 times, 90 times, 80 times, 70 times, 60 times, 50 times, 40 times, 30 times, 25 times, 20 times, 15 times, 10 times, 9 times, 8 times, 7 times, 6 times, 5 times, 4 times, 3 times, 2 times, 1 time, 0.5 times, or 0.1 times. In some cases, the average cross-sectional dimension of the sensing electrode may be greater than the average size of the target molecule by at least 0.1 times, 0.5 times, 1 time, 2 times, 3 times, 4 times, 5 times, 6 times, 7 times, 8 times, 9 times, 10 times, 15 times, 20 times, 25 times, 30 times, 40 times, 50 times, 60 times, 70 times, 80 times, 90 times, or 100 times. In some cases, the average cross-sectional dimension of the sensing electrode may not be greater than at least 0.1 nanometers (nm), 0.5 nm, 1 nm, 2 nm, 3 nm, 4 nm, 5 nm, 10 nm, 50 nm, 100 nm, 500 nm, 1,000 nm, 5,000 nm, 10,000 nm, or greater than the average size of the target molecule. In some cases, the average cross-sectional dimension of the sensing electrode may be at most 10,000 nm, 5,000 nm, 1,000 nm, 500 nm, 100 nm, 50 nm, 10 nm, 5 nm, 4 nm, 3 nm, 2 nm, 1 nm, 0.5 nm, 0.1 nm, or less than the average size of the target molecule.

[0113] Optionally, the average cross-sectional dimension of the sensing electrode may be less than the average size of the target molecule. In some cases, the average cross-sectional dimension of the sensing electrode may be less than the average size of the target molecule by at least 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 15%, 20%, 25%, 30%, 40%, or 50%. In some cases, the average cross-sectional dimension of the sensing electrode may be less than the average size of the target molecule by at most 50%, 40%, 30%, 20%, 15%, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, or 1%.

[0114] The area of the second portion of the sensing electrode surface may not be greater than

[0115] or less. In some cases, the cross-sectional dimension or diameter of the second portion of the sensing electrode surface may be approximately equal to the diameter of the atoms of the conductive material (e.g., twice the van der Waals radius).

[0116] The dielectric material may be a solid layer (e.g., a solid metal or semiconductor material) or a self-assembled monolayer (SAM).

[0117] The conductive material can be a single molecule (e.g., a single conductive polymer chain). In some cases, one or more features of an atomic force microscope (AFM) (e.g., the piezoelectric cantilever probe of the AFM) are used to couple a single molecule to a specific (or random) location within the surface of the sensing electrode. Optionally, the conductive material can be multiple molecules (e.g., multiple identical and / or different conductive polymer chains). The conductive material can consist of at least 1, 5, 10, 50, 100, 500, 1,000, 5,000, 10,000, 50,000, 100,000 or more molecules. The conductive material can consist of at most 100,000, 50,000, 10,000, 5,000, 1,000, 500, 100, 50, 10, 5 or 1 molecule.

[0118] The sensors of the present disclosure can be configured to detect a plurality of signals indicative of, for example, an impedance or a change in impedance between a sensing electrode and a reference electrode when the distance between (i) at least a portion of a target molecule (and / or a tag conjugated to the target molecule) and (ii) the sensing electrode is at least about 0.1 nm, 0.5 nm, 1 nm, 2 nm, 3 nm, 4 nm, 5 nm, 6 nm, 7 nm, 8 nm, 9 nm, 10 nm, 20 nm, 30 nm, 40 nm, 50 nm, 60 nm, 70 nm, 80 nm, 90 nm, 100 nm, 200 nm, 300 nm, 400 nm, 500 nm, 600 nm, 700 nm, 800 nm, 900 nm, 1 μm, 2 μm, 3 μm, 4 μm, 5 μm, 6 μm, 7 μm, 8 μm, 9 μm, 10 μm, 20 μm, 30 μm, 40 μm, 50 μm, 60 μm, 70 μm, 80 μm, 90 μm, 100 μm, 200 μm, 300 μm, 400 μm, 500 μm, 600 μm, 700 μm, 800 μm, 900 μm, 1,000 μm or greater. The sensors disclosed herein can be configured to detect a plurality of signals indicative of, for example, an impedance or a change in impedance between a sensing electrode and a reference electrode when the distance between (i) at least a portion of a target molecule (and / or a tag conjugated to the target molecule) and (ii) the sensing electrode is at most about 1,000 μm, 900 μm, 800 μm, 700 μm, 600 μm, 500 μm, 400 μm, 300 μm, 200 μm, 100 μm, 90 μm, 80 μm, 70 μm, 60 μm, 50 μm, 40 μm, 30 μm, 20 μm, 10 μm, 9 μm, 8 μm, 7 μm, 6 μm, 5 μm, 4 μm, 3 μm, 2 μm, 1 μm, 900 nm, 800 nm, 700 nm, 600 nm, 500 nm, 400 nm, 300 nm, 200 nm, 100 nm, 90 nm, 80 nm, 70 nm, 60 nm, 50 nm, 40 nm, 30 nm, 20 nm, 10 nm, 9 nm, 8 nm, 7 nm, 6 nm, 5 nm, 4 nm, 3 nm, 2 nm, 1 nm, 0.5 nm, 0.1 nm or less.

[0119] The sensors of the present disclosure can be configured to detect a plurality of signals indicative of, for example, an impedance or a change in impedance between a sensing electrode and a reference electrode when a target molecule (and / or a tag conjugated to the target molecule) is within a predetermined space (e.g., a predetermined volume 217 as shown in Figure 2 A) that is close to or adjacent to the sensing electrode. The predetermined space can be characterized as having at least about 0.1 nm 2 , 0.5 nm 2 , 1 nm, 2 nm 2 , 3 nm 2 , 4 nm2 , 5nm 2 , 6nm 2 , 7nm 2 , 8nm 2 , 9nm 2 , 10nm 2 , 20nm 2 , 30nm 2 , 40nm 2 , 50nm 2 , 60nm 2 , 70nm 2 , 80nm 2 , 90nm 2 , 100nm 2 , 200nm 2 , 300nm 2 , 400nm 2 , 500nm 2 , 600nm 2 , 700nm 2 , 800nm 2 , 900nm 2 , 1μm 2 , 2μm 2 , 3μm 2 , 4μm 2 , 5μm 2 , 6μm 2 , 7μm 2 , 8μm 2 , 9μm 2 , 10μm 2 , 20μm 2 , 30μm 2 , 40μm 2 , 50μm 2 , 60μm 2 , 70μm 2 , 80μm 2 , 90μm 2 , 100μm 2 , 200μm 2 , 300μm 2 , 400μm 2 , 500μm 2 , 600μm 2 , 700μm 2 , 800μm 2 , 900μm 2 , 1,000μm 2 or larger. The predetermined space may be characterized by having a volume of at most about 1,000 μm 2 , 900μm 2, 800 μm 2 , 700 μm 2 , 600 μm 2 , 500 μm 2 , 400 μm 2 , 300 μm 2 , 200 μm 2 , 100 μm 2 , 90 μm 2 , 80 μm 2 , 70 μm 2 , 60 μm 2 , 50 μm 2 , 40 μm 2 , 30 μm 2 , 20 μm 2 , 10 μm 2 , 9 μm 2 , 8 μm 2 , 7 μm 2 , 6 μm 2 , 5 μm 2 , 4 μm 2 , 3 μm 2 , 2 μm 2 , 1 μm 2 , 900 nm 2 , 800 nm 2 , 700 nm 2 , 600 nm 2 , 500 nm 2 , 400 nm 2 ] , 300 nm 2 , 200 nm 2 , 100 nm 2 , 90 nm 2 , 80 nm 2 , 70 nm 2 , 60 nm 2 , 50 nm 2 , 40 nm 2 , 30 nm 2 , 20 nm 2 , 10 nm 2 , 9 nm 2 , 8 nm 2 , 7 nm 2 , 6 nm 2 , 5 nm 2 , 4 nm 2 , 3 nm 2 , 2 nm 2 , 1 nm 2 , 0.5 nm 2 , 0.1 nm 2 or a smaller volume.

[0120] The sensors of the present disclosure may not require at least a portion of the target molecule (and / or a tag conjugated to the target molecule) to enter and / or pass through a pore (e.g., a nanopore, such as a protein nanopore or a solid-state nanopore) to detect one or more signals indicative of impedance or impedance changes. For example, the sensor may not include a nanopore or may not be operatively coupled to a nanopore. Optionally, at least a portion of the target molecule (and / or a tag conjugated to the target molecule) may enter and / or pass through a pore of the sensor (e.g., a nanopore, such as a protein nanopore or a solid-state nanopore)) so that the sensor can detect one or more signals indicative of impedance or impedance changes. For example, the sensor may include a nanopore.

[0121] The system may include at least 1, 2, 3, 4, 5 or more additional electric field generators. The system may include at most 5, 4, 3, 2 or 1 additional electric field generators. When multiple additional electric field generators are included, the multiple additional electric field generators may apply multiple electric fields along the same or different directions. Optionally, the system may not include any additional electric field generators.

[0122] The sensors of the present disclosure may be electrodes of a circuit (e.g., a CMOS or FET circuit). The circuit may be coupled to a voltage source. A constant voltage may be applied to the circuit, and changes in current may be measured. Optionally, changes in the voltage required to maintain a steady-state current may be measured. The sensor may be in an electrolytic solution (e.g., 0.5M potassium acetate and 10mM KCl). Optionally, the sensor may not be in an electrolytic solution. In some instances, the sensor may be in an aqueous solution or a gas.

[0123] One or more signals can be a current or voltage measured from a sensing circuit. One or more signals can be a current and a voltage measured from a sensing circuit. The signal can be a tunneling current. Alternatively, the signal can be not a tunneling current. The current can be a Faraday current. Alternatively, the current can be not a Faraday current. The current can be at least 1 picoampere (pA), 10 pA, 100 pA, 1 nanoampere (nA), 10 nA, 100 nA, 1 microampere (mA), 10 mA, 100 mA or higher. The current is at most 100 mA, 10 mA, 1 mA, 100 nA, 10 nA, 1 nA, 100 pA, 10 pA, 1 pA or smaller. The current can be at least in the picoampere (pA) range, tens of pA range, hundreds of pA range, nanoampere (nA) range, tens of nA range, hundreds of nA range, microampere (mA) range, tens of mA range or higher. The current can be at most in the tens of mA range, mA range, hundreds of nA range, tens of nA range, nA range, hundreds of pA range, tens of pA range, pA range or lower. The voltage can be at least 0.1 millivolt (mV), 0.5 mV, 1 mV, 5 mV, 10 mV, 50 mV, 100 mV, 500 mV or higher. The voltage can be at most 500 mV, 100 mV, 50 mV, 10 mV, 5 mV, 1 mV, 0.5 mV, 0.1 mV or lower. The voltage can be at least in the millivolt (mV) range, tens of mV range, hundreds of mV range or higher. The voltage can be at most in the hundreds of mV range, tens of mV range, mV range or lower.

[0124] In some embodiments, the sensors of the present disclosure can be provided as an array, such as an array present on a chip or a biochip. The sensor array can have any suitable number of any sensors of the present disclosure. The array can include about 10, about 20, about 50, about 100, about 200, about 400, about 600, about 800, about 1000, about 1500, about 2000, about 3000, about 4000, about 5000, about 10000, about 15000, about 20000, about 40000, about 60000, about 80000, about 100000, about 200000, about 400000, about 600000, about 800000, about 1000000 or more sensors.

[0125] In another aspect, the present disclosure provides a method for analyzing or identifying a target molecule. The method can include detecting, using a sensor, one or more signals indicative of an impedance or an impedance change in the sensor when at least a portion of the target molecule is bound to at least a portion of the sensor. The method can further include using the one or more signals to analyze or identify the target molecule.

[0126] In another aspect, the present disclosure provides a method for analyzing or identifying a target molecule. The method may include providing a sensor that includes a sensing electrode and a reference electrode in electrical communication with each other. The sensor may further include a dielectric material coupled to the sensing electrode and covering a first portion of the surface of the sensing electrode. The sensor may further include a conductive material coupled to the sensing electrode and covering a second portion of the surface of the sensing electrode. The method may further include detecting one or more signals indicative of an impedance or a change in impedance in the sensor when at least a portion of the target molecule binds to the binding portion. The method may further include using the one or more signals to analyze or identify the target molecule.

[0127] In some embodiments, the sensing electrode and the reference electrode may provide a first electric field. Additionally, the method may further include providing an additional electric field generator. The method may further include using the additional electric field generator to apply a second electric field in a second direction that is approximately perpendicular to the first direction of the first electric field.

[0128] Figure 3 An example process 301 for a method of analyzing or identifying a target molecule is shown. The method includes providing a sensor (process 310). The sensor may include a binding unit. The method includes providing a target molecule to the sensor (process 320). The method includes detecting one or more signals indicative of the impedance or a change in impedance of the sensor when at least a portion of the target molecule binds to the binding unit (process 330). The method includes using the one or more signals to analyze or identify the target molecule.

[0129] Figure 4 Another example process 401 for a method of analyzing or identifying a target nucleic acid molecule is shown. The method includes providing a complex comprising a target nucleic acid molecule and a primer nucleic acid molecule, wherein the primer nucleic acid molecule comprises (i) a region complementary to a portion of the target nucleic acid molecule and (ii) an additional region that is not complementary to the target nucleic acid molecule (process 410). The method further comprises analyzing the target nucleic acid molecule by using the sensor to identify the additional region after the additional region has flowed through a pore of the sensor (process 420).

[0130] Optionally, as provided herein, one or more signals may indicate (i) a resistance or a change in resistance in the sensor, (ii) a capacitance or a change in capacitance in the sensor, or (ii) an inductance or a change in inductance in the sensor. One or more signals may indicate at least two of the following: (i) a resistance or a change in resistance in the sensor, (ii) a capacitance or a change in capacitance in the sensor, or (ii) an inductance or a change in inductance in the sensor. One or more signals may indicate (i) a resistance or a change in resistance in the sensor, (ii) a capacitance or a change in capacitance in the sensor, and (ii) an inductance or a change in inductance in the sensor. One or more signals may be a current or a voltage. One or more signals may be a current and a voltage. One or more signals may not be a tunneling current.

[0131] II. Samples

[0132] The sample for analysis may contain one or more polynucleotides (e.g., multiple polynucleotides). The polynucleotide can be single-stranded DNA, double-stranded DNA, or a combination thereof. The polynucleotide may contain genomic DNA, genomic cDNA, cell-free DNA, cell-free cDNA, or any combination of the foregoing.

[0133] The polynucleotide may include cell-free DNA, circulating tumor DNA, genomic DNA, and DNA from formalin-fixed and paraffin-embedded (FFPE) samples. In some instances, the DNA extracted from FFPE samples may be damaged, and such damaged DNA can be repaired by available FFPE DNA repair kits. The sample may contain any suitable DNA and / or cDNA sample, such as urine, feces, blood, saliva, tissue, biopsy, body fluid, or tumor cells.

[0134] The multiple polynucleotides can be single-stranded or double-stranded.

[0135] The polynucleotide sample can be from any suitable source. For example, the sample can be obtained from a patient, animal, plant, or the environment, such as, for example, a naturally occurring or artificial atmosphere, water system, soil, atmospheric pathogen collection system, subsurface sediment, groundwater, or sewage treatment plant.

[0136] The polynucleotide from the sample may include one or more different polynucleotides, such as, for example, DNA, RNA, ribosomal RNA (rRNA), transfer RNA (tRNA), microRNA (miRNA), messenger RNA (mRNA), fragments of any of the foregoing, or any combination of any of the foregoing. The sample may contain DNA. The sample may contain genomic DNA. The sample may contain mitochondrial DNA, chloroplast DNA, plasmid DNA, bacterial artificial chromosome, yeast artificial chromosome, oligonucleotide tag, or any combination of any of the foregoing.

[0137] The polynucleotide can be single-stranded, double-stranded, or a combination thereof. The polynucleotide can be a single-stranded polynucleotide, in which a double-stranded polynucleotide may or may not be present.

[0138] The starting amount of polynucleotide in the sample can be, for example, less than about 50 nanograms (ng), such as less than about 45 ng, less than about 40 ng, less than about 35 ng, less than about 30 ng, less than about 25 ng, less than about 20 ng, less than about 15 ng, less than about 10 ng, less than about 5 ng, less than about 4 ng, less than about 3 ng, less than about 2 ng, less than about 1 ng, less than about 0.5 ng, less than about 0.1 ng or less. The starting amount of polynucleotide in the sample can be, for example, greater than about 0.1 ng, such as greater than about 0.5 ng, greater than about 1 ng, greater than about 2 ng, greater than about 3 ng, greater than about 4 ng, greater than about 5 ng, greater than about 10 ng, greater than about 15 ng, greater than about 20 ng, greater than about 25 ng, greater than about 30 ng, greater than about 35 ng, greater than about 40 ng, greater than about 45 ng, greater than about 50 ng or more. The amount of starting polynucleotide can be, for example, about 0.1 ng to about 100 ng, about 1 ng to about 75 ng, about 5 ng to about 50 ng or about 10 ng to about 20 ng.

[0139] The polynucleotide in the sample can be single-stranded when obtained, or can be made single-stranded by treatment (e.g., denaturation). For example, further examples of suitable polynucleotides are described herein with respect to various aspects of the present disclosure. The polynucleotide can be subjected to subsequent steps (e.g., circularization and amplification) without an extraction step and / or without a purification step. For example, a fluid sample can be processed to remove cells without an extraction step to produce a purified fluid sample and a cell sample, and then the polynucleotide can be isolated from the purified fluid sample. A variety of procedures for isolating polynucleotides can be used, such as by precipitation or non-specific binding to a substrate followed by washing the substrate to release the bound polynucleotide. If the polynucleotide is isolated from the sample without a cell extraction step, the polynucleotide will be primarily extracellular or "cell-free" polynucleotides, which can correspond to dead or damaged cells. The identity of such cells can be used to characterize the cells or cell populations from which they originated, for example, in a microbial community.

[0140] The sample can be from a subject. The subject can be any suitable organism, including, for example, plants, animals, fungi, protists, prokaryotes without nuclei, viruses, mitochondria, and chloroplasts. The sample polynucleotide can be isolated from the subject, such as a cell sample, a tissue sample, a body fluid sample, or an organ sample, or a cell culture derived from any of these, including, for example, a cultured cell line, a biopsy, a blood sample, a buccal swab, or a fluid sample containing cells such as saliva. The subject can be an animal, such as a cow, a pig, a mouse, a rat, a chicken, a cat, a dog, or a mammal, such as a human. The sample can contain tumor cells, such as in a sample of tumor tissue from a subject.

[0141] The sample may not contain intact cells, may be processed to remove cells, or polynucleotides may be isolated without a cell extraction step, for example to isolate cell-free polynucleotides such as cell-free DNA.

[0142] Other examples of sample sources include blood, urine, feces, nasal cavity, lungs, intestine, other body fluids or excreta, derivatives thereof, or combinations thereof.

[0143] A sample from a single individual may be divided into multiple separate samples, such as 2, 3, 4, 5, 6, 7, 8, 9, 10 or more separate samples, which are independently subjected to the methods of the present disclosure, for example, for duplicate, triplicate, quadruplicate or more analyses. When the sample is from a subject, the reference sequence may also be derived from the subject, such as a consensus sequence from the analyzed sample or a polynucleotide sequence from another sample or tissue of the same subject. For example, ctDNA mutations in a blood sample may be analyzed, and cellular DNA from another sample of the subject (such as a buccal or skin sample) may be analyzed to determine the reference sequence.

[0144] According to any suitable method, polynucleotides may or may not be extracted from cells in the sample.

[0145] Multiple polynucleotides may include cell-free polynucleotides such as cell-free DNA (cfDNA) or circulating tumor DNA (ctDNA). Cell-free DNA circulates in healthy and diseased individuals. CfDNA (ctDNA) from tumors is not limited to any specific cancer type but appears to be a common finding in different malignancies. The concentration of cell-free circulating DNA in the plasma of control subjects may be lower compared to patients with or suspected of having a disease. In one example, the concentration of cell-free circulating DNA in plasma may be, for example, 14 ng / mL to 18 ng / mL in control subjects and 18 ng / mL to 318 ng / mL in patients with neoplasia.

[0146] Apoptosis and necrotic cell death may contribute to the production of cell-free circulating DNA in body fluids. For example, significantly elevated levels of circulating DNA can be observed in the plasma of patients with prostate cancer and other prostate diseases such as benign prostatic hyperplasia and prostatitis. In addition, circulating tumor DNA may be present in fluids from the primary organ of the tumor. In one example, breast cancer detection can be achieved in ductal lavage fluid; colorectal cancer detection can be achieved in feces; lung cancer detection can be achieved in sputum, and prostate cancer detection can be achieved in urine or ejaculate. Cell-free DNA can be obtained from a variety of sources. An exemplary source can be a blood sample of a subject. However, cfDNA or other fragmented DNA can come from a variety of other sources, including, for example, urine and fecal samples can be sources of cfDNA including ctDNA.

[0147] III. Nanopores

[0148] A sequencing system may include a reaction chamber that includes one or more nanopore devices. The nanopore devices may be individually addressable nanopore devices. An individually addressable nanopore may be individually readable. An individually addressable nanopore may be individually writable. An individually addressable nanopore may be individually readable and individually writable. The system may include one or more computer processors for facilitating sample preparation and the various operations of the present disclosure, such as polynucleotide sequencing. The processor may be coupled to the nanopore device.

[0149] The nanopore device may include a plurality of individually addressable sensing electrodes. Each sensing electrode may include a membrane adjacent to the electrode and one or more nanopores in the membrane. The nanopore may be in a membrane such as a lipid bilayer that is adjacent to or in sensing proximity to an electrode that is part of or coupled to an integrated circuit. The nanopore may be associated with a single electrode and a sensing integrated circuit or with a plurality of electrodes and a sensing integrated circuit. The nanopore may include a solid-state nanopore.

[0150] The devices and systems for use in the methods provided by the present disclosure may accurately detect individual nucleotide incorporation events, such as after a nucleotide has been incorporated into a growing strand that is complementary to a template. Enzymes, such as DNA polymerase, RNA polymerase, or ligase, may incorporate nucleotides into a growing polynucleotide chain. Enzymes such as polymerase may generate a polynucleotide chain.

[0151] In some cases, an incorporated nucleotide (e.g., incorporated as part of a growing strand that is coupled to a primer nucleic acid molecule) may be detected by the nanopore as the incorporated nucleotide flows through the nanopore. Thus, there may be a time delay between (i) the time of nucleotide incorporation into the growing strand and (ii) sensing the incorporated nucleotide flowing through the nanopore sensor. The sensed data may not be a real-time measurement of the polymerization step or event, but rather a post-polymerization step measurement when the polymerized nucleotide enters and / or flows through the nanopore. The time delay may be at least or at most about 0.01 milliseconds (ms), at least or at most about 0.02 ms, at least or at most about 0.05 ms, at least or at most about 0.1 ms, at least or at most about 0.2 ms, at least or at most about 0.5 ms, at least or at most about 1 ms, at least or at most about 2 ms, at least or at most about 5 ms, at least or at most about 10 ms, at least or at most about 20 ms, at least or at most about 40 ms, at least or at most about 100 ms, at least or at most about 200 ms, at least or at most about 500 ms, at least or at most about 1 second, at least or at most about 2 seconds, at least or at most about 5 seconds, at least or at most about 10 seconds, at least or at most about 20 seconds, at least or at most about 30 seconds, or at least or at most about 60 seconds.

[0152] In some cases, the sensing provided herein can include obtaining (i) a first signal reading from a sensor obtained during incorporation of a nucleobase (e.g., a labeled or unlabeled nucleobase) into a growing strand, and (ii) a second signal reading from the sensor obtained when the incorporated nucleobase of the growing strand enters, is located within, or flows through a pore of the sensor. Utilizing both the first signal and the second signal (e.g., comparing) can enhance the accuracy of the analysis (e.g., identification) of the incorporated nucleobase, and thus enhance the accuracy of the analysis (e.g., identification) of the target nucleic acid molecule.

[0153] In some cases, the nucleotide to be incorporated into the growing strand can be unlabeled. Alternatively, the nucleotide to be incorporated into the growing strand can be labeled.

[0154] The added nucleotide can be complementary to a corresponding template polynucleotide strand that hybridizes to the growing strand. The nucleotide can include a label or label moiety coupled to any position of the nucleotide, including but not limited to the phosphate of the nucleotide such as the γ-phosphate, the sugar, or the nitrogenous base moiety. In some cases, during the nucleotide labeling incorporation, the label is detected when the label associates with the polymerase. The label can be detected until it translocates through the nanopore after nucleotide incorporation and subsequent cleavage and / or release of the label. The nucleotide incorporation event can release the label from the nucleotide, and the label passes through the nanopore and is detected. The label can be released by the polymerase or cleaved / released in any suitable manner, including but not limited to cleavage by an enzyme located near the polymerase. In this way, since a unique label is released from each type of nucleotide (i.e., adenine, cytosine, guanine, thymine, or uracil), the incorporated base (i.e., A, C, G, T, or U) can be identified. In non-release nucleotide incorporation events, the label coupled to the incorporated nucleotide is detected by means of the nanopore. In some instances, the label can move through or near the nanopore and is detected by means of the nanopore.

[0155] The methods and systems of the present disclosure can achieve the detection of polynucleotide incorporation events, for example, at a resolution of at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30, at least about 40, at least about 50, at least about 100, at least about 500, at least about 1000, at least about 5000, at least about 10000, at least about 50000, or at least about 100000 polynucleotide bases over a given period of time. For example, a nanopore device can be used to detect individual polynucleotide incorporation events, each event being associated with an individual nucleic acid base. In other instances, a nanopore device can be used to detect events associated with multiple bases. For example, the signal sensed by the nanopore device can be a combined signal from at least about 2, at least about 3, at least about 4, or at least about 5 bases. Alternatively, the signal sensed by the nanopore device can be associated with (or indicative of) a single base.

[0156] In some sequencing methods, the tag does not pass through the nanopore. The tag can be detected by the nanopore and exit the nanopore without passing through it, such as exiting from the opposite direction of the tag entering the nanopore. The sequencing device can be configured to actively eject the tag from the nanopore.

[0157] In some sequencing methods, the tag is not released after a nucleotide incorporation event. The nucleotide incorporation event can present the tag to the nanopore without releasing the tag. The tag can be detected by the nanopore without being released. The tag can be attached to the nucleotide by a sufficiently long linker to present the tag to the nanopore for detection. For example, when the tagged nucleobase is now part of the growing strand, it can pass through the nanopore, and the nanopore sensor can measure the signal (e.g., impedance) or its change, thereby indicating the presence identity of the tagged nucleobase in the growing strand.

[0158] When a nucleotide incorporation event occurs, the nanopore can detect it in real time. Enzymes (such as DNA polymerase) attached to or near the nanopore can facilitate the passage of the polynucleotide through or near the nanopore. The nucleotide incorporation event, or the incorporation of multiple nucleotides, can release or present one or more tags, which can be detected by the nanopore. Detection can occur when the tag passes by or near the nanopore, when the tag resides in the nanopore, and / or when presenting the tag to the nanopore. In some cases, enzymes attached to or near the nanopore can assist in detecting the tag after the incorporation of one or more nucleotides.

[0159] The tag can be an atom, a molecule, a collection of atoms, or a collection of molecules. The tag can provide an optical, electrochemical, magnetic, or electrostatic (such as inductive or capacitive) signature that can be detected by means of the nanopore.

[0160] A nanopore can be formed or otherwise embedded in a membrane disposed adjacent to a sensing electrode of a sensing circuit, such as an integrated circuit. The integrated circuit can be an application specific integrated circuit (ASIC). The integrated circuit can be a field effect transistor or a complementary metal oxide semiconductor (CMOS). The sensing circuit can be located in a chip or other device having the nanopore, or outside the chip or device, such as in an off-chip configuration.

[0161] When a nucleic acid or tag passes through or adjacent to the nanopore, the sensing circuit detects an electrical signal associated with the nucleic acid or tag. The nucleic acid can be a subunit of a larger chain. The tag can be a byproduct of a nucleotide incorporation event or other interaction between the tagged nucleic acid and a substance adjacent to or at the nanopore, such as an enzyme that cleaves the tag from the nucleic acid. The tag can remain attached to the nucleotide. The detected signal can be collected and stored in a storage location and then used to construct a nucleic acid sequence. The collected signal can be processed to account for any anomalies in the detected signal, such as errors.

[0162] Nanopores can be used to indirectly sequence polynucleotides, in some cases by electrical detection. Indirect sequencing can be any method in which the nucleotides incorporated into the growing chain do not pass through the nanopore. The polynucleotide can pass at any suitable distance from and / or proximal to the nanopore, such that in some cases, the distance allows for detection in the nanopore of a tag released from a nucleotide incorporation event.

[0163] Byproducts of nucleotide incorporation events can be detected by the nanopore. A nucleotide incorporation event refers to the incorporation of a nucleotide into a growing polynucleotide chain. The byproduct can be associated with the incorporation of a given type of nucleotide. Nucleotide incorporation events can be catalyzed by an enzyme, such as DNA polymerase, and use base pair interactions with a template molecule to select available nucleotides for incorporation at each position.

[0164] A nucleic acid sample can be sequenced using tagged nucleotides or nucleotide analogs. In some instances, methods for sequencing a nucleic acid molecule include: (a) incorporating (e.g., polymerizing) tagged nucleotides, wherein the tag associated with an individual nucleotide is released upon incorporation, and (b) detecting the released tag with a nanopore. In some cases, the method further includes guiding the tag attached to or released from an individual nucleotide through the nanopore. The released or attached tag can be guided by any suitable technique, such as by an enzyme (or molecular motor) and / or a voltage difference across the pore in some cases. Alternatively, the released or attached tag can be guided through the nanopore without using an enzyme. For example, as described herein, the tag can be guided by a voltage difference across the nanopore.

[0165] Labels can be detected by means of a nanopore device having at least one nanopore in a membrane. The label can associate with an individual labeled nucleotide during the incorporation of the individual labeled nucleotide. The nanopore device can detect the label associated with the individual labeled nucleotide during the incorporation process. Labeled nucleotides, whether incorporated into a growing nucleic acid strand or not, can be detected, determined, or distinguished by the nanopore device over a given period of time, in some cases by means of the electrodes and / or nanopores of the nanopore device. The time for the nanopore device to detect the label can be shorter than (in some cases significantly shorter than) the time the label and / or the nucleotide coupled to the label are held by an enzyme, such as an enzyme that facilitates the incorporation of nucleotides into a nucleic acid strand (e.g., a polymerase). During the period when the incorporated labeled nucleotide associates with the enzyme, the label can be detected multiple times by the electrode. For example, during the period when the incorporated labeled nucleotide associates with the enzyme, the label can be detected by the electrode at least about 1 time, at least about 2 times, at least about 3 times, at least about 4 times, at least about 5 times, at least about 6 times, at least about 7 times, at least about 8 times, at least about 9 times, at least about 10 times, at least about 20 times, at least about 30 times, at least about 40 times, at least about 50 times, at least about 60 times, at least about 70 times, at least about 80 times, at least about 90 times, at least about 100 times, at least about 200 times, at least about 300 times, at least about 400 times, at least about 500 times, at least about 1000 times, at least about 10,000 times, at least about 100,000 times, or at least about 1,000,000 times.

[0166] Sequencing can be accomplished using pre-loaded labels. The pre-loaded labels can include guiding at least a portion of the label through at least a portion of the nanopore while the label can be attached to a nucleotide that can have been incorporated into a nucleic acid strand (e.g., a growing nucleic acid strand), is being incorporated into the nucleic acid strand, or has not been incorporated into the nucleic acid strand but can be incorporated into the nucleic acid strand. The pre-loaded labels can include guiding at least a portion of the label through at least a portion of the nanopore before the nucleotide is incorporated into the nucleic acid strand or while the nucleotide is being incorporated into the nucleic acid strand. The pre-loaded labels can include guiding at least a portion of the label through at least a portion of the nanopore after the nucleotide has been incorporated into the nucleic acid strand.

[0167] Labels associated with individual nucleotides can be detected by a nanopore without release from the nucleotide after incorporation. The label can be detected without releasing it from the incorporated nucleotide during synthesis of a nucleic acid strand complementary to the target strand. The label can be attached to the nucleotide via a linker such that the label is presented to the nanopore (e.g., the label hangs in at least a portion of the nanopore or otherwise extends through at least a portion of the nanopore). The linker can be long enough to allow the label to extend into or through at least a portion of the nanopore. In some cases, the label is presented to (i.e., moved into) the nanopore by a voltage difference. Other means of presenting the label into the pore can also be suitable (e.g., using an enzyme, a magnet, an electric field, a pressure difference). In some cases, no active force is applied to the label (i.e., the label diffuses into the nanopore).

[0168] A chip for sequencing a nucleic acid sample can include a plurality of individually addressable nanopores. The individually addressable nanopores among the plurality of individually addressable nanopores can include at least one nanopore formed in a membrane disposed adjacent to an integrated circuit. Each individually addressable nanopore is capable of detecting a label associated with an individual nucleotide. Nucleotides can be incorporated (e.g., polymerized), and the label can not be released from the nucleotide after incorporation.

[0169] The label can be presented to and released from the nucleotide into the nanopore after a nucleotide incorporation event. The released label can pass through the nanopore. In some cases, the label does not pass through the nanopore. The label released after a nucleotide incorporation event is distinguished from a label that can pass through the nanopore but is not fully released during its residence time in the nanopore after the nucleotide incorporation event. In some cases, a label that resides in the nanopore for at least 100 milliseconds (ms) is released after the nucleotide incorporation event, and a label that resides in the nanopore for less than 100 ms is not released after the nucleotide incorporation event. The label can be captured and / or guided through the nanopore by a second enzyme or protein (e.g., a nucleic acid binding protein). The second enzyme can cleave the label at the time of nucleotide incorporation (e.g., during or after this). The linker between the label and the nucleotide can be cleaved.

[0170] Based on the residence time of the tag in the nanopore or based on the signal detected from the nucleotide that has not been incorporated by means of the nanopore, it is possible to distinguish between the tag conjugated to the incorporated nucleotide and the tag associated with the nucleotide that has not been incorporated into the growing complementary strand. The signal generated by the unincorporated nucleotide (e.g., voltage difference, current) can be detectable within a time period between 1 nanosecond (ns) and 100 milliseconds or between 1 ns and 50 ms, while the lifetime of the signal generated by the incorporated nucleotide can be between 50 ms and 500 ms, or between 100 ms and 200 ms. The signal generated by the unincorporated nucleotide can be detectable within a time period between 1 ns and 10 ms, or between 1 ns and 1 ms. The time period (average) during which the nanopore can detect the unincorporated tag is longer than the time period during which the nanopore can detect the incorporated tag.

[0171] Compared with the unincorporated nucleotide, the time period during which the incorporated nucleic acid can be detected by the nanopore is shorter. Alternatively, compared with the unincorporated nucleotide, the time period during which the incorporated nucleic acid can be detected by the nanopore is longer. As described herein, the difference and / or ratio between these times can be used to determine whether the nucleotide detected by the nanopore has been incorporated.

[0172] The detection time can be based on the free flow of the nucleotide through the nanopore; the unincorporated nucleotide can stay in or near the nanopore for a time period between 1 nanosecond (ns) and 100 ms or between 1 ns and 50 ms, while the incorporated nucleotide can stay in or near the nanopore for a time period between 50 ms and 500 ms or between 100 ms and 200 ms. The time period can vary according to the processing conditions; however, the residence time of the incorporated nucleotide can be greater than the residence time of the unincorporated nucleotide.

[0173] The tag or tag substance can include a detectable atom or molecule, or a plurality of detectable atoms or molecules. The tag can include one or more adenine, guanine, cytosine, thymine, uracil, or derivatives thereof, which are linked to any position, including the phosphate group, sugar, or nitrogenous base of the nucleic acid molecule. The tag can include one or more adenine, guanine, cytosine, thymine, uracil, or derivatives thereof covalently linked to the phosphate group of the nucleic acid base.

[0174] The tag can have a length of at least about 0.1 nanometers (nm), at least about 1 nm, at least about 2 nm, at least about 3 nm, at least about 4 nm, at least about 5 nm, at least about 6 nm, at least about 7 nm, at least about 8 nm, at least about 9 nm, at least about 10 nm, at least about 20 nm, at least about 30 nm, at least about 40 nm, at least about 50 nm, at least about 60 nm, at least about 70 nm, at least about 80 nm, at least about 90 nm, at least about 100 nm, at least about 200 nm, at least about 300 nm, at least about 400 nm, at least about 500 nm, or at least about 1000 nm.

[0175] The tag can include a tail of repeating subunits, such as multiple adenines, guanines, cytosines, thymines, uracils, or derivatives thereof. For example, the tag can include a tail having at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30, at least about 40, at least about 50, at least about 60, at least about 70, at least about 80, at least about 90, at least about 100, at least about 1000, at least about 10,000, or 100,000 subunits of adenine, guanine, cytosine, thymine, uracil, or derivatives thereof. These subunits can be linked to each other and linked to the phosphate group of the nucleic acid at the end. Other examples of tag portions include any polymeric material, such as polyethylene glycol (PEG), polysulfonate, amino acids, or any polymer that is fully or partially positively charged, negatively charged, or uncharged.

[0176] The tags disclosed herein can be the labels provided herein.

[0177] IV. Polymerases

[0178] DNA polymerase can bind to the 3′ end of the nicked strand of the polynucleotide at the nicking site. DNA sequencing can be accomplished by amplifying and transcribing the polynucleotide near the nanopore and the tagged nucleotide using an enzyme such as DNA polymerase. The sequencing method can involve incorporating or polymerizing tagged nucleotides using a polymerase such as DNA polymerase or transcriptase. The polymerase can be mutated to receive tagged nucleotides. The polymerase can also be mutated to increase the time for the nanopore to detect the tag.

[0179] For example, the sequencing enzyme can be any suitable enzyme that produces a polynucleotide chain through a phosphodiester bond of a nucleotide. For example, DNA polymerase can be 9°Nm TM polymerase or its variant, Escherichia coli DNA polymerase I, bacteriophage T4 DNA polymerase, Sequenase, Taq DNA polymerase, 9°Nm TMPolymerase (exo-) A485L / Y409V, Φ29 DNA polymerase, Bst DNA polymerase, or a variant, mutant, or homolog of any of the foregoing. The homolog may have any suitable percentage homology, e.g., at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, or at least about 95% sequence identity.

[0180] In some instances, for nanopore sequencing, the polymerase can be attached to or positioned near the nanopore. Suitable methods for attaching the polymerase to the nanopore include crosslinking the enzyme to the nanopore or near the nanopore, such as by forming an intramolecular disulfide bond. The nanopore and the enzyme can also be a fusion, e.g., a fusion encoded by a single polypeptide chain. Methods for generating the fusion protein can include fusing the coding sequence of the enzyme in-frame and adjacent to the coding sequence of the nanopore and expressing the fusion sequence from a single promoter. A polymerase can be attached or coupled to the nanopore using a molecular staple or protein finger. The polymerase can be attached to the nanopore through an intermediate molecule, e.g., biotin conjugated to the enzyme and the nanopore, where a streptavidin tetramer links the two biotins. The intermediate molecule can be referred to as a linker.

[0181] The sequencing enzyme can also be attached to the nanopore with an antibody. Proteins that form covalent bonds with each other can be used to attach the polymerase to the nanopore. A phosphatase or an enzyme that cleaves a tag from a nucleotide can also be attached to the nanopore.

[0182] The polymerase can be mutated relative to a non-mutated polymerase to facilitate and / or increase the efficiency of incorporation of a tagged nucleotide into a growing polynucleotide by the mutated polymerase. The polymerase can be mutated to allow a nucleotide analogue (such as a tagged nucleotide) to better enter the active site region of the polymerase and / or mutated to match the nucleotide analogue in the active region.

[0183] Other mutations, such as amino acid substitutions, insertions, deletions, and / or exogenous features of polymerization, can result in enhanced metal ion coordination, reduced exonuclease activity, reduced reaction rates for one or more steps of the polymerase kinetic cycle, reduced branching ratio, altered cofactor selectivity, increased yield, increased thermal stability, increased accuracy, increased speed, increased read length, increased salt tolerance, relative to the non-mutated polymerase.

[0184] Suitable polymerases can have kinetic rate characteristics suitable for detecting a tag through a nanopore. Rate characteristics generally refer to the overall rate of nucleotide incorporation and / or the rate of any step of nucleotide incorporation, such as the rate of nucleotide addition, enzyme isomerization (e.g., becoming a closed state or isomerizing from a closed state), cofactor binding or release, product release, incorporation of a polynucleotide into a growing polynucleotide, or translocation.

[0185] The polymerase can be adapted to permit detection of a sequencing event. The rate characteristics of the polymerase can be such that a tag is loaded into (and / or detected by) the nanopore for an average of 0.1 milliseconds (ms), 1 ms, 5 ms, 10 ms, 20 ms, 30 ms, 40 ms, 50 ms, 60 ms, 80 ms, 100 ms, 120 ms, 140 ms, 160 ms, 180 ms, 200 ms, 220 ms, 240 ms, 260 ms, 280 ms, 300 ms, 400 ms, 500 ms, 600 ms, 800 ms, or 1000 ms. For example, the rate characteristics of the polymerase can be such that a tag is loaded into and / or detected by the nanopore for at least 5 ms, at least 10 ms, at least 20 ms, at least 30 ms, at least 40 ms, at least 50 ms, at least 60 ms, at least 80 ms, at least 100 ms, at least 120 ms, at least 140 ms, at least 160 ms, at least 180 ms, at least 200 ms, at least 220 ms, at least 240 ms, at least 260 ms, at least 280 ms, at least 300 ms, at least 400 ms, at least 500 ms, at least 600 ms, at least 800 ms, or at least 1000 ms. The nanopore can detect the tag between an average of 80 ms and 260 ms, between 100 ms and 200 ms, or between 100 ms and 150 ms.

[0186] The nanopore / polymerase complex can be configured to permit detection of one or more events associated with amplification and transcription of a circular polynucleotide. The one or more events can be kinetically observable and / or non-kinetically observable, such as nucleotide migration through the nanopore without contacting the polymerase.

[0187] In some cases, the polymerase reaction exhibits two kinetic steps beginning with an intermediate where a nucleotide or polyphosphate product binds to the polymerase, and two kinetic steps beginning with an intermediate where the nucleotide and polyphosphate product do not bind to the polymerase. The two kinetic steps can include enzyme isomerization, nucleotide incorporation, and product release. In some cases, the two kinetic steps are template translocation and nucleotide binding.

[0188] A suitable polymerase can exhibit strong or enhanced strand displacement.

[0189] V. Identification of Sequence Variants

[0190] The method provided by the present disclosure can be used to identify sequence variants in a polynucleotide sample. If a sequence difference occurs in at least two different polynucleotides, for example, two different circular polynucleotides, the sequence difference between the sequencing reads and the reference sequence is referred to as a true sequence variant, which can be distinguished by having different junctions. Since the positions and types of sequence variants caused by amplification or sequencing errors are unlikely to be precisely replicated on two different polynucleotides containing the same target sequence, including this validation parameter can reduce the background of error sequence variants while improving the sensitivity and accuracy of detecting actual sequence variations in the sample. The frequency of sequence variants can be less than 5%, less than 4%, less than 3%, less than 2%, less than 1.5%, less than 1%, less than 0.75%, less than 0.5%, less than 0.25%, less than 0.1%, less than 0.075%, less than 0.05%, less than 0.04%, less than 0.03%, less than 0.02%, less than 0.01%, less than 0.005%, less than 0.001%, or a lower frequency that is sufficiently higher than the background value to allow accurate identification. The occurrence frequency of sequence variants can be less than 0.1%. When the frequency of sequence variants is statistically significantly higher than the background error rate, for example, the p-value is less than 0.05, 0.01, 0.001, or 0.0001, the frequency of the sequence variant can be sufficiently higher than the background. When the frequency is at least 2-fold, 3-fold, 4-fold, 5-fold, 6-fold, 7-fold, 8-fold, 9-fold, 10-fold, 25-fold, 50-fold, 100-fold, or more higher than the background error rate, the frequency of sequence variants can be sufficiently higher than the background. The background error rate for accurately determining the sequence at a given position can be less than 1%, 0.5%, 0.1%, 0.05%, 0.01%, 0.005%, 0.001%, or 0.0005%.

[0191] Identifying sequence variants can include optimally aligning one or more sequencing reads with a reference sequence to identify the differences between the two, and identifying junctions. Alignment can include placing one sequence along another sequence, iteratively introducing gaps along each sequence, scoring the degree of match between the two sequences, and repeating at different positions along the reference sequence. The highest-scoring match is considered the alignment and represents an inference of the degree of relatedness between the sequences.

[0192] The reference sequence for comparison with the sequencing reads is a reference genome, such as the genome of a member of the same species as the subject. The reference genome can be complete or incomplete. The reference genome can only consist of regions containing the target polynucleotide, such as regions from the reference genome or consensus sequences generated from the sequencing reads in the analysis. The reference sequence can contain or be composed of polynucleotide sequences of one or more organisms, such as sequences from one or more bacteria, archaea, viruses, protists, fungi, or other organisms. The reference sequence can consist of only a portion of the reference genome, such as a region corresponding to one or more target sequences in the analysis. For example, for detecting a pathogen, the reference genome can be the entire genome of the pathogen, or a portion thereof for identification, such as a particular strain or serotype. The sequencing reads can be aligned to multiple different reference sequences for screening multiple different organisms or strains.

[0193] VI. Therapeutic Applications

[0194] The methods, systems, and compositions provided herein can be directed to one or more therapeutic applications, such as for characterizing a patient sample and optionally diagnosing a condition of a subject. Therapeutic applications can include, based on the results of the methods provided in the present disclosure, informing a patient of treatment options to which the patient may respond most effectively and / or informing the treatment of a subject in need of therapeutic intervention.

[0195] For example, the methods provided in the present disclosure can be used to diagnose the presence, progression, and / or metastasis of a tumor, such as when the polynucleotide being analyzed contains, consists of, or is composed of cfDNA, ctDNA, or fragmented tumor DNA. For example, the efficacy of tumor treatment of a subject can be monitored by monitoring ctDNA over time, a decrease in ctDNA can be used as an indication of treatment efficacy, and an increase in ctDNA can inform the selection of a different treatment and / or a different dose. Other uses include assessing organ rejection in a transplant recipient, such as using an increase in the amount of circulating DNA corresponding to the transplant donor genome as an early indicator of transplant rejection, and genotyping / haplotyping of pathogen infections (such as viral or bacterial infections). Detection of sequence variants in circulating fetal DNA can be used to diagnose a condition of the fetus.

[0196] The methods provided in the present disclosure can include diagnosing a subject based on sequencing results, such as diagnosing a subject with a disease associated with a detected causal genetic variant, or reporting the likelihood that a patient has or will develop such a disease.

[0197] Causal genetic variants can include sequence variants associated with a specific type or stage of cancer or cancer having specific characteristics such as metastatic potential, drug resistance, and / or drug responsiveness. The methods provided by this disclosure can be used to inform treatment decisions, guide, and monitor cancer treatment. For example, the treatment efficacy can be monitored by comparing ctDNA samples of a patient before, during, and after treatment, which includes specific molecular targeted therapies such as monoclonal drugs, chemotherapy drugs, radiation regimens, and combinations of any of the foregoing methods. For example, ctDNA can be monitored to see if certain mutations increase or decrease after treatment, or if new mutations appear, which can enable a doctor to modify the treatment regimen within a shorter time period than monitoring methods that track a patient's symptoms. The methods can include diagnosing a subject based on the results of polynucleotide sequencing, such as diagnosing a subject with a specific stage or type of cancer associated with the detected sequence variant, or reporting the likelihood that a patient has or will develop such cancer.

[0198] For example, for a therapy specifically targeted at a patient based on molecular markers, the patient can be tested to find out if certain mutations are present in their tumor, and these mutations can be used to predict the response to or drug resistance of the therapy, and to guide the decision of whether to use the therapy. Detecting and monitoring ctDNA during treatment helps to guide treatment selection.

[0199] Sequence variants associated with one or more cancers can be used for diagnosis, prognosis, or treatment decisions. For example, suitable target sequences of oncological significance include alterations in the TP53 gene, ALK gene, KRAS gene, PIK3CA gene, BRAF gene, EGFR gene, and KIT gene. The target sequences can be specifically amplified, and / or whether the sequence variants of the target sequences can be all or part of cancer-related genes can be specifically analyzed.

[0200] The methods provided by this disclosure can be used to discover new rare mutations associated with one or more cancer types, stages, or cancer characteristics. For example, in a population of individuals sharing an analyzed characteristic such as a specific disease, cancer type, and / or cancer stage, the methods provided by this disclosure can be used to identify sequence variants reflecting mutations of a specific gene or gene part. The identified sequence variants that occur at a statistically significantly higher frequency in the group of individuals sharing the characteristic compared to individuals without the characteristic can have an association with the characteristic. Then, the identified sequence variants or types of sequence variants can be used to diagnose or treat individuals found to carry them.

[0201] Additional therapeutic applications can include use in non-invasive fetal diagnosis. Fetal DNA can be found in the blood of pregnant women. The methods provided by the present disclosure can be used to identify sequence variants in circulating fetal DNA and can thus be used to diagnose one or more genetic diseases in the fetus, such as diseases associated with one or more causal genetic variants. Examples of causal genetic variants include trisomies, cystic fibrosis, sickle cell anemia, and Tay-Sachs disease. The mother can provide a control sample and a blood sample for comparison. The control sample can be any suitable tissue and can then be sequenced to provide a reference sequence. The cfDNA sequences corresponding to the fetal genomic DNA can then be identified as sequence variants relative to the maternal reference. The father can also provide a reference sample to assist in identifying fetal sequences and sequence variants.

[0202] Different therapeutic applications can include detecting exogenous polynucleotides, including detection from pathogens such as bacteria, viruses, fungi, and microorganisms, and this information can suggest treatment.

[0203] VII. Computer Systems

[0204] The present disclosure provides computer systems programmed to implement one or more methods of the present disclosure. The computer systems of the present disclosure can be used to regulate various operations of sensors, such as detecting one or more signals indicative of impedance or impedance changes in the sensor when at least a portion of a target nucleotide binds to a binding portion of the sensor.

[0205] Figure 2 Computer system 201 is shown, which is programmed or otherwise configured to communicate with and regulate various aspects of the sequencing of the present disclosure. For example, computer system 201 can communicate with one or more circuits coupled to or including sensors, and one or more devices (e.g., machines) for preparing, processing, or holding one or more reaction mixtures for sensing. Computer system 201 can also communicate with one or more controllers or processors of the present disclosure. Computer system 201 can be an electronic device of the user or a computer system located remotely relative to the electronic device. The electronic device can be a mobile electronic device.

[0206] The computer system 201 includes a central processing unit (CPU, also referred to herein as “processor” and “computer processor”) 205, which can be a single-core or multi-core processor, or multiple processors for parallel processing. The computer system 201 also includes a memory or memory location 210 (e.g., random access memory, read-only memory, flash memory), an electronic storage unit 215 (e.g., hard disk), a communication interface 220 (e.g., network adapter) for communicating with one or more other systems, and peripheral devices 225 such as caches, other memories, data storage, and / or electronic display adapters. The memory 210, storage unit 215, interface 220, and peripheral devices 225 communicate with the CPU 205 via a communication bus (solid lines), such as a motherboard. The storage unit 215 can be a data storage unit (or data repository) for storing data. The computer system 201 can be operably coupled to a computer network (“network”) 230 by means of the communication interface 220. The network 230 can be the Internet, an intranet, and / or an extranet, or an intranet and / or an extranet that communicates with the Internet. In some cases, the network 230 is a telecommunications and / or data network. The network 230 can include one or more computer servers, which can implement distributed computing, such as cloud computing. In some cases, the network 230 can implement a peer-to-peer network with the help of the computer system 201, which can enable devices coupled to the computer system 201 to act as clients or servers.

[0207] The CPU 205 can execute a series of machine-readable instructions, which can be embodied in a program or software. The instructions can be stored in a memory location, such as the memory 210. The instructions can be directed to the CPU 205, and subsequently, the CPU 205 can be programmed or otherwise configured to implement the methods of the present disclosure. Examples of operations performed by the CPU 205 can include fetching, decoding, executing, and writing back.

[0208] The CPU 205 can be part of a circuit, such as an integrated circuit. One or more other components of the system 201 can be included in the circuit. In some cases, the circuit is an application-specific integrated circuit (ASIC).

[0209] The storage unit 215 can store files, such as drivers, libraries, and saved programs. The storage unit 215 can store user data, e.g., user preferences and user programs. In some cases, the computer system 201 can include one or more additional data storage units external to the computer system 201, such as on a remote server that communicates with the computer system 201 via an intranet or the Internet.

[0210] The computer system 201 may communicate with one or more remote computer systems via the network 230. For example, the computer system 201 may communicate with a user's remote computer system. Examples of remote computer systems include personal computers (e.g., portable PCs), tablets or tablet computers (e.g., iPad, Galaxy Tab), telephones, smart phones (e.g., iPhone, Android - enabled devices, ) or personal digital assistants. A user may access the computer system 201 via the network 230.

[0211] The methods described herein may be implemented in the form of machine (e.g., computer processor) - executable code stored on an electronic storage location (e.g., the memory 210 or the electronic storage unit 215) of the computer system 201. The machine - executable or machine - readable code may be provided in software form. During use, the code may be executed by the processor 205. In some cases, the code may be retrieved from the storage unit 215 and stored in the memory 210 for ready access by the processor 205. In some cases, the electronic storage unit 215 may not be included and the machine - executable instructions may be stored in the memory 210.

[0212] The code may be pre - compiled and configured to be used with a machine having a processor suitable for executing the code, or may be compiled at runtime. The code may be provided in a programming language, and the programming language may be chosen such that the code can be executed in a pre - compiled or just - in - time compiled manner.

[0213] Aspects of the systems and methods provided herein, such as computer system 201, may be embodied as programming. Various aspects of technology may be considered "products" or "articles of manufacture", typically in the form of machine (or processor) executable code and / or associated data carried or embodied on a machine-readable medium. Machine-executable code may be stored on an electronic storage unit, such as a memory (e.g., read-only memory, random access memory, flash memory) or a hard disk. "Storage" type media may include any or all tangible memories of a computer, processor, etc., or associated modules thereof, such as various semiconductor memories, tape drives, disk drives, etc., which can provide non-transitory storage for software programming at any time. All or part of the software can sometimes be communicated via the Internet or other various telecommunication networks. For example, such communication can enable the software to be loaded from one computer or processor to another, e.g., from an administrative server or host to a computer platform of an application server. Thus, another type of media that can carry software elements includes light waves, radio waves, and electromagnetic waves, such as those used between physical interfaces of local devices, over wired and optical landline networks, and over various air links. Physical elements that carry such waves, such as wired or wireless links, optical links, etc., can also be considered media that carry software. As used herein, unless restricted to non-transitory, tangible "storage" media, the term such as computer or machine "readable medium" refers to any medium that participates in providing instructions to a processor for execution.

[0214] Thus, machine-readable media, such as computer-executable code, can take many forms, including but not limited to tangible storage media, carrier media, or physical transmission media. Non-volatile storage media includes, for example, optical or magnetic disks, such as any storage device in any computer, etc., which can be used, for example, to implement databases shown in the drawings. Volatile storage media includes dynamic memory, such as the main memory of such computer platforms. Tangible transmission media includes coaxial cables; copper wire and fiber optics, including the wires that make up the internal bus of a computer system. Carrier transmission media can take the form of electrical or electromagnetic signals, or can take the form of acoustic or light waves, such as those generated during radio frequency (RF) and infrared (IR) data communications. Thus, common forms of computer-readable media include, for example: floppy disks, flexible disks, hard disks, magnetic tapes, any other magnetic media, CD-ROMs, DVDs or DVD-ROMs, any other optical media, punched cards, paper tapes, any other physical storage media with punched patterns, RAMs, ROMs, PROMs, and EPROMs, FLASH-EPROMs, any other memory chips or cartridges, carriers that transmit data or instructions, cables or links that transmit such carriers, or any other media from which a computer can read programming code and / or data. Many of these forms of computer-readable media can involve transmitting one or more sequences of one or more instructions to a processor for execution.

[0215] The computer system 201 may include or communicate with an electronic display 235 that includes a user interface (UI) 240 to provide, for example, (i) the progress of a reaction mixture, (ii) the progress of sequencing, and (iii) sequencing information obtained from sequencing. Examples of UIs include, but are not limited to, graphical user interfaces (GUIs) and web-based user interfaces.

[0216] The methods and systems of the present disclosure may be implemented by one or more algorithms. The algorithms may be implemented by software when executed by a central processing unit 205. For example, the algorithms may determine sequence reads of target nucleotides, polynucleotides, peptides, polypeptides, proteins, and the like.

[0217] Examples

[0218] Example 1: Sequencing by Synthesis (SBS)

[0219] In some embodiments, the methods provided herein may be used to determine polymer sequence information (e.g., polynucleotide sequence information), including using a system (e.g., a device or instrument). In some cases, the polymer may be a target nucleic acid (NA) molecule (e.g., a template). In some cases, the determination may include directly reading nanopore SBS.

[0220] (1A) In some embodiments, the method may utilize a system (e.g., a device, an instrument, etc.) that includes an array of pores or wells (e.g., nanopores), the array of pores or wells being configured to receive a fluid for sensing (e.g., sequencing). The fluid may be configured to receive target NA molecules, primer NA molecules provided herein (e.g., comprising a complementary region and a non-complementary region), and / or a complex comprising a target NA molecule and a primer NA molecule.

[0221] (1B) In some embodiments, the system may include a membrane. The membrane may be used to cover at least a portion of the pore / well array, e.g., to provide engagement and / or support for the interaction of the complex of (1A) with one or more pores / wells of the array. The membrane may be configured to separate the pores / wells of (1A) into an upper fluid chamber (e.g., a cis chamber) and a lower fluid chamber (e.g., a trans chamber).

[0222] (1C) In some embodiments, the nanopore array engages (e.g., is at least partially embedded in) the membrane of (1B).

[0223] (1D) In some embodiments, a series of additional complex conjugations (e.g., at least partially embedded) to the membrane of (1B). Each additional complex array contains nanopores, enzymes (e.g., polymerases), and the complex of (1A).

[0224] (1E) In some embodiments, the nanopores of (1C) or (1D) can be synthetic pores or biological pores.

[0225] (1F) In some embodiments, each complex of (1D) can be a template-prime complex.

[0226] (1G) In some embodiments, the complex can bind to an enzyme to form an additional complex, which is an enzyme-template-prime complex.

[0227] (1H) In some embodiments, the complex can bind to an enzyme conjugated to a nanopore to form an additional complex, which is an enzyme-template-prime complex.

[0228] (1I) In some embodiments, in each additional complex of (1D), the enzyme can be conjugated to the nanopore at a specific site of the nanopore (e.g., via hydrogen bonding, affinity binding, etc.).

[0229] (1J) In some embodiments, the conjugation between the enzyme and the nanopore in (1I) can be permanent or can be disrupted, e.g., by an external force or treatment (such as high voltage at high pH).

[0230] (1K) In some embodiments, the formation of each additional complex of (1D) can be achieved by forming a nanopore-enzyme subunit and then adding a template-prime subunit to the nanopore-enzyme subunit to form a nanopore-enzyme-template-prime complex.

[0231] (1L) In some embodiments, the formation of each additional complex of (1D) can be achieved by forming an enzyme-template-prime subunit and then adding the enzyme-template-prime subunit to the nanopore to form a nanopore-enzyme-template-prime complex.

[0232] (1M)In some embodiments, the formation of each additional complex of (1D) can be achieved by breaking the connection (e.g., bond) between the nanopore and the enzyme, and then forming a new nanopore-enzyme-template-prime complex by adding a new enzyme-template-prime subunit to the nanopore, thereby forming a new nanopore-enzyme-template-prime complex. In some cases, the nanopore can be engineered to contain a binding moiety that can bind to the enzyme, and this binding through the binding moiety can be reversible. In some cases, the enzyme can be engineered to contain a binding moiety that can bind to the nanopore, and this binding through the binding moiety can be reversible. In some cases,

[0233] (1N)In some embodiments, the system comprises a sensor electrode and circuitry configured to measure one or more electrical signals from any of the components, compounds, or processes described in (1A)-(1M).

[0234] Example 2: Synthetic Sequencing (SBS)

[0235] In some embodiments, the methods provided herein may include using the system described in (1A)-(1M) of Example 1.

[0236] (2A)In some embodiments, the method may include applying an electric field across a membrane, e.g., when at least a portion of the nanopore is engaged (e.g., embedded) in the membrane. In some cases, the nanopore may not be conjugated to any enzyme. In some cases, the nanopore may be conjugated to an enzyme but not to a template-prime complex. In some cases, the nanopore may be conjugated to an enzyme-template-prime complex.

[0237] (2B)In some embodiments, the method may include engaging at least the nanopore to the membrane to effect engagement of the nanopore-enzyme-template-prime complex to the membrane, applying the electric field as provided in (2A), and subsequently measuring one or more electrical signals from a sensor (e.g., a nanopore sensor) as the polymerization product of the enzyme (e.g., a growing strand conjugated to the primer) is directed into or flows / translocates through the nanopore.

[0238] (2C)In some embodiments, the polymerization product of the enzyme in (2B) can be generated by initiating the enzymatic activity (e.g., polymerase activity) of the enzyme on a template (e.g., a target NA molecule) in the presence of a primer.

[0239] (2D)In some embodiments, the initiation of the enzymatic activity of the enzyme in (2C) is carried out in the presence of an electric field as provided in (2A) or (2B) (e.g., the same electric field) or a different electric field.

[0240] (2E) In some embodiments, the method includes forming an enzyme-template-primer complex having a primer overhang (e.g., comprising a polynucleotide sequence that is not complementary to the template provided herein), and loading the enzyme-template-primer complex onto a nanopore by attracting the primer overhang to the nanopore using an electric field as provided in (2A) or (2B) (e.g., the same electric field) or a different electric field. In some cases, the method further includes initiating enzyme activity to generate a growing strand that will be attracted to and translocate through the nanopore, e.g., by an electric field as provided in (2A) or (2B) or a different electric field.

[0241] (2F) In some embodiments, the method includes forming an enzyme-template-primer without a primer overhang and initiating enzyme activity to expose the end of the growing strand such that it is attracted to and translocates through the nanopore.

[0242] (2G) In some embodiments, the method includes forming an enzyme-template-primer complex having a primer overhang and loading the enzyme-template complex onto the nanopore (or array or nanopore) when the primer overhang is attracted towards the nanopore by an electric field as provided in (2A) or (2B) or a different electric field.

[0243] (2H) In some embodiments, the method includes preventing the template (e.g., target NA molecule) from being attracted to the nanopore by adding (e.g., covalently or non-covalently coupling) a heterologous nucleic acid molecule that is not a primer NA molecule (e.g., a hairpin) to one or more ends (e.g., 5' end or 3' end) of the template. The length of the heterologous NA molecule can be at least or at most about 2, at least or at most about 5, at least or at most about 10, at least or at most about 15, at least or at most about 20, at least or at most about 30, at least or at most about 40, at least or at most about 50, at least or at most about 60, at least or at most about 70, at least or at most about 80, at least or at most about 90, or at least or at most about 100 nucleobases.

[0244] (2I) In some embodiments, the method can include measuring an electrical signal generated when the polymerization product is attracted to and / or translocates through the nanopore.

[0245] (2J) In some embodiments, the method can include measuring an electrical signal as described in Example 1 or Example 2, e.g., by measuring current, voltage, impedance, or a change thereof in a sensor. The measurement can be made by applying direct current or alternating current.

[0246] (2K) In some embodiments, the method can include measuring the electrical signal as described in Example 1 or Example 2 multiple times by driving the polymer back and forth through the nanopore multiple times by alternately varying the current on the sensor multiple times.

[0247] (2L) In some embodiments, the method can include using a motor enzyme (e.g., a polymerase) as a motor to control the translocation rate of a polymerization product through a nanopore.

[0248] (2M) In some embodiments, the method can include using an electrical bias to control the translocation rate of a polymer through a nanopore (e.g., in process (2L)).

[0249] (2N) In some embodiments, the method can include using a nanopore having a sensor region with a sensing resolution for a single nucleobase of a polymerization product or a combination of two or more nucleobases.

[0250] (2O) In some embodiments, in (2N), the method can include using a time-resolved electrical signal to distinguish and resolve sensing data indicative of a combination of two or more nucleobases of a polymerization product, by using direct current or alternating current, thereby achieving single nucleobase resolution.

[0251] (2P) In some embodiments, the method can include in (2N) or (2O), using a deconvolution method to distinguish and resolve a combination of two or more nucleobases of a polymerization product, by using direct current or alternating current, thereby achieving single nucleobase resolution. As used herein, the term "deconvolution" generally refers to a process (e.g., a computerized process) that given an output signal (e.g., sensing data from a combination of two or more nucleobases) determines two or more unknown input signals (e.g., predicted individual sensing data from each of the two or more nucleobases in the combination).

[0252] (2Q) In some embodiments of any of the methods in Example 1 or Example 2, the method can include reducing the variation of the polymerization product or preventing the polymerization product from fully retracting outside the nanopore (e.g., along a direction opposite to the direction in which the polymerization product initially enters the nanopore), e.g., by conjugating a termination moiety (e.g., a small molecule or a heterologous nucleic acid molecule, etc.) to at least a portion of the polymerization product that reaches the trans chamber (e.g., opposite to the cis chamber containing a polymerase conjugated to the nanopore). In some cases, the termination moiety can be streptavidin, which is conjugated to the biotinylated end of the polymerization product or the primer NA molecule). For example, a non-complementary region of the primer NA molecule can contain a conjugation unit for the termination moiety (e.g., biotin) before forming a complex with the template, such that during the complex sequencing process, when at least a portion of the non-complementary region translocates to the trans chamber, the termination moiety (e.g., streptavidin) can conjugate with the conjugation unit. In some cases, the termination moiety can be conjugated to the 5' end of the polymerization product or the primer NA molecule. In some cases, the termination moiety can be conjugated to the 3' end of the polymerization product or the primer NA molecule.

[0253] Although the preferred embodiments of the present invention have been shown and described herein, it will be apparent to those skilled in the art that these embodiments are provided by way of example only. The present invention is not intended to be limited by the specific examples provided in the specification. Although the present invention has been described with reference to the foregoing specification, the description and illustration of the embodiments herein are not meant to be construed in a limiting sense. Various changes, alterations, and substitutions will now occur to those skilled in the art without departing from the present invention. In addition, it should be understood that all aspects of the present invention are not limited to the specific descriptions, configurations, or relative proportions set forth herein that depend on various conditions and variables. It should be understood that various alternatives to the embodiments of the present invention described herein may be used to practice the present invention. Therefore, it is contemplated that the present invention will also cover any such alternatives, modifications, variations, or equivalents. The following claims are intended to define the scope of the present invention and thereby cover the methods and structures within the scope of these claims and their equivalents.

Claims

1. A method for analyzing a target nucleic acid molecule, the method comprising: (a) providing a complex comprising the target nucleic acid molecule and a primer nucleic acid molecule, wherein the primer nucleic acid molecule comprises (i) a region complementary to a portion of the target nucleic acid molecule and (ii) an additional region non-complementary to the target nucleic acid molecule; and (b) identifying the additional region using a sensor after the additional region has flowed through a pore of the sensor, thereby analyzing the target nucleic acid molecule.

2. The method according to claim 1, further comprising contacting the complex with an enzyme operably coupled to the pore.

3. The method according to claim 2, wherein the contacting enables the flow of the additional region through the pore.

4. The method according to claim 1, further comprising, in (b), alternately varying an electrical signal in the sensor to direct the additional region to flow through the pore two or more times.

5. The method according to claim 1, further comprising, in (b): (i) performing an extension reaction of the target nucleic acid molecule from the primer nucleic acid molecule to generate a growing strand coupled to the primer nucleic acid molecule, wherein the growing strand exhibits sequence complementarity to an additional portion of the target nucleic acid molecule; and (ii) obtaining sequence information of at least a portion of the growing strand by directing at least a portion of the growing strand through the pore of the sensor.

6. The method according to claim 5, wherein, In (a), the region of the primer nucleic acid molecule flanks (1) the additional region of the primer nucleic acid molecule and (2) the growing strand.

7. The method according to claim 5, wherein In (b), the obtaining of the sequence information comprises identifying a sequence read comprising characteristic sequence information indicative of at least a portion of the additional region of the primer nucleic acid molecule.

8. The method according to claim 5, wherein In (b), directing the additional region of the primer nucleic acid molecule through the pore before at least a portion of the growing strand.

9. The method according to claim 5, wherein In (b), the directing comprises alternately varying an electrical signal in the sensor to direct at least a portion of the growing strand through the pore two or more times.

10. The method according to claim 1, wherein the additional region comprises at least 5 bases.

11. The method according to claim 1, wherein the additional region comprises at least 10 bases.

12. The method according to claim 1, wherein the additional region comprises at least 20 bases.

13. The method according to claim 1, wherein the additional region comprises a net positive charge.

14. The method according to claim 1, wherein the additional region comprises a net negative charge.

15. The method according to claim 1, wherein the pore is part of a nanopore protein.

16. The method according to claim 1, wherein the pore is part of a solid-state nanopore.

17. The method according to claim 1, wherein the target nucleic acid molecule is a circular nucleic acid molecule.

18. The method according to claim 1, wherein in (c), the obtaining comprises detecting one or more signals indicative of an impedance or a change in impedance in the sensor when guiding at least the additional portion through the pore.

19. The method according to claim 1, wherein the pore is embedded in a membrane.

20. The method according to claim 1, wherein the primer nucleic acid molecule comprises a barcode.

21. A system for analyzing a target nucleic acid molecule, comprising: A primer nucleic acid molecule comprising: (i) a region complementary to a portion of the target nucleic acid molecule and (ii) an additional region non-complementary to the target nucleic acid molecule, wherein the primer nucleic acid molecule is configured to form a complex with the target nucleic acid molecule; A sensor comprising a pore, wherein the sensor is configured to guide the additional region to flow through the pore; and A controller operably coupled to the sensor, wherein the controller is configured to identify the additional region flowing through the pore, thereby analyzing the target nucleic acid molecule.

22. The system according to claim 21, further comprising an enzyme operably coupled to the pore, wherein the enzyme is configured to contact the complex.

23. The system according to claim 22, wherein the contact between the enzyme and the complex is configured to enable the additional region to flow through the pore.

24. The system according to claim 21, wherein the controller is further configured to alternately vary an electrical signal in the sensor to guide the additional region to flow through the pore two or more times.

25. The system according to claim 21, wherein the sensor is further configured to perform an extension reaction of the target nucleic acid molecule from the primer nucleic acid molecule to generate a growing strand coupled to the primer nucleic acid molecule, wherein the growing strand exhibits sequence complementarity to an additional portion of the target nucleic acid molecule, and wherein the controller is further configured to obtain sequence information of at least the portion of the growing strand by guiding at least a portion of the growing strand through the pore of the sensor.

26. The system according to claim 25, wherein the region of the primer nucleic acid molecule flanks (1) the additional region of the primer nucleic acid molecule and (2) the growing strand.

27. The system according to claim 25, wherein the controller is configured to identify a sequence read comprising characteristic sequence information indicative of at least a portion of the additional region of the primer nucleic acid molecule.

28. The system according to claim 25, wherein the additional region of the primer nucleic acid molecule is guided through the pore before at least the portion of the growing strand.

29. The system according to claim 25, wherein the controller is configured to alternately vary an electrical signal in the sensor to guide at least the portion of the growing strand through the pore two or more times.

30. The system according to claim 21, wherein the additional region comprises at least 5 bases.

31. The system according to claim 21, wherein the additional region comprises at least 10 bases.

32. The system according to claim 21, wherein the additional region comprises at least 20 bases.

33. The system according to claim 21, wherein the additional region comprises a net positive charge.

34. The system according to claim 21, wherein the additional region comprises a net negative charge.

35. The system according to claim 21, wherein the pore is part of a nanopore protein.

36. The system according to claim 21, wherein the pore is part of a solid-state nanopore.

37. The system according to claim 21, wherein the target nucleic acid molecule is a circular nucleic acid molecule.

38. The system according to claim 21, wherein the controller is configured to detect one or more signals indicative of impedance or a change in impedance in the sensor when guiding at least the additional portion through the pore.

39. The system according to claim 21, wherein the pore is embedded in a membrane.

40. The system according to claim 21, wherein the primer nucleic acid molecule comprises a barcode.