Methods and systems for nanopore-based multiomics analysis

Solid-state nanopores with structural DNA labels and protein-based strategies enhance nanopore technology for high-speed multiomics analysis, addressing the limitations of current methods in characterizing complex protein mixtures and small molecules, thereby advancing precision medicine diagnostics.

WO2025151716A1PCT designated stage expired Publication Date: 2025-07-17CATALOG TECHNOLOGIES INC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/011070
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-11
Filing Date
2025-01-10
Publication Date
2025-07-17

AI Technical Summary

Technical Problem

Current nanopore technologies are limited in their ability to efficiently characterize complex protein mixtures and small molecules, requiring time-consuming and computationally intensive methods for DNA sequencing and lacking high-throughput capabilities for multiomics analysis.

Method used

The use of solid-state nanopores with structural DNA labels and protein-based labeling strategies, such as DNA dumbbells and enzymatic modifications, allows for high-speed characterization of DNA, proteins, and small molecules by detecting unique current-time signals through nanopores.

Benefits of technology

Enables high-throughput multiomics analysis, improving the speed and accuracy of biomolecule detection and characterization, facilitating precision medicine diagnostics and therapeutics.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025011070_17072025_PF_FP_ABST
    Figure US2025011070_17072025_PF_FP_ABST
Patent Text Reader

Abstract

A methods for extracting information from a particle includes modifying at least a portion of a set of particles with one or more labels, and translocating one or more modified particles through one or more nanopores disposed in a substrate. The method includes reading signals received from the nanopores and translating the signals into the information. The reading step includes detecting a change in electric current through the nanopores corresponding to the translocation. A first current level corresponds to passage of at least an unlabeled portion of a particle and a second current level corresponds to passage of a label through the one or more nanopores. The reading step includes identifying the label from the second voltage level. The reading step includes identifying the information based at least in part on the identified label.
Need to check novelty before this filing date? Find Prior Art

Description

METHODS AND SYSTEMS FOR NANOPORE-BASED MULTIOMICS ANALYSISCROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This Application claims priority to and the benefit of U.S. Provisional Patent Application No. 63 / 619,798, filed on January 11, 2024, titled “Methods and Systems of Nanopore-based multiomics analysis,” the entire contents of which are hereby incorporated by reference.BACKGROUND

[0002] Current state-of-the-art methodologies for the analysis of biological samples focus on single biomolecule characterization, for example, qPCR, next-generation sequencing, and nanopore sequencing for nucleic acids and mass spectrometry for peptides, proteins, and other related small molecules.

[0003] Nanopores, on a primitive level, may act as molecular counters, whereby a molecule of interest translocates across a channel resulting in a transient change in conductance from a baseline, these disruption events demarcate a translocation event (FIG. 1A), thus potentially providing a surrogate for quantifying the number of molecules present in a sample [4], The magnitude and duration of disruption are properties that can be utilized to characterize and distinguish features of a molecule (FIG. IB). Traditionally, nanopore technology has been viewed as a nucleic acid sequencing technology, where, as a molecule (e.g., a nucleic acid molecule) translocates, each nucleotide provides a unique disruption signature, facilitating base-calling and the determination of a previously unknown DNA sequence. The use of nanopores for DNA sequencing is well known and commercially available primarily using biological nanopores. The development of solid-state nanopores provides several advantages over biological pores including high thermal, mechanical, and chemical stability, and the ability to fabricate and tune pore size.

[0004] Apart from nucleic acid characterization, the characterization of folded proteins has also been carried out using nanopores, where generic features, such as protein volume, dipole, and length, can be determined [3], [7], Nanopores can discriminate and quantify hemoglobin directly from blood; in fact, the single amino acid difference between sickle cell anemia hemoglobin and healthy adult hemoglobin was distinguished with >97% accuracy and with near 100% accuracy in fetal blood [8], Complex protein mixtures require more targeted protein detection and quantification approaches. These can be carried out via the use of DNA carriers, which are DNA molecules or nanostructures containing protein bindingmotifs at known intervals, which carry bound proteins across the pore (FIG. 2A). Alternatively, ligands such as biotin, aptamers, protein domains, or antibodies can be directly attached to a nanopore (FIG. 2B), and binding events detected [3], Nanopore technology has also been demonstrated to identify and distinguish between the monosaccharides D-fructose, D-galactose, D-mannose, D-glucose, L-sorbose, D-ribose, D-xylose, L-rhamnose and N- acetyl-D-galactosamine [9] in addition to the detection of metabolites such as vitamin Bl and Uric acid

[0010] ,SUMMARY

[0005] Described in this specification are high-throughput reading technologies for multiomics analysis. The technologies provide high-speed characterization of different types of biological particles, e.g., DNA, small molecules, and / or cellular fragments.

[0006] Described in this specification are technologies including methods for extracting information from a particle. In an aspect, the methods include providing a mixture comprising at least a first set of particles and a second set of particles. The methods include modifying at least a portion of the first set of particles with one or more labels. The methods include translocating one or more modified particles through one or more nanopores disposed in a substrate and configured to receive an input particle. The methods include reading one or more signals received from the one or more nanopores and translating the one or more signals into the information. The reading step includes detecting an electric current signal from the one or more nanopores. The reading step includes detecting a change in current through the one or more nanopores corresponding to the translocation. A first current level corresponds to passage of at least an unlabeled portion of a particle and a second current level corresponds to passage of one of the one or more labels through the one or more nanopores. The reading step includes identifying the one label from the second voltage level. The reading step includes identifying the information based at least in part on the identified label.

[0007] In some implementations, the particles are or include polymers, small molecules, biological particles or combinations thereof. In some implementations, the labels include a DNA structure, a protein, biotin-streptavidin, a bead, a fluorophore, or a combination thereof. In some implementations, the methods include labeling at least a subset of the second set of particles. In some implementations, each labeled particle is labeled with one type of label. In some implementations, each labeled particle is labeled with two or more types of label. In some implementations, each labeled particle is labeled with two or more labels of the same type.

[0008] In an aspect, a system for extracting information from a particle includes one or more second reagents to label one or more particles. The system includes a fluidic device configured for transfer fluid, the fluid comprising one or more of the first or second reagents. The system includes a processor and a memory functionally connected to the fluidic device, the memory comprising instructions that, when executed, cause the processor to actuate one or more components of the device to perform one or more method(s) for extracting information from a particle as described in this specification.BRIEF DESCRIPTION OF THE DRAWINGS

[0009] The novel features of the technologies are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present technologies will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the invention are utilized, and the accompanying drawings (also “Figure” and “FIG.” herein).

[0010] FIG. 1A is a graph illustrating the impact of molecular translocation events across a nanopore on conductance. Disruptions in current from baseline indicate translocation events across a nanopore. The raw signal is shown in blue, whilst a line of best fit is shown in red (adapted from [5]). FIG. IB is a graph illustrating changes in current as a DNA molecule translocates across a nanopore, where each nucleotide provides a different, detectable impact on conductance (adapted from [6]).

[0011] FIG. 2A is a diagram illustrating targeted protein detection from complex mixtures. DNA carriers containing protein binding motifs can be used to detect specific proteins in complex mixtures

[0011] , FIG. 2B Modified nanopores can be fabricated to contain ligands that target proteins of interest, in this example, a biotin-containing pore can detect the binding of streptavidin

[0012] ,

[0012] FIGS. 3A-3C are diagrams illustrating generalized labeling strategies for proteins, nucleic acids, and small molecules utilizing DNA-based labels. FIG. 3A shows Proteins, nucleic acids, and small molecules are introduced for site-specific labeling. FIG 3B illustrates labels for a specific biomolecule contain a structural barcode (proteins -triangle, nucleic acids -circle, and small molecules -pentagram). The barcode used for small molecules can further contain a longer DNA sequence that aids in the translocation through the nanopore. FIG. 3C illustrates site-specifically labeled biomolecules are shown in denatured form to translocate through solid-state nanopores.

[0013] FIGS. 4A and 4B are schematics illustrating the use of catalytically dead restriction enzymes to read example nucleic acid molecules. FIG. 4A illustrates example molecules A and B that can be distinguished based on restriction enzyme site composition. FIG. 4B illustrates example molecules A and B that can be distinguished based on Mnll restriction enzyme patterning.

[0014] FIG. 5A and 5B are schematics illustrating the use of multiple enzymes to label subbase sequences (or components) using different proteins. FIG. 5A illustrates direct labeling. FIG. 5B illustrates the use of indirect labels (e.g., on DNA flaps).

[0015] FIG. 6A is a schematic illustrating an example nucleic acid molecule with secondary structural features (labels). FIG. 6B is a schematic illustrating an example process to generate an example nucleic acid molecule with secondary structural features (labels). FIG. 6C is a schematic illustrating the reading of an example nucleic acid molecule with secondary structural features using nanopores.

[0016] FIG. 7A is a schematic of an example labeled DNA molecule. The example structural label is represented as a dumbbell DNA structure. FIG. 7B is a set of schematics illustrating different forms of dumbbell structures.

[0017] FIG. 8A is a schematic of 7.2 kb single-stranded DNA molecule labeled with short oligos ~40 bp in length and including 29 DNA dumbbells. FIG. 8B shows an agarose gel electrophoresis validation showing: (1) only tiling oligos, (2) 14 DNA dumbbells and tiling oligos, (3) 15 DNA dumbbells and tiling oligos (different position of dumbbells), and (4) 29 DNA dumbbells and tiling oligos.

[0018] FIGS. 9A-9C illustrate translocation of a 20 kb double-stranded DNA without any structural features. FIG. 9A is a graph showing average blockage (current) versus log dwell time of molecules in the nanopore. FIG. 9B is a graph showing an overlayed plot of individual translocation events. FIG. 9C(i-iii) are graphs showing individual translocation events of example 20 kb dsDNA fragments.

[0019] FIGS. 10A and 10B illustrates translocation of a 7.2 kb linear DNA fragment with structural labels (construct number 4 in FIG. 8). FIG. 10A is a graph showing an overlayed plot of individual translocation events. FIG. lOB(i-iii) are graphs showing individual translocation events of example 7.2 kb linear DNA fragments with structural labels.

[0020] FIG. 11A is a graph showing average blockage (current) versus log dwell time of molecules in the nanopore for translocation of an example 7.2 kb linear DNA fragment with structural labels (construct number 4 in FIG. 8). FIG. 11B shows an agarose gelelectrophoresis validation showing example 7.2 kb DNA constructs labeled with 29 DNA dumbbells (directly in the middle -400 bp long) and tiling oligos.

[0021] FIGS. 12A-12C illustrates translocation of a 7.2 kb linear DNA fragment with structural labels (construct number 2 in FIG. 8). FIG. 12A is a graph showing average blockage (current) versus log dwell time of molecules in the nanopore. FIG. 12B is a graph showing an overlayed plot of individual translocation events as described above. FIG. 12C(i- iii) are graphs showing individual translocation events of example 7.2 kb linear DNA fragments.

[0022] FIGS. 13A-13C illustrates translocation of a 7.2 kb linear DNA fragment without any structural labels, i.e., a 7.2kb single-stranded fragment with short — 30-40nt oligo tiles (construct number 1 in FIG. 8). FIG. 13A is a graph showing average blockage (current) versus log dwell time of molecules in the nanopore. FIG. 13B is a graph showing an overlayed plot of individual translocation events. FIG. 13C(i-iii) are graphs showing individual translocation events of 7.2 kb linear DNA fragments with no structural labels.

[0023] FIGS. 14A-14C are diagrams illustrating solid-state nanopore sequencing of a denatured protein labeled with two different amino acids. FIG. 14A illustrates native protein denatured and tagged with DNA-based labels. In this example, two distinct labels are used to label two different amino acids. FIG. 14B illustrates translocation of labeled protein through solid-state nanopore. The DNA handles are negatively charged and thus provide additional charge to aid in translocation. FIG. 14C is a graph showing current versus time signal depicting the three DNA labels translocated through the solid-state nanopore. This process provides a fingerprinting signal corresponding to distances (or time delays) between each labeled amino acid residue that is used to uniquely identify the protein.

[0024] FIGS. 15A and 15B are diagrams illustrating protein-protein interaction measured via antibody-DNA-scaffold. FIG. 15A illustrates protein complex (left) and antibodies targeting the different epitopes on the protein complex, immobilized on a long DNA scaffold (right). FIG. 15B illustrates protein complex and antibody-scaffold interaction. The long scaffold creates loops when the antibodies bind to the proteins creating a unique current-time signal in solid-state nanopore sequencing.

[0025] FIG. 16 is a flow chart illustrating a theoretical real-time pipeline for an example process as described in this specification.DESCRIPTION

[0026] Described in this specification are technologies including systems and methods for multi-analyte detection. Solid-state nanopore-based technology has the potential to provide a platform for the detection and characterization of a wide range of biomolecules. For example, the fabrication of larger pores allows the translocation of labelled nucleic acids. This has advantages in diagnostic applications, where specific DNA motifs can be decorated with, e.g., structural labels (see, e.g., FIGS. 3A-3C). For example, a finite set of uniquely designed DNA labels contain identifiable structural features in the subsequent current versus time signal from the nanopores. The detection of these labels can be utilized in the rapid detection of pathogens and identification of their known variants, allowing the use of targeted therapies in patients. Therefore, a nanopore-based platform may be capable of detecting and characterizing multiple types of analytes.

[0027] Described in this specification are high-throughput reading technologies that can expand the structural DNA-based labeling strategy - e.g., as applicable to DNA data storage - from nucleic acids to other biomolecules (e.g., proteins and / or small molecules) in combination with solid-state nanopores measurements. These new methods can support the development of a wide range of multi omics analytical tools to meet the needs of human health researchers pursuing precision medicine-based diagnostics and therapeutics.

[0028] In some implementations, unique structural labels can be introduced through chemical modifications on the DNA, such as site-specific methylation, nicks

[0013] , deliberately introduced gaps via post modification of double-stranded DNA (dsDNA), or introduction of unnatural bases, which can serve as a substrate for further covalent modifications (e.g., click chemistry). Unique secondary structures and site-specific overhangs (single- or doublestranded) of different lengths can further increase the diversity of engineered features

[0014] , In addition, DNA labels can consist entirely of uniquely decorated DNA nanostructures

[0015] that moreover allow tuning the persistence length of the translocated DNA label. Finally, DNA labels can be paired with site-specific enzyme docking (e.g., dCas9) or sequencespecific DNA-binding proteins

[0016] ,Protein-based labeling

[0029] Described in this specification are technologies for protein-based labeling of, e.g., nucleic acids for the identification or characterization of, e.g., DNA. The technologies described can also be used to label other particles, e.g., polymers, small molecules, or biological structures.

[0030] In some implementations, reading and characterizing DNA relies on DNA sequencing, whereby individual bases are read and, e.g., assembling the sequence based on multiple reads. This process is time-consuming and computationally intensive.Characterizing DNA libraries, however, can be accomplished by distinguishing different subbase sequences (or components) of a DNA molecule. Therefore, an approach as described here is to read regions of nucleic acids, e.g., sub-base sequences, motifs and / or components, instead of individual bases. Sequencing approaches, such as nanopore and optical mappingbased technologies, can be leveraged to detect, e.g., component labels on DNA.

[0031] In some implementations, the technologies utilize proteins that target specific DNA motifs found on specific components. Enzymes, such as restriction endonucleases, and RNA- guided endonucleases, e.g., Cas proteins, can be modified by generating mutations of their nuclease domains to eliminate their ability to cleave DNA. Alternatively, transcription factor recognition sites can be incorporated into component sequences allowing transcription factors to bind along a full-length nucleic acid molecule (FLNA). Overall, these methods result in the mentioned proteins binding to specific sequences (e.g., components or fragments thereof) and decorating FLNAs. These labels can then be identified via nanopore or optical mapping to characterize an FLNA. Rather than reading individual or multiple bases on a DNA molecule, stretches of DNA can be identified by reading the protein label, which can improve reading speed and accuracy. For example, reading speed can be increased because there is a greater distance between two labels than between two bases. Current decoding involves base-calling, component calling by mapping bases to components and then identifying / defining FLNAs by calling components. A protein-labeling technique as described here can include calling components directly.

[0032] Alternatively or additionally, optical mapping using fluorescent labels to detect and distinguish different DNA motifs can be used. These optical methods, however, may be limited due to the limited number of fluorophores available. In contrast, protein labeling has the advantage that a large library of proteins, variants, or combinations can be created to distinguish multiple components. In some implementations, fluorescent labels can be used to tag / label the labeling proteins, or can be used alone or in combination with any other labeling technology described in this specification (e.g., to label one or more structural motifs, e.g., hairpins, dumbbells or other). Fluorescent signals can be detected using, e.g., CCD cameras or other optical detectors. Such cameras or detectors can be combined with one or more nanopore readers as described in this specification into a single device or can be stand-alone devices.

[0033] There are different variants of this approach. In some implementations, a catalytically dead restriction enzyme can be used as illustrated in FIGS. 4A-4C. In FIG. 4A, FLNAs A and B can be distinguished based on restriction enzyme site composition. In FIG. 4B, Mnll restriction enzyme sits on two different identifiers, illustrating how a pattern of Mnll can classify FLNAs.

[0034] Other variations include the use of multiple enzymes to label, e.g., components or motifs using different proteins (FIG. 5A) or the use of indirect labels (e.g., on DNA flaps) as shown in (FIG. 5B).

[0035] The proteins used in the technologies described in this specification can have modifications, e.g., phosphorylation, acetylation, hydroxylation, ubiquitination, or methylation. The modifications can be used as a method of identifying DNA sequences. Protein cross-linking agents, such as formaldehyde, can be added to crosslink proteins to the DNA, preventing them from disassociating.Labeling Strategies for Motif or Component Identification via Nanopore Readout

[0036] In some implementations of the technologies described in this specification, a nick- translation-driven labeling scheme can be used for high-throughput component identification via nanopore readout.

[0037] A feature of the labeling technologies described in this specification is the use of labeling distance and labeling types specific for components that provide a unique signal (e.g., a unique current-time signal pattern or a “fingerprint”), via nanopore-based DNA readout. In some implementations, chemistry optimization and labeling rule determination utilizing nick-translation can be used. In an example implementation, first, nicking enzymes can be used to produce single-stranded “nicks” at sequence-specific locations along the FLNA. Next, DNA Polymerase I can be used to replace some of the nucleotides of a DNA sequence with their labeled analogues. Finally, the original “nick” is sealed by DNA ligase.

[0038] In some implementations, methylation with CpG Methyltransferase (M.SssI), which methylates all cytosine residues (C5) within the double-stranded dinucleotide recognition sequence 5'...CG...3' can be used. Next, the methylated FLNA is run on established DNA sequencing technology, e.g., nanopore sequencing to generate an amplified signal related to the methylation sites that can be compared to a predetermined FLNA set to be tested. A signal processing scheme can be used to classify nanopore signals corresponding to labels.Example labels for increased sensitivity in translocation signal detection using solid-state nanopores

[0039] Unlike biological nanopores, solid-state nanopores do not incorporate proteins into their systems, but use various metal or metal alloy substrates with nanometer sized pores that allow DNA or RNA to pass through. Solid-state nanopore measurements typically have a higher translocation speed that can be ~ 1000-fold higher over standard biological nanopore translocation driven by helicase enzymes (Oxford Nanopore Technologies: ~ 450 bp / s). The increased translocation speed results in a current versus time signal that, depending on the sampling rate, can impede the accurate detection or identification of bases or structural features. To increase the sensitivity of translocation signal resulting from linear DNA molecules decorated with structural DNA labels, a DNA expandomer strategy can be deployed, as described in this specification. The strategy includes ligating individual oligos or more complex nanostructures onto a scaffold (this can be a single-stranded DNA, e.g., an FLNA) that can subsequently be expanded under denaturing conditions (or by any other means that unfolds the DNA expandomer). The expandomer structures are ligated onto or between structural labels. The resulting increase in template length of a label (e.g., anywhere between 10 and 10,000 nucleotides) adds additional DNA that will space out the structural features and allow better detection of structural labels in the current versus time measurements on the solid-state nanopore. The current signal or change can include one or more changes in signal spacing, signal amplitude, or a combination thereof.

[0040] FIGS. 6A-6C illustrate the reading of an example FLNA with secondary structural features using nanopores. FIG. 6A shows a schematic depicting a DNA object with seven unique structural features or labels, in this example a linear double-stranded DNA template with structural labels SO to S6. The length of DNA oligo can vary and can include between 1 and 10 labels, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10. In some implementations, a DNA oligo can include 10 or more labels, e.g., between 10 and 30 labels. Structural labels can be any nucleic acid (e.g., DNA) nanostructure composed of a single DNA strand with a secondary structure or can be composed of multiple strands that form a 2D or 3D DNA nanostructure. Furthermore, structural labels can also include one or more chemical modifications that distinguish the labels by charge or molecular weight. FIG. 6B illustrates an example combinatorial assembly used to assemble structural labels on an FLNA template. An FLNA or other double-stranded molecule can be rendered single-stranded. Subsequently, structural labels are added (and / or ligated), or an FLNA is assembled from single components already containing the structural labels. In addition, in some implementations, an FLNA can alsoconsist of only single-stranded DNA that contains structural labels comprising DNA secondary structures. In some implementations, the features or secondary structures of a label are designed such that the labeled DNA segment almost completely fills the nanopore. FIG. 6C illustrates an example FLNA with structural labels translocating through a (solid-state) nanopore.

[0041] FIG. 7A illustrate an example labeled DNA molecule (e.g., an FLNA). The example structural label is represented as a dumbbell DNA structure. FIG. 7B illustrates different forms of dumbbell structures where t-arm length, t-arm stem, and / or number of arms can vary in nucleotide composition or size, or both. In some implementations, the nucleic acid structure of the labels can vary in one or more of size, shape, stem length, or mechanical properties (flexibility). In some implementations, one or more of the nucleic acids of the labels are decorated, e.g., with one more atoms or molecules. In some implementations, the labels described in this specification, e.g., dumbbells, hairpins, proteins, etc. are fluorescently labeled. Hairpins can have a length of 10-500 bp and can have different loop sizes. In some implementations, hairpins include unnatural or methylated bases and / or are chemically modified. In some implementations, one or more of the nucleic acids of the labels are decorated with one or more fluorescently labeled nucleotides (e.g., dye- or quencher-labeled molecules). In some implementations, one or more of the nucleic acids of the labels include nucleotides decorated with small moieties. Such moieties can be or include Biotin, Desthiobiotin, Digoxigenin, DNP (Dinitrophenol), Photo-labile groups (“Caged”), Triple bonds (Alkyne), DBCO, Azide (-N3), Trans-Cyclooctene (TCO), Vinyl, Free amino group (- NH2), Redox Dyes, Halogen atoms (F, Cl, Br, I), Mercury (Hg), Selenium (Se), Methyl group (-CH3). In some implementations, click chemistry can be used to decorate nucleic acids of the labels, wherein pairs of functional groups rapidly and selectively react (“click”) with each other under mild, aqueous conditions. This method provides a convenient, versatile, and reliable two-step coupling procedure for coupling molecules A and B, where A is or includes DBCO-containing nucleotides or Alkyne-containing nucleotides and where B is or includes Azide-containing fluorescent dye, Azide-containing (desthio) biotinylation reagent, Azide- containing FLAG reagent, or Azide-containing amino acid, or where A is or includes an Azide-containing nucleotides and B is or includes DBCO-containing fluorescent dye, a (desthio) biotinylation reagent, a FLAG reagent, an Alkyne-containing fluorescent dye, or an amino acid. In some implementations one or more of the nucleic acids of the labels include nucleotides decorated using (1) a dead Cas9 system programmed to target and bind — but not cleave — a specific structural label region (~20 nt) by a guide RNA, (2) transcription factors(TF) to bind to a label region that contain specific TF recognition sites (as described above), or (3) catalytically dead restriction enzymes to target binding motifs designed into the label regions.

[0042] FIGS. 8A-8B illustrate an example molecule (FLI) labeled with a set of “dumbbell” DNA labels. FIG. 8A is a schematic of a 7.2 kb single-stranded DNA molecule labeled with short oligos ~40 bp in length and including 29 DNA dumbbells located in the middle (length of labeled region: -400 bp). The labeled portion of the construct is flanked by regions comprising “tiles,” i.e., short single-stranded unlabeled oligos that are separated by nicks (flanking regions -3 kb). FIG. 8B shows results of an agarose gel electrophoresis validation (2% agarose gel, 10 pL SYBR Safe in 160 mL gel (IxTE, 10 mM MgCh), 60 V for 4.5 h, loaded equimolar amounts), showing four different 7.2 kb DNA constructs labeled with (1) only tiling oligos, (2) 14 DNA dumbbells and tiling oligos, (3) 15 DNA dumbbells and tiling oligos, and (4) 29 DNA dumbbells and tiling oligos. The 14 and 15 dumbbells of samples 2 and 3 are located in the same region of sample 4. Sample 2 contains the first 14 dumbbells and sample 3 contains the last 15 dumbbells (assuming counting from left to right). Results indicate little variation in molecular weight. The two lanes to the left show results for circularized and linearized Ml 3 bacteriophage genomes.

[0043] FIGS. 9A-9C illustrate translocation of a 20 kb double-stranded DNA without any structural features. Experimental conditions were: 1 M KC1 + lx TBS Buffer, 15 nm Norcada nanopore, 200 mV, 12.4 nA, 4ml2s experiment duration. FIG. 9A shows average blockage (current) versus log dwell time of molecules in the nanopore. FIG. 9B is an overlayed plot of individual translocation events (All plots were generated using Nanolyzer software). The change in current from about 0 to about -400 indicates entry of a molecule into the nanopore, while the change in current from about -400 to about 0 indicates exit of the molecule from the nanopore. FIG. 9C (i-iii) show individual translocation events of example 20 kb dsDNA fragments. FIG. 9C (iii) shows a step in the signal at about 550ms, indicating translocation of a folded molecule.

[0044] FIGS. 10A-10B illustrates translocation of a 7.2 kb linear DNA fragment with structural labels (construct number 4 in FIG. 8). Experimental conditions: full set of 29 dumbbells present (-400 bp stretch), 3.75 M LiCl + lx TE Buffer, 15 nm Norcada nanopore, 200 mV, 14.5 nA, 3m43s experiment duration. FIG. 16A is an overlayed plot of individual translocation events. FIG. lOB(i-iii) show individual translocation events of example 7.2 kb linear DNA fragments with structural labels. FIG. 10B(i) shows a translocation event with a particularly strong signal from the structural feature in the middle 400 bp stretch of themolecule containing 29 DNA dumbbells (corresponding time interval: approx. 1800-2000). It should be noted that the number of read molecules can be increased to improve signal-to- noise ration (e.g., increasing the likelihood of reading molecules with strong signals).

[0045] FIG. 11A shows average blockage (current) versus log dwell time of molecules in the nanopore for translocation of an example 7.2 kb linear DNA fragment with structural labels (construct number 4 in FIG. 8). Experimental conditions were: All dumbbells present (-400 bp stretch), 3.75 M LiCl + IxTE Buffer, 15 nm Norcada nanopore, 200 mV, 14.5 nA, 3m43s experiment duration. FIG. 11B illustrates an agarose gel electrophoresis validation (2% agarose gel, 10 pL SYBR Safe in 160 mL gel (IxTE, 10 mM MgCh), 60 V for 4.5 h, loaded equimolar amounts), showing example 7.2 kb DNA constructs labeled with 29 DNA dumbbells (directly in the middle -400 bp long) and tiling oligos. The gel also shows individual (not hybridized) tiling oligos and DNA dumbbells running at a lower molecular weight than the constructs.

[0046] FIGS. 12A-12C illustrate translocation of a 7.2 kb linear DNA fragment with structural labels (construct number 2 in FIG. 8), i.e. with only the first 14 DNA dumbbells out of the 29 possible dumbbells (see FIG. 12B). The labeled length is about -200 bp (approx, half of the ~400bp tiled by the 29 dumbbells). Experimental conditions were: 3 pL DNA [10 nM stock] in 50 pL KC1 [1 M] + lx TBS Buffer, 10 nm Norcada nanopore, 150 mV, 2m45s experiment duration. FIG. 12A shows average blockage (current) versus log dwell time of molecules in the nanopore. FIG. 12B is an overlayed plot of individual translocation events as described above. FIG. 12C(i-iii) show individual translocation events of example 7.2 kb linear DNA fragments. The structural labels are clearly visible, e.g. at approximate time intervals 790-840 (i), 800-850 (ii); and 2100-2200. Note that a nanopore is agnostic in terms of direction of a molecule entering the pore. Therefore, different time plots can be seen depending on whether molecule entered with the left or right end first(FIG. 12B)

[0047] FIGS. 13A-13C illustrates translocation of a 7.2 kb linear DNA fragment without any structural labels, i.e., a 7.2kb single-stranded fragment with short — 30-40nt oligo tiles (construct number 1 in FIG. 8). Experimental conditions were: all tiles (no dumbbells), 1.5 pL DNA [10 nM stock] in 50 pL KC1 [1 M] + lx TBS Buffer, 10 nm Norcada nanopore, 150 mV, 8.8 nA, 5m0s experiment duration. FIG. 13A shows average blockage (current) versus log dwell time of molecules in the nanopore. FIG. 13B is an overlayed plot of individual translocation events. FIG. 13C(i-iii) show individual translocation events of 7.2 kb linearDNA fragment with no structural labels. None of the plots show no distinguishable feature / label signals.Prioritized / selective nanopore sequencing

[0048] Throughput in nanopore sequencing is dependent on nanopore-finding rate (e.g., how many molecules, e.g., FLNAs can find a nanopore and / or how fast a molecule can find a nanopore) and nanopore translocation rate. The technologies described in this specification can be used to increase pore-finding rate.

[0049] In some implementations, to increase throughput, the length of molecules (e.g., FLNAs) can be increased, e.g., by concatenating shorter fragments. In some implementations, an example system is configured to have longer molecules reach nanopores first. For example, the charge of DNA or the size of a DNA molecule can be used for size selection. For example, agarose gels can be used for selection, e.g., by running a gel in reverse. Using this method, longer molecules can reach nanopores first. In some implementations, flow- based techniques can be used. In some implementations, (electrical) resistance can be used for size selection.

[0050] In some implementations, the rim of one or more nanopores can be decorated with one or more molecules that attract other molecules. This technique is similar to affinity capture near a nanopore. For example, a charged molecule positioned near a nanopore can attract a DNA molecule. Tuning charge density (e.g., of a DNA molecule, e.g., an FLNA, or of the charged molecule) can improve selection performance. In some implementations, a library of FLNAs includes molecules tagged with beads (e.g., magnetic beads). The beads can be attracted towards a nanopore. In some implementations, nanopores can be pre-loaded with bead-tagged DNA molecules (e.g., a library of FLNAs) given a voltage potential across the cis and trans chambers of the nanopore, or nanopore array. In this configuration, the bead dimension (a sphere with diameter, Db) is larger than the nanopore diameter (DP), which essentially captures the bead on the cis side of the chamber and stretches the FLNA in the pore lumen towards the trans side. After this pre-loading step, beads can be cleaved so that the molecule, e.g., an FLNA, can translocate through the nanopore. In some implementations, a bead can be attached with an ssDNA adapter. In some implementations, beads and / or helicase enzymes (as stoppers) can be used to repeatedly read an FLNA. First, the DNA molecules tagged with a stopper are captured in the nanopore using a voltage potential applied across the nanopore or nanopore array and subsequently identified based on their current blockade signature (an electronic fingerprint). Next, by reversing the voltagepotential, the detected molecule is “kicked out” from the nanopore. Then, by continuously iterating between capturing and “kicking out” the molecule (using voltage reversal), the molecule can be re-attracted to the nanopore and read it again. Using this method, the same DNA molecule can be read / interrogated multiple times, increasing the probability of accurately identifying the molecule or its components based on its repeated current blockade signature. In some implementations, the DNA that has been read and accurately identified can be (enzymatically, and / or chemically) cleaved in a selective manner near a nanopore, which will minimize excessive data acquisition upon molecule classification, especially when using a large (millions of pores) nanopore array.

[0051] In some implementations, a library of DNA molecules (e.g., FLNAs) is prepared with helicases so that the helicases can find a nanopore faster than DNA without helicase. In some implementations, helicase can be split, with one half attached to the nanopore and another attached to the identifier molecule so that the helicase units assemble faster than the DNA finds the nanopore.

[0052] In some implementations, the translocation rate of a molecule can be adaptively tuned based on how many molecules are captured successfully. In other words, given an initial molecular concentration, the translocation rate can be dynamically tuned, e.g., by increasing / decreasing the voltage gradient across the nanopore or nanopore array, which reflects in the increase / decrease of the number of molecules, e.g., FLNAs, captured through the pores in a unit time. Given a constant voltage potential across a nanopore array, the initial concentration of the molecules contained in the cis chamber will decrease as a function of time, as the molecules will translocate to the trans chamber. Thus, a constant translocation rate can be dynamically established based on the actual concentration (number of molecules) in the cis chamber, which is measured as the rate of molecule detection in the nanopore. Therefore, when the rate of molecule detection decreases (the number of molecules in the cis chamber decreases), the voltage gradient can be dynamically adjusted across the pore array to accelerate molecule transport and hence establish a constant translocation rate, e.g., until all molecules have been translocated to the trans side of the nanopore chamberParticle Labeling

[0053] The labels described in this specification, e.g., DNA-based labels (e.g., the DNA structures, e.g., dumbbells described above) can be used to label one or more particles. The particles include polymers (e.g., proteins, lipids, complex carbohydrates, nucleic acid molecules, or a combination thereof), small molecules (e.g., small molecule therapeuticagents, small molecule diagnostic agents, biological metabolites, or a combination thereof), biological structures (e.g., viral particles, cellular fragments, cellular particles, 3D protein structures, 3D nucleic acid structures, liposomes, or a combination thereof), or combinations thereof. The each of the particles can be labeled with one or more labels. Each particle can be labeled with a plurality of labels, where the labels are identical or different. Labels that are used to label the particles include biotin-streptavidin, DNA structures, proteins, fluorophores, or a combination thereof.

[0054] DNA-based labels can be site-specifically introduced via reactive groups to label amino acids, nucleotides, different length sequences, or reactive groups on small molecules

[0017] —

[0019] . In an implementation, the DNA labels are categorized with a structural barcode, according to the biomolecule category (FIGS. 3A-3C) to which each label will be covalently linked. This combinatorial approach allows for a reduction in the number of uniquely engineered labels required for the correct identification of a biomolecule (FIG. 3B) and offers the capability to deploy a multiomics approach suitable for heterogeneous samples. For example, protein fingerprinting can occur when a subset of specific amino acids is labeled, which offers the ability to fingerprint and subsequently identify the protein based on known amino acid distances (FIGS. 14A-14C)

[0017] ,

[0018] , The site-specific labeling of amino acids and subsequent fingerprinting can also serve as an untargeted approach to identify unknown proteins as part of a protein family with already known sequences, e.g., slight variations in the fingerprinting could be mapped back to known sequence variants. Similarly, this approach can be used as a coarse metagenomics approach in nucleic acid identification, where labels are only used to target specific sequences of interest, e.g., counting the number of sequence repeats or identifying if a sequence contains a known pathogenic sequence

[0020] , Small molecule labels can contain longer DNA tags with known sequences to provide additional signals during solid-state sequencing (FIG. 3B). In some implementations, DNA- based labels are designed to contain a reaction group (denoted as R in the figure) that allows site-specific labeling of amino acids, nucleotides, different length sequences, and reactive groups on small molecules. The DNA-based labels can be of different number of nucleotides or strands, both single and double-stranded, contain secondary structures, contain modified nucleotides that can covalently link double-stranded DNA (e.g.,cnvK) or can be used to link other chemical groups to the DNA. Labels for a specific biomolecule contain a structural barcode. The barcode used for small molecules can further contain a longer DNA sequence that aids in the translocation through the nanopore.

[0055] To further expand, known proteins with already discovered antibody binders can be investigated for protein-protein interactions via a DNA scaffold-mediated approach

[0020] ,

[0021] , Antibodies can be immobilized on a DNA scaffold by conjugating handle sequences to the antibodies (FIG. 15A). When a protein complex is introduced, the antibody-epitope interaction loops the DNA scaffold in different length loops that generate a unique signal on the solid-state measurement (FIG. 15B)

[0021] ,

[0056] Signal classification: A state-of-the-art (SOTA) signal detection and processing pipeline for a multiomic biological analyzer can ingest electrical traces from a heterogeneous sample of analytes and produce a breakdown of sample constituents and their relative populations. SOTA signal detection may take place in a real-time computational pipeline, which involves receiving a constant stream of electrical traces from the analyzer and applying labels from a classification algorithm to these traces

[0022] , Multi -barcode classification may require a labeled library of trace-barcode data for training a supervised machine learning model, which would be prepared with known analyte solutions in advance.

[0057] An example process for the addition of an analyte to a trace-barcode library can include the following steps1. Running the analyte sample, generating an ensemble of electrical traces. Analytes may exhibit orientation dependence while passing through the pore

[0023] , necessitating an ensemble of electrical traces per barcode to properly train a model for all instances of an analyte.2. Each trace generates distinct features that describe a barcode or sequence of barcodes.3. The ensemble of trace features is added to the trace-barcode dataset and divided into training and testing sets, e.g., for machine learning.4. The training data is used with a supervised model to find the best-fitted model.5. The best-fitted model is evaluated using testing data to determine its unbiased accuracy.

[0058] Several supervised models, such as neural networks, support vector machines, random forests, and decision trees, have been used for DNA classification

[0024] ,

[0025] and can be extended for multi -barcode classification of electrical signal traces from non-DNA analytes. Unsupervised models can also be used to help identify unknown analytes via their proximity to known analytes in the feature space of the trace

[0025] , The real-time feature generation and labeling of trace data using an arbitrary model can be done with a streaming big data framework, which is easily scalable

[0026] , A SOTA computational pipeline can contain a real-time component that ingests data and applies the latest model, and a periodic batch component that retrains a model with new data (FIG. 16).

[0059] In some implementations, the technologies described in this specification can include models providing the identification of specific combinations of molecular labels to provide the characterization of the specific nucleic acid molecules present in a DNA data sample, e.g., identifiers for storing digital information in nucleic acids. In an implementation, such technologies include a method for writing information into nucleic acid sequence(s), including: (a) translating the information into a string of symbols; (b) mapping the string of symbols to a plurality of identifiers, wherein an individual identifier of the plurality of identifiers comprises one or more components, wherein an individual component of the one or more components comprises a nucleic acid sequence, and wherein the individual identifier of the plurality of identifiers corresponds to an individual symbol of the string of symbols; and (c) constructing an identifier library comprising at least a subset of the plurality of identifiers. This technology can form the basis for signal analysis strategies that naturally extends to multiomic analysis for samples measured with the methods described above.

[0060] Labeling scheme: In some implementations, the technologies described in this specification include a nanopore array-based approach to particle classification, e.g., nucleic acid classification, that requires particles, e.g., double-stranded, labeled DNA, to be read at high speed. The closest label spacing that can be reliably discriminated can determine the limits of multilabel approaches per single biomolecule. To determine the optimal choice of a set of labels producing unique nanopore signals or fingerprints, labels are screened and the best labels are selected that meet the required nanopore signal-to-noise ratio threshold by investigating a variety of structures such as biotin-streptavidin, 3D DNA structures, and enzyme labels. By assessing results on the optimal spacing of labels, amount, and type of labels, it is possible to design the appropriate labeling scheme for the NPFET-based readout.

[0061] Sample preparation: As a first step, standard sample preparation techniques like those for serum or tissue samples, following established protocols for protein sequencing and / or nucleic acid extraction from the genomics or molecular diagnostics field, can be implemented

[0027] , To ensure the orthogonality of labeling, samples can be split and addressed with each biomolecule labeling scheme individually.

[0062] Going beyond preparing samples of the same molecular species, a labeling strategy coupled with ssNP analysis makes it possible to detect multiple classes of analytes (e.g., DNA / RNA, protein, and small-molecule targets) simultaneously directly from human blood serum without amplification or extraction steps. However, a filtration step may be required toremove large (e.g., >10 kDa) proteins to prevent pore blockage (that reduces the overall BM capture rate). Alternatively, real-time nanopore sensing data can be used to monitor the individual pore activities, and dynamically actuate voltage polarity change to unblock pores from contaminants by ejecting them into a cis reservoir. Likewise, this technique can be used to enrich biomolecules of interest (adaptive sampling) by ejecting unwanted molecules - and hence capturing the target molecules - based on real-time signal identification upon their capture in the pore.

[0063] Labeling combinatorics: Theoretical analysis of protein fingerprinting has shown that only a small subset of amino acids (2-3 out of 20) requires unique labels to identify >95% of the whole human proteome. For example, translocating SDS-denatured proteins and recording the counts and order of three unique labels targeting only cysteine, lysine, and methionine results in the identification of -99% of the human proteome (using the human Swiss-Prot database

[0028] ). Each of the three amino acids can be site-specifically labeled with orthogonal chemistries. Lysines can be modified with NHS esters, cysteines with maleimide groups, and methionines with a two-step redox-activated chemical tagging

[0027] ,

[0029] —

[0031] ,

[0064] To ensure DNA labels have high specificity, DNA sequence design can aid in locating suitable target regions and optimizing label sequences. This process can include the removal of secondary structures, selecting sequence unique target regions, modifying the length of the binding region, and / or introducing chemistries for the covalent linkage of a duplex.

[0065] Methods like biotin redox-activated chemical tagging can be used in tagging small molecules conjugated to oligonucleotides that contain a biotin modification

[0019] , For further discrimination of groups of small molecules, or small molecules bound to protein targets previously reported, affinity -based pull-down methods may be adapted for labeling schemes in ssNP sequencing

[0032] ,

[0066] Nanopore fabrication: ssNPs can be produced at quantity down to 15 nm (or smaller), using advanced photolithography and patterning processes. The passive ssNP has excellent noise properties and can be readily used to develop assays using established labeling techniques. Advanced NPFET fabrication processes may be used to bring nanopore signals to the CMOS plane with high signal bandwidth. Nanopore diameter size reduction to <10 nm is expected to be achievable by further optimizing processes, which will improve SNR. In some implementations, single-base resolution may not be required in the proposed detection approach, so <3 nm nanopores may not be required. This design can provide the capabilities for reading of 0.5-30 kbp DNA molecules and other biomolecules (proteins, metabolites, and other small molecules) of equivalent persistence length at very high throughput.

[0067] Scalability: Conventional methods of reading nanopores are limited by the signal bandwidth of individual nanopores and the number of nanopore cells that can be arrayed on one die. The NPFET has the intrinsic advantage of offering larger signal bandwidths than existing solid-state technologies. While 10 MHz and 1,000 nanopores are achievable, technical solutions for higher integration densities may be explored. Moreover, techniques that can reduce the unit cell size, e.g., by separating the critical analog front-end function close to the nanopore from any analog and / or digital processing that can be done more efficiently further away, may help bring down the active area per nanopore. The approach using standard CMOS for ssNP readers combined with a microfluidic cartridge, processed in industry-compatible settings, brings manufacturability to the nanopore community. These technologies can provide a technology roadmap for advancing read-throughput across various emerging technology nodes, which can be translated into a cost estimation for the various multi omic analysis workflows.

[0068] Example approach: There is a wide spectrum of emerging DNA nanotechnologies ranging from synthetic biology to diagnostics that address the grand challenge of enabling digital multi omics. For example, a super-resolution technique, DNA-PAINT, provides a promising way to fingerprint DNA, RNA, and proteins on the level of single molecules [1], Another method for DNA-based protein identification, called DNA proximity recording, generates DNA records that vary in length and abundance according to pairwise distances within a protein [2], Moreover, biological, and solid-state nanopores may serve as a tool to analyze biomolecules and are well-suited to detect proteins, metabolites, and / or other analytes [3], High-throughput single-molecule detection platforms are being developed using solid-state nanopore (ssNP) technology. Since a multiomics characterization of a biological sample would require hundreds of billions, even a trillion reads per experiment, one strategy is to address this challenge by creating a very high-density array of nanopores with individually addressable electrodes. This array of nanopores can be used in conjunction with molecular labeling strategies providing the identification of signatures unique to specific macromolecules.

[0069] Reader: Based on initial calculations for ssNPs, the throughput of a single nanopore field-effect transistor (NPFET) reader can be estimated and extrapolated for an array of NPFETs cointegrated on a complementary metal-oxide-semiconductor (CMOS) chip. NPFETs have the potential to break the translocation speed limitations of regular nanopores due to their improved detection bandwidth (10 MHz versus 5 kHz), while they can also scale to very large arrays (e.g., to the order of millions). To this end, in some implementations,these sensor arrays may analyze up to 109barcode molecules, e.g., biomolecules (BM), per second. Importantly, a single NPFET reader allows much higher (103BM / s) data throughput than regular biological pores working at 10° BM / s read rates. Therefore, the technologies described in this specification provide the capabilities to read the freely translocating biomolecules at a very high translocation speed. The NPFET can achieve this due to its high bandwidth, which can be maintained at very high integration densities. In an example implementation, an array of IM NPFET readers, up to 1012BM / day DNA read throughput may be achieved. In some implementations, the reader can read a plurality of particles, e.g., nucleic acid molecules, at a very high throughput from which a consensus read can be generated. In some implementations, reading nucleic acid molecules includes circularizing one or more nucleic acid molecules and performing rolling circle amplification

[0070] The technologies described in this specification provide a number of advantages: The technologies improve multiplex reading by DNA barcoding of biomolecules solving the problem of nanopore signal classification rather than single-molecule sequencing. The technologies include novel nanopore field-effect transistor (NPFET) arrays to demonstrate the technical feasibility and cost-effectiveness of this novel HT reading method for multi omics analysis. The nanopores can have a diameter of 10-15 nm, or less. The technologies bring the read speed to a potential array -based throughput of up to 1015biomolecules / day. The technologies provide scalability via massive parallelization by combining the precise control of CMOS-based process technology and reader electronics. The technologies improve on technologies in the emerging landscape of single-molecule protein sequencing and fingerprinting technologies. The technologies open new commercialization opportunities for biotechnology and pharmaceutical companies in the life sciences.

[0071] Next-generation sequencing and single-molecule DNA sequencing technologies have revolutionized genomics and transformed precision medicine-based diagnostics. Proteomics anticipates a similar transformative wave - fueled by the emerging landscape of singlemolecule protein sequencing and fingerprinting technologies - with the promise of resolving the full proteome with single-protein resolution, opening the door to unprecedented opportunities in basic science and medical diagnostics.

[0072] These emerging technologies have underlined the need for innovative approaches to attach various functional groups to, e.g., peptides with a high degree of specificity to avoid downstream misidentification of amino acids, which could lead to sequencing errors. The labeling strategies and detection technologies described in this specification can fuel newcommercialization opportunities in the life sciences to realize new generations of products and services. The described strategies also present an opportunity to speed up read times for the genomic community, for example, by improving metagenomic analysis and repeat expansion detection, by utilizing the massive nanopore array reading technology.Example Implementations

[0073] Item 1. A method for extracting information from a particle, comprising: (a) providing a mixture comprising at least a first set of particles and a second set of particles; (b) modifying at least a portion of the first set of particles with one or more labels; (c) translocating one or more modified particles through one or more nanopores disposed in a substrate and configured to receive an input particle; (d) reading one or more signals received from the one or more nanopores and translating the one or more signals into the information, the reading comprising (i) detecting an electric current signal from the one or more nanopores; (ii) detecting a change in current through the one or more nanopores corresponding to the translocation, wherein a first current level corresponds to passage of at least an unlabeled portion of a particle and a second current level corresponds to passage of one of the one or more labels through the one or more nanopores; (iii) identifying the one label from the second voltage level; and (iv) identifying the information based at least in part on the identified label.

[0074] Item 2. The method of item 1, wherein the first set or the second set, or both, comprise a plurality of polymers.

[0075] Item 3. The method as in any one of items 1-2, wherein the first set or the second set, or both, comprise a plurality of small molecules.

[0076] Item 4. The method as in any one of items 1-3, wherein the first set comprises a plurality of polymers and the second set comprises a plurality of small molecules, or vice versa.

[0077] Item 5. The method as in any one of items 1-4, wherein the first set or the second set, or both, comprise a plurality of biological structures.

[0078] Item 6. The method as in any one of items 1-5, wherein the first set comprises a plurality of polymers and the second set comprises a plurality of biological structures, or vice versa.

[0079] Item 7. The method as in any one of items 1-5, wherein the first set comprises a plurality of small molecules and the second set comprises a plurality of biological structures, or vice versa.

[0080] Item 8. The method of item 1, wherein the first set is a plurality of first polymers and the second set is a plurality of second polymers.

[0081] Item 9. The method of item 1, wherein the first set is a plurality of first small molecules and the second set is a plurality of second small molecules.

[0082] Item 10. The method of item 1, wherein the first set is a plurality of first biological structure and the second set is a plurality of second biological structures.

[0083] Item 11. The method of item 1, wherein the first set is a plurality of polymers, and the second set is a plurality of small molecules, or vice versa.

[0084] Item 12. The method of item 1, wherein the first set is a plurality of polymers, and the second set is a plurality of biological structures, or vice versa.

[0085] Item 13. The method of item 1, wherein the first set is a plurality of small molecules, and the second set is a plurality of biological structures, or vice versa.

[0086] Item 14. The method as in any one of items 2-13, wherein the polymers comprise proteins, lipids, complex carbohydrates, nucleic acid molecules, or a combination thereof.

[0087] Item 15. The method as in any one of items 3-14, wherein the small molecules comprise small molecule therapeutic agents, small molecule diagnostic agents, biological metabolites, or a combination thereof.

[0088] Item 16. The method as in any one of items 4-15, wherein the biological structures comprise viral particles, cellular fragments, cellular particles, 3D protein structures, 3D nucleic acid structures, liposomes, or a combination thereof.

[0089] Item 17. The method as in any one of items 1-16, modifying at least a portion of the second set of particles with one or more labels.

[0090] Item 18. The method of item 17, wherein the labeled particles of the first set, the second set, or both are labeled with at least two labels.

[0091] Item 19. The method as in any one of items 1-18, wherein the first set is labeled with a first label producing a first signal and a second label producing a second signal.

[0092] Item 20. The method of item 19, wherein the first signal is distinct from the second signal.

[0093] Item 21. The method of item 19, wherein the first signal is substantially the same as the second signal.

[0094] Item 22. The method as in any one of items 1-18, wherein the first set is labeled with a first label producing a first signal and the second set is labeled with a second label producing a second signal.

[0095] Item 23. The method of item 22, wherein the first signal is distinct from the second signal.

[0096] Item 24. The method of item 22, wherein the first signal is substantially the same as the second signal.

[0097] Item 25. The method as in any one of items 1-24, comprising modifying at least a portion of a third set of particles with one or more labels.

[0098] Item 26. The method of item 25, wherein substantially all particles of the third set are labeled.

[0099] Item 27. The method as in any one of items 5-26, wherein the third set is labeled with a third label producing a third signal, the third signal distinct from the first signal and the second signal.

[0100] Item 28. The method as in any one of items 25-26, wherein the third set is labeled with a third label producing a third signal, the third signal substantially the same as the first signal or the second signal.

[0101] Item 29. The method as in any one of items 1-28, wherein the one or more labels are or comprise biotin-streptavidin, DNA structures, protein labels, or a combination thereof.

[0102] Item 30. The method of item 29, wherein the one or more labels comprise an enzyme label.

[0103] Item 31. The method of item 29, wherein the protein is an endonuclease.

[0104] Item 32. The method as in any one of items 29-31, wherein the protein is modified by phosphorylation, acetylation, hydroxylation, ubiquitination, methylation, or cross-linking.

[0105] Item The method as in any one of items 29-32, wherein the one or more labels comprise two or more proteins.

[0106] Item 34. The method as in any one of items 1-33, wherein the one or more labels comprise an indirect label on a DNA flap.

[0107] Item 35. The method as in any one of items 1-34, wherein the one or more labels are or comprise one or more DNA labels.

[0108] Item The method of item 35, wherein the one or more one or more DNA labels comprise a barcode.

[0109] Item 37. The method of item 35, wherein the one or more one or more DNA labels comprise a hairpin.

[0110] Item 38. The method of item 37, wherein the one or more hairpins have a length of 10-500 bp.[OHl] Item 39. The method as in any one of items 37-38, wherein two or more hairpins have different hairpin stem length.

[0112] Item 40. The method as in any one of items 37-39, wherein two or more hairpins have different numbers of hairpin stems

[0113] Item 41. The method as in any one of items 37-40, wherein two or more hairpins have different hairpin loop sizes

[0114] Item 42. The method as in any one of items 37-40, wherein two or more hairpins incorporate one or more unnatural bases or methylated bases.

[0115] Item 43. The method as in any one of items37-42, wherein the hairpins are chemically modified.

[0116] Item 44. The method as in any one of items 37-43, wherein the one or more hairpins are located on one strand of a double stranded nucleic acid molecule.

[0117] Item 45. The method as in any one of items 1-44, wherein the one or more labels comprise an expandomer.

[0118] Item 46. The method as in any one of items 1-45, wherein the one or more labels comprise a DNA dumbbell.

[0119] Item 47. The method as in any one of items 1-46, comprising reading a plurality of nucleic acid molecules and generating, from the plurality of reads, a consensus read.

[0120] Item The method of item 47, comprising circularizing one or more nucleic acid molecules and performing rolling circle amplification.

[0121] Item 49. The method as in any one of items 1-48, wherein the one or more nanopores are or comprise a Nanopore Field Effect Transistor (NPFET)

[0122] Item 50. The method as in any one of items 1-49, wherein the one or more nanopores are integrated into a complementary metal-oxide-semiconductor (CMOS) circuit.

[0123] Item 51. The method as in any one of items 1-50, wherein the one or more nanopores have a diameter of 10-15 nm.

[0124] Item 52. The method as in any one of items 1-51, wherein the one or more nanopores have a diameter of less than 10 nm.

[0125] Item 53. The method as in any one of items 1-52, wherein a first label causes a first current change in the one or more nanopores, and a second label causes a second current change.

[0126] Item 54. The method of item 53, wherein the first current change is different from the second current change.

[0127] Item 55. The method of item 54, wherein reading comprises detecting a current change fingerprint.

[0128] Item 56. The method as in any one of items 53-55, wherein the current change includes one or more changes in signal spacing, signal amplitude, or a combination thereof.

[0129] Item 57. The method as in any one of items 1-56, wherein one or more nanopores are decorated with one or more molecules to attract a nucleic acid molecule.

[0130] Item 58. The method as in any one of items 1-57, comprising labeling or tagging at least a portion of the particles in the mixture with a bead.

[0131] Item 59. The method as in any one of items 14-58, comprising labeling each of a plurality of nucleic acid molecules with a helicase configured to act as a stopper preventing complete translocation of the nucleic acid molecule.

[0132] Item 60. The method as in any one of items 14-59, wherein translocating comprises repeatedly passing a nucleic acid molecule through the nanopore by iteratively changing nanopore voltage polarity.

[0133] Item 61. The method as in any one of items 1-60, comprising dynamically adjusting nanopore translocation rate based on particle detection rate.

[0134] Item 62. The method as in any one of items 1-61, comprising modifying at least a portion of one of the one or more labels with one or more fluorescent labels, reading one or more fluorescent signals from the one or more fluorescent labels and translating the one or more signals into the information, the reading comprising: (i) detecting a change in optical signal along a length of a particle, wherein a first optical signal corresponds to a first motif and a second optical signal corresponds to a second motif; (ii) identifying a first optical label from the first optical signal and a second optical label from the second optical signal; and (iii) identifying, from the library of nucleic acid molecules, the information based at least in part on the identified optical labels.

[0135] Item 63. A system for extracting information from a particle, the system comprising: one or more second reagents to label one or more particles; a fluidic device configured for transfer fluid, the fluid comprising one or more of the first or second reagents; and a processor and a memory functionally connected to the fluidic device, the memory comprising instructions that, when executed, cause the processor to actuate one or more components of the device to perform one or more method(s) as in any one of items 1-62.REFERENCES[1] J. A. Alfaro et al., “The emerging landscape of single-molecule protein sequencing technologies,” Nat Methods, vol. 18, no. 6, pp. 604-617, 2021, doi: 10.1038 / s41592-021- 01143-1.[2] T. E. Schaus, S. Woo, F. Xuan, X. Chen, and P. Yin, “A DNA nanoscope via autocycling proximity recording,” Nat Commun, vol. 8, no. 1, p. 696, 2017, doi: 10.1038 / s41467- 017-00542-3.[3] Y. L. Ying et al., “Nanopore-based technologies beyond DNA sequencing,” Nature Nanotechnology, vol. 17, no. 11. Nature Research, pp. 1136-1146, Nov. 01, 2022. doi:10.1038 / s41565-022-01193-2.[4] S. M. Bezrukov, I. Vodyanoy, and V. A. Parsegian, “Counting polymers moving through a single ion channel,” Nature, vol. 370, no. 6487, 1994, doi: 10.1038 / 370279a0.[5] C. Raillon, P. Granjon, M. Graf, L. J. Steinbock, and A. Radenovic, “Fast and automatic processing of multi-level events in nanopore translocation experiments,” Nanoscale, vol. 4, no. 16, 2012, doi: 10.1039 / c2nr30951c.[6] Y. K. Wan, C. Hendra, P. N. Pratanwanich, and J. Goke, “Beyond sequencing: machine learning algorithms extract biology hidden in Nanopore signal data,” Trends in Genetics, vol. 38, no. 3. 2022. doi: 10.1016 / j .tig.2021.09.001.[7] E. C. Yusko et al., “Real-time shape approximation and fingerprinting of single proteins using a nanopore,” Nat Nanotechnol, vol. 12, no. 4, 2017, doi:10.1038 / nnano.2016.267.[8] G. Huang et al., “Ply AB Nanopores Detect Single Amino Acid Differences in Folded Haemoglobin from Blood**,” Angewandte Chemie, vol. 134, no. 34, 2022, doi:10.1002 / ange.202206227.[9] S. Zhang et al., “A Nanopore-Based Saccharide Sensor,” Angewandte Chemie, vol. 134, no. 33, 2022, doi: 10.1002 / ange.202203769.

[0010] I. Soldanescu, A. Lobiuc, M. Covasa, and M. Dimian, “Detection of Biological Molecules Using Nanopore Sensing Techniques,” Biomedicines, vol. 11, no. 6. 2023. doi: 10.3390 / biomedicinesl 1061625.

[0011] N. A. W. Bell and U. F. Keyser, “Specific protein detection using designed DNA carriers and nanopores,” J Am Chem Soc, vol. 137, no. 5, 2015, doi: 10.1021 / ja512521w.

[0012] W. Shi, A. K. Friedman, and L. A. Baker, “Nanopore Sensing,” Anal Chem, vol. 89, no. 1, pp. 157-188, Nov. 2016, doi: 10.1021 / acs.analchem.6b04260.

[0013] S. K. Tabatabaei et al., “DNA punch cards for storing data on native DNA sequences via enzymatic nicking,” Nat Commun, vol. 11, no. 1, Dec. 2020, doi: 10.1038 / s41467-020- 15588-z.

[0014] F. Boskovic and U. F. Keyser, “Nanopore microscope identifies RNA isoforms with structural colours,” Nat Chem, vol. 14, no. 11, pp. 1258-1264, 2022, doi: 10.1038 / s41557- 022-01037-5.

[0015] L. He et al., “DNA origami characterized via a solid-state nanopore: insights into nanostructure dimensions, rigidity and yield,” Nanoscale, vol. 15, no. 34, 2023, doi: 10.1039 / d3nr01873c.

[0016] F. Praetorius and H. Dietz, “Self-assembly of genetically encoded DNA-protein hybrid nanoscale shapes,” Science (1979), vol. 355, no. 6331, Mar. 2017, doi:10.1126 / science.aam5488.

[0017] P. Shrestha et al., “Single-molecule mechanical fingerprinting with DNA nanoswitch calipers,” Nat Nanotechnol, vol. 16, no. 12, 2021, doi: 10.1038 / s41565-021-00979-0.

[0018] P. Shrestha, D. Yang, A. Ward, W. M. Shih, and W. P. Wong, “Mapping SingleMolecule Protein Complexes in 3D with DNA Nanoswitch Calipers,” J Am Chem Soc, vol. 145, no. 51, pp. 27916-27921, 2023, doi: 10.1021 / jacs.3cl0262.

[0019] A. D. Cotton, J. A. Wells, and I. B. Seiple, “Biotin as a Reactive Handle to Selectively Label Proteins and DNA with Small Molecules,” ACS Chem Biol, vol. 17, no. 12, 2022, doi: 10.1021 / acschembio.lc00252.

[0020] J. Zhu et al., “Multiplexed Nanopore-Based Nucleic Acid Sensing and Bacterial Identification Using DNA Dumbbell Nanoswitches,” J Am Chem Soc, Jun. 2023, doi: 10.1021 / jacs.3c01649.

[0021] M. A. Koussa, K. Halvorsen, A. Ward, and W. P. Wong, “DNA nanoswitches: A quantitative platform for gel-based biomolecular interaction analysis,” Nat Methods, vol. 12, no. 2, pp. 123-126, Jan. 2015, doi: 10.1038 / nmeth.3209.

[0022] C. Wen, D. Dematties, and S. L. Zhang, “A Guide to Signal Processing Algorithms for Nanopore Sensors,” ACS Sensors, vol. 6, no. 10. American Chemical Society, pp. 3536- 3555, Oct. 22, 2021. doi: 10.1021 / acssensors. lc01618.

[0023] Y. Li, S. E. Sandler, U. F. Keyser, and J. Zhu, “DNA Volume, Topology, and Flexibility Dictate Nanopore Current Signals,” Nano Lett, Jul. 2023, doi: 10.1021 / acs.nanolett.3c01823.

[0024] W. S. Alharbi and M. Rashid, “A review of deep learning applications in human genomics using next-generation sequencing data,” Human Genomics, vol. 16, no. 1. 2022. doi: 10.1186 / s40246-022-00396-x.

[0025] A. Yang, W. Zhang, J. Wang, K. Yang, Y. Han, and L. Zhang, “Review on the Application of Machine Learning Algorithms in the Sequence Data Mining of DNA,” Frontiers in Bioengineering and Biotechnology, vol. 8. 2020. doi: 10.3389 / fbioe.2020.01032.

[0026] S. Ilbeigipour, A. Albadvi, and E. Akhondzadeh Noughabi, “Real-Time Heart Arrhythmia Detection Using Apache Spark Structured Streaming,” J Healthc Eng, vol. 2021, 2021, doi: 10.1155 / 2021 / 6624829.

[0027] S. Ohayon, A. Girsault, M. Nasser, S. Shen-Orr, and A. Meller, “Simulation of singleprotein nanopore sensing shows feasibility for whole-proteome identification,” PLoS Comput Biol, vol. 15, no. 5, 2019, doi: 10.1371 / joumal.pcbi.1007067.

[0028] T. U. Consortium, “UniProt: the Universal Protein Knowledgebase in 2023,” Nucleic Acids Res, vol. 51, no. DI, pp. D523-D531, Jan. 2023, doi: 10.1093 / nar / gkacl052.

[0029] J. Swaminathan, A. A. Boulgakov, and E. M. Marcotte, “A Theoretical Justification for Single Molecule Peptide Sequencing,” PLoS Comput Biol, vol. 11, no. 2, 2015, doi: 10.1371 / journal.pcbi.1004080.

[0030] Y. Yao, M. Docter, J. Van Ginkel, D. De Ridder, and C. Joo, “Single-molecule protein sequencing through fingerprinting: Computational assessment,” Phys Biol, vol. 12, no. 5, 2015, doi: 10.1088 / 1478-3975 / 12 / 5 / 055003.

[0031] S. Lin et al., “Redox-based reagents for chemoselective methionine bioconjugation,” Science (1979), vol. 355, no. 6325, 2017, doi: 10.1126 / science.aal3316.

[0032] Y. Tabana, D. Babu, R. Fahlman, A. G. Siraki, and K. Barakat, “Target identification of small molecules: an overview of the current applications in drug discovery,” BMC Biotechnol, vol. 23, no. 1, p. 44, 2023, doi: 10.1186 / sl2896-023-00815-4.What is claimed is:

Claims

CLAIMS1. A method for extracting information from a particle, comprising:(a) providing a mixture comprising at least a first set of particles and a second set of particles;(b) modifying at least a portion of the first set of particles with one or more labels;(c) translocating one or more modified particles through one or more nanopores disposed in a substrate and configured to receive an input particle;(d) reading one or more signals received from the one or more nanopores and translating the one or more signals into the information, the reading comprising(i) detecting an electric current signal from the one or more nanopores;(ii) detecting a change in current through the one or more nanopores corresponding to the translocation, wherein a first current level corresponds to passage of at least an unlabeled portion of a particle and a second current level corresponds to passage of one of the one or more labels through the one or more nanopores;(iii) identifying the one label from the second voltage level; and(iv) identifying the information based at least in part on the identified label.

2. The method of claim 1, wherein the first set or the second set, or both, comprise a plurality of polymers.

3. The method as in any one of claims 1-2, wherein the first set or the second set, or both, comprise a plurality of small molecules.

4. The method as in any one of claims 1-3, wherein the first set comprises a plurality of polymers and the second set comprises a plurality of small molecules, or vice versa.

5. The method as in any one of claims 1-4, wherein the first set or the second set, or both, comprise a plurality of biological structures.

6. The method as in any one of claims 1-5, wherein the first set comprises a plurality of polymers and the second set comprises a plurality of biological structures, or vice versa.

7. The method as in any one of claims 1-5, wherein the first set comprises a plurality of small molecules and the second set comprises a plurality of biological structures, or vice versa.

8. The method of claim 1, wherein the first set is a plurality of first polymers and the second set is a plurality of second polymers.

9. The method of claim 1, wherein the first set is a plurality of first small molecules and the second set is a plurality of second small molecules.

10. The method of claim 1, wherein the first set is a plurality of first biological structure and the second set is a plurality of second biological structures.

11. The method of claim 1, wherein the first set is a plurality of polymers, and the second set is a plurality of small molecules, or vice versa.

12. The method of claim 1, wherein the first set is a plurality of polymers, and the second set is a plurality of biological structures, or vice versa.

13. The method of claim 1, wherein the first set is a plurality of small molecules, and the second set is a plurality of biological structures, or vice versa.

14. The method as in any one of claims 2-13, wherein the polymers comprise proteins, lipids, complex carbohydrates, nucleic acid molecules, or a combination thereof.

15. The method as in any one of claims 3-14, wherein the small molecules comprise small molecule therapeutic agents, small molecule diagnostic agents, biological metabolites, or a combination thereof.

16. The method as in any one of claims 4-15, wherein the biological structures comprise viral particles, cellular fragments, cellular particles, 3D protein structures, 3D nucleic acid structures, liposomes, or a combination thereof.

17. The method as in any one of claims 1-16, modifying at least a portion of the second set of particles with one or more labels.

18. The method of claim 17, wherein the labeled particles of the first set, the second set, or both are labeled with at least two labels.

19. The method as in any one of claims 1-18, wherein the first set is labeled with a first label producing a first signal and a second label producing a second signal.

20. The method of claim 19, wherein the first signal is distinct from the second signal.

21. The method of claim 19, wherein the first signal is substantially the same as the second signal.

22. The method as in any one of claims 1-18, wherein the first set is labeled with a first label producing a first signal and the second set is labeled with a second label producing a second signal.

23. The method of claim 22, wherein the first signal is distinct from the second signal.

24. The method of claim 22, wherein the first signal is substantially the same as the second signal.

25. The method as in any one of claims 1-24, comprising modifying at least a portion of a third set of particles with one or more labels.

26. The method of claim 25, wherein substantially all particles of the third set are labeled.

27. The method as in any one of claims 5-26, wherein the third set is labeled with a third label producing a third signal, the third signal distinct from the first signal and the second signal.

28. The method as in any one of claims 25-26, wherein the third set is labeled with a third label producing a third signal, the third signal substantially the same as the first signal or the second signal.

29. The method as in any one of claims 1-28, wherein the one or more labels are or comprise biotin-streptavidin, DNA structures, protein labels, or a combination thereof.

30. The method of claim 29, wherein the one or more labels comprise an enzyme label.

31. The method of claim 29, wherein the protein is an endonuclease.

32. The method as in any one of claims 29-31, wherein the protein is modified by phosphorylation, acetylation, hydroxylation, ubiquitination, methylation, or cross-linking.

33. The method as in any one of claims 29-32, wherein the one or more labels comprise two or more proteins.

34. The method as in any one of claims 1-33, wherein the one or more labels comprise an indirect label on a DNA flap.

35. The method as in any one of claims 1-34, wherein the one or more labels are or comprise one or more DNA labels.

36. The method of claim 35, wherein the one or more one or more DNA labels comprise a barcode.

37. The method of claim 35, wherein the one or more one or more DNA labels comprise a hairpin.

38. The method of claim 37, wherein the one or more hairpins have a length of 10-500 bp.

39. The method as in any one of claims 37-38, wherein two or more hairpins have different hairpin stem length.

40. The method as in any one of claims 37-39, wherein two or more hairpins have different numbers of hairpin stems41. The method as in any one of claims 37-40, wherein two or more hairpins have different hairpin loop sizes42. The method as in any one of claims 37-40, wherein two or more hairpins incorporate one or more unnatural bases or methylated bases.

43. The method as in any one of claims37-42, wherein the hairpins are chemically modified.

44. The method as in any one of claims 37-43, wherein the one or more hairpins are located on one strand of a double stranded nucleic acid molecule.

45. The method as in any one of claims 1-44, wherein the one or more labels comprise an expandomer.

46. The method as in any one of claims 1-45, wherein the one or more labels comprise a DNA dumbbell.

47. The method as in any one of claims 1-46, comprising reading a plurality of nucleic acid molecules and generating, from the plurality of reads, a consensus read.

48. The method of claim 47, comprising circularizing one or more nucleic acid molecules and performing rolling circle amplification.

49. The method as in any one of claims 1-48, wherein the one or more nanopores are or comprise a Nanopore Field Effect Transistor (NPFET)50. The method as in any one of claims 1-49, wherein the one or more nanopores are integrated into a complementary metal-oxide-semiconductor (CMOS) circuit.

51. The method as in any one of claims 1-50, wherein the one or more nanopores have a diameter of 10-15 nm.

52. The method as in any one of claims 1-51, wherein the one or more nanopores have a diameter of less than 10 nm.

53. The method as in any one of claims 1-52, wherein a first label causes a first current change in the one or more nanopores, and a second label causes a second current change.

54. The method of claim 53, wherein the first current change is different from the second current change.

55. The method of claim 54, wherein reading comprises detecting a current change fingerprint.

56. The method as in any one of claims 53-55, wherein the current change includes one or more changes in signal spacing, signal amplitude, or a combination thereof.

57. The method as in any one of claims 1-56, wherein one or more nanopores are decorated with one or more molecules to attract a nucleic acid molecule.

58. The method as in any one of claims 1-57, comprising labeling or tagging at least a portion of the particles in the mixture with a bead.

59. The method as in any one of claims 14-58, comprising labeling each of a plurality of nucleic acid molecules with a helicase configured to act as a stopper preventing complete translocation of the nucleic acid molecule.

60. The method as in any one of claims 14-59, wherein translocating comprises repeatedly passing a nucleic acid molecule through the nanopore by iteratively changing nanopore voltage polarity.

61. The method as in any one of claims 1-60, comprising dynamically adjusting nanopore translocation rate based on particle detection rate.

62. The method as in any one of claims 1-61, comprising modifying at least a portion of one of the one or more labels with one or more fluorescent labels,reading one or more fluorescent signals from the one or more fluorescent labels and translating the one or more signals into the information, the reading comprising(i) detecting a change in optical signal along a length of a particle, wherein a first optical signal corresponds to a first motif and a second optical signal corresponds to a second motif;(ii) identifying a first optical label from the first optical signal and a second optical label from the second optical signal; and(iii) identifying, from the library of nucleic acid molecules, the information based at least in part on the identified optical labels.

63. A system for extracting information from a particle, the system comprising: one or more second reagents to label one or more particles; a fluidic device configured for transfer fluid, the fluid comprising one or more of the first or second reagents; and a processor and a memory functionally connected to the fluidic device, the memory comprising instructions that, when executed, cause the processor to actuate one or more components of the device to perform one or more method(s) as in any one of claims 1-62.

Citation Information

Patent Citations

  • Method for characterization of nucleic acid molecules

    US20030104428A1

  • Target Detection with Nanopore

    US20160258939A1

  • Mutant pore

    US20190071721A1