Single-molecule identification with reactive heterogeneous nanopores
Patent Information
- Application Number
- JP2024521196
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-07-08
- Filing Date
- 2022-10-09
- Publication Date
- 2025-09-19
AI Technical Summary
Existing methods for analyzing sugars, RNA modifications, and amino acids are incomplete, expensive, time-consuming, and unable to provide stereochemical information or distinguish isomers, and there is a lack of nanopore methods that can simultaneously identify all natural amino acids and their post-translational modifications.
A protein nanopore with reactive amino acid residues coupled to metal ions through ligands, such as nitrilotriacetic acid, is used to interact with target analytes, allowing for characterization based on ionic current patterns.
The method enables efficient and accurate identification of various sugars, RNA modifications, and amino acids, including distinguishing isomers and post-translational modifications, using machine learning to analyze current patterns through nanopores.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to systems and methods for identifying analytes using nanopores. [Background technology]
[0002] While studies of sugar sequence or structure are known using (micro)arrays, capillary electrophoresis (CE), liquid chromatography (LC), nuclear magnetic resonance (NMR), or mass spectrometry (MS), characterization by any single method provides only an incomplete picture of the glycan analyte. Specifically, MS cannot provide stereochemical information of monosaccharides, nor can it distinguish between isomers. Characterization of sugars by these methods is typically expensive and time-consuming.
[0003] Analysis of RNA modifications can be performed by thin-layer chromatography (TLC), high-performance liquid chromatography-ultraviolet spectrophotometry (HPLC-UV), or high-performance liquid chromatography-mass spectrometry (HPLC-MS). While these methods can simultaneously measure a large number of RNA modifications, they cannot provide sequence information. Strand sequencing strategies are limited by spatial resolution equivalent to an average read of approximately five nucleotides, which still limits the ability to distinguish between all epigenetic modifications by sequencing. This situation becomes even more severe when modified nucleotides are close together. The analysis and detection of alditols is necessary in the medical and food industries, but the similarity of their chemical structures poses significant technical challenges in the design of sensing strategies.
[0004] The analysis and detection of natural amino acids through nanopores is important for achieving nanopore sequencing of peptides or proteins, but no nanopore method is yet available that can simultaneously distinguish all 20 natural amino acids and their post-translational chemical modifications. New detection methods are still needed. Summary of the Invention
[0005] A first aspect of the present invention provides a protein nanopore comprising at least one sensing moiety, wherein the sensing moiety couples the metal ion to a reactive amino acid residue within the nanopore channel and is capable of interacting with a target analyte.
[0006] In some embodiments, the metal ion is coupled to the reactive amino acid residue via a ligand, and the metal ion and the ligand form a coordination complex. In some embodiments, the ligand is nitrilotriacetic acid (NTA).
[0007] In some embodiments, the metal ion is Ni 2+ , Cu 2+ , Co 2+ , Zn 2+ , Cd 2+ , Ag 2+ , Pb 2+ , Fe 2+ , or Fe 3+ Selected from. In some embodiments, the reactive amino acid residue is selected from cysteine, methionine, and lysine.
[0008] In some embodiments, the protein nanopore is a heterogeneous protein nanopore, wherein one or more, but not all, of the monomers comprise the sensing moiety and other monomers do not comprise the sensing moiety.
[0009] In some embodiments, the heterogeneous protein nanopore is a mutant of a nanopore, wherein the nanopore is selected from MspA, α-HL, Aerolysin, ClyA, FhuA, FraC, PlyA / B, CsgG, and Phi 29 linker. In some embodiments, the heterogeneous protein nanopore is a variant of MspA.
[0010] In some embodiments, the protein nanopore comprises a Ni nanopore coupled to the reactive amino acid residue via a ligand. 2+ It is a heterogeneous MspA nanopore containing In some embodiments, Ni 2+ is coupled to the reactive amino acid residue via NTA. In some embodiments, the reactive amino acid residue is located at a position selected from 83 to 111, preferably 90, 91, 92, and 93.
[0011] In some embodiments, the heterogeneous protein nanopore has an N90C, N90M, or N91C mutation in one or more monomers compared to M2 MspA.
[0012] A second aspect of the present invention provides a protein nanopore comprising at least one sensing module, wherein the protein nanopore is heterogeneous MspA, and one or more, but not all, of the monomers comprise the sensing module and other monomers do not comprise the sensing module, and wherein the sensing module is capable of interacting with a target analyte.
[0013] In some embodiments, the sensing module consists of one or more reactive amino acid residues contained in one or more monomers of the heterogeneous MspA. In some embodiments, the reactive amino acid residue is selected from methionine, histidine, cysteine, or lysine, or a combination thereof.
[0014] In some embodiments, the sensing module consists of one or more sensing moieties that couple to one or more reactive amino acid residues contained in one or more monomers of the heterogeneous protein nanopore, and other monomers of the heterogeneous protein nanopore do not contain reactive amino acid residues. In some embodiments, the reactive amino acid residue is selected from cysteine, methionine, and lysine. In some embodiments, the sensing moiety is a boronic acid-containing moiety. In some embodiments, the boronic acid-containing moiety is phenylboronic acid (PBA).
[0015] In some embodiments, the reactive amino acid residues are located at one or more positions selected from 83 to 111, preferably 90, 91, 92 and / or 93.
[0016] In some embodiments, the heterogeneous protein nanopore has N90C, N90M and / or N91C mutations in one or more monomers compared to M2 MspA. A third aspect of the present invention provides a method for characterizing a target analyte, the method comprising: (i) providing any one of the protein nanopores described above; (ii) applying a voltage across the protein nanopore reactor; (iii) allowing the target analyte to pass through the nanopore; (iv) measuring the ionic current passing through the nanopore to provide a current pattern, and characterizing the target analyte based on the current pattern.
[0017] In some embodiments of the third aspect, the target analyte is in a sample, and step (iii) comprises allowing the sample to pass through the nanopore. In some embodiments of the third aspect, the sample is selected from fruit juices, beverages, teas, and herbal extracts. A fourth aspect of the present invention provides the use of any one of the above protein nanopores in the characterisation of a target analyte. In some embodiments of the fourth aspect, the target analyte is in a sample. In some embodiments of the fourth aspect, the sample is selected from fruit juices, beverages, teas, and herbal extracts.
[0018] In some embodiments of the third or fourth aspect, the target analyte may interact with a boronic acid, a metal ion, methionine, histidine, cysteine, lysine, or any combination thereof.
[0019] In some embodiments of the third or fourth aspect, the analyte capable of interacting with the boronic acid is selected from a compound containing a 1,2-diol or a 1,3-diol, an ion containing a metal element, hydrogen peroxide, and any combination thereof; The analyte capable of interacting with a metal ion is a molecule capable of interacting with a metal ion through coordination, and Analytes that can interact with methionine, histidine, cysteine, or lysine are ions that contain metal elements.
[0020] In some embodiments of the third or fourth aspect, The ions containing a metal element are selected from alkaline earth metal ions, transition metal ions, and any combination thereof, and are preferably AuCl4 - , Mg 2+ , Ca 2+ , Ba 2+ , Ni 2+ , Cu 2+ , Co 2+ , Zn 2+ , Cd 2+ , Ag 2+ , Pb 2+ and any combination thereof.
[0021] In some embodiments of the third or fourth aspect, the compound containing 1,2-diol or 1,3-diol is selected from the group consisting of sugars or derivatives thereof, α-hydroxy acids, ribose, nucleoside sugars, alditols, polyphenols, compounds containing catecholamines or catecholamine derivatives, tris(hydroxymethyl)methylaminomethane (Tris), protocatechuic aldehyde, protocatechuic acid, caffeic acid, rosmarinic acid, lithospermic acid, tansinol A, salvianolic acid B, and any combination thereof; In some embodiments of the third or fourth aspect, the sugar is selected from monosaccharides, oligosaccharides, polysaccharides, and any combination thereof; the sugar derivative is selected from N-acetylneuraminic acid (sialic acid), N-acetyl-D-galactosamine, and any combination thereof; the alpha-hydroxy acid is selected from tartaric acid, malic acid, citric acid, isocitric acid, and any combination thereof;
[0022] the ribose-containing compound is selected from a nucleotide or modified nucleotide, a derivative of a nucleotide or modified nucleotide, a nucleoside or nucleoside analog, and any combination thereof;
[0023] the nucleotide sugar is selected from uridine diphosphate glucose (UDPG), uridine diphosphate N-acetylglucosamine, uridine diphosphate glucuronic acid, adenosine diphosphate glucose, uridine diphosphate galactose, uridine diphosphate xylose, guanosine diphosphate mannose, guanosine diphosphate fucose, cytidine monophosphate N-acetylneuraminic acid, uridine diphosphate N-acetylgalactosamine, and any combination thereof;
[0024] the alditol is selected from glycerin, propanetriol, butitol, pentitol, hexitol, erythritol, threitol, arabitol, xylitol, ribitol (adonitol), fucitol, sorbitol such as L-sorbitol or D-sorbitol, mannitol, galactitol, iditol, talitol, allitol (allodulcitol), maltitol, lactitol, isomalt, and any combination thereof;
[0025] The polyphenol is selected from catechin, neochlorogenic acid, anthocyanin, proanthocyanidin, catechol or a derivative thereof, such as catechol, 3-fluorocatechol, 3-chlorocatechol, 3-bromocatechol, 4-fluorocatechol, 4-chlorocatechol, 4-bromocatechol, 3-methylcatechol, 4-methylcatechol, 3-methoxycatechol, 3-propylcatechol, 3-isopropylcatechol, 3,6-dibromocatechol, 4,5-dibromocatechol, 3,6-dichlorocatechol, and any combination thereof; and The catecholamine or catecholamine derivative is selected from epinephrine, norepinephrine, isoproterenol, and any combination thereof.
[0026] In some embodiments of the third or fourth aspect, the monosaccharide is selected from D-glyceraldehyde, D-erythrose, D-ribose, 2'-deoxy-D-ribose, D-xylose, L-arabinose, D-lyxose, D-glucose, D-galactose, D-mannose, D-fructose, L-sorbose, L-fucose, D-allose, D-tagatose, L-rhamnose, D-galactose, and any combination thereof;
[0027] the oligosaccharide is selected from disaccharides (e.g., sucrose, isomaltulose, maltulose, turanose, leucrose, trehalose, lactulose, maltose, etc.), trisaccharides (e.g., raffinose), tetrasaccharides (e.g., stachyose), and complex oligosaccharides (e.g., acarbose), and any combination thereof; The polysaccharides are selected from five sugars, including verbascose. the nucleotides are selected from adenine nucleotides, cytosine nucleotides, uracil nucleotides, guanine nucleotides, and any combination thereof;
[0028] The modified nucleotide is 5-methylcytidine (m 5 C), N6-methyladenosine (m 6 A), pseudouridine (Ψ), inosine (I), N7-methylguanosine (m 7 G), N1-methyladenosine (m 1 A), dihydrouridine (D), N2-methylguanosine (m 2 G), N2, N2-dimethylguanosine
[0029]
number
[0030] The derivative of a nucleotide or modified nucleotide is selected from monophosphate, diphosphate, triphosphate and tetraphosphate derivatives of a nucleotide or modified nucleotide, and any combination thereof, such as ADP, UDP, GDP, CDP, ATP, UTP, GTP, CTP and any combination thereof; and The nucleoside analog is selected from galidesvir, ribavirin, molnupiravir, remdesivir, loxoribine, mizoribine, 5-azacytidine, capecitabine, doxifluridine, 5-fluorouridine, forodesine, kleitosine, pyrazofurin, sangivamycin, pseudouridimycin, and any combination thereof.
[0031] In some embodiments of the third or fourth aspect, Molecules that can interact with metal ions through coordination contain nitrogen, oxygen, sulfur, phosphorus, or carbon atoms that can coordinate with the metal ion.
[0032] In some embodiments of the third or fourth aspect, Molecules that can interact with metal ions through coordination are compounds containing at least one carboxylic acid group or at least one amine group, amino acids, modified amino acids, polymers of amino acids or modified amino acids, compounds containing guanine, adenine, thymine, cytosine or uracil, and any combination thereof.
[0033] In some embodiments of the third or fourth aspect, the amino acid is selected from alanine, cysteine, aspartic acid, glutamic acid, phenylalanine, glycine, histidine, isoleucine, lysine, leucine, methionine, asparagine, proline, glutamine, arginine, serine, threonine, valine, tryptophan, tyrosine, pyrolysine, selenocysteine, and any combination thereof;
[0034] The modified amino acids are selected from phosphorylated amino acids, glycosylated amino acids, acetylated amino acids, methylated amino acids, and any combination thereof, such as O-phosphoserine (pS), N4-(β-N-acetyl-D-glucoseamino)-asparagine (GlcNAc-N), O-acetyl-threonine (Ac-T), Nω,N'ω-dimethyl-arginine (SDMA), and any combination thereof; and
[0035] A compound containing guanine, adenine, thymine, cytosine, or uracil is selected from guanine, adenine, thymine, cytosine, or uracil, or contains any one nucleoside selected from them, or contains any one nucleotide selected from them, and the nucleotide is a ribonucleotide or a deoxyribonucleotide. [Brief explanation of the drawings]
[0036] [Figure 1]This figure shows the preparation of boronated MspA for sugar sensing. (a) The structure of (N90C)1(M2)7. (N90C)1(M2)7 is a heterogeneous MspA octamer consisting of seven units of M2 MspA-D16H6 (gray) and one unit of N90C MspA-H6 (red). The only sulfhydryl group is present in the pore cavity of (N90C)1(M2)7. (b) Gel electrophoresis results showing different types of heterogeneous MspA octamer assembly. Gel electrophoresis was performed on a 10% SDS-PAGE gel. The left lane is the homooctamer M2 MspA-D16H6. The right lane is the homooctamer N90C MspA-H6. The middle lane shows all possible combinations of heterogeneously assembled MspA octamer structures when M2 MspA-D16H6 and N90C MspA-H6 are coexpressed and assembled. The schematic diagram on the right represents the corresponding heterogeneous octameric MspA types. Gray dots represent M2 MspA-D16H6 monomers, and red dots represent N90C MspA-H6. (N90C)1(M2)7, the desired MspA heterogeneous octamer, is shown as the second band in the middle lane (red rectangle). (c) Single molecule display of MPBA [3-(maleimido)phenylboronic acid] functionalization of (N90C)1(M2)7. Single-channel recordings were performed using a single (N90C)1(M2)7 MspA pore, as described in Methods in Example 1. A bias voltage of +100 mV was applied continuously. Additional shot noise was also observed prior to MPBA binding. Upon binding of MPBA to the pore, an irreversible current drop was observed, measured at approximately 53 pA. I0 represents the pore-gating current of (N90C)1(M2)7, and Ip represents the pore-gating current of MPBA-modified (N90C)1(M2)7. (d) shows the mechanism of sugar sensing. L-sorbose is used as a representative sugar, and the reversible binding / dissociation of L-sorbose and phenylboronic acid is the basis for sensing. (e) shows a representative trace of L-sorbose sensing. A bias voltage of +160 mV was continuously applied, and current traces were obtained upon addition of L-sorbose to the cis chamber (final concentration 10 mM). A buffer solution of 1.5 M KCl, 10 mM MOPS, pH 7.0 was used. Is represents the blocking level upon sugar binding to the pore.Upon sensing L-sorbose, a current drop of approximately 102 pA is monitored. (f) A representative event for L-sorbose is shown. The corresponding full-point histogram of the blockade level is located on the right. (g) A scatter plot of ΔI versus the standard deviation (SD) when L-sorbose was sensed as the sole analyte. ΔI = Ip - Is. This figure contains 910 events (n = 910). The scatter plot is color-coded according to the local event density around each data point (ggplot 2, R). (h) A plot of the reciprocal of the mean interevent interval (1 / τon) and mean residence time (1 / τoff) of L-sorbose sensing events versus the L-sorbose concentration in the cis chamber. Three independent measurements were performed for each condition (N = 3). [Figure 2] Single-molecule identification of D-fructose, D-galactose, D-mannose, and D-glucose is shown. Measurements were performed in a 1.5 M KCl, 10 mM MOPS, pH 7.0 buffer solution, as described in Example 1. Different sugars were added to the cis chamber to achieve the desired final concentration. (a, d, g, j) are the chemical structures of D-fructose (Fru, a), D-galactose (Gal, d), D-mannose (Man, g), or D-glucose (Glc, j). The molecular weight (MW) of all mentioned monosaccharides is 180.16. (b, e, h, k) are representative event types for D-fructose (20 mM, b), D-galactose (20 mM, e), D-mannose (20 mM, h), or D-glucose (60 mM, k). All mentioned monosaccharides reported more than one type of event, each represented by a Roman numeral. (c,f,i,l) Scatter plots of the SD of ΔI / Ip for different monosaccharides. To aid in the identification of different event populations, the scatter plots are colored according to the local event density around each data point (ggplot 2, R). Different event types are displayed on the scatter plots, as are those labeled with Roman numerals, and correspond to those shown in (b,e,h,k). [Figure 3]Distinguishing five types of monosaccharides using machine learning. (a) shows the machine learning workflow. Five classes of events are collected from the detection of D-fructose, D-galactose, D-mannose, D-glucose, or L-sorbose, respectively, to form a database. Nine event features are extracted to form a feature matrix. 1,000 samples are then randomly selected from each class to form a labeled dataset. The labeled dataset is randomly divided into a training set (80%) and a test set (20%) for model training and testing. A 10-fold cross-validation is performed to evaluate the model. Validation accuracy is defined as the proportion of correctly identified events across the entire dataset. Six models are employed, and the random forest model reports the highest accuracy score of 0.974. The random forest is then further tuned, resulting in an improved accuracy score of 0.975. (b) shows the feature importance of the trained random forest model. (c) shows the confusion matrix generated by the test set. (d) shows the learning curves with different training set sizes. When the training set size exceeds 508, the validation accuracy reaches 0.95. (e-f) are representative traces containing a sugar-sensing event mixture. The machine learning algorithm automatically predicts the events in the trace as D-fructose (green pentagons), D-galactose (yellow circles), D-mannose (green circles), D-glucose (blue circles), and L-sorbose (orange pentagons). Measurements are performed in a 1.5 M KCl, 10 mM MOPS, pH 7.0 buffer solution, as described in the method in Example 1. D-fructose (5 mM), D-galactose (10 mM), D-mannose (10 mM), D-glucose (60 mM), and L-sorbose (1 mM) are added to the cis chamber to form a mixture. [Figure 4]Single-molecule identification of D-ribose, D-xylose, L-rhamnose, and N-acetyl-D-galactosamine is shown. Measurements were performed in 1.5 M KCl, 10 mM MOPS, pH 7.0 buffer, as described in the method in Example 1. Different sugars were added to the cis chamber to achieve the desired final concentrations. (a, d, g, j) are the chemical structures of D-ribose (Rib, a), D-xylose (Xyl, d), L-rhamnose (L-Rha, g), and N-acetyl-D-galactosamine (GalNAc, j). (b, e, h, k) are representative event types for D-ribose (20 mM, b), D-xylose (20 mM, e), L-rhamnose (40 mM, h), or N-acetyl-D-galactosamine (20 mM, k). All mentioned monosaccharides report more than one type of event, each represented by a Roman numeral. (c, f, i, l) Scatter plots of the ΔI / Ip of different monosaccharides against the SD. To aid in the identification of different event populations, the scatter plots are color-coded according to the local event density around each data point (ggplot 2, R). Each different event type is represented on the scatter plot, as are those labeled with a Roman numeral, consistent with those shown in (b, e, h, k). [Figure 5]Distinguishing nine types of monosaccharides using machine learning. (a) Schematic diagram. Nine types of monosaccharides are tested with MspA-PBA. (b) shows the confusion matrix generated by the test set. (c) is a scatter plot of the results obtained from a mixture of nine monosaccharides. Machine learning predicts tags for each sugar type. (d-f) are representative traces containing events when nine sugars are detected in the mixture. The events in the traces are automatically predicted by machine learning to be D-fructose (green pentagon), D-galactose (yellow circle), D-mannose (green circle), D-glucose (blue circle), L-sorbose (orange pentagon), D-ribose (pink star), D-xylose (orange star), L-rhamnose (green triangle), and N-acetyl-D-galactosamine (yellow square). Measurements are performed with MspA-PBA in 1.5 M KCl, 10 mM MOPS, pH 7.0, as described in the method of Example 1. D-fructose (2.5 mM), D-galactose (5 mM), D-mannose (5 mM), D-glucose (30 mM), L-sorbose (0.5 mM), D-ribose (5 mM), D-xylose (5 mM), L-rhamnose (10 mM), and N-acetyl-D-galactosamine (2.5 mM) are added simultaneously to the cis chamber to form a mixture. [Figure 6] Figure 1 shows the co-expression vector map. The two target genes, N90C MspA-H6 and M2 MspA-D16H6, were custom synthesized by GenScript Biotech Corporation (New Jersey) and simultaneously constructed in the co-expression vector pETDuet-1. Specifically, the gene encoding N90C MspA-H6 was constructed in the first multiple cloning site between the NcoI and HindIII restriction sites. The gene encoding M2 MspA-D16H6 was constructed in the second multiple cloning site between the NdeI and BlpI restriction sites. [Figure 7]This figure shows the production of heterogeneous MspA. (a) is a schematic diagram displaying the heterogeneous MspA octamer. The (N90C)1(M2)7 octamer assembly contains a single cysteine within the pore cavity and is the desired heterogeneous MspA type. (b) is the UV absorption spectrum during column elution. The labeled fractions were further characterized by gel electrophoresis. (c, d) are the results of gel electrophoresis of different elution fractions. Gel electrophoresis was performed on a 4% to 15% gradient SDS-polyacrylamide gel. Lanes M, Precision Plus Protein Standard (Bio-Rad), 1, bacterial lysate supernatant, 2, eluate collected from the nickel affinity column immediately after loading the bacterial lysate supernatant, 3, 5–8, and 10–16, corresponding elution fractions described in (b), and 4 and 9, purified octameric M2 MspA as the standard. According to the electrophoresis results, fractions 10–16 contain the major heterogeneous MspA. These fractions were collected for further characterization and purification (Figure 1b). [Figure 8]Single-molecule characterization of (N90C)1(M2)7 before and after chemical modification. (a) Reaction mechanism for (N90C)1(M2)7 modification. Briefly, the cysteine thiol at site 90 of the pore cavity irreversibly reacts with 3-(maleimido)phenylboronic acid (MPBA) via Michael addition. 1,2 (b) (N90C)1(M2)7 spontaneously inserts into the membrane. (c) MPBA-modified (N90C)1(M2)7 (MspA-PBA) spontaneously inserts into the membrane. (d) Pore-gating current histograms obtained for (N90C)1(M2)7 (red) and MspA-PBA (black). (b-d) Measurements were performed in 1.5 M KCl buffer (1.5 M KCl, 10 mM MOPS, pH 7.0) by applying a continuous voltage of +100 mV. The histograms are overlaid with the corresponding Gaussian fit results. Under these conditions (N = 58), modification with a single MPBA reduces the pore-opening current from I0 (295.23 ± 0.32 pA) to Ip (241.95 ± 0.10 pA). (e) Current-voltage (IV) curves (N = 3) of (N90C)1(M2)7 MspA (red) and (N90C)1(M2)7 MspA-MPBA (black). The IV curves were obtained by applying a gradient voltage between -150 mV and +150 mV in a 1.5 M KCl and 10 mM MOPS buffer at pH 7.0. (N=3 for each condition) [Figure 9] This shows the concentration dependence of L-sorbose sensing. Single-channel recordings were performed using MspA-PBA in a buffer solution of 1.5 M KCl, 10 mM MOPS, pH 7.0. A bias voltage of +160 mV was continuously applied. L-sorbose was added to the cis chamber until the final concentration was between 0 mM and 2 mM. Typically, no sensing events were observed in the absence of L-sorbose. The frequency of events was proportional to the final concentration of L-sorbose. [Figure 10]Definitions of event parameters are shown. A representative trace containing consecutive L-sorbose sensing events is shown. Traces were obtained using MspA-PBA as shown in Figure 1. The pore opening current (Ip), cutoff level (Is), dwell time (toff), inter-event duration (toff), mean value, and standard deviation (SD) are labeled on the trace. The mean dwell time (τoff) or mean inter-event interval (τon) was obtained by fitting the histogram of toff or toff, respectively, with an exponential function. [Figure 11] Figure 1 shows L-sorbose sensing by octameric M2 MspA-D16H6. L-sorbose sensing is performed using octameric M2 MspA-D16H6 in a buffer of 1.5 M KCl, 10 mM MOPS, pH 7.0. A bias voltage of +160 mV is continuously applied. L-sorbose is added to the cis chamber to a final concentration of 50 mM, preventing any L-sorbose sensing events. [Figure 12] L-sorbose sensing at different voltages. (a) Graph of the inverse of the mean inter-event interval (1 / τon) and the inverse of the mean residence time (1 / τoff) when used for sensing L-sorbose at different voltages. (b) Graph of the mean blockage depth when used for sensing L-sorbose at different voltages. The concentration of L-sorbose in the cis chamber is set to 5 mM. All results are from measurements performed using MspA-PBA in a buffer of 1.5 M KCl, 10 mM MOPS, pH 7.0. [Figure 13] Event scatter plots against the SD of ΔI / Ip for L-sorbose. (a) Chemical structure of L-sorbose. (b-d) Scatter plots against the SD of ΔI / Ip for L-sorbose. Events in each figure were obtained as described in Figure 1 and are from a 50-minute continuous recording trace. The results of nanopore sensing of L-sorbose are shown as individual event populations in the scatter plots. Event populations are labeled with Roman numerals. For display purposes, each scatter point is color-coded based on the local density around each point. Scatter plots were generated using the ggplot2 package in R. [Figure 14] Different types of D-fructose events are shown. Measurements are performed using MspA-PBA (method in Example 1). A bias voltage of +160 mV is continuously applied. D-fructose is added to the cis chamber to a final concentration of 20 mM. (a) is a representative trace containing D-fructose-sensing events. Different types of D-fructose events are labeled with Roman numerals I-IV. Repeated occurrences of the same type of event are labeled with Arabic numerals immediately following the Roman numeral. (b-f) are enlarged displays of each type of event. [Figure 15] Event scatter plots of D-fructose versus the SD of ΔI / Ip. (a) Chemical structure of D-fructose. (b-d) Scatter plots of D-fructose versus the SD of ΔI / Ip. Events in each plot are from a 30-minute continuous recording trace obtained as described in Figure 14. Nanopore sensing of D-fructose results in different types of sensing events, shown as distinct event populations in the scatter plots. The different event populations are labeled with Roman numerals and correspond to those defined in Figure 14. For display purposes, each scatter point is color-coded based on the local density around each point. Scatter plots were generated using the ggplot2 package in R. [Figure 16] Different types of D-galactose events are shown. Measurements are performed using MspA-PBA (method in Example 1). A bias voltage of +160 mV is continuously applied. D-galactose is added to the cis chamber to a final concentration of 20 mM. (a) is a representative trace containing D-galactose sensing events. Different types of D-galactose events are labeled with Roman numerals I-IV. Repeated occurrences of the same type of event are labeled with Arabic numerals immediately following the Roman numeral. (b-f) are enlarged displays of each type of event. [Figure 17]Event scatter plots against the SD of ΔI / Ip for D-galactose. (a) Chemical structure of D-galactose. (b-d) Scatter plots against the SD of ΔI / Ip for D-galactose. Events in each figure are from a 30-minute continuous recording trace obtained as described in Figure 16. Nanopore sensing of D-galactose results in different types of sensing events, shown as distinct event populations in the scatter plots. The different event populations are labeled with Roman numerals and correspond to those defined in Figure 16. For display purposes, each scatter point is color-coded based on the local density around each point. Scatter plots were generated using the ggplot2 package in R. [Figure 18] Different types of D-mannose events are shown. Measurements are performed using MspA-PBA (method in Example 1). A bias voltage of +160 mV is continuously applied. D-mannose is added to the cis chamber to a final concentration of 20 mM. (a) is a representative trace containing D-mannose-sensing events. Different types of D-mannose events are labeled with Roman numerals I-III. Repeated occurrences of the same type of event are labeled with Arabic numerals immediately following the Roman numeral. (b-f) are enlarged displays of each type of event. [Figure 19] Event scatter plots against the SD of ΔI / Ip for D-mannose. (a) Chemical structure of D-mannose. (b-d) Scatter plots against the SD of ΔI / Ip for D-mannose. Events in each figure were obtained as described in Figure 18 and are from a 30-minute continuous recording trace. Nanopore sensing of D-mannose results in different types of sensing events, which are shown as distinct event populations in the scatter plots. The different event populations are labeled with Roman numerals and correspond to those defined in Figure 18. For display purposes, each scatter point is color-coded based on the local density around each point. Scatter plots were generated using the ggplot2 package in R. [Figure 20]Different types of D-glucose events are shown. Measurements are performed using MspA-PBA (method in Example 1). A bias voltage of +160 mV is continuously applied. D-glucose is added to the cis chamber to a final concentration of 60 mM. (a) is a representative trace containing D-glucose sensing events. Different types of D-glucose events are labeled with Roman numerals I-IV. Repeated occurrences of the same type of event are labeled with Arabic numerals immediately following the Roman numeral. (b-f) are enlarged displays of each type of event. [Figure 21] Event scatter plots against the SD of ΔI / Ip for D-glucose. (a) Chemical structure of D-glucose. (b-d) Scatter plots against the SD of ΔI / Ip for D-glucose. Events in each plot are from a 50-minute continuous recording trace obtained as described in Figure 20. Nanopore sensing of D-glucose results in different types of sensing events, shown as distinct event populations in the scatter plots. The different event populations are labeled with Roman numerals and correspond to those defined in Figure 20. For display purposes, each scatter point is color-coded based on the local density around each point. Scatter plots were generated using the ggplot2 package in R. [Figure 22] Labeling of monosaccharide-sensing events by machine learning. (a) Scatter plot of ΔI / Ip against SD for events obtained from a mixture of D-fructose (5 mM), D-galactose (10 mM), D-mannose (10 mM), D-glucose (60 mM), and L-sorbose (1 mM). (b) Events labeled by machine learning. The distribution of each monosaccharide type matches the corresponding monosaccharide-sensing event when obtained individually. [Figure 23]Different types of D-ribose events are shown. Measurements are performed using MspA-PBA (method in Example 1). A bias voltage of +160 mV is continuously applied. D-ribose is added to the cis chamber to a final concentration of 20 mM. (a) is a representative trace containing D-ribose sensing events. Different types of D-ribose events are labeled with Roman numerals I-IV. Repeated occurrences of the same type of event are labeled with Arabic numerals immediately following the Roman numeral. (b-f) are enlarged displays of each type of event. [Figure 24] Event scatter plots against the SD of ΔI / Ip for D-ribose. (a) Chemical structure of D-ribose. (b-d) Scatter plots against the SD of ΔI / Ip for D-ribose. Events in each plot are from a 30-minute continuous recording trace obtained as described in Figure 23. Nanopore sensing of D-ribose results in different types of sensing events, shown as distinct event populations in the scatter plots. The different event populations are labeled with Roman numerals and correspond to those defined in Figure 23. For display purposes, each scatter point is color-coded based on the local density around each point. Scatter plots were generated using the ggplot2 package in R. [Figure 25] Different types of D-xylose events are shown. Measurements are performed using MspA-PBA (method in Example 1). A bias voltage of +160 mV is continuously applied. D-xylose is added to the cis chamber to a final concentration of 20 mM. (a) is a representative trace containing D-xylose sensing events. Different types of D-xylose events are labeled with Roman numerals I-IV. Repeated occurrences of the same type of event are labeled with Arabic numerals immediately following the Roman numeral. (b-f) are enlarged displays of each type of event. [Figure 26]Event scatter plots against the SD of ΔI / Ip for D-xylose. (a) Chemical structure of D-xylose. (b-d) Scatter plots against the SD of ΔI / Ip for D-xylose. Events in each plot are from a 50-minute continuous recording trace obtained as described in Figure 25. Nanopore sensing of D-xylose results in different types of sensing events, shown as distinct event populations in the scatter plots. The different event populations are labeled with Roman numerals and correspond to those defined in Figure 25. For display purposes, each scatter point is color-coded based on the local density around each point. Scatter plots were generated using the ggplot2 package in R. [Figure 27] Different types of L-rhamnose events are shown. Measurements are performed using MspA-PBA (method in Example 1). A bias voltage of +160 mV is continuously applied. L-rhamnose is added to the cis chamber to a final concentration of 40 mM. (a) is a representative trace containing L-rhamnose sensing events. Different types of L-rhamnose events are labeled with Roman numerals I-II. Repeated occurrences of the same type of event are labeled with Arabic numerals immediately following the Roman numeral. (b-f) are enlarged displays of each type of event. [Figure 28] Event scatter plots against the SD of ΔI / Ip for L-rhamnose. (a) Chemical structure of L-rhamnose. (b-d) Scatter plots against the SD of ΔI / Ip for L-rhamnose. Events in each figure were obtained as described in Figure 27 and are from a 30-minute continuous recording trace. Nanopore sensing of L-rhamnose results in different types of sensing events, which are shown as distinct event populations in the scatter plots. The different event populations are labeled with Roman numerals and correspond to those defined in Figure 27. For display purposes, each scatter point is color-coded based on the local density around each point. Scatter plots were generated using the ggplot2 package in R. [Figure 29]Different types of N-acetyl-D-galactosamine events are shown. Measurements are performed using MspA-PBA (method in Example 1). A bias voltage of +160 mV is continuously applied. N-acetyl-D-galactosamine is added to the cis chamber to a final concentration of 20 mM. (a) is a representative trace containing N-acetyl-D-galactosamine-sensing events. Different types of N-acetyl-D-galactosamine events are labeled with Roman numerals I-II. Repeated occurrences of the same type of event are labeled with Arabic numerals immediately following the Roman numeral. (b-f) are enlarged displays of each type of event. [Figure 30] Event scatter plots of N-acetyl-D-galactosamine versus SD of ΔI / Ip. (a) Chemical structure of N-acetyl-D-galactosamine. (b-d) Scatter plots of N-acetyl-D-galactosamine versus SD of ΔI / Ip. Events in each plot are from a 50-minute continuous recording trace obtained as described in Figure 29. Nanopore sensing of N-acetyl-D-galactosamine results in different types of sensing events, which are shown as distinct event populations in the scatter plots. The different event populations are labeled with Roman numerals and correspond to those defined in Figure 29. For display purposes, each scatter plot point is color-coded based on the local density around each point. Scatter plots were generated using the ggplot2 package in R. [Figure 31]Model evaluation of a classifier trained to distinguish between nine types of monosaccharides. The model is trained as shown in Figure 5. (a) shows the classification performance of different machine learning algorithms based on nine types of sugars. Validation accuracy is calculated using 10-fold cross-validation. The random forest model outperforms all other models by reporting the highest validation accuracy (marked in red). It is selected for further hyperparameter tuning and model evaluation. (b) shows the feature importance of the trained model based on the random forest model. (c) shows the learning curves with different numbers of training samples. From the results in the image inset, a validation accuracy of 0.92 is achieved when the training set size is greater than 720. [Figure 32] This is a machine learning prediction of the results for nine monosaccharides in a mixture. Measurements were performed using a mixture of nine sugars, including D-fructose (2.5 mM), D-galactose (5 mM), D-mannose (5 mM), D-glucose (30 mM), L-sorbose (0.5 mM), D-ribose (5 mM), D-xylose (5 mM), L-rhamnose (10 mM), and N-acetyl-D-galactosamine (2.5 mM) (Figure 5). (a) is a scatter plot of the SD of ΔI / Ip generated from the extracted events. The tags for each event are currently unknown. (b) is a scatter plot of the SD of ΔI / Ip with the machine learning predicted tags. The distribution of each monosaccharide in the scatter plot matches the respective results obtained. [Figure 33]The PBA-modified MspA is used to distinguish classical NMPs. (a) shows the structure of (N90C)1(M2)7. (N90C)1(M2)7 is a heterogeneous octameric MspA consisting of seven M2 MspA-D16H6 units (gray) and one N90C MspA-H6 unit (pink). (N90C)1(M2)7 contains only one cystine (blue) available for subsequent modification. The block shows a top view of (N90C)1(M2)7. (b) shows the mechanism of NMP identification. Phenylboronic acid (PBA) is introduced into the pore-constriction region by modifying the only cysteine thiol with 3-(maleimido)phenylboronic acid (MPBA) via Michael addition. When an analyte is driven into the pore-constriction region by electrophoresis, the NMP reversibly reacts with the PBA, generating random sensing events. (c) Single-channel observation of MPBA modification. After MPBA addition, a current drop of approximately 100 pA was observed, indicating successful MPBA modification. AMP was then added, and successive binding events immediately appeared. To minimize disruption of the lipid bilayer, the applied bias voltage was switched to +20 mV when the Faraday cage was opened. Significant noise was introduced during the period when the Faraday cage was opened and MPBA or AMP was added. The pore opening current (Ip) and cutoff level (Ib) of MspA-PBA are also labeled. (d) NMPs and their corresponding events are shown. The top diagram shows the chemical structures of CMP (C), UMP (U), AMP (A), and GMP (G), with the nucleobases clearly visible. The bottom diagram shows representative sensing events corresponding to the NMPs listed above. As described in Example 2, measurements were performed in 1.5 M KCl buffer (1.5 M KCl, 10 mM MOPS, pH 7.0). A bias voltage of +200 mV is continuously applied. NMP is added to the cis chamber until the final concentration of each analyte is 300 μM. Ip is labeled with a gray dashed line. The cutoff level is labeled with an ink ribbon. (e) The top panel shows a scatter plot of %Ib versus SD for the results obtained using four types of NMP. The bottom panel shows a histogram of the corresponding events for %Ib.Events are obtained from four separate measurements, in which four NMPs are added to the cis chamber to a final concentration of 300 μM. Statistics are generated using 500 consecutive events for each NMP. (f) A representative trace is shown for simultaneous sensing of four NMPs. NMPs are added to the cis chamber simultaneously to a final concentration of 300 μM for each analyte. Events for different NMPs are distinguished according to their characteristic blockage depth and are labeled C, U, A, and G, respectively. [Figure 34]Epigenetic NMPs identified by MspA-PBA are shown. The upper panel (a) shows the epigenetic NMPs studied in this study. Seven epigenetic NMPs are studied, including the monophosphates of 5-methylcytidine (m5C), N6-methyladenosine (m6A), pseudouridine (Ψ), dihydrouridine (D), inosine (I), N7-methylguanosine (m7G), and N1-methyladenosine (m1A). For ease of presentation, only the nucleobases are shown, and all modifications are highlighted in red. The lower panel shows representative events for the corresponding NMPs. From left to right, the representative events are m5C, m6A, Ψ, D, I, m7G, and m1A, respectively. As described in the method of Example 2, measurements were performed in 1.5 M KCl buffer (1.5 M KCl, 10 mM MOPS, pH 7.0). A transmembrane potential of +200 mV was applied continuously. Epigenetic NMPs were added to the cis chamber until the final concentration of each analyte reached 300 μM. The pore opening current (Ip) of MspA-PBA is indicated by a dashed line. Blockade levels are indicated by ink ribbons. A significant difference was observed between %Ib and SD. (b) Violin plot of %Ib for different NMPs. In addition to U and m5C, different NMPs are generally distinguishable solely by analyzing %Ib. (c) Scatter plot of %Ib versus SD for classical and epigenetic NMPs. When considering %Ib and SD, events from the 11 NMPs are clearly distinguishable. Events in (b) and (c) were obtained from 11 independent measurements, in which the 11 NMPs were added to the cis chamber as the sole analyte to a final concentration of 300 μM. Statistics in (b) and (c) were generated using 500 consecutive events for each NMP. [Figure 35]Distinguishing between classical and epigenetic NMPs. (a) A representative trace obtained from simultaneous sensing of CMP and m5C. m5C events exhibit deeper blockage amplitudes and greater noise than CMP events. (b) A scatterplot of the corresponding SD of %Ib of the results from (a). Statistics are generated using 365 consecutive events. (c) A representative trace obtained from simultaneous sensing of GMP and m7G. m7G events are significantly different from GMP events. (d) A scatterplot of the corresponding SD of %Ib of the results from (c). Statistics are generated using 865 consecutive events. (e) A representative trace obtained from simultaneous sensing of AMP, m6A, and m1A. m1A events exhibit a much deeper blockage depth than AMP and m6A events. (f) A scatterplot of the corresponding SD of %Ib of the results from (e). Statistics are generated using 2230 consecutive events. (g) A representative trace obtained from simultaneous sensing of AMP and I. I events exhibit slightly deeper blockage than AMP events. (h) A scatter plot of the corresponding SD of %Ib of the results from g. Statistics are generated using 558 consecutive events. (i) A representative trace obtained from simultaneous sensing of UMP and ψ. ψ events typically exhibit deeper blockage than UMP events. (j) A scatter plot of the corresponding SD of %Ib of the results from i. Statistics are generated using 710 consecutive events. (k) A representative trace obtained from simultaneous sensing of UMP and D. D events typically exhibit deeper blockage than UMP events. (l) A scatter plot of the corresponding SD of %Ib of the results from k. Statistics are generated using 735 consecutive events. NMP-sensing events are labeled according to their characteristic event features. All measurements were performed as described in the methods of Example 2. The chambers are filled with 1.5 M KCl buffer (1.5 M KCl, 10 mM MOPS, pH 7.0). A potential of +200 mV is applied continuously. NMP is simultaneously added to the chambers to a final concentration of 300 μM for each analyte. [Figure 36]This figure shows machine learning-assisted NMP identification. (a) is a flowchart of the training process. Eleven classes of events (including C, U, A, G, m5C, m6A, Ψ, D, I, m7G, and m1A) are applied as the input dataset. Each class consists of 400 events randomly selected from the event pool obtained separately for each analyte type. The mean and standard deviation of each event are extracted to create a feature matrix. The resulting matrix is further randomly divided into a training set for model training and a validation set for model validation, thereby performing 10-fold cross-validation. All classifiers in MATLAB's Classification Learner Toolbox are evaluated to screen the best-performing model. The SVM and naive Bayes models have already demonstrated the highest accuracy score of 0.996. SVM is chosen for all further studies. (b) shows the confusion matrix for NMP classification generated using the SVM model. 100 events from each NMP class are considered as the test set. (c) shows the decision boundary generated by the SVM model. Each colored region represents the region where the corresponding NMP event is predicted. A scatter plot of the SD of %Ib generated using the test data is overlaid on the decision boundary for display purposes. (d) shows a representative trace obtained by simultaneously sensing 11 NMPs. Measurements were performed as described in the method in Example 2. The chamber was filled with 1.5 M KCl buffer (1.5 M KCl, 10 mM MOPS, pH 7.0). A transmembrane potential of +200 mV was applied continuously. NMPs were simultaneously added to the cis chamber until the final concentration of each analyte reached 100 μM. Characteristic events from different NMPs were automatically predicted by the trained SVM model and labeled with different colored dots (CMP: red, UMP: blue, AMP: green, GMP: purple, m5C: yellow, m6A: light purple, Ψ: orange, I: sunset orange, D: cyan, m7G: blue-green, m1A: pink). [Figure 37]Detection of epigenetic modifications from RNA. (a) Schematic diagram of the identification of NMPs from RNA using MspA-PBA. S1 nuclease (green), an endonuclease insensitive to epigenetic modifications, is used to degrade target RNA into NMPs. The generated NMPs are then characterized by MspA-PBA, enabling quantitative analysis of RNA modifications. (b) The sequence of hsa-miR-21 and the corresponding trace of nanopore sensing of the digestion product are shown. Hsa-miR-21 is reported to contain m5C at position 9. Characteristic events from C, U, A, G, and m5C are clearly detected in the trace and labeled accordingly. The m5C cutoff level is labeled with a yellow dashed line. The m5C cutoff level is close to that of U, but the noise of m5C is significantly larger. (c) Scatter plot of the SD of %Ib of hsa-miR-21 digestion products. The displayed events are obtained from a 60-minute continuous recording. NMP identities are predicted by an SVM model. Five populations are detected from the C, U, A, G, and m5C events, respectively. (d) shows the sequence of hsa-miR-17 and the corresponding trace of nanopore sensing of the digestion products. Hsa-miR-21 is reported to contain m6A at position 13. The characteristic events of C, U, A, G, and m6A are clearly detected and labeled with the corresponding tags. The m6A cutoff level is labeled with a light purple dashed line. (e) is a scatter plot of the SD of %Ib obtained from sensing events of hsa-miR-17 digestion products. The displayed events are obtained from a 60-minute continuous recording. NMP identities are predicted by an SVM model. Five populations are detected from the C, U, A, G, and m6A events. As described in the method of Example 2, all measurements were performed in 1.5 M KCl buffer (1.5 M KCl, 10 mM MOPS, pH 7.0). A transmembrane potential of +200 mV was continuously applied. The microRNA digestion product was added to the cis chamber to a final concentration of 100 ng / µL. [Figure 38]Quantitative detection of epigenetic modifications of yeast tRNAPhe. (a) shows the sequence and modifications of yeast tRNAPhe. The modification sites are m2G = N2-methylguanosine, D = dihydroureosine,
[0037]
number
[0038] =N2, N2-dimethylguanosine, C m = 2'-O-methylcytidine, G m = 2'-O-methylguanosine, Y = wybutosine, ψ = pseudouridine, T = 5-methyluridine, m 5 C = 5-methylcytidine, m 7 G = 7-methylguanosine, and m 1 A = 1-methyladenosine, shown shaded. These modifications can be divided into three classes: D, ψ, m, with known event characteristics. 5 Cm 7 G and m 1 A (shown as a red circle) is identifiable. 2 G.
[0039]
number
[0040] , T, and Y (indicated by blue circle shading) are in principle detectable by MspA-PBA, but are not recognized. m and G m (indicated by gray circle shading) cannot be detected by MspA-PBA in principle. (b) shows the results of gel electrophoresis. Lane 1 is a low molecular weight ssRNA standard (NEB Company (New England Biolabs)). Lane 2 is yeast tRNA. phe and lane 3 is yeast tRNA treated with S1 nuclease. phe The gel results show that yeast tRNAphe The results show that yeast tRNA was completely digested by S1 nuclease treatment. phe The digestion procedure is explained in detail in the method of Example 2. (c) tRNA phe %I of digestion products b Scatter plot of SD of C, U, A, G, m. The events shown are from a 240-minute continuous recording. NMP identity is predicted by a linear SVM model. 5 C, ψ, D, m 7 G and m 1 (d) identifies nine major event populations, each corresponding to an event in A. Four event populations that do not belong to any of the previously identified event types are also detected by the unsupervised learning method DBSCAN (Figure 74). phe The comparison of the measured composition and the true value is shown. The measured values are measured and calibrated according to the method in Example 2. (e) shows the tRNA phe Representative traces obtained during the nanopore sensing period of digestion products are shown. Characteristic events of classical and epigenetic NMPs are clearly detected and labeled with the corresponding tags. UT indicates unidentified events. All measurements are performed in 1.5 M KCl buffer (1.5 M KCl, 10 mM MOPS, pH 7.0), as described in the method of Example 2. A transmembrane potential of +200 mV is continuously applied. Yeast tRNA phe The digestion product is added to the cis chamber to a final concentration of 100 ng / µL. [Figure 39] Construction of a co-expression vector is shown. The vector pETDuet-1 is used to co-express two target genes (N90C MspA-H6 and M2 MspA-D16H6) in the same host cell. Specifically, the gene encoding N90C MspA-H6 is inserted between the NcoI and HindIII restriction sites. The gene encoding M2 MspA-D16H6 is inserted between the NdeI and BlpI restriction sites. [Figure 40]This figure shows the production of heterogeneous octameric MspA. (a) UV absorption spectrum during gradient elution on a nickel column. Two major peaks are observed in this spectrum, near fractions 12 and 26. Their identity is confirmed by subsequent gel electrophoresis. (b-c) Heterogeneous octameric MspA was characterized using SDS-polyacrylamide gel electrophoresis (4%-20% gradient gel). Gel electrophoresis was performed continuously for 30 minutes at an applied voltage of +200 V. Lane M is Precision Plus Protein Standard (Bio-Rad). Lane 1 is the bacterial lysate supernatant. Lane 2 is the bacterial lysate eluate after loading onto the column. Lanes 3-15 are the elution fractions described in (a). The fraction indices are labeled in red on the gel. The gel results show that the first peak (fraction 12) corresponds to host cell-derived protein. The second peak (fractions 21-33) corresponds to the eluted heterooctameric MspA, consisting of different combinations of M2 MspA-D16H6 and N90C MspA-H6 monomers. Fractions containing heterooctameric MspA are collected and further separated on a 10% SDS-PAGE gel (Figure 41). [Figure 41] Gel electrophoresis results. All heterooctameric MspA fragments were further characterized using a 10% SDS-PAGE gel to separate the desired pore assembly type. Gel electrophoresis was performed continuously for 16 h at an applied potential of +160 V. On the left, the results from the 10% SDS-PAGE gel are shown. Lane 1 is the homooctameric M2 MspA-D16H6. Lane 2 is the heterooctameric MspA produced using the co-expression plasmid (Figure 39). Lane 3 is the homooctameric N90C MspA-H6. (N90C)1(M2)7, containing one fraction of N90C MspA-H6 and seven fractions of M2 MspA, is shown as the second band above lane 2 (labeled with a pink dashed block). On the right, a cartoon illustration corresponding to the heterooctameric MspA assembly is shown. Red dots represent N90C MspA-H6 monomers, and grey dots represent M2 MspA-D16H6 monomers. [Figure 42] Single-channel characterization of (N90C)1(M2)7 and MspA-PBA is shown. (a) A representative trace exhibiting the sequential insertion of (N90C)1(M2)7 is shown. (b) A representative trace exhibiting the sequential insertion of MspA-PBA is shown. Measurements in (a-b) were performed in 1.5 M KCl buffer (1.5 M KCl, 10 mM MOPS, pH 7.0) with a continuous bias voltage of +200 mV. A nanopore was added to the cis chamber to induce spontaneous pore insertion into the membrane. (c) IV curves of (N90C)1(M2)7 and MspA-PBA are shown. A gradient voltage from -200 mV to +200 mV was applied to obtain the IV curves. Statistics are based on three independent measurements (N = 3) under each condition. (d) Pore opening current histograms for (N90C)1(M2)7 and MspA-PBA measured at +200 mV. Statistics are based on 50 events for each type of pore (N = 50). Pore modification with a single MPBA results in a significant current decrease from 623 ± 13 (mean ± FWHM) pA to 510 ± 14 (mean ± FWHM) pA according to the Gaussian fit results (red and black lines). Relative frequency represents the relative frequency of event counts within the histogram. [Figure 43] The conductance of (N90C)1(M2)7 and MspA-PBA is shown. The conductance of (N90C)1(M2)7 and MspA-PBA was evaluated at different KCl concentrations. (a) IV curves of (N90C)1(M2)7 measured at different KCl concentrations. (b) Graph of the conductance of (N90C)1(M2)7 versus KCl concentration. (c) IV curves of MspA-PBA measured at different KCl concentrations. (d) Graph of the conductance of MspA-PBA versus KCl concentration. Measurements were performed in 0.15 M, 0.5 M, 1.0 M, 1.5 M, and 2.0 M KCl buffer solutions, respectively. A gradient voltage ranging from -150 mV to +150 mV was applied for each condition. The conductance was derived from the slope of the IV curve. Statistics are based on three independent measurements (N=3) under each condition. [Figure 44]Figure 1 shows NMP sensing by M2 MspA. (a-d) Representative traces of (a) CMP, (b) UMP, (c) AMP, and (d) GMP sensing by M2 MspA. As described in the method in Example 2, all measurements were performed in 1.5 M KCl buffer (1.5 M KCl, 10 mM MOPS, pH 7.0). A bias voltage of +200 mV was continuously applied. NMP was added to the cis chamber until the final concentration of each analyte reached 1 mM. No binding events were observed, suggesting that NMP cannot be directly detected by M2 MspA due to the lack of phenylboronic acid chemical modification. [Figure 45] Figure 1 shows single-molecule sensing of AMP and dAMP by MspA-PBA. (a-b) Chemical structures of (a) adenine ribonucleotide (AMP) and (b) adenine deoxyribonucleotide (dAMP). AMP and dAMP differ only in their sugar subunits (labeled red). (c-d) Representative traces obtained with MspA-PBA when (c) AMP or (d) dAMP was added as the sole analyte. In (c), continuous AMP sensing events are observed. However, in (d), no dAMP sensing events are observed. The results indicate that PBA has affinity only for cis-diols, and deoxyribonucleotides cannot be detected by MspA-PBA in principle. Measurements were performed in 1.5 M KCl buffer (1.5 M KCl, 10 mM MOPS, pH 7.0) as described in Example 2. A bias voltage of +200 mV was continuously applied. AMP and dAMP are each added to the cis chamber to a final concentration of 1 mM. [Figure 46]Event parameter definitions are shown. A representative trace containing an AMP-sensing event is shown as an exhibit. Measurements were performed as described in the methods of Example 2. The chamber was filled with 1.5 M KCl buffer (1.5 M KCl, 10 mM MOPS, pH 7.0). A transmembrane potential of +200 mV was continuously applied. AMP was added to the cis chamber to a final concentration of 300 μM. The pore opening current (Ip), residual current (Ib), dwell time (toff), and inter-event duration (toff) were defined as labeled on the trace. The percentage of block (%Ib) was defined as (Ip-Ib) / Ip. The noise amplitude (SD) was defined as the standard deviation of the block level. [Figure 47] Figure 1 shows the binding dynamics of CMP. (a-e) shows representative traces obtained at different CMP concentrations. Measurements were performed in a buffer solution of 1.5 M KCl, 10 mM MOPS, pH 7.0, as described in the method in Example 2. A potential of +200 mV was continuously applied. CMP was added to the cis chamber to a final concentration of 100 μM to 500 μM. (f) shows 1 / τoff versus CMP concentration. (g) shows 1 / τon versus CMP concentration. Three independent measurements were performed to generate statistics. Typically, 1 / τoff remains approximately constant as CMP concentration is changed. As CMP concentration increases, 1 / τon increases. For each condition, events from a 5-minute recording were used to generate the statistics in (f) and (g). [Figure 48]Figure 1 shows the binding dynamics of UMP. (a-e) shows representative traces obtained at different UMP concentrations. Measurements were performed in a buffer solution of 1.5 M KCl, 10 mM MOPS, pH 7.0, as described in the methods section. A potential of +200 mV was continuously applied. UMP was added to the cis chamber to a final concentration of 100 μM–500 μM. (f) shows 1 / τoff versus UMP concentration. (g) shows 1 / τon versus UMP concentration. Three independent measurements were performed to collect statistics. Typically, 1 / τoff remains approximately constant as UMP concentration changes. As UMP concentration increases, 1 / τon increases. For each condition, events from a 5-minute recording were used to generate the statistics in (f) and (g). [Figure 49] AMP binding dynamics are shown. (a-e) show representative traces obtained at different AMP concentrations. Measurements were performed in a buffer solution of 1.5 M KCl, 10 mM MOPS, pH 7.0, as described in the method in Example 2. A potential of +200 mV was continuously applied. AMP was added to the cis chamber to a final concentration of 100 μM to 500 μM. (f) shows 1 / τoff versus AMP concentration. (g) shows 1 / τon versus AMP concentration. Three independent measurements were performed to generate statistics. Typically, 1 / τoff remains approximately constant as AMP concentration changes. As AMP concentration increases, 1 / τon increases. For each condition, events from a 5-minute recording were used to generate the statistics in (f) and (g). [Figure 50]GMP binding dynamics. (a-e) show representative traces obtained at different GMP concentrations. Measurements were performed in a 1.5 M KCl, 10 mM MOPS, pH 7.0 buffer solution, as described in Example 2. A potential of +200 mV was continuously applied. GMP was added to the cis chamber to a final concentration of 100 μM–500 μM. (f) shows 1 / τoff versus GMP concentration. (g) shows 1 / τon versus GMP concentration. Three independent measurements were performed to generate statistics. Typically, 1 / τoff remains approximately constant with changes in GMP concentration. 1 / τon increases with increasing GMP concentration. For each condition, events from a 5-minute recording were used to generate the statistics in (f) and (g). [Figure 51] Figure 1 shows the binding dynamics of AMP at different voltages. (a-e) shows representative traces of AMP sensing obtained using varying voltages. Measurements were performed in a buffer of 1.5 M KCl, 10 mM MOPS, pH 7.0, as described in the method in Example 2. AMP was added to the cis chamber to a final concentration of 500 μM. The applied transmembrane potential ranged from +40 mV to +200 mV, as indicated on each trace. (f) shows 1 / τ versus voltage. (g) shows 1 / τ versus voltage. Three independent measurements were performed to generate statistics. As the voltage increased, 1 / τ decreased, but 1 / τ increased. The results show that a relatively high applied voltage increased the frequency of events, but the dwell time of the events systematically decreased. For each condition, events obtained from a 5-minute recording were used to generate the statistics in (f) and (g). [Figure 52]Figure 1 shows the division of NMPs at different voltages. (a) Representative binding events of CMP, UMP, AMP, and GMP obtained at +100 mV. (b) Histogram of %Ib obtained from simultaneous sensing of CMP, UMP, AMP, and GMP at +100 mV. Events obtained from a 10-minute recording are used to generate statistics. Under these conditions, it is impossible to distinguish between AMP and GMP events. (c) Representative binding events of four NMPs measured at +150 mV. (d) Histogram of %Ib obtained from simultaneous sensing of four NMPs measured at +150 mV. Events obtained from a 10-minute recording are used to generate statistics. Under these conditions, the blockade amplitudes of AMP and GMP events are almost indistinguishable. (e) Representative binding events of four NMPs measured at +200 mV. (f) shows a histogram of %Ib obtained from the simultaneous sensing of four NMPs measured at +200 mV. Events from a 10-minute recording are used to generate statistics. All four NMPs are now completely distinguishable. All measurements are performed as described in the method of Example 2. The chamber is filled with 1.5 M KCl buffer (1.5 M KCl, 10 mM MOPS, pH 7.0). Different types of NMPs (CMP: 600 μM; UMP: 600 μM; AMP: 300 μM; GMP: 300 μM) are simultaneously added to the cis chamber. The concentrations of CMP and UMP are set high to balance the event frequency during simultaneous sensing. [Figure 53]Nanopore sensing of different NMPs is shown. Measurements were performed as described in the methods section. The chamber was filled with 1.5 M KCl buffer (1.5 M KCl, 10 mM MOPS, pH 7.0). A transmembrane potential of +200 mV was continuously applied. NMPs were added to the cis-chamber until the final concentration of each analyte reached 300 μM. (a-d) Representative traces obtained with (a) CMP, (b) UMP, (c) AMP, or (d) GMP as the sole analyte. (e-h) Event histograms of %Ib obtained from sensing (e) CMP, (f) UMP, (g) AMP, and (h) GMP. The blockade amplitudes between the four NMPs are well differentiated. [Figure 54] Fitting results for NMP sensing events are shown. Four parameters, %Ib, SD, toff, and ton, are evaluated during single-molecule sensing of NMP. (a-d) Histograms of %Ib, SD, toff, and ton for (a) CMP, (b) UMP, (c) AMP, or (d) GMP binding events. The histograms of %Ib and SD follow a Gaussian distribution, from which the average percentage blockade
[0041]
number
[0042]
number
[0043] get t off and t on The histogram of is fitted with a single exponential function. off ) and the mean inter-event interval (τ on) are obtained from the corresponding fitting results. Measurements are performed as described in the Methods. The chambers are filled with 1.5 M KCl buffer (1.5 M KCl, 10 mM MOPS, pH 7.0). A transmembrane potential of +200 mV is continuously applied. NMP is added to each chamber until the final concentration of each analyte reaches 300 μM. [Figure 55] Comparison of characteristic event parameters. (ad) shows the (a) of CMP, UMP, AMP, and GMP.
[0044]
number
[0045]
number
[0046] , (c)τ off and (d) τ on The comparison of the two is shown. Measurements are performed as described in the methods. The chamber is filled with 1.5 M KCl buffer (1.5 M KCl, 10 mM MOPS, pH 7.0). A transmembrane potential of +200 mV is continuously applied. NMP is added to the cis chamber until the final concentration of each analyte is 300 μM. Three independent measurements (N=3) are performed to generate statistics. [Figure 56]Sequential addition of NMP during the nanopore sensing period is shown. Measurements are performed as described in the methods section. The chamber is filled with 1.5 M KCl buffer (1.5 M KCl, 10 mM MOPS, pH 7.0). CMP, UMP, AMP, and GMP are sequentially added to the cis chamber. The final concentration of each NMP is 300 μM. Immediately after each addition, continuous single-channel recordings are performed at +200 mV for 20 min. (a) A representative trace in the presence of CMP is shown. Only CMP events (labeled with the letter "C") are observed. (b) The corresponding histogram of %Ib is shown. The %Ib of CMP events is a Gaussian fit. (c) A representative trace obtained immediately after the addition of UMP is shown. Subsequently, UMP events (labeled "U") are observed. (d) The corresponding histogram of %Ib is shown. The distribution of %Ib shows two populations, each with a Gaussian fit. (e) A representative trace obtained immediately after further addition of AMP. AMP events (labeled "A") are then observed in the trace. (f) The corresponding histogram of %Ib. Three event distributions can be clearly seen in the histogram. (g) A representative trace obtained immediately after further addition of GMP. GMP events (labeled "G") exhibit the deepest blockade amplitude. (h) The corresponding histogram of %Ib. Four completely distinguishable Gaussian distributions are clearly observed, corresponding to CMP, UMP, AMP, and GMP events, respectively. [Figure 57] H NMR spectrum of pseudouridine-5'-monophosphate (ψ). ψ was produced and characterized by Wuxi AppTec. [Figure 58] H NMR spectrum of dihydrouridine-5'-monophosphate (D). Production and characterization of D was provided by Wuxi Pharmaceutical Co., Ltd. [Figure 59]Single-molecule sensing of epigenetic NMPs is shown. Measurements are performed as described in the methods section. The chamber is filled with 1.5 M KCl buffer (1.5 M KCl, 10 mM MOPS, pH 7.0). A transmembrane potential of +200 mV is continuously applied. Epigenetically modified NMPs are added to the cis chamber until the final concentration of each analyte reaches 300 μM. (a–g) Representative traces obtained using (a) m5C, (b) m6A, (c) Ψ, (d) D, (e) I, (f) m7G, or (g) m1A as the sole analyte, respectively. (h–n) Corresponding event histograms of %Ib obtained from the sensing results of (h) m5C, (i) m6A, (j) Ψ, (k) D, (l) I, (m) m7G, or (n) m1A, respectively. [Figure 60] The figures show the display of nonspecific events. (a) A representative trace containing consecutive binding events of ψ. Nonspecific events (red arrows) were rarely observed and represented a deeper blockage than ψ (labeled with an orange dashed line). (b) A scatter plot of the corresponding SD of %Ib of ψ. Nonspecific events were detected, showing that only 0.9% of all events were detected. (c) A representative trace containing consecutive binding events of m7G. Nonspecific events (red arrows) were rarely observed and represented a shallower blockage than m7G (labeled with a blue dashed line). (d) A scatter plot of the corresponding SD of %Ib of m7G. Nonspecific events were detected, showing that 1.7% of all events were detected. These nonspecific events may be caused by trace impurities in the sample. These do not interfere with the measurement because they only generate a small fraction of events. Measurements are performed as described in the method of Example 2. The chamber is filled with 1.5 M KCl buffer (1.5 M KCl, 10 mM MOPS, pH 7.0). A transmembrane potential of +200 mV is continuously applied. For each measurement, ψ and m7G are added to the cis chamber to a final concentration of 300 μM. [Figure 61]Statistics of sensing events of epigenetically modified NMPs are shown. Four parameters were evaluated during single-molecule sensing of modified NMPs, including %Ib, SD, toff, and ton. (a–g) Histograms of %Ib, SD, toff, and ton for (a) m5C, (b) m6A, (c) Ψ, (d) D, (e) I, (f) m7G, or (g) m1A binding events. The histograms of %Ib and SD follow a Gaussian distribution, from which the mean percentage blockade
[0047]
number
[0048]
number
[0049] get t off and t on The histogram of is fitted with a single exponential function. off ) and the mean inter-event interval (τ on ) is obtained from the corresponding fitting results. Measurements are performed as described in the method of Example 2. The chamber is filled with 1.5 M KCl buffer (1.5 M KCl, 10 mM MOPS, pH 7.0). A transmembrane potential of +200 mV is continuously applied. Modified NMP is added to the cis chamber until the final concentration of each analyte is 300 μM. [Figure 62] Comparison of NMP characteristic parameters. (a-d) shows the comparison of NMP characteristic parameters for all 11 types of NMP.
[0050]
number
[0051]
number
[0052] , (c)τ off and (d) τ on Comparison of the results is shown. Measurements are performed as described in the methods. The chamber is filled with 1.5 M KCl buffer (1.5 M KCl, 10 mM MOPS, pH 7.0). A transmembrane potential of +200 mV is continuously applied. NMP is added to the cis chamber until the final concentration of each analyte is 300 μM. To generate statistics, three independent measurements are performed for each condition (N=3). [Figure 63] The sequential addition of epigenetic NMPs is shown. Measurements are performed in 1.5 M KCl buffer (1.5 M KCl, 10 mM MOPS, pH 7.0) as described in Example 2. A potential of +200 mV is continuously applied during the measurement period. m5C, m6A, I, m7G, m1A, Ψ, and D are sequentially added to the cis chamber in the presence of CMP, UMP, AMP, and GMP. The final concentration of each NMP is 100 μM. After each addition, single-channel recordings are continuously performed for 10 min. (a) Representative traces are shown in the presence of CMP, UMP, AMP, and GMP. (b-g) Representative traces are shown after the sequential addition of (b) m5C, (c) m6A, (d) I, (e) m7G, (f) m1A, (g) Ψ, and (h) D. Each event was identified using a trained linear SVM model, and the predicted identity is labeled on top of each event with a corresponding color-coded dot (CMP: red, UMP: blue, AMP: green, GMP: purple, m5C: yellow, m6A: light purple, Ψ: orange, I: sunset orange, D: cyan, m7G: blue-green, m1A: pink). [Figure 64]Identification of NMPs by machine learning. As described in the methods, measurements are performed in 1.5 M KCl buffer (1.5 M KCl, 10 mM MOPS, pH 7.0). m5C, m6A, I, m7G, m1A, Ψ, and D are sequentially added to the cis chamber in the presence of CMP, UMP, AMP, and GMP. The final concentration of each NMP is 100 μM. Immediately after each NMP addition, continuous single-channel recordings are performed at +200 mV for 10 min. Identification of NMPs is performed using a trained linear SVM model. (a) Left: Scatter plot of %Ib vs. SD in the presence of CMP, UMP, AMP, and GMP. CMP, UMP, AMP, and GMP events are displayed as red, blue, green, and purple dots, respectively. (b) Right: The corresponding histogram of the event counter. (b-h) Left: Scatter plots of the SD of %Ib after successive additions of (b) m5C, (c) m6A, (d) I, (e) m7G, (f) m1A, (g) ψ, and (h) D. Right: Corresponding histograms of event counts after each addition of NMP. Using a linear SVM model, each newly added NMP can be accurately identified. The gray circles in the scatter plots and black arrows in the histograms indicate the distribution of events corresponding to newly added NMPs. [Figure 65] This is a learning curve. Different amounts of training samples are fed into the machine learning model to evaluate the model's accuracy score. The training score and validation score are obtained from the results of 10-fold cross-validation. According to the learning curve, when the training set contains more than 176 samples, the validation accuracy already reaches 0.990. When the number of samples exceeds 3124, the accuracy saturates at approximately 0.996. The learning curves generated from the training data or validation data are merged with each other to ensure that the model is not overfitting. [Figure 66]Direct nanopore detection of methylated microRNAs is shown. Measurements are performed with MspA-PBA as described in the methods of Example 2. The chamber is filled with 1.5 M KCl buffer (1.5 M KCl, 10 mM MOPS, pH 7.0). A transmembrane potential of +200 mV is continuously applied. MicroRNAs are added to the cis chamber until a final concentration of each analyte is 200 nM. (a) A representative trace obtained with hsa-miR-21 is shown. (b) A representative trace obtained with hsa-miR-17 is shown. Only short-lived dwell spike events are observed for the two analytes. [Figure 67] Figure 1 shows MspA-PBA sensing of a stock solution of S1 nuclease and glycerol. Measurements were performed as described in the method in Example 2. The chamber was filled with 1.5 M KCl buffer (1.5 M KCl, 10 mM MOPS, pH 7.0). A transmembrane potential of +200 mV was continuously applied. (a) A representative trace obtained with a stock solution of S1 nuclease. The stock solution was added to the cis chamber to a final concentration of 1 U / μL. Continuous binding events were observed. These events may be caused by the glycerol in the stock solution, which is used to minimize damage to S1 nuclease due to repeated freeze-thaw cycles. (b) A representative trace obtained with the addition of glycerol. The same events occurred when 1 μL of 40% glycerol was added to the cis chamber. (c) A scatter plot of the SD of %Ib for the S1 nuclease stock solution and glycerol. A thorough comparison of the event parameters shown in (c) leads to the conclusion that the events observed in the S1 nuclease stock solution are caused by glycerol. Scatter plots are generated using 400 consecutive events for each analyte. [Figure 68]Ultrafiltration of an S1 nuclease solution is shown. Measurements are performed with MspA-PBA as described in the method of Example 2. The chamber is filled with 1.5 M KCl buffer (1.5 M KCl, 10 mM MOPS, pH 7.0). A transmembrane potential of +200 mV is continuously applied. (a) A representative trace obtained with an S1 nuclease stock solution is shown. The S1 nuclease stock solution is added to the cis chamber to a final concentration of 1 U / μL. Glycerol in the solution can report characteristic binding events. To minimize interference from glycerol, the S1 nuclease stock solution can be removed by ultrafiltration. (b) A representative trace obtained with an S1 nuclease solution after two rounds of ultrafiltration is shown. During each centrifugation run, the S1 nuclease solution is added to a centrifugal filter containing a 10 kDa MWCO and centrifuged at 8000 rpm for 60 minutes at 4°C. The remaining solution containing S1 nuclease in the filtration device is then collected (method of Example 2). After two rounds of ultrafiltration, S1 nuclease is added to the cis chamber to a final concentration of 1 U / μL. A significant decrease in glycerol binding events is observed. (c) shows a representative trace obtained with the S1 nuclease solution after four rounds of ultrafiltration. Most glycerol binding events disappear. Therefore, the S1 nuclease stock solution pretreated by four rounds of ultrafiltration is used for all subsequent RNA digestion experiments. [Figure 69]This figure shows the results of gel electrophoresis of microRNA digestion. Briefly, 150 μg of microRNA (hsa-miR-21 or hsa-miR-17), 21 μL of pretreated S1 nuclease solution (180 U / μL), and 6 μL of 10× S1 nuclease buffer (300 mM CH3COONa, 2800 mM NaCl, 10 mM ZnSO4, pH 4.6) were mixed and ultrapure water was added to a final volume of 60 μL. The mixture was then incubated at 23°C for 4 hours before gel electrophoresis. The RNA sample was loaded onto a 15% urea-PAGE gel. Gel electrophoresis was performed continuously for 60 minutes at an applied voltage of +200 V. S1 nuclease was pretreated by ultrafiltration to remove glycerol from its storage solution (method of Example 2). Lane M is a microRNA marker (NEB). Lane 1 is hsa-miR-21, L2 is hsa-miR-21 treated with S1 nuclease, L3 is hsa-miR-17, and L4 is hsa-miR-17 treated with S1 nuclease. The gel shows that both hsa-miR-21 and hsa-miR-17 were completely digested by S1 nuclease treatment. The RNA digestion procedure is described in detail in the method of Example 2. [Figure 70]Figure 1 shows the identification of microRNA modifications by SVM. (a) Scatter plot of %Ib vs. SD of the digestion products of hsa-miR-21. The displayed events are from a 60-minute continuous recording. Events are predicted through the trained SVM model. AMP (red), UMP (blue), CMP (green), GMP (purple), and m5C (yellow) populations are successfully detected, consistent with the composition of hsa-miR-21. Furthermore, events from glycerol remaining in the S1 nuclease storage solution are also accurately identified by the SVM model. (b) The corresponding histogram of the event counter. (c) Scatter plot of %Ib vs. SD of the digestion products of hsa-miR-17. The displayed events are from a 60-minute recording. The event distribution corresponding to m6A is clearly observed. (d) The corresponding histogram of the event counter. Measurements are performed with MspA-PBA as described in the method of Example 2. The chamber is filled with 1.5 M KCl buffer (1.5 M KCl, 10 mM MOPS, pH 7.0). A transmembrane potential of +200 mV is continuously applied. The RNA digestion product is added to the cis chamber to a final concentration of 100 ng / µL. The corresponding NMP modifications are detected by a machine learning algorithm. [Figure 71] Figure 1 shows the identification of microRNA composition by the SVM model. (a) shows the sequences of hsa-miR-21 and hsa-miR-17. (b) shows a comparison of the measured hsa-miR-21 composition with the true value. The measured values were measured and calibrated according to the method in Example 2. The NMP composition of hsa-miR-21 was calculated to be 2.2 CMP, 6.8 UMP, 6.9 AMP, 4.9 GMP, and 1.0 m5C. The C, G, and m5C counters in hsa-miR-21 generally matched the true value. (c) shows a comparison of the measured hsa-miR-17 composition with the true value. The NMP composition of hsa-miR-17 was calculated to be 4.5 CMP, 4.1 UMP, 6.5 AMP, 6.4 GMP, and 1.1 m6C. The m6C counter in hsa-miR-21 also generally matched the true value. [Figure 72]This figure shows the removal of glycerin events using machine learning. During nanopore sensing of yeast tRNAPhe digestion products, both NMP and residual glycerin in the S1 nuclease storage solution report binding events. Given the high resolution of MspA, glycerin events (Figure 67) and NMP events are completely distinguishable from each other. The scatter plot on the left displays all events obtained from a 240-minute continuous recording of yeast tRNAPhe digestion products. Glycerin events and NMP events are observed in the scatter plot. Glycerin events contain highly distinctive negative spikes above the cutoff level. As shown in the scatter plot on the right, by learning the event characteristics of glycerin events, described by multiple event features including the event's %Ib, SD, dwell time, skewness, and kurtosis, glycerin events are automatically identified and removed, generating a new set of data that does not contain glycerin events. The remaining data is further analyzed for NMP identification. Here, a one-class SVM with 400 glycerin events is used for model training. Multiple event features, including the %Ib, SD, dwell time, skewness, and kurtosis of the event, are used for training. The events obtained from the yeast tRNAPhe digestion products are predicted by the model. According to the prediction results, events with a decision score greater than 0 are identified as glycerol events. To avoid interference with further data analysis, glycerol events are removed from the dataset. [Figure 73]Figure 7 shows outlier boundary analysis. In machine learning, outlier detection, also known as anomaly detection, is widely used to detect anomalous events that do not appear to belong to any of the previously trained data types. One-class SVM is an anomaly detection algorithm trained to learn whether an event belongs to a previously trained dataset. If not, the event is labeled as an outlier event. Otherwise, it is labeled as an in-point event. (a) shows the outlier boundaries for 11 NMPs generated by one-class SVM. The left shows a scatter plot of the %Ib of the 11 NMPs against the SD. The events are from 11 independent measurements, each of which adds one NMP as the sole analyte to the cis chamber. The right shows the corresponding outlier boundaries. The boundary separating outliers and in-points occurs when the contour value is 0. (b) shows the detection of in-points and outliers for events generated by tRNAphe digestion products. The left shows a scatter plot of the %Ib of the tRNAphe digestion products against the SD. Glycerol events are removed by machine learning (Figure 72). The right shows the separation of in-points and outliers. In-points are labeled with black dots, which are events that belong to one of the previously trained NMP types. Outliers that do not belong to a previously trained NMP type are labeled with gray dots. [Figure 74]Identification of NMPs using supervised and unsupervised learning is shown. (a) Scatter plot of the %Ib of yeast tRNAPhe digestion products against the SD. Glycerol events are automatically removed from the scatter plot using machine learning (Figure 72). With the help of outlier boundary analysis, events are further divided into interior events (b) and outlier events (c) (Figure 73). (b) Scatter plot of the %Ib of interior events against the SD. (d) Using the linear SVM model, interior events are identified. The AMP (red), UMP (blue), CMP (green), GMP (purple), m5C (yellow), ψ (orange), D (cyan), m7G (blue-green), and m1A (pink) populations were successfully detected, consistent with the composition of tRNAPhe. Very few m6A (purple) events were also detected, which may result from secondary background events that coincidentally share the same event characteristics of m6A or other RNAs. However, the proportion of m6A is extremely small. (c) Scatter plot of the %Ib of outliers against the SD. (e) Cluster analysis of outliers. DBSCAN is an unsupervised machine learning algorithm used to identify cluster events among outlier events. ε is set to 0.12, and min_samples is set to 18. Four event clusters were detected, including m2G in yeast tRNAphe,
[0053]
number
[0054]
number
[0055]
number
[0056] (e) Error bars represent standard deviations obtained from independent measurements (N=3).
[0057]
number
[0058]
number
[0059]
number
[0060]
number
[0061] (e) Error bars represent the standard deviation between independent measurements (N=3).
[0062]
number
[0063]
number
[0064]
number
[0065]
number
[0066]
number
[0067]
number
[0068]
number
[0069]
number
[0070] (e) Error bars represent the standard deviation between independent measurements (N=3). Generally, the change in D-sorbitol concentration
[0071]
number
[0072]
number
[0073] It should be understood that the specific methods and conditions described in the embodiments of the present invention are intended to illustrate specific embodiments and are not intended to be limiting, and that any methods and conditions similar or equivalent to those described herein can be used in the practice or testing of the present invention. Any explanation of theories or mechanisms related to the present invention is intended merely to aid in the understanding of the present invention and should not be construed as limiting the embodiments protected by the present invention.
[0074] Unless otherwise specified, the terms used herein have the meanings that are commonly understood in the art and can be understood by reference to standard textbooks, reference works, and literature known to those skilled in the art.
[0075] Unless otherwise specified, the terms "comprise," "comprise," "contain," and variations of these terms, such as "comprises" and "includes," are not intended to exclude other elements, components, wholes, or steps. These terms also encompass the meaning "consisting of" or "consisting of." The terms "consisting of" or "consisting of" are specific embodiments of the term "comprise," which exclude other unspecified elements, components, wholes, or steps.
[0076] It should be noted that, as used in this specification and the appended claims, the singular forms "a," "an," "one," and "the above" include plural referents unless the context clearly dictates otherwise. It should also be noted that the claims may be drafted to exclude any element. Accordingly, this statement is intended to serve as a precondition for the use of exclusive terminology such as "only" or "only," or the use of a "negative" limitation in connection with the recitation of claim elements. The terms "at least one" or "one or more" refer to one, two, three, four, five, six, seven, eight, nine, ten or more. The term "about" refers to a range of plus or minus ten percent (+ / -10%) of a particular value. The term "and / or" refers to any one, some, or all of the elements connected by that term.
[0077] Unless otherwise defined, the terms "first" and "second," when used in conjunction with an element or feature, are used only to distinguish one element or feature from another, and do not imply any particular meaning or priority to a position or step aspect.
[0078] The term "derivative" of a compound means that the derivative contains a core chemical structure common to that compound, but differs by having at least one structural difference, e.g., one or more added and / or removed and / or substituted substituents, and / or one or more atoms replaced with different atoms.
[0079] The term "analog" refers to a chemical molecule that is structurally and functionally similar to another chemical, which retains the same chemical scaffold and function as the parent chemical, but differs structurally by a single element or group, or by two or more groups (e.g., two, three, or four groups).
[0080] It is understood that the methods of the present invention can be practiced in vivo, in vitro, or ex vivo. The methods of the present invention may not be intended to treat a disease and / or may not be intended to diagnose a disease.
[0081] As used herein, the term "nanopore" refers to a pore, channel, or passageway having a very small diameter, typically on the nanoscale, that extends through a membrane. Nanopores can have a characteristic width or diameter of about 0.1 nanometers (nm) to about 1000 nm.
[0082] The term "protein nanopore" refers to a polypeptide subunit or polymer of polypeptide subunits (each subunit can be referred to as a monomer of the protein nanopore) that can form a channel across a membrane. The term "protein nanopore" includes wild-type nanopores, such as α-hemolysin (α-HL), Mycobacterium smegmatis porin A (MspA), erolysin, curli biogenesis / transport component (CsgG), outer membrane porin F (OmpF), cytolysin A (ClyA), ferric hydroxamate uptake component A (FhuA), Fragaceatoxin C (FraC), pleurothricin A (PlyA) / pleurothricin B (PlyB), curli biogenesis / transport component CsgG (CsgG) or Phi29 connexin, or mutants of wild-type nanopores. The sequence of the wild-type protein nanopore can be found in GenBank at https: / / www.ncbi.nlm.nih.gov / . In recent years, various mutants of the protein nanopore have been established.
[0083] A mutant protein nanopore may have one or more amino acid additions, substitutions, and / or deletions compared to its parent, or have at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity compared to its parent, wherein the parent protein or peptide may be a wild-type protein or peptide, or a homologue or variant thereof, and retains tunnel-forming ability.
[0084] As used herein, the term "sequence identity" refers to the percentage of identical nucleotides or amino acid residues at corresponding positions in two or more sequences when matching sequences by maximizing sequence alignment, i.e., taking into account gaps and insertions. Sequence alignment and calculation of percentage sequence identity can be performed using suitable computer programs known in the art. Such programs include, but are not limited to, BLAST, ALIGN, ClustalW, EMBOSS Needle, and the like. An example of a local alignment program is BLAST (Basic Local Alignment Search Tool), which is available from the National Center for Biotechnology Information webpage, currently found at http: / / www.ncbi.nlm.nih.gov / , and was first described in Altschul et al. (1990) J. Mol. Biol. 215; 403-410. Examples of global alignment programs (which optimize alignment over the entire sequence) are the EMBOSS Needle and EMBOSS Stretcher programs, which are based on the Needleman-Wunsch algorithm (Needleman, Saul B. and Wunsch, Christian D. (1970), "A general method applicable to the search for similarities in the amino acid sequence of two proteins", Journal of Molecular Biology 48(3):443-53), both of which are available in their entirety at http: / / www.ebi.ac.uk / Tools / psa / .
[0085] Preferably, protein nanopores used in the present invention do not spontaneously gate at voltages of 150 mV to 200 mV or higher. "Gating" refers to a spontaneous change in conductance through the protein tunnel, which is typically transient (e.g., lasting from as short as 1-10 milliseconds to as long as 1 second). For some protein nanopores, the probability of gating increases as higher voltages are applied. Typically, the protein becomes less conductive during the gating period, and conductance may cease permanently (i.e., the tunnel may permanently close), making the process irreversible. Preferably, gating refers to a spontaneous change in conductance through the protein tunnel to less than 75% of the open-state current.
[0086] Protein nanopores The protein nanopore of the present invention comprises at least one sensing module within a single protein nanopore, which can interact with an analyte, thereby enabling the protein nanopore to characterize a single molecule of the analyte. In a preferred embodiment, the single protein nanopore comprises only one sensing module.
[0087] As used herein, the term "sensing module" refers to a chemical moiety that can interact with a single molecule of a target analyte. The chemical moiety may include one or more chemical molecules or one or more chemical groups. A sensing module may be composed of one or more (e.g., two or more) sensing units.
[0088] As used herein, the term "moiety" refers to a chemical molecule or any portion of a chemical molecule, such as a functional group. As used herein, the term "sensing moiety" refers to a moiety that can interact with a single molecule of a target analyte.
[0089] As used herein, the term "interaction" may refer to a reaction or binding between a sensing moiety and a target analyte, which may be reversible or irreversible. The interaction between the sensing moiety and the target analyte results in a change in the ionic current flowing through the nanopore, which can be measured.
[0090] The sensing module may consist of only one sensing moiety that can interact solely with a single molecule of the target analyte, and this sensing moiety may be called a non-cooperative sensing moiety. In this case, the sensing module corresponds to the non-cooperative sensing moiety.
[0091] A sensing module may be composed of two, three, four, or more sensing moieties, where the two or more sensing moieties interact with a single molecule of a target analyte, and each sensing moiety interacts with one or two or more binding sites of the single molecule. Two or more sensing moieties interacting with a single molecule of a target analyte may be called cooperative sensing moieties. Some single molecules of target analytes may contain two or more binding sites at which the sensing moieties and the target analyte interact. Two or more cooperative sensing moieties in a sensing module may each interact with two or more binding sites in a single molecule. The two or more cooperative sensing moieties in a sensing module may be the same or different and may be designed based on the binding sites in the target analyte. A sensing module composed of cooperative sensing moieties allows for easier and more robust capture of analyte molecules.
[0092] In some embodiments, a protein nanopore (which may also be referred to as a multimeric nanopore) composed of two or more monomers is used. The at least one sensing module may be contained in one or more monomers. A single sensing module may be contained in a single monomer, and the single monomer may contain all of the sensing moieties of the single sensing module. When a sensing module is composed of two or more sensing moieties, the two or more sensing moieties may each be contained in two or more monomers, and each monomer within these monomers may contain one or more sensing moieties.
[0093] In some embodiments, one or more, but not all, monomers of the multimeric nanopore contain one or more sensing modules (which may be referred to as reactive monomers), and the remaining monomers (which may be referred to as non-reactive monomers) do not contain any sensing modules. Such a multimeric nanopore may be referred to as a heterogeneous protein nanopore in the present invention. In some embodiments, only one monomer of the heterogeneous protein nanopore contains one or more sensing modules (preferably only one sensing module or only one sensing moiety), and none of the remaining monomers contain any sensing modules.
[0094] The term "heterogeneous protein nanopore" refers to a protein nanopore in which at least one of the monomers has a different structure (eg, amino acid sequence or amino acid sequence and modifications thereof) than the other monomers.
[0095] The sensing moiety may be an amino acid residue in the polypeptide of the protein nanopore protein or an amino acid residue in a polypeptide coupled to the protein nanopore. In some embodiments, a single sensing moiety consists of or is coupled to a single amino acid residue. In the present invention, the amino acid residues that function as sensing moieties (first class amino acid residues) and the amino acid residues attached to the sensing moiety (first class amino acid residues) are all referred to as reactive amino acid residues (also referred to as reactive sites). A single sensing module may consist of one or more reactive amino acid residues in the polypeptide of the nanopore protein, or one or more sensing moieties each attached to one or more reactive amino acid residues in the polypeptide of the nanopore protein. In some embodiments, the protein nanopore of the present invention comprises one or more reactive amino acid residues (first class or second class). In some embodiments, the protein nanopore comprises only one reactive amino acid residue.
[0096] In a heterogeneous protein nanopore, one or more reactive amino acid residues may be located in one or more, but not all, monomers, and none of the remaining monomers contain a reactive amino acid, hi some embodiments, the protein nanopore contains only one reactive amino acid residue in a single monomer.
[0097] The term "amino acid" refers to any organic molecule containing at least one amino group and at least one carboxyl group. Typically, at least one amino group is located opposite the carboxyl group. The term "amino acid" includes natural amino acids such as the 20 common amino acids (i.e., alanine, cysteine, aspartic acid, glutamic acid, phenylalanine, glycine, histidine, isoleucine, lysine, leucine, methionine, asparagine, proline, glutamine, arginine, serine, threonine, valine, tryptophan, and tyrosine) and protein amino acids including pyrrolysine or selenocysteine; as well as unnatural amino acids such as modified amino acids. The nanopore proteins of the present invention may contain at least one reactive amino acid residue that functions as a sensing moiety (first class) or that couples to a sensing moiety (second class).
[0098] As used herein, the term "modified" or "modification" refers to a change in the state or structure of a molecule of the invention. Molecules can be modified in many ways, including chemical, structural, and functional modifications, for example, by replacing an original molecule or group with a different molecule or group, or by introducing a molecule or group by covalent attachment.
[0099] The term "reactive" is specific to a particular analyte, a particular sensing moiety, and / or a particular linker. An amino acid residue is considered reactive to a first analyte and unreactive to a second analyte if it can interact with a first analyte but not with a second analyte. An amino acid residue is considered reactive to a first sensing moiety and unreactive to a second analyte if it can couple to a first sensing moiety but not to a second sensing moiety. Two different amino acid residues are considered reactive to the same analyte if they can both interact with the same analyte. An amino acid residue is considered reactive to a first linker but unreactive to a second linker if it can interact with a first linker but not with a second linker.
[0100] The term "attached" includes direct or indirect attachment, such as when a first compound is directly attached to a second compound, and refers to a linkage or connection by a bond or force to hold two or more components together, as well as embodiments in which one or more intermediate compounds (particularly groups) are disposed between the first and second compounds. In some embodiments, the sensing moieties or reactive amino acid residues may be attached to each other via a covalent bond.
[0101] The reactive amino acid residues (first class amino acid residues) that can function as sensing moieties can be natural amino acids. In some embodiments, the amino acids that function as sensing molecules can be selected from methionine, histidine, cysteine, lysine, and any combination thereof. In some embodiments, methionine, histidine, cysteine, or lysine can interact alone with a single molecule of a metal ion, and each of them can be used as a sensing module to characterize the metal ion. In some embodiments, two or more of methionine, histidine, cysteine, and lysine can interact with a single molecule of a metal ion and function together as a sensing module consisting of cooperative sensing moieties to characterize the metal ion. In some embodiments, the protein nanopores (particularly heterogeneous protein nanopores) of the present invention contain a single reactive amino acid residue that functions as a sensing moiety, which can be selected from, for example, methionine, histidine, cysteine, and lysine.
[0102] The sensing moiety is preferably coupled to the second class of reactive amino acid residues via a linker. In some embodiments, the reactive amino acid residues are reactive to the linker. The linker may be coupled and attached to the reactive amino acid residue and may be linked to the sensing moiety. In some embodiments, the linker and the sensing moiety may be linked by a covalent bond or by coordination. In some embodiments, the linker may be a ligand. In some embodiments, the linker and the sensing moiety may form a coordination complex.
[0103] The term "coordination" refers to the interaction of a coordinate bond between a multi-electron pair donor and a metal ion, i.e., "coordination." The term "coordination" refers to the interaction between the electron pair donor and a coordination site on the metal ion, which results in an attractive force between the electron pair donor and the metal ion. A coordinate bond can be formed between the electron pair donor and the metal ion. The electron pair donor can be a non-metallic atom such as nitrogen, sulfur, phosphorus, carbon, or oxygen. A compound containing an electron pair donor can be called a ligand. The term "coordination complex" refers to a complex in which a coordinate bond exists between a metal ion and an electron pair donor, ligand, or chelating group. Thus, a ligand or chelating group is typically an electron pair donor, molecule, or molecular ion that has an unshared electron pair available to donate to a metal ion.
[0104] The sensing moiety or linker can be coupled to the reactive amino acid residue by any suitable method, such as a chemical reaction (eg, a click reaction). Examples of click reactions include, but are not limited to, copper(I)-catalyzed alkyne-azide cycloaddition (CuAAC), e.g., the reaction between an azide and an alkyne; copper-free alkyne-azide cycloaddition, e.g., the reaction between an azide and a difluorinated cyclooctyne; Staudinger ligation, e.g., the reaction between an azide and a phosphine; radical addition, e.g., the reaction between a thiol and an olefin; Michael addition, e.g., the reaction between a thiol and a maleimide; and nucleophilic substitution, e.g., the reaction between an amine and a parafluoroethylene (Becer, Hoogenboom, and Schubert, Click Chemistry beyond Metal-Catalyzed Cycloaddition, Angewandte Chemie International Edition, 2009, 48:490-4908; Rostovtsev, VV et al., 2002, A stepwise Huisgen cycloaddition process: Copper(I)-catalyzed regioselective "ligation" of azides and terminal alkynes.Angew.Chem., Int.Ed.41, 2596-2599;Torne, CW et al., 2002, Peptidotriazoles on solid phase: [1,2,3]-Triazoles by regiospecific copper(I)-catalyzed 1,3-dipolar cycloadditions of terminal alkynes to azides.J. Org.Chem.67, 3057-3064;Agard, NJ et al., 2004, A strainpromoted [3+2] azide-alkyne cycloaddition for covalent modification of blomolecules in living systems.J.Am.Chem.Soc.126, 15046-15047;Kohn, M., and Breinbauer, R.(2004, The Staudinger ligation: A gift to chemical biology. Angew. Chem., Int. Ed. 43, 3106-3116). In some embodiments, a sensing moiety or linker can be coupled to a reactive amino acid residue through a reaction between a pair of reactive sites, where a first reactive site is contained within the reactive amino acid residue and a second reactive site is contained in a chemical molecule that also contains the sensing moiety or linker. When a chemical molecule containing the first reactive site is contacted with the reactive amino acid residue, a reaction occurs between the two reactive sites, coupling the sensing moiety or linker to the reactive amino acid residue. In some embodiments, the reactive sites can be click reaction sites.
[0105] The reactive amino acid residue may be a natural amino acid residue that contains a first reactive site. The first reactive site may be introduced into the reactive amino acid residue by amino acid modification. In some embodiments, the first reactive site may be a thiol or amino group, i.e., an ε-amino group. In some embodiments, the second reactive site may be an olefin or maleimide. In some embodiments, a sensing moiety or linker may be coupled to the reactive amino acid residue via a reaction between a thiol and a maleimide. In some embodiments, the second class of reactive amino acid residues may be selected from cysteine, methionine, and lysine. In some embodiments,
[0106] As used herein, the term "reactive site" refers to a chemical molecule, chemical moiety, or chemical group that is exposed to react with another reactive site. A reactive site pair typically consists of a first reactive site and a second reactive site, where the first reactive site can react with the second reactive site. Reactive site pairs are known to those skilled in the art. Reactive site pairs useful in the present invention include, but are not limited to, click reactive sites. The term "click reactive site" refers to a chemical molecule, chemical moiety, or chemical group that participates in a click reaction.
[0107] In some embodiments, the sensing moiety may be a moiety containing a boronic acid, such as phenylboronic acid (PBA), which may function as a non-cooperative sensing moiety and may be coupled to a reactive amino acid residue by a chemical reaction such as a click reaction (e.g., reaction between a thiol and maleic acid). In some embodiments, the protein nanopores (particularly heterogeneous protein nanopores) of the invention comprise a single moiety containing a boronic acid, such as a single phenylboronic acid (PBA).
[0108] In some embodiments, the sensing moiety is a metal ion (which can function as a non-cooperative sensing moiety), e.g., Ni 2+ , Cu 2+ , Co 2+ , Zn 2+ , Cd 2+ , Ag 2+ Pb 2+ , Fe 2+ or Fe 3+ In some embodiments, the metal ion may be coupled to a reactive amino acid residue via a linker, such as a ligand.
[0109] In some embodiments, the ligand may be a metal chelator such as nitrilotriacetic acid (NTA) or iminodiacetic acid (IDA), which may be coupled to a reactive amino acid residue by a chemical reaction such as a click reaction (e.g., reaction between a thiol and a maleimide).
[0110] In some preferred embodiments, the protein nanopores (particularly heterogeneous protein nanopores) of the present invention contain Ni as the sensing module coupled to reactive amino acid residues via NTA. 2+ Including NTA and Ni 2+form a coordination complex that can be referred to as NTA-Ni. A protein nanopore containing NTA-Ni can also be referred to as an NTA-Ni-modified protein nanopore. In some embodiments, the protein nanopore (particularly a heterogeneous protein nanopore) of the present invention comprises a single reactive amino acid residue and a single sensing moiety attached to the reactive amino acid residue by a single ligand. In a more preferred embodiment, the protein nanopore (particularly a heterogeneous protein nanopore) of the present invention comprises a single reactive amino acid residue and a single NTA-Ni attached to the single reactive amino acid residue, wherein "single NTA-Ni" refers to a single NTA and a single Ni. 2+ It refers to a coordination complex consisting of
[0111] It should be understood that in some cases, the protein nanopore inherently contains suitable reactive amino acid residues as defined herein. In other cases, if the protein nanopore does not contain suitable reactive sites, suitable reactive sites can be obtained by modifying the amino acids of the protein nanopore. The protein nanopore to be modified may be referred to as the parent protein nanopore. The modified protein nanopore may be referred to as a variant of the parent protein nanopore or as a parent-derived protein nanopore. Modifications may include amino acid insertion, substitution, deletion, and / or chemical modification. For example, amino acid residues (e.g., non-reactive amino acid residues) in the parent protein nanopore may be replaced with reactive amino acid residues, which can be achieved by chemical synthesis or genetic recombination. When the parent protein nanopore contains two or more multi-reactive amino acid residues, suitable reactive amino acid residues can be obtained by replacing one or more, but not all, of these reactive amino acid residues with non-reactive amino acid residues. "Chemical modification of an amino acid" refers to adding or changing a group in an amino acid by a chemical method to produce a non-natural amino acid.
[0112] The parent protein nanopore can be a wild-type protein nanopore or a mutant thereof. A mutant multimeric protein is a protein nanopore in which one or more or all of the monomers have been modified compared to the parent protein nanopore.
[0113] In some embodiments, the parent protein nanopore may be selected from α-hemolysin (α-HL), Mycobacterium smegmatisporin A (MspA), erolysin, curli biogenesis assembly / transport component (CsgG), outer membrane porin F (OmpF), cytolysin A (ClyA), ferric hydroxamate uptake component A (FhuA), Fragaceatoxin C (FraC), pleurothricin A (PlyA) / pleurothricin B (PlyB), curli biogenesis assembly / transport component CsgG (CsgG) or Phi29 connexin, and any variants thereof. In some embodiments, the parent protein nanopore is selected from wild-type MspA, M1 MspA, and M2 MspA.
[0114] Wild-type MspA, also called MspA, is an octameric protein nanopore, of which each monomer has the following sequence: GLDNELSLVDGQDRTLTVQQWDTFLNGVFPLDRNRLTREWFHSGRAKYIVAGPGADEFEGTLELGYQIGFPWSLGVGINFSYTTPNILIDDGDITAPPFGLNSVITPNLFPGVSISADLGNGPGIQEVATFSVDVSGAEGGVAVSNAHGTVTGAAGGVLLRPFARLIASTGDSVTTYGEPWNMN (SEQ ID NO: 1).
[0115] Mutants of MspA include, but are not limited to, octameric nanopore proteins, in which each monomer has the D90N / D91N / D93N (M1 MspA) or D93N / D91N / D90N / D118R / D134R / E139K (M2 MspA) mutations compared to wild-type MspA. The expression of a mutation means that the mutant simultaneously contains all of the listed mutations compared to wild-type MspA, and the amino acid numbering refers to wild-type MspA.
[0116] The term "heterogeneous protein nanopore" may be considered a variant of a parent protein nanopore in which one or more, but not all, of the monomers are modified compared to the parent protein nanopore.
[0117] The heterogeneous protein nanopore of the present invention can be fabricated by providing one or more monomers that contain one or more reactive amino acid residues (which may be referred to as reactive monomers) and one or more monomers that do not contain reactive sites (which may be referred to as non-reactive monomers), and then assembling them into a protein nanopore under appropriate conditions (e.g., by mixing them together).
[0118] Monomers containing one or more reactive amino acid residues and monomers without reactive amino acid residues can be produced by modifying the protein nanopore. The monomer to be modified can be referred to as the parent monomer, and the modified monomer can be referred to as a variant of the parent monomer or a parent-derived monomer. Modifications can include amino acid insertion, substitution, deletion, and / or chemical modification. For example, an amino acid residue (e.g., a non-reactive amino acid residue) in the parent monomer can be replaced with a reactive amino acid residue, which can be achieved by chemical synthesis or genetic recombination. When a parent monomer contains two or more reactive amino acid residues, suitable reactive amino acid residues can be obtained by replacing one or more, but not all, of these reactive amino acid residues with non-reactive amino acid residues.
[0119] The parent monomer may be derived from a parent protein nanopore and may be a monomer of a wild-type protein nanopore or a mutant thereof. In some embodiments, the parent monomer may be selected from the following protein nanopore monomers: α-hemolysin (α-HL), Mycobacterium smegmatisporin A (MspA), erolysin, curli biogenesis / transport component (CsgG), outer membrane porin F (OmpF), cytolysin A (ClyA), ferric hydroxamate uptake component A (FhuA), Fragaceatoxin C (FraC), pleurothricin A (PlyA) / pleurothricin B (PlyB), curli biogenesis / transport component CsgG (CsgG), or Phi29 connexin, and any mutant thereof. In some embodiments, the parent monomer may be a monomer of wild-type MspA, M1 MspA, or M2 MspA. When the heterogeneous protein nanopore comprises two or more non-reactive monomers, the two or more non-reactive monomers can be the same or different.
[0120] The reactive amino acid residues (class 1 or class 2) may be located on the surface of the nanopore channel, such as in the constriction region (the narrowest part of the nanopore channel) or in the antechamber (located at one end of the nanopore channel and having a larger diameter than the constriction region).
[0121] When the protein nanopore is derived from MspA or a mutant thereof, or when the monomer of the protein nanopore is derived from a monomer of MspA or a mutant thereof, the one or more reactive amino acid residues are located at one or more positions selected from 83 to 111, preferably 90, 91, 92, and 93, the positions of the amino acid residues being relative to wild-type MspA. In some embodiments, the reactive amino acid residue is a cysteine or methionine located at a position selected from 90, 91, 92, and 93.
[0122] In some embodiments, the heterogeneous protein nanopore of the invention is a mutant of MspA, which comprises at least one amino acid mutation in one or more monomers compared to MspA or M2 MspA, in some embodiments, the mutation comprises a mutation at one or more positions to cysteine, methionine, or lysine, preferably selected from 83-113, and preferably selected from 90, 91, 92, and 93.
[0123] In some embodiments, the heterogeneous protein nanopore of the invention is a mutant of MspA and comprises a single reactive monomer containing a single reactive amino acid residue, the single reactive amino acid residue being selected from stein and methionine at position 90, 91, 92, or 93. In some embodiments, the protein nanopore of the invention has an N90C, N90M, and / or N91C mutation in one or more monomers compared to M2 MspA. In some embodiments, the heterogeneous protein nanopore of the invention has a D90C, D90M, and / or D91C mutation in one or more monomers compared to MspA.
[0124] Characterization of target analytes A protein nanopore comprising at least one sensing module of the present invention can be used to characterize (or identify) an analyte. The term "analyte," which may also be referred to as a "target analyte," is a target molecule detectable by a protein nanopore of the present invention. The target analyte can interact with the sensing module contained within the protein nanopore, thereby causing a measurable change in the ionic current across the nanopore. It should be understood that the target analyte matches the sensing module, i.e., the target analyte can be any molecule that can interact reversibly or irreversibly with the sensing module when contacted with the sensing module within the channel of the protein nanopore.
[0125] In some embodiments, the target analyte may be a boronic acid, a metal ion (e.g., Ni 2+, Cu 2+ , Co 2+ , Zn 2+ , Cd 2+ , Ag 2+ Pb 2+ , Fe 2+ or Fe 3+ ), methionine, histidine, cysteine, lysine, and any combination thereof.
[0126] The analyte capable of interacting with the boronic acid may be selected from compounds containing 1,2-diols or 1,3-diols (which may be cis-diols), ions containing metal elements, hydrogen peroxide, and combinations thereof.
[0127] The 1,2-diol or 1,3-diol-containing compound may be selected from polyols, sugars or derivatives thereof, α-hydroxy acids, ribose, nucleoside sugars, alditols, polyphenols, compounds containing catecholamines or catecholamine derivatives, tris(hydroxymethyl)methylaminomethane (Tris), protocatechuic aldehyde, protocatechuic acid, caffeic acid, rosmarinic acid, lithospermic acid, tansinol A, salvianolic acid B, and any combination thereof.
[0128] Polyols include alditols, polyphenols, vitamins, catecholamines, and nucleotide analogs. The sugars may be selected from monosaccharides, oligosaccharides, polysaccharides, and any combination thereof.
[0129] The monosaccharide may be selected from D-glyceraldehyde, D-erythrose, D-ribose, 2'-deoxy-D-ribose, D-xylose, L-arabinose, D-lyxose, D-glucose, D-galactose, D-mannose, D-fructose, L-sorbose, L-fucose, D-allose, D-tagatose, L-rhamnose, and any combination thereof.
[0130] The oligosaccharides may be selected from disaccharides (e.g., sucrose, isomaltulose, maltulose, turanose, leucrose, trehalose, lactulose, maltose, etc.), trisaccharides (e.g., raffinose), tetrasaccharides (e.g., stachyose), and complex oligosaccharides (e.g., acarbose), and any combination thereof. The polysaccharide may be selected from pentasaccharides such as verbascose. The sugar derivative may be selected from N-acetylneuraminic acid (sialic acid), N-acetyl-D-galactosamine, and any combination thereof. The alpha-hydroxy acid may be selected from tartaric acid, malic acid, citric acid, isocitric acid, and any combination thereof.
[0131] The ribose-containing compound may be selected from a nucleotide or modified nucleotide, a derivative of a nucleotide or modified nucleotide, a nucleoside or nucleoside analog, and any combination thereof.
[0132] The nucleotides may be selected from adenine nucleotides, cytosine nucleotides, uracil nucleotides, guanine nucleotides, and any combination thereof.
[0133] Modified nucleotides include methylated, deaminated, reduced or thiolated nucleotides, as well as nucleotides resulting from isomerization of the ribose or nucleobase of the nucleotide. Modified nucleotides include 5-methylcytidine (m 5 C), N6-methyladenosine (m 6 A), pseudouridine (Ψ), inosine (I), N7-methylguanosine (m 7 G), N1-methyladenosine (m 1 A), dihydrouridine (D), N2-methylguanosine (m 2 G), N2, N2-dimethylguanosine
[0134]
number
[0135] The derivative of a nucleotide or modified nucleotide can be selected from monophosphate, diphosphate, triphosphate, and tetraphosphate derivatives of a nucleotide or modified nucleotide, and any combination thereof, such as ADP, UDP, GDP, CDP, ATP, UTP, GTP, CTP, their derivatives, and any combination thereof. The monophosphate, diphosphate, triphosphate, or tetraphosphate derivative of a nucleotide or modified nucleotide can also be referred to as the monophosphate, diphosphate, triphosphate, or tetraphosphate derivative of a nucleoside or modified nucleoside, which refer to a nucleoside monophosphate, modified nucleoside monophosphate, nucleoside diphosphate, modified nucleoside diphosphate, nucleoside triphosphate, modified nucleoside triphosphate, nucleoside tetraphosphate, modified nucleoside tetraphosphate, or a derivative thereof.
[0136] The nucleoside analog is selected from galidesvir, ribavirin, molnupiravir, remdesivir, loxoribine, mizoribine, 5-azacytidine, capecitabine, doxifluridine, 5-fluorouridine, forodesine, kleitosine, pyrazofurin, sangivamycin, pseudouracilidine, and any combination thereof.
[0137] The nucleotide sugar may be selected from uridine diphosphate glucose (UDPG), uridine diphosphate N-acetylglucosamine, uridine diphosphate glucuronic acid, adenosine diphosphate glucose, uridine diphosphate galactose, uridine diphosphate xylose, guanosine diphosphate mannose, guanosine diphosphate fucose, cytidine monophosphate N-acetylneuraminic acid, uridine diphosphate N-acetylgalactosamine, and any combination thereof.
[0138] The alditol may be selected from glycerin, propanetriol, butitol, pentitol, hexitol, erythritol, threitol, arabitol, xylitol, ribitol (adonitol), fucitol, sorbitol (including L-sorbitol or D-sorbitol), mannitol, galactitol, iditol, talitol, allitol (allodulcitol), maltitol, lactitol, isomalt, and any combination thereof.
[0139] The polyphenol may be selected from catechin, neochlorogenic acid, anthocyanins, proanthocyanidins, catechol or a derivative thereof, such as catechol, 3-fluorocatechol, 3-chlorocatechol, 3-bromocatechol, 4-fluorocatechol, 4-chlorocatechol, 4-bromocatechol, 3-methylcatechol, 4-methylcatechol, 3-methoxycatechol, 3-propylcatechol, 3-isopropylcatechol, 3,6-dibromocatechol, 4,5-dibromocatechol, 3,6-dichlorocatechol, and any combination thereof.
[0140] The catecholamine or catecholamine derivative may be selected from epinephrine, norepinephrine (or noradrenaline), isoproterenol, and any combination thereof.
[0141] The ions containing a metal element are selected from alkaline earth metal ions, transition metal ions, and any combination thereof, and are preferably AuCl4 - , Mg 2+ , Ca 2+ , Ba 2+ , Ni 2+ , Cu 2+ , Co 2+ , Zn 2+ , Cd 2+ , Ag 2+ , Pb 2+ and any combination thereof.
[0142] Metal ions (e.g., Ni 2+ , Cu2+ , Co 2+ , Zn 2+ , Cd 2+ , Ag 2+ Pb 2+ , Fe 2+ or Fe 3+ The analyte capable of interacting with the metal ion may be a compound capable of interacting with the metal ion by any means, such as coordination. Such compounds may contain non-metallic atoms that function as electron donors and can coordinate with the metal ion, such as nitrogen, oxygen, or carbon atoms. The compound contains a suitable chemical group capable of coordinating with the metal ion. For example, it may contain at least one carboxylic acid group or at least one amine group. The group may be selected from amino acids; modified amino acids; unnatural amino acids; polymers of amino acids or modified amino acids; compounds containing guanine, adenine, thymine, cytosine, or uracil; and any combination thereof.
[0143] The amino acids may be selected from alanine, cysteine, aspartic acid, glutamic acid, phenylalanine, glycine, histidine, isoleucine, lysine, leucine, methionine, asparagine, proline, glutamine, arginine, serine, threonine, valine, tryptophan, tyrosine, pyrrolysine, selenocysteine, and any combination thereof.
[0144] The modified amino acids may be selected from phosphorylated amino acids, glycosylated amino acids, acetylated amino acids, methylated amino acids, and any combination thereof, such as O-phosphoserine (pS), N4-(β-N-acetyl-D-glucoseamino)-asparagine (GlcNAc-N), O-acetyl-threonine (Ac-T), Nω,N'ω-dimethyl-arginine (SDMA), and any combination thereof.
[0145] A compound containing guanine, adenine, thymine, cytosine, or uracil may be selected from guanine, adenine, thymine, cytosine, or uracil, and may contain any one nucleoside selected from them, and may contain any one nucleotide selected from them, and the nucleotide may be a ribonucleotide or a deoxyribonucleotide.
[0146] Analytes capable of interacting with methionine, histidine, cysteine, and / or lysine can be ions containing metal elements, for example, as defined above.
[0147] The protein nanopore or method of the present invention can be used to characterize carbohydrate-based drugs, polysaccharides / oligosaccharides, small molecule glycosides and glycomimetics, glycopeptides and glycoproteins that contain 1,2-diols or 1,3-diols (which may be cis-diols).
[0148] The protein nanopore of the present invention can be disposed within a membrane separating a first conductive liquid medium from a second conductive liquid medium, which can be referred to as a nanopore system. The nanopore channel is the only pathway through which the first and second conductive liquid media communicate. Typically, a target analyte is added to at least one of the first and second conductive liquid media. The membrane can be an organic membrane, such as a lipid bilayer, or a synthetic membrane, such as a membrane made of a polymeric material. The thickness of the membrane through which the nanopore extends can range from 1 nm to about 10 μm.
[0149] The fabrication of nanopore systems is well known; for example, for protein nanopore systems, when a porin (e.g., a protein nanopore of the present invention) is placed in either a first or second conducting liquid medium separated by a membrane (e.g., a lipid bilayer), the porin can spontaneously insert into the membrane to form a nanopore.
[0150] The sensing moiety can be coupled to the reactive amino acid residue before or after the porin is inserted into the membrane. For example, the sensing moiety can first be coupled to the reactive amino acid residue of the porin, and then the porin containing the sensing moiety can be inserted into the membrane, whereupon the sensing moiety can be coupled to the reactive amino acid residue by mixing them together under conditions suitable for binding the sensing moiety and the porin. For another example, a porin without the sensing moiety can first be inserted into the membrane, and then a molecule containing the sensing moiety can be added to the first or second conductive liquid medium, and then, as it moves through the nanopore, it will contact the reactive amino acid residue and be coupled to the porin.
[0151] When the sensing moiety is coupled to a reactive amino acid residue via a linker, the linker and the sensing moiety can be coupled to the reactive amino acid residue before or after the porin is inserted into the membrane. For example, the linker and the sensing moiety can be first coupled to the reactive amino acid residue of the porin, and then the porin containing the sensing moiety is inserted into the membrane to form a nanopore, whereupon the linker can be coupled to the reactive amino acid residue by mixing them together under conditions suitable for binding the sensing moiety and the porin, and the sensing moiety can be bound to the linker by mixing them together under conditions suitable for the sensing moiety and the linker to interact with each other. For another example, a porin without a sensing moiety can be first inserted into the membrane to form a nanopore, and then a molecule containing a linker can be added to the first or second conductive liquid medium, whereupon it contacts the reactive amino acid residue as it moves through the nanopore and is coupled to the porin, and then a molecule containing the sensing moiety can be added to the first or second conductive liquid medium, whereon it contacts the linker as it moves through the nanopore and is bound to the linker. A linker can be coupled to the reactive amino acid residue by mixing the sensing moiety and the porin together under conditions suitable for binding them. The sensing moiety can be attached to the linker by mixing them together under conditions suitable for the sensing moiety and the linker to interact.
[0152] The target analyte can be added to either side of the nanopore, i.e., to either the first or second conductive liquid medium. In some embodiments, the final concentration of the added analyte can range from about 0.01 mM to about 100 mM, e.g., from about 0.1 mM to about 50 mM, e.g., from about 0.1 mM to about 40 mM. For example, the final concentration of the added analyte can be from about 0.1 mM to about 0.2 mM, about 300 μM, about 0.4 mM, about 0.5 mM, about 0.8 mM, about 1 mM, about 2 mM, about 4 mM, about 6 mM, about 10 mM, about 20 mM, or about 40 mM. For example, the final concentration of the added analyte can be from about 0.1 mM, about 0.2 mM, about 300 μM, about 0.4 mM, about 0.5 mM, about 0.8 mM, about 1 mM, about 2 mM, about 4 mM, about 6 mM, about 10 mM, or about 20 mM to about 40 mM. Appropriate concentrations of different analytes may vary and can be determined empirically.
[0153] When a potential difference (also called a voltage or electric field) is applied between the first and second conductive liquid media (i.e., when an electric field or voltage is applied across the nanopore), an ionic current is generated through the channel of the nanopore, and the target analyte is driven from the conductive liquid medium into the nanopore and can further spread out by, for example, electrophoretic forces and / or diffusion. The potential difference can be 20 mV or more, 40 mV or more, 60 mV or more, 80 mV or more, 100 mV or more, 120 mV or more, 140 mV or more, 160 mV or more, 180 mV or more, or 200 mV or more, or in the range of about 20 mV to 220 mV, about 40 mV to 200 mV, about 60 mV to 180 mV, about 80 mV to 180 mV, about 100 mV to 180 mV, about 120 mV to 180 mV, about 140 mV to 180 mV, or about 160 mV to 180 mV.
[0154] In some embodiments, the potential difference between the first and second conductive liquid media is varied or held constant. Methods and devices for applying an electric field to a nanopore are known to those skilled in the art. For example, a pair of electrodes is used to apply an electric field to the nanopore. As will be appreciated, the voltage range that can be used may depend on the type of nanopore system and the analyte used.
[0155] A target analyte is delivered to the nanopore and interacts with a sensing module on the nanopore. This interaction results in a blockage, which is measured to characterize the target analyte. The system for characterizing a target analyte may further include the target analyte. Preferably, the system may be one in which the target analyte has already interacted with the sensing module, or the target analyte may not yet have interacted with the sensing module.
[0156] A target analyte can be driven into the nanopore by electrophoretic forces or concentration differences (diffusion effects). The target analyte interacts with a sensing module present within the channel of the nanopore, and this interaction causes a blockage of ionic current, which can be measured, for example, by measuring the current after the target analyte enters the nanopore and comparing this current to the current before the target analyte enters the nanopore. Blockage of ionic current can be related to the identity of the target analyte, the interaction between the target analyte and a reagent (e.g., a sensing moiety), the binding dynamics of the target analyte, etc.
[0157] Typically, a "blocking ionic current," which may also be referred to as a "blocking current," is evidenced by a change in ionic current that is clearly distinguishable from noise fluctuations and typically correlates with the presence of an analyte molecule within the nanopore. The strength of the blockage or change in current depends on the characteristics of the analyte. More specifically, a "blockage" can refer to an interval in which the ionic current decreases to a level about 5% to 100% below the unblocked current level, remains there for a period of time, and then spontaneously returns to the unblocked level. For example, the blocking current level may be about, at least about, or at most about 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 100% lower than the unblocked current level. A blockage may also be referred to as a blocking event or events. The measurement can be carried out at any suitable temperature, for example, from −4° C. to 100° C., for example, from 4° C. to 50° C., from 5° C. to 25° C., or at room temperature.
[0158] Measuring the current through a nanopore is well known in the art and can be performed by optical or electrical signals. For example, one or more measurement electrodes can be used to measure the current through the nanopore. These can be, for example, patch clamp amplifiers or data acquisition devices.
[0159] "Liquid medium" includes aqueous, organic-aqueous, and purely organic liquid media. Organic media include, for example, methanol, ethanol, dimethyl sulfoxide, and mixtures thereof. Liquids that can be used in the methods described herein are well known in the art. Descriptions and examples of such media (including conductive liquid media) are provided in U.S. Patent No. 7,189,503, the entire disclosure of which is incorporated herein by reference. Salts, detergents, or buffers may be added to such media. Such reagents may be used to modify the pH or ionic strength of the liquid medium. In some embodiments, the salt may include KCl. In some embodiments, the salt concentration may be 0.5 M to 2.5 M. In some embodiments, the KCl concentration is about 1.5 M. The buffer may be HEPES, MOPS, CHES, Tris, or the like. The pH of the first and / or second conductive liquid medium can range from about 1.0 to about 13.0, preferably from about 6.0 to about 9.0, preferably from about 6.0 to about 8.0, and preferably from about 7.0 to about 7.4, depending on the desired charge characteristics of the target analyte. In some embodiments, the first and / or second conductive liquid medium does not contain Tris. In some embodiments, the first and / or second conductive liquid medium contains 1.5 M KCl, 10 mM MOPS, and has a pH of about 7.0. In some embodiments, the first and / or second conductive liquid medium contains 1.5 M KCl, 10 mM HEPES, and has a pH of about 7.0. In some embodiments, the first and / or second conductive liquid medium contains 1.5 M KCl, 10 mM CHES, and has a pH of about 9.0.
[0160] As used herein, the terms "current pattern" and "current trace" are used interchangeably to refer to ionic currents that change over time. A current pattern may include one or more types of interruption events, or may include one or more individual interruption events of the same type. Characteristics such as the distribution, frequency, and amplitude of the interruption events may be learned from the current pattern.
[0161] As used herein, the term "event" refers to the blocking of a nanopore by a target analyte (i.e., the interval during which the ionic current drops to a level approximately 5%-100% below the unblocked current level, remains there for a period of time, and then spontaneously returns to the unblocked level), and also refers to the change in current caused by the blocking of the target analyte. Those skilled in the art will know how to determine the occurrence of an event. Various characteristic parameters can be obtained from the current pattern. The characteristic parameters are the pore opening current (I p ), interruption level (I s ), cutoff amplitude (ΔI, ΔI=I p -I s ), the inter-event interval (t on ), event residence time (t off ), average residence time (τ off ), mean inter-event interval (τ on ), blocking percentage (ΔI / I p These characteristic parameters include, but are not limited to, the mean square root of the mean square root of the analyte, ... and the standard deviation (SD) of each event. One or more of these characteristic parameters can be used to characterize (or identify) an analyte.
[0162] Characterization (or identification) of a target analyte can include, but is not limited to, determining the identity of the target analyte, determining whether the target analyte is a specific substance, determining the presence or absence of the target analyte, determining the interaction of the target analyte with a reagent (e.g., the reagent can be a sensing moiety and the systems and methods of the invention can be used to determine whether an interaction exists between the target analyte and the sensing moiety), or measuring the binding dynamics of the target analyte with a reagent (e.g., the reagent can be a sensing moiety and the systems and methods of the invention can be used to determine the binding dynamics of the target analyte with the sensing moiety). Identity includes, but is not limited to, the identity of the analyte, the structure of the analyte, the protonation or deprotonation state of the analyte, the chirality of the analyte, etc. As an example, to determine the identity of a target analyte, the tested current pattern can be compared to a reference current pattern to determine the identity of the target analyte.
[0163] As an example, to determine whether a target analyte and a reagent interact, the reagent may be included as a sensing module in a protein nanopore of the present invention, with the occurrence of an event indicating an interaction between the target analyte and the reagent. As used herein, a test current pattern refers to a current pattern obtained using a test analyte (ie, a target analyte).
[0164] A reference current pattern refers to a current pattern used as a reference for determining at least one characteristic of a target analyte. Different reference current patterns can be used depending on the purpose of the characterization. For example, the reference current pattern can be a current pattern obtained by using a known analyte under the same conditions as the test current pattern. This can determine whether the tested analyte is the same as or different from the reference analyte. In some embodiments, characterization of target analytes according to test current patterns can be achieved by using machine learning algorithms.
[0165] In some embodiments, the test current pattern may be high-pass and / or low-pass filtered to provide the test current pattern from the high-pass and / or low-pass, hi some embodiments, the cutoff frequency of the high-pass and / or low-pass is about 100 Hz.
[0166] The nanopore and method of the present invention can be used to characterize a single molecule of a target analyte. As long as the size of the analyte can fit into the nanopore channel, a large amount of analyte can be characterized by the nanopore and method of the present invention. The analyte can interact with one or more moieties, and the analyte can be characterized by the nanopore and method of the present invention, and the one or more moieties can function as a sensing module.
[0167] The nanopores and methods of the present invention can be used to simultaneously characterize multiple (e.g., two or more) different target analytes. Multiple different target analytes can interact with the same sensing moiety. In some embodiments, multiple different target analytes can be simultaneously driven into the nanopore channel to interact with the sensing module, respectively. Different interactions between different analytes and the sensing module can be measured individually and distinguished from one another according to their respective current patterns.
[0168] The term "different" refers to differences in the structure of the target analytes. The target analytes may have different, similar, or identical molecular weights, physical properties, chemical properties, and / or biological properties. The target analytes may be epimers or isomers of each other.
[0169] The nanopores or methods of the invention can be used to distinguish between two or more different analytes that have similar structure and / or similar or identical molecular weight, such as a compound and its isomer or epimer, or a nucleotide and its epigenetic counterpart. The nanopores and methods of the present invention can be used to characterize one or more analytes in a sample.
[0170] The term "sample" may include blood, serum, plasma, bodily fluids, cerebrospinal fluid, foods, beverages, dietary supplements, environmental samples, water samples, etc. The nanopore or method of the present invention may be used to determine the identity of an analyte contained in a sample.
[0171] The sample is preferably liquid, or preferably dissolvable in a liquid medium such as water or an organic solvent. The sample may be added directly to the nanopore system, or may be diluted or dissolved to an appropriate concentration before being added to the nanopore system.
[0172] For example, the sample can be a fruit juice (e.g., grape juice, prune juice, lemon juice), an unsweetened beverage, tea, or an extract of a medicinal herb (e.g., Danshen). The systems and methods of the present invention can be used to characterize sugars, alpha-hydroxy acids and / or alditols in fruit juices, alditols in unsweetened beverages, polyphenols in tea, or protocatechuic aldehyde, protocatechuic acid, caffeic acid, rosmarinic acid, lithospermic acid, tansinol A / salvianolic acid B in an extract of a Chinese herbal medicine (e.g., Danshen).
[0173] The nanopores and methods of the present invention can be used to characterize nucleotides, including unmodified and modified nucleotides, in RNA (e.g., microRNA or tRNA). The RNA can be digested with a nuclease into individual nucleotides, and these nucleotides can then be added to the nanopore system of the present invention to be characterized as analytes. The present invention further relates to the following solutions:
[0174] Solution 1: A heterogeneous protein nanopore comprising two or more monomers, at least one of which contains a reactive site and the other monomers do not contain a reactive site.
[0175] Solution 2: A heterogeneous protein nanopore as described in Solution 1, wherein the reactive site is an amino acid capable of interacting with a target analyte or being linked to a sensing moiety, and the sensing moiety is capable of interacting with the target analyte.
[0176] Solution 3: The heterogeneous protein nanopore described in Solution 1 or 2, wherein the heterogeneous protein nanopore is a mutant of the nanopore, and the nanopore is selected from MspA, α-HL, aerolysin, ClyA, FhuA, FraC, PlyA / B, CsgG, Phi 29 linker and homologs thereof.
[0177] Solution 4: A heterogeneous protein nanopore described in any one of Solutions 1 to 3, wherein the heterogeneous protein nanopore is a mutant of MspA, and the mutant contains at least one amino acid mutation in at least one monomer compared to MspA or M2 MspA.
[0178] Solution 5: A heterogeneous protein nanopore according to Solution 4, wherein the heterogeneous protein nanopore comprises one monomer containing the reactive site and seven monomers not containing the reactive site.
[0179] Solution 6: The heterogeneous protein nanopore according to Solution 4 or 5, wherein the reactive site is an amino acid located at a position selected from 83 to 111, preferably 90, 91, 92, and 93.
[0180] Solution 7: The heterogeneous protein nanopore of any one of Solutions 1 to 6, wherein the reactive site is selected from cysteine, methionine, lysine and unnatural amino acids.
[0181] Solution 8: A protein nanopore reactor comprising a heterogeneous protein nanopore according to any one of Solutions 1 to 7 and a sensing moiety optionally linked to said reactive site.
[0182] Solution 9: The protein nanopore reactor of Solution 8, wherein the reactive site or the sensing moiety is capable of interacting with a target analyte.
[0183] Solution 10: The protein nanopore reactor according to Solution 9, wherein the sensing moiety is phenylboronic acid (PBA).
[0184] Solution 11: The target analyte is Ions containing a metal element, preferably ions containing an alkaline earth metal or a transition metal, more preferably AuCl4 - , Mg 2+ , Ca 2+ , Ba 2+ , Ni 2+ , Cu2+ , Co 2+ , Zn 2+ , Cd 2+ , Ag 2+ or Pb 2+ , Monosaccharides, preferably D-glyceraldehyde, D-erythrose, D-ribose, 2'-deoxy-D-ribose, D-xylose, L-arabinose, D-lyxose, D-glucose, D-galactose, D-mannose, D-fructose, L-sorbose, L-fucose, D-allose, D-tagatose, L-rhamnose, N-acetylneuraminic acid (sialic acid), Oligosaccharides, preferably disaccharides such as sucrose, isomaltulose, maltulose, turanose, leucrose, trehalose, lactulose, maltose, trisaccharides such as raffinose, tetrasaccharides such as acarbose or stachyose, Polysaccharides such as verbascose, A compound containing a ribose moiety, preferably a nucleotide or modified nucleotide, or a monophosphate, diphosphate, triphosphate, or tetraphosphate derivative of a nucleotide or modified nucleotide, or a nucleoside or nucleoside analogue, preferably the nucleotide comprises an adenine nucleotide, a cytosine nucleotide, a uracil nucleotide, or a guanine nucleotide, preferably the modified nucleotide is 5-methylcytidine (m 5 C), N6-methyladenosine (m 6 A), pseudouridine (Ψ), inosine (I), N7-methylguanosine (m 7 G), or N1-methyladenosine (m 1 A), preferably the nucleoside analog comprises galidesvir, ribavirin, molnupiravir or remdesivir; nucleotide sugars, such as uridine diphosphate glucose, uridine diphosphate N-acetylglucosamine, uridine diphosphate glucuronic acid, adenosine diphosphate glucose, uridine diphosphate galactose, uridine diphosphate xylose, guanosine diphosphate mannose, guanosine diphosphate fucose, cytidine monophosphate N-acetylneuraminic acid, or uridine diphosphate N-acetylgalactosamine; Alditols, such as erythritol, threitol, arabitol, xylitol, ribitol (adonitol), fucitol, sorbitol, mannitol, galactitol, iditol, talitol, allitol (alodulcitol), maltitol, lactitol, or isomalt; polyphenols such as anthocyanins or proanthocyanidins, a catecholamine or catecholamine derivative, preferably epinephrine, norepinephrine, or isoproterenol; Catechol or its derivatives, for example, catechol, 3-fluorocatechol, 3-chlorocatechol, 3-bromocatechol, 4-fluorocatechol, 4-chlorocatechol, 4-bromocatechol, 3-methylcatechol, 4-methylcatechol, 3-methoxycatechol, 3-propylcatechol, 3-isopropylcatechol, 3,6-dibromocatechol, 4,5-dibromocatechol, 3,6-dichlorocatechol, hydrogen peroxide, a buffer reagent, preferably Tris, glycerin, or any combination thereof.
[0185] Solution 12: The protein nanopore reactor of Solution 9, wherein the sensing moiety is a nickel ion, a cobalt ion, or a copper ion.
[0186] Solution 13: The protein nanopore reactor of Solution 9 or 12, wherein the target analyte is selected from natural amino acids, unnatural amino acids, and modified amino acids such as selenocysteine.
[0187] Solution 14: A method for identifying a target analyte, the method comprising: (i) providing a protein nanopore reactor according to any one of Solutions 8 to 13; (ii) applying a voltage across the protein nanopore reactor; (iii) allowing the target analyte to pass through the nanopore; and (iv) measuring the ionic current passing through the nanopore to provide a current pattern, and identifying the target analyte based on the current pattern.
[0188] Solution 15: The target analyte is Ions containing a metal element, preferably ions containing an alkaline earth metal or a transition metal, more preferably AuCl4 - , Mg 2+ , Ca 2+ , Ba 2+ , Ni 2+ , Cu 2+ , Co 2+ , Zn 2+ , Cd 2+ , Ag 2+ or Pb 2+ , Monosaccharides, preferably D-glyceraldehyde, D-erythrose, D-ribose, 2'-deoxy-D-ribose, D-xylose, L-arabinose, D-lyxose, D-glucose, D-galactose, D-mannose, D-fructose, L-sorbose, L-fucose, D-allose, D-tagatose, L-rhamnose, N-acetylneuraminic acid (sialic acid), Oligosaccharides, preferably disaccharides such as sucrose, isomaltulose, maltulose, turanose, leucrose, trehalose, lactulose, maltose, trisaccharides such as raffinose, tetrasaccharides such as acarbose or stachyose, Polysaccharides such as verbascose, A compound containing a ribose moiety, preferably a nucleotide or modified nucleotide, or a monophosphate, diphosphate, triphosphate, or tetraphosphate derivative of a nucleotide or modified nucleotide, or a nucleoside or nucleoside analogue, preferably the nucleotide comprises an adenine nucleotide, a cytosine nucleotide, a uracil nucleotide, or a guanine nucleotide, preferably the modified nucleotide is 5-methylcytidine (m 5 C), N6-methyladenosine (m 6 A), pseudouridine (Ψ), inosine (I), N7-methylguanosine (m 7 G), or N1-methyladenosine (m 1 A), preferably the nucleoside analog comprises galidesvir, ribavirin, molnupiravir or remdesivir; nucleotide sugars, such as uridine diphosphate glucose, uridine diphosphate N-acetylglucosamine, uridine diphosphate glucuronic acid, adenosine diphosphate glucose, uridine diphosphate galactose, uridine diphosphate xylose, guanosine diphosphate mannose, guanosine diphosphate fucose, cytidine monophosphate N-acetylneuraminic acid, or uridine diphosphate N-acetylgalactosamine; Alditols, such as erythritol, threitol, arabitol, xylitol, ribitol (adonitol), fucitol, sorbitol, mannitol, galactitol, iditol, talitol, allitol (alodulcitol), maltitol, lactitol, or isomalt; polyphenols such as anthocyanins or proanthocyanidins, a catecholamine or catecholamine derivative, preferably epinephrine, norepinephrine, or isoproterenol; Catechol or its derivatives, for example, catechol, 3-fluorocatechol, 3-chlorocatechol, 3-bromocatechol, 4-fluorocatechol, 4-chlorocatechol, 4-bromocatechol, 3-methylcatechol, 4-methylcatechol, 3-methoxycatechol, 3-propylcatechol, 3-isopropylcatechol, 3,6-dibromocatechol, 4,5-dibromocatechol, 3,6-dichlorocatechol, hydrogen peroxide, a buffer reagent, preferably Tris, glycerin, or any combination thereof.
[0189] Solution 16: Use of a heterogeneous protein nanopore according to any one of Solutions 1 to 7 or a protein nanopore reactor according to any one of Solutions 8 to 13 in identifying a target analyte.
[0190] Solution 17: The target analyte is Ions containing a metal element, preferably ions containing an alkaline earth metal or a transition metal, more preferably AuCl4 - , Mg 2+ , Ca 2+ , Ba 2+ , Ni 2+ , Cu 2+ , Co 2+ , Zn 2+ , Cd 2+ , Ag 2+ or Pb 2+ , Monosaccharides, preferably D-glyceraldehyde, D-erythrose, D-ribose, 2'-deoxy-D-ribose, D-xylose, L-arabinose, D-lyxose, D-glucose, D-galactose, D-mannose, D-fructose, L-sorbose, L-fucose, D-allose, D-tagatose, L-rhamnose, N-acetylneuraminic acid (sialic acid), Oligosaccharides, preferably disaccharides such as sucrose, isomaltulose, maltulose, turanose, leucrose, trehalose, lactulose, maltose, trisaccharides such as raffinose, tetrasaccharides such as acarbose or stachyose, Polysaccharides such as verbascose, A compound containing a ribose moiety, preferably a nucleotide or modified nucleotide, or a monophosphate, diphosphate, triphosphate, or tetraphosphate derivative of a nucleotide or modified nucleotide, or a nucleoside or nucleoside analogue, preferably the nucleotide comprises an adenine nucleotide, a cytosine nucleotide, a uracil nucleotide, or a guanine nucleotide, preferably the modified nucleotide is 5-methylcytidine (m 5 C), N6-methyladenosine (m 6 A), pseudouridine (Ψ), inosine (I), N7-methylguanosine (m 7 G), or N1-methyladenosine (m 1 A), preferably the nucleoside analog comprises galidesvir, ribavirin, molnupiravir or remdesivir; nucleotide sugars, such as uridine diphosphate glucose, uridine diphosphate N-acetylglucosamine, uridine diphosphate glucuronic acid, adenosine diphosphate glucose, uridine diphosphate galactose, uridine diphosphate xylose, guanosine diphosphate mannose, guanosine diphosphate fucose, cytidine monophosphate N-acetylneuraminic acid, or uridine diphosphate N-acetylgalactosamine; Alditols, such as erythritol, threitol, arabitol, xylitol, ribitol (adonitol), fucitol, sorbitol, mannitol, galactitol, iditol, talitol, allitol (alodulcitol), maltitol, lactitol, or isomalt; polyphenols such as anthocyanins or proanthocyanidins, a catecholamine or catecholamine derivative, preferably epinephrine, norepinephrine, or isoproterenol; Catechol or its derivatives, for example, catechol, 3-fluorocatechol, 3-chlorocatechol, 3-bromocatechol, 4-fluorocatechol, 4-chlorocatechol, 4-bromocatechol, 3-methylcatechol, 4-methylcatechol, 3-methoxycatechol, 3-propylcatechol, 3-isopropylcatechol, 3,6-dibromocatechol, 4,5-dibromocatechol, 3,6-dichlorocatechol, hydrogen peroxide, a buffer reagent, preferably Tris, glycerin, or any combination thereof.
[0191] Solution 18: A method for producing a heterogeneous protein nanopore according to any one of Solutions 1 to 7, the method comprising: (a) expressing modified and unmodified monomers in the same host cell, wherein additional polyamino acids are added to either end of the modified and unmodified monomers, the polyamino acids being sufficient to provide a distinguishable molecular weight difference between the monomers with the polyamino acid and the monomers without the polyamino acid; (b) allowing the modified monomer and the unmodified monomer to self-assemble; (c) using a specific number of modified monomers and a specific number of unmodified monomers to purify the heterogeneous protein nanopore by said molecular weight difference.
[0192] Example Example 1: Single-molecule identification of monosaccharides with boronic acid-modified Mycobacterium smegmatisporin A nanopores Sugars play important roles in many forms of cellular activity, including energy supply, structural organization, and immune recognition. However, sugar structures are highly complex and similar, presenting a technical obstacle to direct identification. Nanopores are emerging single-molecule tools that are sensitive to microstructural differences between analytes and can be engineered to identify sugars. We fabricated heterogeneous octameric Mycobacterium smegmatisporin A (MspA) nanopores containing a unique phenylboronic acid (PBA) and were able to clearly identify nine monosaccharide types, including D-fructose, D-galactose, D-mannose, D-glucose, L-sorbose, D-ribose, D-xylose, L-rhamnose, and N-acetyl-D-galactosamine. Given that the conical structure of MspA provides high resolution, microstructural differences between sugar epimers can also be distinguished. To assist in automated classification of events, machine learning algorithms were developed, achieving a typical accuracy score of 0.96. This sensing strategy may generally be applied to other sugar types or even small oligosaccharides, bringing new insights into nanopore sugar sequencing.
[0193] introduction Sugars, also known as carbohydrates, are important biomolecules for almost all living organisms. 1 As a core component of food, it provides energy for almost all cellular activities. 2 These further constitute the main components of cellulose and pectin, providing structural integrity to the cell. 3 Glycosylation is the process by which glycans are covalently attached to lipids or proteins to form lipopolysaccharides or glycoproteins, which are essential for the physiological and pathological functions of cells. 4-6 The recently discovered glycogen RNA demonstrates that conserved small non-coding RNAs also have sialylated glycans. 7 The diverse functions of sugars stem from their diverse structures, which can be extremely complex and whose mechanisms of action are not fully understood. 8,9 (micro)array 10,11 , capillary electrophoresis (CE) 12,13 , liquid chromatography (LC)14,15 , nuclear magnetic resonance (NMR) 16,17 , and mass spectrometry (MS) 18,19 Although studies of polysaccharide sequence or structure have been performed by various methods, characterization by any single method provides an incomplete picture of the glycan analyte. 20 Specifically, the stereochemical information of monosaccharides cannot be determined by MS, and isomers cannot be distinguished. 20,21 In nature 15 The low abundance of N makes it difficult to determine the structure of amino modifications carried on glycans using NMR 20,22 Characterization of sugars by these methods is usually expensive and time-consuming; large amounts of input material may be required, and the corresponding data are difficult to interpret. 23,24 .
[0194] nucleic acid 25-27 or peptides 28,29 Recent advances in nanopore sequencing of glycans indicate that they have the potential to sequence glycans in a similar manner. However, because the structures of the monosaccharide components are very similar, 30 Therefore, there is an urgent need for nanopores that can completely distinguish between monosaccharides. Solid-state nanopores have previously been used for polysaccharide sensing. 31-34 However, direct identification of monosaccharides using solid-state nanopores has not been reported to date. In aqueous environments, boronic acids react with 1,2- or 1,3-diols. 35 known to form reversible covalent bonds with (including sugars) 36,37 However, designing a boronic acid sensor that can selectively report specific sugar signals is very complicated. 38 Direct sensing of D-glucose, D-fructose, and D-maltose by placing a phenylboronic acid (PBA) adaptor on α-hemolysin (α-HL) has been reported. 39 However, possibly due to its unfavorable cylindrical lumen geometry, it has low resolution and cannot truly distinguish between D-glucose and D-fructose, and other types of sugars have not been tested by this method. 39
[0195] Mycobacterium smegmatisporin A (MspA) has an overall conical lumen structure. 40-42 It is an octameric pore-forming toxin with a .GAMMA.-containing .GAMMA.-containing nanopore, and the first nanopore for which DNA was successfully sequenced. 25 Then, epigenetic modifications during nanopore sequencing 43 and DNA damage 44,45 Manipulation of the pore constriction region further enables MspA to directly monitor chemical reactions with high resolution. 46,47 Recent demonstrations of the use of programmable nanopore reactors further demonstrate that phenylboronic acid can be placed into the pore cavity to permit binding of polyols such as epinephrine and remdesivir. 48 This report suggests the modification of PBA in MspA for sugar sensing. However, to our knowledge, there are no reports on sugar sensing using modified MspA.
[0196] result Before placing the only PBA into MspA, we first produced a heterogeneous octameric MspA. Experimentally, two distinct genes encoding M2 MspA-D16H6 (Table 1) and N90C MspA-H6 (Table 1), respectively, were custom synthesized and simultaneously inserted into the pETDUET-1 co-expression vector (Figure 6). After heat shock transformation with this vector, the two genes were co-expressed using the E. coli BL21(DE3)pLysS strain (method of Example 1, Figure 7), producing an octameric MspA assembly consisting of different amounts of the two protein types (Figure 1b). The MspA assembly consisting of one unit N90C MspA-H6 and seven units M2 MspA-D16H6 is the desired MspA heterogeneous octamer, referred to as (N90C)1(M2)7 (Figure 1a). (N90C)1(M2)7 contains a single cysteine in the pore-constriction region of the N90C MspA-H6 component at site 90. Experimentally, (N90C)1(M2)7 was purified from other types of MspA assemblies by gel separation and used directly in all downstream measurements (Figure 1b). To introduce a phenylboronic acid (PBA) group into (N90C)1(M2)7, 3-(maleimido)phenylboronic acid (MPBA) is chemically attached to the only cysteine of (N90C)1(M2)7 via maleimide-thiol coupling (Figure 1c). Real-time single-molecule characterization of this reaction is achieved by performing single-channel recordings (methods in Example 1). Briefly, electrophysiological measurements are performed using a single (N90C)1(M2)7 in a buffer solution of 1.5 M KCl, 10 mM MOPS, pH 7.0. When a bias voltage of +100 mV is continuously applied, the pore-opening current (I0) of (N90C)1(M2)7 is measured to be approximately 295 pA. At this stage, additional shot noise is also observed (Figure 1c). Cysteine thiol residues in the pore-constriction region may result in the generation of these noises, a phenomenon previously reported by engineered α-hemolysin (α-HL) mutants. 49 Next, MPBA was added to the cis chamber to a final concentration of 2 mM. An irreversible single-step decrease in current was observed, measuring approximately 53 pA. No further decrease in current was observed, and the previously observed shot noise disappeared simultaneously. These phenomena demonstrated that MPBA had already successfully bound to the only cysteine of (N90C)1(M2)7. This PBA-bound MspA heterooctamer was called MspA-PBA, and its pore-opening current was I p is defined as (Figure 1c).
[0197] Prior to single channel recording, the entire preparation of MspA-PBAs is carried out by mixing purified (N90C)1(M2)7 with MPBAs (method of Example 1). (N90C)1(M2)7(I0) and MspA-PBAs (I p Further characterization of the pore opening current of (N90C)1(M2)7 showed that the current difference between the current measured with (N90C)1(M2)7 and that measured with MspA-PBA was constant, indicating that the produced MspA-PBA reported a uniform structure and could be easily distinguished from the unmodified form (N90C)1(M2)7 (Figure 8, Tables 2 and 3). The I values and I pThese values are also consistent with those previously measured during single-channel recording (Figure 1c), demonstrating that the overall fabricated MspA-PBA is the same as that previously characterized during real-time pore modification. Unless otherwise stated, all subsequent measurements are performed using the overall fabricated MspA-PBA. L-sorbose is a monosaccharide and a ketose. It is present in all living species, from bacteria to humans. 50 Commercial production of vitamin C (ascorbic acid) usually begins with L-sorbose. 51 To further verify the presence of PBAs in the pore cavity, we used MspA-PBA to sense L-sorbose. Measurements were performed using a single MspA-PBA and a 1.5 M KCl, 10 mM MOPS, pH 7.0 buffer solution, with a continuous bias voltage of +160 mV. Adding L-sorbose to the cis chamber to a final concentration of 10 mM immediately generated a series of long-lasting resistive pulse events (Figure 1d-f, Supplementary Video 1). However, when testing M2 MspA, adding L-sorbose to the cis chamber to a final concentration of 50 mM instead of MspA-PBA did not result in any resistive pulse events (Figure 11). This demonstrates that introducing PBAs into the pore cavity is crucial for generating sugar-sensing events. To quantitatively describe the events, we define the parameters of each event, such as the event residence time (t off ), inter-event interval (t on ), interruption level (I s ), cutoff amplitude (ΔI=I p -I s ) and standard deviation (SD) are defined in Figure 10. Typically, L-sorbose sensing events report large and uniform blockade amplitudes (ΔI), measured at approximately 100 pA, more than 10 times the value previously reported for monosaccharide sensing by boronic acid-modified α-HL. 39Furthermore, highly distinctive noise features are also consistently observed. This is well illustrated in the scatter plot of ΔI versus the SD (Fig. 1g), where only a single event population of events is observed. ΔI / I p Normalized scatter plots of the SD of the sigma-positive and sigma-negative values also show results from three independent measurements (Fig. 13, N = 3), and similar results are obtained.
[0198] By continuously increasing the cis-L-sorbose concentration during the measurement period, the event frequency increased proportionally (Figure 1h, Figure 9, Table 4). To quantitatively describe the concentration dependence of the binding dynamics, the mean residence time (τ off ) or the average inter-event interval (τ on ) are t off or t on The reciprocal of the mean inter-event interval (1 / τ on ) increases linearly with increasing L-sorbose concentration, which is consistent with the bimolecular model. However, 1 / τ off The value is independent of L-sorbose concentration, consistent with a single-molecule dissociation mechanism. This further demonstrates that the resistance pulse event is the result of reversible binding of individual L-sorbose molecules to the sole phenylboronic acid reactive site of a single MspA-PBA. The same measurements are also performed at different applied voltages. The blockade amplitude (ΔI) increases linearly with applied voltage, whereas τ on and τ offThe change in the pore closure frequency is minimal (Figure 12, Tables 5 and 6). Despite being geometrically confined to the pore contraction region, the reaction between PBA and L-sorbose is not impeded by the local electric field. In contrast, analyte diffusion plays a more important role in regulating the frequency of the event. This is expected because L-sorbose is an electrically neutral molecule under the test conditions. Although not demonstrated here, the binding dynamics of charged sugars are expected to be strongly regulated by the applied voltage. Experimentally, measurements at higher voltages reported larger event amplitudes, but spontaneous pore closure was observed more frequently. Therefore, a bias voltage of +160 mV appears optimal for continuous and long-term measurements.
[0199] The feasibility of sugar sensing by MspA-PBA has been successfully demonstrated using L-sorbose. The same principle can be applied to sensing other sugar analytes as long as the analyte can react with PBA in the pore constriction region. The huge event amplitude and unique fluctuation noise observed from the binding of L-sorbose and MspA-PBA indicate that MspA has the resolution to directly distinguish different sugar types solely through nanopore readouts. To confirm this speculation, D-fructose, D-galactose, D-mannose, and D-glucose are used as analytes. These four sugars also represent the most abundant monosaccharide types in nature. 52 Their molecular weights are identical, meaning that they cannot be directly distinguished by mass spectrometry alone. Specifically, D-mannose and D-galactose are the C2 and C4 epimers of D-glucose, respectively, and have very slight structural differences.
[0200] All subsequent measurements are performed using MspA-PBA in a buffer solution of 1.5 M KCl, 10 mM MOPS, pH 7.0. A bias voltage of +160 mV is continuously applied. D-fructose, D-galactose, D-mannose, or D-glucose is added to the cis chamber to reach the desired concentration. For D-fructose (Figure 2a), the nanopore sensor reports multiple types of events. Representative events of each type are summarized in Figures 2b and 14. This is expected because, in an aqueous environment, D-fructose exists as a mixture of pyranose and furanose isomers, each of which has an α-anomer and a β-anomer. 38 The binding between different combinations of hydroxyl groups and PBAs may also serve to generate different types of sensing events. 36-38 However, each event is further divided by ΔI / I p The scatter plots of events against the SD of the sigma-based ...
[0201] Following the same principle, we also test and evaluate D-galactose (Figure 2d-2f), D-mannose (Figure 2g-2i, Supplementary Video 2), and D-glucose (Figure 2j-2l). Similar to D-fructose, each type of sugar exhibits multiple types of events when sensed by MspA-PBA (Figure 2e, Figure 2h, Figure 2k). These event types exhibit highly distinguishable blockage depths and characteristic event noise, which can be used to identify different sugar types. The scatter plot results for D-galactose (Figure 2f), D-mannose (Figure 2i), and D-glucose (Figure 2l) also show distinct event populations among the different sugar types. A detailed description of the consistency of event types and reproducibility across different tests for each condition (N = 3) is summarized in Figures 16-21. Specifically, when the same type of sugar was tested, the scatter plot results showed a highly consistent pattern of event distribution (Figures 15, 17, 19, and 21), providing information to distinguish between different sugar types. Although events from different sugar types are visually distinguishable, the ΔI / I of each event p It remains difficult to automatically and quantitatively characterize the differences between sugar types by considering only the SD and the saturation. The task can in turn become even more complicated when different sugars are detected in a mixture. Machine learning, which aims to build computerized algorithms that can learn from data rather than focusing on programming, is an important branch of artificial intelligence research. 53,54 Machine learning has already been widely used in previous reports of nanopore research. 32,33,55-60Existing sensing data for D-fructose, D-galactose, D-mannose, D-glucose, and L-sorbose demonstrated highly distinguishable event features from one another, and the high degree of consistency when testing the same sugar type forms the basis for automated event classification using machine learning. The entire machine learning training process includes feature extraction, model training, and model construction. First, nanopore measurements with MspA-PBA were performed using D-fructose (Figures 2a-c, 14, 15), D-galactose (Figures 2d-f, 16, 17), D-mannose (Figures 2g-i, 18, 19), D-glucose (Figures 2j-l, 20-21), or L-sorbose (Figures 1e-h, 9, 13). Events were then extracted from the original current-time traces of sugar sensing to obtain corresponding event features, and the average values (ΔI / I p ), standard deviation (SD), skewness (skew), kurtosis (kurt), minimum (min), maximum (max), peak-to-peak value (pk), median (med), and dwell time (h). Each event tag is assigned to the glycotype being tested (Figure 3a). However, machine learning does not assign different tags to different event types for the same glycotype being tested. For each glycotype, continuous measurements are performed for more than 1 hour to collect enough events for each class. However, events with a dwell time of less than 30 ms are ignored. A minimum of 5,000 events are then collected for each class to form a database.
[0202] Before model training, 1,000 events of each sugar type were randomly selected from the database to assemble a dataset. The dataset was then divided into a training set (80%) for model training and a test set (20%) for model testing. Six common machine learning models were evaluated, including KNN, Xgboost, regression tree (CART), SVM, gradient boosting (GBDT), and random forest. All model evaluations were performed using default hyperparameters. To avoid bias, 10-fold cross-validation was applied during model training and evaluation, from which the validation accuracy of each model was derived and reported (Figure 3a). Specifically, validation accuracy is defined as the percentage of correctly identified events across the entire validation set. Typically, all models exhibited satisfactory performance by reporting a minimum accuracy score of 0.945, indicating that the data quality for sugar detection was sufficient for sugar identification. Specifically, the random forest model reported the highest validation accuracy of 0.974, demonstrating the best performance. Therefore, the models were further tuned by hyperparameter optimization. After fine-tuning the hyperparameters "n_estimators" and "max_Depth", the validation accuracy improves to 0.975. The feature importance of the fine-tuned random forest model is shown in Figure 3b, where all parameters play a significant role. However, the feature contributions of the mean (ΔI / I0), median (med), and standard deviation (SD) are the largest.
[0203] The fine-tuned model was further applied to the test set to generate a confusion matrix (Figure 3c). The accuracies for D-fructose (Fru), D-galactose (Gal), D-mannose (Man), D-glucose (Glc), and L-sorbose (L-sor) were 0.965, 1.000, 0.965, 0.970, and 0.990, respectively. D-galactose and L-sorbose show the highest scores among all five monosaccharides. To estimate the efficiency of model training, learning curves were generated, and accuracy scores for training sets of different sizes were given. The results show that when 508 events randomly selected from the entire training set were fed to the program, the overall judgment accuracy reached 95% (Figure 3d).
[0204] Nanopore measurements are then performed using a mixture of D-fructose, D-galactose, D-mannose, D-glucose, and L-sorbose. The previously trained model is used to predict the unlabeled events obtained from this measurement. Representative traces are shown in Figures 3e and 3f, where tags for events predicted by machine learning are labeled on the trace. The labeled tags match the types of events previously demonstrated when the corresponding sugar was tested as the sole analyte (Figures 1-2). A scatterplot of all events is also shown in Figure 22. After prediction by machine learning, events resulting from the binding of different sugars are clearly distinguishable from one another. The distribution of events for each sugar type in the scatterplot also matches the distribution of events when tested individually (Figures 13, 15, 17, 19, and 21).
[0205] All sugar types tested so far are hexoses. D-ribose and D-xylose are naturally occurring pentoses that are epimers of each other and, in principle, could also be detected by MspA-PBA. Experimentally, both D-ribose (Figures 4a-c) and D-xylose (Figures 4d-f) report multiple types of events. A detailed description of the event types and consistency of repeated measurements is summarized in Figures 23-26. Although D-ribose and D-xylose have the same molecular weight, the event characteristics and event distribution patterns are highly discriminative.
[0206] In the subsequent demonstration, we use L-rhamnose (L-Rha) as a representative deoxy sugar and N-acetyl-D-galactosamine (GalNAc) as a representative amino sugar. Both types of sugars have substituted hydroxyl groups and their overall structures are significantly different from previously tested sugar types, demonstrating that they can also be easily distinguished by the nanopore. Measurements are performed similarly to those described above (Figures 1 and 2). When measured with MspA-PBA, L-rhamnose or N-acetyl-D-galactosamine are tested as analytes, respectively. Both L-rhamnose (Figures 4g and 4i) and N-acetyl-D-galactosamine (Figures 4j and 4l) report both types of events. A detailed description of the event types and the consistency between different tests is summarized in Figures 27 and 32. Specifically, N-acetyl-D-galactosamine reports the largest event amplitude. This may be due to its significantly larger molecular size compared to all other sugars tested.
[0207] We currently have nine classes of input data for machine learning, derived from D-fructose (Fru), D-galactose (Gal), D-mannose (Man), D-glucose (Glc), L-sorbose (L-Sor), D-ribose (Rib), D-xylose (Xyl), L-rhamnose (L-Rha), and N-acetyl-D-galactosamine (GalNAc) (Figure 5a). One thousand events from each class are again randomly selected to form the dataset. The previous six models are evaluated a second time using a relatively large database that now contains nine classes of events (Figure 31). The random forest model outperformed all other models and will be further fine-tuned. The learning curve and feature importance of the fine-tuned model are shown in Figure 31. The confusion matrix results for the test set are shown in Figure 5b, with all nine monosaccharides achieving accuracies higher than 0.915. Although the general prediction accuracy decreases slightly when more sugar types are included in the model, Gal, Man, L-Sor, Rib, L-Rha, and GalNAc all show extremely high accuracy scores above 0.965. Nanopore measurements using mixtures of all nine sugar types are performed as described above (Figures 1-4). Events from the mixture are collected and analyzed to determine the ΔI / I p We generate corresponding scatter plots against the SD of the event labels predicted by the trained machine learning classifier (Figure 5c, Figure 32). The labeled scatter plots show distinguishable sugar-sensing events, consistent with the results of individual tests on each sugar type. Representative traces of sugar sensing in the mixture are also shown in Figure 5d-f, with the corresponding tags predicted by machine learning labeled.
[0208] conclusion In summary, we have already demonstrated the direct identification of nine monosaccharides using PBA-attached heterooctameric MspA. To our knowledge, heterooctameric MspA containing individually attached chemically reactive groups has never been reported as a nanoreactor. By generating a large event amplitude and abundant event features during sugar sensing, MspA exhibits excellent single-molecule sugar identification performance. 31-34,39 Experimental46,57,61 and theoretical evaluation 62,63 According to the study, the geometry of the conical lumen of MspA is the primary contributor to this superior resolution. Discrimination between sugar isomers or epimers was also demonstrated, further demonstrating the structural advantage of MspA in sugar identification. The extracted event features were fed into a machine learning-based classifier, which reported an accuracy of 0.96. Some specific sugar types, such as GalNAc, were even reported with an accuracy score of 0.99. Although only representative monosaccharides were demonstrated, this sensing strategy could in principle be applied to other types of monosaccharides, as long as the analyte can interact with the PBA and fit the size of the pore constriction region. 36-38 , sugar derivative 36 , glycodrugs, or low molecular weight oligosaccharides 64,65 It also applies to 35,66 and recent reports on the use of programmable nanopore reactors. 48 According to the authors, other polyols such as glycerol, vitamins, catechol, catecholamines, and nucleotide analogues can also be detected in MspA-PBA and will be reported in subsequent studies.
[0209] method 1. Production of Homooctameric MspA The genes encoding M2 MspA-D16H6 and N90C MspA-H6 (Table 1) were custom synthesized by GenScript Biotech Corporation (New Jersey). These two genes were inserted between the NdeI and HindIII restriction sites of pET-30a(+) plasmid DNA. The constructed plasmids, designated pET-M2 MspA and pET-N90C, were used to produce the homooctamer M2 MspA-D16H6 and N90C-H6, respectively. During gel electrophoresis, the homooctamer M2 MspA-D16H6 and N90C-H6 served as standards (Figure 1b). The homooctamer M2 MspA-D16H6 also served as a representative nanopore that does not contain any reactive sites within the pore cavity (Figure 11). Experimentally, 1 μL (100 ng / μL) of any plasmid DNA was added to 100 μL of E. coli BL21(DE3)pLysS competent cells (Sangon Biotech (Shanghai) Co., Ltd.) in an Eppendorf tube and shaken to ensure uniform distribution. The tube was incubated on ice for 30 minutes, then at 42°C for 90 seconds, and then on ice for another 3 minutes. 800 μL of Luria-Bertani (LB) medium was then added to the tube. The medium was then incubated at 37°C and 175 rpm for 50 minutes. The medium was then evenly spread onto an agar plate containing 30 μg / mL kanamycin sulfate and 34 μg / mL chloramphenicol and cultured at 37°C for 18 hours. Collect a single colony and add it to a 250 mL Erlenmeyer flask containing 100 mL of LB liquid medium containing 30 μg / mL kanamycin sulfate and 34 μg / mL chloramphenicol. Incubate the Erlenmeyer flask at 37 °C until the OD 600 The cells were then induced by adding isopropyl β-D-thiogalactopyranoside (IPTG) to a final concentration of 0.5 mM and shaken (175 rpm) at 16°C for 16 hours. The cells were then harvested by centrifugation (4500 rpm, 4°C, 20 minutes). The bacterial pellet was resuspended in 40 mL of lysis buffer (100 mM NaHPO / NaHPO, 0.1 mM EDTA, 150 mM NaCl, 0.5% (w / v) Genapol X-80, pH 6.5) and heated to 60°C for 10 minutes. The suspension was first cooled on ice for 10 minutes, then centrifuged at 13,000 rpm for 40 minutes at 4°C to collect the supernatant. The supernatant was then syringe filtered and loaded onto a nickel affinity column (HisTrap TMThe column is loaded onto a 1000-kJ / ml PBS (HP, GE Healthcare). The column is first eluted with buffer A (0.5 M NaCl, 20 mM HEPES, 5 mM imidazole, 0.5% (w / v) Genapol X-80, pH 8.0), followed by a linear gradient of imidazole (5 mM to 500 mM) by mixing buffer A and buffer B (0.5 M NaCl, 20 mM HEPES, 500 mM imidazole, 0.5% (w / v) Genapol X-80, pH 8.0). When purifying N90C MspA-H6, an additional 2 mM tris(2-carboxyethyl)phosphine (TCEP) is added to the buffer to prevent disulfide bond formation between cysteine residues in the homooctameric MspA. The eluted fractions are further characterized by SDS-polyacrylamide gel electrophoresis (PAGE) to identify fractions containing the target protein. A 4%–15% Mini-PROTEAN TGX gel (Bio-Rad, catalog number 4561083) is used for this step. Identified fractions can be used immediately or stored long-term at -80°C. 67 .
[0210] 2. Preparation of Heterogeneous Octameric MspA To produce heterogeneous octameric MspA consisting of M2 MspA-D16H6 and N90C MspA-H6, the two genes are co-located in the co-expression vector pETDuet-1. 68Briefly, the gene encoding N90C MspA-H6 was placed in the first multiple cloning site between the NcoI and HindIII restriction sites. The gene encoding M2 MspA-D16H6 was placed in the second multiple cloning site between the NdeI and BlpI restriction sites. A hexahistidine tag (H6) was engineered at the C-terminus of each gene to aid in nickel affinity chromatography-based purification. A tag consisting of 16 consecutive aspartic acids (D16) was added to the C-terminus of the gene encoding M2 MspA-D16H6, immediately adjacent to the hexahistidine tag (H6). The D16 tag was used to generate molecular weight differences between the heterogeneous octameric MspA fragments consisting of M2 MspA-D16H6 and the distinct fractions of N90C MspA-H6. Therefore, the D16 tag can be used to purify the heterogeneous oligomerized MspA consisting of the desired one N90C MspA-H6 and seven M2 MspA-D16H6, i.e., (N90C)1(M2)7 (Figure 1b).
[0211] 1 μL (100 ng / μL) of plasmid DNA was added to 100 μL of E. coli BL21(DE3)pLysS competent cells (Sangon Biotech (Shanghai) Co., Ltd.) in an Eppendorf tube and shaken to ensure uniform distribution. The tube was incubated on ice for 30 minutes, then at 42°C for 90 seconds, and then on ice for another 3 minutes. 800 μL of Luria-Bertani (LB) medium was then added to the tube. The medium was then incubated at 37°C and 175 rpm for 50 minutes. The medium was then evenly spread onto an agar plate containing 30 μg / mL kanamycin sulfate and 34 μg / mL chloramphenicol and cultured at 37°C for 18 hours. A single colony is then picked and added to a 50 mL test tube containing 10 mL of LB liquid medium containing 50 μg / mL ampicillin and 34 μg / mL chloramphenicol. The tube is incubated at OD 600Shake at 37°C and 175 rpm for 5 hours until OD = 0.7. Then, add the medium to the 1 L system and 600 The medium is further cultured at 37°C and 175 rpm until the RI = 0.6. IPTG is then added to the medium to reach a final concentration of 0.1 mM, and the medium is shaken at 16°C for 24 hours to induce protein overexpression. The cells are then harvested by centrifugation (4500 rpm, 20 minutes, 4°C).
[0212] The collected bacterial pellet was resuspended in 160 mL of lysis buffer (100 mM NaHPO / NaHPO, 0.1 mM EDTA, 150 mM NaCl, 0.5% (w / v) Genapol X-80, pH 6.5) and heated at 60°C for 50 minutes. The suspension was cooled on ice for 30 minutes and centrifuged at 13,000 rpm for 60 minutes at 4°C to collect the supernatant. The supernatant was syringe filtered and loaded onto a nickel affinity column (HisTrap TM The column was loaded with 1000kJ / ml of 1000kcal of 1000kJ / ml PBS (HP, catalog number 17-5248-01, GE Healthcare). The column was first eluted with buffer A (0.5 M NaCl, 20 mM HEPES, 5 mM imidazole, 2 mM TCEP, 0.5% (w / v) Genapol X-80, pH 8.0), followed by a linear gradient of imidazole (5 mM to 500 mM) by mixing buffer A and buffer B (0.5 M NaCl, 20 mM HEPES, 500 mM imidazole, 2 mM TCEP, 0.5% (w / v) Genapol X-80, pH 8.0).
[0213] All elution fractions are characterized by gel electrophoresis on a 4%-15% gradient SDS-polyacrylamide gel. Fractions corresponding to all heterogeneous MspA aggregates are collected for further purification. To further separate the desired (N90C)1(M2)7 pore type from the mixture, gel electrophoresis is performed on the collected fractions on a 10% SDS-polyacrylamide gel containing Tris-Gly buffer. A bias voltage of +160 V is continuously applied for 16 h at room temperature (rt) (Figure 1b). The gel is then stained with Coomassie Brilliant Blue (1.25 g Coomassie Brilliant Blue R250, 225 mL MeOH, 50 mL glacial AcOH, 225 mL ultrapure water) for 4 h. The gel is then destained using elution buffer (400 mL MeOH, 100 mL glacial AcOH, topped up to 1 L with ultrapure water) until protein bands are clearly visible. The gel was immersed in ultrapure water for 10 minutes for imaging. The gel fragment containing the band corresponding to the (N90C)1(M2)7 pore type was excised, crushed, and rehydrated in extraction solution (150 mM NaCl, 15 mM Tris-HCl, pH 7.5, 0.2% DDM, 0.5% Genapol X-80, 5 mM TCEP, 10 mM EDTA). The resulting suspension was left at room temperature for 12 hours, after which the supernatant was collected. The collected (N90C)1(M2)7 was either used immediately or stored at -80°C for long periods.
[0214] 3. Chemical modification of (N90C)1(M2)7. To chemically modify (N90C)1(M2)7, mix 5 μL of freshly prepared (N90C)1(M2)7 with 2.5 μL of 3-(maleimido)phenylboronic acid in DMSO (500 mM), and add 42.5 μL of 1.5 M KCl buffer (1.5 M KCl, 10 mM MOPS, pH 7.0). Allow the mixture to stand at room temperature for 10 min. The chemically modified (N90C)1(M2)7 is used immediately for all downstream nanopore measurements. For simplicity, unless otherwise specified, the modified heterogeneous MspA is referred to as MspA-PBA throughout this specification.
[0215] 4. Nanopore Measurements and Data Analysis This measurement device consists of two custom-made polyoxymethylene chambers separated by a polytetrafluoroethylene membrane approximately 20 μm thick and drilled (approximately 100 μm in diameter). Prior to measurement, the chambers were first treated with a 0.5% (v / v) hexadecane-in-pentane solution and then placed for pentane evaporation. 500 μL of electrolyte buffer was then added to the two chambers. Unless otherwise noted, the buffer used for all electrical recordings consisted of 1.5 M KCl and 10 mM MOPS (pH 7.0). Two custom-made Ag / AgCl electrodes connected to a patch-clamp amplifier were placed in the chambers in contact with the buffer. The electrically grounded chamber is typically defined as the cis chamber, and the opposite chamber as the trans chamber. After adding 100 μL of a pentane solution of DPhPC (5 mg / mL) to the two chambers, a lipid bilayer spontaneously formed when the electrolyte buffer in one of the chambers was manually swiped up and down several times. Once the bilayer was formed, the resulting current immediately dropped to 0 pA, indicating that the pore was now electrically sealed. MspA was added to the cis chamber to induce spontaneous pore insertion. After a single nanopore insertion, the buffer solution in the cis chamber was exchanged to prevent further pore insertion. To avoid interference from external electromagnetic and vibration noise, the device was shielded in a custom-made Faraday cage (34 cm x 23 cm x 15 cm) mounted on an air-levitated vibration-isolated optical platform (Jiangxi Liansheng Technology Co., Ltd.). All electrophysiological measurements were performed using an Axonpatch 200B patch-clamp amplifier paired with a Digidata 1550B digitizer (Molecular Devices). Unless otherwise noted, the applied voltage was +160 mV throughout all measurement periods. All measurements were performed at room temperature (23 °C). All single-channel recordings were sampled at 25 kHz and low-pass filtered with a corner frequency of 1 kHz. Sugar sensing is performed using a single MspA-PBA pore inserted into a planar lipid bilayer, and sugar analytes are added to the cis chamber prior to single-channel recording.All events are detected via the "single channel search" function in Clampfit 10.7. Subsequent analysis, including histogram plotting, scatter plot generation, and curve fitting, is performed in Origin Pro 2018.
[0216] 5. Event Feature Extraction For each class, results from three independent measurements are included. From the original time current trace, the start and end times of each event are identified via Clampfit 10.7. The asterisk and end time serve as indicators for splitting events from the original trace and are used to derive dwell time features for each event. The split event portions are used to extract other event features, including the mean, standard deviation, skewness, kurtosis, peak-to-peak minimum, maximum, and median.
[0217] Specifically, we calculated the average current amplitude before and after the start and end of each event to estimate the MspA-PBA pore-opening current (I p ) The event amplitude is ΔI = I p -I s is derived, where I s represents the average blocking level of each event. To avoid bias between pores, the relative current amplitude (ΔI / I p ) is considered as the average value for each event. ΔI / I p Events with a value less than 0.35 are collected for subsequent analysis. The extracted event features form a feature matrix. Only events with a duration greater than 30 ms are selected. For each glycotype, 1000 events are randomly selected to form a labeled dataset for model training and testing. To extract event features for model prediction, the above process is performed similarly, except that no event tags are assigned.
[0218] 6. Machine Learning The input data is randomly divided into a training set (80% of the tagged dataset) and a test set (20%) for model training and model testing. The data in the training set is first normalized and then applied to training six models, including KNN, Xgboost, regression tree (CART), SVM, gradient boosting (GBDT), and random forest. Random forest is selected for hyperparameter tuning based on 10-fold cross-validation. A confusion matrix is generated using the test set for model evaluation (Figures 3c and 5b). This model is saved to predict untagged data (Figures 22 and 32).
[0219] Author contributions SYZ and SH conceived the project. SYZ, ZYC, LYW, and KFW performed the measurements. PPF designed the machine learning algorithm. SYZ, YQW, YL, and SHY fabricated the MspA nanopore. PKZ installed the instrument. SH and SYZ wrote the paper. SYZ, YQW, and SHY prepared the supplementary video. WDJ, XYD, and CZH provided human-pleasing reviews. SH and HYC supervised the project.
[0220] Data availability declaration All data provided in this study may be obtained from the authors upon reasonable request.
[0221] Code availability declaration A custom machine learning algorithm is submitted as supplementary material and is named "Supplementary Material" and a brief introductory document is provided.
[0222] Competitive Interest Declaration SH and SYZ have already filed a patent describing heterogeneous MspA and its applications.
[0223] Acknowledgments The authors would like to thank Professor Hagan Bayley (University of Oxford) for his valuable advice in preparing the manuscript. The authors would like to thank Professors Zijian Guo, Shaolin Zhu, Congqing Zhu, Jie Li, and Ran Xie from Nanjing University. This project is financially supported by the National Natural Science Foundation of China (Grant Nos. 31972917, 91753108, 21675083), and by the Special Fund for Basic Science Research Work of Central High Schools (Grant Nos. 020514380257, 020514380261), Jiangsu Province High-Level Innovation Entrepreneur Introduction Plan (Individual and Group Projects), Jiangsu Province Natural Science Foundation (Grant No. BK20200009), Nanjing University Outstanding Research Program (Grant No. ZYJH004), Shanghai Municipal-Level Science and Technology Key Project, State Key Laboratory of Life Analytical Chemistry (Grant No. 5431ZZXM1902), and Nanjing University Technology Innovation Fund Project.
[0224] References 1 Varki, A. & Kornfeld, S. in Essentials of Glycobiology [Internet]. 3rd edition. Vol. Chapter 1. (eds A. Varki et al.) (Cold Spring Harbor Laboratory Press; 2015-2017., 2017). 2 Dashty, M. A quick look at biochemistry: Carbohydrate metabolism. Clin. Biochem. 46, 1339-1352, doi:https: / / doi.org / 10.1016 / j.clinbiochem.2013.04.027 (2013). 3 Zeng, Y., Himmel, M. E. & Ding, S.-Y. Visualizing chemical functionality in plant cell walls. Biotechnol. Biofuels 10, 263, doi:10.1186 / s13068-017-0953-3 (2017). 4 Matsuura, M. Structural Modifications of Bacterial Lipopolysaccharide that Facilitate Gram-Negative Bacteria Evasion of Host Innate Immunity. Front. Immunol. 4, doi:10.3389 / fimmu.2013.00109 (2013). 5 Varki, A. Biological roles of oligosaccharides: all of the theories are correct. Glycobiology 3, 97-130, doi:10.1093 / glycob / 3.2.97 (1993). 6 Haltiwanger, R. S. & Lowe, J. B. Role of Glycosylation in Development. Annu. Rev. Biochem. 73, 491-537, doi:10.1146 / annurev.biochem.73.011303.074043 (2004). 7 Flynn, R. A. et al. Small RNAs are modified with N-glycans and displayed on the surface of living cells. Cell 184, 3109-3124.e3122, doi:https: / / doi.org / 10.1016 / j.cell.2021.04.023 (2021). 8 Reily, C., Stewart, T. J., Renfrow, M. B. & Novak, J. Glycosylation in health and disease. Nat. Rev. Nephrol. 15, 346-366, doi:10.1038 / s41581-019-0129-4 (2019). 9 Moremen, K. W., Tiemeyer, M. & Nairn, A. V. Vertebrate protein glycosylation: diversity, synthesis and function. Nat. Rev. Mol. Cell Biol. 13, 448-462, doi:10.1038 / nrm3383 (2012). 10 Puvirajesinghe, T. M. & Turnbull, J. E. Glycoarray Technologies: Deciphering Interactions from Proteins to Live Cell Responses. Microarrays 5, doi:10.3390 / microarrays5010003 (2016). 11 Hu, S. & Wong, D. T. Lectin microarray. Proteomics: Clin. Appl. 3, 148-154, doi:10.1002 / prca.200800153 (2009). 12 Mantovani, V., Galeotti, F., Maccari, F. & Volpi, N. Recent advances in capillary electrophoresis separation of monosaccharides, oligosaccharides, and polysaccharides. Electrophoresis 39, 179-189, doi:10.1002 / elps.201700290 (2018). 13 Rovio, S., Simolin, H., Koljonen, K. & Siren, H. Determination of monosaccharide composition in plant fiber materials by capillary zone electrophoresis. J. Chromatogr. A 1185, 139-144, doi:https: / / doi.org / 10.1016 / j.chroma.2008.01.031 (2008). 14 Nagy, G., Peng, T. & Pohl, N. L. B. Recent Liquid Chromatographic Approaches and Developments for the Separation and Purification of Carbohydrates. Anal. Methods 9, 3579-3593, doi:10.1039 / C7AY01094J (2017). 15 Vreeker, G. C. M. & Wuhrer, M. Reversed-phase separation methods for glycan analysis. Anal. Bioanal. Chem. 409, 359-378, doi:10.1007 / s00216-016-0073-0 (2017). 16 Lundborg, M., Fontana, C. & Widmalm, G. Automatic Structure Determination of Regular Polysaccharides Based Solely on NMR Spectroscopy. Biomacromolecules 12, 3851-3855, doi:10.1021 / bm201169y (2011). 17 Fontana, C., Kovacs, H. & Widmalm, G. NMR structure analysis of uniformly 13C-labeled carbohydrates. J. Biomol. NMR 59, 95-110, doi:10.1007 / s10858-014-9830-6 (2014). 18 Veillon, L. et al. Characterization of isomeric glycan structures by LC-MS / MS. Electrophoresis 38, 2100-2114, doi:https: / / doi.org / 10.1002 / elps.201700042 (2017). 19 Zhou, S., Veillon, L., Dong, X., Huang, Y. & Mechref, Y. Direct comparison of derivatization strategies for LC-MS / MS analysis of N-glycans. Analyst 142, 4446-4455, doi:10.1039 / c7an01262d (2017). 20 Gray, C. J. et al. Advancing Solutions to the Carbohydrate Sequencing Challenge. J. Am. Chem. Soc. 141, 14463-14479, doi:10.1021 / jacs.9b06406 (2019). 21 Aretz, I. & Meierhofer, D. Advantages and Pitfalls of Mass Spectrometry Based Metabolome Profiling in Systems Biology. Int. J. Mol. Sci. 17, 632, doi:10.3390 / ijms17050632 (2016). 22 Emwas, A. H. The strengths and weaknesses of NMR spectroscopy and mass spectrometry with particular focus on metabolomics research. Methods Mol. Biol. 1277, 161-193, doi:10.1007 / 978-1-4939-2377-9_13 (2015). 23 Morimoto, K. et al. GlycanAnalysis Plug-in: a database search tool for N-glycan structures using mass spectrometry. Bioinformatics 31, 2217-2219, doi:10.1093 / bioinformatics / btv110 (2015). 24 Walsh, I. et al. GlycanAnalyzer: software for automated interpretation of N-glycan profiles after exoglycosidase digestions. Bioinformatics 35, 688-690, doi:10.1093 / bioinformatics / bty681 (2019). 25 Manrao, E. A. et al. Reading DNA at single-nucleotide resolution with a mutant MspA nanopore and phi29 DNA polymerase. Nat. Biotechnol. 30, 349-353, doi:10.1038 / nbt.2171 (2012). 26 Yan, S. et al. Direct sequencing of 2'-deoxy-2'-fluoroarabinonucleic acid (FANA) using nanopore-induced phase-shift sequencing (NIPSS). Chem. Sci. 10, 3110-3117, doi:10.1039 / c8sc05228j (2019). 27 Zhang, J. et al. Direct microRNA Sequencing Using Nanopore-Induced Phase-Shift Sequencing. iScience 23, doi:10.1016 / j.isci.2020.100916 (2020). 28 Yan, S. et al. Single Molecule Ratcheting Motion of Peptides in a Mycobacterium smegmatis Porin A (MspA) Nanopore. Nano Lett. 21, 6703-6710, doi:10.1021 / acs.nanolett.1c02371 (2021). 29 Brinkerhoff, H., Kang, A. S. W., Liu, J., Aksimentiev, A. & Dekker, C. Infinite re-reading of single proteins at single-amino-acid resolution using nanopore sequencing. bioRxiv, 2021.2007.2013.452225, doi:10.1101 / 2021.07.13.452225 (2021). 30 Stylianopoulos, C. in Encyclopedia of Human Nutrition (Third Edition) (ed Benjamin Caballero) 265-271 (Academic Press, 2013). 31 Karawdeniya, B. I., Bandara, Y., Nichols, J. W., Chevalier, R. B. & Dwyer, J. R. Surveying silicon nitride nanopores for glycomics and heparin quality assurance. Nat. Commun. 9, 3278, doi:10.1038 / s41467-018-05751-y (2018). 32 Im, J., Lindsay, S., Wang, X. & Zhang, P. Single Molecule Identification and Quantification of Glycosaminoglycans Using Solid-State Nanopores. ACS Nano 13, 6308-6318, doi:10.1021 / acsnano.9b00618 (2019). 33 Xia, K. et al. Synthetic heparan sulfate standards and machine learning facilitate the development of solid-state nanopore analysis. Proc. Natl. Acad. Sci. U. S. A. 118, doi:10.1073 / pnas.2022806118 (2021). 34 Cai, Y. et al. A solid-state nanopore-based single-molecule approach for label-free characterization of plant polysaccharides. Plant Commun. 2, 100106, doi:10.1016 / j.xplc.2020.100106 (2021). 35 Guo, Z., Shin, I. & Yoon, J. Recognition and sensing of various species using boronic acid derivatives. Chem. Commun. (Cambridge, U. K.) 48, 5956-5967, doi:10.1039 / c2cc31985c (2012). 36 Wu, X. et al. Selective sensing of saccharides using simple boronic acids and their aggregates. Chem. Soc. Rev. 42, 8032-8048, doi:10.1039 / c3cs60148j (2013). 37 Peters, J. A. Interactions between boric acid derivatives and saccharides in aqueous media: Structures and stabilities of resulting esters. Coord. Chem. Rev. 268, 1-22, doi:https: / / doi.org / 10.1016 / j.ccr.2014.01.016 (2014). 38 van den Berg, R., Peters, J. A. & van Bekkum, H. The structure and (local) stability constants of borate esters of mono- and di-saccharides as studied by 11B and 13C NMR spectroscopy. Carbohydr. Res. 253, 1-12, doi:10.1016 / 0008-6215(94)80050-2 (1994). 39 Ramsay, W. J. & Bayley, H. Single-Molecule Determination of the Isomers of d-Glucose and d-Fructose that Bind to Boronic Acids. Angew. Chem., Int. Ed. Engl. 57, 2841-2845, doi:10.1002 / anie.201712740 (2018). 40 Butler, T. Z., Pavlenok, M., Derrington, I. M., Niederweis, M. & Gundlach, J. H. Single-molecule DNA detection with an engineered MspA protein nanopore. Proc. Natl. Acad. Sci. U. S. A. 105, 20647, doi:10.1073 / pnas.0807514106 (2008). 41 Niederweis, M. et al. Cloning of the mspA gene encoding a porin from Mycobacterium smegmatis. Mol. Microbiol. 33, 933-945, doi:10.1046 / j.1365-2958.1999.01472.x (1999). 42 Faller, M., Niederweis, M. & Schulz, G. E. The structure of a mycobacterial outer-membrane channel. Science 303, 1189-1192, doi:10.1126 / science.1094114 (2004). 43 Laszlo, A. H. et al. Detection and mapping of 5-methylcytosine and 5-hydroxymethylcytosine with nanopore MspA. Proc. Natl. Acad. Sci. U. S. A. 110, 18904-18909, doi:10.1073 / pnas.1310240110 (2013). 44 Wang, Y. et al. Nanopore Sequencing Accurately Identifies the Mutagenic DNA Lesion O(6) -Carboxymethyl Guanine and Reveals Its Behavior in Replication. Angew. Chem., Int. Ed. Engl. 58, 8432-8436, doi:10.1002 / anie.201902521 (2019). 45 Ma, F. et al. Nanopore Sequencing Accurately Identifies the Cisplatin Adduct on DNA. ACS Sens. 6, 3082-3092, doi:10.1021 / acssensors.1c01212 (2021). 46 Cao, J. et al. Giant single molecule chemistry events observed from a tetrachloroaurate(III) embedded Mycobacterium smegmatis porin A nanopore. Nat. Commun. 10, 5668, doi:10.1038 / s41467-019-13677-2 (2019). 47 Wang, S. et al. Single molecule observation of hard-soft-acid-base (HSAB) interaction in engineered Mycobacterium smegmatis porin A (MspA) nanopores. Chem. Sci. 11, 879-887, doi:10.1039 / c9sc05260g (2019). 48 Jia, W. et al. Programmable Nano-Reactors for Stochastic Sensing. Nat. Commun., doi:10.1038 / s41467-021-26054-9 (2021). 49 Choi, L.-S. & Bayley, H. S-Nitrosothiol Chemistry at the Single-Molecule Level. Angew. Chem., Int. Ed. Engl. 51, 7972-7976, doi:https: / / doi.org / 10.1002 / anie.201202365 (2012). 50 Lehmacher, A. & Bockemuehl, J. l-Sorbose utilization by virulent Escherichia coli and Shigella: Different metabolic adaptation of pathotypes. Int. J. Med. Microbiol. 297, 245-254, doi:https: / / doi.org / 10.1016 / j.ijmm.2007.01.007 (2007). 51 Sugisawa, T., Miyazaki, T. & Hoshino, T. Microbial Production of L-Ascorbic Acid from D-Sorbitol, L-Sorbose, L-Gulose, and L-Sorbosone by Ketogulonicigenium vulgare DSM 4025. Biosci., Biotechnol., Biochem. 69, 659-662, doi:10.1271 / bbb.69.659 (2005). 52 Adair, W. L. in xPharm: The Comprehensive Pharmacology Reference (eds S. J. Enna & David B. Bylund) 1-12 (Elsevier, 2007). 53 Deo, R. C. Machine Learning in Medicine. Circulation 132, 1920-1930, doi:10.1161 / CIRCULATIONAHA.115.001593 (2015). 54 Diaz Carral, A., Ostertag, M. & Fyta, M. Deep learning for nanopore ionic current blockades. J. Chem. Phys. 154, 044111, doi:10.1063 / 5.0037938 (2021). 55 Schreiber, J. et al. Error rates for nanopore discrimination among cytosine, methylcytosine, and hydroxymethylcytosine along individual DNA strands. Proc. Natl. Acad. Sci. U. S. A. 110, 18910, doi:10.1073 / pnas.1310615110 (2013). 56 Misiunas, K., Ermann, N. & Keyser, U. F. QuipuNet: Convolutional Neural Network for Single-Molecule Nanopore Sensing. Nano Lett. 18, 4040-4045, doi:10.1021 / acs.nanolett.8b01709 (2018). 57 Wang, Y. et al. Structural-profiling of low molecular weight RNAs by nanopore trapping / translocation using Mycobacterium smegmatis porin A. Nat. Commun. 12, 3368, doi:10.1038 / s41467-021-23764-y (2021). 58 Doroschak, K. et al. Rapid and robust assembly and decoding of molecular tags with DNA-based nanopore signatures. Nat. Commun. 11, 5454-5454, doi:10.1038 / s41467-020-19151-8 (2020). 59 Wei, Z.-X. et al. Learning Shapelets for Improving Single-Molecule Nanopore Sensing. Anal. Chem. 91, 10033-10039, doi:10.1021 / acs.analchem.9b01896 (2019). 60 Sui, X.-J. et al. Aerolysin Nanopore Identification of Single Nucleotides Using the AdaBoost Model. J. Anal. Test. 3, 134-139, doi:10.1007 / s41664-019-00088-x (2019). 61 Liu, Y. et al. Allosteric Switching of Calmodulin in a Mycobacterium smegmatis porin A (MspA) Nanopore-Trap. Angew. Chem., Int. Ed. Engl. n / a, doi:https: / / doi.org / 10.1002 / anie.202110545 (2021). 62 Zhou, W., Qiu, H., Guo, Y. & Guo, W. Molecular Insights into Distinct Detection Properties of α-Hemolysin, MspA, CsgG, and Aerolysin Nanopore Sensors. J. Phys. Chem. B 124, 1611-1618, doi:10.1021 / acs.jpcb.9b10702 (2020). 63 Yu, M. et al. Unveiling the Microscopic Mechanism of Current Variation in the Sensing Region of the MspA Nanopore for DNA Sequencing. J. Phys. Chem. Lett. 12, 9132-9141, doi:10.1021 / acs.jpclett.1c02414 (2021). 64 Ma, Q., Zhao, X., Shi, A. & Wu, J. Bioresponsive Functional Phenylboronic Acid-Based Delivery System as an Emerging Platform for Diabetic Therapy. Int. J. Nanomed. 16, 297-314, doi:10.2147 / IJN.S284357 (2021). 65 Cambre, J. N. & Sumerlin, B. S. Biomedical applications of boronic acid polymers. Polymer 52, 4631-4643, doi:https: / / doi.org / 10.1016 / j.polymer.2011.07.057 (2011). 66 Bull, S. D. et al. Exploiting the Reversible Covalent Bonding of Boronic Acids: Recognition, Sensing, and Assembly. Acc. Chem. Res. 46, 312-326, doi:10.1021 / ar300130w (2013). 67 Zhang, J. et al. Mapping Potential Engineering Sites of Mycobacterium smegmatis porin A (MspA) to Form a Nanoreactor. ACS Sens. 6, 2449-2456, doi:10.1021 / acssensors.1c00792 (2021). 68 Pavlenok, M. & Niederweis, M. Hetero-oligomeric MspA pores in Mycobacterium smegmatis. FEMS Microbiol. Lett. 363, doi:10.1093 / femsle / fnw046 (2016). Supplementary Information
[0225] Materials 1,2-Diphytanoyl-sn-glycero-3-phosphocholine (DPhPC) was obtained from Avanti Polar Lipids. Pentane, hexadecane, tris(2-carboxyethyl)phosphine hydrochloride (TCEP), ethylenediaminetetraacetic acid (EDTA), Genapol X-80, ammonium persulfate (≥98%), sodium dodecyl sulfate (≥98.5%), N,N,N',N'-tetramethylethylenediamine (99%), and a 30% solution of acrylamide / bis-acrylamide were obtained from Sigma-Aldrich. Potassium chloride, sodium chloride (99.99%), sodium hydroxide (99.9%), disodium hydrogen phosphate, and sodium dihydrogen phosphate were obtained from Shanghai Aladdin Biochemical Technology Co., Ltd. (China). Hydrochloric acid (HCl) was obtained from Sinopharm Group Co., Ltd. (China). 4-(2-Hydroxyethyl)-1-piperazineethanesulfonic acid (HEPES) was obtained from Shanghai Yuanye Bio-Technology Co., Ltd. (China). Dioxane-free isopropyl-β-D-thiogalactopyranoside (IPTG), kanamycin sulfate, imidazole, and tris(hydroxymethyl)aminomethane (Tris) were obtained from Beijing Solebao Technology Co., Ltd. SDS-PAGE electrophoresis buffer powder was obtained from Shanghai Beyotime Biotechnology Co., Ltd. (China). Precision Plus Protein TM Two-color standard product, TGX TM FastCast TMAmide kit (4%-15%), stacking gel buffer (0.5 M Tris-HCl buffer, pH 6.8), and separating gel buffer (1.5 M Tris-HCl buffer, pH 8.8) were obtained from Bio-Rad. LB broth and LB agar were obtained from Hopebiol (China). 3-(Maleimido)phenylboronic acid (MPBA, catalog number sc-352346) was obtained from Santa Cruz Biotechnology (Shanghai) Co., Ltd. All the above-mentioned products were used as received.
[0226] D-(+)-Mannose (≥99%) was obtained from Sigma-Aldrich. D-(+)-Glucose (99%) was obtained from Shanghai Titan Technology Co., Ltd. (China). D-(+)-Galactose (98%), D-(+)-Xylose (98%), L-Rhamnose monohydrate (99%), D-(-)-Ribose (≥99%), and N-acetyl-D-galactosamine (98%) were obtained from Shanghai Aladdin Biochemical Technology Co., Ltd. (China). L-(-)-Sorbose (98%) was obtained from Shanghai Macklin Biochemical Technology Co., Ltd. (China). D-(-)-Fructose (≥98%) was obtained from Shanghai Dibai Bio-Technology (China).
[0227] 1.5 M KCl buffer (1.5 M KCl, 10 mM MOPS, pH 7.0), lysis buffer (100 mM NaHPO / NaHPO, 0.1 mM EDTA, 150 mM NaCl, 0.5% (w / v) Genapol X-80, pH 6.5), buffer A (0.5 M NaCl, 20 mM HEPES, 5 mM imidazole, 0.5% (w / v) Genapol X-80, pH 8.0), and buffer B (0.5 M NaCl, 20 mM HEPES, 500 mM imidazole, 0.5% (w / v) Genapol X-80, pH 8.0) were prepared using Milli-Q water and filtered through a 0.2 μM Whatman membrane.
[0228] [Table 1]
[0229] [Table 2]
[0230] [Table 3]
[0231] [Table 4]
[0232] [Table 5]
[0233] [Table 6]
[0234] References 1 Ramsay, WJ & Bayley, H. Single-Molecule Determination of the Isomers of d-Glucose and d-Fructose that Bind to Boronic Acids. Angew. Chem., Int. Ed. Engl. 57, 2841-2845, doi:10.1002 / anie.201712740 (2018). 2 Alcock, LJ, Perkins, MV & Chalker, JM Chemical methods for mapping cysteine oxidation. Chem. Soc. Rev. 47, 231-268, doi:10.1039 / c7cs00607a (2018). 3 Shin, SH, Luchian, T., Cheley, S., Braha, O. & Bayley, H. Kinetics of a Reversible Covalent-Bond-Forming Reaction Observed at the Single-Molecule Level. Angew. Chem., Int. Ed. Engl. 41, 3707-3709 (2002).
[0235] Example 2: Identification of nucleoside monophosphates and their epigenetic modifications using engineered nanopores Chemical modifications of RNA play a crucial role in regulating RNA biological processes and are associated with many human diseases. However, direct identification of RNA modifications by sequencing remains challenging. Nanopore sequencing may offer a promising solution by directly detecting sequence modifications, but currently available strand sequencing strategies remain complicated by sequence decoding. Alternatively, sequential nanopore identification of enzymatically cleaved nucleoside monophosphates (NMPs) may simultaneously provide accurate sequence and modification information. To this end, we produced a heterogeneous octamer of Mycobacterium smegmatisporin A (MspA) modified with phenylboronic acid (PBA). This MspA cleaved all four classical NMPs, including 5-methylcytidine (mCytidine), and 5-methylcytidine (mCytidine), into the RNA. 5 C), N6-methyladenosine (m 6 A), N7-methylguanosine (m 7 G), N1-methyladenosine (m 1 Direct discrimination against nucleotides (A), inosine (I), pseudouridine (Ψ), and dihydrouridine (D) is achieved. A custom machine learning algorithm was also developed and found to provide a typical accuracy score of 0.996. This method is applied to the quantitative analysis of base modifications in microRNAs and tRNAs. It is typically suitable for sensing various nucleoside or nucleotide derivatives and may bring new insights to epigenetic RNA sequencing.
[0236] introduction Many RNA modifications are chemical modifications caused by enzymes, such as methylation, deamination, reduction and thiolation, or isomerization of ribose or nucleotides to nucleic acid bases. These modifications are carried out by specialized transcriptional proteins during the post-transcriptional stage. According to the MODOMICS database, there are approximately 170 known RNA modifications. 1 , various biological processes, e.g., gene recoding 2 , pre-mRNA splicing 3 , mRNA export 4 , RNA folding 5, and regulation of chromatin states 6 Accumulating evidence suggests that a large number of RNA modifications are necessary for cancer. 7, 8 , neuropathy 9 , and other human diseases 10 Recent reports have shown that RNA modifications are also associated with grain yield. 11 However, there is an urgent and unmet need to accurately map various RNA modifications, which is complicated by the similarity of their chemical structures. 12 .
[0237] Analysis of RNA modifications was performed using thin layer chromatography (TLC) 13 , High-Performance Liquid Chromatography-Ultraviolet Spectrophotometry (HPLC-UV) 14 , or high performance liquid chromatography-mass spectrometry (HPLC-MS) 15 These methods can simultaneously measure a large number of RNA modifications but cannot provide sequence information. Next-generation sequencing (NGS)-based methods measure RNA modifications across the entire transcriptome. 16 Although it is possible to map the modified RNA fragments, antibodies that immunoprecipitate the modified RNA fragments are not 17 or RNA modifications that lead to mutations or truncations during cDNA production. 18 These methods are usually customized for only one specific modification, and lack antibodies or chemical reagents that can detect all RNA modifications, limiting the types of modifications that can be detected by sequencing. 19 , m 6 A 20, 21 , m 5 C 22 , m 1 A 23 , m 7 G 24 , 5-hydroxymethylcytosine (5hmC) 25 , N6,2'-O-dimethyladenosine (m6Am) 17 , N4-acetylcytidine (ac4C) 26, and A~I Editor 27 Third-generation sequencing technologies, including methods developed by Pacific Biosciences (PacBio) or Oxford Nanopore Technologies (ONT), overcome these drawbacks by performing direct RNA sequencing. 28 In PacBio sequencing, RNA modifications are identified by observing the time change between base incorporations. 29 On the other hand, the nanopore sequencing provided by ONT is based on the ionic current 30, 31 or event residence time 32 However, strand sequencing strategies report RNA modifications by identifying changes in 33 is limited by a spatial resolution equivalent to an average read of approximately 5 nucleotides 34 This still leaves sequencing vulnerable to discriminating between all epigenetic modifications, a situation that becomes even more severe when modified nucleotides are in close proximity. 35 .
[0238] Exonuclease sequencing is a different strategy in which nucleotides degraded by nucleic acid exonucleases are sequentially read using a nanopore sensor. However, this requires the presence of a high-resolution nanopore that can unambiguously recognize all nucleotides and their primary modifications. α-Hemolysin (α-HL) embedded in cyclodextrin. 36, 37 has previously been reported to accomplish this task, but the results show insufficient resolution, e.g., an inability to truly distinguish between cytidine diphosphate (CDP) and uridine diphosphate (UDP). Identification of RNA modifications has also not been demonstrated. 36 This low-resolution sensing is due to the geometry of the cylindrical lumen of α-HL. 38 In contrast, Mycobacterium smegmatisporin A (MspA) 39 may be more advantageous, and the MspA may be useful for nanopore sequencing. 40, Single-molecule chemistry 41 and structural analysis of biopolymers 42, 43 Phenylboronic acid (PBA) is known to form reversible covalent bonds with 1,2- or 1,3-diols. 44 Previously, the introduction of PBA into nanopore cavities has already been demonstrated. 45 , epinephrine, and remdesivir 46 However, a heterogeneous octameric MspA nanopore containing a single PBA adaptor has not been reported, nor has nanopore identification of various epigenetically modified NMPs been reported.
[0239] Identification of nucleoside monophosphates (NMPs) using PBA-modified MspA To construct heterogeneous octameric MspA, two distinct genes encoding N90C-MspA-H6 and M2-MspA-D16H6 (Table 7) were custom synthesized and simultaneously inserted into the pETDuet-1 coexpression vector (methods described in Example 2). Specifically, N90C-MspA-H6 encodes an MspA monomer with a unique cysteine in the pore-constricting region. Meanwhile, M2-MspA-D16H6 encodes a monomer that does not contain any cysteines. Heterogeneous octameric MspA, consisting of the distinct parts of the two gene expression products, was produced by prokaryotic coexpression (Figure 39) and characterized by gel electrophoresis (Figures 40-41). The heterooctameric MspA consisting of one unit N90C-MspA-H6 and seven units M2-MspA-D16H6 is the only desired MspA assembly and is designated (N90C)1(M2)7 (Figure 33a). (N90C)1(M2)7 is separated from other MspA heterooctamer structures by high-resolution gel electrophoresis followed by gel extraction (methods in Example 2, Figures 40-41). 3-(maleimido)phenylboronic acid (MPBA) is then reacted with the only cysteine in (N90C)1(M2)7 (Figure 33b). The reaction is observed in real time at the single-molecule level by single-channel recording in a buffer solution of 1.5 M KCl, 10 mM MOPS, pH 7.0 (Figure 33c, methods in Example 2). When a single (N90C)1(M2)7 is inserted into the membrane and a bias voltage of +200 mV is continuously applied, the measured pore opening current of (N90C)1(M2)7(I0) is approximately 620 pA. 47 As shown, an upward noise caused by the cysteine residues in the pore constriction region is also observed at this stage. Immediately after adding MPBA to the cis chamber at a final concentration of 1 mM, a single current drop of approximately 100 pA is observed. The previously observed upward noise also disappears at the same time, indicating that the cysteine residues are occupied and the PBA modification of the pore constriction region is successful. By adding a higher concentration of MPBA, the reaction caused by the diffusion of MPBA into the pore constriction region can be accelerated. For simplicity, this PBA-modified MspA is referred to as MspA-PBA. Under the same conditions, MspA-PBA (Ip The measured pore opening current of ) is approximately 520 pA (Figure 33c).
[0240] MspA-PBA can also be prepared integrally by mixing (N90C)1(M2)7 with MPBA (method in Example 2). Unless otherwise noted, all subsequent measurements are performed using the integrally prepared MspA-PBA. After adding the integrally prepared MspA-PBA to the cis chamber, spontaneous pore insertion was observed, confirming that the high pore-forming activity of MspA-PBA was fully retained (Figure 42). The statistical results, measuring pore-gating currents of (N90C)1(M2)7 and MspA-PBA at 623 ± 13 (mean ± FWHM) pA and 510 ± 14 (mean ± FWHM) pA (Figure 42), are consistent with the pore-gating currents previously measured in single-channel recordings (Figure 1c). Figure 43 further displays the IV curves of (N90C)1(M2)7 and MspA-PBA obtained using different concentrations of KCl (0.15 M to 2 M KCl). Based on the slope of the IV curve, the conductance of MspA-PBA measured in 1.5 M KCl buffer was estimated to be approximately 2.91 nS, which is a sufficiently large value.
[0241] NMP consists of ribose, phosphate group, and nucleic acid base as a monomer unit of RNA. The presence of cis-diol in ribose gives NMP an affinity for PBA. 48These signals can be directly detected by MspA-PBA. Experimentally, single-channel recordings were performed using MspA-PBA in 1.5 M KCl buffer (1.5 M KCl, 10 mM MOPS, pH 7.0) (method of Example 2). A transmembrane potential of +200 mV was continuously applied. Four classical NMPs: adenine mononucleotide (AMP), guanine mononucleotide (GMP), cytosine mononucleotide (CMP), and uracil mononucleotide (UMP) were tested as analytes (Figure 33d). Either type of NMP was added to the cis chamber to a final concentration of 300 μM. Immediately afterward, continuous resistive pulses evoked by either type of NMP were observed (Figure 33d). However, when M2 MspA was tested, no events were observed, confirming that the PBA located in the pore constriction region is important in generating NMP-sensing events (Figure 44). Under these conditions, approximately 1800 events / hour were obtained from each pore. Typically, the pore can withstand continuous measurements for hours. When sensed by MspA-PBA, deoxyribonucleoside monophosphates (dNMPs) fail to report any events (Figure 45). This is expected because dNMPs lack a cis-diol structure and cannot form the boronate ester required for sensing. This demonstrates the molecular sensing mechanism, which occurs through the formation of a reversible covalent bond between NMPs and PBAs in the pore constriction region. This also demonstrates that in a practical measurement scenario, such a sensing strategy would not be subject to interference from dNMPs.
[0242] To quantitatively describe the NMP sensing events, the event residence time (t off ), inter-event interval (t on ), percent interception (%I b =(I p -I b ) / I p ), and noise amplitude (SD). off and t on The histogram of τ shows an exponential distribution and can be fitted to obtain the mean time constant for each. off Or τon . %I b The histograms of and SD show Gaussian distributions, and the mean percentage of blockage for each
[0243]
number
[0244] During the NMP sensing period, the reciprocal of the residence time (1 / τ) can be obtained by changing the NMP concentration in the cis chamber. off ) is held constant, consistent with a single-molecule dissociation mechanism. In contrast, the inverse of the inter-event interval (1 / τ on ) is linearly related to the NMP concentration in the cis chamber (Tables 8-11, Figures 47-50), consistent with the bimolecular model. AMP is used as a representative analyte to study the dependence of the NMP sensing period on the applied voltage. Typically, increasing the voltage increases the τ by 1 / τ off decreases, and 1 / τ on increases (Table 12, Figure 51). This is expected because in pH 7.0 buffer, the phosphate groups of NMP are negatively charged and the electrophoretic force applied to NMP analytes can strongly modulate the frequency of events and the residence time of events. The conical lumen structure of MspA provides excellent resolution, allowing it to distinguish analytes with minute structural differences. 41 Although NMPs differ only in their nucleobase composition, binding of different NMPs to MspA-PBA results in highly distinguishable event signatures (Figure 33d). This difference becomes even greater at higher applied voltages (Figure 52). However, to avoid rupture of the bilayer while maintaining high resolution to distinguish different NMPs, all subsequent measurements are performed at a voltage of +200 mV unless otherwise stated. In this case, events generated by different NMPs are differentiated by %I b The %I of different NMP events forms highly distinguishable clusters in the scatter plot of the %I vs. SD (Fig. 33e). b The histogram of further shows a completely separated Gaussian distribution (Fig. 33e, Fig. 53), among which CMP(
[0245]
number
[0246]
number
[0247]
number
[0248]
number
[0249] = 11.8 ± 0.2%, N = 3) are completely separated without any ambiguity (Table 13, Figures 54 and 55). The event dwell times of different NMPs are widely distributed, producing events with different pulse widths. However, the average event dwell times of different NMPs
[0250]
number
[0251] MspA-PBA was used to simultaneously sense CMP, UMP, AMP, and GMP (Fig. 33f, Fig. 56), from which different NMP identities were directly distinguished according to their distinct blocking characteristics. However, the widely distributed event residence times (Fig. 56) were not considered as a parameter in event recognition. To our knowledge, nanopore discrimination between classical NMPs without any overlap in event distribution has not been reported previously.
[0252] Classification of epigenetic NMPs The above method is in principle suitable for the detection of any nucleoside monophosphate as long as the cis-diol structure of the ribose is preserved. According to the literature, approximately 170 epigenetic NMPs have been discovered to date. 1 They are generated post-transcriptionally and play important roles in many biological activities, including cell differentiation, gene expression, and disease processes. 2 However, these epigenetic NMPs exhibit very small structural differences, posing a significant challenge to their direct identification. Given the high resolution of MspA, this challenge can be resolved by directly monitoring events characteristic of nanopore readouts when epigenetic NMPs bind to the pore constriction region.
[0253] To verify this assumption, m 5 Cm 6 A, m 7 G, m 1 The same measurements were performed using A, I, Ψ, and D as analytes. Due to the lack of commercially available model compounds, Ψ (Figure 57) and D (Figure 58) were custom synthesized and characterized by Wuxi Pharmaceutical Kangde New Drug Development Co., Ltd. These epigenetic NMPs encompass common types of modifications that occur in classical NMPs, such as methylation, deamination, isomerization, and reduction. To more clearly illustrate their chemical structures, their nucleobase components are shown in Figure 34a. Nanopore measurements (method of Example 2) were performed in 1.5 M KCl buffer (1.5 M KCl, 10 mM MOPS, pH 7.0). A transmembrane potential of +200 mV was applied sequentially. Each epigenetic NMP was added to the cis chamber to a final concentration of 300 μM. As shown in Figure 34a and Figure 59, the epigenetic NMP events have significantly different blockage amplitudes. To demonstrate a complete comparison between all NMPs tested to date, the %I of each NMP was plotted in a violin plot. b Exhibits distribution, UMP and m 5Although the event distributions of C still have some overlap, most NMPs have their %I b We show that they can be distinguished only by the analysis of ψ and m (Fig. 34b). 7 The large variation in G comes from the detection of nonspecific events away from the main population of events. These can be generated from impurities introduced during the synthesis of the compounds. However, these nonspecific events only contribute to 0.9% and 1.7% of all detected events, respectively (Figure 60). The noise characteristics of NMP may also be included in the event analysis to improve discrimination performance (Table 13, Figures 61-62). %I including NMP-sensed events obtained from 11 different analytes. b By plotting the scatter plot against the SD of , we generate 11 completely differentiated event populations corresponding to each of the sensed NMPs (Figure 34c). This demonstrates that this detection method is suitable for recognizing epigenetic NMPs and that the events are completely distinguishable. However, to our knowledge, the direct differentiation of these 11 NMPs using nanopores has not been reported to date. To further demonstrate the distinction between epigenetic NMPs and their corresponding classical counterparts, we perform nanopore sensing between epigenetic NMPs and classical NMPs in different sets (Figure 35). Methylated RNA nucleotides are ubiquitous in all species of living organisms, and two-thirds of RNA modifications involve the addition of a methyl group. 49 CMP and m 5 C, GMP and m 7 G, AMP, m 1 A and m 6 Simultaneous sensing of A produced clearly enhanced current blockade and noise upon methylation within the heterocycle (Figures 35a-35f), while methylation at other sites reported the opposite effect (e.g., m 6A, Figure 35e, Figure 35f). Simultaneous detection of AMP and I indicates that deamination can reduce the blockade current (approximately 7.5 pA) and noise (approximately 31.2 pA) (Figure 35g, Figure 35h). This finding is also confirmed by the characteristic events of CMP and UMP. For the isomerization of U to Ψ, a current increase of approximately 8.0 pA is observed (Figure 35i, Figure 35j), which cannot be directly distinguished by mass spectrometry alone, even though U and Ψ have the same molecular weight. For the reduction of U to D, a current increase of approximately 14.0 pA is observed (Figure 35k, Figure 35l). These findings may provide clues for predicting other modification signals. However, because these changes in event characteristics are simultaneously determined by molecular volume, net charge, and other factors, further detailed studies may require the assistance of molecular dynamics simulations.
[0254] Identification of NMPs using machine learning A machine learning algorithm is constructed to automatically identify NMPs. The entire training process includes dataset input, feature extraction, and model building (Fig. 36a, Method of Example 2). Specifically, a dataset is formed using 500 representative events obtained from each NMP. All events in the dataset have known tags because they were obtained during the measurement period by a unique NMP with a known identity. The dataset is then divided into a training set (80%) for model training and a test set (20%) for model testing. MATLAB is used to calculate the %I of each event. bThe SD of the variances is automatically extracted to form a feature matrix. Ten-fold cross-validation is performed to randomly divide the training data into a training set for model training and a validation set for model validation. The model training process is performed using MATLAB's Classification Learner toolbox. Mainstream classifiers include decision trees, discriminant analysis, naive Bayes, support vector machines (SVMs), k-nearest neighbors (KNNs), fusion, and neural networks, all of which are estimated with default parameter settings. The same dataset is repeatedly used for model evaluation. Most models demonstrated satisfactory validation accuracy, indicating the high quality of the input data. In particular, the kernel naive Bayes model and linear SVM model reported the highest accuracy score of 0.996 (Table 14). The trained models were further evaluated using the test set, which showed a slight improvement in the performance of the linear SVM model (Table 14), making it the best model for further evaluation and prediction. The confusion matrix results based on model testing using a linear SVM model are shown in Figure 36b, where most NMP detection results report 99% or 100% accuracy, confirming no significant deviation in identifying different NMPs. Figure 36c also displays the decision boundary plot generated by the linear SVM model. To visually display the event recognition, it is placed on a scatter plot of the test data.
[0255] A previously trained linear SVM model is used to predict events with unknown identity. Measurements are performed as described in the methods of Example 2. The modified NMPs are used as m 5 Cm 6 A, I, m 7 G, m 1A, Ψ, and D were added to the cis chamber in this order, with CMP, UMP, AMP, and GMP already placed in the cis chamber. The final concentration of each NMP in the cis chamber was 100 μM. Using a linear SVM model, newly added NMPs could be accurately identified (Figures 63 and 64). To evaluate the model's training efficiency, learning curves were generated using the training data or validation data, respectively (Figure 65). From these, it can be concluded that 176 events were required for the model to achieve an accuracy of 0.990. When the dataset size exceeded 3124, the accuracy saturated at approximately 0.996. According to the learning curve results, overfitting of the model did not occur. To demonstrate event identification from a mixture, a representative trace containing events from 11 different NMPs is shown in Figure 36d and Supplementary Video 1. Different NMP types could be identified, and the corresponding tags predicted by machine learning are labeled on the traces. This effectively supports automatic nanopore sensing of different NMPs in real measurement scenarios where different NMPs exist as a mixture.
[0256] Detection of epigenetic NMPs in methylated microRNAs We further attempt to demonstrate direct sensing of epigenetic NMPs in RNA. The measurement diagram is shown in Figure 37a. Briefly, RNA is first enzymatically degraded into NMPs by S1 nuclease treatment. The generated NMPs are then sensed by MspA-PBA. The observed nanopore events are identified through a previously trained machine learning model that reports RNA composition, including epigenetic modifications. To experimentally demonstrate its feasibility, we performed a nanopore analysis of known methylation sites. 50 Two microRNAs, including hsa-miR-21 and hsa-miR-17, which have the m at position 9, are applied. 5 hsa-miR-17 contains C and m at position 13 6The results include the following: (Table 15). First, hsa-miR-21 and hsa-miR-17 were detected by MspA-PBA without enzymatic treatment. However, only short-lived spike events with undefined event amplitudes were observed (Figure 66), indicating that this detection method is not sensitive to the template RNA itself. To minimize the interference of glycerol in the S1 nuclease storage solution (Figure 67), the S1 nuclease was pretreated to remove glycerol by ultrafiltration, thereby improving detection efficiency (Method of Example 2, Figure 68). Next, the microRNAs were digested with the pretreated S1 nuclease at 23°C for 4 hours. Gel electrophoresis showed that the two microRNAs were completely degraded (Method of Example 2, Figure 69). Next, the enzymatic treatment product was ultrafiltered to remove the S1 enzyme before nanopore measurement (Method of Example 2). Nanopore measurements were performed in 1.5 M KCl buffer (1.5 M KCl, 10 mM MOPS, pH 7.0) using MspA-PBA (method of Example 2). A transmembrane potential of +200 mV was continuously applied. The hsa-miR-21 digestion product was added to the cis chamber to a final concentration of 100 ng / μL. A representative trace is shown in Figure 37b, demonstrating that numerous NMP binding events were observed and that the generated NMPs were successfully detected by MspA-PBA. Events caused by glycerol binding were completely reduced by the previous ultrafiltration process, but were still evident during the nanopore sensing period. Calling the identity of the NMPs through the algorithm also identified glycerol events, which are highly distinguishable from the exhibited NMP events (Figure 70).
[0257] The results obtained using hsa-miR-21 showed that CMP, UMP, AMP, GMP, and m 5Five NMPs, including C, were detected (Figures 37b and 37c), consistent with the composition of the hsa-miR-21 sequence (Table 15). The abundance of each NMP type in has-miR-21 was also assessed based on calibrated event frequency (Methods in Example 2, Table 16). The relative NMP composition in hsa-miR-21 was 2.17 CMP, 6.81 UMP, 6.88 AMP, 4.92 GMP, 1.03 m 5 The estimated values of C, 0.06 I, 0.01 Ψ, and 0.10 D (Figure 71) are generally consistent with the true values. Although the misidentification of I, ψ, and D is due to the small overlap of the distributions of AMP, ψ, GMP, and I, the rate of misidentification is negligible. Therefore, the feasibility of epigenetic NMP identification by nanopore sensing is demonstrated. To test its generality, hsa-miR-17 (another microRNA containing a different epigenetic NMP in its sequence) is tested in the same manner as shown for hsa-miR-21. A representative trace of nanopore sensing events for digestion products containing hsa-miR-17 is shown in Figure 7d. The scatter plot results show that AMP, UMP, AMP, GMP, and m 6 The five predominant clusters of NMP events corresponding to A (Figure 37e) are consistent with the sequence components of hsa-miR-17 (Table 15). 6 The relative count of A sites was also shown to be 1.08, indicating that hsa-miR-17 has m 6 Only one A site was shown to be present (Figure 71), which is also in line with expectations.
[0258] Brewer's yeast tRNA phe Detection of epigenetic NMPs derived from Transfer RNA (tRNA) is a small RNA used to link messenger RNA sequences to the amino acid sequence of proteins. Mature tRNAs further contain a wealth of chemical modifications. More than 90 types of modifications have been reported in tRNAs. 51Therefore, it is an ideal RNA for evaluating the performance of MspA-PBA in identifying epigenetic modifications in natural samples. As a model RNA, brewer's yeast phenylalanine-specific tRNA (yeast tRNA) was used. phe ) 42, 52 is used to test its feasibility. As reported, mature yeast tRNA Phe contains 14 epigenetic modification sites, and these sites are 2 G = N2-methylguanosine, D = dihydroureosine,
[0259]
number
[0260]
number
[0261] , T, and Y monophosphates can be detected. 5 Cm 7 G, m 1The event parameters of A have been previously obtained and used for model training (Fig. 34a, Fig. 36), so that these events can be identified by machine learning algorithms. 2 G.
[0262]
number
[0263] The monophosphates of , T, and Y can in principle be detected by MspA-PBA, and new event clusters are expected to be observed. However, the corresponding nanopore events are detectable but not distinguishable due to the lack of corresponding pure compounds to generate training events. The C lacks a cis-diol. m and G m In principle, cannot be detected by MspA-PBA. tRNA phe The DNA was first treated with S1 nuclease at 23°C for 15 hours.
[0264] NMP is produced (method of Example 2). Gel electrophoresis results show that tRNA phe This confirms the complete degradation of yeast tRNA (Figure 38b). The enzyme digestion product is then ultrafiltered to remove S1 nuclease and used for subsequent nanopore measurements (method of Example 2). Nanopore measurements are performed in 1.5 M KCl buffer (1.5 M KCl, 10 mM MOPS, pH 7.0) using MspA-PBA (method of Example 2). A transmembrane potential of +200 mV is continuously applied. Yeast tRNA phe The digestion product is added to the cis chamber to a final concentration of 100 ng / μL. The resulting original events are shown in a scatter plot (Figure 72). Glycerol events introduced by the S1 nuclease stock solution are further removed from the dataset by using machine learning to identify their highly distinctive event features (Figure 72). Yeast tRNA pheTo process unknown epigenetic modifications in the nuclei, we combine supervised and unsupervised learning algorithms to identify remaining events in the digested NMPs. Here, we use a one-class SVM to identify events that do not belong to any previously trained event types. These events are considered outliers. Conversely, events that match previously trained event types are considered interior points (Figure 73), and these events are further identified by a trained linear SVM model. However, outlier events are analyzed using a density-based adaptive spatial clustering with noise (DBSCAN) model to detect events that appear as clusters (Figure 74). Non-clustered events randomly distributed in the scatter plot are considered background events and are removed from the dataset without further analysis. yeast tRNA phe The modification profile results are shown in Figure 38c. 5 Cm 7 G and m 1 The detection of A was successful, consistent with previous training results and the literature. 53, 54 A small amount of m 6 A events are observed, which are due to m 6 These events may originate from background events that happen to share similar event characteristics for A or other types of RNA. Four new event clusters were also observed, exhibiting event characteristics distinct from all NMP types previously applied for training. These new event clusters were also observed for yeast tRNA phe m in 2 G.
[0265]
number
Claims
1. a protein nanopore comprising at least one sensing moiety, said sensing moiety coupling said metal ion to a reactive amino acid residue within said nanopore channel and capable of interacting with a target analyte; Protein nanopores.
2. the metal ion is coupled to the reactive amino acid residue via a ligand, and the metal ion and the ligand form a coordination complex; The protein nanopore of claim 1.
3. The ligand is nitrilotriacetic acid (NTA). The nanopore of claim 2.
4. The metal ion is Ni 2+ , Cu 2+ , Co 2+ , Zn 2+ , Cd 2+ , Ag 2+ , Pb 2+ , Fe 2+ , or Fe 3+ Selected from The nanopore according to any one of claims 1 to 3.
5. The reactive amino acid residue is selected from cysteine, methionine, and lysine. The protein nanopore according to any one of claims 1 to 3.
6. the protein nanopore is a heterogeneous protein nanopore, in which one or more, but not all, monomers comprise the sensing moiety and other monomers do not comprise the sensing moiety; The protein nanopore according to any one of claims 1 to 3.
7. the heterogeneous protein nanopore is a variant of the nanopore, and the nanopore is selected from MspA, α-HL, erolysin, ClyA, FhuA, FraC, PlyA / B, CsgG, and Phi 29 linker; The nanopore of claim 6.
8. the heterogeneous protein nanopore is a mutant of MspA; The nanopore of claim 7.
9. The protein nanopore is a heterogeneous MspA nanopore, and the heterogeneous MspA nanopore comprises a Ni nanopore coupled to the reactive amino acid residue via a ligand. 2+ Including, The nanopore of claim 6.
10. Ni 2+ is coupled to the reactive amino acid residue via NTA; The nanopore of claim 9.
11. the reactive amino acid residue is located at a position selected from 83 to 111, preferably 90, 91, 92 and 93; The nanopore of claim 9.
12. the heterogeneous protein nanopore has an N90C, N90M, or N91C mutation in one or more monomers compared to M2 MspA; The nanopore of claim 11.
13. A protein nanopore comprising at least one sensing module, wherein said protein nanopore is heterogeneous MspA, wherein one or more but not all monomers comprise said sensing module and other monomers do not comprise said sensing module, wherein said sensing module can interact with a target analyte; Protein nanopores.
14. the sensing module is composed of one or more reactive amino acid residues contained in one or more monomers of the heterogeneous MspA; The nanopore of claim 13.
15. The reactive amino acid residue is selected from methionine, histidine, cysteine, or lysine, or a combination thereof; 15. The protein nanopore of claim 14.
16. the sensing module consists of one or more sensing moieties, wherein the one or more sensing moieties couple to one or more reactive amino acid residues contained in one or more monomers of the heterogeneous protein nanopore, and other monomers of the heterogeneous protein nanopore do not contain the reactive amino acid residues; 13. The protein nanopore of claim 12.
17. The reactive amino acid residue is selected from cysteine, methionine, and lysine.
17. The protein nanopore of claim 16.
18. the sensing moiety is a boronic acid-containing moiety; 17. The nanopore of claim 16.
19. the boronic acid-containing moiety is phenylboronic acid (PBA); 19. The nanopore of claim 18.
20. the reactive amino acid residues are located at one or more positions selected from 83 to 111, preferably 90, 91, 92 and / or 93; The protein nanopore of any one of claims 13 to 15.
21. the heterogeneous protein nanopore has N90C, N90M and / or N91C mutations in one or more monomers compared to M2 MspA; The nanopore according to any one of claims 13 to 15.
22. 1. A method for characterizing a target analyte, the method comprising: (i) providing a protein nanopore according to claim 1; (ii) applying a voltage across the protein nanopore reactor; (iii) allowing the target analyte to pass through the nanopore; and (iv) measuring the ionic current passing through the nanopore to provide a current pattern, and characterizing the target analyte based on the current pattern.
23. the target analyte is in a sample, and step (iii) comprises allowing the sample to pass through the nanopore; 23. The method of claim 22.
24. The sample is selected from fruit juices, beverages, teas and herbal extracts.
24. The method of claim 23.
25. 10. Use of the protein nanopore of claim 1 in the characterization of a target analyte.
26. the target analyte is in a sample; 26. The use according to claim 25.
27. The sample is selected from fruit juices, beverages, teas and herbal extracts.
27. The use according to claim 26.
28. The target analyte may interact with a boronic acid, a metal ion, a methionine, a histidine, a cysteine, a lysine, or any combination thereof.
26. The method according to claim 22 or the use according to claim 25.
29. the analyte capable of interacting with a boronic acid is selected from a compound containing a 1,2-diol or a 1,3-diol, an ion containing a metal element, hydrogen peroxide, and any combination thereof, and the analyte capable of interacting with a metal ion is a molecule capable of interacting with the metal ion through coordination; and The analyte capable of interacting with methionine, histidine, cysteine, or lysine is an ion containing a metal element.
29. The method or use of claim 28.
30. The ions containing a metal element are selected from alkaline earth metal ions, transition metal ions, and any combination thereof, and are preferably AuCl 4 - , Mg 2+ , Ca 2+ , Ba 2+ , Ni 2+ , Cu 2+ , Co 2+ , Zn 2+ , Cd 2+ , Ag 2+ , Pb 2+ and any combination thereof; 30. The method or use of claim 29.
31. The compound containing 1,2-diol or 1,3-diol is selected from the group consisting of sugars or derivatives thereof, α-hydroxy acids, ribose, nucleoside sugars, alditols, polyphenols, compounds containing catecholamines or catecholamine derivatives, tris(hydroxymethyl)methylaminomethane (Tris), protocatechuic aldehyde, protocatechuic acid, caffeic acid, rosmarinic acid, lithospermic acid, tansinol A, salvianolic acid B, and any combination thereof.
30. The method or use of claim 29.
32. the sugar is selected from monosaccharides, oligosaccharides, polysaccharides, and any combination thereof; the sugar derivative is selected from N-acetylneuraminic acid (sialic acid), N-acetyl-D-galactosamine, and any combination thereof; the alpha-hydroxy acid is selected from tartaric acid, malic acid, citric acid, isocitric acid, and any combination thereof; the compound comprising ribose is selected from a nucleotide or modified nucleotide, a derivative of a nucleotide or modified nucleotide, a nucleoside or nucleoside analogue, and any combination thereof; the nucleotide sugar is selected from uridine diphosphate glucose (UDPG), uridine diphosphate N-acetylglucosamine, uridine diphosphate glucuronic acid, adenosine diphosphate glucose, uridine diphosphate galactose, uridine diphosphate xylose, guanosine diphosphate mannose, guanosine diphosphate fucose, cytidine monophosphate N-acetylneuraminic acid, uridine diphosphate N-acetylgalactosamine, and any combination thereof; the alditol is selected from glycerin, propanetriol, butitol, pentitol, hexitol, erythritol, threitol, arabitol, xylitol, ribitol (adonitol), fucitol, sorbitol such as L-sorbitol or D-sorbitol, mannitol, galactitol, iditol, talitol, allitol (alodulcitol), maltitol, lactitol, isomalt, and any combination thereof; The polyphenol is selected from catechin, neochlorogenic acid, anthocyanin, proanthocyanidin, catechol or a derivative thereof, such as catechol, 3-fluorocatechol, 3-chlorocatechol, 3-bromocatechol, 4-fluorocatechol, 4-chlorocatechol, 4-bromocatechol, 3-methylcatechol, 4-methylcatechol, 3-methoxycatechol, 3-propylcatechol, 3-isopropylcatechol, 3,6-dibromocatechol, 4,5-dibromocatechol, 3,6-dichlorocatechol, and any combination thereof; and The catecholamine or catecholamine derivative is selected from epinephrine, norepinephrine, isoproterenol, and any combination thereof.
32. The method or use of claim 31.
33. The monosaccharides include D-glyceraldehyde, D-erythrose, D-ribose, 2'-deoxy-D-ribose, D-xylose, L-arabinose, D-lyxose, D-glucose, D-galactose, selected from D-mannose, D-fructose, L-sorbose, L-fucose, D-allose, D-tagatose, L-rhamnose, and any combination thereof; the oligosaccharide is selected from disaccharides (e.g., sucrose, isomaltulose, maltulose, turanose, leucrose, trehalose, lactulose, maltose, etc.), trisaccharides (e.g., raffinose), tetrasaccharides (e.g., stachyose), and complex oligosaccharides (e.g., acarbose), and any combination thereof; the polysaccharide is selected from pentasaccharides such as verbascose, the nucleotides are selected from adenine nucleotides, cytosine nucleotides, uracil nucleotides, guanine nucleotides, and any combination thereof; The modified nucleotide is 5-methylcytidine (m 5 C), N6-methyladenosine (m 6 A), pseudouridine (Ψ), inosine (I), N7-methylguanosine (m 7 G), N1-methyladenosine (m 1 A), dihydrouridine (D), N2-methylguanosine (m 2 G), N2, N2-dimethylguanosine [Equation 1] , ybutosine (Y), 5-methyluridine (T), N-acetylcytidine (ac4C), and any combination thereof; The derivative of a nucleotide or modified nucleotide is selected from monophosphate, diphosphate, triphosphate and tetraphosphate derivatives of a nucleotide or modified nucleotide, and any combination thereof, such as ADP, UDP, GDP, CDP, ATP, UTP, GTP, CTP and any combination thereof; and The nucleoside analog is selected from galidesvir, ribavirin, molnupiravir, remdesivir, loxoribine, mizoribine, 5-azacytidine, capecitabine, doxifluridine, 5-fluorouridine, forodesine, kleitosine, pyrazofurin, sangivamycin, pseudouridimycin, and any combination thereof; 33. The method or use of claim 32.
34. The molecule capable of interacting with the metal ion through coordination contains a nitrogen, oxygen, sulfur, phosphorus, or carbon atom capable of coordinating with the metal ion; 30. The method or use of claim 29.
35. The molecule capable of interacting with the metal ion through coordination is a compound containing at least one carboxylic acid group or at least one amine group, an amino acid, a modified amino acid, a polymer of amino acids or modified amino acids, a compound containing guanine, adenine, thymine, cytosine or uracil, and any combination thereof.
35. The method or use of claim 34.
36. the amino acids are selected from alanine, cysteine, aspartic acid, glutamic acid, phenylalanine, glycine, histidine, isoleucine, lysine, leucine, methionine, asparagine, proline, glutamine, arginine, serine, threonine, valine, tryptophan, tyrosine, pyrrolysine, selenocysteine, and any combination thereof; The modified amino acids are selected from phosphorylated amino acids, glycosylated amino acids, acetylated amino acids, methylated amino acids, and any combination thereof, such as O-phosphoserine (pS), N4-(β-N-acetyl-D-glucoseamino)-asparagine (GlcNAc-N), O-acetyl-threonine (Ac-T), Nω,N′ω-dimethyl-arginine (SDMA), and any combination thereof; and The compound containing guanine, adenine, thymine, cytosine or uracil is selected from guanine, adenine, thymine, cytosine or uracil, or contains any one nucleoside selected from them, or contains any one nucleotide selected from them, and the nucleotide is a ribonucleotide or a deoxyribonucleotide.
36. The method or use of claim 35.