Nucleic acid sequencing system
By using NetB nanoporous protein and positively charged molecular markers to optimize NetB protein mutants, the signal interference problem caused by DNA capture by nanoporous proteins was solved, achieving higher accuracy, longer read lengths and higher throughput nucleic acid sequencing.
Patent Information
- Application Number
- PCT/CN2025/108936
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-16
- Filing Date
- 2025-07-16
- Publication Date
- 2026-01-22
AI Technical Summary
In existing nucleic acid sequencing systems, the capture of DNA by nanoporous proteins leads to severe signal interference and short channel lifetime, making it difficult to meet the sequencing requirements for longer read lengths and higher throughput.
The NetB nanoporous protein was used as the detection component, positively charged molecules were used as nucleotide markers, and the ability to capture positively charged molecules and the stability of the channel were improved by mutating and optimizing the NetB protein.
It achieves higher sequencing accuracy, longer read length and higher throughput, and the nanopore channel lifetime can be stably tested for more than 12 hours, reducing DNA interference signals.
Smart Images

Figure CN2025108936_22012026_PF_FP_ABST
Abstract
Description
Nucleic acid sequencing system
[0001] Cross-citation of related applications
[0002] This application claims priority to Chinese patent application CN 202410954729.6, filed on July 16, 2024, the entire contents of which are incorporated herein by reference. Technical Field
[0003] This disclosure pertains to the field of nucleic acid sequencing. It relates to an improved NetB nanoporous protein and its use as a detection component in a nucleic acid sequencing system. It also relates to a positively charged molecule and its use as a labeling agent for substrate nucleotides. Furthermore, it relates to a nucleic acid sequencing system comprising the improved NetB nanoporous protein as a detection component in a nucleic acid sequencing system and a positively charged molecule as a labeling agent for substrate nucleotides. Background Technology
[0004] The earliest patent reports on sequencing-by-synthesis schemes based on nanopore detection were proposed by Pacific Biosiences and Genia (now acquired by Roche). Their working principle involves coupling DNA polymerase with nanopore proteins, then using labeled nucleotides as substrates for the polymerase to synthesize new strands of DNA. The nucleotide labels used are electrically charged macromolecules that can be captured by the nanopores under the influence of an electric field, changing the resistance of the nanopore channels and thus enabling the detection of the labels. By distinguishing four types of nucleotides using different labels, the DNA sequence can be determined. Based on this model, combining long-progressive DNA polymerases (such as Phi29) can achieve single-molecule and long-read sequencing of DNA.
[0005] In existing embodiments, the nanoporin used is either α-hemolysin or γ-hemolysin, with α-hemolysin being the most widely used. Because α-hemolysin and many other widely reported nanoporins, including MspA and CsgG, tend to capture negatively charged molecules and generate characteristic blocking currents under an electric field, the nucleotide markers used in existing embodiments are all negatively charged molecules. Since DNA itself is also a negatively charged macromolecule, in this system, DNA molecules also tend to be captured by α-hemolysin or interact with it under an electric field, generating blocking current signals. Therefore, during sequencing, the nucleotide marker signal corresponding to the base sequence is strongly interfered with by the DNA, severely affecting signal recognition.
[0006] Furthermore, in practical applications, we found that α-hemolysin protein channels have a short lifetime during testing. Typically, a large number of channels close within about an hour or even less, failing to continue supporting sequencing. This limits single-molecule level sequencing read lengths to 10Kb or less, making it difficult to meet the longer read length requirements of many sequencing scenarios. Premature sequencing termination also significantly limits sequencing data output. Although this can be improved by optimizing testing conditions and reshaping the nanopore working environment, the channel sequencing lifetime is still insufficient to support sequencing needs exceeding 50Kb read lengths and higher throughput.
[0007] Therefore, in the context of the continuous pursuit of higher accuracy, longer read length, higher sequencing efficiency and lower sequencing cost in the field of single-molecule sequencing, the industry still needs to improve existing nucleic acid sequencing systems to obtain nucleic acid sequencing systems with less interference in signal recognition and longer nanopore channel lifespan, so as to achieve the goals of higher accuracy, longer read length and higher throughput. Summary of the Invention
[0008] To address the shortcomings of existing technologies, this disclosure utilizes a unique positively charged molecule as a label for substrate nucleotides. To capture this positively charged label, wild-type NetB nanoporous protein or a modified version of NetB nanoporous protein is used as the detection component of the sequencing system. This system effectively avoids interference signals caused by DNA during sequencing. Furthermore, NetB nanoporous protein exhibits an exceptionally long channel lifespan during testing, providing stable testing for over 12 hours, effectively increasing single-channel sequencing throughput.
[0009] This disclosure reveals for the first time the capture properties of the NetB protein for charged analytes. Unlike most commonly used nanoporous proteins, such as α-hemolysin, MspA, and CsgG, which tend to capture negatively charged molecules such as DNA and nucleotides, NetB prefers to capture positively charged molecules under the influence of an electric field. The stable opening current of NetB channels, its ability to capture positively charged molecules, and the stable blocking current generated after capturing positively charged molecules make it a preferred nanoporous protein for supporting nucleic acid sequencing using positively charged labeled nucleotides.
[0010] In one aspect, this disclosure provides a NetB protein mutant comprising an amino acid sequence having, compared to the wild-type NetB protein amino acid sequence shown in SEQ ID NO:1, one or more mutations selected from the group consisting of: deletion of 10 to 20 amino acid residues at the N-terminus of the wild-type NetB protein, and any one or more of K20, K24, H53, K114, K115, N236, D285, Q284 and E289, each independently mutated to D, N, Q, S, T, Y, H or E.
[0011] In some embodiments, the amino acid sequence of the NetB protein mutant, compared to the wild-type NetB protein amino acid sequence shown in SEQ ID NO:1, includes one or more mutations selected from the group consisting of: deletion of 10 amino acid residues from the N-terminus of the wild-type NetB protein, deletion of 11 amino acid residues from the N-terminus of the wild-type NetB protein, deletion of 12 amino acid residues from the N-terminus of the wild-type NetB protein, deletion of 13 amino acid residues from the N-terminus of the wild-type NetB protein, deletion of 14 amino acid residues from the N-terminus of the wild-type NetB protein, deletion of 15 amino acid residues from the N-terminus of the wild-type NetB protein, deletion of 16 amino acid residues from the N-terminus of the wild-type NetB protein, deletion of 17 amino acid residues from the N-terminus of the wild-type NetB protein, deletion of 18 amino acid residues from the N-terminus of the wild-type NetB protein, deletion of 19 amino acid residues from the N-terminus of the wild-type NetB protein, deletion of 20 amino acid residues from the N-terminus of the wild-type NetB protein, and H53D. , H53Y, H53N, H53Q, H53E, N236D, N236E, Q284D, Q284E, K20D, K20N, K20Q, K20S, K20T, K20 Y, K20H, K20E, K24N, K24D, K24Q, K24S, K24T, K24Y, K24H, K24E, K114N, K114D, K114Q, K11 4S, K114T, K114Y, K114H, K114E, K115N, K115D, K115Q, K115S, K115T, K115Y, K115H, K115 E, D285N, D285Q, D285S, D285T, D285Y, D285H, E289N, E289Q, E289S, E289T, E289Y and E289H.
[0012] In some implementations, the mutation includes:
[0013] (1) The N-terminus of wild-type NetB protein is missing 10 to 20 amino acid residues; and
[0014] (2) Select one or more mutations from the group consisting of the following mutations: K20D, K20N, K24D, K24N, K24S, H53D, K114D, K114N, K115D, K115T, N236D, N236E, Q284D, Q284E, D285H, D285N, E289Q and E289Y.
[0015] In some embodiments, the NetB protein mutant includes the amino acid sequence shown in SEQ ID NO:2, SEQ ID NO:3, SEQ ID NO:4, SEQ ID NO:5, SEQ ID NO:6, SEQ ID NO:7, SEQ ID NO:8, SEQ ID NO:9, SEQ ID NO:10, SEQ ID NO:11 or SEQ ID NO:12.
[0016] This disclosure also provides an isolated nucleic acid molecule that encodes the NetB protein mutant described herein.
[0017] This disclosure also provides an expression vector comprising the nucleic acid molecules described herein.
[0018] This disclosure also includes a host cell comprising the expression vector described herein.
[0019] On the other hand, this disclosure provides the use of the NetB protein in the preparation of detection components for nucleic acid sequencing, wherein the NetB protein is selected from:
[0020] (a) Wild-type NetB protein comprising the amino acid sequence shown in SEQ ID NO:1; or
[0021] (b) The NetB protein mutant described in this disclosure.
[0022] In some implementations, the nucleic acid sequencing is sequencing-while-synthesizing, and the detection component includes a nanopore formed by the NetB protein and coupled to a DNA polymerase.
[0023] In another aspect, this disclosure provides a modified NetB protein comprising a NetB protein and a linker for coupling with a DNA polymerase, wherein the NetB protein is selected from:
[0024] (a) Wild-type NetB protein comprising the amino acid sequence shown in SEQ ID NO:1; or
[0025] (b) The NetB protein mutant described in this disclosure.
[0026] In some embodiments, the linker is a chemical moiety capable of participating in covalent or non-covalent protein linking reactions, selected from the group consisting of: SpyTag peptide, SnoopTag peptide, sorting enzyme recognition site, chemically cross-linkable amino acid residue, biotin, HaloTag, or SNAP-tag.
[0027] In another aspect, this disclosure provides NetB nanoporous protein, which is a polymer composed of at least two protein monomers as subunits, wherein each of the at least two protein monomers is independently selected from:
[0028] (a) A wild-type NetB protein containing the amino acid sequence shown in SEQ ID NO:1; or
[0029] (b) The NetB protein mutant described in this disclosure.
[0030] In some embodiments, the polymer is a heptamer, and at least one of the protein monomers contains a linker for coupling with DNA polymerase.
[0031] In some embodiments, the linker is a chemical moiety capable of participating in covalent or non-covalent protein linking reactions, selected from the group consisting of: SpyTag peptide, SnoopTag peptide, sorting enzyme recognition site, chemically cross-linkable amino acid residue, biotin, HaloTag, or SNAP-tag.
[0032] In some alternative embodiments, this disclosure provides a NetB nanoporous protein, which is a polymer composed of at least two NetB protein monomers as subunits, wherein each of the at least two NetB protein monomers is independently selected from:
[0033] (a) Wild-type NetB protein comprising the amino acid sequence shown in SEQ ID NO:1;
[0034] (b) The NetB protein mutant described in this disclosure; or
[0035] (c) The modified NetB protein described in this disclosure.
[0036] In some embodiments, at least one NetB protein monomer is a modified NetB protein containing a linker as described in this disclosure.
[0037] In some embodiments, the multimer is a heptamer, which includes NetB protein subunits with the same or different amino acid sequences.
[0038] In some embodiments, the NetB nanoporous protein comprises:
[0039] (1) Six NetB proteins that do not contain linkers, wherein the linker-free NetB proteins are selected from...
[0040] (a) Wild-type NetB protein comprising the amino acid sequence shown in SEQ ID NO:1; and
[0041] (b) The NetB protein mutants described in this disclosure;
[0042] (2) A modified NetB protein as described in this disclosure.
[0043] In some embodiments, the six NetB proteins that do not contain linkers are the same NetB protein, and the amino acid sequence contained therein is the same as that contained in the modified NetB protein.
[0044] In another aspect, this disclosure provides a protein complex comprising:
[0045] (a) The NetB nanoporous protein described in this disclosure; and
[0046] (b) DNA polymerase coupled to the NetB nanoporous protein via a linker.
[0047] In another aspect, this disclosure provides a positively charged labeled molecule, as shown in formula (I): MX Formula (I)
[0048] in,
[0049] The marked part M contains one or more positive charges; and
[0050] The reactive linker part X is adapted to covalently link with the polyphosphate moiety of a nucleotide.
[0051] In some implementations, the marker portion M includes:
[0052] (a) The head portion of a polypeptide, consisting of 3 to 15 positively charged amino acid residues; and
[0053] (b) Signal identification main part R'.
[0054] In some embodiments, the polypeptide head comprises amino acid residues selected from the group consisting of lysine, arginine, and α,β-diaminopropionic acid.
[0055] In some embodiments, the signal identification body portion R' is a chain polymer whose main chain backbone optionally contains one or more heterohydrocarbon chains selected from N, O, or S.
[0056] In some embodiments, the reactive linker portion X comprises an azide or alkynyl group for click chemistry.
[0057] In another aspect, this disclosure provides a positively charged nucleotide comprising the positively charged molecule described herein, the positively charged molecule being covalently linked by a linker to the terminal phosphate of a polyphosphate nucleotide moiety comprising at least three phosphate groups.
[0058] In some embodiments, the linker is a triazole ring formed by an azide-yne cycloaddition reaction.
[0059] In some embodiments, the polyphosphate nucleotide moiety is nucleoside pentaphosphate or nucleoside hexaphosphate.
[0060] In another aspect, this disclosure provides the use of positively charged nucleotides for nucleic acid replication, amplification, or sequencing, such as as substrate nucleotides in nucleic acid replication, amplification, or sequencing.
[0061] In another aspect, this disclosure provides substrate nucleotides for nucleic acid replication, amplification, or sequencing, which are positively charged nucleotides of this disclosure.
[0062] In another aspect, this disclosure provides a nucleic acid sequencing system comprising:
[0063] (a) a detection unit comprising the NetB nanoporous protein or protein complex described in this disclosure; and
[0064] (b) A group of positively charged nucleotides.
[0065] In some embodiments, the positively charged nucleotide is the positively charged nucleotide described in this disclosure.
[0066] In another aspect, this disclosure provides a method for nucleic acid sequencing, which includes...
[0067] (a) Using the NetB nanoporous protein or protein complex described in this disclosure as a single-molecule detection component; and
[0068] (b) Use a set of positively charged nucleotides as substrates for sequencing reactions.
[0069] In some embodiments, the positively charged nucleotide is the positively charged nucleotide described in this disclosure.
[0070] In another aspect, a kit for nucleic acid sequencing is disclosed, comprising:
[0071] (a) the NetB nanoporous protein or protein complex according to this disclosure; and
[0072] (b) A group of positively charged nucleotides, and
[0073] Optionally, (c) sequencing buffer, and / or (d) instructions for use.
[0074] In some embodiments, the positively charged nucleotide is the positively charged nucleotide according to claims 17-19. Attached Figure Description
[0075] Figure 1 shows a schematic diagram of the structure of the NetB nanoporous protein disclosed herein.
[0076] Figure 2 shows a model diagram of a biochemical system based on nanopore-based sequencing-by-synthesis: 1. DNA polymerase; 2. DNA to be sequenced; 3. Labeled nucleotides; 4. The connection bridge between DNA polymerase and nanopore; 5. Nanopore protein; 6. Artificial phospholipid bilayer; 7. A model diagram of four bases corresponding to the changes in current through the nanopore channel.
[0077] Figure 3 shows the stable opening current of NetB nanoporous protein and its capture signal characteristics for different charged molecules (PA-1-1, PA-2-1, PA-6, PA-8): the test was conducted at 50 mM Tris-HCl pH 8.0, 300 mM Kac buffer, and 120 mV. The collected current signals were clustered and analyzed. 1. Opening current signal band, 2. PA-1-1 signal band, 3. PA-2-1 signal band, 4. PA-6 signal band, 5. PA-8 signal band.
[0078] Figure 4 shows the synthetic route of the positively charged labeled molecule PA-1-1.
[0079] Figure 5 shows the synthetic route of the positively charged labeled molecule PA-2-1.
[0080] Figure 6 shows the synthetic route of the positively charged labeled molecule PA-6.
[0081] Figure 7 shows the synthetic route of the positively charged labeled molecule PA-8.
[0082] Figure 8 shows the synthetic route for the positively charged nucleotide dC5P-PA-1-1.
[0083] Figure 9 shows the synthetic route for the positively charged nucleotide dG5P-PA-2-1.
[0084] Figure 10 shows the synthetic route for the positively charged nucleotide dT5P-PA-6.
[0085] Figure 11 shows the synthetic route for the positively charged nucleotide dA5P-PA-8.
[0086] Figure 12 shows a comparison between the data obtained from nucleic acid sequencing using the nucleic acid sequencing system of this disclosure and the data obtained from nucleic acid sequencing using a negatively charged marker system: A is the negatively charged system, and B is the positively charged system. The figure shows that the positively charged sequencing signal is clear and significantly differentiated, while the negatively charged sequencing system is severely affected by DNA interference, resulting in loss of open-pore signal.
[0087] Figure 13 shows a performance comparison of the sequencing system disclosed herein with existing negatively charged labeling systems in terms of sequencing accuracy (A) and read length (B). As can be seen from the figure, compared with existing sequencing systems based on negatively charged nucleotides, the sequencing system disclosed herein has higher sequencing accuracy and longer read length.
[0088] Figure 14 shows a schematic diagram of the sequencing apparatus of this disclosure.
[0089] Figure 15 shows the pore current state diagram of NetB nanoporous protein composed of wild-type NetB protein (SEQ ID NO:1).
[0090] Figure 16 shows the pore current state of NetB nanoporous protein composed of mutant NetB protein (SEQ ID NO:2).
[0091] Figure 17 shows the pore current state diagram of NetB nanoporous protein composed of mutant NetB protein (SEQ ID NO:3).
[0092] Figure 18 shows the pore current state of NetB nanoporous protein composed of mutant NetB protein (SEQ ID NO:4).
[0093] Figure 19 shows the sequencing signal of NetB nanoporous protein composed of mutant NetB protein (SEQ ID NO:5).
[0094] Figure 20 shows the sequencing signal of NetB nanoporous protein composed of mutant NetB protein (SEQ ID NO:6).
[0095] Figure 21 shows the sequencing signal of NetB nanoporous protein composed of mutant NetB protein (SEQ ID NO:7).
[0096] Figure 22 shows the sequencing signal of NetB nanoporous protein composed of mutant NetB protein (SEQ ID NO:8).
[0097] Figure 23 shows the sequencing signal of NetB nanoporous protein composed of mutant NetB protein (SEQ ID NO:9).
[0098] Figure 24 shows the sequencing signal of NetB nanoporous protein composed of mutant NetB protein (SEQ ID NO:10).
[0099] Figure 25 shows the sequencing signal of NetB nanoporous protein composed of mutant NetB protein (SEQ ID NO:11).
[0100] Figure 26 shows the sequencing signal of NetB nanoporous protein composed of mutant NetB protein (SEQ ID NO:12). Detailed Implementation
[0101] Unless otherwise stated, the technical and scientific terms used in this disclosure have the meanings commonly understood by a person of ordinary skill in the art to which this disclosure pertains.
[0102] The term "or" refers to a single element among the listed optional elements, unless the context explicitly indicates otherwise. The term "and / or" refers to any one, any two, any three, any more, or all of the listed optional elements.
[0103] The terms “including,” “contains,” and “have” mean “including but not limited to,” but also refer to situations consisting only of the listed elements.
[0104] As used herein, the term "wild-type NetB protein" or "NetB protein wild-type" refers to the pore-forming toxin NetB produced by *Clostridium perfringens*, the structure of which is deposited in the Protein Database (PDB) under accession code 4H56, as shown in SEQ ID NO:1. The modifier "wild-type" as used herein does not imply that it is naturally occurring, but rather that the wild-type NetB protein is used herein as a reference for determining the mutations and mutation locations included in the NetB protein mutants provided herein, or can be considered as the parent protein of the NetB protein mutants provided herein. Additionally, in some cases, the wild-type NetB protein is also used as a reference for studying various activities or properties of the NetB protein mutants provided herein.
[0105] As used herein, the term "NetB protein mutant" or "NetB protein strain" refers to a protein that differs from the wild-type NetB protein in its amino acid sequence. This difference can manifest as the substitution, deletion, or insertion of one or more amino acids at one or more positions in the amino acid sequence.
[0106] The “NetB protein mutants” disclosed herein are also intended to include “functional equivalents” of the wild-type NetB protein (SEQ ID NO:1) or mutants described herein. The term “functional equivalent” refers to a polypeptide that differs from a reference protein in its amino acid sequence but substantially retains the key biological functions of the reference protein. In the context of this invention, key biological functions include, but are not limited to: (a) the ability to self-assemble into multimeric nanopores; and (b) the ability of said nanopores to capture positively charged molecules (e.g., positively charged nucleotides of this disclosure) under an electric field and generate a detectable current signal. Functional equivalents can be obtained by modifying the nucleic acid sequence encoding the reference protein, for example, through techniques such as site-directed mutagenesis.
[0107] As used herein, the term "amino acid substitution," also known as "amino acid replacement," refers to the replacement of one amino acid (called the original amino acid) at a specific position in an amino acid sequence by another amino acid (the substituted amino acid). For example, if the second amino acid in a wild-type protein is a glutamic acid residue (E), and the corresponding position in a mutant protein is an alanine residue (A), then the mutant protein is considered to have an amino acid substitution: alanine replaces glutamic acid at position 2. For amino acid substitutions, the following nomenclature is used herein: original amino acid, position, and substituted amino acid, and the amino acid name uses the single-letter abbreviations for amino acids as defined by IUPAC. For the example described above, it can be represented as E2A. When a mutant protein contains two or more amino acid substitutions simultaneously, it is considered to have an amino acid substitution combination. "Amino acid substitution combination" in this document refers to a mutant protein containing two or more positions where amino acid substitutions occur. For example, if the second amino acid in the wild-type protein is a glutamic acid residue (E), and the corresponding position in the mutant protein is an asparagine residue (N); and the 23rd amino acid in the wild-type protein is also a glutamic acid residue (E), and the corresponding position in the mutant protein is an alanine residue (A), then the mutant protein can be considered to contain the amino acid substitution combination E2N and E23A. When multiple amino acid substitutions exist simultaneously in a mutant, the multiple mutations can be separated by "+", for example, "E2N+E23A+Y65R" represents glutamic acid being replaced by asparagine at positions 2, 23, and 65, glutamic acid being replaced by alanine, and tyrosine being replaced by arginine, respectively.
[0108] As used herein, the term "conserved amino acid substitution" refers to the replacement of one amino acid residue in an amino acid sequence with another amino acid residue that has similar physicochemical properties (e.g., similar size, charge, or hydrophobicity). Such substitution typically does not significantly affect the secondary or tertiary structure of a protein. Those skilled in the art will know that functionally similar amino acids can be grouped according to the properties of their side chains. For example, they can be grouped into the following groups:
[0109] Aliphatic side chains (nonpolar): glycine (Gly), alanine (Ala), valine (Val), leucine (Leu), isoleucine (Ile);
[0110] Aromatic side chains: phenylalanine (Phe), tyrosine (Tyr), tryptophan (Trp);
[0111] Sulfur-containing side chains: methionine (Met), cysteine (Cys);
[0112] Acidic side chains (negatively charged): Aspartic acid (Asp), glutamic acid (Glu);
[0113] Basic side chains (positively charged): Lysine (Lys), Arginine (Arg), Histidine (His);
[0114] Containing hydroxyl side chains (polar and uncharged): Serine (Ser), Threonine (Thr);
[0115] Amide side chains (polar and uncharged): Asparagine (Asn), glutamine (Gln).
[0116] Therefore, in some embodiments, the NetB protein mutant of this disclosure may contain one or more conserved amino acid substitutions.
[0117] When used in this document, the term “amino acid sequence” is synonymous with and interchangeable with the terms “polypeptide,” “protein,” and “peptide.” Amino acid residues in an amino acid sequence can be represented by conventional single-letter or three-letter amino acid names, where the amino acid sequence is presented with a standard amino-to-carboxyl terminal orientation (i.e., N→C).
[0118] When used in this paper, the term "sequence identity" (also known as "sequence consistency") refers to the degree of similarity between two amino acid sequences (e.g., a query sequence and a reference sequence), typically expressed as a percentage. Generally, sequence alignment is performed and gaps (if any) are introduced before calculating the percentage of similarity between two amino acid sequences. If amino acid residues in the two sequences are the same at a given alignment position, the two sequences are considered to be consistent or matched at that position; if amino acid residues in the two sequences are different, they are considered inconsistent or mismatched at that position. In some algorithms, sequence consistency is obtained by dividing the number of matched positions by the total number of positions in the alignment window. In other algorithms, the number of gaps and / or gap lengths are also taken into account. Commonly used sequence alignment algorithms or software include DANMAN, CLUSTALW, MAFFT, BLAST, MUSCLE, etc. For the purposes of this disclosure, the publicly available alignment software BLAST (available from https: / / www.ncbi.nlm.nih.gov / ) can be used to obtain the best sequence alignment and calculate the sequence identity between the two amino acid sequences by using the default settings.
[0119] As used herein, the term "nanoporin" refers to a protein polymer with a nanoscale central pore formed by the self-assembly of one or more protein subunits (monomers). In applications, this polymer is typically embedded in an insulating membrane (such as an artificial phospholipid bilayer) to separate two electrolyte solution pools. When a voltage is applied, ions pass through the pores, forming a stable open-pore current. When a single analyte molecule enters or passes through the pore, it causes a temporary blockage of the ion flow, producing a characteristic signal change. Therefore, nanoporins can serve as highly sensitive single-molecule random sensors. The nanoporins involved in this invention specifically refer to NetB protein and its mutants, but also encompass other biologically derived nanoporins with similar functions.
[0120] As used herein, the term "positively charged labeled molecule" refers to a molecular marker that carries a net positive charge, either wholly or partially, under specific sequencing buffer conditions (e.g., pH 4.0–9.0). This molecule typically possesses a sufficiently long and flexible linear structure, with one end covalently linked to the polyphosphate terminus of the substrate nucleotide. During nanopore sequencing, when the attached nucleotide is recognized by polymerase, the labeled molecule enters the trapping region of the nanopore under the influence of an electric field, generating a characteristic blocking current signal. By recognizing the unique signals generated by different labels, the corresponding nucleotide species can be identified. The "positively charged labeled molecule" of this invention differs fundamentally in charge properties from the negatively charged labels commonly used in the prior art, aiming to solve the signal interference problem caused by the negatively charged DNA template.
[0121] As used herein, the term "heteroalkyl group" refers to a substituted or unsubstituted straight-chain and / or branched, saturated or unsaturated hydrocarbon group containing a heteroatom selected from N, O, and S, including straight-chain, branched, or cyclic alkyl, alkenyl, and ynyl groups. In some embodiments, the heteroatom contained in the heteroalkyl group may form the backbone of a heteroaliphatic group together with a carbon atom, such as, but not limited to, group structures such as -CNC-, -COC-, -COOC, -CSC-, -CSSC, or any combination thereof. In some embodiments, the heteroatom contained in the heteroalkyl group may be a substituent attached to a carbon atom, such as, but not limited to, substituted structures such as -C≡N, -C=N-, -CN=, -C=O, -C-OH, -C=S, -C-SH, etc. In some of the embodiments described, the heteroatom contained in the heteroalkyl group may be any combination of the group structures listed above. In some embodiments, the heteroalkyl group comprises one or more, such as 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11 or 12 unsaturated carbon-carbon double bonds (-C=C-) or carbon-carbon triple bonds (-C≡C-). -NH-, -NH2-, OH, -OR m -O-, -C(O)-, -C(OR) n -, -C(O)O-, -SH, -SR o -S-, -C(S)-, -C(SR) p -, -C(S)O-, -P(O)-, and / or any combination thereof, where R m R n R o and R p Each is independently either substituted or unsubstituted C1-C 14 Aliphatic hydrocarbon groups, such as C1-C 12 C1-C 10 C1-C8, C1-C6, C1-C4 aliphatic hydrocarbon groups, substituted or unsubstituted C1-C 14 Non-limiting specific examples of aliphatic hydrocarbon groups include, but are not limited to, methyl, ethyl, n-propyl and isopropyl, n-butyl, sec-butyl, isobutyl, tert-butyl, neopentyl, cyclopropyl, cyclobutyl, cyclopentyl, cyclohexyl, ethylene (vinyl), propenyl, butenyl, pentenyl, 1-methyl-2-buten-1-yl, 5-hexenyl, ethynyl, 1-propynyl, 2-propynyl, etc.
[0122] As used herein, the term "click chemistry" refers to one or more covalent bonds and associated functional groups formed through a "click chemistry" reaction, which stably link one molecule (such as a positively charged labeled molecule) to another molecule (such as a substrate nucleotide), forming a linker. Click chemistry is characterized by high yields, high selectivity, mild reaction conditions (e.g., it can be carried out in an aqueous phase), harmless byproducts, and easy separation. Preferred click chemistry reactions in this invention include, but are not limited to, copper-catalyzed azido-alkyne cycloaddition reactions (CuAAC) and metal-free strain-promoted azido-alkyne cycloaddition reactions (SPAAC). Therefore, "click chemistry" can specifically refer to stable chemical structures such as triazole rings formed by these reactions.
[0123] As used herein, the term "substrate nucleotide" specifically refers to a modified deoxynucleoside polyphosphate used in the sequencing methods described herein. Its structure typically includes: (a) a nucleobase (A, G, C, or T); (b) a deoxyribose sugar; and (c) a polyphosphate chain containing three or more phosphate groups (e.g., triphosphate, tetraphosphate, pentaphosphate, hexaphosphate, etc.). Crucially, the terminal phosphate group (γ-position or further) of this polyphosphate chain is covalently linked to a "positively charged labeling molecule" as described herein via a linker (e.g., a click chemistry). In polymerase-catalyzed DNA synthesis, when the substrate nucleotide pairs with the template strand, the polymerase cleaves its α-β phosphate bond, incorporating the monophosphate portion of the nucleotide into the newly synthesized DNA strand, while simultaneously releasing the labeling molecule portion linked to the β-position and distal phosphate groups.
[0124] As used herein, the term "blocking current" refers to the reduced ionic current level exhibited by a nanopore after it has trapped an analyte molecule. In a typical nanopore detection system, the current measured when there is no analyte within the pore is called the "open-pore current." When an analyte (such as the positively charged labeled molecule of this invention) enters and temporarily occupies the sensing region within the pore, it physically impedes the free flow of ions, causing the measured current to drop to a new, lower, stable level, which is called the "blocking current." The characteristics of the blocking current, including its percentage of the residual current relative to the open-pore current and the duration of the blocking event, are closely related to the size, shape, charge, and interaction with the analyte's inner wall of the pore, and thus can serve as a "fingerprint" signal for identifying the molecule.
[0125] I. NetB protein and NetB nanoporous protein, and their applications in nucleic acid sequencing
[0126] NetB protein mutants
[0127] Through extensive and innovative work, the inventors have surprisingly achieved the application of NetB protein in the detection of nucleotides with positively charged molecular tags, and further in nucleic acid sequencing, for example, as a detection component. This application, based on NetB's preferred trapping properties for positively charged molecules, differs from existing negatively charged nanopores (such as α-hemolysin), thus improving signal purity and channel lifetime (experiments show >12 hours, see Table 2 for details).
[0128] The core of this invention lies in the first-ever disclosure that the NetB protein family, including wild-type NetB proteins, can be creatively used to construct high-performance nucleic acid sequencing systems. This invention further provides a series of optimized NetB protein mutants, modified NetB proteins for coupling with DNA polymerase, nanopores assembled from these proteins, and ultimately, nanopore-polymerase functional complexes.
[0129] In this disclosure, the NetB protein includes wild-type NetB protein and NetB protein mutants.
[0130] The inventors optimized the biochemical properties of the wild-type NetB protein with the amino acid sequence SEQ ID NO:1 and investigated its potential application on a single-molecule sequencing platform. This NetB protein is a pore-forming toxin produced by *Clostridium perfringens*, and its structure is deposited in the Protein Database (PDB) under accession code 4H56.
[0131] The present disclosure describes the NetB protein mutant obtained by modifying the wild-type NetB protein.
[0132] In some embodiments, the NetB protein mutants of this disclosure comprise an amino acid sequence with at least 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity compared to the wild-type NetB protein amino acid sequence shown in SEQ ID NO:1. In some embodiments, the NetB protein mutants comprise an amino acid sequence with at least 1, 2, 3, 4, or 5, up to at least 20, 21, 22, 23, 24, or 25 amino acid residues substituted, deleted, or inserted compared to the wild-type NetB protein amino acid sequence shown in SEQ ID NO:1.
[0133] In some embodiments, the amino acid sequence of the NetB protein mutant, compared to the amino acid sequence of the wild-type NetB protein shown in SEQ ID NO:1, includes one or more mutations selected from the group consisting of: deletion of 10 or 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 amino acid residues at the N-terminus of the wild-type NetB protein, and any one or more mutations of K20, K24, H53, K114, K115, N236, D285, Q284, and E289 to D, N, Q, S, T, Y, H, or E.
[0134] When described herein, the number of N-terminal deletion mutations is based on the wild-type NetB protein shown in SEQ ID NO:1 of this disclosure, calculated starting from the leftmost end of the sequence shown in SEQ ID NO:1. When described herein, the amino acid residue numbering is based on the complete sequence of the PDB crystal structure and follows its amino acid residue numbering; in other words, numbering begins at the 5th position from the leftmost end of the sequence shown in SEQ ID NO:1 of this disclosure.
[0135] In some embodiments, the NetB protein mutant of this disclosure includes, compared with the amino acid sequence of the wild-type NetB protein shown in SEQ ID NO:1, one or more mutations selected from the group consisting of: deletion of 10 amino acid residues from the N-terminus of the wild-type NetB protein, deletion of 11 amino acid residues from the N-terminus of the wild-type NetB protein, deletion of 12 amino acid residues from the N-terminus of the wild-type NetB protein, deletion of 13 amino acid residues from the N-terminus of the wild-type NetB protein, deletion of 14 amino acid residues from the N-terminus of the wild-type NetB protein, deletion of 15 amino acid residues from the N-terminus of the wild-type NetB protein, deletion of 16 amino acid residues from the N-terminus of the wild-type NetB protein, deletion of 17 amino acid residues from the N-terminus of the wild-type NetB protein, deletion of 18 amino acid residues from the N-terminus of the wild-type NetB protein, deletion of 19 amino acid residues from the N-terminus of the wild-type NetB protein, deletion of 20 amino acid residues from the N-terminus of the wild-type NetB protein, H53D , H53Y, H53N, H53Q, H53E, N236D, N236E, Q284D, Q284E, K20D, K20N, K20Q, K20S, K20T, K20 Y, K20H, K20E, K24N, K24D, K24Q, K24S, K24T, K24Y, K24H, K24E, K114N, K114D, K114Q, K11 4S, K114T, K114Y, K114H, K114E, K115N, K115D, K115Q, K115S, K115T, K115Y, K115H, K115 E, D285N, D285Q, D285S, D285T, D285Y, D285H, E289N, E289Q, E289S, E289T, E289Y, E289H.
[0136] In some embodiments, the amino acid sequence of the NetB protein mutant of this disclosure, compared with the wild-type NetB protein amino acid sequence shown in SEQ ID NO:1, includes any one of the following two mutations, or both:
[0137] (1) Deletion mutations, including deletions of 10 to 20 amino acid residues from the N-terminus of the wild-type NetB protein; and
[0138] (2) Substitution mutations, including one or more mutations selected from groups consisting of the following mutations, for example, 1 to 6, 1 to 5, 1 to 4, 2 to 6, or 2 to 5 mutations: H53D, H53Y, H53N, H53Q, H53E, N236D, N236E, Q284D, Q284E, K20D, K20N, K20Q, K20S, K20T, K20Y, K20H, K20E, K24N, K24D, K24Q, K24S, K24T, K24Y. K24H, K24E, K114N, K114D, K114Q, K114S, K114T, K114Y, K114H, K114E, K115N, K115D, K115Q, K115S, K115T, K115Y, K115H, K115E, D258N, D258Q, D258S, D258T, D258Y, D258H, E289N, E289Q, E289S, E289T, E289Y, and E289H.
[0139] In some embodiments, the NetB protein mutant of this disclosure includes an amino acid sequence that, compared to the wild-type NetB protein amino acid sequence shown in SEQ ID NO:1, includes the deletion of 10 to 20 amino acid residues from the N-terminus of the wild-type NetB protein.
[0140] In some preferred embodiments, deletion mutations include the deletion of 10-20, 11-20, 13-20, or 18-20 amino acid residues at the N-terminus of the wild-type NetB protein.
[0141] In some embodiments, the amino acid sequence of the NetB protein mutant of this disclosure, compared with the amino acid sequence of the wild-type NetB protein shown in SEQ ID NO:1, includes one or more mutations (substitution mutations) selected from the group consisting of the following mutations: H53D, H53Y, H53N, H53Q, H53E, N236D, N236E, Q284D, Q284E, K20D, K20N, K20Q, K20S, K20T, K20Y, K20H, K20E, K24N, K24D, K24Q, K24S, K24T, K24Y, K24H. K24E, K114N, K114D, K114Q, K114S, K114T, K114Y, K114H, K114E, K115N, K115D, K115Q, K115S, K115T, K115Y, K115H, K115E, D258N, D258Q, D258S, D258T, D258Y, D258H, E289N, E289Q, E289S, E289T, E289Y, and E289H.
[0142] In some preferred embodiments, the substitution mutation includes one or more mutations selected from the group consisting of the following mutations, for example, 1 to 6, 1 to 5, 1 to 4, 2 to 6, or 2 to 5 mutations: K20D, K20N, K24D, K24N, K24S, H53D, K114D, K114N, K115D, K115T, N236D, N236E, Q284D, Q284E, D285H, D285N, E289Q, and E289Y.
[0143] In some embodiments, the NetB protein mutant of this disclosure includes the following two mutations compared to the amino acid sequence shown in SEQ ID NO:1:
[0144] (1) Deletion mutations, including deletions of 10 to 20 amino acid residues from the N-terminus of the wild-type NetB protein; and
[0145] (2) Substitution mutations, including one or more mutations selected from groups consisting of the following mutations, for example, 1 to 6, 1 to 5, 1 to 4, 2 to 6, or 2 to 5 mutations: H53D, H53Y, H53N, H53Q, H53E, N236D, N236E, Q284D, Q284E, K20D, K20N, K20Q, K20S, K20T, K20Y, K20H, K20E, K24N, K24D, K24Q, K24S, K24T, K24Y. K24H, K24E, K114N, K114D, K114Q, K114S, K114T, K114Y, K114H, K114E, K115N, K115D, K115Q, K115S, K115T, K115Y, K115H, K115E, D258N, D258Q, D258S, D258T, D258Y, D258H, E289N, E289Q, E289S, E289T, E289Y, and E289H.
[0146] In some embodiments, the amino acid sequence of the NetB protein mutant of this disclosure includes:
[0147] 1) The amino acid sequence shown in SEQ ID NO:2, SEQ ID NO:3, SEQ ID NO:4, SEQ ID NO:5, SEQ ID NO:6, SEQ ID NO:7, SEQ ID NO:8, SEQ ID NO:9, SEQ ID NO:10, SEQ ID NO:11 or SEQ ID NO:12; or
[0148] 2) An amino acid sequence having at least about 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with the amino acid sequences shown in SEQ ID NO:2, SEQ ID NO:3, SEQ ID NO:4, SEQ ID NO:5, SEQ ID NO:6, SEQ ID NO:7, SEQ ID NO:8, SEQ ID NO:9, SEQ ID NO:10, SEQ ID NO:11, or SEQ ID NO:12; or
[0149] 3) An amino acid sequence having at least 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 amino acid residues substituted, deleted or inserted compared to the amino acid sequences shown in SEQ ID NO:2, SEQ ID NO:3, SEQ ID NO:4, SEQ ID NO:5, SEQ ID NO:6, SEQ ID NO:7, SEQ ID NO:8, SEQ ID NO:9, SEQ ID NO:10, SEQ ID NO:11 or SEQ ID NO:12.
[0150] This disclosure provides a series of specific NetB protein mutants with superior performance, the amino acid sequences of which are shown in SEQ ID NO:2 to SEQ ID NO:12. Specific mutation combinations of these sequences (relative to SEQ ID NO:1) are shown in Table 1 below.
[0151] Table 1. Specific NetB protein mutants
[0152] These mutations improve the stability and positive charge trapping efficiency of NetB, for example, the orifice current of SEQ ID NO:10 is more stable (see Figure 24 for details).
[0153] Upstream coding nucleic acid and production system
[0154] This disclosure also provides an isolated nucleic acid molecule encoding any of the NetB protein mutants described herein. The nucleic acid molecule may be DNA or RNA.
[0155] This disclosure further provides an expression vector, such as a plasmid, which comprises the aforementioned nucleic acid molecule and optionally includes regulatory elements such as a promoter for regulating the expression of the nucleic acid molecule.
[0156] This disclosure also provides a host cell, such as an Escherichia coli or yeast cell, which contains the aforementioned expression vector and is capable of expressing the NetB protein mutant.
[0157] Modified NetB protein for coupling
[0158] It is important to understand that the nanoporous proteins themselves, formed by the self-assembly of NetB proteins (including wild-type or mutants) described in this disclosure, can serve as a general single-molecule sensor, for example, to detect the passage of specific small molecules, peptides, or ions. However, for sequencing-by-synthesis technology, a "naked" nanopore composed solely of NetB proteins cannot independently complete the sequencing task because it lacks the functional module to recognize the DNA template and catalyze the synthesis of new strands.
[0159] To achieve sequencing-while-synthesizing, NetB nanopores need to be coupled to DNA polymerase. Therefore, this disclosure provides a modified NetB protein comprising a NetB protein and a linker for coupling with DNA polymerase. The NetB protein can be a wild-type NetB protein (SEQ ID NO:1) or any of the NetB protein mutants described in this disclosure. The linker is a chemical moiety capable of participating in covalent or non-covalent protein linkage reactions.
[0160] In some embodiments, the linker is selected from the group consisting of: SpyTag peptide, SnoopTag peptide, sorting enzyme recognition site, chemically cross-linkable amino acid residues (such as cysteine), biotin, HaloTag, or SNAP-tag, etc. This invention is not limited to employing any particular coupling technique.
[0161] NetB nanoporin
[0162] When used in this document, the terms “NetB nanopore” and “NetB nanopore protein” are used interchangeably.
[0163] The NetB nanoporous protein disclosed herein is a polymer composed of one or more NetB protein monomers as subunits, for example, a polymer composed of at least 2, 3, 4, 5, 6, 7, or 8 NetB protein monomers as subunits, preferably a 7-mer composed of 7 NetB proteins as subunits, as shown in Figure 1. The NetB protein monomers may be wild-type NetB proteins including the amino acid sequence shown in SEQ ID NO:1, NetB protein mutants as described in this disclosure, or modified NetB proteins as described in this disclosure.
[0164] The NetB nanoporous protein disclosed herein can be a homomer, meaning all NetB protein subunits are identical; or it can be a heteromer, meaning it is composed of different types of NetB protein subunits. In some embodiments, the multiple NetB protein monomers of this disclosure that compose the NetB nanoporous protein can be the same or different. For example, when the NetB nanoporous protein is composed of 7 NetB protein monomers, these 7 NetB protein monomers can be completely identical, not completely identical, or completely different, for example, including 1, 2, 3, 4, 5, 6, or 7 types of NetB protein monomers. Preferably, the NetB nanoporous protein is composed of 7 identical NetB protein monomers, optionally the NetB protein monomers are wild-type NetB protein or NetB protein mutants, such as the NetB protein mutants of this disclosure.
[0165] When multiple NetB protein monomers form a polymer, i.e., NetB nanoporous protein, as subunits, the connection between NetB protein monomers is non-covalent and depends on the interaction forces between proteins, such as hydrogen bonds, electrostatics, van der Waals forces, dipole moment forces, etc.
[0166] In some embodiments, at least one NetB protein monomer in the NetB nanoporous protein is a modified NetB protein containing a linker as disclosed herein. The resulting NetB nanoporous protein can be coupled to DNA polymerase via the linker on the modified NetB protein, thereby enabling sequencing tasks to be performed in a "synthesis-while-sequencing" process.
[0167] In some implementations, the NetB nanoporous protein is a heptamer, comprising:
[0168] (1) Six NetB proteins that do not contain linkers, wherein the linker-free NetB proteins are selected from...
[0169] (a) Wild-type NetB protein comprising the amino acid sequence shown in SEQ ID NO:1; and
[0170] (b) The NetB protein mutant disclosed herein;
[0171] (2) A modified NetB protein disclosed herein.
[0172] In some embodiments, the six linkerless NetB proteins may be the same or different. In some preferred embodiments, the six linkerless NetB proteins are the same NetB protein, such as a wild-type NetB protein or a mutant of the same NetB protein disclosed herein.
[0173] In some embodiments, the amino acid sequence contained in the modified NetB protein may be the same as or different from the amino acid sequences of the six NetB proteins that do not contain linkers.
[0174] In some preferred embodiments, the six NetB proteins without linkers are the same NetB protein, and the modified NetB protein contains the same amino acid sequence as the six NetB proteins without linkers. That is, in some preferred embodiments, the NetB nanoporous protein comprises seven NetB protein monomers with the same amino acid sequence, one of which contains a linker that couples with a DNA polymerase.
[0175] In some embodiments, this disclosure provides a NetB nanoporous protein, which is a polymer composed of at least two protein monomers as subunits, wherein each of the at least two protein monomers is independently selected from:
[0176] (a) A wild-type NetB protein containing the amino acid sequence shown in SEQ ID NO:1; or
[0177] (b) The NetB protein mutant described in this disclosure.
[0178] In some embodiments, the polymer is a heptamer, and at least one of the protein monomers contains a linker for coupling with DNA polymerase.
[0179] In some embodiments, the linker is a chemical moiety capable of participating in covalent or non-covalent protein linking reactions, selected from the group consisting of: SpyTag peptide, SnoopTag peptide, sorting enzyme recognition site, chemically cross-linkable amino acid residue, biotin, HaloTag, or SNAP-tag.
[0180] Nanopore-polymerase complex
[0181] This disclosure also provides a protein complex comprising the aforementioned NetB nanoporous protein and a DNA polymerase coupled thereto via a linker.
[0182] In some implementations, to match the sequencing-by-synthesis technique, a special coupling scheme is used to form a stable complex of the NetB multimer and the nucleic acid polymerase in a 1:1 ratio on the functional multimer of the NetB protein. In some implementations, such a complex can be achieved based on the widely reported spytag-spycatcher coupling mechanism (cited in: Proc Natl Acad Sci US A. 2016 Nov 1; 113(44):E6749-E6756). By fusion expression, a Spytag peptide is added to the C-terminus of one of the NetB multimer monomers (the NetB protein with the spytag peptide is hereby referred to as the coupling monomer), and a spycatcher domain is added to the C-terminus of the nucleic acid polymerase. Then, under mild conditions, the NetB nanopore and the polymerase can form a stable covalently coupled complex.
[0183] This invention is not limited to using the SpyTag / SpyCatcher system to couple NetB nanopores with polymerase. Those skilled in the art will understand that any technique capable of achieving stable, site-directed, or non-site-directed linkages between two proteins can be used to construct the nanopore-polymerase complex of this invention. These optional coupling techniques include, but are not limited to: other protein covalent linkage systems (such as SnoopTag / SnoopCatcher), enzyme-catalyzed linkages (such as sorting enzyme-mediated linkages), chemical cross-linking, linkages based on high-affinity non-covalent interactions (such as the biotin-streptavidin system), and direct gene fusion, etc.
[0184] In some implementations, an additional purification step is required to obtain NetB nanopores containing only one conjugated monomer. This ensures that each nanopore is conjugated to only one polymerase to sequence a single nucleic acid molecule, avoiding interference between sequencing signals from different nucleic acid molecules. In these implementations, the conjugated monomer has a His tag that carries a positive charge in a low pH environment; while the corresponding unconjugated monomer NetB has its His tag removed using a protease. Therefore, when the conjugated and unconjugated NetB monomers are mixed together to spontaneously form complete heptamer nanopores, a random ratio distribution of 0:7, 1:6, 2:5, etc., will occur. In some implementations, a 1:6 heptamer NetB nanopore protein is preferred for sequencing applications. In some implementations, the NetB nanopore protein is conjugated to the polymerase in a 1:1 ratio.
[0185] Separating polymers with a specific ratio from heterogeneous mixtures containing different monomer ratios (e.g., 0:7, 1:6, 2:5) can utilize the property that His tags, retained on coupled monomers but removed on uncoupled monomers, are protonated in acidic buffers (e.g., pH < 6.0) to carry a stable positive charge. The net positive charge carried by the heptamer complex at low pH is proportional to the number of coupled monomers it contains. This property can be used to specifically separate 1:6 NetB heptamers using ion exchange chromatography.
[0186] II. Positively charged labeled molecules and positively charged labeled nucleotides
[0187] Given the shortcomings of existing technologies that use negatively charged labeling molecules to label nucleotides, resulting in strong interference of the label signal by DNA and severely affecting signal recognition, the inventors have developed a novel positively charged labeling molecule and combined it with nucleotides to obtain positively charged labeled nucleotides, thereby overcoming the signal recognition interference problem in single-molecule nucleic acid sequencing.
[0188] Positively charged labeled molecules
[0189] One aspect of this disclosure relates to a positively charged labeling molecule that can be used as a tool molecule to link to a nucleotide, thereby giving the nucleotide a positive charge.
[0190] This disclosure provides a positively charged labeled molecule, as shown in the following formula (I): MX Formula (I)
[0191] in,
[0192] The labeled portion M, containing one or more positive charges, is used to be trapped by the nanopore in an electric field and generate a characteristic signal; and
[0193] The reactive linker part X is a chemical group suitable for covalently linking with the polyphosphate moiety of nucleotides.
[0194] In some embodiments, the label portion M itself is a chain polymer containing multiple positively charged side chains, which can be represented as R. The positive charge of R is distributed on its polymer chain. This indicates the corresponding positively charged form. Therefore, in some embodiments, the positively charged labeled molecule of this disclosure, or its positively charged form, is as shown in the following formula (Ia):
[0195] RX, or Formula (Ia)
[0196] in,
[0197] R represents the chain-like structure marker portion with an amino side chain. This indicates the corresponding positively charged chain-like structure with amino side chains.
[0198] X indicates the ability to convert R or The linker portion connected to the terminal phosphate group of a dNNP can also be referred to as a reactive linker portion suitable for covalently linking to the polyphosphate portion of a nucleotide. Those skilled in the art will understand that the meanings of these two definitions are the same.
[0199] In some embodiments, the main chain of the chain-like structural portion R is optionally a heterohydrocarbon chain containing one or more heteroatoms selected from N, O or S, with a main chain length of 50-300 atoms.
[0200] In some embodiments, the chain-like structural portion R has one or more amino side chains, such as 1-12, 1-10, 1-8, 1-6, or 1-4 amino side chains. When multiple amino side chains are present, they can be the same or different.
[0201] As used herein, the term "amino side chain" refers to a short chain structure containing one or more amino groups (-NH2) attached to the main chain of a chain-like structure. In some embodiments, the amino side chain has a chain length of 3-12, for example 3-10, 3-8, 3-6, or 1-2 atoms. In some embodiments, the amino side chain is an amino group (-NH2). In some embodiments, the amino side chain is selected from the side chains of amino acids, such as the side chains of lysine, arginine, or α,β-diaminopropionic acid (Dap). In some preferred embodiments, R is a polylysine, polyarginine chain, or polyDap. It will be understood by those skilled in the art that within the scope of this disclosure, during the polymerization of amino acids, either the α-amino group or the amino group on the side chain can form an amide bond with the carboxyl group.
[0202] In the positively charged labeled molecules of this disclosure, the positively charged amino side chains are sufficient to enable the chain-like labeled portion R to be effectively captured by the NetB nanoporous protein of this disclosure and generate a blocking current within a typical nanopore detection voltage range (e.g., 50-200 mV). As illustrated in the examples of this disclosure, the specific structure and density of the amino side chains result in different blocking currents upon interaction with NetB, which can be used as characteristic values to label and distinguish different dNNPs.
[0203] In some implementations, the reactive linker portion X is capable of, for example, transferring R or... Chemical structure of the polyphosphate moiety attached to dNNP.
[0204] In some embodiments, to provide broader applicability, the labeled portion M consists of two parts: (i) a positively charged head portion, whose primary function is to provide driving force under an electric field, for example, a polypeptide composed of 3-15 lysine, arginine, or Dap residues; and (ii) a chain-like structure portion (R'), also known as the signal discrimination body portion, whose primary function is to generate a distinguishable blocking current signal, for example, whose structure can be a polypeptide chain, a polyethylene glycol (PEG) chain, an alkane chain, or a combination thereof. Therefore, in some embodiments, the positively charged labeled molecule of this disclosure, or its positively charged form, is as shown in the following formula (Ib):
[0205] Head-R'-X, or Formula (Ib)
[0206] in,
[0207] Head refers to the head portion that can carry a positive charge. This represents the corresponding positively charged form.
[0208] R' represents the chain-like structure, which can also be called the main part for signal discrimination.
[0209] X represents the reactive linker part X, which is suitable for covalent linking with the polyphosphate moiety of nucleotides.
[0210] In some embodiments, the head portion is a polymeric group composed of multiple amino acid residues or amino groups, formed by the polymerization of multiple amino acid monomers via amide bonds, and may be represented as Poly(AA) in some embodiments. In some embodiments, at least a portion of the amino acid residues are positively charged amino acid residues.
[0211] Therefore, in some embodiments, the positively charged labeled molecule of this disclosure, or its positively charged form, is as shown in the following formula (Ic):
[0212] Poly(AA)-R'-X, or Formula (Ic)
[0213] in,
[0214] Poly(AA) represents the head portion of a polymer group consisting of three or more amino acid residues or amino groups. This represents the corresponding positively charged form.
[0215] R' represents the chain-like structure, which can also be called the main part for signal discrimination.
[0216] X represents a linker moiety capable of connecting R' to the terminal phosphate group of a dNNP, and can also be referred to as a reactive linker moiety suitable for covalently linking to the polyphosphate group of a nucleotide, as defined elsewhere in this document. Those skilled in the art will understand that the meanings of these two definitions are the same.
[0217] (i) A positively charged head group (also referred to as Poly(AA) in some embodiments of this disclosure) primarily functions to provide driving force under an electric field. The head group is a polymer formed by the polymerization of multiple amino acid monomers via amide bonds, wherein at least a portion of the amino acid monomers are positively charged amino acids.
[0218] In embodiments of this disclosure, it is not necessary to distribute positively charged groups throughout the nucleotide marker. Instead, by simply adding three or more positively charged groups to the ends of the chain molecule, it is possible to effectively capture the nucleotide within a voltage range of 60-180 mV using NetB nanopores.
[0219] In this disclosure, the formation of the "amide bond" is varied. In some embodiments, the amide bond is a standard peptide bond formed by the α-amino group of one amino acid and the carboxyl group of another amino acid, thereby constituting an α-peptide polymer. In other embodiments, the amide bond can be an isopeptide bond formed by the side-chain amino group of one amino acid (e.g., the ε-amino group of lysine) and the carboxyl group of another amino acid, thereby constituting an isopeptide polymer, such as an ε-peptide. This invention covers any polymer formed from standard peptide bonds, isopeptide bonds, or a mixture of both.
[0220] The term "positively charged amino acid" refers to an amino acid whose side chain can be protonated to carry a positive charge. A common characteristic of these amino acids is that their side chain contains at least one group selected from primary, secondary, and guanidinium groups. This covers not only proteinogenic amino acids, such as lysine (with a primary amino group at the end of its side chain) and arginine (with a guanidinium group in its side chain), but also non-proteinogenic amino acids with similar side chain characteristics, such as α,β-diaminopropionic acid (Dap), α,γ-diaminobutyric acid (Dab), and ornithine (Orn).
[0221] In some embodiments, the head group or Poly (AA) may be a polymer formed by polymerizing three or more, for example, three to fifteen, such as three, four, five, six, seven, eight, nine, ten, eleven, twelve, thirteen, fourteen, or fifteen such positively charged amino acid monomers (which may be the same or different, and are polymerized by any one or more of the above-mentioned amide bonds).
[0222] For ease of synthesis and uniformity of product structure, in some preferred embodiments, the head group or Poly(AA) is a homopolymer. This means that it is polymerized from a single class of positively charged amino acid monomers (e.g., only lysine) through the same amide bond formation mode (e.g., all α-peptide bonds, or all ε-peptide bonds).
[0223] In some embodiments, Poly(AA) consists of three or more amino acid residues, for example, selected from lysine, arginine, α,β-diaminopropionic acid (Dap), α,γ-diaminobutyric acid (Dab), or ornithine (Orn), preferably lysine, arginine, or α,β-diaminopropionic acid (Dap). In some embodiments, Poly(AA) is PolyK, i.e., polylysine. In some embodiments, Poly(AA) is poly(Dap).
[0224] (ii) The chain-like structure (R') has the core function of generating a characteristic blocking current signal that can be recognized and distinguished by the detector through its own size, shape, charge distribution and other physicochemical properties when the labeled molecule passes through the nanopore, thereby realizing the identification of different nucleotides. Therefore, it can also be called the main part of signal identification in this disclosure, and the two can be used interchangeably.
[0225] Structurally, the chain-like structure portion R' is a chain-like polymer, whose main chain backbone optionally contains one or more heteroatoms selected from N, O or S. The main chain length is preferably 50-300 atoms, corresponding to a physical length of about 10-20 nm.
[0226] In some embodiments, the main chain comprises a combination selected from one or more of the following structural units: (a) amino acid residues linked by amide bonds; (b) alkylene structural units containing ether bonds, such as polyethylene glycol (PEG) units; and (c) alkylene units.
[0227] One of the core advantages of this invention lies in the high designability of the chain-like structural portion R', which allows for the systematic adjustment and optimization of the generated blocking current signal, thereby enabling the identification of the types of nucleotides it is linked to. To obtain four or more distinct signals, the structure of R' can be designed and modified in one or more of the following ways:
[0228] Main chain composition: The main chain of R' can be composed of homopolymers or block copolymers such as polypeptide chains, alkylene chains containing ether bonds (e.g., polyethylene glycol (PEG) chains), and alkane chains. In some embodiments, R' can be a hybrid polymer formed by polymerizing different structural units (e.g., different amino acids, PEG units of different lengths, etc.) through solid-state synthesis or other methods. Optionally, these structural units are linked by amide bonds.
[0229] Topology: R' can be linear or branched. By introducing branched structures, its spatial occupancy within the nanopore can be significantly altered, thereby modulating the signal.
[0230] Radial Scale and Side Chains: R' can synthesize the main chain by selecting structural units with different radial scales (i.e., "thickness"), or by introducing side chains of different lengths and properties onto the main chain. For example, introducing different amino acid side chains (such as lysine, arginine, asparagine, glutamine, etc.) or adjusting the distribution density of side chain units can effectively change the blocking rate.
[0231] Charge distribution: R' itself can be net positive, net negative, or electrically neutral. By introducing charged groups into the structure, its interaction with the inner wall of the nanopore can be finely tuned, thereby obtaining distinctive signals.
[0232] Hybrid-signal design (electro-optical): In addition to altering the electrical signal by adjusting the aforementioned physicochemical properties, optically active groups can be introduced into the structure of R'. For example, amino acid side chains with intrinsic fluorescence (such as tryptophan) can be selected, or optical reporter groups such as fluorescent groups or Raman probes can be introduced onto specific amino acid side chains through chemical modification. This design makes it possible to develop hybrid-mode (electro-optical) detection systems, i.e., simultaneously detecting the current blocking signal and exciting and detecting the corresponding optical signal (such as fluorescence) through an external light source. Through this dual-signal cross-validation, the accuracy and reliability of base recognition can be further improved.
[0233] In some preferred embodiments, R' is a polymer composed of one or more structural modules combined in a linear or branched manner. The structural modules are selected from the group consisting of:
[0234] a) Peptide module: A peptide segment consisting of 1 to 20 amino acid residues. These amino acid residues can be selected to impart specific properties to R'. For example:
[0235] Polar and neutral amino acids, such as glutamine (Gln), asparagine (Asn), and serine (Ser), are used to regulate hydrophilicity and the formation of hydrogen bonds.
[0236] Aromatic amino acids, such as tyrosine (Tyr) and tryptophan (Trp), have large side chain volumes that can be used to generate significant blocking signals;
[0237] Charged amino acids: such as aspartic acid (Asp), glutamic acid (Glu), lysine (Lys), and diaminopropionic acid (Dap).
[0238] Used for fine-tuning the overall charge and signal of R'.
[0239] Special structural amino acids, such as proline (Pro), are used to introduce rigid kinks into polymer chains.
[0240] b) Spacer modules: These are non-peptide chain structures with flexible or rigid structures, used to adjust the distance between modules and the overall conformation of the polymer. For example:
[0241] Alkylene modules containing ether bonds: used to provide flexibility and modulate hydrophilicity / hydrophobicity. For example, the module may be a polyethylene glycol (PEG) module having a -(O-CH2-CH2-)n- repeating unit, where n can be 1 to 24, for example 2-20, 2-16, 2-12, or 2-10, or preferably 4-18, 4-16, 4-14, 4-12, 4-10, or 4-8, or any value within the above ranges and ranges consisting of these arbitrary values. In other embodiments, the module may also be a non-PEG ether-containing structure, for example containing a -(CH2)xO-(CH2)y- structural unit, where x and y are each independently integers from 1 to 10, to provide different chain rigidities and spatial conformations.
[0242] Alkylene modules: those with -(CH2-) m - Repeating unit, where m can be 1 to 20, for example m is 4-16, 4-12, or 4-8, or preferably 6-12, 6-10, or 6-8, or any value within the above range and a range of such arbitrary values, to provide hydrophobicity and different chain stiffness.
[0243] By combining these modules in different orders and proportions, the physicochemical properties of R' (such as size, charge, hydrophilicity / hydrophobicity, rigidity / flexibility) can be systematically tuned to obtain a set (e.g., four or more) of R' structures that can generate mutually distinguishable characteristic blocking current signals when passing through nanopores. The PA-1 to PA-52 molecules, as shown in the following specific examples, are concrete examples of the combination of the signal discrimination host (R') and the reactive linker X. They fully demonstrate that by combining different amino acids (such as glutamine, proline, tyrosine, tryptophan, etc.) and PEG chains as structural units, a rich variety of R'-X structures with different signal characteristics can be obtained through the above strategy.
[0244] It should be understood that although the structure of the signal identification body (R') exhibits significant diversity, its scope is clear and well-defined. This disclosure provides a clear set of modular design principles. Those skilled in the art, upon reading this disclosure, will appreciate that R' can be constructed and optimized by: (1) selecting suitable building blocks from the disclosed "peptide modules" and "spacer modules"; (2) combining these modules in a linear or branched manner; and (3) systematically modulating the physicochemical properties of the polymer (such as size, charge, hydrophilicity, etc.) to obtain target molecules with specific blocking current signals.
[0245] Therefore, this disclosure provides a widely applicable technical solution that allows various chain structures, regardless of whether they carry a positive charge, to be effectively captured by nanopores, greatly enriching the design ideas for positively charged labeled nucleotides and enabling the acquisition of labeled molecules with different blocking signals.
[0246] (iii) A reactive linker portion X, which is a chemical group suitable for covalently linking to the polyphosphate moiety of a nucleotide. This disclosure preferably employs bioorthogonal chemistry to achieve the linking of the labeled portion to the nucleotide moiety, which has the advantages of high selectivity, high yield, and mild reaction conditions.
[0247] In some embodiments, the linking reaction is click chemistry. Therefore, the reactive linker moiety X is one of a pair of reactive groups capable of participating in a click chemistry reaction. Specifically, the reactive linker moiety X can be selected from the group consisting of:
[0248] Azide group (-N3): Used to react with nucleotides containing an alkyne group.
[0249] Alkyne group: Used for reaction with nucleotides containing an azide group. In some embodiments, the alkynyl group is a terminal alkynyl group (-C≡CH), which participates in copper (I)-catalyzed azido-alkynyl cycloaddition reactions (CuAAC). In other embodiments, the alkynyl group is a cyclic strained alkynyl group, such as dibenzocyclooctyne (DBCO) and its derivatives, which can participate in metal-free stress-promoted azido-alkynyl cycloaddition reactions (SPAAC).
[0250] Those skilled in the art should know that, through mature chemical synthesis techniques, any of the aforementioned reactive linker portions X can be introduced to the end of the signal identification host portion R', while simultaneously introducing the paired reactive group to the nucleotide polyphosphate portion.
[0251] The click chemistry reactions of azides and alkyne-containing structures are shown below:
[0252] 1.
[0253] 2.
[0254] R1 and R2 are the parts containing chain-like R' and the parts containing dNNP, respectively.
[0255] In some implementations, the structure formed by the signal discrimination main portion R' and the reactive connection sub-portion X includes, but is not limited to, the following exemplary structures:
[0256] PA-1
[0257] PA-2-1
[0258] PA-2-2
[0259] PA-3
[0260] PA-5
[0261] PA-6
[0262] PA-7
[0263] PA-8
[0264] PA-9
[0265] PA-10
[0266] PA-11
[0267] PA-12
[0268] PA-13
[0269] PA-14
[0270] PA-15
[0271] PA-16
[0272] PA-17
[0273] PA-18
[0274] PA-19
[0275] PA-20
[0276] PA-21
[0277] PA-22
[0278] PA-23
[0279] PA-24
[0280] PA-25
[0281] PA-26
[0282] PA-27
[0283] PA-28
[0284] PA-29
[0285] PA-31
[0286] PA-33
[0287] In some embodiments, the two forms of positively charged labeled molecules disclosed in this invention can be combined.
[0288] In some embodiments, exemplary examples of positively charged labeled molecules of this disclosure include, but are not limited to, the following:
[0289] PA-1-1
[0290] PA-2-1
[0291] PA-6
[0292] PA-8
[0293] Positively charged nucleotides
[0294] Another aspect of this disclosure relates to a final product that can be directly used for nucleic acid sequencing, consisting of any of the positively charged labeled molecules described in this disclosure linked to the polyphosphate portion of a nucleotide. This product is a “positively charged labeled nucleotide”.
[0295] This disclosure provides a positively charged labeled nucleotide, as shown in formula (II): M-X'-dNNP Formula (II)
[0296] in,
[0297] The labeled portion M contains one or more positive charges for being captured by the nanopore in an electric field and generating a characteristic signal, as defined above.
[0298] The linker part (X') is formed by the reaction of the reactive linker part X as defined above with the polyphosphate part of the nucleotide. For details, please refer to the definition of the reactive linker part X above.
[0299] dNNP represents a polyphosphate nucleotide moiety with three or more phosphate groups.
[0300] In some embodiments, the label portion M itself is a chain polymer containing multiple positively charged side chains, which can be represented as R. The positive charge of R is distributed on its polymer chain. This indicates the corresponding positively charged form. Therefore, in some embodiments, the positively charged labeled molecule of this disclosure, or its positively charged form, is as shown in the following formula (IIa):
[0301] R-X'-dNNP, or Formula (IIa)
[0302] in,
[0303] R, As defined above, dNNP represents a polyphosphate nucleotide moiety with three or more phosphate groups.
[0304] In some embodiments, the marker portion M consists of two parts: (i) a positively charged head portion; and (ii) a chain-like structure portion (R'), also referred to as the signal identification body portion. Therefore, in some embodiments, the positively charged marker molecule of this disclosure, or its positively charged form, is as shown in the following formula (IIb):
[0305] Head-R'-X'-dNNP, or Equation (IIb)
[0306] Among them, Head, As defined above, dNNP represents a polyphosphate nucleotide moiety with three or more phosphate groups.
[0307] In some embodiments, the head portion is a polymeric group composed of multiple amino acid residues or amino groups, formed by the polymerization of multiple amino acid monomers via amide bonds, and may be represented as Poly(AA) in some embodiments. In some embodiments, at least a portion of the amino acid residues are positively charged amino acid residues.
[0308] Therefore, in some embodiments, the positively charged labeled molecule of this disclosure, or its positively charged form, is as shown in the following formula (IIc):
[0309] Poly(AA)-R'-X'-dNNP, or Formula (IIc)
[0310] in,
[0311] Poly(AA) R' and X' are defined above, and dNNP represents a polyphosphate nucleotide moiety with three or more phosphate groups.
[0312] In some embodiments, the positively charged nucleotides of this disclosure are selected from the following:
[0313] dC5P-PA-1-1
[0314] dG5P-PA-2-1
[0315] dT5P-PA-6
[0316] dA5P-PA-8
[0317] The positively charged nucleotides disclosed herein are carefully designed macromolecular structures possessing both amphiphilic and zwitterionic properties. Structurally, these molecules exhibit the following significant characteristics:
[0318] The "dumbbell-shaped" or "functionalized at both ends" structure: This molecule exhibits two functionally distinct ends in space. One end is a positively charged labeled portion, and the other end is a negatively charged nucleotide polyphosphate portion. These two portions are connected by a linker of a specific length and flexibility, forming a dumbbell-like configuration.
[0319] Amphoteric properties: Because the labeled portion is positively charged at physiological pH, while the nucleotide polyphosphate portion is negatively charged, the entire molecule possesses both positive and negative charge centers within the same structure, making it a typical zwitterion. This characteristic is the key difference between this invention and purely negatively charged labels in the prior art. It is worth noting that although the labeled nucleotide molecule of this disclosure exhibits zwitterionic properties, experimental results show that this 'dumbbell-shaped' structure demonstrates good solubility and stability under sequencing conditions, without leading to severe electrostatic aggregation between molecules or frequent blockage of nanopores, ensuring a smooth and efficient sequencing process.
[0320] Tunable physicochemical properties: As mentioned earlier, through modular design of the labeling part, the size, shape, charge distribution and hydrophilicity / hydrophobicity of the entire molecule can be systematically tuned, thereby designing four distinct labeled nucleotides that can produce four clearly distinguishable signals for four different bases (A, T, C, G).
[0321] This unique structure endows the positively charged nucleotides disclosed herein with significant advantages in applications, particularly in nanopore sequencing:
[0322] Key Advantages: Avoiding DNA Interference and Improving Signal-to-Noise Ratio: In traditional nanopore sequencing systems, negatively charged labels and equally negatively charged DNA template strands compete to enter the nanopore, severely interfering with the signal. The fundamental innovation of this invention lies in its "reversed electric field polarity" application logic: by using positively charged labels in conjunction with a nanopore that preferentially captures positively charged molecules (such as the NetB nanopore protein described in this disclosure), the sequencing system can "directionally" capture the labels while "actively" rejecting and ignoring the DNA template strands also located near the pore opening. This fundamentally eliminates background noise from the DNA, resulting in a purer and clearer detected signal, significantly improving the signal-to-noise ratio and accuracy of sequencing.
[0323] Main applications: As a substrate for sequencing-by-synthesis: The positively charged labeled nucleotides of this disclosure are specifically designed as substrates for nanopore-based sequencing-by-synthesis technologies. During sequencing, DNA polymerase binds to the DNA template to be tested and captures the paired labeled nucleotides of this disclosure from the solution. When the nucleotide is incorporated into a newly synthesized DNA strand, its attached labeled portion is pulled through the nanopore by an electric field, generating a characteristic current blocking signal. By recognizing this signal, the type of base incorporated can be determined, thereby enabling real-time reading of the DNA sequence.
[0324] Other potential applications: In addition to nanopore sequencing, the positively charged nucleotides disclosed herein can also be used in any technology platform that requires real-time monitoring of nucleic acid replication or amplification, such as for molecular diagnostics, single-molecule dynamics studies, etc.
[0325] In summary, the positively charged nucleotides disclosed herein, through their innovative zwitterionic structure, solve the long-standing signal interference problem in existing technologies, providing a key chemical basis for developing next-generation nucleic acid sequencing systems with higher precision and higher signal-to-noise ratio.
[0326] Preparation of positively charged labeled nucleotides
[0327] The "positively charged nucleotide" disclosed herein is prepared by covalently linking the "positively charged molecule" described above with a "reactive nucleotide polyphosphate".
[0328] 1. Reactants: Reactive nucleotide polyphosphates
[0329] In order to achieve efficient and specific ligation with the aforementioned positively charged labeled molecules (MX) having a reactive linker motif X, this disclosure also relates to a class of chemically modified reactive nucleotide polyphosphates.
[0330] In its general structure, this reactive nucleotide polyphosphate can be represented as YN, where:
[0331] N is the polyphosphate moiety of the nucleotide. It consists of a nucleobase (such as adenine A, guanine G, cytosine C, or thymine T), a deoxyribose, and a polyphosphate chain containing 3 or more phosphate groups (e.g., 3 to 10, 3 to 8, preferably 5 to 6).
[0332] Y is a paired reactive group. This group is attached to the terminal phosphate group of the polyphosphate chain and is selected to specifically react with the reactive linker moiety X on the labeled molecule.
[0333] To achieve efficient and selective linking, groups Y and X are preferably a pair of paired groups capable of participating in bioorthogonal reactions (especially click chemistry). Therefore, when the reactive linker portion X on the labeled molecule is an azide group (-N3), the paired reactive group (Y) on the nucleotide is preferably an alkynyl group.
[0334] In some embodiments, in order to achieve a connection without metal catalysis and under milder reaction conditions, the paired reactive group (Y) is a cyclic strained alkyne, such as dibenzocyclooctylene (DBCO) or a derivative thereof.
[0335] This disclosure covers four reactive nucleotide polyphosphate derivatives corresponding to A, G, C, and T, respectively. In some embodiments, examples of the structures of these derivatives are shown in the figure below, wherein the nucleotide moiety is a deoxyhexaphosphate nucleotide, and its terminal phosphate is linked to a DBCO group via an alkyl chain linker:
[0336] It should be understood that the DBCO-nucleotide structure shown in the figure above is merely an exemplary structure for illustrating the inventive concept and is not intended to limit the scope of the invention. Those skilled in the art will recognize that various feasible variations of the structure of the reactive nucleotide polyphosphate (YN) exist within the framework of this invention. For example:
[0337] The length of the polyphosphate chain: The number of phosphate groups is not limited to six, but varies between three and ten, to regulate its binding affinity to DNA polymerase.
[0338] Chemical chain connecting Y and N: The chemical chain connecting the paired reactive group (Y) to the terminal phosphate group can have its length and chemical properties (such as hydrophilicity / hydrophobicity) modified. For example, the butylamide structure in the figure can be replaced by a longer or shorter alkyl chain, or by a flexible chain containing polyethylene glycol (PEG) units.
[0339] Other reactive groups: Although DBCO is preferred, other alkynyl derivatives that can undergo click chemistry with the azide group, or functional groups that can participate in other bioorthogonal reactions, can also be used to modify nucleotides.
[0340] Based on the principles taught in this disclosure, those skilled in the art can prepare these different reactive nucleotide derivatives using chemical synthesis methods known in the art (e.g., terminal phosphate activation and modification techniques for polyphosphate nucleosides), or purchase them commercially from suppliers specializing in custom nucleic acid chemical synthesis services.
[0341] 2. Connection reaction
[0342] The core step in preparing the final product of this invention is the covalent linkage of a positively charged labeled molecule (MX) with a reactive nucleotide polyphosphate (YN). This invention covers any chemical reaction capable of stably linking these two components.
[0343] In a preferred embodiment, the ligation reaction is a bioorthogonal chemical reaction, which has the advantages of high selectivity, high yield and mild reaction conditions (e.g., carried out in an aqueous buffer), and can maximize the preservation of the bioactivity of the labeled molecule and nucleotide moieties.
[0344] Click chemistry is a preferred type of bioorthogonal reaction in this disclosure. Specifically, this disclosure covers at least the following two mature click chemistry linkage strategies:
[0345] Copper(I)-catalyzed azido-yne cycloaddition reaction (CuAAC): In this scheme, one reactant carries an azide group (-N3), and the other reactant carries a terminal alkyne group (-C≡CH). Under the catalysis of copper(I) ions, the two undergo an efficient cycloaddition reaction to form a stable 1,4-substituted triazole ring as a linker (X').
[0346] Stress-Promoted Azide-Alkyne Cycloaddition Reaction (SPAAC): This is a particularly preferred embodiment of the present disclosure. In this embodiment, one reactant carries an azide group (-N3), and the other reactant carries a cyclic strained alkyne (e.g., DBCO or a derivative thereof). Due to the presence of ring strain, the reaction proceeds spontaneously and rapidly without any metal catalyst, thereby forming a stable triazole ring as a linker (X').
[0347] Therefore, in some preferred embodiments of the present invention, the final product can be efficiently prepared by mixing a positively charged labeled molecule with an azide group (X = -N3) with a reactive nucleotide with a DBCO group (Y = DBCO) in a mild aqueous buffer.
[0348] It should be understood that, regardless of the specific linking chemistry used, the stable chemical structure formed by the reaction of the reactive linker part X and the paired reactive group (Y) constitutes the linker (X') in the final product.
[0349] 3. Final product: Positively charged nucleotides
[0350] After the above ligation reaction, the final product of the present invention, a positively charged nucleotide, as described above, can be obtained.
[0351] 4. Modular combination and customized sequencing reagent library
[0352] A key advantage of this invention lies in its high degree of modularity and combinability. This disclosure provides a "toolbox" comprising a large library of "positively charged labeled molecules" with different signal characteristics (e.g., the aforementioned PA-1 to PA-52 molecules), and a set of four "reactive nucleotide libraries" corresponding to A, G, C, and T, respectively (e.g., dATP-DBCO, dGTP-DBCO, dCTP-DBCO, dTTP-DBCO).
[0353] It should be understood that the scope of this invention is not limited to pairing a specific marker molecule with a specific nucleotide. Rather, those skilled in the art, after understanding the concept of this invention, can freely select any four marker molecules from the marker molecule library according to the needs of actual applications (e.g., the type of nanopore used, detector sensitivity, target signal-to-noise ratio, etc.) and link them to four reactive nucleotides respectively, thereby constructing a customized and complete sequencing reagent combination.
[0354] The principle behind this combination is that the four positively charged nucleotides (e.g., M1-X'-dATP, M2-X'-dGTP, M3-X'-dCTP, M4-X'-dTTP) that are ultimately formed must produce distinctive blocking current signals that are mutually distinguishable when passing through a nanopore. This means that a signal can be definitively attributed to one of A, G, C, or T with a sufficiently high confidence using signal analysis algorithms.
[0355] Therefore, this disclosure provides a powerful methodology that enables researchers and developers to create countless combinations of sequencing reagents with different performance characteristics through rational design and screening, thereby greatly promoting the development and application of nucleic acid sequencing and related molecular diagnostic technologies.
[0356] III. Nucleic Acid Sequencing Systems and Devices
[0357] Another aspect of this disclosure relates to a nucleic acid sequencing system and apparatus that integrates the core biological and chemical components of this invention to achieve high-precision, high-throughput single-molecule sequencing. This system physically utilizes a single nanopore-polymerase complex as a sensing unit to monitor electrical signals generated during DNA synthesis in real time.
[0358] The nucleic acid sequencing system disclosed herein includes:
[0359] (a) A detection unit comprising the NetB nanoporous protein of the present disclosure comprising NetB protein or a mutant thereof;
[0360] (b) a set of positively charged nucleotides. The labeling portion of each nucleotide produces a signal that distinguishes them from one another.
[0361] In some embodiments, the positively charged nucleotide is a nucleotide according to this disclosure.
[0362] The system disclosed herein fundamentally solves the problem of interference to the detection signal caused by the DNA template in the prior art by integrating the NetB nanopore with a preference for capturing positive charges and positively charged nucleotide markers, thereby enabling sequencing with a high signal-to-noise ratio.
[0363] In some implementations, the system or its associated sequencing device may further include one or more of the following components, as shown in Figure 14:
[0364] Core Detection Unit: This unit is the core of the system and includes:
[0365] Electrically insulating membrane and microporous structure: An electrically insulating membrane, such as an artificial phospholipid bilayer, has micron- or nanometer-sized pores formed or embedded on it. This membrane separates the device into two independent electrolyte solution pools.
[0366] The core of nanopore detection: A single NetB nanoporous protein, as described in this disclosure, is embedded and spans the micropores on the insulating membrane, making it the sole ion pathway connecting the two solution pools. In some preferred embodiments, the NetB nanoporous protein is coupled to DNA polymerase to form a nanoporous complex. In other preferred embodiments, the nanoporous complex further includes a pre-bound nucleic acid template to be tested, thereby forming a ternary complex that can be directly sequenced.
[0367] The fluid control subsystem includes microfluidic channels, pumps, and valves for the precise delivery and replacement of buffer solutions and sequencing reagents. This subsystem ensures the stability of the chemical environment required for the sequencing reaction and performs a washing step to remove unbound complexes.
[0368] Electronic detection and control subsystem:
[0369] Electrodes: A pair of electrodes are placed in the cis cell and the trans cell, respectively, to apply transmembrane potential across the membrane.
[0370] Voltage control source: used to apply and maintain a stable DC or AC voltage (e.g., 150 mV in some embodiments), which is the power source for driving the ion flow and capturing positively charged markers.
[0371] High-sensitivity current amplifier: Connected to electrodes, it can detect and amplify weak ion currents at the picoampere (pA) level that pass through a single nanopore.
[0372] Data acquisition and processing subsystem:
[0373] Data acquisition card: Converts analog current signals from amplifiers into digital signals at a high sampling frequency.
[0374] Data processing unit: Typically a computer running specialized software to receive digital signals in real time, perform baseline correction, identify blocking events, extract event features (such as depth and duration), and finally interpret the electrical signal sequence into a DNA base sequence through a base recognition algorithm.
[0375] In operation, the system utilizes the chemical conditions described in this disclosure. For example, the cis and trans cells are filled with a conductive electrolyte solution, such as Buffer-Seq (50 mM Tris-HCl pH 8.0, 300 mM KAc). Upon sequencing initiation, four positively charged nucleotides (e.g., each at a final concentration of 5 μM) as described in this disclosure, along with cofactors required for the polymerase (e.g., 1 mM MnCl2), are added to the cis cell. When a voltage is applied, the system is able to begin monitoring and recording sequencing events.
[0376] Compared with existing technologies, the system of this invention fundamentally solves the problem of DNA template interference with signals by integrating NetB nanopores that preferentially capture positive charges with positively charged nucleotide markers. This results in a system with extremely high signal-to-noise ratio and stability, making it an ideal physical platform for achieving high accuracy and ultra-long read sequencing.
[0377] IV. Nucleic Acid Sequencing Methods
[0378] Another aspect of this disclosure relates to a nucleic acid sequencing method that solves a fundamental problem in existing technologies through a unique positive charge detection system. The nucleic acid sequencing method of this disclosure includes:
[0379] (a) Using nanopores containing NetB protein or its mutants as single-molecule detection components; and
[0380] (b) Use a set of positively charged nucleotides as substrates for sequencing reactions.
[0381] The innovation of the nucleic acid sequencing method disclosed herein lies in the fact that it utilizes the specific capture capability of the NetB nanopore for positively charged markers, while simultaneously using an electric field to repel negatively charged DNA templates, thereby achieving high signal-to-noise ratio sequencing.
[0382] In some embodiments, the positively charged nucleotide is a nucleotide according to this disclosure.
[0383] In some embodiments, the nucleic acid sequencing method of this disclosure includes: (a) providing a detection unit comprising the NetB nanoporous protein or protein complex described in this disclosure; (b) introducing a nucleic acid template to be tested and a set of positively charged nucleotides into the environment of the detection unit; (c) detecting, under the action of an electric field, a characteristic electrical signal corresponding to the positively charged labeling generated by the incorporation of nucleotides directed by the nucleic acid template; and (d) determining the sequence of the nucleic acid template based on the characteristic electrical signal.
[0384] In some embodiments, the nucleic acid sequencing method of this disclosure includes: (a) providing the protein complex described in this disclosure; (b) providing a nucleic acid template to be tested and at least one positively charged substrate nucleotide in the presence of the protein complex; and (c) detecting a signal generated by the interaction of the positively charged substrate nucleotide with the protein complex to determine the sequence of the nucleic acid template to be tested.
[0385] In some implementations, the nucleic acid sequencing method disclosed herein is a sequencing-by-synthesis method, which typically includes the following steps:
[0386] 1. A detection unit is provided: wherein the detection unit comprises a nanoporous complex embedded in an electrically insulating membrane. Specifically, the NetB nanoporous complex of this disclosure, optionally coupled with DNA polymerase, is embedded in an electrically insulating membrane.
[0387] 2. Setting up the reaction system: Introduce the nucleic acid template to be tested and the charged nucleotide substrates. Specifically, in the reaction environment containing the detection unit, introduce the nucleic acid template to be tested and a set of positively charged nucleotide substrates corresponding to different bases.
[0388] 3. Sequencing Reaction and Signal Detection: Under the influence of an electric field, the charged label is selectively bound and captured, generating an electrical signal. Specifically, a transmembrane electric field configured to attract positive charges is applied. When the DNA polymerase selects a complementary nucleotide substrate according to the template, the NetB nanopore efficiently captures the positively charged label of that nucleotide under the drive of the electric field, generating a characteristic ion current blocking signal. Crucially, this electric field configuration simultaneously actively repels negatively charged DNA templates, thereby fundamentally eliminating signal interference caused by the template strand in existing technologies. After signal generation, the polymerase completes the incorporation of the nucleotide and releases the labeled portion, restoring the current to baseline levels and preparing for the detection of the next event.
[0389] 4. Determine the nucleic acid sequence: Based on the electrical signal, process and determine the nucleic acid sequence. Specifically, by recording and processing a series of high signal-to-noise ratio current blocking events generated in step 3 in real time, and comparing them with the characteristic signals of different markers, the sequence of the nucleic acid to be tested can be accurately reconstructed.
[0390] In some alternative implementations, the nucleic acid template to be tested can first be combined with a nanopore complex to form a ternary complex, and then embedded in an electrically insulating membrane. Specifically, the nucleic acid sequencing method of this disclosure includes the following steps:
[0391] 1. A detection unit is provided: wherein the detection unit comprises a ternary complex of a nanopore embedded in an electrically insulating membrane, a DNA polymerase, and a nucleic acid template to be tested. Specifically, the NetB nanopore complex of this disclosure (i.e., a nanopore coupled with a DNA polymerase) is pre-mixed and incubated with the nucleic acid template to be tested in a suitable buffer solution. Through this step, a stable ternary complex comprising a nanopore, a DNA polymerase, and a nucleic acid template to be tested is formed.
[0392] In an exemplary specific instance, to connect the nanopore and the polymerase, an I27C mutation was additionally introduced into the previously described NetB mutant to modify GMBS-Biotin, and a Spy-tag tag was introduced at the C-terminus (hereinafter referred to as the coupling monomer). The coupling monomer and the corresponding mutant were mixed and assembled, and after separation and purification, a 1:6 nanopore protein containing a single coupling monomer and six mutants was obtained (hereinafter referred to as the coupled nanopore). Monomeric Rhizavidin and SpyCatcher protein were simultaneously fused to the polymerase protein, which can be linked to Biotin and Spy-tag, respectively (hereinafter referred to as the fused polymerase). The coupled nanopore, coupled polymerase, and test DNA were mixed, and after separation and purification, a ternary complex formed by the coupled nanopore, coupled polymerase, and test DNA was obtained (hereinafter referred to as the DPN complex).
[0393] 2. Setting up the reaction system: Introducing charged nucleotide substrates. Specifically, in the reaction environment containing the detection unit, a set of positively charged nucleotide substrates corresponding to different bases are introduced.
[0394] Steps 3 and 4 are as described above.
[0395] In some implementations, for various purposes, such as to further improve the performance of the sequencing method, one or more of the following steps may optionally be included:
[0396] DNA template modification: Before sequencing, the DNA template to be tested can be modified, for example, by attaching specific adapter sequences or hairpin structures to one or both ends of the template. Attaching hairpin structures allows the polymerase to continue synthesizing using the complementary strand as a template after the polymerase has completed the synthesis of the first strand, thereby achieving double-stranded sequencing of the same DNA molecule and improving the accuracy of the final sequence.
[0397] Polymerase rate regulation: This can be achieved by adjusting the conditions of the sequencing reaction, such as the presence of divalent cations in the buffer (e.g., Mn). 2+ or Mg 2+The type and concentration of DNA polymerase, or by modifying the DNA polymerase itself through protein engineering, can be used to regulate the rate of chain elongation in order to match the data acquisition frequency of the detection system and ensure that each base incorporation event can be accurately captured.
[0398] Complex formation: In some embodiments, a pre-formed nanopore-polymerase-DNA ternary complex is directly loaded and embedded into an insulating membrane. In other embodiments, the nanopore-polymerase complex is first embedded into the insulating membrane, and then the DNA template to be tested is introduced, allowing it to bind to the polymerase on the membrane surface.
[0399] Signal model calibration: Optionally, a calibration DNA template with a known sequence can be introduced before or during sequencing to obtain a standard blocking current signal model for the four positively charged nucleotides in this system. This model can be used for more accurate base identification in subsequent sequencing data from unknown samples.
[0400] In summary, the method disclosed herein, through its unique positive charge detection system and combined with one or more of the above-mentioned preferred technical features, not only solves the fundamental problems of the prior art, but also achieves a significant improvement in sequencing accuracy and read length as shown in Figure 13, providing a more reliable and efficient sequencing solution for various scientific research and clinical applications.
[0401] V. Kits for nucleic acid sequencing
[0402] Furthermore, this disclosure provides a kit for performing nucleic acid sequencing methods. This kit pre-packages the key biological and chemical components required to perform the nucleic acid sequencing methods of this disclosure, providing convenience for users.
[0403] In some implementations, the kit includes:
[0404] (a) NetB Nanopores: Nanopores assembled from NetB protein or its mutants according to this disclosure. In some preferred embodiments, the NetB nanopores are coupled to a DNA polymerase to form a NetB nanopore complex. In other preferred embodiments, the kit may provide a pre-formed ternary complex comprising the nanopore complex and a pre-bound DNA template of a known sequence for system calibration or as a positive control. The component may be provided in a stable form, such as a lyophilized powder or a buffer solution containing a stabilizer.
[0405] (b) A group of positively charged nucleotides: This group comprises at least four independent, positively charged nucleotide substrates, each corresponding to a different nucleic acid base (e.g., A, G, C, T). In some preferred embodiments, the positively charged nucleotides are nucleotides according to this disclosure. The labeled portion of each nucleotide in this group is specifically designed to generate a distinct characteristic electrical signal upon passing through a nanopore.
[0406] (c) Sequencing buffer: One or more optimized reaction buffers containing salts, buffers, and other cofactors required to maintain the activity of nanopore functional complexes and promote polymerase chain elongation reactions. For example, the buffer may be Buffer-Seq containing Tris-HCl, KAc, and MnCl2.
[0407] (d) Optional instruction manual: which provides instructions to users on how to use the components in the kit to perform nucleic acid sequencing methods.
[0408] Optionally, the kit may also include reagents for preparing the user's own DNA template for testing.
[0409] The technical solutions of this disclosure will be further described in detail below through embodiments and in conjunction with the accompanying drawings. Unless otherwise stated, the methods and materials of the embodiments described below are all conventional products that can be purchased from the market. Those skilled in the art to which this disclosure pertains will understand that the methods and materials described below are merely exemplary and should not be considered as limiting the scope of this disclosure.
[0410] Example
[0411] The present invention will be further described in detail below with reference to the embodiments, but the scope of protection of the present invention is not limited to these embodiments.
[0412] The abbreviations and their corresponding English translations are as follows:
[0413] EA: Ethyl acetate;
[0414] dGMP: 2'-deoxyguanosine-5'-monophosphate;
[0415] DMF: N,N-dimethylformamide;
[0416] TEAA: Triethylamine Acetate Buffer;
[0417] TEAB: Triethylamine carbonate buffer;
[0418] Example 1. Preparation of NetB protein mutants
[0419] Mutant protein expression: pET26b was used as the expression vector. When constructing the NetB protein mutant expression vector, TEV and His-Tag tags were added to the C-terminus. The vector was transformed into Escherichia coli BL21 strain. Single clones were picked and inoculated into 10 mL of LB medium containing 50 μg / mL kanamycin and cultured at 37°C and 250 rpm for 16 hours. Then the culture was transferred to 400 mL of TB self-induction medium containing 50 μg / mL kanamycin and cultured at 25°C and 250 rpm for another 16 hours.
[0420] The culture was transferred to a centrifuge flask and centrifuged at 4000 rpm, 4°C for 15 min to pellet the cells. The cell pellet was then resuspended in 40 mL of Binding Buffer 1 (50 mM Tris-HCl pH 8.0, 200 mM NaCl, 10% glycerol). The resuspended cell pellet was then sonicated in an ice-water bath, followed by centrifugation at 15000 rpm, 4°C for 15 min to pellet residual cell debris. The supernatant was added to a 1 mL pre-equilibrated Ni-agarose gel gravity column. After addition, the column was washed with 10 mL of Washing Buffer 1 (50 mM Tris-HCl pH 8.0, 200 mM NaCl, 10% glycerol, 30 mM imidazole). Finally, the target protein was eluted with 3–10 mL of Elution Buffer-1 (50 mM Tris-HCl pH 8.0, 200 mM NaCl, 10% glycerol, 300 mM imidazole).
[0421] After protein elution, the OD280 concentration was determined using a spectrophotometer. Then, TEV protease was added at a concentration of 100:1, and the mixture was incubated at room temperature for 16 hours to remove the C-terminal His-Tag tag. The digested protein was dialyzed into 1 L of Binding Buffer-1, dialyzed three times at room temperature for 1 hour each time. Finally, the dialyzed sample was added again to a 1 mL pre-equilibrated Ni-agarose gel gravity column, and the flow-through sample was collected to obtain the NetB protein mutant. The sample was then examined by 12% SDS-PAGE gel electrophoresis for further testing.
[0422] The NetB protein mutant SEQ ID NO:2-12 was prepared according to the above preparation method.
[0423] The preparation method described in Example 1 is for preparing NetB protein and mutants, including the general method for preparing NetB protein mutants as disclosed herein.
[0424] Example 2. Preparation of NetB protein mutant conjugate monomers
[0425] The expression of the conjugate monomer is the same as that of the NetB protein mutant.
[0426] After expression, the culture was transferred to a centrifuge flask and centrifuged at 4000 rpm, 4°C for 15 min to pellet the cells. The cell pellet was then resuspended in 40 mL of Binding Buffer 2 (50 mM PB (pH 8.0), 200 mM NaCl, 10% glycerol). The resuspended cell pellet was then sonicated in an ice-water bath, followed by centrifugation at 15000 rpm, 4°C for 15 min to pellet residual cell debris. The supernatant was added to a 1 mL pre-equilibrated Ni-agarose gel gravity column. After addition, the column was washed with 10 mL of Washing Buffer 2 (50 mM PB (pH 8.0), 200 mM NaCl, 10% glycerol, 30 mM imidazole). Finally, the target protein was eluted with 3–10 mL of Elution Buffer-2 (50 mM PB (pH 8.0), 200 mM NaCl, 10% glycerol, 300 mM imidazole).
[0427] After protein elution, the sample was dialyzed into 1 L of Binding buffer 2, and dialyzed three times at room temperature for 1 hour each time. Then, GMBS-Biotin was added to the dialyzed sample to a final concentration of 400 μM, and the sample was incubated at room temperature for 3 hours. Finally, DTT was added to a final concentration of 1 mM to terminate the reaction. The sample was then examined by 12% SDS-PAGE gel electrophoresis and used for further testing.
[0428] Example 3. Preparation of NetB-coupled nanoporous proteins
[0429] To determine the protein concentration for OD280 assay, the mutant protein was mixed with the conjugated monomer at a 5:1 ratio, and DPHPC phospholipids were added to a final concentration of 0.75 mg / mL. The mixture was incubated at 37°C for 16 hours, and the heptamer formation was detected by 12% SDS-PAGE. Then, β-OG (n-octyl-β-D-glucopyranoside) was added to a final concentration and mixed thoroughly to completely dissolve the phospholipids. Finally, Tween-20 was added to a final concentration of 0.2%, and the mixture was pipetted and centrifuged at 3500 rpm for 3 minutes to remove the precipitate. The supernatant was then dialyzed into 1 L of Buffer A (20 mM acetate - pH 5.0, 0.1% Tween-20) and dialyzed three times at room temperature for 1 hour each time. The dialyzed sample was loaded into a cation exchange column (RESOURSE S1ML) and eluted with a gradient of 10-30% Buffer B (20mM acetate-pH 5.0, 2M NaCl, 0.1% Tween-20). The fraction collected at the 40S conductivity value was the desired 1:6 coupled nanopore.
[0430] Example 4. Preparation of DPN complex
[0431] In 100 μL of Buffer C (50 mM Tris-HCl pH 8.0, 300 mM KAc, 0.1% Tween-20, 1 mM DTT), the DNA template to be tested, the coupling polymerase, and the coupling nanopore were added at a molar ratio of 2:2:1 (the total volume added should exceed 10 μL). After mixing, the mixture was incubated at 37 °C for 30 min. Then, it was centrifuged at 15000 rpm and 4 °C for 15 min to precipitate the coupling polymerase-coupling nanopore protein that was not bound to DNA. The supernatant contained the desired DPN complex.
[0432] Example 5. Preparation of positively charged labeled molecules
[0433] Preparation of positively charged molecule PA-1-1
[0434] The synthesis route of PA-1-1 is shown in Figure 4;
[0435] 1. Synthesize compound A
[0436] (1) Weigh 2-CTC Resin resin (8g, 10.48mmol, 1eq, SD = 1.31mmol / g) into a 500ml peptide synthesis reactor, add 50-200ml DCM to swell the resin for 1-10 minutes, filter, add 50-200ml DMF to wash the resin 1-6 times, and dry. Weigh Boc-DAP(Fmoc)-OH (5-36g, 1-8eq) into a 500ml beaker, add 50-200ml DMF to dissolve, then measure DIPEA (4-30ml, 1-8eq) and add it to the beaker and mix well. Add the mixture to the reactor containing the resin and place it on a shaker at room temperature for overnight (10-24h) with shaking.
[0437] (2) After the reaction is complete, filter the mixture and wash it 1-6 times with DMF (50-200ml), 1-6 times with DCM (50-200ml), and 1-6 times with MeOH (50-200ml). Dry the mixture to obtain compound A (10-20g, SD = 0.5-1.0mmol / g).
[0438] 2. Synthesize compound D
[0439] (1) Add 50-200 ml of 20% piperidine / DMF solution to the resin obtained above, place it on a shaker and shake at room temperature for 5-20 min, and then dry it. Add 50-200 ml of 20% piperidine / DMF solution again, place it on a shaker and shake at room temperature for 5-20 min, and after the reaction is complete, wash it 1-6 times with DMF (50-200 ml), 1-6 times with DCM (50-200 ml), and 1-6 times with DMF (50-200 ml) to obtain compound B.
[0440] (2) Weigh Boc-DAP(Fmoc)-OH (5-36g, 1-8eq) and HOBt (4-20g, 1-8eq) into a 500ml beaker, add 50-200ml of DMF to dissolve, then measure DIC (4-25ml, 1-8eq) and add it to the beaker and mix well. Add the mixture to the reactor in the previous step, place it on a constant temperature shaker at 40-60℃ and shake for 20-60min. After the reaction is complete, wash with DMF (50-200ml) 1-6 times and dry to obtain compound C.
[0441] (3) The other 10 Boc-DAP(Fmoc)-OH molecules were sequentially coupled using the methods described in (1) and (2) to obtain compound D (30-40g, SD = 0.1-0.3mmol / g).
[0442] 3. Synthesis of compound E
[0443] (1) Weigh compound D (1.0 g, 1 eq, SD = 0.1-0.3 mmol / g) into a 10 ml synthesis reactor, add 3-10 ml of DCM to swell the resin for 1-10 minutes, filter, add 3-10 ml of DMF to wash the resin 1-6 times, and dry under vacuum. Add 20% piperidine / DMF solution (3-10 ml), place on a shaker and shake at room temperature for 5-20 minutes, and dry under vacuum. Add another 20% piperidine / DMF solution (3-10 ml), place on a shaker and shake at room temperature for 5-20 minutes. After the reaction is complete, wash with DMF (3-10 ml) 1-6 times, DCM (3-10 ml) 1-6 times, and DMF (3-10 ml) 1-6 times in sequence, and filter under vacuum.
[0444] (2) Weigh Fmoc-PEG6-OH (0.1-1.0g, 1-8eq) and HOBt (0.1-1.0g, 1-8eq) into a 50ml centrifuge tube, add 3-20ml of DMF to dissolve them, then add DIC (0.2-1.0ml, 1-8eq) to the centrifuge tube and mix well. Add the mixture to a reactor containing resin, place it on a constant temperature shaker at 40-60℃ and shake for 20-60min, filter, wash with DMF (3-10ml) 1-6 times, and dry. Add a combination solution of DMF, acetic anhydride and DIPEA in a volume ratio of 1-20:0.1-2:0.2-4 and block for 5-30min, filter, wash with DMF (3-10ml) 1-6 times, and dry.
[0445] (3) Using the same method, six more Fmoc-PEG6-OH molecules were coupled to obtain compound E.
[0446] 4. Synthesis of compound F
[0447] Compound F was obtained by sequentially coupling Fmoc-Tyr(tBu)-OH (0.1-0.8 g, 1-8 eq), Fmoc-PEG1-OH (0.1-0.6 g, 1-8 eq), and N3-C3-COOH (0.05-0.50 g, 1-8 eq) using the method described in step 3.
[0448] 5. Synthesis of compound G
[0449] Cutting process:
[0450] (1) Add a cleavage buffer with a DCM:TFA volume ratio of 1-4:4-16 (V / V) to compound F obtained in the previous step, react at room temperature for 10-30 min, filter, and collect the filtrate. Repeat this operation 1-7 times and collect the combined filtrates. Wash the resin peptide 1-6 times with cleavage buffer (3-10 ml) and DCM (3-10 ml) 1-6 times, filter, collect the filtrate in a 50 ml concentration flask, and concentrate under reduced pressure at 40℃~60℃.
[0451] (2) Add a cutting solution with a volume ratio of TFA:TIS:water of 1-4:0.01-0.04:0.05-0.2 (V / V / V) to the concentration flask of the previous step, stir the reaction at room temperature for 30-180 min, concentrate under reduced pressure at 40℃~60℃, add pure water to dissolve and obtain crude J(PA-1-1), identify the product by LC-MS and proceed with preparation.
[0452] (3) The crude product was purified by HPLC (purification conditions: column size 30*250mm 10um, mobile phase A: volume fraction 0.1% TFA / H2O, B: volume fraction 0.1% TFA / MeOH, detection wavelength 272nm, elution gradient: B phase 5% 3min, B phase 5-30% 20min, B phase 30% constant flow 5min) and freeze-dried to obtain pure G(PA-1-1) (300-1000mg, yield 10-30%).
[0453] Preparation of positively charged molecule PA-2-1
[0454] The synthesis route of PA-2-1 is shown in Figure 5;
[0455] 1. Synthesize compound A
[0456] (1) Weigh 2-CTC Resin resin (8g, 10.48mmol, 1eq, SD = 1.31mmol / g) into a 500ml peptide synthesis reactor, add 50-200ml DCM to swell the resin for 1-10 minutes, filter, add 50-200ml DMF to wash the resin 1-6 times, and dry. Weigh Boc-DAP(Fmoc)-OH (5-36g, 1-8eq) into a 500ml beaker, add 50-200ml DMF to dissolve, then measure DIPEA (4-30ml, 1-8eq) and add it to the beaker and mix well. Add the mixture to the reactor containing the resin and place it on a shaker at room temperature for overnight (10-24h) with shaking.
[0457] (2) After the reaction is complete, filter the mixture and wash it 1-6 times with DMF (50-200ml), 1-6 times with DCM (50-200ml), and 1-6 times with MeOH (50-200ml). Dry the mixture to obtain compound A (10-20g, SD = 0.5-1.0mmol / g).
[0458] 2. Synthesize compound D
[0459] (1) Add 50-200 ml of 20% piperidine / DMF solution to the resin obtained above, place it on a shaker and shake at room temperature for 5-20 min, and then dry it. Add 50-200 ml of 20% piperidine / DMF solution again, place it on a shaker and shake at room temperature for 5-20 min, and after the reaction is complete, wash it 1-6 times with DMF (50-200 ml), 1-6 times with DCM (50-200 ml), and 1-6 times with DMF (50-200 ml) to obtain compound B.
[0460] (2) Weigh Boc-DAP(Fmoc)-OH (5-36g, 1-8eq) and HOBt (4-20g, 1-8eq) into a 500ml beaker, add 50-200ml of DMF to dissolve, then measure DIC (4-25ml, 1-8eq) and add it to the beaker and mix well. Add the mixture to the reactor in the previous step, place it on a constant temperature shaker at 40-60℃ and shake for 20-60min. After the reaction is complete, wash with DMF (50-200ml) 1-6 times and dry to obtain compound C.
[0461] (3) The other 10 Boc-DAP(Fmoc)-OH molecules were sequentially coupled using the methods described in (1) and (2) to obtain compound D (30-40g, SD = 0.1-0.3mmol / g).
[0462] 3. Synthesis of compound E
[0463] (1) Weigh compound D (1.0 g, 1 eq, SD = 0.1-0.3 mmol / g) into a 10 ml synthesis reactor, add 3-10 ml of DCM to swell the resin for 1-10 minutes, filter, add 3-10 ml of DMF to wash the resin 1-6 times, and dry under vacuum. Add 20% piperidine / DMF solution (3-10 ml), place on a shaker and shake at room temperature for 5-20 minutes, and dry under vacuum. Add another 20% piperidine / DMF solution (3-10 ml), place on a shaker and shake at room temperature for 5-20 minutes. After the reaction is complete, wash with DMF (3-10 ml) 1-6 times, DCM (3-10 ml) 1-6 times, and DMF (3-10 ml) 1-6 times in sequence, and filter under vacuum.
[0464] (2) Weigh Fmoc-PEG6-OH (0.1-1.0g, 1-8eq) and HOBt (0.1-1.0g, 1-8eq) into a 50ml centrifuge tube, add 3-20ml of DMF to dissolve them, then add DIC (0.2-1.0ml, 1-8eq) to the centrifuge tube and mix well. Add the mixture to a reactor containing resin, place it on a constant temperature shaker at 40-60℃ and shake for 20-60min, filter, wash with DMF (3-10ml) 1-6 times, and dry. Add a combination solution of DMF, acetic anhydride and DIPEA in a volume ratio of 1-20:0.1-2:0.2-4 and block for 5-30min, filter, wash with DMF (3-10ml) 1-6 times, and dry.
[0465] (3) Using the same method, 20 more Boc-Lys(Fmoc)-OH molecules were coupled to obtain compound E.
[0466] 4. Synthesis of compound F
[0467] Compound F was obtained by sequentially coupling Fmoc-Tyr(tBu)-OH (0.1-1.0 g, 1-8 eq) and N3-C3-COOH (20-200 mg, 1-8 eq) using the method described in step 3.
[0468] 5. Synthesis of compound G
[0469] Cutting process:
[0470] (1) Add a cleavage buffer with a DCM:TFA volume ratio of 1-4:4-16 (V / V) to compound F obtained in the previous step, react at room temperature for 10-30 min, filter, and collect the filtrate. Repeat this operation 1-7 times and collect the combined filtrates. Wash the resin peptide 1-6 times with cleavage buffer (3-10 ml) and DCM (3-10 ml) 1-6 times, filter, collect the filtrate in a 50 ml concentration flask, and concentrate under reduced pressure at 40℃~60℃.
[0471] (2) Add a cutting solution with a volume ratio of TFA:TIS:water of 1-4:0.01-0.04:0.05-0.2 (V / V / V) to the concentration flask of the previous step, stir the reaction at room temperature for 30-180 min, concentrate under reduced pressure at 40℃~60℃, add pure water to dissolve and obtain crude J(PA-2-1), identify the product by LC-MS and proceed with the preparation.
[0472] (3) The crude product was purified by HPLC (purification conditions: column size 30*250mm 10um, mobile phase A: volume fraction 0.1% TFA / H2O, B: volume fraction 0.1% TFA / MeOH, detection wavelength 272nm, elution gradient: B phase 5% 3min, B phase 5-30% 20min, B phase 30% constant flow 5min) and freeze-dried to obtain pure G(PA-2-1) (20-80mg, yield 3-10%).
[0473] Preparation of positively charged molecules PA-6
[0474] The synthesis route of PA-6 is shown in Figure 6;
[0475] 1. Synthesize compound A
[0476] (1) Weigh 2-CTC Resin resin (8g, 10.48mmol, 1eq, SD = 1.31mmol / g) into a 500ml peptide synthesis reactor, add 50-200ml DCM to swell the resin for 1-10 minutes, filter, add 50-200ml DMF to wash the resin 1-6 times, and dry. Weigh Fmoc-Lys(Boc)-OH (5-36g, 1-8eq) into a 500ml beaker, add 50-200ml DMF to dissolve, then measure DIPEA (4-30ml, 1-8eq) and add it to the beaker and mix well. Add the mixture to the reactor containing the resin, place it on a shaker and shake overnight at room temperature (10-24h).
[0477] (2) After the reaction is complete, filter the mixture and wash it 1-6 times with DMF (50-200ml), 1-6 times with DCM (50-200ml), and 1-6 times with MeOH (50-200ml). Dry the mixture to obtain compound A (10-20g, SD = 0.5-1.0mmol / g).
[0478] 2. Synthesize compound D
[0479] (1) Add 50-200 ml of 20% piperidine / DMF solution to the resin obtained above, place it on a shaker and shake at room temperature for 5-20 min, and then dry it. Add 50-200 ml of 20% piperidine / DMF solution again, place it on a shaker and shake at room temperature for 5-20 min, and after the reaction is complete, wash it 1-6 times with DMF (50-200 ml), 1-6 times with DCM (50-200 ml), and 1-6 times with DMF (50-200 ml) to obtain compound B.
[0480] (2) Weigh Fmoc-Lys(Boc)-OH (5-36g, 1-8eq) and HOBt (4-20g, 1-8eq) into a 500ml beaker, add 50-200ml of DMF to dissolve, then measure DIC (4-25ml, 1-8eq) and add it to the beaker and mix well. Add the mixture to the reactor in the previous step, place it on a constant temperature shaker at 40-60℃ and shake for 20-60min. After the reaction is complete, wash with DMF (50-200ml) 1-6 times and dry to obtain compound C.
[0481] (3) The other 10 Fmoc-Lys(Boc)-OH molecules were sequentially coupled using the methods described in (1) and (2) to obtain compound D (30-40g, SD = 0.1-0.3mmol / g).
[0482] 3. Synthesis of compound G
[0483] (1) Weigh compound D (1.0 g, 1 eq, SD = 0.1-0.3 mmol / g) into a 10 ml synthesis reactor, add 3-10 ml of DCM to swell the resin for 1-10 minutes, filter, add 3-10 ml of DMF to wash the resin 1-6 times, and dry under vacuum. Add 20% piperidine / DMF solution (3-10 ml), place on a shaker and shake at room temperature for 5-20 minutes, and dry under vacuum. Add another 20% piperidine / DMF solution (3-10 ml), place on a shaker and shake at room temperature for 5-20 minutes. After the reaction is complete, wash with DMF (3-10 ml) 1-6 times, DCM (3-10 ml) 1-6 times, and DMF (3-10 ml) 1-6 times in sequence, and filter under vacuum.
[0484] (2) Weigh Fmoc-PEG1-OH (0.05-0.5g, 1-8eq) and HOBt (0.1-1.0g, 1-8eq) into a 50ml centrifuge tube, add 3-20ml of DMF to dissolve them, then add DIC (0.2-1.0ml, 1-8eq) to the centrifuge tube and mix well. Add the mixture to a reactor containing resin, place it on a constant temperature shaker at 40-60℃ and shake for 20-60min, filter, wash with DMF (3-10ml) 1-6 times, and dry. Add a combination solution of DMF, acetic anhydride and DIPEA in a volume ratio of 1-20:0.1-2:0.2-4 to block for 5-30min, filter, wash with DMF (3-10ml) 1-6 times, and dry.
[0485] (3) Using the methods described in (1) and (2), couple two more Fmoc-Gln(Trt)-OH molecules (0.5-4.0, 1-8 eq) sequentially.
[0486] (4) Repeat the method and linking group described in (1)(2)(3) 11 times.
[0487] (5) The method described herein was then sequentially coupled with Fmoc-PEG6-OH (0.1-1.0 g, 1-8 eq), Fmoc-Tyr(tBu)-OH (0.1-1.0 g, 1-8 eq), and N3-C3-COOH (20-200 mg, 1-8 eq) to obtain compound G.
[0488] 4. Synthesis of compound I
[0489] Cutting process:
[0490] (1) Add a cleavage solution with a DCM:TFA volume ratio of 1-4:4-16 (V / V) to compound G obtained in the previous step, react at room temperature for 10-30 min, filter, and collect the filtrate. Repeat this operation 1-7 times and collect the combined filtrates. Wash the resin peptide 1-6 times with cleavage solution (3-10 ml) and DCM (3-10 ml) 1-6 times, filter, collect the filtrate in a 50 ml concentration flask, and concentrate under reduced pressure at 40℃~60℃.
[0491] (2) Add a cutting solution with a volume ratio of TFA:TIS:water of 1-4:0.01-0.04:0.05-0.2 (V / V / V) to the concentration flask of the previous step, stir the reaction at room temperature for 30-180 min, concentrate under reduced pressure at 40℃~60℃, add pure water to dissolve and obtain crude J(PA-6), identify the product by LC-MS and proceed with preparation.
[0492] (3) The crude product was purified by HPLC (purification conditions: column size 30*250mm 10um, mobile phase A: volume fraction 0.1% TFA / H2O, B: volume fraction 0.1% TFA / MeOH, detection wavelength 272nm, elution gradient: B phase 5% 3min, B phase 5-30% 20min, B phase 30% constant flow 5min) and freeze-dried to obtain pure I (PA-6) (200-700mg, yield 10-30%).
[0493] Preparation of positively charged molecules PA-8
[0494] The synthesis route of PA-8 is shown in Figure 7;
[0495] 1. Synthesize compound A
[0496] (1) Weigh 2-CTC Resin resin (8g, 10.48mmol, 1eq, SD = 1.31mmol / g) into a 500ml peptide synthesis reactor, add 50-200ml DCM to swell the resin for 1-10 minutes, filter, add 50-200ml DMF to wash the resin 1-6 times, and dry. Weigh Fmoc-Lys(Boc)-OH (5-36g, 1-8eq) into a 500ml beaker, add 50-200ml DMF to dissolve, then measure DIPEA (4-30ml, 1-8eq) and add it to the beaker and mix well. Add the mixture to the reactor containing the resin, place it on a shaker and shake overnight at room temperature (10-24h).
[0497] (2) After the reaction is complete, filter the mixture and wash it 1-6 times with DMF (50-200ml), 1-6 times with DCM (50-200ml), and 1-6 times with MeOH (50-200ml). Dry the mixture to obtain compound A (10-20g, SD = 0.5-1.0mmol / g).
[0498] 2. Synthesize compound D
[0499] (1) Add 50-200 ml of 20% piperidine / DMF solution to the resin obtained above, place it on a shaker and shake at room temperature for 5-20 min, and then dry it. Add 50-200 ml of 20% piperidine / DMF solution again, place it on a shaker and shake at room temperature for 5-20 min, and after the reaction is complete, wash it 1-6 times with DMF (50-200 ml), 1-6 times with DCM (50-200 ml), and 1-6 times with DMF (50-200 ml) to obtain compound B.
[0500] (2) Weigh Fmoc-Lys(Boc)-OH (5-36g, 1-8eq) and HOBt (4-20g, 1-8eq) into a 500ml beaker, add 50-200ml of DMF to dissolve, then measure DIC (4-25ml, 1-8eq) and add it to the beaker and mix well. Add the mixture to the reactor in the previous step, place it on a constant temperature shaker at 40-60℃ and shake for 20-60min. After the reaction is complete, wash with DMF (50-200ml) 1-6 times and dry to obtain compound C.
[0501] (3) The other 10 Fmoc-Lys(Boc)-OH molecules were sequentially coupled using the methods described in (1) and (2) to obtain compound D (30-40g, SD = 0.1-0.3mmol / g).
[0502] 3. Synthesis of compound G
[0503] (1) Weigh compound D (1.0 g, 1 eq, SD = 0.1-0.3 mmol / g) into a 10 ml synthesis reactor, add 3-10 ml of DCM to swell the resin for 1-10 minutes, filter, add 3-10 ml of DMF to wash the resin 1-6 times, and dry under vacuum. Add 20% piperidine / DMF solution (3-10 ml), place on a shaker and shake at room temperature for 5-20 minutes, and dry under vacuum. Add another 20% piperidine / DMF solution (3-10 ml), place on a shaker and shake at room temperature for 5-20 minutes. After the reaction is complete, wash with DMF (3-10 ml) 1-6 times, DCM (3-10 ml) 1-6 times, and DMF (3-10 ml) 1-6 times in sequence, and filter under vacuum.
[0504] (2) Weigh Fmoc-PEG1-OH (0.05-0.5g, 1-8eq) and HOBt (0.1-1.0g, 1-8eq) into a 50ml centrifuge tube, add 3-20ml of DMF to dissolve them, then add DIC (0.2-1.0ml, 1-8eq) to the centrifuge tube and mix well. Add the mixture to a reactor containing resin, place it on a constant temperature shaker at 40-60℃ and shake for 20-60min, filter, wash with DMF (3-10ml) 1-6 times, and dry. Add a combination solution of DMF, acetic anhydride and DIPEA in a volume ratio of 1-20:0.1-2:0.2-4 to block for 5-30min, filter, wash with DMF (3-10ml) 1-6 times, and dry.
[0505] (3) Use the method described in (1) and (2) to sequentially couple one Fmoc-Dap[PEG2(Boc)]-OH (0.8-7.0, 1-8 eq).
[0506] (4) Repeat the method and linking group described in (1)(2)(3) 11 times.
[0507] (5) The method described therein was followed by coupling Fmoc-Tyr(tBu)-OH (0.1-1.0 g, 1-8 eq) and N3-C3-COOH (20-200 mg, 1-8 eq) to obtain compound G.
[0508] 4. Synthesis of compound I
[0509] Cutting process:
[0510] (1) Add a cleavage solution with a DCM:TFA volume ratio of 1-4:4-16 (V / V) to compound G obtained in the previous step, react at room temperature for 10-30 min, filter, and collect the filtrate. Repeat this operation 1-7 times and collect the combined filtrates. Wash the resin peptide 1-6 times with cleavage solution (3-10 ml) and DCM (3-10 ml) 1-6 times, filter, collect the filtrate in a 50 ml concentration flask, and concentrate under reduced pressure at 40℃~60℃.
[0511] (2) Add a cutting solution with a volume ratio of TFA:TIS:water of 1-4:0.01-0.04:0.05-0.2 (V / V / V) to the concentration flask of the previous step, stir the reaction at room temperature for 30-180 min, concentrate under reduced pressure at 40℃~60℃, add pure water to dissolve and obtain crude J(PA-8), identify the product by LC-MS, and proceed with the preparation.
[0512] (3) The crude product was purified by HPLC (purification conditions: column size 30*250mm 10um, mobile phase A: volume fraction 0.1% TFA / H2O, B: volume fraction 0.1% TFA / MeOH, detection wavelength 272nm, elution gradient: B phase 5% 3min, B phase 5-30% 20min, B phase 30% constant flow 5min) and freeze-dried to obtain pure I (PA-8) (250-750mg, yield 10-30%).
[0513] Example 6. Preparation of positively charged labeled nucleotides
[0514] Preparation of dC5P-PA-1-1
[0515] The synthetic route of dC5P-PA-1-1 is shown in Figure 8:
[0516] Unless otherwise specified, all reagents used in the synthesis method of dC5P-DBCO disclosed herein are commercially available.
[0517] Synthesis of dC5P-DBCO:
[0518] (1) Preparation of dCMP·tri-n-propylamine salt
[0519] Dissolve 1-3g of dCMP disodium salt in 5-20mL of deionized water, mix with H-type cation exchange resin, shake on a shaker for 10-60min, filter off the resin, add 4-15mL of tri-n-propylamine to the filtrate, shake for half a minute, concentrate by rotary evaporation for 15min, and freeze dry in a freeze dryer to obtain a white powdery dCMP·tri-n-propylamine salt solid.
[0520] (2) Synthesis of DBCO-OH
[0521] Take 2.5-7.5 g of dibenzocyclooctyne-N-hydroxysuccinimino ester into a round-bottom flask, add 15-80 mL of dichloromethane to completely dissolve the raw material, and lower the temperature to -40 to -10 °C for 6-15 minutes. Slowly add 1-4 mL of 4-aminobutanol to the above solution and react for 1-5 hours. After complete conversion of dibenzocyclooctyne-N-hydroxysuccinimino ester by TLC, quench with 40-100 mL of purified water, wash with 1-5 M HCl (100-200 mL x 3) and water to remove 4-aminobutanol, wash 1-7 times, and then rotary evaporate the organic phase to obtain the intermediate DBCO-OH in 70-95% yield.
[0522] (3) Synthesis of DBCO-P
[0523] Add 2-5 g of DBCO-OH to a round-bottom flask, dissolve in 25-70 ml of acetonitrile, and cool at -40 to -10 °C for 10 minutes. Slowly add 1-4 ml of phosphorus oxychloride to the reaction system and react at room temperature for 25-50 minutes. After complete conversion of DBCO-OH as monitored by liquid chromatography-mass spectrometry, quench with ice water at low temperature. Adjust the pH to approximately 2 with triethylamine, and then separate the product using HPLC to obtain the pure product. Analyze the product using liquid chromatography and mass spectrometry. The obtained product is lyophilized, with a yield of 60-80%.
[0524] (4) Synthesis of dC5P-DBCO
[0525] ① Add 0.1-0.2 mg of trimetaphosphoric acid and 20-40 mg of 1-methylimidazole to a 50 mL round-bottom flask and dissolve them in 0.2-0.6 mL of DMF. Separately weigh 20-40 mg of trimethylbenzoyl chloride and dissolve it in 0.4-0.6 mL of DMF. Add these to the flask and stir simultaneously. React at room temperature for 20-40 min.
[0526] ② Add a solution of dCMP·tri-n-propylamine salt in DMF (100-200 mg dissolved in 0.5-0.8 ml DMF) to a 50 mL three-necked flask under nitrogen protection, and react at room temperature for 3 h.
[0527] ③ At -10 to 0℃, add 20-40mg of DABCO, 10-15mg of MgCl2, and a DMF solution of DBCO-P (100-200mg dissolved in 0.5-0.8ml of DMF) to the three-necked flask in ②. After the addition is complete, place it at room temperature and stir to react overnight.
[0528] ④ At -10 to 0℃, add 20-40 mL of 200 mM TEAA buffer to the three-necked flask in ③ to quench the effluent, stir for 5-10 min, and extract 1-5 times with 30-60 mL of EA. The aqueous phase is purified by C18 reversed-phase HPLC (purification conditions: column specifications are reversed-phase polymer column 30*250 mm, mobile phase A is 100 mM TEAA, mobile phase B is acetonitrile, detection wavelength is 260 nm, elution gradient is: B phase 5-40% 30 min) to obtain pure dC5P-DBCO with a yield of 25-45%.
[0529] (5) Synthesis of dC5P-PA-1-1
[0530] Dissolve 0.022 mmol (1.1 eq) of dC5P-DBCO in 11 mL of water, add 1.1 mL of 1 M PBS buffer and 0.02 mmol (1 eq) of PA-1-1, adjust the pH to 8.0, and after reacting for 1 hour, LCMS monitoring showed that PA-1-1 was completely converted. The reaction solution was then directly prepared by HPLC.
[0531] Preparation of dG5P-PA-2-1
[0532] The synthetic route of dG5P-PA-2-1 is shown in Figure 9:
[0533] Unless otherwise specified, all reagents used in the synthesis method of dG5P-DBCO disclosed herein are commercially available.
[0534] Synthesis of dG5P-DBCO:
[0535] (1) Preparation of dGMP tri-n-propylamine salt
[0536] Dissolve 1-3g of dGMP disodium salt in 5-20mL of deionized water, mix with H-type cation exchange resin, shake on a shaker for 10-60min, filter off the resin, add 4-15mL of tri-n-propylamine to the filtrate, shake for half a minute, concentrate by rotary evaporation for 15min, and freeze dry in a freeze dryer to obtain a white powdery dGMP tri-n-propylamine salt solid.
[0537] (2) Synthesis of DBCO-OH
[0538] Take 2.5-7.5 g of dibenzocyclooctyne-N-hydroxysuccinimino ester into a round-bottom flask, add 15-80 mL of dichloromethane to completely dissolve the raw material, and lower the temperature to -40 to -10 °C for 6-15 minutes. Slowly add 1-4 mL of 4-aminobutanol to the above solution and react for 1-5 hours. After complete conversion of dibenzocyclooctyne-N-hydroxysuccinimino ester by TLC, quench with 40-100 mL of purified water, wash with 1-5 M HCl (100-200 mL x 3) and water to remove 4-aminobutanol, wash 1-7 times, and then rotary evaporate the organic phase to obtain the intermediate DBCO-OH in 70-95% yield.
[0539] (3) Synthesis of DBCO-P
[0540] Add 2-5 g of DBCO-OH to a round-bottom flask, dissolve in 25-70 ml of acetonitrile, and cool at -40 to -10 °C for 10 minutes. Slowly add 1-4 ml of phosphorus oxychloride to the reaction system and react at room temperature for 25-50 minutes. After complete conversion of DBCO-OH as monitored by liquid chromatography-mass spectrometry, quench with ice water at low temperature. Adjust the pH to approximately 2 with triethylamine, and then separate the product using HPLC to obtain the pure product. Analyze the product using liquid chromatography and mass spectrometry. The obtained product is lyophilized, with a yield of 60-80%.
[0541] (4) Synthesis of dG5P-DBCO
[0542] ① Add 0.1-0.2 mg of trimetaphosphoric acid and 20-40 mg of 1-methylimidazole to a 50 mL round-bottom flask and dissolve them in 0.2-0.6 mL of DMF. Separately weigh 20-40 mg of trimethylbenzoyl chloride and dissolve it in 0.4-0.6 mL of DMF. Add these to the flask and stir simultaneously. React at room temperature for 20-40 min.
[0543] ② Add a DMF solution of dGMP tri-n-propylamine salt (100-200 mg dissolved in 0.5-0.8 ml DMF) to a 50 mL three-necked flask under nitrogen protection, and react at room temperature for 3 h.
[0544] ③ At -10 to 0℃, add 20-40mg of DABCO, 10-15mg of MgCl2, and a DMF solution of DBCO-P (100-200mg dissolved in 0.5-0.8ml of DMF) to the three-necked flask in ②. After the addition is complete, place it at room temperature and stir to react overnight.
[0545] ④ At -10 to 0℃, add 20-40 mL of 200 mM TEAA buffer to the three-necked flask in ③ to quench the effluent, stir for 5-10 min, and extract 1-5 times with 30-60 mL of EA. The aqueous phase is purified by C18 reversed-phase HPLC (purification conditions: column specifications are reversed-phase polymer column 30*250 mm, mobile phase A is 100 mM TEAA, mobile phase B is acetonitrile, detection wavelength is 260 nm, elution gradient is: B phase 5-40% 30 min) to obtain pure dG5P-DBCO with a yield of 25-45%.
[0546] (5) Synthesis of dG5P-PA-2-1
[0547] Dissolve 0.022 mmol (1.1 eq) of dG5P-DBCO in 11 mL of water, add 1.1 mL of 1 M PBS buffer and 0.02 mmol (1 eq) of PA-2-1, adjust the pH to 8.0, and after reacting for 1 hour, LCMS monitoring showed that PA-2-1 was completely converted. The reaction solution was then directly prepared by HPLC.
[0548] Preparation of dT5P-PA-6
[0549] The synthetic route of dT5P-PA-6 is shown in Figure 10:
[0550] Unless otherwise specified, all reagents used in the synthesis method of dT5P-DBCO provided by this invention are commercially available.
[0551] Example 1 of synthesizing dT5P-DBCO:
[0552] (1) Preparation of dTMP·tri-n-propylamine salt
[0553] Dissolve 1-3g of dTMP disodium salt in 5-20mL of deionized water, mix with H-type cation exchange resin, shake on a shaker for 10-60min, filter off the resin, add 4-15mL of tri-n-propylamine to the filtrate, shake for half a minute, concentrate by rotary evaporation for 15min, and freeze dry in a freeze dryer to obtain a white powdery dTMP·tri-n-propylamine salt solid.
[0554] (2) Synthesis of DBCO-OH
[0555] Take 2.5-7.5 g of dibenzocyclooctyne-N-hydroxysuccinimino ester into a round-bottom flask, add 15-80 mL of dichloromethane to completely dissolve the raw material, and lower the temperature to -40 to -10 °C for 6-15 minutes. Slowly add 1-4 mL of 4-aminobutanol to the above solution and react for 1-5 hours. After complete conversion of dibenzocyclooctyne-N-hydroxysuccinimino ester by TLC, quench with 40-100 mL of purified water, wash with 1-5 M HCl (100-200 mL x 3) and water to remove 4-aminobutanol, wash 1-7 times, and then rotary evaporate the organic phase to obtain the intermediate DBCO-OH in 70-95% yield.
[0556] (3) Synthesis of DBCO-P
[0557] Add 2-5 g of DBCO-OH to a round-bottom flask, dissolve in 25-70 ml of acetonitrile, and cool at -40 to -10 °C for 10 minutes. Slowly add 1-4 ml of phosphorus oxychloride to the reaction system and react at room temperature for 25-50 minutes. After complete conversion of DBCO-OH as monitored by liquid chromatography-mass spectrometry, quench with ice water at low temperature. Adjust the pH to approximately 2 with triethylamine, and then separate the product using HPLC to obtain the pure product. Analyze the product using liquid chromatography and mass spectrometry. The obtained product is lyophilized, with a yield of 60-80%.
[0558] (4) Synthesis of dT5P-DBCO
[0559] ① Add 0.1-0.2 mg of trimetaphosphoric acid and 20-40 mg of 1-methylimidazole to a 50 mL round-bottom flask and dissolve them in 0.2-0.6 mL of DMF. Separately weigh 20-40 mg of trimethylbenzoyl chloride and dissolve it in 0.4-0.6 mL of DMF. Add these to the flask and stir simultaneously. React at room temperature for 20-40 min.
[0560] ② Add a solution of dTMP·tri-n-propylamine salt in DMF (100-200 mg dissolved in 0.5-0.8 ml DMF) to a 50 mL three-necked flask under nitrogen protection, and react at room temperature for 3 h.
[0561] ③ At -10 to 0℃, add 20-40mg of DABCO, 10-15mg of MgCl2, and a DMF solution of DBCO-P (100-200mg dissolved in 0.5-0.8ml of DMF) to the three-necked flask in ②. After the addition is complete, place it at room temperature and stir to react overnight.
[0562] ④ At -10 to 0℃, add 20-40 mL of 200 mM TEAA buffer solution to the three-necked flask in ③ to quench the effluent, stir for 5-10 min, and extract with 30-60 mL of EA 1-5 times. The aqueous phase is purified by C18 reversed-phase HPLC (purification conditions: column specifications are reversed-phase polymer column 30*250 mm, mobile phase A is 100 mM TEAA, mobile phase B is acetonitrile, detection wavelength is 260 nm, elution gradient is: B phase 5-40% 30 min) to obtain pure dT5P-DBCO product with a yield of 25-45%.
[0563] (5) Synthesis of dT5P-PA-6
[0564] Dissolve 0.022 mmol (1.1 eq) of dT5P-DBCO in 11 mL of water, add 1.1 mL of 1 M PBS buffer and 0.02 mmol (1 eq) of PA-6-N3, adjust the pH to 8.0, and after reacting for 1 hour, LCMS monitoring showed that PA-6-N3 was completely converted. The reaction solution was then directly prepared by HPLC.
[0565] Preparation of dA5P-PA-8
[0566] The synthetic route for dA5P-PA-8 is shown in Figure 11:
[0567] Unless otherwise specified, all reagents used in the synthesis method of dA5P-DBCO disclosed herein are commercially available.
[0568] Example 1 of synthesizing dA5P-DBCO:
[0569] (1) Preparation of dAMP·tri-n-propylamine salt
[0570] Dissolve 1-3g of dAMP disodium salt in 5-20mL of deionized water, mix with H-type cation exchange resin, shake on a shaker for 10-60min, filter off the resin, add 4-15mL of tri-n-propylamine to the filtrate, shake for half a minute, concentrate by rotary evaporation for 15min, and freeze dry in a freeze dryer to obtain a white powdery dAMP·tri-n-propylamine salt solid.
[0571] (2) Synthesis of DBCO-OH
[0572] Take 2.5-7.5 g of dibenzocyclooctyne-N-hydroxysuccinimino ester into a round-bottom flask, add 15-80 mL of dichloromethane to completely dissolve the raw material, and lower the temperature to -40 to -10 °C for 6-15 minutes. Slowly add 1-4 mL of 4-aminobutanol to the above solution and react for 1-5 hours. After complete conversion of dibenzocyclooctyne-N-hydroxysuccinimino ester by TLC, quench with 40-100 mL of purified water, wash with 1-5 M HCl (100-200 mL x 3) and water to remove 4-aminobutanol, wash 1-7 times, and then rotary evaporate the organic phase to obtain the intermediate DBCO-OH in 70-95% yield.
[0573] (3) Synthesis of DBCO-P
[0574] Add 2-5 g of DBCO-OH to a round-bottom flask, dissolve in 25-70 ml of acetonitrile, and cool at -40 to -10 °C for 10 minutes. Slowly add 1-4 ml of phosphorus oxychloride to the reaction system and react at room temperature for 25-50 minutes. After complete conversion of DBCO-OH as monitored by liquid chromatography-mass spectrometry, quench with ice water at low temperature. Adjust the pH to approximately 2 with triethylamine, and then separate the product using HPLC to obtain the pure product. Analyze the product using liquid chromatography and mass spectrometry. The obtained product is lyophilized, with a yield of 60-80%.
[0575] (4) Synthesis of dA5P-DBCO
[0576] ① Add 0.1-0.2 mg of trimetaphosphoric acid and 20-40 mg of 1-methylimidazole to a 50 mL round-bottom flask and dissolve them in 0.2-0.6 mL of DMF. Separately weigh 20-40 mg of trimethylbenzoyl chloride and dissolve it in 0.4-0.6 mL of DMF. Add these to the flask and stir simultaneously. React at room temperature for 20-40 min.
[0577] ② Add dAMP·tri-n-propylamine salt DMF solution (100-200 mg dissolved in 0.5-0.8 ml DMF) to a 50 mL three-necked flask under nitrogen protection, and react at room temperature for 3 h.
[0578] ③ At -10 to 0℃, add 20-40mg of DABCO, 10-15mg of MgCl2, and a DMF solution of DBCO-P (100-200mg dissolved in 0.5-0.8ml of DMF) to the three-necked flask in ②. After the addition is complete, place it at room temperature and stir to react overnight.
[0579] ④ At -10 to 0℃, add 20-40 mL of 200 mM TEAA buffer to the three-necked flask in ③ to quench the effluent, stir for 5-10 min, and extract 1-5 times with 30-60 mL of EA. The aqueous phase is purified by C18 reversed-phase HPLC (purification conditions: column specifications are reversed-phase polymer column 30*250 mm, mobile phase A is 100 mM TEAA, mobile phase B is acetonitrile, detection wavelength is 260 nm, elution gradient is: B phase 5-40% 30 min) to obtain pure dA5P-DBCO with a yield of 25-45%.
[0580] (5) Synthesis of dA5P-PA-8
[0581] Dissolve 0.022 mmol (1.1 eq) of dA5P-DBCO in 11 mL of water, add 1.1 mL of 1 M PBS buffer and 0.02 mmol (1 eq) of PA-8-N3, adjust the pH to 8.0, and after reacting for 1 hour, LCMS monitoring showed that PA-8-N3 was completely converted. The reaction solution was then directly prepared by HPLC.
[0582] Example 7. Nucleic acid sequencing
[0583] As shown in Figure 14, the prepared DPN complex samples (including the NetB protein mutant SEQ ID NO:2-12 prepared according to Example 1 as the protein monomer constituting the nanopore) were diluted to an appropriate concentration using Buffer-Seq and added to the device. After the DPN complex was embedded into the artificial phospholipid membrane, excess DPN complex was washed away using Buffer-Seq. Finally, the prepared sequencing reagents were added to start the test and recording. The test used an AC excitation voltage of 150mV and 500Hz. The detector measured the current value passing through the nanopore every 200us. The results of comparing the current signals and generated sequences obtained from some tests with the DNA template sequences are shown in Figure 3. The nanopore protein composed of NetB wild-type protein and the NetB protein mutant SEQ ID NO:2-12 prepared in Example 1 was verified by nucleic acid sequencing experiments. Compared with existing negatively charged nanopores (such as α-hemolysin), it improved signal purity and channel lifetime (experiments showed >12 hours, see Table 2 for details).
[0584] Verification using the aforementioned nucleic acid sequencing methods showed that the NetB nanopores constructed from both wild-type NetB protein and NetB protein mutants exhibited good sequencing performance, achieving beneficial effects such as high sequencing accuracy, high average read length, and high average channel lifetime. These nanopores can be widely applied in nucleic acid sequencing work.
[0585] Table 2
[0586] The embodiments disclosed herein are not limited to those described above. Without departing from the spirit and scope of the invention, those skilled in the art can make various changes and improvements to the invention in form and detail, and all of these are considered to fall within the protection scope of the invention.
[0587] sequence list
Claims
1. A mutant NetB protein comprising an amino acid sequence having one or more mutations selected from the group consisting of: deletion of 10 to 20 amino acid residues from the N-terminus of the wild-type NetB protein, and mutation of any one or more of K20, K24, H53, K114, K115, N236, D285, Q284, and E289, each independently to D, N, Q, S, T, Y, H, or E, as compared to the wild-type NetB protein amino acid sequence set forth in SEQ ID NO:
1.
2. The mutant NetB protein of claim 1, wherein, The mutant NetB protein comprises an amino acid sequence having one or more mutations selected from the group consisting of: deletion of 10 amino acid residues from the N-terminus of the wild-type NetB protein, deletion of 11 amino acid residues from the N-terminus of the wild-type NetB protein, deletion of 12 amino acid residues from the N-terminus of the wild-type NetB protein, deletion of 13 amino acid residues from the N-terminus of the wild-type NetB protein, deletion of 14 amino acid residues from the N-terminus of the wild-type NetB protein, deletion of 15 amino acid residues from the N-terminus of the wild-type NetB protein, deletion of 16 amino acid residues from the N-terminus of the wild-type NetB protein, deletion of 17 amino acid residues from the N-terminus of the wild-type NetB protein, deletion of 18 amino acid residues from the N-terminus of the wild-type NetB protein, deletion of 19 amino acid residues from the N-terminus of the wild-type NetB protein, deletion of 20 amino acid residues from the N-terminus of the wild-type NetB protein, H53D, H53Y, H53N, H53Q, H53E, N236D, N236E, Q284D, Q284E, K20D, K20N, K20Q, K20S, K20T, K20Y, K20H, K20E, K24N, K24D, K24Q, K24S, K24T, K24Y, K24H, K24E, K114N, K114D, K114Q, K114S, K114T, K114Y, K114H, K114E, K115N, K115D, K115Q, K115S, K115T, K115Y, K115H, K115E, D285N, D285Q, D285S, D285T, D285Y, D285H, E289N, E289Q, E289S, E289T, E289Y, and E289H, as compared to the wild-type NetB protein amino acid sequence set forth in SEQ ID NO:
1.
3. The mutant of NetB protein according to claim 1 or 2, wherein, The mutant NetB protein comprises an amino acid sequence having one or more mutations selected from the group consisting of: deletion of 10 to 20 amino acid residues from the N-terminus of the wild-type NetB protein; and one or more mutations selected from the group consisting of: K20D, K20N, K24D, K24N, K24S, H53D, K114D, K114N, K115D, K115T, N236D, N236E, Q284D, Q284E, D285H, D285N, E289Q, and E289Y, as compared to the amino acid sequence set forth in SEQ ID NO:
1. 4. The mutant NetB protein according to any one of claims 1-3, comprising: the amino acid sequence set forth in SEQ ID NO: 2, SEQ ID NO: 3, SEQ ID NO: 4, SEQ ID NO: 5, SEQ ID NO: 6, SEQ ID NO: 7, SEQ ID NO: 8, SEQ ID NO: 9, SEQ ID NO: 10, SEQ ID NO: 11, or SEQ ID NO:
12.
5. An isolated nucleic acid molecule encoding the NetB protein mutant according to any one of claims 1 to 4.
6. An expression vector comprising the nucleic acid molecule according to claim 5.
7. A host cell comprising the expression vector according to claim 6.
8. A NetB nanopore protein which is a multimer composed of at least two protein monomers as subunits, wherein each of the at least two protein monomers is independently selected from: (a) a wild-type NetB protein comprising the amino acid sequence set forth in SEQ ID NO: 1; or (b) the NetB protein mutant according to any one of claims 1-3.
9. The NetB nanopore protein of claim 8, wherein, the multimer is a heptamer, and wherein at least one protein monomer comprises a linker for coupling to a DNA polymerase.
10. The NetB nanopore protein of claim 9, wherein, the linker is a chemical moiety capable of participating in a protein covalent or non-covalent ligation reaction selected from the group consisting of a SpyTag peptide, a SnoopTag peptide, a sortase recognition site, an amino acid residue that can undergo chemical cross-linking, biotin, a HaloTag, or a SNAP-tag.
11. A protein complex comprising: (a) the NetB nanopore protein according to any one of claims 8-10; and (b) a DNA polymerase coupled to the NetB nanopore protein via a linker.
12. A positive charge label molecule as represented by formula (I): M-X Formula (I) wherein, a label moiety M comprising one or more positive charges; and a reactive linker moiety X adapted for covalent ligation to a nucleotide polyphosphate moiety.
13. The positive charge tag molecule of claim 12, wherein, the label moiety M comprises: (a) a polypeptide head portion (Head) consisting of 3 to 15 positively charged amino acid residues; and (b) a signal discrimination body portion R’.
14. The positive charge tag molecule of claim 13, wherein, the polypeptide head comprises an amino acid residue selected from the group consisting of lysine, arginine, and a, b-diaminopropionic acid.
15. The positive charge label molecule according to claim 13 or 14, wherein, the signal discrimination body portion R’ is a chain polymer with a main chain backbone being a heterohydrocarbyl chain optionally comprising one or more heteroatoms selected from N, O, or S.
16. The positive charge tag molecule of any one of claims 12-15, wherein, the reactive linker moiety X comprises an azido group or an alkyne group for a click chemistry reaction.
17. A positively charged labeled nucleotide comprising the positive charge label molecule according to any one of claims 12-16 covalently linked via a linker to a terminal phosphate of a polyphosphate nucleotide moiety comprising at least 3 phosphate groups.
18. The positively charged labeled nucleotide according to claim 17, wherein the linker is a triazole ring formed by an azido-alkyne cycloaddition reaction.
19. A positively charged labeled nucleotide according to claim 17 or 18, wherein the polyphosphate nucleotide moiety is a nucleoside pentaphosphate or a nucleoside hexaphosphate.
20. A nucleic acid sequencing system comprising: (a) a detection unit comprising a NetB nanopore protein or protein complex according to any one of claims 8-11; and (b) a set of positively charged labeled nucleotides. The positively charged labeled nucleotides are according to claims 17-19.
21. The nucleic acid sequencing system of claim 20, wherein, 22. A method of nucleic acid sequencing comprising: (a) using a NetB nanopore protein or protein complex according to any one of claims 8-11 as a single molecule detection component; and (b) using a set of positively charged labeled nucleotides as substrates for a sequencing reaction. The positively charged labeled nucleotides are according to claims 17-19.
24. A kit for nucleic acid sequencing comprising:
23. The sequencing method of claim 22, wherein, (a) a NetB nanopore protein or protein complex according to any one of claims 8-11; and (b) a set of positively charged labeled nucleotides, and optionally, (c) a sequencing buffer, and / or (d) instructions for use. The positively charged labeled nucleotides are according to claims 17-19. 25. The kit for nucleic acid sequencing according to claim 24, wherein,
Citation Information
Patent Citations
Clostridial toxin NETB
CN101903399A
Method of preparation of nanopore and uses thereof
CN104379761A
Polypeptide tagged nucleotides and use thereof in nucleic acid sequencing by nanopore detection
CN108350017A
Tagged nucleotides useful for nanopore detection
CN109863250A
Polypeptide tagged nucleotides and use thereof in nucleic acid sequencing by nanopore detection
US20190002968A1