Methods of characterizing analytes and adapters used therein

By designing a new adapter with a double-stranded region and a convex ring region, the problem of sequencing efficiency reduction caused by motor protein shedding in the prior art is solved, and more efficient sequencing and longer sequencing time are achieved.

CN120230833APending Publication Date: 2025-07-01BEIJING QITAN TECH CO LTD
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
CN202311865452.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-29
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

In existing nanopore sequencing methods, sequencing linkers have common motor proteins such as helicase shedding during sequencing preparation and actual sequencing, resulting in reduced sequencing efficiency.

Method used

A new type of adapter was designed. By forming a double-stranded region with local close interaction and a convex ring region with a range restricted by the double-stranded region, it stabilizes the motor protein in the convex ring region, reducing the frequency of motor protein shedding.

Benefits of technology

By reducing the frequency of motor protein shedding, more stable complexes are formed, sequencing efficiency is improved, sequencing time is extended, and the output of sequencing data is increased.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120230833A_ABST
    Figure CN120230833A_ABST
Patent Text Reader

Abstract

A novel adaptor for characterizing a biopolymer is provided. Also provided are novel complexes formed by the adapters and displacement proteins, methods of characterizing biopolymers using the adapters and complexes, and kits prepared using the adapters and complexes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of next-generation sequencing, in particular to transmembrane nanopore sequencing technology, and to a method for characterizing biological polymers and an adapter used in the method. Background Art

[0002] Nanopore sequencing technology, as a new generation of sequencing, has the characteristics of single-molecule sequencing, long read length, fast sequencing speed, and real-time data monitoring. It can directly read / detect modification information of analytes such as nucleic acids (for example, methylation and acetylation of nucleic acids, myristoylation of polypeptides, etc.). Compared with other sequencing technologies, it has significant advantages and is increasingly used in research fields such as species genome assembly, epigenetics, transcriptomics and metagenomics.

[0003] In the currently common nanopore sequencing method, sequencing adapters (also known as adaptors) are generally used to connect the analytes to be sequenced, such as polynucleotide molecules, and to bind and arrest motor proteins such as helicases. The sequencing adapters of the prior art are open structures, generally comprising a Y-shaped composite structure comprising a functional region fragment capable of binding to a motor protein such as a helicase and a functional region fragment capable of temporarily blocking the movement of the enzyme (see CN114457145A). In the process of realizing the present invention, the inventors found that there are at least the following problems in the prior art: this type of adapter is often associated with the phenomenon of motor proteins such as helicases falling off during sequencing preparation and actual sequencing, resulting in reduced sequencing efficiency. Therefore, there is an urgent need in the art for adapters that can effectively reduce the easy falling off of motor proteins during preparation and actual sequencing. Summary of the Invention

[0004] The present application proposes a new type of adapter, and a method for characterizing biopolymers by nanopore technology using the adapter. The adapter disclosed herein forms a double-stranded region with local close interaction and a convex loop region whose range is restricted by the double-stranded region, which stably keeps the motor protein required for controlling the displacement of the analyte in the convex loop region, effectively reducing the frequency of motor protein shedding in the preparation stage before characterization and during the characterization process, and forming a more stable complex with the adapter, thereby being able to more efficiently characterize the analyte, increase sequencing time, and improve the output of sequencing data. The structure formed by the convex loop region and the double-stranded region in the adapter disclosed herein can also be controllably untied, and therefore also allows the accurate setting of the analyte's through-hole timing, effectively preventing missed readings.

[0005] Therefore, an object of the present invention is to provide a complex comprising an adaptor having at least one double-stranded region formed by non-covalent interactions and a bulge loop region connected to the double-stranded region, and a motor protein that binds to and is arrested in the bulge loop region.

[0006] In some embodiments, the complex disclosed herein also comprises at least one motor protein binding motif in the bulge loop region. In some embodiments, the double-stranded region is formed by at least two segments of reverse complementary nucleotide sequences, or is formed by at least two segments of amino acid sequences comprising a coiled coil region. In some embodiments, the bulge loop region and / or the double-stranded region further comprises an element for blocking the displacement of the motor protein relative to the adapter. In a preferred embodiment, the blocking element is located downstream of the bulge loop region and / or the double-stranded region in terms of the direction of movement of the motor protein. In an optional embodiment, the blocking element comprises any one selected from organic oligocations, iSpC3, iSp18, iSp9, nitroindole, inosine, acridine, 2-aminopurine, 2-6-diaminopurine, 5-bromo-deoxyuracil, inverted thymidine (inverted dT), inverted dideoxythymidine (ddT), dideoxycytidine (ddC), 5-methylcytidylic acid, 5-hydroxymethylcytidine, 2'-O-methyl RNA base, isodeoxycytidine (iso-dC), isodeoxyguanosine (iso-dG), photocleavable (PC) group, hexanediol, locked nucleic acid (LNA), peptide nucleic acid (PNA), methoxy (OMe), bicyclic nucleoside (BNA), glycerol nucleic acid (GNA), threose nucleic acid (TNA), cyclic pyrrolimidazole polyamide (cPIP) or double-stranded binding protein modified nucleotides, nucleotides with a benzo structure as the nucleobase, or any combination thereof.

[0007] In some embodiments, the double-stranded region and the bulge loop region in the complex disclosed herein form a stem-loop structure, a bulge structure, or an internal loop structure, or any combination thereof.

[0008] In some embodiments, the complexes disclosed herein further comprise a region in the adapter located outside the bulge loop region, preferably located on either side or both sides of the structure formed by the bulge loop region and the double-stranded region. In some embodiments, the region on either side or both sides may comprise an element selected from a leader sequence, a linker binding region, and a library connection region. In some embodiments, the region on either side or both sides preferably comprises a modification or element that can hinder the binding of the motor protein to a sequence outside the bulge loop region, and the modification is preferably selected from 2'-methoxylated ribose.

[0009] In some embodiments, the complex disclosed herein further comprises a connecting strand, the connecting strand comprising a region that binds to an external connector and a complementary region that is complementary to a portion of the region on either side or both sides of the adaptor, such that the connecting strand can bind to the adaptor by annealing. In a preferred embodiment, the connecting strand binds to a portion of the region in the adaptor disclosed herein in a complementary pairing manner.

[0010] In some embodiments, the adapter in the complex disclosed herein consists of a first nucleotide single strand. In other embodiments, the adapter in the complex disclosed herein comprises a first nucleotide single strand and a second nucleotide single strand, wherein the first nucleotide single strand comprises a stem-loop structure, and the second nucleotide single strand comprises a pairing region that is reverse complementary to a partial region of the first nucleotide single strand that is located outside the stem-loop structure and located on the downstream side along the direction of movement of the motor protein. In some embodiments, the adapter in the complex disclosed herein comprises a first nucleotide single strand and a second nucleotide single strand, wherein the first nucleotide single strand comprises a stem-loop structure, and the second nucleotide single strand comprises a pairing region that is reverse complementary to a partial region of the first nucleotide single strand that is located on both sides of the stem-loop structure, and the pairing region is optionally composed of continuous or discontinuous nucleotides in the second nucleotide single strand. In some embodiments, the linker in the complex disclosed herein comprises a first nucleotide single strand, a second nucleotide single strand, and a third nucleotide single strand, wherein the first nucleotide single strand comprises a stem-loop structure, the second nucleotide single strand comprises a pairing region that is reversely complementary to a partial region of the first nucleotide single strand that is located outside the stem-loop structure and located on the downstream side along the direction of movement of the motor protein, and the third nucleotide single strand comprises a pairing region that is reversely complementary to a partial region of the first nucleotide single strand that is located on the other side outside the stem-loop structure.

[0011] In some embodiments, the motor protein in the complex disclosed herein is a polynucleotide binding protein. In an optional embodiment, the polynucleotide binding protein is selected from at least one of a DNA polymerase, an RNA polymerase, a helicase, an endonuclease, an exonuclease, a DNA protease, a topoisomerase, a nucleic acid translocase, a nucleic acid nicking enzyme, or any fusion protein thereof.

[0012] In some embodiments, the complex disclosed herein further comprises an analyte, which is directly or indirectly linked to the adaptor.

[0013] Another object of the present invention is to provide a method for preparing a complex disclosed herein, comprising (a) providing an adapter disclosed herein; (b) allowing the motor protein to bind to and arrest in the convex loop region of the adapter to form a complex. In some embodiments, the method for preparing a complex disclosed herein comprises (a) providing an adapter disclosed herein; (b) allowing the motor protein to bind to and arrest in the convex loop region of the adapter to form a complex; (b') connecting the adapter to an analyte. In some embodiments, the method for preparing a complex disclosed herein performs step (b') before, simultaneously with, or after performing step (b).

[0014] Another object of the present invention is to provide a method for characterizing an analyte or controlling the movement of a motor protein. The method comprises the steps of providing an adaptor disclosed herein, causing the motor protein to bind to and stop in the convex loop region of the adaptor under certain conditions to form a complex, and changing the conditions so that the motor protein is displaced relative to the convex loop region and completely passes through the double-stranded region. In some embodiments of the present invention, the changed conditions can be selected from the direction of the electric field, the strength of the electric field and / or the concentration of nucleotide triphosphates, preferably the concentration of ATP. In some embodiments of the present invention, the method disclosed herein further comprises the step of causing the analyte to pass through a hole for characterization and to displace relative to the hole to obtain one or more measurements, wherein the measurements represent one or more characteristics of the analyte. In some embodiments of the present invention, the one or more characteristics are selected from (i) the length of the analyte, (ii) the sequence identity of the analyte, (iii) the sequence of the analyte, (iv) the secondary structure of the analyte and (v) the modification of the analyte. In some embodiments of the present invention, the one or more characteristics can be measured by electrical measurement and / or optical measurement.

[0015] Another object of the present invention is to provide a kit for characterizing an analyte or controlling the movement of a motor protein, which comprises the complex disclosed herein, or comprises an independently packaged adaptor and an independently packaged motor protein, wherein the motor protein is bound to and retained in the protruding loop region of the adaptor before characterization begins.

[0016] Another object of the present invention is to provide use of the complex disclosed herein in preparing a reagent or a kit for characterizing an analyte or controlling the movement of a motor protein.

[0017] Another object of the present invention is to provide a device for characterizing an analyte or controlling the movement of a motor protein, the device comprising a pore, an insulating layer supporting the pore and penetrated on both sides by the pore, and a complex disclosed herein distributed on either or both sides of the insulating layer. In optional embodiments, the device disclosed herein further comprises an electrical component for providing a potential difference between the two sides of the insulating layer supporting the pore. In some embodiments of the present invention, the pore is selected from a nanopore, a transmembrane pore, a biological pore, a solid-state pore, or a biological and solid-state hybrid pore. In a preferred embodiment of the present invention, the biological pore is derived from hemolysin, leukocidin, CsGG, Mycobacterium smegmatis porin A (MspA), porin B, porin C, porin D, outer membrane porin F, outer membrane porin G, outer membrane phospholipase A, Neisseria autotransporter lipoprotein, and WZA. In other preferred embodiments, the solid-state pore is derived from a graphene nanopore, a MoS2 nanopore, a BN nanopore, or a PA63 nanopore. In some embodiments, the analyte is selected from a polynucleotide, a polypeptide, a polysaccharide, or a lipid. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 The structures of complexes 1 and 2 formed by the SL-type adapter (1A) and the classic Y-type adapter (1B) in Example 1 and the motor protein are shown as examples.

[0019] Figure 2 The gel electrophoresis diagram of the complex 1 in Example 1 is shown. In the figure, the "nucleic acid" lane and the "nucleic acid and enzyme incubation" lane represent samples before and after incubation of the SL-type adapter 1 with the RM1 protein, respectively. The "nucleic acid protein complex" band corresponds to the SL-type adapter 1-RM1 protein complex, and the "nucleic acid" band represents free SL-type adapter 1.

[0020] Figure 3 Column chromatography elution peaks of the SL-type adapter-enzyme complex 1 in Example 1 and the Y-type adapter-enzyme complex 2 in Comparative Example 1 are shown. Figure 3 In A, E1 represents the SL adaptor-enzyme complex 1, and E2 represents the free SL adaptor. Figure 3 In B, E3 represents the Y-shaped adaptor-enzyme complex 2.

[0021] Figure 4 The TBE PAGE electrophoresis results of each elution peak fraction after the SL-type adapter-enzyme complex 1 in Example 1 and the Y-type adapter-enzyme complex 2 in Comparative Example 1 were separated and purified by column chromatography are shown. Figure 4 The meanings of the legends in Figures A and 4B are respectively Figure 3 The meanings in A and 3B are the same. In addition, the "nucleic acid" band is Figure 4A and 4B represent free SL-type adapters and Y-type adapters, respectively. The “nucleic acid-protein complex” band is Figure 4 The bands in A and 4B represent complex 1 and complex 2, respectively.

[0022] Figure 5 The diagram shows the changes in the dissociation of RM1 protein from the SL-type adapter-enzyme complex 1 and the Y-type adapter-enzyme complex 2 during quality control. Figure 5 In A (left) and 5B (right), the band indicated by "nucleic acid" is the free adaptor, and the band indicated by "nucleic acid-protein complex" is the complex formed by the adaptor and the translocase.

[0023] Figure 6 The signal profile of RNA sequencing using SL-type adapter-enzyme complex 1 is shown. Figure 6 B (bottom) is a local amplification of the current signal in 6A (top).

[0024] Figure 7 The structure of the complex 3 formed by the SL-type adaptor and the motor protein in Example 5 is exemplarily shown.

[0025] Figure 8 The detection results of SL-type adapter-enzyme complex 3 are shown. Figure 8 A shows the column chromatography elution results of complex 3, Figure 8 B shows the gel electrophoresis results of complex 3. In the figure, E2 represents SL-type adaptor-R protein complex 3.

[0026] Figure 9 The enzyme dissociation status of the SL-type adaptor-enzyme complex 3 immediately after preparation (0 hours) and after long-term storage (24 hours) is shown.

[0027] Figure 10 The signal profile of DNA sequencing using SL-type adapter-enzyme complex 3 is shown. Figure 10 (Bottom) is a local amplification of the current signal at a specific time period during the sequencing process of 10 (top).

[0028] Figure 11 Shown are the exemplary structure of the single-chain SL-type adapter-enzyme complex 4 in Example 8 and its detection results. Figure 11 C shows the column chromatography elution results of complex 4, Figure 11 D shows the gel electrophoresis results of complex 4. In the figure, "incubation" represents the incubation product of the single-stranded SL-type adapter and the enzyme before elution, E1 represents the single-stranded SL-type adapter-enzyme complex 4, and E2 represents the single-stranded SL-type adapter.

[0029] Figure 12The structure of the complex formed by the double-stranded SL-type adaptor and the motor protein in Example 9 is exemplarily shown. DETAILED DESCRIPTION

[0030] definition

[0031] Unless otherwise defined herein, the scientific and technical terms used in association with the present disclosure will have the meaning commonly understood by those of ordinary skill in the art. The meaning and scope of the terms should be clear, but in the case of any potential ambiguity, the definitions provided herein take precedence over any dictionary or external definition. Unless otherwise stated, the operating methods specifically adopted in this application (including: preparation technology, experimental procedures, detection means, etc.) adopt the conventional techniques of biochemical experiments, cell biology experiments, molecular biology experiments related fields conventional in the field of this technology. These technologies have been fully described in the existing literature, specifically, for example, referring to Sam brook et al. Molecular Cloning: a Laboratory Manual 4th edition, Cold Spring Harbor Laboratory Press, 2012; Ausubel et al., Current Protocols in Molecular Biology, Wiley Online Press, updated from time to time.

[0032] As used herein, the terms "comprising" or "including" mean that sequences, compositions, and methods include the recited components or steps, but do not exclude other components or steps. "Consisting essentially of," when used to define compositions and methods, should be construed to exclude any other components or other steps that are clearly important for the technical effect to be achieved. "Consisting of" should be construed to exclude other components and steps not mentioned.

[0033] It should be understood that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting. As used in this specification and the appended claims, the singular forms "a," "an," and "the" include plural referents unless the context clearly indicates otherwise. Thus, for example, reference to "a cell" includes a combination of two or more cells, or an entire culture of cells. Unless expressly specified or obvious from the context, as used herein, the term "or" is understood to be inclusive.

[0034] Unless expressly provided or obvious from the context, otherwise as used herein, the term "about" should be understood to be within the normal tolerance range in the field, for example, within 2 standard deviations of the mean. "About" can be understood to be within 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, 1%, 0.9%, 0.8%, 0.7%, 0.6%, 0.5%, 0.4%, 0.3%, 0.2%, 0.1%, 0.09%, 0.08%, 0.07%, 0.06%, 0.05%, 0.04%, 0.03%, 0.02% or 0.01% of the value. Unless otherwise obvious from the context, all numerical values ​​provided herein are modified by the term "about".

[0035] As used herein, the terms "A binds to B," "A (covalently / non-covalently) binds to B," or similar expressions refer to the molecular coupling of the indicated element A to the indicated element B in a controlled manner, or the physical proximity, affinity, adsorption, or attraction of another indicated element B, or to the placement of the controlled element A at an appropriate spatial position relative to element B, or the spatial proximity of the controlled element A to element B, indicating a positional / functional relationship between the two elements. This binding does not occur ab initio, but rather occurs under certain conditions or time (i.e., in a controlled manner). By way of example and not limitation, in the context of "a helicase binds to a nucleic acid chain," "binding" may mean that, under certain conditions (e.g., providing a suitable buffer to provide a suitable ionic environment), the helicase is loaded into the DNA. For example, if the helicase is a ring of two hexamer, the nucleic acid chain passes through the hexamer ring, and this loading is binding. By way of example and not limitation, in the context of "capable of binding to the main chain", "binding" may mean that under certain conditions (for example, providing a desired buffer to provide a suitable ionic environment), at least a portion of the connecting chain can undergo non-covalent binding (for example, base complementary pairing, etc.) with the main chain.

[0036] As used herein, the term "stall" or similar expressions may refer to the fact that the relative spatial position of a motor protein and a characterized biological molecule (e.g., a polynucleotide binding protein and a nucleic acid chain) does not change to a certain extent or range, and such stalling is relatively stable and reversible under certain conditions (e.g., such stalling can be broken, released, or overcome by changing ionic conditions or potential conditions).

[0037] As used herein, the term "motor protein" refers to any protein that can control the movement of polynucleotides (such as but not limited to dsDNA, RNA-DNA hybrid duplexes) waiting for the sequence to be characterized through the nanopore, and may also be referred to as a rate-controlling protein. The motor protein may include modifications that, for example, help the molecule to be characterized to bind, prevent its detachment, and / or contribute to its activity at high salt concentrations and / or room temperature. As a preferred example, the "motor protein" of the present invention can list any known or unknown polynucleotide binding protein, including but not limited to DNA / RNA polymerases, helicases, endonucleases, exonucleases, DNA proteases, topoisomerases, nucleic acid translocases, nucleic acid nicking enzymes, etc., but are not limited thereto. For examples of polynucleotide binding proteins, see International Patent Application Publication No. WO2021253410A. .

[0038] Those skilled in the art can also easily understand that the "motor protein" and "polynucleotide binding protein" of the present invention can be a combination of their variants, modifications, motifs, functional regions or domains, as long as they can achieve binding to the biological molecules to be characterized and can control their relative movement.

[0039] As used herein, the term "helicase" or similar expressions may refer to an enzyme that unravels hydrogen bonds between nucleic acid bases. Helicases generally include members of helicase superfamily 1 or superfamily 2. Non-limiting examples of members of helicase superfamily 1 include, for example, members belonging to the Pif1-like family, the Upf1-like family, and the UvrD / Rep family. More preferably, members belonging to, for example, the RecD subfamily, the Upf1 subfamily, the PcrA subfamily, the Rep subfamily, and the UvrD subfamily may be cited. Non-limiting examples of members of the helicase superfamily 2 include, for example, members belonging to the Ski-like family, Rad3 / XPD family, NS3 / NPH-II family, DEAD family, DEAH / RHA family, RecG-like family, REcQ-like family, T1R-like family, Swi / Snf-like family, and Rig-I-like family, and more preferably, members belonging to the Hel308 subfamily, Mtr4 subfamily, XPD subfamily, NS3 subfamily, Mss116 subfamily, Prp43 subfamily, RecG subfamily, RecQ subfamily, T1R subfamily, RapA subfamily, and Hef subfamily. The helicases of the present invention include DNA helicases and RNA helicases. In some embodiments, the DNA helicase can be selected from members belonging to the RecD subfamily, the PcrA subfamily, the Rep subfamily, the UvrD subfamily, the Hel308 subfamily, the XPD subfamily, the RecG subfamily, the RecQ subfamily, the T1R subfamily, and the RapA subfamily. In some embodiments, the RNA helicase can be selected from members belonging to the Upf1 subfamily, the Mtr4 subfamily, the NS3 subfamily, the Mss116 subfamily, the Prp43 subfamily, and the Hef subfamily. In a preferred embodiment, the helicase of the present invention can also be a Hel308 helicase, a RecD helicase, a Tral helicase, a TrwC helicase, an XPD helicase, a Dda helicase, a pif1-like helicase, a T4 helicase, an UvrD helicase, a rep protein, a PriA (n' protein), TraY / I, TFII-F, or TFII-H. In a more preferred embodiment, the helicase of the present invention is selected from Pif1-like helicase, Dda helicase, Prp43 helicase or NS3 helicase. The helicase of the present invention can also be a functional variant of a helicase obtained by natural or artificial mutation or modification, or an artificial enzyme or enzyme analog having an enzymatic active center of a helicase, or any combination of motifs, functional regions or domains having helicase activity.

[0040] As used herein, the term "nucleic acid analog" or similar expressions refers to nucleic acid-derived substances with properties similar to those of RNA and DNA. While RNA and DNA are primarily composed of phosphate, pentose sugars, and bases, nucleic acid analogs replace at least one of these with other substances, such as peptide nucleic acid (PNA), morpholino (MNA), bridged nucleic acid (BNA), locked nucleic acid (LNA), glycol nucleic acid (GNA), and threose nucleic acid (TNA). The side chain regions can be specifically configured as ribonucleotides (RNA) and / or nucleic acid analogs to retain the original base specificity, allowing them to bind to the linker while reducing their ability to bind to the motor protein (in some embodiments, the polynucleotide binding protein). Of course, the side chain regions can also be modified deoxyribonucleotides (DNA).

[0041] As used herein, the term "transmembrane nanopore" or similar expressions may refer to a structure that allows hydrated ions driven by an applied potential to flow from one side of the membrane to the other side of the membrane. A transmembrane nanopore is a polypeptide or a collection of polypeptides that allows hydrated ions (e.g., analytes) to flow from one side of the membrane to the other side of the membrane, which can form a pore that allows hydrated ions driven by an applied potential to flow from one side of the membrane to the other side. Transmembrane nanopores preferably allow analytes (e.g., nucleotides) to flow from one side of a membrane (e.g., a lipid bilayer) to the other side. Transmembrane protein pores allow polynucleotides (e.g., DNA or RNA) to move through the pore. Transmembrane nanopores may be monomers or oligomers. The pore is preferably composed of several repeating subunits (e.g., 6, 7, or 8 subunits). Transmembrane nanopores typically comprise a barrel or channel through which ions can flow. Transmembrane nanopore subunits typically surround a central axis and provide chains for a transmembrane β barrel or channel or a transmembrane α-helical bundle or channel. The barrel or channel of a transmembrane nanopore typically contains amino acids that promote interaction with the analyte (e.g., nucleotides, polynucleotides, nucleic acids, polypeptides). These amino acids are preferably located near the constriction of the barrel or channel. Transmembrane nanopores typically contain one or more positively charged amino acids (e.g., arginine, lysine, or histidine) or aromatic amino acids (e.g., tyrosine or tryptophan). These amino acids typically promote interaction between the pore and the nucleotides, polynucleotides, or nucleic acids.

[0042] As used herein, the term "blocking element" or similar expressions may refer to an oligonucleotide region lacking a nucleobase and a sugar. For specific examples, synthesis methods, and composition forms, please refer to the prior art WO2014135838A1.

[0043] As used herein, the term "characterize" or similar expressions can refer to measuring one, two, three, four, or five or more characteristics of a target biomolecule. For example, the length, identity, sequence, secondary structure, and whether structural units of the biomolecule are modified can be measured. The length of the biomolecule can be determined, for example, by determining the number of interactions between the target molecule and the pore and the duration between interactions between the target biomolecule and the pore. Identity can be determined in conjunction with or without determining the sequence of the target biomolecule. The former is direct; the biomolecule is sequenced and identified thereby. The latter can be accomplished in several ways. For example, the presence of a specific motif in a biomolecule can be determined (without determining the rest of the sequence of the molecule). Alternatively, a specific electrical and / or optical signal determined in the method can identify the target biomolecule as originating from a specific source. The sequence can be determined as previously described. Suitable sequencing methods, particularly those using electrical measurements, are described in Stoddart D et al., Proc Natl Acad Sci, 12; 106(19): 7702-7, Lieberman KR et al, J Am Chem Soc. 2010; 132(50): 17961-72, and International Application WO 2000 / 28312. Secondary structure can be measured in a variety of ways. For example, if the method includes electrical measurements, secondary structure can be measured using changes in residence time or current across the pore. This allows regions of single-stranded and double-stranded polynucleotides to be identified. The presence or absence of any modification can be determined. The method preferably includes determining whether the target polynucleotide has been modified by methylation, oxidation, damage, modification with one or more proteins or one or more markers, tags or blocking chains, whether the target polypeptide has been modified by glycosylation, acetylation, etc. Specific modifications will result in specific interactions with the pore, which can be determined using the methods described below. For example, cytosine and methylated cytosine can be distinguished based on the current passing through the pore during its interaction with each nucleotide.

[0044] As used herein, the term "nucleic acid strand to be characterized" or similar expressions may refer to a nucleic acid strand whose sequence-related characteristics are to be determined, such as a library, a single-stranded / double-stranded DNA, RNA, or a hybrid thereof. The term "target polynucleotide" or similar expressions may refer to a nucleic acid molecule obtained by ligating a biological molecule to be characterized to a sequencing adapter.

[0045] Characterization methods

[0046] The method of the present invention can be used to characterize biopolymers such as polynucleotides or polypeptides through nanopores. The method of the present invention uses an adapter (or main chain) that can form a reversible convex loop region and a motor protein that can recognize and bind to the convex loop region, so that the motor protein can be prevented or inhibited from moving relative to the adapter through the interaction force of the convex loop region and / or the double-stranded region without the need for other additional blocking elements, thereby preventing the adapter chain (main chain), especially the molecule to be characterized (such as the biopolymer to be characterized) connected thereto from undergoing undesirable translocation through the pore. When needed (for example, when sequencing begins), the double-stranded region can also be easily opened, thereby overcoming or eliminating the resistance to relative displacement, allowing the main chain and the molecule to be characterized connected thereto to move through the nanopore and complete the characterization. In some exemplary embodiments, when characterization is desired, it is only necessary to apply an electric potential to either or both sides of the molecular membrane of the nanopore so that the electrostatic force of the molecule to be characterized is opposite to the direction of the above-mentioned resistance. When the magnitude of the electrostatic force exceeds the resistance, the resistance is overcome and the molecule to be characterized passes through the nanopore for characterization. In other exemplary embodiments, non-covalent binding such as complementary pairing can be released by changing the ionic environment.

[0047] Thus, in the method of the present invention, the adapter and the motor protein can be combined in any order, as long as the motor protein can be bound to the convex loop region formed in the adapter before the characterization begins. For example, the adapter can be annealed to form a stem-loop structure before the motor protein is added, or after the motor protein is added, the mixture containing the motor protein and the adapter can be annealed to form a stem-loop structure bound to the motor protein. In the method of the present invention, the adapter and the motor protein can also be combined with the biopolymer to be characterized in any order. For example, the binding of the adapter to the biopolymer to be characterized can occur before, at the same time, or after the binding of the motor protein. In the present invention, the adapter and the molecule to be characterized can be connected in any manner known in the art, as long as it can keep the adapter connected to the molecule to be characterized before passing through the nanopore. Specific connection methods can be found in CN113736778A, which is incorporated herein by reference in its entirety.

[0048] Adapters and adapter complexes

[0049] The adapters provided herein can be used to characterize biopolymers such as polynucleotides in nanopore sequencing. The adapters comprise at least one double-stranded region formed by non-covalent interactions and a bulge region connected to the double-stranded region. For example, the adapters comprise at least a main strand capable of forming at least one stem-loop structure, or a main strand and an antisense strand capable of forming at least one bulge or internal loop structure.

[0050] As used herein, the meaning of "stem-loop structure" is substantially consistent with that known in the art, and refers to a nucleic acid, polypeptide, or a combination thereof having a secondary structure, wherein the secondary structure includes a paired region known or predicted to form a local double strand (stem), and the local double strand (stem) is connected on one side by a predominantly single strand region (bulge loop region). "Stem-loop structure" can also be used interchangeably with the terms "hairpin" and "fold-back" structure, and such structures are all well known in the art. For example, as non-limiting examples of amino acid sequences forming stem-loop structures, peptide chains having coiled coil (coiled coil peptide) structures, such as leucine zippers, can be cited (see Jody M. Mason et al., 2004, ChemBioChem, 5, 170-176; and Jonathan AR Worrall et al., 2011, FEBS Jornal, 278, 663-672).

[0051] In this article, "protrusions" primarily refer to structures formed by expansion at unpaired residues due to imperfect pairing in the secondary structure of nucleic acids, polypeptides, or a combination of the two. "Internal loops" primarily refer to bulges in the secondary structure of nucleic acids, polypeptides, or a combination of the two, resulting from the inability of opposite residues to form pairing (e.g., non-covalent interactions). Internal loops can be symmetrical or asymmetrical, while protrusions are always asymmetrical.

[0052] There are no particular restrictions on the length and composition of the double-stranded region, as long as the paired residues in the region can form a tight interaction, thereby forming a secondary structure that can be stably maintained before starting characterization. For example, it is known in the art that the stem portion of the stem-loop structure does not require accurate base pairing and can contain one or more base mispairings. Alternatively, the double-stranded region can also be accurately paired, i.e., does not include any mispairing. In the present disclosure, there are no particular restrictions on the length of the double-stranded region. From the perspective of being conducive to the stability of the stem-loop structure, the stem portion preferably has more than 5 pairs of mutually paired structural units.

[0053] In a preferred embodiment, the convex loop region includes at least one binding motif that can be recognized and bound by the motor protein. Those skilled in the art can select the binding motif in the convex loop region according to the motor protein used. For example, when using a helicase as a motor protein, it is preferred to use a single-stranded polynucleotide as a binding motif contained in the convex loop region. The number of binding motifs is not particularly limited, as long as it can be recognized and bound by the motor protein without hindering the formation of the convex loop region and the double-stranded region. Therefore, one or more binding motifs can be included in the convex loop region disclosed herein. When using a helicase as a motor protein, those skilled in the art can reasonably select the length of the single-stranded polynucleotide based on the length required for binding to different helicases and the number of expected binding helicases, preferably 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15 bp or more single-stranded polynucleotides as binding motifs for use in the convex loop region. The convex loop region and / or the double-stranded region may also additionally include one or more elements that can block the displacement of the motor protein relative to the adapter. In terms of the direction in which the motor protein is expected to move relative to the motor protein, the element is preferably located downstream of the motor protein binding motif. However, it will be understood by those skilled in the art that the blocking element is not required for the adaptor disclosed herein.

[0054] In one example, the adapter disclosed herein can be used to characterize the constructed library. Alternatively, after the library is end-repaired (converted to a flat-end DNA with 5'-phosphate and 3'-hydroxyl groups), it is connected to the adapter (flat end) by a ligase, or after the library is phosphorylated and repaired, it has sticky ends and is connected to an adapter with complementary sticky ends. The adapter can be connected to either side of the molecule to be characterized (for example, as shown in CN114134142A) or both sides.

[0055] In an optional embodiment, the adapter disclosed herein may further include a connecting chain. The connecting chain is used to bind to a connector provided on a molecular membrane of the nanopore protein or its vicinity, so that the molecule to be characterized is close to the molecular membrane. To achieve the above purpose, the connecting chain generally comprises a region capable of binding to the main chain and a region capable of binding to an external connector, and the binding is preferably non-covalent. For example, Chinese Patent Application Publication No. CN114134142A illustrates a connecting chain that can be used in the present invention. By further including a connecting chain, the adapter disclosed herein can also enrich the molecule to be characterized connected thereto near the nanopore protein or molecular membrane, increase the frequency of the pore passage of the molecule to be characterized, and improve the characterization efficiency. It can be understood that the connecting chain is not a necessary component of the adapter disclosed herein. In addition, there is no special limitation on the position of the region where the connecting chain is bound to the main chain, as long as it does not interfere with the formation of the convex loop region of the main chain and does not interfere with the binding of the motor protein to the convex loop region. In a preferred embodiment, the connecting chain is bound downstream of the stem-loop structure in the main chain along the direction in which the motor protein is expected to move, which may help reduce the interference of sequences located outside the bulge loop region and the double-stranded region in the formation of the bulge loop region and the double-stranded region in the main chain.

[0056] Those skilled in the art will understand that in the present application, the main chain and the connecting chain may include various chemical modification groups that can be used to modify the sequencing adapter, and those skilled in the art can choose according to different sequencing platforms, different sequencing requirements, and different requirements for molecules to be characterized. Including but not limited to: modification of abasic groups, especially groups that cannot form base pairs such as iSp18 groups or iSpC3, or nucleotides or nucleosides that lack a nucleobase at the 1' position of the ribose moiety (see, for example, U.S. Patent No. 5,998,203, Chinese Patent CN113462764B); modification of ribose, such as nucleotides with 2'-methoxy modification; LNA modification, such as iXNA, etc. Non-limiting examples of these modifications include, for example, organic oligocations, iSpC3, iSp18, iSp9, nitroindole, inosine, acridine, 2-aminopurine, 2-6-diaminopurine, 5-bromo-deoxyuracil, inverted thymidine (inverted dT), inverted dideoxythymidine (ddT), dideoxycytidine (ddC), 5-methylcytidylic acid, 5-hydroxymethylcytidine, 2'-O-methyl RNA base, isodeoxycytidine (iso-dC), isodeoxyguanosine (iso-dG), photocleavable (PC) groups, hexanediol, locked nucleic acid (LNA), peptide nucleic acid (PNA), methoxy (OMe), bicyclic nucleoside (BNA), glycerol nucleic acid (GNA), threose nucleic acid (TNA), cyclic pyrrolimidazole polyamide (cPIP) or double-stranded binding protein-modified nucleotides. These modifications can act as blocking elements, acting in conjunction with the bulge loop region and double-stranded region disclosed herein to keep the motor protein on the adapter. In some embodiments, when the adapter is double-stranded, these modifications can have the effect of stabilizing the double strand. For modifications in the main strand, please refer to Chinese Patent Application Publication No. CN116200477A.

[0057] In an optional embodiment, the adapter disclosed herein may further comprise a leader sequence in the main strand. The leader sequence is used to preferentially enter the nanopore, thereby facilitating the movement of the molecule to be characterized through the pore. The present invention may use any leader sequence suitable for achieving the above purpose.

[0058] The leader sequence generally includes a charged polymer. The polymer can be positively charged or negatively charged, preferably negatively charged. The charged polymer can be a polynucleotide, such as DNA or RNA, modified nucleotides (such as base-free DNA), PNA, LNA, polyethylene glycol (PEG), or polypeptides such as negatively charged polypeptides (such as multiple glutamine or asparagine residues) or positively charged polypeptides (such as arginine, lysine), or other negatively charged synthetic monomers, such as dextran sulfate. The leader sequence preferably comprises a polynucleotide, more preferably a single-stranded polynucleotide. The leader sequence can be any length, but it is generally 10 to 150 nucleotides in length, such as 20 to 150 nucleotides in length. The length of the leader sequence generally depends on the transmembrane pore used in the method.

[0059] Connection Chain

[0060] In the characterization of biomolecules, the more free biomolecules to be tested are near the transmembrane nanopore, the more efficiently the biomolecules to be tested pass through the transmembrane nanopore, and thus the higher the detection efficiency. Therefore, by setting a connector in or near the transmembrane nanopore and connecting it to the connector through a connecting chain, the concentration of the biomolecules to be tested enriched around the transmembrane nanopore and the efficiency of passing through the transmembrane nanopore can be increased.

[0061] If the molecular membrane is an amphiphilic layer, such as a copolymer membrane or a lipid bilayer, the connector may include a polypeptide anchor and / or a hydrophobic anchor embedded in or on the surface of the molecular membrane, and a tether connected to the polypeptide anchor and / or the hydrophobic anchor. The tether is a polynucleotide chain with a certain specificity, and the tether can be connected to the connector by base pairing with the side chain region of the connecting chain.

[0062] It will be understood by those skilled in the art that a tethered cholesterol can be provided on an external matrix / connector (e.g., a nanopore protein or a molecular membrane near the nanopore), and a portion of the connecting chain in the adapter can bind to the tether (e.g., base complementary pairing), and can be connected to the molecular membrane of the transmembrane nanopore based on the tethered cholesterol. The region of the connecting chain that binds to the tether can be a charge-dense region that can wander near the transmembrane nanopore. When it binds to the tether (e.g., base complementary pairing), under the action of an external potential, the biomolecule to be tested connected to the adapter enters the transmembrane nanopore to characterize the biomolecule to be tested through changes in electrical signals.

[0063] Reagent test kit

[0064] The present application also provides a kit for characterizing a biopolymer, comprising the adaptor and motor protein disclosed herein. In the kit of the present application, the above components can be physically separated and independently contained in different containers, or can be combined in any manner or all contained in the same container.

[0065] The kit of the present application may also include any of the nanopores and nanopore proteins discussed above. The kit may also include components of the membrane, such as phospholipids such as lipid bilayers required to form an amphiphilic layer. The kit may additionally include one or more other reagents or instruments that enable any of the embodiments mentioned above to be implemented. Such reagents or instruments include one or more of the following reagents or instruments: a suitable buffer (aqueous solution), a tool for obtaining a sample from a subject (e.g., a container or instrument comprising a needle), a tool for amplifying and / or expressing biomolecules, a membrane or pressure clamp or patch clamp device as defined above. The reagents may be present in a dry state in the kit so that the fluid sample is resuspended in the reagent.

[0066] Optionally, the kit may further include instructions for using the kit in the method of the present invention, or detailed information on which patients the method can be used for. The kit may optionally include the necessary components to promote the movement of the helicase (e.g., ATP and Mg). 2+ ).

[0067] The present invention is further described below by way of specific examples, but this is not intended to limit the present invention. Those skilled in the art may make various modifications or adjustments based on the teachings of the present invention without departing from the spirit and scope of the present invention.

[0068] Example

[0069] Unless otherwise specified, the experimental methods in the following examples are conventional methods and were performed according to the techniques or conditions described in the literature in the field or according to the product instructions. The materials, reagents, and instruments used in the following examples are all commercially available unless otherwise specified.

[0070] Experimental materials, reagents, instruments and experimental methods

[0071] The amino acid sequence of RM1 helicase is shown in SEQ ID NO: 1; the amino acid sequence of M1 helicase T4 Dda-M1G / E94C / C109A / C136A / A360C is shown in SEQ ID NO: 2;

[0072] The main chain modification method can be found in CN114134142A, the entire text of which is incorporated herein by reference.

[0073] Unless otherwise explicitly stated herein, other reagents, consumables, and methods used herein are specifically described in the instruction manual of the QNome-3841 sequencer (or its equivalent, newer, or upgraded model sequencer), as well as the instructions for the dedicated library preparation kit, sequencing kit, and sequencing chip specified in the instruction manual of the QNome-3841 sequencer (or its equivalent, newer, or upgraded model sequencer).

[0074] Example 1:

[0075] This example describes an exemplary SL-type adapter capable of forming a stable complex with a translocase and its preparation. The adapter can be used for RNA sequencing.

[0076] First, the main chain S1, antisense chain S2, and antisense chain S3 were chemically synthesized respectively.

[0077] Main strand S1: 5'- G CGACAACGACTTATG -8888-TTTTTTTTTTTT- CATAAGTCGTTGTCG C -33333333333333333333333333333-3';

[0078] Antisense strand S2: 5'- -3';

[0079] Antisense strand S3: 5'-TTT T-3'.

[0080] The multiple 3s at the 3' end of S1 represent the iSpC3 modification, a leader sequence that increases pore access; the multiple 8s within the strand represent the iSp18 modification, which blocks enzyme migration; and the multiple Ts allow for preferential binding of the translocase protein RM1. Upon annealing, the main strand S1 forms a stem-loop structure. The two nucleotide sequences marked with single underscores in S1 are reverse-complementary, forming the stem, while the portion of sequence between the reverse-complementary sequences forms the bulge region. Antisense strands S2 and S3 complement each other with sequences outside the stem-loop structure at the 3' and 5' ends of the main strand S1, respectively.

[0081] In a suitable buffer (10mM HEPES (pH7.0), 100mM NaAC), the main chain S1, antisense chain S2 and antisense chain S3 were mixed in a ratio of 1:1.1:1.1, and annealed by slowly cooling from 95°C to 25°C with a cooling amplitude of no more than 0.1°C / s to obtain SL-type adapter 1. Then, 7.5 times the amount of RM1 protein and a cross-linker 1,8-bismaleimido-diethylene glycol (BM(PEG)2 (purchased from Thermo Scientific, product number 22336) with a final concentration of 200μM were added to 1μM annealed SL-type adapter 1, and incubated at 30°C for 30min to obtain a Figure 1 A shows the structure of SL-type adaptor-enzyme complex 1.

[0082] The prepared complex 1 was subjected to TBE PAGE gel electrophoresis at 160V for 50 min. The results are shown in Figure 2 In. By Figure 2 It can be seen that after incubation of SL-type linker 1 with RM1 protein, a part of the linker binds to RM1 protein to form a nucleic acid-protein complex, resulting in its migration rate in the electric field being significantly slower than that of the free linker not bound to RM1 protein.

[0083] Comparative Example 1:

[0084] This comparative example describes a classic Y-type adapter.

[0085] The main strand S4, antisense strand S5, and antisense strand S6 were chemically synthesized respectively.

[0086] The main chain S4 is: 5'- -8888-TTTTTTTTTT- ACTGCTCAT TCGGTCCTGCTGACT -33333333333333333333333333333-3';

[0087] Antisense strand S5 is: 5'- AGTCAGCAGGACCGAATGAGCAGT -3';

[0088] Antisense strand S6: 5'-TTT T-3'.

[0089] The 3' end of the main strand S4 also features multiple iSpC3 modifications, which act as a leader sequence to enhance pore access. The strand also contains a poly-T sequence for binding to the RM1 protein and multiple iSp18 modifications to block RM1 protein migration. Antisense strands S5 and S6 also complementarily bind to the single-stranded sequences at either end of the main strand S4. However, the main strand S4 lacks the reverse complementary sequences flanking the translocase binding site, preventing the formation of a stem-loop structure.

[0090] Based on the same method as in Example 1, the main strand S4, the antisense strand S5 and the antisense strand S6 were mixed in a ratio of 1:1.1:1.1, and annealed by slowly cooling from 95°C to 25°C with a cooling amplitude of no more than 0.1°C / s to obtain a Y-shaped adapter 1, thereby preparing a Figure 1 B shows the structure of the Y-shaped adapter-enzyme complex 2.

[0091] Example 2:

[0092] This example describes the differences between complexes prepared using SL-type adapters and classic Y-type adapters and translocases.

[0093] The complex 1 prepared in Example 1 and the complex 2 prepared in Comparative Example 1 were loaded onto a DNAPacPA200 column (Thermo Scientific, Cat. No. 082510) and washed with 5 to 10 column volumes of elution buffer A (20 mM Na-CHES, 250 mM NaCl, 4% (W / V) glycerol, pH 8.6) to remove the enzyme and nucleic acid not bound to the adapter. Then, elution was performed with a mixture of elution buffer A and buffer B (20 mM Na-CHES, 1 M NaCl, 4% (W / V) glycerol, pH 8.6) for 10 column volumes. The elution results are shown in FIG. Figure 3 A and 3B. Collect Figure 3 The three elution peaks in A and Figure 3 The fractions corresponding to the five elution peaks in B were detected by TBE PAGE gel electrophoresis (160 V, 50 min). The results are shown in Figure 4 A and 4B.

[0094] Compare Figure 3 A. 4A and Figure 3 B and 4B show that during the preparation and purification process of the complex obtained by combining the classic Y-type adapter with the translocase RM1 protein, the enzyme seriously detached from the adapter, forming more miscellaneous peaks and miscellaneous bands. Using ImageJ software, the grayscale value of the target complex band relative to the grayscale value of the total band was calculated to determine the purity of the nucleic acid-protein complex in the sample (purity = grayscale of the enzyme-linker complex / (grayscale of the linker nucleic acid + grayscale of the enzyme-linker complex)). The purity of the nucleic acid-protein complex (E3) in the eluted sample of the final Y-type adapter can only reach about 55%. In contrast, the adapter-enzyme complex (E1) formed by the SL-type adapter has a purity of more than 90%, which is significantly higher than that of the Y-type adapter.

[0095] Example 3:

[0096] This example compares the stability of complexes prepared from an exemplary SL-type adaptor of the present invention and a commonly used Y-type adaptor.

[0097] Approximately 10 ng of the purified complex 1 and complex 2 prepared in Example 2 were added to 50 μL of Seq buffer (10 mM HEPES pH 7.0, 600 mM KCl, 25 mM ATP, 25 mM MgCl2), mixed, and allowed to stand at 40°C. A control group was treated with only RNAase-free water added to the buffer. TBE PAGE gel electrophoresis (160 V, 50 min) was performed at 0, 2, 4, and 20 hours after the start of the stand, and the rate of enzyme shedding from the complex was measured over time. The results are shown in Figure 2. Figure 5 .

[0098] Figure 5 It shows that during the 20-hour standing process, only a small amount of translocase fell off from the complex formed with the SL-type adapter; in contrast, for the complex formed with the Y-type adapter, most of the translocase had fallen off from the complex after about two hours of standing. At the 20th hour, basically all the Y-type adapters existed in a free state without any translocase bound to them.

[0099] According to the manufacturer's manual, ImageJ software (open source software) was used for statistical analysis. Figure 5 The grayscale values ​​of each band are summarized in Table 1 below. As can be seen from Table 1, the vast majority (more than 85%) of SL-type adapter-enzyme complexes can exist stably for at least 2 hours. Even after standing for 20 hours, about 70% of the SL-type adapters remain bound to the translocase. This shows that the complex prepared using the SL-type adapter exhibits significantly higher stability than the complex prepared using the classic Y-type adapter over a long time window. In the practice of long sequencing during nanopore sequencing, the library to be tested usually needs to wait in the sequencing buffer for a long time before it can be tested through the pore. Therefore, keeping the motor protein such as the translocase bound to the adapter is conducive to the successful determination of the sequence, thereby improving the sequencing efficiency.

[0100] Table 1

[0101]

[0102] Example 4:

[0103] This example describes the results of nanopore sequencing using exemplary SL-type adapters.

[0104] The SL-type adapter-enzyme complex 1 prepared in Example 2 was covalently linked to the homemade RNA sequencing library (RNA containing poly A tail) Figure 1The 5' end of the adapter main chain shown in A, and according to the literature "Yunhao Wang et al., 2021, Nanopore sequencing technology, bioinformatics and applications, Nature Biotechnology" Figure 3 The RNA sequencing library was constructed by the method shown in the figure. The sequencing was performed using the QNome-3841 sequencer of Qi Carbon Technology Co., Ltd. The results are shown in Figure 6 .

[0105] Figure 6 The results show that the SL-type adapters of the present invention can generate sequencing signals normally, with clear signal steps, allowing for subsequent sequence analysis. This indicates that after sequencing begins, the stable complex formed by the adapters of the present invention and the translocase can release the enzyme blockage as needed, allowing subsequent through-hole sequencing to proceed normally.

[0106] Example 5:

[0107] This example describes another exemplary SL-type adapter that can be used for DNA sequencing.

[0108] The main chain S7 and antisense chain S8 were chemically synthesized respectively.

[0109] Main chain S7: 5'-33333333 / i2OMeU / / i2OMeU / / i2OMeU / / i2OMeU / / i2OMeU / / i2OMeU / / i2OMe U / / i2OMeU / / i2OMeU / / i2OMeU / / i2OMeU / / i2OMeU / / i2OMeU / / i2OMeU / / i2OMeU / / i2OMeU / / i2 OMeU / / i2OMeU / / i2OMeU / / i2OMeU / / i2OMeU / / i2OMeU / / i2OMeU / / i2OMeU / / i2OMeU / / i2OMeU / TTTTTTTTTTTTTTT8888 / i2OMeU / / i2OMeU / / i2OMeU / / i2OMeU / ACTGCTCATTCGGTCCTGCTGAC T-3';

[0110] Antisense strand S8: 5'-P- GTCAGCAGGACCGAATGA / i2OMeG / / i2OMeC / / i2OMeA / / i2OMeG / / i2OMeU / / i2OMeA / / i2OMeG / / i2OMeU / / i2OMeC / / i2OMeC / / i2OMeA / / i2OMeG / / i2OMeC / / i2OMeA / / i2OMeC / / i2OMeC / / i2OMeG / / i2OMeA / / i2OMeC / / i2OMeC / -3'.

[0111] In S7, the multiple 3s at the 5' end represent iSpC3 modifications, which are used to increase the leading sequence for entering the pore; the multiple 8s within the chain represent iSp18 modifications, which are used to block enzyme movement; i2Ome represents a nucleotide with a 2'-methoxy modification on the ribose ring. The modified nucleotide cannot bind to the helicase, but does not affect the binding with the complementary base pairing; iXNA is an LNA modification that can be combined with the iSp18-modified nucleotide to block the movement of the enzyme. S7 also forms a stem-loop structure after annealing. The double underline marks the reverse complementary sequence that forms the stem, and the convex loop region is located between the two reverse complementary sequences. The antisense chain S8 can be complementary to the 3' end sequence outside the stem-loop structure in the main chain S7 (shown by a single underline).

[0112] The main strand S7 and the antisense strand S8 were mixed at a ratio of 1:1.1 and annealed by slowly cooling from 95°C to 25°C with a cooling rate of no more than 0.1°C / s to obtain SL-type adapter 2. Then, 7.5 times the amount of M1 protein and 1 mM TMAD cross-linker (Sigma, catalog number 8.21101) were added to 1 μM adapter 2 and incubated at 30°C for 30 min to obtain a 1 μM adapter. Figure 7 The structure of the SL-type adaptor-enzyme complex 3 is shown.

[0113] The prepared complex 3 was loaded onto a DNAPac PA200 column and washed with 5 to 10 column volumes of elution buffer A (20 mM Na-CHES, 250 mM NaCl, 4% (W / V) glycerol, pH 8.6) to remove impurities such as enzymes and nucleic acids that were not bound to the adapter. Then, the complex was eluted with a mixture of buffer A (20 mM Na-CHES, 250 mM NaCl, 4% (W / V) glycerol, pH 8.6) and buffer B (20 mM Na-CHES, 1 M NaCl, 4% (W / V) glycerol, pH 8.6) for 10 column volumes. The elution results are shown in Figure 8 A. Collection Figure 8 The corresponding fractions of the three elution peaks in A were detected by TBE PAGE gel electrophoresis (160 V, 50 min), and the results are shown in Figure 8 B.

[0114] Figure 8It was shown that the complex formed by another exemplary SL-type adaptor of the present invention and translocase M1 protein also had a purity of more than 90%.

[0115] Example 6:

[0116] About 10 ng of the purified complex 3 prepared in Example 5 was added to 10 μL of Seq buffer (10 mM HEPES pH 7.0, 500 mM KCl, 100 mM ATP, 100 mM MgCl2), mixed, and allowed to stand at 40°C. The control group was prepared by adding the same amount of Y-type adapter AMX (component LA of the Oxford Nanopore (ONT) kit SQK-LSK114) to the buffer. At 0, 0.5, 2, 3, 4, and 24 hours after the start of standing, TBE PAGE gel electrophoresis (160 V, 50 min) was performed to measure the ratio of enzyme shedding from the complex over time. Among them, the results at 0 and 24 hours are shown in Figure 9 .

[0117] Figure 9 The results showed that another exemplary SL-type adaptor can also form a stable complex with the translocase (M1 protein) and maintain it for a long time. Under the same conditions, the currently available ONT adaptor with a classic Y-shaped structure showed much lower stability.

[0118] ImageJ software was used for statistical analysis in the same manner as in Example 3. Figure 9 The band intensity in Table 2 was obtained. Table 2 shows more intuitively that the stability of the complex formed by the SL-type adapter disclosed in this article and the translocase is much higher than that of the commonly used Y-type adapter currently on sale. The ratio of enzyme shedding of the complex prepared by the Y-type ONT adapter is close to half after being placed for only 30 minutes. In contrast, the nucleic acid-enzyme complex obtained using the SL-type adapter of the present invention still has more than 85% of the enzyme remaining bound to the adapter even after being placed for 24 hours, indicating that the SL-type adapter of the present invention can effectively reduce the ratio of enzyme shedding during sequencing.

[0119] Table 2

[0120]

[0121] Example 7:

[0122] This example describes the results of nanopore sequencing using another exemplary SL-type adapter.

[0123] According to the manufacturer's instructions, the SL-type adapter-enzyme complex 3 prepared in Example 5 was covalently linked to a homemade DNA library (a 10 kb DNA library prepared by end repair) using the library construction kit QLK-V1.1.1 (Qican Technology Co., Ltd.) to construct a test library. The sequence was then sequenced using a QNome-3841 sequencer (Qican Technology Co., Ltd.). The signal is shown in Figure 10 .

[0124] Figure 10 The results show that the SL-type adapter of the present invention can effectively complete the DNA sequencing process and generate a clear signal that can be used for subsequent analysis. This shows that the SL-type adapter that forms a stable complex with the helicase before sequencing begins can relieve the blockage of the relative displacement of the helicase after sequencing begins, and will not hinder the normal subsequent through-hole sequencing. At the same time, it suggests that the SL-type adapter of the present invention can achieve accurate control of the start time of nanopore sequencing. In addition, Figure 10 It also shows that the peaks of the sequencing signals are relatively uniform in time sequence, and there are no dense peak areas or sparse peak areas in time sequence, indicating that the SL-type adapter will not interfere with the helicase moving at a stable relative rate on the sequence to be tested.

[0125] Example 8:

[0126] This example describes an example of a single-stranded SL-type adaptor that is able to form a stable complex with a translocase.

[0127] First, the following single chains were synthesized by chemical synthesis:

[0128] SL single strand: 5'- GCGACAACGACTTATG -TTTTTTTTTTTT- CATAAGTCGTTGTCGC -3'

[0129] The SL single-strand solution was heated to 95°C and then slowly cooled to 25°C with the cooling rate not exceeding 0.1°C / s, so that the SL single-stranded solution formed a stem-loop structure (the underlined sequences at both ends complement each other to form the stem, and the polyT sequence in the middle formed a convex loop).

[0130] To 1 μM annealed SL single chain, 7.5 times the amount of RM1 protein and 75 times the amount of cross-linker 1,8-bismaleimido-diethylene glycol (BM(PEG)2; purchased from Thermo, product number 22336) were added and incubated at 30°C for 30 minutes. Figure 11 B shows the structure of SL-type adaptor-enzyme complex 4.

[0131] The prepared complex 4 was subjected to TBE PAGE gel electrophoresis (160V, 30min), and the results are shown in Figure 11 A. Figure 11In A, the RM1 protein binds to the convex loop region of the stem-loop structure of the SL single chain, resulting in the migration rate of the formed nucleic acid-protein complex in the electric field being significantly slower than that of the free single-chain adapter not bound to the RM1 protein.

[0132] The resulting complex 4 was loaded onto a DNAPac PA200 column (Thermo Scientific; Cat. No. 082510) for purification to remove unbound enzymes and nucleic acids. It was then eluted with a mixture of 10 column volumes of buffer A (20 mM Na-CHES, 250 mM NaCl, 4% (W / V) glycerol, pH 8.6) and buffer B (20 mM Na-CHES, 1 M NaCl, 4% (W / V) glycerol, pH 8.6). The elution peaks were collected and measured on a TBE PAGE gel (160 V, 30 min). The results are shown in Figure 2. Figure 11 C and 11D. Figure 11 As measured in Figures C and D, the complex prepared using the SL single chain and the RM1 protein was essentially free of detachment during the preparation and purification process, and had a purity of over 99%.

[0133] Example 9:

[0134] This example describes a double-stranded SL-type adaptor capable of forming a stable complex with a translocase, such as Figure 12 shown.

[0135] The adapter is composed of two single-stranded nucleotide chains. One is the main chain, which has a sequence structure such as the main chain S4 in Example 1, that is, it includes a 5'-terminal single-stranded region, a 5'-terminal reverse complementary region, a bulge loop region, a 3'-terminal reverse complementary region, and a 3'-terminal single-stranded region in the direction from 5' to 3', wherein the main chain forms a stem-loop structure disclosed herein through the 5'-terminal reverse complementary region, the bulge loop region, and the 3'-terminal reverse complementary region after annealing. The other is an antisense chain, which includes a 3'-terminal pairing region and a 5'-terminal pairing region, which can complement and pair with a portion of the 5'-terminal single-stranded region or the 3'-terminal single-stranded region in the main chain, thereby increasing the stability of the adapter forming the stem-loop structure.

[0136] If necessary, the double-stranded SL-type adapter may further comprise elements known to be suitable for use in nanopore adapters, such as a blocking element, a leader sequence, a linker binding region, and a library connection region.

Claims

1. A complex for characterizing an analyte or controlling the movement of a motor protein, comprising an adaptor and a motor protein, wherein the adaptor comprises at least one double-stranded region formed by non-covalent interactions and a loop region connected to the double-stranded region, and the motor protein binds to and stalls at the loop region.

2. The composite according to claim 1, wherein, The loop region comprises at least one motor protein binding motif.

3. The composite according to any one of the preceding claims, wherein, The double-stranded region is selected from reverse complementary nucleotide sequences or amino acid sequences containing a coiled coil region.

4. The composite according to any one of the preceding claims, wherein, The loop region and / or the double-stranded region further comprises a blocking element for blocking the displacement of the motor protein relative to the adaptor, and the blocking element is preferably located downstream of the loop region and / or the double-stranded region; Optionally, the blocking element comprises any one or any combination selected from organic oligocations, iSpC3, iSp18, iSp9, nitroindole, inosine, acridine, 2-aminopurine, 2,6-diaminopurine, 5-bromo-deoxyuridine, reverse thymidine (reverse dT), reverse dideoxythymidine (ddT), dideoxycytidine (ddC), 5-methylcytidylate, 5-hydroxymethylcytidine, 2'-O-methyl RNA base, isodeoxycytidine (iso-dC), isodeoxyguanosine (iso-dG), photocleavable (PC) group, hexanediol, locked nucleic acid (LNA), peptide nucleic acid (PNA), methoxy (OMe), bicyclic nucleoside (BNA), glycerol nucleic acid (GNA), threose nucleic acid (TNA), cyclopyrrole imidazole polyamide (cPIP) or double-stranded binding protein-modified nucleotides, nucleotides with a benzostructure as the nucleobase.

5. The composite according to any one of the preceding claims, wherein, The double-stranded region and the loop region form a stem-loop structure, a bulge structure or an internal loop structure, or any combination thereof.

6. The composite according to any one of the preceding claims, wherein, The adaptor further comprises a region located outside the loop region, preferably on either one or both sides of the structure formed by the loop region and the double-stranded region, and the region on either one or both sides preferably comprises a modification or element capable of hindering the motor protein from binding to the sequence outside the loop region, and the modification is preferably selected from 2'-methoxylated ribose; Optionally, the region on either one or both sides may comprise an element selected from a leader sequence, a linker-binding region, a library-linking region.

7. The complex according to any one of the preceding claims, further comprising a linker chain, the linker chain comprising a region for binding to an external linker and comprising a complementary region complementary to a partial region in the region on either one or both sides in the adaptor, enabling the linker chain to bind to the adaptor by annealing, preferably in a complementary pairing manner.

8. The composite according to any one of the preceding claims, wherein, The adaptor consists of a first single-stranded nucleotide; Optionally, the adaptor comprises a first single-stranded nucleotide and a second single-stranded nucleotide, wherein the first single-stranded nucleotide comprises a stem-loop structure, and the second single-stranded nucleotide comprises a pairing region reverse complementary to a partial region outside the stem-loop structure in the first single-stranded nucleotide and located downstream along the moving direction of the motor protein; Optionally, the adaptor comprises a first nucleotide single strand and a second nucleotide single strand, wherein the first nucleotide single strand comprises a stem-loop structure, and the second nucleotide single strand comprises a pairing region that is reverse complementary to partial regions on both outer sides of the stem-loop structure in the first nucleotide single strand, and the pairing region is optionally composed of consecutive or non-consecutive nucleotides in the second nucleotide single strand; Optionally, the adaptor comprises a first nucleotide single strand, a second nucleotide single strand and a third nucleotide single strand, wherein the first nucleotide single strand comprises a stem-loop structure, the second nucleotide single strand comprises a pairing region that is reverse complementary to a partial region outside the stem-loop structure in the first nucleotide single strand and on the downstream side along the moving direction of the motor protein, and the third nucleotide single strand comprises a pairing region that is reverse complementary to a partial region on the other side outside the stem-loop structure in the first nucleotide single strand.

9. The composite according to any one of the preceding claims, wherein, The motor protein is a polynucleotide-binding protein, Optionally, the polynucleotide-binding protein is selected from at least one of DNA polymerase, RNA polymerase, helicase, endonuclease, exonuclease, DNA protease, topoisomerase, nucleic acid translocase, nucleic acid nickase, or any fusion protein thereof.

10. The complex according to any one of the preceding claims, further comprising an analyte, wherein the analyte is directly or indirectly linked to the adaptor.

11. A method for preparing the complex according to any one of claims 1 to 10, the method comprising: (a) providing the adaptor according to any one of claims 1 to 10; (b) binding and stalling the motor protein at the loop region of the adaptor to form a complex; Optionally, the method further comprises: (b’) connecting the adaptor to the analyte; wherein step (b’) occurs before, simultaneously with, or after step (b).

12. A method for characterizing an analyte or controlling the movement of a motor protein, the method comprising: (i) providing the complex according to any one of claims 1 to 10 or the complex prepared according to claim 11; (ii) providing conditions for the motor protein to displace relative to the loop region and partially or completely pass through the double-stranded region.

13. The method according to claim 12, wherein, The conditions in step (ii) are optionally selected from the direction of the electric field, the intensity of the electric field, and / or the concentration of nucleoside triphosphates, preferably the concentration of ATP.

14. The method according to any one of claims 12 to 13, further comprising: (iii) passing the analyte through a pore for characterization and displacing relative to the pore to obtain one or more measurement values, which represent one or more characteristics of the analyte.

15. The method according to any one of claims 12 to 14, wherein The one or more characteristics are selected from (i) the length of the analyte, (ii) the sequence identity of the analyte, (iii) the sequence of the analyte, (iv) the secondary structure of the analyte, and (v) the modification status of the analyte.

16. The method according to any one of claims 12 to 15, wherein, The one or more characteristics can be measured by electrical measurement and / or optical measurement.

17. A kit for characterizing an analyte or controlling the movement of a motor protein, comprising the complex according to any one of claims 1 to 10, or comprising the adaptor in an independent package and the motor protein in an independent package, wherein the motor protein binds and stalls at the lug region of the adaptor before starting characterization or movement; Optionally, the kit further comprises a protein for forming a pore and / or the membrane component for carrying the pore.

18. Use of the complex according to any one of claims 1 to 10 in the preparation of a reagent or a kit for characterizing an analyte or controlling the movement of a motor protein.

19. A device for characterizing an analyte and controlling the movement of a motor protein, the device comprising: A pore; An insulating layer for carrying the pore and penetrated by the pore on both sides; And The complex according to any one of claims 1 to 10, distributed on either one or both sides of the insulating layer; Optionally, the device further comprises an electrical component for providing a potential difference across the insulating layer carrying the pore.

20. The apparatus according to claim 19, wherein The pore is selected from a nanopore, a transmembrane pore, a biological pore, a solid-state pore, or a pore formed by hybridization of a biological and a solid-state material; Preferably, the biological pore is derived from hemolysin, leukocidin, CsGG, Mycobacterium smegmatis porin A (MspA), porin B, porin C, porin D, outer membrane porin F, outer membrane porin G, outer membrane phospholipase A, Neisseria autotransporter lipoprotein, and WZA; Preferably, the solid-state pore is derived from a graphene nanopore, a MoS2 nanopore, a BN nanopore, or a PA63 nanopore.

21. The device according to claim 19 or 20, wherein, The analyte is selected from polynucleotides, polypeptides, polysaccharides, or lipids.

Citation Information

Patent Citations

  • Hairpin-like adaptors, constructs, and methods for characterizing double-stranded target polynucleotides

    CN113462764B

  • Sequencing joint, construction method, nano-hole library building kit and application

    CN113736778A

  • Linker, complex, single-chain molecule, kit, method and application

    CN114134142A

  • Linker comprising joint blocking element, constructs, methods and uses

    CN116200477A

  • Enzymatic nucleic acids containing 5'-and / or 3'-cap structures

    US5998203A