Methods of sequencing double-stranded DNA
By forming interstrand crosslinks between the first and second strands of a double-stranded nucleic acid and breaking them at or near the crosslinks, single-molecule sequencing technology is used to sequence the crosslinked nucleic acid constructs. This solves the problems of complexity and contamination risk in sequencing both strands of double-stranded nucleic acids in existing technologies, and achieves simplified and efficient sequencing.
Patent Information
- Application Number
- CN202080081357.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-11-22
- Filing Date
- 2020-11-20
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2040-11-20
AI Technical Summary
Existing technologies struggle to efficiently sequence both strands of double-stranded nucleic acids, especially due to the increased experimental complexity and sample preparation time caused by the use of hairpin adaptors, which may also lead to contamination and degradation risks.
Cross-linked nucleic acid constructs are formed by creating interstrand crosslinks between the first and second strands of a double-stranded nucleic acid and breaking them at or near the crosslinking site using a crosslinking agent. These constructs are then sequenced using single-molecule sequencing technology.
It provides orthogonal alignment sequence information for the two strands of double-stranded nucleic acids, avoiding the use of hairpin adapters, simplifying the experimental process, and reducing sample contamination and complexity.
Smart Images

Figure CN114945679B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to a method for sequencing a target double-stranded nucleic acid by forming an interstrand crosslink between the first and second strands of the double-stranded nucleic acid. This disclosure also relates to kits, systems, and apparatus for performing such methods. Background Technology
[0002] Nucleic acid sequencing is an increasingly important aspect of medical, biomedical, and biotechnological applications. It allows for the study of genomes and the proteins they encode, for example, allowing for the correlation between nucleic acid mutations and observable phenomena (e.g., disease indicators). Sequencing can be used in evolutionary biology to study relationships between organisms. Metagenomics involves identifying organisms present in a sample by sequencing nucleic acids that allow identification of these organisms, such as microbes in the microbiome. In medicine, genetic testing of a subject may highlight the risk of a genetic disease or allow for the selection of the best treatment for the disease. DNA sequencing is also a key technology in forensic medicine. Given the many different applications of DNA sequencing, there is a need for improved nucleic acid sequencing methods, especially for double-stranded nucleic acids.
[0003] Nucleic acids can be sequenced using many known techniques. Single-molecule techniques have proven particularly attractive because they offer high fidelity, avoid amplification bias, and can potentially achieve extremely long read lengths. Single-molecule sequencing can also provide information about the presence of characteristic features such as base modifications, oxidation, reduction, decarboxylation, and deamination.
[0004] One promising approach for single-molecule nucleic acid sequencing is nanopore sensing. Nanopore sensing is a method for analyte detection and characterization that relies on the observation of individual binding or interaction events between analyte molecules and ion conduction channels. Nanopore sensors can be created by placing nanoscale monopores within an electrically insulating membrane and measuring the voltage-driven ion current flowing through the pore in the presence of analyte molecules. The presence of an analyte inside or near the nanopore will alter the ion flow through the pore, resulting in changes in the ions or current measured on the channel. The identity of the analyte is revealed by its unique current characteristics, particularly the duration and extent of the current block and the changes in current level during interaction with the pore. Nanopore sensing may allow for rapid and inexpensive polynucleotide sequencing, providing single-molecule sequence reads of polynucleotides ranging from tens to tens of thousands of bases in length.
[0005] Other known single-molecule sequencing methods for nucleic acids include real-time sequencing based on the detection of base incorporation catalyzed by enzymes such as DNA polymerase. In one known method, the polymerase catalyzes the formation of a complementary strand of the target nucleic acid strand to be sequenced using labeled nucleotides. Labeled nucleotides are sequentially incorporated into the synthesizing strand, releasing a label (e.g., a fluorescent label) attached to each nucleotide, and the detection of the label allows the identification of the nucleotide, thereby determining the sequence of the template strand.
[0006] Single-molecule sequencing of nucleic acids typically relies on the controlled processing of polynucleotides. For example, in nanopore sensing of polynucleotides, controlling the movement of polynucleotides relative to the pore is crucial. Uncontrolled movement can prevent or hinder accurate characterization of polynucleotides. For instance, accurately distinguishing each nucleotide in a homopolymer polynucleotide is problematic when the movement of polynucleotides relative to the pore is not controlled. Similarly, in sequencing-by-synthesis techniques, controlling the synthetic reactions with respect to the sensor used is important. Uncontrolled movement can prevent or hinder accurate characterization of polynucleotides because bases may be skipped or processed too quickly, and signals are not easily deconvolved.
[0007] To address this problem, it is known to use polynucleotide processing enzymes (also referred to as motor proteins in some embodiments) to control and characterize the movement of polymers (e.g., polynucleotides). Suitable enzymes include polynucleotide processing enzymes such as helicases, exonucleases, topoisomerases, etc. This protein processes polynucleotides in a controlled manner. For example, in nanopore sequencing, motor proteins can be used to control the movement of polymers such as polynucleotides relative to the pore. In single-molecule real-time sequencing, polymerases control the movement of the target polynucleotide as it moves relative to the polymerase during processing.
[0008] One challenge in single-molecule nucleic acid sequencing is improving the accuracy of base calling. In conventional techniques, only one strand of a double-stranded nucleic acid (e.g., the template) is sequenced, while the other strand (e.g., the complement) is either omitted from the sequencing reaction or its information is discarded. However, the inventors have recognized that sequencing both strands of a double-stranded polynucleotide can provide valuable orthogonal alignment information. Therefore, sequencing information obtained for the first strand of a double-stranded nucleic acid can be supplemented by sequencing information obtained for the second strand.
[0009] Several methods are known in the art to allow sequential sequencing of the two strands of a double-stranded nucleic acid. One known method involves attaching a so-called hairpin adapter to the double-stranded nucleic acid. The hairpin adapter links the two nucleic acid strands together through their 5' and 3' ends, thus forming a continuous strand that can be sequenced using single-molecule techniques.
[0010] While hairpin adaptors offer some advantages, new methods for sequencing both strands of double-stranded nucleic acids are still needed. For example, modifications to the 3' or 5' ends of the double-stranded nucleic acid to be sequenced can hinder hairpin adaptor attachment. The requirements for the hairpin adaptor itself increase the complexity of experimental design and may require specialized ligation reagents. The requirements for attaching the adaptor may increase sample preparation time or complexity, thus posing a risk of sample contamination or degradation. Therefore, methods for sequencing both strands of double-stranded nucleic acids that do not require hairpin adaptors would be valuable. Summary of the Invention
[0011] This disclosure relates to a method for sequencing double-stranded nucleic acids. The method includes forming an interstrand crosslink between the first and second strands of the double-stranded nucleic acid. The first and second strands of the double-stranded nucleic acid are typically referred to as the template and complement, respectively.
[0012] The methods disclosed herein include forming interstrand crosslinks between the first and second strands of a double-stranded nucleic acid by exposing it to a crosslinking agent. In some embodiments, the method further includes allowing the target double-stranded nucleic acid to break at or near the interstrand crosslink formed by the crosslinking agent. This breakage results in the formation of at least one double-stranded nucleic acid construct, which is crosslinked at or near its ends. The method includes sequencing the thus formed double-stranded nucleic acid construct using single-molecule sequencing technology.
[0013] Therefore, this paper provides a method for sequencing target double-stranded nucleic acids, which includes:
[0014] (a) By exposing the target double-stranded nucleic acid to a crosslinking agent, an interstrand crosslink is formed between the first and second strands of the target double-stranded nucleic acid; thereby forming at least one crosslinked double-stranded nucleic acid construct; and
[0015] (b) Sequencing of the cross-linked construct formed in (a) using single-molecule sequencing technology.
[0016] In some embodiments, the method further includes allowing the interstrand crosslinks formed in step (a) or near them to break, such that the interstrand crosslinks are located at or near the end of the construct.
[0017] In some embodiments, breakage at or near the inter-chain crosslinks occurs naturally due to deformation. In other embodiments, breakage at or near the inter-chain crosslinks is caused by physical agitation.
[0018] In some embodiments, single-molecule sequencing technology includes contacting a cross-linked construct with an enzyme that sequentially processes a first and second strand of the resulting cross-linked construct. In some embodiments, the sequencing technology provides orthogonal alignment sequence information.
[0019] In some embodiments, the cross-linking agent is or includes electromagnetic radiation. In some embodiments, the cross-linking agent is or includes a chemical reagent. In some embodiments, the cross-linking agent is or includes a nucleic acid cross-linking enzyme.
[0020] In some embodiments, interstrand crosslinks can form at random locations on the target double-stranded nucleic acid. In some embodiments, interstrand crosslinks target specific base compositions on the target double-stranded nucleic acid.
[0021] In some embodiments, interchain crosslinks are formed between nucleobases in the first and second strands of the target double-stranded nucleic acid. In some embodiments, interchain crosslinks are formed between glycogroups in the first and second strands of the target double-stranded nucleic acid. In some embodiments, interchain crosslinks are formed between nucleobases in the first and second strands of the target double-stranded nucleic acid and between glycogroups in the first and second strands of the target double-stranded nucleic acid. In some embodiments, interchain crosslinks are formed between nucleobase adducts in the first and second strands of the target double-stranded nucleic acid.
[0022] In some embodiments, the crosslinking is located within 100 bases at the end of the crosslinked construct formed in part (b) of claim 1.
[0023] In some implementations, single-molecule sequencing technology includes sequencing using nanopore sensors.
[0024] This article also provides cross-linked nucleic acid constructs formed by exposing double-stranded nucleic acids to cross-linking agents as described herein.
[0025] Systems comprising cross-linked nucleic acid constructs formed by exposing double-stranded nucleic acids to cross-linking agents as described herein; one or more polynucleotide processing enzymes as described herein; and one or more nanopores.
[0026] Kits containing one or more cross-linking agents and one or more polynucleotide processing enzymes are also provided. Attached Figure Description
[0027] Figure 1Schematic diagrams illustrating the formation of interstrand crosslinks in double-stranded nucleic acids, followed by optional breaking of the interstranded double-stranded nucleic acids to form at least one interstranded double-stranded construct. A: One interstrand crosslink can be formed in a double-stranded nucleic acid, allowing both strands of the nucleic acid to break at or near the interstrand crosslink to form an interstranded construct and a non-crosslinked double-stranded polynucleotide. B: One interstrand crosslink can be formed in a double-stranded nucleic acid, allowing one strand of the nucleic acid to break at or near the interstrand crosslink to form an interstranded construct and a single-stranded polynucleotide. C: Multiple interstrand crosslinks can be formed in a double-stranded nucleic acid (two are shown), allowing both strands of the nucleic acid to break between two of the multiple crosslinked strands to form two interstranded constructs. D: Multiple interstrand crosslinks can be formed in a double-stranded nucleic acid (two are shown), allowing both strands of the nucleic acid to break at or near one of the multiple interstrand crosslinks to form an interstranded construct and a non-crosslinked double-stranded polynucleotide.
[0028] Figure 2 A schematic diagram of cross-linking and preparation of double-stranded nucleic acids for single-molecule sequencing. Step (a): A double-stranded polynucleotide (e.g., dsDNA) is exposed to conditions that create interstrand cross-links between the two strands of the double-stranded substrate (e.g., by covalently cross-linking adjacent bases of the strands). Step (b): A cross-linked substrate can be prepared for sequencing, such as nanopore sequencing, by breaking the strands at or near the cross-linking site. Step (c): Sequencing adaptors can be attached to facilitate sequencing of the cross-linked construct (e.g., by ligation or transposition). As a non-limiting example, the diagram illustrates the preparation of a cross-linked substrate using a dA-tailing method, leaving a 3' dA-tail, to which a complementary sequencing adaptor can be attached via ligation as shown in the diagram.
[0029] Figure 3 This illustrates the use of nanoporous systems for, for example Figure 2 A schematic diagram of a method for sequencing the interstranded cross-linked construct obtained as described herein. As a non-limiting example, it illustrates the capture and translocation cross-linking of double-stranded polynucleotides through nanopores in a membrane, with movement controlled by a motor protein. Step (A): The substrate is captured in the nanopore from the cis side (e.g., under a transmembrane voltage force) until the motor enzyme contacts the top of the nanopore. Step (B): The motor protein passes through the stagnation region of the sequencing adaptor and begins to control the movement of the polynucleotide through the nanopore from the cis side to the trans side. The motor protein first passes through the template portion of the double-stranded molecule (“A portion”), then [step (C)] continues by controlling the movement of the second inverse complementary portion (“B portion”) through the nanopore via the cross-linking connecting the two strands.
[0030] Figure 4This diagram illustrates how a double-stranded nucleic acid (e.g., double-stranded DNA) crosslinks at multiple sites and is then used to prepare repeating circular synthetic products. Step (a): The double-stranded nucleic acid (e.g., a double-stranded DNA strand) crosslinks at two sites, creating a “circular” region defined by the two interstrand crosslinks. Step (b): The substrate can be processed in various ways, such as unwinding and invading the primer (strands with arrows). Step (c): Polymerase-mediated strand replication (dashed lines) can be initiated. The polymerase cannot remain on the same strand through the interstrand crosslink sites, so when it encounters an interstrand crosslink, the polymerase switches to the adjacent strand (the synthesized strand is therefore complementary to the adjacent strand) and continues replication around the “circular” region, as shown. The synthesized strand product can then be processed and sequenced, for example, by attaching a sequencing adaptor and sequencing on a nanopore system.
[0031] Figure 5 Exemplary current (y-axis) versus time (x-axis) data from a single nanopore sequencing of a λDNA sample using baseline sequencing parameters on an Oxford Nanopore MinION nanopore sequencer. Two typical data portions from a single experiment are shown. The sample was cross-linked prior to sequencing as described in the examples. The nanopore has an opening current of approximately 250 pA, which is blocked to a lower current when capturing the DNA analyte. Short-lived events originate from free (unbound) sequencing adaptors in the sample, while long-lived events in DNA sequencing are characterized by portions A and B, corresponding to the template (A) of the cross-linked dsDNA substrate and the linked reverse complementary portion (B). The data described above are discussed in the examples.
[0032] Figure 6 Exemplary current (y-axis) versus time (x-axis) data (left column) obtained from single-nanopore sequencing of cross-linked λ DNA using baseline sequencing parameters on an Oxford Nanopore MinION nanopore sequencer. The left side shows the current versus time signals, and the right side shows the alignment of base recognitions of these strands aligned to a λ phage genome reference. The above data are discussed in the examples. Detailed Implementation
[0033] This invention will be described with reference to specific embodiments and certain accompanying drawings, but is not limited thereto by the claims. Any reference numerals in the claims should not be construed as limiting the scope. It should be understood, of course, that not all aspects or advantages can be achieved according to any particular embodiment of the invention. Therefore, for example, those skilled in the art will recognize that the invention may be embodied or practiced in a manner that achieves or optimizes one or more advantages as taught herein, without necessarily achieving other aspects or advantages as may be taught or suggested herein.
[0034] The invention (both in terms of organization and operation) and its features and advantages can be best understood by referring to the following detailed description when read in conjunction with the accompanying drawings. Aspects and advantages of the invention will become apparent from one or more embodiments described below, and will be set forth with reference to said embodiments. Throughout this specification, reference to “one embodiment” or “embodiment” means that a particular feature, structure, or characteristic described in connection with an embodiment is included in at least one embodiment of the invention. Therefore, the phrases “in one embodiment” or “in an embodiment” appearing in various places throughout this specification do not necessarily refer to the same embodiment, but may refer to the same embodiment. Similarly, it should be understood that in the description of exemplary embodiments of the invention, various features of the invention are sometimes grouped together in a single embodiment, drawing, or description thereof for the purpose of simplifying this disclosure and aiding in the understanding of one or more of the various inventive aspects. However, the method of this disclosure should not be construed as reflecting an intention to reflect more features required by the claimed invention than expressly recited in each claim. Rather, as reflected in the following claims, the inventive aspect lies in fewer features than all of the features of a single foregoing disclosed embodiment.
[0035] It should be understood that, unless the context otherwise requires, the “implementations” of this disclosure can be specifically combined together. Specific combinations of all disclosed embodiments (unless the context otherwise implies) constitute further disclosed embodiments of the claimed invention.
[0036] Furthermore, as used in this specification and the appended claims, unless otherwise expressly indicated, the singular forms "a / an" and "the" both encompass the plural objects. Thus, for example, a reference to "polynucleotide" includes two or more polynucleotides; a reference to "motor protein" includes two or more such proteins; a reference to "helicase" includes two or more helicases; a reference to "monomer" refers to two or more monomers; a reference to "pore" includes two or more pores, etc.
[0037] All publications, patents, and patent applications cited in this article, whether mentioned above or below, are incorporated herein by reference in their entirety.
[0038] definition
[0039] When referring to singular nouns (e.g., "a / an," "the"), the use of indefinite or definite articles includes the plural form of the noun unless specifically stated otherwise. The use of the term "comprising" in this specification and claims does not exclude other elements or steps. Furthermore, the terms first, second, third, etc., in the specification and claims are used to distinguish similar elements and are not necessarily used to describe order or chronological sequence. It should be understood that the terms thus used are interchangeable where appropriate, and embodiments of the invention described herein can operate in orders other than those described or illustrated herein. The following terms or definitions are provided only to aid in understanding the invention. Unless specifically defined herein, all terms used herein have the same meaning to those skilled in the art to which this invention pertains. For the definitions and terminology used in this field, practicing physicians have specifically referred to Sambrook et al., Molecular Cloning: A Laboratory Manual, 4th Edition, Cold Spring Harbor Press, Plainsview, New York (2012); and Ausubel et al., Current Protocols in Molecular Biology (Supplement 114), John Wiley & Sons, New York (2016). The definitions provided herein should not be construed as having a scope less than that understood by one of ordinary skill in the art.
[0040] When referring to measurable values such as quantity or duration, the term “about” as used herein means to encompass deviations from the specified value of ±20% or ±10%, more preferably ±5%, even more preferably ±1%, and still more preferably ±0.1%, as such deviations are suitable for performing the disclosed method.
[0041] As used herein, the terms “nucleotide sequence,” “DNA sequence,” or “one or more nucleic acid molecules” refer to a polymer of nucleotides of any length, whether ribonucleotides or deoxyribonucleotides. This term refers only to the primary structure of the molecule. Therefore, this term encompasses both double-stranded and single-stranded DNA, as well as RNA. As used herein, the term “nucleic acid” is a single-stranded or double-stranded covalently linked sequence of nucleotides, wherein the 3' and 5' ends of each nucleotide are linked by a phosphodiester bond. Polynucleotides may consist of deoxyribonucleotide bases or ribonucleotide bases. Nucleic acids can be synthesized in vitro or isolated from natural sources. Nucleic acids may further comprise modified DNA or RNA, such as methylated DNA or RNA, or RNA that has undergone post-translational modifications, such as 5' capping with 7-methylguanosine, 3' processing such as cleavage and polyadenylation, and splicing. Nucleic acids can also include synthetic nucleic acids (XNAs), such as hexitol nucleic acid (HNA), cyclohexene nucleic acid (CeNA), threonine nucleic acid (TNA), glycerol nucleic acid (GNA), locked nucleic acid (LNA), and peptide nucleic acid (PNA). The size of a nucleic acid (also referred to herein as a “polynucleotide”) is typically expressed as the number of base pairs (bp) in a double-stranded polynucleotide, or, in the case of a single-stranded polynucleotide, as the number of nucleotides (nt). One thousand bp or nt equals one thousand bases (kb). Polynucleotides shorter than approximately 40 nucleotides are often referred to as “oligonucleotides” and may include primers used for manipulating DNA, such as via polymerase chain reaction (PCR).
[0042] In the context of this disclosure, the term "amino acid" is used in its broadest sense and refers to organic compounds containing amine (NH2) and carboxyl (COOH) functional groups, as well as side chains (e.g., R groups) specific to each amino acid. In some embodiments, amino acid refers to naturally occurring Lα-amino acids or residues. One and three commonly used letter abbreviations for naturally occurring amino acids are used herein: A = Ala; C = Cys; D = Asp; E = Glu; F = Phe; G = Gly; H = His; I = Ile; K = Lys; L = Leu; M = Met; N = Asn; P = Pro; Q =
[0043] Gln; R = Arg; S = Ser; T = Thr; V = Val; W = Trp; and Y = Tyr (Lehninger, AL, (1975) Biochemistry, 2nd ed., pp. 71-92, Worth Publishers, New York). The general term “amino acid” further includes D-amino acids, trans-amino acids, and chemically modified amino acids (such as amino acid analogs), naturally occurring amino acids that are not typically incorporated into proteins (such as ortholeucine), and chemically synthesized compounds (such as β-amino acids) that have properties known in the art as characteristics of amino acids. For example, analogs or mimics of phenylalanine or proline that allow conformational restrictions to be the same as those of natural Phe or Pro are included within the definition of amino acids. Such analogs and mimics are referred to herein as “functional equivalents” of the corresponding amino acids. Other examples of amino acids are listed by Roberts and Vellaccio, The Peptides: Analysis, Synthesis, Biology, edited by Gross and Meiehofer, Vol. 5, p. 341, Academic Press, Inc., NY, 1983, which is incorporated herein by reference.
[0044] The terms “polypeptide” and “peptide” are used interchangeably herein to refer to polymers containing amino acid residues, as well as their variants and synthetic analogs. Therefore, these terms apply to amino acid polymers where one or more amino acid residues are synthetic, non-naturally occurring amino acids, such as chemical analogs of corresponding naturally occurring amino acids, and to polymers containing naturally occurring amino acids. Polypeptides may also undergo maturation or post-translational modification processes, which may include, but are not limited to, glycosylation, proteolytic cleavage, lipolysis, signal peptide cleavage, propeptide cleavage, phosphorylation, etc. Peptides can be prepared using recombinant techniques, for example, by expressing recombinant or synthetic polynucleotides. Recombinant peptides are typically substantially free of culture medium, for example, the culture medium comprises less than about 20% of the volume of the protein formulation, more preferably less than about 10%, and most preferably less than about 5%.
[0045] The term "protein" is used to describe folded polypeptides that have secondary or tertiary structures. Proteins can consist of a single polypeptide or can comprise multiple polypeptides that assemble to form a multimer. The multimer can be a homooligomer or a heterooligomer. Proteins can be naturally occurring or wild-type proteins, or modified or non-natural proteins. Proteins can differ from wild-type proteins, for example, through the addition, substitution, or deletion of one or more amino acids.
[0046] Protein “variants” encompass peptides, oligopeptides, polypeptides, proteins, and enzymes that have amino acid substitutions, deletions, and / or insertions relative to the unmodified or wild-type protein in question, and possess biological and functional activities similar to those of the unmodified protein from which they are derived. As used herein, the term “amino acid identity” refers to the degree to which sequences are identical on an amino acid-to-amino acid basis within a comparison window. Thus, the “sequence identity percentage” is calculated by comparing two optimally aligned sequences within a comparison window, determining the number of positions in which identical amino acid residues (e.g., Ala, Pro, Ser, Thr, Gly, Val, Leu, Ile, Phe, Tyr, Trp, Lys, Arg, His, Asp, Glu, Asn, Gln, Cys, and Met) appear in both sequences to produce the number of matching positions, dividing the number of matching positions by the total number of positions in the comparison window (i.e., the window size), and multiplying the result by 100 to produce the sequence identity percentage.
[0047] For all aspects and embodiments of the invention, the "variant" has at least 50%, 60%, 70%, 80%, 90%, 95%, or 99% complete sequence identity with the corresponding wild-type protein's amino acid sequence. Sequence identity can also be a fragment or portion of a full-length polynucleotide or polypeptide. Thus, a sequence may have only 50% overall sequence identity with a full-length reference sequence, but sequences of specific regions, domains, or subunits may share 80%, 90%, or up to 99% sequence identity with the reference sequence.
[0048] The term "wild-type" refers to a gene or gene product isolated from a naturally occurring source. Wild-type genes are the most frequently observed genes in a population and are therefore arbitrarily engineered to be in their "normal" or "wild-type" form. Conversely, the terms "modified," "mutant," or "variant" refer to a gene or gene product that exhibits sequence modifications (e.g., substitution, truncation, or insertion), post-translational modifications, and / or functional characteristics (e.g., altered properties) compared to a wild-type gene or gene product. Note that naturally occurring mutants can be isolated; these mutants are identified by the fact that they possess altered properties compared to a wild-type gene or gene product. Methods for introducing or substituting naturally occurring amino acids are well known in the art. For example, this can be achieved by substituting the codon of methionine (ATG) with the codon of arginine (CGT) at the relevant position in the polynucleotide encoding the mutant monomer, and by substituting methionine (M) with arginine (R). Methods for introducing or substituting non-naturally occurring amino acids are also well known in the art. For example, non-naturally occurring amino acids can be introduced by including synthetic aminoacyl-tRNA in the IVTT system used to express mutant monomers. Alternatively, this can be introduced by expressing mutant monomers in *E. coli* that are auxotrophic for the specific amino acid in the presence of synthetic (i.e., non-naturally occurring) analogs of those specific amino acids. If the mutant monomer is produced using partial peptide synthesis, it can also be produced via naked ligation. Conservative substitution replaces an amino acid with another amino acid having a similar chemical structure, similar chemical properties, or similar side chain volume. The introduced amino acid can have similar polarity, hydrophilicity, hydrophobicity, basicity, acidity, neutrality, or charge as the substituted amino acid. Alternatively, conservative substitution can introduce another aromatic or aliphatic amino acid to replace a pre-existing aromatic or aliphatic amino acid. Conservative amino acid changes are well known in the art and can be selected based on the properties of the 20 major amino acids defined in Table 1 below. In the case of amino acids having similar polarity, this can also be determined with reference to the hydrophilicity scale of the amino acid side chains in Table 2.
[0049] Table 1 - Chemical properties of amino acids
[0050]
[0051] Table 2 - Hydrophilicity Scale
[0052]
[0053] Mutants or modified proteins, monomers, or peptides can also be chemically modified in any manner and at any site. Preferably, mutants or modified monomers are chemically modified by attaching the molecule to one or more cysteine residues (cysteine linkage), attaching the molecule to one or more lysine residues, attaching the molecule to one or more non-natural amino acids, enzymatic modification of epitopes, or terminal modification. Suitable methods for performing such modifications are well known in the art. Mutants of modified proteins, monomers, or peptides can be chemically modified by attaching any molecule. For example, mutants of modified proteins, monomers, or peptides can be chemically modified by linking dyes or fluorophores.
[0054] The disclosed method
[0055] This disclosure relates to a method for sequencing a target double-stranded nucleic acid by forming an interstrand crosslink between the first and second strands of the target double-stranded nucleic acid and sequencing the formed crosslinked construct.
[0056] The inventors have surprisingly discovered that a polynucleotide processing enzyme (also referred to in some embodiments as a motor protein) can sequentially process the first and second strands of a double-stranded nucleic acid when the first and second strands are linked by interstrand crosslinks. The inventors have also surprisingly discovered that it is not necessary to link the 5' and 3' ends of the double-stranded nucleic acid together. Furthermore, the inventors have surprisingly discovered that the polynucleotide processing enzyme can process from the first strand of the double-stranded polynucleotide to the second strand at the interstrand crosslink site, rather than simply through a polynucleotide bridge provided, for example, by a hairpin adaptor. The method provided herein utilizes this capability. For example, by sequencing a crosslinked construct obtained by forming an interstrand crosslink between the first and second strands of a target double-stranded nucleic acid, the sequential processing of the first and second strands can provide orthogonal proofreading sequence information. The polynucleotide processing enzyme is described in more detail herein.
[0057] Therefore, this paper provides a method for sequencing target double-stranded nucleic acids, which includes:
[0058] (a) By exposing the target double-stranded nucleic acid to a crosslinking agent, an interstrand crosslink is formed between the first and second strands of the target double-stranded nucleic acid; thereby forming at least one crosslinked double-stranded nucleic acid construct; and
[0059] (b) Sequencing of the cross-linked construct formed in (a) using single-molecule sequencing technology.
[0060] Interchain crosslinks between the first and second chains can be formed by any suitable method. Some exemplary methods are described in more detail herein.
[0061] Typically, in the disclosed methods, interstrand crosslinks are not connections between the 5' and 3' ends of the double-stranded nucleic acid. For example, as used herein, interstrand crosslinks are not the result of using hairpin adaptors. While hairpin adaptors do connect the first and second strands of a double-stranded nucleic acid, they do not form interstrand crosslinks as used herein.
[0062] Interstrand crosslinks can form at any suitable location on the double-stranded nucleic acid strand. For example, interstrand crosslinks can form at random locations on the double-stranded nucleic acid. Alternatively, interstrand crosslinks can target a specific base composition on the double-stranded nucleic acid. In other words, in some embodiments, the disclosed method may include forming interstrand crosslinks at random locations. In other embodiments, the disclosed method includes determining the location of interstrand crosslinks using the base composition of the double-stranded nucleic acid. This base composition may be selected, determined, or controlled to be a specific composition for targeting interstrand crosslinks.
[0063] In some implementations, the base composition of the double-stranded nucleic acid can be determined by introducing a target sequence into the double-stranded nucleic acid. Upon contact with a cross-linking agent, the target sequence can be pre-determined to facilitate interstrand cross-linking. Any suitable technique can be used to introduce the target sequence into the double-stranded nucleic acid. For example, CRISPR can be used, such as by using CRISPR-Cas9 genome editing to introduce the target sequence into the genome. Conventional cloning techniques such as those described in Sambrook Molecular Cloning: A Laboratory Manual, 4th Edition, Cold Spring Harbor Press, Plainsview, New York (2012); and Ausubel et al., Current Protocols in Molecular Biology (Supplement 114), John Wiley & Sons, New York (2016) can be used to introduce the target sequence into the double-stranded nucleic acid. Other suitable techniques will be apparent to those skilled in the art.
[0064] In some embodiments, the method includes attaching one or more adaptors to a target double-stranded nucleic acid (e.g., by ligation). In such embodiments, the disclosed method may include forming interstrand crosslinks in the adaptors.
[0065] Those skilled in the art will understand that the location of interstrand crosslinks determines the obtained sequencing results. Therefore, the location of interstrand crosslinks can be controlled to provide the desired information. For example, if it is desired to sequence only the starting portion of a long double-stranded nucleic acid, an interstrand crosslink can be introduced downstream of that starting portion. The disclosed method will allow selective sequencing of that portion of the double-stranded nucleic acid. On the other hand, if it is desired to sequence the entire double-stranded nucleic acid, an interstrand crosslink can be introduced at the ends of the double-stranded nucleic acid. The disclosed method will then allow selective sequencing of the entire double-stranded nucleic acid.
[0066] The random location of interstrand crosslinks can also provide useful information. For example, populations of double-stranded nucleic acids can undergo interstrand crosslinking reaction conditions. The characteristic distribution of interstrand crosslinks will result in a group of polynucleotide sequences of a specific length. Therefore, this length distribution may be characteristic of a population of double-stranded nucleic acids. In this way, different groups can be compared. Alternatively, the prevalence of interstrand crosslinks can provide information about the sequence of double-stranded nucleic acids before the formation of interstrand crosslinks.
[0067] In another implementation, the random location of interstrand crosslinks can be used for genome sequencing. The incorporation of random interstrand crosslinks (e.g., at low levels in double-stranded nucleic acids) can allow for an equal balance of reads (i.e., sequencing information) from the entire genome.
[0068] These and other advantages will be obvious to those skilled in the art.
[0069] As discussed in more detail below, some embodiments of the disclosed methods include allowing the target double-stranded nucleic acid to break at or near the interstrand crosslinking site. As described in more detail herein, the breakage can occur naturally as a result of interstrand crosslinking. For example, in some embodiments, interstrand crosslinking results in distortion of the target double-stranded nucleic acid structure, and this distortion causes the target double-stranded nucleic acid to break at or near the interstrand crosslinking site. In other embodiments, the methods disclosed herein may include causing the target double-stranded nucleic acid to break at or near the interstrand crosslinking site. For example, in some embodiments, the method includes physically agitating the crosslinked double-stranded nucleic acid to cause the double-stranded nucleic acid to break at or near the interstrand crosslinking site. As described in more detail herein, breaking the target double-stranded nucleic acid at or near the interstrand crosslinking site forms a double-stranded nucleic acid construct crosslinked at or near the ends of the construct.
[0070] As described in more detail herein, the disclosed methods include sequencing cross-linked double-stranded nucleic acid constructs using single-molecule sequencing technology. Any suitable sequencing technology can be used. Exemplary suitable sequencing technologies are discussed in more detail herein. For example, in some preferred embodiments, the sequencing technology is a nanopore sensing method. Nanopore sensing methods are described in detail herein. However, the methods disclosed herein are not limited to nanopore sensing. Other single-molecule sequencing technologies are applicable to the methods disclosed herein.
[0071] In some embodiments, the disclosed method includes contacting a cross-linked double-stranded nucleic acid construct (“cross-linked construct”) with an enzyme that sequentially processes a first and second strand of the cross-linked construct. For example, in some embodiments, the enzyme may process the first strand and then proceed to process the second strand. In some embodiments, the enzyme may process the first strand, perform inter-strand crosslinking, and then process the second strand. In some embodiments, the enzyme may process from the first strand to the second strand via inter-strand crosslinking. In some embodiments, the enzyme may process the first and second strands sequentially without inter-strand crosslinking. Inter-strand crosslinking holds the first and second strands together, allowing the enzyme to process the first and second strands sequentially. Polynucleotide processing enzymes are described in more detail below. In some embodiments, the method of contacting a double-stranded nucleic acid construct with an enzyme that sequentially processes the first and second strands of the cross-linked construct can be used to provide additional information about the construct, such as orthogonal alignment sequence information.
[0072] The methods disclosed herein are advantageous because they provide additional information compared to conventional methods. As mentioned above, conventional methods sequence only one strand of a double-stranded nucleic acid. If both strands need to be sequenced, an adaptor, such as a hairpin adaptor, is typically required. The methods disclosed herein do not rely on the use of hairpin adaptors and can therefore be performed without the disadvantages associated with using hairpin adaptors, such as increased complexity, the possibility of contamination, and the time required for adaptor ligation to the target double-stranded nucleic acid.
[0073] Interchain crosslinking
[0074] As used herein, the term interstrand crosslink refers to the chemical bond between the first and second strands of a double-stranded nucleic acid. This chemical bond can be, for example, a covalent bond or an ionic bond. Typically, the term interstrand crosslink refers to a covalent connection between the first and second strands of a double-stranded nucleic acid. Therefore, this connection can be located between the template strand and the complementary strand of the double-stranded nucleic acid. Double-stranded nucleic acids are typically double-stranded DNA.
[0075] Typically, the interchain crosslinks formed in the methods disclosed herein are irreversible. However, in some embodiments, the interchain crosslinks are formed reversibly. If desired, reversible interchain crosslinks can be used to reverse the linking process.
[0076] In some embodiments, interchain crosslinks are formed between nucleobases in the first and second strands of the double-stranded nucleic acid. In some embodiments, interchain crosslinks are formed between nucleobases on the first strand selected from adenine (A), cytosine (C), guanine (G), and thymine (T); and between nucleobases on the second strand selected from adenine (A), cytosine (C), guanine (G), and thymine (T). In some embodiments, interchain crosslinks are formed between guanine residues on the first strand and guanine residues on the second strand. In some embodiments, interchain crosslinks are formed between guanine residues on the first strand and cytosine residues on the second strand, or between cytosine residues on the first strand and guanine residues on the second strand. In some embodiments, interchain crosslinks are formed between thymine residues on the first strand and thymine residues on the second strand.
[0077] In some embodiments, interstrand crosslinks are formed between sugar groups in the first and second strands of the double-stranded nucleic acid. For example, interstrand crosslinks can be formed between deoxyribose groups in the first and second strands.
[0078] In some embodiments, interstrand crosslinks are formed between nucleobases in the first and second strands of the double-stranded nucleic acid and between sugar groups in the first and second strands of the double-stranded nucleic acid.
[0079] In some embodiments, interchain crosslinks are formed between the nucleobases of the first strand of the double-stranded nucleic acid and the sugar groups of the second strand. In some embodiments, interchain crosslinks are formed between the backbone of the first strand of the double-stranded nucleic acid and the nucleobases of the second strand. In some embodiments, interchain crosslinks are formed between the backbone of the first strand of the double-stranded nucleic acid and the sugar groups of the second strand. In some embodiments, interchain crosslinks are formed between the backbone of the first strand of the double-stranded nucleic acid and the backbone of the second strand. In the methods disclosed herein, interchain crosslinks forming between the nucleobases of the first and second strands of the double-stranded nucleic acid are most common.
[0080] In the methods disclosed herein, inter-strand crosslinks are formed by contacting double-stranded nucleic acids with a crosslinking agent. In some embodiments, the crosslinking agent is a crosslinking condition, such as exposure to electromagnetic radiation. As used herein, the term "crosslinking agent" includes conditions that promote the formation of inter-strand crosslinks, such as exposure to electromagnetic radiation, like UV light.
[0081] Any suitable crosslinking agent can be used in the methods disclosed herein. For example, in some embodiments, the crosslinking agent is or includes electromagnetic radiation (e.g., the crosslinking agent is an optical crosslinking agent). In some embodiments, the crosslinking agent is or includes a chemical reagent. Many suitable chemical reagents are known to those skilled in the art. In some embodiments, the chemical reagent is a polynucleotide interacting molecule. In some embodiments, the crosslinking agent is or includes a polynucleotide interacting enzyme. In some embodiments, the crosslinking agent is or includes a nucleic acid crosslinking enzyme. The disclosed methods may include the use of a variety of such crosslinking agents, for example, combinations of such crosslinking agents may be used. For example, in some embodiments, the formation of interchain crosslinks includes the use of electromagnetic radiation and a chemical reagent. In some embodiments, the formation of interchain crosslinks includes the use of a chemical reagent and a nucleic acid crosslinking enzyme.
[0082] In some embodiments, crosslinking agents are selected or determined to promote interchain crosslinking rather than the formation of unwanted reactions (e.g., chain breakage or cleavage, base damage, intrachain crosslinking, etc.). In some embodiments, the location and / or extent of interchain crosslinking is controlled by using suitable crosslinking agents and / or by controlling reaction conditions (e.g., the concentration of the interchain crosslinking agent used).
[0083] In some embodiments, the crosslinking agent is or includes a clastogenic chemical reagent. In some embodiments, the crosslinking agent is a chemical reagent that is a bifunctional reactive group.
[0084] In some embodiments, the chemical reagent is a bifunctional alkylating agent. Bifunctional alkylating agents include nitrogen mustard and lipid peroxidation products.
[0085] Nitrogen mustard is a bifunctional alkylating agent. Nitrogen mustard typically contains a reactive N,N-bis(2-chloroethyl)amine functional group with a variable R group. Nitrogen mustard usually reacts with the N7 position of guanine.
[0086] Non-limiting examples of nitrogen mustards include chlorambucil, dichloromethyldiethylamine, and phosphoramide mustard. Some nitrogen mustards are shown below:
[0087] (Where R = alkyl, alkenyl, alkynyl, aryl, aralkyl, heteroaryl, heteroarylalkyl, phosphate ester, etc.); for example...
[0088]
[0089] In some embodiments, the chemical reagent is nitrogen mustard, and interstrand crosslinks are formed between guanine residues on the first and second strands of the double-stranded nucleic acid. In some embodiments, the guanine residues are contained in the 5'-GNC-3' sequence. In some embodiments, the formation of such interstrand crosslinks results in a twisting of the helix formed by the double-stranded nucleic acid in the crosslinked regions.
[0090] In some embodiments, the chemical reagent is mitomycin C or an analogue thereof. Mitomycin C is a bifunctional alkylating agent. Mitomycin C is shown below:
[0091]
[0092] In some embodiments, the chemical agent is mitomycin C or an analogue thereof, and interchain crosslinks are formed between guanine residues on the first and second strands of the double-stranded nucleic acid. In some embodiments, the guanine residues are contained in the 5'-CG-3' sequence. In some embodiments, the formation of such interchain crosslinks results in a twisting of the helix formed by the double-stranded nucleic acid in the crosslinked regions.
[0093] In some embodiments, the chemical reagent is a lipid peroxidation product (typically an aldehyde). Non-limiting examples of lipid peroxidation products include acrolein, crotonaldehyde, and malondialdehyde. Other aldehydes that may be used include trans-4-hydroxynonenal, acetaldehyde, and formaldehyde. In some embodiments, the chemical reagent is a lipid peroxidation product or other aldehyde, and interchain crosslinks are formed between guanine residues on the first and second strands of the double-stranded nucleic acid. In some embodiments, the guanine residues are contained in a 5'-CG-3' or 5'-GC-3' sequence. In some embodiments, the formation of such interchain crosslinks results in a twisting of the helix formed by the double-stranded nucleic acid in the crosslinked regions.
[0094] In some embodiments, the chemical reagent is a platinum compound. In some embodiments, the platinum compound is cisplatin or a derivative thereof. In some embodiments, the chemical reagent is cis-diammineplatinum dichloride. In some embodiments, the chemical reagent is cis-diammineplatinum dichloride, and interchain crosslinks are formed between guanine residues on the first and second strands of the double-stranded nucleic acid. In some embodiments, the guanine residues are contained in the 5'-GC-3' sequence. In some embodiments, the formation of such interchain crosslinks results in a twisting of the helix formed by the double-stranded nucleic acid in the crosslinked region. In some embodiments, the chemical reagent is trans-diammineplatinum dichloride. In some embodiments, the chemical reagent is trans-diammineplatinum dichloride, and interchain crosslinks are formed between guanine residues on the first strand of the double-stranded nucleic acid and cytosine residues on the second strand. Typically, the guanine residues on the first strand and the cytosine residues on the second strand are base-paired before the formation of the interchain crosslinks. In some embodiments, the use of a platinum compound to form interchain crosslinks results in a twisting of the helix formed by the double-stranded nucleic acid in the crosslinked region.
[0095] In some embodiments, the chemical reagent is chloroethylnitrosourea, for example, carmustine or an analogue thereof. Carmustine is shown below:
[0096]
[0097] Carmustine is typically produced by alkylating the O of the guanine group. 6 Position formation O 6 -Vinylguanine (O 6 -ethanoguanine), and then crosslinked to the base-paired cytosine. Therefore, in some embodiments, the chemical agent is carmustine or an analogue, and the interstrand crosslink is formed between guanine residues on the first strand of the double-stranded nucleic acid and cytosine residues on the second strand. Typically, the guanine residues on the first strand and the cytosine residues on the second strand are base-paired before the interstrand crosslink is formed.
[0098] In some embodiments, the chemical reagent is psoralen or a derivative thereof, such as methoxypsoralen. Psoralen and methoxypsoralen are shown below:
[0099]
[0100] For example, when activated in the presence of light, such as under ultraviolet-A (UV-A) radiation, typical psoralen inserts into DNA and may form covalent interchain crosslinks. Such adducts are typically formed by linking the 3',4' (pyranone) or 4',5' (furan) edges of psoralen to the 5',6' double bond of thymine. Therefore, in some embodiments, the chemical agent is psoralen or a derivative thereof, and the interchain crosslinks are formed between thymine residues on the first and second strands of the double-stranded nucleic acid. In some embodiments, the thymine residues are contained in the 5'-AT-3' sequence. In some embodiments, the crosslinking agent comprises psoralen and electromagnetic radiation, such as UV radiation, e.g., UV-A radiation. In some embodiments, the use of psoralen to form interchain crosslinks results in a twisting of the helix formed by the double-stranded nucleic acid in the crosslinked region.
[0101] In some embodiments, the chemical reagent is nitrous acid. Nitrous acid can form interstrand crosslinks in double-stranded nucleic acids by converting amino groups in DNA to carbonyl groups. In some embodiments, the chemical reagent is nitrous acid and the interstrand crosslinks are formed between guanosine residues. In some embodiments, the guanosine residues are contained in the 5'-CG-3' sequence. In some embodiments, the use of nitrous acid to form interstrand crosslinks results in twisting of the helix formed by the double-stranded nucleic acids in the crosslinked regions.
[0102] In some embodiments, the chemical reagent is an insert dye. Insert dyes are well known in the art.
[0103] In some embodiments, the chemical reagent is one used to promote the formation of free radicals. Such chemical reagents include metals that react with peroxides to form or promote cross-linking between free radical chains.
[0104] In some embodiments, the cross-linking agent comprises a reagent that reacts with both the first and second strands of a double-stranded nucleic acid to enable the reaction product to further react, thereby generating interstrand cross-links. Such reagents are known in the art.
[0105] In some embodiments, the chemical reagent is a click chemistry reagent. This reagent can promote or participate in click chemistry reactions between nucleobases on the first and second strands of a double-stranded nucleic acid. Click chemistry reactions can form interchain crosslinks between side groups on the nucleobases of the first and second strands of the double-stranded nucleic acid. Suitable side groups can be introduced into the first and / or second strands of the double-stranded nucleic acid by any suitable means, such as by polymerase incorporation or by using one or more nucleic acid-modifying enzymes. For example, methyltransferases can be used to modify nucleic acids (e.g., DNA) with reactive groups, which can then be linked together to form interchain crosslinks. Therefore, interchain crosslinks can be formed between nucleobase adducts. Thus, some embodiments include forming interchain crosslinks between nucleobase adducts in the first and second strands of the target double-stranded nucleic acid.
[0106] Many suitable click chemistry reagents are known in the art. Suitable examples of click chemistry include, but are not limited to, the following:
[0107] (a) Copper (I)-catalyzed azide-alkyne cycloaddition (azide-alkyne Huisgen cycloaddition);
[0108] (b) Strain-promoted azide-alkyne cycloaddition; including olefin and azide [3+2] cycloaddition; olefin and tetrazine reverse demand Diels-Alder reaction; and olefin and tetrazolium photoclick reaction;
[0109] (c) Copper-free variants of 1,3-dipolar cycloaddition reactions, wherein the azide reacts with an alkyne under strain (e.g., in a cyclooctane ring);
[0110] (d) The reaction of an oxygen nucleophile at one junction with an epoxide or aziridine reactive moiety at the other junction; and
[0111] (e) Staudinger ligation, in which the alkyne moiety can be replaced by arylphosphine.
[0112] This leads to a specific reaction with azides, resulting in amide bonds.
[0113] Any reactive group can be used to form inter-chain crosslinks. Some suitable reactive groups include [1,4-bis[3-(2-pyridyldithio)propamido]butane; 1,1,1-bis-maleimide triethylene glycol; 3,3'-dithiodipropionate di(N-hydroxysuccinimide); ethylene glycol-bis(N-hydroxysuccinimide succinate); 4,4'-diisothiocyanate-2,2'-stilbenesulfonate disodium salt; bis[2-(4-azidosalicylic acid)ethyl]disulfide; 3-(2-pyridinyldithio)propionate N-succinimide; 4-maleimidebutyrate-N-hydroxysuccinimide; iodoacetic acid N-hydroxysuccinimide; S-acetylsethioglycolic acid N-hydroxysuccinimide; azide-PEG-maleimide; and alkyne-PEG-maleimide. The reactive group can be any of those groups disclosed in WO2010 / 086602 (especially in Table 3 of that application).
[0114] In some embodiments, the crosslinking agent is or includes electromagnetic radiation. In some embodiments, the electromagnetic radiation is optical radiation. In some embodiments, the optical radiation is visible light or UV (ultraviolet) radiation. In some embodiments, the radiation has a wavelength of about 100 nm to about 600 nm, more typically about 200 nm to about 500 nm, such as about 300 nm to about 340 nm, such as about 340 nm to about 380 nm, such as about 360 nm to about 370 nm, such as about 365 nm. In some embodiments, the optical radiation is UV radiation. In some embodiments, the UV radiation is UVA radiation. In some embodiments, the UV radiation is provided by a UV laser (e.g., a gas laser, a laser diode, or a solid-state laser). In some embodiments, UV radiation is provided using a UV lamp. In some embodiments, the UV radiation is high-intensity UV radiation.
[0115] In some embodiments, light radiation is applied to a double-stranded nucleic acid sample for about 1 second to about 1 hour, for example, about 1 minute to about 20 minutes, for example, about 5 minutes to about 10 minutes, for example, about 8 minutes. Those skilled in the art will be able to readily select an appropriate exposure time based on the sample being processed and the desired degree of interstrand crosslinking.
[0116] As is apparent from the above discussion, in some embodiments, the crosslinking agent may include both chemical reagents and electromagnetic radiation. For example, as described above, a combination of UV radiation and psoralen can be used to generate interstrand crosslinks. In some embodiments, crosslinking double-stranded nucleic acids using an optical crosslinking agent (i.e., electromagnetic radiation as described above) may include the use of mediators, such as photosensitizers and / or free radical sources.
[0117] In some implementations, the cross-linking agent is a nucleic acid cross-linking enzyme. Any suitable enzyme capable of forming interstrand cross-links between the first and second strands of a double-stranded nucleic acid can be used.
[0118] Suitable enzymes are commercially available. For example, nucleic acid-modifying enzymes, including nucleic acid cross-linking enzymes, are available from New England Biolabs (NEB, USA). For example, in some embodiments, the nucleic acid cross-linking enzyme is a prokaryotic telomerase. In some embodiments, the prokaryotic telomerase is TelN (NEB catalog number M0651S). The TelN prokaryotic telomerase, derived from bacteriophage N15, cleaves dsDNA at the TelN recognition sequence (56 bp) and leaves a covalently closed end at the cleavage site.
[0119] Nucleic acid cross-linking enzymes can also function by catalyzing the formation of interstrand cross-links using chemical reagents. For example, the DNA interstrand cross-linking agent (5-(aziridin-1-yl)-4-hydroxyamino-2-nitrobenzamide) can be formed from 5-(aziridin-1-yl)-2,4-dinitrobenzamide (CB 1954) via nitroreductase.
[0120] Optionally allow double-stranded nucleic acid strand breaks
[0121] As discussed, the methods disclosed herein include forming interstrand crosslinks between the first and second strands of a double-stranded nucleic acid. In some embodiments, the method may further include allowing the double-stranded nucleic acid to break at or near the interstrand crosslink site. Allowing the double-stranded nucleic acid to break at or near the interstrand crosslink site forms at least one double-stranded nucleic acid construct that is crosslinked at or near the end of the construct.
[0122] In some implementations, double-stranded nucleic acid breaks are allowed to form a single double-stranded nucleic acid construct. In other implementations, double-stranded nucleic acid breaks are allowed to form two double-stranded nucleic acid constructs.
[0123] For example, in some embodiments, forming an interstrand crosslink includes forming a bond (interstrand crosslink) between the first and second strands of the double-stranded nucleic acid. Therefore, breaking the double-stranded nucleic acid at or near the interstrand crosslink forms a portion of the double-stranded nucleic acid that includes the interstrand crosslink between the first and second strands, as well as a portion that includes the first and second strands not connected by an interstrand crosslink. The portion containing the interstrand crosslink between the first and second strands corresponds to a crosslinked double-stranded nucleic acid construct. The portion not containing the interstrand crosslink between the first and second strands can remain a double-stranded nucleic acid or can be dissociated into a single-stranded nucleic acid.
[0124] In some implementations, forming interstrand crosslinks includes forming two or more bonds (interstrand crosslinks) between the first and second strands of a double-stranded nucleic acid.
[0125] In some embodiments, breaking the double-stranded nucleic acid at or near the interstrand crosslinking site includes breaking the double-stranded nucleic acid between two of two or more interstrand crosslinks to form two partial nucleic acid segments of a double strand, each containing an interstrand crosslink between the first and second strands. One or both of the double-stranded nucleic acid segments containing the interstrand crosslink between the first and second strands correspond to a crosslinked double-stranded nucleic acid construct.
[0126] In some embodiments, a portion of a double-stranded nucleic acid is formed by breaking the double-stranded nucleic acid at or near the interstrand crosslinks, which contains two or more bonds (interstrand crosslinks) between the first and second strands, as well as a portion in which the first and second strands are not connected by interstrand crosslinks. The portion containing the interstrand crosslinks between the first and second strands corresponds to a crosslinked double-stranded nucleic acid construct. The portion not containing interstrand crosslinks between the first and second strands may remain a double-stranded nucleic acid or may be dissociated into a single-stranded nucleic acid.
[0127] In embodiments that include forming multiple interchain crosslinks and forming interchain crosslinked portions containing at least two interchain crosslinks, the resulting constructs can be used to prepare repeatable cyclic synthetic products. In some embodiments, unwinding and intrusion primers can be used to provide initiation sites for polymerase-mediated chain replication. In some embodiments, a polynucleotide processing enzyme (e.g., polymerase) cannot continue through the interchain crosslink sites while remaining on the same strand, so that when it encounters an interchain crosslink, the polymerase switches to an adjacent strand. Thus, the synthesized strand becomes complementary to the adjacent strand. The polynucleotide processing enzyme (e.g., polymerase) can then continue replicating around the “circular” region formed between the interchain crosslinks. This process... Figure 4 This is illustrated schematically. Those skilled in the art will understand that in such embodiments, the formation of the repeating cyclic synthetic product does not necessarily allow for the breaking of the double-stranded nucleic acid. However, it is not excluded that the double-stranded nucleic acid may break outside the portion bound by the interstrand cross-links. Therefore, such embodiments are applicable to methods that include allowing for the breaking of the double-stranded nucleic acid, and also to those embodiments that do not allow for the breaking of the double-stranded nucleic acid.
[0128] In some embodiments, breaking a double-stranded nucleic acid includes breaking both strands of the double-stranded nucleic acid. In some embodiments, breaking a double-stranded nucleic acid includes breaking one strand of the double-stranded nucleic acid.
[0129] By way of non-limiting examples, some of these implementations are... Figure 1 It is shown schematically in the middle. Figure 1 A illustrates the formation of interstrand crosslinks between the first and second strands of a double-stranded nucleic acid, followed by breakage of the two strands of the double-stranded nucleic acid near the interstrand crosslinks to form a crosslinked double-stranded nucleic acid construct and a non-crosslinked portion near the ends of the construct. Figure 1 B illustrates the formation of interstrand crosslinks between the first and second strands of a double-stranded nucleic acid, followed by breakage in one strand of the double-stranded nucleic acid near the interstrand crosslinks to form a crosslinked double-stranded nucleic acid construct and a non-crosslinked (single-stranded) portion near the end of the construct. Figure 1C illustrates the formation of multiple interstrand crosslinks between the first and second strands of a double-stranded nucleic acid (two crosslinks of the multiple crosslinks are shown), followed by breakage of the two strands of the double-stranded nucleic acid between two of the multiple crosslinks to form two double-stranded nucleic acid constructs, each construct having crosslinks near the ends of the constructs. Figure 1 D illustrates the formation of multiple interstrand crosslinks between the first and second strands of a double-stranded nucleic acid (two of the multiple crosslinks are shown), followed by breakage in both strands of the double-stranded nucleic acid to form a crosslinked double-stranded nucleic acid construct and a non-crosslinked portion near the ends of the construct.
[0130] As described above, some embodiments of the methods disclosed herein include allowing the target double-stranded nucleic acid to break at or near the interstrand crosslinks. In some embodiments, the method includes breaking the target double-stranded nucleic acid at or near the interstrand crosslinks. In some embodiments, the method includes inducing the target double-stranded nucleic acid to break at or near the interstrand crosslinks.
[0131] In some embodiments, the disclosed method includes allowing the target double-stranded nucleic acid to break within about 500 nucleotides of the interstrand crosslink. In some embodiments, the disclosed method includes allowing the target double-stranded nucleic acid to break between about 1 nucleotide and about 200 nucleotides of the interstrand crosslink. In some embodiments, the method includes allowing the target double-stranded nucleic acid to break between about 5 nucleotides and about 150 nucleotides of the interstrand crosslink, for example, between about 10 nucleotides and about 100 nucleotides, for example, between about 20 nucleotides and about 80 nucleotides, for example, between about 30 nucleotides and about 50 nucleotides. In such embodiments, allowing the double-stranded nucleic acid to break at or near the interstrand crosslink forms an "overhanging region" adjacent to the interstrand crosslink. The length of the overhanging region corresponds to the number of nucleotides between the interstrand crosslink and the break in the double-stranded nucleic acid.
[0132] Therefore, in some embodiments, the crosslinks formed in the double-stranded nucleic acid are located within about 500 nucleotides at the end of the crosslinked construct formed by contacting the double-stranded nucleic acid with a crosslinking agent and subsequently breaking it. In some embodiments, the crosslinks are located within about 1 nucleotide to about 200 nucleotides at the thus formed end. In some embodiments, the crosslinks are located between about 5 nucleotides and about 150 nucleotides at the end of the crosslinked construct formed by contacting the double-stranded nucleic acid with a crosslinking agent and subsequently breaking it, for example, between about 10 nucleotides and about 100 nucleotides, for example, between about 20 nucleotides and about 80 nucleotides, for example, between about 30 nucleotides and about 50 nucleotides.
[0133] In some implementations, the method includes breaking the double-stranded nucleic acid at the site of interstrand crosslinking to prevent the formation of overhanging regions.
[0134] Some embodiments include forming interstrand crosslinks at or near the ends of the target double-stranded nucleic acid. In such embodiments, it may not be necessary to break the double-stranded nucleic acid. Therefore, in some embodiments, the crosslinks formed in the target double-stranded nucleic acid are located within about 500 nucleotides from the end of the target double-stranded nucleic acid. In some embodiments, the crosslinks are located within about 1 nucleotide to about 200 nucleotides from the end of the target double-stranded nucleic acid. In some embodiments, the crosslinks are located between about 5 nucleotides and about 150 nucleotides from the end of the target double-stranded nucleic acid, for example, between about 10 nucleotides and about 100 nucleotides, such as between about 20 nucleotides and about 80 nucleotides, for example, between about 30 nucleotides and about 50 nucleotides. In other words, in some embodiments, the end of the target double-stranded nucleic acid is within about 500 nucleotides of the interstrand crosslinks formed in the disclosed method. In some embodiments, the end of the target double-stranded nucleic acid is within about 1 nucleotide to about 200 nucleotides of the interstrand crosslinks formed in the disclosed method. In some embodiments, the ends of the target double-stranded nucleic acid are between about 5 nucleotides and about 150 nucleotides of interstrand crosslinks formed in the disclosed method, for example between about 10 nucleotides and about 100 nucleotides, such as between about 20 nucleotides and about 80 nucleotides, for example between about 30 nucleotides and about 50 nucleotides.
[0135] In embodiments including the disclosed method of allowing double-stranded nucleic acid breakage, any suitable method can be used to break the double-stranded nucleic acid. In some embodiments, breakage occurs spontaneously. In some embodiments, breakage occurs spontaneously due to twisting. In some embodiments, twisting is formed by inter-strand crosslinks. Non-limiting examples of reagents for producing inter-strand crosslinks that generate twisted double-stranded nucleic acid structures are provided above. Therefore, in some embodiments, allowing double-stranded nucleic acid breakage includes allowing the double-stranded nucleic acid to break due to twisting within the double-stranded nucleic acid. In some embodiments, twisting is caused by inter-strand crosslinks.
[0136] In some embodiments, the breakage is caused by a cross-linking reaction. For example, in some embodiments, the cross-linking reaction forms inter-strand cross-links between the backbones of the first and second strands of the target double-stranded nucleic acid and simultaneously causes the double-stranded nucleic acid to break.
[0137] In some implementations, the breakage is caused by physical agitation. Therefore, in some implementations, the breakage of double-stranded nucleic acids is permitted to include physical agitation of the double-stranded nucleic acids. Any suitable form of physical agitation can be used. For example, breakage can be induced by pipetting, vortexing, sonication, stirring, acoustic shearing, nebulization, point-sink shearing, etc.
[0138] In some embodiments, the breakage is induced by a chemical or biological reagent. For example, in some embodiments, allowing double-stranded nucleic acid breakage involves contacting the double-stranded nucleic acid with a chemical or biological reagent. In some embodiments, the biological reagent is an enzyme. In some embodiments, the enzyme is an enzyme involved in DNA repair in vivo. Repair enzymes are known to detect damage and cleave double-stranded nucleic acids at interstrand crosslinks. Natural or engineered enzymes can be used to perform cleavage at or near interstrand crosslinks while maintaining the integrity of the interstrand crosslinks. Suitable reagents include, for example, restriction enzymes. Suitable restriction enzymes are commercially available, for example from New England Biolabs (USA). https: / / international.neb.com / products / restriction-endonucleases lists some suitable restriction enzymes.
[0139] In some embodiments, the interstrand cross-linked constructs formed in the disclosed methods are processed with one or more enzymes (e.g., one or more exonucleases) after the cross-linking process to degrade the exposed ends of the double-stranded polynucleotides. In some such embodiments, the method does not include breaking the double-stranded nucleic acid. In some such embodiments, the method further includes breaking the double-stranded nucleic acid.
[0140] Non-limiting examples of exonucleases that can be used to specifically degrade dsDNA ends up to cross-linking include exonuclease III from Escherichia coli; T7 exonuclease; exonuclease V (RecBCD); exonuclease VIII, truncated; and λ exonuclease.
[0141] In some embodiments, processing the cross-linked double-stranded nucleic acid construct with one or more exonucleases leaves little or no overhanging strands. In embodiments where overhanging strands are left, they can be removed using a suitable enzyme (e.g., mung bean endonuclease).
[0142] In some implementations, the interstrand-crosslinked constructs can be further processed before sequencing using single-molecule sequencing technology. For example, in some implementations, the ends of the interstrand-crosslinked constructs can be repaired if desired, for example, using NEBNext from New England Biolabs. (R) End-repair modules or equivalent reagents. In some embodiments, the inter-chain crosslinked construct may have one or more tails attached thereto, for example, by using NEBNext from New England Biolabs. (R) A dA-tailing module or equivalent reagent can be used to attach a dA tail to the construct.
[0143] Single-molecule sequencing
[0144] The methods disclosed in this paper include sequencing cross-linked constructs formed by cross-linking the first and second strands of double-stranded nucleic acids using single-molecule sequencing technology.
[0145] The methods disclosed herein are applicable to use with any suitable technology. As explained in more detail herein, in some embodiments, the methods provide orthogonal alignment sequence information. Therefore, they are suitable for single-molecule technologies where orthogonal alignment sequence information is useful. Suitable single-molecule sequencing technologies include nanopore chain sequencing (i.e., sequencing using nanopore sensors) and sequencing-on-synthesis. These technologies are described in more detail herein.
[0146] In nanopore sequencing, the target analyte (i.e., the cross-linked construct) moves relative to a transmembrane nanopore, as described in more detail herein, such as entering or passing through the nanopore. The signals recorded as the cross-linked construct moves relative to the nanopore allow for the determination of the sequence of the target double-stranded nucleic acid. Other characteristics of the construct can also be determined, such as (i) the length of the polynucleotide, (ii) the identity of the polynucleotide, (iii) the secondary structure of the polynucleotide, and (iv) whether the polynucleotide is modified.
[0147] In some implementations of nanopore chain sequencing, the binding of molecules (e.g., target polynucleotides) within the nanopore channel influences the open-channel ion flow through the pore. This is the essence of pore-channel “molecular sensing.” For example, changes in open-channel ion flow can be measured by variations in current using suitable measurement techniques (e.g., WO2000 / 28312 and D. Stoddart et al., Proc. Natl. Acad. Sci., 2010, 106, 7702-7 or WO2009 / 077734). The degree of reduction in ion flow measured by a decrease in current is related to the size of the barrier within or near the pore. Similar information can be obtained using optical methods, e.g., disclosed in Huang et al., Nature Nanotechnology 10, 986–991 (2015). Thus, the binding of molecules of interest (e.g., target polynucleotides) within or near the pore provides detectable and measurable events, forming the basis of a “biosensor.”
[0148] When nucleic acid molecules or individual bases move relative to a nanopore (e.g., through a channel within the nanopore), the size difference between the bases causes a directly related reduction in the ion flow through the channel. Changes in ion flow can be recorded. Suitable electrical measurement techniques for recording these changes are described, for example, in WO2000 / 28312 and D. Stoddart et al., Proc. Natl. Acad. Sci., 2010, 106, pp. 7702-7 (Single-channel recording devices); and, for example, WO2009 / 077734 (Multi-channel recording techniques). With proper calibration, this characteristic reduction in ion flow can be used to identify specific nucleotides and associated bases passing through the channel in real time. In typical nanopore nucleic acid sequencing, the open channel ion flow is reduced as individual nucleotides of the nucleic acid sequence of interest sequentially pass through the nanopore channel due to partial nucleotide blockage. This reduction in ion flow is precisely what is measured using the suitable recording techniques described above. A reduction in ion current can be calibrated to a measured reduction in ion current of a known nucleotide passing through the channel, thereby providing a means for determining which nucleotide is crossing the channel, and thus, when proceeding sequentially, a way to determine the nucleotide sequence of the nucleic acid passing through the nanopore. To accurately determine individual nucleotides, it is generally necessary to directly correlate the reduction in ion current through the channel with the size of the individual nucleotide crossing the constriction (or “readhead”). It should be understood that, for example, sequencing can be performed on intact nucleic acid polymers that “pass through” the pore, for example, by the action of associated motor proteins such as polymerases or helicases. Applicable motor proteins are described in more detail herein. Alternatively, the sequence can be determined by a pathway that allows nucleotide triphosphates to be sequentially removed from the target nucleic acid in adjacent pores (see, for example, WO2014 / 187924).
[0149] In other implementations, the sequence of the construct is determined using sequencing-by-synthesis. One example of sequencing-by-synthesis is single-molecule real-time sequencing.
[0150] Single-molecule real-time sequencing is a parallelized single-molecule DNA sequencing method. In embodiments of the disclosed methods that use single-molecule real-time sequencing to determine the sequence of a double-stranded nucleic acid construct, nucleic acid processing enzymes such as DNA polymerases are typically confined in a zero-mode waveguide. For example, the zero-mode waveguide may comprise a nanophotonic confinement structure comprising pores typically about 70 nm in diameter and about 100 nm deep in a coating film (e.g., an aluminum coating film) deposited on a substrate such as silica. Suitable enzymes (e.g., suitable polymerases) are described in more detail herein.
[0151] The single strand of the double-stranded nucleic acid construct comes into contact with a polymerase. The polymerase catalyzes the incorporation of a labeled nucleotide (e.g., optically (e.g., fluorescently) labeled nucleotide) to synthesize a new strand complementary to the construct strand processed by the polymerase. The incorporation of the labeled nucleotide generates a detectable signal characteristic of the incorporated nucleotide and, therefore, is complementary to the nucleotide at that position in the template (construct) strand. Multiple zero-mode waveguides can be used in the array to provide improved data throughput.
[0152] Process and proofread sequentially
[0153] When single-molecule sequencing technology is used to sequence double-stranded nucleic acid constructs, the methods disclosed herein typically allow for improved information. For example, in some embodiments, orthogonal alignment sequence information is provided by contacting the cross-linked construct with an enzyme that sequentially processes the first strand, optionally the cross-linked strand, and the second strand of the construct.
[0154] Those skilled in the art will understand that the methods disclosed herein can be used to obtain proofread sequence information because if both strands of a double-stranded nucleic acid construct linked by interstrand crosslinks are sequenced, each position of the double-stranded nucleic acid is not only sequenced once, but also queried twice. In other words, for example, if the strands of a double-stranded nucleic acid are not linked together by interstrand crosslinks, conventional techniques typically only sequence one strand of the double-stranded nucleic acid. The other strand is usually discarded. However, in embodiments of the disclosed method, the two strands of the double-stranded nucleic acid construct are processed sequentially because they are attached together by interstrand crosslinks, with each base pair being probed twice—once as the template strand and then again as the linked complementary strand. This “double” sequencing of both the template strand and the complementary strand, attached together by interstrand crosslinks, provides orthogonal sequence information. This information is orthogonal because the first nucleotide in the first strand is complementary to the corresponding nucleotide in the second strand. For example, when a C base is sequenced in the first strand, the corresponding position in the second strand will be a G base. Base calling algorithms can leverage this relationship to use information obtained when sequencing the second strand of a construct to "correct" the information obtained from the first strand when it was sequenced.
[0155] This ability to interrogate each location twice is particularly important when sequencing nucleic acids using random sensing methods such as nanopore sensing. Such sequencing typically relies on detecting each base one by one, for example, by capturing it through transmembrane nanopores, and requires a sufficiently high sampling rate to accurately identify the captured bases. The ability to efficiently interrogate each base twice reduces the need to capture each base at a sufficiently high rate.
[0156] The ability to query each position twice also helps distinguish similar bases, such as methylcytosine and thymine. When these two bases are characterized by random sequencing (e.g., using transmembrane pores), they sometimes produce similar signals. Therefore, it is difficult to distinguish between them. However, querying each position in a nucleic acid twice will allow this distinction because the complementary base of methylcytosine is guanine, while the complementary base of thymine is adenine. Methylcytosine is associated with a variety of diseases, including cancer.
[0157] Target double-stranded nucleic acid
[0158] As explained in more detail in this article, the method presented herein is a method for sequencing target double-stranded nucleic acids. Double-stranded nucleic acids can also be called double-stranded polynucleotides.
[0159] In some implementations, the double-stranded nucleic acid is secreted by the cell. Alternatively, the double-stranded nucleic acid may be an analyte present within the cell, thus requiring extraction from the cell for sequencing.
[0160] The double-stranded nucleic acids characterized in the methods described herein can be provided as an impure mixture of one or more target analytes and one or more impurities. Impurities may include truncated forms of target polynucleotide analytes that differ from the target analytes. For example, the target analyte may be genomic DNA, and impurities may include portions of genomic DNA, plasmids, etc. Target polynucleotides may be coding regions of genomic DNA, and unwanted polynucleotides may include non-coding regions of DNA.
[0161] Examples of polynucleotides include DNA and RNA. The bases in DNA and RNA can be distinguished by their physical size.
[0162] Polynucleotides, or nucleic acids, can include any combination of any nucleotides. Nucleotides can be naturally occurring or artificial. One or more nucleotides in a polynucleotide can be oxidized or methylated. One or more nucleotides in a polynucleotide can be damaged. For example, polynucleotides can include pyrimidine dimers. Such dimers are commonly associated with UV damage and are a major cause of melanoma.
[0163] One or more nucleotides in a polynucleotide may be modified (e.g., with a tag or label), suitable embodiments of which are known to those skilled in the art. A polynucleotide may include one or more spacers. An adaptor, such as a sequencing adaptor, may be included in the polynucleotide. Adaptors, tags, and spacers are described in more detail herein.
[0164] Examples of modified bases are disclosed herein and can be incorporated into target double-stranded nucleic acids in ways known in the art, such as by incorporation with a polymerase that modifies the nucleotide triphosphate during strand replication (e.g., in PCR) or by a polymerase-filling method. In some embodiments, one or more bases can be chemically modified using reagents known in the art.
[0165] Nucleotides typically contain a nucleobase, a sugar, and at least one phosphate group. The nucleobase and sugar form a nucleoside. The nucleobase is typically heterocyclic. Nucleobases include, but are not limited to, purines and pyrimidines, and more specifically, adenine (A), guanine (G), thymine (T), uracil (U), and cytosine (C). The sugar is typically a pentose sugar. Nucleotide sugars include (but are not limited to) ribose and deoxyribose. The sugar is preferably deoxyribose. Polynucleotides preferably contain the following nucleosides: deoxyadenosine (dA), deoxyuridine (dU), and / or thymidine (dT), deoxyguanosine (dG), and deoxycytidine (dC). Nucleotides are typically ribonucleotides or deoxyribonucleotides. Nucleotides typically contain monophosphate, diphosphate, or triphosphate. Nucleotides may include more than three phosphates, such as four or five phosphates. Phosphates may be attached to the 5' or 3' side of the nucleotide. Nucleotides in polynucleotides may be attached to each other in any manner. Nucleotides are typically attached by their sugar and phosphate groups, as in nucleic acids. Nucleotides can also be linked by their nucleobases, as in pyrimidine dimers.
[0166] The target polynucleotide is double-stranded. The target polynucleotide can be double-stranded DNA. The target polynucleotide can be double-stranded RNA. The target polynucleotide can be a DNA-RNA hybrid. A DNA-RNA hybrid can be prepared from single-stranded RNA by reverse transcription of a cDNA complement.
[0167] The preferred polynucleotides are double-stranded deoxyribonucleic acid (DNA) or double-stranded ribonucleic acid (RNA).
[0168] In addition to determining the sequence of the double-stranded nucleic acid, the methods disclosed herein may include determining one or more characteristics of the double-stranded nucleic acid selected from: (i) the length of the polynucleotide, (ii) the identity of the polynucleotide, (iii) the secondary structure of the polynucleotide and (iv) whether the polynucleotide is modified.
[0169] Polynucleotides can be of any length (i). For example, the length of a polynucleotide can be at least 10, at least 50, at least 100, at least 150, at least 200, at least 250, at least 300, at least 400, or at least 500 nucleotides or nucleotide pairs. The length of a polynucleotide can be 1000 or more nucleotides or nucleotide pairs, 5000 or more nucleotides or nucleotide pairs, or 100,000 or more nucleotides or nucleotide pairs. Any number of polynucleotides can be studied. For example, methods can involve characterizing 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 50, 100, or more polynucleotides. If two or more polynucleotides are characterized, they can be different polynucleotides or two examples of the same polynucleotide. Polynucleotides can be naturally occurring or artificial. For example, methods can be used to verify the sequence of a manufactured oligonucleotide. Methods are typically performed in vitro.
[0170] Nucleotides can have any identity (ii) and include, but are not limited to, adenosine monophosphate (AMP), guanosine monophosphate (GMP), thymidine monophosphate (TMP), uridine monophosphate (UMP), 5-methylcytidine monophosphate, 5-hydroxymethylcytidine monophosphate, cytidine monophosphate (CMP), cyclic adenosine monophosphate (cAMP), cyclic guanosine monophosphate (cGMP), deoxyadenosine monophosphate (dAMP), deoxyguanosine monophosphate (dGMP), deoxythymidine monophosphate (dTMP), deoxyuridine monophosphate (dUMP), deoxycytidine monophosphate (dCMP), and deoxymethylcytidine monophosphate. The nucleotide is preferably selected from AMP, TMP, GMP, CMP, UMP, dAMP, dTMP, dGMP, dCMP, and dUMP. Nucleotides can be baseless (i.e., lacking a nucleobase). Nucleotides may also lack a nucleobase and a sugar (i.e., be a C3 spacer). The sequence of nucleotides in a double-stranded nucleic acid is determined by the sequential identity of the following nucleotides attached to each other in the 5' to 3' direction of the strand throughout the polynucleotide strain.
[0171] The target double-stranded nucleic acid (DDNA) may include the product of a PCR reaction, genomic DNA, products of endonuclease digestion, and / or a DNA library. The target DDNA can be obtained or extracted from any organism or microorganism. Target DDNA is typically obtained from humans or animals, such as from urine, lymph, saliva, mucus, semen, or amniotic fluid, or from whole blood, plasma, or serum. Target DDNA can be obtained from plants, such as cereals, legumes, fruits, or vegetables. The target DDNA may contain genomic DNA. Genomic DNA can be fragmented. DNA can be fragmented by any suitable method. For example, methods for fragmenting DNA are known in the art, and such methods may use transposases, such as MuA transposase. Often, genomic DNA is not fragmented. In some embodiments, the target DDNA may be DNA, RNA, and / or a DNA / RNA hybrid.
[0172] Labeling analytes with molecular markers is within the scope of the methods provided herein. Molecular markers can be modifications of the analytes that facilitate detection of the analytes in the methods provided herein. For example, a marker can modify the signal obtained when determining the sequence of a double-stranded nucleic acid. For instance, in embodiments that include determining the sequence of a double-stranded nucleic acid as it moves relative to a nanopore, the marker can interfere with the ion flow through the nanopore. In this way, the marker can improve the sensitivity of the method.
[0173] connector
[0174] In some embodiments of the methods provided herein, the double-stranded nucleic acid to be sequenced has a polynucleotide adaptor to which it is attached. The adaptor typically comprises a polynucleotide chain capable of attaching to the end of the target polynucleotide.
[0175] In some embodiments, the adaptor is attached to the double-stranded nucleic acid before interstrand crosslinks are formed. In some embodiments, the adaptor is attached to a crosslinked double-stranded nucleic acid construct. In some embodiments of the disclosed method that includes allowing double-stranded nucleic acid breakage, the adaptor is attached to the double-stranded nucleic acid after interstrand crosslinks are formed but before allowing double-stranded nucleic acid breakage. In some embodiments of the disclosed method that includes allowing double-stranded nucleic acid breakage, the adaptor is attached to a crosslinked construct formed by breaking the interstranded crosslinks of the double-stranded nucleic acid.
[0176] Therefore, in some embodiments, the method includes attaching an intransitive linker (e.g., the intransitive linker described herein) to a target double-stranded nucleic acid and forming interstrand crosslinks within the intransitive linker. In some embodiments, the intransitive linker may be selected or modified to provide specific sites for interstrand crosslinking. For example, in some embodiments, the intransitive linker may contain components that are particularly sensitive to crosslinking conditions, such that only the intransitive linker crosslinks at the desired time.
[0177] An adaptor may be attached to only one end of a double-stranded nucleic acid or construct. Polynucleotide adaptors may be added to both ends of a double-stranded nucleic acid or construct. Alternatively, different adaptors may be added to both ends of a double-stranded nucleic acid or construct. The following discussion of "double-stranded polynucleotides" refers to target double-stranded nucleic acids or their cross-linked constructs.
[0178] An adaptor can be added to both strands of a double-stranded polynucleotide. An adaptor can also be added to only one strand of a polynucleotide. Methods for adding an adaptor to a polynucleotide are known in the art. The adaptor can be attached to the polynucleotide, for example, by ligation, by click chemistry, by labeling, by topoisomerization, or by any other suitable method.
[0179] In one embodiment, the adaptor or each adaptor is synthetic or artificial. Typically, the adaptor or each adaptor comprises a polymer as described herein. In some embodiments, the adaptor or each adaptor comprises a spacer as described herein. In some embodiments, the adaptor or each adaptor comprises a polynucleotide. The polynucleotide adaptor or each polynucleotide adaptor may comprise DNA, RNA, modified DNA (e.g., base-free DNA), RNA, PNA, LNA, BNA, and / or PEG. Typically, the adaptor or each adaptor comprises single-stranded and / or double-stranded DNA or RNA. The adaptor may contain a polynucleotide of the same type as the polynucleotide chain to which it is attached. The adaptor may contain a polynucleotide of a different type than the polynucleotide chain to which it is attached. In some embodiments, the polynucleotide chain evaluated and characterized in the methods described herein is a double-stranded DNA chain and the adaptor comprises DNA or RNA, such as double-stranded or single-stranded DNA.
[0180] In some embodiments, the adapter may be a bridging moiety. The bridging moiety can be used to connect the two strands of a double-stranded polynucleotide. For example, in some embodiments, the bridging moiety is used to connect the template strand of the double-stranded polynucleotide to its complementary strand.
[0181] The bridging portion typically covalently links the two strands of the target polynucleotide. The bridging portion can be anything capable of linking the two strands of the target polynucleotide, provided that it does not interfere with the movement of the single-stranded polynucleotide through the transmembrane pore. Suitable bridging portions include, but are not limited to, polymeric linkers, chemical linkers, polynucleotides, or peptides. Preferably, the bridging portion includes DNA, RNA, modified DNA (e.g., baseless DNA), RNA, PNA, LNA, or PEG. More preferably, the bridging portion is DNA or RNA.
[0182] In some embodiments, the bridging portion is a hairpin adaptor. A hairpin adaptor is an adaptor comprising a single polynucleotide chain, wherein the ends of the polynucleotide chains are capable of hybridizing to or being hybridized to each other, and wherein the middle segment of the polynucleotide forms a loop. Suitable hairpin adaptors can be designed using methods known in the art. In some embodiments, the length of the hairpin loop is typically 4 to 100 nucleotides, for example 4 to 50, 4 to 20, or 4 to 8 nucleotides. In some embodiments, the bridging portion (e.g., the hairpin adaptor) is attached to one end of the target polynucleotide. The bridging portion (e.g., the hairpin adaptor) is typically not attached to either end of the target polynucleotide.
[0183] In some embodiments, the adaptor is a linear adaptor. A linear adaptor can bind to either end or both ends of a single-stranded polynucleotide. When the polynucleotide is a double-stranded polynucleotide, the linear adaptor can bind to either end or both ends of either strand or both strands of the double-stranded polynucleotide. The linear adaptor may contain a leader sequence as described herein. The linear adaptor may contain a portion for hybridization with a tag (e.g., a pore tag) as described herein. The length of the linear adaptor can be from 10 to 150 nucleotides, such as 20 to 120, such as 30 to 100, such as 40 to 80, such as 50 to 70 nucleotides. The linear adaptor can be single-stranded. The linear adaptor can be double-stranded.
[0184] In some embodiments, the adaptor may be a Y-adaptor. A Y-adaptor is typically a polynucleotide adaptor. A Y-adaptor is typically double-stranded and includes (a) a region at one end where the two strands hybridize, and (b) a region at the other end where the two strands are not complementary. The non-complementary portions of the strands form overhangs. The presence of non-complementary regions in the Y-adaptor gives it a Y-shape because the two strands are not typically non-hybridized to each other as in a double-stranded portion. The two single-stranded portions of the Y-adaptor may be of the same length or different lengths. For example, one single-stranded portion of the Y-adaptor may be 10 to 150 nucleotides long, such as 20 to 120, 30 to 100, 40 to 80, or 50 to 70 nucleotides long, and the other single-stranded portion of the Y-adaptor may independently be 10 to 150 nucleotides long, such as 20 to 120, 30 to 100, 40 to 80, or 50 to 70 nucleotides long. The length of the double-stranded "stem" portion of the Y-connector can be, for example, 10 to 150 nucleotides, such as 20 to 120, such as 30 to 100, such as 40 to 80, such as 50 to 70 nucleotides.
[0185] The adaptor can be linked to the target polynucleotide by any suitable method known in the art. The adaptor can be synthesized separately and linked to the target polynucleotide by chemical attachment or enzymatic linkage. Alternatively, the adaptor can be generated during the processing of the target polynucleotide. In some embodiments, the adaptor is linked to the target polynucleotide at or near one end of the target polynucleotide. In some embodiments, the adaptor is linked to the target polynucleotide within 50 nucleotides, such as within 20 nucleotides, or even within 10 nucleotides from the end of the target polynucleotide. In some embodiments, the adaptor is linked to the end of the target polynucleotide. When the adaptor is linked to the target polynucleotide, the adaptor may comprise a nucleotide of the same type as the target polynucleotide or may comprise a nucleotide different from the target polynucleotide.
[0186] spacer
[0187] In some embodiments of the methods provided herein, the target double-stranded nucleic acid to be sequenced, the construct formed by its cross-linking, or the adaptor described herein may contain spacers. For example, one or more spacers may be present in a polynucleotide adaptor. For example, a polynucleotide adaptor may contain one to about 10 adaptors, such as one to about five spacers, such as one, two, three, four, or five spacers. Spacers may include any suitable number of spacer units. Spacers can provide an energy barrier that impedes the movement of motor proteins. For example, spacers can impede motor proteins by reducing the traction force of motor proteins on the polynucleotide. This can be achieved, for example, by using a base-free spacer, i.e., a spacer in which one or more nucleotides from the polynucleotide adaptor have been removed. Spacers can physically block the movement of motor proteins, for example by introducing a large chemical group to physically impede the movement of motor proteins.
[0188] In some implementations, particularly those in which nanopores are used to sequence double-stranded nucleic acids, one or more spacers are included in polynucleotides or constructs or adaptors as used in the methods claimed herein, so as to provide a unique signal as they pass through or across the nanopore (i.e., as they move relative to the nanopore).
[0189] In some embodiments, the spacer may comprise a linear molecule, such as a polymer. Typically, such spacers have a structure different from the target polynucleotide. For example, if the target polynucleotide is DNA, the spacer, or each spacer, typically does not contain DNA. Specifically, if the target polynucleotide is deoxyribonucleic acid (DNA) or ribonucleic acid (RNA), the spacer preferably comprises peptide nucleic acid (PNA), glycerol nucleic acid (GNA), threonine nucleic acid (TNA), locked nucleic acid (LNA), or a synthetic polymer with nucleotide side chains. In some embodiments, the spacer may comprise one or more nitroindole, one or more inosine, one or more acridine, one or more 2-aminopurine, one or more 2-6-diaminopurine, one or more 5-bromo-deoxyuridine, one or more reverse thymidine (reverse dT), one or more reverse dideoxythymidine (ddT), one or more dideoxycytidine (ddC), one or more 5-methylcytidine, one or more 5-hydroxymethylcytidine, one or more 2'-O-methylRNA bases, one or more isodeoxycytidine (iso -dC), one or more isodeoxyguanosine (Iso-dG), one or more C3 (OC3H6OPO3) groups, one or more optically cleavable (PC) [OC3H6-C(O)NHCH2-C6H3NO2-CH(CH3)OPO3] groups, one or more hexanediol groups, one or more spacer 9 (iSp9) [(OCH2CH2)3OPO3] groups or one or more spacer 18 (iSp18) [(OCH2CH2)6OPO3] groups; or one or more thiols are linked. The spacers can include any combination of these groups. Many of these groups can be derived from... (Integrated DNA Commercially available. For example, C3, iSp9, and iSp18 spacers can all be obtained from... Obtained. Spacers may include any number of the above-mentioned groups as spacer units.
[0190] In some embodiments, the spacer may contain one or more chemical groups that cause motor protein arrest. In some embodiments, suitable chemical groups are one or more chemical side groups. One or more chemical groups may be attached to one or more nucleobases in a polynucleotide, construct, or adaptor. One or more chemical groups may be attached to the backbone of a polynucleotide adaptor. Any number of suitable chemical groups may be present, such as 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, or more. Suitable groups include, but are not limited to, fluorophores, streptavidin and / or biotin, cholesterol, methylene blue, dinitrophenol (DNP), digoxigenin and / or anti-digoxigenin and diphenylcyclooctynyl groups. In some embodiments, the spacer may comprise a polymer. In some embodiments, the spacer may comprise a polymer, said polymer being a polypeptide or polyethylene glycol (PEG).
[0191] In some embodiments, the spacer may comprise one or more abasic nucleotides (i.e., nucleotides lacking a nucleobase), such as 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, or more abasic nucleotides. In the abasic nucleotides, the nucleobase may be replaced by -H (idSp) or -OH. The abasic spacer can be inserted into the target polynucleotide by removing a base from one or more adjacent nucleotides. For example, the polynucleotide may be modified to contain 3-methyladenine, 7-methylguanine, 1,N6-vinylidene adenine inosine, or hypoxanthine, and the nucleobase may be removed from these nucleotides using human alkyladenine DNA glycosidase (hAAG). Alternatively, the polynucleotide may be modified to contain uracil, and the nucleobase may be removed using uracil-DNA glycosidase (UDG). In one embodiment, one or more spacers do not include any abasic nucleotides.
[0192] anchor
[0193] In some embodiments, the double-stranded nucleic acid, its construct, or the adapter to which it is attached includes a membrane anchor or transmembrane pore anchor, for example, attached to the adapter. In one embodiment, the anchor facilitates the characterization of the double-stranded nucleic acid according to the methods disclosed herein. For example, the membrane anchor or transmembrane pore anchor can facilitate the localization of the double-stranded nucleic acid around nanopores in a membrane.
[0194] The anchor can be a peptide anchor and / or a hydrophobic anchor that can be inserted into the membrane. In one embodiment, the hydrophobic anchor is a lipid, fatty acid, sterol, carbon nanotube, peptide, protein, or amino acid, such as cholesterol, palmitate, or tocopherol. The anchor may include thiols, biotin, or surfactants.
[0195] On the one hand, the anchor can be biotin (for binding to streptavidin), amylose (for binding to maltose-binding proteins or fusion proteins), Ni-NTA (for binding to polyhistidine or polyhistidine-labeled proteins), or peptides (such as antigens).
[0196] In one embodiment, the anchor may include a linker, or two, three, four or more linkers. Preferred linkers include, but are not limited to, polymers such as polynucleotides, polyethylene glycol (PEG), polysaccharides, and peptides. These linkers may be linear, branched, or cyclic. For example, the linker may be a cyclic polynucleotide. The linker may hybridize to a complementary sequence on the cyclic polynucleotide linker. One or more anchors or one or more linkers may include components that can be cleaved or broken down, such as restriction sites or photostable groups. The linker may be functionalized with maleimide groups to attach to cysteine residues in a protein. Suitable linkers are described in WO2010 / 086602.
[0197] In one embodiment, the anchor is a cholesterol or fatty acyl chain. For example, any fatty acyl chain having a length of 6 to 30 carbon atoms, such as hexadecanoic acid, can be used. Examples of suitable anchors and methods for attaching anchors to connectors are disclosed in WO2012 / 164270 and WO2015 / 150786.
[0198] Controlling the movement of the analyte relative to the detector
[0199] As explained in more detail above, some embodiments of the methods provided herein involve contacting a cross-linked construct with an enzyme that sequentially processes the first and second strands of the construct. In some embodiments, the enzyme is used to control the movement of the construct relative to a detector such as a transmembrane nanopore. In some embodiments, the enzyme is used to process the construct in next-generation sequencing methods, such as sequencing-by-synthesis methods, such as single-molecule real-time sequencing.
[0200] In some embodiments, the enzyme can process the first and second chains of the cross-linked construct sequentially. In some embodiments, the enzyme can process the first chain and then the second chain. In some embodiments, the enzyme can process the first chain, inter-chain crosslinks, and the second chain. In some embodiments, the enzyme can process from the first chain to the second chain via inter-chain crosslinks. In some embodiments, the enzyme can process the first and second chains sequentially without inter-chain crosslinks.
[0201] The movement of the construct relative to the detector used in sequencing technology can be controlled in any suitable manner. In some embodiments, the movement of the construct is driven by physical or chemical forces (potentials). In some embodiments, the physical forces are provided by electric potential (e.g., voltage potential) or temperature gradients, etc.
[0202] In some embodiments, the detector is a nanopore, and the construct moves relative to the nanopore when a potential is applied across the nanopore. Polynucleotides are negatively charged, so applying a potential across the nanopore will cause the polynucleotide to move relative to the nanopore under the influence of the applied potential. For example, if a positive voltage potential is applied relative to the cis side of the nanopore to the trans side, this will induce a negatively charged analyte to move from the cis side to the trans side. Similarly, if a positive voltage potential is applied relative to the cis side of the nanopore to the trans side, this will prevent a negatively charged analyte from moving from the trans side to the cis side. The opposite occurs if a negative voltage potential is applied relative to the cis side of the nanopore to the trans side. Apparatus and methods for applying appropriate voltages are described in more detail herein.
[0203] In some implementations, the chemical force is provided by a concentration (e.g., pH) gradient.
[0204] In some embodiments, the methods provided herein include contacting the construct with a polynucleotide processing enzyme.
[0205] Suitable polynucleotide processing enzymes are sometimes referred to as motor proteins or polynucleotide processing enzymes. Suitable polynucleotide processing enzymes are known in the art. Therefore, in some embodiments, the provided method includes contacting the construct with a motor protein, wherein the motor protein controls the movement of the construct relative to the nanopore.
[0206] In some implementations, the motor protein may be present on the construct prior to its contact with the nanopore. For example, the motor protein may be present on an adaptor containing a portion of the construct analyte, or it may be present on a portion of the construct itself.
[0207] In some embodiments, the motor protein is modified to prevent it from detaching from the polynucleotide, construct, or adaptor (other than by removing the ends of the polynucleotide / construct / adaptor). Such modified motor proteins are particularly suitable for the disclosed methods.
[0208] Motor proteins can be modified in any suitable manner. For example, a motor protein can be loaded onto a polynucleotide, construct, or adaptor and then modified to prevent its detachment. Alternatively, a motor protein can be modified to prevent its detachment before being loaded onto a polynucleotide, construct, or adaptor. Modification of motor proteins to prevent their detachment from polynucleotides, constructs, or adaptors can be achieved using methods known in the art, such as those discussed in WO2014 / 013260 (which is incorporated herein by reference in its entirety), and with particular reference to the paragraph describing the modification of motor proteins (polynucleotide-binding proteins) such as helicases to prevent their detachment from polynucleotide chains.
[0209] For example, motor proteins may have polynucleotide unwinding openings; for example, cavities, crevices, or gaps through which the polynucleotide chain can pass when the motor protein is detached from the chain. In some embodiments, the polynucleotide unwinding opening for a given motor protein (polynucleotide-binding protein) can be determined by referring to its structure, such as its X-ray crystal structure. The X-ray crystal structure can be obtained in the presence and / or absence of a polynucleotide substrate. In some embodiments, the location of the polynucleotide unwinding opening in a given motor protein can be inferred or confirmed by molecular modeling using standard packages known in the art. In some embodiments, the polynucleotide unwinding opening can be transiently generated by the movement of one or more portions of the motor protein, such as one or more domains.
[0210] Motor proteins can be modified by closing polynucleotide helical openings. Therefore, closing polynucleotide helical openings prevents the motor protein from dissociating from the polynucleotide or adaptor. For example, motor proteins can be modified by covalently closing polynucleotide helical openings. In some embodiments, as described herein, the motor protein used for addressing in this manner is a helicase.
[0211] In one embodiment, the motor protein is or is derived from a polynucleotide processing enzyme. A polynucleotide processing enzyme is a polypeptide capable of interacting with and modifying at least one property of a polynucleotide. The enzyme can modify a polynucleotide by cleaving it to form individual nucleotides or shorter nucleotide chains such as dinucleotides or trinucleotides. The enzyme can also modify a polynucleotide by orienting or moving it to a specific location.
[0212] Motor proteins can be selected or chosen based on the double-stranded nucleic acid target to be sequenced in the methods disclosed herein. Alternatively, double-stranded nucleic acid targets can be selected or chosen based on motor proteins (if any) used to process the first and second strands of the terminal crosslinked construct. For example, when the target double-stranded nucleic acid is double-stranded DNA, DNA motor proteins can typically be used. When the target double-stranded nucleic acid is double-stranded RNA, RNA motor proteins can be used. When the target double-stranded nucleic acid is a DNA-RNA hybrid, motor proteins capable of processing both DNA and RNA can be used.
[0213] In one implementation, the motor protein is derived from members of any enzyme classification (EC) group: 3.1.11, 3.1.13, 3.1.14, 3.1.15, 3.1.16, 3.1.21, 3.1.22, 3.1.25, 3.1.26, 3.1.27, 3.1.30, and 3.1.31.
[0214] In some embodiments, the motor protein is a helicase, polymerase, exonuclease, topoisomerase, or a variant thereof.
[0215] In one embodiment, the motor protein is an exonuclease. Suitable enzymes include, but are not limited to, exonuclease I (SEQ ID NO:1) from *Escherichia coli*, exonuclease III (SEQ ID NO:2) from *Escherichia coli*, RecJ (SEQ ID NO:3) from *T. thermophilus*, bacteriophage λ exonuclease (SEQ ID NO:4), TatD exonuclease, and variants thereof. Three subunits comprising the sequence shown in SEQ ID NO:3, or variants thereof, interact to form a trimer exonuclease.
[0216] In one implementation, the motor protein is a polymerase. The polymerase can be... 3173 DNA polymerase (which is commercially available) (Company), SD polymerase (commercially available) The enzyme is Klenow from NEB or a variant thereof. In one embodiment, the enzyme is Phi29 DNA polymerase (SEQ ID NO:5) or a variant thereof. A modified version of the Phi29 polymerase that can be used in the disclosed methods is disclosed in U.S. Patent No. 5,576,204.
[0217] In the embodiments provided herein, including methods for sequencing constructs using sequencing-by-synthesis reactions, the enzyme is typically a polymerase, such as the polymerase described herein.
[0218] In one embodiment, the motor protein is a topoisomerase. In one embodiment, the topoisomerase is a member of either group 5.99.1.2 or 5.99.1.3 of the partial classification (EC). The topoisomerase can be a reverse transcriptase, which is an enzyme capable of catalyzing the formation of cDNA from an RNA template. These can be derived from, for example, New England... and Acquired through commercial purchase.
[0219] In one embodiment, the motor protein is a helicase. Any suitable helicase can be used according to the methods provided herein. For example, the motor protein used according to this disclosure, or each motor protein, can be independently selected from Hel308 helicase, RecD helicase, TraI helicase, TrwC helicase, XPD helicase, and Dda helicase, or variants thereof. Monomeric helicases can include several domains attached together. For example, TraI helicase and TraI subgroup helicases can contain two RecD helicase domains, a release enzyme domain, and a C-terminal domain. These domains typically form a monomeric helicase capable of functioning without forming oligomers. Specific examples of suitable helicases include Hel308, NS3, Dda, UvrD, Rep, PcrA, Pif1, and TraI. These helicases typically act on single-stranded DNA. Examples of helicases that can move along both strands of double-stranded DNA include FtfK and hexamethylenetetramer complexes, or multi-subunit complexes such as RecBCD. NS3 helicases are particularly suitable for the disclosed methods because they are capable of processing both DNA and RNA, and therefore can be used in embodiments of the disclosed methods in which the target double-stranded nucleic acid is a DNA-RNA hybrid.
[0220] Hel308 helicase is described in publications such as WO2013 / 057495, the entire contents of which are incorporated herein by reference. RecD helicase is described in publications such as WO2013 / 098562, the entire contents of which are incorporated herein by reference. XPD helicase is described in publications such as WO2013 / 098561, the entire contents of which are incorporated herein by reference. Dda helicase is described in publications such as WO2015 / 055981 and WO2016 / 055777, the entire contents of which are incorporated herein by reference.
[0221] In one embodiment, the helicase comprises the sequence shown in SEQ ID NO:6 (Trwc Cba) or a variant thereof, the sequence shown in SEQ ID NO:7 (Hel308 Mbu) or a variant thereof, or the sequence shown in SEQ ID NO:8 (Dda) or a variant thereof. The variants may differ from the natural sequence in any of the ways discussed below. An example variant of SEQ ID NO:8 includes E94C / A360C. Another example variant of SEQ ID NO:8 includes E94C / A360C, followed by (ΔM1)G1G2 (i.e., the deletion of M1, followed by the addition of G1 and G2).
[0222] In some embodiments, motor proteins (e.g., helicases) can operate in at least two modes of activity (when the motor protein has all the necessary components for promoting movement, such as fuels and cofactors discussed herein, such as ATP and Mg).2+ It controls the movement of the construct in an inactive mode of operation (when the motor protein does not have the necessary components to promote movement).
[0223] When all the necessary components are provided to facilitate movement (i.e., in active mode), motor proteins (e.g., helicases) move along the polynucleotide construct in a 5' to 3' or 3' to 5' direction (depending on the motor protein). In embodiments in which motor proteins are used to control the movement of the polynucleotide chains of the construct relative to the nanopore, the motor protein can be used to move the construct away from (e.g., out of) the pore (e.g., against an applied field) or to move the construct toward (e.g., into) the pore (e.g., using an applied field). For example, when the end of the construct moved by the motor protein is trapped in the pore, the motor protein works against the direction of the force and pulls the constructed through out of the pore (e.g., into the cis chamber). However, when the far end of the construct moved by the motor protein is trapped in the pore, the motor protein works in the direction of the force and pushes the constructed through into the pore (e.g., into the trans chamber).
[0224] When motor proteins (such as helicases) do not provide the necessary components to promote movement (i.e., they are in an inactive mode), they can bind to the construct and act as a brake, slowing down the construct's movement relative to the nanopore, for example, by being pulled into the pore by force. In the inactive mode, it is not important which end of the construct is captured; the applied force determines the construct's movement relative to the pore, and the polynucleotide-binding protein acts as the brake. The control of construct movement by polynucleotide-binding proteins in the inactive mode can be described in several ways (including ratcheting, sliding, and braking).
[0225] Motor proteins typically require fuel to process polynucleotides. This fuel is usually a free nucleotide or a free nucleotide analogue. Free nucleotides can be, but are not limited to, adenosine monophosphate (AMP), adenosine diphosphate (ADP), adenosine triphosphate (ATP), guanosine monophosphate (GMP), guanosine diphosphate (GDP), guanosine triphosphate (GTP), thymidine monophosphate (TMP), thymidine diphosphate (TDP), thymidine triphosphate (TTP), uridine monophosphate (UMP), uridine diphosphate (UDP), uridine triphosphate (UTP), cytidine monophosphate (CMP), cytidine diphosphate (CDP), cytidine triphosphate (CTP), cyclic adenosine monophosphate (cAMP), cyclic guanosine monophosphate (cGMP), and deoxygenated adenosine monophosphate (dAMP). Deoxyadenosine monophosphate (dAMP), deoxyadenosine diphosphate (dADP), deoxyadenosine triphosphate (dATP), deoxyguanosine monophosphate (dGMP), deoxyguanosine diphosphate (dGDP), deoxyguanosine triphosphate (dGTP), deoxythymidine monophosphate (dTMP), deoxythymidine diphosphate (dTDP), deoxythymidine triphosphate (dTTP), deoxyuridine monophosphate (dUMP), deoxyuridine diphosphate (dUDP), deoxyuridine triphosphate (dUTP), deoxycytidine monophosphate (dCMP), deoxycytidine diphosphate (dCDP), and deoxycytidine triphosphate (dCTP). Free nucleotides are typically selected from AMP, TMP, GMP, CMP, UMP, dAMP, dTMP, dGMP, or dCMP. The most common free nucleotide is adenosine triphosphate (ATP).
[0226] Cofactors of motor proteins are factors that allow motor proteins to function. Cofactors are preferably divalent metal cations. The preferred divalent metal cation is Mg. 2+ Mn 2+ Ca 2+ or Co 2+ The most preferred cofactor is Mg. 2+ .
[0227] Nanopores
[0228] As described above, some implementations of the methods provided herein include sequencing cross-linked constructs using nanopore sensors.
[0229] In the disclosed method embodiments involving nanopores, any suitable nanopore can be used. In one embodiment, the nanopore is a transmembrane pore.
[0230] A transmembrane pore is a structure that spans the membrane to some extent. It allows hydrated ions to flow across or within the membrane, driven by an applied potential. A transmembrane pore typically extends across the entire membrane, allowing hydrated ions to flow from one side to the other. However, a transmembrane pore does not necessarily extend across the membrane. It may be closed at one end. For example, a pore can be a hole, gap, channel, groove, or slit in the membrane, allowing hydrated ions to flow into or into the membrane.
[0231] Any transmembrane pore can be used in the methods described herein. The pore can be biological or artificial. Suitable pores include, but are not limited to, protein pores, polynucleotide pores, and solid pores. The pore can be a DNA origami pore (Langecker et al., Science, 2012; 338:932-936). Suitable DNA origami pores are disclosed in WO2013 / 083983.
[0232] In one embodiment, the nanopore is a transmembrane protein pore. A transmembrane protein pore is a polypeptide or aggregate of polypeptides that allows hydrated ions (e.g., polynucleotides) to flow from one side of a membrane to the other. In the methods provided herein, transmembrane protein pores are capable of forming pores that allow hydrated ions, driven by an applied potential, to flow from one side of a membrane to the other. Transmembrane protein pores preferably allow polynucleotides to flow from one side of a membrane (e.g., a triblock copolymer membrane) to the other. Transmembrane protein pores allow polynucleotides to move through the pore.
[0233] In one embodiment, the nanopore is a transmembrane protein pore, which is a monomer or oligomer. The pore is preferably composed of a plurality of repeating subunits, such as at least 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, or 16 subunits. The pore is preferably a hexamer, heptamer, octamer, or non-merchanmeric pore. The pore can be a homooligomer or a heterooligomer.
[0234] In one embodiment, a transmembrane protein pore comprises a barrel or channel through which ions can flow. The subunits of the pore typically surround a central axis and provide chains for transmembrane β-barrels or channels, or transmembrane α-helical bundles or channels.
[0235] Typically, the barrels or channels of transmembrane protein pores comprise amino acids that facilitate interaction with analytes, such as target polynucleotides (as described herein). These amino acids are preferably located near the constriction of the barrel or channel. Transmembrane protein pores typically contain one or more positively charged amino acids, such as arginine, lysine, or histidine, or aromatic amino acids, such as tyrosine or tryptophan. These amino acids typically facilitate interactions between the pore and nucleotides, polynucleotides, or nucleic acids.
[0236] In one embodiment, the nanopores are transmembrane protein pores derived from β-barrel pores or α-helical bundle pores. β-barrel pores comprise barrels or channels formed by β-chains. Suitable β-barrel pores include, but are not limited to, β-toxins such as α-hemolysin, anthrax toxin, and leukocytoxin, as well as bacterial outer membrane proteins / porins such as Mycobacterium smegmatis porin (Msp), such as MspA, MspB, MspC, or MspD, CsgG, outer membrane porin F (OmpF), outer membrane porin G (OmpG), outer membrane phospholipase A, and Neisseria self-transporter lipoprotein (NalP), and other pores such as lysenin. α-helical bundle pores comprise barrels or channels formed by α-helices. Suitable α-helical bundle pores include, but are not limited to, inner membrane proteins and α-outer membrane proteins such as WZA and ClyA toxins.
[0237] In one embodiment, the nanopore is a transmembrane pore derived from or based on Msp, α-hemolysin (α-HL), cytolysin, CsgG, ClyA, Sp1, and the hemolysin fragaceatoxin C (FraC).
[0238] In one embodiment, the nanopores are derived from CsgG, for example, from CsgG derived from the Escherichia coli strain K-12 substrain MC4100. Such pores are oligomeric and typically comprise 7, 8, 9, or 10 monomers derived from CsgG. The pores can be homopolymeric oligomeric pores derived from CsgG comprising the same monomers. Alternatively, the pores can be heteropolymeric oligomeric pores derived from CsgG comprising at least one monomer different from the others. Examples of suitable pores derived from CsgG are disclosed in WO2016 / 034591.
[0239] In one embodiment, the nanopores are transmembrane pores derived from lysozyme. Examples of suitable pores derived from lysozyme are disclosed in WO2013 / 153359.
[0240] In one embodiment, the nanopore is derived from or based on transmembrane pore α-hemolysin (α-HL). Wild-type α-hemolysin pores are formed by seven identical monomers or subunits (i.e., they are heptameric). α-hemolysin pores can be α-hemolysin-NN or variants thereof. Variants preferably include N residues at positions E111 and K147.
[0241] In one embodiment, the nanopore is a transmembrane protein pore derived from Msp (e.g., derived from MspA). Examples of suitable pores derived from MspA are disclosed in WO2012 / 107778.
[0242] In one implementation, the nanopores are transmembrane pores derived from or based on ClyA.
[0243] Label
[0244] In some embodiments of the methods provided herein, as described above, sequencing of cross-linked double-stranded nucleic acid constructs involves contacting the constructs with a transmembrane nanopore. In some embodiments, for example, tags on the nanopores may be used to facilitate the capture of analytes by the nanopores.
[0245] The interaction between the tag on the nanopore and the binding site on the polynucleotide (e.g., a binding site present in the adaptor attached to the polynucleotide, wherein the binding site may be provided by the anchor or leader sequence of the adaptor or by the capture sequence within the double-stranded stem of the adaptor) can be reversible. For example, the polynucleotide can bind to the tag on the nanopore, for example, through its adaptor, and be released at certain points, for example, during the characterization of the polynucleotide through the nanopore and / or during motor protein processing. Strong non-covalent binding (e.g., biotin / avidin) remains reversible and can be used in some embodiments of the methods described herein. For example, a pair of pore tags and polynucleotide adaptors can be designed to provide sufficient interaction between the complement of the double-stranded polynucleotide (or a portion of the adaptor attached to the complement) and the nanopore, such that the complement remains close to the nanopore (without dissociating from and diffusing from the nanopore) but is able to be released from the nanopore upon processing.
[0246] The pore tag and polynucleotide adaptor can be configured such that the binding strength or affinity of the binding site on the polynucleotide (e.g., a binding site provided by the anchor or leader sequence of the adaptor or by a capture sequence within the double-stranded stem of the adaptor) to the tag on the nanopore is sufficient to maintain the coupling between the nanopore and the polynucleotide until an applied force is placed thereon to release the bound polynucleotide from the nanopore. In some embodiments where the analyte is a double-stranded polynucleotide, the applied force can be the complementary strand processed by a polymerase.
[0247] In some implementations, the tag or tether is uncharged. This ensures that the tag or tether will not be pulled into the nanopore by a potential difference (if present).
[0248] One or more molecules that attract or bind to polynucleotides or adaptors can be attached to a detector (e.g., a pore). Any molecule that hybridizes to the adaptor and / or target polynucleotide can be used. The molecules attached to the pore can be selected from PNA tags, PEG linkers, short oligonucleotides, positively charged amino acids, and aptamers. Pores with such molecules attached to them are known in the art. For example, pores with short oligonucleotides attached thereto are disclosed in Howarka et al. (2001), Nature Biotech, 19:636-639 and WO2010 / 086620, and pores including PEG attached to the lumen of the pore are disclosed in Howarka et al. (2000), J.Am.Chem.Soc., 122(11):2411-2416.
[0249] Short oligonucleotides attached to nanopores include sequences complementary to a leader sequence in an adaptor or another single-stranded sequence, which can be used to enhance the capture of target polynucleotides in the methods described herein.
[0250] In some embodiments, the tag or tie may comprise or may be an oligonucleotide (e.g., DNA, RNA, LNA, BNA, PNA, or morpholino). The oligonucleotide may be about 10-30 nucleotides long or about 10-20 nucleotides long. In some embodiments, the oligonucleotide may have at least one end (e.g., a 3' or 5' end) modified for conjugation to other modified or solid substrate surfaces (including, for example, beads). The end modifier may add reactive functional groups that can be used for conjugation. Examples of functional groups that can be added include, but are not limited to, amino, carboxyl, thiol, maleimide, aminooxy, and any combination thereof. The functional group may be combined with spacers of different lengths (e.g., C3, C9, C12, spacers 9 and 18) to increase the physical distance between the functional group and the end of the oligonucleotide sequence.
[0251] Examples of modifications to the 3' and / or 5' ends of oligonucleotides include, but are not limited to, 3' affinity tags and functional groups for chemical linking (including, for example, 3'-biotin, 3'-primary amine, 3'-disulfide amide, 3'-pyridyl disulfide, and any combination thereof); 5' end modifications (including, for example, 5'-primary amine and / or 5'-fluorescein); modifications for click chemistry (including, for example, 3'-azide, 3'-alkynyl, 5'-azide, 5'-alkynyl) and any combination thereof.
[0252] In some embodiments, the tag or tether may further include a polymeric connector, for example, to facilitate coupling to the nanopore. Exemplary polymeric connectors include, but are not limited to, polyethylene glycol (PEG). The molecular weight of the polymeric connector may be from about 500 Da to about 10 kDa (including end values), or from about 1 kDa to about 5 kDa (including end values). The polymeric connector (e.g., PEG) may be functionalized with different functional groups, including, but not limited to, maleimide, NHS ester, dibenzocyclooctylene (DBCO), azide, biotin, amine, alkyne, aldehyde, and any combination thereof.
[0253] Other embodiments of the tag or tether include, but are not limited to, His tags, biotin or streptavidin, antibodies that bind to the analyte, aptamers that bind to the analyte, analyte-binding domains such as DNA-binding domains (including, for example, peptide zippers, such as leucine zippers, single-stranded DNA-binding proteins (SSBs)) and any combination thereof.
[0254] Tags or tethers can be attached to the outer surface of a nanopore, for example, on the cis side of a membrane, using any method known in the art. For example, one or more tags or tethers can be attached to a nanopore via one or more cysteine residues (cysteine bonds), one or more primary amines (such as lysine), one or more non-natural amino acids, one or more histidine residues (His tags), one or more biotin or streptavidin residues, one or more antibody-based tags, one or more enzymatic modifications of epitopes (including, for example, acetyltransferases), and any combination thereof. Suitable methods for making such modifications are well known in the art. Suitable non-natural amino acids include, but are not limited to, 4-azido-L-phenylalanine (Faz), and [the following is a list of amino acids and their components]. (Liu C.C. and Schultz PG, AnnuRev Biochem, 2010, 79, 413-444) Figure 1 Any one of the amino acids numbered 1-71 in the Chinese language.
[0255] In some embodiments where one or more tags or chains are attached to nanopores via cysteine bonds, one or more cysteine residues may be introduced into one or more monomers that form nanopores by substitution. In some embodiments, the nanopores may be chemically modified by attaching: (i) maleimides, including dibromomaleimides such as 4-benzodiazepine, 1,N-(2-hydroxyethyl)maleimide, N-cyclohexylmaleimide, 1,3-maleimide propionic acid, 1,1-4-aminophenyl-1H-pyrrole,2,5,dione, 1,1-4-hydroxyphenyl-1H-pyrrole,2,5,dione, N-ethylmaleimide, N-methoxycarbonylmaleimide. Imide, N-tert-butylmaleimide, N-(2-aminoethyl)maleimide, 3-maleimide-propoxy, N-(4-chlorophenyl)maleimide, 1-[4-(dimethylamino)-3,5-dinitrophenyl]-1H-pyrrole-2,5-dione, N-[4-(2-benzimidazolyl)phenyl]maleimide, N-[4-(2-benzoxazolyl)phenyl]maleimide, N-(1-naphthyl)maleimide, N-(2,4-dimethyl)maleimide Phenyl)maleimide, N-(2,4-difluorophenyl)maleimide, N-(3-chloro-p-tolyl)-maleimide, 1-(2-amino-ethyl)-pyrrole-2,5-dione hydrochloride, 1-cyclopentyl-3-methyl-2,5-dihydro-1H-pyrrole-2,5-dione, 1-(3-aminopropyl)-2,5-dihydro-1H-pyrrole-2,5-dione hydrochloride, 3-methyl-1-[2-oxo-2-(piperazin-1-yl)ethyl]-2 5-Dihydro-1H-pyrrole-2,5-dione hydrochloride, 1-benzyl-2,5-dihydro-1H-pyrrole-2,5-dione, 3-methyl-1-(3,3,3-trifluoropropyl)-2,5-dihydro-1H-pyrrole-2,5-dione, 1-[4-(methylamino)cyclohexyl]-2,5-dihydro-1H-pyrrole-2,5-dione trifluoroacetic acid, SMILESO=C1C=CC(=O)N1CC=2C=CN=CC2, SMILES O=C1C=CC(=O)N1CN2CCNCC2、1-Benzyl-3-methyl-2,5-dihydro-1H-pyrrole-2,5-dione、1-(2-fluorophenyl)-3-methyl-2,5-dihydro-1H-pyrrole-2,5-dione、N-(4-phenoxyphenyl)maleimide、N-(4-nitrophenyl)maleimide、(ii)iodoacetamide, such as 3-(2-iodoacetamido)-propoxy、N-(cyclopropylmethyl)-2-iodoacetamide、2-iodo-N-(2-phenylethyl)acetamide、2-iodo-N-(2,2,2-trifluoroethyl)acetamide、N-(4-acetylphenyl)-2-iodoacetamide、N-(4-(aminosulfonyl)phenyl)-2-iodoacetamide、N-(1,3-Benzothiazol-2-yl)-2-iodoacetamide, N-(2,6-diethylphenyl)-2-iodoacetamide, N-(2-benzoyl-4-chlorophenyl)-2-iodoacetamide, (iii) bromoacetamides: such as N-(4-(acetamido)phenyl)-2-bromoacetamide, N-(2-acetylphenyl)-2-bromoacetamide, 2-bromo-n-(2-cyanophenyl)acetamide, 2-bromo-N-(3-(trifluoromethyl)phenyl)acetamide, N-(2-benzoylphenyl)-2-bromoacetamide, 2-bromo-N-(4-fluorophenyl)-3-methylbutyramide, N-benzyl-2-bromo-N-phenylpropionamide, N-(2-bromo-butyryl)-4-chloro-benzenesulfonamide, 2-bromo-N-methyl-N-phenylacetamide, 2-bromo- N-Phenylacetamide, 2-adamantane-1-yl-2-bromo-N-cyclohexylacetamide, 2-bromo-N-(2-methylphenyl)butyramide, acetyl-p-bromoaniline; (iv) disulfides, such as aldrithiol-2, aldrithiol-4, isopropyl disulfide, 1-(isobutyldithioalkyl)-2-methylpropane, dibenzyl disulfide, 4-aminophenyl disulfide, 3-(2-pyridyldithio)propionic acid, 3-(2-pyridyldithio)propionic acid hydrazide, 3-(2-pyridyldithio)propionic acid N-succinimide, am6amPDP1-βCD; and (v) thiols, such as 4-phenylthiazolyl-2-thiol, Pulpald, 5,6,7,8-tetrahydro-quinazolin-2-thiol.
[0256] In some embodiments, the tag or tether can be attached to the nanopore directly or via one or more adapters. The tag or tether can be attached to the nanopore using the hybridization adapters described in WO2010 / 086602. Alternatively, peptide adapters can be used. Peptide adapters are amino acid sequences. The length, flexibility, and hydrophilicity of peptide adapters are generally designed so that they do not interfere with the function of the monomer and the pore. Preferred flexible peptide adapters are 2 to 20, such as 4, 6, 8, 10, or 16 serine and / or glycine extensions. More preferred flexible adapters comprise (SG)1, (SG)2, (SG)3, (SG)4, (SG)5, and (SG)8, where S is serine and G is glycine. Preferred rigid adapters are 2 to 30, such as 4, 6, 8, 16, or 24 proline extensions. More preferred rigid adapters comprise (P) 12 , where P is proline.
[0257] membrane
[0258] In embodiments of the disclosed methods that include the use of transmembrane nanopores, the transmembrane nanopores are typically present within the membrane. Any suitable membrane can be used in the system.
[0259] The membrane is preferably an amphiphilic layer. An amphiphilic layer is a layer formed by amphiphilic molecules such as phospholipids, possessing both hydrophilic and lipophilic properties. The amphiphilic molecules can be synthetic or naturally occurring. Non-naturally occurring amphiphiles and amphiphiles forming monolayers are known in the art, including, for example, block copolymers (Gonzalez-Perez et al., Langmuir, 2009, 25, 10447-10450). A block copolymer is a polymeric material in which two or more monomer subunits polymerized together form a single polymer chain. Block copolymers typically possess properties contributed by each monomer subunit. However, block copolymers can possess unique properties not found in polymers formed from individual subunits. Block copolymers can be engineered so that one of the monomer subunits is hydrophobic (i.e., lipophilic) in an aqueous medium, while the other subunits are hydrophilic. In this case, the block copolymer can possess amphiphilic properties and can form a structure mimicking a biological membrane. Block copolymers can be diblock (consisting of two monomer subunits), but can also be constructed from more than two monomer subunits, forming a more complex arrangement exhibiting amphiphilic behavior. The copolymer can be triblock, tetrablock, or pentablock copolymers. The membrane is preferably a triblock copolymer membrane.
[0260] Archaea bipolar tetraether lipids are naturally occurring lipids that are constructed to form monolayer membranes. These lipids are generally found in extremophiles, thermophiles, halophiles, and acidophiles that survive in harsh biological environments. Their stability is thought to stem from the fusion properties of the final bilayer. A straightforward approach is to construct block copolymer materials that mimic these biological entities by generating triblock polymers with a general motif of hydrophilic-hydrophobic-hydrophilic properties. These materials can form monomeric membranes that exhibit lipid bilayer-like behavior and encompass a range of stages from vesicles to lamellar membranes. Membranes formed from these triblock copolymers retain several advantages over biological lipid membranes. Because of the synthesis of triblock copolymers, precise construction can be carefully controlled to provide the correct chain lengths and properties required for membrane formation and interaction with pores and other proteins.
[0261] Block copolymers can also be constructed from subunits not classified as lipid submaterials; for example, hydrophobic polymers can be made from siloxanes or other non-hydrocarbon-based monomers. The hydrophilic subsegments of the block copolymers can also possess low protein-binding properties, allowing for the creation of highly resistant membranes when exposed to pristine biological samples. This head group unit can also be derived from non-classical lipid head groups.
[0262] Compared to bio-lipid membranes, triblock copolymer membranes also exhibit increased mechanical and environmental stability, such as a much wider operating temperature or pH range. The synthetic properties of block copolymers provide a platform for customizing polymer-based membranes for a wide range of applications.
[0263] In some embodiments, the membrane is one of the membranes disclosed in International Application No. WO2014 / 064443 or WO2014 / 064444.
[0264] Amphiphilic molecules can be chemically modified or functionalized to facilitate the coupling of polynucleotides. The amphiphilic layer can be monolayer or bilayer. The amphiphilic layer is typically planar. The amphiphilic layer can be curved. The amphiphilic layer can be supportive.
[0265] Amphiphilic membranes are usually naturally mobile, essentially at about 10 -8 cm s -1 The lipid diffusion rate acts as a two-dimensional liquid. This means that pores and coupled polynucleotides can normally move within the amphiphilic membrane.
[0266] The membrane can be a lipid bilayer. Lipid bilayers are models of the cell membrane and serve as excellent platforms for a range of experimental studies. For example, lipid bilayers can be used for in vitro studies of membrane proteins via single-channel recording. Alternatively, lipid bilayers can be used as biosensors to detect the presence of a range of substances. A lipid bilayer can be any lipid bilayer. Suitable lipid bilayers include, but are not limited to, planar lipid bilayers, supported bilayers, or liposomes. Preferably, a flat lipid bilayer is used. Suitable lipid bilayers are disclosed in WO2008 / 102121, WO2009 / 077734, and WO2006 / 100484.
[0267] Methods for forming lipid bilayers are known in the art. Lipid bilayers are typically formed by the method of Montal and Mueller (Proc. Natl. Acad. Sci. USA., 1972; 69:3561-3566), in which a lipid monolayer is carried on an aqueous / air interface across an opening perpendicular to the interface. The lipid is typically added to the surface of an aqueous electrolyte solution by first dissolving the lipid in an organic solvent and then evaporating a drop of solvent from the surface of the aqueous solution on either side of the opening. Once the organic solvent has evaporated, the solution / air interface on either side of the opening physically moves back and forth through the opening until a bilayer is formed. Planar lipid bilayers can be formed across openings in a membrane or across openings in a groove.
[0268] The Montal and Mueller method is commonly used because it is cost-effective and a relatively straightforward method for forming a good-quality lipid bilayer suitable for protein pore insertion. Other common methods for bilayer formation include tip immersion, bilayer brushing, and patch clamping.
[0269] Tip-immersion bilayer formation requires contact between the open end surface (e.g., a pipette tip) and the surface of the test solution carrying the lipid monolayer. Similarly, a lipid monolayer is first generated at the solution / air interface by evaporating a drop of lipid dissolved in an organic solvent at the solution surface. The bilayer is then formed via a Langmuir-Schaefer process, requiring mechanical automation to move the open end relative to the solution surface.
[0270] For the brush-coated bilayer, a drop of lipid dissolved in an organic solvent is applied directly to an opening immersed in an aqueous test solution. Using a brush or equivalent, the lipid solution is thinly diffused within the opening. This solvent dilution allows the formation of a lipid bilayer. However, completely removing the solvent from the bilayer is very difficult, and therefore the bilayer formed by this method is less stable and more prone to noise during electrochemical measurements.
[0271] Patch clamping is commonly used in biological cell membrane research. The cell membrane is clamped to the end of a pipette by suction, and the membrane patch becomes attached within an opening. This method is suitable for generating lipid bilayers by clamping and then bursting liposomes to detach from the lipid bilayer sealed within an opening of the pipette. This method requires stable, large, monolayer liposomes and the fabrication of small openings in a material with a glass surface.
[0272] Liposomes can be formed by sonication, extrusion or the Mozafari method (Colas et al. (2007) Micron 38:841-847).
[0273] In some embodiments, a lipid bilayer is formed as described in International Application WO2009 / 077734. It is advantageous in this method to form the lipid bilayer from dried lipids. In a most preferred embodiment, the lipid bilayer is formed across an opening, as described in WO2009 / 077734.
[0274] A lipid bilayer is formed by two opposing lipid layers. The two lipid layers are arranged such that their hydrophobic tail groups face each other, forming a hydrophobic interior. The hydrophilic head groups of the lipids face outwards towards the aqueous environment on each side of the bilayer. The bilayer can exist in various lipid stages, including but not limited to liquid disordered stages (liquid sheets), liquid ordered stages, solid ordered stages (sheet-gel stages, interleaved gel stages), and planar bilayer crystals (sheet-subgel stages, sheet-crystalline stages).
[0275] Any lipid composition that forms a lipid bilayer can be used. The lipid composition is selected such that the lipid bilayer has desired properties, such as surface charge, ability to support membrane proteins, filling density, or the mechanical properties formed. The lipid composition may include one or more different lipids. For example, the lipid composition may contain up to 100 lipids. The lipid composition preferably contains 1 to 10 lipids. The lipid composition may include naturally occurring lipids and / or artificial lipids.
[0276] Lipids typically consist of a head group, an interfacial portion, and two hydrophobic tail groups that may be the same or different. Suitable head groups include (but are not limited to): neutral head groups, such as diacylglycerol esters (DG) and ceramides (CM); zwitterionic head groups, such as phosphatidylcholine (PC), phosphatidylethanolamine (PE), and sphingomyelin (SM); negatively charged head groups, such as phosphatidylglycerol (PG); phosphatidylserine (PS), phosphatidylinositol (PI), phosphatidic acid (PA), and cardiolipin (CA); and positively charged head groups, such as trimethylammonium propane (TAP). Suitable interfacial portions include, but are not limited to, naturally occurring interfacial portions, such as glycerol-based or ceramide-based portions. Suitable hydrophobic tail groups include, but are not limited to: saturated hydrocarbon chains, such as lauric acid (n-dodecanoic acid), myristic acid (n-tetradecanoic acid), palmitic acid (n-hexadecanoic acid), stearic acid (n-octadecanoic acid), and arachidic acid (n-eicosanoic acid); unsaturated hydrocarbon chains, such as oleic acid (cis-9-octadecanoic acid); and branched hydrocarbon chains, such as phytanoyl groups. The chain length and the position and number of double bonds in the unsaturated hydrocarbon chain can vary. The chain length and the position and number of branches (such as methyl groups) in the branched hydrocarbon chain can also vary. The hydrophobic tail group can be attached to the interfacial portion as an ether or ester. Lipids can be mycolic acids.
[0277] Lipids can also be chemically modified. The head or tail groups of lipids can be chemically modified. Suitable lipids with chemically modified head groups include, but are not limited to: PEG-modified lipids, such as 1,2-diacyl-sn-glycerol-3-phosphate ethanolamine-N-[methoxy(polyethylene glycol)-2000]; functionalized PEG lipids, such as 1,2-distearate-sn-glycerol-3-phosphate ethanolamine-N-[biotinyl(polyethylene glycol)2000]; and lipids for conjugation modification, such as 1,2-dioleoyl-sn-glycerol-3-phosphate ethanolamine-N-(succinyl) and 1,2-dispalmitoyl-sn-glycerol-3-phosphate ethanolamine-N-(biotinyl). Suitable lipids with chemically modified tails include, but are not limited to: polymerizable lipids, such as 1,2-bis(10,12-tetracarbadiynyl)-sn-glycerol-3-phosphate choline; fluorinated lipids, such as 1-palmitoyl-2-(16-fluoropalmitoyl)-sn-glycerol-3-phosphate choline; deuterated lipids, such as 1,2-dipalmitoyl-D62-sn-glycerol-3-phosphate choline; and ether-linked lipids, such as 1,2-di-O-phytyl-sn-glycerol-3-phosphate choline. Lipids may be chemically modified or functionalized to facilitate coupling with polynucleotides.
[0278] Amphiphilic layers, such as lipid compositions, typically include one or more additives that will affect the properties of the layer. Suitable additives include, but are not limited to: fatty acids, such as palmitic acid, myristic acid, and oleic acid; fatty alcohols, such as palmitol, myristicol, and oleyl alcohol; sterols, such as cholesterol, ergosterol, lanosterol, sitosterol, and stigmasterol; lysophospholipids, such as 1-acyl-2-hydroxy-sn-glycerol-3-phosphocholine; and ceramides.
[0279] In another embodiment, the membrane includes a solid layer. The solid layer can be formed of both organic and inorganic materials, including, but not limited to: microelectronic materials, insulating materials (such as Si3N4, Al2O3, and SiO), organic and inorganic polymers (such as polyamides), and plastics (such as...). The membrane may be composed of an amphiphilic membrane or layer, such as a two-component addition-cured silicone rubber, or glass. The solid layer may be formed from graphene. Suitable graphene layers are disclosed in WO2009 / 035647. If the membrane includes a solid layer, pores are typically present in the amphiphilic membrane or layer contained within the solid layer, for example, in pores, holes, gaps, channels, trenches, or slots within the solid layer. Suitable solid / amphiphilic hybrid systems can be prepared by those skilled in the art. Suitable systems are disclosed in WO2009 / 020682 and WO2012 / 005857. Any of the amphiphilic membranes or layers discussed above may be used.
[0280] The methods disclosed herein are typically carried out using: (i) an artificial amphiphilic layer comprising a pore, (ii) a separated, naturally occurring lipid bilayer comprising a pore, or (iii) a cell into which the pore is inserted. Artificial amphiphilic layers (such as artificial triblock copolymer layers) are typically used to perform the methods. The layers may include other transmembrane and / or intramembrane proteins and other molecules besides the pore. Suitable apparatus and conditions are discussed below. The disclosed methods are typically performed in vitro.
[0281] condition
[0282] As described above, and as further described in detail herein, some exemplary embodiments of this disclosure involve sequencing a target double-stranded nucleic acid as it moves (e.g., enters or passes through) relative to a transmembrane nanopore.
[0283] Characterization methods can be performed using any apparatus suitable for studying membrane / pore systems in which pores are inserted into the membrane. Characterization methods can be performed using any apparatus suitable for transmembrane pore sensing. For example, the apparatus may include a chamber containing an aqueous solution and a barrier dividing the chamber into two sections. The barrier typically has openings in which a pore-containing membrane is formed. Transmembrane pores are described herein.
[0284] The characterization method can be performed using the apparatus described in WO2008 / 102120, WO2010 / 122293 or WO00 / 28312.
[0285] Characterization methods may involve measuring the ion current flowing through the pore, typically by measuring the current. Alternatively, the ion current through the pore may be measured optically, as disclosed in Heron et al.: J. Am. Chem. Soc. 9 Vol. 131, No. 5, 2009. Therefore, the apparatus may also include circuitry capable of applying a potential and measuring the electrical signal across the membrane and the pore. Characterization methods may be performed using patch clamps or voltage clamps. Characterization methods preferably involve the use of voltage clamps.
[0286] Characterization methods can be performed on silicon-based aperture arrays, each array comprising 128, 256, 512, 1024, 2000, 3000, 4000, 6000, 10000, 12000, 15000 or more apertures.
[0287] Characterization methods may involve measuring the current flowing through the pore. This method is typically performed with a voltage applied across the membrane and through the pore. The voltage used is typically +2V to -2V, and usually -400mV to +400mV. The voltage used is preferably within a range having a lower limit selected from -400mV, -300mV, -200mV, -150mV, -100mV, -50mV, -20mV, and 0mV, and the upper limit is independently selected from +10mV, +20mV, +50mV, +100mV, +150mV, +200mV, +300mV, and +400mV. More preferably, the voltage used is within the range of 100mV to 240mV, and most preferably within the range of 120mV to 220mV. By using an increased applied potential, the distinguishability between different nucleotides can be increased through the pore.
[0288] Characterization methods are typically performed in the presence of any charge carriers, such as metal salts, for example alkali metal salts; halide salts, such as chloride salts, like alkali metal chloride salts. The charge carriers may comprise ionic liquids or organic salts, such as tetramethylammonium chloride, trimethylphenylammonium chloride, phenyltrimethylammonium chloride, or 1-ethyl-3-methylimidazolium chloride. In the exemplary apparatus discussed above, the salt is present in an aqueous solution within the chamber. Potassium chloride (KCl), sodium chloride (NaCl), or cesium chloride (CsCl) is typically used. KCl is preferred. The salt may be an alkaline earth metal salt, such as calcium chloride (CaCl2). The salt concentration may be saturated. The salt concentration may be 3 M or lower, and is typically 0.1 M to 2.5 M, 0.3 M to 1.9 M, 0.5 M to 1.8 M, 0.7 M to 1.7 M, 0.9 M to 1.6 M, or 1 M to 1.4 M. The salt concentration is preferably 150 mM to 1 M. The characterization method preferably uses a salt concentration of at least 0.3 M, such as at least 0.4 M, at least 0.5 M, at least 0.6 M, at least 0.8 M, at least 1.0 M, at least 1.5 M, at least 2.0 M, at least 2.5 M, or at least 3.0 M. High salt concentrations provide a high signal-to-noise ratio and allow identification of currents indicating binding / non-binding against a background of normal current fluctuations.
[0289] Characterization methods are typically performed in the presence of a buffer solution. In the exemplary apparatus discussed above, the buffer solution is present in an aqueous solution within the chamber. Any suitable buffer solution can be used. Typically, the buffer solution is HEPES. Another suitable buffer solution is Tris-HCl buffer. The methods are typically performed at the following pH values: 4.0 to 12.0, 4.5 to 10.0, 5.0 to 9.0, 5.5 to 8.8, 6.0 to 8.7, or 7.0 to 8.8, or 7.5 to 8.5. The pH value used is preferably about 7.5.
[0290] Characterization methods can be performed at the following temperatures: 0°C to 100°C, 15°C to 95°C, 16°C to 90°C, 17°C to 85°C, 18°C to 80°C, 19°C to 70°C, or 20°C to 60°C. Characterization methods are typically performed at room temperature. Optionally, characterization methods can be performed at temperatures that support enzyme function, such as approximately 37°C.
[0291] Other methods
[0292] This article also discloses a method for sequencing target double-stranded nucleic acids, which includes:
[0293] (a') Interstrand crosslinks are formed between the first and second strands of the target double-stranded nucleic acid by exposing the target double-stranded nucleic acid to a crosslinking agent;
[0294] (b') Break the target double-stranded nucleic acid at or near the interstrand crosslinking site to form at least one double-stranded nucleic acid construct crosslinked at or near the end of the construct; and
[0295] (c') The cross-linked construct formed in (b) was sequenced using single-molecule sequencing technology.
[0296] Preferred embodiments of such methods include contacting the cross-linked construct with an enzyme that sequentially processes the first and second strands of the terminally cross-linked construct formed in (a'). In some embodiments, this provides orthogonal alignment sequence information. In some embodiments, the enzyme processes the cross-links. In some embodiments, the enzyme does not process the cross-links. In some embodiments, the enzyme is a polynucleotide processing enzyme disclosed herein.
[0297] Such methods are particularly suitable for sequencing cross-linked constructs using sequencing-by-synthesis techniques, such as real-time single-molecule sequencing based on the detection of base incorporation catalyzed by enzymes such as DNA polymerase. This article describes suitable polymerases.
[0298] In such methods, double-stranded nucleic acids, interstrand crosslinks, and sequencing techniques are typically as described herein. Breakage of double-stranded nucleic acids is described in more detail herein.
[0299] Construct
[0300] This document also provides cross-linked nucleic acid constructs formed by exposing double-stranded nucleic acids to a cross-linking agent as described herein. The double-stranded nucleic acids and cross-linking agents are preferably as described in more detail herein.
[0301] The constructs provided herein may further include one or more adaptors and / or anchors and / or spacers as described herein. The constructs may include polynucleotide processing enzymes attached thereto as described herein.
[0302] system
[0303] Systems comprising cross-linked nucleic acid constructs formed by exposing double-stranded nucleic acids to cross-linking agents as described herein; one or more polynucleotide processing enzymes as described herein; and one or more nanopores are also provided. Such systems are typically provided for characterizing double-stranded nucleic acids.
[0304] In one embodiment, a system for characterizing a target double-stranded nucleic acid is provided, comprising:
[0305] - A construct formed by exposing a target double-stranded nucleic acid to a cross-linking agent;
[0306] - Polynucleotide processing enzymes; and
[0307] - Nanopores, which are used to characterize target polynucleotides as they move relative to the nanopores.
[0308] Typically, constructs, enzymes, and nanopores are described in more detail in this article.
[0309] Reagent test kit
[0310] Kits containing one or more cross-linking agents and one or more polynucleotide processing enzymes are also provided. In some embodiments, the kits provided herein also contain one or more adaptors and / or anchors and / or spacers as described herein. In some embodiments, the cross-linking agents and polynucleotide processing enzymes are generally as described herein.
[0311] In some embodiments, the kit further includes a detector for sequencing constructs formed by contacting a target double-stranded nucleic acid with a cross-linking agent. In some embodiments, the detector is a nanopore.
[0312] In some embodiments, the kit also includes reagents such as buffer solutions, fuel molecules (e.g., those mentioned above, such as nucleoside triphosphates, such as ATP), and instructions for use.
[0313] This kit can be configured for use with an algorithm, also provided herein, adapted to run on a computer system. The algorithm can be adapted to detect sequence information from the first strand of a double-stranded nucleic acid attached to the second strand of the double-stranded nucleic acid via interstrand crosslinking, and selectively process signals obtained, for example, during sequential sequencing of the first and second strands using a nanopore sensor, to obtain orthogonal alignment sequence information. A system including a computing device configured to detect sequence information from the first strand of a double-stranded nucleic acid attached to the second strand of the double-stranded nucleic acid via interstrand crosslinking, and selectively process signals obtained, for example, during sequential sequencing of the first and second strands using a nanopore sensor, to obtain orthogonal alignment sequence information. In some embodiments, the system includes a receiving device for receiving data from detecting the first and second strands of the crosslinked double-stranded nucleic acid, a processing device for processing signals obtained during sequential sequencing of the first and second strands, and an output device for outputting alignment sequence information.
[0314] It should be understood that although specific embodiments, configurations, and materials and / or molecules according to the invention have been discussed herein, various changes or modifications in form and detail may be made without departing from the scope and spirit of the invention. The following examples are provided to better illustrate specific embodiments and should not be construed as limiting the scope of this application. This application is limited only by the claims.
[0315] Example
[0316] This embodiment demonstrates that double-stranded nucleic acids can be cross-linked according to the methods disclosed herein, and the resulting interstrand-crosslinked constructs can be sequenced using single-molecule sequencing.
[0317] Two λ genomic DNA samples were prepared for sequencing. Sample A was a control sample containing 1000 ng of λ DNA (obtained from New England Biolabs, product NEB N3013), which was not cross-linked prior to sequencing. Sample B, also containing 1000 ng of λ DNA (NEB N3013), was incubated on ice under a 365 nm UV lamp for 8 minutes.
[0318] Two samples were prepared separately for DNA sequencing. First, following the manufacturer's instructions, using... Ultra TM The DNA sample was end-repaired and dA-tailed using the II End Repair / dA-Tailing Module (New England Biolabs, product NEB E7546). The sample was then purified with 1x SPRI using Agencourt AMPure beads (Beckman Coulter) following the manufacturer's instructions.
[0319] During sample preparation, especially in the dA tailing and SPRI steps, the chain is subjected to pipetting shear.
[0320] Then, following the manufacturer's instructions, sequence the sample using the Oxford Nanopore Technologies sequencing kit (Oxford Nanopore Technologies, product SQK-LSK108). This ligates the sequencing adaptor preloaded with E8 motors to the dA-tailed ends of the dsDNA in the sample. Then, following the manufacturer's instructions, run the prepared sample on an Oxford Nanopore MinION flow cell using the baseline sequencing protocol.
[0321] The relationship between current and time was analyzed from a large amount of raw data sequencing files. Figure 5 A typical example of current (y-axis, in pA) versus time (x-axis, in seconds) data is shown, highlighting events corresponding to the sequencing chain. Figure 5 An example of a sequencing template-complement construct is shown, where the portion labeled "A" represents the template and "B" represents the linked reverse complement.
[0322] The presence of a template and an inverse complementary moiety in a strand can sometimes be identified by changes in the current signal between the two moieties, which may occur when the template and inverse complementary moieties re-anneal on the other side of the nanopore during translocation. Changes in the current level can provide a useful indication for sequencing two linked strands. However, in the method disclosed herein, it is not necessary to change the current level. Depending on the sequence of the double-stranded nucleic acid being sequenced in the method disclosed herein, changes in the current level may or may not be observed when sequencing constructs with interstrand crosslinks. Therefore, while changes in the current level can be a useful indicator, the absence of a change in the current level is not an indicator that the method is invalid.
[0323] Data corresponding to individual sequencing strand events were extracted and identified using the Oxford Nanopore Guppy base recognizer, and aligned with a complete λ phage genome reference using standard sequence alignment tools. Strands with both the aligned template and the reverse complement were identified as “2D,” while strands with only the template portion were identified as “1D.” The percentage of strands identified as 2D in the crosslinked and control samples is shown below. Clearly, the percentage of strands identified as 2D is significantly increased by crosslinking double-stranded nucleic acid samples according to the method presented herein.
[0324]
[0325] Figure 6The diagram illustrates signal and sequence alignment of the template and the inverse complementary portion of the link in some example 2D chains. As mentioned above, the difference between the template and the inverse complementary signal is usually evident in current versus time signals. However, even without observing this change in the current signal, the base alignment confirms that the chain simultaneously contains both the template and the inverse complementary portion.
[0326] Sequence List Description
[0327] SEQ ID NO:1 shows the amino acid sequence of (hexahistine-labeled) exonuclease I (EcoExo I) from Escherichia coli.
[0328] SEQ ID NO:2 shows the amino acid sequence of exonuclease III from Escherichia coli.
[0329] SEQ ID NO:3 shows the amino acid sequence of the RecJ enzyme (TthRecJ-cd) of T. thermophilus.
[0330] SEQ ID NO:4 shows the amino acid sequence of the phage λ exonuclease. This sequence is one of three identical subunits that assemble into the trimer. (http: / / www.neb.com / nebecomm / products / productM0262.asp).
[0331] SEQ ID NO:5 shows the amino acid sequence of Phi29 DNA polymerase from Bacillus subtilis.
[0332] SEQ ID NO:6 shows the amino acid sequence of the Trwc Cba (Citromicrobium bathyomarinum) helicase.
[0333] SEQ ID NO:7 shows the amino acid sequence of the helicase from Hel308 Mbu (Methanococcoides burtonii).
[0334] SEQ ID NO:8 shows the amino acid sequence of Dda helicase 1993 from Enterobacter T4 bacteriophage.
[0335] sequence list
[0336] SEQ ID NO:1 - Exonuclease I from Escherichia coli
[0337] MMNDGKQQSTFLFHDYETFGTHPALDRPAQFAAIRTDSEFNVIGEPEVFYCKPADDYLPQPGAVLITGITPQEARAKGENEAAFAARIHSLFTVPKTCILGYNNVRFDDEVTRNIFYRNFYDPYAWSWQHDNSRWDLLDVMRACYALRPEGINWPENDDGLPSFRLEHLTKANGIEHSNAHDAMADVYATIAMAKLVKTRQPRLFDYLFTHRNKHKLMALIDVPQMKPLVHVSGMFGAWRGNTSWVAPLAWHPENRNAVIMVDLAGDISPLLELDSDTLRERLYTAKTDLGDNAAVPVKLVHINKCPVLAQANTLRPEDADRLGINRQHCLDNLKILRENPQVREKVVAIFAEAEPFTPSDNVDAQLYNGFFSDADRAAMKIVLETEPRNLPALDITFVDKRIEKLLFNYRARNFPGTLDYAEQQRWLEHRRQVFTPEFLQGYADELQMLVQQYADDKEKVALLKALWQYAEEIVSGSGHHHHHH
[0338] SEQ ID NO: 2 - Exonuclease III from Escherichia coli
[0339] MKFVSFNINGLRARPHQLEAIVEKHQPDVIGLQETKVHDDMFPLEEVAKLGYNVFYHGQKGHYGVALLTKETPIAVRRGFPGDDEEAQRRIIMAEIPSLLGNVTVINGYFPQGESRDHPIKFPAKAQFYQNLQNYLETELKRDNPVLIMGDMNISPTDLDIGIGEENRKRWLRTGKCSFLPEEREWMDRLMSWGLVDTFRHANPQTADRFSWFDYRSKGFDDNRGLRIDLLLASQPLAECCVETGIDYEIRSMEKPSDHAPVWATFRR
[0340] SEQ ID NO:3 - RecJ enzyme from Thermus thermophilus MFRRKEDLDPPLALLPLKGLREAAALLEEALRQGKRIRVHGDYDADGLTGTAILVRGLAALGADVHPFIPHRLEEGYGVLMERVPEHLEASDLFLTVDCGITNHAELRELLENGVEVIVTDHHTPGKTPPPGLVVHPALTPDLKEKPTGAGVAFLLLWALHERLGLPPPLEYADLAAVGTIADVAPLWGWNRALVKEGLARIPASSWVGLRLLAEAVGYTGKAVEVAFRIAPRINAASRLGEAEKALRLLLTDDAAEAQALVGELHRLNARRQTLEEAMLRKLLPQADPEAKAIVLLDPEGHPGVMGIVASRILEATLRPVFLVAQGKGTVRSLAPISAVEALRSAEDLLLRYGGHKEAAGFAMDEALFPAFKARVEAYAARFPDPVREVALLDLLPEPGLLPQVFRELALLEPYGEGNPEPLFL
[0341] SEQ ID NO:4 - Exonuclease of bacteriophage λ
[0342] MTPDIILQRTGIDVRAVEQGDDAWHKLRLGVITASEVHNVIAKPRSGKKWPDMKMSYFHTLLAEVCTGVAPEVNAKALAWGKQYENDARTLFEFTSGVNVTESPIIYRDESMRTACSPDGLCSDGNGLELKCPFTSRDFMKFRLGGFEAIKSAYMAQVQYSMWVTRKNAWYFANYDPRMKREGLHYVVIERDEKYMASFDEIVPEFIEKMDEALAEIGFVFGEQWR
[0343] SEQ ID NO: DNA polymerase
[0344] MKHMPRKMYSCAFETTTKVEDCRVWAYGYMNIEDHSEYKIGNSLDEFMAWVLKVQADLYFHNLKFDGAFIINWLERNGFKWSADGLPNTYNTIISRMGQWYMIDICLGYKGKRKIHTVIYDSLKKLPFPVKKIAKDFKLTVLKGDIDYHKERPVGYKITPEEYAYIKNDIQIIAEALLIQFKQGLDRMTAGSDSLKGFKDIITTKKFKKVFPTLSLGLDKEVRYAYRGGFTWLNDRFKEKEIGEGMVFDVNSLYPAQMYSRLLPYGEPIVFEGKYVWDEDYPLHIQHIRCEFELKEGYIPTIQIKRSRFYKGNEYLKSSGGEIADLWLSNVDLELMKEHYDLYNVEYISGLKFKATTGLFKDFIDKWTYIKTTSEGAIKQLAKLMLNSLYGKFASNPDVTGKVPYLKENGALGFRLGEEETKDPVYTPMGVFITAWARYTTITAAQACYDRIIYCDTDSIHLTGTEIPDVIKDIVDPKKLGYWAHESTFKRAKYLRQKTYIQDIYMKEVDGKLVEGSPDDYTDIKFSVKCAGMTDKIKKEVTFENFKVGFSRKMKPKPVQVPGGVVLVDDTFTIKSGGSAWSHPQFEKGGGSGGGSGGSAWSHPQFEK
[0345] SEQ ID NO:6 - Trwc Cba Helicase
[0346] MLSVANVRSPSAAASYFASDNYYASADADRSGQWIGDGAKRLGLEGKVEARAFDALLRGELPDGSSVGNPGQAHRPGTDLTFSVPKSWSLLALVGKDERIIAAYREAVVEALHWAEKNAAETRVVEKGMVVTQATGNLAIGLFQHDTNRNQEPNLHFHAVIANVTQGKDGWRTLKNDRLWQLNTTLNSIAMARFRVAVEKLGYEPGPVLKHGNFEARGISREQVMAFSTRRKEVLEARRGPGGLDAGRIAALDTRASKEGIEDRATLSKQWSEAAQSIGLDLKPLVDRARTKALGQGMEATRIGSLVERGRAWLSRFAAHVRGDPADPLVPPSVLKQDRQTIAAAQAVASAVRHLSQREAAFERTALYKAALDFGLPTTIADVEKRTRALVRSGDLIAGKGEHKGWLASRDAVVTEQRILSEVAAGKGDSSPAITPQKAAASVQAAALTGQGFRLNEGQLAAARLILISKDRTIAVQGIAGKS SVLKPVAEVLRDEGHPVIGLAIQNTLVQMLERDTGIGSQTLARFLGGWNKLLDDPGNVALRAEAQASLKDHVLVLDEASMVSNEDKEKLVRLANLAGVHRLVLIGDRKQLGAVDAGKPFALLQRAGIARAEMATNLRARDPVVREAQAAAQAGDVRKARLHLKSHTVEARGDGAQVAAETWLALDKETRARTSIYASGRAIRSAVNAAVQQGLLASREIGPAKMKLEVLDRVNTTREELRHLPAYRAGRVLEVSRKQQALGLFIGEYRVIGQDRKGKLVEVEDKRGKRFRFDPARIRAGKGDDNLTLLEPRKLEIHEGDRIRWTRNDHRRGLFNADQARVVEIANGKVTFETSKGDLVELKKDDPMLKRIDLAYALNVHMAQGLTSDRGIAVMDSRERNLSNQKTFLVTVTRLRDHLTLVVDSADKLGAAVARNKGEKASAIEVTGSVKPTATKGSGVDQPKSVEANKAEKELTRSKSKTLDFGI
[0347] SEQ ID NO:7 - Hel308 Mbu helicase MMIRELDIPRDIIGFYEDSGIKELYPPQAEAIEMGLLEKKNLLAAIPTASGKTLLAELAMIKAIREGGKALYIVPLRALASEKFERFKELAPFGIKVGISTGDLDSRADWLGVNDIIVATSEKTDSLLRNGTSWMDEITTVVVDEIHLLDSKNRGPTLEVTITKLMRLNPDVQVVALSATVGNAREMADWLGAALVLSEWRPTDLHEGVLFGDAINFPGSQKKIDRLEKDDAVNLVLDTIKAEGQCLVFESSRRNCAGFAKTASSKVAKILDNDIMIKLAGIAEEVESTGETDTAIVLANCIRKGVAFHHAGLNSNHRKLVENGFRQNLIKVISSTPTLAAGLNLPARRVIIRSYRRFDSNFGMQPIPVLEYKQMAGRAGRPHLDPYGESVLLAKTYDEFAQLMENYVEADAEDIWSKLGTENALRTHVLSTIVNGFASTRQELFDFFGATFFAYQQDKWMLEEVINDCLEFLIDKAMVSETEDIEDASKLFLRGTRLGSLVSMLYIDPLSGSKIVDGFKDIGKSTGGNMGSLEDDKGDDITVTDMTLLHLVCSTPDMRQLYLRNTDYTIVNEYIVAHSDEFHEIPDKLKETDYEWFMGEVKTAMLLEEWVTEVSAEDITRHFNVGEGDIHALADTSEWLMHAAAKLAELLGVEYSSHAYSLEKRIRYGSGLDLMELVGIRGVGRVRARKLYNAGFVSVAKLKGADISVLSKLVGPKVAYNILSGIGVRVNDKHFNSAPISSNTLDTLLDKNQKTFNDFQ
[0348] SEQ ID NO:8 - Dda helicase
[0349] MTFDDLTEGQKNAFNIVMKAIKEKKHHVTINGPAGTGKTTLTKFIIEALISTGETGIILAAPTHAAKKILSKLSGKEASTIHSILKINPVTYEENVLFEQKEVPDLAKCRVLICDEVSMYDRKLFKILLSTIPPWCTIIGIGDNKQIRPVDPGENTAYISPFFTHKDFYQCELTEVKRSNAPIIDVATDVRNGKWIYDKVVDGHGVRGFTGDTALRDFMVNYFSIVKSLDDLFENRVMAFTNKSVDKLNSIIRKKIFETDKDFIVGEIIVMQEPLFKTYKIDGKPVSEIIFNNGQLVRIIEAEYTSTFVKARGVPGEYLIRHWDLTVETYGDDEYYREKIKIISSDEELYKFNLFLGKTAETYKNWNKGGKAPWSDFWDAKSQFSKVKALPASTFHKAQGMSVDRAFIYTPCIHYADVELAQQLLYVGVTRGRYDVFYV
Claims
1. A method for sequencing a target double-stranded nucleic acid, comprising: (a) By exposing the target double-stranded nucleic acid to a crosslinking agent, an interstrand crosslink is formed between the first and second strands of the target double-stranded nucleic acid; This results in the formation of at least one cross-linked double-stranded nucleic acid construct; and (b) Sequencing the cross-linked construct formed in (a) using single-molecule sequencing technology. The method further includes allowing the interstrand crosslinks formed in step (a) or near the site of the crosslinks to break, such that the interstrand crosslinks are located at or near the end of the construct.
2. The method according to claim 1, wherein, Fractures at or near the inter-chain crosslinking points occur naturally due to deformation.
3. The method according to claim 1, wherein, The breakage at or near the inter-chain crosslinking points is caused by physical agitation.
4. The method according to any one of the preceding claims, wherein, The single-molecule sequencing technology includes contacting the cross-linked construct with an enzyme that sequentially processes the first and second strands of the cross-linked construct.
5. The method according to claim 4, wherein, The sequencing technology provides orthogonal proofreading sequence information.
6. The method according to any one of claims 1-3, wherein, The crosslinking agent is or contains electromagnetic radiation.
7. The method according to any one of claims 1-3, wherein, The crosslinking agent is or contains a chemical reagent.
8. The method according to any one of claims 1-3, wherein, The cross-linking agent is or contains a nucleic acid cross-linking enzyme.
9. The method according to any one of claims 1-3, wherein, The interstrand crosslinks are formed at random locations on the target double-stranded nucleic acid.
10. The method according to any one of claims 1-3, wherein, The interstrand crosslinks are targeted to specific base compositions on the target double-stranded nucleic acid.
11. The method according to any one of claims 1-3, wherein, The inter-strand crosslinks are formed between the nucleobases in the first and second strands of the target double-stranded nucleic acid.
12. The method according to any one of claims 1-3, wherein, The inter-chain crosslinks are formed between the sugar groups in the first and second strands of the target double-stranded nucleic acid.
13. The method according to any one of claims 1-3, wherein, The interchain crosslinks are formed between the nucleobases in the first and second strands of the target double-stranded nucleic acid and between the glycogroups in the first and second strands of the target double-stranded nucleic acid.
14. The method according to any one of claims 1-3, wherein, The inter-chain crosslinks are formed between the nucleobase adducts in the first and second strands of the target double-stranded nucleic acid.
15. The method according to any one of claims 1-3, wherein, The crosslinking is located within 100 bases at the end of the crosslinked construct formed in part (b) of claim 1.
16. The method according to any one of claims 1-3, wherein, The single-molecule sequencing technology includes sequencing using nanopore sensors.
Citation Information
Patent Citations
phi 29 DNA polymerase
US5576204A
A miniature support for thin films containing single channels or nanopores and methods for using same
WO2000028312A1
Deliver of molecules to a li id bila
WO2006100484A2
Lipid bilayer sensor system
WO2008102120A1
Formation of lipid bilayers
WO2008102121A1