Method for assembling a protein having multiple subunits
By utilizing the time difference in association between labeled nucleotides and polymerase and nanopore detection, the problem of insufficient differentiation between purine and pyrimidine nucleotides in nucleic acid sequencing has been solved, achieving high sensitivity and high accuracy in nucleic acid sequence identification.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- F HOFFMANN LA ROCHE & CO AG
- Filing Date
- 2013-11-07
- Publication Date
- 2026-05-05
AI Technical Summary
Existing nucleic acid sequencing methods lack sufficient sensitivity in distinguishing between purine and pyrimidine nucleotides, resulting in low accuracy and efficiency of sequence information, which is insufficient to meet the needs of diagnosis and treatment.
Nucleic acid sequencing is performed using labeled nucleotides through nanopores. By utilizing the time difference in association between labeled nucleotides and polymerase, incorporated and unincorporated nucleotides are detected and distinguished through nanopores. Combined with alternating current waveforms and electrode detection, accurate identification of nucleotides is achieved.
It improves the sensitivity and accuracy of nucleic acid sequencing, enabling the differentiation between incorporated and unincorporated labeled nucleotides with an accuracy of over 99%, thereby enhancing the efficiency and precision of nucleic acid sequence identification.
Smart Images

Figure CN112480218B_ABST
Abstract
Description
[0001] This application is a divisional application of the invention patent application filed on November 7, 2013, with application number 201380058224.1 and invention title "Using Labeled Nucleic Acid Sequencing".
[0002] Cross-referencing
[0003] This application claims the benefit of U.S. Provisional Patent Application Serial No. 61 / 724,869, filed November 9, 2012; U.S. Provisional Patent Application Serial No. 61 / 737,621, filed December 14, 2012; and U.S. Provisional Patent Application Serial No. 61 / 880,407, filed September 20, 2013, each of which is incorporated herein by reference in its entirety. Technical Field
[0004] This invention relates to the field of biotechnology, and more particularly to a method for assembling proteins having multiple subunits. Background Technology
[0005] Nucleic acid sequencing is a process that can provide sequence information from nucleic acid samples. This sequence information can aid in the diagnosis and / or treatment of subjects. For example, a subject's nucleic acid sequence can be used to identify and diagnose genetic diseases and potentially develop treatments for them. As another example, research on pathogens can lead to treatments for infectious diseases.
[0006] There are available methods for sequencing nucleic acids. However, such methods are expensive and may not provide sequence information for a certain period of time and with the accuracy required for diagnosing and / or treating subjects. Summary of the Invention
[0007] Methods for sequencing single-stranded nucleic acid molecules through nanopores may have insufficient or inadequate sensitivity to provide signals for diagnostic and / or therapeutic purposes. The nucleic acid bases that make up nucleic acid molecules (e.g., adenine (A), cytosine (C), guanine (G), thymine (T), and / or uracil (U)) may not provide signals that are significantly different from each other. In particular, purines (i.e., A and G) have similar size, shape, and charge and provide signals that are not significantly different in some cases. Additionally, pyrimidines (i.e., C, T, and U) have similar size, shape, and charge and provide signals that are not significantly different in some cases. This paper recognizes the need for improved methods for nucleic acid molecule recognition and sequencing.
[0008] In some implementations, nucleotide incorporation events (e.g., incorporation of a nucleotide into a nucleic acid strand complementary to the template strand) present a label to the nanopore and / or release the label from the nucleotide detected through the nanopore. The incorporated base (i.e., A, C, G, T, or U) can be identified because a unique label is released and / or presented for each type of nucleotide (i.e., A, C, G, T, or U).
[0009] In some implementations, the label is attributed to a successfully incorporated nucleotide based on the time period during which the label interacts with the nanopore. This time period can be longer than the time period associated with the free flow of the nucleotide label through the nanopore. The detection time period for successfully incorporated nucleotide labels can also be longer than the detection time period for unincorporated nucleotides (e.g., nucleotides mismatched with the template strand).
[0010] In some cases, the polymerase associates with the nanopore (e.g., covalently linked to the nanopore) and the polymerase performs a nucleotide incorporation event. When a labeled nucleotide associates with the polymerase, the label can be detected through the nanopore. In some cases, unincorporated labeled nucleotides pass through the nanopore. The method can distinguish between labels associated with unincorporated nucleotides and labels associated with incorporated nucleotides based on the length of time it takes to detect the labeled nucleotide through the nanopore. In one embodiment, unincorporated nucleotides are detected through the nanopore in less than about 1 millisecond and incorporated nucleotides are detected through the nanopore in at least about 1 millisecond.
[0011] In some implementations, the polymerase has a slow kinetic step, wherein the label can be detected through a nanopore within at least 1 millisecond, with an average detection time of approximately 100 ms. The polymerase may be a mutant phi29 DNA polymerase.
[0012] Polymerases can be mutated to reduce the rate at which they incorporate nucleotides into nucleic acid chains (e.g., growing nucleic acid chains). In some cases, the rate of nucleotide incorporation into nucleic acid chains can be reduced by functionalizing nucleotides and / or the template chain to provide steric hindrance (e.g., by methylating the template nucleic acid chain). In other cases, the rate is reduced by incorporating methylated nucleotides.
[0013] On the one hand, a method for sequencing nucleic acid samples using nanopores adjacent to sensing electrodes in a membrane includes: (a) providing a labeled nucleotide to a reaction chamber containing a nanopore, wherein a single labeled nucleotide of the labeled nucleotide contains a label coupled to the nucleotide, the label being detectable by means of the nanopore; (b) performing a polymerization reaction using a polymerase to incorporate a single labeled nucleotide of the labeled nucleotide into a growth chain complementary to a single-stranded nucleic acid molecule from the nucleic acid sample; and (c) detecting the label associated with the single labeled nucleotide by means of the nanopore during and / or at the time of incorporation of the single labeled nucleotide, wherein the label is detected by means of the nanopore when the nucleotide associates with the polymerase.
[0014] In some implementations, the marker is detected multiple times when associated with polymerase.
[0015] In some implementations, the electrodes are recharged between labeling detection periods.
[0016] In some implementations, the method distinguishes between incorporated and unincorporated labeled nucleotides based on the time it takes to detect the labeled nucleotide through the nanopore.
[0017] In some embodiments, the ratio of the time for detecting incorporated labeled nucleotides via nanopores to the time for detecting unincorporated labeled nucleotides via nanopores is at least about 1.5, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 100, 1000, or 10,000.
[0018] In some embodiments, the ratio of the time period during which the label associated with the incorporated nucleotide interacts with the nanopore (and by means of nanopore detection) to the time period during which the label associated with the unincorporated nucleotide interacts with the nanopore is at least about 1.5, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 100, 1000 or 10,000.
[0019] In some implementations, the nucleotide associates with the polymerase for at least an average (or mean) time period of about 1 millisecond.
[0020] In some implementations, when the nucleotide does not associate with the polymerase, the labeled nucleotide passes through the nanopore in less than 1 millisecond (ms).
[0021] In some implementations, the markers have a selected length that can be detected by nanopores.
[0022] In some implementations, the incorporation of the first labeled nucleotide does not interfere with the detection of the labeled nanopore associated with the second labeled nucleotide.
[0023] In some implementations, the detection of the labeled nanopore associated with the first labeled nucleotide does not interfere with the incorporation of the second labeled nucleotide.
[0024] In some implementations, the nanopores are able to distinguish between incorporated and unincorporated labeled nucleotides with at least 95% accuracy.
[0025] In some implementations, the nanopores are able to distinguish between incorporated and unincorporated labeled nucleotides with at least 99% accuracy.
[0026] In some implementations, the tag associated with the single tagged nucleotide is detected when the tag is released from the single tagged nucleotide.
[0027] On the one hand, a method for sequencing nucleic acid samples using nanopores adjacent to sensing electrodes in a membrane includes: (a) providing a labeled nucleotide to a reaction chamber containing a nanopore, wherein a single labeled nucleotide of the labeled nucleotide contains a label coupled to the nucleotide, the label being detectable by means of the nanopore; (b) incorporating a single labeled nucleotide of the labeled nucleotide into a growth chain complementary to a single-stranded nucleic acid molecule derived from the nucleic acid sample by means of an enzyme; and (c) during the incorporation of the single labeled nucleotide, distinguishing by means of the nanopore a label associated with the single labeled nucleotide and one or more labels associated with one or more unincorporated single labeled nucleotides.
[0028] In some embodiments, the enzyme is a nucleic acid polymerase or any enzyme that can extend new synthetic chains based on template polymers.
[0029] In some implementations, the single labeled nucleotide incorporated in (b) is distinguished from the single labeled nucleotide not incorporated in (b) based on the time length and / or time ratio of detecting the single labeled nucleotide incorporated in (b) and the single labeled nucleotide not incorporated in (b) by means of nanopores.
[0030] On the one hand, a method for sequencing nucleic acids using nanopores in a membrane includes: (a) providing a labeled nucleotide to a reaction chamber containing a nanopore, wherein a single labeled nucleotide of the labeled nucleotide contains a label that can be detected by the nanopore; (b) incorporating the labeled nucleotide into a growing nucleic acid chain, wherein the label associated with a single labeled nucleotide of the labeled nucleotide remains at or near at least a portion of the nanopore during incorporation, wherein the ratio of the time during which the incorporated labeled nucleotide can be detected by the nanopore to the time during which the unincorporated label can be detected by the nanopore is at least 1.1, 1.2, 1.3, 1.4, 1.5, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 1000, or 10,000; and (c) detecting the label using the nanopore.
[0031] In some embodiments, the ratio of the time when the incorporated labeled nucleotide can be detected by the nanopore to the time when the unincorporated label can be detected by the nanopore is at least about 1, 1.2, 1.3, 1.4, 1.5, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 1000, or 10,000.
[0032] In some implementations, the label remains associated with a single nucleotide when incorporated into the nucleotide.
[0033] In some implementations, a tag associated with a single nucleotide is released during nucleotide incorporation.
[0034] In some implementations, the method further includes removing labels from nanopores.
[0035] In some implementations, the markers are expelled in the opposite direction to the direction in which they entered the nanopores.
[0036] In some implementations, the label remains in the nanopore for at least about 100 ms.
[0037] In some implementations, the label remains in the nanopore for at least about 10 ms.
[0038] In some implementations, the label remains in the nanopore for at least about 1 ms.
[0039] In some implementations, the labeled nucleotide is incorporated at a rate of up to about one nucleotide per second.
[0040] In some implementations, labeled molecules are ejected from nanopores using voltage pulses.
[0041] In some implementations, at least 99% of the labeled molecules are likely to be ejected by voltage pulses.
[0042] In some implementations, the nanopore evicts the labeled molecules over a certain period of time so that the two labeled molecules do not coexist in the nanopore.
[0043] In some implementations, the probability of the nanopore expelling the labeled molecules within a certain time period so that two labeled molecules are simultaneously present in the nanopore is at most 1%.
[0044] In some implementations, the marker has a diameter of less than about 1.4 nm.
[0045] In some implementations, each tag associated with the nucleotide is detected by means of a nanopore when the tag is attached to the incorporated nucleotide.
[0046] In some implementations, the tag associated with the single tagged nucleotide is detected when the tag is released from the single tagged nucleotide.
[0047] On one hand, a chip for sequencing nucleic acid samples includes: a plurality of nanopores, wherein said plurality of nanopores have at least one nanopore arranged in a membrane adjacent to or near an electrode, wherein each nanopore detects a label associated with a single labeled nucleotide during the incorporation of a labeled nucleotide into a growing nucleic acid chain. In some embodiments, the nanopores are individually addressable.
[0048] In some implementations, a label associated with a nucleotide is detected in a single nanopore during subsequent label passage or during the passage of a label adjacent to the nanopore.
[0049] In some embodiments, the chip includes at least 500 individually addressable electrodes per square millimeter. In other embodiments, the chip includes at least 50 individually addressable electrodes per square millimeter.
[0050] In some implementations, the chip distinguishes between incorporated and unincorporated labeled nucleotides based at least in part on the time it takes to detect labeled nucleotides through nanopores.
[0051] In some implementations, the ratio of the time when the incorporated labeled nucleotide can be detected by the nanopore to the time when the unincorporated label can be detected by the nanopore is at least about 1.5.
[0052] In some implementations, incorporating the first labeled nucleotide does not interfere with the detection of the labeled nanopore associated with the second labeled nucleotide.
[0053] In some implementations, the detection of the labeled nanopore associated with the first labeled nucleotide does not interfere with the incorporation of the second labeled nucleotide.
[0054] In some implementations, the nanopores are able to distinguish between incorporated and unincorporated labeled nucleotides with at least 95% accuracy.
[0055] In some implementations, the nanopores are able to distinguish between incorporated and unincorporated labeled nucleotides with at least 99% accuracy.
[0056] In some implementations, the electrodes are portions of an integrated circuit.
[0057] In some implementations, the electrodes are coupled to the integrated circuit.
[0058] In some implementations, each tag associated with the nucleotide is detected by means of a nanopore when the tag is attached to the incorporated nucleotide.
[0059] In some implementations, the tag associated with the single tagged nucleotide is detected when the tag is released from the single tagged nucleotide.
[0060] On one hand, a chip for sequencing nucleic acid samples includes: a plurality of nanopores, wherein the plurality of nanopores contains at least one nanopore arranged adjacent to an electrode in a membrane, wherein each nanopore is capable of detecting a label type when or during the incorporation of a nucleic acid molecule containing a label type into a growing nucleic acid chain, wherein the ratio of the time at which the incorporated label nucleotide can be detected by the nanopore to the time at which the unincorporated label can be detected by the nanopore is at least about 1.5. In some embodiments, the plurality of nanopores are individually addressable.
[0061] In some implementations, the label species do not pass through the nanopores at the time of incorporation.
[0062] In some implementations, the chip is configured to expel the label species from the nanopores.
[0063] In some implementations, nanopores are ejected using voltage pulses to remove the label species.
[0064] In some implementations, the electrodes are portions of an integrated circuit.
[0065] In some implementations, the electrodes are coupled to the integrated circuit.
[0066] In some implementations, each tag associated with the nucleotide is detected by means of a nanopore when the tag is attached to the incorporated nucleotide.
[0067] In some implementations, the marker species of nucleic acid molecules are detected even when no marker species are released from the nucleic acid molecules.
[0068] On one hand, a system for sequencing nucleic acid samples includes: (a) a chip containing one or more nanopore devices, each of the one or more nanopore devices containing a nanopore in a membrane adjacent to an electrode, wherein the nanopore device detects a label associated with a single labeled nucleotide during incorporation of a labeled nucleotide by a polymerase; and (b) a processor coupled to the chip, wherein the processor is programmed to facilitate the characterization of the nucleic acid sequence of the nucleic acid sample based on electrical signals received from the nanopore devices.
[0069] In some implementations, the nanopore device detects the tag associated with a single tagged nucleotide during subsequent tagging advance through or near the nanopore.
[0070] In some implementations, the nanopore device includes individually addressable nanopores.
[0071] In some embodiments, the chip includes at least 500 individually addressable electrodes per square millimeter. In other embodiments, the chip includes at least 50 individually addressable electrodes per square millimeter.
[0072] In some implementations, the chip distinguishes between incorporated and unincorporated labeled nucleotides based at least in part on the time it takes to detect labeled nucleotides through nanopores.
[0073] In some implementations, the ratio of the time when the incorporated labeled nucleotide can be detected by the nanopore to the time when the unincorporated label can be detected by the nanopore is at least about 1.5.
[0074] In some implementations, the incorporation of the first labeled nucleotide does not interfere with the detection of the labeled nanopore associated with the second labeled nucleotide.
[0075] In some implementations, the detection of the labeled nanopore associated with the first labeled nucleotide does not interfere with the incorporation of the second labeled nucleotide.
[0076] In some implementations, the nanopores are able to distinguish between incorporated and unincorporated labeled nucleotides with at least 95% accuracy.
[0077] In some implementations, the nanopores are able to distinguish between incorporated and unincorporated labeled nucleotides with at least 99% accuracy.
[0078] In some implementations, the electrodes are portions of an integrated circuit.
[0079] In some implementations, the electrodes are coupled to the integrated circuit.
[0080] In some implementations, each tag associated with the nucleotide is detected by means of a nanopore when the tag is attached to the incorporated nucleotide.
[0081] In some implementations, the tag associated with the single tagged nucleotide is detected when the tag is released from the single tagged nucleotide.
[0082] In some embodiments, a marker is guided into and through at least a portion of the nanopore using a given driving force, such as a potential (V+) applied to the nanopore or the membrane containing the nanopore. The marker can be guided into the nanopore from the opening. The driving force can then be reversed (e.g., by applying a potential of opposite polarity, or V-) to expel at least a portion of the marker from the nanopore through the opening. The driving force (e.g., V+) can then be applied again to drive at least a portion of the marker into the nanopore through the opening. Optionally, the polarity of the marker can be reversed and a potential sequence including V-, V+, and V- can be used. This can increase the time period during which the marker can be detected by means of the nanopore.
[0083] In some implementations, nanopores and / or labels are configured to provide an energy landscape so that, within the nanopore, labels associated with nucleotides are more likely to move in one direction (e.g., into the pore) rather than in another direction (e.g., out of the pore).
[0084] In some implementations, the detection of modified bases (e.g., methylated) in the template sample chain can be achieved by measuring the difference in labeling time of the labeled nucleotide during and / or at the time of association with the polymerase, detected through a nanopore. In some cases, the nucleotide labeling and enzyme association time is longer when the relative nucleotides of the sample sequence are methylated compared to unmethylated nucleotides.
[0085] Examples of labeled nucleotides described herein can be any naturally occurring nucleotide modified with a cleavable label or a synthetic non-natural nucleotide analog modified with a cleavable label. For example, universal bases modified with cleavable or uncleavable labels can be used for simple counting of the number of bases in a sample chain.
[0086] Examples of labeled nucleotides described herein may be dimeric nucleotides or dimeric nucleotide analogs that can be extended as dimeric units, and the labeling is reported as a dimeric composition of the combined dimeric nucleotides based on the time of association with polymerase and the signal level detected by a nanopore device.
[0087] When the timing of label association with enzyme can be used to distinguish between incorporated and unincorporated nucleotides, unique current levels and / or the electrical response of the label in the nanopore to an applied potential or a changing applied potential allow for differentiation of labels associated with different nucleotides.
[0088] On one hand, a method for nucleic acid sequencing includes applying an alternating current (AC) waveform to a circuit near a nanopore and a sensing electrode, wherein a marker associated with a nucleotide incorporated into a growing nucleic acid chain complementary to a template nucleic acid chain is detected when the waveform has a first polarity, and the electrode is recharged when the waveform has a second polarity.
[0089] In another aspect, a method for sequencing nucleic acid molecules includes: (a) providing one or more labeled nucleotides to a nanopore in a membrane adjacent to an electrode; (b) incorporating a single labeled nucleotide of the one or more labeled nucleotides into a strand complementary to the nucleic acid molecule; and (c) detecting a label associated with the labeled nucleotide once or multiple times by means of an alternating current (AC) waveform applied to the electrode, wherein the label is detected when the label is attached to the single labeled nucleotide incorporated into the strand.
[0090] In some implementations, the waveform ensures that the electrode does not deplete over a period of time of at least about 1 second, 10 seconds, 30 seconds, 1 minute, 10 minutes, 20 minutes, 30 minutes, 1 hour, 2 hours, 3 hours, 4 hours, 5 hours, 6 hours, 12 hours, 24 hours, 1 day, 2 days, 3 days, 4 days, 5 days, 6 days, 1 week, 2 weeks, 3 weeks, or 1 month.
[0091] In some implementations, the identity of the tag is determined by the relationship between the current measured at different voltages and the voltage applied through the waveform.
[0092] In some embodiments, the nucleotide comprises adenine (A), cytosine (C), thymine (T), guanine (G), uracil (U), or any derivative thereof.
[0093] In some implementations, the methylation of bases in the template nucleic acid strand is determined by the time period during which the label is detected when the base is methylated compared to the time period during which the base is not methylated.
[0094] On one hand, a method for determining the length of a nucleic acid or a segment thereof by means of a nanopore in a membrane adjacent to a sensing electrode includes (a) providing a labeled nucleotide to a reaction chamber containing the nanopore, wherein the nucleotide having at least two different bases contains the same label coupled to the nucleotide, the label being detectable by means of the nanopore; (b) performing a polymerization reaction by means of a polymerase, thereby incorporating a single labeled nucleotide of the labeled nucleotide into a growth chain complementary to a single-stranded nucleic acid molecule from the nucleic acid sample; and (c) detecting the label associated with the single labeled nucleotide by means of the nanopore during or after the incorporation of the single labeled nucleotide.
[0095] On one hand, a method for determining the length of a nucleic acid or a segment thereof using a nanopore in a membrane adjacent to a sensing electrode includes (a) providing a labeled nucleotide to a reaction chamber containing the nanopore, wherein a single labeled nucleotide of the labeled nucleotide contains a label coupled to the nucleotide, the label being capable of reducing the magnitude of the current flowing through the nanopore compared to the current in the absence of the label; (b) performing a polymerization reaction using a polymerase, thereby incorporating the single labeled nucleotide of the labeled nucleotide into a growth chain complementary to a single-stranded nucleic acid molecule from the nucleic acid sample and reducing the magnitude of the current flowing through the nanopore; and (c) detecting a time period between the incorporation of the single labeled nucleotide using the nanopore. In some embodiments, the magnitude of the current flowing through the nanopore returns to at least 80% of the maximum current during the time period between the incorporation of the single labeled nucleotide.
[0096] In some embodiments, all nucleotides have the same label coupled to the nucleotide. In some embodiments, at least some of the nucleotides have a label that recognizes the nucleotide. In some embodiments, up to 20% of the nucleotides have a label that recognizes the nucleotide. In some embodiments, all nucleotides are identified as adenine (A), cytosine (C), guanine (G), thymine (T), and / or uracil (U). In some embodiments, all nucleic acids or segments thereof are short tandem repeat sequences (STRs).
[0097] On one hand, a method for assembling a protein having multiple subunits includes (a) providing a plurality of first subunits; (b) providing a plurality of second subunits, wherein the second subunits are modified relative to the first subunits; (c) contacting the first subunits and the second subunits at a first ratio to form a plurality of proteins having first subunits and second subunits, wherein the plurality of proteins have a plurality of ratios of first subunits and second subunits; and (d) fractionating the plurality of proteins to enrich proteins having a second ratio of first subunits and second subunits, wherein the second ratio is one second subunit for every (n-1) first subunits, where 'n' is the number of subunits constituting the protein.
[0098] In some embodiments, the protein is a nanopore.
[0099] In some implementations, the nanopores are at least 80% homologous to α-hemolysin.
[0100] In some implementations, the first or second subunit contains a purification marker.
[0101] In some implementations, the purification label is a multihistidine label.
[0102] In some implementations, fractionation is performed using ion exchange chromatography.
[0103] In some implementations, the second ratio is 1 second subunit for every 6 first subunits.
[0104] In some implementations, the second ratio is 2 second subunits for every 5 first subunits and a single polymerase is attached to each second subunit.
[0105] In some embodiments, the second subunit comprises a chemically active portion and the method further includes (e) reacting to cause the entity to attach to the chemically active portion.
[0106] In some embodiments, the protein is a nanopore and the entity is a polymerase.
[0107] In some implementations, the first subunit is wild-type.
[0108] In some implementations, the first subunit and / or the second subunit are recombinant.
[0109] In some implementations, the first ratio is approximately equal to the second ratio.
[0110] In some implementation schemes, the first ratio is greater than the second ratio.
[0111] In some embodiments, the method further includes inserting a protein having a second ratio subunit into the bilayer.
[0112] In some implementations, the method further includes sequencing nucleic acid molecules using proteins having a second ratio subunit.
[0113] On the other hand, the nanopore contains multiple subunits, wherein a polymerase is attached to one of the subunits and at least one but fewer than all of the subunits contain a first purification label.
[0114] In some implementations, the nanopores are at least 80% homologous to α-hemolysin.
[0115] In some implementations, all said subunits contain a first purification marker or a second purification marker.
[0116] In some implementations, the first purification label is a multihistidine label.
[0117] In some implementations, the first purification label is placed on a subunit of the polymerase with the attachment.
[0118] In some implementations, the first purification label is placed on a subunit that does not have an attached polymerase.
[0119] On the other hand, a method for sequencing a nucleic acid sample by means of a nanopore in a membrane adjacent to a sensing electrode includes: (a) providing a labeled nucleotide to a reaction chamber containing the nanopore, wherein a single labeled nucleotide of the labeled nucleotide contains a label coupled to the nucleotide, the label being detectable by means of the nanopore; (b) performing a polymerization reaction by means of a polymerase to incorporate a single labeled nucleotide of the labeled nucleotide into a growth chain complementary to a single-stranded nucleic acid molecule from the nucleic acid sample; (c) detecting the label associated with the single labeled nucleotide by means of the nanopore during incorporation of the single labeled nucleotide, wherein the label is detected by means of the nanopore when the nucleotide is associated with the polymerase, and wherein the detection includes (i) applying an applied voltage across the nanopore, and (ii) measuring a current at the applied voltage using the sensing electrode; and (d) calibrating the applied voltage.
[0120] In some implementations, the calibration includes (i) measuring multiple escape voltages of the labeled molecule, (ii) calculating the difference between the measured escape voltages and a reference point, and (iii) altering the calculated difference by applying a voltage.
[0121] In some implementations, the distribution of the desired escape voltage over time is estimated.
[0122] In some implementations, the reference point is the average or median value of the measured escape voltage.
[0123] In some implementations, the method removes the detected changes in the desired escape voltage distribution.
[0124] In some implementations, the method is performed on a plurality of independently addressable nanopores, each nanopore being adjacent to a sensing electrode.
[0125] In some implementations, the applied voltage decreases over time.
[0126] In some implementations, the presence of labels in the nanopores reduces the current measured by the sensing electrodes under applied voltage.
[0127] In some implementations, the labeled nucleotide comprises multiple different labels and the method detects each of the multiple different labels.
[0128] In some implementations, (d) increases the accuracy of the method when compared with steps (a)-(c).
[0129] In some implementations, (d) compensation is made for changes in electrochemical conditions over time.
[0130] In some implementations, (d) compensation is provided for different nanopores with different electrochemical conditions in a device having multiple nanopores.
[0131] In some implementations, (d) different electrochemical conditions are compensated for each performance of the method.
[0132] In some implementations, the method further includes (e) calibrating changes in current gain and / or current offset.
[0133] In some implementations, the marker is detected multiple times when associated with the polymerase.
[0134] In some implementations, the electrodes are recharged between labeling detection periods.
[0135] In some implementations, the method distinguishes between incorporated and unincorporated labeled nucleotides based on the time it takes to detect the labeled nucleotide through the nanopore.
[0136] On the other hand, a method for sequencing nucleic acid samples using nanopores adjacent to sensing electrodes in a membrane includes: (a) removing repetitive nucleic acid sequences from the nucleic acid sample to provide single-stranded nucleic acid molecules for sequencing; (b) providing a labeled nucleotide to a reaction chamber containing the nanopore, wherein a single labeled nucleotide of the labeled nucleotide contains a label coupled to the nucleotide, the label being detectable by means of the nanopore; (c) performing a polymerization reaction using a polymerase to incorporate a single labeled nucleotide of the labeled nucleotide into a growth chain complementary to the single-stranded nucleic acid molecule; and (d) detecting the label associated with the single labeled nucleotide by means of the nanopore during incorporation of the single labeled nucleotide, wherein the label is detected by means of the nanopore when the nucleotide associates with the polymerase.
[0137] In some implementations, the repetitive nucleic acid sequence contains at least 20 consecutive nucleic acid bases.
[0138] In some implementations, the repetitive nucleic acid sequence contains at least 200 consecutive nucleic acid bases.
[0139] In some implementations, the repetitive nucleic acid sequence contains at least 20 consecutive repetitive subunits of nucleic acid bases.
[0140] In some implementations, the repetitive nucleic acid sequence contains at least 200 consecutive repetitive subunits of nucleic acid bases.
[0141] In some implementations, the repetitive nucleic acid sequence is removed by hybridization with a nucleic acid sequence complementary to the repetitive nucleic acid sequence.
[0142] In some implementations, a nucleic acid sequence complementary to the repetitive nucleic acid sequence is fixed on a solid support.
[0143] In some implementations, the solid support is the surface.
[0144] In some implementations, the solid support is beads.
[0145] In some implementations, the nucleic acid sequence complementary to the repetitive nucleic acid sequence contains Cot-1 DNA.
[0146] In some implementations, Cot-1 DNA is enriched in repetitive nucleic acid sequences with lengths between approximately 50 and approximately 100 nucleic acid bases.
[0147] On the other hand, a method for sequencing nucleic acid samples using nanopores adjacent to sensing electrodes in a membrane includes: (a) providing a labeled nucleotide to a reaction chamber containing the nanopore, wherein a single labeled nucleotide of the labeled nucleotide contains a label coupled to the nucleotide, the label being detectable by means of the nanopore; (b) incorporating a single labeled nucleotide of the labeled nucleotide into a growth chain complementary to a single-stranded nucleic acid molecule from the nucleic acid sample by means of a polymerase attached to the nanopore via a polymerase; and (c) detecting the label associated with the single labeled nucleotide by means of the nanopore during incorporation of the single labeled nucleotide, wherein the label is detected by means of the nanopore when the nucleotide associates with the polymerase.
[0148] In some implementations, the joint is flexible.
[0149] In some implementations, the connector is at least 5 nanometers long.
[0150] In some implementations, the connector is a direct attachment.
[0151] In some implementations, the linker contains amino acids.
[0152] In some implementations, the nanopores and polymerases contain a single polypeptide.
[0153] In some implementations, the connector contains nucleic acid or polyethylene glycol (PEG).
[0154] In some implementations, the connector contains non-covalent bonds.
[0155] In some implementations, the connector contains biotin and streptavidin.
[0156] In some embodiments, at least one of the following is present: (a) the C-terminus of the polymerase is attached to the N-terminus of the nanopore; (b) the C-terminus of the polymerase is attached to the C-terminus of the nanopore; (c) the N-terminus of the polymerase is attached to the N-terminus of the nanopore; (d) the N-terminus of the polymerase is attached to the C-terminus of the nanopore; and (e) the polymerase is attached to the nanopore, wherein at least one of the polymerase and the nanopore is not terminally attached.
[0157] In some implementations, the connector is oriented relative to the nanopore to the polymerase so that the label can be detected by means of the nanopore.
[0158] In some implementations, the polymerase is attached to the nanopore via two or more adapters.
[0159] In some implementations, the adapter comprises one or more SEQ ID NO 2-35, or PCR products generated therefrom.
[0160] In some embodiments, the adapter comprises a peptide encoded by one or more SEQ ID NO 1-35, or a PCR product generated therefrom.
[0161] In some implementations, the nanopores are at least 80% homologous to α-hemolysin.
[0162] In some implementations, the polymerase is at least 80% homologous to phi-29.
[0163] On the other hand, the marker molecule comprises (a) a first polymer chain comprising a first segment and a second segment, wherein the second segment is narrower than the first segment; and (b) a second polymer chain comprising two ends, wherein the first end is attached to the first polymer chain adjacent to the second segment and the second end is not attached to the first polymer chain, wherein the marker molecule is capable of passing through a nanopore in a first direction aligned with the second polymer chain adjacent to the second segment.
[0164] In some embodiments, the marker molecule is unable to pass through the nanopore in a second direction in which the second polymer chain is not aligned adjacent to the second segment.
[0165] In some implementations, the second polymer chain pairs with the bases of the first polymer chain when the second polymer chain is not aligned with the second segment.
[0166] In some implementations, the first polymer chain is immobilized with a nucleotide.
[0167] In some implementations, when nucleotides are incorporated into the growing nucleic acid chain, the first polymer chain is released from the nucleotides.
[0168] In some implementations, the first polymer chain is fixed to the terminal phosphate of the nucleotide.
[0169] In some implementations, the first polymer chain contains nucleotides.
[0170] In some implementations, the second segment contains debaseted nucleotides.
[0171] In some implementations, the second segment contains a carbon chain.
[0172] On the other hand, a method for sequencing nucleic acid samples by means of a nanopore in a membrane adjacent to a sensing electrode includes: (a) providing a labeled nucleotide to a reaction chamber containing the nanopore, wherein a single labeled nucleotide of the labeled nucleotide contains a label coupled to the nucleotide, the label being detectable by means of the nanopore, wherein the label comprises (i) a first polymer chain comprising a first segment and a second segment, wherein the second segment is narrower than the first segment, and (ii) a second polymer chain comprising two ends, wherein the first end is fixed to the first polymer chain adjacent to the second segment and the second end is not fixed to the first polymer chain, wherein the label molecule is capable of passing through the nanopore in a first direction aligned with the second polymer chain adjacent to the second segment; (b) performing a polymerization reaction by means of a polymerase to incorporate a single labeled nucleotide of the labeled nucleotide into a growth chain complementary to a single-stranded nucleic acid molecule from the nucleic acid sample; and (c) detecting the label associated with the single labeled nucleotide by means of the nanopore during incorporation of the single labeled nucleotide, wherein the label is detected by means of the nanopore when the nucleotide associates with the polymerase.
[0173] In some embodiments, the marker molecule is not able to pass through the nanopore in a second direction in which the second polymer chain is not aligned with the second segment.
[0174] In some implementations, the marker is detected multiple times when associated with the polymerase.
[0175] In some implementations, the electrodes are recharged between labeling detection periods.
[0176] In some embodiments, the label penetrates the nanopore during incorporation of the single labeled nucleotide, and the label does not exit the nanopore when the counter electrode is recharged.
[0177] In some implementations, the method distinguishes between incorporated and unincorporated labeled nucleotides based on the time it takes to detect the labeled nucleotide through the nanopore.
[0178] In some embodiments, the ratio of the time for detecting incorporated labeled nucleotides through the nanopore to the time for detecting unincorporated labeled nucleotides through the nanopore is at least 1.5.
[0179] On the other hand, a method for nucleic acid sequencing includes: (a) providing a single-stranded nucleic acid to be sequenced; (b) providing a plurality of probes, wherein the probes comprise (i) a hybridization portion capable of hybridizing with the single-stranded nucleic acid, (ii) a loop structure having two ends, wherein each end is attached to the hybridization portion, and (iii) a cleavable group of the hybridization portion located between the ends of the loop structure, wherein the loop structure includes a gate preventing the loop structure from passing through a nanopore in the opposite direction; and (c) polymerizing the plurality of probes in an order determined by hybridizing the hybridization portion with the single-stranded nucleic acid to be sequenced.
[0180] (d) cleaving the cleavable group to provide an extended strand to be sequenced; (e) passing the extended strand through a nanopore, wherein the gate prevents the extended strand from passing through the nanopore in the opposite direction; and (f) sequencing the single-stranded nucleic acid to be sequenced by means of the nanopore to detect the loop structure of the extended strand in the order determined by hybridization of the hybridized portion with the single-stranded nucleic acid to be sequenced.
[0181] In some embodiments, the ring structure includes a narrow segment and the gate is a polymer with two ends, wherein a first end is fixed to the ring structure adjacent to the narrow segment and a second end is not fixed to the ring structure, wherein the ring structure is capable of passing through the nanopore in a first direction aligned with the narrow segment adjacent to the gate.
[0182] In some implementations, the ring structure cannot pass through the nanopore in the opposite direction to where the gate is not aligned with the adjacent narrow segment.
[0183] In some implementations, the gate pairs with the bases of the ring structure when the gate is not aligned with adjacent narrow segments.
[0184] In some implementations, the gate comprises a nucleotide.
[0185] In some implementations, the narrow segment contains debaseted nucleotides.
[0186] In some implementations, the narrow segment contains a carbon chain.
[0187] In some implementations, the electrodes are recharged between detection periods.
[0188] In some implementations, when the electrode is recharged, the extended chains do not pass through the nanopore in the reverse direction.
[0189] Further aspects and advantages of this disclosure will become apparent to those skilled in the art from the following detailed description, which shows and describes only illustrative embodiments of the disclosure. It will be appreciated that other and different embodiments are possible with respect to this disclosure, and that certain details may vary in various obvious respects without departing from this disclosure. Therefore, the drawings and description are to be regarded as illustrative in nature and not restrictive.
[0190] Incorporate by reference
[0191] All publications, patents and patent applications mentioned in this specification are incorporated herein by reference to the extent that each individual publication, patent or patent application is expressly and individually incorporated by reference. Attached Figure Description
[0192] The novel features of the invention are particularly set forth in the appended claims. A better understanding of the features and advantages of the invention will be obtained by referring to the following detailed description, which sets forth exemplary embodiments utilizing the principles of the invention, and the accompanying drawings, in which:
[0193] Figure 1 The steps of the method are illustrated schematically;
[0194] Figure 2A , 2B Examples of nanopore detectors are shown in 2C, where Figure 2A It has nanopores arranged on the electrode, Figure 2B It has nanopores inserted into the membrane on the pore, and Figure 2C It has nanopores on the protruding electrodes;
[0195] Figure 3 Examples illustrate the components of the device and method;
[0196] Figure 4 An example illustrates a method for nucleic acid sequencing in which the released label is detected through a nanopore when it associates with polymerase.
[0197] Figure 5 An example illustrates a method for nucleic acid sequencing in which a marker is not released at the time of a nucleotide incorporation event and is detected through a nanopore;
[0198] Figure 6 An example of a signal generated by a tag that is briefly retained in a nanopore is shown;
[0199] Figure 7 An array of nanopore detectors is shown;
[0200] Figure 8 An example of a chip setup containing nanopores rather than pores is shown;
[0201] Figure 9 Showing the following Figure 9A The relative position of -D Figure 9A -D indicates an example of a test chip cell array configuration;
[0202] Figure 10 An example of a unit analog circuit is shown;
[0203] Figure 11 An example of an ultra-compact measurement circuit is shown;
[0204] Figure 12 An example of an ultra-compact measurement circuit is shown;
[0205] Figure 13 Examples of labeled molecules attached to phosphate groups of nucleotides are shown;
[0206] Figure 14 Examples of alternative marker locations are shown;
[0207] Figure 15 The detectable markers are shown – polyphosphates and other detectable markers;
[0208] Figure 16 A computer system configured to control a sequencer is shown;
[0209] Figure 17 This demonstrates the docking of phi29 polymerase with hemolysin nanopores;
[0210] Figure 18 The probability density of residence time for polymerases exhibiting two restrictive kinetic steps is shown.
[0211] Figure 19 The markings showing the association with the binding coupler on the side of the nanopore closer to the detection circuit are shown.
[0212] Figure 20 The barbed markers show that it is easier for the flow to pass through the nanopores than to flow out of the nanopores;
[0213] Figure 21A -C shows an example of a waveform;
[0214] Figure 22 A graph showing the extracted signals of four nucleic acid bases—adenine (A), cytosine (C), guanine (G), and thymine (T)—compared to the applied voltage is shown.
[0215] Figure 23 A graph showing the extracted signals from multiple runs of four nucleic acid bases—adenine (A), cytosine (C), guanine (G), and thymine (T)—compare the applied voltages.
[0216] Figure 24 A graph showing the percentage of reference conductance difference (RCD%) compared to the applied voltage for multiple runs of four nucleic acid bases: adenine (A), cytosine (C), guanine (G), and thymine (T).
[0217] Figure 25 This demonstrates the use of oligonucleotide speed-bumps to slow down the progression of nucleic acid polymerases;
[0218] Figure 26 This demonstrates the use of a second enzyme or protein other than polymerase, such as helicase or nucleic acid-binding protein;
[0219] Figure 27 Examples of methods for forming multimeric proteins having a specified number of modified subunits are shown;
[0220] Figure 28 Examples of multiple nanopores with hierarchical separation of different numbers of modified subunits are shown;
[0221] Figure 29 Examples of multiple nanopores with hierarchical separation of different numbers of modified subunits are shown;
[0222] Figure 30 An example of calibration with applied voltage is shown;
[0223] Figure 31 Examples of labeled nucleotides with gates are shown;
[0224] Figure 32 An example of nucleic acid sequencing using labeled nucleotides with gates is shown;
[0225] Figure 33 The probes used for expandant sequencing are shown.
[0226] Figure 34 An example of an extended polymer probe for polymerization is shown;
[0227] Figure 35 Examples are shown of cleaving cleavable groups to provide extended strands for sequencing;
[0228] Figure 36 An example is shown where the extended chain passes through a nanopore;
[0229] Figure 37 An example of non-Faradaic conduction is shown;
[0230] Figure 38 An example of capturing two labeled molecules is shown;
[0231] Figure 39 An example of a ternary complex formed between the nucleic acid to be sequenced, the labeled nucleotide, and the fusion of the nanopore and the polymerase is shown;
[0232] Figure 40 An example is shown of an electric current flowing through a nanopore in the absence of labeled nucleotides;
[0233] Figure 41 Examples of using current levels to distinguish different labeled nucleotides are shown;
[0234] Figure 42 Examples of using current levels to distinguish different labeled nucleotides are shown;
[0235] Figure 43 Examples of using current levels to distinguish different labeled nucleotides are shown;
[0236] Figure 44 An example of using current levels to sequence nucleic acid molecules using labeled nucleotides is shown;
[0237] Figure 45 An example of using current levels to sequence nucleic acid molecules using labeled nucleotides is shown;
[0238] Figure 46 Examples of using current levels to sequence nucleic acid molecules using labeled nucleotides are shown; and
[0239] Figure 47 An example of using current levels to sequence nucleic acid molecules using labeled nucleotides is shown. Detailed Implementation Plan
[0240] Although various embodiments of the invention have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Many changes, variations, and substitutions may occur to those skilled in the art without departing from the invention. It should be understood that various alternatives to the embodiments of the invention described herein may be employed.
[0241] As used herein, the term "nanopore" generally refers to a pore, channel, or pathway formed or otherwise provided in a membrane. The membrane can be an organic membrane, such as a lipid bilayer, or a synthetic membrane, such as a membrane formed from a polymeric material. The membrane can be a polymeric material. Nanopores can be located adjacent to or near electrode arrangements coupled to sensing circuits, such as, for example, complementary metal-oxide-semiconductor (CMOS) or field-effect transistor (FET) circuits. In some instances, nanopores have a characteristic width or diameter of approximately 0.1 nanometers (nm) to approximately 1000 nm. Some nanopores are proteins. α-hemolysin is an example of a protein nanopore.
[0242] As used herein, the term "nucleic acid" generally refers to a molecule containing one or more nucleic acid subunits. Nucleic acids may include one or more subunits selected from adenosine (A), cytosine (C), guanine (G), thymine (T), and uracil (U), or variants thereof. Nucleotides may include A, C, G, T, or U, or variants thereof. Nucleotides may include any subunit that can be incorporated into a growing nucleic acid chain. Such subunits may be A, C, G, T, or U, or specific to one or more complementary A, C, G, T, or U, or any other subunit complementary to purines (i.e., A or G, or variants thereof) or pyrimidines (i.e., C, T, or U, or variants thereof). Subunits enable the resolution of individual nucleic acid bases or base sets (e.g., AA, TA, AT, GC, CG, CT, TC, GT, TG, AC, CA, or their uracil counterparts). In some instances, nucleic acids are deoxyribonucleic acid (DNA) or ribonucleic acid (RNA), or derivatives thereof. Nucleic acids may be single-stranded or double-stranded.
[0243] As used herein, the term "polymerase" generally refers to any enzyme capable of catalyzing polymerization reactions. Examples of polymerases include, but are not limited to, nucleic acid polymerases, transcriptases, or ligases. A polymerase can be a polymerase.
[0244] Methods and systems for sequencing samples
[0245] This document describes methods, apparatus, and systems for nucleic acid sequencing using or by means of one or more nanopores. One or more nanopores may be arranged in a membrane (e.g., a lipid bilayer), adjacent to or sensing proximity to electrodes that are part of or coupled to an integrated circuit.
[0246] In some instances, nanopore devices comprise a single nanopore in a membrane adjacent to or sensing proximity electrodes. In other instances, nanopore devices comprise at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 1000, or 10,000 nanopores of proximity sensor circuitry or sensing electrodes. One or more nanopores may be associated with a single electrode and a sensing integrated circuit, or multiple electrodes and a sensing integrated circuit.
[0247] The system may include a reaction chamber comprising one or more nanopore devices. The nanopore device may be an individually addressable nanopore device (e.g., a device capable of detecting signals and providing outputs independent of other nanopore devices in the system). An individually addressable nanopore may be individually readable. In some cases, an individually addressable nanopore may be individually writable. Alternatively, an individually addressable nanopore may be both individually readable and individually writable. The system may include one or more computer processors for facilitating sample preparation and various operations of this disclosure, such as nucleic acid sequencing. The processor may be coupled to the nanopore device.
[0248] Nanopore devices may include multiple individually addressable sensing electrodes. Each sensing electrode may include a membrane adjacent to the electrode, and one or more nanopores within the membrane.
[0249] The methods, apparatus, and systems disclosed herein can accurately detect single nucleotide incorporation events, such as when a nucleotide is incorporated into a growing chain complementary to a template. Enzymes (e.g., DNA polymerases, RNA polymerases, ligases) can incorporate nucleotides into growing polynucleotide chains. Enzymes (e.g., polymerases) provided herein can produce polymer chains.
[0250] The added nucleotide is complementary to the corresponding template nucleic acid strand that hybridizes with the growth chain (e.g., polymerase chain reaction (PCR)). The nucleotide may include a tag (or tag type) coupled to any position of the nucleotide, including but not limited to the phosphate (e.g., γ-phosphate), sugar, or nitrogenous base moiety of the nucleotide. In some cases, the tag is detected when it associates with the polymerase during nucleotide incorporation. Following nucleotide incorporation and subsequent cleavage and / or release of the tag, the tag may continue to be detected until it moves through the nanopore. In some cases, the nucleotide incorporation event releases the tag from the nucleotide, which passes through the nanopore and is detected. The tag may be released by the polymerase, or cleaved or released in any suitable manner (including but not limited to cleavage by an enzyme adjacent to the polymerase). Thus, the incorporated base can be identified (i.e., A, C, G, T, or U) because a unique tag is released from each type of nucleotide (i.e., adenine, cytosine, guanine, thymine, or uracil). In some cases, the nucleotide incorporation event does not release the tag. In this context, the label coupled to the incorporated nucleotide is detected using a nanopore. In some instances, the label is movable through or near the nanopore and detected using the nanopore.
[0251] The methods and systems disclosed herein can detect nucleic acid incorporation events within a given time period, such as at a resolution of at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 100, 500, 1000, 5000, 10000, 50000, or 100000 nucleic acid bases ("bases"). In some instances, nanopore devices are used to detect single nucleic acid incorporation events, where each event is associated with a single nucleic acid base. In other instances, nanopore devices are used to detect events associated with multiple bases. For example, the signal sensed by the nanopore device can be a combination signal from at least 2, 3, 4, or 5 bases.
[0252] In some cases, the marker does not pass through the nanopore. The marker can be detected by the nanopore and leave the nanopore without passing through it (e.g., leaving in the opposite direction from where the marker entered the nanopore). The chip can be configured to actively expel the marker from the nanopore.
[0253] In some cases, the label is not released during the nucleotide incorporation event. In other cases, the nucleotide incorporation event "presents" the label to the nanopore (i.e., without label release). The label can be detected by the nanopore without release. The label can be presented to the nanopore for detection by attaching a linker of sufficient length to the nucleotide.
[0254] Nucleotide incorporation events can be detected in real time (i.e., as they occur) and by means of nanopores. In some cases, enzymes attached to or near nanopores (e.g., DNA polymerases) can facilitate the flow of nucleic acid molecules through or near the nanopores. Nucleotide incorporation events, or the incorporation of multiple nucleotides, can release or present one or more types of tags (also referred to herein as "tags") that can be detected by nanopores. Detection can occur when a tag flows through or near a nanopore, when a tag remains in a nanopore, and / or when a tag is presented to a nanopore. In some cases, enzymes attached to or near nanopores can facilitate the detection of tags at the time of incorporation of one or more nucleotides.
[0255] The markings disclosed herein can be atoms or molecules, or collections of atoms or molecules. The markings can provide optical, electrochemical, magnetic, or electrostatic (e.g., inductive, capacitive) signatures that can be detected by means of nanopores.
[0256] The methods described herein can be single-molecule methods. That is, the detected signal is generated by a single molecule (i.e., a single nucleotide incorporation) and not by multiple cloned molecules. The method may not require DNA amplification.
[0257] Nucleotide incorporation events can occur from a mixture containing multiple nucleotides (e.g., deoxyribonucleoside triphosphates (dNTPs, where N is adenosine (A), cytidine (C), thymidine (T), guanosine (G), or uridine (U))). Nucleotide incorporation events do not necessarily occur from a solution containing a single type of nucleotide (e.g., dATP). Nucleotide incorporation events do not necessarily occur from alternating solutions of multiple nucleotides (e.g., dATP, followed by dCTP, followed by dGTP, followed by dTTP, followed by dATP). In some cases, dimers of multiple nucleotides (e.g., AA, AG, AC, AT, GA, GG, GG, GC, GT, CA, etc.) are incorporated via ligases.
[0258] Methods for nucleic acid identification and sequencing
[0259] Methods for sequencing nucleic acids may include retrieving a biological sample containing the nucleic acid to be sequenced, extracting or otherwise isolating a nucleic acid sample from a biological sample, and in some cases preparing a nucleic acid sample for sequencing.
[0260] Figure 1 The illustrative examples illustrate methods for sequencing nucleic acid samples. These methods include isolating nucleic acid molecules from biological samples (e.g., tissue samples, fluid samples) and preparing nucleic acid samples for sequencing. In some cases, nucleic acid samples are extracted from cells. Examples of techniques used for nucleic acid extraction include the use of lysozyme, sonication, extraction, high pressure, or any combination thereof. In some cases, the nucleic acids are cell-free and do not require extraction from cells.
[0261] In some cases, nucleic acid samples can be prepared for sequencing using processes that include the removal of proteins, cell wall debris, and other components from the nucleic acid samples. Many commercially available products are available for this purpose, such as, for example, rotating columns. Ethanol precipitation and centrifugation can also be used.
[0262] Nucleic acid samples can be segmented (or broken) into multiple fragments, which can facilitate nucleic acid sequencing, such as with devices comprising multiple nanopores in an array. However, breaking the nucleic acid molecules to be sequenced may not be necessary.
[0263] In some cases, long sequences can be determined (i.e., shotgun sequencing may not be necessary). Nucleic acid sequences of any suitable length can be determined. For example, sequences of at least approximately 5, approximately 10, approximately 20, approximately 30, approximately 40, approximately 50, approximately 100, approximately 200, approximately 300, approximately 400, approximately 500, approximately 600, approximately 700, approximately 800, approximately 800, approximately 1000, approximately 1500, approximately 2000, approximately 2500, approximately 3000, approximately 3500, approximately 4000, approximately 4500, approximately 5000, approximately 6000, approximately 7000, approximately 8000, approximately 9000, approximately 10000, approximately 20000, approximately 40000, approximately 60000, approximately 80000, or approximately 100000 bases can be sequenced. In some cases, sequencing is performed on bases of at least 5, at least 10, at least 20, at least 30, at least 40, at least 50, at least 100, at least 200, at least 300, at least 400, at least 500, at least 600, at least 700, at least 800, at least 800, at least 1000, at least 1500, at least 2000, at least 2500, at least 3000, at least 3500, at least 4000, at least 4500, at least 5000, at least 6000, at least 7000, at least 8000, at least 9000, at least 10000, at least 20000, at least 40000, at least 60000, at least 80000, and at least 100000. In some cases, the sequenced bases are sequential. In some cases, the sequenced bases are not sequential. For example, sequencing can be performed on rows of a given number of bases. In another instance, one or more sequenced bases may be separated from one or more blocks whose sequence information is not determined and / or unavailable. In some implementations, the template may be sequenced multiple times (e.g., using a circular nucleic acid template), in some cases generating redundant sequence information. In some cases, software is used to provide the sequence. In some cases, the nucleic acid sample may be split before sequencing. In certain cases, the nucleic acid sample strand may be processed to make a given double-helix DNA or RNA / DNA region circular, thereby including the corresponding sense and antisense portions of the double-helix DNA or RNA / DNA region within the circular DNA or circular DNA / RNA molecule. In this case, sequenced bases from such molecules allow for easier data combination and verification of base position reads.
[0264] Nanopore sequencing and molecular detection
[0265] This article provides systems and methods for sequencing nucleic acid molecules using nanopores. Nanopores can be formed or otherwise embedded in a film of a sensing electrode arrangement adjacent to a sensing circuit, such as an integrated circuit. The integrated circuit can be an application-specific integrated circuit (ASIC). In some instances, the integrated circuit is a field-effect transistor or a complementary metal-oxide-semiconductor (CMOS). The sensing circuit can be located in a chip or other device having nanopores, or external to the chip or device, such as in an off-chip configuration. The semiconductor can be any semiconductor, including but not limited to group IV (e.g., silicon) and group III-V semiconductors (e.g., gallium arsenide).
[0266] In some cases, sensing circuitry detects electrical signals associated with the nucleic acid or label as it flows through or near a nanopore. The nucleic acid can be a subunit of a larger chain. The label can be a byproduct of a nucleotide incorporation event or other interactions between the labeled nucleic acid and the nanopore or adjacent nanopores (such as enzymes that cleave the label from the nucleic acid). The label can remain attached to the nucleotide. The detected signal can be collected and stored in a memory location and subsequently used to construct the nucleic acid sequence. The collected signal can be processed to interpret any anomalies in the detected signal, such as errors.
[0267] Figure 2 illustrates an example of a temperature-controlled nanopore detector (or sensor), which can be prepared according to the method described in U.S. Patent Application Publication No. 2011 / 0193570 (which is incorporated herein by reference in its entirety). Reference Figure 2A The nanopore detector includes a top electrode 201 in contact with a conductive solution (e.g., a salt solution) 207. A bottom conductive electrode 202 is adjacent to, near, or close to a nanopore 206, which is inserted into a membrane 205. In some cases, the bottom conductive electrode 202 is embedded in a semiconductor 203, which is circuitically embedded in a semiconductor substrate 204. The surface of the semiconductor 203 may be treated to be hydrophobic. The sample being detected passes through the pores in the nanopore 206. The semiconductor chip sensor is placed in a package 208, which is in turn located in proximity to a temperature control element 209. The temperature control element 209 may be a thermoelectric heating and / or cooling device (e.g., a Peltier device). Multiple nanopore detectors may form a nanopore array.
[0268] refer to Figure 2B The same numbers represent the same elements, and the membrane 205 can be arranged on the hole 210, wherein the sensor 202 forms part of the surface of the hole. Figure 2C An example is shown in which the electrode 202 protrudes from the treated semiconductor surface 203.
[0269] In some instances, film 205 is formed on the bottom conductive electrode 202, rather than on the semiconductor 203. In this case, film 205 can form a coupling interaction with the bottom conductive electrode 202. However, in some cases, film 205 is formed on both the bottom conductive electrode 202 and the semiconductor 203. Alternatively, film 205 can be formed on the semiconductor 203 and not on the bottom conductive electrode 202, but can extend on the bottom conductive electrode 202.
[0270] Nanopores can be used for indirect sequencing of nucleic acid molecules, and in some cases, for electrical detection. Indirect sequencing can be any method in which nucleotides incorporated into the growing chain do not pass through the nanopore. Nucleic acid molecules can pass through at any suitable distance from and / or near the nanopore, and in some cases, at a distance that allows the label released from the nucleotide incorporation event to be detected within the nanopore.
[0271] Byproducts of nucleotide incorporation events can be detected through nanopores. A "nucleotide incorporation event" is the incorporation of a nucleotide into a growing polynucleotide chain. Byproducts can be associated with the incorporation of a given type of nucleotide. Nucleotide incorporation events are generally catalyzed by enzymes such as DNA polymerases and utilize base-pairing interactions with the template molecule to select from available nucleotides to be incorporated at each position.
[0272] Nucleic acid samples can be sequenced using labeled nucleotides or nucleotide analogs. In some instances, methods for sequencing nucleic acid molecules include (a) incorporating (e.g., polymerizing) a labeled nucleotide, wherein the label associated with a single nucleotide is released upon incorporation, and (b) detecting the released label by means of a nanopore. In some cases, the method further includes guiding the label to attach to a single nucleotide or to release it from a single nucleotide through a nanopore. The released or attached label can be guided by any suitable technique, in some cases by means of an enzyme (or molecular motor) and / or a voltage difference across the pore. Alternatively, the released or attached label can be guided through a nanopore without the use of an enzyme. For example, the label can be guided by a voltage difference across a nanopore as described herein.
[0273] Sequencing with preloaded markers
[0274] Tags released without being loaded into a nanopore can diffuse out of the nanopore and go undetected. This can introduce errors into sequencing nucleic acid molecules (e.g., missing nucleic acid sites or detecting tags in the wrong order). This paper provides a method for sequencing nucleic acid molecules in which the tag molecules are “preloaded” into the nanopore before the tags are released from the nucleotides. Preloaded tags are more likely to be detected by the nanopore than unpreloaded tags (e.g., possibly at least about 100 times more). Additionally, preloaded tags provide a method for determining whether the tag nucleotide has been incorporated into the growing nucleic acid chain. Tags associated with incorporated nucleotides can associate with the nanopore for a longer period of time (e.g., at least about 50 milliseconds on average) compared to tags that pass through the nanopore (and are detected by the nanopore) without incorporation (e.g., less than about 1 ms on average). In some instances, the label associated with the incorporated nucleotide may associate with or remain with the nanopore or otherwise couple to an enzyme (e.g., polymerase) adjacent to the nanopore for an average time period of at least about 1 ms, 20 ms, 30 ms, 40 ms, 50 ms, 100 ms, 200 ms, or greater than 250 ms. In some instances, the label signal associated with the incorporated nucleotide may have an average detection lifetime of at least about 1 ms, 20 ms, 30 ms, 40 ms, 50 ms, 100 ms, 200 ms, or greater than 250 ms. The label may be coupled to the incorporated nucleotide. Label signals with an average detection lifetime less than that attributable to the average detection lifetime of the incorporated nucleotide (e.g., less than about 1 ms) may be attributable to unincorporated nucleotides coupled with the label. In some cases, at least 'x' of the average detection lifetime may be attributable to the incorporated nucleotide, and average detection lifetimes less than 'x' may be attributable to unincorporated nucleotides. In some instances, 'x' can be 0.1 ms, 1 ms, 20 ms, 30 ms, 40 ms, 50 ms, 100 ms, or 1 second.
[0275] The label can be detected using a nanopore device having at least one nanopore in a membrane. The label can associate with a single labeled nucleotide during incorporation. The methods provided herein may involve distinguishing, by means of a nanopore, a label associated with a single labeled nucleotide and one or more labels associated with one or more unincorporated single labeled nucleotides. In some cases, the nanopore device detects a label associated with a single labeled nucleotide during incorporation. Using a nanopore device, in some cases, the electrodes and / or nanopores of the nanopore device can detect, identify, or distinguish labeled nucleotides (whether incorporated into a growing nucleic acid chain or not) within a given time period. The time period during which the label and / or the nucleotide coupled to the label is held by an enzyme, such as an enzyme that promotes nucleotide incorporation into a nucleic acid chain (e.g., a polymerase), may be shorter, in some cases much shorter, than the time period during which the nanopore device detects the label. In some instances, the label can be detected multiple times by the electrodes within the time period during which the incorporated labeled nucleotide associates with the enzyme. For example, the label can be detected by the electrode at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 1000, 10,000, 100,000, or 1,000,000 times during the time period in which the incorporated labeled nucleotide associates with the enzyme.
[0276] Any enumeration of detection times or average detection times may allow for a proportion of detection times that are higher or lower than the stated time or average time. In some cases, when sequencing multiple nucleic acid bases, the detection times follow a statistical distribution (e.g., an exponential distribution or a Gaussian distribution). For example, an exponential distribution may have a relatively large percentage of detection times that are lower than the average detection time, as in... Figure 18 middle.
[0277] In some instances, preloading the label includes guiding at least a portion of the label through at least a portion of a nanopore when the label is attached to a nucleotide, which has been incorporated into a nucleic acid chain (e.g., a growing nucleic acid chain), is undergoing incorporation into a nucleic acid chain, or has not yet been incorporated into a nucleic acid chain but may undergo incorporation. In some instances, preloading the label includes guiding at least a portion of the label through at least a portion of the nanopore before or while the nucleotide is being incorporated into the nucleic acid chain. In some cases, preloading the label may include guiding at least a portion of the label through at least a portion of the nanopore after the nucleotide has been incorporated into the nucleic acid chain.
[0278] Figure 3The main components of the method are shown. Here, a nanopore 301 is formed in a membrane 302. An enzyme 303 (e.g., a polymerase such as DNA polymerase) is associated with the nanopore. In some cases, the enzyme is covalently linked to the nanopore as described below. The polymerase is associated with a single-stranded nucleic acid molecule 304 to be sequenced. The single-stranded nucleic acid molecule is circular in some cases, but this is not necessary. In some cases, the nucleic acid molecule is linear. In some embodiments, a nucleic acid primer 305 hybridizes with a portion of the nucleic acid molecule. In some cases, the primer has a hairpin (e.g., to prevent the newly generated nucleic acid strand from passing through the nanopore after the first passage around the circular template). Using the single-stranded nucleic acid molecule as a template, the polymerase catalyzes the incorporation of nucleotides into the primer. Nucleotides 306 contain a type of label ("label") 307 as described herein.
[0279] Figure 4 The illustration provides an example of a nucleic acid sequencing method utilizing the "preload" tag. Part A shows... Figure 3 The main components are described. Part C illustrates the tag loaded into the nanopore. A "loaded" tag can be located at and / or maintained at or immediately adjacent to the nanopore for a perceptible time period, such as, for example, at least 0.1 milliseconds (ms), at least 1 ms, at least 5 ms, at least 10 ms, at least 50 ms, at least 100 ms, or at least 500 ms or at least 1000 ms. In some cases, a "preloaded" tag is loaded into the nanopore prior to release from the nucleotide. In some cases, the tag is preloaded if the probability that, at the time of the nucleotide incorporation event, the tag passes through the nanopore after release (and / or is detected by the nanopore) is suitably high, such as, for example, at least 90%, at least 95%, at least 99%, at least 99.5%, at least 99.9%, at least 99.99%, or at least 99.999%.
[0280] In the transition from part A to part B, the nucleotide has associated with the polymerase. The associated nucleotide pairs with the bases of the single-stranded nucleic acid molecule (e.g., A with T and G with C). It should be recognized that many nucleotides can become transiently associated with the polymerase, even if they are not paired with the bases of the single-stranded nucleic acid molecule. Unpaired nucleotides may be rejected by the polymerase, and nucleotide incorporation generally only begins when the nucleotide bases are paired. Unpaired nucleotides are generally rejected within a shorter timescale than the timescale during which correctly paired nucleotides remain associated with the polymerase. Unpaired nucleotides can be rejected for at least about 100 nanoseconds (ns), 1 ms, 10 ms, 100 ms, or 1 second (on average), while correctly paired nucleotides remain associated with the polymerase for a longer timescale, such as an average timescale of at least about 1 millisecond (ms), 10 ms, 100 ms, 1 second, or 10 seconds. Figure 4The current passing through the nanopore during parts A and B can be between 3 and 30 picoamperes (pA) in some cases.
[0281] Figure 4 Part C describes the docking of polymerase to nanopores. The polymerase can be pulled toward the nanopore by means of a voltage (e.g., DC or AC voltage) applied to the membrane or nanopore to which it resides. Figure 17 Polymerase docking with nanopores is also depicted, in this case, as a phi29 DNA polymerase (φ29 DNA polymerase) with α-hemolysin nanopores. The label can be pulled into the nanopore during docking by electrodynamic forces, such as forces generated in the presence of an electric field produced by applying a voltage to the membrane and / or nanopore. In some embodiments, in Figure 4 The current flowing through the nanopores during some C phases is approximately 6 pA, approximately 8 pA, approximately 10 pA, approximately 15 pA, or approximately 30 pA. The polymerase undergoes isomerization and transphosphorylation reactions to incorporate nucleotides into the growing nucleic acid molecule and release the labeled molecule.
[0282] In section D, the labeling is depicted through a nanopore. The labeling was detected through the nanopore as described herein. Repeated cycles (i.e., sections A through E or A through F) then allow for the sequencing of nucleic acid molecules.
[0283] In some cases, labeled nucleotides that are not incorporated into the growing nucleic acid molecules will also pass through the nanopores, such as in... Figure 4 As seen in part F. Unincorporated nucleotides can be detected by nanopores in some cases, but this method provides a device for distinguishing between incorporated and unincorporated nucleotides based at least in part on the time it takes for the nucleotide to be detected in the nanopore. A label bound to an unincorporated nucleotide rapidly passes through the nanopore and is detected for a short time period (e.g., less than 100 ms), while a label bound to an incorporated nucleotide is loaded into the nanopore and detected for a longer time period (e.g., at least 100 ms).
[0284] In some embodiments, the method distinguishes between incorporated (e.g., polymerized) and unincorporated labeled nucleotides based on the length of time it takes to detect the labeled nucleotide through a nanopore. The label can remain near the nanopore for a longer time upon incorporation compared to when it is unincorporated. In some cases, the polymerase is mutated to increase the time difference between incorporated and unincorporated labeled nucleotides. The ratio of the time for detecting the incorporated labeled nucleotide through the nanopore to the time for detecting the unincorporated label through the nanopore can be any suitable value. In some embodiments, the ratio of the time for detecting the incorporated labeled nucleotide through the nanopore to the time for detecting the unincorporated label through the nanopore is about 1.5, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 12, about 14, about 16, about 18, about 20, about 25, about 30, about 40, about 50, about 100, about 200, about 300, about 400, about 500, or about 1000. In some embodiments, the ratio of the time for detecting incorporated (e.g., polymerized) labeled nucleotides via nanopores to the time for detecting unincorporated labeled nucleotides via nanopores is at least about 1.5, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 12, at least about 14, at least about 16, at least about 18, at least about 20, at least about 25, at least about 30, at least about 40, at least about 50, at least about 100, at least about 200, at least about 300, at least about 400, at least about 500, or at least about 1000.
[0285] The time for the label to be applied to the nanopore (and / or detected through the nanopore) is any suitable value. In some cases, the label is detected through the nanopore at an average of approximately 10 ms, approximately 20 ms, approximately 30 ms, approximately 40 ms, approximately 50 ms, approximately 60 ms, approximately 80 ms, approximately 100 ms, approximately 120 ms, approximately 140 ms, approximately 160 ms, approximately 180 ms, approximately 200 ms, approximately 220 ms, approximately 240 ms, approximately 260 ms, approximately 280 ms, approximately 300 ms, approximately 400 ms, approximately 500 ms, approximately 600 ms, approximately 800 ms, or approximately 1000 ms. In some cases, the label is detected by nanopores for an average of at least about 10 milliseconds (ms), at least about 20 ms, at least about 30 ms, at least about 40 ms, at least about 50 ms, at least about 60 ms, at least about 80 ms, at least about 100 ms, at least about 120 ms, at least about 140 ms, at least about 160 ms, at least about 180 ms, at least about 200 ms, at least about 220 ms, at least about 240 ms, at least about 260 ms, at least about 280 ms, at least about 300 ms, at least about 400 ms, at least about 500 ms, at least about 600 ms, at least about 800 ms, or at least about 1000 ms.
[0286] In some instances, markers generating signals for time periods of at least about 1 ms, at least about 10 ms, at least about 50 ms, at least about 80 ms, at least about 100 ms, at least about 120 ms, at least about 140 ms, at least about 160 ms, at least about 180 ms, at least about 200 ms, at least about 220 ms, at least about 240 ms, or at least about 260 ms are attributed to nucleotides incorporated into the growth chain that are complementary to at least a portion of the template. In some cases, markers generating signals for time periods of less than about 100 ms, less than about 80 ms, less than about 60 ms, less than about 40 ms, less than about 20 ms, less than about 10 ms, less than about 5 ms, or less than about 1 ms are attributed to nucleotides not incorporated into the growth chain.
[0287] Nucleic acid molecules can be linear (e.g.) Figure 5 (As seen in the text). In some cases, such as Figure 3As seen, nucleic acid molecule 304 is circular (e.g., circular DNA, circular RNA). Circular (e.g., single-stranded) nucleic acids can be sequenced multiple times (e.g., re-sequencing portions of the template as polymerase 303 proceeds completely around the loop). For the same genomic location linked together, circular DNA can be sense and antisense strands (allowing for more robust and accurate reads in some cases). Circular nucleic acids can be sequenced up to a suitable accuracy (e.g., at least 95%, at least 99%, at least 99.9%, or at least 99.99% accuracy). In some cases, nucleic acids are sequenced at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 12, at least 15, at least 20 times, at least 40 times, at least 50 times, at least 100 times, or at least 1000 times.
[0288] On one hand, the methods and devices described herein are partially based on the fact that the time period during which incorporated nucleotides are detected and / or can be detected by nanopores is longer than that of unincorporated nucleotides to distinguish between incorporated and unincorporated nucleotides. In some cases, replacing the second nucleic acid strand that hybridizes with the sequencing nucleic acid strand (double-stranded nucleic acid) increases the time difference between the detection of incorporated and unincorporated nucleotides. Reference Figure 3 After the first sequencing, the polymerase may encounter double-stranded nucleic acids (e.g., starting when it encounters primer 305), and the second nucleic acid strand may need to be replaced from the template to continue sequencing. Compared to when the template is single-stranded, this replacement can slow down the rate of polymerase and / or nucleotide incorporation events.
[0289] In some cases, template nucleic acid molecules become double-stranded by hybridization of oligonucleotides with single-stranded templates. Figure 25 An example is shown in which multiple oligonucleotides 2500 are hybridized. Polymerase 2501 moves in the direction indicated by 2502, replacing the oligonucleotide from the template. The polymerase can move more slowly compared to when the oligonucleotide is not present. In some cases, the resolution of incorporated and unincorporated nucleotides is improved by using oligonucleotides as described herein, based on the fact that the time period during which the incorporated nucleotide is and / or can be detected by the nanopore is longer than that of the unincorporated nucleotide. The oligonucleotide can be of any suitable length (e.g., about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 12, about 14, about 16, about 18, or about 20 bases long). Oligonucleotides may contain natural bases (e.g., adenine (A), cytosine (C), guanine (G), thymine (T) and / or uracil (U)), universal bases (e.g., 5-nitroindole, 3-nitropyrrole, 3-methyl-7-propynyl isoquinolone (PIM), 3-methyl isoquinolone (MICS) and / or 5-methyl isoquinolone (5MICS)), or any combination thereof in any proportion.
[0290] In some cases, nucleic acid polymerases proceed more slowly with methylated nucleic acid templates than with unmethylated ones. On the one hand, the methods and / or devices described herein use methylated nucleic acids and / or methylate nucleic acid molecules. In some cases, the methyl group is present and / or added to the 5-position of the cytosine ring and / or the 6-position of the adenine ring. The nucleic acid to be sequenced can be isolated from an organism that has methylated the nucleic acid. In some cases, the nucleic acid can be methylated in vitro (e.g., by using DNA methyltransferases). Time can be used to distinguish methylated bases from unmethylated bases. This enables epigenetic studies.
[0291] In some cases, methylated bases can be distinguished from unmethylated bases based on the characteristic current or the characteristic shape of the current / time plot. For example, labeled nucleotides can produce different blocking currents depending on whether the nucleic acid template has a methylated base at a given position (e.g., due to conformational differences in polymerases). In some cases, C and / or A bases are methylated and the incorporation of corresponding G and / or T labeled nucleotides alters the current.
[0292] Enzymes used for nucleic acid sequencing
[0293] The method described herein allows sequencing of nucleic acid molecules using enzymes (e.g., polymerases, transcriptases, or ligases) through nanopores and labeled nucleotides as described herein. In some cases, the method involves incorporating (e.g., polymerizing) labeled nucleotides using a polymerase (e.g., DNA polymerase). In some cases, the polymerase has been mutated to allow it to accept labeled nucleotides. The polymerase may also be mutated to increase the time required to detect the label through the nanopore (e.g., ...). Figure 4 Part of C's time).
[0294] In some embodiments, the enzyme is any enzyme that produces a nucleic acid chain by phosphate bonding of nucleotides. In some cases, the DNA polymerase is a 9°N polymerase or a variant thereof, *E. coli* DNA polymerase I, bacteriophage T4 DNA polymerase, sequencing enzyme, Taq DNA polymerase, 9°N polymerase (external-) A485L / Y409V, phi29 DNA polymerase (φ29 DNA polymerase), Bct polymerase, or a variant, mutant, or homologue thereof. Homologues may have any suitable percentage of homology, including but not limited to at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, or at least about 95% sequence identity.
[0295] refer to Figure 3Enzyme 303 can be attached to nanopore 301. Suitable methods for attaching the enzyme to the nanopore include cross-linking, such as forming intramolecular disulfide bonds. The nanopore and enzyme can also be fusion proteins (i.e., encoded by a single polypeptide chain). Methods for generating fusion proteins can include the coding sequences within and adjacent to the nanopore (without a stop codon between them), the coding sequence of the fusion enzyme, and expression of the fusion sequence from a single promoter. In some instances, enzyme 303 can be attached to or otherwise coupled to nanopore 301 using molecular staples or fingers. In some cases, the enzyme is attached via intermediate molecules, such as, for example, biotin conjugated to both the enzyme and the nanopore, and streptavidin tetramers linked to both biotin. The enzyme can also be attached to the nanopore using antibodies. In some cases, proteins that form covalent bonds with each other (e.g., the SpyTag™ / SpyCatcher™ system) are used to attach the polymerase to the nanopore. In some cases, phosphatases or enzymes labeled from nucleotide cleavage are also attached to the nanopore.
[0296] In some cases, the DNA polymerase is the phi29 DNA polymerase. Mutations can be made to the polymerase to promote and / or improve the efficiency of the mutant polymerase in incorporating labeled nucleotides into growing nucleic acid molecules compared to the non-mutant polymerase. Mutations can also be made to the polymerase to improve the entry of nucleotide analogs (e.g., labeled nucleotides) into the active site region of the polymerase and / or to mutate it to coordinate with nucleotide analogs in the active region.
[0297] In some embodiments, the polymerase has an active region having an amino acid sequence that is homologous (e.g., at least 70%, at least 80%, or at least 90% of the same amino acid position) to the active region of the polymerase that receives the nucleotide analog (e.g., VentA 488L).
[0298] Suitable mutations for phi29 DNA polymerase include, but are not limited to, deletions of residues 505-525, deletions within residues 505-525, K135A mutations, E375H mutations, E375S mutations, E375K mutations, E375R mutations, E375A mutations, E375Q mutations, E375W mutations, E375Y mutations, E375F mutations, E486A mutations, E486D mutations, K512A mutations, and combinations thereof. In some cases, the DNA polymerase also contains the L384R mutation. Suitable DNA polymerases are described in U.S. Patent Publication No. 2011 / 0059505, which is incorporated herein by reference in its entirety. In some embodiments, the polymerase is a phi29 DNA polymerase having mutations N62D, L253A, E375Y, A484E, and / or K512Y.
[0299] Suitable mutations for phi29 polymerase are not limited to mutations that confer improved incorporation of labeled nucleotides. Compared to non-mutant (wild-type) phi29 DNA polymerase, other mutations (e.g., amino acid substitutions, insertions, deletions, and / or exogenous features) can confer (but are not limited to) enhanced metal ion coordination, decreased exonuclease activity, decreased reaction rate of one or more steps in the polymerase kinetic cycle, reduced branching ratio, altered cofactor selectivity, increased yield, increased thermostability, increased accuracy, increased speed, increased read length, increased salt tolerance, etc.
[0300] Suitable mutations for phi29 DNA polymerase include, but are not limited to, mutations at position E375, mutations at position K512, and mutations at one or more positions selected from the group consisting of L253, A484, V250, E239, Y224, Y148, E508, and T368.
[0301] In some embodiments, the mutation at position E375 includes an amino acid substitution selected from the group consisting of E375Y, E375F, E375R, E375Q, E375H, E375L, E375A, E375K, E375S, E375T, E375C, E375G, and E375N. In some cases, the mutation at position K512 includes an amino acid substitution selected from the group consisting of K512Y, K512F, K512I, K512M, K512C, K512E, K512G, K512H, K512N, K512Q, K512R, K512V, and K512H. In one embodiment, the mutation at position E375 includes an E375Y substitution and the mutation at position K512 includes a K512Y substitution.
[0302] In some cases, mutant phi29 polymerases contain selections from L253A, L253C, L253S, A484E, A484Q, A484N, A484D, A484K, V250I, V250Q, V250L, V250M, V250C, V250F, V250N, V250R, V250T, V250Y, E239G, Y224K, Y224Q, and Y2 One or more amino acid substitutions in the group consisting of 24R, Y148I, Y148A, Y148K, Y148F, Y148C, Y148D, Y148E, Y148G, Y148H, Y148K, Y148L, Y148M, Y148N, Y148P, Y148Q, Y148R, Y148S, Y148T, Y148V, Y148W, E508R, and E508K.
[0303] In some cases, the phi29 DNA polymerase contains mutations selected from one or more positions in the group consisting of D510, E515, and F526. Mutations may include one or more amino acid substitutions selected from the group consisting of D510K, D510Y, D510R, D510H, D510C, E515Q, E515K, E515D, E515H, E515Y, E515C, E515M, E515N, E515P, E515R, E515S, E515T, E515V, E515A, F526L, F526Q, F526V, F526K, F526I, F526A, F526T, F526H, F526M, F526V, and F526Y. Examples of DNA polymerases that can be used in the methods of this disclosure are described in U.S. Patent Publication No. 2012 / 0034602, which is incorporated herein by reference in its entirety.
[0304] Polymerases can have kinetic rate characteristics suitable for the detection of labels via nanopores. Rate characteristics can refer to the overall rate of nucleotide incorporation and / or the rate of any step of nucleotide incorporation (such as nucleotide addition, enzyme isomerization to or from a closed state, cofactor binding or release, product release, and incorporation or translocation of nucleic acids into growing nucleic acids).
[0305] The system disclosed herein allows for the detection of one or more sequencing-related events. These events can be kinetically observed and / or non-kinetically observed (e.g., nucleotides migrating through nanopores without contact with polymerase).
[0306] The polymerase can be adapted to allow the detection of sequencing events. In some implementations, the rate characteristics of the polymerase can enable the label to be loaded into the nanopore (and / or detected through the nanopore) in an average of approximately 0.1 ms, approximately 1 ms, approximately 5 ms, approximately 10 ms, approximately 20 ms, approximately 30 ms, approximately 40 ms, approximately 50 ms, approximately 60 ms, approximately 80 ms, approximately 100 ms, approximately 120 ms, approximately 140 ms, approximately 160 ms, approximately 180 ms, approximately 200 ms, approximately 220 ms, approximately 240 ms, approximately 260 ms, approximately 280 ms, approximately 300 ms, approximately 400 ms, approximately 500 ms, approximately 600 ms, approximately 800 ms, or approximately 1000 ms. In some implementations, the rate characteristics of the polymerase can enable the label to be loaded into the nanopore (and / or detected by the nanopore) for an average of at least about 5 milliseconds (ms), at least about 10 ms, at least about 20 ms, at least about 30 ms, at least about 40 ms, at least about 50 ms, at least about 60 ms, at least about 80 ms, at least about 100 ms, at least about 120 ms, at least about 140 ms, at least about 160 ms, at least about 180 ms, at least about 200 ms, at least about 220 ms, at least about 240 ms, at least about 260 ms, at least about 280 ms, at least about 300 ms, at least about 400 ms, at least about 500 ms, at least about 600 ms, at least about 800 ms, or at least about 1000 ms. In some cases, the markers were detected via nanopores at average intervals between approximately 80 ms and 260 ms, between approximately 100 ms and 200 ms, or between approximately 100 ms and 150 ms.
[0307] In some cases, polymerase reactions exhibit two kinetic steps starting from an intermediate (where the nucleotide or polyphosphate product binds to the polymerase) and two kinetic steps starting from an intermediate (where the nucleotide or polyphosphate product is not bound to the polymerase). The two kinetic steps can include enzyme isomerization, nucleotide incorporation, and product release. In some cases, the two kinetic steps are template translocation and nucleotide binding.
[0308] Figure 18 For example, if there is a kinetic step, the probability of a given residence time in the nanopore decreases exponentially as the residence time increases by 1800, providing a distribution with a relatively high probability of a shorter residence time in the nanopore (and therefore possibly not passing nanopore detection). Figure 18The example also illustrates that for the case with two or more kinetic steps (e.g., observable or "slow" steps) 1802, the probability of a very rapid residence time of the label in the nanopore is relatively low compared to the case with one slow step 1801 1803. In other words, adding two exponential functions produces a Gaussian function or distribution 1802. Furthermore, the probability distribution of the two slow steps shows a peak in the probability density versus residence time plot 1802. This type of residence time distribution is advantageous for nucleic acid sequencing as described herein (e.g., where it is desirable to detect a high proportion of incorporated labels). A relatively large number of nucleotide incorporation events allows the label to be loaded into the nanopore for a time greater than the minimum (T0). min In some cases, the time period can be greater than 100 ms.
[0309] In some cases, the phi29 DNA polymerase is mutated relative to the wild-type enzyme to provide two kinetically slow steps and / or to provide rate characteristics suitable for the detection of labels via nanopores. In some cases, the phi29 DNA polymerase has at least one amino acid substitution or combination of substitutions selected from positions 484, 198, and 381. In some embodiments, the amino acid substitution is selected from E375Y, K512Y, T368F, A484E, A484Y, N387L, T372Q, T372L, K478Y, 1370W, F198W, L381A, and any combination thereof. Suitable DNA polymerases are described in U.S. Patent No. 8,133,672, which is incorporated herein by reference in its entirety.
[0310] Enzyme kinetics can also be influenced and / or controlled by manipulating the contents of the solution in contact with the enzyme. For example, non-catalytic divalent ions (e.g., ions that do not promote polymerase function, such as strontium (Sr)) can be manipulated. 2+ It can react with catalytic divalent ions (e.g., ions that promote polymerase function, such as magnesium (Mg)). 2+ ) and / or manganese (Mn 2 + Mixing is used to slow down the polymerase. The ratio of catalytic to non-catalytic ions can be any suitable value, including approximately 20, 15, 10, 9, 8, 7, 6, 5, 4, 3, 2, 1, 0.5, 0.2, or 0.1. In some cases, the ratio depends on the concentration of the monovalent salt (e.g., potassium chloride (KCl)), temperature, and / or pH. In one example, the solution contains 1 micromolar Mg 2+ and 0.25 micromoles of Sr 2+ In another example, the solution contained 3 micromoles of Mg. 2+ and 0.7 micromoles of Sr 2+ Magnesium (Mg) 2+ ) and manganese (Mn 2+The concentration of Mg2+ can be any suitable value and can be varied to affect enzyme kinetics. In one example, the solution contains 1 micromolar of Mg2+. 2+ and 0.25 micromolar Mn 2+ In another example, the solution contained 3 micromoles of Mg. 2+ and 0.7 micromolar Mn 2+ .
[0311] Nanopore sequencing with preloaded labeled molecules
[0312] The label can be detected during the synthesis of a nucleic acid strand complementary to the target strand without being released from the incorporated nucleotide. The label can be attached to the nucleotide using a linker to present the label into a nanopore (e.g., the label dangles into or otherwise extends through at least a portion of the nanopore). The linker can be long enough to allow the label to extend into or through at least a portion of the nanopore. In some cases, the label is presented into (i.e., moved into) the nanopore by a voltage difference. Other methods for presenting the label into the pore may also be suitable (e.g., using enzymes, magnets, electric fields, voltage differences). In some cases, no force is applied to the label (i.e., the label diffuses into the nanopore).
[0313] This invention provides a method for sequencing nucleic acids. The method includes incorporating (e.g., polymerizing) a labeled nucleotide. The label associated with a single nucleotide can be detected through a nanopore if it is not released from the nucleotide at the time of incorporation.
[0314] A chip for sequencing nucleic acid samples may contain multiple individually addressable nanopores. These multiple individually addressable nanopores may include at least one nanopore formed in a membrane adjacent to an integrated circuit arrangement. Each individually addressable nanopore is capable of detecting a tag associated with a single nucleotide. Nucleotides may be incorporated (e.g., polymerized), and the tag may not be released from the nucleotide upon incorporation.
[0315] The instance of the method is described in Figure 5 In this context, nucleic acid strand 500 crosses or approaches (but does not pass through, as indicated by the arrow passing through 501) nanopore 502. Enzyme 503 (e.g., DNA polymerase) uses the first nucleic acid molecule as a template 500 to extend the growing nucleic acid strand 504 by incorporating one nucleotide at a time (i.e., the enzyme catalyzes the nucleotide incorporation event). A label is detected through nanopore 502. The label can remain in the nanopore for a certain period of time.
[0316] Enzyme 503 can attach to nanopore 502. Suitable methods for attaching enzymes and nanopores include cross-linking such as forming intramolecular disulfide bonds and / or generating fusion proteins as described above. In some cases, phosphatases also attach to nanopores. These enzymes can also bind to residual phosphate on the cleavage marker and generate a clearer signal by further increasing the residence time in the nanopore. Suitable DNA polymerases include Phi29 DNA polymerase (φ29 DNA polymerase) and also include, but are not limited to, those described above.
[0317] Continue to refer to Figure 5 The enzyme is pulled from an aggregate of nucleotides (solid circles at indicator 505) attached to a label molecule (a hollow circle at indicator 505). Each type of nucleotide is attached to a different label molecule so that they can be distinguished from each other based on signals generated in or associated with the nanopore 502 when the label remains in the nanopore.
[0318] In some cases, the tag is presented to the nanopore and released from the nucleotide at the time of nucleotide incorporation. In some cases, the released tag passes through the nanopore. In some cases, the tag does not pass through the nanopore. In some cases, tags that have been released at the time of nucleotide incorporation are distinguished from tags that can flow through the nanopore, but have at least partially passed through the nanopore for a residence time and are not released at the time of nucleotide incorporation. In some cases, tags that remain in the nanopore for at least about 100 milliseconds (ms) are released at the time of nucleotide incorporation, while tags that remain in the nanopore for less than about 100 ms are not released at the time of nucleotide incorporation. In some cases, the tag is captured and / or directed through the nanopore by a second enzyme or protein (e.g., a nucleic acid-binding protein). The second enzyme can cleave the tag at the time of nucleotide incorporation (e.g., during or after). The linker between the tag and the nucleotide can be cleaved.
[0319] like Figure 26 As seen, the second enzyme or protein 2600 may be attached to the polymerase. In some embodiments, the second enzyme or protein is a nucleic acid helicase that promotes the dissociation of the double-stranded template from the single-stranded template. In some cases, the second enzyme or protein is not attached to the polymerase. The second enzyme or protein may be a nucleic acid-binding protein that binds to the single-stranded nucleic acid template to help maintain the template single strand. The nucleic acid-binding protein may slide along the single-stranded nucleic acid molecule.
[0320] Based on the time length of time it takes for the tag associated with the nucleotide to be detected by means of a nanopore, the incorporated nucleotide can be distinguished from the unincorporated nucleotide. In some instances, the tag associated with the nucleotide already incorporated into the nucleic acid chain ("incorporated nucleotide") is detected, either by means of a nanopore, for an average time period of at least about 5 milliseconds (ms), 10 ms, 20 ms, 30 ms, 40 ms, 50 ms, 60 ms, 70 ms, 80 ms, 90 ms, 100 ms, 200 ms, 300 ms, 400 ms, or 500 ms. Labels associated with unincorporated (e.g., free-flowing) nucleotides were detected via nanopores for average durations of less than approximately 500 ms, 400 ms, 300 ms, 200 ms, 100 ms, 90 ms, 80 ms, 70 ms, 60 ms, 50 ms, 40 ms, 30 ms, 20 ms, 10 ms, 5 ms, or 1 ms. In some cases, labels associated with incorporated nucleotides were detected via nanopores for an average duration of at least approximately 100 ms, and labels associated with unincorporated nucleotides were detected via nanopores for an average duration of less than 100 ms.
[0321] In some instances, a label coupled to an incorporated nucleotide is distinguished from a label associated with an unincorporated complementary strand during growth, based on the residence time of the label in the nanopore or by means of a signal detected by the nanopore from an unincorporated nucleotide. Unincorporated nucleotides may produce signals (e.g., voltage difference, current) detectable in time intervals between about 1 nanosecond (ns) and 100 ms, or about 1 ns and 50 ms, while incorporated nucleotides may produce signals with lifetimes between about 50 ms and 500 ms, or about 100 ms and 200 ms. In some instances, unincorporated nucleotides may produce signals detectable in time intervals between about 1 ns and 10 ms, or about 1 ns and 1 ms. In some cases, the time interval (on average) during which unincorporated labels can be detected by the nanopore is longer than that during which incorporated labels can be detected by the nanopore.
[0322] In some cases, the time period during which incorporated nucleotides pass through and / or can be detected via nanopores is shorter than that of unincorporated nucleotides. These time differences and / or ratios can be used to determine whether a nucleotide detected via a nanopore has been incorporated, as described herein.
[0323] The detection period can be based on the free flow of nucleotides through the nanopores; undoped nucleotides may remain in or near the nanopores for a time period between approximately 1 nanosecond (ns) and 100 ms, or approximately 1 ns and 50 ms, while doped nucleotides may remain in or near the nanopores for a time period between approximately 50 ms and 500 ms, or approximately 100 ms and 200 ms. The time periods can vary depending on process conditions; however, doped nucleotides can have a longer residence time than undoped nucleotides.
[0324] Polymerization (e.g., incorporation) and detection can both occur without interference. In some embodiments, the polymerization of the first labeled nucleotide does not significantly interfere with the nanopore detection of the label associated with the second labeled nucleotide. In some embodiments, the nanopore detection of the label associated with the first labeled nucleotide does not interfere with the polymerization of the second labeled nucleotide. In some cases, the label is long enough to be detected through the nanopore and / or to be detected without preventing nucleotide incorporation events.
[0325] A label (or label type) may include detectable atoms or molecules, or multiple detectable atoms or molecules. In some cases, a label includes one or more adenine, guanine, cytosine, thymine, uracil, or derivatives thereof attached to any position (including phosphate groups, sugars, or nitrogenous bases of nucleic acid molecules). In some instances, a label includes one or more adenine, guanine, cytosine, thymine, uracil, or derivatives thereof covalently attached to phosphate groups of nucleic acid bases.
[0326] The marker may have a length of at least about 0.1 nanometers (nm), 1 nm, 2 nm, 3 nm, 4 nm, 5 nm, 6 nm, 7 nm, 8 nm, 9 nm, 10 nm, 20 nm, 30 nm, 40 nm, 50 nm, 60 nm, 70 nm, 80 nm, 90 nm, 100 nm, 200 nm, 300 nm, 400 nm, 500 nm or 1000 nm.
[0327] The label may include a tail portion of repeating subunits such as adenine, guanine, cytosine, thymine, uracil, or derivatives thereof. For example, the label may include a tail portion having at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 1000, 10,000, or 100,000 subunits of adenine, guanine, cytosine, thymine, uracil, or derivatives thereof. The subunits may be linked to each other and terminated with a phosphate group of a nucleic acid. Other examples of the label portion include any polymeric material, such as polyethylene glycol (PEG), polysulfonate, amino acids, or any polymer that is wholly or partially positively charged, negatively charged, or uncharged.
[0328] The label type may have an electronic signature that is unique to the type of nucleic acid molecule incorporated during incorporation. For example, nucleic acid bases for adenine, guanine, cytosine, thymine, or uracil may have label types that have one or more distinct types for adenine, guanine, cytosine, thymine, or uracil, respectively.
[0329] Figure 6 Examples of the different signals generated by different labels when detected through nanopores are shown. Four different signal intensities (601, 602, 603, and 604) were detected. These could correspond to four different labels. For example, a label presented to a nanopore and / or released by incorporation of adenosine (A) could produce a signal with an amplitude of 601. A label presented to a nanopore and / or released by incorporation of cytosine (C) could produce a signal with a higher amplitude of 603; a label presented to a nanopore and / or released by incorporation of guanine (G) could produce a signal with an even higher amplitude of 604; and a label presented to a nanopore and / or released by incorporation of thymine (T) could produce a signal with an even higher amplitude of 602. Figure 6 The detection of labeled molecules that have been released from nucleotides and / or presented to nanopores at the time of nucleotide incorporation events is also demonstrated. The methods described herein are able to distinguish between labels inserted into nanopores and subsequently cleaved (see, for example, Figure 4 , D) and free-floating uncut markers (see, for example, Figure 4 F).
[0330] The method described herein can distinguish between released (or cut) markers and unreleased (or uncut) markers with an accuracy of at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 99%, at least about 99.5%, at least about 99.9%, at least about 99.95%, or at least about 99.99%, or at least about 99.999%, or at least about 99.9999%.
[0331] refer to Figure 6 The magnitude of the current can be reduced by any suitable amount, including about 5%, about 10%, about 15%, about 20%, about 25%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, about 90%, about 95%, or about 99%. In some embodiments, the magnitude of the current is reduced by at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, or at least 99%. In some embodiments, the magnitude of the current is reduced by at most 5%, at most 10%, at most 15%, at most 20%, at most 25%, at most 30%, at most 40%, at most 50%, at most 60%, at most 70%, at most 80%, at most 90%, at most 95%, or at most 99%.
[0332] The method may also include detecting time intervals between the incorporation of a single labeled nucleotide (e.g., Figure 6 The time period between the incorporation of a single labeled nucleotide can have a high current magnitude. In some embodiments, the current magnitude flowing through the nanopore between nucleotide incorporation events is (e.g., back) about 50%, about 60%, about 70%, about 80%, about 90%, about 95%, or about 99% of the maximum current (e.g., when no label is present). In some embodiments, the current magnitude flowing through the nanopore between nucleotide incorporation events is at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, or at least 99% of the maximum current. Detecting and / or observing the current during the time period between the incorporation of a single labeled nucleotide can improve sequencing accuracy in certain situations (e.g., when sequencing repeating extensions of nucleic acids such as 3 or more identical bases). The time period between nucleotide incorporation events can be used as a clock signal to indicate the length of the nucleic acid molecule or its segment being sequenced.
[0333] The method described herein can distinguish between incorporated (e.g., polymerized) labeled nucleotides and non-polymerized labeled nucleotides (e.g., Figure 5 (506 and 505 in the example). In some instances, the incorporated labeled nucleotides can be distinguished from unincorporated labeled nucleotides with an accuracy of at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 99%, at least about 99.5%, at least about 99.9%, at least about 99.95%, or at least about 99.99%, or at least about 99.999%, or at least about 99.9999%.
[0334] Association between markers and nanopores
[0335] On the one hand, the methods and apparatus described herein are partially based on the time (or time ratio) at which the label associates with the nanopore and / or the label can be detected through the nanopore to distinguish between labeled nucleotides incorporated into nucleic acid molecules and unincorporated labeled nucleotides. In some cases, the interaction between the nucleotide and the polymerase increases the time at which the label associates with the nanopore and / or the label can be detected through the nanopore. In some cases, the label interacts with and / or associates with the nanopore.
[0336] Inserting a marker into a nanopore may be relatively easier than removing it. In some cases, marker entry into a nanopore is faster and / or requires less force than marker exiting the nanopore. Once associated with the nanopore, marker passage through the nanopore can be faster and / or require less force than marker exiting the nanopore from the direction in which it entered.
[0337] The association between the label and the nanopore can be any suitable force or interaction, such as non-covalent bonds, reversible covalent bonds, electrostatic or electrodynamic forces, or any combination thereof. In some cases, the label is designed to interact with the nanopore, the nanopore is abruptly generated or designed to interact with the label, or both the label and the nanopore are designed or selected to associate with each other.
[0338] The association between the tag and the nanopore can be of any suitable strength. In some cases, the association is strong enough to allow the electrode to be recharged without the tag being expelled from the nanopore. In other cases, the voltage polarity across the nanopore can be reversed to recharge the electrode and then reversed again to detect the tag in certain cases where the tag has not left the nanopore.
[0339] Figure 19 Examples are shown in which the labeled portion of the labeled nucleotide 1902 is bound to and / or interacts with an affinity pair (e.g., an affinity molecule) or a binding pair 1903 on the nanopore 1901 side opposite the polymerase. The affinity molecule or binding pair 1903 may be detached from the nanopore 1901 but attached to it, or alternatively, may be a portion of the nanopore 1901. The binding pair may be attached to any suitable surface such as a nanopore or membrane. In some instances, any suitable combination of the labeled molecule and the binding pair may be used. In some cases, the labeled molecule and the binding pair comprise nucleic acid molecules that hybridize with each other. In some cases, the labeled molecule and the binding pair comprise streptavidin and biotin bound to each other. Alternatively, the binding pair 1903 may be a portion of the nanopore 1901.
[0340] Figure 20An example is shown in which the labeled nucleotide comprises a nucleotide portion 2001 and a label portion 2002, wherein the label portion is barbed. The label flows through the nanopore 2003 more easily (e.g., faster and / or with less force) than it exits the nanopore in the direction it entered, thus shaping the label portion in such a way as to have barbs. In one embodiment, the label portion comprises a single-stranded nucleic acid and bases (e.g., A, C, T, G) are attached to the backbone of the nucleic acid label at an angle pointing towards the nucleotide portion 2001 of the labeled nucleotide (i.e., barbed). Alternatively, the nanopore may comprise a flap or other obstruction that allows the label portion to flow in a first direction (e.g., out of the nanopore) and prevents the label portion from flowing in a second direction (e.g., in the opposite direction). The flap may be any hinged obstruction.
[0341] The marker may be designed or selected (e.g., using directed evolution) to bind to and / or associate with nanopores (e.g., within the pore portion of the nanopore). In some embodiments, the marker is a peptide having an arrangement of hydrophilic, hydrophobic, positively charged, and negatively charged amino acid residues that bind to the nanopore. In some embodiments, the marker is a nucleic acid having an arrangement of bases that bind to the nanopore.
[0342] Nanopores can be mutated to associate with labeled molecules. For example, nanopores can be (e.g., using directed evolution) designed or selected to have an arrangement of hydrophilic, hydrophobic, positively charged, and negatively charged amino acid residues that bind to labeled molecules. The amino acid residues can be located in the vestibule and / or pore of the nanopore.
[0343] Labeling extrusion from nanopores
[0344] This disclosure provides a method for removing labeled molecules from nanopores. For example, a chip may be adapted to remove labeled molecules if the label remains in the nanopore or is presented to the nanopore at the time of a nucleotide incorporation event, such as during sequencing. The label may be removed in the opposite direction of its entry into the nanopore (e.g., in the case where the label does not pass through the nanopore)—for example, the label may be pointed from a first opening toward the nanopore and removed from the nanopore through a second opening different from the first opening. Alternatively, the label may be removed from the opening through which it enters the nanopore—for example, the label may be pointed from a first opening toward the nanopore and removed from the nanopore through the first opening.
[0345] This invention provides a chip for sequencing nucleic acid samples, the chip comprising a plurality of individually addressable nanopores, each having at least one nanopore formed in a membrane adjacent to an integrated circuit arrangement, each individually addressable nanopore being adapted to eject labeled molecules from the nanopore. In some embodiments, the chip is adapted to eject (or eject) the label in the direction in which the label enters the nanopore. In some cases, the nanopores eject labeled molecules using a voltage pulse or a series of voltage pulses. The voltage pulse may have a duration of about 1 nanosecond to 1 minute, or 10 nanoseconds to 1 second.
[0346] The nanopore is suitable for ejecting (or by the method described above) labeled molecules over a certain period of time so that two labeled molecules do not coexist in the nanopore. In some embodiments, the probability of two molecules coexisting in the nanopore is at most 1%, at most 0.5%, at most 0.1%, at most 0.05%, or at most 0.01%.
[0347] In some cases, nanopores are suitable for eluting labeled molecules within approximately 0.1 ms, 0.5 ms, 1 ms, 5 ms, 10 ms, or 50 ms of the time interval when the label enters the nanopore (within time intervals of less than approximately 0.1 ms, 0.5 ms, 1 ms, 5 ms, 10 ms, or 50 ms).
[0348] The marker can be expelled from the nanopore using a potential (or voltage). In some cases, the voltage may have the opposite polarity to that used to pull the marker into the nanopore. The voltage can be applied using an alternating current (AC) waveform having a period of at least about 1 nanosecond, 10 nanoseconds, 100 nanoseconds, 500 nanoseconds, 1 microsecond, 100 microseconds, 1 millisecond (ms), 5 ms, 10 ms, 20 ms, 30 ms, 40 ms, 50 ms, 100 ms, 200 ms, 300 ms, 400 ms, 500 ms, 600 ms, 700 ms, 800 ms, 900 ms, 1 second, 2 seconds, 3 seconds, 4 seconds, 5 seconds, 6 seconds, 7 seconds, 8 seconds, 9 seconds, 10 seconds, 100 seconds, 200 seconds, 300 seconds, 400 seconds, 500 seconds, or 1000 seconds.
[0349] Alternating current (AC) waveform
[0350] Sequencing nucleic acid molecules by passing nucleic acid chains through nanopores requires the application of direct current (DC) (e.g., ensuring that the direction of molecule movement through the nanopore is not reversed). However, operating nanopore sensors with DC for extended periods can alter the electrode composition, causing imbalances in ion concentrations across the nanopore and other undesirable effects. Applying an alternating current (AC) waveform avoids these undesirable effects and offers some of the advantages described below. The nucleic acid sequencing method described herein using labeled nucleotides is fully compatible with AC applied voltage and is therefore used to achieve these advantages.
[0351] The ability to recharge the electrode during a detection cycle can be advantageous when using sacrificial electrodes or electrodes that alter molecular characteristics in a current-carrying reaction (e.g., electrodes containing silver). Electrodes may be depleted during a detection cycle, although in some cases they may not. Recharging prevents the electrode from reaching a given depletion limit, such as becoming completely depleted, which can be a problem when the electrode is small (e.g., when the electrode is small enough to provide an array with at least 500 electrodes per square millimeter). Electrode lifetime scales proportionally and depends at least in part on the width of the electrode in some cases.
[0352] In some cases, the electrode is porous and / or "sponge-like". Porous electrodes can have enhanced bilayer-to-bulk liquid capacitance compared to non-porous electrodes. Porous electrodes can be formed by electroplating a metal (e.g., a noble metal) onto a surface in the presence of a detergent. The electroplated metal can be any suitable metal. The metal can be a noble metal (e.g., palladium, silver, osmium, iridium, platinum, silver, or gold). In some cases, the surface is a metallic surface (e.g., palladium, silver, osmium, iridium, platinum, silver, or gold). In some cases, the surface diameter is about 5 micrometers and is smooth. The detergent can create nanoscale intercellular spaces in the surface, making it porous or "sponge-like". Another method for producing porous and / or sponge-like electrodes is to deposit a metal oxide (e.g., platinum oxide) and expose it to a reducing agent (e.g., 4% H2). The reducing agent can reduce the metal oxide (e.g., platinum oxide) back to the metal (e.g., platinum), and by doing so, a sponge-like and / or porous electrode is provided. (For example, palladium) sponges can absorb electrolytes and generate a large effective surface area (e.g., the top-to-bottom region of an electrode with 33 pFarads / μm²). Increasing the electrode surface area by making the electrode porous as described herein can produce an electrode with capacitance that does not become completely depleted.
[0353] In some cases, the need to maintain a voltage difference of conserved polarity across the nanopore for a long period during detection (e.g., when sequencing nucleic acids by passing them through the nanopore) depletes the electrode and can limit the duration of detection and / or the size of the electrode. The devices and methods described herein allow for longer detection times (e.g., indefinite) and / or electrodes that can be scaled down to arbitrarily small sizes (e.g., as limited by considerations other than electrode depletion during detection). As described herein, only a portion of the time during which the label associates with the polymerase can be detected. Switching the polarity and / or magnitude of the voltage across the nanopore between detection periods (e.g., applying an AC waveform) allows for recharging of the electrode. In some cases, the label is detected multiple times (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 100, 1000, 10,000, 100,000, 1,000,000 or more times within a 100 ms period).
[0354] In some cases, the polarity of the voltage across the nanopore reverses periodically. The voltage polarity can reverse after a detection period of any suitable duration (e.g., approximately 1 ms, approximately 5 ms, approximately 10 ms, approximately 15 ms, approximately 20 ms, approximately 25 ms, approximately 30 ms, approximately 40 ms, approximately 50 ms, approximately 60 ms, approximately 80 ms, approximately 100 ms, approximately 125 ms, approximately 150 ms, approximately 200 ms, etc.). The duration of electrode recharging and the strength of the electric field (i.e., when the voltage polarity is opposite to that used for marker detection) allow the electrode to return to its state (e.g., electrode mass) before detection. The net voltage across the nanopore is zero in some cases (e.g., a positive voltage period minus a negative voltage period over a reasonably long timescale such as 1 second, 1 minute, or 5 minutes). In some cases, the voltage applied to the nanopore is balanced to create a net zero current that can be detected by sensing electrodes adjacent to or near the nanopore.
[0355] In some instances, an alternating current (AC) waveform is applied to a nanopore in the membrane or to an electrode adjacent to the membrane to pull a tag through or near the nanopore and to release the tag. The AC waveform may have frequencies of at least 10 microseconds, 1 millisecond (ms), 5 ms, 10 ms, 20 ms, 100 ms, 200 ms, 300 ms, 400 ms, and 500 ms. The waveform may facilitate the alternating and sequential capture and release of the tag, or the movement of the tag in multiple directions (e.g., opposite directions), which can increase the total time the tag associates with the nanopore. This balance of charging and discharging allows for the generation of longer signals from the nanopore electrodes and / or a given tag.
[0356] In some instances, an AC waveform is applied to repeatedly guide at least a portion of the label associated with the labeled nucleotide (e.g., an incorporated labeled nucleotide) into the nanopore and at least a portion of the label out of the nanopore. The label or the nucleotide coupled to the label may be retained by an enzyme (e.g., a polymerase). This repetitive loading and eviction of a single label retained by the enzyme advantageously provides more opportunities for label detection. For example, if the label is retained by the enzyme for 40 milliseconds (ms) and a high AC waveform is applied for 5 ms (to guide the label into the nanopore) and a low AC waveform is applied for 5 ms (to guide the label out of the nanopore), the nanopore can be used to read the label approximately four times. Multiple reads allow for correction of errors, such as errors associated with the label penetrating and / or exiting the nanopore.
[0357] Waveforms can have any suitable shape, including regular shapes (e.g., repeating over a certain period of time) and irregular shapes (e.g., not repeating over any reasonably long period of time, such as 1 hour, 1 day, or 1 week). Figure 21 shows some suitable (regular) waveforms. Examples of waveforms include triangle waves (Figure A), sine waves (Figure B), sawtooth waves, square waves, etc.
[0358] When an alternating current (AC) waveform is applied, the reversal of the polarity of the voltage across the nanopore (i.e., from positive to negative or from negative to positive) can be performed for any reason, including but not limited to (a) recharging the electrode (e.g., charging the chemical composition of the metal electrode), (b) rebalancing the ion concentrations on the cis and antis sides of the membrane, (c) re-establishing a non-zero applied voltage across the nanopore and / or (d) changing the bilayer capacitance (e.g., resetting the voltage or charge present at the interface between the metal electrode and the analyte to a desired level, such as zero).
[0359] Figure 21CThe horizontal dashed line is shown at zero potential difference across the nanopore, where positive voltages extend upwards proportionally to their magnitude and negative voltages extend downwards proportionally to their magnitude. Regardless of the waveform shape, the "duty cycle" compares the combined area under the curve on the positive direction 2100 and the combined area under the curve on the negative direction 2101 of the voltage-time plot. In some cases, the positive region 2100 is equal to the negative region 2101 (i.e., the net duty cycle is zero); however, AC waveforms can have any duty cycle. In certain circumstances, the fair use of an AC waveform with an optimal duty cycle can be used to achieve one or more of the following: (a) the electrodes are electrochemically balanced (e.g., neither charged nor depleted), (b) ion concentrations are balanced between the cis and antis sides of the membrane, (c) the voltage applied across the nanopore is known (e.g., because the capacitive bilayer on the electrode is periodically reset and the capacitor discharges to the same extent as each polarity flip), (d) the labeled molecules are identified multiple times in the nanopore (e.g., by expelling and recapturing the labels with each polarity flip), (e) additional information is captured by each reading of the labeled molecules (e.g., because the measured current can be a different function of the applied voltage for each labeled molecule), (f) high-density nanopore sensors are achieved (e.g., because the composition of the metal electrodes does not change, so it is not limited by the amount of metals constituting the electrodes), and / or (g) low power consumption of the chip. These benefits can allow for continuous extended operation of the device (e.g., at least 1 hour, at least 1 day, at least 1 week).
[0360] In some cases, a first current is measured when a positive potential is applied across the nanopore, and a second current is measured when a negative potential (e.g., having an absolute magnitude equal to the positive potential) is applied across the nanopore. The first current may be equal to the second current, although in some cases the first and second currents may be different. For example, the first current may be less than the second current. In some cases, only one of the positive and negative currents is measured.
[0361] In some cases, the nanopore detects labeled nucleotides for a relatively long period at lower voltages (e.g., Figure 21, indicator 2100) and recharges the electrode for a relatively short period at higher voltages (e.g., Figure 21, indicator 2101). In some cases, the detection period is at least 2, at least 3, at least 4, at least 5, at least 6, at least 8, at least 10, at least 15, at least 20, or at least 50 times longer than the period used for electrode recharging.
[0362] In some cases, the waveform changes in response to the input. In some cases, the input is the electrode depletion level. In some cases, the polarity and / or magnitude of the voltage are variable, at least in part, based on electrode depletion or depletion of current-carrier ions, and the waveform is irregular.
[0363] The ability to repeatedly detect and recharge the electrode over short time periods (e.g., less than about 5 seconds, less than about 1 second, less than about 500 ms, less than about 100 ms, less than about 50 ms, less than about 10 ms, or less than about 1 ms) allows for the use of electrodes that are much smaller than those used for sequencing multinucleotides through nanopores and which maintain a constant direct current (DC) potential and current. Smaller electrodes allow for a higher number of detection sites on the surface (e.g., incorporating the electrode, sensing circuitry, nanopores, and polymerase).
[0364] The surface contains discrete sites of any suitable density (e.g., a density suitable for sequencing nucleic acid samples over a given time period or for a given cost). In one embodiment, the surface has greater than or equal to about 500 sites / 1 mm. 2 The density of discrete sites. In some embodiments, the surface has approximately 100, approximately 200, approximately 300, approximately 400, approximately 500, approximately 600, approximately 700, approximately 800, approximately 900, approximately 1000, approximately 2000, approximately 3000, approximately 4000, approximately 5000, approximately 6000, approximately 7000, approximately 8000, approximately 9000, approximately 10000, approximately 20000, approximately 40000, approximately 60000, approximately 80000, approximately 100000, or approximately 500000 sites / 1 mm. 2 The density of discrete sites. In some embodiments, the surface has at least about 200, at least about 300, at least about 400, at least about 500, at least about 600, at least about 700, at least about 800, at least about 900, at least about 1000, at least about 2000, at least about 3000, at least about 4000, at least about 5000, at least about 6000, at least about 7000, at least about 8000, at least about 9000, at least about 10000, at least about 20000, at least about 40000, at least about 60000, at least about 80000, at least about 100000, or at least about 500000 sites / 1 mm. 2 The density of discrete sites.
[0365] The electrode can be recharged before, between, during, or after a nucleotide incorporation event. In some cases, the electrode is recharged within approximately 20 milliseconds (ms), approximately 40 ms, approximately 60 ms, approximately 80 ms, approximately 100 ms, approximately 120 ms, approximately 140 ms, approximately 160 ms, approximately 180 ms, or approximately 200 ms. In other cases, the electrode is recharged within less than approximately 20 milliseconds (ms), less than approximately 40 ms, less than approximately 60 ms, less than approximately 80 ms, less than approximately 100 ms, less than approximately 120 ms, less than approximately 140 ms, less than approximately 160 ms, less than approximately 180 ms, approximately 200 ms, less than approximately 500 ms, or less than approximately 1 second.
[0366] Chips capable of distinguishing between cut and uncut markings
[0367] On the other hand, a chip for sequencing nucleic acid samples is provided. In one example, the chip includes multiple individually addressable nanopores. The multiple individually addressable nanopores may have at least one nanopore formed in a membrane adjacent to an integrated circuit arrangement. Each individually addressable nanopore may be adapted to determine whether a label molecule is bound to or not bound to a nucleotide, or to read changes between different labels.
[0368] In some cases, the chip may contain multiple individually addressable nanopores. These multiple individually addressable nanopores may have at least one nanopore formed in the membrane adjacent to an integrated circuit arrangement. Each individually addressable nanopore may be adapted to determine whether a labeling molecule binds to an incorporated (e.g., polymerized) nucleotide or an unincorporated nucleotide.
[0369] The chip described in this article can distinguish between released and unreleased tags (e.g., Figure 4 (D vs. F). In some implementations, the chip is capable of distinguishing between released and unreleased tags with an accuracy of at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 97%, at least about 99%, at least about 99.5%, at least about 99.9%, at least about 99.95%, or at least about 99.99%. This level of accuracy is achieved when detecting groups of about 5, 4, 3, or 2 consecutive nucleotides. In some cases, accuracy is achieved for single-base resolution (i.e., 1 consecutive nucleotide).
[0370] The chip described herein can distinguish between incorporated and unincorporated labeled nucleotides (e.g., Figure 5(referring to 506 and 505 in the original text). In some embodiments, the chip is capable of distinguishing between incorporated and unincorporated labeled nucleotides with an accuracy of at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 97%, at least about 99%, at least about 99.5%, at least about 99.9%, at least about 99.95%, or at least about 99.99%. Accuracy levels are achieved when detecting groups of about 5, 4, 3, or 2 consecutive nucleotides. In some cases, accuracy is achieved for single-base resolution (i.e., 1 consecutive nucleotide).
[0371] Nanopores can help determine, at least in part, whether a labeled molecule is bound to or not of a nucleotide based on differences in electrical signals. In some cases, nanopores can help determine, at least in part, whether a labeled molecule is bound to or not of a nucleotide based on residence time within the nanopore. Nanopores can also help determine, at least in part, whether a labeled molecule is bound to or not of a nucleotide based on fall-out voltage (the voltage at which the label or labeled nucleotide leaves the nanopore).
[0372] Capable of capturing a high proportion of chips with cutting marks
[0373] On the other hand, a chip for sequencing nucleic acid samples is provided. In one example, the chip includes multiple individually addressable nanopores. The multiple individually addressable nanopores may include at least one nanopore formed in a membrane adjacent to an integrated circuit arrangement. Each individually addressable nanopore is adapted to capture most of the labeled molecules released upon incorporation (e.g., polymerization) of labeled nucleotides.
[0374] The chip can be configured to capture any reasonably high percentage of the marker (e.g., to determine nucleic acid sequences with reasonably high accuracy). In some embodiments, the chip captures at least 90%, at least 99%, at least 99.9%, or at least 99.99% of the marked molecules.
[0375] In some implementations, the nanopore captures multiple different labeled molecules at a single current level (e.g., four different labeled molecules released upon incorporation of four nucleotides). The chip can be adapted to capture labeled molecules in the same order as the order in which they are released.
[0376] Equipment settings
[0377] Figure 8A nanopore device 100 (or sensor) is illustrated illustratively for use in nucleic acid sequencing and / or detection of labeled molecules as described herein. The nanopore containing the lipid bilayer can be characterized by resistance and capacitance. The nanopore device 100 includes a lipid bilayer 102 formed on a lipid bilayer-compatible surface 104 of a conductive solid substrate 106, wherein the lipid bilayer-compatible surface 104 may be isolated by a lipid bilayer-incompatible surface 105, and the conductive solid substrate 106 may be electrically isolated by an insulating material 107, wherein the lipid bilayer 102 may be surrounded by amorphous lipids 103 formed on the lipid bilayer-incompatible surface 105. The lipid bilayer 102 may embed a single nanopore structure 108 having a size large enough to allow the labeled molecule to be characterized and / or small ions (e.g., Na+) between the two sides of the lipid bilayer 102. + K + Ca 2+ Cl - The device includes a nanopore 110 through which water molecules 114 can be adsorbed onto a lipid bilayer compatible surface 104, sandwiched between the lipid bilayer 102 and the lipid bilayer compatible surface 104. The aqueous film 114 adsorbed onto the hydrophilic lipid bilayer compatible surface 104 promotes the ordering of lipid molecules and facilitates the formation of a lipid bilayer on the lipid bilayer compatible surface 104. A sample chamber 116 containing a solution of nucleic acid molecules 112 and labeled nucleotides can be provided on the lipid bilayer 102. The solution may be an aqueous solution containing an electrolyte and can be buffered to an optimal ion concentration and maintained at an optimal pH to keep the nanopore 110 open. The device includes a pair of electrodes 118 (including a negative electrode 118a and a positive electrode 118b) coupled to a variable voltage source 120 for providing... Electrical stimulation (e.g., bias voltage) across the lipid bilayer and sensing of the lipid bilayer's electrical properties (e.g., resistance, capacitance, and ionic current). The surface of the positive electrode 118b is or forms part of a lipid bilayer-compatible surface 104. A conductive solid substrate 106 may be coupled to or formed as part of one of the electrodes 118. The device 100 may also include circuitry 122 for controlling the electrical stimulation and for processing the detected signal. In some embodiments, a variable voltage source 120 is included as part of circuitry 122. Circuitry 122 may include amplifiers, integrators, noise filters, feedback control logic, and / or various other components. Circuitry 122 may be an integrated circuit integrated within a silicon substrate 128 and may also be coupled to a computer processor 124 coupled to a memory 126.
[0378] The lipid bilayer-compatible surface 104 can be formed from a variety of materials suitable for ion conduction and gas formation to facilitate lipid bilayer formation. In some embodiments, conductive or semi-conductive hydrophilic materials may be used, as they allow for better detection of changes in the electrical properties of the lipid bilayer. Example materials include Ag-AgCl, Au, Pt, or doped silicon or other semiconductor materials. In some cases, the electrode is not a sacrificial electrode.
[0379] The lipid bilayer incompatible surface 105 can be formed of various materials unsuitable for lipid bilayer formation, and these are typically hydrophobic. In some embodiments, a non-conductive hydrophobic material is preferred because it not only separates the lipid bilayer regions from each other but also electrically insulates them. Examples of lipid bilayer incompatible materials include, for example, silicon nitride (e.g., Si3N4) and Teflon, and silicon oxide (e.g., SiO2) silanized with hydrophobic molecules.
[0380] In one instance, Figure 8 The nanoporous device 100 is an α-hemolysin (aHL) nanoporous device having a single α-hemolysin (aHL) protein 108 embedded in a diphyllylphosphatidylcholine (DPhPC) lipid bilayer 102 formed on a lipid bilayer-compatible silver (Ag) surface 104 coated on an aluminum material 106. The lipid bilayer-compatible Ag surface 104 is isolated by a lipid bilayer-incompatible silicon nitride surface 105, and the aluminum material 106 is electrically insulated by a silicon nitride material 107. The aluminum 106 is coupled to a circuit 122 integrated in a silicon substrate 128. A silver-silver chloride electrode, positioned on the chip or extending downward from the cover plate 128, is in contact with an aqueous solution containing nucleic acid molecules.
[0381] The aHL nanopore is an assembly of seven individual peptides. The entrance, or vestibule, of the aHL nanopore has a diameter of approximately 26 angstroms, wide enough to accommodate a portion of a dsDNA molecule. Starting from the vestibule, the aHL nanopore first widens and then narrows into a barrel with a diameter of approximately 15 angstroms, wide enough to allow a single ssDNA molecule (or a smaller marker molecule) to pass through but not wide enough to allow a dsDNA molecule (or a larger marker molecule) to pass through.
[0382] Besides DPhPC, the lipid bilayer of the nanoporous device can also be assembled from various other suitable amphiphilic materials, selected based on considerations such as the type of nanopore used, the type of molecule characterized, and the different physical, chemical, and / or electrical properties of the formed lipid bilayer, such as the stability and permeability, resistance, and capacitance of the formed lipid bilayer. Examples of amphiphilic materials include various phospholipids, such as palmitoyl-oleoyl-phosphatidyl-choline (POPC) and dioleoyl-phosphatidyl-methyl ester (DOPME), diphyranoyl-phosphatidylcholine (DPhPC), 1,2-di-O-phyranoyl- sn 3-glycerol-3-phosphate choline (DoPhPC), dipalmitoylphosphatidylcholine (DPPC), phosphatidylcholine, phosphatidylethanolamine, phosphatidylserine, phosphatidic acid, phosphatidylinositol, phosphatidylglycerol, and sphingomyelin.
[0383] Besides the aHL nanopores shown above, nanopores can also possess various other types. Examples include γ-hemolysin, leukocidin, melitoxin, mycobacterium smegmatis porin A (MspA), and various other naturally occurring, naturally modified, and synthetic nanopores. Suitable nanopores can be selected based on different characteristics of the analyte molecule, such as the size of the analyte molecule relative to the pore size of the nanopore. For example, aHL nanopores have a confining pore size of approximately 15 angstroms.
[0384] Current measurement
[0385] In some cases, current can be measured under different applied voltages. To achieve this, a desired potential can be applied to the electrodes, and the applied potential can then be maintained throughout the measurement. In one implementation, an operational amplifier integrator topology can be used for this purpose, as described below. The integrator maintains the voltage potential at the electrodes by means of capacitive feedback. Integrator circuits offer outstanding linearity, inter-cell matching, and offset characteristics. Operational amplifier integrators typically require large size to achieve the desired performance. A more compact integrator topology is described below.
[0386] In some cases, a voltage potential "Vliquid" can be applied to a chamber that provides a common potential (e.g., 350 mV) for all cells on the chip. An integrator circuit initializes the electrode (which is electrically the top plate of an integrating capacitor) to a potential greater than the common liquid potential. For example, applying a bias voltage of 450 mV can provide a positive 100 mV potential between the electrode and the liquid. This positive voltage potential allows current to flow from the electrode to the liquid chamber contact. In this case, the carriers are: (a) K+ ions flowing from the (opposite) side of the double-layer electrode through a via to the (forward) side of the double-layer reservoir and (b) chloride (Cl-) ions reacting with the silver electrode on the opposite side according to the following electrochemical reaction: Ag + Cl- → AgCl + e-.
[0387] In some cases, K+ flows out of the closed unit (from the reverse side to the cis side of the bilayer), while Cl- is converted to silver chloride. As a result of the current, the electrode side of the bilayer can become dilute. In some cases, a silver / silver chloride liquid sponge material or matrix can be used as a reservoir to supply Cl- ions in the reverse reaction, which appear at the electrical chamber contacts to complete the circuit.
[0388] In some cases, electrons eventually flow to the top side of the integrating capacitor, which generates a current that is measured. An electrochemical reaction converts silver to silver chloride, and the current continues to flow provided there is available silver to be converted. In some cases, the limited supply of silver results in a current-dependent electrode lifetime. In some embodiments, undepleted electrode materials (e.g., platinum) are used.
[0389] When a constant potential is applied to a nanopore detector, the tag can modulate the ion current through the nanopore, allowing current recording to determine the tag's identity. However, a constant potential may be insufficient to distinguish different tags (e.g., tags associated with A, C, T, or G). On the other hand, the applied voltage can be varied (e.g., frequency sweeping over a voltage range) to identify the tag (e.g., with at least 90%, at least 95%, at least 99%, at least 99.9%, or at least 99.99% confidence).
[0390] The applied voltage can be varied in any suitable manner (including any waveform shown in Figure 21). The voltage can be varied within any suitable range, including from about 120 mV to about 150 mV, and from about 40 mV to about 150 mV.
[0391] Figure 22 The extracted signals (e.g., differential logarithmic conductance (DLC)) of the nucleotides adenine (A, green), cytosine (C, blue), guanine (G, black), and thymine (T, red) are shown in contrast to the applied voltage. Figure 23The same information is shown for multiple nucleotides (as demonstrated in many experimental studies). As observed here, cytosine is relatively easily distinguished from thymine at 120 mV, but becomes difficult to distinguish at 150 mV (e.g., because the extracted signals for C and T are approximately equal at 150 mV). Furthermore, thymine is difficult to distinguish from adenine at 120 mV, but is relatively easy to distinguish at 150 mV. Therefore, in one embodiment, the applied voltage can be varied from 120 mV to 150 mV to distinguish each of nucleotides A, C, G, and T.
[0392] Figure 24 Nucleotides adenine (A, green), cytosine (C, blue), guanine (G, black), and thymine (T, red) are shown as the percentage of reference conductivity difference (RCD%) as a function of applied voltage. Plotting RCD% (which is essentially the difference in conductivity per molecule compared to a 30T reference molecule) removes offsets and adds variation between experiments. Figure 24 Includes individual DNA waveforms from the first block of the 17 / 20 trials. For all 17 good trials, the RCD% of all individual nucleotide DNA was captured from 50 to 200. The voltage indicating the distinguishable value of each nucleotide is shown.
[0393] although Figure 22-24 The response of a nucleotide to a varying applied voltage is shown, but the concept of a varying applied voltage can be used to distinguish labeled molecules (e.g., those attached to labeled nucleotides).
[0394] Unit circuit
[0395] Examples of unit circuits are shown in Figure 12 In the middle. An applied voltage Va is applied to operational amplifier 1200 before MOSFET current transmitter gate 1201. Electrode 1202 and the resistance of nucleic acid and / or label detected by device 1203 are also shown here.
[0396] An applied voltage Va drives the current transmitter gate 1201. The voltage generated on the electrode is then Va - Vt, where Vt is the MOSFET's threshold voltage. In some cases, this results in limited control over the actual voltage applied to the electrode because the MOSFET threshold voltage can vary considerably across processes, voltages, temperatures, and even within the device itself. This Vt variation can be significant at low current levels, where subthreshold leakage effects can begin to take effect. Therefore, to provide better control over the applied voltage, an operational amplifier can be used in a follower feedback configuration with the current transmitter device. This ensures that the voltage applied to the electrode is Va, independent of variations in the MOSFET threshold voltage.
[0397] Another example of a unit circuit is shown in Figure 10 It includes an integrator, a comparator, and a shift-in control bit, as well as digital logic that simultaneously shifts the state out of the comparator output. The unit circuitry is adaptable for use with the systems and methods provided herein. The B0-B1 link can be generated from a shift register. Analog signals are shared by all units within the group, while digital links can be daisy-chained between units.
[0398] The unit's digital logic comprises a 5-bit Data Shift Register (DSR), a 5-bit Parallel Load Register (PLR), control logic, and an analog integrator circuit. Using the LIN signal, control data shifted into the DSR is loaded into the PLR in parallel. These 5-bit control digital "break-before-make" sequential logic controls switching within the control unit. Furthermore, the digital logic includes a set-reset (SR) latch to record comparator output switching.
[0399] The architecture delivers a variable sampling rate proportional to the current in a single cell. Higher currents can produce more samples per second than lower currents. The resolution of current measurements is related to the current being measured. Smaller currents can be measured with better resolution compared to larger currents, which may be an advantage over fixed-resolution measurement systems. Analog inputs allow users to adjust the sampling rate by changing the voltage swing of the integrator. The sampling rate can be increased to analyze fast biological processes or decreased (and thus increased in precision) to analyze slow biological processes.
[0400] The integrator output is initialized to the voltage LVB (low voltage offset) and integrated to the voltage CMP. A sample is generated whenever the integrator output swings between these two levels. Therefore, a larger current results in a faster integrator output swing, and thus a faster sampling rate. Similarly, if the CMP voltage decreases, the integrator output swing required to generate a new sample decreases, and therefore the sampling rate increases. Thus, simply reducing the voltage difference between LVB and CMP provides a mechanism to increase the sampling rate.
[0401] Nanopore-based sequencing chips can incorporate a large number of autonomously operating or individually addressable units configured in arrays. For example, an array of one million units can be constructed from 1000 rows × 1000 columns. This array enables parallel sequencing of nucleic acid molecules, for instance, by measuring the difference in conductivity when a tag released at the time of nucleotide incorporation is detected through the nanopore. Furthermore, this circuit implementation allows for the determination of the conductivity characteristics of the pore-molecule complex, thereby identifying which tags might be valuable in distinguishing them.
[0402] Integrated nanopore / bilayer electronic unit structures allow for the application of appropriate voltages for current measurements. For example, both of the following may be necessary: (a) controlling the electrode voltage potential and (b) monitoring the electrode current simultaneously for proper measurement.
[0403] Furthermore, it may be necessary to control cells that are independent of each other. Managing a large number of cells that may be in different physical states may require independent control of the cells. Precise control of the piecewise linear voltage waveform stimulation applied to the electrodes can be used for the transition between the physical states of the cells.
[0404] To reduce circuit size and complexity, it may be sufficient to provide logic to apply two discrete voltages. This allows for two independent groupings of cells and the application of corresponding state transition stimuli. State transitions are inherently random with a relatively low probability of occurrence. Therefore, the ability to maintain an appropriate control voltage and subsequently measure to determine whether the desired state transition has occurred can be highly useful. For example, an appropriate voltage can be applied to the cells, and then the current can be measured to determine whether a double layer has formed. The cells are divided into two groups: (a) cells that already have a double layer and no longer require an applied voltage. These cells may be biased at 0V to produce a no-op (NOP) – they remain in the same state; and (b) cells that have not formed a double layer. These cells will then be subjected to a double-layer forming voltage.
[0405] Substantial simplification and circuit size reduction can be achieved by limiting the permissible applied voltage to 2 and repeatedly switching batches of cells between physical states. For example, reductions of at least 1.1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, or 100 times can be achieved by limiting the permissible applied voltage.
[0406] Another implementation of the invention using a compact measurement circuit is shown in Figure 11 In some cases, compact measurement circuitry can be used to achieve the high array density described herein. This circuit is also designed to apply voltage to the electrodes while simultaneously measuring low-level currents.
[0407] The cell operates as an ultra-compact integrator (UCI), and its basic operation is described herein. The cell is electrically connected to an electrochemically active electrode (e.g., AgCl) via an Electro-Sense (ELSNS) connection. The NMOS transistor M11 performs two independent functions: (1) operating as a source follower to apply a voltage to the ELSNS node given by (Vg1-Vt1), and (2) operating as a current conveyor to move electrons from capacitor C1 to the ELSNS node (and vice versa).
[0408] In some cases, a control voltage potential can be applied to the ELSNS electrode, and this can be varied simply by changing the voltage on the gate of the electrode source follower M11. Furthermore, any current from the source pin of M11 propagates directly and accurately to the drain pin of M11, where it accumulates on the capacitor C0. Therefore, M11 and C0 function together as an ultra-compact integrator. This integrator can be used to determine the current flowing out of the electrode and the current flowing into the electrode by measuring the change in voltage integrated on the capacitor according to the following equation: I*t = C*V, where I is the current, t is time, C is the capacitance, and V is the voltage change.
[0409] In some cases, voltage changes are measured at fixed intervals t (e.g., every 1 ms).
[0410] Transistor M2 can be configured as a source follower to buffer the capacitor voltage and provide a low-impedance representation of the integrated voltage. This prevents charge sharing from altering the voltage across the capacitor.
[0411] Transistor M3 can be used as a row access device, where the analog voltage output AOUT is connected in a column configuration shared with many other cells. Only a single row of column connections for the AOUT signal is implemented to allow the voltage of a single cell to be measured.
[0412] In an alternative implementation, transistor M3 can be omitted by connecting the drain of transistor M2 to a row-selectable "switched rail".
[0413] Transistor M4 can be used to reset the cell to a predetermined starting voltage, from which voltage integration is performed. For example, applying a high voltage (e.g., to VDD = 1.8V) to both RST and RV will pull the capacitor up to the pre-charged (VDD - Vt5) value. Due to reset switching thermal noise (sqrt(KTC) noise), the precise starting value can vary between cells and between measurements (due to Vt variations in M4 and M2). Therefore, correlated double sampling (CDS) technique is used to measure the integrator's starting and ending voltages to determine the actual voltage changes during the integration period.
[0414] It is also noted that the drain of transistor M4 can be connected to the control voltage RV (reset voltage). In normal operation, this can be driven to VDD; however, it can also be driven to a low voltage. If the "drain" of M4 is actually driven to ground, then the current can be reversed (i.e., current can flow into the circuit from the electrodes through M1 and M4, and the concepts of source and drain can be interchanged). In some cases, when the circuit is operated in this mode, the negative voltage applied to the electrodes (relative to a liquid reference) is controlled by this RV voltage (assuming Vg1 and Vg5 are at least greater than the threshold of RV). Therefore, the ground voltage on RV can be used to apply a negative voltage to the electrodes (e.g., to complete electroporation or double-layer formation).
[0415] The analog-to-digital converter (ADC, not shown) measures the AOUT voltage immediately after reset and again after the integration period (when CDS measurement is performed) to determine the integrated current over a fixed time period. The ADC can be implemented as an analog multiplexer based on each column or discrete transistors used for each column to share a single ADC across multiple columns. The column multiplexer factor can be varied according to noise, accuracy, and throughput requirements.
[0416] At any given time, each unit can be in one of four different physical states: (1) short-circuited to liquid (2) bilayer formation (3) bilayer + pore (4) bilayer + pore + nucleic acid and / or labeled molecules.
[0417] In some cases, a voltage is applied to cause the cell to move between states. The NOP operation is used to keep a cell in a specific desired state, while other cells are stimulated by the application of a potential to move from one state to another.
[0418] This can be accomplished by having two (or more) different voltages with gate voltages that can be applied to the source follower M1, which can be indirectly used to control the voltage applied to the electrode relative to the liquid potential. Therefore, transistor M5 is used to apply voltage A, while transistor M6 is used to apply voltage B. Thus, M5 and M6 together operate as an analog multiplexer, where SELA or SELB is driven high to select the voltage.
[0419] Since each cell can be in a different state and because SELA and SELB are complementary, a memory element can be used in each cell to select between voltage A and B. This memory element can be a dynamic element (capacitor) refreshed in each cycle or a simple cheat-latch memory element (a cross-coupled inverter).
[0420] Operational amplifier test chip structure
[0421] In some instances, the test chip comprises an array of 264 sensors arranged in four discrete groups (also called clusters), with 66 sensor units in each discrete group. Each group, in turn, is divided into three "columns," with 22 sensor "units" in each column. The term "unit" is appropriate because, ideally, a virtual unit consisting of a lipid bilayer and inserted nanopores is formed above each of the 264 sensors in the array (although the device can be successfully operated with only a portion of such a densely packed array of sensor units).
[0422] A single analog I / O pad applies a voltage potential to a liquid contained within a conductive cylinder mounted to the surface of a template. This "liquid" potential is applied to the top side of the orifice and is common to all units in the detector array. The bottom side of the orifice has exposed electrodes, and each sensor unit can have a different bottom-side potential applied to its electrodes. The current between the top liquid connection and the electrode connection of each unit on the bottom side of the orifice is then measured. The sensor unit measures the current traveling through the orifice, regulated by labeled molecules passing within it.
[0423] In some cases, five bits control the mode of each sensor unit. See further reference. Figure 9 Each of the 264 cells in the array can be controlled individually. Values are applied to groups of 66 cells. The mode of each of the 66 cells in a group is controlled by continuously shifting 330 (66 * 5 bits / cell) values into the DataShiftRegister (DSR). These values are shifted into the array using the KIN (clock) and DIN ((dat in)) pins, with a discrete pin pair for each group of 66 cells.
[0424] Therefore, 330 clock cycles are used to shift 330 bits into the DSR shift register. When the corresponding LIN... When the (load input) is asserted as high, a second 330-bit parallel load register (PLR) is loaded in parallel from this shift register. While the PLR is being loaded in parallel, the cell's status value is simultaneously loaded into the DSR.
[0425] A complete operation consists of 330 clock cycles to shift 330 data bits into the DSR, a single clock cycle (with the LIN signal asserted high), followed by 330 clock cycles to read the captured status data shifted out of the DSR. Pipelining the operation allows new 330 bits to be shifted into the DSR while simultaneously reading 330 bits from the array. Therefore, at a 50MHz clock frequency, the cycle time for reading is 331 / 50MHz = 6.62µs.
[0426] Arrays of nanopores for sequencing
[0427] This disclosure provides an array of nanopore detectors (or sensors) for nucleic acid sequencing. Reference Figure 7 Multiple nucleic acid molecules can be sequenced on an array of nanopore detectors. Here, each nanopore location (e.g., 701) contains a nanopore, which in some cases is attached to a polymerase and / or phosphatase. Sensors are also generally present at each array location as described elsewhere in this document.
[0428] In some instances, an array of nanopores attached to a nucleic acid polymerase is provided, and labeled nucleotides are incorporated into the polymerase. During polymerization, the label is detected through the nanopores (e.g., by release and permeation through the nanopores, or by presentation to the nanopores). The array of nanopores can have any suitable number of nanopores. In some cases, the array contains approximately 200, approximately 400, approximately 600, approximately 800, approximately 1000, approximately 1500, approximately 2000, approximately 3000, approximately 4000, approximately 5000, approximately 10000, approximately 15000, approximately 20000, approximately 40000, approximately 60000, approximately 80000, approximately 100000, approximately 200000, approximately 400000, approximately 600000, approximately 800000, approximately 1000000, and so on. In some cases, the array contains at least 200, at least 400, at least 600, at least 800, at least 1000, at least 1500, at least 2000, at least 3000, at least 4000, at least 5000, at least 10000, at least 15000, at least 20000, at least 40000, at least 60000, at least 80000, at least 100000, at least 200000, at least 400000, at least 600000, at least 800000, or at least 1,000000 nanopores.
[0429] In some cases, a single label is released and / or presented upon incorporation into a single nucleotide and detected through a nanopore. In other cases, multiple labels are released and / or presented upon incorporation into multiple nucleotides. A nanopore sensor adjacent to the nanopore can detect a single label or multiple labels. One or more signals associated with multiple labels can be detected and processed to obtain an average signal.
[0430] The markers can be detected over time by sensors. Time-detected markers can be used to determine the nucleic acid sequence of a nucleic acid sample, such as by means of a computer system programmed to record sensor data and generate sequence information from the data (see, for example, Figure 16 ).
[0431] Nanopore detector arrays can have high-density discrete sites. For example, a relatively large number of sites per unit area (i.e., density) allows for the construction of smaller devices that are portable, low-cost, or have other advantageous features. Individual sites in the array can be individually addressable. The large number of sites, including nanopores and sensing circuitry, allows for the sequencing of a relatively large number of nucleic acid molecules at once, such as through parallel sequencing. Such systems can increase throughput and / or reduce the cost of sequencing nucleic acid samples.
[0432] Nucleic acid samples can be sequenced using sensors (or detectors) with a substrate (whose surface contains discrete sites), each individual site having a nanopore, a polymerase, and in some cases, at least one phosphatase attached to the nanopore and sensing circuitry adjacent to the nanopore. The system may also include a flow cell in fluid communication with the substrate, the flow cell being adapted to deliver one or more reagents to the substrate.
[0433] The surface contains discrete sites of any suitable density (e.g., a density suitable for sequencing nucleic acid samples at a given time or cost). Each discrete site can include a sensor. The surface may have approximately 500 sites / 1 mm or more. 2 The density of discrete sites. In some embodiments, the surface has approximately 200, approximately 300, approximately 400, approximately 500, approximately 600, approximately 700, approximately 800, approximately 900, approximately 1000, approximately 2000, approximately 3000, approximately 4000, approximately 5000, approximately 6000, approximately 7000, approximately 8000, approximately 9000, approximately 10000, approximately 20000, approximately 40000, approximately 60000, approximately 80000, approximately 100000, or approximately 500000 sites / 1 mm. 2 The density of discrete sites. In some cases, the surface has at least 200, at least 300, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, at least 1000, at least 2000, at least 3000, at least 4000, at least 5000, at least 6000, at least 7000, at least 8000, at least 9000, at least 10000, at least 20000, at least 40000, at least 60000, at least 80000, at least 100000, or at least 500000 sites / 1 mm. 2 The density of discrete sites.
[0434] Labeled nucleotides
[0435] In some cases, labeled nucleotides contain a tag that can be cleaved during a nucleotide incorporation event and detected by means of a nanopore. The tag may be attached to the 5'-phosphate of the nucleotide. In some cases, the tag is not a fluorophore. The tag is detectable by its charge, shape, size, or any combination thereof. Examples of tags include various polymers. Each type of nucleotide (i.e., A, C, G, T) generally contains a unique tag.
[0436] The marker can be placed at any suitable position on the nucleotide. Figure 13 Examples of labeled nucleotides are provided. Here, while R1 is generally OH and R2 is H (i.e., for DNA) or OH (i.e., for RNA), other modifications are also acceptable. Figure 13 In this context, X represents any suitable linker. In some cases, the linker is cleavable. Examples of linkers include, but are not limited to, O, NH, S, or CH2. Examples of suitable chemical groups for position Z include O, S, or BH3. The base is any base suitable for incorporation into nucleic acids (including adenine, guanine, cytosine, thymine, uracil, or derivatives thereof). In some cases, universal bases are also acceptable.
[0437] The number of phosphates (n) is any suitable integer value (e.g., many phosphates allow nucleotides to be incorporated into nucleic acid molecules). In some cases, all types of labeled nucleotides have the same number of phosphates, but this is not necessary. In some applications, there are different labels for each type of nucleotide and the number of phosphates is not necessarily used to distinguish the various labels. However, in some cases, more than one type of nucleotide (e.g., A, C, T, G, or U) have the same labeled molecule and the ability to distinguish one nucleotide from another is determined at least in part by the number of phosphates (where the various types of nucleotides have different values for n). In some embodiments, the value of n is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or greater.
[0438] Suitable markers are described below. In some cases, the charge sign of the marker is opposite to that of the rest of the compound. When the marker is attached, the charge on the entire compound can be neutral. The release of the marker produces two molecules: a charged marker and a charged nucleotide. The charged marker passes through nanopores and is detected in some cases.
[0439] More examples of suitable labeled nucleotides are shown in Figure 14 The label can be attached to sugar molecules, base molecules, or any combination thereof. (See reference) Figure 13 Y is the label and X is the connector (which is cleavable in some cases). Furthermore, R1 (if present) is generally OH, -OCH2N3, or -O-2-nitrobenzyl, and R2 (if present) is generally H. Additionally, Z is generally O, S, or BH3, and n is any integer including 1, 2, 3, or 4. In some cases, A is O, S, CH2, CHF, CFF, or NH.
[0440] Continue to refer to Figure 14 Each dNPP analog generally has a different base type than each of the other three dNPP analogs, and the labeling type on each dNPP analog generally differs from the labeling type on each of the other three dNPP analogs. Suitable bases include, but are not limited to, adenine, guanine, cytosine, uracil, or thymine, or derivatives of each thereof. In some cases, the base is one of 7-denitroguanine, 7-denitroadenine, or 5-methylcytosine.
[0441] When R1 is -O-CH2N3, the method may further include treating the incorporated dNPP analog to remove -CH2N3 and generate an OH group attached to the 3' position, thereby allowing the incorporation of another dNPP analog.
[0442] When R1 is -O-2-nitrobenzyl, the method may further include treating the incorporated nucleotide analog to remove the -2-nitrobenzyl group and generate an OH group attached to the 3' position, thereby allowing the incorporation of another dNPP analog.
[0443] Instances of tags
[0444] The label can be any chemical group or molecule that can be detected in nanopores. In some cases, the label contains one or more of the following: ethylene glycol, amino acid, carbohydrate, peptide, dye, chemiluminescent compound, mononucleotide, dinucleotide, trinucleotide, tetranucleotide, pentanucleotide, hexanucleotide, aliphatic acid, aromatic acid, alcohol, thiol group, cyano group, nitro group, alkyl group, alkenyl group, alkynyl group, azido group, or combinations thereof.
[0445] It should also be assumed that the label contains an appropriate amount of lysine or arginine to balance the amount of phosphate in the compound.
[0446] In some cases, the label is used to refer to a polymer. Polyethylene glycol (PEG) is an example of a polymer and has the following structure:
[0447]
[0448] Any number of ethylene glycol units (W) can be used. In some cases, W is an integer from 0 to 100. In some cases, the number of ethylene glycol units is different for each type of nucleotide. In one embodiment, the four types of nucleotides contain labels with 16, 20, 24, or 36 ethylene glycol units. In some cases, the label also contains additional identifiable portions, such as coumarin-based dyes. In some cases, the polymer is charged. In some cases, the polymer is uncharged and the label is detected in high concentrations of salt (e.g., 3-4 M).
[0449] As used herein, the term "alkyl" includes a branched and straight-chain saturated aliphatic hydrocarbon group having a specified number of carbon atoms and may be unsubstituted or substituted. As used herein, "alkenyl" refers to a straight-chain or branched non-aromatic hydrocarbon group containing at least one carbon-carbon double bond and may contain up to a maximum possible number of non-aromatic carbon-carbon double bonds, and may be unsubstituted or substituted. The term "alkynyl" refers to a straight-chain or branched hydrocarbon group containing at least one carbon-carbon triple bond and may contain up to a maximum possible number of non-aromatic carbon-carbon triple bonds, and may be unsubstituted or substituted. The term "substituted" refers to a functional group such as an alkyl or hydrocarbon group as described above, in which at least one bond with a hydrogen atom contained therein is replaced by a bond with a non-hydrogen or non-carbon atom, provided that the normal valence state is maintained and said substitution produces a stable compound. Substituted groups also include groups in which one or more bonds with carbon or hydrogen atoms are replaced by one or more bonds (including double or triple bonds) with heteroatoms.
[0450] In some cases, the marker can only pass through the nanopore in one direction (e.g., without reversing the direction). The marker may have a hinged gate attached to it, and the marker may be thin enough to pass through the nanopore when the gate is aligned with the marker in one direction, rather than the other. Reference Figure 31 This disclosure provides a marker molecule comprising a first polymer chain 3105 having a first segment 3110 and a second segment 3115, wherein the second segment is narrower than the first segment. The second segment may have a width smaller than the narrowest opening of the nanopore. The marker molecule may include a second polymer chain 3120 comprising two ends, wherein a first end is attached to the first polymer chain adjacent to the second segment and a second end is not attached to the first polymer chain. The marker molecule is capable of passing through the nanopore in a first direction, wherein the second polymer chain is aligned adjacent to the second segment 3125. In some cases, the marker molecule is not capable of passing through the nanopore in a second direction in which the second polymer chain is not aligned adjacent to the second segment 3130. The second direction may be opposite to the first direction.
[0451] The first and / or second polymer chain may contain nucleotides. In some cases, the second polymer chain pairs with the bases of the first polymer chain when the second polymer chain is not aligned with the adjacent second segment. In some cases, the first polymer chain is immobilized at nucleotide 3135 (e.g., immobilized at the terminal phosphate of the nucleotide). The first polymer chain may be released from the nucleotide when the nucleotide is incorporated into the growing nucleic acid chain.
[0452] The second segment may contain any polymer or other molecule that is thin enough to pass through the nanopore when aligned with the gate (second polymer). For example, the second segment may contain debased nucleotides (i.e., nucleic acid chains without any nucleic acid bases) or carbon chains.
[0453] This disclosure also provides a method for sequencing nucleic acid samples using nanopores adjacent to sensing electrodes in a membrane. (Reference) Figure 32 The method may include providing a labeled nucleotide 3205 to a reaction chamber containing a nanopore, wherein a single labeled nucleotide contains a label coupled to the nucleotide, wherein the label is detectable by means of the nanopore. The label comprises a first polymer chain containing a first segment and a second segment (wherein the second segment is narrower than the first segment), and a second polymer chain containing two ends, wherein the first end is attached to the first polymer chain adjacent to the second segment and the second end is not attached to the first polymer chain. The labeled molecule is capable of passing through the nanopore in a first direction 3210, wherein the second polymer chain is aligned adjacent to the second segment.
[0454] The method includes a polymerization reaction using polymerase 3215 to incorporate a single labeled nucleotide into a growth chain 3220 complementary to a single-stranded nucleic acid molecule 3225 from a nucleic acid sample. The method may include detecting the label associated with the single labeled nucleotide during incorporation using a nanopore 3230, wherein the label is detected by means of the nanopore when the nucleotide associates with the polymerase.
[0455] In some cases, the labeled molecules are unable to pass through the nanopore in a second direction in which the second polymer chain is not aligned with the second segment.
[0456] The label can be detected multiple times upon association with the polymerase. In some embodiments, the electrode is recharged between label detection periods. In some cases, the label penetrates the nanopore during the incorporation of a single labeled nucleotide and does not exit the nanopore when the electrode is recharged.
[0457] Methods for attaching tags
[0458] Any suitable method for attaching the label can be used. In one example, the label can be attached to the terminal phosphate by: (a) contacting the nucleoside triphosphate with dicyclohexylcarbodiimide / dimethylformamide under conditions that allow the formation of a cyclic trimetaphosphate; (b) contacting the product obtained from step a) with a nucleophile to form a -OH or -NH2 functionalized compound; and (c) reacting the product of step b) with a label having a -COR group attached thereto under conditions that allow the label to be indirectly bound to the terminal phosphate, thereby forming a nucleoside triphosphate analog.
[0459] In some cases, the nucleophile is H₂N-R-OH, H₂N-R-NH₂, R'S-R-OH, R'S-R-NH₂, or
[0460] .
[0461] In some cases, the method includes, in step b), contacting the product obtained from step a) with a compound having the following structure:
[0462]
[0463] The product is then or simultaneously contacted with NH4OH to form a compound having the following structure:
[0464] ,
[0465] The product of step b) can then be reacted with a label having a -COR group attached thereto, under conditions that allow the label to be indirectly bound to the terminal phosphate, thereby forming a nucleoside triphosphate analog having the following structure:
[0466]
[0467] Wherein R1 is OH, and R2 is H or OH, wherein the base is adenine, guanine, cytosine, thymine, uracil, 7-deazopurine, or 5-methylpyrimidine.
[0468] Release of the marker
[0469] The label can be released in any manner. The label can be released during or after the incorporation of the labeled nucleotide into the growing nucleic acid chain. In some cases, the label is attached to a polyphosphate (e.g., Figure 13 Incorporating nucleotides into nucleic acid molecules leads to the release of polyphosphates with attached labels. Incorporation can be catalyzed by at least one polymerase that can attach to the nanopore. In some cases, at least one phosphatase also attaches to the pore. The phosphatase can cleave the label from the polyphosphate to release it. In some cases, the phosphatase is positioned such that the pyrophosphate generated by the polymerase in the polymerase reaction interacts with the phosphatase before entering the pore.
[0470] In some cases, the marker is not attached to the polyphosphate (see, for example, Figure 14 In these cases, the marker is attached via a cleavable adapter (X). A method for generating cleavable capped and / or cleavable linked nucleotide analogs is disclosed in U.S. Patent No. 6,664,079, which is incorporated herein by reference in its entirety. The adapter is not necessarily cleavable.
[0471] The connector can be any suitable connector and can be cut in any suitable manner. The connector can be photocuttable. In one embodiment, the photocuttable connector and portion are cut using UV light via photochemical cutting. In one embodiment, the photocuttable connector is a 2-nitrobenzyl portion.
[0472] The -CH2N3 group can be treated with TCEP (tris(2-carboxyethyl)phosphine) to remove it from the 3' O atom of dNPP or rNPP analogs, thereby producing a 3' OH group.
[0473] Label detection
[0474] In some cases, the polymerase pulls from an aggregate of labeled nucleotides containing multiple different bases (e.g., A, C, G, T, and / or U). It is also possible to repeatedly contact the polymerase with various types of labeled bases. In this case, each type of nucleotide may not necessarily have a unique base, but in some cases, cycling between different base types increases the cost and complexity of the process; however, this invention covers such embodiments.
[0475] Figure 15 This demonstrates that incorporating labeled nucleotides into nucleic acid molecules (e.g., using polymerases to extend primers that pair with template bases) can release labeled polyphosphates that are detectable in some embodiments. In some cases, the labeled polyphosphates are detected when they pass through a nanopore. In some embodiments, the labeled polyphosphates are detected when they remain in the nanopore.
[0476] In some cases, the method distinguishes nucleotides based on the number of phosphate groups that make up the polyphosphate (e.g., even when the tags are the same). However, each type of nucleotide generally has a unique tag.
[0477] refer to Figure 15 Labeled polyphosphate compounds can be treated with phosphatases (e.g., alkaline phosphatases) before labeling and / or passing through nanopores and before measuring ion currents.
[0478] The tags can flow through nanopores after they are released from nucleotides. In some cases, a voltage is applied to pull the tags through the nanopores. At least about 85%, at least 90%, at least 95%, at least 99%, at least 99.9%, or at least 99.99% of the released tags can be displaced through the nanopores.
[0479] In some cases, the tags remain in the nanopores for a certain period of time and are detected within the nanopores. In other cases, a voltage is applied to pull the tags into the nanopores, the tags are detected, the tags are expelled from the nanopores, or any combination thereof. The tags may be released at the time of a nucleotide incorporation event or remain bound to the nucleotide.
[0480] The label can be detected in nanopores (at least in part) because of its charge. In some cases, the labeled compound is optionally charged, having a first net charge and a different second net charge after a chemical, physical, or biological reaction. In some cases, the charge on the label is equal to the charge on the rest of the compound. In one embodiment, the label has a positive charge, and removing the label changes the charge of the compound.
[0481] In some cases, when a label is introduced and / or passes through a nanopore, it can produce electronic changes. In some cases, these electronic changes are changes in current magnitude, changes in nanopore conductivity, or any combination thereof.
[0482] Nanopores can be biological or synthetic. It is also envisioned that the pores are proteins, for example, α-hemolysin. Examples of synthetic nanopores are solid-state pores or graphene.
[0483] In some cases, polymerases and / or phosphatases are attached to nanopores. Fusion proteins or disulfide crosslinks are examples of methods for attaching proteins to nanopores. In the case of solid nanopores, attachment to the surface immediately adjacent to the nanopore can be via biotin-streptavidin bonding. In one example, DNA polymerase is attached to a solid surface via a monolayer modified with an amino-functionalized alkylthiol-based self-assembly, wherein the amino groups are modified into NHS esters for attachment to the amino groups on the DNA polymerase.
[0484] The method can be carried out at any suitable temperature. In some embodiments, the temperature is between 4°C and 10°C. In some embodiments, the temperature is room temperature.
[0485] The method can be performed in any suitable solution and / or buffer. In some cases, the buffer is 300 mM KCl buffered to pH 7.0 to 8.0 with 20 mM HEPES. In some embodiments, the buffer does not contain divalent cations. In some cases, the method is not affected by the presence of divalent cations.
[0486] Computer systems for sequencing nucleic acid samples
[0487] The nucleic acid sequencing system and method disclosed herein can be controlled using a computer system. Figure 16 A system 1600 is shown, including a computer system 1601 coupled to a nucleic acid sequencing system 1602. The computer system 1601 may be a server or multiple servers. The computer system 1601 may be programmed to regulate sample preparation and processing, as well as nucleic acid sequencing via the sequencing system 1602. The sequencing system 1602 may be a nanopore-based sequencer (or detector), as described elsewhere herein.
[0488] A computer system can be programmed to implement the methods of the present invention. Computer system 1601 includes a central processing unit (CPU, also referred to herein as a "processor") 1605, which may be a single-core or multi-core processor, or multiple processors for parallel processing. Computer system 1601 also includes memory 1610 (e.g., random access memory, read-only memory, flash memory), electronic storage unit 1615 (e.g., hard disk), communication interface 1620 for communicating with one or more other systems (e.g., network adapter), and peripheral devices 1625, such as cache memory, other memory, data storage, and / or electronic display adapter. Memory 1610, storage unit 1615, interface 1620, and peripheral devices 1625 communicate with CPU 1605 via a communication bus (solid line), such as a motherboard. Storage unit 1615 may be a data storage unit (or data repository) for storing data. Computer system 1601 may be operatively coupled to a computer network ("network") by means of communication interface 1620. A network can be the Internet, an intranet and / or an extranet, or a local area network (LAN) and / or an extranet that communicates with the Internet. A network may include one or more computer servers that enable distributed computing.
[0489] The method of the present invention can be implemented by machine (or computer processor) executable code (or software) stored in an electronic storage location of computer system 1601, such as, for example, in memory 1610 or electronic storage unit 1615. During use, the code can be executed by processor 1605. In some cases, the code can be retrieved by storage unit 1615 and stored by processor 1605 in memory 1610 for fast access. In some cases, electronic storage unit 1615 can be excluded, and machine executable instructions are stored in memory 1610.
[0490] The code may be pre-compiled and configured for use with a machine having a processor suitable for executing the code, or it may be compiled at runtime. The code may be provided in a programming language, which may be selected to implement the code in a pre-compiled or compiled state.
[0491] Computer system 1601 may be adapted to store user profile information, such as, for example, name, physical address, email address, telephone number, instant messaging (IM) processing, educational information, work information, social interests and / or dislikes, and other information potentially relevant to the user or other users. Such profile information may be stored on storage unit 1615 of computer system 1601.
[0492] The aspects of the systems and methods provided herein, such as computer system 1601, can be embodied in programming. Various aspects of the technology can be considered as "products" or "articles of art," typically carried on or embodied in a type of machine-readable medium containing machine (or processor) executable code and / or associated data. Machine-executable code can be stored in electronic storage units, such memories (e.g., ROM, RAM), or hard disks. "Storage" type media can include any or all tangible memories of computers, processors, etc., or their associated modules, such as various semiconductor memories, magnetic tape drives, disk drives, etc., which can provide non-volatile memory for software programs at any time. All or part of the software can sometimes be communicated via the Internet or various other telecommunications networks. For example, such communication enables software to be loaded from one computer or processor to another, such as from a management server or host computer platform to an application server. Therefore, another type of medium that can carry software elements includes light waves, radio waves, and electromagnetic waves used in physical interfaces between local devices, such as via wired and optical route networks and via various air links. Physical elements carrying such waves, such as wired or wireless links, optical links, etc., can also be considered as media carrying software. As used herein, unless limited to non-volatile, tangible "memory" media, the term "readable medium" for a computer or machine refers to any medium that participates in providing instructions to a processor for execution.
[0493] Therefore, machine-readable media, such as computer-executable code, can take many forms, including but not limited to tangible storage media, carrier media, or physical transmission media. Non-volatile storage media include, for example, optical discs or disks, any storage device such as any computer, such as those used to implement databases as shown in the accompanying drawings. Volatile storage media include dynamic memory, such as the main memory of a computer platform. Tangible transmission media include coaxial cables; copper wires and optical fibers, including wires contained within a bus in a computer system. Carrier transmission media can take the form of electrical or electromagnetic signals, or sound or light waves such as those generated during radio frequency (RF) and infrared (IR) data communication. Common forms of computer-readable media therefore include, for example: floppy disks, floppy hard disks, hard disks, magnetic tapes, any other magnetic media, CD-ROMs, DVDs or DVD-ROMs, any other optical media, punched cardstock, any other physical storage media with a perforated pattern, RAM, ROM, PROM and EPROM, FLASH-EPROM, any other memory chips or tape cartridges, carrier waves for transmitting data or instructions, cables or links for transmitting such carrier waves, or any other media from which a computer can read programming code and / or data. Many of these forms of computer-readable media may involve transmitting one or more sequences of one or more instructions to a processor for execution.
[0494] The systems and methods disclosed herein can be used for sequencing various types of biological samples, such as nucleic acids (e.g., DNA, RNA) and proteins. In some embodiments, the methods, devices, and systems described herein can be used to sort biological samples (e.g., proteins or nucleic acids). The sorted samples and / or molecules can be directed to various bins for further analysis.
[0495] sequencing accuracy
[0496] The method provided herein accurately distinguishes single nucleotide incorporation events (e.g., single-molecule events). This method can accurately distinguish single nucleotide incorporation events in a single pass—that is, without necessarily re-sequencing a given nucleic acid molecule. In some cases, the method provided herein can be used for sequencing and re-sequencing nucleic acid molecules, or for sensing a label associated with a labeled molecule, once or multiple times. For example, a label can be sensed at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, or 10,000 times using a nanopore. The label can be sensed and re-sensed, for example, by applying a voltage to a membrane having a nanopore, which can pull the label into or expel the label from the nanopore.
[0497] Methods for nucleic acid sequencing include distinguishing individual nucleotide incorporation events with an accuracy greater than about 4σ. In some cases, nucleotide incorporation events are detected by means of nanopores. A tag associated with the nucleotide may be released at the time of incorporation and the tag may pass through the nanopore. In some cases, the tag is not released (e.g., presented to the nanopore). In many other embodiments, the tag is released but remains in the nanopore (e.g., does not pass through the nanopore). Different tags may be associated with and / or released from each type of nucleotide (e.g., A, C, T, G) and detected through the nanopore. Errors include, but are not limited to, (a) failure to detect a tag, (b) misidentification of a tag, (c) detection of a tag in the absence of a tag, (d) detection of a tag in an incorrect order (e.g., two tags are released in the first order but detected in the second order), (e) detection of a tag that has not yet been released from the nucleotide at the time of release, (f) detection of a tag not attached to the incorporated nucleotide when incorporated into a growing nucleotide chain, or any combination thereof. In some implementations, the accuracy of distinguishing individual nucleotide incorporation events is 100% minus the rate at which errors occur (i.e., the error rate).
[0498] The accuracy for distinguishing individual nucleotide incorporation events is any suitable percentage. Accuracy for distinguishing individual nucleotide incorporation events can be approximately 93%, approximately 94%, approximately 95%, approximately 96%, approximately 97%, approximately 98%, approximately 99%, approximately 99.5%, approximately 99.9%, approximately 99.99%, approximately 99.999%, approximately 99.999%, etc. In some cases, the accuracy for distinguishing individual nucleotide incorporation events is at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, at least 99.9%, at least 99.99%, at least 99.999%, at least 99.999%, etc. In some cases, the accuracy for distinguishing individual nucleotide incorporation events is reported in sigma(σ). σ is a statistical variable, sometimes used in business management and manufacturing strategies to report error rates such as the percentage of defect-free products. Here, the σ value can be used interchangeably with the accuracy according to the following relationship: 4σ is 99.38% accuracy, 5σ is 99.977% accuracy, and 6σ is 99.99966% accuracy.
[0499] Distinguishing individual nucleotide incorporation events as described herein can be used to accurately determine nucleic acid sequences. In some cases, the determination of nucleic acid sequences (e.g., DNA and RNA) involves errors. Examples of errors include, but are not limited to, deletions (failure to detect nucleic acid), insertions (detection of nucleic acid when no nucleic acid is actually present), and substitutions (detection of incorrect nucleic acid). The accuracy of nucleic acid sequencing can be determined by comparing the measured nucleic acid sequence with a true nucleic acid sequence (e.g., according to bioinformatics techniques) and the percentage of nucleic acid sites identified as deletions, insertions, and / or substitutions. Errors are any combination of deletions, insertions, and substitutions. Accuracy ranges from 0% to 100%, where 100% represents a completely correct determination of the nucleic acid sequence. Similarly, an error rate of 100% represents accuracy and ranges from 0% to 100%, where 0% error represents a completely correct determination of the nucleic acid sequence.
[0500] Nucleic acid sequencing performed according to the methods described herein and / or using the devices described herein achieves high accuracy. Accuracy is any reasonably high value. In some cases, accuracy is approximately 95%, approximately 95.5%, approximately 96%, approximately 96.5%, approximately 97%, approximately 97.5%, approximately 98%, approximately 98.5%, approximately 99%, approximately 99.5%, approximately 99.9%, approximately 99.99%, approximately 99.999%, approximately 99.999%, etc. In other cases, accuracy is at least 95%, at least 95.5%, at least 96%, at least 96.5%, at least 97%, at least 97.5%, at least 98%, at least 98.5%, at least 99%, at least 99.5%, at least 99.9%, at least 99.99%, at least 99.999%, at least 99.999%, at least 99.999%, etc. In some cases, the accuracy is between approximately 95% and 99.9999%, between approximately 97% and 99.9999%, between approximately 99% and 99.9999%, between approximately 99.5% and 99.9999%, between approximately 99.9% and 99.9999%, etc.
[0501] High accuracy can be achieved by performing multiple passes (i.e., sequencing nucleic acid molecules multiple times, for example by passing the nucleic acid through or near a nanopore and sequencing the bases of the nucleic acid molecule). Data from multiple passes can be combined (e.g., deletions, insertions, and / or substitutions from the first pass are corrected using data from other repeated passes). In some cases, the accuracy of label detection can be increased by passing the label through or near the nanopore multiple times, such as by reversing the voltage applied to the nanopore or membrane (e.g., DC or AC voltage). The method provides high accuracy with very few passes (also known as reads, the multiplicity of sequencing coverage). The number of passes is any suitable number and does not have to be an integer. In some embodiments, nucleic acid molecules are sequenced 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 12, 14, 16, 18, 20, 25, 30, 35, 40, 45, 50, etc. In some implementations, nucleic acid molecules are sequenced up to 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 12, 14, 16, 18, 20, 25, 30, 35, 40, 45, or 50 times. In some implementations, nucleic acid molecules are sequenced approximately 1 to 10 times, approximately 1 to 5 times, or approximately 1 to 3 times. Accuracy levels can be achieved by combining data collected from up to 20 passes. In some implementations, accuracy levels are achieved by combining data collected from up to 10 passes. In some implementations, accuracy levels are achieved by combining data collected from up to 5 passes. In some cases, accuracy levels are achieved in a single pass.
[0502] The error rate is any appropriately low rate. In some cases, the error rate is approximately 10%, approximately 5%, approximately 4%, approximately 3%, approximately 2%, approximately 1%, approximately 0.5%, approximately 0.1%, approximately 0.01%, approximately 0.001%, approximately 0.0001%, etc. In some cases, the error rate is at most 10%, at most 5%, at most 4%, at most 3%, at most 2%, at most 1%, at most 0.5%, at most 0.1%, at most 0.01%, at most 0.001%, at most 0.0001%, etc. In some cases, the error rate is between 10% and 0.0001%, between 3% and 0.0001%, between 1% and 0.0001%, between 0.01% and 0.0001%, etc.
[0503] Removal of repeating sequences
[0504] Genomic DNA may contain repetitive sequences that are not of interest in some cases when performing nucleic acid sequencing reactions. This article provides methods for removing these repetitive sequences (e.g., by hybridization with sequences complementary to repetitive sequences such as Cot-1 DNA).
[0505] On one hand, a method for sequencing nucleic acid samples using nanopores adjacent to sensing electrodes in a membrane includes removing repetitive nucleic acid sequences from the nucleic acid sample to provide single-stranded nucleic acid molecules for sequencing. The method may further include providing a labeled nucleotide to a reaction chamber containing a nanopore, wherein a single labeled nucleotide contains a label coupled to the nucleotide, the label being detectable via the nanopore. In some cases, the method includes a polymerization reaction using a polymerase to incorporate a single labeled nucleotide into a growth chain complementary to the single-stranded nucleic acid molecule. The method may include detecting the label associated with the single labeled nucleotide via the nanopore during incorporation of the single labeled nucleotide, wherein the label is detected via the nanopore when the nucleotide associates with the polymerase.
[0506] In some cases, repetitive sequences are not physically removed from the reaction, but are rendered unsequential and remain in the reaction mixture (e.g., by hybridization with Cot-1 DNA, which causes the repetitive sequences to double-strand and effectively "remove" them from the sequencing reaction). In other cases, the repetitive sequences are made to double-strand.
[0507] Repetitive sequences can have any suitable length. In some cases, repetitive nucleic acid sequences contain approximately 20, 40, 60, 80, 100, 200, 400, 600, 800, 1000, 5000, 10000, or 50000 nucleic acid bases. In other cases, repetitive nucleic acid sequences contain at least approximately 20, 40, 60, 80, 100, 200, 400, 600, 800, 1000, 5000, 10000, or 50000 nucleic acid bases. In some cases, the bases are continuous.
[0508] The repetitive nucleic acid sequence may have any number of repeating subunits. In some cases, the repeating subunits are continuous. In some embodiments, the repetitive nucleic acid sequence contains about 20, about 40, about 60, about 80, about 100, about 200, about 400, about 600, about 800, or about 1000 repeating subunits of nucleic acid bases. In some cases, the repetitive nucleic acid sequence contains at least about 20, at least about 40, at least about 60, at least about 80, at least about 100, at least about 200, at least about 400, at least about 600, at least about 800, or at least about 1000 repeating subunits of nucleic acid bases.
[0509] In some cases, the repetitive nucleic acid sequence is removed by hybridization with a nucleic acid sequence complementary to the repetitive nucleic acid sequence. The nucleic acid sequence complementary to the repetitive nucleic acid sequence can be fixed onto a solid support such as a surface or beads. In some cases, the nucleic acid sequence complementary to the repetitive nucleic acid sequence contains Cot-1 DNA (which is an example of a repetitive nucleic acid sequence with a length of about 50 to about 100 nucleic acid bases).
[0510] Nanopore assembly and insertion
[0511] The methods described herein can utilize nanopores having polymerases attached to them. In some cases, it is desirable to have one and only one polymerase per nanopore (e.g., to enable sequencing of only one nucleic acid molecule at each nanopore). However, many nanopores, including α-hemolysin (αHL), can be multimeric proteins with multiple subunits (e.g., seven subunits for αHL). The subunits can be identical copies of the same polypeptide. This document provides multimeric proteins (e.g., nanopores) with a specified ratio of modified to unmodified subunits. Methods for producing multimeric proteins (e.g., nanopores) with a specified ratio of modified to unmodified subunits are also provided herein.
[0512] refer to Figure 27 A method for assembling a protein having multiple subunits includes providing a plurality of first subunits 2705 and a plurality of second subunits 2710, wherein the second subunits are modified relative to the first subunits. In some cases, the first subunits are wild-type (e.g., purified from a natural source or recombinantly produced). The second subunits can be modified by any suitable method. In some cases, the second subunit has an attached protein (e.g., a polymerase) (e.g., as a fusion protein). The modified subunits may contain a chemically active moiety (e.g., an azide or alkynyl group suitable for forming bonds). In some cases, the method further includes performing a reaction (e.g., click cycloaddition) to attach the entity (e.g., a polymerase) to the chemically active moiety.
[0513] The method may further include contacting the first subunit and the second subunit 2715 at a first ratio to form a plurality of proteins 2720 having the first subunit and the second subunit. For example, a partially modified αHL subunit having a reactive group suitable for attaching a polymerase may be mixed with six partially wild-type αHL subunits (i.e., at a first ratio of 1:6). The plurality of proteins may have multiple ratios of first subunits to second subunits. For example, the mixed subunits may form several nanopores with a stoichiometric distribution of modified and unmodified subunits (e.g., 1:6, 2:5, 3:4).
[0514] In some cases, proteins are formed simply by mixing subunits. In the case of αHL nanopores, for example, detergents (e.g., deoxycholic acid) can trigger αHL monomers to adopt a pore conformation. Nanopores can also be formed using lipids (e.g., 1,2-diphydanyl-sn-glycerol-3-phosphocholine (DPhPC) or 1,2-di-O-phydanyl- sn 3-glycerol-3-phosphocholine (DoPhPC) is formed with moderate temperatures (e.g., below about 100°C). In some cases, mixing DoPhPC with a buffer solution produces larger multilayered vesicles (LMVs), and the addition of the αHL subunit to this solution and incubation of the mixture at 40°C for 30 minutes can lead to pore formation.
[0515] If two different types of subunits are used (e.g., a natural wild-type protein and a second αHL monomer, which may contain a single point mutation), the resulting protein may have a mixed stoichiometry (e.g., a wild-type and mutant protein). The stoichiometry of these proteins can follow a formula that depends on the ratio of the concentrations of the two proteins used in the pore-forming reaction. This formula is as follows:
[0516] ,in
[0517] P m =Probability of a pore having m number of mutant subunits
[0518] n = total number of subunits (e.g., 7 for αHL)
[0519] m = the number of "mutated" subunits
[0520] f mut = The fraction or ratio of the mutant subunits mixed together
[0521] f wt = The fraction or ratio of wild-type subunits mixed together
[0522] The method may further include fractional separation of the plurality of proteins to enrich proteins 2725 having a second subunit to a second subunit ratio. For example, nanoporous proteins having one and only one modified subunit can be separated (e.g., a 1:6 second ratio). However, any second ratio is suitable. The distribution of the second ratio can also be fractionally separated, such as enriching proteins having one or two modified subunits. The total number of subunits forming the protein is not always 7 (e.g., different nanopores can be used or α-hemolysin nanopores with six subunits can be formed), such as... Figure 27 The description is as follows. In some cases, proteins with only one modified subunit are enriched. In such cases, the second ratio is 1 second subunit for every (n-1) first subunits, where n is the number of subunits constituting the protein.
[0523] The first ratio can be the same as the second ratio; however, this is not necessary. In some cases, proteins with mutant monomers are less efficiently formed than proteins without mutant subunits. If this is the case, then the first ratio can be greater than the second ratio (e.g., if a second ratio of 1 mutant subunit to 6 non-mutant subunits is desired in the nanopore, then forming a suitable 1:6 protein may require mixing the subunits at a ratio greater than 1:6).
[0524] Proteins with different second-component ratios can exhibit different behaviors during separation (e.g., different retention times). In some cases, chromatography, such as ion-exchange chromatography or affinity chromatography, is used to fractionate proteins. Because the first and second subunits can be identical (except for modifications), the number of modifications on the protein can serve as a basis for separation. In some cases, the first or second subunit has a purification label (e.g., in addition to modifications) to allow or improve the efficiency of fractionation. In some cases, multihistidine labels (His-labels), streptavidin labels (Strep-labels), or other peptide labels are used. In some cases, the first and second subunits each contain different labels, and the fractionation step is performed based on each label. In the case of His-labels, a charge is generated on the label at a low pH (histidine residues become positively charged, below the pKa of the side chain). Utilizing the significant difference in charge of one αHL molecule compared to others, ion-exchange chromatography can be used to separate oligomers with 0, 1, 2, 3, 4, 5, 6, or 7 "charged-labeled" αHL subunits. In principle, this charged label can be any string of amino acids carrying a uniform charge. Figure 28 and Figure 29 An example of hierarchical separation of nanopores based on His-labeling is shown. Figure 28 The graph shows the UV absorbance at 280 nm, UV absorbance at 260 nm, and conductivity. The peaks correspond to nanopores with different ratios of modified and unmodified subunits. Figure 29 The hierarchical separation of αHL nanopores and their mutants using His- and Strep-labels is shown.
[0525] In some cases, an entity (e.g., a polymerase) is attached to the protein after fractionation 2730. The protein can be a nanopore and the entity can be a polymerase. In some cases, the method also includes inserting a protein with a second ratio subunit into the bilayer.
[0526] In some cases, nanopores may contain multiple subunits. A polymerase may attach to one of the subunits, and at least one but fewer than all subunits may contain a first purification label. In some instances, the nanopore is α-hemolysin or a variant thereof. In some cases, all said subunits may contain a first purification label or a second purification label. The first purification label may be a multihistidine label (e.g., on the subunit with the attached polymerase).
[0527] connector
[0528] The methods described herein can be used for nanopore detection, including nucleic acid sequencing, using enzymes (e.g., polymerases) attached to nanopores. In some cases, the connection between the enzyme and the nanopore can affect the performance of the system. For example, engineering the attachment of DNA polymerase to a pore (α-hemolysin) can increase the effective concentration of labeled nucleotides, thereby reducing the entropy barrier. In some cases, the polymerase is directly attached to the nanopore. In other cases, a linker is used between the polymerase and the nanopore.
[0529] The labeled sequencing described herein may benefit from the efficient capture of specific labeled nucleotides into αHL pores via potential-induced specificity. Capture can occur during or after DNA template-based polymerase primer extension. One approach to improve capture efficiency is to optimize the ligation between the polymerase and the αHL pore. Without limitation, three characteristics of the ligation to be optimized are: (a) ligation length (which can increase the effective concentration of labeled nucleotides, influence capture kinetics, and / or alter entropy barriers); (b) ligation flexibility (which can influence the kinetics of conjugate conformational changes); and (c) the number and location of ligations between the polymerase and the nanopore (which can reduce the number of available conformational states, thereby increasing the likelihood of suitable pore-polymerase orientation, increasing the effective concentration of labeled nucleotides, and reducing entropy barriers).
[0530] Enzymes and polymerases can be linked in any suitable manner. In some cases, open reading frames (ORFs) are fused directly or with amino acid linkers. Fusion can occur in any order. In some cases, chemical bonds are formed (e.g., via click chemistry). In others, the linkage is non-covalent (e.g., molecular staples, via biotin-streptavidin interactions, or via protein-protein tags such as PDZ, GBD, SpyTag, halogenated tags, or SH3 ligands).
[0531] In some cases, the linker is a polymer such as peptides, nucleic acids, or polyethylene glycol (PEG). The linker can be of any suitable length. For example, the linker can be about 5 nanometers (nm), about 10 nm, about 15 nm, about 20 nm, about 40 nm, about 50 nm, or about 100 nm long. In some cases, the linker is at least about 5 nanometers (nm), at least about 10 nm, at least about 15 nm, at least about 20 nm, at least about 40 nm, at least about 50 nm, or at least about 100 nm long. In some cases, the linker is at most about 5 nanometers (nm), at most about 10 nm, at most about 15 nm, at most about 20 nm, at most about 40 nm, at most about 50 nm, or at most about 100 nm long. The linker can be rigid, flexible, or any combination thereof. In some cases, no linker is used (e.g., polymerase is directly attached to the nanopore).
[0532] In some cases, more than one linker connects the enzyme to the nanopore. The number and location of the links between the polymerase and the nanopore can vary. Examples include: αHL C-terminus to polymerase N-terminus; αHL N-terminus to polymerase C-terminus; and links between amino acids not at the terminus.
[0533] On one hand, a method for sequencing nucleic acid samples using nanopores adjacent to sensing electrodes in a membrane includes providing a labeled nucleotide to a reaction chamber containing a nanopore, wherein a single labeled nucleotide of the labeled nucleotide contains a label coupled to the nucleotide, the label being detectable by means of the nanopore. The method may include incorporating a single labeled nucleotide of the labeled nucleotide into a growth chain complementary to a single-stranded nucleic acid molecule from the nucleic acid sample via a polymerization reaction using a polymerase attached to the nanopore via a linker. The method may also include detecting the label associated with the single labeled nucleotide during incorporation using the nanopore, wherein the label is detected by means of the nanopore when the nucleotide associates with the polymerase.
[0534] In some cases, the polymerase is directed relative to the nanopore via a linker to enable the detection of the label by means of the nanopore. In other cases, the polymerase is attached to the nanopore via two or more linkers.
[0535] In some cases, the adapter contains one or more of SEQ ID NO 2-35 or PCR products generated therefrom. In other cases, the adapter contains a peptide encoded by one or more of SEQ ID NO 2-35 or PCR products generated therefrom.
[0536]
[0537]
[0538]
[0539]
[0540] Calibration of applied voltage
[0541] Molecular-specific output signals from single-molecule nanopore sensor devices can originate from the presence of an electrochemical potential difference across an ion-impermeable membrane surrounded by an electrolyte solution. This transmembrane potential difference determines the intensity of the nanopore-specific electrochemical current, which can be detected by electronics within the device via sacrificial (i.e., Faraday) or non-sacrificial (i.e., capacitive) reactions occurring at the electrode surface.
[0542] For any given nanopore state (i.e., open channel, trapped state, etc.), a time-dependent transmembrane potential can function as an input signal, which determines the resulting current flowing through the nanopore complex as a function of time. This nanopore current can provide a specific molecular signal output through a nanopore sensor device. The open-channel nanopore current can be modulated to varying degrees by the interaction between the nanopore, which partially blocks the flow of ions through the channel, and the trapped molecules.
[0543] These modulations can exhibit type-specificity of the captured molecules, allowing some molecules to be directly identified from their nanopore current modulation. For a given molecule type and fixed set of device conditions, the degree of modulation of the open-channel nanopore current by the captured molecule type can vary depending on the applied transmembrane potential, mapping each molecule type to a specific current-to-voltage (IV) curve.
[0544] A systemically variable offset between the applied voltage setting and the transmembrane potential can introduce a horizontal shift in the IV curve along the horizontal voltage axis, potentially reducing the accuracy of molecule identification based on the measured current signal reported as an output signal by the nanopore sensor device. Therefore, an uncontrolled offset between the applied potential and the transmembrane potential can be problematic for accurately comparing measurements of the same molecule under identical conditions.
[0545] This so-called "potential shift" between the externally applied potential and the actual transmembrane potential can vary both within and between experiments. Changes in potential shift can be caused by variations in initial conditions and time-dependent changes (drifts) in electrochemical conditions within the nanopore sensor device.
[0546] These measurement errors can be eliminated by calibrating the time-dependent shift between the applied voltage and the transmembrane potential for each experiment, as described herein. Physically, the probability of observing escape events of molecules trapped in nanopores can depend on the applied transmembrane potential, and this probability distribution can be the same for the same molecular sample under identical conditions (e.g., the sample can be a mixture of different types of molecules, provided their proportions do not vary between samples). In some cases, the voltage distribution at which escape events occur for a fixed sample type provides a measurement of the shift between the applied voltage and the transmembrane potential. This information can be used to calibrate the applied voltage across the nanopore, eliminating systematic sources of error caused by intra- and inter-experimental potential shifts and improving the accuracy of molecule recognition and other measurements.
[0547] For a given nanopore sensor device operated with the same molecular sample and reagents, the expected value of the escape voltage distribution can be estimated by statistically analyzing a sample of single-molecule escape events (although each individual event may be a stochastic process subject to random fluctuations). This estimation can be time-dependent to account for the time drift of potential shifts within the experiment. This can correct for the variable difference between the applied voltage setting and the actual voltage sensed at the pore, effectively “comparing” all measurements in the horizontal direction when plotted in IV space.
[0548] In some cases, potential (i.e., voltage) offset calibration does not account for changes in current gain and current offset, and it can also be used to calibrate for improved accuracy and reproducibility of nanopore current measurements. However, potential offset calibration is generally performed before gain and offset correction to prevent errors in estimating changes in current gain and current offset, as these, in turn, involve fitting current-to-voltage (IV) curves, and the results of these fittings are affected by changes in voltage offset. That is, shifting data left to right (horizontally) in IV space can introduce errors in subsequent current gain and current offset fitting.
[0549] Figure 30 The graph shows the current (solid line) through the nanopore versus the applied voltage (dashed line) over time. The current decreases as the molecule is trapped in the nanopore (3005). As the applied voltage decreases over time (3010), the current decreases until the molecule detaches from the nanopore (3015), at which point the current increases to the desired level at the applied voltage. The applied voltage at molecule detachment can depend on the molecule's length. For example, a label with 30 bases may detach at approximately 40 mV, while a label with 50 bases may detach at approximately 10 mV. The detachment voltage can vary for different nanopores or for different measurements over time on the same nanopore (3020). Adjusting this detachment voltage to a desired value makes the data easier to interpret and / or more accurate.
[0550] On one hand, this paper provides a method for sequencing nucleic acid samples using nanopores adjacent to sensing electrodes in a membrane. The method may include providing a labeled nucleotide to a reaction chamber containing a nanopore, wherein a single labeled nucleotide contains a label coupled to the nucleotide, the label being detectable by means of the nanopore. The method may include a polymerization reaction using a polymerase to incorporate a single labeled nucleotide into a growth chain complementary to a single-stranded nucleic acid molecule from the nucleic acid sample. The method may then include detecting the label associated with the single labeled nucleotide during incorporation using the nanopore, wherein the label is detected by means of the nanopore when the nucleotide associates with the polymerase. In some cases, the detection includes applying an applied voltage across the nanopore and measuring a current using a sensing electrode at the applied voltage.
[0551] In some cases, the applied voltage is calibrated. Calibration may include estimating the desired escape voltage distribution against time for the sensing electrodes. The calibration may then calculate the difference between the desired escape voltage distribution and a reference point (e.g., any reference point, such as zero). The calibration may then change the applied voltage to alter the calculated difference. In some cases, the applied voltage decreases over time.
[0552] In some cases, the distribution of the desired escape voltage over time is estimated. In some cases, the reference point is zero volts. The method can remove detected variations in the distribution of the desired escape voltage. In some cases, the method is performed on multiple independently addressable nanopores, each adjacent to a sensing electrode.
[0553] In some embodiments, the presence of the label in the nanopore reduces the current measured by the sensing electrode under an applied voltage. In some cases, the labeled nucleotide comprises multiple different labels, and the method detects each of the multiple different labels.
[0554] In some cases, calibration increases the accuracy of a method when compared to an uncalibrated method. In some cases, calibration compensates for changes in electrochemical conditions over time. In some cases, calibration compensates for different nanopores with different electrochemical conditions in devices with multiple nanopores. In some embodiments, calibration compensates for different electrochemical conditions for each performance characteristic of the method. In some cases, the method also includes calibrating for changes in current gain and / or current offset.
[0555] Extended polymerase chain sequencing method
[0556] This disclosure provides a method for sequencing nucleic acid molecules using expanded polymer sequencing. Expanded polymer sequencing involves several steps to generate an expanding polymer that is longer than the nucleic acid to be sequenced and has a sequence derived from the nucleic acid molecule to be sequenced. The expanding polymer can be passed through a nanopore to determine its sequence. As described herein, the expanding polymer may have gates thereon so that the expanding polymer can pass through the nanopore in only one direction. The steps of this method are described in... Figures 33 to 36 The description is provided below. Further information regarding nucleic acid sequencing via expansion (i.e., expanded aggregates) can be found in U.S. Patent 8,324,360, which is incorporated herein by reference in its entirety.
[0557] refer to Figure 33 On one hand, methods for nucleic acid sequencing include providing a single-stranded nucleic acid to be sequenced and providing multiple probes. The probes comprise a hybridization moiety 3305 capable of hybridizing with the single-stranded nucleic acid, a loop structure 3310 having two ends (each end attached to the hybridization moiety), and a cleavable group 3315 located between the ends of the loop structure. The loop structure includes a gate 3320 that prevents the loop structure from passing through the nanopore in the opposite direction.
[0558] refer to Figure 34 The method may include polymerizing multiple probes 3405 in a sequence determined by hybridization with the hybridization moiety 3410 of the single-stranded nucleic acid to be sequenced. (Reference) Figure 35 The method may include cleaving the 3505 cleavable group to provide an extended strand for sequencing.
[0559] refer to Figure 36 The method may include passing the extended strand 3610 through a 3605 nanopore 3615, wherein the gate prevents the extended strand from passing through the nanopore in the opposite direction 3620. The method may also include sequencing the single-stranded nucleic acid to be sequenced by means of the nanopore to detect the loop structure of the extended strand in a sequence determined by hybridization with the hybridization moiety.
[0560] In some cases, the ring structure comprises a narrow segment and the gate is a polymer with two ends, wherein a first end is attached to the ring structure adjacent to the narrow segment and a second end is not attached to the ring structure. The ring structure is capable of passing through the nanopore in a first direction aligned with the gate adjacent to the narrow segment. In some embodiments, the ring structure is not capable of passing through the nanopore in the opposite direction, where the gate is not aligned with the narrow segment.
[0561] In some cases, the gate contains a nucleotide. When the gate is not aligned with adjacent narrow segments, it can pair with ring bases.
[0562] The narrow segment can contain any polymer or molecule that is narrow enough to allow the polymer to pass through the nanopore when the gate is aligned with the narrow segment. In some cases, the narrow segment contains debased nucleotides (i.e., nucleic acid side chains without attached nucleotide bases) or carbon chains.
[0563] In some cases, the electrodes are recharged between detection periods. When the electrodes are recharged, the extended chains generally do not pass through the nanopores in the reverse direction.
[0564] Non-sequencing methods and applications
[0565] The devices and methods described herein can be used to measure certain properties of nucleic acid samples and / or nucleic acid molecules other than their nucleic acid sequences (e.g., any measurement of sequence length or the quality of a nucleic acid sample, including but not limited to the degree of cross-linking of nucleic acids in the sample). In some cases, it may be desirable to have an uncertain sequence of nucleic acid molecules. For example, a human individual (or other organisms such as horses) can be identified by determining the length of certain repetitive sequences in the genome (e.g., referred to as microsatellites, simple sequence repeats (SSRs), or short tandem repeats (STRs)). One may wish to know the length of one or more STRs (e.g., to identify the origin of a crime or the perpetrator) without knowing the STR sequence and / or the DNA sequence found before (5') or after (3') the STR (e.g., to identify a person's race, the likelihood of exposure to a disease, etc.).
[0566] On the one hand, the method identifies one or more STRs present in the genome. Any number of STRs can be identified (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30 or more). STRs may contain repeating segments (e.g., 'AGGTCT' of the sequence SEQ. ID. No. 1-AGGTCT AGGTCTAGGTCT AGGTCT AGGTCT AGGTCT AGGTCT AGGTCT), which have any number of nucleic acid bases (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30 or more bases). STR can contain any number of repeating segments, which are generally repeated consecutively (e.g., repeated 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30 or more times).
[0567] The number of nucleotide incorporation events and / or the length of nucleic acids or their segments can be determined by using nucleotides with the same label attached to some, most, or all of the labeled nucleotides. Detection of the label (pre-loaded into the nanopore before release or guided into the nanopore after release from the labeled nucleotide) indicates that a nucleotide incorporation event has occurred, but in this case, it does not identify which nucleotide has been incorporated (e.g., no sequence information is determined).
[0568] In some embodiments, all nucleotides (e.g., all adenine (A), cytosine (C), guanine (G), thymine (T), and / or uracil (U) nucleotides) have the same tag coupled to the nucleotide. In some cases, however, this may not be necessary. At least some of the nucleotides may have a tag that identifies the nucleotide (e.g., to allow some sequence information to be determined). In some cases, about 5%, about 10%, about 15%, about 20%, about 25%, about 30%, about 40%, or about 50% of the nucleotides have a tag that identifies the nucleotide (e.g., to allow some nucleic acid positions to be sequenced). The sequenced nucleic acid positions may be randomly distributed along the nucleic acid strand. In some cases, all nucleotides of a single type have a tag that identifies the nucleotide (e.g., to allow all adenine to be sequenced). In some cases, at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 40%, or at least 50% of the nucleotides have a tag that identifies the nucleotide. In some implementations, up to 5%, up to 10%, up to 15%, up to 20%, up to 25%, up to 30%, up to 40%, or up to 50% of the nucleotides have markers for recognizing the nucleotides. In some implementations, all nucleic acids or segments thereof are short tandem repeat sequences (STRs).
[0569] On one hand, a method for determining the length of a nucleic acid or a segment thereof using a nanopore in a membrane adjacent to a sensing electrode includes providing a labeled nucleotide to a reaction chamber containing a nanopore. The nucleotide may have different bases containing the same label coupled to the nucleotide, such as at least two different bases, the label being detectable by means of the nanopore. The method may further include a polymerization reaction using a polymerase to incorporate a single labeled nucleotide into a growth chain complementary to a single-stranded nucleic acid molecule from a nucleic acid sample. The method may also include detecting the label associated with the single labeled nucleotide during or after incorporation using a nanopore.
[0570] On one hand, a method for determining the length of a nucleic acid or its segments using nanopores adjacent to sensing electrodes in a membrane includes providing a labeled nucleotide to a reaction chamber containing the nanopore. A single labeled nucleotide may contain a label coupled to the nucleotide, which reduces the current flowing through the nanopore compared to the current when the label is absent.
[0571] In some embodiments, the method further includes a polymerization reaction using a polymerase to incorporate a single labeled nucleotide into a growth chain complementary to a single-stranded nucleic acid molecule from a nucleic acid sample and to reduce the magnitude of the current flowing through the nanopore. The magnitude of the current can be reduced by any suitable amount, including about 5%, about 10%, about 15%, about 20%, about 25%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, about 90%, about 95%, or about 99%. In some embodiments, the magnitude of the current is reduced by at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, or at least 99%. In some implementations, the current is reduced by up to 5%, up to 10%, up to 15%, up to 20%, up to 25%, up to 30%, up to 40%, up to 50%, up to 60%, up to 70%, up to 80%, up to 90%, up to 95%, or up to 99%.
[0572] The method may also include detecting time intervals between the incorporation of a single labeled nucleotide using nanopores (e.g., Figure 6 The time period between the incorporation of a single labeled nucleotide can have a high current magnitude. In some embodiments, the current magnitude flowing through the nanopore between nucleotide incorporation events is (e.g., back) about 50%, about 60%, about 70%, about 80%, about 90%, about 95%, or about 99% of the maximum current (e.g., when no label is present). In some embodiments, the current magnitude flowing through the nanopore between nucleotide incorporation events is at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, or at least 99% of the maximum current.
[0573] In some cases, partial sequencing of nucleic acids is performed before (5') or after (3') the STR to identify which STR has its length determined in the nanopore (e.g., in the case of multiplexing where multiple primers point to multiple STRs). In other cases, approximately 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15 or more nucleic acids are sequenced before (5') or after (3') the STR.
[0574] Pattern matching
[0575] This disclosure also provides an electronic reader for matching signal patterns detected by a nanopore device with known (or reference) signals. The nanopore device may include nanopores in a membrane, as described elsewhere herein. The known signal may be stored in a memory location, such as a remote database or memory location located on a chip containing the nanopore device. The electronic reader may match patterns using a pattern matching algorithm, which may be implemented using the computer processor of the electronic reader. The electronic reader may be located on a chip.
[0576] Pattern matching can be achieved in real time, such as while data is being collected through a nanopore device. Alternatively, pattern matching can be achieved by first collecting data and then processing it to match patterns.
[0577] In some cases, the reader contains a list of one or more nucleic acid sequences of interest to the user (also referred to as the "white list" in this document), and a list (or lists of) one or more other nucleic acid sequences of no interest to the user (also referred to as the "black list" in this document). During nucleic acid detection (including nucleic acid incorporation events), the reader may detect and record nucleic acid sequences in the white list, and may not detect or record nucleic acid sequences in the black list. Example
[0578] Example 1 - Illegal Radical Conduction
[0579] Figure 37 This illustrates how non-Radidatic conduction can decouple the nanopore from regulation. The vertical axis of this graph is the current measured in picoamperes (pA), ranging from -30 to 30. The horizontal axis is the time measured in seconds, ranging from 0 to 2. The waveform has a 40% duty cycle. Data point 3705 is the current measured above and below the bilayer using a sponge platinum working electrode in the presence of 150 mM KCl, pH 7.5, and 20 mM HEPES buffer, and 3 mM SrCl2. A 240 nM viscous polymerase and a 5GS sandwich structure with a 0.0464 OD are present. The lipids are 75% phosphatidylethanolamine (PE) and 25% phosphatidylcholine (PC). The simulated voltage 3710 across the working electrode and counter electrode (AgCl pellet) is multiplied by 100 and fitted to the graph. The simulated electrochemical potential across the nanopore-polymerase complex 3715 is multiplied by 100 and fitted to the graph. Simulate current 3720 using an integrated circuit enhancement (SPICE) model and a simulation program.
[0580] Example 2 - Tag Capture
[0581] Figure 38 Two markers captured in a nanopore within an alternating current (AC) system are shown. The vertical axis of the figure represents current measured in picoamperes (pA), ranging from 0 to 25. The horizontal axis represents time measured in seconds, ranging from approximately 769 to 780. The first marker, 3805, was captured at approximately 10 pA. The second marker, 3810, was captured at approximately 5 pA. The open-channel current, 3815, was approximately 18 pA. Due to the rapid capture of the markers, few data points were observed at the open-channel current. The waveform ranged from 0 to 150 mV at 10 Hz and a 40% duty cycle. The solution contained 150 mM KCl.
[0582] Example 3 - Label Sequencing
[0583] Figure 39 An example of a ternary complex 3900 is shown, formed between a fusion of a template DNA molecule 3905 to be sequenced, a hemolysin nanopore 3910, and a DNA polymerase 3915, and a labeled nucleotide 3920. The polymerase 3915 is attached to the nanopore 3910, which has a protein linker 3925. The nanopore / polymerase construct is formed such that only one of the seven polypeptide monomers of the nanopore has an attached polymerase. The labeled nucleotide portion penetrates the 3920 nanopore and affects the current flowing through the nanopore.
[0584] Figure 40 The diagram illustrates the current flowing through a nanopore in the presence of template DNA to be sequenced but without labeled nucleotides. The solution in contact with the nanopore contains 150 mM KCl, 0.7 mM SrCl2, 3 mM MgCl2, and 20 mM HEPES buffer (pH 7.5) at an applied voltage of 100 mV. With a few exceptions, the current is maintained at approximately 18 picoamperes (pA). Exceptions can be electronic noise and can be a single data point on the transverse time axis. Algorithms that distinguish noise from signal, such as adaptive signal processing algorithms, can be used to mitigate electronic noise.
[0585] Figure 41 , Figure 42 and Figure 43 Different labels provide different current levels. In all examples, the solutions in contact with the nanopores were in 150 mM KCl, 0.7 mM SrCl2, 3 mM MgCl2, and 20 mM HEPES buffer (pH 7.5) at an applied voltage of 100 mV. Figure 41 Guanine (G) 4105 is shown to be distinct from thymine (T) 4110. It is labeled as dT6P-T6-dSp8-T16-C3 (for T) with a current level of about 8 to 10 pA and dG6P-Cy3-30T-C6 (for G) with a current level of about 4 or 5 pA. Figure 42 Guanine (G) 4205 is shown to be distinct from adenine (A) 4210. It is labeled as dA6P-T4-(Sp18)-T22-C3 (for A) with a current level of about 6 to 7 pA and dG6P-Cy3-30T-C6 (for G) with a current level of about 4 or 5 pA. Figure 43 Guanine (G) 4305 is shown to be distinct from cytosine (C) 4310. It is labeled as dC6P-T4-(Sp18)-T22-C3 (for C) with a current level of about 1 to 3 pA and dG6P-Cy3-30T-C6 (for G) with a current level of about 4 or 5 pA.
[0586] Figure 44 , Figure 45 , Figure 46 and Figure 47 Examples of sequencing using labeled nucleotides are shown. The DNA molecule to be sequenced is single-stranded and has the sequence AGTCAGTC (SEQ. ID. No: 36) and is stabilized by two flanking hairpin structures. In all examples, the solution in contact with the nanopore contains 150 mM KCl, 0.7 mM SrCl2, 3 mM MgCl2, and 20 mM HEPES buffer (pH 7.5) at an applied voltage of 100 mV. Four labels corresponding to guanine (dG6P-Cy3-30T-C6), adenine (dA6P-T4-(Sp18)-T22-C3), cytosine (dC6P-T4-(Sp18)-T22-C3), and thymine (dT6P-T6-dSp8-T16-C3) are included in the solution. Figure 44 An example is shown in which four consecutive labeled nucleotides (i.e., C4405, A4410, G4415, and T4420) corresponding to the sequence GTCA in SEQ ID No: 36 are identified. The labels can be introduced and passed through the nanopore several times before incorporation into the growth chain (e.g., so that for each incorporation event, the current level can switch several times between an open channel current and a reduced current level that distinguishes the labels).
[0587] For any reason, including but not limited to different numbers of labels entering and leaving the nanopore and / or labels being temporarily held by the polymerase but not fully incorporated into the growing nucleic acid chains, the duration of current reduction between experiments may vary. In some embodiments, the duration of current reduction between experiments is approximately constant (e.g., varying by no more than about 200%, 100%, 50%, or 20%). In some cases, the selection of the enzyme, the applied voltage waveform, the concentration of divalent and / or monovalent ions, the temperature, and / or pH make the duration of current reduction between experiments approximately constant. Figure 45 The identical sequence GTCA in SEQ ID No: 36 is shown as follows. Figure 44 In this case, the identified marker nucleotides are C4505, A4510, G4515, and T4520. In some cases, the current remains reduced for an extended period of time (e.g., about 2 seconds, as shown in 4520).
[0588] Figure 46 The five consecutive labeled nucleotides identified (i.e., T4605, C4610, A4615, G4620, T4625) correspond to the sequence AGTCA in SEQ ID No: 36. Figure 47 The five consecutive labeled nucleotides identified (i.e., T4705, C4710, A4715, G4720, T4725, C4730) correspond to the sequence AGTCAG in SEQ ID No: 36.
[0589] Although preferred embodiments of the invention have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. The invention is not intended to be limited by the specific embodiments provided in the specification. Although the invention has been described with reference to the foregoing description, the description and illustration of embodiments herein are not intended to be construed as limiting. Many variations, alterations, and substitutions will now occur to those skilled in the art without departing from the invention. Furthermore, it should be understood that all aspects of the invention are not limited to the specific descriptions, configurations, or relative proportions set forth herein, as this depends on a variety of conditions and variables. It should be understood that various alternatives to the embodiments of the invention described herein may be employed in the practice of the invention. Therefore, it is contemplated that the invention should also cover any such alternatives, modifications, variations, or equivalents. The following claims are intended to define the scope of the invention and thereby cover the methods and structures within the scope of these claims and their equivalents.
Claims
1. A method for assembling nanopores having multiple subunits, wherein the method comprises: (a) Provide multiple first subunits; (b) Providing a plurality of second subunits, wherein the second subunits are modified relative to the first subunit; (c) Contacting the first subunit and the second subunit at a first ratio to form a plurality of nanopores having the first subunit and the second subunit, wherein the plurality of nanopores have a plurality of ratios of the first subunit to the second subunit; as well as (d) The plurality of nanopores are fractionally separated by ion exchange chromatography to enrich nanopores having a second ratio of first subunits to second subunits, wherein the second ratio is 1 second subunit for every 6 first subunits; The first subunit is wild-type or recombinant α-hemolysin; The second subunit is recombinant α-hemolysin; The second subunit contains a purification label, and the method further includes (e) performing a reaction to attach the polymerase to the purification label using non-covalent interactions.
2. A method for assembling nanopores having multiple subunits, the method comprising: (a) Provide multiple first subunits; (b) Providing a plurality of second subunits, wherein the second subunits are modified relative to the first subunit; (c) Contacting the first subunit and the second subunit at a first ratio to form a plurality of nanopores having the first subunit and the second subunit, wherein the plurality of nanopores have a plurality of ratios of the first subunit to the second subunit; as well as (d) The plurality of nanopores are fractionated by ion exchange chromatography to enrich nanopores having a second ratio of first subunits to second subunits, wherein the second ratio is 2 second subunits for every 5 first subunits and a single polymerase is attached to each second subunit. The first subunit is wild-type or recombinant α-hemolysin; The second subunit is recombinant α-hemolysin; The second subunit contains a purification label, and the method further includes (e) performing a reaction to attach the polymerase to the purification label using non-covalent interactions.
3. The method according to any one of claims 1-2, wherein the nanopores are at least 80% homologous to α-hemolysin.
4. The method of any one of claims 1-2, wherein the first subunit comprises a multihistidine tag.
5. The method of any one of claims 1-2, wherein the first subunit is wild-type α-hemolysin.
6. The method of any one of claims 1-2, wherein the first subunit and the second subunit are recombinant α-hemolysin.
7. The method of any one of claims 1-2, wherein the first ratio is equal to the second ratio.
8. The method of any one of claims 1-2, wherein the first ratio is greater than the second ratio.
9. The method of any one of claims 1-2, further comprising inserting a nanopore having subunits having the second ratio into a lipid bilayer.
10. A method for sequencing nucleic acid molecules using a nanopore assembled by means of any one of claims 1-2, comprising sequencing nucleic acid molecules using a nanopore having the second ratio subunit.
11. A nanopore comprising a plurality of first subunits and second subunits, wherein the ratio between the first subunits and the second subunits is one second subunit for every six first subunits or two second subunits for every five first subunits; The first subunit is wild-type or recombinant α-hemolysin; wherein the second subunit is recombinant α-hemolysin; The second subunit contains a purification label; and the polymerase is non-covalently attached to the purification label.
12. The nanopore of claim 11, wherein the nanopore is at least 80% homologous to α-hemolysin.
13. The nanopore of claim 11, wherein the first subunit comprises a multihistidine label.
Citation Information
Patent Citations
Polymerases for nucleotide analogue incorporation
US20110059505A1
Systems and methods for characterizing a molecule
US20110193570A1
Recombinant Polymerases For Improved Single Molecule Sequencing
US20120034602A1
Massive parallel method for decoding DNA and RNA
US6664079B2
Two slow-step polymerase enzyme systems and methods
US8133672B2