Single-molecule sequencing of plasma DNA
Concatemerization and enrichment techniques improve nanopore sequencing efficiency and accuracy for small DNA fragments by creating longer molecules and enhancing alignment to reference genomes.
Patent Information
- Application Number
- JP2024092950
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2015-08-12
- Filing Date
- 2024-06-07
- Publication Date
- 2026-02-03
- Estimated Expiration
- 2036-08-12
AI Technical Summary
Current nanopore sequencing technologies are inefficient for analyzing small DNA fragments found in samples like plasma due to low concentration and short fragment lengths, leading to inaccurate and inefficient sequencing results.
Concatemerization of DNA fragments to create longer molecules for sequencing, combined with enrichment techniques to increase fragment concentration, and bioinformatics methods to identify original fragments within concatemers.
Enhances the sequencing efficiency and accuracy of small DNA fragments by increasing the likelihood of fragments being processed through nanopores and improving alignment to reference genomes.
Smart Images

Figure 0007810455000001 
Figure 0007810455000002 
Figure 0007810455000003
Abstract
Description
[Technical Field]
[0001] Cross-incorporation of related applications This application claims the benefit of U.S. Provisional Patent Application No. 62 / 204,396, filed August 12, 2015, the contents of which are incorporated herein by reference for all purposes. [Background technology]
[0002] Noninvasive prenatal testing (NIPT) using maternal plasma DNA sequencing is now clinically available for screening for fetal chromosomal aneuploidies (1). Unlike amniocentesis, maternal plasma DNA sequencing does not pose any risk of miscarriage. These tests have a sensitivity of nearly 99% and a specificity of nearly 99% (1). As a result, clinical demand for NIPT has increased significantly since it was first commercially available in 2011.
[0003] Massively parallel sequencing is a central component of most experimental protocols currently used for NIPT of chromosomal aneuploidies (2). Due to high instrumentation costs, these tests are currently performed in reference laboratories. Oxford Nanopore Technologies has developed a nanopore-based DNA sequencing platform (3). Nanopore sequencers have relatively low capital costs and a small footprint. Each flow cell costs between $500 and $900 and can be used multiple times for up to 48 hours. Sequencing speed is also relatively fast, with each nanopore capable of reading 30 bases per second. These features would be advantageous for use in clinical laboratories. However, current nanopore technology is inefficient for samples typically used for NIPT (e.g., plasma). Summary of the Invention
[0004] overview Embodiments improve the efficiency of single-molecule sequencing techniques when analyzing samples with relatively small DNA fragments. For example, the concentration of DNA fragments in a sample can be significantly increased, allowing more DNA fragments to interact with a sequencing device (e.g., a nanopore). As another example, DNA fragments can be bound to concatemers to read longer molecules, thereby efficiently reading multiple DNA fragments (i.e., DNA fragments originally present in a sample) via the sequence of a single molecule. Bioinformatics techniques can be used to detect different DNA samples that are part of the same concatemer. Embodiments can combine two techniques.
[0005] Embodiments may include a method of determining a nucleic acid sequence. The method may include receiving a plurality of DNA fragments. The method may also include concatemerizing a first set of DNA fragments to obtain first concatemers. The method may include performing single-molecule sequencing of the first concatemers to obtain a first sequence of the first concatemers. In some embodiments, the single-molecule sequencing is performed using a nanopore, and the method may include passing the first concatemers through the first nanopore. A first electrical signal may be detected as the first concatemers pass through the first nanopore. The first electrical signal may correspond to the first sequence of the first concatemers.
[0006] Another embodiment may include a method for determining a nucleic acid sequence. The method may include receiving a plurality of DNA fragments. A first set of the DNA fragments may be concatemerized to obtain a first concatemer. Fluorescently labeled nucleotides may be hybridized to the concatemers. A first fluorescent signal may be detected, the first fluorescent signal corresponding to the specific nucleotide. The fluorescent label may then be cleaved, another fluorescently labeled nucleotide may be added, and the process may be repeated.
[0007] Embodiments may include a method implemented by a computer system. The method may include receiving a first sequence of first concatemers generated by concatemerizing a first set of DNA fragments. The method may also include aligning subsequences of the first sequence to identify fragment sequences corresponding to each DNA fragment of the first set of DNA fragments.
[0008] Some embodiments may include a method for sequencing cell-free DNA fragments. The cell-free DNA fragments may include plasma DNA fragments. The method may include receiving a biological sample containing a plurality of DNA fragments. The biological sample may have a first concentration of DNA fragments. The method may also include enriching the corresponding biological sample to have a second concentration of DNA fragments. The second concentration of DNA fragments may be five times or more greater than the first concentration of DNA fragments. The method may further include passing the plurality of DNA fragments through a nanopore on a substrate. For each of the plurality of DNA fragments, an electrical signal may be detected as the DNA fragment passes through the nanopore. The electrical signal may correspond to a sequence of the DNA fragment.
[0009] Embodiments may include a method for cell-free DNA fragment sequencing. The method may include enriching a corresponding biological sample to have a second concentration of DNA fragments that is at least five times higher than the initial concentration of DNA fragments. The method may further include single molecule sequencing techniques. The DNA fragments may be hybridized with fluorescently labeled nucleotides. The method may also include detecting a signal from the fluorescently labeled nucleotides corresponding to the nucleotides. The fluorescently labeled nucleotides may be cleaved, and the process may repeat to identify additional nucleotides and sequence the DNA fragments.
[0010] Embodiments may also include computer products including a computer-readable medium having instructions stored thereon for performing the operations of any of the methods of DNA sequencing described herein. Some embodiments include one or more processors and computer products for executing the instructions stored on the computer-readable medium. Additional embodiments include systems for performing any of the methods.
[0011] Other embodiments are directed to systems, portable consumer devices, and computer-readable media related to the methods described herein.
[0012] A better understanding of the nature and advantages of embodiments of the present invention may be obtained by reference to the following detailed description and accompanying drawings. [Brief explanation of the drawings]
[0013] [Figure 1A] FIG. 1A shows a simplified diagram of a nanopore device and a nucleic acid according to an embodiment of the present invention. [Figure 1B] FIG. 1B illustrates a process for concatemerizing DNA fragments according to an embodiment of the present invention. [Figure 2] FIG. 2 illustrates the concatemerization process of a DNA fragment and a spacer DNA fragment according to an embodiment of the present invention. [Figure 3A] FIG. 3A shows a simplified block flow diagram of a method for sequencing DNA fragments by concatemerizing the DNA fragments and using single molecule sequencing according to an embodiment of the present invention. [Figure 3B] FIG. 3B shows a simplified block flow diagram of a method for sequencing DNA fragments using nanopore sequencing according to an embodiment of the present invention. [Figure 4] FIG. 4 shows a simplified block flow diagram of a method for analyzing the sequences of concatemers according to an embodiment of the present invention. [Figure 5]FIG. 5 shows a simplified block flow diagram of a method for more efficiently sequencing DNA fragments by increasing the concentration of the DNA fragments multiple times according to an embodiment of the present invention. [Figure 6] Figure 6 shows the size distribution of a plasma DNA pool sequenced by a nanopore sequencing device according to an embodiment of the present invention. The frequency distribution of sequenced plasma DNA fragments ranging from 0 to 500 base pairs is plotted. [Figure 7] FIG. 7 shows the size profile of plasma DNA from maternal plasma with a female fetus from data obtained by nanopore sequencing and Illumina sequencing platforms, according to an embodiment of the present invention. [Figure 8] Figure 8 shows chromosome read distributions compared to the distribution expected from a mappable human genome according to an embodiment of the present invention. The proportional distribution of nanopore sequence reads to each chromosome for each sample pool. The solid gray bars represent the proportion of nucleotides derived from each human chromosome based on the mappable portion of the reference human genome hg19. The remaining colored bars represent the proportion of sequenced reads that aligned to each human chromosome for the plasma DNA sample. [Figure 9] FIG. 9 shows a Circos plot of plasma DNA sequencing results from a cancer patient using nanopore sequencing and massively parallel sequencing, according to an embodiment of the present invention. [Figure 10] FIG. 10 shows the size distribution of concatemerized DNA molecules according to an embodiment of the present invention. [Figure 11] FIG. 11 shows the size distribution of plasma DNA molecules derived from concatemerized segments from non-pregnant women, according to an embodiment of the present invention. [Figure 12] FIG. 12 shows a genomic representation of aligned segments from concatenated plasma DNA from a non-pregnant woman according to an embodiment of the present invention. [Figure 13] FIG. 13 shows the size distribution of plasma DNA molecules derived from concatemerized segments from a pregnant woman carrying a male fetus according to an embodiment of the present invention. [Figure 14]FIG. 14 shows a genomic representation of aligned segments from enriched concatenated plasma DNA from a pregnant woman carrying a male fetus, according to an embodiment of the present invention. [Figure 15] FIG. 15 shows a block diagram of a system for implementing an embodiment of the present invention. [Figure 16] FIG. 16 illustrates a block diagram of an exemplary computer system that may be used with systems and methods according to embodiments of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0014] term A "tissue" corresponds to a group of cells grouped together as a functional unit. Two or more types of cells can be found within a single tissue. Different types of tissue can consist of different types of cells (e.g., liver cells, alveolar cells, or blood cells), but may also correspond to tissues from different organisms (maternal or fetal) or healthy versus tumor cells.
[0015] A "biological sample" refers to any sample obtained from a subject (e.g., a human, such as a pregnant woman, a person with or suspected of having cancer, an organ transplant recipient, or a subject suspected of having a disease process involving an organ (e.g., the heart in a myocardial infarction, or the brain in a stroke, or the hematopoietic system in anemia)) and containing one or more nucleic acid molecules of interest. Biological samples include blood, plasma, serum, urine, vaginal fluid, fluid from the crystalline lens (e.g., testes), vaginal washings, pleural effusions, peritoneal fluids, cerebrospinal fluid, saliva, sweat, tears, sputum, bronchoalveolar lavage fluid, nipple discharge, aspirates from different parts of the body (e.g., thyroid, breast), etc. Stool samples may also be used. In various embodiments, the majority of the DNA in a biological sample enriched for cell-free DNA (e.g., a plasma sample obtained by a centrifugation protocol) is cell-free, e.g., 50%, 60%, 70%, 80%, 90%, 95%, or 99% of the DNA may be cell-free. The centrifugation protocol may include, for example, 3,000 g for 10 minutes to obtain a fluid portion, which may then be centrifuged, for example, at 30,000 g for an additional 10 minutes to remove residual cells. The cell-free DNA in a sample may be derived from cells of various tissues, and thus the sample may contain a mixture of cell-free DNA.
[0016] "Nucleic acid" can refer to deoxyribonucleotides or ribonucleotides and their polymers in either single-stranded or double-stranded form. This term encompasses nucleic acids containing known nucleotide analogs or modified backbone residues or linkages, which are synthetic, naturally occurring, non-naturally occurring, have similar binding properties to the reference nucleic acid, and are metabolized in the same manner as the reference nucleotide. Examples of such analogs include, but are not limited to, phosphorothioates, phosphoramidites, methyl phosphonates, chiral methyl phosphonates, 2-O-methyl ribonucleotides, and peptide-nucleic acids (PNAs).
[0017] Unless otherwise specified, a particular nucleic acid sequence also implicitly encompasses conservatively modified variants thereof (e.g., degenerate codon substitutions) and complementary sequences, as well as the sequence explicitly set forth. Specifically, degenerate codon substitutions can be achieved by generating sequences in which the third position of one or more selected (or all) codons is substituted with mixed-base and / or deoxyinosine residues. (Batzer et al., Nucleic Acid Res. 19: 5081 (1991); Ohtsuka et al., J. Biol. Chem. 260: 2605-2608 (1985); Rossolini et al., Mol. Cell. Probes 8: 91-98 (1994)) The term nucleic acid is used interchangeably with gene, cDNA, mRNA, oligonucleotide, and polynucleotide.
[0018] In addition to referring to naturally occurring ribonucleotide or deoxyribonucleotide monomers, the term "nucleotide" can be understood to refer to related structural variants (including derivatives and analogs) thereof that are functionally equivalent for the particular context in which the nucleotide is used (e.g., hybridization to a complementary base), unless the context clearly indicates otherwise.
[0019] A "concatemer" is a continuous DNA molecule consisting of separate DNA fragments linked together in a single molecule. The various separate DNA fragments of a concatemer may or may not have the same sequence. At least some of the DNA fragments in a concatemer may have different sequences. The separate DNA fragments used to create a concatemer can be derived from various tissues present in a biological sample, such as, for example, if the DNA fragments are cell-free DNA fragments, they may be present in plasma and other cell-free DNA mixtures.
[0020] A "sequence read" refers to a string of nucleotides sequenced from any part or all of a nucleic acid molecule. For example, the sequence read may be the entire nucleic acid fragment present in a biological sample. A sequence read can be obtained from single-molecule sequencing. "Single-molecule sequencing" refers to sequencing a single template DNA molecule to obtain a sequence read without having to interpret base sequence information from clonal copies of the template DNA molecule. Single-molecule sequencing can sequence the entire molecule or only a portion of the DNA molecule. A majority of the DNA molecule is sequenced, for example, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 99% or more.
[0021] In the case of sequencing, a "signal" can refer to a measurement taken at a certain moment. Examples of such signals include optical or electrical signals. An optical signal can provide an image corresponding to a particular color corresponding to a particular base, for example, when hybridization occurs. A single image can be an array of sequencing devices (e.g., nanopores or other sequencing volumes), whereby a single image has many optical signals. An electrical signal can be a measurement across an electrode at an instant. An electrical signal over a period of time can provide one or more bases in the sequence of a DNA molecule. An electrical signal can include a current signal or a voltage signal.
[0022] "Nanopore" refers to an opening into which a molecule or portion of a molecule may be disposed, such that a signal may be detected based on one or more properties of the portion of the molecule within the nanopore. Nanopores can be constructed from a variety of materials, such as polymers, metals or other solid materials, proteins, or combinations thereof.
[0023] "Classification" refers to any number or other character associated with a particular characteristic of a sample. For example, a "+" sign (or the word "positive") can indicate that the sample has been classified as having a deletion or amplification. Classifications may be binary (e.g., positive or negative) or may have more classification levels (e.g., a scale of 1 to 10 or 0 to 1). The terms "cutoff" and "threshold" refer to a predetermined number used in an operation. For example, a cutoff size refers to the size at which fragments are excluded. A threshold may be a value above or below which a particular classification is applied. Either of these terms can be used in any of these contexts.
[0024] As used herein, the term "chromosomal aneuploidy" refers to a change in chromosomal dose from that of the diploid genome. The variation may be a gain or loss. This may involve an entire chromosome or a region of a chromosome.
[0025] As used herein, the term "sequence imbalance" or "aberration" refers to a significant deviation in the amount of a clinically relevant chromosomal region from a reference amount, as defined by at least one cutoff value. Sequence imbalance can include chromosomal dosage imbalance, allelic imbalance, mutational dosage imbalance, copy number imbalance, haplotype dosage imbalance, and other similar imbalances. As an example, allelic imbalance can occur when a tumor has one allele of a deleted gene or an amplification or differential amplification of two alleles in its genome, resulting in an imbalance at a specific locus in the sample. As another example, a patient may have a genetic mutation in a tumor suppressor gene. The patient may then develop a tumor in which the non-mutated allele of the tumor suppressor gene is deleted. Thus, there is a mutation dosage imbalance within the tumor. When a tumor releases its DNA into a patient's plasma, the tumor DNA mixes with the patient's constituent DNA (derived from normal cells) in the plasma. Using the methods described herein, it is possible to detect mutation dosage imbalances in this DNA mixture in the plasma. Aberrations can include deletions or amplifications of chromosomal regions.
[0026] The term "size profile" generally relates to the size of DNA fragments in a biological sample. A size profile can be a histogram that provides the distribution of the amount of DNA fragments of various sizes. To distinguish one size profile from another, various statistical parameters (also called size parameters or simply parameters) can be used. One parameter is the ratio of DNA fragments of a certain size or range of sizes to all DNA fragments or DNA fragments of other sizes or ranges.
[0027] Detailed Description of the Invention Embodiments can improve the efficiency of single-molecule sequencing when applied to cell-free DNA fragments in biological samples (e.g., plasma or serum) that are relatively short fragments of approximately 200 bases. Because cell-free DNA fragments, such as plasma DNA fragments, are small or short and typically present at low concentrations in biological samples, the use of nanopores to sequence cell-free DNA fragments is unlikely to yield accurate results. Small fragments passing through a small pore make it difficult to sequence the fragments while they are within the nanopore. Furthermore, small fragments may make alignment to a reference genome more difficult, given expected sequencing errors. Nanopore sequencing can have a sequencing error of approximately 10-15%. In some embodiments, efficiency can be improved by concatenating DNA fragments to generate longer molecules to be sequenced. In other embodiments, efficiency can be improved by increasing the concentration of DNA fragments in the sample.
[0028] Without these developments, single-molecule sequencing of plasma DNA might not have been practical. These approaches demonstrated that a range of plasma DNA abnormalities could be detected by single-molecule sequencing. The embodiments can be applied to many areas of medicine, including prenatal diagnosis, cancer assessment, inflammatory disease management, autoimmune disease assessment, and acute care such as trauma. Furthermore, it has been reported that methylated cytosines can be distinguished from unmethylated cytosines by single-molecule sequencing (12). We previously reported that detecting the methylation profile of plasma DNA could determine the fetal DNA fraction, detect pregnancy-associated diseases with abnormal placental methylation profiles, and detect aberrant methylation associated with cancer and systemic lupus erythematosus (7, 9, 11). Therefore, the embodiments can also be used for methylation detection applications.
[0029] I. Single-molecule sequencing of DNA fragments Plasma or serum DNA is a cell-free nucleic acid molecule found in the circulatory system of human subjects and released during the process of cell death, as part of natural turnover, or as part of pathological processes. Because DNA molecules are released into the circulation as part of the cell degradation process, they circulate in the form of short fragments (<200 bp) and are present in low concentrations. It is known that the majority of plasma DNA in healthy subjects originates from blood cells (6).
[0030] During pregnancy, the placenta contributes plasma DNA molecules to the maternal circulation, providing a means of accessing fetal DNA for noninvasive prenatal diagnosis (7). Tumors and cancers have high cell turnover rates and contribute DNA to plasma, allowing chromosomal and genetic abnormalities to be detected noninvasively, serving as liquid biopsies for cancer diagnosis, monitoring, screening, and prognosis (8, 9). Inflammatory conditions such as myocardial infarction, stroke, and hepatitis also increase plasma DNA contribution from inflamed organs (10). Autoimmune diseases such as systemic lupus erythematosus are also associated with plasma DNA abnormalities (11). In these diseases, plasma DNA profiles can reveal single-nucleotide variants associated with disease, chromosomal or chromosomal copy number abnormalities (increased or decreased copy number, aneuploidy), abnormal methylation (hypermethylation or hypomethylation), and size profile abnormalities (excessive short or excessive long DNA molecules). In summary, circulating cell-free DNA analysis has become an important tool for the development of molecular diagnostic tests for a wide range of diseases.
[0031] Single-molecule sequencing, not limited to the format manufactured by Oxford Nanopore Technologies, is an attractive platform for DNA sequencing analysis due to the high speed at which individual DNA bases can be identified. The omission of the amplification step simplifies the library construction workflow. It can sequence continuous stretches of DNA up to hundreds of kilobases long. However, such advantages cannot be easily applied to plasma DNA analysis. First, plasma DNA molecules are mostly short fragments, <200 bp in length. The DNA concentration in plasma and serum is substantially lower than that in tissue biopsies. Therefore, when plasma DNA samples were applied to a nanopore sequencer, sequencing efficiency was low. In other words, the DNA in the sample was dilute, and plasma DNA molecules were found to be too frequently directed toward the nanopore. Even if plasma DNA molecules reached the nanopore and were sequenced due to their short sequence length, the sequencing information provided only a very small amount of information about the human genome. The haploid human genome is 3.3 × 10 9Furthermore, nanopore sequencing error rates predict that sequencing short DNA fragments will be less accurate than sequencing long DNA fragments. Therefore, we aimed to develop an approach that would increase the frequency, or chance, that plasma DNA molecules could be processed by nanopore and other single-molecule sequencing devices.
[0032] A. Nanopore Sequencing Nanopore sequencing is a form of single-molecule sequencing in which DNA bases are detected through the use of nanopores. Oxford Nanopore Technologies uses a protein pore, α-hemolysin. A matrix of such pores is fabricated to sit on a membrane (5). In addition to protein pores, nanopores can be solid-state nanopores, where the nanopores are fabricated in semiconductor materials, including silicon compounds such as silicon nitride and graphene. Each pore may be connected to an electrical circuit.
[0033] FIG. 1A shows a simplified diagram of a nanopore 102 and a DNA molecule 104 according to an embodiment of the present technology. Electrodes 106 and 108 may define a portion of the nanopore or may be positioned near the nanopore. Electrodes 106 and 108 are shown as having conical or triangular ends. In some embodiments, the electrodes may have different shapes. For example, the electrodes may have flat, hemispherical, or rounded ends. Both electrodes do not have to have the same shape. For example, one electrode may have a conical end and the other electrode may be flat. The distance between the electrodes may be equal to, less than, or greater than the diameter or width of the nanopore. Electrodes 106 and 108 may be connected to a power source 110. Current 112 may tunnel from electrode 106 to electrode 108. The power source 110 may be in electrical communication with multiple nanopores and their respective electrode pairs.
[0034] As a DNA molecule 104 passes through the nanopore 102, there is a change in current, which can be measured by the meter 114. Different DNA bases, i.e., A, C, G, and T, will induce different magnitudes in the current change. By observing the voltage or current pattern at each nanopore, the sequence of the DNA molecule 104 that passed through the nanopore 102 could be determined. Because the sequencing is sensitive enough to detect DNA bases on a single molecule of DNA, i.e., no amplification of the DNA sequencing library is required, the speed of sequence detection is substantially improved.
[0035] However, a drawback is that the accuracy of DNA base identification is relatively poorer than sequencing techniques such as Illumina sequencing, which detect a consensus sequence from amplified clones of each DNA molecule. However, it has been shown that the accuracy of base detection can be substantially improved when 2D reads are interpreted (3). When preparing DNA samples for nanopore sequencing, a hairpin adapter is added to one end of each double-stranded DNA molecule. As such a DNA molecule approaches the nanopore, the double strand is unwound, and one end of the now single-stranded DNA molecule passes through the nanopore when a current change is detected. As sequencing approaches the end of this single-end, the complementary strand connected by the hairpin continues to pass through the nanopore and is sequenced. When both strands of the DNA molecule are sequenced, a consensus sequence can be derived, which is called a 2D read. As reported in the literature, 2D reads have higher base calling accuracy than 1D reads (sequences interpreted from only one strand).
[0036] B. Other Single-Molecule Sequencing This embodiment may also include single-molecule sequencing techniques other than those using nanopores. For example, DNA molecules can be hybridized onto a glass surface using a Helicos single-molecule sequencing system (SeqLL). Fluorescently labeled nucleotides can then be added to each DNA molecule, and images can be captured. The fluorescent molecules can then be cleaved and washed away, and another fluorescently labeled nucleotide can be added, repeating the process. Each nucleotide can have a different fluorescent label, and the DNA can be sequenced.
[0037] Another example can include the Pacific Biosciences single-molecule real-time (SMRT) sequencing method. In this method, a DNA polymerase enzyme can be immobilized at the bottom of a zero-mode waveguide (ZWM). The polymerase can then capture a single molecule of DNA. Fluorescently labeled nucleotides can then be incorporated into the DNA-enzyme complex. A detector can then detect the fluorescent signal and perform a base call. The fluorescent tag can then be cleaved and diffused away from the ZWM. This process can be repeated so that each type of nucleotide has a different fluorescent label, allowing for DNA sequencing.
[0038] II. Growth efficiency To increase efficiency, one option is to amplify the plasma DNA pool to increase the amount of genetic material in the sample within the finite volume applied to the sequencing chamber or flow cell. However, adopting such a protocol means that single-molecule sequencing is no longer performed, since both the original DNA template and the replicated DNA fragments are sequenced. This use of replicated DNA fragments can introduce errors when obtaining quantitative information about the original DNA fragments, such as the copy number of the original DNA fragment in a given genomic region, such as an entire chromosome.
[0039] The following sections describe two approaches to increasing efficiency. In one approach, DNA fragments can be ligated to form concatemers before single-molecule sequencing is performed. In another approach, DNA fragments are highly enriched in the sample before single-molecule sequencing (e.g., before nanopore sequencing).
[0040] A. Concatemer In some embodiments, DNA fragments (e.g., plasma DNA fragments) may be concatemerized, as shown in FIG. 1B. Concatemerization can link a series of DNA fragments together. For example, DNA fragment 152 and DNA fragment 154 may be present in a sample. As a first step to attach the DNA fragments to each other, the ends of DNA fragment 152 and DNA fragment 154 are prepared by adding phosphate groups to the ends to yield DNA fragment 156 and DNA fragment 158. The DNA fragments can then be linked together by blunt-end ligation using a ligase enzyme to form a long DNA molecule, such as concatemer 160. Concatemer 160 can be considered a new molecule that is a combination of DNA fragments 152 and 154. Concatemer 160 has a sequence that is a combination of the subsequences corresponding to DNA fragments 152 and 154.
[0041] Figure 2 illustrates a similar approach to forming concatemers with spacer fragments (also referred to as "spacers") between DNA fragments. DNA fragment 202 and DNA fragment 204 can be "A-tailed," adding an A nucleotide to one end of each strand of the DNA fragments to form DNA fragment 206 and DNA fragment 208. Spacer DNA fragment 210 can have a known sequence with a T nucleotide at one end of each strand of the spacer DNA fragment. The known sequence of the spacer DNA fragment can be 20 base pairs or less, including 4-10 base pairs and 10-20 base pairs, without A or T nucleotides added to the end of the spacer fragment. Other complementary nucleotides can be used at the ends of the spacer DNA fragment and the plasma DNA fragment as well. Spacer DNA fragment 210 can then be ligated to DNA fragment 206 and DNA fragment 208 by a ligase enzyme to form a long DNA molecule, shown as concatemer 212. The spacer DNA fragment can be placed between two DNA fragments that are not spacers. For example, one spacer DNA fragment may be between two DNA fragments extracted from a biological sample. Concatamer 212 may have a sequence that is a combination of sequences corresponding to DNA fragments 206 and 208 along with one or more spacer DNA fragments 210.
[0042] In some embodiments, if the known sequence of the spacer fragment does not appear in the reference genome (e.g., the human genome) corresponding to the subject, or appears less than a certain number of times (e.g., less than two or three), the spacer fragment can indicate the end of one DNA fragment and the beginning of another DNA fragment. A computer system can identify the sequence of the spacer fragment by comparing the subsequence with an expected set of one or more known sequences used in spacer fragments. In some embodiments, the spacer fragment can simply be a combination of AT at one position. In such an example, the use of A and T may be primarily for attachment, as opposed to identifying the beginning and end of the DNA fragment.
[0043] The method using a spacer can include identifying a starting base of a subsequence corresponding to a DNA fragment of a biological sample. The identification of the starting base can be based on identifying one of the known sequences of the spacer DNA fragment. The end of the subsequence can be identified based on identifying one of the known sequences of the spacer DNA fragment after the subsequence.
[0044] In these approaches, when long molecules (concatamers) reach the nanopore, many plasma DNA fragments can be sequenced because the concatemers can contain several plasma DNA fragments within a single long molecule. Because plasma DNA molecules can be naturally double-stranded, hairpin adapters are applied to the ends of the concatemers to generate 2D reads; that is, both strands are read. After sequencing, the original plasma DNA molecule units incorporated into the concatemers can be identified. Creating long molecules from several DNA fragments and sequencing the long molecules within the nanopore is more efficient than waiting for several identical molecules to move separately into the nanopore and be sequenced.
[0045] 1. Use in Single Molecule Sequencing FIG. 3A shows a simplified block flow diagram of a method 300 for sequencing DNA fragments by concatemerizing the DNA fragments and using single molecule sequencing according to an embodiment of the present invention.
[0046] At block 302, the method 300 can include receiving a plurality of DNA fragments. The plurality of DNA fragments can be cell-free DNA fragments from a biological sample. The biological sample can be plasma or serum. The DNA fragments can be separated from other components of the biological sample, for example, when plasma is separated from other components of blood. The DNA fragments can be any DNA fragments described herein.
[0047] In block 304, method 300 can include concatemerizing the first set of DNA fragments to obtain a first concatemer. Concatemerization can be by any method described herein. The first concatemer can include DNA fragments other than the first set. Thus, the first set of DNA fragments may not include all of the DNA fragments that make up the first concatemer. Other DNA fragments in the first concatemer can include spacer DNA fragments.
[0048] Block 306 indicates that method 300 can also include performing single-molecule sequencing of the first concatemer to obtain a first sequence of the first concatemer. Single-molecule sequencing can include nanopore sequencing. The first concatemer can be provided to a sequencing device as part of the single-molecule sequencing. The sequencing device can be a nanopore device, an optical waveguide, a flow cell configured to hybridize the concatemer onto the flow cell, or any reaction cell or location in which sequence detection of a single DNA molecule occurs. An example of a flow cell can include a flow cell having oligonucleotides attached to its surface, which can hybridize with the first concatemer. Method 300 can further include detecting a plurality of signals corresponding to the first concatemer using the sequencing device. The plurality of signals can correspond to the first sequence of the first concatemer. The signals can include signals from fluorescent labels or other labels having optically detectable signals attached to the first concatemer. The signal from the fluorescent label can be detected by a photodetector, laser, charge-coupled device, zero-mode waveguide, other optically sensitive element capable of determining the presence or absence of an optical event, or a combination thereof.
[0049] 2. Detection of electrical signals in nanopores 3B shows a simplified block flow diagram of a method 350 for sequencing DNA fragments using nanopore sequencing according to an embodiment of the present invention. In method 350, the sequencing device is a nanopore, which may be present in an array of nanopores on a substrate. Aspects of method 350 may be performed when performing method 300.
[0050] At block 352, the method 350 can include receiving a plurality of DNA fragments. The plurality of DNA fragments can be cell-free DNA fragments from a biological sample. The biological sample can be plasma or serum. The DNA fragments can be separated from other components of the biological sample, for example, when plasma is separated from other components of blood.
[0051] The DNA fragments can be received in a container used for concatemerization. The container can be a vial or a tube, such as an Eppendorf tube. The DNA fragments can be mixed with a ligase enzyme and a buffer solution in the container.
[0052] Block 354 indicates that method 350 also includes concatemerizing the first set of DNA fragments to obtain a first concatemer. Concatemerization can be by any method described herein. The first concatemer can include DNA fragments other than the first set. Thus, the first set of DNA fragments may not be all of the DNA fragments that make up the first concatemer.
[0053] Multiple concatemerization processes can be performed in parallel, and each process forms a separate concatemer. Various concatemers can be of different lengths, for example, because different DNA fragments can be of different lengths, and therefore different numbers of DNA fragments can be incorporated into the concatemer. For example, one concatemer can be composed of three DNA fragments, and another concatemer can be composed of 100 DNA fragments.
[0054] At block 356, the method 350 may further include passing the first concatemer through a first nanopore. The first nanopore may be one of multiple nanopores on the substrate. Passing the first concatemer through the first nanopore may include passing a first strand of the first concatemer through the nanopore. After the first strand passes through the nanopore, a second strand of the first concatemer may pass through the nanopore. The first and second strands may be connected by a hairpin adapter. In this manner, both strands can be sequenced, thereby increasing the accuracy of base calling. The electrical signals of both strands can be compared as part of the base calling.
[0055] In block 358, a first electrical signal can be detected when the first concatemer passes through the first nanopore. The first electrical signal can correspond to the first sequence of the first concatemer. The first electrical signal can include a current or a voltage. In the absence of the first concatemer passing through the first nanopore, an ionic current can pass between the electrodes. When a biomolecule, such as a concatemer, passes between the electrodes, the biomolecule can affect the passage of ions or electrons between the electrodes. As a result, the current or voltage can decrease. The magnitude of the change can be related to the portion of the biomolecule between the electrodes. For example, when a specific nucleotide or functional group passes between the electrodes, the current or voltage can have a specific electrical signal signature.
[0056] In some embodiments, the nanopore can be part of an electrical circuit that includes two electrodes. The current between the two electrodes can vary based on which nucleotide (base) or corresponding tag is present in the nanopore. The first electrical signal can be detected using any suitable technique for measuring voltage or current in the circuit.
[0057] In block 360, the method 350 can include analyzing the first electrical signal to determine a first sequence. The analysis can include comparing the pattern of the electrical signal to known patterns corresponding to specific bases. Nanopore base calling can include using different models, including hidden Markov models, as described in Schreiber J. and Karplus K., “Analysis of nanopore data using hidden Markov models,” Bioinformatics 2015 31: 1897-1903, incorporated herein by reference for all purposes. The analyzing can include analyzing the first electrical signals of the first and second strands by a computer system to determine a first sequence. For example, the sequences can be determined for the first and second strands, and the two sequences can be compared to each other. The sequences must be complementary. Non-complementary positions can be, for example, ignored or reanalyzed.
[0058] In some embodiments, analysis of the first electrical signal can be used to determine methylation classifications for various sites, such as CpG sites, of the first concatemer. Methylation classifications can include whether the base is methylated, whether aberrant methylation (hypermethylation or hypomethylation) is present (e.g., whether a region such as a CpG island has aberrant methylation), and whether the concatemer is hydroxymethylated.
[0059] At block 362, the method 350 may include aligning subsequences of the first sequence to identify fragment sequences corresponding to each of the first set of DNA fragments. The subsequences may be any set of contiguous bases of the first sequence, for example, as specified by a sliding window. The alignment may be performed against a reference genome that allows for mismatches in the alignment to the reference genome. Aligning the subsequences is described in more detail below.
[0060] Concatemerization can be performed on a second set of DNA fragments to obtain a second concatemer, and on other sets of DNA fragments to obtain other concatemers. Each concatemer can have a different combination or permutation of DNA fragments than the other concatemers. Each concatemer can pass through a first nanopore or other nanopores of a plurality of nanopores, which can include a sequencing device. As the other concatemers pass through the nanopores, other electrical signals can be obtained. The other electrical signals can correspond to the respective sequences of the other concatemers. Details of the method for generating the other concatemers can be similar to the method for generating the first concatemer.
[0061] Method 350 can also include determining the size of each of the first set of DNA fragments in the first concatemer. The sizes of the DNA fragments in other sets of DNA fragments in other concatemers can also be determined. As an example, the size of the DNA fragments can be determined by aligning subsequences with a reference genome or known spacer sequences. For example, if spacer sequences can be identified, the length of the DNA fragments can be determined as the number of bases between two spacer sequences. In such embodiments using spacer sequences, the identified sequences of the DNA fragments do not need to be aligned to a reference genome to identify them, as the spacer sequences can provide such information. Furthermore, the sequences of the DNA fragments can be assembled after or instead of alignment to a reference genome. When aligning to a reference genome, determining the size of the DNA fragments can include determining the length of the longest subsequence that aligns to a region of the reference genome.
[0062] 3. Subsequence Alignment As shown in Figure 4, an embodiment may include a method 400 performed by a computer system. Figure 4 shows a simplified block flow diagram of a method for analyzing sequences of concatemers according to an embodiment of the present invention.
[0063] At block 402, the method 400 can include receiving a first sequence of a first concatemer generated by concatemerizing a first set of DNA fragments. In some embodiments, the first concatemer can be generated by concatemerizing a first set of DNA fragments and a second set of DNA fragments. The first sequence can be received from a base calling routine that can reside on a sequencing device on which the computer system can reside. As another example, the computer system can be separate from the sequencing device, and the first sequence can be received via a network connection or a removable memory device. The concatemer can be, for example, any concatemer as described herein.
[0064] At block 404, method 400 may also include aligning subsequences of the first sequence to identify fragment sequences corresponding to each DNA fragment of the first set of DNA fragments. In some embodiments, aligning the subsequences may include aligning the subsequences to a second set of DNA fragments. The second set of DNA fragments may be spacer DNA as described herein. The spacer DNA may be 20 nucleotides or less, including 15-20, 10-15, and 5-10 nucleotides, or a known sequence equivalent thereto.
[0065] In some embodiments, method 400 may include aligning a subsequence of the first sequence to a reference genome. The reference genome may be the human genome. To identify the original unit of the plasma DNA molecule, embodiments may align a long DNA sequence to the human genome over a window. For example, a sliding window (e.g., 100-300 bases) may be selected from the long sequence of the concatemer, and the window sequence (subsequence) may be aligned to the reference genome. The reference genome may be a derivative of the reference human genome, including, but not limited to, a subset of the human genome sequence, a repeat-masked genome, exons, or a portion of the genome with moderate or balanced GC content.
[0066] The window can be advanced or retreated (slid) by an amount less than the window length (e.g., 20–50 bases). The window in its new position can be considered a second window. Subsequences in this second window can also be aligned to the reference genome. If a subsequence in the second window aligns to a subsequence in the reference genome that overlaps with a previous subsequence in the reference genome, the two subsequences can be considered to be part of the same DNA fragment. If two sliding windows align to different, non-contiguous, or non-overlapping regions of the reference genome, the sequences of the two DNA fragments can be distinguished.
[0067] When a subsequence does not align with the genome, but the preceding and following subsequences do, a crossover (edge) between two DNA fragments can be identified. The two aligned DNA fragments can be analyzed to determine the specific point of crossover (e.g., the starting base of one DNA fragment and the ending base of the other). This crossover may be the end or the beginning of a subsequence of the DNA fragment. This approach can be particularly useful when spacers are not used to construct concatemers. In other embodiments, specific sequences (e.g., specific barcodes or spacers) can be added to the ends of the original molecules, and these specific sequences can indicate the end of one molecule and the beginning of another.
[0068] Thus, the present invention can find stretches or segments of DNA bases (generally up to several hundred bases in length) that belong to different regions of the human genome. Each adjacent stretch or segment of DNA bases can represent one original plasma DNA molecule. Adjacent juxtaposed DNA segments or stretches aligned to different distant parts of the human genome can belong to other plasma DNA molecules assembled into concatemers.
[0069] The size of the window and the step size by which the window is moved or slid can be adjusted based on the desired resolution of the DNA fragment size. A smaller step size can increase the resolution of the determined size of the DNA fragment while increasing the computational intensity. A larger window size may not recognize DNA fragments smaller than the window size, but a small window size may not obtain a unique alignment with the genome.
[0070] The window size and step size may be dynamically adjusted. For example, a large step size can be used to narrow the region of potential matches, and a smaller step size can be used to more accurately identify matches. Aligning subsequences can be to chromosomes or chromosomal regions of the reference genome. Aligning subsequences can include allowing a large number or frequency of mismatches to account for sequencing errors. For example, aligning sequences can allow for no more than about 10-15% mismatches.
[0071] As described below, this approach was effective in sequencing plasma DNA molecules, identifying human chromosomes, detecting proportional differences in the nucleotide content of each chromosome, and determining the size profile of the original plasma DNA sample before concatenation. Concatemers can contain DNA fragments from the entire genome. The DNA fragments in concatemers may be randomly distributed across all or most of the chromosomes.
[0072] B. increased concentration Another embodiment can include increasing the concentration of the plasma DNA library loaded into the sample chamber of the flow cell. Increasing the concentration of plasma DNA is not typically expected to increase the efficiency of nanopore sequencing. Nanopores have relatively high sequencing error, and given the small DNA fragments required for analysis from plasma and other cell-free samples, as well as the low DNA concentration in such samples, nanopores are not expected to effectively sequence the fragments. Increasing the concentration of DNA fragments is not expected to address this issue. However, concentrating the extracted DNA or input sequencing library enhances the chances that DNA molecules will reach the nanopore or other single-molecule analysis technology. Concentrating the extracted DNA includes concentrating it beyond the level required to obtain a volume compatible with the sequencing technology. In some cases, the concentration of the extracted DNA can be increased by 10-fold or more. In other words, the volume may be reduced to less than 10% of its original volume.
[0073] 5 shows a simplified block flow diagram of a method 500 for more efficiently sequencing DNA fragments by increasing the concentration of the DNA fragments multiple-fold, according to an embodiment of the present invention. Method 500 can be used with a variety of single-molecule sequencing platforms. In the example provided, a nanopore sequencing platform is described.
[0074] At block 502, the method 500 may include receiving a biological sample containing a plurality of DNA fragments. The biological sample may have a first concentration of DNA fragments in a starting volume. The biological sample may be of various types, for example, as described herein. For example, the biological sample may be plasma or serum.
[0075] At block 504, the method 500 can also include concentrating the biological sample to have a second concentration of DNA fragments. In various examples, the second concentration of DNA fragments can be increased by 5 or more times, 6 or more times, 7 or more times, 8 or more times, 9 or more times, 10 or more times, 50 or more times, 100 or more times, 500 or more times, or 1000 or more times greater than the first concentration of DNA fragments. The concentration can be measured per volume or per mass.
[0076] Concentration can be achieved in a variety of ways, as will be understood by those skilled in the art. For example, biological samples can be concentrated by vacuum drying, fluid removal by percolation or filtration, or other concentration techniques known to those skilled in the art. Filtration or percolation can be combined with centrifugation to force the fluid through a size filter or molecular sieve. Semipermeable membranes that allow unidirectional fluid flow can also be used for concentration. The volume of a biological sample after concentration may decrease inversely with an increase in concentration. For example, a five-fold increase in concentration can result in a five-fold decrease in volume.
[0077] The concentration can be significantly increased compared to conventional processes. In some conventional processes, a small amount of plasma DNA is extracted from a single volume of plasma. The volume can be further concentrated to reduce the volume to meet the reaction volume or other requirements of the analytical device. For example, in conventional processes, 210 μL of plasma DNA can be extracted from 4 mL of plasma. To provide a total reaction volume of 100 μL, 210 μL of plasma DNA can be concentrated to 85 μL. In conventional processes, the increase in concentration is less than three-fold, and the plasma DNA is not concentrated to improve sequencing accuracy or precision. In the present method, the increase in concentration results in more frequent passage of DNA fragments through the nanopore or other sequencing device, thus improving detection and analysis.
[0078] At block 506, the method 500 can further include passing the plurality of DNA fragments through a nanopore on the substrate. In some embodiments, the method can include a single-molecule sequencing technology other than a nanopore. For example, the sequencing technology can include SMRT technology by Pacific Biosciences or Helicos sequencing by SeqLL. The single-molecule sequencing technology can include the technology described in Eid J. et al., "Real-time DNA sequencing from single polymerase molecules," Science 2009 323: 133-138, which is incorporated herein by reference for all purposes.
[0079] In block 508, for each of the plurality of DNA fragments, an electrical signal can be detected as the DNA fragment passes through the nanopore. The electrical signal can correspond to a sequence or subsequence of the DNA fragment. The electrical signal can include a current or voltage or any electrical signal described herein. If other sequencing techniques are used, a fluorescent signal can be used instead of an electrical signal.
[0080] At block 510, method 500 can include analyzing the electrical signals to determine the sequence or subsequence of the DNA fragments. Method 500 can include determining the size of the DNA fragments and the sizes of the DNA fragments, for example, by using the alignment information. As a result, the size distribution of the DNA fragments can also be determined. Chromosomal differences can also be determined based on the size distribution of the DNA fragments, as described, for example, in U.S. patent application Ser. No. 12 / 940,992, filed Nov. 5, 2010, entitled "Size-based genomics," U.S. patent application Ser. No. 13 / 308,473, filed Nov. 30, 2011, entitled "Detection of genetic or molecular aberrations associated with cancer," and U.S. patent application Ser. No. 13 / 789,553, filed Mar. 7, 2013, entitled "Size-based analysis of fetal DNA fraction in maternal plasma."
[0081] In some embodiments, the electrical signal can correspond to a methylation classification of the first concatemer, which can include whether the base is methylated, whether aberrant methylation (hypermethylation or hypomethylation) is present, and whether the concatemer is hydroxymethylated.
[0082] The method can include aligning a sequence or subsequence of the DNA fragment with a reference genome. In particular, the alignment can be to a particular chromosome or chromosomal region of the reference genome.
[0083] The same DNA fragment can pass through the same nanopore multiple times. With each pass, an electrical signal can be detected. The electrical signals from different passes can be compared to help identify the sequence. Increasing the concentration of DNA fragments can be used with single-molecule sequencing techniques other than nanopores.
[0084] III. Examples using increased concentrations Examples demonstrate that plasma DNA can be concentrated to increase the efficiency of nanopore sequencing while still providing accurate results. Sequencing using increased concentrations is further described in Cheng SH et al., "Noninvasive prenatal testing by nanopore sequencing of maternal plasma DNA: feasibility assessment," Clin. Chem. 61: 10 (2015).
[0085] A. Materials and Methods Plasma samples were obtained from four groups of individuals recruited with informed consent and institutional approval: pregnant women in their third trimester carrying a male fetus, pregnant women in their third trimester carrying a female fetus, adult men, and non-pregnant women. EDTA plasma samples were pooled within each group to provide at least 20 mL of plasma per group. Pooled plasma samples were extracted using a QIAamp DSP DNA Blood Mini Kit (Qiagen, Germany) (2). 1,050 μL of eluted plasma DNA per pool was concentrated to 85 μL using a Speedvac concentrator (Thermo Fisher Scientific, Waltham, MA). Each concentrated plasma DNA pool was then completely consumed for DNA library preparation using an end-repair and A-tailing module (New England Biolabs, Ipswich, MA) and a genomic DNA sequencing kit (SQK-MAP-005, Oxford Nanopore Technologies, UK). Each library (150 μL) was fully loaded into a MinION Flow Cell (v7.3) (Nanopore) and sequenced. The output data files were base-called using METRICHOR™ software (Nanopore). 2D reads were extracted and aligned to the reference genome hg19 using LAST Genome-Scale Sequence Comparison software (Computational Biology Research Consortium, Japan).
[0086] Each library was sequenced until exhaustion, which took 6–24 hours. 26.9%–32.5% of reads passed the basecaster. For plasma pools from pregnant women carrying male fetuses, pregnant women carrying female fetuses, adult males, and non-pregnant women, the number of 2D reads was 56,844, 50,268, 35,878, and 36,167, respectively. The average observed identity, the percentage of bases in the reads that aligned to matching bases in the reference sequence (3), was 82.7% (81.4–84.5%). Of the 2D reads, 16.9% (15.6–23.9%) aligned to unique genomic locations and were further analyzed.
[0087] The sequenced plasma DNA fragments aligned to the human genome ranged in length from 76 to 5,776 bp, peaking at 162 bp (155–168 bp) (Figure 6). Each graph in Figure 6 shows the size of the sequenced plasma DNA fragment in base pairs on the x-axis and the frequency of that plasma DNA fragment size as a percentage of all sequenced plasma DNA fragments. The graphs represent DNA sequenced from maternal plasma carrying a male fetus, maternal plasma carrying a female fetus, adult male plasma, and non-pregnant female plasma. The peak plasma DNA sizes from the four plasmas are consistent with previous findings based on Illumina sequencing platforms (4). Trace amounts (0.06–0.3%) of long plasma DNA fragments (>1,000 bp) were observed in the nanopore sequencing data but not in previous data analyses of other sequencing platforms (4).
[0088] Figure 7 shows the size profile of plasma DNA from maternal plasma carrying a female fetus obtained by nanopore sequencing and from data acquired by an Illumina sequencing platform. The nanopore sequencing data are the same as the maternal plasma carrying a female fetus in Figure 6. The size profiles in Figure 7 have similar shapes, with peaks of approximately the same size. For example, the size ratio of nanopore sequencing for fragments below 150 bp to fragments between 161 bp and 170 bp is 1.21 compared to 1.10 for Illumina sequencing. These results demonstrate that the size profile of plasma DNA can be accurately determined using nanopore and enriched plasma DNA. The peak in the 250-400 bp range is more prominent in data obtained by nanopore sequencing than by Illumina sequencing. This peak corresponds to cell-free DNA derived from dinucleosomes, and its presence varies between individuals. Furthermore, Illumina sequencing is not very efficient at sequencing fragments in this size range.
[0089] B. Analysis of readings Figure 8 shows the chromosomal read distribution compared to the distribution expected from a mappable human genome, labeled hg19. Chromosomes are listed on the x-axis. On the y-axis, the proportional distribution of reads to each chromosome (genomic representation) was calculated for each sample by counting the number of reads aligned to each chromosome relative to the total number of uniquely aligned reads sequenced from that sample and expressed as frequencies and percentages. Plotted in hg19 are results from plasma of mothers carrying male fetuses, plasma of mothers carrying female fetuses, male plasma, and non-pregnant female plasma. The read distribution to autosomes for all four plasma DNA pools was comparable to that expected for a mappable human genome.
[0090] Differences between chromosome X and chromosome Y were observed in the read distribution. The percentage of reads mapping to chromosome X was lower in male plasma (2.70%) compared to female plasma (5.22%). Chromosome Y sequences were detected in the adult male (0.30%) plasma DNA pool but not in the non-pregnant female plasma DNA pool.
[0091] The plasma DNA pool from women carrying male fetuses had 0.11% reads aligned to chromosome Y. Consistent with previous data (2), 0.018% of reads aligned to chromosome Y sequences in the plasma DNA pool from women carrying female fetuses. The presence of chromosome Y-aligned reads in women carrying female fetuses may be the result of known errors in aligning to the male genome. Maternal plasma DNA pools carrying male fetuses had approximately 1% fewer chromosome X sequences than women carrying female fetuses.
[0092] The relative distribution of reads for chromosome X and chromosome Y is observed to be similar to expected results. Male plasma had more chromosome Y than plasma from both pregnant and non-pregnant women. Male plasma had less chromosome X than plasma from pregnant and non-pregnant women. Male plasma contained approximately half the amount of X chromosome as non-pregnant women, which predicts that males have one X chromosome and females have two X chromosomes. Plasma from mothers carrying male fetuses had a higher distribution of chromosome Y reads than plasma from mothers carrying female fetuses and women carrying female fetuses.
[0093] Therefore, fetal DNA sequencing and differences in chromosome X dosage between male and female fetuses are detectable by nanopore sequencing and plasma DNA enrichment. The amount of chromosome X in male fetuses is equivalent to that in female fetuses with monosomy X or Turner syndrome. This observation suggests the potential feasibility of nanopore sequencing for the noninvasive detection of fetal chromosomal aneuploidies, such as monosomy X, or copy number abnormalities. Because monosomy X represents a loss of one chromosome copy in the genome, the degree of copy number change is equivalent to that of trisomy, where one chromosome copy is added to the genome. These data also demonstrate that our protocol can be applied noninvasively to the detection of fetal trisomy 21, trisomy 18, trisomy 13, and other fetal chromosomal aneuploidies. These data suggest the feasibility of nanopore sequencing-based NIPT and point-of-care NIPT.
[0094] IV. Cancer Sequences with Enhanced Concentrations Circulating cell-free DNA can be used as a "liquid biopsy" for real-time monitoring of cancer. Cell-free DNA reveals genetic abnormalities found in the underlying tumor, which can be detected by massively parallel sequencing techniques. These chromosomal abnormalities in the plasma DNA of cancer patients can also be detected using nanopore sequencing. Plasma DNA samples from two patients with hepatocellular carcinoma (HCC) were analyzed by both nanopore sequencing and massively parallel sequencing on an Illumina platform.
[0095] A. Materials and Methods Twenty milliliters of peripheral blood was collected from two patients diagnosed with HCC before surgery. Plasma was isolated by centrifugation at 1600 × g for 10 minutes followed by centrifugation at 16,000 × g for 10 minutes. DNA was extracted from 8 mL of plasma using a QIAamp DSP DNA Blood Mini Kit (Qiagen). Three-quarters of the plasma DNA was subjected to nanopore sequencing, and the remaining DNA was sequenced using a NextSeq 500 (Illumina) system.
[0096] Nanopore sequencing libraries were prepared using the End Repair / dA Tailing Module (NEB) and a Nanopore Sequencing Kit (SQK-NSK007, Oxford Nanopore Technologies). The libraries were fully loaded into a Minion Flow Cell (R9 version) and sequenced on a Minion Mk1B sequencer (Nanopore). The output data files were base-called using METRICHOR™ software (Nanopore). 2D reads were extracted and aligned to the reference genome hg19 using LAST software. Illumina sequencing of plasma DNA was performed as previously described (8).
[0097] B. Results The proportional distribution of aligned reads to each chromosome arm (genomic representation, GR) was calculated for each sample. In other words, the number of high-quality pass filter reads aligned to the chromosome arm, p or q, was expressed as a percentage of the high-quality pass filter reads sequenced from the sample. The difference in GR relative to the plasma DNA sample of a normal individual was then calculated. If the GR of a chromosome arm was three standard deviations above the mean of the control group, the region was considered to exhibit copy number gain. If the GR of a chromosome arm was three standard deviations below the mean of the control group, the region was considered to exhibit copy number loss.
[0098] Figure 9 shows GR differences in two cases of HCC (HOT530 and HOT536) by nanopore sequencing (outer ring) and the Illumina platform (inner ring). The differing chromosomes are shown on the outer ring. The analyzed regions were chromosome arms. Chromosome gains are represented by green bars extending outward from the center of each ring. Chromosome losses are represented by red bars extending inward from the center of each ring. As shown in Figure 9, the nanopore sequencing results were nearly consistent with those generated by the Illumina platform. The sequenced samples also tended to have longer DNA compared to plasma DNA from non-cancer subjects. This example demonstrates that nanopore sequencing with increased concentration can be used to analyze plasma DNA from cancer patients.
[0099] V. Examples Using Concatamers Plasma DNA molecules are typically short (<200 bp), and nanopore sequencing can typically be used to sequence longer DNA molecules. The efficiency of sequencing short plasma DNA molecules can be improved by linking or joining individual molecules to construct longer molecules called concatemers.
[0100] A. Materials and Methods 1. Generation of Plasma DNA Concatemers Samples from two different subjects were tested. Twenty milliliters of peripheral venous blood was collected from both a non-pregnant female subject and a female subject carrying a male fetus. Plasma was collected after centrifugation at 1600 × g for 10 minutes and then centrifuged at 16,000 × g for another 10 minutes. DNA was then extracted from 8 mL of plasma using a QIAamp DSP DNA Blood Mini Kit (Qiagen), yielding a plasma DNA volume of 420 μL. The extracted DNA was concentrated to 85 μL using a SpeedVac concentrator (Thermo Scientific) and end-repaired using the NEBNext End Repair Module (New England Biolabs, NEB). The end-repaired DNA was purified using a MinElute Reaction Cleanup Kit (Qiagen) and eluted with 20 μL of Buffer EB. Next, 20 μL of plasma DNA was ligated by adding Blunt / TA Ligase Master Mix (NEB) and incubating at 25°C for 4 hours, and after incubation, the DNA was purified using the MinElute Reaction Cleanup Kit.
[0101] 2. Nanopore Sequencing The enriched DNA was then used for nanopore sequencing library preparation using the End Repair and A-Tailing Module (NEB) and a genomic DNA sequencing kit (SQK-MAP-005, Oxford Nanopore Technologies). The library was fully loaded into a Minion Flow Cell (v7.3) (Nanopore) and sequenced. The output data file was base-called using METRICHOR™ software (Nanopore). 2D reads were extracted.
[0102] 3. Alignment Alignments were performed using the LAST software. The first match for each possible starting position was found. Matches were limited to a minimum length match occurring the most times in the reference genome, or matches were limited to a specific, predetermined length. From these initial matches, additional alignments of sequences longer than these initial matches were performed, and those with a certain gap score were retained based on the desired tolerance for sequencing errors. If multiple alignments shared the same endpoint, the alignment with the highest score was retained. In this way, fragments aligned to the human genome were identified. The parameters and rules used in LAST may be modified depending on considerations of precision and accuracy.
[0103] B. Results The ligated plasma DNA was sequenced on the MinION for 6 hours until the library was exhausted.
[0104] 1. Non-pregnant women Base calling resulted in 2,234 2D reads, with read lengths ranging from 86 to 8,672 bp. Figure 10 shows the frequency distribution curve for concatemer sizes sequenced from non-pregnant women. Concatemer size in base pairs is on the x-axis. The frequency (expressed as a percentage) at which a given concatemer size is present in the sequenced sample is plotted on the y-axis. To improve the readability of the graph, the outlier data point for the 8,672 bp outlier data point is not shown. Approximately 20–50 long DNA molecules were detected for each sequenced sample.
[0105] The size of the sequenced molecules is much longer than the length of typical plasma DNA fragments, which are generally less than 200 bp. These data indicate that the plasma DNA fragments were successfully assembled as concatemers. The reads were then aligned to the human genome (hg19). Stretches of bases or segments belonging to separate regions of the human genome were separated and considered as a single plasma DNA fragment. Overall, 3,801 uniquely mapped segments of sequenced plasma DNA fragments were obtained, and 80.6% of the uniquely mapped segments showed sequence identity to the reference genome.
[0106] Figure 11 shows the size distribution of aligned segments from non-pregnant women. The size of these aligned segments ranged from 78 bp to 560 bp, with most being less than 200 bp. The peak size was 162 bp. The average size was 173 bp. The median size of 162 bp is consistent with previous observations of plasma DNA size by massively parallel sequencing on the Illumina platform and by nanopore sequencing using increased concentrations of DNA fragments.
[0107] Figure 12 shows the calculated distribution of aligned segments compared to a reference genome from a male. The proportional distribution (genomic representation) of aligned segments to each chromosome was calculated. The calculated distribution was similar to the distribution of hg19 among all autosomes. For the sex chromosomes, the genomic representation of chromosome X was 5.79%, which is expected for a female sample. No misalignments were found with respect to chromosome Y. Therefore, the method using concatemers with nanopores can distinguish between male and female plasma.
[0108] 2. A pregnant woman carrying a male fetus The male fetus from the pregnant woman was 38 weeks gestational age and 4 days gestational age when the blood sample was collected. The pregnancy was considered normal. Plasma DNA fragments were concatenated before processing for nanopore sequencing. Concatemer sizes ranged from 100 to 13,466 bp, with a median size of 676 bp and an average size of 965 bp.
[0109] Figure 13 shows the size profile of aligned segments of concatemers for a woman carrying a male fetus. The size profile is typical of plasma DNA. The median fragment size was 196 bp. The peak size was 174 bp, with sizes ranging from 92 to 2,934 bp. Only approximately 0.4% of the fragments had a size greater than 2,000 bp. This distribution is similar to other size distributions obtained using concatemer methods, enhanced concentration, or massively parallel sequencing. For example, the distribution is similar to the size distribution found using enhanced concentration methods for maternal plasma carrying a male fetus, as shown in Figure 6.
[0110] Figure 14 shows the calculated chromosome distribution of aligned segments for a female pregnant woman carrying a male fetus compared to the reference genomes of a male and a non-pregnant woman. The distribution of concatemers shows chromosome X levels between those of males and non-pregnant women. This suggests that the chromosome X data reflect monosomy of chromosome X in a normal male fetus. Additionally, the concatemers show evidence of fetal-derived chromosome Y. Some deviations in the read distribution obtained by sequencing from the reference genome may be the result of limiting the analysis to only the alignable portion of the reference genome. Nevertheless, as demonstrated by these results, methods using concatemers can distinguish the plasma of mothers carrying a male fetus from that of males or non-pregnant women. These methods are likely to be used to distinguish the plasma of mothers carrying a male fetus from that of mothers carrying a female fetus.
[0111] VI. Further Embodiments Embodiment 1 includes a method including receiving a plurality of DNA fragments; concatemerizing a first set of the plurality of DNA fragments to obtain first concatemers; passing the first concatemers through the first nanopore; and detecting a first electrical signal as the first concatemers pass through the first nanopore, wherein the electrical signal corresponds to the first sequence of the first concatemers.
[0112] Embodiment 2 includes the method of embodiment 1, further including performing, by a computer system, analyzing the first electrical signal to determine the first sequence and aligning subsequences of the first sequence to a reference genome to identify fragment sequences corresponding to each of the first set of DNA fragments.
[0113] Embodiment 3 includes the method of embodiment 1, wherein passing the first concatemer through the first nanopore includes passing a first strand of the first concatemer through the nanopore, followed by passing a second strand of the first concatemer through the nanopore.
[0114] Example 4 includes the method of Example 3, further including analyzing the first electrical signals for the first strand and the second strand with a computer system to determine the first sequence.
[0115] Embodiment 5 includes the method of embodiment 1, wherein the plurality of DNA fragments are cell-free DNA fragments from a biological sample.
[0116] Embodiment 6 includes the method of embodiment 2, wherein the biological sample is plasma or serum.
[0117] Example 7 includes the method of Example 1, wherein the first nanopore is one of a plurality of nanopores on the substrate.
[0118] Example 8 includes the method of Example 7, further including concatemerizing a second set of the DNA fragments to obtain second concatemers; passing the second concatemers through a second nanopore of the plurality of nanopores; and detecting an electrical signal as the second concatemers pass through the second nanopore, the electrical signal corresponding to a second sequence of the second concatemers.
[0119] Example 9 includes the method of Example 7, further including concatemerizing a second set of the DNA fragments to obtain second concatemers; passing the second concatemers through the first nanopore; and detecting an electrical signal as the second concatemers pass through the first nanopore, the electrical signal corresponding to a second sequence of the second concatemers.
[0120] Embodiment 10 includes a method including receiving a biological sample containing a plurality of DNA fragments, concentrating the biological sample to obtain a higher concentration of the DNA fragments, passing the plurality of DNA fragments through a nanopore on a substrate, and detecting, for each of the plurality of DNA fragments, an electrical signal as the DNA fragment passes through the nanopore, the electrical signal corresponding to the sequence of the DNA fragment.
[0121] Embodiment 11 includes a method, implemented by a computer system, comprising receiving a first sequence of first concatemers generated by concatemerizing a first set of DNA fragments and aligning subsequences of the first sequence to a reference genome to identify fragment sequences corresponding to each of the first set of DNA fragments. Embodiment 12 includes the method of Embodiment 11, wherein the aligning the subsequences includes aligning windows of the first sequence to the reference genome and identifying when two windows align to different regions of the reference genome. Embodiment 13 includes the method of Embodiment 12, wherein one or more windows between the two windows are identified when they are not aligned to the reference genome.
[0122] Embodiment 14 includes a computer product including a computer-readable medium storing a plurality of instructions for controlling a computer system to perform the operations of any one of embodiments 11 to 13. Embodiment 15 includes a system including the computer product of embodiment 14 and one or more processors for executing the instructions stored on the computer-readable medium. Embodiment 16 includes a system including means for performing any one of embodiments 11 to 13. Embodiment 17 includes a system configured to perform any one of embodiments 11 to 13. Embodiment 18 includes a system including modules for performing the steps of any one of embodiments 11 to 13, respectively.
[0123] VII. Example Sequencing Systems FIG. 15 shows a block diagram of a system 1500 for performing single-molecule sequencing according to an embodiment of the present invention. A biological sample can be obtained from a patient 1502 by an extraction device 1504. The biological sample can be any bodily fluid or any biological sample described herein. The extraction device 1504 can include a syringe, lancet, swab, or container or vial for collecting a sample, such as urine. The extraction device 1504 can include a QIAamp DSP DNA Blood Mini Kit (Qiagen, Germany) (2). The biological sample can contain cell-free DNA fragments, which are sent to a preparation device 1506. The preparation device 1506 can include a device that produces a form of cell-free DNA fragments that is more efficient for single-molecule sequencing. For example, the output of the preparation device 1506 can be a concentrated sample of concatemers or cell-free DNA fragments.
[0124] If concatemers are generated, the preparation device 1506 may include a vacuum dryer to increase concentration, such as a SpeedVac concentrator (Thermo Scientific) or an ultraconcentrator, which removes fluid by percolation or filtration. For concatemers, the concentration may not increase by more than fivefold. The preparation device 1506 may include a module for repairing the ends of DNA fragments, such as the NEBNext End Repair Module (New England Biolabs, NEB). Additionally, the preparation device 1506 may include a kit for purifying end-repaired DNA, such as the MinElute Reaction Cleanup Kit (Qiagen). The preparation device 1506 may also include an incubator capable of incubating a ligase enzyme (e.g., Blunt / TA Ligase Master Mix (NEB)) for several hours at a certain temperature (e.g., 25°C for 4 hours). The preparation device 1506 may include an additional kit for purifying the concatemers after the reaction, which may include the MinElute Reaction Cleanup Kit (Qiagen). The preparation device 1506 can also include a robotic liquid handler that can mix and transfer fluids.
[0125] If increased concentrations of cell-free DNA fragments are produced, the preparation device 1506 can include a vacuum dryer (e.g., a Speedvac concentrator (Thermo Fisher Scientific, Waltham, MA)) or an ultraconcentrator that removes fluids by percolation or filtration. The preparation device 1506 can also include an end repair and A-tailing module (New England Biolabs, Ipswich, MA). The preparation device 1506 can also include a robotic liquid handler that can mix and transfer fluids.
[0126] The concatemers or concentrated DNA fragments from the preparation device 1506 may be sent to a single molecule sequencing device 1508 to obtain sequence reads. The single molecule sequencing device 1508 may also include devices such as the MinION Flow Cell (v7.3) (Nanopore), SMRT Technology (Pacific Biosciences), or the Helicos Single Molecule Sequencing Device (SeqLL). The single molecule sequencing device 1508 may also include nanopore-related kits such as the Genomic DNA Sequencing Kit (SQK-MAP-005, Oxford Nanopore Technologies, UK) or the Nanopore Sequencing Kit (SQK-NSK007, Oxford Nanopore Technologies).
[0127] The single molecule sequencer 1508 can output sequence reads from the concatemers or enriched DNA fragments. The sequence reads can be output as a data file and analyzed by a computer system 1510. The computer system can be a specialized computer system equipped with software for analyzing sequence reads. The output data file can be base-called using METRICHOR™ software (Nanopore). 2D reads can be extracted and aligned using LAST software. The computer system can be the computer system 10 of FIG. 16, as described below.
[0128] VIII. Computer Systems Any of the computer systems referred to herein may utilize any suitable number of subsystems. Examples of such subsystems in computer system 10 are shown in FIG. 16. In some embodiments, a computer system includes a single computer device, and the subsystems may be components of the computer device. In other embodiments, a computer system may include multiple computer devices, each of which is a subsystem with internal components. Computer systems may include desktop and laptop computers, tablets, mobile phones, and other mobile devices.
[0129] The subsystems shown in FIG. 6 are interconnected via a system bus 75. Additional subsystems are shown, such as a printer 74, a keyboard 78, a storage device 79, and a monitor 76 coupled to a display adapter 82. Peripherals and input / output (I / O) devices coupled to the I / O controller 71 can be connected to the computer system by any number of means known in the art, such as an input / output (I / O) port 77 (e.g., USB, FireWire®). For example, the I / O port 77 or an external interface 81 (e.g., Ethernet, Wi-Fi, etc.) can be used to connect the computer system 10 to a wide area network, such as the Internet, a mouse input device, or a scanner. The interconnection via the system bus 75 allows the central processor 73 to communicate with each subsystem and control the execution of instructions from the system memory 72 or storage device 79 (e.g., a hard drive or fixed disk such as an optical disk), as well as the exchange of information between the subsystems. The system memory 72 and / or storage device 79 can embody computer-readable media. Another subsystem is a data collection device 85, such as a camera, microphone, or accelerometer. Any of the data described herein can be output from one component to another, or to a user.
[0130] A computer system may include multiple identical components or subsystems connected to each other, for example, by an external interface 81, by an internal interface, or through a removable storage device that can be connected and disconnected from one component to another. In some embodiments, computer systems, subsystems, or devices may communicate over a network. In such cases, one computer may be considered a client and another computer may be a server, each part of the same computer system. The client and server may each include multiple systems, subsystems, or components.
[0131] Aspects of the embodiments may be implemented in the form of control logic using hardware (e.g., application specific integrated circuits or field programmable gate arrays) and / or using computer software with a generally programmable processor in a modular or integrated manner. As used herein, a processor includes a single-core processor, a multi-core processor on the same integrated chip, or multiple processing units on a single circuit board, or networked. Based on the disclosure and teachings provided herein, those skilled in the art will understand and appreciate other ways and / or methods of implementing embodiments of the present invention using hardware and combinations of hardware and software.
[0132] Any of the software components or functions described in this application may be implemented as software code executed by a processor using any suitable computer language, such as Java, C, C++, C#, Objective-C, Swift, or a scripting language using conventional or object-oriented techniques, such as Perl or Python. The software code may be stored as a series of instructions or commands on a computer-readable medium for storage and / or transmission. Suitable non-transitory computer-readable media may include random access memory (RAM), read-only memory (ROM), magnetic media such as a hard drive or floppy disk, or optical media such as a compact disc (CD) or DVD (digital versatile disc), flash memory, etc. The computer-readable medium may also be any combination of such storage or transmission devices.
[0133] Such programs may be encoded and transmitted using carrier signals suitable for transmission over wired, optical, and / or wireless networks conforming to various protocols, including the Internet. In this manner, computer-readable media may be created using data signals encoded with such programs. Computer-readable media encoded with program code may be packaged with a compatible device or provided separately from other devices (e.g., via Internet download). Such computer-readable media may reside on or within a single computer product (e.g., a hard drive, CD, or an entire computer system) or may reside on or within different computer products within a system or network. A computer system may include a monitor, printer, or other suitable display for providing a user with any of the results discussed herein.
[0134] Any of the methods described herein can be performed in whole or in part using a computer system including one or more processors that can be configured to perform the steps. Accordingly, embodiments can be directed to a computer system configured to perform any of the steps of the methods described herein, possibly using different components that perform each step or each group of steps. Although presented as numbered steps, steps of the methods herein can be performed simultaneously or in a different order. Furthermore, some of these steps can be used with some of the other steps of other methods. Also, all or some of the steps may be optional. Furthermore, any steps of the methods can be performed using a module, unit, circuit, or other means for performing those steps.
[0135] The specific details of particular embodiments may be combined in any suitable manner without departing from the spirit and scope of the embodiments of the invention. However, other embodiments of the invention may be directed to particular embodiments relating to each individual aspect or particular combinations of these individual aspects.
[0136] The foregoing description of exemplary embodiments of the invention has been presented for purposes of illustration and description. It is not intended to be exhaustive or to limit the invention to the precise form described, and many modifications and variations are possible in light of the above teaching.
[0137] The enumeration of "a," "an," or "the" is intended to mean "one or more" unless otherwise specified. The use of "or" is intended to mean "inclusive," not "exclusive," unless specifically indicated to the contrary. A reference to a "first" element does not necessarily require that a second element be provided. Furthermore, a reference to a "first" or "second" element does not limit the referenced element to a particular location unless explicitly stated.
[0138] All patents, patent applications, publications and descriptions mentioned herein are incorporated by reference in their entirety for all purposes. None is admitted to be prior art.
[0139] IX. References: 1. Dondorp W, de Wert G, Bombard Y, Bianchi DW, Bergmann C, Borry P, et al. Non-invasive prenatal testing for aneuploidy and beyond: Challenges of responsible innovation in prenatal screening. Eur J Hum Genet 2015. 2. Chiu RWK, Chan KCA, Gao Y, Lau VYM, Zheng W, Leung TY, et al. Noninvasive prenatal diagnosis of fetal chromosomal aneuploidy by massively parallel genomic sequencing of dna in maternal plasma. Proc Natl Acad Sci USA 2008; 105: 20458-63. 3. Jain M, Fiddes IT, Miga KH, Olsen HE, Paten B, Akeson M. Improved data analysis for the MinION nanopore sequencer. Nat Methods 2015; 12: 351-6. 4. Lo YMD, Chan KCA, Sun H, Chen EZ, Jiang P, Lun FMF, et al. Maternal plasma DNA sequencing reveals the genome-wide genetic and mutational profile of the fetus. Sci Transl Med 2010; 2: 61ra91-61ra91. 5. Bayley H. Nanopore sequencing: From imagination to reality. Clin Chem 2015; 61: 25-31. 6. Zheng YW, Chan KCA, Sun H, Jiang P, Su X, Chen EZ, et al. Nonhematopoietically derived DNA is shorter than hematopoietically derived DNA in plasma: A transplantation model. Clin Chem 2012; 58: 549-58. 7. Lun FMF, Chiu RWK, Sun K, Leung TY, Jiang P, Chan KCA, et al. Noninvasive prenatal methylomic analysis by genomewide bisulfite sequencing of maternal plasma DNA. Clin Chem 2013; 59: 1583-94. 8. Chan KCA, Jiang P, Zheng YW, Liao GJW, Sun H, Wong J, et al. Cancer genome scanning in plasma: Detection of tumor-associated copy number aberrations, single-nucleotide variants, and tumoral heterogeneity by massively parallel sequencing. Clin Chem 2013; 59: 211-24. 9. Chan KCA, Jiang P, Chan CWM, Sun K, Wong J, Hui EP, et al. Noninvasive detection of cancer-associated genome-wide hypomethylation and copy number aberrations by plasma DNA bisulfite sequencing. Proc Natl Acad Sci U S A 2013; 110: 18761-8. 10. Chan RWY, Wong J, Chan HLY, Mok TSK, Lo WYW, Lee V, et al. Aberrant concentrations of liver-derived plasma albumin mrna in liver pathologies. Clin Chem 2010; 56: 82-9. 11. Chan RWY, Jiang P, Peng X, Tam LS, Liao GJW, Li EK, et al. Plasma DNA aberrations in systemic lupus erythematosus revealed by genomic and methylomic sequencing. Proc Natl Acad Sci U S A 2014; 111: E5302-11. 12. Schreiber J, Wescoe ZL, Abu-Shumays R, Vivian JT, Baatar B, Karplus K, Akeson M. Error rates for nanopore discrimination among cytosine, methylcytosine, and hydroxymethylcytosine along individual DNA strands. Proc Natl Acad Sci U S A 2013; 110: 18910-5.
Claims
1. extracting a plurality of cell-free DNA fragments from a biological sample obtained from a subject, wherein the biological sample comprises a bodily fluid containing the cell-free DNA fragments; concatemerizing the set of cell-free DNA fragments to obtain concatemers, wherein the set of cell-free DNA fragments have different sequences and originate from different locations in the genome of the subject; passing the concatemers through a nanopore of a sequencing device; detecting an electrical signal as the concatemer passes through the nanopore, wherein the electrical signal corresponds to the sequence of the concatemer; analyzing the electrical signals to determine the methylation class of the concatemers; A method comprising:
2. 2. The method of claim 1, wherein the bases of the concatemer indicate whether they are hypermethylated or hypomethylated relative to a reference genome.
3. The method of claim 1 , wherein the methylation classification is a methylation classification for a CpG site of the concatemer.
4. 4. The method of claim 3, wherein the methylation classification indicates whether the CpG sites of the concatemer are hypermethylated or hypomethylated relative to a reference genome.
5. 2. The method of claim 1, wherein the methylation classification indicates whether the concatemer is hydroxymethylated or not.
6. 10. The method of claim 1, further comprising analyzing the electrical signal to determine the sequence of the concatemer.
7. 7. The method of claim 6, further comprising aligning a subsequence of the sequence to a reference genome to identify a fragment sequence corresponding to each DNA fragment of the set of multiple cell-free DNA fragments.
8. The method of claim 7, further comprising determining a size of each of the plurality of sets of cell-free DNA fragments based on the alignment of the subsequences.
9. Concatamerizing the set of cell-free DNA fragments includes concatemerizing the set of cell-free DNA fragments and a set of DNA fragments having known sequences, wherein the set of DNA fragments having known sequences is dispersed among the set of cell-free DNA fragments; 8. The method of claim 7, wherein aligning subsequences of the known sequence comprises aligning subsequences to the known sequence to identify the location of a set of DNA fragments having the known sequence in the concatemer.
10. 10. The method of claim 9, wherein the concatemerization places one DNA fragment of the set of DNA fragments having a known sequence between two cell-free DNA fragments of the set of multiple cell-free DNA fragments.
11. Identifying a starting base of a first subsequence corresponding to a first DNA fragment of the set of multiple cell-free DNA fragments based on identifying one of the known sequences of the set of DNA fragments having the known sequence prior to the first subsequence; Identifying an end base of the first subsequence corresponding to the first DNA fragment of the set of cell-free DNA fragments based on identifying one of the known sequences of the set of DNA fragments having the known sequence after the first subsequence; and The method of claim 10 further comprising:
12. 10. The method of claim 9, wherein each DNA fragment of the set of DNA fragments with known sequences has a known sequence of 7 nucleotides or less.
13. 10. The method of claim 9, wherein each DNA fragment of the set of DNA fragments with known sequences has the same known sequence.
14. The alignment of the subsequences is aligning sliding windows of the sequence to the reference genome, wherein each sliding window corresponds to a subsequence aligned to the reference genome; Distinguishing sequences of two DNA fragments in the set of multiple DNA fragments when two sliding windows are aligned to different regions of the reference genome; The method of claim 7, comprising:
15. 15. The method of claim 14, further comprising determining a first subsequence of an end and a start of a first DNA fragment of the set of multiple cell-free DNA fragments based on two sliding windows that align to the different regions of the reference genome.
16. 15. The method of claim 14, wherein the different regions of the reference genome comprise regions on different chromosomes.
17. 15. The method of claim 14, wherein one or more windows between the two sliding windows are identified as not aligned to the reference genome.
18. determining a size of a first DNA fragment of the set of cell-free DNA fragments; 18. The method of claim 17, wherein determining the size of the first DNA fragment comprises determining the length of a longest subsequence that aligns to a region of the reference genome.
19. passing the concatemers through a nanopore; passing a first strand of the concatemer through a nanopore; then passing a second strand of the concatemer through a nanopore; The method of claim 1 , comprising:
20. 20. The method of claim 19, wherein analyzing the electrical signals, the methylation classification comprises analyzing the electrical signals of the first strand and the second strand.
21. 10. The method of claim 1, further comprising determining the fetal DNA fraction based on the methylation classification.
22. 10. The method of claim 1, further comprising detecting based on a methylation classification, a pregnancy-associated disease, an abnormal pregnancy methylation profile, aberrant methylation associated with cancer, or systemic lupus erythematosus.
23. 10. The method of claim 1, wherein the biological sample is plasma or serum.
24. A computer readable medium storing a plurality of instructions which, when executed by a processor, control a system for performing the method of any one of claims 1 to 23.
25. The computer-readable medium of claim 24; and one or more processors for executing instructions stored on said computer-readable medium; Including, the system.
26. an extraction device configured to perform extraction of a plurality of cell-free DNA fragments from a biological sample; a preparation device configured to perform concatemerization of a set of multiple cell-free DNA fragments; and Sequencing equipment, 26. The system of claim 25, further comprising:
Citation Information
Patent Citations
High-Resolution Analysis Devices and Related Methods
US20130092541A1
Mutational analysis of plasma DNA for cancer detection
WO2013190441A2