Reagents and methods for analyzing associated nucleic acids
Through the associated fragment method, the nucleic acid fragments in a single microparticle are linked together and sequenced, solving the problem of detecting remote genetic information and distinguishing fetal cfDNA from maternal DNA in the prior art, and achieving high sensitivity and accuracy analysis results.
Patent Information
- Application Number
- CN202111280417.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2016-12-23
- Filing Date
- 2017-12-19
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2037-12-19
AI Technical Summary
The prior art is difficult to reliably detect remote genetic information and distinguish fetal cfDNA from maternal DNA, especially in non-invasive prenatal detection and cancer diagnosis.
The nucleic acid fragments from individual particles are linked together by using the associated fragment method to generate a related sequence read group, and these sequence reads are sequenced to analyze the nucleic acid fragments in the circulating particles.
A highly sensitive cfDNA analysis is achieved, able to detect remote genetic information, and improve the ability to distinguish fetal cfDNA from maternal DNA, enhancing the accuracy of non-invasive prenatal detection and cancer diagnosis.
Smart Images

Figure CN114150041B_ABST
Abstract
Description
[0001] This application is a divisional application of the Chinese patent application with the application number 201780087136.2 and the invention title "Reagents and Methods for Analyzing Associated Nucleic Acids". The original application is the application that entered the Chinese national phase on August 21, 2019, of the PCT international application PCT / GB2017 / 053820 filed on December 19, 2017. Technical Field
[0002] The present invention relates to the analysis of cell-free nucleic acids (such as cell-free DNA). In particular, it relates to the analysis of cell-free DNA contained within microparticles derived from blood. Reagents and methods are provided for linking the nucleic acids of individual microparticles. Methods are also provided for analyzing linked nucleic acid fragment sets from individual microparticles. Background Art
[0003] Cell-free DNA (cfDNA) in circulation is typically fragmented (usually with a length of 100 - 200 base pairs), and thus methods for cfDNA analysis have traditionally focused on the biological signals that can be detected using these short DNA fragments. For example, detecting single nucleotide variants within individual molecules or "molecular counting" of large numbers of sequenced fragments to indirectly infer the presence of large-scale chromosomal abnormalities, such as tests for fetal chromosomal trisomies using fetal DNA in maternal circulation (in the form of so-called "non-invasive prenatal testing" or NIPT).
[0004] Numerous methods for analyzing circulating cell-free DNA have been previously described. Depending on the specific application area, these assays may use different terms for broadly similar groups of sample types and technical methods, such as circulating tumor DNA (ctDNA), cell-free foetal DNA (cffDNA), and / or liquid biopsy, or non-invasive prenatal testing. Generally speaking, these methods include laboratory protocols for preparing a sample of circulating cell-free DNA for sequencing, the sequencing reaction itself, and subsequently an information framework for analyzing the resulting sequences to detect relevant biological signals. The method involves DNA purification and isolation steps prior to sequencing, which means that subsequent analysis must rely solely on the information contained within the DNA itself. After sequencing, such methods typically use one or more information or statistical frameworks to analyze multiple aspects of the sequence data, such as detecting specific mutations therein, and / or detecting selective enrichment or selective deletion of specific chromosomal or sub-chromosomal regions (e.g., which may indicate chromosomal aneuploidy in a developing fetus).
[0005] Many of these methods are used for NIPT (e.g., in U.S. Patents 6258540 B1, 8296076 B2, 8318430 B2, 8195415 B2, 9447453 B2, and 8442774 B2). The most commonly used methods for performing non-invasive prenatal testing for detecting fetal chromosomal abnormalities (e.g., trisomies and / or subchromosomal abnormalities, such as microdeletions) involve sequencing a large number of cfDNA molecules, mapping the resulting sequences to the genome (i.e., determining which chromosome and / or which part of a given chromosome the sequences are from), and subsequently, for one or more such chromosomal or subchromosomal regions, determining the amount of sequences mapped thereto (e.g., in the form of the absolute number of reads or the relative number of reads), and subsequently comparing it with one or more normal or abnormal thresholds or cut-off values, and / or performing statistical tests to determine whether the region is likely to be overexpressed in terms of sequence amount (which may, for example, correspond to chromosomal trisomy) and / or whether the region is likely to be underexpressed in terms of sequence amount (which may, for example, correspond to a microdeletion).
[0006] A variety of additional or modified methods for analyzing cell-free DNA using data from unlinked individual molecules are also described (e.g., WO2016094853 A1, US2015344970 A1, and US20150105267 A1).
[0007] Despite such a wide range of methods, there remains a need for new cfDNA analysis methods that allow for reliable detection of remote genetic information (e.g., phasing) and methods with higher sensitivity. For example, in the case of NIPT, fetal cfDNA represents only a small fraction of the overall cfDNA in a pregnant individual (most of the circulating DNA is normal maternal DNA). Thus, a significant technical challenge in NIPT revolves around distinguishing fetal cfDNA from maternal DNA. Similarly, in patients with cancer, cfDNA represents only a small fraction of the overall circulating DNA. Thus, there are similar technical challenges in using cfDNA analysis for diagnosing or monitoring cancer. SUMMARY OF THE INVENTION
[0008] The present invention provides methods for analyzing nucleic acid fragments in circulating microparticles (or microparticles derived from blood). The present invention is based on a linked-fragment approach, in which nucleic acid fragments from a single microparticle are linked together. This linkage enables the generation of a set of linked sequence reads corresponding to the sequences of the fragments from a single microparticle.
[0009] The linked-fragment method provides highly sensitive cfDNA analysis and also enables the detection of remote genetic information. The method is based on a combination of insights. First, these methods utilize the insight that individual circulating microparticles (e.g., individual circulating apoptotic bodies) will contain many genomic DNA fragments produced by the same somatic cells (somewhere in the body) undergoing apoptosis. Second, a portion of such genomic DNA fragments within an individual microparticle will preferentially contain sequences from one or more specific chromosomal regions. Cumulatively, such circulating microparticles thus serve as data-rich and multi-featured "molecular stethoscopes" to observe the potentially very complex genetic events occurring in a limited somatic tissue space somewhere in the body; importantly, since such microparticles mostly enter the circulation before being cleared or metabolized, they can be detected non-invasively. The present invention describes experimental and informational methods for using these "stethoscopes" - i.e., groups of linked fragments and linked sequence reads (in the form of a single individual microparticle or, in many embodiments, a complex sample containing a large number of individual circulating microparticles) to perform analytical and diagnostic tasks.
[0010] The present invention provides a method for analyzing a sample comprising microparticles derived from blood, wherein the microparticles comprise at least two target nucleic acid fragments, and wherein the method comprises: (a) preparing a sample for sequencing, which comprises linking at least two of the at least two target nucleic acid fragments to produce a group of at least two linked target nucleic acid fragments; and (b) sequencing each linked fragment in the group to produce at least two (informationally) linked sequence reads.
[0011] The present invention provides a method for analyzing a sample comprising circulating microparticles, wherein the circulating microparticles comprise at least two target nucleic acid fragments, and wherein the method comprises: (a) preparing a sample for sequencing, which comprises linking at least two of the at least two target nucleic acid fragments to produce a group of at least two linked target nucleic acid fragments; and (b) sequencing each linked fragment in the group to produce at least two (informationally) linked sequence reads.
[0012] The present invention provides a method for analyzing a sample comprising microparticles derived from blood, wherein the microparticles comprise at least two genomic DNA fragments, and wherein the method comprises: (a) preparing a sample for sequencing, which comprises linking at least two of the at least two genomic DNA fragments to produce a group of at least two linked genomic DNA fragments; and (b) sequencing each linked fragment in the group to produce at least two linked sequence reads.
[0013] The present invention provides a method for analyzing a sample containing circulating particles, wherein the circulating particles contain at least two genomic DNA fragments, and wherein the method comprises: (a) preparing a sample for sequencing, which includes linking at least two of the at least two genomic DNA fragments to produce a set of at least two linked genomic DNA fragments; and (b) sequencing each linked fragment in the set to produce at least two linked sequence reads.
[0014] In the method, at least 3, at least 4, at least 5, at least 10, at least 50, at least 100, at least 500, at least 1000, at least 5000, at least 10,000, at least 100,000, or at least 1,000,000 target nucleic acid fragments of the particles can be linked into a set and then sequenced to produce at least 3, at least 4, at least 5, at least 10, at least 50, at least 100, at least 500, at least 1000, at least 5000, at least 10,000, at least 100,000, or at least 1,000,000 linked sequence reads.
[0015] Preferably, at least 5 target nucleic acid fragments of the particles can be linked into a set and then sequenced to produce at least 5 linked sequence reads.
[0016] In the method, each linked sequence read can provide the sequence of at least 1 nucleotide, at least 5 nucleotides, at least 10 nucleotides, at least 20 nucleotides, at least 30 nucleotides, at least 50 nucleotides, at least 100 nucleotides, at least 200 nucleotides, at least 500 nucleotides, at least 1000 nucleotides, or at least 10,000 nucleotides of the linked fragment. Preferably, each linked sequence read can provide the sequence of at least 20 nucleotides of the linked fragment.
[0017] In the method, a total of at least 2, at least 10, at least 100, at least 1000, at least 10,000, at least 100,000, at least 1,000,000, at least 10,000,000, at least 100,000,000, at least 1,000,000,000, at least 10,000,000,000, at least 100,000,000,000, or at least 1,000,000,000,000 sequence reads can be produced. Preferably, a total of at least 500,000 sequence reads are produced.
[0018] A sequence read may comprise at least 5, at least 10, at least 25, at least 50, at least 100, at least 250, at least 500, at least 1000, at least 2000, at least 5000, or at least 10,000 nucleotides from a target nucleic acid (e.g., genomic DNA). Preferably, each sequence read comprises at least 5 nucleotides from the target nucleic acid.
[0019] A sequence read may comprise the raw sequence read generated by a sequencer or a portion thereof, e.g., the raw sequence read of a 50-nucleotide-long sequence generated by an Illumina sequencer. A sequence read may comprise a fused sequence of two reads from a paired-end sequencing run, e.g., a concatenated or fused sequence of both the first and second reads from a paired-end sequencing run on an Illumina sequencer. A sequence read may comprise a portion of the raw sequence read generated by a sequencer, e.g., 20 consecutive nucleotides within a 150-nucleotide raw sequence read generated by an Illumina sequencer. A single raw sequence read may comprise at least two associated sequence reads generated by the methods of the present invention.
[0020] Sequence reads may be generated by any method known in the art. For example, by chain termination or Sanger sequencing. Preferably, sequencing is performed by next-generation sequencing methods such as sequencing by synthesis, sequencing by synthesis using reversible terminators (e.g., Illumina sequencing), pyrosequencing (e.g., 454 sequencing), sequencing by ligation (e.g., SOLiD sequencing), single molecule sequencing (e.g., single molecule real-time (SMRT) sequencing, Pacific Biosciences), or by nanopore sequencing (e.g., on the Minion or Promethion platforms, Oxford Nanopore Technologies). Most preferably, sequence reads are generated by sequencing by synthesis using reversible terminators (e.g., Illumina sequencing).
[0021] The method may include an additional step of mapping each associated sequence read to a reference genomic sequence. Associated sequence reads may comprise sequences mapped to the same chromosome of the reference genomic sequence or sequences mapped to two or more different chromosomes of the reference genomic sequence.
[0022] The diameter of the microparticles can be at least 100 nm, at least 110 nm, at least 125 nm, at least 150 nm, at least 175 nm, at least 200 nm, at least 250 nm or at least 500 nm. Preferably, the diameter of the microparticles is at least 200 nm. The diameter of the microparticles can be 100 - 5000 nm. The diameter of the microparticles can be 10 - 10,000 nm (such as 100 - 10,000 nm, 110 - 10,000 nm), 50 - 5000 nm, 75 - 5,000 nm, 100 - 3,000 nm. The diameter of the microparticles can be 10 - 90 nm, 50 - 100 nm, 90 - 200 nm, 100 - 200 nm, 100 - 500 nm, 100 - 1000 nm, 1000 - 2000 nm, 90 - 5000 nm, or 2000 - 10,000 nm. Preferably, the diameter of the microparticles is 100 to 5000 nm. Most preferably, the diameter of the microparticles is 200 to 5000 nm. The sample can include at least two different sizes, or at least three different sizes or a series of different sizes of microparticles.
[0023] The associated genomic DNA fragment can be derived from a single genomic DNA molecule.
[0024] The method can further include the step of estimating or determining the genomic sequence length of the associated genomic DNA fragment. Optionally, this step can be carried out by sequencing substantially the entire sequence of the associated fragment (i.e., from its approximate 5' end to its approximate 3' end) and counting the number of nucleotides sequenced therein. Optionally, this can be done by: sequencing a sufficient number of nucleotides at the 5' end of the sequence of the associated fragment to map the 5' end to a locus within a reference genomic sequence (such as the human genomic sequence), and similarly sequencing a sufficient number of nucleotides at the 3' end of the sequence of the associated fragment to map the 3' end to a locus within the reference genomic sequence, and then using the reference genomic sequence to determine the genomic sequence length of the associated fragment (i.e., the number of nucleotides sequenced at the 3' end of the associated fragment + the number of nucleotides sequenced at the 5' end of the associated fragment + the number of nucleotides between these sequences in the reference genome (i.e., the unsequenced portion)).
[0025] Preferably, the sample is isolated from blood, plasma or serum. The microparticles can be isolated from blood, plasma or serum. The method can further include the step of isolating the microparticles from blood, plasma or serum. This step can be carried out before or during step (a).
[0026] The microparticles can be isolated by centrifugation, size exclusion chromatography and / or filtration.
[0027] The separation step may include centrifugation. Particles may be separated by precipitation, which utilizes a centrifugation step and / or an ultracentrifugation step, or a series of two or more centrifugation steps and / or ultracentrifugation steps at two or more different speeds, wherein the pellet and / or supernatant from one centrifugation / ultracentrifugation step is further processed in a second centrifugation / ultracentrifugation step and / or differential centrifugation.
[0028] The centrifugation or ultracentrifugation step may be carried out at a speed of 100 - 500,000 G, 100 - 1000 G, 1000 - 10,000 G, 10,000 - 100,000 G, 500 - 100,000 G, or 100,000 - 500,000 G. The centrifugation or ultracentrifugation step may be carried out for a duration of at least 5 seconds, at least 10 seconds, at least 30 seconds, at least 60 seconds, at least 5 minutes, at least 10 minutes, at least 30 minutes, at least 60 minutes, or at least 3 hours.
[0029] The separation step may include size - exclusion chromatography, such as column - based size - exclusion chromatography, such as size - exclusion chromatography including columns with an agarose - based matrix or a sephacryl - based matrix.
[0030] The size - exclusion chromatography may include using a matrix or filter having the following pore sizes: a size or diameter of at least 50 nm, at least 100 nm, at least 200 nm, at least 500 nm, at least 1.0 μm, at least 2.0 μm, or at least 5.0 μm.
[0031] The separation step may include filtering the sample. The filtrate may provide the particles to be analyzed in the method. Optionally, a filter is used to separate particles below a certain size, and wherein the filter preferentially or completely removes particles having a size greater than 100 nm, size greater than 200 nm, size greater than 300 nm, size greater than 500 nm, size greater than 1.0 μm, size greater than 2.0 μm, size greater than 3.0 μm, size greater than 5.0 μm, or size greater than 10.0 μm. Optionally, two or more such filtration steps may be carried out using filters with the same size filtration parameters or different size filtration parameters. Optionally, the filtrate from one or more filtration steps contains particles, and associated sequence reads are generated therefrom.
[0032] In the method, the sample may contain first and second particles derived from blood, wherein each particle contains at least two target nucleic acid (e.g., genomic DNA) fragments, and wherein the method includes performing step (a) to generate a first set of associated target nucleic acid fragments of the first particles and a second set of associated target nucleic acid fragments of the second particles, and performing step (b) to generate a first set of associated sequence reads of the first particles and a second set of associated sequence reads of the second particles.
[0033] In the method, the associated sequence read sets generated for the first microparticles can be distinguished from the associated sequence read sets generated for the second microparticles.
[0034] In the method, the sample can comprise n microparticles derived from blood, wherein each microparticle comprises at least two target nucleic acid (e.g., genomic DNA) fragments, and wherein the method comprises performing step (a) to generate n sets of associated target nucleic acid fragments, one set for each of the n microparticles, and performing step (b) to generate n sets of associated sequence reads, one set for each of the n microparticles.
[0035] In the method, n can be at least 3, at least 5, at least 10, at least 50, at least 100, at least 1000, at least 10,000, at least 100,000, at least 1,000,000, at least 10,000,000, at least 100,000,000, at least 1,000,000,000, at least 10,000,000,000, or at least 100,000,000,000. Preferably, n is at least 100,000 microparticles.
[0036] In the method, the nucleic acid sample can comprise at least 3, at least 5, at least 10, at least 50, at least 100, at least 1000, at least 10,000, at least 100,000, at least 1,000,000, at least 10,000,000, at least 100,000,000, at least 1,000,000,000, at least 10,000,000,000, or at least 100,000,000,000 microparticles, wherein during any step of the method, such as any step of contacting the sample with a library of multimeric barcoding reagents, and / or any step of appending a barcode sequence to a target nucleic acid, and / or any step of appending a coupling sequence to a target nucleic acid, and / or any step of crosslinking or permeabilizing, the microparticles are contained within a single continuous aqueous volume.
[0037] The associated sequence read sets generated for each microparticle can be distinguished from the associated sequence read sets generated for other microparticles.
[0038] Before step (a), the method can further comprise the step of partitioning the sample into at least two different reaction volumes.
[0039] In the present invention, two sequences or sequence reads (e.g., determined by a sequencing reaction) can be informationally linked by any means that allows such sequences to be related or associated with each other in any way within a computer system, within an algorithm, or within a data set. Such a linkage can comprise, be determined by, or be represented by: a discrete identifying linkage, or a shared attribute, or any indirect method that links, correlates, or relates two or more such sequences.
[0040] The linkage can comprise, and / or be determined by, and / or be represented by: sequences within the sequencing reaction itself (e.g., in the form of barcode sequences determined by the sequencing reaction, or in the form of two different portions or segments of a single determined sequence that together comprise a first and a second linked sequence), or independent of such sequence determination, comprise or represent (e.g., by being contained within the same flow cell, or the same lane of a flow cell, or the same chamber or region of a sequencer, or being contained within the same sequencing run of a sequencer, or being contained within a biological sample with a degree of spatial proximity, and / or with a degree of spatial proximity within a sequencer or sequencing flow cell). The linkage can comprise, and / or be determined by, and / or be represented by: a measure or parameter corresponding to a physical location or partition within a sequencer, such as a pixel or pixel location within an image and / or a multi-pixel camera or multi-pixel charge-coupled device, and / or a nanopore or the location of a nanopore within a nanopore sequencer or nanopore membrane.
[0041] The linkage can be absolute (i.e., two sequences are linked or not linked, with no quantitative, semi-quantitative, or qualitative / categorical relationship beyond this). The linkage can also be relative, probabilistic, or determine, comprise, or represent in terms of a degree, probability, or extent of linkage, e.g., relative to one or more parameters (or represented by them) that can have one of a range of quantitative, semi-quantitative, or qualitative / categorical values. For example, two (or more) sequences can be informationally linked by a quantitative, semi-quantitative, or qualitative / categorical parameter that represents, comprises, estimates, or embodies the proximity of the two (or more) sequences within a sequencer, or the proximity of the two (or more) sequences within a biological sample.
[0042] For any analysis involving two or more sequences that are informationally linked by any such means, the presence (or absence) of the linkage can be used as a parameter in any analysis or evaluation step or in any algorithm used for its performance. For any analysis involving two or more sequences that are informationally linked by any such means, the degree, probability, or extent of the linkage can be used as a parameter in any analysis or evaluation step or in any algorithm used for its performance.
[0043] In one form of such association, two or more associated sequences of a given set may be associated with a specific identifier (e.g., an alphanumeric identifier), or a barcode or barcode sequence. In another form, two or more associated sequences of a given set may be associated with a barcode or barcode sequence, where the barcode or barcode sequence is included within the sequence determined by a sequencing reaction. For example, each sequence determined in a sequencing reaction may include both a barcode sequence and a sequence corresponding to a genomic DNA sequence. Optionally, certain sequences or associated sequences may be represented or associated with two or more barcodes or identifiers.
[0044] In another form of association, two or more associated sequences may be maintained within discrete partitions within a computer or computer network, within a hard drive, or within any type of storage medium or any other device that stores sequence data. Optionally, certain sequences or associated sequences may be maintained within two or more partitions within such a computer or data medium.
[0045] Sequences that are associated informationally may include one or more sets of sequences that are associated informationally. The sequences within an associated sequence set may all share the same associative function or its representation; for example, all sequences within an associated set may be associated with the same barcode or with the same identifier, or may be included within the same partition within a computer or storage medium; all sequences may share any other form of association, relationship, and / or correlation. One or more sequences within an associated set may be exclusive members of that set and thus not members of any other set. Alternatively, one or more sequences within an associated set may be non-exclusive members of that set and thus the sequences may be represented and / or associated with two or more different associated sequence sets.
[0046] 1. A sample containing microparticles
[0047] The sample for use in the method of the present invention contains at least one microparticle derived from blood (e.g., human blood). The microparticle may be derived from maternal blood. The microparticle may be derived from the blood of a patient suffering from a disease (e.g., cancer). The sample may be, for example, a blood sample, a plasma sample, or a serum sample. The sample may be a mammalian sample. Preferably, the sample is a human sample.
[0048] A variety of cell-free microparticles have been found in blood, plasma, and / or serum from humans and other animals (Orozco et al, Cytometry Part A (2010). 77A: 502-514, 2010). These microparticles are diverse in the tissues and cells from which they originate, the biophysical processes underlying their formation, and their respective size, molecular structure, and composition. Microparticles can contain components from cell membranes (e.g., incorporated phospholipid components) as well as some intracellular or nuclear components. Microparticles include exosomes, apoptotic bodies (also known as apoptotic vesicles), and extracellular microvesicles.
[0049] Microparticles can be defined as membrane vesicles containing at least two fragments of target nucleic acids (e.g., genomic DNA). The diameter of the microparticles can be 100 - 5000 nm. Preferably, the diameter of the microparticles is 100 - 3000 nanometers.
[0050] Exosomes are one of the smallest circulating microparticles, typically 50 to 100 nanometers in diameter, and are thought to originate from the cell membranes of living intact cells and contain both protein and RNA components (including mRNA molecules and / or degraded mRNA molecules, as well as small regulatory RNA molecules such as microRNA molecules) within an outer phospholipid component. Exosomes are thought to be formed by the exocytosis of cytoplasmic multivesicular bodies (Gyorgy et al, Cell. Mol. Life Sci. (2011) 68: 2667 - 2688). Exosomes are thought to play diverse roles in cell - cell signaling as well as extracellular functions (Kanada et al, PNAS (2015) 1418401112). Techniques for quantifying or sequencing microRNA and / or mRNA molecules found in exosomes have been previously described (e.g., U.S. Patent Application 13 / 456,121, European Application EP2626433A1).
[0051] Microparticles also include apoptotic bodies (also known as apoptotic vesicles) and extracellular microvesicles, which can have a total diameter of up to 1 micrometer or even 2 to 5 micrometers and are generally considered to have a diameter greater than 100 nm (Lichtenstein et al, Ann N Y Acad Sci. (2001); 945: 239 - 49). All types of circulating microparticles are thought to be produced by a large number and variety of cells in the body (Thierry et al, Cancer Metastasis Rev 35(3), 347 - 376.9 (2016) / s10555 - 016 - 9629 - x).
[0052] Preferably, the microparticles are not exosomes, e.g., the microparticles are any microparticles with a diameter greater than that of exosomes.
[0053] A number of methods for separating circulating microparticles (and / or specific subgroups, classes or fractions of circulating microparticles) have been previously described. European Patent ES2540255 (B1) and US Patent 9005888B2 describe methods for separating specific circulating microparticles (such as apoptotic bodies) based on centrifugation operations. A number of methods for separating different types of cell-free microparticles by centrifugation, ultracentrifugation and other techniques have been previously well described and developed (Gyorgy et al, Cell. Mol. Life Sci. (2011) 68: 2667-2688).
[0054] The microparticles contain at least two target nucleic acid fragments (such as molecules of fragmented genomic DNA). These fragmented genomic DNA molecules and / or the sequences contained within these fragmented genomic DNA molecules can be associated by any of the methods described herein.
[0055] The fragments of the target nucleic acid can be DNA fragments (such as molecules of fragmented genomic DNA) or RNA fragments (such as mRNA fragments). Preferably, the target nucleic acid fragments are genomic DNA fragments.
[0056] The DNA fragments can be fragments of mitochondrial DNA. The DNA fragments can be mitochondrial DNA fragments from maternal cells or tissues. The DNA fragments can be mitochondrial DNA fragments from fetal or placental tissues. The DNA fragments can be fragments of mitochondrial DNA from diseased tissues and / or cancerous tissues.
[0057] The microparticles can contain platelets. The microparticles can contain tumour-educated platelets. The target nucleic acid can contain platelet RNA (such as fragments of platelet RNA, and / or fragments of tumour-educated platelet RNA). A sample containing one or more platelets can contain platelet-rich plasma (such as platelet-rich plasma containing tumour-educated platelets).
[0058] The target nucleic acid fragments can contain double-stranded or single-stranded nucleic acids. The genomic DNA fragments can contain double-stranded DNA or single-stranded DNA. The target nucleic acid fragments can contain partially double-stranded nucleic acids. The genomic DNA fragments can contain partially double-stranded DNA.
[0059] The target nucleic acid fragments can be fragments derived from a single nucleic acid molecule, or fragments derived from two or more nucleic acid molecules. For example, the genomic DNA fragments can be derived from a single genomic DNA molecule.
[0060] As understood by those skilled in the art, the term target nucleic acid fragment as used herein refers to the original fragment present in the microparticle and its copies or amplicons. For example, the term gDNA fragment refers to the original gDNA fragment present in the microparticle, and DNA molecules that can be prepared from the original genomic DNA fragment by, for example, primer extension reactions. As another example, the term mRNA fragment refers to the original mRNA fragment present in the microparticle, and cDNA molecules that can be prepared from the original mRNA fragment by, for example, reverse transcription.
[0061] Target nucleic acid (such as genomic DNA) fragments can be at least 10 nucleotides, at least 15 nucleotides, at least 20 nucleotides, at least 25 nucleotides, or at least 50 nucleotides. Target nucleic acid (such as genomic DNA) fragments can be 15 to 100,000 nucleotides, 20 to 50,000 nucleotides, 25 to 25,000 nucleotides, 30 to 10,000 nucleotides, 35 - 5,000 nucleotides, 40 - 1,000 nucleotides, or 50 - 500 nucleotides. Target nucleic acid (such as genomic DNA) fragments can be 20 to 200 nucleotides in length, 100 to 200 nucleotides in length, 200 to 1,000 nucleotides in length, 50 to 250 nucleotides in length, 1,000 to 10,000 nucleotides in length, 10,000 to 100,000 nucleotides in length, or 50 to 100,000 nucleotides in length. Preferably, the fragmented genomic DNA molecule is 50 to 500 nucleotides in length.
[0062] In a sample, the concentration of microparticles can be less than 0.001 microparticles / μL, less than 0.01 microparticles / μL, less than 0.1 microparticles / μL, less than 1.0 microparticles / μL, less than 10 microparticles / μL, less than 100 microparticles / μL, less than 1,000 microparticles / μL, less than 10,000 microparticles / μL, less than 100,000 microparticles / μL, less than 1,000,000 microparticles / μL, less than 10,000,000 microparticles / μL, or less than 100,000,000 microparticles / μL.
[0063] In a sample, the concentration of nucleic acid (such as genomic DNA) fragments can be less than 1.0 picogram DNA / μL, less than 10 picogram DNA / μL, less than 100 picogram DNA / μL, less than 1.0 nanogram DNA / μL, less than 10 nanogram DNA / μL, less than 100 nanogram DNA / μL, or less than 1,000 nanogram DNA / μL
[0064] 2. Linked by barcoding
[0065] The present invention provides a method for preparing a sample for sequencing, wherein the sample comprises microparticles derived from blood, wherein the microparticles comprise at least two target nucleic acid (e.g., genomic DNA) fragments, and wherein the method comprises attaching at least two target nucleic acid fragments of the microparticles to different barcode sequences of a barcode sequence or a set of barcode sequences to produce an associated set of target nucleic acid fragments.
[0066] The present invention provides a method for preparing a sample for sequencing, wherein the sample comprises circulating microparticles, wherein the circulating microparticles comprise at least two target nucleic acid (e.g., genomic DNA) fragments, and wherein the method comprises attaching at least two target nucleic acid fragments of the circulating microparticles to different barcode sequences of a barcode sequence or a set of barcode sequences to produce an associated set of target nucleic acid fragments.
[0067] Before the step of attaching at least two target nucleic acid fragments of the microparticles to different barcode sequences of a barcode sequence or a set of barcode sequences, the method may comprise attaching a coupling sequence to each target nucleic acid (e.g., genomic DNA) fragment of the microparticles, wherein the coupling sequence is subsequently attached to different barcode sequences of a barcode sequence or a set of barcode sequences to produce the associated set of target nucleic acid fragments.
[0068] In the method, the sample may comprise first and second microparticles derived from blood, wherein each microparticle comprises at least two target nucleic acid (e.g., genomic DNA) fragments, and wherein the method may comprise attaching at least two target nucleic acid fragments of the first microparticle to different barcode sequences of a first barcode sequence or a first set of barcode sequences to produce a first set of associated target nucleic acid fragments, and attaching at least two target nucleic acid fragments of the second microparticle to different barcode sequences of a second barcode sequence or a second set of barcode sequences to produce a second set of associated target nucleic acid fragments.
[0069] The first barcode sequence may be different from the second barcode sequence. The barcode sequences of the first set of barcode sequences may be different from the barcode sequences of the second set of barcode sequences.
[0070] In the method, the sample may comprise n microparticles derived from blood, wherein each microparticle comprises at least two target nucleic acid (e.g., genomic DNA) fragments, and wherein the method comprises performing step (a) to produce n sets of associated target nucleic acid fragments, one set for each of the n circulating microparticles.
[0071] In the method, n can be at least 3, at least 5, at least 10, at least 50, at least 100, at least 1000, at least 10,000, at least 100,000, at least 1,000,000, at least 10,000,000, at least 100,000,000, at least 1,000,000,000, at least 10,000,000,000, or at least 100,000,000,000. Preferably, n is at least 100,000 particles.
[0072] Preferably, each set of associated sequence reads is associated by a different barcode sequence or a different set of barcode sequences. Each barcode sequence of the set of barcode sequences can be different from the barcode sequences of at least 1, at least 4, at least 9, at least 49, at least 99, at least 999, at least 9,999, at least 99,999, at least 999,999, at least 9,999,999, at least 99,999,999, at least 999,999,999, at least 9,999,999,999, at least 99,999,999,999, or at least 999,999,999,999 other sets of barcode sequences in the library. Each barcode sequence of the set of barcode sequences can be different from the barcode sequences of all other sets of barcode sequences in the library. Preferably, each barcode sequence of the set of barcode sequences is different from the barcode sequences of at least 9 other sets of barcode sequences in the library.
[0073] The present invention provides a method for analyzing a sample comprising particles derived from blood, wherein the particles comprise at least two target nucleic acid fragments, and wherein the method comprises: (a) preparing a sample for sequencing, which comprises attaching at least two target nucleic acid (e.g., genomic DNA) fragments of the particles to a barcode sequence to generate a set of associated target nucleic acid fragments; and (b) sequencing each associated fragment in the set to generate at least two associated sequence reads, wherein the at least two associated sequence reads are associated by the barcode sequence.
[0074] The barcode sequence can comprise a unique sequence. Each barcode sequence can comprise at least 5, at least 10, at least 15, at least 20, at least 25, at least 50, or at least 100 nucleotides. Preferably, each barcode sequence comprises at least 5 nucleotides. Preferably, each barcode sequence comprises deoxyribonucleotides, optionally all nucleotides in the barcode sequence are deoxyribonucleotides. One or more deoxyribonucleotides can be modified deoxyribonucleotides (e.g., deoxyribonucleotides modified with a biotin moiety or deoxyuridine nucleotides). The barcode sequence can comprise one or more degenerate nucleotides or sequences. The barcode sequence can not comprise any degenerate nucleotides or sequences.
[0075] In the method, prior to the step of attaching at least two target nucleic acid fragments of the microparticle to a barcode sequence, the method may include attaching a coupling sequence to each nucleic acid fragment of the microparticle, wherein the coupling sequence is subsequently attached to the barcode sequence to generate a set of associated fragments.
[0076] In the method, the sample may comprise first and second microparticles derived from blood, wherein each microparticle comprises at least two target nucleic acid (e.g., genomic DNA) fragments, and wherein the method includes performing step (a) to generate a first set of associated target nucleic acid fragments of the first microparticle and a second set of associated target nucleic acid fragments of the second microparticle, and performing step (b) to generate a first set of associated sequence reads of the first microparticle and a second set of associated sequence reads of the second microparticle, wherein at least two associated sequence reads of the first microparticle are associated by a different barcode sequence relative to at least two associated sequence reads of the second microparticle.
[0077] The first set of associated fragments may be associated by a different barcode sequence relative to the second set of associated fragments.
[0078] In the method, the sample may comprise n microparticles derived from blood, wherein each microparticle comprises at least two target nucleic acid (e.g., genomic DNA) fragments, and wherein the method includes performing step (a) to generate n sets of associated target nucleic acid fragments, one set for each of the n microparticles, and performing step (b) to generate n sets of associated sequence reads, one set for each of the n microparticles.
[0079] In the method, n may be at least 3, at least 5, at least 10, at least 50, at least 100, at least 1000, at least 10,000, at least 100,000, at least 1,000,000, at least 10,000,000, at least 100,000,000, at least 1,000,000,000, at least 10,000,000,000, or at least 100,000,000,000. Preferably, n is at least 100,000 microparticles.
[0080] Preferably, each set of associated sequence reads is associated by a different barcode sequence.
[0081] In the method, different barcode sequences can be provided as a barcode sequence library. The library used in the method can contain at least 2, at least 5, at least 10, at least 50, at least 100, at least 1000, at least 10,000, at least 100,000, at least 1,000,000, at least 10,000,000, at least 100,000,000, at least 1,000,000,000, at least 10,000,000,000, at least 100,000,000,000, or at least 1,000,000,000,000 different barcode sequences. Preferably, the library used in the method contains at least 1,000,000 different barcode sequences.
[0082] In the method, each barcode sequence of the library can be attached to a fragment from only a single particle.
[0083] The method can be deterministic, i.e., one barcode sequence can be used to identify a sequence read from a single particle, or probabilistic, i.e., one barcode sequence can be used to identify a sequence read that may be from a single particle. In certain embodiments, one barcode sequence can be attached to genomic DNA fragments from two or more particles.
[0084] The method can include: (a) preparing a sample for sequencing, which includes attaching each of at least two target nucleic acid (e.g., genomic DNA) fragments of the particle to a different barcode sequence of a barcode sequence group to produce an associated group of target nucleic acid fragments; and (b) sequencing each associated fragment in the group to produce at least two associated sequence reads, wherein the at least two associated sequence reads are associated by the barcode sequence group.
[0085] In the method, before the step of attaching each of at least two target nucleic acid fragments of the particle to different barcode sequences, the method can include attaching a coupling sequence to each target nucleic acid fragment of the particle, wherein each of at least two target nucleic acid fragments of the particle is attached to a different barcode sequence of the barcode sequence group through its coupling sequence.
[0086] In the method, the sample can contain first and second particles derived from blood, wherein each particle contains at least two target nucleic acids (e.g., genomic DNA) fragments, and wherein the method can include performing step (a) to produce a first group of associated target nucleic acid fragments of the first particle and a second group of associated target nucleic acid DNA fragments of the second particle, and performing step (b) to produce a first group of associated sequence reads of the first particle and a second group of associated sequence reads of the second particle, wherein the first group of associated sequence reads is associated by a different barcode sequence group relative to the second group of associated sequence reads.
[0087] In the method, the sample can comprise n microparticles derived from blood, where each microparticle comprises at least two target nucleic acid (e.g., genomic DNA) fragments, and where the method can comprise performing step (a) to generate n sets of associated target nucleic acid fragments, one set for each of the n microparticles, and performing step (b) to generate n sets of associated sequence reads, one set for each of the n microparticles.
[0088] In the method, n can be at least 3, at least 5, at least 10, at least 50, at least 100, at least 1000, at least 10,000, at least 100,000, at least 1,000,000, at least 10,000,000, at least 100,000,000, at least 1,000,000,000, at least 10,000,000,000, or at least 100,000,000,000. Preferably, n is at least 100,000 microparticles.
[0089] Preferably, each set of associated sequence reads is associated by a different set of barcode sequences.
[0090] In the method, the different sets of barcode sequences can be provided as a barcode sequence set library. The library used in the method can comprise at least 2, at least 5, at least 10, at least 50, at least 100, at least 1000, at least 10,000, at least 100,000, at least 1,000,000, at least 10,000,000, at least 100,000,000, at least 1,000,000,000, at least 10,000,000,000, at least 100,000,000,000, or at least 1,000,000,000,000 different sets of barcode sequences. Preferably, the library used in the method comprises at least 1,000,000 different sets of barcode sequences.
[0091] Each barcode sequence of a set of barcode sequences can be different from at least 1, at least 4, at least 9, at least 49, at least 99, at least 999, at least 9,999, at least 99,999, at least 999,999, at least 9,999,999, at least 99,999,999, at least 999,999,999, at least 9,999,999,999, at least 99,999,999,999, or at least 999,999,999,999 other sets of barcode sequences in the library. Each barcode sequence of a set of barcode sequences can be different from all other sets of barcode sequences in the library. Preferably, each barcode sequence in a set of barcode sequences is different from the barcode sequences of at least 9 other sets of barcode sequences in the library.
[0092] In the method, barcode sequences from a barcode sequence group of a library can be attached only to fragments from a single particle.
[0093] The method can be deterministic, i.e., the barcode sequence group can be used to identify sequence reads from a single particle, or probabilistic, i.e., the barcode sequence group can be used to identify sequence reads that may be from a single particle.
[0094] The method can include preparing first and second samples for sequencing, where each sample contains at least one particle derived from blood, where each particle contains at least two target nucleic acid (e.g., genomic DNA) fragments, and where the barcode sequences each contain a sample identifier region, and where the method includes: (i) performing step (a) on each sample, where the barcode sequences attached to the target nucleic acid fragments from the first sample have a different sample identifier region from the barcode sequences attached to the target nucleic acid fragments from the second sample; (ii) performing step (b) on each sample, where each associated sequence read contains the sequence of the sample identifier region; and (iii) determining the sample from which each associated sequence read is obtained by its sample identifier region.
[0095] In the method, before, during, and / or after the step of attaching barcode sequences and / or coupling sequences, the method can include a step of crosslinking genomic DNA fragments in the particles.
[0096] In the method, before, during, and / or after the step of attaching barcode sequences and / or coupling sequences, and / or optionally after the step of crosslinking genomic DNA fragments in the particles, the method can include a step of permeabilizing the particles. Before the transfer step and optionally after the crosslinking step, the method includes permeabilizing the particles.
[0097] Barcode sequences can be contained within barcode oligonucleotides in a solution of barcode oligonucleotides; such barcode oligonucleotides can be single-stranded, double-stranded, or single-stranded with one or more double-stranded regions. The barcode oligonucleotides can be ligated to target nucleic acid fragments in a single-stranded or double-stranded ligation reaction. The barcode oligonucleotides can contain single-stranded 5' or 3' regions capable of ligating to target nucleic acid fragments. Each barcode oligonucleotide can be ligated to a target nucleic acid fragment in a single-stranded ligation reaction. Alternatively, the barcode oligonucleotides can contain blunt-ended, recessed-ended, or overhanging 5' or 3' regions capable of ligating to target nucleic acid fragments. Each barcode oligonucleotide can be ligated to a target nucleic acid fragment in a double-stranded ligation reaction.
[0098] In some methods, the ends of the target nucleic acid fragment can be converted to blunt double-stranded ends in a blunt-ending reaction, and the barcoded oligonucleotides can contain blunt double-stranded ends. Each barcoded oligonucleotide can be ligated to the target nucleic acid fragment in a blunt-end ligation reaction. In some methods, the ends of the target nucleic acid fragment can convert its ends to blunt double-stranded ends in a blunt-ending reaction and then convert its ends to a form with a single 3'-adenosine overhang, and wherein the barcoded oligonucleotide contains a double-stranded end with a single 3'-thymine overhang that is capable of annealing to the single 3'-adenosine overhang of the target nucleic acid fragment. Each barcoded oligonucleotide can be ligated to the target nucleic acid fragment in a double-stranded A / T ligation reaction.
[0099] In some methods, the barcoded oligonucleotide contains a target region at its 3' or 5' end that is capable of annealing to a target region in the target nucleic acid and / or the coupling sequence, and the barcode sequence can be attached to the target nucleic acid by annealing the barcoded oligonucleotide to the target nucleic acid and / or the coupling sequence and optionally extending and / or ligating the barcoded oligonucleotide to the nucleic acid target and / or the coupling sequence.
[0100] In some methods, the coupling sequence can be attached to the genomic DNA fragment before attaching the barcoded oligonucleotide.
[0101] Before the attachment step, the method can include the step of dispensing the nucleic acid sample into at least two different reaction volumes.
[0102] 3. Linking by barcoding using a polymeric barcoding reagent
[0103] The present invention provides a method for preparing a sample for sequencing, wherein the sample contains microparticles derived from blood, and wherein the microparticles contain at least two target nucleic acid (e.g., genomic DNA) fragments, and wherein the method comprises the steps of: (a) contacting the sample with a library containing a polymeric barcoding reagent, wherein the polymeric barcoding reagent contains first and second barcode regions linked together, wherein each barcode region contains a nucleic acid sequence; and (b) attaching a barcode sequence to each of the first and second target nucleic acid fragments of the microparticle to produce first and second barcoded target nucleic acid molecules of the microparticle, wherein the first barcoded target nucleic acid molecule contains the nucleic acid sequence of the first barcode region and the second barcoded target nucleic acid molecule contains the nucleic acid sequence of the second barcode region.
[0104] The present invention provides a method for preparing a sample for sequencing, wherein the sample comprises microparticles derived from blood, and wherein the microparticles comprise at least two target nucleic acid (e.g., genomic DNA) fragments, and wherein the method comprises the steps of: (a) contacting the sample with a polymeric barcoding reagent, wherein the polymeric barcoding reagent comprises first and second barcoding oligonucleotides linked together, and wherein each barcoding oligonucleotide comprises a barcode region; and (b) annealing or ligating the first and second barcoding oligonucleotides to the first and second target nucleic acid fragments of the microparticles to produce first and second barcoded target nucleic acid molecules.
[0105] The present invention provides a method for preparing a sample for sequencing, wherein the sample comprises first and second microparticles derived from blood, and wherein each microparticle comprises at least two target nucleic acid (e.g., genomic DNA) fragments, and wherein the method comprises the steps of: (a) contacting the sample with a library comprising at least two polymeric barcoding reagents, wherein each polymeric barcoding reagent comprises first and second barcode regions linked together, wherein each barcode region comprises a nucleic acid sequence, and wherein the first and second barcode regions of the first polymeric barcoding reagent are different from the first and second barcode regions of the second polymeric barcoding reagent of the library; and (b) attaching barcode sequences to each of the first and second target nucleic acid fragments of the first microparticle to produce first and second barcoded target nucleic acid molecules of the first microparticle, wherein the first barcoded target nucleic acid molecule comprises the nucleic acid sequence of the first barcode region of the first polymeric barcoding reagent and the second barcoded target nucleic acid molecule comprises the nucleic acid sequence of the second barcode region of the first polymeric barcoding reagent, and attaching barcode sequences to each of the first and second target nucleic acid fragments of the second microparticle to produce first and second barcoded target nucleic acid molecules of the second microparticle, wherein the first barcoded target nucleic acid molecule comprises the nucleic acid sequence of the first barcode region of the second polymeric barcoding reagent and the second barcoded target nucleic acid molecule comprises the nucleic acid sequence of the second barcode region of the second polymeric barcoding reagent.
[0106] The present invention provides a method for preparing a sample for sequencing, wherein the sample comprises first and second particles derived from blood, and wherein each particle comprises at least two target nucleic acid (e.g., genomic DNA) fragments, and wherein the method comprises the steps of: (a) contacting the sample with a library comprising at least two polymeric barcoding reagents, wherein each polymeric barcoding reagent comprises a first and a second barcoding oligonucleotide linked together, wherein the barcoding oligonucleotides each comprise a barcode region, and wherein the barcode regions of the first and second barcoding oligonucleotides of the first polymeric barcoding reagent of the library are different from the barcode regions of the first and second barcoding oligonucleotides of the second polymeric barcoding reagent of the library; and (b) annealing or ligating the first and second barcoding oligonucleotides of the first polymeric barcoding reagent to the first and second target nucleic acid fragments of the first particle to produce first and second barcoded target nucleic acid molecules, and annealing or ligating the first and second barcoding oligonucleotides of the second polymeric barcoding reagent to the first and second target nucleic acid fragments of the second particle to produce first and second barcoded target nucleic acid molecules.
[0107] The barcoding oligonucleotides can be ligated to the target nucleic acid fragments in a single-stranded or double-stranded ligation reaction.
[0108] In the method, the barcoding oligonucleotides can comprise a single-stranded 5' or 3' region capable of ligating to the target nucleic acid fragments. Each barcoding oligonucleotide can be ligated to the fragment of the target nucleic acid in a single-stranded ligation reaction.
[0109] In the method, the barcoding oligonucleotides can comprise a blunt, recessed or protruding 5' or 3' region capable of ligating to the target nucleic acid fragments. Each barcoding oligonucleotide can be ligated to the fragment of the target nucleic acid in a double-stranded ligation reaction.
[0110] In the method, the ends of the target nucleic acid fragments can be converted to blunt double-stranded ends in a blunt-ending reaction, and the barcoding oligonucleotides can comprise blunt double-stranded ends. Each barcoding oligonucleotide can be ligated to the target nucleic acid fragment in a blunt-end ligation reaction.
[0111] In the method, the ends of the target nucleic acid fragments can convert their ends to blunt double-stranded ends in a blunt-ending reaction and then convert their ends to a form having a single 3' adenosine overhang, and wherein the barcoding oligonucleotides comprise double-stranded ends having a single 3' thymine overhang capable of annealing to the single 3' adenosine overhang of the target nucleic acid fragment. Each barcoding oligonucleotide can be ligated to the fragment of the target nucleic acid in a double-stranded A / T ligation reaction.
[0112] In the method, the ends of the target nucleic acid fragments may be contacted with a restriction enzyme, where the restriction enzyme digests each fragment at a restriction site to produce ligation junctions at these restriction sites, and where the barcoded oligonucleotides comprise ends that are compatible with these ligation junctions. Each barcoded oligonucleotide may be ligated to a target nucleic acid fragment at the ligation junction in a double-stranded ligation reaction. Optionally, the restriction enzyme may be EcoRI, HindIII, or BglII.
[0113] In the method, prior to the step of annealing or ligating the first and second barcoded oligonucleotides to the first and second target nucleic acid fragments, the method may comprise attaching a coupling sequence to each target nucleic acid fragment, where the first and second barcoded oligonucleotides are then annealed or ligated to the coupling sequences of the first and second target nucleic acid fragments.
[0114] In the method, step (b) may comprise: (i) annealing the first and second barcoded oligonucleotides of a first polymeric barcoding reagent to the first and second target nucleic acid fragments of a first particle, and annealing the first and second barcoded oligonucleotides of a second polymeric barcoding reagent to the first and second target nucleic acid fragments of a second particle; and
[0115] (ii) extending the first and second barcoded oligonucleotides of the first polymeric barcoding reagent to produce first and second different barcoded target nucleic acid molecules, and extending the first and second barcoded oligonucleotides of the second polymeric barcoding reagent to produce first and second different barcoded target nucleic acid molecules, where each barcoded target nucleic acid molecule comprises at least one nucleotide synthesized using the target nucleic acid fragment as a template.
[0116] The method may include: (a) contacting a sample with a library comprising at least two polymeric barcoding reagents, wherein each polymeric barcoding reagent comprises a first and a second barcoding oligonucleotide linked together, wherein the barcoding oligonucleotides each comprise a target region and a barcode region in the 5' to 3' direction, wherein the barcode regions of the first and second barcoding oligonucleotides of the first polymeric barcoding reagent of the library are different from the barcode regions of the first and second barcoding oligonucleotides of the second polymeric barcoding reagent of the library, and wherein the sample is also contacted with first and second target primers for each polymeric barcoding reagent; and (b) for each microparticle, performing the following steps: (i) annealing the target region of the first barcoding oligonucleotide to a first subsequence of a first target nucleic acid (e.g., genomic DNA) fragment of the microparticle, and annealing the target region of the second barcoding oligonucleotide to a first subsequence of a second target nucleic acid (e.g., genomic DNA) fragment of the microparticle; (ii) annealing the first target primer to a second subsequence of the first target nucleic acid fragment of the microparticle, wherein the second subsequence is 3' to the first subsequence, and annealing the second target primer to a second subsequence of the second target nucleic acid fragment of the microparticle, wherein the second subsequence is 3' to the first subsequence; (iii) extending the first target primer using the first target nucleic acid fragment of the microparticle as a template until it reaches the first subsequence to produce a first extended target primer, and extending the second target primer using the second target nucleic acid fragment of the microparticle until it reaches the first subsequence to produce a second extended target primer; and (iv) ligating the 3' end of the first extended target primer to the 5' end of the first barcoding oligonucleotide to produce a first barcoded target nucleic acid molecule, and ligating the 3' end of the second extended target primer to the 5' end of the second barcoding oligonucleotide to produce a second barcoded target nucleic acid molecule, wherein the first and second barcoded target nucleic acid molecules are different and each comprises at least one nucleotide synthesized using the target nucleic acid as a template.
[0117] Each polymeric barcoding reagent may comprise: (i) a first and a second hybridization molecule linked together, wherein each hybridization molecule comprises a nucleic acid sequence containing a hybridization region; and (ii) a first and a second barcoding oligonucleotide, wherein the first barcoding oligonucleotide anneals to the hybridization region of the first hybridization molecule, and wherein the second barcoding oligonucleotide anneals to the hybridization region of the second hybridization molecule.
[0118] Each polymeric barcoding reagent may comprise: (i) a first and a second barcoded molecule linked together, wherein each barcoded molecule comprises a nucleic acid sequence containing a barcode region; and (ii) a first and a second barcoding oligonucleotide, wherein the first barcoding oligonucleotide comprises a barcode region that anneals to the barcode region of the first barcoded molecule, and wherein the second barcoding oligonucleotide comprises a barcode region that anneals to the barcode region of the second barcoded molecule.
[0119] In the method, before step (b), the method may include the steps of transferring the first and second barcoded oligonucleotides of a first polymeric barcoding reagent into a first microparticle of a sample and transferring the first and second barcoded oligonucleotides of a second polymeric barcoding reagent into a second microparticle of the sample. Optionally, before step (b), the method further includes the step of transferring a target primer into the first and second microparticles. Optionally, before step (b), the method further includes the step of transferring the first polymeric barcoding reagent into the first microparticle and transferring the second polymeric barcoding reagent into the second microparticle.
[0120] The present invention provides a method for preparing a sample for sequencing, wherein the sample comprises at least two microparticles derived from blood, wherein each microparticle comprises at least two target nucleic acid fragments, and wherein the method comprises the following steps: (a) contacting the sample with a library comprising a first polymeric barcoding reagent and a second polymeric barcoding reagent, wherein each polymeric barcoding reagent comprises a first and a second barcode molecule linked together, and wherein each barcode molecule comprises a nucleic acid sequence optionally comprising a barcode region and an adaptor region in the 5' to 3' direction; (b) attaching a coupling sequence to the first and second target nucleic acid (e.g., genomic DNA) fragments of the first and second microparticles; (c) for each polymeric barcoding reagent, annealing the coupling sequence of the first fragment to the adaptor region of the first barcode molecule and annealing the coupling sequence of the second fragment to the adaptor region of the second barcode molecule; and (d) for each polymeric barcoding reagent, attaching a barcode sequence to each of at least two target nucleic acid fragments of the microparticle to produce first and second different barcoded target nucleic acid molecules, wherein the first barcoded target nucleic acid molecule comprises the nucleic acid sequence of the barcode region of the first barcode molecule, and the second barcoded target nucleic acid molecule comprises the nucleic acid sequence of the barcode region of the second barcode molecule.
[0121] In the method, each barcode molecule may comprise a nucleic acid sequence comprising a barcode region and an adaptor region in the 5' to 3' direction, and wherein step (d) comprises, for each polymeric barcoding reagent, using the barcode region of the first barcode molecule as a template to extend the coupling sequence of the first fragment to produce a first barcoded target nucleic acid molecule, and using the barcode region of the second barcode molecule as a template to extend the coupling sequence of the second fragment to produce a second barcoded target nucleic acid molecule, wherein the first barcoded target nucleic acid molecule comprises a sequence complementary to the barcode region of the first barcode molecule, and the second barcoded target nucleic acid molecule comprises a sequence complementary to the barcode region of the second barcode molecule.
[0122] In the method, each barcode molecule may comprise a nucleic acid sequence containing an adaptor region and a barcode region in the 5' to 3' direction, wherein step (d) comprises, for each polymeric barcoding reagent, (i) annealing and extending a first extension primer using the barcode region of the first barcode molecule as a template to generate a first barcoded oligonucleotide, and annealing and extending a second extension primer using the barcode region of the second barcode molecule as a template to generate a second barcoded oligonucleotide, wherein the first barcoded oligonucleotide contains a sequence complementary to the barcode region of the first barcode molecule, and the second barcoded oligonucleotide contains a sequence complementary to the barcode region of the second barcode molecule, (ii) ligating the 3' end of the first barcoded oligonucleotide to the 5' end of the coupling sequence of the first fragment to generate a first barcoded target nucleic acid molecule, and ligating the 3' end of the second barcoded oligonucleotide to the 5' end of the coupling sequence of the second fragment to generate a second barcoded target nucleic acid molecule.
[0123] In the method, each barcode molecule may comprise a nucleic acid sequence containing an adaptor region, a barcode region and a primer region in the 5' to 3' direction, wherein step (d) comprises, for each polymeric barcoding reagent, (i) annealing a first extension primer to the primer region of the first barcode molecule and extending the first extension primer using the barcode region of the first barcode molecule as a template to generate a first barcoded oligonucleotide, and annealing a second extension primer to the primer region of the second barcode molecule and extending the second extension primer using the barcode region of the second barcode molecule as a template to generate a second barcoded oligonucleotide, wherein the first barcoded oligonucleotide contains a sequence complementary to the barcode region of the first barcode molecule and the second barcoded oligonucleotide contains a sequence complementary to the barcode region of the second barcode molecule, and (ii) ligating the 3' end of the first barcoded oligonucleotide to the 5' end of the coupling sequence of the first fragment to generate a first barcoded target nucleic acid molecule, and ligating the 3' end of the second barcoded oligonucleotide to the 5' end of the coupling sequence of the second fragment to generate a second barcoded target nucleic acid molecule.
[0124] Before step (b) or step (c), the method may comprise transferring the first polymeric barcoding reagent, the coupling sequence and / or the extension primer into the first microparticle and transferring the second polymeric barcoding reagent, the coupling sequence and / or the primer extension into the second microparticle.
[0125] The method may comprise: (a) contacting a sample with a library comprising first and second polymeric barcoding reagents, wherein each polymeric barcoding reagent comprises first and second barcode molecules linked together, wherein each barcode molecule comprises a nucleic acid sequence comprising a barcode region and an adaptor region in the 5' to 3' direction, and wherein the sample is also contacted with first and second adaptor oligonucleotides of each polymeric barcoding reagent, wherein the first and second adaptor oligonucleotides each comprise an adaptor region; and (b) ligating the first and second adaptor oligonucleotides of the first polymeric barcoding reagent to first and second target nucleic acid fragments of a first particle, and ligating the first and second adaptor oligonucleotides of the second polymeric barcoding reagent to first and second target nucleic acid fragments of a second particle; (c) for each polymeric barcoding reagent, annealing the adaptor region of the first adaptor oligonucleotide to the adaptor region of the first barcode molecule, and annealing the adaptor region of the second adaptor oligonucleotide to the adaptor region of the second barcode molecule; and (d) for each polymeric barcoding reagent, extending the first adaptor oligonucleotide using the barcode region of the first barcode molecule as a template to generate a first barcoded target nucleic acid molecule, and extending the second adaptor oligonucleotide using the barcode region of the second barcode molecule as a template to generate a second barcoded target nucleic acid molecule, wherein the first barcoded target nucleic acid molecule comprises a sequence complementary to the barcode region of the first barcode molecule, and the second barcoded target nucleic acid molecule comprises a sequence complementary to the barcode region of the second barcode molecule.
[0126] The method may include the following steps: (a) contacting a sample with a library comprising first and second polymeric barcoding reagents, wherein each polymeric barcoding reagent comprises: (i) first and second barcode molecules linked together, wherein each barcode molecule comprises a nucleic acid sequence optionally comprising an adaptor region and a barcode region in the 5' to 3' direction, and (ii) first and second barcoded oligonucleotides, wherein the first barcoded oligonucleotide comprises a barcode region annealing to the barcode region of the first barcode molecule, wherein the second barcoded oligonucleotide comprises a barcode region annealing to the barcode region of the second barcode molecule, and wherein the barcode regions of the first and second barcoded oligonucleotides of the first polymeric barcoding reagent of the library are different from the barcode regions of the first and second barcoded oligonucleotides of the second polymeric barcoding reagent of the library; wherein the sample is also contacted with first and second adaptor oligonucleotides of each polymeric barcoding reagent, wherein the first and second adaptor oligonucleotides each comprise an adaptor region; (b) annealing or ligating the first and second adaptor oligonucleotides of the first polymeric barcoding reagent to first and second target nucleic acid (e.g., genomic DNA) fragments of a first particle, and annealing or ligating the first and second adaptor oligonucleotides of the second polymeric barcoding reagent to first and second target nucleic acid (e.g., genomic DNA) fragments of a second particle, (c) for each polymeric barcoding reagent, annealing the adaptor region of the first adaptor oligonucleotide to the adaptor region of the first barcode molecule, and annealing the adaptor region of the second adaptor oligonucleotide to the adaptor region of the second barcode molecule; and (d) for each polymeric barcoding reagent, ligating the 3' end of the first barcoded oligonucleotide to the 5' end of the first adaptor oligonucleotide to produce a first barcoded target nucleic acid molecule, and ligating the 3' end of the second barcoded oligonucleotide to the 5' end of the second adaptor oligonucleotide to produce a second barcoded target nucleic acid molecule.
[0127] In the method, step (b) may include annealing the first and second adapter oligonucleotides of a first polymeric barcoding reagent to first and second target nucleic acid (e.g., genomic DNA) fragments of a first microparticle, and annealing the first and second adapter oligonucleotides of a second polymeric barcoding reagent to first and second target nucleic acid (e.g., genomic DNA) fragments of a second microparticle, and wherein: (i) for each polymeric barcoding reagent, step (d) includes ligating the 3'-end of a first barcoding oligonucleotide to the 5'-end of the first adapter oligonucleotide to produce a first barcoded-adapter oligonucleotide and ligating the 3'-end of a second barcoding oligonucleotide to the 5'-end of the second adapter oligonucleotide to produce a second barcoded-adapter oligonucleotide, and extending the first and second barcoded-adapter oligonucleotides to produce first and second distinct barcoded target nucleic acid molecules, each comprising at least one nucleotide synthesized using the target nucleic acid fragment as a template, or (ii) for each polymeric barcoding reagent, prior to step (d), the method includes extending the first and second adapter oligonucleotides to produce first and second distinct target nucleic acid molecules, each comprising at least one nucleotide synthesized using the target nucleic acid fragment as a template.
[0128] In the method, prior to the step of annealing or ligating the first and second adapter oligonucleotides to the first and second target nucleic acid fragments, the method may include attaching a coupling sequence to each target nucleic acid fragment, wherein the first and second adapter oligonucleotides are subsequently annealed to or ligated to the coupling sequences of the first and second target nucleic acid fragments.
[0129] In the method, prior to step (b) or step (c), the method may include transferring the first and second adapter oligonucleotides of a first polymeric barcoding reagent into a first microparticle and transferring the first and second adapter oligonucleotides of a second polymeric barcoding reagent into a second microparticle, optionally wherein the step further includes transferring the first polymeric barcoding reagent into the first microparticle and transferring the second polymeric barcoding reagent into the second microparticle.
[0130] In any of the methods described herein, the method can include a step of crosslinking a target nucleic acid (e.g., genomic DNA) fragment in the microparticles. This step can be carried out using a chemical crosslinking agent such as formaldehyde, paraformaldehyde, glutaraldehyde, disuccinimidyl glutarate, ethylene glycol bis(succinimidyl succinate), homobifunctional crosslinking agent or heterobifunctional crosslinking agent. This step is carried out before any permeabilization step, after any permeabilization step, before any dispensing step, before any step of attaching a coupling sequence, after any step of attaching a coupling sequence, before any step of attaching a barcode sequence (e.g., before step (b)), after any step of attaching a barcode sequence (e.g., after step (d)), simultaneously with attaching a barcode sequence, or any combination thereof. For example, a sample containing microparticles can be crosslinked before contacting the sample containing microparticles with a library of two or more polymeric barcoded reagents. Any such crosslinking step can be further ended by a quenching step, for example, quenching a formaldehyde crosslinking step by mixing with a glycine solution. Any such crosslinking can be removed before a specific subsequent step of the protocol, for example, before a primer extension, PCR or nucleic acid purification step.
[0131] In the method, during step (b), (c) and / or (d) (i.e., the step of attaching a barcode sequence), the microparticles and / or the target nucleic acid fragment can be contained within a gel or hydrogel, such as an agarose gel, a polyacrylamide gel or any covalently crosslinked gel, such as a covalently crosslinked poly(ethylene glycol) gel, or a covalently crosslinked gel comprising a mixture of thiol-functionalized poly(ethylene glycol) and acrylate-functionalized poly(ethylene glycol).
[0132] In any of the methods described herein, optionally after a crosslinking step, the method may include permeabilizing the microparticles. The microparticles may be permeabilized by an incubation step. The incubation step may be carried out in the presence of a chemical surfactant. Optionally, this permeabilization step may occur before attaching the barcode sequence (e.g., before step (b)), after attaching the barcode sequence (e.g., after step (d)), or both before and after attaching the barcode sequence. The incubation step may be carried out at a temperature of at least 20 degrees Celsius, at least 30 degrees Celsius, at least 37 degrees Celsius, at least 45 degrees Celsius, at least 50 degrees Celsius, at least 60 degrees Celsius, at least 65 degrees Celsius, at least 70 degrees Celsius, or at least 80 degrees Celsius. The incubation step may be at least 1 second long, at least 5 seconds long, at least 10 seconds long, at least 30 seconds long, at least 1 minute long, at least 5 minutes long, at least 10 minutes long, at least 30 minutes long, at least 60 minutes long, or at least 3 hours long. This step may be carried out after any crosslinking step, before any permeabilization step, after any permeabilization step, before any dispensing step, before any step of attaching a coupling sequence, after any step of attaching a coupling sequence, before any step of attaching a barcode sequence (e.g., before step (b)), after any step of attaching a barcode sequence (e.g., after step (d)), simultaneously with attaching the barcode sequence, or any combination thereof. For example, before contacting a sample comprising microparticles with a library of two or more multimeric barcoded reagents, the sample comprising microparticles may be crosslinked and subsequently permeabilized in the presence of a chemical surfactant.
[0133] In any of the methods described herein, a sample of the microparticles may be digested with a protease digestion step (e.g., digestion with Proteinase K). Optionally, this protease digestion step may be at least 10 seconds long, at least 30 seconds long, at least 60 seconds long, at least 5 minutes long, at least 10 minutes long, at least 30 minutes long, at least 60 minutes long, at least 3 hours long, at least 6 hours long, at least 12 hours long, or at least 24 hours long. This step may be carried out after any crosslinking step, before any permeabilization step, after any permeabilization step, before any dispensing step, before any step of attaching a coupling sequence, after any step of attaching a coupling sequence, before any step of attaching a barcode sequence (e.g., before step (b)), after any step of attaching a barcode sequence (e.g., after step (d)), simultaneously with attaching the barcode sequence, or any combination thereof. For example, before contacting a sample comprising microparticles with a library of two or more multimeric barcoded reagents, the sample comprising microparticles may be crosslinked and subsequently partially digested with a Proteinase K digestion step.
[0134] In the method, barcoded oligonucleotides, adapter oligonucleotides, and / or multimeric barcoded reagents may be transferred into the microparticles by complexing with a transfection reagent or a lipid carrier (e.g., liposome or micelle).
[0135] The transfection reagent can be a lipid transfection reagent, such as a cationic lipid transfection reagent. Optionally, the cationic lipid transfection reagent comprises at least two alkyl chains. Optionally, the cationic lipid transfection reagent can be a commercially available cationic lipid transfection reagent, such as Lipofectamine.
[0136] In the method, the barcoded oligonucleotide of the first polymer barcoding reagent can be comprised in a first lipid carrier, and wherein the barcoded oligonucleotide of the second polymer barcoding reagent can be comprised in a second lipid carrier. The lipid carrier can be a liposome or a micelle.
[0137] In the method, steps (a) and (b) and optionally (c) and (d) can be performed on at least two microparticles in a single reaction volume.
[0138] Before step (b), the method can further comprise the step of distributing the nucleic acid sample into at least two different reaction volumes.
[0139] The present invention provides a method for analyzing a sample comprising microparticles derived from blood, wherein the microparticles comprise at least two target nucleic acid (e.g., genomic DNA) fragments, and wherein the method comprises: (a) preparing a sample for sequencing, which includes (i) contacting the sample with a polymer barcoding reagent comprising a first and a second barcode region linked together, wherein each barcode region comprises a nucleic acid sequence, and (ii) attaching a barcode sequence to each of at least two target nucleic acid fragments of the microparticle to produce first and second different barcoded target nucleic acid molecules, wherein the first barcoded target nucleic acid molecule comprises the nucleic acid sequence of the first barcode region and the second barcoded target nucleic acid molecule comprises the nucleic acid sequence of the second barcode region; and (b) sequencing each barcoded target nucleic acid molecule to produce at least two associated sequence reads.
[0140] In the method, before the step of attaching a barcode sequence to each of at least two genomic DNA fragments of the microparticle, the method can comprise attaching a coupling sequence to each genomic DNA fragment of the microparticle, wherein the barcode sequence is subsequently attached to the coupling sequence of each of at least two genomic DNA fragments of the microparticle to produce first and second different barcoded target nucleic acid molecules.
[0141] The method can further comprise, optionally before step (a)(i) or (a)(ii), the step of transferring the first and second barcode regions of the polymer barcoding reagent into the microparticle
[0142] Before the transfer step, any method described herein may further include a step of crosslinking genomic DNA fragments in the microparticles. The crosslinking step can be carried out using a chemical crosslinking agent such as formaldehyde, paraformaldehyde, glutaraldehyde, disuccinimidyl glutarate, ethylene glycol bis(succinimidyl succinate), homobifunctional crosslinking agents or heterobifunctional crosslinking agents.
[0143] During step (a), the microparticles and / or the target nucleic acid fragments may be contained within a gel or hydrogel, such as an agarose gel, a polyacrylamide gel, or any covalently crosslinked gel, such as a covalently crosslinked poly(ethylene glycol) gel, or a covalently crosslinked gel comprising a mixture of thiol-functionalized poly(ethylene glycol) and acrylate-functionalized poly(ethylene glycol).
[0144] Before the transfer step and optionally after the crosslinking step, the method may further include a step of permeabilizing the microparticles. The microparticles can be permeabilized by an incubation step. The incubation step can be carried out in the presence of a chemical surfactant. Optionally, this permeabilization step can occur before attachment of the barcode sequence (e.g., before step (a)(ii)), after attachment of the barcode sequence (e.g., after step (a)(ii)), or both before and after attachment of the barcode sequence. The incubation step can be carried out at a temperature of at least 20 °C, at least 30 °C, at least 37 °C, at least 45 °C, at least 50 °C, at least 60 °C, at least 65 °C, at least 70 °C, or at least 80 °C. The incubation step can be at least 1 second long, at least 5 seconds long, at least 10 seconds long, at least 30 seconds long, at least 1 minute long, at least 5 minutes long, at least 10 minutes long, at least 30 minutes long, at least 60 minutes long, or at least 3 hours long.
[0145] The microparticle sample can be digested with a protease digestion step (e.g., digestion with proteinase K). Optionally, this protease digestion step can be at least 10 seconds long, at least 30 seconds long, at least 60 seconds long, at least 5 minutes long, at least 10 minutes long, at least 30 minutes long, at least 60 minutes long, at least 3 hours long, at least 6 hours long, at least 12 hours long, or at least 24 hours long. This step can be carried out before permeabilization, after permeabilization, before attachment of the barcode sequence (e.g., before step (a)(ii)), after attachment of the barcode sequence (e.g., after step (a)(ii)), simultaneously with attachment of the barcode sequence, or any combination thereof.
[0146] The first and second barcode regions of the polymeric barcoding reagent can be transferred into the microparticles by complexing with a transfection reagent or a lipid carrier (such as a liposome or a micelle).
[0147] The transfection reagent can be a lipid transfection reagent, such as a cationic lipid transfection reagent. Optionally, the cationic lipid transfection reagent contains at least two alkyl chains. Optionally, the cationic lipid transfection reagent can be a commercially available cationic lipid transfection reagent, such as Lipofectamine.
[0148] Step (a) of the method can be carried out by any method for preparing a sample (or nucleic acid sample) for sequencing as described herein.
[0149] The method can include preparing first and second samples for sequencing, where each sample contains at least one particle derived from blood, where each particle contains at least two target nucleic acid (e.g., genomic DNA) fragments, and where the barcode sequences each contain a sample identifier region, and where the method includes: (i) performing step (a) on each sample, where the barcode sequence attached to the nucleic acid fragment from the first sample has a different sample identifier region from the barcode sequence attached to the target nucleic acid fragment from the second sample; (ii) performing step (b) on each sample, where each sequence read contains the sequence of the sample identification region; and (iii) determining the sample from which each sequence read is obtained by its sample identifier region.
[0150] The method can include analyzing a sample containing at least two particles derived from blood, where each particle contains at least two target nucleic acid (e.g., genomic DNA) fragments, and where the method includes the following steps: (a) preparing a sample for sequencing, which includes: (i) contacting the sample with a library of polymeric barcoding reagents containing a polymeric barcoding reagent for each of two or more particles, where each polymeric barcoding reagent is as defined herein; and (ii) attaching a barcode sequence to each of at least two target nucleic acid fragments of each particle, where at least two barcoded target nucleic acid molecules are produced from each of the at least two particles, and where at least two barcoded target nucleic acid molecules produced from a single particle each contain a nucleic acid sequence from the barcode region of the same polymeric barcoding reagent; and (b) sequencing each barcoded target nucleic acid molecule to produce at least two associated sequence reads for each particle.
[0151] The barcode sequence can be attached to the genomic DNA fragments of the particles in a single reaction volume, i.e., step (a) of the method can be carried out in a single reaction volume.
[0152] Before the attachment step (step (a)(ii)), the method can further include the step of dispensing the sample into at least two different reaction volumes.
[0153] In any method, prior to the step of attaching barcode sequences, the polymeric barcoding reagent can be separated, fractionated, or solubilized into two or more components, such as releasing barcoded oligonucleotides.
[0154] In any method, the concentration of the polymeric barcoding reagent can be less than 1.0 femtomolar, less than 10 femtomolar, less than 100 femtomolar, less than 1.0 picomolar, less than 10 picomolar, less than 100 picomolar, less than 1 nanomolar, less than 10 nanomolar, less than 100 nanomolar, or less than 1.0 micromolar.
[0155] 4. Linking by bringing the segments together
[0156] The present invention provides a method for analyzing a sample comprising microparticles derived from blood, wherein the microparticles comprise at least two target nucleic acid (e.g., genomic DNA) segments, and wherein the method comprises: (a) preparing a sample for sequencing, which includes bringing together at least two target nucleic acid segments of the microparticles to produce a single nucleic acid molecule comprising the sequences of at least two target nucleic acid segments; and (b) sequencing each segment in the single nucleic acid molecule to produce at least two linked sequence reads.
[0157] The at least two target nucleic acid (e.g., genomic DNA) segments can be contiguous in the single nucleic acid molecule.
[0158] The at least two linked sequence reads can be provided within a single raw sequence read.
[0159] The method can include, prior to the linking step, attaching a coupling sequence to at least one target nucleic acid (e.g., genomic DNA) segment and subsequently bringing together at least two target nucleic acid segments via the coupling sequence.
[0160] Target nucleic acid (e.g., genomic DNA) segments can be brought together by a solid support, where two or more segments are linked to the same solid support (directly or indirectly, e.g., via a coupling sequence). Optionally, the solid support is a bead, such as a Styrofoam bead, a superparamagnetic bead, or an agarose bead.
[0161] Target nucleic acid (e.g., genomic DNA) segments can be brought together by a ligation reaction (e.g., a double-stranded ligation reaction or a single-stranded ligation reaction).
[0162] The ends of the target nucleic acid segments can be converted to blunt-ended ligatable double-stranded ends in a blunt-ending reaction, and the method can include ligating two or more segments to each other by a blunt-end ligation reaction.
[0163] The ends of the target nucleic acid fragments can be contacted with a restriction enzyme, which digests the fragments at restriction sites to create ligation junctions at these restriction sites, and wherein the method can include ligating two or more fragments to each other by a ligation reaction at the ligation junctions. Any target nucleic acid can be contacted with a restriction enzyme, which digests the fragments at restriction sites to create ligation junctions at these restriction sites, and wherein the method can include ligating two or more fragments to each other by a ligation reaction at the ligation junctions. Optionally, the restriction enzyme can be EcoRI, HindIII or BglII.
[0164] Prior to bringing the fragments together, coupling sequences can be attached to two or more target nucleic acid fragments. Optionally, two or more different coupling sequences are attached to a population of target nucleic acid fragments.
[0165] The coupling sequences can include ligation junctions at at least one end, and wherein a first coupling sequence is attached to a first target nucleic acid fragment, and wherein a second coupling sequence is attached to a second target nucleic acid fragment, and wherein the two coupling sequences are ligated to each other, thereby bringing the two target nucleic acid fragments together.
[0166] The coupling sequences can include annealing regions at at least one 3'-end, and wherein a first coupling sequence is attached to a first target nucleic acid fragment, and wherein a second coupling sequence is attached to a second target nucleic acid fragment, and wherein the two coupling sequences are complementary to each other and anneal along a segment of at least one nucleotide in length, and wherein a DNA polymerase is used to extend at least one 3'-end of the first coupling sequence by at least one nucleotide into the sequence of the second target nucleic acid fragment, thereby bringing the two target nucleic acid (e.g., genomic DNA) fragments together.
[0167] Prior to bringing at least two fragments together, the method can further include the step of cross-linking the microparticles, for example using a chemical cross-linking agent such as formaldehyde, paraformaldehyde, glutaraldehyde, disuccinimidyl glutarate, ethylene glycol bis(succinimidyl succinate), homobifunctional cross-linking agents or heterobifunctional cross-linking agents.
[0168] Prior to bringing at least two fragments together, the method can further include partitioning the microparticles into two or more compartments.
[0169] The method can further include permeabilizing the microparticles during the incubation step. This step can be performed before partitioning (if performed), after partitioning (if performed), before bringing the fragments together and / or after bringing the fragments together.
[0170] The incubation step can be performed in the presence of a chemical surfactant, such as Triton X-100 (C 14 H 22 O(C2H4O)n (n = 9 - 10)), NP - 40, Tween 20, Tween 80, saponin, digitonin, or sodium dodecyl sulfate.
[0171] The incubation step is carried out at a temperature of at least 20 °C, at least 30 °C, at least 37 °C, at least 45 °C, at least 50 °C, at least 60 °C, at least 65 °C, at least 70 °C, at least 80 °C, at least 90 °C, or at least 95 °C.
[0172] The incubation step can be at least 1 second long, at least 5 seconds long, at least 10 seconds long, at least 30 seconds long, at least 1 minute long, at least 5 minutes long, at least 10 minutes long, at least 30 minutes long, at least 60 minutes long, or at least 3 hours long.
[0173] The method may include digesting a sample of the microparticles with a protease digestion step (e.g., digestion with proteinase K). Optionally, the protease digestion step can be at least 10 seconds long, at least 30 seconds long, at least 60 seconds long, at least 5 minutes long, at least 10 minutes long, at least 30 minutes long, at least 60 minutes long, at least 3 hours long, at least 6 hours long, at least 12 hours long, or at least 24 hours long. This step can be carried out before (if performed) the dispensing, after (if performed) the dispensing, before and / or after linking the fragments together.
[0174] The method may include amplifying the (original) target nucleic acid fragment and subsequently linking two or more of the resulting nucleic acid molecules together.
[0175] The step of linking the fragments together may produce concatamerised nucleic acid molecules that contain at least 3, at least 5, at least 10, at least 50, at least 100, at least 500, or at least 1000 nucleic acid molecules that have been attached to each other in a single continuous nucleic acid molecule.
[0176] The method can be used to generate linked sequence reads for at least 3 microparticles, at least 5 microparticles, at least 10 microparticles, at least 50 microparticles, at least 100 microparticles, at least 1000 microparticles, at least 10,000 microparticles, at least 100,000 microparticles, at least 1,000,000 microparticles, at least 10,000,000 microparticles, at least 100,000,000 microparticles, at least 1,000,000,000 microparticles, at least 10,000,000,000 microparticles, or at least 100,000,000,000 microparticles.
[0177] The sample can comprise at least two blood-derived particles, where each particle comprises at least two target nucleic acid (e.g., genomic DNA) fragments, and where the method comprises performing step (a) to generate a single nucleic acid molecule comprising the sequences of at least two target nucleic acid fragments of each particle, and performing step (b) to generate associated sequence reads for each particle.
[0178] Before, during, and / or after the step of ligating at least two target nucleic acid (e.g., genomic DNA) fragments together, the method can comprise a step of cross-linking the target nucleic acid fragments in the particles. The cross-linking step can be carried out with a chemical cross-linking agent such as formaldehyde, paraformaldehyde, glutaraldehyde, disuccinimidyl glutarate, ethylene glycol bis(succinimidyl succinate), homobifunctional cross-linking agents, or heterobifunctional cross-linking agents.
[0179] Before, during, and / or after the step of ligating at least two target nucleic acid (e.g., genomic DNA) fragments together, and / or optionally after the step of cross-linking the target nucleic acid fragments in the particles, the method comprises a step of permeabilizing the particles.
[0180] Before step (a), the method can further comprise a step of partitioning a nucleic acid sample into at least two different reaction volumes.
[0181] In one embodiment of a method of linking together at least two target nucleic acid fragments of a circular particle to produce a single nucleic acid molecule comprising the sequences of at least two target nucleic acid fragments, a sample comprising at least one circular particle (e.g., wherein the sample is obtained and / or purified by any method disclosed herein) is cross-linked in a 1% formaldehyde solution for 10 minutes at room temperature and the formaldehyde cross-linking step is subsequently quenched with glycine. The particles are pelleted by a centrifugation step (e.g., 5 minutes at 3000×G) and resuspended in 1×NEBuffer 2 (New England Biolabs) containing 1.0% sodium dodecyl sulfate (SDS) and incubated for 10 minutes at 45 degrees Celsius to permeabilize the particles. The SDS is quenched by the addition of Triton X-100 and the solution is incubated overnight at 37 degrees Celsius with AluI (New England Biolabs) to generate blunt-end ligatable ends. The enzyme is inactivated by the addition of SDS to a final concentration of 1.0% and incubated for 15 minutes at 65 degrees Celsius. The SDS is quenched by the addition of Triton X-100 and the solution is diluted at least 10-fold in 1× buffer for T4 DNA ligase and to a total DNA concentration of at most 1.0 nanogram of DNA per microliter. The diluted solution is incubated overnight at 16 degrees Celsius with T4 DNA ligase to ligate together the fragments from the circular particles. Cross-linking is subsequently reversed and the protein component is degraded by incubating overnight in a solution of proteinase K at 65 degrees Celsius. The ligated DNA is subsequently purified (e.g., with a Qiagen spin-column PCR Purification Kit and / or Ampure XP beads). Subsequently, Illumina sequencing adapter sequences are attached using the Nextera in vitro transposition method (Illumina; according to the manufacturer's protocol), an appropriate number of PCR cycles are performed to amplify the ligated material; and the amplified and purified DNA of appropriate size is subsequently sequenced using an Illumina sequencer (e.g., an Illumina NextSeq 500 or MiSeq) with paired-end reads of at least 50 bases each. Each end of the paired-end sequences is independently mapped to a reference human genome to elucidate linked sequence reads (e.g., wherein both ends contain reads of sequences of different genomic DNA fragments from a single circular particle).
[0182] Methods for linking together at least two target nucleic acid fragments of a particle to produce a single nucleic acid molecule containing the sequences of at least two target nucleic acid fragments can have a variety of unique properties and characteristics that make them desirable for linking sequences from one or more circulating particles. In one aspect, such methods enable the linking of sequences from circulating particles without the need for complex instrumentation (e.g., microfluidics for partition-based methods). Additionally, the method is (broadly) capable of being performed in a single, individual reaction that can contain a large number of circulating particles (e.g., hundreds, or thousands or more), and can thus process a large number of circulating particles without the need for multiple reactions, which may be necessary in other methods, such as in combinatorial indexing methods. Further, since the method does not necessarily require the use of barcodes and / or polymeric barcoding reagents, it is not limited by the size of the barcode library (and / or polymeric barcoding reagent library) to achieve useful molecular measurements of linked sequences from circulating particles.
[0183] 5. Linking by partitioning
[0184] The method can be performed on a nucleic acid sample of at least two particles that have been partitioned into at least two different reaction volumes (or partitions).
[0185] In any method, a nucleic acid sample of at least two particles can be partitioned into at least two different reaction volumes (or partitions). The different reaction volumes (or partitions) can be provided by different reaction vessels (or different physical reaction vessels). The different reaction volumes (or partitions) can be provided by different aqueous droplets, e.g., different aqueous droplets within an emulsion or different aqueous droplets on a solid support (e.g., a glass slide).
[0186] For example, the nucleic acid sample can be partitioned before attaching a barcode sequence to the target nucleic acid fragment of the particle. Alternatively, the nucleic acid sample can be partitioned before linking together at least two target nucleic acid fragments of the particle.
[0187] For any method involving a partitioning step, any step of the method after the partitioning step (e.g., any step of attaching a barcode sequence or an attachment conjugate sequence, or any step of ligation, annealing, primer extension, or PCR) can be performed independently on each partition. Reagents (e.g., oligonucleotides, enzymes, and buffers) can be added directly to each partition. In methods where the partitions comprise aqueous droplets within an emulsion, such addition steps can be performed by a process of fusing the aqueous droplets within the emulsion, e.g., using a microfluidic droplet-merging channel and optionally using mechanical or thermal mixing steps.
[0188] The partition contains different aqueous solution droplets within an emulsion, and wherein the emulsion is a water-in-oil emulsion, and wherein the droplets are generated by physical shaking or vortexing steps, or wherein the droplets are generated by fusing an aqueous solution with an oil solution within a microfluidic conduit or junction.
[0189] For methods in which the partition contains aqueous droplets within an emulsion, such water-in-oil emulsions can be generated by any method or tool known in the art. Optionally, this can include commercially available microfluidic systems such as the Chromium system available from 10×Genomics Inc or other systems, digital droplet generators from Raindance Technologies or Bio-Rad, and component-based systems for microfluidic generation and manipulation such as Drop-Seq (Macosko et al., 2015, Cell 161, 1202-1214) and inDrop (Klein et al., 2015, Cell 161, 1187-1201).
[0190] The partition can contain different physically non-overlapping spatial volumes within a gel or hydrogel, such as an agarose gel, a polyacrylamide gel, or any covalently crosslinked gel, such as a covalently crosslinked poly(ethylene glycol) gel, or a covalently crosslinked gel containing a mixture of thiol-functionalized poly(ethylene glycol) and acrylate-functionalized poly(ethylene glycol).
[0191] The particulate sample can be partitioned into a total of at least 10, at least 100, at least 1000, at least 10,000, at least 100,000, at least 1,000,000, at least 10,000,000, at least 100,000,000, or at least 1,000,000,000 partitions. Preferably, the particulate solution is partitioned into a total of at least 1000 partitions.
[0192] The particulate sample can be partitioned into the partitions such that there is an average of less than 0.0001, less than 0.001, less than 0.01, less than 0.1, less than 1.0, less than 10, less than 100, less than 1000, less than 10,000, less than 100,000, less than 1,000,000, less than 10,000,000, or less than 100,000,000 particulates per partition. Preferably, there is an average of less than 1.0 particulate per partition.
[0193] A solution of the particles can be dispensed into partitions such that each partition contains less than 1.0 attogram of DNA, less than 10 attograms of DNA, less than 100 attograms of DNA, less than 1.0 femtogram of DNA, less than 10 femtograms of DNA, less than 100 femtograms of DNA, less than 1.0 picogram of DNA, less than 10 picograms of DNA, less than 100 picograms of DNA, or less than 1.0 nanogram of DNA. Preferably, each partition contains less than 10 picograms of DNA.
[0194] The volume of the partition can be less than 100 femtoliters, less than 1.0 picoliter, less than 10 picoliters, less than 100 picoliters, less than 1.0 nanoliter, less than 10 nanoliters, less than 100 nanoliters, less than 1.0 microliter, less than 10 microliters, less than 100 microliters, or less than 1.0 milliliter.
[0195] A barcode sequence can be provided in each partition. For each of two or more partitions containing a barcode sequence, the barcode sequence contained therein can include multiple copies of the same barcode sequence, or different barcode sequences from the same barcode sequence group.
[0196] After dispensing the particles into two or more partitions, the particles can be permeabilized by an incubation step of any of the methods described herein.
[0197] The sample of the particles can be digested with a protease digestion step (e.g., digestion with proteinase K). Optionally, the protease digestion step can be at least 10 seconds long, at least 30 seconds long, at least 60 seconds long, at least 5 minutes long, at least 10 minutes long, at least 30 minutes long, at least 60 minutes long, at least 3 hours long, at least 6 hours long, at least 12 hours long, or at least 24 hours long. This step can be performed before dispensing, after dispensing, before attaching the barcode sequence, after attaching the barcode sequence, and / or simultaneously with attaching the barcode sequence.
[0198] Attaching sequences by a combinatorial barcoding method
[0199] Methods of attaching barcode sequences can include combining at least two steps of a barcoding method, where a first barcoding step is performed, where a particulate sample is partitioned into two or more partitions, where each partition contains a different barcode sequence or a different group of barcode sequences, which are subsequently attached to sequences of target nucleic acids (e.g., genomic DNA) from the particulates contained within that partition, and where subsequently the barcoded nucleic acid molecules from at least two partitions are combined into a second sample mixture, and where subsequently the second sample mixture is partitioned into two or more new partitions, where each new partition contains a different barcode sequence or a different group of barcode sequences, which are subsequently attached to sequences of target nucleic acids (e.g., genomic DNA) from the particulates contained within the two or more new partitions.
[0200] Optionally, the combined barcoding method can include a first barcoding step, where: A) a first sample mixture containing at least first and second cycling particulates is partitioned into at least first and second original partitions (e.g., where at least the first cycling particulates from the sample are partitioned into the first original partition, and where at least the second cycling particulates from the sample are partitioned into the second original partition), where the first original partition contains a barcode sequence (or group of barcode sequences) different from the barcode sequence (or group of barcode sequences) contained within the second original partition, and where the barcode sequence (or barcode sequence from the group of barcode sequences) contained within the first original partition is attached to at least first and second target nucleic acid fragments of the first cycling particulates, and where the barcode sequence (or barcode sequence from the group of barcode sequences) contained within the second original partition is attached to at least first and second target nucleic acid fragments of the second cycling particulates; and where at least one cycling particulate contained within the first original partition is combined with at least one cycling particulate contained within the second original partition to produce a second sample mixture, and a second barcoding step, where: B) the particulates contained within the second sample mixture are partitioned into at least first and second new partitions (e.g., where at least the first cycling particulates from the second sample mixture are partitioned into the first new partition, and where at least the second cycling particulates from the second sample mixture are partitioned into the second new partition), where the first new partition contains a barcode sequence (or group of barcode sequences) different from the barcode sequence (or group of barcode sequences) contained within the second new partition, and the barcode sequence (or barcode sequence from the group of barcode sequences) contained within the first new partition is attached to at least first and second target nucleic acid fragments of the first cycling particulates, and where the barcode sequence (or barcode sequence from the group of barcode sequences) contained within the second new partition is attached to at least first and second target nucleic acid fragments of the second cycling particulates.
[0201] Optionally, the combinatorial barcoding method may include a first barcoding step, wherein: A) a first sample mixture containing at least first and second cyclic microparticles is dispensed into at least first and second original partitions (e.g., wherein at least the first cyclic microparticle from the sample is dispensed into the first original partition and wherein at least the second cyclic microparticle from the sample is dispensed into the second original partition), wherein the first original partition contains a barcode sequence (or group of barcode sequences) contained within a barcoded oligonucleotide that is different from the barcode sequence (or group of barcode sequences) contained within the barcoded oligonucleotide contained in the second original partition, and wherein the barcoded oligonucleotide contained within the first original partition is attached to at least first and second target nucleic acid fragments of the first cyclic microparticle, and wherein the barcoded oligonucleotide contained within the second original partition is attached to at least first and second target nucleic acid fragments of the second cyclic microparticle; and wherein at least one cyclic microparticle contained within the first original partition is combined with at least one cyclic microparticle contained within the second original partition to produce a second sample mixture, and a second barcoding step, wherein: B) the microparticles contained within the second sample mixture are dispensed into at least first and second new partitions (e.g., wherein at least the first cyclic microparticle from the second sample mixture is dispensed into the first new partition and wherein at least the second cyclic microparticle from the second sample mixture is dispensed into the second new partition), wherein the first new partition contains a barcode sequence (or group of barcode sequences) contained within a barcoded oligonucleotide that is different from the barcode sequence (or group of barcode sequences) contained within the barcoded oligonucleotide contained in the second new partition, and the barcoded oligonucleotide contained within the first new partition is attached to at least first and second target nucleic acid fragments of the first cyclic microparticle, and wherein the barcoded oligonucleotide contained within the second new partition is attached to at least first and second target nucleic acid fragments of the second cyclic microparticle.
[0202] Optionally, the combinatorial barcoding method may include a first barcoding step, wherein: A) a first sample mixture comprising at least first and second circular particles is dispensed into at least first and second original partitions (e.g., wherein at least first circular particles from the sample are dispensed into the first original partition, and wherein at least second circular particles from the sample are dispensed into the second original partition), wherein the first original partition contains a barcode sequence (or group of barcode sequences) contained within a barcoded oligonucleotide, which is different from the barcode sequence (or group of barcode sequences) contained within the barcoded oligonucleotide contained within the second original partition, and wherein the barcoded oligonucleotide contained within the first original partition is ligated to at least first and second target nucleic acid fragments of the first circular particle, and wherein the barcoded oligonucleotide contained within the second original partition is ligated to at least first and second target nucleic acid fragments of the second circular particle; and wherein at least one circular particle contained within the first original partition is combined with at least one circular particle contained within the second original partition to produce a second sample mixture, and a second barcoding step, wherein: B) the particles contained within the second sample mixture are dispensed into at least first and second new partitions (e.g., wherein at least first circular particles from the second sample mixture are dispensed into the first new partition, and wherein at least second circular particles from the second sample mixture are dispensed into the second new partition), wherein the first new partition contains a barcode sequence (or group of barcode sequences) contained within a barcoded oligonucleotide, which is different from the barcode sequence (or group of barcode sequences) contained within the barcoded oligonucleotide contained within the second new partition, and wherein the barcoded oligonucleotide contained within the first new partition is ligated to at least first and second target nucleic acid fragments of the first circular particle, and wherein the barcoded oligonucleotide contained within the second new partition is ligated to at least first and second target nucleic acid fragments of the second circular particle.
[0203] Optionally, the combinatorial barcoding method may include: A) a chemical crosslinking step, wherein a sample comprising at least first and second circular particles is crosslinked using a chemical crosslinking agent (e.g., formaldehyde), and subsequently optionally wherein the crosslinking step is ended by a quenching step, e.g., by mixing the sample with a glycine solution to quench the formaldehyde crosslinking step, and / or subsequently optionally permeabilizing the crosslinked particles (i.e., making genomic DNA (and / or other target nucleic acids) fragments physically accessible such that they can be further manipulated; e.g., such that they can be barcoded in the barcoding step); optionally, wherein any such permeabilization is carried out by incubation with a chemical surfactant (e.g., a non-ionic detergent); and B) a first barcoding step, wherein a first sample mixture comprising at least first and second circular particles is dispensed into at least first and second original partitions (e.g., wherein at least first circular particles from the sample are dispensed into the first original partition, and wherein at least second circular particles from the sample are dispensed into the second original partition), wherein the first original partition contains a barcode sequence (or group of barcode sequences) comprised within a barcoded oligonucleotide that is different from the barcode sequence (or group of barcode sequences) comprised within a barcoded oligonucleotide contained within the second original partition, and wherein the barcoded oligonucleotide contained within the first original partition is ligated to at least first and second target nucleic acid fragments of the first circular particles, and wherein the barcoded oligonucleotide contained within the second original partition is ligated to at least first and second target nucleic acid fragments of the second circular particles; and wherein at least one circular particle contained within the first original partition is combined with at least one circular particle contained within the second original partition to produce a second sample mixture, and C) a second barcoding step, wherein the particles contained within the second sample mixture are dispensed into at least first and second new partitions (e.g., wherein at least first circular particles from the second sample mixture are dispensed into the first new partition, and wherein at least second circular particles from the second sample mixture are dispensed into the second new partition), wherein the first new partition contains a barcode sequence (or group of barcode sequences) comprised within a barcoded oligonucleotide that is different from the barcode sequence (or group of barcode sequences) comprised within a barcoded oligonucleotide contained within the second new partition, and the barcoded oligonucleotide contained within the first new partition is ligated to at least first and second target nucleic acid fragments of the first circular particles, and wherein the barcoded oligonucleotide contained within the second new partition is ligated to at least first and second target nucleic acid fragments of the second circular particles.
[0204] Optionally, in any combinatorial barcoding method, prior to the first and / or second (and / or additional) barcoding steps, the method may include a step of crosslinking the cyclic particles and / or the target nucleic acid fragments (e.g., genomic DNA fragments) within one or more of the cyclic particles. This step can be carried out using a chemical crosslinking agent such as formaldehyde, paraformaldehyde, glutaraldehyde, disuccinimidyl glutarate, ethylene glycol bis(succinimidyl succinate), homobifunctional crosslinking agents or heterobifunctional crosslinking agents. This step can be carried out before any permeabilization step, after any permeabilization step, before any dispensing step, before any step of attaching barcode sequences, after any step of attaching barcode sequences, simultaneously with the attachment of barcode sequences, or any combination thereof. Any such crosslinking step can be further ended by a quenching step, for example, by mixing with a glycine solution to quench the formaldehyde crosslinking step. Any such crosslinking can be further removed prior to specific subsequent steps of the laboratory protocol, such as prior to primer extension, PCR or nucleic acid purification steps. The step of crosslinking by a chemical crosslinking agent is used to keep the genomic DNA (and / or other target nucleic acids) fragments within each particle physically close to each other, so that the sample can be manipulated and processed while maintaining the basic structural properties of the particles (i.e., while maintaining the physical proximity of the genomic DNA fragments from the same particle).
[0205] Optionally, in any combinatorial barcoding method, in a step after the chemical crosslinking step, the crosslinked particles can be permeabilized (i.e., such that the genomic DNA (and / or other target nucleic acid) fragments are physically accessible so that they can be further manipulated; for example, such that they can be barcoded in a barcoding step); this permeabilization can be carried out, for example, by incubation with a chemical surfactant (e.g., a non-ionic detergent). Optionally, the chemical surfactant used for such a permeabilization step may include Triton X-100 (C 14 H 22 O(C2H4O) n (n = 9 - 10)), NP-40, Tween 20, Tween 80, saponin, digitonin or sodium dodecyl sulfate.
[0206] Optionally, in any combination barcoding method, in any one or more steps after the chemical crosslinking step, the crosslinking can be partially or fully reversed (e.g., such that genomic DNA (and / or other target nucleic acid) fragments are physically more accessible and can be further manipulated; e.g., such that they can be barcoded in the barcoding step; this crosslinking reversal can be carried out, for example, by incubation at a high temperature, such as at least 45°C, at least 50°C, at least 55°C, at least 60°C, at least 65°C, at least 70°C, at least 75°C, at least 80°C, at least 85°C, or at least 90°C; furthermore, this crosslinking reversal can be carried out, for example, for a specific duration, such as at least 1 minute, at least 5 minutes, at least 10 minutes, at least 20 minutes, at least 30 minutes, at least 60 minutes, at least 2 hours, at least 3 hours, at least 5 hours, or at least 24 hours.
[0207] Optionally, in any combination barcoding method, in any one or more steps of attaching barcode sequences (e.g., any step of attaching and / or ligating barcoded oligonucleotides), and / or in any one or more steps of partitioning one or more samples (e.g., circulating microparticles) into different partitions, and / or in any one or more steps of combining two or more circulating microparticles into a single partition, and / or in any one or more chemical crosslinking steps, and / or in any one or more other steps, a purification treatment can be employed, wherein the microparticles are preferentially purified and separated relative to other components in the solution used in said steps. Any one or more such purification steps can include size exclusion chromatography methods. Any one or more such purification steps can include size centrifugation (e.g., differential centrifugation) methods.
[0208] Optionally, in any combination barcoding method, barcode sequences can be attached by any one or more of the methods described herein (e.g., single-strand ligation, double-strand ligation, blunt-end ligation, A-tailing ligation, sticky-end mediated ligation, hybridization, hybridization and extension, hybridization and extension and ligation, and / or transposition).
[0209] Optionally, during any step of any combination barcoding method, at least 2, at least 3, at least 5, at least 10, at least 20, at least 50, at least 100, at least 200, at least 500, at least 1000, at least 2000, at least 5000, at least 10,000, at least 50,000, at least 100,000, at least 500,000, or at least 1,000,000 circulating microparticles can be contained within a partition (and / or within each of at least the first and second partitions; and / or within any greater number of partitions). Preferably, at least 50 circulating microparticles can be contained within a partition (and / or within each of at least the first and second partitions; and / or within any greater number of partitions).
[0210] Optionally, during any step of any combinatorial barcoding method, at least 2, at least 3, at least 5, at least 10, at least 20, at least 50, at least 100, at least 200, at least 500, at least 1000, at least 2000, at least 5000, at least 10,000, at least 50,000, at least 100,000, at least 500,000, at least 1,000,000, at least 10,000,000, or at least 100,000,000 partitions may be employed (e.g., circulating microparticles may be assigned to the number of partitions). Preferably, during any step of any combinatorial barcoding method, at least 24 partitions may be used (e.g., circulating microparticles may be assigned to the number of partitions).
[0211] Optionally, in any step of any combinatorial barcoding method, a microparticle sample may be divided into partitions such that there is an average of less than 0.0001 microparticle, less than 0.001 microparticle, less than 0.01 microparticle, less than 0.1 microparticle, less than 1.0 microparticle, less than 10 microparticles, less than 100 microparticles, less than 1000 microparticles, less than 10,000 microparticles, less than 100,000 microparticles, less than 1,000,000 microparticles, less than 10,000,000 microparticles, or less than 100,000,000 microparticles in each partition. Preferably, there is an average of less than 1.0 microparticle in each partition.
[0212] Optionally, in any step of any combinatorial barcoding method, a microparticle solution may be divided into partitions such that there is an average of less than 1.0 attogram of DNA, less than 10 attograms of DNA, less than 100 attograms of DNA, less than 1.0 femtogram of DNA, less than 10 femtograms of DNA, less than 100 femtograms of DNA, less than 1.0 picogram of DNA, less than 10 picograms of DNA, less than 100 picograms of DNA, or less than 1.0 nanogram of DNA in each partition. Preferably, there is less than 10 picograms of DNA in each partition.
[0213] Optionally, in any step of any combinatorial barcoding method, the volume of a partition may be less than 100 femtoliters, less than 1.0 picoliter, less than 10 picoliters, less than 100 picoliters, less than 1.0 nanoliter, less than 10 nanoliters, less than 100 nanoliters, less than 1.0 microliter, less than 10 microliters, less than 100 microliters, or less than 1.0 milliliter.
[0214] Optionally, any combination of barcoding methods may include at least 2, at least 3, at least 4, at least 5, at least 10, at least 20, at least 30, at least 40, at least 50, at least 100, at least 500, or at least 1000 different barcoding steps. Each barcoding step may be as described herein for the first and second barcoding steps.
[0215] Optionally, in any combination of barcoding methods, any one or more of the dispensing steps may include random features; e.g., an estimated number (rather than an exact or precise number) of circular particles may be dispensed into one or more partitions; i.e., the number of circular particles in each partition may have statistical or probabilistic uncertainty (e.g., subject to Poisson loading and / or distribution statistics).
[0216] Optionally, in any combination of barcoding methods, a set of barcodes attached to a specific sequence (e.g., a sequence attached to a genomic DNA fragment; e.g., a set comprising a first barcode attached to the sequence during the first barcoding step and a second barcode attached to the sequence during the second barcoding step) may be used to associate sequences from a single particle and / or to associate sequences from a group of two or more particles. Optionally, in any combination of barcoding methods, a set of two (or more than two) identical barcodes may be attached to a specific sequence (e.g., a sequence attached to a genomic DNA fragment) from two or more circular particles (e.g., where the two or more circular particles are dispensed into first and second partitions of the same series during the first and second barcoding steps, respectively). Optionally, in any combination of barcoding methods, a set of two (or more than two) identical barcodes may be attached to a specific sequence (e.g., a sequence attached to a genomic DNA fragment) from only one circular particle (e.g., where only one circular particle is dispensed into first and second partitions of the same series during the first and second barcoding steps, respectively).
[0217] Optionally, in any combinatorial barcoding method, the number of partitions used in any one or more barcoding steps and the number of different barcoding steps can be combined combinatorially such that, on average, each group of two (or more) barcodes is attached to a sequence from only one circular particle. For example, for a sample containing 1000 circular particles, the first and second barcoding steps can each employ 100 partitions (and the associated barcodes contained therein); the total number of subsequent different barcode groups will then equal (100×100 =) 10,000 different barcode groups; compared to the 1000 circular particles contained in the original sample, each barcode group is thus, on average, attached to a sequence from only one (or conceptually less than one) circular particle. In some different embodiments of any combinatorial barcoding method, the number of partitions used in any one or more barcoding steps and / or the number of different barcoding steps can be increased and / or decreased to achieve a desired level of resolution and / or sensitivity (e.g., taking into account the need to analyze samples containing different numbers of circular particles, and / or different barcoding specificity requirements for different applications). Optionally, in certain applications, having an imperfect and / or inefficient barcoding method (e.g., where only a small fraction of the sequences from a particular particle are attached to barcodes in one or more barcoding steps; and / or e.g., where the same set of barcode sequences are attached to sequences from two or more circular particles) can enable sufficient molecular and / or information resolution to achieve a desired signal and / or sequencing readout.
[0218] Combinatorial barcoding methods can offer advantages over alternative barcoding methods in the form of reduced need for sophisticated and / or complex equipment to achieve a higher number of potential identification barcode groups for attaching barcodes to sequences from circular particles (e.g., from genomic DNA fragments). For example, a combinatorial barcoding method employing 96 different partitions in two different barcoding steps (e.g., readily achievable using a standard 96-well plate widely used in molecular biology) can achieve a net (96×96 =) 9216 different barcode groups; compared to alternative non-combinatorial methods, this significantly reduces the number of partitions required to perform such indexing. By increasing the number of barcoding steps and / or increasing the number of partitions used in one or more such barcoding steps, significantly higher levels of combinatorial indexing resolution can be further achieved. Additionally, combinatorial barcoding methods can eliminate the need for complex instrumentation (e.g., microfluidic instrumentation (e.g., 10×Genomics Chromium System)) used for alternative barcoding methods.
[0219] 6. Linked by spatial sequencing or in situ sequencing or in situ library construction
[0220] The present invention provides a method for preparing a sample for sequencing, wherein the sample comprises microparticles derived from blood, and wherein the microparticles comprise at least two target nucleic acid (e.g., genomic DNA) fragments, and wherein the method comprises: (a) preparing a sample for sequencing, wherein at least two target nucleic acid fragments of the microparticles are associated with each other by their proximity on a sequencing device to produce a group of at least two associated target nucleic acid fragments; and (b) sequencing each of the associated target nucleic acid fragments using a sequencing device to produce at least two associated sequence reads.
[0221] The nucleic acid sample may comprise at least two microparticles derived from blood, wherein each microparticle comprises at least two target nucleic acid (e.g., genomic DNA) fragments, and wherein the method comprises performing step (a) to produce a group of associated target nucleic acid fragments for each microparticle, and wherein the target nucleic acid fragments of each microparticle are spatially distinct on the sequencing device, and performing step (b) to produce associated sequence reads for each microparticle.
[0222] At least two fragments from the microparticles may maintain physical proximity to each other within or on the sequencing device itself, and wherein this physical proximity is known or determinable or observable by or during the operation of the sequencing device, and wherein this measure of physical proximity is used to associate at least two sequences.
[0223] The method may comprise performing sequencing using an in situ library construction method. In the method, intact or partially intact microparticles from the sample may be placed on a sequencer, and wherein two or more target nucleic acid (e.g., genomic DNA) fragments are processed into sequencing-ready templates within the sequencer, i.e., performing sequencing using an in situ library construction method. In situ library construction is described in Schwartz et al (2012) PNAS 109(46):18749-54).
[0224] The method may comprise in situ sequencing. In the method, the sample may remain intact (e.g., mostly or partially intact), and the target nucleic acid (e.g., genomic DNA) fragments within the microparticles are directly sequenced, e.g., using the 'FISSEQ' fluorescence in situ sequencing technique described in Lee et al (2014) Science, 343, 6177, 1360-1363).
[0225] Optionally, the particulate sample can be cross-linked with a chemical cross-linking agent and subsequently placed in or on a sequencing device and then maintained in physical proximity to each other. Optionally, two or more target nucleic acid (e.g., genomic DNA) fragments from the particulate placed in or on the sequencing device can subsequently have all or part of their sequences determined by a sequencing method. Optionally, such fragments can be sequenced by fluorescence in situ sequencing techniques, wherein the sequence of the fragment is determined by an optical sequencing method. Optionally, one or more coupling sequences, adapter sequences, or amplification sequences can be attached to the target nucleic acid fragment. Optionally, the fragment can be amplified during an amplification process, wherein the amplification products remain in physical proximity or physical contact with the fragments that amplified them. Optionally, these amplification products are subsequently sequenced by an optical sequencing method. Optionally, the amplification products are attached to a planar surface, such as a sequencing flow cell. Optionally, the amplification products generated from a single fragment each constitute a single cluster within the flow cell. Optionally, in any of the methods described above, the distance between any two or more sequencing molecules is known a priori by the configuration within the sequencing device or can be determined or observed during the sequencing process. Optionally, each sequencing molecule is mapped within a cluster field or a pixel array, wherein the distance between any two or more sequencing molecules is determined by the distance between the clusters or pixels. Optionally, any measurement or estimation of distance or proximity can be used to relate any two or more determined sequences.
[0226] Optionally, the sequences determined by any of the methods above can be further evaluated, wherein a measurement of the distance or proximity between two or more sequencing molecules is compared to one or more cut-off values or thresholds, and only the molecules within a specific range or above or below a determined specific threshold or cut-off value are determined to be informationally related. Optionally, two or more such cut-off values or thresholds or groups of ranges thereof can be used such that different degrees and / or categories and / or classifications of the relatedness of any two or more sequencing molecules can be determined.
[0227] 7. Linked by individual sequential methods
[0228] The present invention provides a method for preparing a sample for sequencing, wherein the sample comprises particulates derived from blood, and wherein the particulates comprise at least two target nucleic acid (e.g., genomic DNA) fragments, and wherein the method comprises: (a) preparing a sample for sequencing, wherein at least two target nucleic acid (e.g., genomic DNA) fragments of each particulate are linked by loading into separate sequencing processes to produce at least two groups of linked target nucleic acid fragments; and (b) sequencing each group of linked target nucleic acid fragments using a sequencing device to produce at least two groups of linked sequence reads.
[0229] The sample can comprise at least two blood-derived microparticles, wherein each microparticle comprises at least two target nucleic acid (e.g., genomic DNA) fragments, and the method can include performing step (a) to generate associated target nucleic acid fragments for each microparticle, wherein at least two target nucleic acid fragments for each microparticle are associated by loading into separate sequencing processes, and performing step (b) on each sequencing process to generate associated sequence reads for each microparticle.
[0230] In the method, the fragments of the first single microparticle (or group of microparticles) can be sequenced independently of the fragments of other microparticles, and the resulting sequence reads are informatically associated; the fragments contained within the second single microparticle (or group of microparticles) are sequenced independently of the first microparticle or group, and the resulting sequence reads are informatically associated.
[0231] Optionally, the first and second sequencing processes (of all the sequencing processes) are performed on different sequencers, and / or on the same sequencer but at two different times or within two different sequencing runs. Optionally, the first and second sequencing processes are performed on the same sequencer, but within two different regions, partitions, compartments, ducts, flow cells, lanes, nanopores, microcarriers, microcarrier arrays, or integrated circuits of the sequencer. Optionally, 3 or more, 10 or more, 1000 or more, 1,000,000 or more, or 1,000,000,000 or more microparticles or groups of microparticles can be associated by the above method.
[0232] 8. Amplify the original fragments prior to association
[0233] As would be understood by one of ordinary skill in the art, the term "fragment" as used herein (e.g., "genomic DNA fragment", or "target nucleic acid fragment" or "genomic DNA fragment from a microparticle") refers to the original fragment present in the microparticle, as well as portions, copies, or amplicons thereof, including copies of only a portion of the original fragment (e.g., its amplicon), and modified fragments or copies (e.g., fragments to which a coupling sequence has been attached). For example, the term genomic DNA fragment refers to the original genomic DNA fragment present in the microparticle, and, for example, a DNA molecule that can be prepared from the original genomic DNA fragment by primer extension reaction. As another example, the term mRNA fragment refers to the original mRNA fragment present in the microparticle, and, for example, a cDNA molecule that can be prepared from the original mRNA fragment by reverse transcription.
[0234] Prior to the step of attaching the barcode sequence, the method can further include the step of amplifying the original target nucleic acid fragments of the microparticle, e.g., by a primer extension step or a polymerase chain reaction step. The barcode sequence can subsequently be attached to the amplicon or copy of the original target nucleic acid fragment using any method described herein.
[0235] The primer extension step or polymerase chain reaction step can be carried out using one or more primers comprising a segment with one or more degenerate bases.
[0236] The primer extension step or polymerase chain reaction step can be carried out using one or more primers specific for a particular target nucleic acid sequence (such as a particular target genomic DNA sequence).
[0237] The amplification step can be carried out by a strand displacement polymerase (such as Phi29 DNA polymerase, or Bst polymerase or Bsm polymerase, or a modified derivative of phi29, Bst or Bsm polymerase). Amplification can be carried out by multiple displacement amplification reaction and a primer set comprising a region with one or more degenerate bases. Optionally, random hexamers, random heptamers, random octamers, random nonamers or random decamers are used.
[0238] The amplification step can include extending a single-strand nick in a fragment of the original target nucleic acid by a DNA polymerase. The nick can be generated by an enzyme with single-stranded DNA cleavage behavior or by a sequence-specific nicking restriction endonuclease.
[0239] The amplification step can include introducing at least one or more dUTP nucleotides into a DNA strand synthesized by replicating or amplifying at least a portion of one or more genomic DNA fragments by a DNA polymerase, and wherein a nick is generated by a uracil excision enzyme (such as uracil DNA glycosylase).
[0240] The amplification step can include generating a primer sequence on a nucleic acid containing a genomic DNA fragment, wherein the primer sequence is generated by a primase (such as Thermus Thermophilus PrimPol polymerase or TthPrimPol polymerase), and wherein a DNA polymerase is used to replicate at least one nucleotide of the sequence of the genomic DNA fragment using the primer sequence as a primer.
[0241] The amplification step can be carried out by a linear amplification reaction, such as an RNA amplification process carried out by an in vitro transcription process.
[0242] The amplification step can be carried out by a primer extension step or polymerase chain reaction step, and thus wherein the primers used are universal primers corresponding to one or more universal primer sequences. The universal primer sequences can be attached to the genomic DNA fragment by a ligation reaction, by primer extension or polymerase chain reaction or by an in vitro transposition reaction.
[0243] 9. Attaching a coupling sequence to the fragment before ligation
[0244] In any method, barcode sequences can be attached directly or indirectly (e.g., by annealing or ligation) to target nucleic acid (e.g., gDNA) fragments of the microparticles. The barcode sequences can be attached to a coupling sequence (e.g., a synthetic sequence) that has been attached to the fragment.
[0245] In methods that include bringing together at least two target nucleic acid fragments of the microparticles to produce a single nucleic acid molecule, a coupling sequence can first be attached to each of the at least two fragments, and the fragments can then be brought together by the coupling sequences.
[0246] The coupling sequence can be attached to the original target nucleic acid fragment of the microparticle or a copy or amplicon thereof.
[0247] The coupling sequence can be added to the 5' or 3' end of two or more fragments of the nucleic acid sample. In this method, the target region (of the barcoded oligonucleotide) can contain a sequence complementary to the coupling sequence.
[0248] The coupling sequence can be contained within a double-stranded coupling oligonucleotide or a single-stranded coupling oligonucleotide. The coupling oligonucleotide can be attached to the target nucleic acid by a double-stranded ligation reaction or a single-stranded ligation reaction. The coupling oligonucleotide can contain a single-stranded 5' or 3' region capable of ligating to the target nucleic acid, and the coupling sequence can be attached to the target nucleic acid by a single-stranded ligation reaction.
[0249] The coupling oligonucleotide can contain a blunt, recessed, or protruding 5' or 3' region capable of ligating to the target nucleic acid, and the coupling sequence can be attached to the target nucleic acid by a double-stranded ligation reaction.
[0250] The ends of the target nucleic acid fragment can be converted to blunt double-stranded ends in a blunt-ending reaction, and the coupling oligonucleotide can contain blunt double-stranded ends, and wherein the coupling oligonucleotide can be ligated to the target nucleic acid fragment in a blunt-end ligation reaction.
[0251] The ends of the target nucleic acid fragment can be converted to blunt double-stranded ends in a blunt-ending reaction and then to a form with a single 3' adenosine overhang, and wherein the coupling oligonucleotide can contain a double-stranded end with a single 3' thymine overhang capable of annealing to the single 3' adenosine overhang of the target nucleic acid fragment, and wherein the coupling oligonucleotide is ligated to the fragment of the target nucleic acid in a double-stranded A / T ligation reaction.
[0252] The target nucleic acid can be contacted with a restriction enzyme that digests the target nucleic acid at a restriction site to produce a ligation junction at the restriction site, and wherein the coupling oligonucleotide contains ends compatible with these ligation junctions, and wherein the coupling oligonucleotide is then ligated to the target nucleic acid in a double-stranded ligation reaction.
[0253] The coupling oligonucleotide can be attached by primer extension or polymerase chain reaction steps.
[0254] One or more oligonucleotides comprising a primer extension or polymerase chain reaction step may be used to attach a conjugate oligonucleotide, the primer extension or polymerase chain reaction step comprising one or more degenerate bases.
[0255] One or more oligonucleotides further comprising a primer or hybridization segment specific for a particular target nucleic acid sequence may be used to attach a conjugate oligonucleotide by primer extension or polymerase chain reaction steps.
[0256] A conjugate sequence may be added by a polynucleotide tailing reaction. The conjugate sequence may be added by a terminal transferase (e.g., terminal deoxynucleotidyl transferase). The conjugate sequence may be attached by a polynucleotide tailing reaction with terminal deoxynucleotidyl transferase, and wherein the conjugate sequence comprises at least two consecutive nucleotides of a homopolymeric sequence.
[0257] The conjugate sequence may comprise a homopolymeric 3' tail (e.g., a poly(A) tail). Optionally, in such a method, the target region (of the barcoded oligonucleotide) comprises a complementary homopolymeric 3' tail (e.g., a poly(T) tail).
[0258] The conjugate sequence may be contained within a synthetic transposon and may be attached by an in vitro transposition reaction.
[0259] The conjugate sequence may be attached to the target nucleic acid, and wherein the barcode oligonucleotide is attached to the target nucleic acid by at least one primer extension step or polymerase chain reaction step, and wherein the barcode oligonucleotide comprises a region complementary to the conjugate sequence having a length of at least one nucleotide. Optionally, the complementary region is located at the 3' end of the barcode oligonucleotide. Optionally, the complementary region has a length of at least 2 nucleotides, at least 5 nucleotides, at least 10 nucleotides, at least 20 nucleotides, or at least 50 nucleotides.
[0260] 10. Optional additional steps of the method
[0261] The method may include determining the presence or absence of at least one modified nucleotide or nucleobase in one or more genomic DNA fragments from a sample comprising one or more circulating microparticles. The method may include measuring a modified nucleotide or nucleobase in a genomic DNA fragment of a circulating microparticle (e.g., measuring a modified nucleotide or nucleobase). The measured value may be the total value of the analyzed genomic DNA fragment (i.e., the associated genomic DNA fragment) of the circulating microparticle and / or the measured value may be the value for each analyzed genomic DNA fragment. The modified nucleotide or nucleobase may be 5-methylcytosine or 5-hydroxymethylcytosine.
[0262] The measurement of modified nucleotides or nucleobases in one or more genomic DNA fragments from circulating microparticles enables a variety of molecular and informational analyses, which can complement the measurement of the sequences of the fragments themselves. In one aspect, the measurement of so-called "epigenetic" marks within genomic DNA fragments from circulating microparticles (i.e., the measurement of the "epigenome") enables comparison (and / or mapping relative thereto) with a reference epigenetics sequence and / or a list of reference epigenetics sequences. This enables an "orthogonal" form of analysis of the sequences of genomic fragments from circulating microparticles compared to measuring only the standard 4 (unmodified) bases and / or their conventional "genetic" sequences. In addition, the measurement of modified nucleotides and / or nucleobases can enable more precise determination and / or estimation of the type of cell and / or tissue from which one or more circulating microparticles are derived. Since different cell types in vivo exhibit different epigenetic signatures, measurement of the epigenome of genomic DNA fragments from circulating microparticles can thus permit more precise mapping of such microparticles to cell types. In the method, the epigenetic measurement of genomic DNA fragments from circulating microparticles can be compared (e.g., mapped thereto) with a list (or lists) of reference epigenetics sequences corresponding to methylation and / or hydroxymethylation within a particular specific tissue. This can enable the elucidation and / or enrichment of microparticles (e.g., the associated sequence sets from specific microparticles) from a particular tissue type and / or specific healthy and / or diseased tissues (e.g., cancerous tissues). For example, the measurement of modified nucleotides or nucleobases in genomic DNA fragments of circulating microparticles can enable the identification of associated sequences (or associated sequence reads) of genomic DNA fragments derived from cancer cells. In another example, the measurement of modified nucleotides or nucleobases in genomic DNA fragments of circulating microparticles can enable the identification of associated sequences (or associated sequence reads) of genomic DNA fragments derived from fetal cells. The absolute amount of a particular modified nucleotide or nucleobase can be associated with the health status and / or disease within a particular tissue. For example, the level of 5-hydroxymethylcytosine in cancerous tissues is strongly altered compared to normal healthy tissues; thus, the measurement of 5-hydroxy-methylcytosine in genomic DNA fragments from circulating microparticles can enable more precise detection and / or analysis of circulating microparticles derived from cancer cells.
[0263] The method can include the measurement of 5-methylcytosine in genomic DNA fragments of circulating microparticles (e.g., measuring 5-methylcytosine in genomic DNA fragments of circulating microparticles). The method can include the measurement of 5-hydroxy-methylcytosine in genomic DNA fragments of circulating microparticles (e.g., measuring 5-hydroxy-methylcytosine in genomic DNA fragments of circulating microparticles).
[0264] The method may include measuring 5-methylcytosine in genomic DNA fragments of circulating microparticles (e.g., measuring 5-methylcytosine in genomic DNA fragments of circulating microparticles), wherein the measurement is performed using an enrichment probe that specifically or preferentially binds 5-methylcytosine in genomic DNA fragments as compared to other modified or unmodified bases. The method may include measuring 5-hydroxymethylcytosine in genomic DNA fragments of circulating microparticles (e.g., measuring 5-hydroxymethylcytosine in genomic DNA fragments of circulating microparticles), wherein the measurement is performed using an enrichment probe that specifically or preferentially binds 5-hydroxymethylcytosine in genomic DNA fragments as compared to other modified or unmodified bases.
[0265] The method may include measuring 5-methylcytosine in genomic DNA fragments of two or more circulating microparticles (e.g., measuring 5-methylcytosine in genomic DNA fragments of a first circulating microparticle and measuring 5-methylcytosine in genomic DNA fragments of a second circulating microparticle). The method may include measuring 5-hydroxymethylcytosine in genomic DNA fragments of two or more circulating microparticles (e.g., measuring 5-hydroxymethylcytosine in genomic DNA fragments of a first circulating microparticle and measuring 5-hydroxymethylcytosine in genomic DNA fragments of a second circulating microparticle).
[0266] The method may include measuring 5-methylcytosine in genomic DNA fragments of two or more circulating microparticles (e.g., measuring 5-methylcytosine in genomic DNA fragments of a first circulating microparticle and measuring 5-methylcytosine in genomic DNA fragments of a second circulating microparticle), wherein the measurement is performed using an enrichment probe that specifically or preferentially binds 5-methylcytosine in genomic DNA fragments as compared to other modified or unmodified bases. The method may include measuring 5-hydroxymethylcytosine in genomic DNA fragments of two or more circulating microparticles (e.g., measuring 5-hydroxymethylcytosine in genomic DNA fragments of a first circulating microparticle and measuring 5-hydroxymethylcytosine in genomic DNA fragments of a second circulating microparticle), wherein the measurement is performed using an enrichment probe that specifically or preferentially binds 5-hydroxymethylcytosine in genomic DNA fragments as compared to other modified or unmodified bases.
[0267] The method may include measuring 5-methylcytosine in genomic DNA fragments of circulating microparticles (e.g., measuring 5-methylcytosine in genomic DNA fragments of circulating microparticles), wherein the measurement is performed using a bisulfite conversion method or an oxidative bisulfite conversion method. The method may include measuring 5-hydroxymethylcytosine in genomic DNA fragments of circulating microparticles (e.g., measuring 5-hydroxymethylcytosine in genomic DNA fragments of circulating microparticles), wherein the measurement is performed using a bisulfite conversion method or an oxidative bisulfite conversion method.
[0268] The method may include measuring 5-methylcytosine in genomic DNA fragments of two or more circulating microparticles (e.g., measuring 5-methylcytosine in genomic DNA fragments of a first circulating microparticle and measuring 5-methylcytosine in genomic DNA fragments of a second circulating microparticle), wherein the measurement is performed using a bisulfite conversion method or an oxidative bisulfite conversion method. The method may include measuring 5-hydroxymethylcytosine in genomic DNA fragments of two or more circulating microparticles (e.g., measuring 5-hydroxymethylcytosine in genomic DNA fragments of a first circulating microparticle and measuring 5-hydroxymethylcytosine in genomic DNA fragments of a second circulating microparticle), wherein the measurement is performed using a bisulfite conversion method or an oxidative bisulfite conversion method.
[0269] Optionally, sequences of two or more components from a sample comprising one or more circular particles can be determined as associated to determine the presence or absence of at least one modified nucleotide or nucleobase in one or more genomic DNA fragments from said sample. For example, an enrichment step can be performed to enrich genomic DNA fragments in the sample that contain modified bases (such as 5-methylcytosine or 5-hydroxymethylcytosine), where the first component of the sample containing genomic fragments that has been enriched by said enrichment step can be sequenced, and also the second component of the sample containing genomic fragments that has not been enriched by said enrichment step can be sequenced (such as in an independent sequencing reaction). Optionally, the second component of the sample can comprise non-enriched and / or supernatant fractions generated during the enrichment process (such as fractions that were not bound by enrichment probes or affinity probes during the enrichment process). Optionally, the original sample can be divided into a first and a second subsample, where the first subsample is used to perform the enrichment step to generate the first component of the sample, and where the second component of the sample can comprise the second non-enriched subsample. Any combination of two or more enriched and / or non-enriched and / or transformed (such as bisulfite transformation and / or oxidative bisulfite transformation) and / or untransformed components of the sample can be sequenced. For example, a sample comprising one or more circular particles can be used to generate three components, such as a component enriched in 5-methylcytosine DNA (or, a component that has been bisulfite-transformed), a component enriched in 5-hydroxymethylcytosine (or, a component that has been oxidative bisulfite-transformed), and a non-enriched (and / or untransformed) component. Optionally, any such two or more components of the sample can be sequenced separately in independent sequencing reactions (such as in independent flow cells, or in independent lanes of a single flow cell). Optionally, any such two or more parts of the sample can be attached to an identifying barcode sequence (such as, which identifies a given sequence within an enriched or non-enriched component of the sample), and subsequently sequenced in the same sequencing process (such as in the same flow cell or lane of a flow cell).
[0270] Optionally, any method of associating sequences described herein (such as, by attaching a barcode sequence, such as by attaching a barcode sequence from a polymeric barcoding reagent or by attaching barcode sequences from libraries of two or more polymeric barcoding reagents) can be performed prior to any such enrichment and / or molecular transformation step (such as, where such an associating method is performed on an original sample comprising at least one circular particle or at least two circular particles, where the associated sequence is subsequently used as an input sequence for the enrichment or molecular transformation process).
[0271] For example, a sample containing two or more circular particles can be attached to barcode sequences from a library of two or more polymeric barcoding reagents, where the first and second barcode sequences from a first polymeric barcoding reagent are attached to the first and second genomic DNA fragments from a first circular particle, and where the first and second barcode sequences from a second polymeric barcoding reagent are attached to the first and second genomic DNA fragments from a second circular particle, and where the resulting barcode-attached genomic DNA fragments are enriched for 5-methylcytosine (and / or 5-hydroxymethylcytosine), and where the enriched genomic DNA fragments are subsequently sequenced, where the barcode sequences are subsequently used to determine which of the enriched fragments are attached to barcodes from the same polymeric barcoding reagent and thereby predict (or determine) which of the enriched fragments are contained within the same circular particle. In this example, a second sequencing reaction can also be performed on the non-enriched genomic DNA fragments (e.g., by sequencing genomic fragments within the supernatant fraction of the enrichment step (i.e., the uncaptured non-enriched fraction)), where the barcode sequences are subsequently used to determine which of the non-enriched fragments are attached to barcodes from the same polymeric barcoding reagent and thereby predict (or determine) which of the non-enriched fragments are contained within the same circular particle. In this example, if both the enriched and non-enriched genomic DNA fragments are so sequenced, then it is thus possible to predict (or determine) which of the enriched and which of the non-enriched fragments are attached to barcodes from the same polymeric barcoding reagent and thereby predict (or determine) which of the enriched and which of the non-enriched fragments are contained within the same circular particle. Methods similar to this example can also be used, e.g., by employing one or more molecular conversion processes, and / or e.g., by preparing, analyzing, or sequencing three or more components of a sample (e.g., a 5-methylcytosine-enriched component, a 5-hydroxymethylcytosine-enriched component, and a non-enriched component).
[0272] Optionally, any method of associating sequences described herein (e.g., by attaching barcode sequences, e.g., by attaching barcode sequences from a library of polymeric barcoding reagents or two or more polymeric barcoding reagents) can be performed after any such enrichment and / or molecular conversion step (e.g., where an enrichment step is performed to enrich genomic DNA fragments containing 5-methylcytosine or containing 5-hydroxymethylcytosine, and where the genomic DNA fragments enriched by this process are associated by any method described herein).
[0273] The method may include determining the presence or absence of at least one modified nucleotide or nucleobase in a genomic DNA fragment, wherein an enrichment step is performed to enrich the genomic DNA fragments containing the modified base. Such a modified base may include one or more of 5-methylcytosine, or 5-hydroxymethylcytosine, or any other modified base. Such an enrichment step may be performed by an enrichment probe that binds specifically or preferentially to the modified base as compared to other modified or unmodified bases, such as an antibody, an enzyme, an enzyme fragment, or other protein, or an adaptor, or any other probe. Such an enrichment step may be performed by an enzyme capable of enzymatically modifying a DNA molecule containing the modified base, such as a glucosyltransferase, such as 5-hydroxymethylcytosine glucosyltransferase. Optionally, 5-hydroxymethylcytosine glucosyltransferase may be used to determine the presence of 5-hydroxymethylcytosine within a genomic DNA fragment, wherein the 5-hydroxymethylcytosine glucosyltransferase is used to transfer a glucose moiety from uridine diphosphate glucose to the modified base within the genomic DNA fragment to produce a glucosyl-5-hydroxymethylcytosine base, optionally wherein the glucosyl-5-hydroxymethylcytosine base is subsequently detected, such as by using a glucosyl-5-hydroxymethylcytosine-sensitive restriction enzyme, wherein genomic DNA fragments resistant to digestion by the glucosyl-5-hydroxymethylcytosine-sensitive restriction enzyme are considered to contain the modified 5-hydroxymethylcytosine base; optionally, the genomic DNA fragments resistant to digestion may be sequenced by any method described herein to determine their sequence. Optionally, if barcode sequences are attached, the enrichment step may be performed before or after the step of attaching the barcode sequences. Optionally, if two or more sequences of genomic DNA fragments from microparticles are attached to each other, the enrichment step may be performed before or after the step of attaching these sequences to each other. Any method of measuring at least one modified nucleotide or nucleobase in a genomic DNA fragment using an enrichment probe may be performed using commercially available enrichment probes or other products, such as commercially available antibodies, such as anti-5-hydroxymethylcytosine antibody ab178771 (Abcam), or such as anti-5-methylcytosine antibody ab10805 (Abcam). In addition, commercially available products and / or kits may also be used for other steps of such methods, such as Protein A or Protein G Dynabeads (ThermoFisher) for binding, recovering, and processing / washing antibodies and / or fragments bound thereto.
[0274] The method may include determining the presence or absence of at least one modified nucleotide or nucleobase in a genomic DNA fragment, wherein a molecular conversion step is performed to convert the modified base into a different modified or unmodified nucleobase, which can be detected during determination of the nucleic acid sequence. The conversion step may include a bisulfite conversion step, an oxidative bisulfite conversion step, or any other molecular conversion step. Optionally, if a barcode sequence is attached, the enrichment step may be performed before or after the step of attaching the barcode sequence. Optionally, if two or more sequences of genomic DNA fragments from a microparticle are attached to each other, the enrichment step may be performed before or after the step of attaching these sequences to each other. Any method of measuring at least one modified nucleotide or nucleobase in a genomic DNA fragment using a molecular conversion step can be performed using commercially available molecular conversion kits, such as the EpiMark Bisulfite Conversion Kit (New England Biolabs) or the TruMethyl Seq Oxidative Bisulfite Sequencing Kit (Cambridge Epigenetix).
[0275] In any method in which a molecular conversion step is performed, one or more adapter oligonucleotides may be attached to one or both ends of a genomic DNA fragment (and / or a collection of genomic DNA fragments within a sample) after the molecular conversion process. For example, a single-stranded adapter oligonucleotide (e.g., containing a binding site for a primer for amplification (e.g., by PCR amplification)) may be ligated to one or both ends of the converted genomic DNA fragment (and / or a collection of genomic DNA fragments in a sample) using a single-stranded ligase. Optionally, a barcode sequence and / or an adapter sequence (e.g., within a barcoded oligonucleotide) may be attached to one end of a genomic DNA fragment (and / or a collection of genomic DNA fragments within a sample) before the molecular conversion step, and subsequently an adapter oligonucleotide may be attached to the second end of the genomic DNA fragment after the molecular conversion process. Optionally, the second end may include an end generated during the molecular conversion process (i.e., where a fragment of genomic DNA has undergone a fragmentation process, thus generating one or more new ends of the fragment relative to its original fragment). This method of attaching adapter oligonucleotides may have the benefit of allowing genomic DNA fragments that have been fragmented and / or degraded during the molecular conversion process to be further amplified and / or analyzed and / or sequenced.
[0276] In any method performing a molecular conversion step, any adapter oligonucleotide, and / or barcoded oligonucleotide, and / or barcode sequence, and / or any coupling sequence and / or any coupling oligonucleotide may comprise one or more synthetic 5-methylcytosine nucleotides. Optionally, any adapter oligonucleotide, and / or barcoded oligonucleotide, and / or barcode sequence, and / or any coupling sequence and / or any coupling oligonucleotide may be configured such that any or all of the cytosine nucleotides contained therein are synthetic 5-methylcytosine nucleotides. Optionally, any adapter oligonucleotide, and / or barcoded oligonucleotide, and / or barcode sequence, and / or any coupling sequence and / or any coupling oligonucleotide comprising one or more synthetic 5-methylcytosine nucleotides may be attached to a genomic DNA fragment prior to the molecular conversion step; alternatively and / or additionally, it may be attached to the genomic DNA fragment after the molecular conversion step. Such synthetic 5-methylcytosine nucleotides within the adapter and / or oligonucleotide and / or sequence may have the benefit of reducing or minimizing their degradation and / or fragmentation during the molecular conversion process (such as a bisulfite conversion process), as they are resistant to degradation during such a process.
[0277] The method may include determining the presence or absence of at least one modified nucleotide or nucleobase in a genomic DNA fragment, wherein the modified nucleotide or nucleobase (such as 5-methylcytosine or 5-hydroxymethylcytosine) is determined or detected by a sequencing reaction. Optionally, the sequencing reaction may be performed by a nanopore-based sequencer, such as the Minion, GridION X5, PromethION, and / or SmidgION sequencers produced by Oxford Nanopore Technologies, wherein the presence of the modified nucleotide or nucleobase is determined during the translocation of the genomic DNA fragment through the nanopore within the sequencer and by analyzing the current signal passing through the nanopore device during the translocation of the genomic DNA fragment. Optionally, the sequencing reaction may be performed by a zero-mode-waveguide-based sequencing instrument, such as the Sequel or RSII sequencers produced by Pacific Biosciences, wherein the presence of the modified nucleotide or nucleobase is determined during the process of synthesizing a copy of at least a portion of the genomic DNA fragment within the zero-mode waveguide in the sequencer and by analyzing the optical signal from the zero-mode waveguide during the process of replicating at least a portion of the genomic DNA fragment.
[0278] In any method that performs an enrichment step and / or a molecular conversion step, the enrichment and / or conversion can be incomplete and / or have an efficiency of less than 100%. For example, a molecular conversion process can be performed such that less than 100% of a particular class of target modified nucleotides (e.g., 5-methylcytosine or 5-hydroxymethylcytosine) are converted by the molecular conversion process (e.g., bisulfite conversion or oxidative bisulfite conversion). For example, about 99%, or about 95%, or about 90%, or about 80%, or about 70%, or about 60%, or about 50%, or about 40%, or about 25%, or about 10% of such target modified nucleotides can be converted during such a molecular conversion process. Such an incomplete molecular conversion process can be carried out by limiting the duration of the molecular conversion process (e.g., by making the duration shorter than the standard time for achieving complete or near-complete efficiency of the molecular conversion process) such that, on average, the target conversion efficiency is achieved. Such an incomplete molecular conversion process can have the benefit of reducing the amount of sample degradation / fragmentation and / or sample loss, which is, for example, characteristic of many molecular conversion processes (e.g., bisulfite conversion).
[0279] Similarly, in any method of performing an enrichment step, the enrichment can be incomplete and / or less than 100% efficient. For example, an enrichment step for 5-methylcytosine (and / or 5-hydroxymethylcytosine) can be performed, where about 99%, or about 95%, or about 90%, or about 80%, or about 70%, or about 60%, or about 50%, or about 40%, or about 25%, or about 10% of the genomic DNA fragments containing such a target modified nucleotide are captured and recovered during the enrichment step (e.g., an enrichment step using an affinity probe (e.g., an antibody specific for the target modified nucleotide)). Optionally, the incomplete enrichment can be performed by limiting and / or reducing the amount and / or concentration of the affinity probe used during the enrichment process (e.g., by empirically testing the efficiency of such capture using different amounts and / or concentrations of the affinity probe, and optionally by using a DNA sequence containing a known modified nucleotide profile as an evaluation metric for the empirical test). Optionally, the incomplete enrichment can be performed by limiting and / or reducing the duration, where the affinity probe is used to bind and / or capture target genomic DNA fragments within the enrichment process (i.e., by using different incubation times where the affinity probe is capable of interacting with potential target genomic DNA fragments within the sample); e.g., by empirically testing the efficiency of such capture using different incubation durations, and optionally by using a DNA sequence containing a known modified nucleotide profile as an evaluation metric for the empirical test). Such incomplete enrichment can have the benefit of reducing false positive molecular signals (e.g., where fragments of genomic DNA are captured during the enrichment process, but where the fragments do not have the desired target modified nucleotide). Additionally, the incomplete enrichment can have the benefit of reducing the cost and complexity of the enrichment process itself.
[0280] The method can include performing a sequence enrichment or sequence capture step, where one or more specific genomic DNA sequences are enriched from genomic DNA fragments. This step can be performed by any method of performing sequence enrichment, such as using a DNA oligonucleotide complementary to the sequence, or an RNA oligonucleotide complementary to the sequence, or by a step of performing a primer extension target enrichment step, or by a step of using a set of molecular inversion probes, or by a step of using a set of padlock probes. Optionally, if barcode sequences are attached, this enrichment step can be performed before or after the step of attaching the barcode sequences. Optionally, if two or more sequences of genomic DNA fragments from microparticles are attached to each other, this enrichment step can be performed before or after the step of attaching these sequences to each other.
[0281] The method can include enriching at least 1, at least 5, at least 10, at least 50, at least 100, at least 500, at least 1000, at least 5000, at least 10,000, at least 100,000, at least 1,000,000, or at least 10,000,000 different genomic DNA fragments.
[0282] In the method, each unique input molecule can be sequenced on average at least 1.0 times, at least 1.5 times on average, at least 2.0 times on average, at least 3.0 times on average, at least 5.0 times on average, at least 10.0 times on average, at least 20.0 times on average, at least 50.0 times on average, or at least 100 times in the sequencing reaction. Optionally, unique input molecules that are sequenced at least twice (i.e., redundant sequencing using at least two sequence reads) in the sequencing reaction are used to detect and / or remove errors or inconsistencies in the sequencing between the at least two sequence reads generated by the sequencing reaction.
[0283] Before performing the sequencing reaction and / or before performing the amplification reaction, a nucleotide repair reaction can be carried out, in which damaged and / or excised bases or oligonucleotides are removed and / or repaired. Optionally, the repair reaction can be carried out in the presence of one or more of the following: Thermus aquaticus DNA ligase, Escherichia coli endonuclease IV, Bacillus stearothermophilus DNA polymerase, Escherichia coli formamidopyrimidine [fapy]-DNA glycosylase, Escherichia coli uracil-DNA glycosylase, T4 endonuclease V, and Escherichia coli endonuclease VIII.
[0284] In the method, before the sequencing step and / or before the amplification step (such as a PCR amplification step), a universal adapter sequence (such as one or two universal adapter sequences) can be attached. Optionally, one or more such universal adapter sequences can be added by random priming or gene-specific primer extension steps, by an in vitro transposition reaction (where one or more of the universal adapter sequences are included within a synthetic transposome), by a double-stranded or single-stranded ligation reaction (with or without a previous fragmentation step, such as a chemical fragmentation step, a sonic or mechanical fragmentation step, or an enzymatic fragmentation step; and optionally with or without blunt-ending and / or 3’A-tailing steps).
[0285] Barcode sequences comprising enzymatically generated copies or enzymatically generated complementary sequences
[0286] One or more barcode sequences may be included within an oligonucleotide that contains enzymatically produced copies or enzymatically produced complementary sequences of the barcode sequences (e.g., within a barcoded oligonucleotide).
[0287] Optionally, one or more barcode sequences may be included within a barcoded oligonucleotide, wherein the barcode region of the barcoded oligonucleotide contains enzymatically produced copies or enzymatically produced complementary sequences of the barcode sequences. Optionally, one or more barcode sequences may be included within a barcoded oligonucleotide, wherein the barcode region of the barcoded oligonucleotide contains enzymatically produced complementary sequences of the barcode sequences contained within a barcode molecule. Optionally, one or more barcode sequences may be included within a barcoded oligonucleotide, wherein the barcode region of the barcoded oligonucleotide contains enzymatically produced copies of the barcode sequences contained within a barcode molecule.
[0288] Optionally, one or more barcode sequences may be included within a barcoded oligonucleotide, wherein the barcode region of the barcoded oligonucleotide contains enzymatically produced complementary sequences of the barcode sequences contained within a polymeric barcode molecule. Optionally, one or more barcode sequences may be included within a barcoded oligonucleotide, wherein the barcode region of the barcoded oligonucleotide contains enzymatically produced copies of the barcode sequences contained within a polymeric barcode molecule.
[0289] Optionally, one or more barcode sequences may be included within a first barcoded oligonucleotide, wherein the barcode region of the barcoded oligonucleotide contains enzymatically produced complementary sequences of the barcode sequences contained within a second barcoded oligonucleotide. Optionally, one or more barcode sequences may be included within a first barcoded oligonucleotide, wherein the barcode region of the barcoded oligonucleotide contains enzymatically produced copies of the barcode sequences contained within a second barcoded oligonucleotide.
[0290] Any enzymatic method for copying, replicating, and / or synthesizing nucleic acid sequences may be used to produce enzymatically produced copies or enzymatically produced complementary sequences of the barcode sequences. Optionally, a primer extension method may be employed. Optionally, a primer extension method may be employed, wherein the barcode sequences contained within a barcode molecule (and / or contained within a polymeric barcode molecule, and / or contained within a barcoded oligonucleotide) are replicated during the primer extension step, and wherein the resulting primer extension product of the primer extension step contains all or a portion of the barcode sequence (e.g., contains all or a portion of the barcoded oligonucleotide), which is subsequently attached to a nucleic acid sequence from a circulating particle (e.g., attached to a sequence of a genomic DNA fragment from a circulating particle).
[0291] Optionally, a polymerase chain reaction (PCR) method may be employed. Optionally, a polymerase chain reaction (PCR) method may be employed, wherein the barcode sequence contained within the barcode molecule (and / or within the polymeric barcode molecule, and / or within the barcoded oligonucleotide) is replicated during the PCR extension step, and wherein the resulting extension product of the PCR extension step contains all or a portion of the barcode sequence (e.g., contains all or a portion of the barcoded oligonucleotide), which is subsequently attached to a nucleic acid sequence from a circulating particle (e.g., attached to the sequence of a genomic DNA fragment from a circulating particle). Optionally, a polymerase chain reaction (PCR) method may be employed, wherein the barcode sequence contained within the barcode molecule (and / or within the polymeric barcode molecule, and / or within the barcoded oligonucleotide) is replicated in at least two consecutive PCR extension steps (e.g., replicated using at least a first PCR cycle and subsequently a second PCR cycle), and wherein each of the at least two resulting PCR extension products contains all or a portion of the barcode sequence (e.g., contains all or a portion of the barcoded oligonucleotide), which is subsequently attached to a nucleic acid sequence from a circulating particle (e.g., attached to the sequence of a genomic DNA fragment from a circulating particle).
[0292] Optionally, a rolling-circle amplification (RCA) method may be employed. Optionally, a rolling-circle amplification (RCA) method may be employed, wherein the barcode sequence contained within the barcode molecule (and / or within the polymeric barcode molecule, and / or within the barcoded oligonucleotide) is replicated during the rolling-circle amplification step, and wherein the resulting extension product of the rolling-circle amplification step contains all or a portion of the barcode sequence (e.g., contains all or a portion of the barcoded oligonucleotide, and / or contains all or a portion of the barcode molecule, and / or contains all or a portion of the polymeric barcode molecule), which is subsequently attached to a nucleic acid sequence from a circulating particle (e.g., attached to the sequence of a genomic DNA fragment from a circulating particle).
[0293] Optionally, a rolling-circle amplification (RCA) method may be employed, wherein the barcode sequence contained within the polymeric barcode molecule is replicated during the rolling-circle amplification step, and wherein the resulting extension product of the rolling-circle amplification step contains a second polymeric barcode molecule, and wherein the second polymeric barcode molecule is used as a template to synthesize at least one barcoded oligonucleotide (wherein such barcoded oligonucleotide may be generated by any method described herein; e.g., generating at least one barcoded oligonucleotide by primer extension using the second polymeric barcode molecule as a template or by primer extension and ligation using the second polymeric barcode molecule as a template), which is subsequently attached to a nucleic acid sequence from a circulating particle (e.g., attached to the sequence of a genomic DNA fragment from a circulating particle).
[0294] Optionally, any such method for the enzymatic production of copies or enzymatically produced complementary sequences of the barcode sequences can be carried out in a single reaction volume. Optionally, any such method for the enzymatic production of copies or enzymatically produced complementary sequences of the barcode sequences can be carried out in two or more different reaction volumes (i.e., in two or more different partitions). Optionally, any such method for the enzymatic production of copies or enzymatically produced complementary sequences of the barcode sequences can be carried out in at least 3, at least 5, at least 10, at least 50, at least 100, at least 500, at least 1000, at least 10,000, at least 100,000, at least 1,000,000, at least 10,000,000, or at least 100,000,000 different reaction volumes (and / or partitions).
[0295] Optionally, any such method for the enzymatic production of copies or enzymatically produced complementary sequences of the barcode sequences can be carried out in a reaction volume containing nucleic acid sequences from one or more circular microparticles (e.g., in a reaction volume containing one or more circular microparticles). Optionally, the method for the enzymatic production of copies or enzymatically produced complementary sequences of the barcode sequences can be carried out in a first reaction volume containing nucleic acid sequences from the first circular microparticles of the sample (e.g., genomic DNA fragments from the first circular microparticles of the sample, and / or the first circular microparticles from the sample), and in a second reaction volume containing nucleic acid sequences from the second circular microparticles of the sample (e.g., genomic DNA fragments from the second circular microparticles of the sample, and / or the second circular microparticles from the sample).
[0296] Optionally, the method for the enzymatic production of copies or enzymatically produced complementary sequences of the barcode sequences can be carried out in N different reaction volumes, where each such reaction volume contains at least one barcode sequence and also contains nucleic acid sequences from the circular microparticles of the sample (e.g., also contains genomic DNA fragments from the circular microparticles of the sample, and / or also contains the circular microparticles from the sample), where N is at least 2, at least 3, at least 5, at least 10, at least 50, at least 100, at least 500, at least 1000, at least 10,000, at least 100,000, at least 1,000,000, at least 10,000,000, or at least 100,000,000. Optionally, the barcode sequences contained in the N different reaction volumes can together contain at least 2, at least 3, at least 5, at least 10, at least 50, at least 100, at least 500, at least 1000, at least 10,000, at least 100,000, at least 1,000,000, at least 10,000,000, or at least 100,000,000 different barcode sequences.
[0297] Optionally, a method for generating enzymatically produced copies or enzymatically produced complementary sequences of barcode sequences can be performed in a first reaction volume that contains a first barcode sequence and also contains the nucleic acid sequence of the first circulating microparticles of the sample (e.g., also contains a genomic DNA fragment from the first circulating microparticles of the sample, and / or also contains the first circulating microparticles from the sample), and in a second reaction volume that contains a second barcode sequence and also contains the nucleic acid sequence of the second circulating microparticles of the sample (e.g., also contains a genomic DNA fragment from the second circulating microparticles of the sample, and / or also contains the second circulating microparticles from the sample), wherein the first barcode sequence is different from the second barcode sequence.
[0298] Optionally, a method for generating enzymatically produced copies or enzymatically produced complementary sequences of barcode sequences can be performed in a first reaction volume that contains the nucleic acid sequence of the first circulating microparticles of the sample (e.g., contains a genomic DNA fragment of the first circulating microparticles of the sample), wherein at least the first and second enzymatically produced copies or enzymatically produced complementary sequences of the barcode sequence from the first reaction volume are attached to the nucleic acid sequence of the first circulating microparticles of the sample, and in a second reaction volume that contains the nucleic acid sequence of the second circulating microparticles from the sample (e.g., contains a genomic DNA fragment from the second circulating microparticles of the sample), wherein at least the first and second enzymatically produced copies or enzymatically produced complementary sequences of the barcode sequence from the second reaction volume are attached to the nucleic acid sequence of the second circulating microparticles of the sample.
[0299] Optionally, any method for generating enzymatically produced copies or enzymatically produced complementary sequences of barcode sequences can be performed for (and / or using or with) a library that contains two or more barcode sequences. Optionally, any method for generating enzymatically produced copies or enzymatically produced complementary sequences of barcode sequences can be performed for (and / or using or with) a library that contains two or more barcode molecules. Optionally, any method for generating enzymatically produced copies or enzymatically produced complementary sequences of barcode sequences can be performed for (and / or using or with) a library that contains two or more polymeric barcode molecules. Optionally, any method for generating enzymatically produced copies or enzymatically produced complementary sequences of barcode sequences can be performed for (and / or using or with) a library that contains two or more polymeric barcoding reagents. Optionally, any method for generating enzymatically produced copies or enzymatically produced complementary sequences of barcode sequences can be performed for (and / or using or with) a library that contains two or more barcoded oligonucleotides.
[0300] Optionally, any method for the enzymatic production of copies or enzymatically produced complementary sequences of barcode sequences may further comprise attaching one or more enzymatically produced copies or enzymatically produced complementary sequences of the barcode sequences to each of one or more nucleic acid sequences of the circulating particles (e.g., attaching to the genomic DNA sequence of the circulating particles). Optionally, any one or more such attachment steps may comprise a hybridization step (e.g., the step of hybridizing a barcoded oligonucleotide to a nucleic acid sequence), a step of hybridizing and extending the hybridization (e.g., the step of hybridizing a barcoded oligonucleotide to a nucleic acid sequence and subsequently extending the hybridized barcoded oligonucleotide with a polymerase), and / or a ligation step (e.g., the step of ligating a barcoded oligonucleotide to a nucleic acid sequence). After any one or more such attachment steps, a sequencing step may be performed on the nucleic acid sequence containing the barcode sequence and the nucleic acid sequence from the circulating particles to which it has been attached.
[0301] Optionally, any method for the enzymatic production of copies or enzymatically produced complementary sequences of barcode sequences may further comprise attaching one or more enzymatically produced copies or enzymatically produced complementary sequences of the barcode sequences to each of one or more nucleic acid sequences of the circulating particles, wherein the nucleic acid sequences of the circulating particles further comprise coupling sequences. Any coupling sequences and / or methods for attaching coupling sequences described herein, and / or methods for attaching barcode sequences to coupling sequences (and / or oligonucleotides containing coupling sequences) may be employed.
[0302] Optionally, any method for the enzymatic production of copies or enzymatically produced complementary sequences of barcode sequences and further comprising attaching one or more enzymatically produced copies or enzymatically produced complementary sequences of the barcode sequences to the nucleic acid sequences of the circulating particles may further comprise the step of chemically crosslinking the circulating particles (and / or chemically crosslinking a sample comprising two or more circulating particles). Optionally, the chemical crosslinking step may be performed before and / or after the step of partitioning the circulating particles and / or barcode molecules into two or more different partitions. Optionally, the step following the chemical crosslinking step may be a step of reversing the crosslinking, e.g., by a high-temperature thermal incubation step. Optionally, any method for the enzymatic production of copies or enzymatically produced complementary sequences of barcode sequences and further comprising attaching one or more enzymatically produced copies or enzymatically produced complementary sequences of the barcode sequences to the nucleic acid sequences of the circulating particles may further comprise the step of permeabilizing the circulating particles, e.g., by a high-temperature incubation step and / or a chemical surfactant.
[0303] Optionally, any method of generating enzymatically produced copies or enzymatically produced complementary sequences of barcode sequences can be carried out with any number and / or type and / or volume of the partitions described herein. Optionally, any method of generating enzymatically produced copies or enzymatically produced complementary sequences of barcode sequences in one or more partitions can include one or more partitions that contain any number of the cyclic microparticles as described herein. Optionally, any method of generating enzymatically produced copies or enzymatically produced complementary sequences of barcode sequences in one or more partitions can include one or more partitions that contain any number (or average number) of the cyclic microparticles as described herein. Optionally, any method of generating enzymatically produced copies or enzymatically produced complementary sequences of barcode sequences in one or more partitions can include one or more partitions that contain any mass (or average mass) of nucleic acids from the cyclic microparticles described herein (e.g., any mass of genomic DNA fragments).
[0304] Any method of generating enzymatically produced copies or enzymatically produced complementary sequences of barcode sequences can have a variety of desirable features and characteristics for analyzing associated sequences from cyclic microparticles. In a first case, generating enzymatically produced copies or enzymatically produced complementary sequences of barcode sequences enables the production of a large absolute mass of barcode sequences (e.g., a large absolute mass of barcode molecules or barcoded oligonucleotides) using only a small amount of starting barcode sequence material (e.g., PCR and RCA processing can produce a large exponential amplification of the input material for subsequent use and manipulation).
[0305] Furthermore, generating enzymatically produced copies or enzymatically produced complementary sequences of barcode sequences, where such barcode sequences are included in a library (e.g., included in a library of barcode molecules, a library of polymeric barcode molecules, a library of polymeric barcoding reagents, and / or a library of barcoded oligonucleotides), enables the production of a large absolute mass of barcode sequences with defined sequence characteristics (e.g., where a large absolute mass of barcode sequences contain sequences from a previously established and / or previously characterized library).
[0306] In addition, many enzymatic replication and amplification methods (e.g., rolling circle amplification by phi29 polymerase, and primer extension and / or PCR amplification by a thermostable polymerase (e.g., Phusion polymerase)) exhibit high molecular accuracy during the replication (in terms of the probability of generating errors in the newly replicated sequences), and thus show favorable accuracy characteristics of the resulting barcode sequences (e.g., the resulting barcode molecules, polymeric barcode molecules, and / or barcoded oligonucleotides) compared to non-enzymatic methods (e.g., compared to standard chemical oligonucleotide synthesis methods, such as phosphoramidite oligonucleotide synthesis).
[0307] In addition, enzymatic replication and amplification methods (such as primer extension and PCR methods) are highly suitable for subsequent modification, processing, and functionalization steps of said sequences, and may themselves also have further benefits that can be achieved in a relatively simple manner on substrates of large absolute weight. For example, primer extension products are readily configured and / or configurable for subsequent ligation processes (e.g., as in primer extension and ligation processes, such as may be carried out to generate barcoded oligonucleotides and / or polymeric barcoded reagents). And for a further example, the direct products of the enzymatic replication process itself (e.g., where the complementary sequence / copy of a barcode sequence anneals to the barcode sequence itself) may have desired functional and / or structural properties. For example, barcoded oligonucleotides generated by an enzymatic primer extension process are structurally tethered (by annealed nucleotide sequences) to a barcode molecule (e.g., a polymeric barcode molecule) in a single macromolecular complex during their production, which can then be further processed and / or functionalized in solution as a single intact reagent.
[0308] 11. General properties of polymeric barcoded reagents
[0309] The use of polymeric barcoded reagents exhibits a variety of available features and functions for associating sequences from circulating microparticles. In a first instance, such reagents (and / or their libraries) can contain a very well-defined and fully characterized set of barcodes, which can inform and enhance subsequent bioinformatics analysis (e.g., involving polymeric barcode molecules and / or polymeric barcoded reagents using known and / or empirically determined sequences). Additionally, such reagents are extremely easy to dispense and / or subject to other molecular or biophysical processing of multiple barcode sequences at once (i.e., since multiple barcode sequences are contained within each such reagent, they automatically "move together" within solution and during liquid handling and / or processing steps). Furthermore, the proximity between multiple barcode sequences within these reagents themselves enables new forms of functional assays, such as crosslinking circulating microparticles and subsequently attaching sequences from such polymeric reagents to genomic DNA fragments contained therein (including, for example, in their solution-phase reactions, i.e., two or more microparticles within a single partition).
[0310] The present invention provides polymeric barcoded reagents for labeling one or more target nucleic acids. The polymeric barcoded reagents include two or more barcode regions associated (directly or indirectly) together.
[0311] Each barcode region contains a nucleic acid sequence. The nucleic acid sequence can be single-stranded DNA, double-stranded DNA, or single-stranded DNA having one or more double-stranded regions.
[0312] Each barcode region may contain a sequence that identifies a polymeric barcoded reagent. For example, the sequence may be a constant region common to all barcode regions of a single polymeric barcoded reagent. Each barcode region may contain a unique sequence that is not present in other regions and can thus be used to uniquely identify each barcode region. Each barcode region may contain at least 5, at least 10, at least 15, at least 20, at least 25, at least 50, or at least 100 nucleotides. Preferably, each barcode region contains at least 5 nucleotides. Preferably, each barcode region contains deoxyribonucleotides, and optionally all nucleotides in the barcode region are deoxyribonucleotides. One or more deoxyribonucleotides may be modified deoxyribonucleotides (e.g., deoxyribonucleotides modified with a biotin moiety or deoxyuridine nucleotides). The barcode region may contain one or more degenerate nucleotides or sequences. The barcode region may not contain any degenerate nucleotides or sequences.
[0313] A polymeric barcoded reagent may contain at least 5, at least 10, at least 20, at least 25, at least 50, at least 75, at least 100, at least 200, at least 500, at least 1000, at least 5000, or at least 10,000 barcode regions. Preferably, the polymeric barcoded reagent contains at least 5 barcode regions.
[0314] A polymeric barcoded reagent may contain at least 2, at least 3, at least 4, at least 5, at least 10, at least 20, at least 25, at least 50, at least 75, at least 100, at least 200, at least 500, at least 1000, at least 5000, at least 10 4 at least 10 5 or at least 10 6 unique or distinct barcode regions. Preferably, the polymeric barcoded reagent includes at least 5 unique or distinct barcode regions.
[0315] A polymeric barcoded reagent may contain: a first and a second barcode molecule (i.e., a polymeric barcode molecule) linked together, wherein each barcode molecule contains a nucleic acid sequence containing a barcode region.
[0316] The barcode molecules of the polymeric barcode molecules can be linked on a nucleic acid molecule. The barcode molecules of the polymeric barcode molecules can be contained within a (single) nucleic acid molecule. The polymeric barcode molecules can comprise a single continuous nucleic acid sequence containing two or more barcode molecules. The polymeric barcode molecules can be single-stranded nucleic acid molecules (e.g., single-stranded DNA), double-stranded nucleic acid molecules, or single-stranded molecules containing one or more double-stranded regions. The polymeric barcode molecules can comprise one or more phosphorylated 5′ ends capable of ligating to the 3′ ends of other nucleic acid molecules. Optionally, within a double-stranded region or between two different double-stranded regions, the polymeric barcode molecules can comprise one or more nicks or one or more gaps, where the polymeric barcode molecule itself is separated or split. The length of any such gap can be at least one, at least 2, at least 5, at least 10, at least 20, at least 50, or at least 100 nucleotides. The nicks and / or gaps can be used for the purpose of increasing the molecular flexibility of the polymeric barcode molecules and / or polymeric barcoding reagents, e.g., increasing the accessibility of the molecules or reagents to interact with target nucleic acid molecules. The nicks and / or gaps can also enable more efficient purification or removal of the molecules or reagents. The molecules and / or reagents containing the nicks and / or gaps can maintain the linkage between different barcode molecules by having complementary DNA strands that co-hybridize to regions of two or more separated portions of the polymeric barcode molecule.
[0317] The barcode molecules can be linked via, for example, a support (e.g., a macromolecule, a solid support, or a semi-solid support). The sequence by which the barcode molecules are linked to each support can be known. The barcode molecules can be linked to the support directly or indirectly (e.g., via a linker molecule). The barcode molecules can be linked by binding to the support and / or by binding or annealing to a linker molecule bound to the support. The barcode molecules can bind to the support (or to the linker molecule) by covalent linkage, non-covalent linkage (e.g., protein-protein interaction or streptavidin-biotin bond), or nucleic acid hybridization. The linker molecule can be a biopolymer (e.g., a nucleic acid molecule) or a synthetic polymer. The linker molecule can comprise one or more ethylene glycol and / or poly(ethylene glycol) (e.g., hexaethylene glycol or pentaethylene glycol) units. The linker molecule can comprise one or more ethyl groups, e.g., a C3 (three-carbon) spacer, a C6 spacer, a C12 spacer, or a C18 spacer.
[0318] The barcode molecules can be linked via a macromolecule by binding to and / or annealing to the macromolecule.
[0319] Barcode molecules can be directly or indirectly (e.g., via linker molecules) associated with macromolecules. Barcode molecules can be associated by binding to the macromolecule and / or by binding or annealing to a linker molecule bound to the macromolecule. Barcode molecules can bind to the macromolecule (or linker molecule) via covalent linkage, non-covalent linkage (e.g., protein-protein interaction or streptavidin-biotin bond), or nucleic acid hybridization. The linker molecule can be a biopolymer (e.g., a nucleic acid molecule) or a synthetic polymer. The linker molecule can comprise one or more ethylene glycol and / or poly(ethylene glycol) (e.g., hexaethylene glycol or pentaethylene glycol) units. The linker molecule can comprise one or more ethyl groups, e.g., a C3 (three-carbon) spacer, a C6 spacer, a C12 spacer, or a C18 spacer.
[0320] The macromolecule can be a synthetic polymer (e.g., a dendrimer) or a biopolymer such as a nucleic acid (e.g., a single-stranded nucleic acid, e.g., single-stranded DNA), a peptide, a polypeptide, or a protein (e.g., a multimeric protein).
[0321] The dendrimer can comprise at least 2 generations, at least 3 generations, at least 5 generations, or at least 10 generations.
[0322] The macromolecule can be a nucleic acid comprising two or more nucleotides, each nucleotide capable of binding to the barcode molecule. As a supplement or alternative, the nucleic acid can comprise two or more regions, each region capable of hybridizing to the barcode molecule.
[0323] The nucleic acid can comprise a first modified nucleotide and a second modified nucleotide, wherein each modified nucleotide comprises a binding moiety (e.g., a biotin moiety, or an alkyne moiety useful for click chemistry reactions) capable of binding to the barcode molecule. Optionally, the first and second modified nucleotides can be separated by an intervening nucleic acid sequence of at least one, at least two, at least 5, or at least 10 nucleotides.
[0324] The nucleic acid can comprise a first hybridization region and a second hybridization region, wherein each hybridization region comprises a sequence complementary to and capable of hybridizing to the sequence of at least one nucleotide within the barcode molecule. The complementary sequence can be at least 5, at least 10, at least 15, at least 20, at least 25, or at least 50 consecutive nucleotides. Preferably, the complementary sequence is at least 10 consecutive nucleotides. Optionally, the first and second hybridization regions can be separated by an intervening nucleic acid sequence of at least one, at least two, at least 5, or at least 10 nucleotides.
[0325] The macromolecule can be a protein, e.g., a multimeric protein, e.g., a homomeric protein or a heteromeric protein. For example, the protein can comprise streptavidin, e.g., tetrameric streptavidin.
[0326] The support can be a solid support or a semi-solid support. The support may comprise a flat surface. The support can be, for example, a slide, such as a glass slide. The slide can be a flow cell for sequencing. If the support is a slide, the first and second barcode molecules can be immobilized in discrete regions on the slide. Optionally, the barcode molecules of each polymeric barcoding reagent in the library are immobilized in different discrete regions on the slide relative to the barcode molecules of other polymeric barcoding reagents in the library. The support can be a plate containing wells, optionally wherein the first and second barcode molecules are immobilized in the same well. Optionally, the barcode molecules of each polymeric barcoding reagent in the library are immobilized in different wells of the plate relative to the barcode molecules of other polymeric barcoding reagents in the library.
[0327] Preferably, the support is a bead (such as a gel bead). The beads can be agarose beads, silica beads, styrofoam beads, gel beads (such as those obtainable from ), antibody-conjugated beads, oligo-dT-conjugated beads, streptavidin beads or magnetic beads (such as superparamagnetic beads). The beads can have any size and / or molecular structure. For example, the beads can be from 10 nanometers to 100 micrometers in diameter, from 100 nanometers to 10 micrometers in diameter, or from 1 micrometer to 5 micrometers in diameter. Optionally, the beads are about 10 nanometers in diameter, about 100 nanometers in diameter, about 1 micrometer in diameter, about 10 micrometers in diameter, or about 100 micrometers in diameter. The beads can be solid, or alternatively the beads can be hollow or partially hollow or porous. For certain barcoding methods, certain sizes of beads can be most preferred. For example, beads less than 5.0 micrometers or less than 1.0 micrometers in diameter can be most useful for barcoding nucleic acid targets within individual cells. Preferably, the barcode molecules of each polymeric barcoding reagent in the library are linked together relative to the barcode molecules of other polymeric barcoding reagents in the library to different beads.
[0328] The support can be functionalized to enable the attachment of two or more barcode molecules. Such functionalization can be achieved by adding chemical moieties (such as carboxyl groups, alkynes, azides, acrylate groups, amino groups, sulfate groups or succinimide groups) and / or protein-based moieties (such as streptavidin, avidin or protein G) to the support. The barcode molecules can be directly or indirectly (such as through a linker molecule) attached to the moieties.
[0329] The functionalized support (such as a bead) can be contacted with a solution of barcode molecules (to produce a polymeric barcoding reagent) under conditions that facilitate the attachment of two or more individual barcode molecules to each bead in the solution.
[0330] In a library of polymeric barcoded reagents, the barcode molecules of each polymeric barcoded reagent in the library can be associated together on different supports relative to the barcode molecules of other polymeric barcoded reagents in the library.
[0331] The polymeric barcoded reagent can comprise: at least 2, at least 3, at least 4, at least 5, at least 10, at least 20, at least 25, at least 50, at least 75, at least 100, at least 200, at least 500, at least 1000, at least 5000, at least 10 4 at least 10 5 at least 10 6 barcode molecules associated together, where each barcode molecule is as defined herein; and a barcoded oligonucleotide annealed to each barcode molecule, where each barcoded oligonucleotide is as defined herein. Preferably, the polymeric barcoded reagent comprises at least 5 barcode molecules associated together, where each barcode molecule is as defined herein; and a barcoded oligonucleotide annealed to each barcode molecule, where each barcoded oligonucleotide is as defined herein.
[0332] The polymeric barcoded reagent can comprise: at least 2, at least 3, at least 4, at least 5, at least 10, at least 20, at least 25, at least 50, at least 75, at least 100, at least 200, at least 500, at least 1000, at least 5000, at least 10 4 at least 10 5 at least 10 6 unique or different barcode molecules associated together, where each barcode molecule is as defined herein; and a barcoded oligonucleotide annealed to each barcode molecule, where each barcoded oligonucleotide is as defined herein. Preferably, the polymeric barcoded reagent comprises at least 5 unique or different barcode molecules associated together, where each barcode molecule is as defined herein; and a barcoded oligonucleotide annealed to each barcode molecule, where each barcoded oligonucleotide is as defined herein.
[0333] The polymeric barcoded reagent may comprise two or more barcoded oligonucleotides as defined herein, wherein each barcoded oligonucleotide comprises a barcode region. The polymeric barcoded reagent may comprise: at least 2, at least 3, at least 4, at least 5, at least 10, at least 20, at least 25, at least 50, at least 75, at least 100, at least 200, at least 500, at least 1000, at least 5000, at least 10,000, at least 100,000, or at least 1,000,000 unique or distinct barcoded oligonucleotides. Preferably, the polymeric barcoded reagent comprises at least 5 unique or distinct barcoded oligonucleotides.
[0334] The barcoded oligonucleotides of the polymeric barcoded reagent are linked together (directly or indirectly). The barcoded oligonucleotides of the polymeric barcoded reagent are linked together by a support (e.g., a macromolecule, a solid support, or a semi-solid support) as described herein. The polymeric barcoded reagent may comprise one or more polymers to which the barcoded oligonucleotides are annealed or attached. For example, the barcoded oligonucleotides of the polymeric barcoded reagent may be annealed to a polymeric hybridization molecule (e.g., a polymeric barcode molecule). Alternatively, the barcoded oligonucleotides of the polymeric barcoded reagent may be linked together by a macromolecule (e.g., a synthetic polymer such as a dendrimer, or a biopolymer such as a protein) or a support (e.g., a solid support or a semi-solid support, such as a gel bead). As a supplement or alternative, the barcoded oligonucleotides of a (single) polymeric barcoded reagent may be linked together by being contained within a (single) lipid carrier (e.g., a liposome or a micelle).
[0335] The polymeric barcoded reagent may comprise: a first and a second hybridization molecule (i.e., a polymeric hybridization molecule) linked together, wherein each hybridization molecule comprises a nucleic acid sequence containing a hybridization region; and a first and a second barcoded oligonucleotide, wherein the first barcoded oligonucleotide is annealed to the hybridization region of the first hybridization molecule, and wherein the second barcoded oligonucleotide is annealed to the hybridization region of the second hybridization molecule.
[0336] The hybridization molecule comprises deoxyribonucleotides or consists of deoxyribonucleotides. One or more deoxyribonucleotides may be modified deoxyribonucleotides (e.g., deoxyribonucleotides modified with a biotin moiety or deoxyuridine nucleotides). The hybridization molecule may comprise one or more degenerate nucleotides or sequences. The hybridization molecule may not comprise any degenerate nucleotides or sequences.
[0337] The hybridizing molecules of the multimeric hybrid molecules can be associated on a nucleic acid molecule. Such a nucleic acid molecule can provide a backbone that can anneal to a single-stranded barcoded oligonucleotide. The hybridizing molecules of the multimeric hybrid molecules can be contained within a (single) nucleic acid molecule. The multimeric hybrid molecule can comprise a single continuous nucleic acid sequence containing two or more hybridizing molecules. The multimeric hybrid molecule can be a single-stranded nucleic acid molecule (e.g., single-stranded DNA) containing two or more hybridizing molecules. The multimeric hybrid molecule can comprise one or more double-stranded regions. Optionally, within a double-stranded region or between two different double-stranded regions, the multimeric hybrid molecule can comprise one or more nicks or one or more gaps, wherein the multimeric barcode molecule itself is separated or split. The length of any such gap can be at least one, at least 2, at least 5, at least 10, at least 20, at least 50, or at least 100 nucleotides. The nick and / or gap can be used for the purpose of increasing the molecular flexibility of the multimeric hybrid molecule and / or the multimeric barcoded reagent, e.g., increasing the accessibility of the molecule or reagent to interact with a target nucleic acid molecule. The nick and / or gap can also enable more efficient purification or removal of the molecule or reagent. The molecule and / or reagent containing the nick and / or gap can maintain the association between different hybridizing molecules by having complementary DNA strands that co-hybridize to regions of two or more separated portions of the multimeric hybrid molecule.
[0338] The hybridizing molecules can be associated through a macromolecule by binding to and / or annealing to the macromolecule.
[0339] The hybridizing molecules can be directly or indirectly (e.g., through a linker molecule) associated with a macromolecule. The hybridizing molecules can be associated by binding to the macromolecule and / or by binding to or annealing to a linker molecule bound to the macromolecule. The hybridizing molecules can bind to the macromolecule (or linker molecule) through covalent linkage, non-covalent linkage (e.g., protein-protein interaction or streptavidin-biotin bond), or nucleic acid hybridization. The linker molecule can be a biopolymer (e.g., nucleic acid molecule) or a synthetic polymer. The linker molecule can comprise one or more ethylene glycol and / or poly(ethylene glycol) (e.g., hexaethylene glycol or pentaethylene glycol) units. The linker molecule can comprise one or more ethyl groups, e.g., a C3 (three-carbon) spacer, a C6 spacer, a C12 spacer, or a C18 spacer.
[0340] The macromolecule can be a synthetic polymer (e.g., dendrimer) or a biopolymer such as a nucleic acid (e.g., single-stranded nucleic acid, e.g., single-stranded DNA), a peptide, a polypeptide, or a protein (e.g., multimeric protein).
[0341] The dendrimer can comprise at least 2 generations, at least 3 generations, at least 5 generations, or at least 10 generations.
[0342] The macromolecule can be a nucleic acid containing two or more nucleotides, each nucleotide being capable of binding to a hybridization molecule. As a supplement or alternative, the nucleic acid can contain two or more regions, each region being capable of hybridizing to a hybridization molecule.
[0343] The nucleic acid can contain a first modified nucleotide and a second modified nucleotide, where each modified nucleotide contains a binding moiety (e.g., a biotin moiety, or an alkyne moiety that can be used in a click chemical reaction) capable of binding to a hybridization molecule. Optionally, the first and second modified nucleotides can be separated by an intervening nucleic acid sequence of at least one, at least two, at least 5, or at least 10 nucleotides.
[0344] The nucleic acid can contain a first hybridization region and a second hybridization region, where each hybridization region contains a sequence complementary to and capable of hybridizing to the sequence of at least one nucleotide within the hybridization molecule. The complementary sequence can be at least 5, at least 10, at least 15, at least 20, at least 25, or at least 50 consecutive nucleotides. Optionally, the first hybridization region and the second hybridization region can be separated by an intervening nucleic acid sequence of at least one, at least two, at least 5, or at least 10 nucleotides.
[0345] The macromolecule can be a protein, such as a multimeric protein, such as a homomeric protein or a heteromeric protein. For example, the protein can contain streptavidin, such as tetrameric streptavidin.
[0346] The hybridization molecule can be associated with a support. The hybridization molecule can be directly or indirectly (e.g., via a linker molecule) associated with the support. The hybridization molecule can be associated by binding to the support and / or by binding or annealing to a linker molecule bound to the support. The hybridization molecule can bind to the support (or linker molecule) by covalent linkage, non-covalent linkage (e.g., protein-protein interaction or streptavidin-biotin bond), or nucleic acid hybridization. The linker molecule can be a biopolymer (e.g., a nucleic acid molecule) or a synthetic polymer. The linker molecule can contain one or more ethylene glycol and / or poly(ethylene glycol) (e.g., hexaethylene glycol or pentaethylene glycol) units. The linker molecule can contain one or more ethyl groups, such as a C3 (three-carbon) spacer, a C6 spacer, a C12 spacer, or a C18 spacer.
[0347] The support can be a solid support or a semi-solid support. The support may comprise a flat surface. The support can be, for example, a slide, such as a glass slide. The slide can be a flow cell for sequencing. If the support is a slide, the first and second hybridization molecules can be immobilized in discrete regions on the slide. Optionally, the hybridization molecules of each polymeric barcoding reagent in the library are immobilized in different discrete regions on the slide relative to the hybridization molecules of other polymeric barcoding reagents in the library. The support can be a plate containing wells, optionally wherein the first and second hybridization molecules are immobilized in the same well. Optionally, the hybridization molecules of each polymeric barcoding reagent in the library are immobilized in different wells of the plate relative to the hybridization molecules of other polymeric barcoding reagents in the library.
[0348] Preferably, the support is a bead (such as a gel bead). The beads can be agarose beads, silica beads, styrofoam beads, gel beads (such as those obtainable from ), antibody-conjugated beads, oligo-dT-conjugated beads, streptavidin beads, or magnetic beads (such as superparamagnetic beads). The beads can have any size and / or molecular structure. For example, the beads can be from 10 nanometers to 100 micrometers in diameter, from 100 nanometers to 10 micrometers in diameter, or from 1 micrometer to 5 micrometers in diameter. Optionally, the beads are approximately 10 nanometers in diameter, approximately 100 nanometers in diameter, approximately 1 micrometer in diameter, approximately 10 micrometers in diameter, or approximately 100 micrometers in diameter. The beads can be solid, or alternatively the beads can be hollow or partially hollow or porous. For certain barcoding methods, certain sizes of beads may be most preferred. For example, beads less than 5.0 micrometers or less than 1.0 micrometers in diameter may be most useful for barcoding nucleic acid targets within individual cells. Preferably, the hybridization molecules of each polymeric barcoding reagent in the library are associated together on different beads relative to the hybridization molecules of other polymeric barcoding reagents in the library.
[0349] The support can be functionalized to enable the attachment of two or more hybridization molecules. Such functionalization can be achieved by adding chemical moieties (such as carboxyl groups, alkynes, azides, acrylate groups, amino groups, sulfate groups, or succinimide groups) and / or protein-based moieties (such as streptavidin, avidin, or protein G) to the support. The hybridization molecules can be attached directly or indirectly (such as through a linker molecule) to the moieties.
[0350] The functionalized support (such as a bead) can be contacted with a solution of hybridization molecules (yielding a polymeric barcoding reagent) under conditions that promote the attachment of two or more individual hybridization molecules to each bead in solution.
[0351] In a library of polymeric barcoding reagents, the hybridization molecules of each polymeric barcoding reagent in the library can be associated together on different supports relative to the hybridization molecules of other polymeric barcoding reagents in the library.
[0352] Optionally, the hybrid molecule is linked to the bead by covalent linkage, non-covalent linkage (e.g., streptavidin-biotin bond), or nucleic acid hybridization.
[0353] The polymeric barcoding reagent can comprise: at least 2, at least 3, at least 4, at least 5, at least 10, at least 20, at least 25, at least 50, at least 75, at least 100, at least 200, at least 500, at least 1000, at least 5000, or at least 10,000 hybrid molecules associated together, where each hybrid molecule is as defined herein; and barcoded oligonucleotides annealed to each hybrid molecule, where each barcoded oligonucleotide is as defined herein. Preferably, the polymeric barcoding reagent comprises at least 5 hybrid molecules associated together, where each hybrid molecule is as defined herein; and barcoded oligonucleotides annealed to each hybrid molecule, where each barcoded oligonucleotide is as defined herein.
[0354] The polymeric barcoding reagent can comprise: at least 2, at least 3, at least 4, at least 5, at least 10, at least 20, at least 25, at least 50, at least 75, at least 100, at least 200, at least 500, at least 1000, at least 5000, or at least 10,000 unique or distinct hybrid molecules associated together, where each hybrid molecule is as defined herein; and barcoded oligonucleotides annealed to each hybrid molecule, where each barcoded oligonucleotide is as defined herein. Preferably, the polymeric barcoding reagent comprises at least 5 unique or distinct hybrid molecules associated together, where each hybrid molecule is as defined herein; and barcoded oligonucleotides annealed to each hybrid molecule, where each barcoded oligonucleotide is as defined herein.
[0355] The polymeric hybrid molecule can be a polymeric barcode molecule, where the first hybrid molecule is the first barcode molecule and the second hybrid molecule is the second barcode molecule. The polymeric barcoding reagent can comprise: the first and second barcode molecules (i.e., polymeric barcode molecule) associated together, where each barcode molecule comprises a nucleic acid sequence containing a barcode region; and the first and second barcoded oligonucleotides, where the first barcoded oligonucleotide anneals to the barcode region of the first barcode molecule and where the second barcoded oligonucleotide anneals to the barcode region of the second barcode molecule.
[0356] The barcoded oligonucleotides of the polymeric barcoding reagent can comprise: a first barcoded oligonucleotide that optionally comprises, in the 5' to 3' direction, a barcode region and a target region capable of annealing or ligating to a first target nucleic acid fragment; and a second barcoded oligonucleotide that optionally comprises, in the 5' to 3' direction, a barcode region and a target region capable of annealing or ligating to a second target nucleic acid fragment.
[0357] The barcoded oligonucleotides of the polymeric barcoding reagent can include: a first barcoded oligonucleotide that includes a barcode region and a target region capable of ligating to a first target nucleic acid fragment; and a second barcoded oligonucleotide that includes a barcode region and a target region capable of ligating to a second target nucleic acid fragment.
[0358] The barcoded oligonucleotides of the polymeric barcoding reagent can include: a first barcoded oligonucleotide that includes, in the 5' to 3' direction, a barcode region and a target region capable of annealing to a first target nucleic acid fragment; and a second barcoded oligonucleotide that includes, in the 5' to 3' direction, a barcode region and a target region capable of annealing to a second target nucleic acid fragment.
[0359] 12. General properties of barcoded oligonucleotides
[0360] The barcoded oligonucleotide includes a barcode region. Optionally, the barcoded oligonucleotide can include, in the 5' to 3' direction, a barcode region and a target region. The target region is capable of annealing or ligating to a target nucleic acid fragment. Alternatively, the barcoded oligonucleotide can consist essentially of or consist of a barcode region.
[0361] The 5' end of the barcoded oligonucleotide can be phosphorylated. This can enable the 5' end of the barcoded oligonucleotide to ligate to the 3' end of the target nucleic acid. Alternatively, the 5' end of the barcoded oligonucleotide can be non-phosphorylated.
[0362] The barcoded oligonucleotide can be a single-stranded nucleic acid molecule (e.g., single-stranded DNA). The barcoded oligonucleotide can include one or more double-stranded regions. The barcoded oligonucleotide can be a double-stranded nucleic acid molecule (e.g., double-stranded DNA).
[0363] The barcoded oligonucleotide can include or consist of deoxyribonucleotides. One or more of the deoxyribonucleotides can be modified deoxyribonucleotides (e.g., deoxyribonucleotides modified with a biotin moiety or deoxyuridine nucleotides). The barcoded oligonucleotide can include one or more degenerate nucleotides or sequences. The barcoded oligonucleotide can not include any degenerate nucleotides or sequences.
[0364] The barcode region of each barcoded oligonucleotide may contain different sequences. The barcode region of each barcoded oligonucleotide may contain a sequence that identifies the polymeric barcoding reagent. For example, the sequence may be a constant region common to all barcode regions of a single polymeric barcoding reagent. The barcode region of each barcoded oligonucleotide may contain a unique sequence that is not present in other barcoded oligonucleotides and can thus be used to uniquely identify each barcoded oligonucleotide. The barcode region of each barcoded oligonucleotide may contain at least 5, at least 10, at least 15, at least 20, at least 25, at least 50, or at least 100 nucleotides. Preferably, each barcode region contains at least 5 nucleotides. Preferably, each barcode region contains deoxyribonucleotides, and optionally all nucleotides in the barcode region are deoxyribonucleotides. One or more deoxyribonucleotides may be modified deoxyribonucleotides (e.g., deoxyribonucleotides modified with a biotin moiety or deoxyuridine nucleotides). The barcode region may contain one or more degenerate nucleotides or sequences. The barcode region may not contain any degenerate nucleotides or sequences.
[0365] The target region of each barcoded oligonucleotide may contain different sequences. Each target region may contain a sequence capable of annealing to only a single target nucleic acid fragment within a nucleic acid sample (i.e., a target-specific sequence). Each target region may contain one or more random sequences or one or more degenerate sequences such that the target region is capable of annealing to more than one target nucleic acid fragment. The target region of each barcoded oligonucleotide may contain at least 5, at least 10, at least 15, at least 20, at least 25, at least 50, or at least 100 nucleotides. Preferably, each target region contains at least 5 nucleotides. The target region of each barcoded oligonucleotide may contain 5 to 100 nucleotides, 5 to 10 nucleotides, 10 to 20 nucleotides, 20 to 30 nucleotides, 30 to 50 nucleotides, 50 to 100 nucleotides, 10 to 90 nucleotides, 20 to 80 nucleotides, 30 to 70 nucleotides, or 50 to 60 nucleotides. Preferably, each target region contains 30 to 70 nucleotides. Preferably, each target region contains deoxyribonucleotides, and optionally all nucleotides in the target region are deoxyribonucleotides. One or more deoxyribonucleotides may be modified deoxyribonucleotides (e.g., deoxyribonucleotides modified with a biotin moiety or deoxyuridine nucleotides). Each target region may contain one or more universal bases (e.g., inosine), one or more modified nucleotides, and / or one or more nucleotide analogs.
[0366] The target region can be used to anneal the barcoded oligonucleotide to the target nucleic acid fragment and can subsequently be used as a primer for a primer extension reaction or an amplification reaction (such as a polymerase chain reaction). Alternatively, the target region can be used to ligate the barcoded oligonucleotide to the target nucleic acid fragment. The target region can be located at the 5' end of the barcoded oligonucleotide. Such a target region can be phosphorylated. This can enable the 5' end of the target region to ligate to the 3' end of the target nucleic acid fragment.
[0367] The barcoded oligonucleotide can also include one or more linker regions. The linker region can be between the barcode region and the target region. The barcoded oligonucleotide can, for example, include a linker region 5' to the barcode region (5' linker region) and / or a linker region 3' to the barcode region (3' linker region). Optionally, the barcoded oligonucleotide includes the barcode region, the linker region, and the target region in the 5' to 3' direction.
[0368] The linker region of the barcoded oligonucleotide can include a sequence complementary to the linker region of the polymeric barcode molecule or a sequence complementary to the hybridization region of the polymeric hybridization molecule. The linker region of the barcoded oligonucleotide can enable the barcoded oligonucleotide to ligate to a macromolecule or a support (such as a bead). The linker region can be used to manipulate, purify, recover, amplify, or detect the barcoded oligonucleotide and / or the target nucleic acid to which it can anneal or ligate.
[0369] The linker region of each barcoded oligonucleotide can include a constant region. Optionally, all of the linker regions of the barcoded oligonucleotides of each polymeric barcoding reagent are substantially the same. The linker region can include at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 8, at least 10, at least 15, at least 20, at least 25, at least 50, at least 100, or at least 250 nucleotides. Preferably, the linker region includes at least 4 nucleotides. Preferably, each linker region includes deoxyribonucleotides, and optionally, all of the nucleotides in the linker region are deoxyribonucleotides. One or more deoxyribonucleotides can be modified deoxyribonucleotides (such as deoxyribonucleotides modified with a biotin moiety or deoxyuridine nucleotides). Each linker region can include one or more universal bases (such as inosine), one or more modified nucleotides, and / or one or more nucleotide analogs.
[0370] The barcoded oligonucleotide can be synthesized by chemical oligonucleotide synthesis methods. The barcoded oligonucleotide synthesis process can include one or more of the following steps: an enzymatic production process, an enzymatic amplification process, or an enzymatic modification operation, such as an in vitro transcription process, a reverse transcription process, a primer extension process, or a polymerase chain reaction process.
[0371] These general properties of the barcoded oligonucleotides apply to any polymeric barcoding reagent described herein.
[0372] 13. General Properties of the Polymeric Barcoding Reagent Library
[0373] The present invention provides a library of polymeric barcoding reagents, which comprises first and second polymeric barcoding reagents as defined herein, wherein the barcode regions of the first polymeric barcoding reagent are different from the barcode regions of the second polymeric barcoding reagent.
[0374] The library of polymeric barcoding reagents may comprise at least 5, at least 10, at least 20, at least 25, at least 50, at least 75, at least 100, at least 250, at least 500, at least 10 3 at least 10 4 at least 10 5 at least 10 6 at least 10 7 at least 10 8 at least 10 9 or at least 10
[0375] as defined herein polymeric barcoding reagents. Preferably, the library comprises at least 10 polymeric barcoding reagents as defined herein. Preferably, the first and second barcode regions of each polymeric barcoding reagent are different from the barcode regions of at least 9 other polymeric barcoding reagents in the library. 3 4 5 6 7 8 9 4 5 6 7 8 9 -1) or at least 10
[0376] The barcode regions of each polymeric barcoding reagent may be different from the barcode regions of at least 4, at least 9, at least 19, at least 24, at least 49, at least 74, at least 99, at least 249, at least 499, at least 999 (i.e., 103 -1), at least 10 4 -1), at least 10 5 -1), at least 10 6 -1), at least 10 7 -1), at least 10 8 -1 or at least 10 9 -1 barcode regions of other polymeric barcoding reagents. The barcode region of each polymeric barcoding reagent can be different from the barcode regions of all other polymeric barcoding reagents in the library. Preferably, the barcode region of each polymeric barcoding reagent is different from the barcode regions of at least 9 other polymeric barcoding reagents in the library.
[0377] The present invention provides a library of polymeric barcoding reagents comprising a first and a second polymeric barcoding reagent as defined herein, wherein the barcode region of the barcoded oligonucleotide of the first polymeric barcoding reagent is different from the barcode region of the barcoded oligonucleotide of the second polymeric barcoding reagent.
[0378] Different polymeric barcoding reagents within the library of polymeric barcoding reagents can comprise different numbers of barcoded oligonucleotides.
[0379] The library of polymeric barcoding reagents can comprise at least 5, at least 10, at least 20, at least 25, at least 50, at least 75, at least 100, at least 250, at least 500, at least 10 3 ), at least 10 4 ), at least 10 5 ), at least 10 6 ), at least 10 7 ), at least 10 8 ), or at least 10 9 ), polymeric barcoding reagents as defined herein. Preferably, the library comprises at least 10 polymeric barcoding reagents as defined herein. Preferably, the barcode regions of the first and second barcoded oligonucleotides of each polymeric barcoding reagent are different from the barcode regions of the barcoded oligonucleotides of at least 9 other polymeric barcoding reagents in the library.
[0380] The barcode regions of the first and second barcoded oligonucleotides of each polymeric barcoding reagent can be different from the barcode regions of at least 4, at least 9, at least 19, at least 24, at least 49, at least 74, at least 99, at least 249, at least 499, at least 999 (i.e., 10 3 -1), at least 10 4 -1), at least 10 5 -1), at least 10 6 -1), at least 10 7 -1), at least 108 -1 or at least 10 9 -the barcode region of the barcode oligonucleotide of one or more other polymeric barcoding reagents. The barcode regions of the first and second barcode oligonucleotides of each polymeric barcoding reagent may be different from the barcode regions of the barcode oligonucleotides of all other polymeric barcoding reagents in the library. Preferably, the barcode regions of the first and second barcode oligonucleotides of each polymeric barcoding reagent are different from the barcode regions of the barcode oligonucleotides of at least 9 other polymeric barcoding reagents in the library.
[0381] The barcode region of the barcode oligonucleotide of each polymeric barcoding reagent may be different from the barcode regions of the barcode oligonucleotides of at least 4, at least 9, at least 19, at least 24, at least 49, at least 74, at least 99, at least 249, at least 499, at least 999 (i.e., 10 3 -1), at least 10 4 -1, at least 10 5 -1, at least 10 6 -1, at least 10 7 -1, at least 10 8 -1 or at least 10 9 -the barcode region of the barcode oligonucleotide of one or more other polymeric barcoding reagents. The barcode region of the barcode oligonucleotide of each polymeric barcoding reagent may be different from the barcode regions of the barcode oligonucleotides of all other polymeric barcoding reagents in the library. Preferably, the barcode region of the barcode oligonucleotide of each polymeric barcoding reagent is different from the barcode regions of the barcode oligonucleotides of at least 9 other polymeric barcoding reagents in the library.
[0382] These general properties of the polymeric barcoding reagent library apply to any polymeric barcoding reagent described herein.
[0383] 14. The polymeric barcoding reagent comprises a barcode oligonucleotide annealed to a polymeric barcode molecule
[0384] The present invention provides a polymeric barcoding reagent for labeling a target nucleic acid, wherein the reagent comprises: a first and a second barcode molecule (i.e., a polymeric barcode molecule) linked together, wherein each barcode molecule comprises a nucleic acid sequence containing a barcode region; and a first and a second barcode oligonucleotide, wherein the first barcode oligonucleotide optionally comprises, in the 5' to 3' direction, a barcode region annealed to the barcode region of the first barcode molecule and a target region capable of annealing or ligating to a first target nucleic acid fragment, and wherein the second barcode oligonucleotide optionally comprises, in the 5' to 3' direction, a barcode region annealed to the barcode region of the second barcode molecule and a target region capable of annealing or ligating to a second target nucleic acid fragment.
[0385] The present invention provides a polymeric barcoding reagent for labeling a target nucleic acid, wherein the reagent comprises: a first and a second barcode molecule (i.e., polymeric barcode molecule) linked together, wherein each barcode molecule comprises a nucleic acid sequence containing a barcode region; and a first and a second barcoded oligonucleotide, wherein the first barcoded oligonucleotide comprises a barcode region annealing to the barcode region of the first barcode molecule and a target region capable of ligating to a first target nucleic acid fragment, and wherein the second barcoded oligonucleotide comprises a barcode region annealing to the barcode region of the second barcode molecule and a target region capable of ligating to a second target nucleic acid fragment.
[0386] The present invention provides a polymeric barcoding reagent for labeling a target nucleic acid, wherein the reagent comprises: a first and a second barcode molecule (i.e., polymeric barcode molecule) linked together, wherein each barcode molecule comprises a nucleic acid sequence containing a barcode region; and a first and a second barcoded oligonucleotide, wherein the first barcoded oligonucleotide comprises, in the 5' to 3' direction, a barcode region annealing to the barcode region of the first barcode molecule and a target region capable of annealing to a first target nucleic acid fragment, and wherein the second barcoded oligonucleotide comprises, in the 5' to 3' direction, a barcode region annealing to the barcode region of the second barcode molecule and a target region capable of annealing to a second target nucleic acid fragment.
[0387] The present invention provides a polymeric barcoding reagent for labeling a target nucleic acid, wherein the reagent comprises: a first and a second barcode molecule (i.e., polymeric barcode molecule) linked together, wherein each barcode molecule comprises a nucleic acid sequence containing a barcode region; and a first and a second barcoded oligonucleotide, wherein the first barcoded oligonucleotide comprises a barcode region annealing to the barcode region of the first barcode molecule and capable of ligating to a first target nucleic acid fragment, and wherein the second barcoded oligonucleotide comprises a barcode region annealing to the barcode region of the second barcode molecule and capable of ligating to a second target nucleic acid fragment.
[0388] Each barcoded oligonucleotide may consist essentially of or consist of a barcode region.
[0389] Preferably, the barcode molecule comprises or consists of deoxyribonucleotides. One or more deoxyribonucleotides may be modified deoxyribonucleotides (e.g., deoxyribonucleotides modified with a biotin moiety or deoxyuridine nucleotides). The barcode molecule may comprise one or more degenerate nucleotides or sequences. The barcode molecule may not contain any degenerate nucleotides or sequences.
[0390] The barcode region can uniquely identify each barcode molecule. Each barcode region can contain sequences that identify the polymeric barcoding reagent. For example, the sequence can be a constant region common to all barcode regions of a single polymeric barcoding reagent. Each barcode region can contain at least 5, at least 10, at least 15, at least 20, at least 25, at least 50, or at least 100 nucleotides. Preferably, each barcode region contains at least 5 nucleotides. Preferably, each barcode region contains deoxyribonucleotides, and optionally all nucleotides in the barcode region are deoxyribonucleotides. One or more deoxyribonucleotides can be modified deoxyribonucleotides (e.g., deoxyribonucleotides modified with a biotin moiety or deoxyuridine nucleotides). The barcode region can contain one or more degenerate nucleotides or sequences. The barcode region can not contain any degenerate nucleotides or sequences.
[0391] Preferably, the barcode region of the first barcoding oligonucleotide contains a sequence that is complementary to and anneals with the barcode region of the first barcode molecule, and the barcode region of the second barcoding oligonucleotide contains a sequence that is complementary to and anneals with the barcode region of the second barcode molecule. The complementary sequence of each barcoding oligonucleotide can be at least 5, at least 10, at least 15, at least 20, at least 25, at least 50, or at least 100 consecutive nucleotides.
[0392] The target region of the barcoding oligonucleotide (which does not anneal to the polymeric barcode molecule) may not be complementary to the polymeric barcode molecule.
[0393] The barcoding oligonucleotide can contain a linker region between the barcode region and the target region. The linker region can contain one or more consecutive nucleotides that do not anneal to the polymeric barcode molecule and are not complementary to the target nucleic acid fragment. The linker can contain 1 to 100, 5 to 75, 10 to 50, 15 to 30, or 20 to 25 non-complementary nucleotides. Preferably, the linker contains 15 to 30 non-complementary nucleotides. Using such a linker region enhances the efficiency of the barcoding reaction using the polymeric barcoding reagent.
[0394] The barcode molecule can also contain one or more nucleic acid sequences that are not complementary to the barcode region of the barcoding oligonucleotide. For example, the barcode molecule can contain one or more adaptor regions. The barcode molecule can, for example, contain an adaptor region (5' adaptor region) at the 5' of the barcode region and / or an adaptor region (3' adaptor region) at the 3' of the barcode region. The adaptor region (and / or one or more parts of the adaptor region) can be complementary to and anneal with an oligonucleotide (e.g., the adaptor region of the barcoding oligonucleotide). Alternatively, the adaptor region (and / or one or more parts of the adaptor region) of the barcode molecule may not be complementary to the sequence of the barcoding oligonucleotide. The adaptor region can be used for manipulating, purifying, retrieving, amplifying, and / or detecting the barcode molecule.
[0395] The polymeric barcoded reagent can be configured such that: each barcode molecule comprises a nucleic acid sequence containing an adaptor region and a barcode region in the 5' to 3' direction; the first barcoded oligonucleotide optionally comprises, in the 5' to 3' direction, a barcode region annealing to the barcode region of the first barcode molecule, an adaptor region annealing to the adaptor region of the first barcode molecule, and a target region capable of annealing to the first target nucleic acid region; and the second barcoded oligonucleotide optionally comprises, in the 5' to 3' direction, a barcode region annealing to the barcode region of the second barcode molecule, an adaptor region annealing to the adaptor region of the second barcode molecule, and a target region capable of annealing to the second target nucleic acid region.
[0396] The adaptor region of each barcode molecule may comprise a constant region. Optionally, all adaptor regions of the polymeric barcoded reagent are substantially the same. The adaptor region may comprise at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 8, at least 10, at least 15, at least 20, at least 25, at least 50, at least 100, or at least 250 nucleotides. Preferably, the adaptor region comprises at least 4 nucleotides. Preferably, each adaptor region comprises deoxyribonucleotides, and optionally, all nucleotides in the adaptor region are deoxyribonucleotides. One or more deoxyribonucleotides may be modified deoxyribonucleotides (e.g., deoxyribonucleotides modified with a biotin moiety or deoxyuridine nucleotides). Each adaptor region may comprise one or more universal bases (e.g., inosine), one or more modified nucleotides, and / or one or more nucleotide analogs.
[0397] The barcoded oligonucleotide may comprise a linker region between the adaptor region and the target region. The linker region may comprise one or more consecutive nucleotides that do not anneal to the polymeric barcode molecule and are not complementary to the target nucleic acid fragment. The linker may comprise 1 to 100, 5 to 75, 10 to 50, 15 to 30, or 20 to 25 non-complementary nucleotides. Preferably, the linker comprises 15 to 30 non-complementary nucleotides. The use of such a linker region enhances the efficiency of the barcoding reaction using the polymeric barcoded reagent.
[0398] The barcode molecules of the polymeric barcode molecule may be linked on a nucleic acid molecule. Such a nucleic acid molecule may provide a backbone that can anneal to the single-stranded barcoded oligonucleotide. Alternatively, the barcode molecules of the polymeric barcode molecule may be linked together by any other means described herein.
[0399] The polymeric barcoded reagent may comprise: at least 2, at least 3, at least 4, at least 5, at least 10, at least 20, at least 25, at least 50, at least 75, at least 100, at least 200, at least 500, at least 1000, at least 5000, or at least 10,000 barcoded molecules linked together, where each barcoded molecule is as defined herein; and a barcoded oligonucleotide annealed to each barcoded molecule, where each barcoded oligonucleotide is as defined herein. Preferably, the polymeric barcoded reagent comprises at least 5 barcoded molecules linked together, where each barcoded molecule is as defined herein; and a barcoded oligonucleotide annealed to each barcoded molecule, where each barcoded oligonucleotide is as defined herein.
[0400] The polymeric barcoded reagent may comprise: at least 2, at least 3, at least 4, at least 5, at least 10, at least 20, at least 25, at least 50, at least 75, at least 100, at least 200, at least 500, at least 1000, at least 5000, at least 10 4 at, least 10 5 at, least 10 6 unique or distinct barcoded molecules linked together, where each barcoded molecule is as defined herein; and a barcoded oligonucleotide annealed to each barcoded molecule, where each barcoded oligonucleotide is as defined herein. Preferably, the polymeric barcoded reagent comprises at least 5 unique or distinct barcoded molecules linked together, where each barcoded molecule is as defined herein; and a barcoded oligonucleotide annealed to each barcoded molecule, where each barcoded oligonucleotide is as defined herein.
[0401] The polymeric barcoded reagent may comprise: at least 5, at least 10, at least 20, at least 25, at least 50, at least 75, at least 100, at least 200, at least 500, at least 1000, at least 5000, or at least 10,000 barcode regions, where each barcode region is as defined herein; and a barcoded oligonucleotide annealed to each barcode region, where each barcoded oligonucleotide is as defined herein. Preferably, the polymeric barcoded reagent comprises at least 5 barcode regions, where each barcode region is as defined herein; and a barcoded oligonucleotide annealed to each barcode region, where each barcoded oligonucleotide is as defined herein.
[0402] The polymeric barcoded reagent may comprise: at least 2, at least 3, at least 4, at least 5, at least 10, at least 20, at least 25, at least 50, at least 75, at least 100, at least 200, at least 500, at least 1000, at least 5000, at least 10 4one, at least 10 5 one, or at least 10 6 unique or distinct barcode regions, each barcode region as defined herein; and barcoded oligonucleotides annealed to each barcode region, each barcoded oligonucleotide as defined herein. Preferably, the polymeric barcoding reagent comprises at least 5 unique or distinct barcode regions, each barcode region as defined herein; and barcoded oligonucleotides annealed to each barcode region, each barcoded oligonucleotide as defined herein.
[0403] Figure 1 shows a polymeric barcoding reagent that comprises first (D1, E1, and F1) and second (D2, E2, and F2) barcode molecules, each barcode molecule comprising a nucleic acid sequence that contains a barcode region (E1 and E2). These first and second barcode molecules are linked together, for example, by a linking nucleic acid sequence (S). The polymeric barcoding reagent also comprises first (A1, B1, C1, and G1) and second (A2, B2, C2, and G2) barcoded oligonucleotides. These barcoded oligonucleotides each comprise a barcode region (B1 and B2) and a target region (G1 and G2).
[0404] The barcode regions within the barcoded oligonucleotides can each comprise a unique sequence that is not present in other barcoded oligonucleotides and can thus be used to uniquely identify each such barcode molecule. The target regions can be used to anneal the barcoded oligonucleotides to target nucleic acid fragments and can subsequently be used as primers for primer extension reactions or amplification reactions (such as polymerase chain reaction).
[0405] Each barcode molecule can also optionally comprise a 5' adaptor region (F1 and F2). The barcoded oligonucleotides can then also comprise 3' adaptor regions (C1 and C2) that are complementary to the 5' adaptor regions of the barcode molecules.
[0406] Each barcode molecule may also optionally contain 3' regions (D1 and D2), which may contain the same sequence within each barcode molecule. The barcoded oligonucleotide may then also contain 5' regions (A1 and A2) that are complementary to the 3' regions of the barcode molecule. These 3' regions can be used to manipulate or amplify nucleic acid sequences, such as sequences generated by labeling nucleic acid targets with barcoded oligonucleotides. The 3' region may contain at least 4, at least 5, at least 6, at least 8, at least 10, at least 15, at least 20, at least 25, at least 50, at least 100, or at least 250 nucleotides. Preferably, the 3' region contains at least 4 nucleotides. Preferably, each 3' region contains deoxyribonucleotides, and optionally, all nucleotides in the 3' region are deoxyribonucleotides. One or more deoxyribonucleotides can be modified deoxyribonucleotides (e.g., deoxyribonucleotides modified with a biotin moiety or deoxyuridine nucleotides). Each 3' region may contain one or more universal bases (e.g., inosine), one or more modified nucleotides, and / or one or more nucleotide analogs.
[0407] The present invention provides a library of polymeric barcoded reagents that includes at least 10 polymeric barcoded reagents for labeling target nucleic acids for sequencing, wherein each polymeric barcoded reagent includes: a first and a second barcode molecule contained within a (single) nucleic acid molecule, wherein each barcode molecule contains a nucleic acid sequence that includes a barcode region; and a first and a second barcoded oligonucleotide, wherein the first barcoded oligonucleotide optionally includes, in a 5' to 3' direction, a barcode region that is complementary to and annealed to the barcode region of the first barcode molecule and a target region that is capable of annealing or ligating to a first target nucleic acid fragment, and wherein the second barcoded oligonucleotide optionally includes, in a 5' to 3' direction, a barcode region that is complementary to and annealed to the barcode region of the second barcode molecule and a target region that is capable of annealing or ligating to a second target nucleic acid fragment. Preferably, the barcode regions of the first and second barcoded oligonucleotides of each polymeric barcoded reagent are different from the barcode regions of the barcoded oligonucleotides of at least 9 other polymeric barcoded reagents in the library.
[0408] 15. The polymeric barcoded reagent contains a barcoded oligonucleotide that anneals to a polymeric hybridization molecule
[0409] The present invention provides a polymeric barcoded reagent for labeling a target nucleic acid, wherein the reagent comprises: a first and a second hybridization molecule (i.e., a polymeric hybridization molecule) linked together, wherein each hybridization molecule comprises a nucleic acid sequence containing a hybridization region; and a first and a second barcoded oligonucleotide, wherein the first barcoded oligonucleotide optionally comprises, in the 5' to 3' direction, an adaptor region that anneals to the hybridization region of the first hybridization molecule, a barcode region, and a target region capable of annealing or ligating to a first target nucleic acid fragment, and the second barcoded oligonucleotide optionally comprises, in the 5' to 3' direction, an adaptor region that anneals to the hybridization region of the second hybridization molecule, a barcode region, and a target region capable of annealing or ligating to a second target nucleic acid fragment.
[0410] Optionally, the first and second barcoded oligonucleotides each comprise an adaptor region and a target region in a single continuous sequence that is complementary to and anneals to the hybridization region of the hybridization molecule and is also capable of annealing or ligating to the target nucleic acid fragment.
[0411] The present invention provides a polymeric barcoded reagent for labeling a target nucleic acid, wherein the reagent comprises: a first and a second hybridization molecule (i.e., a polymeric hybridization molecule) linked together, wherein each hybridization molecule comprises a nucleic acid sequence containing a hybridization region; and a first and a second barcoded oligonucleotide, wherein the first barcoded oligonucleotide optionally comprises, in the 5' to 3' direction, a barcode region, an adaptor region that anneals to the hybridization region of the first hybridization molecule, and a target region capable of annealing or ligating to a first target nucleic acid fragment, and wherein the second barcoded oligonucleotide optionally comprises, in the 5' to 3' direction, a barcode region, an adaptor region that anneals to the hybridization region of the second hybridization molecule, and a target region capable of annealing or ligating to a second target nucleic acid fragment.
[0412] Optionally, the first and second barcoded oligonucleotides each comprise an adaptor region and a target region in a single continuous sequence that is complementary to and anneals to the hybridization region of the hybridization molecule and is also capable of annealing or ligating to the target nucleic acid fragment.
[0413] The present invention provides a polymeric barcoded reagent for labeling a target nucleic acid, wherein the reagent comprises: a first and a second hybridization molecule (i.e., a polymeric hybridization molecule) linked together, wherein each hybridization molecule comprises a nucleic acid sequence containing a hybridization region; and a first and a second barcoded oligonucleotide, wherein the first barcoded oligonucleotide comprises (in the 5'-3' or 3'-5' direction) an adaptor region that anneals to the hybridization region of the first hybridization molecule, a barcode region, and a target region capable of ligating to a first target nucleic acid fragment, and wherein the second barcoded oligonucleotide comprises (in the 5'-3' or 3'-5' direction) an adaptor region that anneals to the hybridization region of the second hybridization molecule, a barcode region, and a target region capable of ligating to a second target nucleic acid fragment.
[0414] Optionally, each of the first and second barcoded oligonucleotides comprises a linker region and a target region in a single continuous sequence that is complementary to and anneals with the hybridization region of the hybridization molecule and is also capable of ligating to a target nucleic acid fragment.
[0415] The present invention provides a polymeric barcoded reagent for labeling a target nucleic acid, wherein the reagent comprises: a first and a second hybridization molecule (i.e., a polymeric hybridization molecule) linked together, wherein each hybridization molecule comprises a nucleic acid sequence containing a hybridization region; and a first and a second barcoded oligonucleotide, wherein the first barcoded oligonucleotide comprises (in the 5'-3' or 3'-5' direction) a barcode region, a linker region that anneals with the hybridization region of the first hybridization molecule, and a target region capable of ligating to a first target nucleic acid fragment, and wherein the second barcoded oligonucleotide comprises (in the 5'-3' or 3'-5' direction) a barcode region, a linker region that anneals with the hybridization region of the second hybridization molecule, and a target region capable of ligating to a second target nucleic acid fragment
[0416] Optionally, each of the first and second barcoded oligonucleotides comprises a linker region and a target region in a single continuous sequence that is complementary to and anneals with the hybridization region of the hybridization molecule and is also capable of ligating to a target nucleic acid fragment.
[0417] The present invention provides a polymeric barcoded reagent for labeling a target nucleic acid, wherein the reagent comprises: a first and a second hybridization molecule (i.e., a polymeric hybridization molecule) linked together, wherein each hybridization molecule comprises a nucleic acid sequence containing a barcode region; and a first and a second barcoded oligonucleotide, wherein the first barcoded oligonucleotide comprises in the 5' to 3' direction a linker region that anneals with the hybridization region of the first hybridization molecule, a barcode region, and a target region capable of annealing to a first target nucleic acid fragment, and wherein the second barcoded oligonucleotide comprises in the 5' to 3' direction a linker region that anneals with the hybridization region of the second hybridization molecule, a barcode region, and a target region capable of annealing to a second target nucleic acid fragment.
[0418] The present invention provides a polymeric barcoded reagent for labeling a target nucleic acid, wherein the reagent comprises: a first and a second hybridization molecule (i.e., a polymeric hybridization molecule) linked together, wherein each hybridization molecule comprises a nucleic acid sequence containing a barcode region; and a first and a second barcoded oligonucleotide, wherein the first barcoded oligonucleotide comprises in the 5' to 3' direction a barcode region, a linker region that anneals with the hybridization region of the first hybridization molecule, and a target region capable of annealing to a first target nucleic acid fragment, and wherein the second barcoded oligonucleotide comprises in the 5' to 3' direction a barcode region, a linker region that anneals with the hybridization region of the second hybridization molecule, and a target region capable of annealing to a second target nucleic acid fragment.
[0419] Optionally, each of the first and second coded oligonucleotides comprises a linker region and a target region in a single continuous sequence that is complementary to and anneals with a hybridization region of a hybridization molecule and is also capable of annealing with a target nucleic acid.
[0420] Preferably, the linker region of the first coded oligonucleotide comprises a sequence that is complementary to and anneals with a hybridization region of a first hybridization molecule, and the linker region of the second coded oligonucleotide comprises a sequence that is complementary to and anneals with a hybridization region of a second hybridization molecule. The complementary sequence of each coded oligonucleotide can be at least 5, at least 10, at least 15, at least 20, at least 25, at least 50, or at least 100 contiguous nucleotides.
[0421] The hybridization region of each hybridization molecule may comprise a constant region. Preferably, all hybridization regions of the polymeric barcoded reagent are substantially the same. Optionally, all hybridization regions of the polymeric barcoded reagent library are substantially the same. The hybridization region may comprise at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 8, at least 10, at least 15, at least 20, at least 25, at least 50, at least 100, or at least 250 nucleotides. Preferably, the hybridization region comprises at least 4 nucleotides. Preferably, each hybridization region comprises deoxyribonucleotides, and optionally, all nucleotides in the hybridization region are deoxyribonucleotides. One or more deoxyribonucleotides may be modified deoxyribonucleotides (e.g., deoxyribonucleotides modified with a biotin moiety or deoxyuridine nucleotides). Each hybridization region may comprise one or more universal bases (e.g., inosine), one or more modified nucleotides, and / or one or more nucleotide analogs.
[0422] The target region of the coded oligonucleotide may not anneal with the polymeric hybridization molecule. The target region of the coded oligonucleotide may not be complementary to the polymeric hybridization molecule.
[0423] The coded oligonucleotide may comprise a linker region between the linker region and the target region. The linker region may comprise one or more contiguous nucleotides that do not anneal with the polymeric hybridization molecule and are not complementary to the target nucleic acid fragment. The linker may comprise 1 to 100, 5 to 75, 10 to 50, 15 to 30, or 20 to 25 non-complementary nucleotides. Preferably, the linker comprises 15 to 30 non-complementary nucleotides. The use of such a linker region enhances the efficiency of the barcoding reaction using the polymeric barcoded reagent.
[0424] The hybrid molecule may also comprise one or more nucleic acid sequences that are not complementary to the barcoded oligonucleotide. For example, the hybrid molecule may comprise one or more linker regions. The hybrid molecule may, for example, comprise a linker region 5' (5' linker region) and / or a linker region 3' (3' linker region) of the hybridization region. The linker region may be used to manipulate, purify, retrieve, amplify, and / or detect the hybrid molecule.
[0425] The linker region of each hybrid molecule may comprise a constant region. Optionally, all linker regions of the multimeric hybridization reagent are substantially identical. The linker region may comprise at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 8, at least 10, at least 15, at least 20, at least 25, at least 50, at least 100, or at least 250 nucleotides. Preferably, the linker region comprises at least 4 nucleotides. Preferably, each linker region comprises deoxyribonucleotides, and optionally, all nucleotides in the linker region are deoxyribonucleotides. One or more deoxyribonucleotides may be modified deoxyribonucleotides (e.g., deoxyribonucleotides modified with a biotin moiety or deoxyuridine nucleotides). Each linker region may comprise one or more universal bases (e.g., inosine), one or more modified nucleotides, and / or one or more nucleotide analogs.
[0426] The barcoded oligonucleotide may comprise a linker region between the linker region and the target region. The linker region may comprise one or more consecutive nucleotides that do not anneal to the multimeric hybridization molecule and are not complementary to the target nucleic acid fragment. The linker may comprise 1 to 100, 5 to 75, 10 to 50, 15 to 30, or 20 to 25 non-complementary nucleotides. Preferably, the linker comprises 15 to 30 non-complementary nucleotides. The use of such a linker region enhances the efficiency of the barcoding reaction using the multimeric barcoding reagent.
[0427] The present invention provides a library of multimeric barcoding reagents comprising at least 10 multimeric barcoding reagents for labeling target nucleic acid polymers for sequencing, wherein each multimeric barcoding reagent comprises: a first and a second hybrid molecule contained within a (single) nucleic acid molecule, wherein each hybrid molecule comprises a nucleic acid sequence containing a hybridization region; and a first and a second barcoded oligonucleotide, wherein the first barcoded oligonucleotide optionally comprises, in the 5' to 3' direction, a linker region complementary to and annealing to the hybridization region of the first hybrid molecule, a barcode region, and a target region capable of annealing or ligating to a first target nucleic acid fragment, and wherein the second barcoded oligonucleotide optionally comprises, in the 5' to 3' direction, a linker region complementary to and annealing to the hybridization region of the second hybrid molecule, a barcode region, and a target region capable of annealing or ligating to a second target nucleic acid fragment.
[0428] Preferably, the barcode regions of the first and second barcoding oligonucleotides of each polymeric barcoding reagent are different from the barcode regions of the barcoding oligonucleotides of at least nine other polymeric barcoding reagents in the library.
[0429] The present invention provides a library of polymeric barcoding reagents comprising at least 10 polymeric barcoding reagents for labeling target nucleic acids for sequencing, wherein each polymeric barcoding reagent comprises: a first and a second hybridization molecule contained within a (single) nucleic acid molecule, wherein each hybridization molecule comprises a nucleic acid sequence containing a hybridization region; and a first and a second barcoding oligonucleotide, wherein the first barcoding oligonucleotide optionally comprises, in the 5' to 3' direction, a barcode region, an adaptor region complementary to and annealed to the hybridization region of the first hybridization molecule, and a target region capable of annealing or ligating to a first target nucleic acid fragment, and wherein the second barcoding oligonucleotide optionally comprises, in the 5' to 3' direction, a barcode region, an adaptor region complementary to and annealed to the hybridization region of the second hybridization molecule, and a target region capable of annealing or ligating to a second target nucleic acid fragment. Preferably, the barcode regions of the first and second barcoding oligonucleotides of each polymeric barcoding reagent are different from the barcode regions of the barcoding oligonucleotides of at least nine other polymeric barcoding reagents in the library.
[0430] 16. A polymeric barcoding reagent comprises barcoding oligonucleotides linked by a macromolecule
[0431] The present invention provides a polymeric barcoding reagent for labeling a target nucleic acid, wherein the reagent comprises a first and a second barcoding oligonucleotide linked together by a macromolecule, and wherein the barcoding oligonucleotides each comprise a barcode region.
[0432] The first barcoding oligonucleotide may further comprise a target region capable of annealing or ligating to a first target nucleic acid fragment, and the second barcoding oligonucleotide may further comprise a target region capable of annealing or ligating to a second target nucleic acid fragment.
[0433] The first barcoding oligonucleotide may comprise, in the 5'-3' direction, a barcode region and a target region capable of annealing to a first target nucleic acid fragment, and the second barcoding oligonucleotide may comprise, in the 5'-3' direction, a barcode region and a target region capable of annealing to a second target nucleic acid fragment.
[0434] The barcoding oligonucleotide may further comprise any of the features described herein.
[0435] The barcoding oligonucleotides may be linked by a macromolecule by binding to and / or annealing to the macromolecule.
[0436] Barcoded oligonucleotides can be associated with macromolecules directly or indirectly (e.g., via a linker molecule). Barcoded oligonucleotides can be associated by binding to a macromolecule and / or by binding or annealing to a linker molecule that is bound to a macromolecule. Barcoded oligonucleotides can bind to a macromolecule (or linker molecule) by covalent linkage, non-covalent linkage (e.g., protein-protein interaction or streptavidin-biotin bond), or nucleic acid hybridization. The linker molecule can be a biopolymer (e.g., a nucleic acid molecule) or a synthetic polymer. The linker molecule can comprise one or more ethylene glycol and / or poly(ethylene glycol) (e.g., hexaethylene glycol or pentaethylene glycol) units. The linker molecule can comprise one or more ethyl groups, e.g., a C3 (three-carbon) spacer, a C6 spacer, a C12 spacer, or a C18 spacer.
[0437] The macromolecule can be a synthetic polymer (e.g., a dendrimer) or a biopolymer such as a nucleic acid (e.g., a single-stranded nucleic acid, e.g., single-stranded DNA), a peptide, a polypeptide, or a protein (e.g., a multimeric protein).
[0438] The dendrimer can comprise at least 2 generations, at least 3 generations, at least 5 generations, or at least 10 generations.
[0439] The macromolecule can be a nucleic acid comprising two or more nucleotides, each nucleotide capable of binding to a barcoded oligonucleotide. As a supplement or alternative, the nucleic acid can comprise two or more regions, each region capable of hybridizing to a barcoded oligonucleotide.
[0440] The nucleic acid can comprise a first and a second modified nucleotide, wherein each modified nucleotide comprises a binding moiety (e.g., a biotin moiety, or an alkyne moiety useful for click chemistry reactions) capable of binding to a barcoded oligonucleotide. Optionally, the first and second modified nucleotides can be separated by an intervening nucleic acid sequence of at least one, at least two, at least 5, or at least 10 nucleotides.
[0441] The nucleic acid can comprise a first hybridization region and a second hybridization region, wherein each hybridization region comprises a sequence complementary to and capable of hybridizing to the sequence of at least one nucleotide within a barcoded oligonucleotide. The complementary sequence can be at least 5, at least 10, at least 15, at least 20, at least 25, or at least 50 consecutive nucleotides. Optionally, the first hybridization region and the second hybridization region can be separated by an intervening nucleic acid sequence of at least one, at least two, at least 5, or at least 10 nucleotides.
[0442] The macromolecule can be a protein, e.g., a multimeric protein, e.g., a homomeric protein or a heteromeric protein. For example, the protein can comprise streptavidin, e.g., tetrameric streptavidin.
[0443] Also provided is a library of polymeric barcoding reagents that include barcoded oligonucleotides linked by a macromolecule. Such libraries can be based on the general properties of the polymeric barcoding reagent libraries described herein. In the library, each polymeric barcoding reagent can include a different macromolecule.
[0444] 17. The polymeric barcoding reagent includes barcoded oligonucleotides linked by a solid support or a semi-solid support
[0445] The present invention provides a polymeric barcoding reagent for labeling a target nucleic acid, wherein the reagent includes a first and a second barcoded oligonucleotide linked together by a solid support or a semi-solid support, and wherein the barcoded oligonucleotides each include a barcode region.
[0446] The first barcoded oligonucleotide may further include a target region capable of annealing or ligating to a first target nucleic acid fragment, and the second barcoded oligonucleotide may further include a target region capable of annealing or ligating to a second target nucleic acid fragment.
[0447] The first barcoded oligonucleotide may include the barcode region and a target region capable of annealing to a first target nucleic acid fragment in the 5'-3' direction, and the second barcoded oligonucleotide may include the barcode region and a target region capable of annealing to a second target nucleic acid fragment in the 5'-3' direction.
[0448] The barcoded oligonucleotides may further include any of the features described herein.
[0449] The barcoded oligonucleotides may be linked by a solid support or a semi-solid support. The barcoded oligonucleotides may be directly or indirectly (e.g., through a linker molecule) linked to the support. The barcoded oligonucleotides may be linked by binding to the support and / or by binding or annealing to a linker molecule bound to the support. The barcoded oligonucleotides may be bound to the support (or linker molecule) by covalent linkage, non-covalent linkage (e.g., protein-protein interaction or streptavidin-biotin bond), or nucleic acid hybridization. The linker molecule may be a biopolymer (e.g., a nucleic acid molecule) or a synthetic polymer. The linker molecule may include one or more ethylene glycol and / or poly(ethylene glycol) (e.g., hexaethylene glycol or pentaethylene glycol) units. The linker molecule may include one or more ethyl groups, such as a C3 (three-carbon) spacer, a C6 spacer, a C12 spacer, or a C18 spacer.
[0450] The support may comprise a flat surface. The support can be, for example, a slide, such as a glass slide. The slide can be a flow cell for sequencing. If the support is a slide, the first and second barcoded oligonucleotides can be immobilized in discrete regions on the slide. Optionally, the barcoded oligonucleotides of each polymeric barcoding reagent in the library are immobilized in different discrete regions on the slide relative to the barcoded oligonucleotides of other polymeric barcoding reagents in the library. The support can be a plate containing wells, optionally wherein the first and second barcoded oligonucleotides are immobilized in the same well. Optionally, the barcoded oligonucleotides of each polymeric barcoding reagent in the library are immobilized in different wells of the plate relative to the barcoded oligonucleotides of other polymeric barcoding reagents in the library.
[0451] Preferably, the support is a bead (such as a gel bead). The bead can be an agarose bead, a silica bead, a styrofoam bead, a gel bead (such as those obtainable from ), an antibody-conjugated bead, an oligo-dT-conjugated bead, a streptavidin bead, or a magnetic bead (such as a superparamagnetic bead). The bead can have any size and / or molecular structure. For example, the bead can be from 10 nanometers to 100 micrometers in diameter, from 100 nanometers to 10 micrometers in diameter, or from 1 micrometer to 5 micrometers in diameter. Optionally, the bead is about 10 nanometers in diameter, about 100 nanometers in diameter, about 1 micrometer in diameter, about 10 micrometers in diameter, or about 100 micrometers in diameter. The bead can be solid, or alternatively the bead can be hollow or partially hollow or porous. For certain barcoding methods, certain sizes of beads can be most preferred. For example, beads less than 5.0 micrometers or less than 1.0 micrometer in diameter can be most useful for barcoding nucleic acid targets within individual cells. Preferably, the barcoded oligonucleotides of each polymeric barcoding reagent in the library are associated together on different beads relative to the barcoded oligonucleotides of other polymeric barcoding reagents in the library.
[0452] The support can be functionalized to enable the attachment of two or more barcoded oligonucleotides. Such functionalization can be achieved by adding a chemical moiety (such as a carboxyl group, an alkyne, an azide, an acrylate group, an amino group, a sulfate group, or a succinimide group) and / or a protein-based moiety (such as streptavidin, avidin, or protein G) to the support. The barcoded oligonucleotide can be attached directly or indirectly (such as through a linker molecule) to the moiety.
[0453] The functionalized support (such as a bead) can be contacted with a solution of barcoded oligonucleotides under conditions that promote the attachment of two or more individual barcoded oligonucleotides to each bead in the solution (yielding a polymeric barcoding reagent).
[0454] Also provided is a library of polymeric barcoding reagents comprising barcoded oligonucleotides linked by a support. Such libraries can be based on the general properties of the polymeric barcoding reagent libraries described herein. In the library, each polymeric barcoding reagent can comprise a different support (e.g., different labeled beads). In a library of polymeric barcoding reagents, the barcoded oligonucleotides of each polymeric barcoding reagent in the library can be linked to different supports relative to the barcoded oligonucleotides of other polymeric barcoding reagents in the library.
[0455] 18. The polymeric barcoding reagent comprises barcoded oligonucleotides linked together by inclusion in a lipid carrier
[0456] The present invention provides a polymeric barcoding reagent for labeling a target nucleic acid, wherein the reagent comprises a first and a second barcoded oligonucleotide and a lipid carrier, wherein the first and second barcoded oligonucleotides are linked together by inclusion in the lipid carrier, and wherein each of the barcoded oligonucleotides comprises a barcode region.
[0457] The first barcoded oligonucleotide can further comprise a target region capable of annealing or ligating to a first target nucleic acid fragment, and the second barcoded oligonucleotide can further comprise a target region capable of annealing or ligating to a second target nucleic acid fragment..
[0458] The first barcoded oligonucleotide can comprise a barcode region and a target region capable of annealing to a first target nucleic acid fragment in the 5'-3' direction, and the second barcoded oligonucleotide can comprise a barcode region and a target region capable of annealing to a second target nucleic acid fragment in the 5'-3' direction.
[0459] The barcoded oligonucleotides can further comprise any of the features described herein.
[0460] The present invention provides a library of polymeric barcoding reagents, which comprises first and second polymeric barcoding reagents as defined herein, wherein the barcoded oligonucleotides of the first polymeric barcoding reagent are contained in a first lipid carrier, and wherein the barcoded oligonucleotides of the second polymeric barcoding reagent are contained in a second lipid carrier, and wherein the barcode region of the barcoded oligonucleotides of the first polymeric barcoding reagent is different from the barcode region of the barcoded oligonucleotides of the second polymeric barcoding reagent.
[0461] The library of polymeric barcoding reagents can comprise at least 5, at least 10, at least 20, at least 25, at least 50, at least 75, at least 100, at least 250, at least 500, at least 10 3 at least 10 4 at least 10 5 at least 10 6 at least 10 7one, at least 10 8 or at least 10 9 polyplex barcode reagents as defined herein. Preferably, the barcode regions of the first and second barcode oligonucleotides of each polyplex barcode reagent are different from the barcode regions of the barcode oligonucleotides of at least 9 other polyplex barcode reagents in the library.
[0462] The barcode oligonucleotides of each polyplex barcode reagent are contained in different lipid carriers.
[0463] The lipid carrier can be a liposome or a micelle. The lipid carrier can be a phospholipid carrier. The lipid carrier can comprise one or more amphiphilic molecules. The lipid carrier can comprise one or more phospholipids. The phospholipid can be phosphatidylcholine. The lipid carrier can comprise one or more of the following components: phosphatidylethanolamine, phosphatidylserine, cholesterol, cardiolipin, dicetyl phosphate, stearylamine, phosphatidylglycerol, dipalmitoyl phosphatidylcholine, distearyl phosphatidylcholine, and / or any related and / or derivative molecules thereof. Optionally, the lipid carrier can comprise any combination of two or more of the above components, with or without other components.
[0464] The lipid carrier (e.g., liposome or micelle) can be monolayer or multilayer. The polyplex barcode reagent library can comprise both monolayer lipid carriers and multilayer lipid carriers. The lipid carrier can comprise a copolymer, e.g., a block copolymer.
[0465] The lipid carrier can comprise at least 2, at least 3, at least 5, at least 10, at least 50, at least 100, at least 500, at least 1000, at least 10,000, or at least 100,000 barcode oligonucleotides, or any greater number of barcode oligonucleotides.
[0466] Any lipid carrier (e.g., liposome or micelle, and / or liposome reagent or micelle reagent) can be complexed with 1, or less than 1, or more than 1 polyplex barcode reagent, on average, to form a library of such polyplex barcode reagents.
[0467] The present invention provides a library of polyplex barcode reagents comprising at least 10 polyplex barcode reagents as defined herein, wherein each polyplex barcode reagent comprises first and second barcode oligonucleotides contained in different lipid carriers, and wherein the barcode regions of the first and second barcode oligonucleotides of each polyplex barcode reagent are different from the barcode regions of the barcode oligonucleotides of at least 9 other polyplex barcode reagents in the library.
[0468] Methods for preparing polymeric barcoded reagents include loading barcoded oligonucleotides and / or polymeric barcoded reagents into lipid carriers (such as liposomes or micelles). The methods may include steps of passive, active, and / or remote loading. Preformed lipid carriers (such as liposomes and / or micelles) may be loaded by contacting them with a solution of barcoded oligonucleotides and / or polymeric barcoded reagents. Lipid carriers (such as liposomes and / or micelles) may be loaded by contacting them with a solution of barcoded oligonucleotides and / or polymeric barcoded reagents before and / or during the formation or synthesis of the lipid carrier. The methods may include passive encapsulation and / or trapping of barcoded oligonucleotides and / or polymeric barcoded reagents in the lipid carrier.
[0469] Lipid carriers (such as liposomes and / or micelles) may be prepared by sonication-based methods, French press-based methods, reverse phase methods, solvent evaporation methods, extrusion-based methods, mechanical mixing-based methods, freeze / thaw-based methods, dehydration / rehydration-based methods, and / or any combination thereof.
[0470] Lipid carriers (such as liposomes and / or micelles) may be stabilized and / or stored using known methods prior to use.
[0471] Any polymeric barcoded reagent or kit described herein may comprise a lipid carrier.
[0472] 19. A kit comprising a polymeric barcoded reagent and an adaptor oligonucleotide
[0473] The present invention also provides a kit comprising one or more components defined herein. The present invention also provides a kit that is particularly suitable for implementing any method defined herein.
[0474] The present invention also provides a kit for labeling a target nucleic acid, wherein the kit comprises: (a) a polymeric barcoding reagent, which comprises (i) a first and a second barcoding molecule (i.e., polymeric barcoding molecules) linked together, wherein each barcoding molecule comprises a nucleic acid sequence optionally containing an adaptor region and a barcode region in the 5' to 3' direction, and (ii) a first and a second barcoded oligonucleotide, wherein the first barcoded oligonucleotide comprises a barcode region annealing to the barcode region of the first barcoding molecule, and wherein the second barcoded oligonucleotide comprises a barcode region annealing to the barcode region of the second barcoding molecule; and (b) a first and a second adaptor oligonucleotide, wherein the first adaptor oligonucleotide optionally contains an adaptor region capable of annealing to the adaptor region of the first barcoding molecule and a target region capable of annealing or ligating to a first target nucleic acid fragment in the 5' to 3' direction, and wherein the second adaptor oligonucleotide optionally contains an adaptor region capable of annealing to the adaptor region of the second barcoding molecule and a target region capable of annealing or ligating to a second target nucleic acid fragment in the 5' to 3' direction.
[0475] The present invention also provides a kit for labeling a target nucleic acid, wherein the kit comprises: (a) a polymeric barcoding reagent, which comprises (i) a first and a second barcoding molecule (i.e., polymeric barcoding molecules) linked together, wherein each barcoding molecule comprises a nucleic acid sequence containing an adaptor region and a barcode region, and (ii) a first and a second barcoded oligonucleotide, wherein the first barcoded oligonucleotide comprises a barcode region annealing to the barcode region of the first barcoding molecule, and wherein the second barcoded oligonucleotide comprises a barcode region annealing to the barcode region of the second barcoding molecule; and (b) a first and a second adaptor oligonucleotide, wherein the first adaptor oligonucleotide comprises an adaptor region capable of annealing to the adaptor region of the first barcoding molecule and a target region capable of ligating to a first target nucleic acid fragment, and wherein the second adaptor oligonucleotide comprises an adaptor region capable of annealing to the adaptor region of the second barcoding molecule and a target region capable of ligating to a second target nucleic acid fragment.
[0476] The present invention also provides a kit for labeling a target nucleic acid, wherein the kit comprises: (a) a polymeric barcoding reagent, which comprises (i) a first and a second barcode molecule (i.e., polymeric barcode molecule) linked together, wherein each barcode molecule comprises a nucleic acid sequence optionally comprising a linker region and a barcode region in the 5'-to-3' direction, and (ii) a first and a second barcoded oligonucleotide, wherein the first barcoded oligonucleotide comprises a barcode region annealing to the barcode region of the first barcode molecule, and wherein the second barcoded oligonucleotide comprises a barcode region annealing to the barcode region of the second barcode molecule; and (b) a first and a second linker oligonucleotide, wherein the first linker oligonucleotide comprises, in the 5'-to-3' direction, a linker region capable of annealing to the linker region of the first barcode molecule and a target region capable of annealing to a first target nucleic acid fragment, and wherein the second linker oligonucleotide comprises, in the 5'-to-3' direction, a linker region capable of annealing to the linker region of the second barcode molecule and a target region capable of annealing to a second target nucleic acid fragment.
[0477] The present invention also provides a kit for labeling a target nucleic acid, wherein the kit comprises: (a) a polymeric barcoding reagent, which comprises (i) a first and a second barcode molecule (i.e., polymeric barcode molecule) linked together, wherein each barcode molecule comprises a nucleic acid sequence optionally comprising a linker region and a barcode region in the 5'-to-3' direction, and (ii) a first and a second barcoded oligonucleotide, wherein the first barcoded oligonucleotide comprises a barcode region annealing to the barcode region of the first barcode molecule, and wherein the second barcoded oligonucleotide comprises a barcode region annealing to the barcode region of the second barcode molecule; and (b) a first and a second linker oligonucleotide, wherein the first linker oligonucleotide comprises a linker region capable of annealing to the linker region of the first barcode molecule and capable of ligating to a first target nucleic acid fragment, and wherein the second linker oligonucleotide comprises a linker region capable of annealing to the linker region of the second barcode molecule and capable of ligating to a second target nucleic acid fragment.
[0478] Each linker oligonucleotide may consist essentially of or consist of a linker region. Each linker oligonucleotide may not comprise a target region.
[0479] Preferably, the linker region of the first linker oligonucleotide comprises a sequence complementary to and capable of annealing to the linker region of the first barcode molecule, and the linker region of the second linker oligonucleotide comprises a sequence complementary to and capable of annealing to the linker region of the second barcode molecule. The complementary sequence of each linker oligonucleotide may be at least 5, at least 10, at least 15, at least 20, at least 25, at least 50 or at least 100 consecutive nucleotides.
[0480] The target region of the adapter oligonucleotide may not be able to anneal to the polymeric barcode molecule. The target region of the adapter oligonucleotide may not be complementary to the polymeric barcode molecule.
[0481] The target region of each adapter oligonucleotide may comprise a different sequence. Each target region may comprise a sequence capable of annealing to only a single target nucleic acid fragment within the nucleic acid sample. Each target region may comprise one or more random sequences or one or more degenerate sequences such that the target region is capable of annealing to more than one target nucleic acid fragment. Each target region may comprise at least 5, at least 10, at least 15, at least 20, at least 25, at least 50 or at least 100 nucleotides. Preferably, each target region comprises at least 5 nucleotides. Each target region may comprise 5 to 100 nucleotides, 5 to 10 nucleotides, 10 to 20 nucleotides, 20 to 30 nucleotides, 30 to 50 nucleotides, 50 to 100 nucleotides, 10 to 90 nucleotides, 20 to 80 nucleotides, 30 to 70 nucleotides or 50 to 60 nucleotides. Preferably, each target region comprises 30 to 70 nucleotides. Preferably, each target region comprises deoxyribonucleotides, optionally all of the nucleotides in the target region are deoxyribonucleotides. One or more deoxyribonucleotides may be modified deoxyribonucleotides (e.g., deoxyribonucleotides modified with a biotin moiety or deoxyuridine nucleotides). Each target region may comprise one or more universal bases (e.g., inosine), one or more modified nucleotides and / or one or more nucleotide analogs.
[0482] The target region can be used to anneal the adapter oligonucleotide to a fragment of the target nucleic acid and can subsequently be used as a primer for a primer extension reaction or an amplification reaction (e.g., polymerase chain reaction). Alternatively, the target region can be used to ligate the adapter oligonucleotide to a fragment of the target nucleic acid. The target region may be located at the 5' end of the adapter oligonucleotide. Such a target region may be phosphorylated. This enables the 5' end of the target region to be ligated to the 3' end of the target nucleic acid fragment.
[0483] The adapter oligonucleotide may comprise a linker region between the adapter region and the target region. The linker region may comprise one or more consecutive nucleotides that do not anneal to the first and second barcode molecules (i.e., the polymeric barcode molecule) and are not complementary to the target nucleic acid fragment. The linker may comprise 1 to 100, 5 to 75, 10 to 50, 15 to 30 or 20 to 25 non-complementary nucleotides. Preferably, the linker comprises 15 to 30 non-complementary nucleotides. The use of such a linker region enhances the efficiency of the barcoding reaction performed using the kits described herein.
[0484] Each component of the kit may be in any form defined herein.
[0485] The polymeric barcoded reagent and the adapter oligonucleotide can be provided in the kit as physically separate components.
[0486] The kit can comprise: (a) a polymeric barcoded reagent comprising at least 5, at least 10, at least 20, at least 25, at least 50, at least 75 or at least 100 barcoded molecules linked together, wherein each barcoded molecule is as defined herein; and (b) an adapter oligonucleotide capable of annealing to each barcoded molecule, wherein each adapter oligonucleotide is as defined herein.
[0487] Figure 2 Shown is a kit comprising a polymeric barcoded reagent and an adapter oligonucleotide for labeling a target nucleic acid. More specifically, the kit comprises first (D1, E1 and F1) and second (D2, E2 and F2) barcoded molecules, each barcoded molecule incorporating a barcode region (E1 and E2) and a 5'-adapter region (F1 and F2). In this embodiment, these first and second barcoded molecules are linked together by a linking nucleic acid sequence (S).
[0488] The kit further comprises first (A1 and B1) and second (A2 and B2) barcoded oligonucleotides, each comprising a barcode region (B1 and B2), and a 5'-region (A1 and A2). The 5'-region of each barcoded oligonucleotide is complementary to the 3'-region (D1 and D2) of the barcoded molecule and can thus anneal thereto. The barcode regions (B1 and B2) are complementary to the barcode regions (E1 and E2) of the barcoded molecule and can thus anneal thereto.
[0489] The kit further comprises first (C1 and G1) and second (C2 and G2) adapter oligonucleotides, wherein each adapter oligonucleotide comprises an adapter region (C1 and C2) complementary to the 5'-adapter region (F1 and F2) of the barcoded molecule and capable of annealing thereto. These adapter oligonucleotides can be synthesized to contain a 5'-terminal phosphate group. Each adapter oligonucleotide further comprises a target region (G1 and G2), which can be used to anneal the barcoded-adapter oligonucleotides (A1, B1, C1 and G1, and A2, B2, C2 and G2) to the target nucleic acid and can subsequently be used as a primer for a primer extension reaction or a polymerase chain reaction.
[0490] The kit can comprise two or more libraries of polymeric barcoded reagents, wherein each polymeric barcoded reagent is as defined herein, and adapter oligonucleotides for each polymeric barcoded reagent, wherein each adapter oligonucleotide is as defined herein. The barcode regions of the first and second barcoded oligonucleotides of the first polymeric barcoded reagent are different from the barcode regions of the first and second barcoded oligonucleotides of the second polymeric barcoded reagent.
[0491] The kit may comprise a library of at least 5, at least 10, at least 20, at least 25, at least 50, at least 75, at least 100, at least 250, at least 500, at least 10 3 ones, at least 10 4 ones, at least 10 5 ones, at least 10 6 ones, at least 10 7 ones, at least 10 8 ones or at least 10 9 ones of the polynucleotide barcoded reagents as defined herein. Preferably, the kit comprises a library of at least 10 polynucleotide barcoded reagents as defined herein. The kit may further comprise an adapter oligonucleotide for each polynucleotide barcoded reagent, wherein each adapter oligonucleotide may take th...
Claims
1. A method for measuring modified nucleotides or nucleobases in a sample containing circulating microparticles, wherein the circulating microparticles are membrane vesicles, wherein the circulating microparticles contain at least two genomic DNA fragments, and wherein the method comprises: (a) Link at least two of the at least two genomic DNA fragments to produce a set of at least two linked genomic DNA fragments; And (b) Measure the modified nucleotides or nucleobases in each linked fragment in the set to produce at least two linked measurements of the modified nucleotides or nucleobases, wherein the modified nucleotides or nucleobases are 5-methylcytosine or 5-hydroxymethylcytosine.
2. The method of claim 1, wherein step (b) comprises an enrichment process to enrich genomic DNA fragments containing the modified nucleotide or nucleobase, and wherein the measurement is performed using an enrichment probe that specifically or preferentially binds 5-methylcytosine or 5-hydroxymethylcytosine in the at least two genomic DNA fragments.
3. The method of claim 1, wherein the measurement is performed using a bisulfite conversion process in step (b).
4. The method of claim 1, wherein step (b) comprises sequencing each associated fragment in the group to generate at least two associated sequence reads.
5. The method of claim 1 or claim 2, wherein at least 3, at least 4, at least 5, at least 10, at least 50, at least 100, at least 500, at least 1000, at least 5000, at least 10,000, at least 100,000, or at least 1,000,000 genomic DNA fragments of the circulating microparticles are associated to generate at least 3, at least 4, at least 5, at least 10, at least 50, at least 100, at least 500, at least 1000, at least 5000, at least 10,000, at least 100,000, or at least 1,000,000 associated measurement values of the modified nucleotide or nucleobase.
6. The method of claim 1 or claim 2, wherein the sample contains first and second circulating microparticles, wherein each circulating microparticle contains at least two genomic DNA fragments, and wherein the method comprises performing step (a) to generate a first set of associated genomic DNA fragments of the first circulating microparticle and a second set of associated genomic DNA fragments of the second circulating microparticle, and performing step (b) to generate a first set of associated measurement values of the modified nucleotide or nucleobase of the first circulating microparticle and a second set of associated measurement values of the modified nucleotide or nucleobase of the second circulating microparticle.
7. The method according to claim 1 or claim 2, wherein the sample comprises n circular particles, wherein each circular particle contains at least two genomic DNA fragments, and wherein the method comprises performing step (a) to generate n sets of associated genomic DNA fragments, one set for each of the n circular particles, and performing step (b) to generate n sets of associated measurements of the modified nucleotide or nucleobase, one set for each of the n circular particles.
8. The method according to claim 7, wherein n is at least 3, at least 5, at least 10, at least 50, at least 100, at least 1000, at least 10,000, at least 100,000, at least 1,000,000, at least 10,000,000, or at least 100,000,000 circular particles.
9. The method according to claim 6, wherein prior to step (a), the method further comprises the step of dispensing the sample into at least two different reaction volumes.
10. The method according to claim 1, wherein step (a) comprises attaching at least two genomic DNA fragments of the circular particle to a barcode sequence to generate a set of associated genomic DNA fragments.
11. The method according to claim 10, wherein prior to the step of attaching at least two genomic DNA fragments of the circular particle to a barcode sequence, the method comprises attaching a coupling sequence to each genomic DNA fragment of the circular particle, wherein the coupling sequence is subsequently attached to the barcode sequence to generate the set of associated genomic DNA fragments.
12. The method according to claim 10, wherein the sample comprises first and second circular particles, wherein each circular particle contains at least two genomic DNA fragments, and wherein the method comprises performing step (a) to generate a first set of associated genomic DNA fragments of the first circular particle and a second set of associated genomic DNA fragments of the second circular particle, and performing step (b) to generate a first set of associated measurements of the modified nucleotide or nucleobase of the first circular particle and a second set of associated measurements of the modified nucleotide or nucleobase of the second circular particle, wherein at least two of the associated measurements of the modified nucleotide or nucleobase of the first circular particle are associated by a barcode sequence different from at least two of the associated measurements of the modified nucleotide or nucleobase of the second circular particle.
13. The method according to claim 12, wherein prior to the attachment step, the method further comprises the step of dispensing the sample into at least two different reaction volumes.
14. The method according to claim 1, wherein step (a) comprises attaching each of at least two genomic DNA fragments of the circulating particles to a different barcode sequence of a barcode sequence set to produce a set of associated genomic DNA fragments.
15. The method according to claim 14, wherein prior to the step of attaching each of at least two genomic DNA fragments of the circulating particles to different barcode sequences, the method comprises attaching a coupling sequence to each genomic DNA fragment of the circulating particles, wherein each of at least two genomic DNA fragments of the circulating particles is attached to a different barcode sequence of the barcode sequence set through its coupling sequence.
16. The method according to claim 14, wherein the sample comprises first and second circulating particles, each circulating particle containing at least two genomic DNA fragments, and wherein the method comprises performing step (a) to produce a first set of associated genomic DNA fragments of the first circulating particles and a second set of associated genomic DNA fragments of the second circulating particles, and performing step (b) to produce a first set of associated measurements of modified nucleotides or nucleobases of the first circulating particles and a second set of associated measurements of modified nucleotides or nucleobases of the second circulating particles, wherein the first set of associated measurements of modified nucleotides or nucleobases is associated through a barcode sequence different from the barcode sequence associated with the second set of associated measurements of modified nucleotides or nucleobases.
17. The method according to claim 16, wherein prior to the attachment step, the method further comprises the step of dispensing the sample into at least two different reaction volumes.
Citation Information
Patent Citations
Method for linking and characterising linked nucleic acids in a composition
EP2626433A1
Methods and compositions for nucleic acid analysis
US20120316074A1
Whole genome sequencing of a human fetus
US20150105267A1
Personalized Tumor Biomarkers
US20150344970A1
Non-invasive prenatal diagnosis
US6258540B1
Cited By
Reagents and methods for analysis of associated nucleic acids
CN120665986A