Long-indexed linked read generation on transposome-bound beads
The use of bead-bound transposomes and droplet generators for indexed linked read generation addresses the challenge of interpreting large genomic data by providing efficient nucleic acid indexing and amplification, resulting in improved DNA quality and reduced fragmentation for long-read sequencing.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- ILLUMINA INC
- Filing Date
- 2026-04-30
- Publication Date
- 2026-07-29
AI Technical Summary
Next-generation sequencers generate large amounts of genomic data that are difficult to interpret and analyze, and existing methods for nucleic acid sequencing lack efficient techniques for indexing and amplification, particularly for long-read sequencing.
A system and method using bead-bound transposomes and droplet generators for indexed linked read generation, involving hydrogel beads with crosslinking agents, transposomes, and droplet indexing to perform on-bead tagmentation and PCR, enabling controlled insert size and efficient DNA transfer without fragmentation.
This approach allows for the generation of long-indexed reads with improved control over library inserts, reduced DNA fragmentation, and efficient DNA preparation, facilitating simultaneous assays on millions of nucleic acid molecules with enhanced DNA quality and compatibility with lysates.
Smart Images

Figure 2026123197000003 
Figure 2026123197000004 
Figure 2026123197000005
Abstract
Description
Technical Field
[0001] (Cross - reference to Related Applications) This application claims the benefit of U.S. Provisional Patent Application No. 63 / 145,902, filed Feb. 4, 2021, entitled "LONG INDEXED - LINKED READ GENERATION ON TRANSPOSOME BOUND BEADS", which is hereby incorporated by reference in its entirety.
[0002] (Reference to Sequence Listing) This application is filed with a sequence listing in electronic format. The sequence listing is provided as a file named Sequence_Listing_ILLINC_406WO, created on Feb. 3, 2022, and having a size of 2.7 kilobytes. The information in the electronic format of the sequence listing is hereby incorporated by reference in its entirety.
[0003] (Field of the Invention) The systems, methods, and compositions provided herein relate to compositions, systems, and methods for spatial indexing sequencing and nucleic acid library preparation.
Background Art
[0004] [[ID=ID=26]]The detection of specific nucleic acid sequences present in biological samples has been used, for example, as a method for identifying and classifying microorganisms, diagnosing infectious diseases, detecting and characterizing genetic abnormalities, identifying genetic changes associated with cancer, examining genetic susceptibility to diseases, and measuring responses to various types of treatments.General techniques for detecting specific nucleic acid sequences in biological samples are nucleic acid sequencing.
[0005] Next - generation sequencers are powerful tools that generate large amounts of genomic data for each sequencing run. Interpreting and analyzing this large amount of data can be difficult.
Summary of the Invention
Means for Solving the Problems
[0006] This disclosure relates to a system, method, and composition for producing indexed linked leads using bead-bound transposomes and droplet generators.
[0007] Several embodiments provided herein relate to systems for nucleic acid indexed amplification. In some embodiments, the system includes a plurality of continuum beads, an indexed primer pool, and a detector for obtaining sequencing data. In some embodiments, each continuum bead is associated with a transposome. In some embodiments, each continuum bead contains a bead-bound nucleic acid molecule. In some embodiments, the indexed primer pool contains a plurality of primer beads and a solution primer. In some embodiments, each primer bead contains an adapter, a barcode, and a primer. In some embodiments, the continuum beads and primer beads are distributed together in a droplet. In some embodiments, the primer is a P5 primer. In some embodiments, the solution primer includes an adapter and a primer. In some embodiments, the solution primer includes a B15 adapter and a P7 primer. In some embodiments, the transposome includes a transposase and a transposon.
[0008] In some embodiments, the continuous beads and / or primer beads are hydrogel beads comprising a hydrogel polymer and a crosslinking agent. In some embodiments, the hydrogel polymer is polyethylene glycol (PEG)-thiol / PEG-acrylate, acrylamide / N,N'-bis(acryloyl)cystamine (BACy), PEG / polypropylene oxide (PPO), polyacrylic acid, poly(hydroxyethyl methacrylate) (PHEMA), poly(methyl methacrylate) (PMMA), poly(N-isopropylacrylamide) (PNIPAAm), poly(lactic acid) (PLA), poly(lactic-co-glycolic acid) (PLGA), polycaprolactone (PCL), poly(vinylsulfonic acid) (poly(vinylsulfonic acid) The crosslinking agent includes (acid), PVSA), poly(L-aspartic acid), poly(L-glutamic acid), polylysine, agar, agarose, alginate, heparin, sulfated alginic acid, dextran sulfate, hyaluronan, pectin, carrageenan, gelatin, chitosan, cellulose, or collagen. In some embodiments, the crosslinking agent includes bisacrylamide, diacrylate, diallylamine, triallylamine, divinyl sulfone, diethylene glycol diallyl ether, ethylene glycol diacrylate, polymethylene glycol diacrylate, polyethylene glycol diacrylate, trimethylopropanthrimethacrylate, ethoxylated trimethylol triacrylate, or ethoxylated pentaerythritol tetraacrylate. In some embodiments, the nucleic acid is a DNA molecule with 50,000 base pairs or more.
[0009] Several embodiments provided herein relate to flow cell devices for nucleic acid indexed amplification. In some embodiments, the device includes a solid support containing a plurality of distributed droplets. In some embodiments, the plurality of distributed droplets associate with transposomes and include continuous beads containing bead-bound nucleic acid molecules and primer beads containing adapters, barcodes, and primers. In some embodiments, the plurality of distributed droplets are distributed along the surface of the solid support.
[0010] In some embodiments, the solid support is functionalized with a surface polymer. In some embodiments, the surface polymer is poly(N-(5-azidoacetamidylpentyl)acrylamide-co-acrylamide) (PAZAM) or silane-free acrylamide (SFA). In some embodiments, the flow cell includes a patterned surface. In some embodiments, the patterned surface includes wells. In some embodiments, the wells have a diameter of approximately 10 μm to approximately 50 μm, for example, a diameter of 10 μm, 15 μm, 20 μm, 25 μm, 30 μm, 35 μm, 40 μm, 45 μm, or 50 μm, or within the range defined by any two of the aforementioned values, and the wells have a depth of approximately 0.5 μm to approximately 1 μm, for example, a depth of 0.5 μm, 0.6 μm, 0.7 μm, 0.8 μm, 0.9 μm, or 1 μm, or within the range defined by any two of the aforementioned values. In some embodiments, the wells contain a hydrophobic material. In some embodiments, the hydrophobic material contains an amorphous fluoropolymer such as CYTOP, Fluoropel®, or Teflon®. In some embodiments, the nucleic acid is a DNA molecule with 50,000 base pairs or more. In some embodiments, the transposome contains a transposase and a transposon.
[0011] Some embodiments provided herein relate to methods for nucleic acid indexing. In some embodiments, the method comprises generating a plurality of continuity beads for on-bead tagmentation, each bead being linked to a transposome and containing a bead-bound nucleic acid molecule; performing a tagmentation reaction on the nucleic acid molecule; generating a plurality of primer beads, each primer bead containing an adapter, a barcode, and a primer; distributing the continuity beads and primer beads together with solution primers into droplets; amplifying the nucleic acid molecule in the distributed droplets; and indexing the nucleic acid molecule in each droplet.
[0012] In some embodiments, the nucleic acid is a DNA molecule with 50,000 or more base pairs. In some embodiments, the method further includes nucleic acid amplification of the nucleic acid molecule before performing the tagmentation reaction. In some embodiments, the amplification reaction includes multiple displacement amplification (MDA). In some embodiments, the tagmentation reaction includes contacting the nucleic acid with a transposase mixture containing an adapter sequence and transposomes. In some embodiments, indexing is performed by polymerase chain reaction (PCR). In one embodiment, droplets are distributed into more than 900,000 different indexed PCR compartments. In some embodiments, the method further includes distributing droplets on a solid support. In some embodiments, the solid support is a flow cell device. [Brief explanation of the drawing]
[0013] [Figure 1] This is a schematic diagram of one embodiment of a microfluidic droplet generator system that can be used to generate droplets distributed onto bead-transposomes. [Figure 2]A schematic diagram shows an exemplary method for performing linked long-read indexing, including continuous preserved transposition sequencing (CPT-seq) on beads (Step 1), distribution / indexed PCR (Step 2), and indexed linked reads (Step 3). [Figure 3] A schematic diagram shows an exemplary method for performing long-read indexing, including CPT-seq on beads, droplet distribution, and indexed primer pool indexing. [Figure 4] A schematic diagram of chromosome-level phasing results using the method described herein is shown. [Figure 5] The results of comparing the number of islands with the island length using the long-read indexing method described herein are shown. [Figure 6] The results of mutant calling and phase blocking using the long-read indexing method described herein (left) compared with 10X sequencing (right) are shown. [Figure 7] The results of the long-read indexing method described herein, applied to the human leukocyte antigen (HLA) region, are shown. [Figure 8A] The results of the long-read indexing method described herein, applied to the HLA-DPA1 (Figure 8A) and HLA-A (Figure 8B) regions are shown. [Figure 8B] The results of the long-read indexing method described herein, applied to the HLA-DPA1 (Figure 8A) and HLA-A (Figure 8B) regions are shown. [Modes for carrying out the invention]
[0014] The following detailed description refers to the accompanying drawings, which form part of this specification. In the drawings, similar symbols typically identify similar components unless otherwise indicated in context. The exemplary embodiments described in the detailed description, drawings, and claims are not intended to limit the scope. Other embodiments may be utilized and other modifications made without departing from the spirit or scope of the subject matter presented herein. It will be readily apparent that the aspects of this disclosure may be arranged, substituted, combined, separated, and designed in a wide variety of different configurations, as described herein and illustrated in the drawings, all of which are expressly contemplated herein.
[0015] The embodiments provided herein relate to long-read indexing systems, devices, and methods. The system includes on-bead tagmentation. The beads may include any of the beads disclosed herein, to which transpososomes are bound and nucleic acid molecules are bound.
[0016] In some embodiments, the beads comprise a hydrogel polymer and a crosslinking agent that are mixed in the presence of transposomes and form beads bound to the transposomes. In some embodiments, the beads are prepared, then mixed with transposomes, and subsequently bound to the beads. In some embodiments, the beads are prepared in the presence of nucleic acid molecules that encapsulate or associate with the beads. In some embodiments, the beads are first prepared, mixed with transposomes, and then mixed with nucleic acid molecules. In some embodiments, the beads allow for simultaneous assays on the same sample while maintaining continuity. Specifically, the methods, systems, and compositions provided herein allow for the confinement and access of biomolecules bound to the beads. Thus, in some embodiments, the beads described herein are referred to herein as continuous particles. Therefore, the term “continuous particles” as used herein refers to beads used in continuous preserved transposition sequencing (CPT-seq).
[0017] The continuous particles described herein are used for next-generation partitioning approaches and can enable multi-analyte assays to be performed on nucleic acid molecules. The continuous particles and methods of use described herein enable millions of nucleic acid molecules to be individually analyzed efficiently, thereby reducing the cost of sample preparation and maintaining sample continuity. The compositions and methods described herein maintain continuity without using external compartmentalization strategies (microfluidic techniques) such as emulsions, immobilization, or other microcompartments.
[0018] In some embodiments, the continuous particles described herein can be used in assays for analyzing target nucleic acid molecules. Assays that can be performed on nucleic acid molecules include, for example, DNA analysis, RNA analysis, nucleic acid sequencing, tagmentation, nucleic acid amplification, DNA library preparation, assays for transposase-accessible chromatin use sequencing (ATAC-seq), continuous preservation translocation sequencing (CPT-seq), or any combination thereof performed sequentially.
[0019] The use of continuous particles for performing one or more assays on nucleic acid molecules can be used simultaneously on multiple continuous particles to perform simultaneous assays on several nucleic acid molecules, for example, 10,000 to 1 million nucleic acid molecules, such as 10,000, 50,000, 100,000, 500,000, or 1 million nucleic acid molecules.
[0020] In some embodiments, the methods described herein include methods for producing indexed linked reads for various applications, including phase and assembly. In some embodiments, the method includes combining on-bead tagmentation with droplet indexing. In some embodiments, droplet indexing includes any physical compartment indexing, including emulsion or plate. In some embodiments, the beads provided herein include transposomes that enable on-bead tagmentation to nucleic acid molecules. In some embodiments, each nucleic acid molecule encapsulates a bead, generating a bead-bound fragment of the nucleic acid molecule. In some embodiments, the method is combined with indexing. In some embodiments, the method includes distributing beads into droplets. In some embodiments, the method includes performing indexed PCR in each droplet. In some embodiments, each fragment derived from individual nucleic acid molecules receives the same barcode, thereby generating an indexed linked read.
[0021] The embodiments of the methods, systems, and devices described herein have a number of advantages over conventional methods. For example, the methods, systems, and devices described herein provide a controlled insert size, move DNA to physical partitioning without fragmenting the DNA, where the DNA is uniquely indexed by PCR. In addition, CPT-seq on beads can be performed on over 900,000 different indexed PCR partitions. Furthermore, CPT-seq on beads results in improved control of library inserts (transposon density), more efficient DNA transfer to droplets, robust DNA preparation, assay steps that can be performed prior to droplet formation (e.g., including Tn5 removal), and less DNA fragmentation. Embodiments of the method enable elution of the template from the beads, which results in release of biotinylated products in droplets after heating. In addition, the enzymes associated with the method provide high amplification by strand displacement polymerase and increased amounts of enzyme. Finally, the embodiments of the methods provided herein result in improved DNA quality and enable compatibility with lysates.
[0022] The methods, systems, and devices provided herein combine transposition on beads with transfer to physical partitioning of tagged nucleic acids. In some embodiments, generation of long-indexed reads includes generation of bead-bound transposomes and droplet generators. In some embodiments, droplet generators include microfluidic devices or emulsions. In some embodiments, transposition on beads includes nucleic acid tagging and frequency, which can be controlled by the density of transposomes on the beads. In some embodiments, beads are used to transfer tagged nucleic acids to physical partitioning.
[0023] As used herein, the term “reagent” refers to an active substance or mixture of two or more active substances useful for reacting with, interacting with, diluting, or adding to a sample, and may include active substances used in assays described herein, including active substances for dissolution, nucleic acid analysis, nucleic acid amplification reactions, protein analysis, tagmentation reactions, ATAC-seq, CPT-seq, or SCI-seq reactions, or other assays. Thus, reagents may include, for example, buffers, chemicals, enzymes, polymerases, primers having a size of less than 50 base pairs, template nucleic acids, nucleotides, labels, dyes, or nucleases. In some embodiments, reagents may include lysozyme, proteinase K, random hexamer, polymerase (e.g., Φ29 DNA polymerase, Taq polymerase, Bsu polymerase), transposase (e.g., Tn5), primer (e.g., P5 and P7 adapter sequences), ligase, catalytic enzyme, deoxynucleotide triphosphate, buffers, or divalent cations.
[0024] Continuous particles In some embodiments, the beads have a polymer shell prepared from a hydrogel composition. As used herein, the term "hydrogel" refers to a material formed when organic polymers (natural or synthetic) are crosslinked via covalent, ionic, or hydrogen bonds to produce a three-dimensional open lattice structure that traps water molecules and forms a gel. In some embodiments, the hydrogel may be a biocompatible hydrogel. As used herein, the term "biocompatible hydrogel" refers to a polymer that forms a gel that is not toxic to biological materials.In some embodiments, the hydrogel material may be alginate, acrylamide, or polyethylene glycol (PEG), PEG-acrylate, PEG-amine, PEG-carboxylate, PEG-dithiol, PEG-epoxide, PEG-isocyanate, PEG-maleimide, polyacrylic acid (PAA), poly(methyl methacrylate) (PMMA), polystyrene (PS), polystyrene sulfonate (PSS), polyvinylpyrrolidone (PVPON), N,N'-bis(acryloyl)cystamine, polypropylene oxide (PPO), poly(hydroxyethyl methacrylate) (PHEMA), poly(N-isopropylacrylamide) (PNIPAAm), or poly(lactic acid) (poly(lactic)). This includes (1) acid, (1) PLA, poly(lactic acid-co-glycolic acid) (PLGA), polycaprolactone (PCL), poly(vinyl sulfonic acid) (PVSA), poly(L-aspartic acid), poly(L-glutamic acid), polylysine, agar, agarose, heparin, sulfated alginic acid, dextran sulfate, hyaluronan, pectin, carrageenan, gelatin, chitosan, cellulose, collagen, bisacrylamide, diacrylate, diallylamine, triallylamine, divinyl sulfone, diethylene glycol diallyl ether, ethylene glycol diacrylate, polymethylene glycol diacrylate, polyethylene glycol diacrylate, trimethylopropanthrimethacrylate, ethoxylated trimethylol triacrylate, or ethoxylated pentaerythritol tetraacrylate, or combinations or mixtures thereof. In some embodiments, the hydrogel is an alginate, acrylamide, or PEG-based material. In some embodiments, the hydrogel is a PEG-based material having an acrylate-dithiol, epoxide-amine reaction chemistry.In some embodiments, the hydrogel forms a polymer shell comprising PEG-maleimide / dithiol oil, PEG-epoxide / amine oil, PEG-epoxide / PEG-amine, or PEG-dithiol / PEG-acrylate. In some embodiments, the hydrogel material is selected to avoid the generation of free radicals that may damage intracellular biomolecules. In some embodiments, the hydrogel polymer comprises 60–90% fluid, such as water, and 10–30% polymer. In certain embodiments, the water content of the hydrogel is about 70–80%. As used herein, the terms “about” or “approximately” refer to the variation that may occur in a numerical value when modifying a numerical value. For example, variation may occur due to differences in the manufacture of a particular substrate or component. In one embodiment, the term “about” means within 1%, 5%, or up to 10% of the enumerated numerical value.
[0025] As used herein, the polymer shell is the polymer surface of the beads. Due to the properties of the beads described herein, continuous particles can retain genetic material after multiple assays and can be released by physical force, by cleavage chemicals, or by causing osmotic imbalance depending on the thickness of the polymer shell.
[0026] Hydrogels can be prepared by crosslinking hydrophilic biopolymers or synthetic polymers. Therefore, in some embodiments, the hydrogel may contain a crosslinking agent. As used herein, the term “crosslinking agent” refers to a molecule that can form a three-dimensional network when reacted with a suitable base monomer. Examples of hydrogel polymers that may contain one or more crosslinking agents include hyaluronan, chitosan, agar, heparin, sulfate, cellulose, alginate (including sulfated alginic acid), collagen, dextran (including sulfated dextran), pectin, carrageenan, polylysine, gelatin (including type A gelatin), agarose, (meth)acrylate-oligolactide-PEO-oligolactide-(meth)acrylate, PEO-PPO-PEO copolymer (Pluronic), poly(phosphazene), poly(methacrylate), poly(N-vinylpyrrolidone), PL(G)A-PEO-PL(G)A copolymer, poly(ethyleneimine), polyethylene glycol (PEG)-thiol, PEG-acrylate, acrylamide, N,N'-bis(acryloyl)cystamine, PEG, polypropylene oxide (PPO), polyacrylic acid, poly(hydro) Examples include, but are not limited to, roxyethyl methacrylate (PHEMA), poly(methyl methacrylate) (PMMA), poly(N-isopropylacrylamide) (PNIPAAm), poly(lactic acid) (PLA), poly(lactic acid-co-glycolic acid) (PLGA), polycaprolactone (PCL), poly(vinyl sulfonic acid) (PVSA), poly(L-aspartic acid), poly(L-glutamic acid), bisacrylamide, diacrylate, diallylamine, triallylamine, divinyl sulfone, diethylene glycol diallyl ether, ethylene glycol diacrylate, polymethylene glycol diacrylate, polyethylene glycol diacrylate, trimethylopropanthrimethacrylate, ethoxylated trimethylol triacrylate, or ethoxylated pentaerythritol tetraacrylate, or combinations thereof.Therefore, for example, the combination may include a polymer and a crosslinking agent, such as polyethylene glycol (PEG)-thiol / PEG-acrylate, acrylamide / N,N'-bis(acryloyl)cystamine (BACy), or PEG / polypropylene oxide (PPO). In some embodiments, the polymer shell includes 4-arm polyethylene glycol (PEG). In some embodiments, the 4-arm polyethylene glycol (PEG) is selected from the group consisting of PEG-acrylate, PEG-amine, PEG-carboxylate, PEG-dithiol, PEG-epoxide, PEG-isocyanate, and PEG-maleimide.
[0027] In some embodiments, the crosslinking agent is either an immediate-type crosslinking agent or a delayed-type crosslinking agent. An immediate-type crosslinking agent is a crosslinking agent that immediately crosslinks the hydrogel polymer and is also referred to herein as click chemistry. Immediate-type crosslinking agents may include dithiol oil + PEG-maleimide or PEG epoxide + amine oil. A delayed-type crosslinking agent is a crosslinking agent that slowly crosslinks the hydrogel polymer and may include PEG-epoxide + PEG-amine or PEG-dithiol + PEG-acrylate. A delayed-type crosslinking agent may take several hours or more to crosslink, for example, more than 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, or 12 hours to crosslink. In some embodiments provided herein, continuous particles are formulated with an immediate-type crosslinking agent, thereby better preserving the state of nucleic acid molecules compared to a delayed-type crosslinking agent.
[0028] In some embodiments, the crosslinking agent forms disulfide bonds in the hydrogel polymer, thereby linking the hydrogel polymers. In some embodiments, the hydrogel polymer forms a hydrogel matrix having pores (e.g., a porous hydrogel matrix). These pores can hold sufficiently large particles, such as nucleic acids extracted within the polymer shell, but allow other materials, such as reagents, to pass through the pores and thereby enter and exit the droplets. In some embodiments, the pore size of the polymer shell is fine-tuned by varying the ratio of the polymer concentration to the crosslinking agent concentration. In some embodiments, the polymer-to-crosslinker ratio is 30:1, 25:1, 20:1, 19:1, 18:1, 17:1, 16:1, 15:1, 14:1, 13:1, 12:1, 11:1, 10:1, 9:1, 8:1, 7:1, 6:1, 5:1, 4:1, 3:1, 2:1, 1:1, 1:2, 1:3, 1:4, 1:5, 1:6, 1:7, 1:8, 1:9, 1:10, 1:15, 1:20, or 1:30, or a ratio within the range defined by any two of the aforementioned ratios. In some embodiments, further functionalities such as DNA primers or charged chemical groups can be incorporated into the polymer matrix to meet the requirements of different applications.
[0029] As used herein, the term "porosity" means the dimensionless volume fraction of a hydrogel composed of open spaces, such as pores or other openings. Thus, porosity measures the voids in a material and is the fraction of void volume relative to the total volume, such as a ratio of 0 to 100% (or 0 to 1). The porosity of a hydrogel may range from 0.5 to 0.99, about 0.75 to about 0.99, or about 0.8 to about 0.95.
[0030] The polymer shell can have any pore size that allows sufficient diffusion of reagents while simultaneously retaining nucleic acids. As used herein, the term “pore size” refers to the diameter or effective diameter of the cross-section of a pore. The term “pore size” may also refer to the average diameter or average effective diameter of the cross-section of a pore based on measurements of multiple pores. The effective diameter of a non-circular cross-section is equal to the diameter of a circular cross-section having the same cross-sectional area as the non-circular cross-section. In some embodiments, the hydrogel may swell when the hydrogel is hydrated. The size of the pore size may then change depending on the water content of the hydrogel. In some embodiments, the pores of the hydrogel may have pores that are large enough to allow reagents to pass through while maintaining the integrity of the hydrogel. In some embodiments, the interior of the polymer shell is an aqueous environment. In some embodiments, nucleic acid molecules associated with the polymer shell do not interact with the polymer shell and / or come into contact with the polymer shell.
[0031] In some embodiments, the continuous particles have diameters ranging from approximately 20 μm to approximately 200 μm, for example, 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, or 200 μm, or within a range defined by any two of the aforementioned values. The size of the continuous particles may vary due to environmental factors. In some embodiments, the continuous particles expand when separated from the continuous oil phase and immersed in the aqueous phase. In some embodiments, the expansion of the continuous particles increases the efficiency of performing assays on genetic material. In some embodiments, the expansion of the continuous particles creates a larger environment for indexed inserts to be amplified during PCR, which may otherwise be limited in current cell-based assays.
[0032] In some embodiments, the pore size allows the extracted nucleic acids to diffuse through the polymer shell. In some embodiments, the pore size of the continuous particles can be controlled by altering the crosslinking chemistry. The final crosslinked pore size may be further altered by changing the environment of the continuous particles, for example, by changing the salt concentration, pH, or temperature, thereby releasing immobilized molecules from the continuous particles.
[0033] In some embodiments, the crosslinking agent is a reversible crosslinking agent. In some embodiments, the reversible crosslinking agent can reversibly crosslink a hydrogel polymer and cannot be crosslinked in the presence of a cleaving agent. In some embodiments, the crosslinking agent can be cleaved by the presence of a reducing agent, by high temperature, or by an electric field. In some embodiments, the reversible crosslinking agent may be a reversible crosslinking agent for N,N'-bis(acryloyl)cystamine, polyacrylamide gel, and the disulfide bonds can be broken in the presence of a suitable reducing agent. The porosity of the beads can be increased by temperature or chemical means, thereby releasing contact between the crosslinking agent and the reducing agent, cleaving the disulfide bonds of the crosslinking agent, and decomposing the beads. The beads decompose, releasing contents such as nucleic acids held within them. In some embodiments, the crosslinking agent is cleaved by raising the temperature to over 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, or 100°C. In some embodiments, the crosslinking agent is cleaved by bringing the beads into contact with a reducing agent. In some embodiments, the reducing agent includes phosphine compounds, water-soluble phosphines, nitrogen-containing phosphines, and their salts and derivatives, dithioerythritol (DTE), dithiothreitol (DTT) (cis and trans isomers of 2,3-dihydroxy-1,4-dithiolbutane, respectively), 2-mercaptoethanol or β-mercaptoethanol (BME), 2-mercaptoethanol or aminoethanethiol, glutathione, thioglycolate or thioglycolic acid, 2,3-mercaptopropanol, tris(2-carboxyethyl)phosphine (TCEP), tris(hydroxymethyl)phosphine (THP), or p-[tris(hydroxymethyl)phosphine]propionic acid (THPP).
[0034] In some embodiments, the crosslinking agent is decomposed by increasing the temperature to increase diffusion or by contact with a reducing agent.
[0035] In some embodiments, pores are established in continuous particles by crosslinking with a crosslinking agent. In some embodiments, the pore size in the polymer shell is adjustable and formulated to associate with transpososomes. Onbees transpososomes are formulated to associate with nucleic acids of more than about 300 base pairs. In some embodiments, Onbees transpososomes and nucleic acids can be subjected to various reagents to carry out various reactions. In some embodiments, reagents include reagents for processing genetic material, e.g., reagents for isolating nucleic acids from cells, reagents for amplifying or sequencing nucleic acids, or reagents for preparing nucleic acid libraries. In some embodiments, reagents include, for example, lysozyme, proteinase K, random hexamer, polymerase (e.g., Φ29 DNA polymerase, Taq polymerase, Bsu polymerase), transposase (e.g., Tn5), primer (e.g., P5 and P7 adapter sequences), ligase, catalytic enzyme, deoxynucleotide triphosphate, buffer, or divalent cation.
[0036] Method for creating continuous particles Several embodiments provided herein relate to methods for producing continuous particles of associated transposomes. In some embodiments, the continuous particles are prepared by static means such as microwell / microarray methods or microdissection, without requiring microfluidic devices. Thus, in some embodiments, the continuous particles described herein are prepared by device-free methods. The initiation of polymerization may occur by a chemical reaction between an active group on a monomer unit and a specific portion of a membrane protein, glycan, or other small molecule. Following the initial step of monomer polymerization, one or more monomer unit depositions may follow, facilitated by either electrostatic or hydrophobic forces. Some of the monomer layers may contain functional groups such as biotin or other ligands that can be later used for specific association with transposomes.
[0037] In some embodiments, continuous particles are prepared by dynamic means such as vortex-assisted emulsion, microfluidic droplet generation, or valve-based microfluidic technology. As used herein, vortex-assisted emulsion refers to vortexing a hydrogel polymer with transposomes in a container such as a tube, vial, or reaction vessel. The components can be mixed, for example, by manual or mechanical vortexing or shaking. In some embodiments, manual mixing yields beads that associate with transposomes, having sizes of 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, or 200 μm in diameter, or within a range defined by any two of the aforementioned values. In some embodiments, the bead sizes are heterogeneous, and therefore the bead sizes include beads of various diameters.
[0038] In some embodiments, continuous particles are prepared by microfluidic flow techniques. Microfluidic flow involves the use of a microfluidic device for assisted gel emulsion generation, as shown in Figure 1. In some embodiments, the microfluidic device includes microchannels configured to generate continuous particles of a desired size and to associate with transposomes at a desired density. In some embodiments, the microfluidic device has a height of 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, or 200 μm, or a height within a range defined by any two of the aforementioned values. In some embodiments, the microfluidic device includes one or more channels. In some embodiments, the microfluidic device includes channels for introducing reagents that associate with continuous particles such as transposomes or nucleic acids introduced into the polymer, channels for introducing crosslinking agents, and channels for immiscible fluids. In some embodiments, the widths of one or more channels are identical. In some embodiments, the widths of one or more channels are different. In some embodiments, the width of one or more channels is within the range defined by 20, 30, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, or 150 μm, or any two of the aforementioned values. The channel width and height are not necessarily limited to the values described herein, and those skilled in the art will recognize that the size of the continuity particles depends in part on the size of the channels in the microfluidic device. Thus, the size of the continuity particles can be adjusted in part by modifying the channel size. In addition to the size of the microfluidic device and the channel width, the channel flow rate may also affect the size of the continuity particles and the density of the transposomes associated with each continuity particle.
[0039] In some embodiments, the flow rate of transposomes in the polymer through the microfluidic channel is within the range defined by 1, 2, 5, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, or 150 μL / min, or any two of the aforementioned values. In some embodiments, the flow rate of the crosslinking agent in the microfluidic channel is within the range defined by 1, 2, 5, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, or 150 μL / min, or any two of the aforementioned values. In some embodiments, the flow velocity of the immiscible fluid in the microfluidic channel is within the range defined by 20, 30, 50, 80, 100, 150, 160, 170, 180, 190, 200, 225, 250, 275, 300, 325, 350, 375, or 400 μL / min, or any two of the aforementioned values. In some embodiments, the transposomes mixed with the polymer and the crosslinking agent come into contact with each other in a microfluidic droplet generator upstream of the immiscible fluid. Upon contact with the crosslinking agent, continuous particles begin to form and associate with the transposomes. The resulting continuous particles continue to flow through the microfluidic droplet generator into the immiscible fluid, such as a spacer oil and / or crosslinking oil, at a flow velocity less than that of the immiscible fluid, thereby forming droplets. In some embodiments, the immiscible fluid is introduced in two stages, for example, as a spacer oil and a crosslinking agent oil, as shown in Figure 1. In some embodiments, the spacer oil is mineral oil, hydrocarbon oil, silicone oil, fluorocarbon oil, or polydimethylsiloxane oil, or a mixture thereof. The spacer oil used herein is used to avoid crosslinking of polymers between the channel water-oil phase.
[0040] In some embodiments, continuous particles are formed immediately by crosslinking with an immediate-acting crosslinking agent. For example, transposomes can associate with beads having polymers such as 4-arm PEG maleimide or epoxide using a microfluidic droplet generator and be immediately crosslinked with an oil-soluble crosslinking agent such as mineral oil or fluorocarbon oil such as HFE-7500 to form a crosslinked oil. In some embodiments, the crosslinked oil includes toluene, acetone, and tetrahydrofuran having amine functional groups, such as toluene, 3,4-dithiol, 2,4-diaminotoluene, and hexanedithiol, which readily diffuse into the resulting droplets, thereby immediately crosslinking the continuous particles.
[0041] In some embodiments, continuous particles are formulated with a uniform size distribution. In some embodiments, the size of the continuous particles is fine-tuned by adjusting the size of the microfluidic device, the size of one or more channels, or the flow velocity through the microfluidic channels. In some embodiments, the resulting continuous particles have diameters in the range of 20 to 200 μm, for example, diameters of 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, or 200 μm, or within the range defined by any two of the aforementioned values.
[0042] In some embodiments, the size and uniformity of continuous particles can be further controlled by contacting the hydrogel polymer with a fluid modifier, such as an alcohol containing isopropyl alcohol, before particle formation. In the absence of isopropyl alcohol, continuous particles are formed with a larger diameter than those formed in the presence of isopropyl alcohol. Isopropyl alcohol affects the fluid properties of the hydrogel polymer, allowing for the control of continuous particle size.
[0043] As will be recognized by those skilled in the art, the microfluidic device shown in Figure 1 is an example of a three-channel microfluidic device, but this microfluidic device may be modified, varied, or altered to produce continuous particles of a specific size, or to produce continuous particles formed from various hydrogel materials or crosslinking agents.
[0044] In some embodiments, continuous particles are prepared by a vortex-assisted emulsion or by a microfluidic inertial flow-assisted emulsion. In some embodiments, the density of transposomes associated with continuous particles can be controlled by diluting or concentrating the transposome-containing solution in the introduced sample. The transposome-containing sample is mixed with a hydrogel polymer, and the transposome-containing hydrogel polymer is subjected to the vortex-assisted emulsion or microfluidic flow-assisted emulsion described herein.
[0045] In some embodiments, the continuous particles are functionalized with nucleotides. In some embodiments, the nucleotides are oligonucleotides or poly-T nucleotides. In some embodiments, the nucleotides are bound to the continuous particles, and the functionalized continuous particles can be used for targeted capture of the desired nucleotides.
[0046] In some embodiments, the continuous particles associated with transposomes are cured to maintain the execution of multiple simultaneous assays on a single continuous particle, including multiple buffer washes, multiple reagent exchanges, and multiple analyses based on the assay being performed. Formulated continuous particles, prepared by any of the methods described herein, including surface-initiated polymerization techniques, vortexing, or microfluidic techniques, can be loaded or seeded onto patterned flow cells, microarrays, plates with wells, etched surfaces, microfluidic channels, beads, columns, or other surfaces for performing multiple simultaneous assays.
[0047] Method for creating linked long-read indexes Several embodiments provided herein relate to methods for performing linked long-read indexing using on-bead tagged continuous particles. In some embodiments, the method comprises obtaining the continuous particles described herein, associating the particles with transpososomes and nucleic acid molecules, performing on-bead tagmentation, distributing the on-bead tagged beads, and performing indexed PCR. In some embodiments, the method includes the steps outlined in Figure 2, which shows a schematic diagram of an exemplary method for performing linked long-read indexing, comprising continuous conserved transposition sequencing (CPT-seq) on beads (Step 1), distribution / indexed PCR (Step 2), and indexed linked reads (Step 3). The beads from Step 1 are associated with transpososomes containing transposons and transposases. Nucleic acid molecules associate with transposases and undergo tagmentation, e.g., CPT-seq. After tagmentation, the on-bead tagmentation product is distributed using a droplet generator or other method. The distributed sample containing the transposition DNA on beads can be subjected to solution primers and indexed primer beads for indexing the transposition DNA. In some embodiments, the indexed DNA is subjected to indexed PCR to obtain indexed ligated reads.
[0048] In addition to continuous particles (also referred to herein as continuous beads), primer beads were also prepared. Primer beads may include hydrogel beads of the materials, compositions, and formulations described herein with respect to continuous particles, but instead of associating with transposomes, they include primers (such as P5 primers), barcodes, and adapters (such as Nextera adapters). A single primer bead may be dispensed together with a single continuous particle in a dispensed droplet, and together with a primer solution mix which may include adapters (such as B15 adapters) and primers (such as P7 primers). The dispensed droplets may be used for barcode indexing together with a bead pool of more than 900,000 indexed barcodes.
[0049] In some embodiments, the method includes steps outlined in Figure 3, which shows a schematic diagram of an exemplary method for performing long-read indexing, including CPT-seq on beads, droplet distribution, and indexed primer pool indexing, using non-barcoded dislocations.
[0050] In some embodiments, continuous particles are prepared as described herein, and droplets are distributed onto a surface such as a flow cell device, a plate well, a slide, or a patterned surface. In some embodiments, the surface is a flow cell device and includes an insert having microwells or micropillars in an array for the distribution of continuous particles for spatial indexing within the flow cell device. In some embodiments, droplets are distributed onto a well plate having a single continuous particle in each well. The well plate may be, for example, a 12-well plate, a 24-well plate, a 48-well plate, a 96-well plate, a 384-well plate, a 1536-well plate, a 3456-well plate, or a 9600-well plate, or any number of wells in the plate, having a single continuous particle, and the continuous particle is distributed for droplet indexing for nucleic acid indexing. In some embodiments, continuous particles are sequentially subjected to multiple simultaneous assays, including, for example, buffer washing, lysis, DNA analysis, RNA analysis, protein analysis, tagmentation, nucleic acid amplification, nucleic acid sequencing, DNA library preparation, assays for sequencing using transposase-accessible chromatin (ATAC-seq), continuous preservation transposition sequencing (CPT-seq), or any combination thereof performed sequentially.
[0051] In some embodiments, the continuum particles associated with transpososomes are processed to associate the nucleic acid of interest with them. In some embodiments, the nucleic acid of interest is isolated from cells. For example, cells may be brought into contact with a lysis buffer. As used herein, “lysis” means perturbation or modification of the cell wall or viral particles that facilitates access to or release of cellular RNA or DNA. Neither complete destruction nor breakage of the cell wall is a necessary requirement for lysis. The term “lysis buffer” means a buffer containing at least one lysis agent. Typical enzymatic lysis agents include, but are not limited to, lysozyme, glucolase, thymolase, lithicase, proteinase K, proteinase E, and viral endolysin and exolysin. Thus, for example, cell lysis may be performed by introducing a lysis agent such as lysozyme and proteinase K. Subsequently, the cell-derived gDNA associates with the continuum particles. In some embodiments, after the lysis treatment, the isolated nucleic acid may be retained on the continuum particles and used for further processing.
[0052] DNA analysis refers to any technique used to amplify, sequence, or otherwise analyze DNA. DNA amplification can be achieved using PCR techniques or pyrosequencing. DNA analysis may also include non-targeted, non-PCR-based DNA sequencing techniques (e.g., metagenomics). Non-limiting examples of DNA analysis include sequencing the hypervariable region of 16S rDNA (ribosomal DNA) and using sequencing for DNA-mediated species identification.
[0053] RNA analysis refers to any technique used to amplify, sequence, or otherwise analyze RNA. RNA can be amplified and sequenced using the same techniques used to analyze DNA. RNA, being less stable than DNA, is the translation of DNA in response to stimuli. Therefore, RNA analysis can provide a more accurate picture of the metabolically active members of a community and may be used to provide information about the community function of organisms in a sample. Nucleic acid sequencing refers to the use of sequencing to determine the order of nucleotides in the sequence of nucleic acid molecules such as DNA or RNA.
[0054] As used herein, the term “sequencing” refers to a method for obtaining the identification of at least 10 consecutive nucleotides of a polynucleotide (e.g., the identification of at least 20, at least 50, at least 100, or at least 200 consecutive nucleotides).
[0055] The terms “next-generation sequencing,” “high-throughput sequencing,” or “NGS” generally refer to high-throughput sequencing technologies, including but not limited to large-scale parallel signature sequencing, high-throughput sequencing, ligation-based sequencing (e.g., SOLiD sequencing), proton-ion semiconductor sequencing, DNA nanoball sequencing, single-molecule sequencing, and nanopore sequencing, and may also refer to synthetic parallel sequencing or ligation-based sequencing platforms currently employed by companies such as Illumina, Life Technologies, or Roche. Next-generation sequencing methods may also include nanopore sequencing methods or electron detection-based methods, such as Ion Torrent technology commercialized by Life Technologies or single-molecule fluorescence-based methods commercialized by Pacific Biosciences.
[0056] Protein analysis refers to the study of proteins and may include proteomic analysis, determination of post-translational modifications of a target protein, determination of protein expression levels, or determination of protein interactions with other proteins or other molecules, including nucleic acids.
[0057] As used herein, the term “tagmentation” refers to the modification of DNA by a transposomal complex containing a transposase enzyme complexed with an adapter containing a transposon terminal sequence. Tagmentation results in simultaneous fragmentation of DNA and ligation of the adapter to the 5' ends of both strands of the double fragment. After a purification step to remove the transposase enzyme, additional sequences can be added to the ends of the adapted fragment, for example, by PCR, ligation, or any other suitable method known to those skilled in the art.
[0058] The method of the present invention can use any transposase that can accept a transposase terminal sequence and fragment a target nucleic acid, and that binds the transposase end but not the non-transposase end. A “transposome” consists of at least a transposase enzyme and a transposase recognition site. In some such systems called “transpososomes,” the transposase can form a functional complex with a transposon recognition site that can catalyze the transposition reaction. A transposase or integrase can bind to a transposase recognition site and insert the transposase recognition site into the target nucleic acid in a process sometimes called “tagging.” In some such insertion events, a single strand of the transposase recognition site can be transferred to the target nucleic acid.
[0059] In standard sample preparation methods, each template includes an adapter at one end of the insert, and often requires several steps to perform both DNA or RNA modification and purification of the desired product of the modification reaction. These steps are performed in solution before the fitted fragments are added to the flow cell, where they are bonded to the surface by a primer extension reaction, which copies the hybridized fragment to the ends of a primer covalently bonded to the surface. These "seeding" templates are then copied through several amplification cycles. It generates monoclonal clusters in the template.
[0060] The number of steps required to convert DNA to an adapter-modified template in a solution ready for cluster formation and sequencing can be minimized by using transposase-mediated fragmentation and tagging.
[0061] In some embodiments, transposon-based techniques can be used for DNA fragmentation, for example, as exemplified in the workflow of the Nextera® DNA Sample Preparation Kit (Illumina, Inc.), where genomic DNA is fragmented by engineered transpososomes that simultaneously fragment and tag ("tagmentation") the input DNA, thereby generating a collection of fragmented nucleic acid molecules containing adapter sequences specific to the ends of the fragments.
[0062] Some embodiments may include the use of hyperactive Tn5 transposase and Tn5-type transposase recognition sites (Goryshin and Reznikoff, J. Biol. Chem., 273:7367 (1998)), or MuA transposase and Mu transposase recognition sites containing R1 and R2 terminal sequences (Mizuuchi, K., Cell, 35:785, 1983; Savilahti, H, et al., EMBO J., 14:4893, 1995). Exemplary transposase recognition sites that form complexes with hyperactive Tn5 transposase (e.g., EZ Tn5® Transposase, Epicentre Biotechnologies, Madison, Wis.).
[0063] Further examples of rearrangement systems that can be used with certain embodiments provided herein include Staphylococcus aureus Tn552 (Colegio et al., J. Bacteriol., 183:2384-8, 2001; Kirby C et al., Mol. Microbiol., 43:173-86, 2002), Ty1 (Devine & Boeke, Nucleic Acids Res., 22:3765-72, 1994, and International Publication No. 95 / 23875), transposon Tn7 (Craig, NL, Science. 271:1512, 1996; Craig, NL, Review in: Curr Top Microbiol Immunol., 204:27-48, 1996), Tn / O and IS10 (Kleckner N, et al., Curr Top Microbiol Immunol., 204:49-82, 1996), Mariner transposase (Lampe DJ, et al., EMBO J., 15:5470-9, 1996), Tc1 (Plasterk RH, Curr. Topics Microbiol. Immunol., 204:125-43, 1996), P element (Gloor, GB, Methods) Examples include Mol. Biol., 260:97-114, 2004, Tn3 (Ichikawa & Ohtsubo, J Biol. Chem. 265:18829-32, 1990), bacterial insertion sequences (Ohtsubo & Sekine, Curr. Top. Microbiol. Immunol. 204:1-26, 1996), retroviruses (Brown, et al, Proc Natl Acad Sci USA, 86:2525-9, 1989), and yeast retrotransposons (Boeke & Corces, Annu Rev Microbiol. 43:403-34, 1989). Further examples include genetically modified versions of IS5, Tn10, Tn903, IS911, and the transposase family enzymes (Zhang et al., (2009) PLoS Genet. 5:e1000689. Epub 2009 Oct. 16, Wilson C. et al (2007) J. Microbiol. Methods 71:332-5).
[0064] Transposase-Accessible Chromatin Sequencing Assay (ATAC-seq) refers to a rapid and highly sensitive method for integrated epigenomic analysis. ATAC-seq captures open chromatin regions, revealing interactions between open chromatin, DNA-binding proteins, and individual nucleosome locations within the genome, as well as higher-order compaction in regulatory regions with nucleotide resolution. Classes of DNA-binding factors have been discovered that tightly avoid, tolerate, or tend to overlap with nucleosomes. Using ATAC-seq, the feasibility of measuring the sequential daily epigenome of resting human T cells, evaluating it from probands via standard blood sampling, and reading an individual's epigenome on a clinical timescale for monitoring health and disease has been demonstrated. More specifically, ATAC-seq can be performed by processing chromatin derived from a given cell with an insertion enzyme complex to generate tagged fragments of genomic DNA. In this process, chromatin is tagged using insertion enzymes such as Tn5 or MuA, which cleave genomic DNA in the open regions of chromatin and attach adapters to both ends of the fragments (e.g., fragmentation and tagging in the same reaction).
[0065] In some cases, conditions can be adjusted to obtain a desired level of insertion into chromatin (e.g., insertions occurring on average every 50–200 base pairs in open regions). The chromatin used in this method can be prepared by any suitable method. In some embodiments, nuclei can be isolated and lysed, and the chromatin can be further purified, for example, from the nuclear envelope. In other embodiments, chromatin can be isolated by contacting the isolated nuclei with a reaction buffer. In these embodiments, the isolated nuclei can be lysed when in contact with the reaction buffer (containing the insertion enzyme complex and other necessary reagents), thereby allowing the insertion enzyme complex to access the chromatin. In these embodiments, the method may include isolating nuclei from a population of cells and combining the isolated nuclei with a transposase and adapter, which results in both the lysis of the nuclei and the release of the chromatin, and the generation of a fragment of genomic DNA tagged with the adapter. The chromatin does not require crosslinking, as in other methods (e.g., ChIP-SEQ).
[0066] After chromatin is fragmented and tagged to generate tagged fragments of genomic DNA, at least some of the tagged fragments are sequenced using an adapter to generate multiple sequence reads. The fragments can be sequenced using any suitable method. For example, the fragments can be sequenced using Illumina's reversible terminator method, Roche's pyrosequencing method (454), Life Technologies' ligation sequencing (SOLiD platform), or Life Technologies' Ion Torrent platform. Examples of such methods are described in the following references: Margulies et al. (Nature 2005 437:376-80), Ronaghi et al. (Analytical Biochemistry 1996 242:84-9), Shendure et al. (Science 2005 309:1728-32), Imelfort et al. (Brief Bioinform. 2009 10:609-18), Fox et al. (Methods Mol Biol. 2009;553:79-108), Appleby et al. (Methods Mol Biol. 2009;513:19-39), and Morozova et al. (Genomics. 2008 92:255-64), which are incorporated herein by reference to general descriptions of the methods and specific steps of the methods, including all starting products, methods for library preparation, reagents, and final products at each step. As is evident, forward and reverse sequencing primer sites compatible with a selected next-generation sequencing platform can be added to the ends of the fragment during the amplification step. In certain embodiments, the fragment may be amplified using PCR primers that hybridize to the tag added to the fragment, and the primers used for PCR have a 5' tail compatible with a specific sequencing platform. A method for performing ATAC-seq is described in PCT Patent Application No. PCT / US2014 / 038825, which is incorporated herein by reference in its entirety.
[0067] As used herein, the term “chromatin” refers to a complex of molecules containing proteins and polynucleotides (e.g., DNA, RNA), such as those found in the nucleus of eukaryotic cells. Chromatin is composed in part of histone proteins that form nucleosomes, genomic DNA, and other DNA-binding proteins (e.g., transcription factors) that commonly bind to genomic DNA.
[0068] Continuity-preserving transposition sequencing (CPT-seq) refers to a sequencing method that preserves continuity information by using a transposase to maintain the association of template nucleic acid fragments adjacent to a target nucleic acid. For example, CPT can be performed on nucleic acids such as DNA or RNA. CPT nucleic acids can be captured by hybridization of complementary oligonucleotides having unique indices or barcodes and immobilized on a solid support. In some embodiments, the oligonucleotides immobilized on the solid support may further include primer binding sites and unique molecular indices in addition to barcodes. Advantageously, such use of transpososomes to maintain the physical proximity of fragmented nucleic acids increases the likelihood that fragmented nucleic acids from the same original molecule, e.g., chromosomes, will receive the same unique barcode and index information from the oligonucleotides immobilized on the solid support. This results in a continually linked sequencing library having unique barcodes. The continually linked sequencing library can be sequenced to derive continuity sequence information. The continuity particles described herein may be contacted with CPT-seq reagents to perform CPT-seq on nucleic acids extracted from cells.
[0069] As used herein, the term “continuity information” refers to the spatial relationships between two or more DNA fragments based on shared information. The shared information may relate to adjacency, compartmentalization, and distance. Information regarding these relationships facilitates the hierarchical assembly or mapping of sequence reads derived from DNA fragments. Because traditional assembly or mapping methods used in conjunction with conventional shotgun sequencing do not consider the relative genomic origin or coordinates of individual sequence reads in relation to the spatial relationships between two or more DNA fragments from which each sequence read originates, this continuity information improves the efficiency and accuracy of such assembly or mapping.
[0070] Accordingly, according to the embodiments described herein, methods for capturing continuity information can be achieved by short-range continuity methods for determining adjacent spatial relationships, intermediate-range continuity methods for determining partition spatial relationships, or long-range continuity methods for determining distance-spatial relationships. These methods enhance the accuracy and quality of DNA sequence assembly or mapping and can be used in conjunction with any sequencing method as described herein.
[0071] Continuity information includes the relative genomic origin or coordinates of individual sequence reads, relating to the spatial relationship between two or more DNA fragments from which each sequence read originates. In some embodiments, continuity information includes sequence information from non-overlapping sequence reads.
[0072] In some embodiments, the continuity information of the target nucleic acid sequence represents haplotype information. In some embodiments, the continuity information of the target nucleic acid sequence represents genomic variants.
[0073] Single-cell combinatorial index sequencing (SCI-seq) is a sequencing method for simultaneously creating thousands of low-pass single-cell libraries for somatic copy number variant detection.
[0074] Therefore, multiple simultaneous assays, including the assays described herein, can be performed on continuous particles, either alone or in combination with any other assay, for the purpose of analyzing nucleic acids.
[0075] Indexed continuous particles can also be directly loaded onto a flow cell held through a post / microwell array. The indexed library is released from the continuous particles (chemical / thermal release) and bound to the flow cell. This allows for an electric indexing approach where the first level of indexing originates from spatial location, and the next level originates from an indexed library from a single continuous particle. Alternatively, an indexed library extracted from continuous particles can be collectively loaded into the flow cell.
[0076] In some embodiments, the continuous particles come into contact with one or more reagents for nucleic acid processing. In some embodiments, the reagents may include solubilants, nucleic acid purifiers, DNA amplifiers, tagging agents, PCR agents, or other agents used for processing genetic material. Thus, the continuous particles provide a microenvironment for controlled nucleic acid reactions.
[0077] In some embodiments, whole-DNA library preparation can be achieved seamlessly within continuous particles by performing multiple reagent exchanges by passing gDNA and its library products through a porous hydrogel while retaining them within a polymer shell. The hydrogel can withstand high temperatures up to 95°C for several hours to support different biochemical reactions.
[0078] As used herein, the terms “isolated,” “isolate,” “purified,” “purify,” and “refined,” and their grammatical equivalents as used herein, refer to a reduction in the amount of at least one contaminant (such as proteins and / or nucleic acid sequences) from a sample or from a source from which the material was isolated (e.g., cells), unless otherwise specified. Thus, purification results in “enrichment,” for example, an increase in the amount of a desired protein and / or nucleic acid sequence in the sample.
[0079] Following the lysis and isolation of nucleic acids, amplification can be performed, such as multiple substitution amplification (MDA), a technique widely used to amplify small amounts of DNA, particularly from single cells. In some embodiments, nucleic acids are amplified, sequenced, or used for the preparation of nucleic acid libraries. As used herein, the terms “amplify,” “amplified,” and “to amplify” as used with respect to nucleic acids or nucleic acid reactions refer to in vitro methods for producing copies of a particular nucleic acid, such as a target nucleic acid or a nucleic acid associated with a continuous particle, according to, for example, one embodiment of the present invention. Many methods for amplifying nucleic acids are known in the art, and amplification reactions include polymerase chain reactions, ligase chain reactions, strand substitution amplification reactions, rolling circle amplification reactions, multiple annealing and looping based amplification cycles (MALBAC), transcription-mediated amplification methods such as NASBA, and loop-mediated amplification methods (e.g., “LAMP” amplification using loop-forming sequences). The nucleic acid to be amplified may consist of, or derive from, DNA or RNA, or a mixture of DNA and RNA including modified DNA and / or RNA. The product obtained from the amplification of one or more nucleic acid molecules (e.g., “amplification product”) may be either DNA or RNA, or a mixture of both DNA and RNA nucleosides or nucleotides, regardless of whether the starting nucleic acid is DNA, RNA, or both, or may include modified DNA or RNA nucleosides or nucleotides. “Copy” does not necessarily mean complete sequence complementarity or identity with respect to the target sequence. For example, a copy may include nucleotide analogs such as deoxyinosine or deoxyuridine, intentional sequence modifications (such as sequence modifications introduced via primers containing sequences that are hybridizable to the target sequence but not complementary), and / or sequence errors that occur during amplification.
[0080] Nucleic acids associated with continuous particles can be amplified according to any suitable amplification method known in the art. In some embodiments, the nucleic acids are amplified on the continuous particles. In some embodiments, the continuous particles are captured and degraded on a solid support, the nucleic acids are released onto the solid support, and the nucleic acids are amplified on the solid support.
[0081] It will be understood that nucleic acids can be amplified using any of the amplification methods described herein or commonly known in the art, together with universal or target-specific primers. Preferred amplification methods include polymerase chain reaction (PCR), strand displacement amplification (SDA), transcription-mediated amplification (TMA), and nucleic acid sequence-based amplification (NASBA), as described in U.S. Patent No. 8,003,354, which is incorporated herein by reference in its entirety. The above amplification methods can be used to amplify one or more nucleic acids of interest. For example, nucleic acids can be amplified using PCR, including multiplex PCR, SDA, TMA, NASBA, etc. In some embodiments, primers specifically directed to the nucleic acid of interest are included in the amplification reaction.
[0082] Other suitable methods for nucleic acid amplification include oligonucleotide extension and ligation, rolling circle amplification (RCA) (incorporated herein by reference, Lizardi et al., Nat. Genet. 19:225-232 (1998)), and oligonucleotide ligation assay (OLA) techniques (all incorporated by reference, see generally U.S. Patents Nos. 7,582,420, 5,185,243, 5,679,524, and 5,573,907, European Patent Nos. 0320308(B1), 0336731(B1), 0439182(B1), International Publication Nos. 90 / 01069, 89 / 12696, and 89 / 09835). It will be understood that these amplification methods may be designed to amplify nucleic acids. For example, in some embodiments, the amplification method may include a ligation probe amplification or oligonucleotide ligation assay (OLA) reaction comprising a primer specifically directed to the nucleic acid of interest. In some embodiments, the amplification method may include a primer extension ligation reaction comprising a primer specifically directed to the nucleic acid of interest and capable of passing through hydrogel pores. As non-limiting examples of primer extension and ligation primers that may be specifically designed to amplify the nucleic acid of interest, the amplification may include primers used in GoldenGate assays (Illumina, Inc., San Diego, Calif.), as exemplified in U.S. Patents 7,582,420 and 7,611,869, respectively, which are incorporated herein by reference in their entirety. In each of the methods described, the reagents and components involved in the nucleic acid reaction can pass through the pores of the continuous particles while retaining the nucleic acid itself within the continuous particles.
[0083] In some embodiments, nucleic acids are amplified using cluster amplification methods, as illustrated by the disclosures of U.S. Patents 7,985,565 and 7,115,400, the contents of which are incorporated herein by reference in their entirety. The incorporated materials of U.S. Patents 7,985,565 and 7,115,400 describe nucleic acid amplification methods for immobilizing amplification products onto a solid support to form an array consisting of clusters or “colonies” of immobilized nucleic acid molecules. Each cluster or colony on such an array is formed from multiple identical immobilized polynucleotide chains and multiple identical immobilized complementary polynucleotide chains. The array thus formed is generally referred herein to as a “clustered array.” Products of solid-phase amplification reactions, such as those described in U.S. Patents 7,985,565 and 7,115,400, are so-called "crosslinked" structures formed by annealing a pair of immobilized polynucleotide chains and an immobilized complementary chain, both chains being immobilized on a solid support, preferably via covalent bonds at their 5' ends. Cluster amplification methodologies are an example of a method for producing immobilized amplicons using an immobilized nucleic acid template. Immobilized amplicons can also be produced from immobilized DNA fragments produced according to the methods provided herein using other suitable methodologies. For example, one or more clusters or colonies may be formed by solid-phase PCR, regardless of whether one or both primers of each pair of amplification primers are immobilized. In some embodiments, nucleic acids are amplified on continuous particles and then deposited in an array in clusters or on a solid support.
[0084] Further amplification methods include isothermal amplification. Exemplary isothermal amplification methods that can be used include, but are not limited to, multi-substitution amplification (MDA) as exemplified by Dean et al., Proc. Natl. Acad. Sci. USA 99:5261-66 (2002), or isothermal chain substitution nucleic acid amplification as exemplified by U.S. Patent No. 6,214,587 (each of which is incorporated herein by reference in whole). Other non-PCR-based methods that may be used in this disclosure include, for example, strand displacement amplification (SDA) described in Walker et al., Molecular Methods for Virus Detection, Academic Press, Inc., 1995, U.S. Patents 5,455,166 and 5,130,238, and Walker et al., Nucl. Acids Res. 20:1691-96 (1992), or superbranched strand displacement amplification described in, for example, Lage et al., Genome Research 13:294-307 (2003), each of which is incorporated herein by reference in its entirety. Isothermal amplification can be used with large fragments of strand displacement Phi 29 polymerase or Bst DNA polymerase, 5'→3' exo-, for random primer amplification of genomic DNA. The use of these polymerases takes advantage of their high processability and strand displacement activity. Due to their high processability, polymerases are capable of producing fragments 10–20 kb in length. As described above, smaller fragments can be produced under isothermal conditions using polymerases with low processability and strand displacement activity, such as Klenow polymerase. Further descriptions of amplification reactions, conditions, and components are detailed in the disclosure of U.S. Patent No. 7,670,810, which is incorporated herein by reference in its entirety. In some embodiments, the polymerases, reagents, and components required to carry out these amplification reactions pass through the pores of continuous particles and interact with nucleic acids, thereby amplifying the nucleic acids in the continuous particles. In some embodiments, random hexamers are annealed to denatured DNA, followed by strand displacement synthesis at a constant temperature in the presence of the catalytic enzyme Phi 29.This results in DNA amplification in the continuous particles, as confirmed by the increase in fluorescence intensity after MDA (DNA stained with SYTOX). Independently, Nextera-based tagmentation and subsequent PCR-based gDNA amplification can also be performed after lysis and cleanup, as indicated by the substantial increase in fluorescence intensity in the continuous particles after Nextera tagmentation and PCR. After this Nextera library preparation, the continuous particles can be heated to 80°C for 3 minutes to release the contents of the continuous particles, i.e., the library product ready for sequencing, from the cells.
[0085] Another nucleic acid amplification method useful in this disclosure is tagged PCR using a population of two-domain primers, each consisting of a constant 5' region followed by a random 3' region, as described, for example, in Grothues, et al. Nucleic Acids Res. 21(5):1321-2 (1993), which is incorporated herein by reference in its entirety. An initial round of amplification is performed to allow numerous initiations in the thermally denatured DNA based on individual hybridization from randomly synthesized 3' regions. Due to the nature of the 3' region, the start sites are considered to be random across the genome. Subsequently, unbound primers can be removed, and further replication can be performed using primers complementary to the constant 5' region.
[0086] In some embodiments, nucleic acids are fully or partially sequenced on continuous particles. Nucleic acids can be sequenced according to any suitable sequencing method, including direct sequencing, such as synthesis sequencing, ligation sequencing, hybridization sequencing, and nanopore sequencing.
[0087] One sequencing method is sequencing-by-synthesis (SBS). In SBS, the sequence of nucleotides in a nucleic acid template is determined by monitoring the extension of nucleic acid primers along a nucleic acid template (e.g., a target nucleic acid or its amplicon). The underlying chemical process may be polymerization (e.g., catalyzed by a polymerase enzyme). In certain polymer-based embodiments of SBS, fluorescently labeled nucleotides are attached to the primers in a template-dependent manner (thus extending the primers) so that the sequence of the template can be determined by detecting the order and type of nucleotides attached to the primers.
[0088] One or more amplified nucleic acids can be subjected to SBS or other detection techniques, which involve repeated delivery of reagents during a cycle. For example, to initiate a first SBS cycle, one or more labeled nucleotides, DNA polymerase, etc., can be introduced into / passed through a continuous particle containing one or more amplified nucleic acid molecules. The site where the labeled nucleotides are incorporated by primer extension can be detected. Optionally, the nucleotides may further include reversible termination properties, where the addition of the nucleotide to the primer terminates further primer extension. For example, a nucleotide analog with a reversible terminator moiety can be added to the primer so that further extension does not occur until a deblocking agent is delivered and that moiety is removed. Thus, in embodiments using reversible termination, a deblocking reagent can be delivered to the flow cell (before or after detection occurs). Washing can be performed between various delivery steps. The cycle can then be repeated n times to extend the primer with n nucleotides, thereby enabling the detection of a sequence of length n. Exemplary SBS procedures, fluid systems, and detection platforms that can be readily adapted for use with amplicons produced by the methods of the present disclosure are described, for example, in Bentley et al., Nature 456:53-59 (2008), International Publication No. 04 / 018497, U.S. Patent No. 7,057,026, International Publication No. 91 / 06678, International Publication No. 07 / 123744, U.S. Patent No. 7,329,492, U.S. Patent No. 7,211,414, U.S. Patent No. 7,315,019, U.S. Patent No. 7,405,281, and U.S. Patent Application Publication No. 2008 / 0108082, each of which is incorporated herein by reference.
[0089] Other sequencing procedures using cyclic reactions, such as pyrosequencing, can be used. Pyrosequencing detects the release of inorganic pyrophosphate (PPi) when specific nucleotides are incorporated into nascent nucleic acid chains (Ronaghi et al., Analytical Biochemistry 242(1),84-9(1996), Ronaghi, Genome Res. 11(1),3-11(2001), Ronaghi et al. Science 281(5375),363(1998), U.S. Patent No. 6,210,891, U.S. Patent No. 6,258,568, and U.S. Patent No. 6,274,320, each incorporated herein by reference). In pyrosequencing, the released PPi can be detected by its immediate conversion to adenosine triphosphate (ATP) by ATP sulfurylase, and the level of the generated ATP can be detected via luciferase-produced photons. Therefore, the sequencing reaction can be monitored via a luminescence detection system. Excitation radiation sources used in fluorescence-based detection systems are not required for the pyrosequencing procedure. Useful fluid systems, detectors, and procedures that can be adapted for the application of pyrosequencing to amplicons produced according to this disclosure are described, for example, in International Application US11 / 57111, U.S. Patent Application Publication 2005 / 0191698(A1), U.S. Patent No. 7,595,883, and U.S. Patent No. 7,244,559, each of which is incorporated herein by reference.
[0090] Several embodiments can utilize methods involving real-time monitoring of DNA polymerase activity. For example, nucleotide incorporation can be detected via fluorescence resonance energy transfer (FRET) interactions between fluorophore-supported polymerase and γ-phosphate-labeled nucleotides, or using a zero-mode waveguide (ZMW). Techniques and reagents for FRET-based sequencing are described, for example, in Levene et al. Science 299, 682-686 (2003), Lundquist et al. Opt. Lett. 33, 1026-1028 (2008), and Korlach et al. Proc. Natl. Acad. Sci. USA 105, 1176-1181 (2008), the disclosures of which are incorporated herein by reference.
[0091] Some SBS embodiments include the detection of protons released during the incorporation of nucleotides into the extension product. For example, sequencing based on the detection of released protons can use commercially available electrodetectors and related technologies. Examples of such sequencing systems include pyrosequencing (e.g., a commercially available platform from 454 Life Sciences, a subsidiary of Roche), sequencing using γ-phosphate-labeled nucleotides (e.g., a commercially available platform from Pacific Biosciences), sequencing using proton detection (e.g., a commercially available platform from Ion Torrent, a subsidiary of Life Technologies), or sequencing methods and systems described in U.S. Patent Application No. 2009 / 0026082(A1), U.S. Patent Publication No. 2009 / 0127589(A1), U.S. Patent Publication No. 2010 / 0137143(A1), or U.S. Patent Publication No. 2010 / 0282617(A1), each of which is incorporated herein by reference. The method described herein for amplifying a target nucleic acid using dynamic exclusion can be readily applied to substrates used for proton detection. More specifically, the method described herein can be used to generate a clonal population of amplicons used for proton detection.
[0092] Another sequencing technique is nanopore sequencing (see, for example, Deamer et al. Trends Biotechnol. 18, 147-151 (2000), Deamer et al. Acc. Chem. Res. 35: 817-825 (2002), and Li et al. Nat. Mater. 2: 611-615 (2003), the disclosures of which are incorporated herein by reference). In some embodiments of the nanopore, the target nucleic acid or individual nucleotides removed from the target nucleic acid pass through the nanopore. As the nucleic acid or nucleotides pass through the nanopore, the type of each nucleotide can be identified by measuring the variation in the electrical conductance of the pore. (U.S. Patent No. 7,001,792, Soni et al. Clin. Chem. 53, 1996-2001 (2007), Healy, Nanomed. 2, 459-481 (2007), Cockroft et al. J. Am. Chem. Soc. 130, 818-820 (2008), these disclosures are incorporated herein by reference).
[0093] Exemplary methods for array-based expression and genotyping analysis that can be applied to the detections described herein are described in U.S. Patent Nos. 7,582,420, 6,890,741, 6,913,884, or 6,355,431, or U.S. Patent Publication Nos. 2005 / 0053980(A1), 2009 / 0186349(A1), or 2005 / 0181440(A1), each of which is incorporated herein by reference.
[0094] In the methods for isolating, amplifying, and sequencing nucleic acids described herein, various reagents are used for the isolation and preparation of nucleic acids. Such reagents may include, for example, lysozyme, proteinase K, random hexamer, polymerase (e.g., Φ29 DNA polymerase, Taq polymerase, Bsu polymerase), transposase (e.g., Tn5), primer (e.g., P5 and P7 adapter sequences), ligase, catalytic enzyme, deoxynucleotide triphosphate, buffer, or divalent cation. These reagents pass through the pores of the continuous particles, while the genetic material is retained within the continuous particles. An advantage of the methods described herein is that they provide a microenvironment for processing nucleic acids on continuous particles.
[0095] The adapter may include sequencing primer sites, amplification primer sites, and an index. As used herein, “index” may include a sequence of nucleotides that can be used as a molecular identifier and / or barcode for tagging nucleic acids and / or identifying the source of nucleic acids. In some embodiments, the index can be used to identify a single nucleic acid or a subpopulation of nucleic acids. In some embodiments, a nucleic acid library can be prepared in continuity particles. In some embodiments, a single cell may be processed to obtain nucleic acids that associate with continuity particles and then used for combined indexing of nucleic acids, for example, using a continuity-preserved transposition sequencing (CPT-seq) approach. In some embodiments, DNA derived from a single cell may be barcoded by encapsulation of the single cell after WGA amplification with a barcoded transposon.
[0096] The embodiments of the “spatial indexing” methods and techniques described herein shorten data analysis and simplify the library preparation process from single cells and long DNA molecules. Existing protocols for single-cell sequencing require efficient physical separation of cells, unique barcoding of each isolated cell, and pooling all together for sequencing. Current protocols for synthetic long reads also require a cumbersome barcoding process to distinguish the genetic information arising from each barcoded cell, and pooling each barcoded fragment together for sequencing and data analysis. During these lengthy processes, there is also loss of genetic material that causes sequencing errors. The embodiments described herein not only shorten the process but also increase the data resolution of single cells. Furthermore, the embodiments provided herein simplify the assembly of genomes of novel organisms. The embodiments described herein may be used to reveal the co-occurrence of rare genetic differences and mutations. In some embodiments, DNA libraries confined to continuous particles until release offer the opportunity to control the size of fragments released onto the surface by controlling the release process and hydrogel formulation.
[0097] In some embodiments, the library may be amplified using primer sites in the adapter sequence and sequenced using sequencing primer sites in the adapter sequence. In some embodiments, the adapter sequence may include an index for identifying the nucleic acid source. The efficiency of the subsequent amplification step may be reduced by the formation of primer dimers. To improve the efficiency of the subsequent amplification step, the unligated single-stranded adapter may be removed from the ligation product.
[0098] Preparation of nucleic acid libraries using continuous particles Some embodiments of the systems, methods, and compositions provided herein include a method by which an adapter is ligated to a target nucleic acid. The adapter may include a sequencing primer binding site, an amplification primer binding site, and an index. For example, the adapter may include a P5 sequence, a P7 sequence, or their complements. When used herein, the P5 sequence includes the sequence defined by SEQ ID NO: 1 (AATGATACGGCGACCACCGA), and the P7 sequence includes the sequence defined by SEQ ID NO: 2 (CAAGCAGAAGACGGCATACGA). In some embodiments, the P5 or P7 sequence may further include a spacer polynucleotide, which may be 1 to 20, e.g., 1 to 15, or 1 to 10 nucleotides, e.g., 2, 3, 4, 5, 6, 7, 8, 9, or 10 nucleotides in length. In some embodiments, the spacer includes 10 nucleotides. In some embodiments, the spacer is a poly-T spacer, such as a 10T spacer. The spacer nucleotide may be contained in the 5' end of the polynucleotide or may be attached to a suitable support via ligation to the 5' end of the polynucleotide. Adhesion can be achieved by a sulfur-containing nucleophile, such as a phosphorothioate, present at the 5' end of the polynucleotide. In some embodiments, the polynucleotide includes a poly-T spacer and a 5' phosphorothioate group. Thus, in some embodiments, the P5 sequence is 5'phosphorothioate-TTTTTTTTTTAATGATACGGCGACCACCGA-3' (SEQ ID NO: 3), and in some embodiments, the P7 sequence is 5'phosphorothioate-TTTTTTTTTTCAAGCAGAAGACGGCATACGA-3' (SEQ ID NO: 4).
[0099] The index can be useful in identifying the source of nucleic acid molecules. In some embodiments, the adapter can be modified to prevent concatemer formation by adding blocking groups that prevent the adapter from elongating, for example, at one or both ends. Examples of 3' blocking groups include 3' spacer C3, dideoxynucleotides, and attachment to the substrate. Examples of 5' blocking groups include dephosphorylated 5' nucleotides and attachment to the substrate.
[0100] The adapter contains nucleic acids, such as single-stranded nucleic acids. The adapter may contain short nucleic acids having lengths of approximately 5 nucleotides, 10 nucleotides, 20 nucleotides, 30 nucleotides, 40 nucleotides, 50 nucleotides, 60 nucleotides, 70 nucleotides, 80 nucleotides, 90 nucleotides, 100 nucleotides, less than or greater than these lengths, or in a range between any two of the aforementioned sizes. In some embodiments, the adapter is large enough to pass through the pores of continuous particles. The target nucleic acid may include DNA such as genome or cDNA, RNA such as mRNA, sRNA, or rRNA, or a hybrid of DNA and RNA. The nucleic acid can be isolated from a single cell. The nucleic acid may contain phosphodiester bonds and other types of backbone, including, for example, phosphoramides, phosphorothioates, phosphorodithioates, O-methylphosphoramidites, and peptide nucleic acid backbones and bonds. Nucleic acids can include any combination of deoxyribonucleotides and ribonucleotides, as well as any combination of bases including uracil, adenine, thymine, cytosine, guanine, inosine, xanthine, hypoxanthine, isocytosine, isoguanine, and base analogs such as nitropyrrole (including 3-nitropyrrole) and nitroindole (including 5-nitroindole). In some embodiments, nucleic acids can include at least one indiscriminate base. Indiscriminate bases can base-pair with multiple different types of bases and may be useful, for example, when included in oligonucleotide primers or inserts used for random hybridization in complex nucleic acid samples such as genomic DNA samples. Examples of indiscriminate bases include inosine, which can base-pair with adenine, thymine, or cytosine. Other examples include hypoxanthine, 5-nitroindole, acrylic 5-nitroindole, 4-nitropyrazole, 4-nitroimidazole, and 3-nitropyrrole. Nonspecific bases can be used that can base-pair with at least two, three, four or more types of bases.
[0101] The target nucleic acid may include samples in which the average size of the nucleic acid in the sample is approximately 2kb, 1kb, 500bp, 400bp, 200bp, 100bp, 50bp, less than or greater than these sizes, or in a range between any two of the aforementioned sizes. In some embodiments, the average size of the nucleic acid in the sample is approximately 2000 nucleotides, 1000 nucleotides, 500 nucleotides, 400 nucleotides, 200 nucleotides, 100 nucleotides, 50 nucleotides, less than or greater than these sizes, or in a range between any two of the above sizes. In some embodiments, the nucleic acid is large enough to be trapped within the continuous particles so that the nucleic acid cannot pass through the pores of the continuous particles.
[0102] An exemplary method includes dephosphorylating the 5' end of a target nucleic acid to prevent concatemer formation in a subsequent ligation step; ligating the first adapter to the 3' end of the dephosphorylated target using a ligase in which the 3' end of the first adapter is blocked; rephosphorylating the 5' end of the ligated target; and ligating the second adapter to the 5' end of the dephosphorylated target using a single-stranded ligase in which the 5' end of the second adapter is dephosphorylated.
[0103] Another example involves partially digesting a nucleic acid with a 5' exonuclease to form a double-stranded nucleic acid with a single-stranded 3' overhang. An adapter containing a 3' blocking group can be ligated to the 3' end of the double-stranded nucleic acid with a 3' overhang. The double-stranded nucleic acid with a 3' overhang and the ligated adapter can be dehybridized to form a single-stranded nucleic acid. An adapter containing an unphosphorylated 5' end can be ligated to the 5' end of a single-stranded nucleic acid.
[0104] Methods for dephosphorylating nucleic acids, such as the 5' nucleotide, include contacting the nucleic acid with a phosphatase. Examples of phosphatases include bovine intestinal phosphatase, shrimp alkaline phosphatase, Antarctic phosphatase, and APEX alkaline phosphatase (Epicentre).
[0105] Methods for ligating nucleic acids include contacting the nucleic acid with a ligase. Examples of ligases include T4 RNA ligase 1, T4 RNA ligase 2, RtcB ligase, Methanobacterium RNA ligase, and TS2126 RNA ligase (CIRCLIGASE).
[0106] Methods for phosphorylating nucleic acids, such as the 5' nucleotide, include contacting the nucleic acid with a kinase. An example of a kinase is T4 polynucleotide kinase.
[0107] Embodiments provided herein relate to preparing nucleic acid libraries in continuous particles such that the nucleic acid library is prepared in a single reaction volume.
[0108] Embodiments of systems and methods provided herein include kits comprising one or more hydrogel polymers, crosslinkers, or microfluidic devices for preparing continuous particles, further comprising components useful for processing genetic material, as described herein and used for each processing of genetic material, including lysozyme, proteinase K, random hexamer, polymerase (e.g., Φ29 DNA polymerase, Taq polymerase, Bsu polymerase), transposase (e.g., Tn5), primer (e.g., P5 and P7 adapter sequences), and reagents for cell lysis, nucleic acid amplification and sequencing, or nucleic acid library preparation, comprising ligase, catalytic enzyme, deoxynucleotide triphosphate, buffer, or divalent cation. [Examples]
[0109] Example 1 - Preparation of continuous particles The following examples demonstrate an embodiment for preparing continuous particles associated with transposomes using a microfluidic droplet generator.
[0110] Samples containing cells stored at -80°C were thawed at room temperature. 100 μL of each sample was transferred to a sterile 1.7 mL tube, and the sample was washed once with 1 mL of 0.85% NaCl. The samples were pelletized, and the washing solution was removed. The cell pellets were mixed with hydrogel solution, and the cells were resuspended in the hydrogel solution.
[0111] To generate continuous particles with a uniform size distribution, a microfluidic droplet generator, such as the one shown in Figure 1, was used. A solution containing the hydrogel polymer and cells was introduced into the first channel of the microfluidic droplet generator. Mineral oil, used as a spacer oil, was added to the second channel, and a crosslinking agent was added to the third channel. Upon contact with the crosslinking agent in the third channel, the hydrogel instantaneously formed continuous particles associated with transposomes. The type of crosslinking oil was selected to adjust the crosslinking rate, including slow-speed or instantaneous crosslinking agents, as shown in Table 1.
[0112] [Table 1]
[0113] Example 2 - Simultaneous assay performed on continuous particles The following examples demonstrate exemplary assays performed on continuous particles from Example 1, including SCI-seq, ATAC-seq, combined indexing, and single-cell whole-genome amplification.
[0114] Continuum particles were obtained from Example 1 and deposited on a plate having wells such that single continuous particles were deposited in single wells. The cells were lysed by introducing lysis buffer, followed by washing, thereby extracting nucleic acids from the cells. The continuous particles containing the lysed cells were then exposed to the series of assays described below.
[0115] Indexed transposomes (TSMs) were used to tag genomic DNA and generate ATAC-seq fragments. After proteinase / SDS treatment, oligo-T with the same index as the TSMs was added to each well to initiate cDNA synthesis by reverse transcription (RT). A PCR adapter on the other end was introduced by randomizer extension, generating index 1. Continuity particles from each well were pooled together and then split into indexed PCR plates to generate index 2. This two-layer indexing can be scaled up to 150,000 (384 × 384) cells. The resulting final library is a mixture of ATAC-seq and RNA-seq, where gDNA and cDNA from the same cells are grouped by the same index, and the oligo-T-UMI pattern served as an internal marker to distinguish RNA signals from ATAC signals.
[0116] In addition, random extension was also used for full-length RNA-seq. In this case, the TSM index differed from the randommer index. The one-to-one matching between these two index sets helped to distinguish reads from single cells as well as differentiated DNA and RNA signals. UMI was also applied in this manner to improve the accuracy of read analysis.
[0117] A three-layer combined indexed assay was also performed using two indexed sprint ligation and one indexed PCR. In this method, TSM and oligo-T were universal, both containing sprint 1 fragments, enabling indexing by sprint ligation. Three-layer indexing achieved a throughput of up to 1 million cells (96×96×96). To increase cell throughput, three-layer combined indexing for ATAC-seq was performed using indexed TSM. Instead of using universal TSM, the TSM contained its own index on its B7G and A7G sides. Two different sprints were used to attach the indexed adapter required for indexed PCR. The indexing includes the following components: B15 adapter sequence (GTCTCGTGGGCTCGG, SEQ ID NO: 5), N6, and link 1, together forming the B15_N6_link1 sequence (GTCTCGTGGGCTCGGNNNNNNGACTTGTC, SEQ ID NO: 11), Phos_link 2, A7G sequence (TGGTAGAGAGGGTG, SEQ ID NO: 9), and ME sequence (AGATGTGTATAAGAGACAG, SEQ ID NO: 7), together forming the Phos_link2_A7G_ME sequence (TAGAGCATNNNNNNTGGTAGAGAGGGTGAGATGTGTATAAGAGA CAG (sequence number 12), and the A14 adapter sequence (TCGTCGGCAGCGTC, sequence number 6), N6, and link 1 together form the A14_N6_link1 sequence (TCGTCGGCAGCGTCNNNNNNGTAATCAC, sequence number 13), and Phos_link 2, B7G (TACTACTCACCTCCC, sequence number 10), and the ME sequence (AGATGTGTATAAGAGACAG, sequence number 7), together form the Link2_B7G_ME sequence (CATCATCCNNNNNNTACTACTCACCTCCCAGATGTGTATAAGAGACAG, sequence number 14).This embodiment also provides the ME complementary sequence (TCTACACACATTCTCTGTC, SEQ ID NO: 8), the sprint 1 sequence (ATGCTCTAGACAAGT, SEQ ID NO: 15), and the sprint 2 sequence (GGATGATGGTGATTA, SEQ ID NO: 16). In the sequences, N is A, C, T, or G.
[0118] Single-cell whole-genome amplification was also performed using continuous particles. This was done by using indexed T7 transpositions of continuous particles in individual wells, followed by pooling and extension via T7 in vitro transcription (IVT) linear amplification. The beads were again separated and pooled for indexed random extension, and then split again for final indexed PCR.
[0119] The continuous particles exhibited efficient tagmentation when targeted by FAM-labeled transposomes. Nuclei were stained with Hoechst (a blue-luminescent DNA stain), while transposition nuclei were stained fluorescent green with FAM (a fluorescent dye). Lysing cells with SDS increased the background fluorescence of the continuous particles, but no signal was observed from continuous particles without any cells. The results demonstrate that an ATAC-seq library can be generated for nucleic acid molecules associated with continuous particles, even if only a small portion of the short fragments leaked from tagmentation are present.
[0120] Example 3 - Preparation of a nucleic acid library in continuous particles The following examples demonstrate a method for nucleic acid preparation using continuous particles.
[0121] Contiguity particles prepared in Example 1 were obtained. The contiguity particles (CPs) were loaded onto a 45 μm cell strainer and washed multiple times with PBS or Tris-Cl to remove unencapsulated cells. One advantage of encapsulating cells in contiguity particles is the improved ability to handle and process them. One simple way to do this is through the use of a spin column or filter plate. The pore size of the filter can be smaller than the bead diameter. Examples of filter plates include, but are not limited to, Millipore's MultiScreen-Mesh Filter Plates with 20, 40, and 60 μm pore sizes, Millipore's MultiScreen Migration Invasion and Chemotaxis Filter Plate with 8.0 μm pores, or Pall's AcroPrep Advance 96-Well Filter Plates for Aqueous Filtration with 30-40 μm pores. These filter plates allowed for easy separation of bead-encapsulated cells from solution and enabled multiple buffer exchanges.
[0122] The washed beads were suspended in buffer and removed from the filter. Aliquots of the beads were visualized under a microscope to estimate the final bead concentration, the efficiency of cell loading, and to ensure that no unencapsulated cells remained.
[0123] For Nextera tagmentation, a Millipore 20 μm Nylon MultiScreen-Mesh Filter Plate was used. To limit bead adhesion to the filter, the beads were pre-moistened with Pluronic F-127. After centrifugation of the plate at 500 g for 30 seconds, the buffer flowed through the filter would retain the beads. The beads were washed twice with 200 μL of Tris-Cl buffer and then suspended in lysis buffer (0.1% SDS). The beads were incubated in lysis buffer for 1 minute before removal by centrifugation. Two additional 200 μL Tris-Cl washes were performed to remove residual lysis buffer. The cells were then suspended in 45 μL of 1x tagmentation buffer by pipetting up and down and then transferred to strip tubes. 45 μL of beads were mixed with 5 μL of Tagmentation DNA Enzyme (TDE, Illumina Inc.), and incubated in a thermal cycler at room temperature for 1 hour and at 55°C for 30 minutes. Alternatively, tagmentation could be performed on a filter plate by suspending bead-encapsulated cells in a tagmentation master mix and incubating on a heat block. After tagmentation, aliquots of the beads were stained with Hoechst dye and visualized under a microscope. Visualization confirmed that DNA remained on the beads.
[0124] Tagged beads (25 μL) were PCR amplified using Illumina's Nextera PCR MM (NPM, Illumina) and PCR primers. 0.1% SDS was added to the master PCR mix to remove DNA-bound Tn5. A preliminary incubation was performed at 75°C to facilitate Tn5 removal in the presence of SDS. Eleven PCR cycles were performed to generate the final library. After PCR, the library was purified using 0.9x SPRI, quantified using dsDNA qubit and / or BioAnalyzer, and sequenced on MiSeq and / or NextSeq.
[0125] The method described in this embodiment can be extended to perform single-cell sequencing using an approach similar to sciSEQ (Vitak et al. Nat Meth. 2017;14:302-308). For example, the use of a 96-well filter plate for simultaneous performance of 96 indexed tagmentation reactions can be performed in the extension process. After adding the beads to the plate, continuous buffer is added and then removed by centrifugation or vacuum. After tagmentation, the beads are collected from the filter and pooled. The pooled beads are redistributed into a second 96-well PCR plate for multiplexed PCR. In this dual-level indexing scheme (tagmentation and PCR indexing), all DNA fragments derived from single-bead encapsulated cells receive the same barcode that can be deconvoluted to reconstruct the cell's genome.
[0126] The process outlined in this embodiment is suitable for automation on a liquid processing platform by adding a vacuum manifold such as Millipore's MultiScreen HTS Vacuum Manifold or Orochem's 96-well Plate Vacuum Manifold. These vacuum manifolds can be added to many liquid processing platforms, including Biomek FX, Microlab Star, and Tecan Genesis. Tagmentation can be automated by transferring the filter plate to a heat block. The plate is then transferred back to the vacuum manifold for post-tagmentation cleaning.
[0127] Example 4 - Long-read indexing The following examples demonstrate method chromosome-level phasing using the long-read indexing methods and systems provided herein.
[0128] Transposome-associated continuous particles (also referred to herein as continuous beads) were prepared and subjected to the linked long-read method described herein, which is used for analysis on human leukocyte antigens (HLA). In addition to the continuous particles, primer beads were also prepared, each primer bead having an adapter, a barcode, and a primer. The continuous beads and primer beads were distributed together in droplets along with a mixed solution primer containing the adapter and primer (one continuous bead and one primer bead per droplet). On-bead tagmentation, amplification, and indexing were performed using the distributed droplets to index the bead pool of over 900,000 barcodes, ensuring that each barcode was represented relatively equally.
[0129] As shown in Figure 4, chromosomal-level phasing was obtained for HLA. Using this assay, chromosomal-level phasing up to 26 Mb was achieved, with only one switch error per 50 Mb length, covering over 99% of SNPs. The analysis required a one-day assay (over a 5.7-hour period) compared to a typical two-day assay required for 10X sequencing. Furthermore, as shown in Figure 5, the number of islands compared to island length using the long-read indexing method described herein indicates high DNA quality for the phasing metric.
[0130] Figure 6 shows the results of variant calling and phase blocking using the long-read indexing method described herein (left) compared with 10X sequencing (right). These results demonstrate that the method provided herein provides better average coverage and higher INDEL accuracy than 10X sequencing.
[0131] Finally, as shown in Figures 7, 8A, and 8B, the long-read indexing method described herein was applied to the HLA region (Figure 7), HLA-DPA1 (Figure 8A), and HLA-A (Figure 8B). Detailed results of the long-indexed read method described herein, compared to 10X sequencing, are provided in Table 2.
[0132] [Table 2]
[0133] The embodiments, examples, and drawings described herein provide compositions, methods, and systems for retaining genetic material in a physically confined space during the process from lysis to library generation. Some embodiments provide libraries resulting from a single long DNA molecule or a single cell released onto the surface of a flow cell in a confined space. As libraries from single DNA molecules or single cells in individual compartments are released onto the surface of the flow cell, the libraries from each compartment are seeded in very close proximity to one another.
[0134] As used herein, the term “includes” is synonymous with “contains,” “includes,” or “characterizes,” and is comprehensive or non-exclusive, not excluding any elements or methods or processes not further enumerated.
[0135] The above description discloses some methods and materials of the present invention. The present invention is open to modifications of methods and materials, as well as changes in manufacturing methods and equipment. Such modifications will become apparent to those skilled in the art by considering this disclosure or its disclosure or practice as disclosed herein. Accordingly, the present invention is not intended to be limited to any specific embodiment disclosed herein, but rather to encompass all modifications and substitutions that fall within the true scope and spirit of the invention.
[0136] All references cited herein, including but not limited to published and unpublished applications, patents, and references to documents, are incorporated herein by reference in their entirety and form part of this Specified. If any publication or patent or patent application incorporated by reference conflicts with any disclosure contained herein, this Specified is intended to take precedence and / or supersede such conflicting material. The present invention provides, for example, the following items: (Item 1) A system for nucleic acid indexed amplification, Multiple continuous beads, each continuous bead associating with a transposome and containing a bead-bound nucleic acid molecule, An indexed primer pool, Multiple primer beads, each primer bead comprising an adapter, a barcode, and a primer, and A solution primer, an indexed primer pool, and a solution primer, The continuous beads and primer beads are distributed together in the droplet, and the system further A system including a detector for obtaining sequencing data. (Item 2) The system according to item 1, wherein the continuous beads and / or the primer beads are hydrogel beads comprising a hydrogel polymer and a crosslinking agent. (Item 3) The system according to item 2, wherein the hydrogel polymer comprises polyethylene glycol (PEG)-thiol / PEG-acrylate, acrylamide / N,N'-bis(acryloyl)cystamine (BACy), PEG / polypropylene oxide (PPO), polyacrylic acid, poly(hydroxyethyl methacrylate) (PHEMA), poly(methyl methacrylate) (PMMA), poly(N-isopropylacrylamide) (PNIPAAm), poly(lactic acid) (PLA), poly(lactic acid-co-glycolic acid) (PLGA), polycaprolactone (PCL), poly(vinyl sulfonic acid) (PVSA), poly(L-aspartic acid), poly(L-glutamic acid), polylysine, agarose, alginate, heparin, sulfated alginic acid, dextran sulfate, hyaluronan, pectin, carrageenan, gelatin, chitosan, cellulose, or collagen. (Item 4) The system according to item 2, wherein the crosslinking agent comprises bisacrylamide, diacrylate, diallylamine, triallylamine, divinyl sulfone, diethylene glycol diallyl ether, ethylene glycol diacrylate, polymethylene glycol diacrylate, polyethylene glycol diacrylate, trimethylopropanotrimethacrylate, ethoxylated trimethylol triacrylate, or ethoxylated pentaerythritol tetraacrylate. (Item 5) The system described in item 1, wherein the nucleic acid is a DNA molecule with 50,000 base pairs or more. (Item 6) The system according to item 1, wherein the polymer is a P5 primer. (Item 7) The system according to item 1, wherein the solution primer includes an adapter and a primer. (Item 8) The system according to item 1, wherein the solution primer comprises a B15 adapter and a P7 primer. (Item 9) The system according to item 1, wherein the transposome comprises a transposase and a transposon. (Item 10) A flow cell device for nucleic acid indexed amplification, A solid support containing multiple distributed droplets, It associates with transposomes and forms continuous beads containing bead-bound nucleic acid molecules, A solid support comprising an adapter, a barcode, and primer beads including a primer, A flow cell device in which the plurality of distributed droplets are distributed along the surface of the solid support. (Item 11) The flow cell device according to item 10, wherein the solid support is functionalized with a surface polymer. (Item 12) The flow cell device according to item 11, wherein the surface polymer is poly(N-(5-azidoacetamidylpentyl)acrylamide-co-acrylamide) (PAZAM) or silane-free acrylamide (SFA). (Item 13) The flow cell device according to item 10, wherein the flow cell includes a patterned surface. (Item 14) The flow cell device according to item 13, wherein the patterned surface includes wells. (Item 15) The flow cell device according to item 14, wherein the well has a diameter of approximately 10 μm to approximately 50 μm, for example, a diameter of 10 μm, 15 μm, 20 μm, 25 μm, 30 μm, 35 μm, 40 μm, 45 μm, or 50 μm, or within the range defined by any two of the aforementioned values, and the well has a depth of approximately 0.5 μm to approximately 1 μm, for example, a depth of 0.5 μm, 0.6 μm, 0.7 μm, 0.8 μm, 0.9 μm, or 1 μm, or within the range defined by any two of the aforementioned values. (Item 16) The flow cell device according to item 14, wherein the well is made of a hydrophobic material. (Item 17) The flow cell device according to item 15, wherein the hydrophobic material comprises an amorphous fluoropolymer such as CYTOP, Fluoropel®, or Teflon®. (Item 18) The flow cell device according to item 10, wherein the nucleic acid is a DNA molecule with 50,000 base pairs or more. (Item 19) The flow cell device according to item 10, wherein the transposome comprises a transposase and a transposon. (Item 20) A nucleic acid indexing method, A series of beads for on-bead tagmentation, in which each bead is linked to a transposome and contains a bead-bound nucleic acid molecule, thereby generating a series of beads. The aforementioned nucleic acid molecule is subjected to a tagmentation reaction, Multiple primer beads, each primer bead containing an adapter, a barcode, and a primer, thereby generating multiple primer beads. The continuous beads and the primer beads are distributed together into a droplet along with the solution primer. The process involves amplifying nucleic acid molecules in the distributed droplets, A method comprising indexing the nucleic acid molecules in each droplet. (Item 21) The method according to item 20, wherein the nucleic acid is a DNA molecule with 50,000 base pairs or more. (Item 22) The method according to item 20, further comprising performing nucleic acid amplification on a nucleic acid molecule before performing the tagmentation reaction described above. (Item 23) The method according to item 22, wherein the amplification reaction includes multi-substitution amplification (MDA). (Item 24) The method according to item 20, wherein the tagmentation reaction comprises contacting the nucleic acid with a transposase mixture comprising an adapter sequence and a transposome. (Item 25) The method according to item 20, wherein the indexing is performed by polymerase chain reaction (PCR). (Item 26) The method according to item 20, wherein the droplets are distributed into more than 900,000 different indexed PCR compartments. (Item 27) The method according to item 20, further comprising distributing the aforementioned droplets onto a solid support. (Item 28) The method according to item 27, wherein the solid support is a flow cell device.
Claims
[Claim 1] The invention described in the specification.