Methods and compositions for tracking barcodes in partitions

By introducing transposase-mediated fragmentation and adapter insertion into the partition, the data fragmentation problem caused by multi-barcode beads is solved, and the conversion rate and accuracy of the sequencing data are improved.

CN120417997APending Publication Date: 2025-08-01BIO RAD LABORATORIES INC
View PDF 106 Cites 0 Cited by

Patent Information

Application Number
CN202380088316.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-10-20
Filing Date
2023-10-18
Publication Date
2025-08-01

AI Technical Summary

Technical Problem

When more than one barcode beads exist in the partition, the substrate and data are divided, resulting in fragmented data points, affecting the conversion rate of sequencing data.

Method used

Hybrid molecular fragments containing universal sequences and unique molecular identifier (UMI) barcodes are formed by introducing transposase-mediated fragmentation and adapter insertion in fixed and permeabilized cells, which are subsequently amplified in partitions and identify different bead-linked oligonucleotides in the same partitions by detection of the same UMI sequence in the sequencing reads.

Benefits of technology

It effectively solves the problem of data fragmentation caused by multiple barcode beads, improves the conversion rate and accuracy of sequencing data, and ensures that data from the same partition is correctly attributed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120417997A_ABST
    Figure CN120417997A_ABST
Patent Text Reader

Abstract

Methods and compositions for generating sequencing readings by initiating partitions. Unique molecular identifiers and bead-specific barcodes can be introduced into targeted nucleic acid fragments, and combinations of bead-specific barcodes, UMIs, and fragments can be used to identify when multiple bead-specific barcodes originate from the same partition, thereby improving deconvolution of partition-based sequence analysis.
Need to check novelty before this filing date? Find Prior Art

Description

Cross - reference to related patent applications

[0001] This patent application claims the benefit of priority of U.S. Provisional Patent Application No. 63 / 417,897, filed on October 20, 2022, which is incorporated herein by reference for all purposes. Background Art

[0002] Labeling biological substrates in partitions with molecular barcodes can provide new biological insights into substrates co - localized to discrete partitions through sequencing and analysis of the molecular barcodes. Increasing the number of barcoded effective partitions, such as droplets, increases the number of sequencing - based data points and converts a larger fraction of the input substrate into data. Beads can be used as delivery agents to deliver barcodes to partitions such as droplets. Thus, overloading barcoded beads in partitions, which results in partitions having more than one bead and increases the percentage of barcoded effective partitions, provides a higher substrate - to - sequencing data conversion rate. However, when two or more barcodes appear in a discrete partition, the substrate and data are split between the two barcodes, resulting in fragmented data points. The present disclosure provides solutions to problems arising when there are more than one barcoded bead in a partition, such as problems associated with fragmented data points. Summary of the Invention

[0003] In some embodiments, a method of sorting sequencing reads by initiating partitions is provided. In some embodiments, the method includes: Providing RNA / cDNA or DNA / cDNA hybrid molecules in fixed and permeabilized cells, wherein the fixed and permeabilized cells contain cross - linked molecules; Generating random breaks in the hybrid molecules and randomly inserting a first adapter oligonucleotide or a second adapter oligonucleotide at the breaks, thereby forming hybrid molecule fragments, the hybrid molecule fragments comprising (i) a first 5' end linked to the first adapter oligonucleotide and (ii) a first 3' end and (iii) a second 5' end linked to the second adapter oligonucleotide and (iv) a second 3' end, wherein the first adapter oligonucleotide comprises a first universal sequence, and the second adapter oligonucleotide comprises a second universal sequence, and wherein the first adapter oligonucleotide, the second adapter oligonucleotide, or both further comprise unique molecular identifier (UMI) sequences; Partition the cells together with one or more beads into partitions, wherein each bead is linked to multiple copies of a bead-specific barcoded oligonucleotide having the same 3' end comprising the first universal sequence or the second universal sequence, wherein bead-specific barcoded oligonucleotides linked to different beads can be identified by unique bead-specific barcodes in the bead-specific barcoded oligonucleotides, and wherein at least some of the partitions contain at least two different beads; Optionally, in the partition, reverse at least some of the crosslinks in the crosslinked molecules in the cells and / or lyse the cells; Before, during, or after the reversal, use a polymerase to (i) extend the first 3' end using the second adapter oligonucleotide as a template such that the first 3' end is linked to the reverse complementary sequence of the second universal sequence, and (ii) extend the second 3' end using the first adapter oligonucleotide as a template such that the second 3' end is linked to the reverse complementary sequence of the first universal sequence, thereby forming a gap-filled hybrid molecular fragment; Amplify the gap-filled hybrid molecular fragment in the partition by annealing the bead-specific barcoded oligonucleotide to the cDNA fragment and extending the bead-specific barcoded oligonucleotide to the reverse complementary sequence of the first universal sequence or the reverse complementary sequence of the second universal sequence to produce an amplicon, the amplicon comprising: a bead-specific barcoded oligonucleotide, a cDNA fragment, a second universal terminal sequence, and a UMI sequence, under the following conditions: wherein if a first bead and a second bead are present in the partition, the bead-specific barcoded oligonucleotides from the first bead and the second bead are each separately extended using the same cDNA hybrid molecular fragment as a template to form (i) an amplicon comprising a first bead-specific barcode and a first UMI sequence and (ii) an amplicon comprising a second bead-specific barcode and a first UMI sequence; Sequence the amplified nucleotide sequencing amplicons to generate sequencing reads; and Sort the sequencing reads from different partitions, wherein (i) the amplicon comprising the first bead-specific barcode and the first UMI sequence and (ii) the amplicon comprising the second bead-specific barcode and the first UMI sequence, and (iii) optionally the same fragment breakpoints are from the same partition.

[0004] In some embodiments, the hybrid molecule is an RNA / cDNA hybrid molecule, and the cDNA is first-strand cDNA. In some embodiments, the RNA / first-strand cDNA hybrid molecule is formed by reverse transcribing RNA from cells using a poly-A, random, or gene-specific reverse transcription primer. In some embodiments, the hybrid molecule is a DNA / cDNA hybrid molecule. In some embodiments, the DNA / first-strand cDNA hybrid molecule is formed by polymerase chain reaction or primer extension.

[0005] In some embodiments, the first adapter oligonucleotide comprises a UMI sequence. In some embodiments, the second adapter oligonucleotide comprises a UMI sequence. In some embodiments, both the first adapter oligonucleotide and the second adapter oligonucleotide comprise a UMI sequence.

[0006] In some embodiments, the first adapter oligonucleotide, the second adapter oligonucleotide, or both further comprise a sample barcode sequence.

[0007] In some embodiments, generation includes contacting the RNA / cDNA hybrid molecule with a transposase that introduces an adapter oligonucleotide into the RNA / cDNA hybrid molecule.

[0008] In some embodiments, the bead-specific barcoded oligonucleotide comprises a 3' end comprising a first universal sequence, and amplification includes annealing and extending the bead-specific barcoded oligonucleotide to the reverse complementary sequence of the first universal sequence. In some embodiments, amplification further includes using the amplicon as a template to extend a reverse primer having a 3' end comprising the reverse complementary sequence of a second universal sequence.

[0009] In some embodiments, the bead-specific barcoded oligonucleotide comprises a 3' end comprising a second universal sequence, and amplification includes annealing and extending the bead-specific barcoded oligonucleotide to the reverse complementary sequence of the second universal sequence. In some embodiments, amplification further includes using the amplicon as a template to extend a reverse primer having a 3' end comprising the first universal sequence.

[0010] In some embodiments, the partitioning is into droplets or microwells in an emulsion.

[0011] In some embodiments, the cell is a mammalian cell. In some embodiments, the cell is a prokaryotic cell. In some embodiments, the cell is a eukaryotic cell.

[0012] In some embodiments, a plurality of partitions as described herein are provided. In some embodiments, at least some of the partitions comprise: Fixed and permeabilized cells that contain crosslinked molecules and contain A gap-filled hybrid molecular fragment formed by: Random breaks are generated in an RNA / cDNA or DNA / cDNA hybrid molecule, and a first adapter oligonucleotide or a second adapter oligonucleotide is randomly inserted at the break, thereby forming a hybrid molecular fragment, the hybrid molecular fragment comprising (i) a first 5' end linked to the first adapter oligonucleotide and (ii) a first 3' end and (iii) a second 5' end linked to the second adapter oligonucleotide and (iv) a second 3' end, wherein the first adapter oligonucleotide comprises a first universal sequence, and the second adapter oligonucleotide comprises a second universal sequence, and wherein the first adapter oligonucleotide, the second adapter oligonucleotide, or both further comprise a unique molecular identifier (UMI) sequence; The cells are partitioned into partitions together with one or more beads, wherein each bead is linked to multiple copies of a bead-specific barcoded oligonucleotide having the same 3' end comprising the first universal sequence or the second universal sequence, wherein the bead-specific barcoded oligonucleotides linked to different beads can be identified by unique bead-specific barcodes in the bead-specific barcoded oligonucleotides, and wherein at least some of the partitions contain at least two different beads; and Using a polymerase (i) to extend the first 3' end using the second adapter oligonucleotide as a template such that the first 3' end comprises the reverse complementary sequence of the second universal sequence, and (ii) to extend the second 3' end using the first adapter oligonucleotide as a template such that the second 3' end comprises the reverse complementary sequence of the first universal sequence, thereby forming a gap-filled hybrid molecular fragment; Wherein said at least some of the partitions further comprise said one or more beads.

[0013] In some embodiments, the partition is a droplet or a microwell in an emulsion.

[0014] In some embodiments, the cell is a mammalian cell. In some embodiments, the cell is a prokaryotic cell. In some embodiments, the cell is a eukaryotic cell.

[0015] In some embodiments, the first adapter oligonucleotide comprises a UMI sequence. In some embodiments, the second adapter oligonucleotide comprises a UMI sequence. In some embodiments, both the first adapter oligonucleotide and the second adapter oligonucleotide comprise a UMI sequence.

[0016] In some embodiments, the first adapter oligonucleotide, the second adapter oligonucleotide, or both further comprise a sample barcode sequence. Brief Description of the Drawings

[0017] Figure 1 Illustrates exemplary steps in the method described herein. In step 1, the cells are fixed and permeabilized in bulk (i.e., not in partitions). In step 2, polyA and non-polyA RNAs are reverse transcribed in situ within the bulk-permeabilized cells. In step 3, the reverse-transcribed RNA:DNA hybrid molecules are tagged with transposase labeled with UMI and sample indices (i.e., the transposase introduces breaks and inserts adapter oligonucleotides at the breaks). The UMI is used in bead pooling; the sample index is used for combinatorial indexing to increase cell throughput. In step 4, the crosslinking of the tagged cells is reversed either before or after partitioning into bead-overloaded droplets.

[0018] Figure 2 Illustrates Figure 1 a continuation of the exemplary steps. In step 4, the crosslinking of the tagged cells is reversed (the same as the last step of Figure 1 ) either before or after partitioning into bead-overloaded droplets. In step 5, after the in-droplet cell lysis step, the tagged RNA:DNA hybrids are gap-filled to generate cDNA fragments with adapter sequences at both ends. In step 6, the (transcriptome / gene-specific) fragments of the droplet partitioning are labeled with a set of unique UMIs and subsequently with unique cell barcodes (CBCs) within the same partition. In step 7, after the droplet PCR process, the fragments associated with a set of unique droplet-specific UMIs are ligated to a set of unique cell barcodes within the same droplet.

[0019] Figure 3 Illustrates Figure 2 a continuation of the exemplary steps. In step 8, bead pooling is performed by calculating the Jaccard similarity between library fragments consisting of UMIs and barcodes. The set of barcodes associated with the same UMI set is pooled and interpreted as a unique partition and thus from a single cell.

[0020] Figure 4 Illustrates one aspect of the method described herein, where the starting nucleic acid is an RNA / cDNA hybrid that has been treated with tagged transposase to fragment the hybrid and introduce adapter oligonucleotides into the resulting fragments.

[0021] Figure 5 Illustrates exemplary adapter sequences.

[0022] Figure 6 Depicts a flow chart as detailed in Example 1.

[0023] Figure 7 Illustrates a sorted plot of the detected partitions (after barcode pooling) in descending order of read counts, as explained in Example 2.

[0024] Figure 8 A graph showing the total unique read counts for each detected partition having various numbers of detected barcodes as described in Example 2. Definition

[0025] Unless otherwise defined, all technical and scientific terms used herein generally have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. Generally, the nomenclature used herein and the laboratory procedures in cell culture, molecular genetics, organic chemistry, and nucleic acid chemistry and hybridization described below are those well known and commonly employed in the art. Standard techniques are used for nucleic acid and peptide synthesis. Techniques and procedures are generally performed according to conventional methods in the art and various general references (see generally, Sambrook et al., MOLECULAR CLONING: A LABORATORY MANUAL, 2nd ed. (1989) Cold Spring Harbor Laboratory Press, Cold Spring Harbor, N.Y.), which are incorporated herein by reference), and such conventional methods and various general references are provided throughout this application. The nomenclature used herein and the laboratory procedures in analytical chemistry and organic synthesis described below are well known and commonly used in the art.

[0026] The term "amplification reaction" refers to any in vitro means for amplifying copies of a target sequence of nucleic acid in a linear or exponential manner. Such methods include, but are not limited to, polymerase chain reaction (PCR); DNA ligase chain reaction (see U.S. Pat. Nos. 4,683,195 and 4,683,202; PCR Protocols: A Guide to Methods and Applications (Innis et al., eds., 1990)) (LCR); Qβ RNA replicase and RNA transcription-based amplification reactions (e.g., amplification involving T7, T3, or SP6 primed RNA polymerases), such as transcription-based amplification system (TAS), nucleic acid sequence-based amplification (NASBA), and self-sustained sequence replication (3SR); isothermal amplification reactions (e.g., single primer isothermal amplification (SPIA)); and other techniques known to those of skill in the art.

[0027] "Amplification" refers to the step of placing a solution under conditions sufficient to amplify a polynucleotide if all components of the reaction are intact. Components of an amplification reaction include, for example, primers, polynucleotide templates, polymerases, nucleotides, etc. The term "amplification" generally refers to an "exponential" increase in the target nucleic acid. However, as used herein, "amplification" can also refer to a linear increase in the number of selected target sequences of a nucleic acid, such as obtained by cycle sequencing or linear amplification. In an exemplary embodiment, amplification refers to PCR amplification using a first amplification primer and a second amplification primer.

[0028] The term "amplification reaction mixture" refers to an aqueous solution containing various reagents for amplifying a target nucleic acid. These amplification reaction mixtures include enzymes, aqueous buffers, salts, amplification primers, target nucleic acids, and nucleoside triphosphates. The amplification reaction mixture can further contain stabilizers and other additives to optimize efficiency and specificity. Depending on the context, the mixture can be a complete or incomplete amplification reaction mixture.

[0029] "Polymerase chain reaction" or "PCR" refers to a method of amplifying a specific segment or subsequence of a target double-stranded DNA in geometric progression. PCR is well known to those skilled in the art; see, for example, U.S. Patent Nos. 4,683,195 and 4,683,202; and "PCR Protocols: A Guide to Methods and Applications", edited by Innis et al., 1990. Exemplary PCR reaction conditions typically include two-step or three-step cycles. The two-step cycle has a denaturation step followed by a hybridization / extension step. The three-step cycle includes a denaturation step followed by a hybridization step followed by a separate extension step.

[0030] A "primer" refers to a polynucleotide sequence that hybridizes to a sequence on a target nucleic acid and serves as a starting point for nucleic acid synthesis. Primers can have various lengths, and the length is generally less than 50 nucleotides, for example, 12 to 30 nucleotides in length. The length and sequence of primers for PCR can be designed based on principles known to those skilled in the art; see, for example, Innis et al., supra. Primers can be DNA, RNA, or chimeras of DNA and RNA moieties. In some cases, primers can include one or more modified or non-natural nucleobases. In some cases, primers are labeled.

[0031] "Primer extension" refers to any method in which a primer is extended in a template-specific manner. Examples of primer extension include, for example, methods in which a primer hybridizes to a template nucleic acid and a polymerase extends the primer in a template-specific manner. In some embodiments, the template is DNA and the polymerase is a DNA polymerase. In some embodiments, the template is RNA and the polymerase is a reverse transcriptase. Primer extension can also include, for example, template switching (see, for example, Zhu YY, Machleder EM, et al., (2001) Biotechniques, 30(4): 892–897; Ramskold D, Luo S, et al., (2012) Nat Biotechnol, 30(8): 777–78) and nick polymerization (also known as nick translation), the latter involving nicking one strand of a nucleic acid duplex and using the nicked strand as a primer that extends using the other strand as a template (see, for example, Dr. Leonard G. Davis et al., in Basic Methods in Molecular Biology, 1986).

[0032] A nucleic acid or a portion thereof "hybridizes" to another nucleic acid under conditions such that non-specific hybridization at a defined temperature in a physiological buffer (e.g., pH 6–9, 25–150 mM chloride salt) is minimized. In some cases, the nucleic acid or a portion thereof hybridizes to a conserved sequence common to a set of target nucleic acids. In some cases, a primer or a portion thereof can hybridize to a primer binding site if there are at least about 6, 8, 10, 12, 14, 16, or 18 consecutive complementary nucleotides, including "universal" nucleotides that are complementary to more than one nucleotide partner. Alternatively, a primer or a portion thereof can hybridize to a primer binding site if there are less than 1 or 2 complementary mismatches over at least about 12, 14, 16, or 18 consecutive complementary nucleotides. In some embodiments, the defined temperature at which specific hybridization occurs is room temperature. In some embodiments, the defined temperature at which specific hybridization occurs is above room temperature. In some embodiments, the defined temperature at which specific hybridization occurs is at least about 37°C, 40°C, 42°C, 45°C, 50°C, 55°C, 60°C, 65°C, 70°C, 75°C, or 80°C. In some embodiments, the defined temperature at which specific hybridization occurs is 37°C, 40°C, 42°C, 45°C, 50°C, 55°C, 60°C, 65°C, 70°C, 75°C, or 80°C.

[0033] "Template" refers to a polynucleotide sequence of a polynucleotide to be amplified that includes a side primer hybridization site or a pair of primer hybridization sites. Thus, "target template" includes a target polynucleotide sequence adjacent to at least one hybridization site of a primer. In some cases, "target template" includes a target polynucleotide sequence that flanks the hybridization sites of a "forward" primer and a "reverse" primer.

[0034] As used herein, "nucleic acid" means DNA, RNA, single-stranded, double-stranded or more highly aggregated hybridization motifs, and any chemical modifications thereof. Modifications include, but are not limited to, modifications that provide chemical groups that add additional charges, polarities, hydrogen bonds, electrostatic interactions, attachment points, and functionality to the nucleic acid ligand bases or to the entire nucleic acid ligand. Such modifications include, but are not limited to, peptide nucleic acids (PNAs), phosphodiester group modifications (e.g., phosphorothioates, methylphosphonates), 2'-sugar modifications, 5-position pyrimidine modifications, 8-position purine modifications, modifications at exocyclic amines, substitution of 4-thiouridine, substitution of 5-bromo or 5-iodo-uracil; backbone modifications, methylation, unusual base pairing combinations such as isobases, isocytidine, and isoguanine, etc. Nucleic acids can also include unnatural bases, such as nitroindole. Modifications can also include 3' and 5' modifications that include, but are not limited to, capping with a fluorophore (e.g., quantum dots) or another moiety.

[0035] "Polymerase" refers to an enzyme that performs template-directed synthesis of polynucleotides (e.g., DNA and / or RNA). The term encompasses both full-length polypeptides and domains having polymerase activity. DNA polymerases are well known to those of skill in the art and include, but are not limited to, DNA polymerases isolated or derived from Pyrococcus furiosus, Thermococcus litoralis, and Thermotoga maritime, or modified versions thereof. Additional examples of commercially available polymerases include, but are not limited to: Klenow fragment (New England Biolabs), Taq DNA polymerase (QIAGEN), 9°N TM DNA polymerase (NewEngland Biolabs), Deep Vent TM DNA polymerase (New England Biolabs), Manta DNA polymerase Bst DNA polymerase (New England Biolabs) and phi29 DNA polymerase (New England Co., Ltd.)

[0036] Polymerases include both DNA-dependent polymerases and RNA-dependent polymerases, such as reverse transcriptase. At least five families of DNA-dependent DNA polymerases are known, although most belong to families A, B, and C. Other types of DNA polymerases include phage polymerases. Similarly, RNA polymerases generally include eukaryotic RNA polymerases I, II, and III, and bacterial RNA polymerases, as well as phage and viral polymerases. RNA polymerases can be DNA-dependent and RNA-dependent.

[0037] As used herein, the term "partition" or "partitioned" refers to separating a sample into multiple parts or "compartments". Compartments are typically physical, such that the sample in one compartment does not or substantially does not mix with the sample in an adjacent compartment. Compartments can be solid or fluid. In some embodiments, the compartment is a solid compartment, e.g., a microchannel. In some embodiments, the compartment is a fluid compartment, e.g., a droplet. In some embodiments, the fluid compartment (e.g., a droplet) is a hybrid mixture of immiscible fluids (e.g., water and oil). In some embodiments, the fluid compartment (e.g., a droplet) is an aqueous droplet surrounded by an immiscible carrier fluid (e.g., oil).

[0038] As used herein, a "barcode" is a short nucleotide sequence (e.g., at least about 4, 6, 8, 10, or 12 nucleotides in length) that identifies the molecule to which it is conjugated. Barcodes can be used, for example, to identify molecules in partitions. A partition-specific barcode should be unique to the partition as compared to barcodes present in other partitions. For example, a partition containing target RNA from a single cell can undergo reverse transcription conditions in which primers containing different partition-specific barcode sequences in each partition are used, thereby incorporating copies of unique "cell barcodes" into the reverse-transcribed nucleic acids in each partition. Thus, due to the unique "cell barcodes," nucleic acids from each cell can be distinguished from nucleic acids from other cells. In some cases, the cell barcodes are provided as "bead barcodes" that are present on oligonucleotides conjugated to beads, where the bead barcodes are shared (e.g., identical or substantially identical) by all or substantially all of the oligonucleotides conjugated to the beads. Thus, cell barcodes and bead barcodes can be present in a partition, attached to beads, or bound to cell nucleic acids as multiple copies of the same barcode sequence. Cell barcodes or bead barcodes of the same sequence can be identified as originating from the same cell, partition, or bead. Such partition-specific barcodes, cell barcodes, or bead barcodes can be generated using a variety of methods that can produce barcodes conjugated to or incorporated into a solid support or hydrogel support (e.g., solid beads or particles or hydrogel beads or particles). In some cases, partition-specific barcodes, cell barcodes, or bead barcodes are generated using a split-and-pool (also known as split-and-mix) synthesis scheme. A partition-specific barcode can be a cell barcode and / or a bead barcode. Similarly, a cell barcode can be a partition-specific barcode and / or a bead barcode. Additionally, a bead barcode can be a cell barcode and / or a partition-specific barcode.

[0039] In other cases, the barcode uniquely identifies the molecule to which it is conjugated and is referred to as a unique molecular identifier (UMI). The number of nucleotides that can be a continuous or discontinuous UMI will depend on the number of UMI sequences desired. In some embodiments, the number of available UMIs is many-fold higher than the number of possible conjugation partners (e.g., 2X, 10X, 100X, etc.), thereby reducing the likelihood of rare repeats ligating to different molecules. In some embodiments, a set of different UMIs is present in a partition, and the composition of the set serves as an identifier for the partition, where some UMIs are common to some other partitions, but the total set of UMIs is unique or substantially unique between partitions. UMI sequences can be generated, for example, as random sequences of a set length and, in some embodiments, are identified by flanking known sequences.

[0040] The length of the barcode sequence can determine how many unique samples can be distinguished. For example, a 1-nucleotide barcode can distinguish 4 or fewer different samples or molecules; a 4-nucleotide barcode can distinguish 4 4 or 256 samples or fewer; a 6-nucleotide barcode can distinguish 4096 different samples or fewer; and an 8-nucleotide barcode can index 65,536 different samples or fewer. Additionally, barcodes can be attached to both strands via barcoded primers used for first-strand and second-strand synthesis, by ligation, or in a tagging reaction.

[0041] Barcodes are typically synthesized and / or polymerized (e.g., amplified) using inherently imprecise methods. Thus, barcodes that are intended to be uniform (e.g., cell-, particle-, or partition-specific barcodes shared among all barcoded nucleic acids in a single partition, cell, or bead) can contain various N-1 deletions or other mutations from the canonical barcode sequence. Accordingly, barcodes referred to as "identical" or "substantially identical" copies are barcodes that differ due to one or more errors such as synthesis, polymerization, or purification errors, and thus contain various N-1 deletions or other mutations from the canonical barcode sequence. In addition, the use of, for example, split-and-pool methods and / or random coupling of barcode nucleotides with an equal mixture of nucleotide precursor molecules as described herein during synthesis can result in low-probability events where the barcode is not absolutely unique (e.g., different from all other barcodes in the population, or different from barcodes in different partitions, cells, or beads). However, such minor variations from the theoretically ideal barcode do not interfere with the high-throughput sequencing analysis methods, compositions, and kits described herein. Thus, as used herein, the term "unique" in the context of particle-, cell-, partition-specific, or molecular barcodes encompasses various inadvertent N-1 deletions and mutations from the ideal barcode sequence. In some cases, problems due to the imprecise nature of barcode synthesis, polymerization, and / or amplification are overcome by oversampling the possible barcode sequences compared to the number of barcode sequences to be distinguished (e.g., at least about 2-fold, 5-fold, 10-fold, or more possible barcode sequences). For example, cell barcodes with 9 barcode nucleotides can be used to analyze 10,000 cells, representing 262,144 possible barcode sequences. The use of barcode technology is well known in the art, see, for example, Katsuyuki Shiroguchi et al., Proc Natl Acad Sci U S A., January 24, 2012; 109(4):1347-52; and Smith, AM et al., Nucleic Acids Research Can 11, (2010). Other methods and compositions using barcode technology include those described in US2016 / 0060621.

[0042] "Transposase" or "Tagmentase" means an enzyme that can form a functional complex with a composition containing transposon ends and catalyze the insertion or transposition of the composition containing transposon ends into double-stranded target DNA incubated with it in an in vitro transposition reaction.

[0043] The term "transposon end" means double-stranded DNA that only presents a nucleotide sequence ("transposon end sequence"), which is necessary for forming a complex with a transposase functioning in an in vitro transposition reaction. The transposon end forms a "complex" or "associated complex" or "transpososome complex" or "transpososome composition" with a transposase or integrase that recognizes and binds to the transposon end, and the complex can insert or transpose the transposon end into the target DNA incubated with it in an in vitro transposition reaction. The transposon ends exhibit two complementary sequences consisting of a "transferred transposon end sequence" or "transferred strand" and an "untransferred transposon end sequence" or "untransferred strand". For example, a transposon end that forms a complex with a highly active Tn5 transposase (e.g., EZ-Tn5 TM transposase, EPICENTRE Biotechnologies, Madison, Wis., USA), which is active in an in vitro transposition reaction) includes a transferred strand presenting the following "transferred transposon end sequence": 5′AGATGTGTATAAGAGACAG 3′ (SEQ ID NO:1), and an untransferred strand presenting the following "untransferred transposon end sequence": 5′CTGTCTCTTATACACATCT 3′ (SEQ ID NO:2).

[0044] In an in vitro transposition reaction, the 3' end of the transferred strand is ligated or transferred to the target DNA. In an in vitro transposition reaction, the untransferred strand presenting a transposon end sequence complementary to the transferred transposon end sequence is not ligated or transferred to the target DNA.

[0045] The term "solid support" refers to beads, microtiter wells, or other surfaces that can be used to attach nucleic acids, such as oligonucleotides or polynucleotides. The surface of the solid support can be treated to facilitate the attachment of nucleic acids, such as single-stranded nucleic acids.

[0046] The term "bead" refers to any solid support that can be in a partition, such as a small particle or other solid support. In some embodiments, the beads comprise polyacrylamide. For example, in some embodiments, the beads are chemically modified with acryloyl phosphoramidite monomers attached to each oligonucleotide to incorporate barcoded oligonucleotides into the gel matrix. Exemplary beads can comprise hydrogel beads. In some cases, the hydrogel is in sol form. In some cases, the hydrogel is in gel form. An exemplary hydrogel is an agarose hydrogel. Other hydrogels include, but are not limited to, those described in, for example, U.S. Patent Nos. 4,438,258; 6,534,083; 8,008,476; 8,329,763; U.S. Patent Application Nos. 2002 / 0,009,591; 2013 / 0,022,569; 2013 / 0,034,592; and International Patent Publications WO / 1997 / 030092; and WO / 2001 / 049240.

[0047] It should be understood that any numerical range disclosed herein can include the endpoints of the range, as well as any value or sub-range between the endpoints. For example, the range from 1 to 10 includes the endpoints 1 and 10, as well as any value between 1 and 10. The values typically include one significant digit.

[0048] The term "sample" refers to a biological composition that contains a target nucleic acid, such as a cell.

[0049] The term "deconvolution" refers to the assignment of two barcodes and their attached beads as being from the same partition or having originally occupied the same partition. Deconvolution can be determined by detecting two barcodes on a single nucleic acid fragment during sequencing.

[0050] The term "about" refers to the usual error range of the corresponding value known to a person of ordinary skill in the art. For example, a range of ±10%, ±5%, or ±1% can encompass the value, even if the value is not modified by the term "about".

[0051] All ranges described herein can include the endpoint values of the range, as well as any sub-range that includes values between the endpoints, where the values include the first significant digit. For example, the range from 1 to 10 includes ranges such as 2 to 9, 3 to 8, 4 to 7, 5 to 6, 1 to 5, 2 to 5, 2 to 10, 3 to 10, etc. Detailed Description

[0052] Barcode bead overloading in partitions such as droplets increases droplet utilization such that >90% of the droplets contain at least one solid support (such as a bead) and are active during barcode encoding. However, to prevent fragmentation of substrate (i.e., cell) data and / or overrepresentation of substrate (i.e., cell) when a partition has more than one solid support, the solid supports need to be co-localized to a single partition. When this occurs, the data for each co-localized barcode can be computationally combined to maintain data integrity.

[0053] In general, it may be desirable to use partition-specific barcodes to barcode all target nucleic acids in a partition with the same barcode such that, given different reads with different partition-specific barcodes, the contents of different partitions can be later combined and still traced to their origin in different partitions. Many multiple copies of the partition-specific barcode can be linked to beads and delivered to the partitions. However, the methods of preparing partition-specific barcodes and delivering them to the partitions (e.g., attaching to beads) are generally controlled by a Poisson distribution, which means that some partitions do not receive any partition-specific barcode (e.g., any beads), some partitions receive only one partition-specific barcode (e.g., one bead), and some partitions receive two or more partition-specific barcodes (e.g., two or more beads). The latter creates problems because if two different partition-specific barcodes appear in a partition, sequence reads with different partition-specific barcodes will be misinterpreted as originating from different partitions. As described herein, methods are provided for determining when two partition-specific barcodes originate from the same partition and thus allowing the use of data from more partitions, effectively allowing partitions to be "overloaded" with beads such that the presence of partitions lacking any beads is significantly reduced.

[0054] The problems associated with more than one barcoded bead per partition can be solved by the methods and compositions described herein. The methods described herein involve initiating transposase-mediated fragmentation and adapter insertion in fixed and permeabilized cells to form "tagged" fragments that contain a universal sequence as well as unique molecular identifier (UMI) barcodes in one or two adapter sequences. The permeabilized cells carrying the tagged fragments are partitioned with or added to partitions containing bead-linked oligonucleotides. The tagged fragments in a single partitioned cell are then gap-filled after cell lysis, and the bead-linked oligonucleotides are then ligated to the gap-filled tagged fragments. The resulting product is sequenced in bulk (after partition pooling). Two different bead-linked oligonucleotides in the same partition can be identified by detecting the same UMI sequence in the sequencing reads that are linked to different bead-linked oligonucleotides, indicating that they both likely originated from the same partition because it is unlikely for a UMI to appear in different partitions. In some embodiments, further confirmation can also be made bySame Fragments, i.e., having the same fragmented ends formed by tagging, obtain the confidence that sequencing reads with different bead linkages of oligonucleotide sequences and the same UMI originate from the same partition. Once it is determined that two different bead-linked oligonucleotides occur in the same partition, it can be assumed that all sequencing reads containing the oligonucleotide sequences of either bead linkage are from the same partition, and thus, for example, all sequencing reads from the two bead-linked oligonucleotides can be combined to represent the sequencing results from the same partition.

[0055] The method can involve RNA / first-strand cDNA hybrid molecules or DNA / cDNA hybrid molecules. RNA / first-strand cDNA hybrid molecules can be generated by any means of forming first-strand cDNA, such as using reverse transcriptase to form first-strand cDNA based on an RNA template. DNA / cDNA hybrid molecules can be formed, for example, using a DNA strand as a template for the cDNA strand with a polymerase. Alternatively, any source of double-stranded DNA can be used. In this case, a "hybrid" refers to a double-stranded nucleic acid duplex.

[0056] In the context of the methods described herein, RNA / first-strand cDNA hybrid molecules or DNA / cDNA hybrid molecules are provided in fixed and permeabilized cells. See, for example Figure 1 , item 1. Thus, RNA / first-strand cDNA hybrid molecules or DNA / cDNA hybrid molecules can be formed before or, in many embodiments, after the cells have been fixed and permeabilized to allow reagents (such as polymerases) to diffuse into the cells while substantially retaining the nucleic acids within the fixed and permeabilized cells.

[0057] In some embodiments, the sample comprising the target nucleic acid is a biological sample. The biological sample can be obtained from any biological organism, such as, for example, an animal, a plant, a fungus, a pathogen (such as a bacterium or a virus), or any other organism. In some embodiments, the biological sample is from an animal, such as a mammal (such as a human or a non-human primate, a cow, a horse, a pig, a sheep, a cat, a dog, a mouse, or a rat), a bird (such as a chicken), or a fish. The biological sample can be any tissue or body fluid obtained from a biological organism, such as, for example, blood, blood fractions or blood products (such as serum, plasma, platelets, red blood cells, etc.), sputum or saliva, tissue (such as kidney, lung, liver, heart, brain, nerve tissue, thyroid, eye, skeletal muscle, cartilage, or bone tissue); cultured cells, such as primary cultures, explants, and transformed cells, stem cells, feces, urine, etc. In some embodiments, the sample is a sample comprising cells. In some embodiments, the sample is a single-cell sample.

[0058] The formation of RNA / first-strand cDNA hybrid molecules by reverse transcription in permeabilized cells can be achieved using a poly-T primer, a random primer (e.g., random hexamer), or a gene-specific primer complementary to the target RNA. See, for example Figure 1 Item 2 of. In some embodiments, the 3′ sequence is a poly-T sequence comprising at least five consecutive thymines. In some embodiments, the 3′ sequence is a random sequence of at least 5 (e.g., at least 8, at least 10, at least 12, e.g., 6-30) consecutive nucleotides. In some embodiments, the 3′ sequence is a target gene-specific sequence of at least 5 consecutive nucleotides. The target RNA can be mRNA or non-poly-A RNA. In some embodiments, the first-strand cDNA generated is full-length based on the corresponding RNA from which the cDNA is generated. A variety of reverse transcriptases can be used to form RNA / cDNA hybrids in cells.

[0059] Any method of fixing and permeabilizing cells can be used. For example, in some embodiments, the cells are formalin-fixed paraffin-embedded (FFPE) samples. In fixed and permeabilized cells, reagents can diffuse into the cells, for example to cause reverse transcription, cleavage, annealing, and / or ligation to occur within the fixed cells. Exemplary fixing reagents can include the use of digitonin or fixatives such as methanol (see, for example, Alles, J. et al., BMC Biol 15, 44 (2017)) or paraformaldehyde. Permeabilizing reagents can include, for example, Triton X-100. In some embodiments, the sample comprises a target nucleic acid isolated from tissue or cells. In some embodiments, the cells will have intact chromatin such that some chromosomal regions are more accessible to transposase than other chromosomal regions, allowing the generation of ATACseq results. Exemplary conditions can include those, for example, as described by Nesterenko et al., Proc. Natl. Acad. Sci. USA Vol. 118 No. 3 (2021). As described below, in some embodiments, the fixing reagent is selected such that at least some of the crosslinks that occur during fixation (e.g., protein-protein crosslinks) can be reversed at a later point, for example by heat less than 110°C and / or a reversible reagent, such as a reducing agent, or both.

[0060] Fragmentation and attachment of terminal adapters to RNA / first-strand cDNA hybrid molecules or DNA / cDNA hybrid molecules can be achieved using a transposase. See, for example Figure 1Item 3. The action of a transposase is sometimes referred to as "tagging" and can involve introducing different adaptor sequences on different sides of the DNA breakpoints generated by the transposase, or the added adaptor sequences can be the same. In either case, one or both of the adaptor sequences are common adaptor sequences in that the adaptor sequences are the same with respect to the diversity of the DNA fragments. A tagging enzyme loaded with homotypic adaptors is a tagging enzyme containing only one sequence of adaptors that is added to both ends of the tagging enzyme-induced breakpoint in genomic DNA. A tagging enzyme loaded with heteroadaptors is a tagging enzyme containing two different adaptors such that different adaptor sequences are added to the two DNA ends created by the tagging enzyme-induced breakpoint in DNA. Tagging enzymes loaded with adaptors are further described, for example, in U.S. Patent Publication Nos. 2010 / 0120098; 2012 / 0301925; and 2015 / 0291942; and U.S. Patent Nos. 5,965,443; 6,437,109; 7,083,980; 9,005,935; and 9,238,671, the respective contents of which are incorporated by reference in their entirety for all purposes. The tagging of RNA / DNA hybrids is described, for example, in Bo LuLiting et al., eLife 9:e54919 (2020).

[0061] A tagging enzyme is an enzyme that can form a functional complex with a composition containing a transposon end and catalyze the insertion or transposition of the composition containing the transposon end into double-stranded target DNA incubated with it in an in vitro transposition reaction. Exemplary transposases include, but are not limited to, modified Tn5 transposases with high activity compared to wild-type Tn5, which may have one or more mutations selected from E54K, M56A, or L372P. The wild-type Tn5 transposon is a composite transposon in which two nearly identical insertion sequences (IS50L and IS50R) flank three antibiotic resistance genes (Reznikoff WS Annu Rev Genet 42:269-286 (2008)). Each IS50 contains two inverted 19-bp end sequences (ES), an outer end (OE), and an inner end (IE). However, the wild-type ES has relatively low activity and is replaced in vitro by a highly active mosaic end (ME) sequence. Thus, a complex of the transposase with the 19-bp ME is required for transposition, provided that the intervening DNA is long enough for two of these sequences to come close together to form an active Tn5 transposase homodimer (Reznikoff WS., Mol Microbiol 47:1199-1206 (2003)). Transposition is a very rare event in vivo, and highly active mutants have historically been derived by introducing three missense mutations (E54K, M56A, L372P) into the 476 residues of the Tn5 protein, which is encoded by IS50R (Goryshin IY, Reznikoff WS. 1998. J Biol Chem 273:7367-7374 (1998)). Transposition works by a "cut and paste" mechanism, in which Tn5 excises itself from the donor DNA and inserts into the target sequence, generating a 9-bp repeat of the target (Schaller H. Cold Spring Harb Symp Quant Biol 43:401–408 (1979); Reznikoff WS., Annu Rev Genet 42:269-286 (2008)). In current commercial protocols (Nextera TM DNA kits, Illumina), free synthetic ME adapters are ligated to the 5' end of the target DNA by the transposase (tagging enzyme) ends.

[0062] In some embodiments, the length of the (multiple) adaptors is at least 19 nucleotides, such as 19 - 100 nucleotides. In some embodiments, the adaptor is double-stranded and has a 5' overhang, where the 5' overhang sequence varies between heterologous adaptors while the double-stranded portion (usually 19bp) is the same. In some embodiments, the adaptor comprises TCGTCGGCAGCGTC (SEQ ID NO:1) or GTCTCGTGGGCTCGG (SEQ ID NO:2). In some embodiments of the tagging enzyme loaded with heterologous adaptors, the tagging enzyme is loaded with a first adaptor comprising TCGTCGGCAGCGTC (SEQ ID NO:1) and a second adaptor comprising GTCTCGTGGGCTCGG (SEQ ID NO:2). In some embodiments, the adaptor comprises AGATGTGTATAAGAGACAG (SEQ ID NO:3) and its complementary sequence (this is the mosaic end and this is the only cis-acting sequence essential for Tn5 transposition). In some embodiments, the adaptor comprises TCGTCGGCAGCGTCAGATGTGTATAAGAGACAG (SEQ ID NO:4) having the complementary sequence of AGATGTGTATAAGAGACAG (SEQ ID NO:3) or GTCTCGTGGGCTCGGAGATGTGTATAAGAGACAG (SEQ ID NO:5) having the complementary sequence of AGATGTGTATAAGAGACAG (SEQ ID NO:3). In some embodiments of the tagging enzyme loaded with heterologous adaptors, the tagging enzyme is loaded with a first adaptor comprising TCGTCGGCAGCGTCAGATGTGTATAAGAGACAG (SEQ ID NO:4) having the complementary sequence of AGATGTGTATAAGAGACAG (SEQ ID NO:3) and GTCTCGTGGGCTCGGAGATGTGTATAAGAGACAG (SEQ ID NO:5) having the complementary sequence of AGATGTGTATAAGAGACAG (SEQ ID NO:3). See, for example, Figure 5 。

[0063] Transposases generate random breaks in the hybrid molecule and randomly insert a first adaptor oligonucleotide or a second adaptor oligonucleotide at the break, thereby forming hybrid molecule fragments, the hybrid molecule fragments comprising (i) a first 5' end ligated to the first adaptor oligonucleotide and (ii) a first 3' end and (iii) a second 5' end ligated to the second adaptor oligonucleotide and (iv) a second 3' end, wherein the first adaptor oligonucleotide comprises a first universal sequence, and the second adaptor oligonucleotide comprises a second universal sequence, and wherein the first adaptor oligonucleotide, the second adaptor oligonucleotide, or both further comprise unique molecular identifier (UMI) sequences. See, e.g., Figure 4 In some embodiments, transposase-loaded heterologous adaptors are used such that two different adaptors can be attached to the breaks in the target hybrid nucleic acid. Alternatively, two transposases can be used, where the first transposase is loaded with a first set of adaptor oligonucleotides (e.g., for ease of description, "first adaptor oligonucleotide") and the second transposase is loaded with a second set of adaptor oligonucleotides, ("second adaptor oligonucleotide"), thereby allowing, through random fragmentation by the two transposases, the formation of fragments having the first adaptor oligonucleotide at the first end and the second adaptor oligonucleotide at the second end. An "adaptor oligonucleotide" refers to an oligonucleotide carrying a universal sequence, where the universal sequence is common to the end sequences between different fragments to which the oligonucleotide adaptor is attached and allows them to be used as PCR handle sequences, thereby allowing a pair of universal primers to amplify different fragments having the universal sequence.

[0064] The product of transposase-based fragmentation is a hybrid molecule fragment (a "hybrid" herein means a double-stranded nucleic acid) comprising: in the first strand, (i) a first 5' end ligated to the first adaptor oligonucleotide and (ii) a first 3' end, and in the second strand, (iii) a second 5' end ligated to the second adaptor oligonucleotide and (iv) a second 3' end. Because at least one and optionally both of the first adaptor oligonucleotide and the second adaptor oligonucleotide comprise UMI sequences, at least some hybrid molecule fragments comprising UMI sequences will be produced. The first adaptor oligonucleotide and the second adaptor oligonucleotide comprise universal sequences and optionally UMI sequences. The universal sequence can be of any length as needed and will typically be used as a PCR handle sequence, thereby allowing subsequent amplification of hybrid molecule fragments having the appropriate adaptor sequences using universal primers that hybridize to the universal sequence. In some embodiments, the universal sequence is 4-20 nucleotides in length, e.g., 6-12 nucleotides in length. The universal sequence in each first adaptor oligonucleotide will be the same; similarly, the universal sequence in each second adaptor oligonucleotide will be the same. But in some embodiments, the universal sequences of the first adaptor oligonucleotide and the second adaptor oligonucleotide will be different.

[0065] In contrast, the UMI sequences will ideally be different for each adapter oligonucleotide. The number of nucleotides in the UMI sequence will depend on the number of fragments used. Ideally, the number of unique UMI sequences will be greater than the number of fragments (e.g., 2x, 10x, 50, etc.). The UMI sequences can be contiguous or formed from non-contiguous nucleotides. Exemplary UMI sequences are, for example, 5-20 nucleotides in length, such as 8-16 nucleotides in length.

[0066] After transposition (tagging), permeabilized cells containing transposase-treated nucleic acids having adapter sequences at their ends can be added to partitions, or partitioning can be performed in the presence of permeabilized cells. Ideally, most of the cells in a partition are the only cells in the partition, and this can be achieved, for example, by allowing a certain number of empty (cell-free) partitions according to the Poisson distribution. Based on the Poisson distribution, the probability of having single-cell partitions can be calculated based on the ratio of cells to available partitions. In some embodiments, a ratio of 1 cell to 10 partitions is used, e.g., 0.5-2 cells per 10 partitions.

[0067] In some embodiments, in the partition, cell fixation can be at least partially reversed, thus allowing improved accessibility of the tagged nucleic acids for priming, and / or the cells can be lysed, thus enhancing the capture efficiency of the barcodes. Preferably, the reversal of cell crosslinking occurs in the partition such that the cell contents remain in the partition for barcoding. The reversal of fixation can occur before, during, or after gap filling. The reversal of fixation can involve, but is not limited to, heating (e.g., 95-98 °C), applying a protease to degrade crosslinked proteins, adding a reducing agent, or a combination thereof. See, for example, Namimatsu et al., J Histochem Cytochem. January 2005; 53(1):3-11. Cell lysis can include, for example, contacting the cells with a detergent compatible with the partition, heat, or other conditions to cause lysis.

[0068] In some embodiments, the tagging enzyme is stripped from the nucleic acid before proceeding. For example, in some embodiments, the tagged nucleic acids in permeabilized cells are not gap-filled until the transposase protein bound to them is stripped. In some embodiments, the transposase protein can be stripped from the nucleic acid using, for example, SDS or heating (e.g., to about 80-82 °C), thus exposing the gap / 3' end for polymerization.

[0069] Subsequent gap filling steps can be performed on the hybrid molecular fragments formed by tagging. See, for example Figure 2Item 5. For example, a polymerase, such as a DNA polymerase, can be used to fill any gaps, such as at the ends of hybrid molecular fragments having adaptor sequences at their ends. Generally, in embodiments involving RNA / cDNA hybrids, the extension from RNA is less efficient than the extension of cDNA (using RNA as a template). Thus, the cDNA strand is the backbone amplified in downstream reactions, as for example Figure 4 shown by the bold dashed line in

[0070] Methods and compositions for partitioning are described, for example, in the published patent applications WO 2010 / 036,352, US2010 / 0173,394, US2011 / 0092,373, and US2011 / 0092,376, the contents of which are incorporated herein by reference in their entirety. The multiple mixture partitions can be located in multiple emulsion droplets, or multiple microwells, etc.

[0071] In some embodiments, the primers and other reagents can be partitioned into multiple mixture partitions, and then the ligated DNA segments can be introduced into the multiple mixture partitions. Methods and compositions for delivering reagents to one or more mixture partitions include microfluidic methods known to those skilled in the art; the merging, coalescing, fusing, rupturing, or degradation of droplets or microcapsules (e.g., see U.S.2015 / 0027,892; US2014 / 0227,684; WO 2012 / 149,042; and WO 2014 / 028,537); droplet injection methods (e.g., see WO 2010 / 151,776); and combinations thereof.

[0072] As described herein, the mixture partitions can be picoliter wells, nanoliter wells, or microwells. The mixture partitions can be picoliter, nanoliter, or microreaction chambers, such as picocapsules, nanocapsules, or microcapsules. The mixture partitions can be picoliter, nanoliter, or microchannels. The mixture partitions can be droplets, such as emulsion droplets.

[0073] In some embodiments, the partition is a droplet. In some embodiments, the droplet comprises an emulsion composition, i.e., a mixture of immiscible liquids (such as water and oil). In some embodiments, the droplet is an aqueous-phase droplet surrounded by an immiscible carrier liquid (such as oil). In some embodiments, the droplet is an oil-phase droplet surrounded by an immiscible carrier liquid (such as an aqueous solution). In some embodiments, the droplets described herein have high stability and there is very little coalescence between two or more droplets. In some embodiments, among the droplets generated from a sample, less than 0.0001%, 0.0005%, 0.001%, 0.005%, 0.01%, 0.05%, 0.1%, 0.5%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9% or 10% of the droplets coalesce with other droplets. The emulsion may also have limited flocculation, i.e., the process by which the dispersed phase precipitates in the form of flakes. In some cases, the stability or very little coalescence can be maintained for up to 4, 6, 8, 10, 12, 24 or 48 hours or longer (such as at room temperature or about 0, 2, 4, 6, 8, 10 or 12 °C). In some embodiments, the droplets are formed by flowing an oil phase through an aqueous sample or reagent.

[0074] The oil phase may comprise a fluorinated base oil, which can be combined with a fluorinated surfactant (such as perfluoropolyether) to achieve stability. In some embodiments, the base oil comprises one or more of HFE 7500, FC-40, FC-43, FC-70 or other common fluorinated oils. In some embodiments, the oil phase comprises an anionic fluorosurfactant. In some embodiments, the anionic fluorosurfactant is Ammonium Krytox (Krytox-AS), the ammonium salt of Krytox FSH or the morpholine derivative of Krytox FSH. Krytox-AS may be present at a concentration of about 0.1%, 0.2%, 0.3%, 0.4%, 0.5%, 0.6%, 0.7%, 0.8%, 0.9%, 1.0%, 2.0%, 3.0% or 4.0% (w / w). In some embodiments, the concentration of Krytox-AS is about 1.8%. In some embodiments, the concentration of Krytox-AS is about 1.62%. The morpholine derivative of Krytox FSH may be present at a concentration of about 0.1%, 0.2%, 0.3%, 0.4%, 0.5%, 0.6%, 0.7%, 0.8%, 0.9%, 1.0%, 2.0%, 3.0% or 4.0% (w / w). In some embodiments, the concentration of the morpholine derivative of Krytox FSH is about 1.8%. In some embodiments, the concentration of the morpholine derivative of Krytox FSH is about 1.62%.

[0075] In some embodiments, the oil phase further comprises additives for modulating the properties of the oil, such as vapor pressure, viscosity, or surface tension. Non-limiting examples include perfluorooctanol and 1H,1H,2H,2H-perfluorodecanol. In some embodiments, 1H,1H,2H,2H-perfluorodecanol is added at a concentration of about 0.05%, 0.06%, 0.07%, 0.08%, 0.09%, 0.1%, 0.2%, 0.3%, 0.4%, 0.5%, 0.6%, 0.7%, 0.8%, 0.9%, 1.0%, 1.25%, 1.50%, 1.75%, 2.0%, 2.25%, 2.5%, 2.75%, or 3.0% (w / w). In some embodiments, the concentration is about 0.18% (w / w).

[0076] In some embodiments, the emulsion is formulated to form highly monodisperse droplets with a liquid interfacial film that can be converted to microcapsules with a solid interfacial film by heating; these microcapsules can function as bioreactors and retain their contents during incubation. The conversion process can be achieved by heating, for example, at a temperature above about 40°C, 50°C, 60°C, 70°C, 80°C, 90°C, or 95°C. During the heating process, a liquid or mineral oil overlay can be used to prevent evaporation. The excess continuous phase oil can be removed before heating or can be retained. The microcapsules can exhibit anti-fusion and / or anti-flocculation properties under a wide range of thermal processing or mechanical treatment conditions.

[0077] After the droplets are converted to microcapsules, the microcapsules can be stored at conditions of about -70°C, -20°C, 0°C, 3°C, 4°C, 5°C, 6°C, 7°C, 8°C, 9°C, 10°C, 15°C, 20°C, 25°C, 30°C, 35°C, or 40°C. In some embodiments, these capsules are useful for the storage or transportation of mixture partitions. For example, a sample can be collected at one location, partitioned into droplets containing enzymes, buffers, and / or primers or other probes, optionally subjected to one or more polymerization reactions, then the droplets are heated for microencapsulation, and the microcapsules are stored or transported for subsequent analysis.

[0078] In some embodiments, the sample is partitioned into, or at least partitioned into 500, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 10000, 15000, 20000, 30000, 40000, 50000, 60000, 70000, 80000, 90000, 100000, 200000, 300000, 400000, 500000, 600000, 700000, 800000, 900000, 1000000, 2000000, 3000000, 4000000, 5000000, 10000000, 20000000, 30000000, 40000000, 50000000, 60000000, 70000000, 80000000, 90000000, 100000000, 150000000 or 200000000 partitions.

[0079] In some embodiments, the generated droplets are substantially uniform in shape and / or size. For example, in some embodiments, the droplets are substantially consistent in average diameter. In some embodiments, the average diameter of the generated droplets is about 0.001 microns, about 0.005 microns, about 0.01 microns, about 0.05 microns, about 0.1 microns, about 0.5 microns, about 1 micron, about 5 microns, about 10 microns, about 20 microns, about 30 microns, about 40 microns, about 50 microns, about 60 microns, about 70 microns, about 80 microns, about 90 microns, about 100 microns, about 150 microns, about 200 microns, about 300 microns, about 400 microns, about 500 microns, about 600 microns, about 700 microns, about 800 microns, about 900 microns or 1000 microns. In some embodiments, the average diameter of the droplets is less than about 1000 microns, less than about 900 microns, less than about 800 microns, less than about 700 microns, less than about 600 microns, less than about 500 microns, less than about 400 microns, less than about 300 microns, less than about 200 microns, less than about 100 microns, less than about 50 microns or less than about 25 microns. In some embodiments, the generated droplets are non-uniform in shape and / or size.

[0080] In some embodiments, the generated droplets are substantially uniform in volume. For example, the standard deviation of the droplet volume can be less than about 1 picoliter, about 5 picoliters, about 10 picoliters, about 100 picoliters, about 1 nL, or less than about 10 nL. In some cases, the standard deviation of the droplet volume can be less than about 10-25% of the average droplet volume. In some embodiments, the average volume of the generated droplets is about 0.001 nL, about 0.005 nL, about 0.01 nL, about 0.02 nL, about 0.03 nL, about 0.04 nL, about 0.05 nL, about 0.06 nL, about 0.07 nL, about 0.08 nL, about 0.09 nL, about 0.1 nL, about 0.2 nL, about 0.3 nL, about 0.4 nL, about 0.5 nL, about 0.6 nL, about 0.7 nL, about 0.8 nL, about 0.9 nL, about 1 nL, about 1.5 nL, about 2 nL, about 2.5 nL, about 3 nL, about 3.5 nL, about 4 nL, about 4.5 nL, about 5 nL, about 5.5 nL, about 6 nL, about 6.5 nL, about 7 nL, about 7.5 nL, about 8 nL, about 8.5 nL, about 9 nL, about 9.5 nL, about 10 nL, about 11 nL, about 12 nL, about 13 nL, about 14 nL, about 15 nL, about 16 nL, about 17 nL, about 18 nL, about 19 nL, about 20 nL, about 25 nL, about 30 nL, about 35 nL, about 40 nL, about 45 nL, or about 50 nL.

[0081] Also included in the partition (resulting from the partitioning step or by adding to an already formed partition) are one or more beads, wherein the beads are linked to copies of an oligonucleotide that enables bead-specific (and to some extent partition-specific) barcoding. See Figure 2 , item 4. The beads can be linked to multiple copies of the same oligonucleotide, for example, at least about 10, 50, 100, 500, 1000, 5000, 10,000, 50,000, 100,000, 500,000, 1,000,000, 5,000,000, 10,000,000, 10 8 、10 9 、10 10 or more copies of the same or substantially the same oligonucleotide can be linked to one (e.g., the same) bead.

[0082] Since the beads enter the partitions in a Poisson distribution, at least some partitions contain at least two different beads. See Figure 1 or Figure 2, Item 4. During the partitioning process, affected by Poisson distribution statistics, a partition may encapsulate multiple barcode beads. In fact, since the present method allows detection and resolution of multiple bead-specific barcodes present in a partition, in the method described herein, the number of beads delivered into a partition can be set such that a large number of partitions contain multiple beads. This is advantageous because if a method can only accurately utilize partitions containing only one bead, there will be a relatively large number of empty partitions; while if data of multiple beads in a partition can be resolved, the number of partitions containing at least one bead can be increased. Therefore, in some embodiments, conditions can be selected such that at least 10%, 20%, 30%, 40%, 50% or more partitions contain at least one, or contain more than one bead.

[0083] Each oligonucleotide can be linked to the bead at its 5'-end or other positions of the oligonucleotide. In some embodiments, a cleavable group can be included to remove the oligonucleotide from the bead, for example, before the oligonucleotide is used for primer extension. In some embodiments, the cleavable linker contains a site incorporating uridine, located in a portion of the nucleotide sequence. The site incorporating uridine can be cleaved by using uracil glycosylase (such as uracil N-glycosylase or uracil DNA glycosylase (UDG)). In some embodiments, the cleavable linker contains a photocleavable nucleotide. Photocleavable nucleotides include, for example, photocleavable fluorescent nucleotides and photocleavable biotinylated nucleotides. See Li et al., PNAS, 2003, 100:414-419; Luo et al., Methods Enzymol, 2014, 549:115-131. In some cases, the oligonucleotide is linked to the bead via a disulfide bond (such as a disulfide bond formed between sulfur on the solid support and sulfur covalently linked to the 5' or 3'-end or the middle nucleic acid of the oligonucleotide). In such cases, the oligonucleotide can be cleaved from the solid support by contacting the solid support with a reducing agent (such as a thiol or phosphine reagent, including but not limited to β-mercaptoethanol, dithiothreitol (DTT) or tris(2-carboxyethyl)phosphine (TCEP)).

[0084] The oligonucleotide from the bead will include, for example, a bead-specific barcode such that the bead-specific barcode sequence on the first oligonucleotide can be used to distinguish it from the bead-specific barcode of the second oligonucleotide on another bead. The 3'-end of the oligonucleotide will contain the reverse complement of a universal linker sequence, which is one of the linker sequences added to the fragment, so that the oligonucleotide from the bead can be used as a primer for a primer extension reaction (such as PCR), which uses a single-stranded gap-filled hybrid molecular fragment with the linker sequence as a template. See Figure 2, item 6, wherein the bead - specific barcode is labeled as "CBC". For example, in some embodiments, DNA polymerase and reagents that permit polymerase activity (such as salts) are included in the partitions to effect the above - mentioned process. Thus, the resulting template will be ligated to the bead - specific barcode from the oligonucleotide within the partition. Optionally, a reverse primer may be included in the partition. The reverse primer may hybridize to the second universal sequence at the other end of the gap - filled hybrid molecule fragment or its reverse - complementary sequence. See Figure 2 , item 6, "index adaptor", see also Figure 4 the RNA / cDNA - related content in

[0085] After primer extension using the bead - specific barcode oligonucleotide, the gap - filled hybrid molecule fragment will also include the bead - specific barcode. See Figure 2 , item 5. Since some partitions contain two or more beads and thus contain bead - specific barcode oligonucleotides with different sequences, different extension products (such as amplification products) will have different bead - specific barcodes. See Figure 2 , item 7. The amplification product will contain, for example: a bead - specific barcode, a first universal end sequence (or its complementary sequence), a cDNA fragment, a second universal end sequence (or its complementary sequence), and a UMI sequence (or its complementary sequence).

[0086] Once the above - mentioned amplification products are formed, the contents of each partition can be mixed to form an overall solution containing the contents of multiple partitions. The amplification products in the overall solution can be subjected to nucleotide sequencing. Any nucleotide sequencing method can be employed as needed to obtain sequencing reads that contain the bead - specific barcode, the first universal end sequence (or its complementary sequence), the cDNA fragment, the second universal end sequence (or its complementary sequence), and the UMI sequence (or its complementary sequence). A variety of high - throughput sequencing and genotyping methods are known in the art. For example, such sequencing techniques include, but are not limited to, pyrosequencing, ligation - based sequencing, single - molecule sequencing, sequencing - by - synthesis (SBS), massively parallel cloning sequencing, massively parallel single - molecule SBS, massively parallel single - molecule real - time sequencing, massively parallel single - molecule real - time nanopore sequencing technology, etc. Some of these techniques were reviewed by Morozova and Marra in Genomics, 92:255 (2008), the content of which is incorporated herein by reference in its entirety.

[0087] Exemplary DNA sequencing techniques include fluorescence-based sequencing methods (see Birren et al., Genome Analysis: Analyzing DNA, Volume 1, Cold Spring Harbor Press, New York; the contents of which are incorporated herein by reference in their entirety). In some embodiments, automated sequencing techniques understood in the art are used. In some embodiments, the present technology provides parallel sequencing of partitioned amplification products (see PCT Publication No. WO 2006 / 084132, the contents of which are incorporated herein by reference in their entirety). In some embodiments, DNA sequencing is achieved by parallel oligonucleotide extension (see U.S. Patent Nos. US 5,750,341 and US 6,306,597, both of which are incorporated herein by reference in their entirety). Other examples of sequencing techniques include Church's polony technology (Mitra et al., 2003, Analytical Biochemistry 320:55-65; Shendure et al., 2005, Science 309:1728-1732; and U.S. Patent Nos. US 6,432,360; US 6,485,944; US 6,511,803, all of which are incorporated herein by reference in their entirety), 454 picotiter pyrosequencing technology (Margulies et al., 2005, Nature 437:376-380; U.S. Publication No. US2005 / 0130173, both of which are incorporated herein by reference in their entirety), Solexa single-base addition technology (Bennett et al., 2005, Pharmacogenomics 6:373-382; U.S. Patent Nos. US 6,787,308 and US 6,833,246, both of which are incorporated herein by reference in their entirety), Lynx massively parallel signature sequencing technology (Brenner et al., 2000, Nat. Biotechnol. 18:630-634; U.S. Patent Nos. US 5,695,934 and US 5,714,330, both of which are incorporated herein by reference in their entirety), and Adessi's PCR colony technology (Adessi et al., 2000, Nucleic Acid Res. 28:E87; WO 2000 / 018957, the contents of which are incorporated herein by reference in their entirety).

[0088] High-throughput sequencing methods generally share the common features of large-scale parallelism and high throughput, with the goal of reducing costs compared to traditional sequencing methods (see Voelkerding et al., Clinical Chem., 55:641-658, 2009; MacLean et al., Nature Rev. Microbiol., 7:287-296; each incorporated herein by reference in its entirety). Such methods can be broadly divided into methods that generally require template amplification and methods that do not require template amplification. Methods that require amplification include pyrosequencing commercialized by Roche (as the 454 technology platform, such as GS20 and GS FLX), the Solexa platform commercialized by Illumina, and the supported oligonucleotide ligation and detection (SOLiD) platform commercialized by Applied Biosystems. Methods that do not require amplification, also known as single-molecule sequencing, include the HeliScope platform commercialized by Helicos BioSciences and platforms commercialized by VisiGen, Oxford Nanopore Technologies Ltd., Life Technologies / Ion Torrent, and Pacific Biosciences, respectively.

[0089] In pyrosequencing (Voelkerding et al., Clinical Chem., 55:641-658, 2009; MacLean et al., Nature Rev. Microbiol., 7:287-296; U.S. Patent Nos. US 6,210,891 and US 6,258,568, each incorporated herein by reference in its entirety), template DNA is fragmented, end-repaired, ligated with adapters, and single template molecules are clonally amplified by beads that capture oligonucleotides complementary to the adapters. Beads each carrying a single template type are partitioned in water-in-oil microcapsules and clonally amplified by a technique called emulsion PCR. After amplification, the emulsion is broken, and the beads are deposited into separate wells of a picotiter plate, which serves as a flow cell in the sequencing reaction. In the presence of a sequencing enzyme and a fluorescent reporter enzyme (such as luciferase), the four dNTP reagents are introduced into the flow cell sequentially and iteratively. When the appropriate dNTP is added to the 3' end of the sequencing primer, the resulting ATP will excite fluorescence in the well, which is recorded by a CCD camera. Read lengths of over or equal to 400 bases can be achieved, generating 10 6 sequencing reads, and ultimately a sequencing output of up to 500 million base pairs (Mb).

[0090] In the Solexa / Illumina platform (Voelkerding et al., Clinical Chem., 55:641-658, 2009; MacLean et al., Nature Rev. Microbiol., 7:287-296; U.S. Patent Nos. 6,833,246, 7,115,400, and 6,969,488, which are incorporated herein by reference in their entirety), sequencing data are generated in the form of short reads. In this method, single-stranded fragmented DNA is end-repaired to generate 5'-phosphorylated blunt ends, and then an A base is added to the 3' end of the fragment by the action of Klenow enzyme. A addition facilitates the ligation of T-overhang adapter oligonucleotides, and then the template-adapter molecules are captured on the surface of a flow cell with anchored oligonucleotides. The anchored sequences serve as PCR primers, but due to the template length and its proximity to other anchored oligonucleotides, PCR amplification "arches" the molecules to hybridize with adjacent anchored oligonucleotides, forming a bridge structure on the surface of the flow cell. These DNA loops are denatured and cleaved. Subsequently, the forward strand is sequenced using reversible fluorescent terminators. The sequence of the incorporated bases is determined by detecting the fluorescence after incorporation, and the fluorescent group and blocking group are removed after each round of incorporation before entering the next round of dNTP addition. The sequencing read lengths range from 36 bases to more than 50 bases, and each analysis run can generate data of more than 1 billion base pairs.

[0091] Sequencing of nucleic acid molecules using the SOLiD technology (Voelkerding et al., Clinical Chem., 55:641-658, 2009; MacLean et al., Nature Rev. Microbiol., 7:287-296; U.S. Patent Nos. 5,912,148 and 6,130,073, which are incorporated herein by reference in their entirety) also includes fragmentation of the template, ligation to oligonucleotide adapters, attachment to beads, and clonal amplification by emulsion PCR. Subsequently, the beads carrying the template are immobilized on the modified surface of a glass flow cell, and a primer complementary to the adapter oligonucleotide is annealed. However, instead of using this primer for 3'-extension, a probe containing two probe-specific bases, six degenerate bases, and a fluorescent label is ligated through the 5'-phosphate group provided by it. In the SOLiD system, the probe has 16 possible base combinations at the 3' end and carries one of four fluorescent labels at the 5' end. The fluorescent colors correspond to a specific color space encoding scheme for the identity of the probe. Multiple cycles (usually 7 rounds) of probe annealing, ligation, and fluorescence detection are performed, followed by denaturation, and a second primer that is offset by one base relative to the initial primer is used for the second round of sequencing. In this way, the template sequence can be reconstructed by calculation, and the template bases are detected twice, thereby improving the accuracy. The average sequencing read length is 35 bases, and the output of a single sequencing run exceeds 4 billion bases.

[0092] In certain embodiments, nanopore sequencing is employed (see, e.g., Astier et al., J. Am. Chem. Soc., Feb. 8, 2006; 128(5):1705-10, which is incorporated herein by reference in its entirety). The principle of nanopore sequencing is related to the phenomenon that occurs when a nanopore is immersed in a conductive liquid and an electric potential (voltage) is applied across it. Under these conditions, ion conduction through the nanopore generates a small current, the magnitude of which is extremely sensitive to the size of the nanopore. When the individual bases of a nucleic acid pass through the nanopore, they cause changes in the current intensity passing through the nanopore, and the changes caused by each of the four bases are distinguishable, thereby allowing the determination of the sequence of the DNA molecule.

[0093] In certain embodiments, the HeliScope platform developed by Helicos BioSciences is employed (Voelkerding et al., Clinical Chem., 55:641-658, 2009; MacLean et al., Nature Rev. Microbiol., 7:287-296; U.S. Patent Nos. US 7,169,560; US 7,282,337; US 7,482,120; US7,501,245; US 6,818,395; US 6,911,345; and US 7,501,245, which are incorporated herein by reference in their entirety). The template DNA is fragmented, a polyadenosine tail is added to the 3' end, and the last adenosine carries a fluorescent label. The denatured polyadenosine template fragments are ligated to poly(dT) oligonucleotides on the surface of the flow cell. The initial physical positions of the captured template molecules are recorded by a CCD camera, and then the fluorescent labels are cleaved and eluted. Sequencing is achieved by adding polymerase and sequentially adding fluorescently labeled dNTP reagents. Base incorporation events generate fluorescent signals corresponding to the dNTPs, and these signals are captured by the CCD camera before each round of dNTP addition. The read length ranges from 25 to 50 bases, and the output per analysis run exceeds 1 billion base pairs.

[0094] Ion Torrent technology is a DNA sequencing method based on the detection of hydrogen ions released during DNA polymerization (see, e.g., Science 327(5970):1190(2010); U.S. Patent Publication Nos. US2009 / 0026082; US2009 / 0127589; US2010 / 0301398; US 2010 / 0197507; US2010 / 0188073; US2010 / 0137143, all incorporated by reference in their entirety for all purposes). A microwell contains template DNA strands to be sequenced. Below the microwell layer is a highly sensitive ISFET ion sensor. All layers are encapsulated in a CMOS semiconductor chip similar to those used in the electronics industry. When a dNTP is incorporated into the nascent complementary strand, a hydrogen ion is released, triggering the highly sensitive ion sensor. If there is a homopolymer repeat sequence in the template sequence, multiple dNTP molecules will be incorporated in one cycle. This will result in the release of a corresponding number of hydrogen ions and a proportionally higher electronic signal. This technology is distinguished from other sequencing technologies in that it does not use modified nucleotides or optical systems. The Ion Torrent sequencer has an accuracy of approximately 99.6% per base at 50-base reads and can generate approximately 100 Mb of data in a single run. The read length is 100 base pairs. For 5-base-long homopolymer repeats, the accuracy is approximately 98%. Advantages of ion semiconductor sequencing include fast sequencing speed and low upfront and running costs.

[0095] Another nucleic acid sequencing method that can be used in the present invention was developed by Stratos Genomics, Inc. and involves the use of Xpandomer. The sequencing process generally includes providing a daughter strand generated by template-directed synthesis. The daughter strand generally contains multiple subunits that are sequentially linked and correspond to all or part of the continuous nucleotide sequence of the target nucleic acid. Each subunit contains a linker arm, at least one probe or base residue, and at least one selectively cleavable bond. The cleavable bond is cleaved to generate Xpandomer, which is longer than the number of subunits in the daughter strand. The Xpandomer generally contains a linker arm and a reporting element for resolving genetic information, in an order corresponding to the continuous nucleotide sequence of the target nucleic acid. Subsequently, the reporting element of the Xpandomer is detected. For more details on the Xpandomer-based method, see, e.g., U.S. Patent Publication No. US2009 / 0035777, which is incorporated herein by reference in its entirety.

[0096] Other single molecule sequencing methods include real-time synthesis sequencing using the VisiGen platform (Voelkerding et al., Clinical Chem., 55:641-658, 2009; U.S. Patent No. US 7,329,492; and U.S. Patent Application Nos. 11 / 671,956; 11 / 781,166, which are incorporated herein by reference in their entirety), in which immobilized primer-bearing DNA templates are extended by a polymerase with a fluorescent label and a fluorescent acceptor molecule to generate a detectable fluorescence resonance energy transfer (FRET) signal upon base incorporation.

[0097] Another real-time single molecule sequencing system developed by Pacific Biosciences (Voelkerding et al., Clinical Chem., 55:641-658, 2009; MacLean et al., Nature Rev. Microbiol., 7:287-296; U.S. Patent Nos. US 7,170,050; US 7,302,146; US 7,313,308; US 7,476,503, which are incorporated herein by reference in their entirety) utilizes reaction pores with a diameter of 50–100 nm and a reaction volume of approximately 20 zeptoliter (10-21 L). The sequencing reaction is carried out using a fixed template, a modified phi29 DNA polymerase, and a high local concentration of fluorescently labeled dNTPs. The high local concentration and continuous reaction conditions allow real-time capture of the fluorescence signal of incorporation events by laser excitation, an optical waveguide, and a CCD camera.

[0098] In some embodiments, a single molecule real-time (SMRT) DNA sequencing method using zero-mode waveguides (ZMWs) developed by Pacific Biosciences or a similar method is employed. This technology performs DNA sequencing on an SMRT chip, each chip containing thousands of zero-mode waveguides (ZMWs). A ZMW is a hole with a diameter of dozens of nanometers, fabricated in a 100-nm metal film deposited on a silica substrate. Each ZMW constitutes a nanophotonic visualization chamber with a detection volume of only 20 zeptoliter (10-21 L). At this volume, the activity of a single molecule can be detected against a background of hundreds of thousands of labeled nucleotides. The ZMW provides a window to observe the process of DNA polymerase performing synthesis sequencing. In each chamber, a single DNA polymerase molecule is immobilized on the bottom surface so that it is always within the detection volume. Nucleotides with a phosphate linkage, each type carrying a different-colored fluorescent group, are added to the reaction solution at a high concentration to enhance the speed, accuracy, and continuity of the enzyme. Due to the extremely small volume of the ZMW, even at high concentrations, nucleotides only enter the detection volume for a small fraction of the time. In addition, due to the extremely short diffusion distance, the process of nucleotides reaching the detection volume is very fast, lasting only a few microseconds, resulting in a very low background signal.

[0099] Processes and systems for such real-time sequencing useful in the present invention are described, for example, in U.S. Patent Nos. 7,405,281; 7,315,019; 7,313,308; 7,302,146; 7,170,050; and U.S. Patent Publication Nos. US2008 / 0212960; US2008 / 0206764; US 2008 / 0199932; US2008 / 0199874; US2008 / 0176769; US2008 / 0176316; US 2008 / 0176241; US2008 / 0165346; US2008 / 0160531; US2008 / 0157005; US 2008 / 0153100; US2008 / 0153095; US2008 / 0152281; US2008 / 0152280; US 2008 / 0145278; US2008 / 0128627; US2008 / 0108082; US2008 / 0095488; US 2008 / 0080059; US2008 / 0050747; US2008 / 0032301; US2008 / 0030628; US 2008 / 0009007; US2007 / 0238679; US2007 / 0231804; US2007 / 0206187; US 2007 / 0196846; US2007 / 0188750; US2007 / 0161017; US2007 / 0141598; US 2007 / 0134128; US2007 / 0128133; US2007 / 0077564; US2007 / 0072196; US 2007 / 0036511; and Korlach et al. (2008) "Selective aluminum passivation for targeted immobilization of single DNA polymerase molecules in zero-mode waveguide nanostructures" PNAS 105(4):1176–81, each incorporated by reference in its entirety.

[0100] Sequencing reads can be attributed to a specific partition, and thus can be identified as coming from a cell within that partition based on the bead-specific barcode linked to the read. However, since a partition can contain multiple beads and thus multiple bead-specific barcodes, this method can be used to determine when, for example, two or more beads are from the same partition, and accordingly apply the sequencing reads containing any one of the relevant bead-specific barcodes. Sequencing reads with different bead-specific barcodes but from the same partition can be identified by the presence of the same UMI in these reads. Since the UMI is unique or nearly unique, if the same UMI is present in different sequencing reads and is linked to different bead-specific barcodes, it indicates that these reads are from the same partition, and thus they can be aggregated with other sequencing reads from the same partition. Note that once it is confirmed that two bead-specific barcodes appear in the same partition, all sequencing reads with these two bead-specific barcodes can be aggregated within that partition, regardless of whether they have the same UMI. This is because only a part (possibly a small amount) of the events are initiated by these two different bead-specific barcode oligonucleotides and extended from the same fragment with the same UMI.

[0101] Optionally, it can also be further determined by the exact fragment sequence that two sequencing reads with different bead-specific barcodes and the same UMI indeed come from the same partition. In some embodiments, the UMI sequence can be randomly generated, and the number of UMIs can be more than the number of copies they are linked to, so that the probability of two different fragments being linked to the same UMI is low but still exists. In other embodiments, a UMI sequence library can be used, and a partition can be identified by the presence or absence of one or more UMIs. However, if the number of UMIs is reduced to save reagents, the probability that two different bead-specific barcodes are linked to the same UMI or UMI library will increase (although still relatively low) due to random events. To rule out these rare events, the cDNA fragment itself can also be considered. If two sequencing reads have different bead-specific barcodes and the same UMI (or UMI library), and the fragments are also the same (the ends are the same, and optionally the sequences are also the same), then these reads can be considered to come from the same partition; conversely, if the fragment ends or sequences are different, these reads can be discarded, or at least cannot be considered to come from the same partition.

[0102] Thus, in some embodiments, the method includes classifying sequencing reads from different partitions, where (i) an amplification product containing a first bead-specific barcode and a first UMI sequence and (ii) an amplification product containing a second bead-specific barcode and the first UMI sequence are from the same partition. See Figure 3 the examples of data deconvolution.

[0103] In another aspect, provided are compositions comprising partitions containing hybrid molecular fragments, said fragments being the hybrid molecular fragments described herein (i.e., double-stranded nucleic acids), wherein the first strand comprises: (i) a first 5'-end linked to a first adapter oligonucleotide and (ii) a first 3'-end, and the second strand comprises: (iii) a second 5'-end linked to a second adapter oligonucleotide and (iv) a second 3'-end. In other aspects, provided are compositions comprising partitions containing the gap-filled hybrid molecular fragments described herein. In other aspects, provided are compositions comprising partitions of permeabilized cells containing gap-filled hybrid molecular fragments with adapter sequences. In other aspects, provided are compositions comprising partitions containing amplification products (e.g., amplicons) that contain bead-specific barcodes. In some embodiments, at least a portion of the partitions contain amplification products (e.g., amplicons) with different bead-specific barcodes. Any of the above partitions may further contain other reagents described herein, such as polymerases, one or more beads, or other reagents.

[0104] In addition, reaction mixtures are provided herein. For example, a bulk solution containing amplification products (e.g., amplicons) with different bead-specific barcodes. Examples Example 1

[0105] To simulate the barcode merging process for multi-barcode partition deconvolution in a computer, a computer-based experimental procedure was designed to describe and validate the present invention. This involved single-cell transcriptome creation, random fragmentation and UMI tagging, ligation of barcode sequences, and barcode merging and determination of single-cell partitions.

[0106] Single-cell transcriptomes were generated from 4000 hypothetical genes. The copy number of each gene followed a truncated normal distribution with a minimum of 20 copies and a maximum of 4000 copies, and a standard deviation of 400 copies. Using this transcriptome, five libraries were created representing five separate partitions containing single cells, where each library had every transcript in the transcriptome that was randomly cut somewhere between 1 - 2 kb and then UMI was added. After randomly sampling fragments from each library, different forms of barcode combinations were added to these fragments to simulate the situation where single or multiple barcodes were present in the partitions, as shown in the process flow chart. Then all the fragments were pooled before the single-cell partition deconvolution step.

[0107] Identify partitions containing single cells by merging barcodes that are most likely to coexist in the barcoding step. During this process, count the number of identical segments between barcodes, where identical is defined as the same cleavage site and UMI, and plot it as an Identical Fragment - Rank plot, where each point represents a pair of barcodes and the y - axis represents the number of identical segments between that pair of barcodes (i.e., the same gene, cleavage site, and UMI). Based on the inflection point, determine a threshold of 5 such that barcodes with more than 5 identical segments are grouped into corresponding bins representing partitions containing cells, and these partitions are shown as cell identities in the table. Example 2

[0108] To validate the barcode - merging process for deconvolution of multi - barcode partitions in single - cell RNAseq experiments, we constructed a transpososome complex loaded with UMI - indexed adapters as described herein. After cell fixation and permeabilization, the transcriptome of a single cell was converted to cDNA in situ and kept in the form of cDNA:RNA hybrids, and then labeled using the UMI - indexed transpososome. Labeled single cells from human and mouse species were mixed in equal molar ratios and then co - partitioned with barcode beads in a droplet - based microfluidic device for single - cell barcoding. Each partition carried various numbers of barcode beads, and according to the Poisson distribution, there were zero or one labeled cell in each partition. During barcoding, the labeled cDNA fragments were barcoded by the barcodes in the partition. After the barcoding step, the barcoded fragments were recovered, merged, and then subjected to bulk - sequencing analysis. Following the same method described in the computer simulation test, single - cell partitions were identified by bioinformatics methods by merging barcodes that were likely to coexist during the barcoding step.

[0109] The success of barcode merging for deconvolution of multi - barcode partitions is demonstrated by the detection of single cells through RNAseq experiments. Figure 7 A rank plot of the detected partitions (after barcode merging) sorted in descending order of read counts is shown. The upper - left dashed box represents the partitions containing single cells. The lower - right dashed box represents the background and empty partitions. The vertical line is the threshold cutoff for determining the total number of single cells. By plotting the total number of unique reads corresponding to each detected partition in the presence of different numbers of barcodes ( Figure 8 ), the effectiveness of barcode merging for deconvolution of multi - barcode partitions was further confirmed. The success of barcode merging was manifested as each partition having a similar number of unique reads, regardless of the number of barcode beads coexisting in that partition.

[0110] Although the foregoing invention has been described in some detail by way of illustration and example for the purposes of clarity of understanding, those skilled in the art will appreciate that certain changes and modifications may be practiced within the scope of the appended claims. Additionally, each reference provided herein is incorporated by reference in its entirety as if individually incorporated item by item, and if there is any conflict between this application and the incorporated references, this application shall prevail.

Claims

1. A method for sorting sequencing reads by an initiation partition, the method comprising: Providing RNA / cDNA or DNA / cDNA hybrid molecules in fixed and permeabilized cells, wherein the fixed and permeabilized cells contain cross-linked molecules; Generating random breaks in the hybrid molecules and randomly inserting a first adapter oligonucleotide or a second adapter oligonucleotide at the breaks, thereby forming hybrid molecule fragments, the hybrid molecule fragments comprising (i) a first 5' end linked to the first adapter oligonucleotide and (ii) a first 3' end and (iii) a second 5' end linked to the second adapter oligonucleotide and (iv) a second 3' end, wherein the first adapter oligonucleotide comprises a first universal sequence, and the second adapter oligonucleotide comprises a second universal sequence, and wherein the first adapter oligonucleotide, the second adapter oligonucleotide, or both further comprise a unique molecular identifier (UMI) sequence; Partitioning the cells together with one or more beads into partitions, wherein each bead is linked to multiple copies of a bead-specific barcode oligonucleotide having the same 3' end comprising the first universal sequence or the second universal sequence, wherein the bead-specific barcode oligonucleotides linked to different beads can be identified by unique bead-specific barcodes in the bead-specific barcode oligonucleotides, and wherein at least some of the partitions contain at least two different beads; Reversing at least partial cross-linking in the cross-linked molecules in the cells; Before, during, or after the reversal, extending with a polymerase: (i) the first 3' end, using the second adapter oligonucleotide as a template, such that the first 3' end is linked to the reverse complementary sequence of the second universal sequence, and (ii) the second 3' end, using the first adapter oligonucleotide as a template, such that the second 3' end is linked to the reverse complementary sequence of the first universal sequence, thereby forming gap-filled hybrid molecule fragments; Amplifying the gap-filled hybrid molecule fragments in the partitions by annealing and extending the bead-specific barcode oligonucleotides to the reverse complementary sequence of the first universal sequence or the second universal sequence on the cDNA fragments to produce amplicons, the amplicons comprising: The bead-specific barcode oligonucleotides, cDNA fragments, the second universal terminal sequence, and the UMI sequence, Under the condition that a first bead and a second bead are present in the partition, respectively using the same cDNA hybrid molecule fragment as a template to extend the bead-specific barcode oligonucleotides from the first bead and the second bead to form (i) an amplicon comprising a first bead-specific barcode and a first UMI sequence and (ii) an amplicon comprising a second bead-specific barcode and a first UMI sequence; Performing nucleotide sequencing on the amplified amplicons to generate sequencing reads; And Sort sequencing reads from different partitions, where: (i) amplicons containing a first bead-specific barcode and a first UMI sequence, and (ii) the amplicons containing a second bead-specific barcode and the first UMI sequence, and (iii) optionally, the same fragment breakpoint, are from the same partition.

2. The method according to claim 1, wherein the hybrid molecule is an RNA / cDNA hybrid molecule, and the cDNA is first-strand cDNA.

3. The method according to claim 2, wherein the RNA / first-strand cDNA hybrid molecule is formed by reverse transcribing RNA from the cell using a poly-A, random, or gene-specific reverse transcription primer.

4. The method according to claim 1, wherein the hybrid molecule is a DNA / cDNA hybrid molecule.

5. The method according to claim 4, wherein the DNA / first-strand cDNA hybrid molecule is formed by polymerase chain reaction or primer extension.

6. The method according to claim 1, wherein the first adapter oligonucleotide contains a UMI sequence.

7. The method according to claim 1, wherein the second adapter oligonucleotide contains a UMI sequence.

8. The method according to claim 1, wherein the first adapter oligonucleotide and the second adapter oligonucleotide contain UMI sequences.

9. The method according to claim 1, wherein the first adapter oligonucleotide, the second adapter oligonucleotide, or both further contain a sample barcode sequence.

10. The method according to claim 1, wherein the generating comprises: Contact the RNA / cDNA hybrid molecule with a transposase, which introduces the adapter oligonucleotide into the RNA / cDNA hybrid molecule.

11. The method according to claim 1, wherein the 3' end of the bead-specific barcode oligonucleotide contains a first universal sequence, and the amplification includes annealing and extending the bead-specific barcode oligonucleotide to the reverse complementary sequence of the first universal sequence.

12. The method according to claim 11, wherein the amplification further comprises: Use the amplicon as a template to extend a reverse primer having a 3' end that contains the reverse complementary sequence of a second universal sequence.

13. The method according to claim 1, wherein the 3' end of the bead-specific barcode oligonucleotide contains a second universal sequence, and the amplification includes annealing and extending the bead-specific barcode oligonucleotide to the reverse complementary sequence of the second universal sequence.

14. The method according to claim 13, wherein said amplification further comprises: Use the amplicon as a template to extend a reverse primer having a 3' end that contains the reverse primer of a first universal sequence.

15. The method according to claim 1, wherein the partition is a micro-well or a droplet in an emulsion.

16. The method according to claim 1, wherein the cell is a mammalian cell.

17. The method according to claim 1, wherein the cell is a prokaryotic cell.

18. The method according to claim 1, wherein the cell is a eukaryotic cell.

19. A plurality of partitions, at least some of the partitions comprising: Fixed and permeabilized cells, the fixed and permeabilized cells containing cross-linked molecules and containing: Hybrid molecule fragments with gap filling, which are formed by: Random breaks are generated in RNA / cDNA or DNA / cDNA hybrid molecules, and a first adaptor oligonucleotide or a second adaptor oligonucleotide is randomly inserted at the break sites, thereby forming hybrid molecule fragments, the hybrid molecule fragments comprising: (i) a first 5' end linked to the first adaptor oligonucleotide, and (ii) a first 3' end, and (iii) a second 5' end linked to the second adaptor oligonucleotide, and (iv) a second 3' end, wherein the first adaptor oligonucleotide comprises a first universal sequence, and the second adaptor oligonucleotide comprises a second universal sequence, and wherein the first adaptor oligonucleotide, the second adaptor oligonucleotide, or both further comprise unique molecular identifier (UMI) sequences; The cells and one or more beads are partitioned into compartments, wherein each bead is linked to multiple copies of a bead-specific barcode oligonucleotide having the same 3' end comprising the first universal sequence or the second universal sequence, wherein the bead-specific barcode oligonucleotides linked to different beads can be identified by the unique bead-specific barcodes in the bead-specific barcode oligonucleotides, and wherein at least some of the compartments contain at least two different beads; and Polymerase extension is used to: (i) extend the first 3' end, using the second adaptor oligonucleotide as a template, such that the first 3' end comprises the reverse complementary sequence of the second universal sequence, and (ii) extend the second 3' end, using the first adaptor oligonucleotide as a template, such that the second 3' end comprises the reverse complementary sequence of the first universal sequence, thereby forming gap-filled hybrid molecule fragments; wherein at least some of the compartments further comprise the one or more beads.

20. The plurality of compartments according to claim 19, wherein the compartments are micro-wells or droplets in an emulsion.

21. The plurality of compartments according to claim 19, wherein the cells are mammalian cells.

22. The plurality of compartments according to claim 19, wherein the cells are prokaryotic cells.

23. The plurality of compartments according to claim 19, wherein the cells are eukaryotic cells.

24. The plurality of compartments according to claim 19, wherein the first adaptor oligonucleotide comprises a UMI sequence.

25. The plurality of compartments according to claim 19, wherein the second adaptor oligonucleotide comprises a UMI sequence.

26. The plurality of compartments according to claim 19, wherein both the first adaptor oligonucleotide and the second adaptor oligonucleotide comprise UMI sequences.

27. The plurality of compartments according to claim 19, wherein the first adaptor oligonucleotide, the second adaptor oligonucleotide, or both further comprise sample barcode sequences.

Citation Information

Patent Citations

  • hydrogels

    US20020009591A1

  • Methods of amplifying and sequencing nucleic acids

    US20050130173A1

  • Methods and systems for monitoring multiple optical signals from a single source

    US20070036511A1

  • Fluorescent nucleotide analogs and uses therefor

    US20070072196A1

  • Reactive surfaces, substrates and methods of producing same

    US20070077564A1